From 4448c8c9d69cec112c56e106c95abcc6d7d59d8c Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Thu, 10 Sep 2026 11:26:30 +0200 Subject: [PATCH 001/181] docs(stories): confirm the loop-rule consolidation profile MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Daniel confirmed the proposed profile (risk high, security none, validation battery+check+verification) on 2026-09-10. The header loses its proposed markers, the DRAFT note is removed, and the profile log records the adoption. Prose only under docs/ — no gate. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../stories/2026-08-29-loop-rule-consolidation-story.md | 9 ++++----- 1 file changed, 4 insertions(+), 5 deletions(-) diff --git a/docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md b/docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md index b094ec5..f9e826b 100644 --- a/docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md +++ b/docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md @@ -1,12 +1,10 @@ # §5 loop-rule consolidation: exits, duties, and the decline — Story **Date:** 2026-08-29 · **Size:** story -**Risk:** *(proposed)* high · **Security:** *(proposed)* none · **Validation:** *(proposed)* battery+check+verification +**Risk:** high · **Security:** none · **Validation:** battery+check+verification -> **DRAFT — the profile above is proposed, not confirmed.** Per §5 a profile is confirmed by the -> human, and until it is this story is not executable. Nothing depends on it yet: the work it -> describes is split out of a cycle that is still running, and the successor starts when someone -> picks it up. The proposal's reasons are in §5. +**Profile log:** +- 2026-09-10 · adoption · proposed 2026-08-29 at the split from the parent cycle, confirmed by Daniel as proposed after the parent shipped (PR #26) · gates now read this header ## 1. Problem statement @@ -147,6 +145,7 @@ ceiling. cycle. **No named `high` trigger matches literally**, so this is a judgement call under intake's "surfaces, not words", and the human decides it. The parent's experience is evidence for rather than against: eleven passes and three mandatory stops on this material. + *(Resolved 2026-09-10: confirmed as proposed — see the profile log.)* - **How much of the ordering is new text versus reference.** §5 already contains the two sentences from which "only clean completion closes" follows; whether the ordering is stated fresh or assembled from what is there changes the old-conditions accounting and the parity surface. From 510f6ce1851188e4116e00a22f634e302f0deee4 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Thu, 10 Sep 2026 12:48:35 +0200 Subject: [PATCH 002/181] docs(specs): loop-rule consolidation design (Gate A pending) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit One closure ordering for the §5 review loop, stated once in both prompt copies: clean completion closes, the scope stop, clearly-stuck exit and two-tell stop suspend; the four standing duties classified; a decline rule and commit-body record; the raw-severity answer for the loop-health counts; the pass-4 report without prior-pass history; the deferred slot discriminator dissolved. Old-conditions accounting against a 135-condition inventory. Docs-only, under docs/ — the Gate-A spec cycle runs next; no gate here. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- ...26-09-10-loop-rule-consolidation-design.md | 462 ++++++++++++++++++ 1 file changed, 462 insertions(+) create mode 100644 docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md new file mode 100644 index 0000000..832260a --- /dev/null +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -0,0 +1,462 @@ +# §5 loop-rule consolidation: one closure ordering — Design + +**Date:** 2026-09-10 · **Status:** DRAFT — Gate A not yet run +**Story:** `docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md` +**Profile:** read from that header at every pass, never from here (it is `high / none / +battery+check+verification` as this is written, and the header is the only writable copy). + +Prompt-only, in two mirrored copies: `CLAUDE.md` §5 (**C** below) and the inline template in +`plugins/dev-workflow/commands/workflow-init.md` (**W** below). No file under +`plugins/dev-workflow/hooks/` changes. Line numbers cite the tree at `7c0d475` (main) and are +re-read at execution; the plan carries the `grep -n` sites. + +Two research inventories were taken before this spec and the plan re-reads them: a +135-condition inventory of the passages this spec rewrites (ids `a1`…`j4`, cited below), and a +site map of every sentence that enumerates the records a commit body carries. + +--- + +## 1. Intent + +**What ships.** One closure ordering, stated once in each copy, that says which exit *closes* a +cycle and which merely *suspend* it, what two suspensions at once do, and which of the four +standing duties participate in that ordering versus gate it as preconditions. On top of it: a +decline rule with a commit-body record, so a user's answer on a surfaced finding has one rule in +both directions; the answer to what a severity demotion does to the loop-health counts; the +answer to the pass-4 report when prior-pass history is unavailable; and the deferred slot +discriminator, dissolved rather than shipped. + +**What does not.** The pass floor and severity semantics (the parent shipped them and this spec +reads them as given); hook code; and the items the story's §2 parks — the pass-counter anomaly, +the CodeRabbit plan-metadata contradiction, and the fixture-per-predicate question. + +--- + +## 2. Settled inputs + +The story's §4 table, decisions 1–10, is the design's starting point and is not restated here; +each is cited below as **D1**…**D10** (with **D9b**, **D9c**). Two implementation facts the parent +cycle established are read as given: a pass's cleanliness is a fact about what that pass found +and is **never rewritten** — an answer changes whether the *cycle* may close; and the findings +files establish the **inventory** of findings, not their resolutions. + +This spec adds four things the table does not settle, each decided in the section that uses +it: the definition of a clean pass under a decline (§3); what the user's answer to a stuck or +two-tell surface produces (§3, last paragraph); the wording and the nonce status of the decline +record (§4); and the raw-severity rule for the counts (§5, passage (g)). + +--- + +## 3. The closure ordering — the block that ships + +It sits in §5 **immediately before** the paragraph "**What a loop absorbs, and what stops it**", +in both copies, byte-identical. It states the ordering once; the paragraphs after it keep their +triggers and point at it. Verbatim as it will ship: + +``` +**How a cycle ends — one ordering, stated here and referenced everywhere else.** A running +cycle stops spending passes in exactly four ways, and only one of them closes it. **Clean +completion closes**: a clean pass at or above the derived floor, or the one early exit below it — +a pass with **zero** findings, which surfaces nothing and can trip none of the stops below. The +other three are **suspensions**: the **scope stop** (a finding outside the assigned fix set, or +one opening a new structural or contract question), the **clearly-stuck exit**, and the +**two-tell stop**. A suspension surfaces to the user with the finding still open and every duty +standing — the floor, the Blocker/Major filter and the clean-final-pass rule are not waived by +stopping — and the loop resumes on the answer. **Nothing else ends a cycle**: a cycle the user +stops without a clean pass is not closed, it is left open, and none of its records is a closing +commit. + +**Four duties, and what each one is.** Two are **preconditions on closure**: the **derived +floor** gates closing and is discharged by the count of valid passes, the last of them clean, +or by the zero-finding exit above; the **Blocker/Major-resolve duty** gates closing and the cleanliness of every pass, is discharged +by repair followed by a clean pass, and applies to **in-set** findings only. Two are the +ordering itself: **a surfaced finding holds closure** until the user answers — that hold is what +makes a suspension a suspension — and **any answer ends it, in either direction**; and **no pass +that raised a hold is credited as clean**, a fact about that pass that no later answer rewrites. +So a decline never closes the cycle on the pass that surfaced the finding: the next pass is +clean or not on its own findings, and that further pass is one more pass, not a licence to +close. + +**A pass is clean** when its validated findings file carries no Blocker or Major except one +matched by a decline recorded in this cycle (Mechanics, the decline record) — such a finding is +outside the set by the user's own decision and reopens nothing — and when the pass raised no +hold. Accepting a surfaced finding puts it in the fix set: a Blocker or Major then owes +resolution before any pass can be clean, a Minor or Nit is collected and never iterated. + +**When two apply at once.** Two suspensions compose: one surface, both reasons in the report, +resume when every question raised is answered. Clean completion outranks the two-tell stop, +exactly as the clearly-stuck paragraph already says of its own exit — a clean pass at or above +the floor closes, and a plateau or the tells go into the closing report rather than blocking +it; reporting "will not converge" on a converged loop is a false report. Clean completion and the scope stop **cannot co-occur**, and +no rule ranks them: the scope stop is raised by a surfaced finding, and a pass that surfaced +one is not clean. Neither can a zero-finding pass meet any suspension: it has nothing to +surface, nothing regenerating, no cluster and no require↔withdraw pair. The Gate-B triviality +skip is outside this ordering — a skipped cycle runs no passes and ends by its own rule. + +**The answer to a clearly-stuck or two-tell surface** is continue or stop. Continue resumes +the loop, on the artifact as revised and the fix set as the story or plan now assigns it — a +finding that narrowing has put outside the set surfaces at the next pass as a scope stop and +can be declined there. Stop closes nothing, as above. +``` + +Why each part is there, briefly. The first paragraph is **D1** and **D3** with the three +suspensions named as the existing paragraphs name them (`b11`/`b13`, `c4`, `e7`). The +"nothing else ends a cycle" sentence exists for story AC 4: every stop now has an answer that +changes state, and the answer that is not "continue" is named rather than left to be inferred. +The duties paragraph is the story's four duties, classified as AC 2 requires; the "further +pass" sentence carries the never-rewritten fact and mirrors the profile-change rule's "one more +pass, not a licence to close" (`o16`), so the two rules read alike. The clean-pass paragraph is +**D4**–**D7** turned into one definition; it is where the cycle would otherwise be unable to +close after a decline. The conflicts paragraph is **D2**, **D3**, and AC 1's demand that +unreachable conflicts be named with their reason instead of legislated (the `fic2` record +documents the instrument that legislating built). **D3** preserves the clearly-stuck +paragraph's sentence "a clean completion takes precedence over this exit" verbatim, so the +block defers to it for that exit rather than stating the same ranking twice; the ordering is +still stated once, and the kept sentence agrees with it. The last paragraph is the one new +answer this spec supplies beyond the table — what "hand the decision to the user" produces — +and it uses only mechanisms already shipped: resumption on the revised artifact (`b18`), and +the scope stop (`b11`). + +What the block deliberately does not do: it does not define the floor, the tells, or the +stuck reading — those stay in their paragraphs (`o24`'s one-definition requirement, `o2`'s +independence of the other predicates from the floor). It does not restate any rule it does not +own (`a13`). + +--- + +## 4. The decline rule and record + +**The rule.** A decline is the user's recorded answer on a finding surfaced by the scope stop — +that it stays outside the assigned fix set — and it is available for no other finding (**D6**): +an in-set Blocker or Major cannot be declined, and a finding surfaced by the stuck or two-tell +stop is not a scope-stop finding until a later pass raises it as one (§3, last paragraph). It +binds for the remainder of its cycle and has no effect in any later one (**D7**); it must be an +explicit, attributable decision on that specific finding — never silence, a general remark +about scope, or an inference (**D8**). A later pass's finding is **the same finding** when all +five of location, defect, severity, consequence and suggested fix match (**D9b**); the match is +read on meaning, not on bytes, because a reviewer rewrites its sentences between passes — and +any difference, severity included, or any genuine uncertainty, makes it a new finding and the +hold applies. A decline releases the hold and never qualifies the Blocker/Major-resolve duty, +which the finding never reached (**D5**). + +**The record**, in the closing commit body, reusing the human-exception transport as a distinct +record type (**D9**): + +``` +Declined: · · cycle +Finding: | | | | +``` + +The `Finding:` line is the finding line from the pass's findings file with its confidence +field removed — the five fields the sameness test reads, in the file's order, a literal pipe +escaped as `\|` exactly as there. **Which commit:** the same rule as the human exception — a +Gate-A cycle in the spec or plan commit, a Gate-B cycle in the WIP commit restated by the +closing amend; several accumulate, order means nothing. **It carries the cycle nonce**, because +**D7** keys it to one cycle and a record that cannot be attributed to its cycle cannot bind to +it; it therefore joins the named set of records the nonce appears in, and the human-exception +record stays outside that set. **What it is worth:** an unverified assertion of the same kind +as the human exception — nothing checks the handle, the asking, or the reason (**D9c**). +**How it differs in force:** it releases a hold, which the human-exception form never does; +narrowness bounds what a false one can do — one fully identified finding, one cycle — and that +is not the same as making it safe. Where the human-exception form is never the answer to a +`STOP and surface`, the decline record is the answer to exactly one of them, the scope stop, +and to nothing else on that list: never a below-floor pass, an unclean final pass, a stuck or +two-tell surface, a Gate-A or Gate-B obligation, or an evidence requirement. + +**Which rules the answer modifies** (story AC 3): the scope stop's resume sentence (`b12`, +which gains "accepting or declining"), the hold and the clean-pass definition (both in the §3 +block, where the answer's two directions are one rule). It modifies **no other rule**, and +none other carries the qualification: not the Blocker/Major-resolve duty (`o19`, **D5** — its +"both must resolve" stays unqualified because a declined finding never enters its scope), not +the floor, not the tells, not the stuck reading. A reader finding "decline" at any rule +outside those three has found a defect. + +**Existing sentences that must name it**, both copies, old → new. Line numbers are C's; W's +are in the site map and re-read at execution. + +1. **Squash carry** (C:892 / W:1076, `j1`). OLD: "…copy every evidence entry, every + human-exception record, the provenance lines, the curves and any skipped cycle's skip + record TOGETHER WITH THE SKIP REASON IT POINTS AT…". NEW: "…copy every evidence entry, + every human-exception record, **every decline record**, the provenance lines, the curves + and any skipped cycle's skip record TOGETHER WITH THE SKIP REASON IT POINTS AT…". Six + members; the rest of the sentence unchanged. +2. **The named nonce set** (C:384–387 / W:578–581). OLD: "…and that set is named rather than + left open: the provenance line, the per-pass curve (including a skip record standing in for + one), the cycle's findings slots, and its advisory working record." NEW: "…and that set is + named rather than left open: the provenance line, the per-pass curve (including a skip + record standing in for one), **any decline record (Mechanics)**, the cycle's findings + slots, and its advisory working record." +3. **The nonce exemption** (C:397–399 / W:591–593). OLD: "The nonce is not required in + records this change neither introduces nor keys to a cycle — the evidence entry and a + human-exception record among them." NEW: "The nonce is not required in records that are + not keyed to a cycle — the evidence entry and a human-exception record among them; **a + decline record is keyed to its cycle and carries it**." (The old sentence's "this change" + dated it to the parent; the new one states the criterion.) +4. **"Both shipped records below"** (C:367 / W:561). OLD: "Both shipped records below carry a + **cycle field**, because…". NEW: "Both shipped records below carry a **cycle field** — and + so does the decline record in Mechanics — because…". A load-bearing count that a third + cycle-attributed record would otherwise falsify. +5. **The unknown-start fallback** (C:153–167 / W:360–374, `i4`–`i8`, extended at `i12`'s + invitation). OLD list: "…at minimum floor 3, severity classified without the demotion, the + provenance-line duty owed, the curve duty owed, and the nonce duties at their strictest…". + NEW: "…at minimum floor 3, severity classified without the demotion, the provenance-line + duty owed, the curve duty owed, the nonce duties at their strictest, **every suspension + binding, and decline records treated as absent so that no hold is released**…". `i12`'s + sentence stays as written; this is the addition it invites. +6. **The human-exception "which commit" rule** (C:988–990 / W:1172–1174, `h3`–`h6`). + Unchanged in wording; the decline record's own paragraph (below the human-exception block, + §5 passage (h)) says "which commit: the same rule as the human exception" and points here, + so the rule is stated once. +7. **The gate-off surface** (C:143–151 / W:350–358). OLD tail: "…silencing reminders; or not + running a pass and reporting that it ran." NEW tail: "…silencing reminders; not running a + pass and reporting that it ran; **or recording a decline nobody made, or one on an in-set + finding**." The list says it is not complete; this change creates a route and names it, + as the parent did for the stated floor. + +The shorter "Copy every record into the squash body" sentence inside the human-exception block +(C:1004–1007) is generic and already covers a decline record; it is not edited. The "records +every cycle owes" list (C:698) enumerates unconditional records only; a decline record is +conditional, like the human exception, and is not added. + +--- + +## 5. Edits to the existing passages, with the old-conditions accounting + +Ids are the inventory's. "Kept" means the sentence stays in its passage; "moved" means it now +lives in the §3 block and leaves the passage; "dropped" carries its reason. + +**(a) The floor paragraphs** (C:72–136 / W:279–343) — **trim** `a17`–`a19` and point at the +block. OLD: "Your final pass must be clean — if the pass at the floor still finds +Blocker/Major, keep going until clean or clearly stuck → then STOP and surface to the user. The +only early exit below the floor is a pass with **zero** findings; don't manufacture findings to +pad." NEW: "Your final pass must be clean; how a cycle closes, and what stops it short of +closing, is the closure ordering below — don't manufacture findings to pad." Accounting: +`a1`–`a16`, `a20`–`a22` kept; `a17` kept (the clause survives); `a18`, `a19` moved. `a13`'s +prohibition on restating is what the pointer form obeys. + +**(b) What a loop absorbs** (C:195–223 / W:402–426) — **two sentence edits**, the triggers stay. +`b12` OLD: "…and it resumes the moment the user says whether the set now includes it." NEW: +"…and it resumes the moment the user says whether the set now includes it — accepting or +declining it, under the closure ordering above." `b17`–`b18` OLD: "Stopping this way is **not +an exit from the gate**: the floor, the Blocker/Major filter and the clean-final-pass rule all +stand, and the loop resumes on the revised artifact once the question is answered — what the +stop prevents…" NEW: "Stopping this way is a **suspension** under the closure ordering above, +and the loop resumes on the revised artifact once the question is answered — what the stop +prevents…". Accounting: `b1`–`b16`, `b18` kept; `b17` moved (the "every duty standing" sentence +of the block). W's three wording differences in this passage are untouched (§6). + +**(c) Recognizing clearly stuck** (C:225–246 / W:428–449) — **trim** the closure sentences and +the surfacing block. OLD, from "That third condition is what makes a plateau…" to the end of +"…and then nothing could satisfy both.": replaced by "That third condition is what makes a +plateau rather than a finish, and it is why **a clean completion takes precedence over this +exit**. This exit is a **suspension** under the closure ordering above: you surface with the +findings still open, and the resolve rule is not waived by surfacing." +Accounting: `c1`–`c8` kept; `c9` kept verbatim (**D3**); `c10`, `c12`, `c13`, `c14` moved +(at-or-above-floor clean pass closes; below the floor nothing closes; the zero-finding +exception; the Minor-below-floor case is the block's "clean pass at or above the derived +floor" read in the negative); `c11` moved (the false-report sentence); `c15`, `c16`, `c17` +kept in the pointer sentence; `c18` moved (no pass that raised a hold is credited as clean); +`c19` moved (resumes on the answer); `c20` **dropped** — it argued that reading the +exit as "stop instead of fixing" would put it in competition with the resolve duty, and the +block now classifies the exit as a suspension that waives nothing, which is that argument's +conclusion stated as a rule. A rationale for a competition the ordering no longer permits +would be prose about a rule that no longer applies. + +**(d) From pass 4 onward** (C:255–261 / W:459–465) — **add** the Q6 sentence (§7). `d1`–`d7` +kept, unchanged. + +**(e) The five tells** (C:263–268 / W:467–472) — **add** one pointer sentence after `e10`: +"This stop is a **suspension** under the closure ordering above." `e1`–`e11` kept. + +**(f) The two rules above do not compete** (C:275–287 / W:473–483) — **unchanged**. "The two +rules above" still names the absorb rule and the stuck reading; the block sits before both and +adds no third rule between them. `f1` remains true: the block composes the two suspensions +and ranks neither over the other. + +**(g) Mechanics · Severity · the handed-over question** (C:810–815 / W:996–999) — **replace** +the whole paragraph, both copies, with the answer. NEW: + +``` + **The demotion changes what a cycle must resolve, never what it counts.** The per-pass + counts, the finding clusters and the tell thresholds read the severity the reviewer wrote in + the findings file, before the ceiling is applied: a demoted finding still counts in the + finding total and in its cluster, and a Blocker demoted to Minor is still a Blocker to the + curve. Two reasons. The curve must stay derivable from the findings files alone — counting + finding lines and leading `BLOCKER` fields per pass reproduces it, which is the only thing + that makes a self-reported curve checkable. And the demotion is the author's judgement about + the fix set; a loop spending passes on findings the author keeps demoting is exactly what the + prose-cluster tell exists to surface, and lowering the counts by that same judgement would + hide it. +``` + +Accounting: `g1` **dropped** (the unsettled statement, now settled); `g2`, `g3` **dropped** +(the interim report-and-stop duty existed only until the question was settled, and its trigger +no longer exists); `g4` **dropped** in C (the ownership sentence, discharged by this change), +and W, which never carried it, gets the same replacement paragraph — so the one deliberate +story-path difference between the copies is removed. + +**(h) Recording a human exception** (C:977–1033 / W:1161–1217) — **unchanged**, and a new +block **"Recording a decline."** is added immediately after it (before "**Do not expect +silence from the gate hook**"), carrying §4's rule, form, which-commit pointer, worth, and force +paragraphs. `h1`–`h26` kept; `h17` stays true of the human-exception form, and the decline +block says which one item of that list it *is* the answer to. + +**(i) When these rules bind** (C:153–167 / W:360–374) — **extend** the strict-reading list +(§4 item 5). `i1`–`i16` kept. + +**(j) The squash carry** (C:892 / W:1076) — **extend** (§4 item 1). `j1`–`j4` kept. + +Also touched, outside the inventoried passages: the nonce set, the nonce exemption, the "Both +shipped records" count and the gate-off list (§4 items 2, 3, 4, 7). + +--- + +## 6. Parity + +The two copies must agree on every rule this spec changes. The block (§3), the decline block +(§4), the (g) replacement, the Q6 sentence and every list extension ship byte-identical in C +and W. The pre-existing divergences the inventory found are handled as follows: + +| Divergence | Kind | This change | +|---|---|---| +| (b) cross-reference "Mechanics" (C) vs "the severity rule" (W) | deliberate: W has no section by that name | **as-is**, stated | +| (b) "exactly how" (C) vs "is how" (W); closing rationale reworded; C-only `infinite-portfolio-canvas` parenthetical | rationale and field citation | **as-is**, stated | +| (e) "you report" (C) vs "report" (W); C-only "Recorded rationale" Bricks paragraph | rationale | **as-is**, stated | +| (f) evidence framing; C-only hypothesis qualifier; punctuation of the late-Blockers clause | rationale | **as-is**, stated; (f) is not edited | +| (c)/(d) paragraph break: C fuses the surfacing block onto "Every pass report states three things" (no blank line at C:246/247); W separates them | structural | **aligned**: the (c) edit re-paragraphs the surfacing block, and C gains the blank line, so the floor-report paragraph stands alone in both | +| (e)/(f) paragraph break: W runs the five-tells paragraph into "The two rules above" (no blank line at W:472/473); C has the Bricks paragraph between | structural | **aligned**: W gains a blank line before "The two rules above"; the Bricks paragraph stays C-only | + +The parity check at execution: extract each edited passage from both files by its lead phrase +and `diff` them; the only differences permitted are the rows marked as-is above, and the (g) +row is gone. Any other difference is a defect, not a wording choice. + +The decline block and the human-exception block it follows are byte-identical between copies +today (the site map confirmed C:840–1033 = W:1024–1217), and stay so. + +--- + +## 7. Q6 — the pass-4 report without prior-pass history + +Appended to the "**From pass 4 onward**" paragraph, both copies, after "…demanding what an +earlier pass had removed.": + +``` +**Where the earlier passes' findings files are unavailable** — a fresh checkout, a cleared +`.context/`, a cycle resumed elsewhere — the report states which of the three lines it can +compute from the files it has, names the ones it cannot and why, and says that the two-tell +threshold is being read on that reduced record. It is not a stop of its own, and it does not +make the working record mandatory: a report that says what it could not see is the duty; a +report that invents the trend, or omits the line without saying so, is the failure. +``` + +This is **D10**. The trend and the require↔withdraw pair are derivable from the mandated +findings files alone (the `fic2` record verified that derivation reproduces the reported +figures), so unavailability is a property of the workspace, not of the format, and the answer +is disclosure rather than a new stop or a new mandatory artifact. + +--- + +## 8. The slot discriminator — dissolved + +Plan C's Tasks 19 and 20 (`docs/superpowers/plans/2026-08-30-review-loop-economics-plan-c-rollout.md`, +the drop note under Task 19) deferred "a general production" for a short deterministic +discriminator in the nonce's slot position to this story. It ships nothing here, because the +case it served no longer exists: every cycle started after the parent's rules bind holds a +nonce, and a cycle whose start cannot be established mints one rather than claiming `none +(pre-rule)` (`i11`). A rule for a no-nonce cycle would legislate for an unreachable state, +which is AC 1's prohibition. The `rle` naming that cycle used stays what its closing body +recorded it as — a plan-local exception under the old rules. This section and the story's §5 +are the record; no prompt text changes. + +--- + +## 9. Verification — mode `battery+check+verification` + +**Battery.** The quality command in `AGENTS.md` § Commands runs once, green, at the Gate-B WIP +commit (`check-version-bump.sh main` needs the committed bump, §10). + +**The check.** Inline shell asserts in the plan, run against the working tree and against the +parent tree, so the same assertion is observed passing where the change exists and failing +where it does not: + +- lead phrase of the §3 block, `**How a cycle ends — one ordering`, count **1** in C and **1** + in W; the same greps against `git show 7c0d475:CLAUDE.md` and + `git show 7c0d475:plugins/dev-workflow/commands/workflow-init.md` count **0**; +- the (g) sentence "That question is owned by the loop-rule consolidation work" count **0** in + both; against the parent tree, **1** in C and **0** in W; +- lead phrase `**Recording a decline.**` count **1** in each; **0** in the parent tree; +- the six-member squash-carry sentence: `every decline record` count **1** in each; **0** in + the parent tree. + +If the claim "the ordering and the record ship in both copies" were false, one of the +working-tree counts would be **0** or the parent-tree counts would not differ from it. The +wiring can produce that observation: each grep reads the file bytes at the named revision, +nothing supplies its own input, and the parent tree is the actual prior state. **The +counterfactual is ABSENT, and is claimed as absent**: the parent carries no ordering block and +no decline record, and the (g) count is the one site where the parent is present and the change +removes it. No count against the parent is claimed as "contradictory". + +**The named verification of the risk path** (story AC 4). Walk every stop the shipped text +names — scope stop, clearly-stuck exit, two-tell stop, a hold awaiting its answer, accept, +decline, two suspensions at once, a below-floor unclean pass, a zero-finding pass, the +unknown-start fallback — and write the **next-state table**: for each, the input that ends +it and the state the cycle is in afterwards, citing the shipped line the row reads, in both +copies. The table lives in the plan and is quoted by the closing commit body, not here. What +would be observed if the claim "no path leaves a cycle unable to close and unable to stop" +were false: a row whose next state is the same stop with no input consumed — the shape the +parent cycle shipped once (a stop whose only answer resumed a cycle that immediately stopped +again). The wiring can produce it because every row is filled from the shipped text, not from +this spec, and the two inputs the `fic2` instrument omitted — the user's answer, and the +decline — are input columns here. + +**What this is not.** It is not the `fic2` decision matrix: that instrument scored N states +against old and new text with an expected output each, and Gate B found two defects in the +technique — a state's inputs must include every input the rule reads (the user's answer was +never a column), and a counterfactual must distinguish ABSENT from CONTRADICTORY. The check +above is a presence test whose parent state is absent by inspection; the walk above takes the +answer as an input and produces a next state rather than a scored expected output. **No fixture +per predicate is built** — that question is parked in the story's §2 and is not reopened. + +**Evidence entry**, in the closing commit body, names: the battery run; the four assert pairs +with their working-tree and parent-tree counts; and the next-state table's location in the plan +plus its row count. It is revalidated before every Gate-B re-review and before the closing +amend, as §5 requires. + +--- + +## 10. AGENTS.md invariants touched + +- **Invariant 11 — `docs/prompt-standards.md`, all 12 items.** Most at risk here: item 6 + (rules carry their why — every sentence of the §3 block carries its reason inline, and the + (g) replacement gives two); item 8 (token-lean — the block replaces closure sentences rather + than adding beside them, and `a13` is the reason the old sentences leave); item 10 + (diagnostic states name their causes — the Q6 sentence names what could not be computed and + why); and item 3 (stop conditions defined — the block's last paragraph is what "hand the + decision to the user" now produces). +- **Don't: "Never replace a decision procedure without accounting for its old conditions."** + Satisfied by §5: every inventory id for every edited passage is marked kept, moved or + dropped with a reason. +- **Don't: "Never rename or delete a doc section without grepping for references first."** The + (g) sentence names the story path; the grep finds it at `CLAUDE.md:815` (the site itself), + `docs/superpowers/plans/2026-08-29-review-loop-economics-plan-a-rules.md:878`, + `docs/superpowers/specs/2026-08-28-review-loop-economics-design.md:33` and + `docs/superpowers/stories/2026-08-28-review-loop-economics-pass-floor-story.md:78`. All + four cite the story file, which continues to exist; none cites the sentence. Nothing breaks. +- **Invariant 12 — a plugin change requires a version bump.** `workflow-init.md` is under + `plugins/`, so `plugins/dev-workflow/.claude-plugin/plugin.json` goes `0.11.0 → 0.12.0` with a + `CHANGELOG.md` entry: a minor bump, because the template gains a record type and a rule. + The entry also notes that the squash-carry sentence now lists six record kinds (the 0.9.0 + entry described two, and history entries are not edited). +- **Invariant 4 / the hook.** Untouched: `codex-gate.sh` is not edited, and the §5 heading it + greps (`Cross-Model Review`) does not move. + +--- + +## 11. Out of scope / parked + +- The pass-counter anomaly (`fic2` record). +- The CodeRabbit plan-metadata contradiction (`fic2` record). +- The fixture-per-predicate question (`fic2` record; story §2). +- Hook code under `plugins/dev-workflow/hooks/`. +- `todos.md`: both-branches-misread-each-other; self-consuming-deletion (prompt-standards + item 11); the three bot findings in resolved plans. From 07c88a1cfaef7b7aaaf0177884cffb4c68ff3572 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Thu, 10 Sep 2026 13:18:33 +0200 Subject: [PATCH 003/181] =?UTF-8?q?docs(specs):=20loop-rule=20consolidatio?= =?UTF-8?q?n=20=E2=80=94=20Gate-A=20pass=201=20revision?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Pass 1 (cycle awsf1ec771): 24 findings, 5 Blocker / 15 Major / 4 Minor, all in-set, all applied. The §3 block gains an evaluation order (clean completion first, at effective severity; suspensions only on a non-closing pass), the scope stop splits into membership and question stops, a stop answer is a standing resumable suspension, and the decline record ships as a verbatim Mechanics block recorded when made. Q6 names its state and its reduced sensitivity. Docs only — no gate. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- ...26-09-10-loop-rule-consolidation-design.md | 508 +++++++++++------- 1 file changed, 308 insertions(+), 200 deletions(-) diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index 832260a..31bc741 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -1,9 +1,9 @@ # §5 loop-rule consolidation: one closure ordering — Design -**Date:** 2026-09-10 · **Status:** DRAFT — Gate A not yet run +**Date:** 2026-09-10 · **Status:** DRAFT — Gate A running (pass 1 revised) **Story:** `docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md` -**Profile:** read from that header at every pass, never from here (it is `high / none / -battery+check+verification` as this is written, and the header is the only writable copy). +**Profile:** read from that header at every pass, never from here — the header is the only +writable copy, and a value copied into this file would be a remembered value. Prompt-only, in two mirrored copies: `CLAUDE.md` §5 (**C** below) and the inline template in `plugins/dev-workflow/commands/workflow-init.md` (**W** below). No file under @@ -11,19 +11,22 @@ Prompt-only, in two mirrored copies: `CLAUDE.md` §5 (**C** below) and the inlin re-read at execution; the plan carries the `grep -n` sites. Two research inventories were taken before this spec and the plan re-reads them: a -135-condition inventory of the passages this spec rewrites (ids `a1`…`j4`, cited below), and a -site map of every sentence that enumerates the records a commit body carries. +135-condition inventory of the passages this spec rewrites (ids `a1`…`j4`, cited below; +22 + 18 + 20 + 7 + 11 + 7 + 4 + 26 + 16 + 4), and a site map of every sentence that +enumerates the records a commit body carries. Where a sentence outside those passages is cited, +it is quoted by its lead phrase and C line. --- ## 1. Intent **What ships.** One closure ordering, stated once in each copy, that says which exit *closes* a -cycle and which merely *suspend* it, what two suspensions at once do, and which of the four -standing duties participate in that ordering versus gate it as preconditions. On top of it: a -decline rule with a commit-body record, so a user's answer on a surfaced finding has one rule in -both directions; the answer to what a severity demotion does to the loop-health counts; the -answer to the pass-4 report when prior-pass history is unavailable; and the deferred slot +cycle and which merely *suspend* it, in what order a pass is read so the ranking is executable +rather than asserted, what any set of suspensions at once does, and which of the four standing +duties participate in that ordering versus gate it as preconditions. On top of it: a decline +rule with a commit-body record, so a user's answer on a surfaced finding has one rule in both +directions; the answer to what a severity demotion does to the loop-health counts; the answer +to the pass-4 report when prior-pass history is unavailable; and the deferred slot discriminator, dissolved rather than shipped. **What does not.** The pass floor and severity semantics (the parent shipped them and this spec @@ -40,10 +43,12 @@ cycle established are read as given: a pass's cleanliness is a fact about what t and is **never rewritten** — an answer changes whether the *cycle* may close; and the findings files establish the **inventory** of findings, not their resolutions. -This spec adds four things the table does not settle, each decided in the section that uses -it: the definition of a clean pass under a decline (§3); what the user's answer to a stuck or -two-tell surface produces (§3, last paragraph); the wording and the nonce status of the decline -record (§4); and the raw-severity rule for the counts (§5, passage (g)). +This spec adds what the table does not settle, each decided in the section that uses it: the +evaluation order and the definition of a clean pass under a decline, read at effective +severity (§3); the two triggers of the scope stop and what each answer does (§3); what the +user's answer to a stuck or two-tell surface produces, including a stop (§3); the wording, +the nonce status and the recording point of the decline record (§4); and the raw-severity +rule for the health measures (§5, passage (g)). --- @@ -54,122 +59,172 @@ in both copies, byte-identical. It states the ordering once; the paragraphs afte triggers and point at it. Verbatim as it will ship: ``` -**How a cycle ends — one ordering, stated here and referenced everywhere else.** A running -cycle stops spending passes in exactly four ways, and only one of them closes it. **Clean -completion closes**: a clean pass at or above the derived floor, or the one early exit below it — -a pass with **zero** findings, which surfaces nothing and can trip none of the stops below. The -other three are **suspensions**: the **scope stop** (a finding outside the assigned fix set, or -one opening a new structural or contract question), the **clearly-stuck exit**, and the -**two-tell stop**. A suspension surfaces to the user with the finding still open and every duty -standing — the floor, the Blocker/Major filter and the clean-final-pass rule are not waived by -stopping — and the loop resumes on the answer. **Nothing else ends a cycle**: a cycle the user -stops without a clean pass is not closed, it is left open, and none of its records is a closing -commit. - -**Four duties, and what each one is.** Two are **preconditions on closure**: the **derived -floor** gates closing and is discharged by the count of valid passes, the last of them clean, -or by the zero-finding exit above; the **Blocker/Major-resolve duty** gates closing and the cleanliness of every pass, is discharged -by repair followed by a clean pass, and applies to **in-set** findings only. Two are the -ordering itself: **a surfaced finding holds closure** until the user answers — that hold is what -makes a suspension a suspension — and **any answer ends it, in either direction**; and **no pass -that raised a hold is credited as clean**, a fact about that pass that no later answer rewrites. -So a decline never closes the cycle on the pass that surfaced the finding: the next pass is -clean or not on its own findings, and that further pass is one more pass, not a licence to -close. - -**A pass is clean** when its validated findings file carries no Blocker or Major except one -matched by a decline recorded in this cycle (Mechanics, the decline record) — such a finding is -outside the set by the user's own decision and reopens nothing — and when the pass raised no -hold. Accepting a surfaced finding puts it in the fix set: a Blocker or Major then owes -resolution before any pass can be clean, a Minor or Nit is collected and never iterated. - -**When two apply at once.** Two suspensions compose: one surface, both reasons in the report, -resume when every question raised is answered. Clean completion outranks the two-tell stop, -exactly as the clearly-stuck paragraph already says of its own exit — a clean pass at or above -the floor closes, and a plateau or the tells go into the closing report rather than blocking -it; reporting "will not converge" on a converged loop is a false report. Clean completion and the scope stop **cannot co-occur**, and -no rule ranks them: the scope stop is raised by a surfaced finding, and a pass that surfaced -one is not clean. Neither can a zero-finding pass meet any suspension: it has nothing to -surface, nothing regenerating, no cluster and no require↔withdraw pair. The Gate-B triviality -skip is outside this ordering — a skipped cycle runs no passes and ends by its own rule. - -**The answer to a clearly-stuck or two-tell surface** is continue or stop. Continue resumes -the loop, on the artifact as revised and the fix set as the story or plan now assigns it — a -finding that narrowing has put outside the set surfaces at the next pass as a scope stop and -can be declined there. Stop closes nothing, as above. +**How a cycle ends — one ordering, stated here and referenced everywhere else.** A pass is +read in a fixed order, because every rule below bears on one decision — may this cycle close — +and a stated order is what stops them qualifying each other. **First, clean completion**, read +on the pass's validated findings file at **effective** severity, after the Mechanics severity +ceiling: cleanliness is about what the cycle must repair, and the ceiling is what decides that. +**A pass is clean** when it carries no in-set Blocker or Major at effective severity that is +not matched by a decline recorded in this cycle (Mechanics, the decline record), and it +surfaced no scope stop; a pass with **zero** findings is clean whatever the floor. **A clean +pass at or above the derived floor, or a zero-finding pass, closes the cycle** — a plateau or +tells present on that pass go into the closing report and never block it, because reporting +"will not converge" on a converged loop is a false report, and the clearly-stuck paragraph +says the same of its own exit. Only the named health measures — the per-pass counts, the +clusters and the tells — read the reviewer-written field before the ceiling (Mechanics, +Severity). **Nothing else closes a cycle.** + +**Second, only a pass that is not a clean completion can suspend** — that order is what makes +"clean completion outranks the two-tell stop" executable rather than asserted. Three +suspensions, by the names their paragraphs use: the **scope stop**, with two triggers — a +**membership stop** (a finding outside the assigned fix set) and a **question stop** (a +finding, in-set or not, opening a new structural or contract question); the **clearly-stuck +exit**; and the **two-tell stop**. A suspension waives nothing — the floor, the Blocker/Major +filter and the clean-final-pass rule stand while it does. Any non-empty set of them can apply +to one pass: **one surface, every reason reported, every question asked**, because a reason +left out is a decision made by omission. A finding surfaced solely by the stuck or two-tell +reading is not a scope-stop finding; one that is also outside the set, or also opens a +question, takes the scope stop's answers at that same surface — it is not asked twice. + +**What a suspension asks, and what ends it.** Only a scope stop raises a **hold**, on its +finding, and a hold ends with **any** answer, in either direction — a hold only accepting could +end would be the resolve duty under another name. At a membership stop the answer is +**accept** (the finding joins the fix set and resolves by its effective severity: Blocker or +Major before any pass can be clean, Minor or Nit collected and never iterated) or **decline** +(the finding stays outside, recorded, binding for the rest of this cycle — Mechanics, the +decline record). At a question stop the answer is the user's decision on the question, and +membership does not change: an in-set finding then resolves under that decision or is +dismissed with its one-line why; an out-of-set finding that opened the question is a +membership stop as well and takes accept or decline. **Decline is available only at a +membership stop.** The stuck and two-tell readings raise no hold: each asks one question, +**continue or stop**. Continue resumes the loop on the artifact as revised and the fix set as +the governing artifacts now assign it — where several plans or stories govern one cycle, the +union of the scopes they assign plus the obligations already accepted — and a finding that a +narrowing has put outside the set surfaces at the next pass as a membership stop. **Stop +leaves the suspension standing**: the cycle stays open under its nonce, resumable by a later +continue, nothing it wrote is a closing commit, and an open cycle nobody resumes is a human's +to resolve, exactly as the nonce rules already say of open cycles. + +**Composition, and what cannot happen.** Every question is answered on its own, and the loop +resumes only when every answer resumes it — accept or decline at a membership stop, a decision +at a question stop, continue at the stuck or two-tell reading; one stop answer leaves the +whole suspension standing, because a loop resumed over an unanswered question decides it by +running. **No pass that surfaced a scope stop is credited as clean** — a fact about that pass +no later answer rewrites — so a decline never closes the cycle on the surfacing pass; the +next pass is clean or not on its own findings. Two states cannot co-occur, and no rule ranks +them: clean completion and a scope stop, since a pass that surfaced one is not clean; and a +zero-finding pass and any suspension, since it has nothing to surface, nothing regenerating, +no cluster and no require↔withdraw pair. The Gate-B triviality skip is outside this ordering: +a skipped cycle runs no passes and ends by its own rule. ``` -Why each part is there, briefly. The first paragraph is **D1** and **D3** with the three -suspensions named as the existing paragraphs name them (`b11`/`b13`, `c4`, `e7`). The -"nothing else ends a cycle" sentence exists for story AC 4: every stop now has an answer that -changes state, and the answer that is not "continue" is named rather than left to be inferred. -The duties paragraph is the story's four duties, classified as AC 2 requires; the "further -pass" sentence carries the never-rewritten fact and mirrors the profile-change rule's "one more -pass, not a licence to close" (`o16`), so the two rules read alike. The clean-pass paragraph is -**D4**–**D7** turned into one definition; it is where the cycle would otherwise be unable to -close after a decline. The conflicts paragraph is **D2**, **D3**, and AC 1's demand that -unreachable conflicts be named with their reason instead of legislated (the `fic2` record -documents the instrument that legislating built). **D3** preserves the clearly-stuck -paragraph's sentence "a clean completion takes precedence over this exit" verbatim, so the -block defers to it for that exit rather than stating the same ranking twice; the ordering is -still stated once, and the kept sentence agrees with it. The last paragraph is the one new -answer this spec supplies beyond the table — what "hand the decision to the user" produces — -and it uses only mechanisms already shipped: resumption on the revised artifact (`b18`), and -the scope stop (`b11`). +Why each part is there, briefly. The evaluation order is the answer to the objection that a +ranking between clean completion and the two-tell stop cannot fire if the stop can make the +pass unclean: clean completion is read first, so a suspension is only ever evaluated on a pass +that did not close. Reading cleanliness at effective severity keeps the resolve duty and the +clean-pass test on one field, and leaves the health measures — passage (g) — on the other; +a demoted finding cannot then keep a pass unclean while owing no repair. The first paragraph +is **D1**, **D2** and **D3**, the three suspensions named as their paragraphs name them +(`b11`/`b13`, `c4`, `e7`); "nothing else closes a cycle" exists for story AC 4. The two +triggers of the scope stop are `b11` (membership) and `b13` (question) read separately, +because an in-set finding that opens a question can neither "join" nor "stay outside" the set +and needs its own answer. The composition paragraph is AC 1's demand that every reachable +conflict be covered and every unreachable one named with its reason (the `fic2` record +documents the instrument that legislating built). The stop answer is the one new state this +spec supplies beyond the table — a standing suspension, resumable, never a close — and it +uses only mechanisms already shipped: resumption on the revised artifact (`b18`), the +membership stop (`b11`), and the nonce rules' existing treatment of open cycles ("they stay +open, keep their own nonces, and are a human's to resolve", C:423–424). **D3** keeps the +clearly-stuck paragraph's own precedence sentence verbatim (§5, passage (c)); the block cites +it rather than restating its consequence clause. The "further pass" reading mirrors the +profile-change rule — "it is one more pass, not a licence to close on the next one" (C:763–764) +— so the two read alike. What the block deliberately does not do: it does not define the floor, the tells, or the -stuck reading — those stay in their paragraphs (`o24`'s one-definition requirement, `o2`'s -independence of the other predicates from the floor). It does not restate any rule it does not -own (`a13`). +stuck reading — those stay in their paragraphs ("exactly one definition of the floor must be +present", C:176; "the other closure and stop predicates … not required to derive from the +floor", C:181–185). It does not restate any rule it does not own (`a13`). --- ## 4. The decline rule and record -**The rule.** A decline is the user's recorded answer on a finding surfaced by the scope stop — -that it stays outside the assigned fix set — and it is available for no other finding (**D6**): -an in-set Blocker or Major cannot be declined, and a finding surfaced by the stuck or two-tell -stop is not a scope-stop finding until a later pass raises it as one (§3, last paragraph). It -binds for the remainder of its cycle and has no effect in any later one (**D7**); it must be an -explicit, attributable decision on that specific finding — never silence, a general remark -about scope, or an inference (**D8**). A later pass's finding is **the same finding** when all -five of location, defect, severity, consequence and suggested fix match (**D9b**); the match is -read on meaning, not on bytes, because a reviewer rewrites its sentences between passes — and -any difference, severity included, or any genuine uncertainty, makes it a new finding and the -hold applies. A decline releases the hold and never qualifies the Blocker/Major-resolve duty, -which the finding never reached (**D5**). - -**The record**, in the closing commit body, reusing the human-exception transport as a distinct -record type (**D9**): - -``` -Declined: · · cycle -Finding: | | | | -``` - -The `Finding:` line is the finding line from the pass's findings file with its confidence -field removed — the five fields the sameness test reads, in the file's order, a literal pipe -escaped as `\|` exactly as there. **Which commit:** the same rule as the human exception — a -Gate-A cycle in the spec or plan commit, a Gate-B cycle in the WIP commit restated by the -closing amend; several accumulate, order means nothing. **It carries the cycle nonce**, because -**D7** keys it to one cycle and a record that cannot be attributed to its cycle cannot bind to -it; it therefore joins the named set of records the nonce appears in, and the human-exception -record stays outside that set. **What it is worth:** an unverified assertion of the same kind -as the human exception — nothing checks the handle, the asking, or the reason (**D9c**). -**How it differs in force:** it releases a hold, which the human-exception form never does; -narrowness bounds what a false one can do — one fully identified finding, one cycle — and that -is not the same as making it safe. Where the human-exception form is never the answer to a -`STOP and surface`, the decline record is the answer to exactly one of them, the scope stop, -and to nothing else on that list: never a below-floor pass, an unclean final pass, a stuck or -two-tell surface, a Gate-A or Gate-B obligation, or an evidence requirement. +**The rule.** A decline is the user's recorded answer at a membership stop — that the surfaced +finding stays outside the assigned fix set — and it is available for no other finding +(**D6**): not for an in-set Blocker or Major, and not as the answer to a question stop, a stuck +or two-tell surface, or any obligation. It binds for the remainder of its cycle and has no +effect in any later one (**D7**); it must be an explicit, attributable decision on that specific +finding — never silence, a general remark about scope, or an inference (**D8**). A later pass's +finding is **the same finding** when all five of location, defect, severity, consequence and +suggested fix match (**D9b**); the match is read on meaning, not on bytes, because a reviewer +rewrites its sentences between passes — and any difference, severity included, or any genuine +uncertainty, makes it a new finding and the hold applies. A decline releases the hold and never +qualifies the Blocker/Major-resolve duty, which the finding never reached (**D5**). + +**When it is recorded.** When made, into the cycle's next commit body on the branch — a spec or +plan revision commit for a Gate-A cycle, the `WIP:` amend for a Gate-B cycle — and restated in +every later body of that cycle, the closing one included. The advisory working record may +carry it too. A cycle that resumes with no commit body carrying a decline treats it as absent +and the hold applies again — the reading the unknown-start fallback gives, and the safe +direction: a lost decline costs a repeated question, an invented one releases a hold nobody +chose to release. **D9**'s "closing commit body" is therefore the last of the bodies that carry +it, not the first. + +**The block that ships**, in Mechanics, placed **immediately after the whole human-exception +passage** — after its last paragraph "**What the record is worth.**" (C:1028–1033 / +W:1212–1217) and before the next bullet "- **Timeout / abort:**" (C:1034 / W:1218), at the same +two-space indent. Verbatim, byte-identical in both copies: + +```` + **Recording a decline.** A decline is the user's answer at a **membership stop** (the + closure ordering above) that the surfaced finding stays outside the assigned fix set. It is + available there and nowhere else: not for an in-set Blocker or Major, which owes resolution, + and not as the answer to a question stop, a stuck or two-tell surface, a below-floor pass, an + unclean final pass, or any Gate-A, Gate-B or evidence obligation — of the list the + human-exception form is never the answer to, the membership stop is the one item this + record answers. It must be an **explicit, attributable decision on that specific finding** — + never silence, never a general remark about scope, never inferred — because a hold released + by inference is a hold nobody chose to release. It **releases the hold** and **binds for the + remainder of this cycle**, with no effect in any later one; it never qualifies the + Blocker/Major-resolve duty, which the finding never reached. A later pass raises **the same + finding** when all five of location, defect, severity, consequence and suggested fix match, + read on meaning rather than bytes, since a reviewer rewrites its sentences between passes; + **any difference — severity included — or any genuine uncertainty makes it a new finding, + and the hold applies.** + + ``` + Declined: · · cycle + Finding: | | | | + ``` + + The `Finding:` line is the finding line from the pass's findings file with its confidence + field removed — the five fields the sameness test reads, in the file's order, a literal pipe + escaped as `\|` exactly as there. **It carries the cycle nonce**, because it binds to one + cycle and a record that cannot be attributed to its cycle cannot bind to it. **When and + where:** written when made, into the cycle's next commit body on the branch — a spec or plan + revision commit for a Gate-A cycle, the `WIP:` amend for Gate B — and restated in every later + body of that cycle, the closing one included, because a body that drops it releases nothing + and re-asks the question; the advisory working record may carry it as well. A cycle that + resumes with no commit body carrying a decline **treats it as absent and the hold applies + again** — the reading the unknown-start fallback gives, and the safe direction. It is copied + on squash-merge with the other records (the carry rule above). **What it is worth:** an + unverified assertion of the same kind as the human exception — nothing checks that the + handle belongs to whoever decided, that a human was asked, or that the reason is honest. + **How it differs in force:** it releases a hold, which the human-exception form never does; + narrowness bounds what a false one can do — one fully identified finding, one cycle — and + that is not the same as making it safe. **What the body does not record, said here rather + than discovered:** an accepted finding, a stuck or two-tell surface and its continue-or-stop + answer leave no record of their own — from a closing body a reader can infer the close and + any declines, and nothing else about which suspensions the cycle passed through. +```` **Which rules the answer modifies** (story AC 3): the scope stop's resume sentence (`b12`, -which gains "accepting or declining"), the hold and the clean-pass definition (both in the §3 -block, where the answer's two directions are one rule). It modifies **no other rule**, and -none other carries the qualification: not the Blocker/Major-resolve duty (`o19`, **D5** — its -"both must resolve" stays unqualified because a declined finding never enters its scope), not -the floor, not the tells, not the stuck reading. A reader finding "decline" at any rule -outside those three has found a defect. +which gains "accepting or declining"), and the hold and the clean-pass definition (both in the +§3 block, where the answer's two directions are one rule). It modifies **no other rule**, and +none other carries the qualification: not the Blocker/Major-resolve duty (Mechanics, Severity: +"both must resolve", C:784 — **D5**: it stays unqualified because a declined finding never +enters its scope), not the floor, not the tells, not the stuck reading. A reader finding +"decline" at any rule outside those three has found a defect. **Existing sentences that must name it**, both copies, old → new. Line numbers are C's; W's are in the site map and re-read at execution. @@ -180,38 +235,66 @@ are in the site map and re-read at execution. every human-exception record, **every decline record**, the provenance lines, the curves and any skipped cycle's skip record TOGETHER WITH THE SKIP REASON IT POINTS AT…". Six members; the rest of the sentence unchanged. -2. **The named nonce set** (C:384–387 / W:578–581). OLD: "…and that set is named rather than - left open: the provenance line, the per-pass curve (including a skip record standing in for - one), the cycle's findings slots, and its advisory working record." NEW: "…and that set is - named rather than left open: the provenance line, the per-pass curve (including a skip +2. **The named nonce set** (C:385–387 / W:579–581). OLD: "**and that set is named rather than + left open**: the provenance line, the per-pass curve (including a skip record standing in + for one), the cycle's findings slots, and its advisory working record." NEW: "**and that set + is named rather than left open**: the provenance line, the per-pass curve (including a skip record standing in for one), **any decline record (Mechanics)**, the cycle's findings slots, and its advisory working record." -3. **The nonce exemption** (C:397–399 / W:591–593). OLD: "The nonce is not required in - records this change neither introduces nor keys to a cycle — the evidence entry and a - human-exception record among them." NEW: "The nonce is not required in records that are - not keyed to a cycle — the evidence entry and a human-exception record among them; **a - decline record is keyed to its cycle and carries it**." (The old sentence's "this change" - dated it to the parent; the new one states the criterion.) +3. **The nonce exemption** (C:397–399 / W:591–593). OLD: "The nonce is not required in records + this change neither introduces nor keys to a cycle — the evidence entry and a + human-exception record among them." NEW: "The nonce is not required in records that are not + keyed to a cycle — the evidence entry and a human-exception record among them; **a decline + record is keyed to its cycle and carries it**." (The old sentence's "this change" dated it + to the parent; the new one states the criterion.) 4. **"Both shipped records below"** (C:367 / W:561). OLD: "Both shipped records below carry a **cycle field**, because…". NEW: "Both shipped records below carry a **cycle field** — and so does the decline record in Mechanics — because…". A load-bearing count that a third cycle-attributed record would otherwise falsify. 5. **The unknown-start fallback** (C:153–167 / W:360–374, `i4`–`i8`, extended at `i12`'s - invitation). OLD list: "…at minimum floor 3, severity classified without the demotion, the + invitation). OLD: "…at minimum floor 3, severity classified without the demotion, the provenance-line duty owed, the curve duty owed, and the nonce duties at their strictest…". NEW: "…at minimum floor 3, severity classified without the demotion, the provenance-line duty owed, the curve duty owed, the nonce duties at their strictest, **every suspension binding, and decline records treated as absent so that no hold is released**…". `i12`'s sentence stays as written; this is the addition it invites. -6. **The human-exception "which commit" rule** (C:988–990 / W:1172–1174, `h3`–`h6`). - Unchanged in wording; the decline record's own paragraph (below the human-exception block, - §5 passage (h)) says "which commit: the same rule as the human exception" and points here, - so the rule is stated once. +6. **The human-exception "which commit" rule** (C:988–990 / W:1172–1174, `h3`–`h6`). Unchanged + in wording. The decline block carries its own rule, which deliberately differs — written + when made and restated in every later body, not written once into the closing body — + because a decline must survive to the next pass while the human exception only has to + survive to history. 7. **The gate-off surface** (C:143–151 / W:350–358). OLD tail: "…silencing reminders; or not running a pass and reporting that it ran." NEW tail: "…silencing reminders; not running a pass and reporting that it ran; **or recording a decline nobody made, or one on an in-set finding**." The list says it is not complete; this change creates a route and names it, as the parent did for the stated floor. +8. **The closing message and the soft-reset path** (C:834–838 / W:1018–1022, with the + `git reset --soft` sentence at C:829–830 / W:1013–1014). OLD: "**The closing message + carries the validated evidence entry for every cited profiled story** — one each, and none + for a cited unprofiled story, which owes no entry. The amend replaces the WIP message + wholesale, so an entry written only into the WIP body is destroyed exactly when the cycle + closes." NEW: "**The closing message carries the validated evidence entry for every cited + profiled story** — one each, and none for a cited unprofiled story, which owes no entry — + **and every decline record the cycle made, including any held only in WIP bodies a + `git reset --soft` collapsed**: the single commit after the reset carries all of them, + because a body the reset discards is unreachable from the commit that replaces it. The + amend replaces the WIP message wholesale, so an entry or record written only into the WIP + body is destroyed exactly when the cycle closes." The soft-reset sentence itself is + unchanged; this sentence is where the carry duty already lives. +9. **The one-contract paragraph** (C:879–890 / W:1063–1074). OLD opening: "**These records are + one contract, and a partial adoption breaks it.** The nonce, the slot naming, the + provenance line, the curve, this carry rule **and the unknown-start activation semantics + that say what a cycle owes when its starting rules cannot be established** depend on one + another," NEW opening: "**These records are one contract, and a partial adoption breaks + it.** The nonce, the slot naming, the provenance line, the curve, this carry rule, **the + closure ordering and the decline record**, and **the unknown-start activation semantics + that say what a cycle owes when its starting rules cannot be established** depend on one + another," and after "and a carry rule naming records a project does not produce is + inert." (C:885) add: "an ordering without the decline record is a hold whose declined + direction has no release rule, a decline record without the ordering is a release with no + hold to release, and a decline record missing from the squash carry or the nonce set is a + record that cannot survive a merge or be attributed." The stop sentence that follows is + unchanged and now covers these states. The shorter "Copy every record into the squash body" sentence inside the human-exception block (C:1004–1007) is generic and already covers a decline record; it is not edited. The "records @@ -223,7 +306,8 @@ conditional, like the human exception, and is not added. ## 5. Edits to the existing passages, with the old-conditions accounting Ids are the inventory's. "Kept" means the sentence stays in its passage; "moved" means it now -lives in the §3 block and leaves the passage; "dropped" carries its reason. +lives in the §3 block and leaves the passage; "replaced" means the condition changes and says +how; "dropped" carries its reason. **(a) The floor paragraphs** (C:72–136 / W:279–343) — **trim** `a17`–`a19` and point at the block. OLD: "Your final pass must be clean — if the pass at the floor still finds @@ -242,25 +326,32 @@ an exit from the gate**: the floor, the Blocker/Major filter and the clean-final stand, and the loop resumes on the revised artifact once the question is answered — what the stop prevents…" NEW: "Stopping this way is a **suspension** under the closure ordering above, and the loop resumes on the revised artifact once the question is answered — what the stop -prevents…". Accounting: `b1`–`b16`, `b18` kept; `b17` moved (the "every duty standing" sentence -of the block). W's three wording differences in this passage are untouched (§6). - -**(c) Recognizing clearly stuck** (C:225–246 / W:428–449) — **trim** the closure sentences and -the surfacing block. OLD, from "That third condition is what makes a plateau…" to the end of -"…and then nothing could satisfy both.": replaced by "That third condition is what makes a -plateau rather than a finish, and it is why **a clean completion takes precedence over this -exit**. This exit is a **suspension** under the closure ordering above: you surface with the -findings still open, and the resolve rule is not waived by surfacing." -Accounting: `c1`–`c8` kept; `c9` kept verbatim (**D3**); `c10`, `c12`, `c13`, `c14` moved -(at-or-above-floor clean pass closes; below the floor nothing closes; the zero-finding -exception; the Minor-below-floor case is the block's "clean pass at or above the derived -floor" read in the negative); `c11` moved (the false-report sentence); `c15`, `c16`, `c17` -kept in the pointer sentence; `c18` moved (no pass that raised a hold is credited as clean); -`c19` moved (resumes on the answer); `c20` **dropped** — it argued that reading the -exit as "stop instead of fixing" would put it in competition with the resolve duty, and the -block now classifies the exit as a suspension that waives nothing, which is that argument's -conclusion stated as a rule. A rationale for a competition the ordering no longer permits -would be prose about a rule that no longer applies. +prevents…". Accounting: `b1`–`b16`, `b18` kept; `b17` moved (the block's "a suspension waives +nothing" sentence). W's three wording differences in this passage are untouched (§6). + +**(c) Recognizing clearly stuck** (C:225–246 / W:428–449) — **keep** the precedence sentence +in full and **trim** what follows it. The sentence kept byte-for-byte (**D3**): "That third +condition is what makes a plateau rather than a finish, and it is why **a clean completion +takes precedence over this exit**: a Blocker/Major-free pass **at or above the floor** has +satisfied the clean-final-pass rule — collect the Minors and Nits and close — and reporting +"will not converge" on a converged loop is a false report." OLD, from "**Below the floor +nothing closes**" to the end of "…and then nothing could satisfy both.": replaced by +"**Below the floor nothing closes**, exactly as the closure ordering above says. This exit is +a **suspension** under that ordering: you surface with the findings still open, the resolve +rule is not waived by surfacing, and the loop resumes on a continue." Accounting: `c1`–`c8` +kept; `c9`, `c10`, `c11` kept verbatim; `c12` kept (pointer form); `c13`, `c14` moved (the +zero-finding exception; the Minor-below-floor case is the block's "clean pass at or above the +derived floor" read in the negative); `c15`, `c16`, `c17` kept in the pointer sentence; `c18` +moved (no pass that surfaced a scope stop is credited as clean — narrowed to the scope stop, +because a stuck or two-tell surface on a below-floor clean pass leaves that pass's cleanliness +as it was); `c19` **replaced** — old: the loop resumes on whatever the user decides, +unconditionally; new: it resumes on a resuming answer and a stop leaves the suspension +standing — because an unconditional resume is the stop-with-no-transition path AC 4 forbids; +`c20` **dropped** — it argued that reading the exit as "stop instead of fixing" would put it +in competition with the resolve duty, and the block classifies the exit as a suspension that +waives nothing, which is that argument's conclusion stated as a rule; a rationale for a +competition the ordering no longer permits would be prose about a rule that no longer +applies. **(d) From pass 4 onward** (C:255–261 / W:459–465) — **add** the Q6 sentence (§7). `d1`–`d7` kept, unchanged. @@ -270,8 +361,8 @@ kept, unchanged. **(f) The two rules above do not compete** (C:275–287 / W:473–483) — **unchanged**. "The two rules above" still names the absorb rule and the stuck reading; the block sits before both and -adds no third rule between them. `f1` remains true: the block composes the two suspensions -and ranks neither over the other. +adds no third rule between them. Accounting: `f1`–`f7` kept, unchanged; `f1` remains true +because the block composes the suspensions and ranks none over another. **(g) Mechanics · Severity · the handed-over question** (C:810–815 / W:996–999) — **replace** the whole paragraph, both copies, with the answer. NEW: @@ -281,12 +372,13 @@ the whole paragraph, both copies, with the answer. NEW: counts, the finding clusters and the tell thresholds read the severity the reviewer wrote in the findings file, before the ceiling is applied: a demoted finding still counts in the finding total and in its cluster, and a Blocker demoted to Minor is still a Blocker to the - curve. Two reasons. The curve must stay derivable from the findings files alone — counting - finding lines and leading `BLOCKER` fields per pass reproduces it, which is the only thing - that makes a self-reported curve checkable. And the demotion is the author's judgement about - the fix set; a loop spending passes on findings the author keeps demoting is exactly what the - prose-cluster tell exists to surface, and lowering the counts by that same judgement would - hide it. + curve. Cleanliness and the resolve duty read the effective severity, after the ceiling (the + closure ordering above). Two reasons for the split. The curve must stay derivable from the + findings files alone — counting finding lines and leading `BLOCKER` fields per pass + reproduces it, which is the only thing that makes a self-reported curve checkable. And the + demotion is the author's judgement about the fix set; a loop spending passes on findings the + author keeps demoting is exactly what the prose-cluster tell exists to surface, and lowering + the counts by that same judgement would hide it. ``` Accounting: `g1` **dropped** (the unsettled statement, now settled); `g2`, `g3` **dropped** @@ -295,11 +387,10 @@ no longer exists); `g4` **dropped** in C (the ownership sentence, discharged by and W, which never carried it, gets the same replacement paragraph — so the one deliberate story-path difference between the copies is removed. -**(h) Recording a human exception** (C:977–1033 / W:1161–1217) — **unchanged**, and a new -block **"Recording a decline."** is added immediately after it (before "**Do not expect -silence from the gate hook**"), carrying §4's rule, form, which-commit pointer, worth, and force -paragraphs. `h1`–`h26` kept; `h17` stays true of the human-exception form, and the decline -block says which one item of that list it *is* the answer to. +**(h) Recording a human exception** (C:977–1033 / W:1161–1217) — **unchanged**, and the +"**Recording a decline.**" block (§4, verbatim) is added immediately after its last paragraph, +before "- **Timeout / abort:**". `h1`–`h26` kept; `h17` stays true of the human-exception form, +and the decline block says which one item of that list it *is* the answer to. **(i) When these rules bind** (C:153–167 / W:360–374) — **extend** the strict-reading list (§4 item 5). `i1`–`i16` kept. @@ -307,7 +398,8 @@ block says which one item of that list it *is* the answer to. **(j) The squash carry** (C:892 / W:1076) — **extend** (§4 item 1). `j1`–`j4` kept. Also touched, outside the inventoried passages: the nonce set, the nonce exemption, the "Both -shipped records" count and the gate-off list (§4 items 2, 3, 4, 7). +shipped records" count, the gate-off list, the closing-message sentence and the one-contract +paragraph (§4 items 2, 3, 4, 7, 8, 9). --- @@ -328,7 +420,8 @@ and W. The pre-existing divergences the inventory found are handled as follows: The parity check at execution: extract each edited passage from both files by its lead phrase and `diff` them; the only differences permitted are the rows marked as-is above, and the (g) -row is gone. Any other difference is a defect, not a wording choice. +row is gone. Any other difference is a defect, not a wording choice. The check's result is part +of the evidence entry (§9). The decline block and the human-exception block it follows are byte-identical between copies today (the site map confirmed C:840–1033 = W:1024–1217), and stay so. @@ -341,18 +434,27 @@ Appended to the "**From pass 4 onward**" paragraph, both copies, after "…deman earlier pass had removed.": ``` -**Where the earlier passes' findings files are unavailable** — a fresh checkout, a cleared -`.context/`, a cycle resumed elsewhere — the report states which of the three lines it can -compute from the files it has, names the ones it cannot and why, and says that the two-tell -threshold is being read on that reduced record. It is not a stop of its own, and it does not -make the working record mandatory: a report that says what it could not see is the duty; a -report that invents the trend, or omits the line without saying so, is the failure. +**Where an earlier pass's findings file is unavailable** — its slot path absent from +`.context/codex-reviews/` in this workspace, as after a fresh checkout, a cleared `.context/` +or a cycle resumed elsewhere; a file that is present but fails validation is an incomplete +pass, already excluded, and a path resolved against the wrong root is the existing stop — the +report first reads the cycle's working record if one exists and says whether it used it, then +states which of the three lines it computed from the files it has, names the ones it could +not and why, and says that the two-tell threshold is being read on that reduced record. The +sensitivity is reduced concretely: the trend and the require↔withdraw comparison read only +the passes present, so a rising count or a pair that spans a missing pass cannot be seen. It +is not a stop of its own, and it does not make the working record mandatory: a report that +says what it could not see is the duty; a report that invents the trend, or omits the line +without saying so, is the failure. ``` This is **D10**. The trend and the require↔withdraw pair are derivable from the mandated findings files alone (the `fic2` record verified that derivation reproduces the reported figures), so unavailability is a property of the workspace, not of the format, and the answer -is disclosure rather than a new stop or a new mandatory artifact. +is disclosure rather than a new stop or a new mandatory artifact. The three causes named are +partitioned by what the agent observes — path absent, file present but invalid, path resolved +against the wrong root — and only the first is this state; the other two already have their +own rules, which is why the sentence points at them rather than restating them. --- @@ -386,7 +488,9 @@ where it does not: both; against the parent tree, **1** in C and **0** in W; - lead phrase `**Recording a decline.**` count **1** in each; **0** in the parent tree; - the six-member squash-carry sentence: `every decline record` count **1** in each; **0** in - the parent tree. + the parent tree; +- the Q6 lead phrase `**Where an earlier pass's findings file is unavailable**` count **1** in + each; **0** in the parent tree. If the claim "the ordering and the record ship in both copies" were false, one of the working-tree counts would be **0** or the parent-tree counts would not differ from it. The @@ -397,17 +501,19 @@ no decline record, and the (g) count is the one site where the parent is present removes it. No count against the parent is claimed as "contradictory". **The named verification of the risk path** (story AC 4). Walk every stop the shipped text -names — scope stop, clearly-stuck exit, two-tell stop, a hold awaiting its answer, accept, -decline, two suspensions at once, a below-floor unclean pass, a zero-finding pass, the -unknown-start fallback — and write the **next-state table**: for each, the input that ends -it and the state the cycle is in afterwards, citing the shipped line the row reads, in both -copies. The table lives in the plan and is quoted by the closing commit body, not here. What -would be observed if the claim "no path leaves a cycle unable to close and unable to stop" -were false: a row whose next state is the same stop with no input consumed — the shape the -parent cycle shipped once (a stop whose only answer resumed a cycle that immediately stopped -again). The wiring can produce it because every row is filled from the shipped text, not from -this spec, and the two inputs the `fic2` instrument omitted — the user's answer, and the -decline — are input columns here. +names — membership stop, question stop, clearly-stuck exit, two-tell stop, a hold awaiting its +answer, accept, decline, a stop answer, two or three suspensions at once, a below-floor clean +pass, a zero-finding pass, the unknown-start fallback — and write the **next-state table**: for +each, the input that ends it and the state the cycle is in afterwards, citing the shipped line +the row reads, in both copies. The table lives in the plan and is quoted by the closing commit +body, not here. What would be observed if the claim "no path leaves a cycle unable to close +and unable to suspend" were false: a row whose next state is the same stop with no input +consumed — the shape the parent cycle shipped once (a stop whose only answer resumed a cycle +that immediately stopped again). The wiring can produce it because every row is filled from +the shipped text, not from this spec, and the inputs the `fic2` instrument omitted — the user's +answer, and the decline — are input columns here; so is the stop answer, whose next state is a +standing suspension by design and must read as one in the table, not as the same stop +re-raised. **What this is not.** It is not the `fic2` decision matrix: that instrument scored N states against old and new text with an expected output each, and Gate B found two defects in the @@ -417,25 +523,27 @@ above is a presence test whose parent state is absent by inspection; the walk ab answer as an input and produces a next state rather than a scored expected output. **No fixture per predicate is built** — that question is parked in the story's §2 and is not reopened. -**Evidence entry**, in the closing commit body, names: the battery run; the four assert pairs -with their working-tree and parent-tree counts; and the next-state table's location in the plan -plus its row count. It is revalidated before every Gate-B re-review and before the closing -amend, as §5 requires. +**Evidence entry**, in the closing commit body, names: the battery run; the five assert pairs +with their working-tree and parent-tree counts; the §6 parity diff — the passages extracted, +the differences observed, and that each is one of the permitted rows; and the next-state +table's location in the plan plus its row count. It is revalidated before every Gate-B +re-review and before the closing amend, as §5 requires. --- ## 10. AGENTS.md invariants touched - **Invariant 11 — `docs/prompt-standards.md`, all 12 items.** Most at risk here: item 6 - (rules carry their why — every sentence of the §3 block carries its reason inline, and the - (g) replacement gives two); item 8 (token-lean — the block replaces closure sentences rather - than adding beside them, and `a13` is the reason the old sentences leave); item 10 - (diagnostic states name their causes — the Q6 sentence names what could not be computed and - why); and item 3 (stop conditions defined — the block's last paragraph is what "hand the - decision to the user" now produces). + (rules carry their why — each constraint in the §3 block and the decline block carries its + reason in the same sentence, and the (g) replacement gives two; the plan's review reads each + sentence for one); item 8 (token-lean — the block replaces closure sentences rather than + adding beside them, and `a13` is the reason the old sentences leave); item 10 (diagnostic + states name their causes — the Q6 sentence partitions the unavailable state from its two + neighbours and points each at its rule); and item 3 (stop conditions defined — the block's + stop answer is a named, resumable state). - **Don't: "Never replace a decision procedure without accounting for its old conditions."** - Satisfied by §5: every inventory id for every edited passage is marked kept, moved or - dropped with a reason. + Satisfied by §5: every inventory id for every edited passage is marked kept, moved, replaced + or dropped with a reason. - **Don't: "Never rename or delete a doc section without grepping for references first."** The (g) sentence names the story path; the grep finds it at `CLAUDE.md:815` (the site itself), `docs/superpowers/plans/2026-08-29-review-loop-economics-plan-a-rules.md:878`, From 8d20b2ea669f22ae471d5a4842f5f2d4cf02f22d Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Thu, 10 Sep 2026 13:58:21 +0200 Subject: [PATCH 004/181] =?UTF-8?q?docs(specs):=20loop-rule=20consolidatio?= =?UTF-8?q?n=20=E2=80=94=20Gate-A=20pass=202=20revision?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Pass 2 (cycle awsf1ec771): 17 findings, 2 Blocker, 13 Major, 2 Minor; 16 fixed, one fixed in part (the mandatory-record demand dismissed: the existing no-identity rule already answers it). The clean predicate now names the fix set only; a matching declined finding raises no stop; a hold ends when every required answer is given; the four duties are classified in the block; every predicate reads the logical pass's branch files as one set; the decline is recorded before the loop resumes; Q6 gains a root check and a three-state partition. Docs-only — no gate. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- ...26-09-10-loop-rule-consolidation-design.md | 377 ++++++++++++------ 1 file changed, 249 insertions(+), 128 deletions(-) diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index 31bc741..c3b0b5d 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -1,6 +1,6 @@ # §5 loop-rule consolidation: one closure ordering — Design -**Date:** 2026-09-10 · **Status:** DRAFT — Gate A running (pass 1 revised) +**Date:** 2026-09-10 · **Status:** DRAFT — Gate A running (pass 2 revised) **Story:** `docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md` **Profile:** read from that header at every pass, never from here — the header is the only writable copy, and a value copied into this file would be a remembered value. @@ -45,10 +45,12 @@ files establish the **inventory** of findings, not their resolutions. This spec adds what the table does not settle, each decided in the section that uses it: the evaluation order and the definition of a clean pass under a decline, read at effective -severity (§3); the two triggers of the scope stop and what each answer does (§3); what the +severity over every branch file of the logical pass (§3); the classification of the four +duties (§3); the two triggers of the scope stop and what each answer does (§3); what the user's answer to a stuck or two-tell surface produces, including a stop (§3); the wording, -the nonce status and the recording point of the decline record (§4); and the raw-severity -rule for the health measures (§5, passage (g)). +the nonce status and the recording point of the decline record, and what happens to the +state the body does not record when a cycle loses it (§4); and the raw-severity rule for +the health measures (§5, passage (g)). --- @@ -61,18 +63,38 @@ triggers and point at it. Verbatim as it will ship: ``` **How a cycle ends — one ordering, stated here and referenced everywhere else.** A pass is read in a fixed order, because every rule below bears on one decision — may this cycle close — -and a stated order is what stops them qualifying each other. **First, clean completion**, read -on the pass's validated findings file at **effective** severity, after the Mechanics severity -ceiling: cleanliness is about what the cycle must repair, and the ceiling is what decides that. -**A pass is clean** when it carries no in-set Blocker or Major at effective severity that is -not matched by a decline recorded in this cycle (Mechanics, the decline record), and it -surfaced no scope stop; a pass with **zero** findings is clean whatever the floor. **A clean -pass at or above the derived floor, or a zero-finding pass, closes the cycle** — a plateau or -tells present on that pass go into the closing report and never block it, because reporting -"will not converge" on a converged loop is a false report, and the clearly-stuck paragraph -says the same of its own exit. Only the named health measures — the per-pass counts, the -clusters and the tells — read the reviewer-written field before the ceiling (Mechanics, -Severity). **Nothing else closes a cycle.** +and a stated order is what stops them qualifying each other. Every predicate here reads the +validated findings file **or files** of the logical pass as one set — a `full` Gate-B pass +has two, and one branch alone is already an incomplete pass — at **effective** severity, after +the Mechanics severity ceiling, because cleanliness is about what the cycle must repair and +the ceiling is what decides that; only the named health measures — the per-pass counts, the +clusters and the tells — read the reviewer-written field before it (Mechanics, Severity). + +**First, clean completion.** A pass is clean when it carries no Blocker or Major at effective +severity that is **in the assigned fix set**, and it surfaced no scope stop; a pass with +**zero** findings is clean whatever the floor. A finding matched by a decline recorded in +this cycle is outside the set by that decision and is not surfaced when re-raised (Mechanics, +the decline record) — a decline keeps a finding out and never excuses one that is in, so a +declined finding the set later comes to include owes resolution like any other. **A clean +pass at or above the derived floor, or a zero-finding pass, closes the cycle**, under the +final-acceptance preconditions the floor section already states and this ordering does not +restate — the cited set and every profile re-read before the pass is accepted as final, a +header or profile that changed during it making the pass not final and costing the further +pass that section requires. A plateau or tells present on the closing pass go into the +closing report and never block it, because reporting "will not converge" on a converged loop +is a false report, and the clearly-stuck paragraph says the same of its own exit. **Nothing +else closes a cycle.** + +**The four standing duties, classified.** The **derived floor** is a precondition on +closure: it gates closing, and is discharged by the count of valid logical passes reaching +the floor with the last of them clean, or by the zero-finding exit. The **Blocker/Major-resolve +duty** is a precondition on closure and on any pass being clean: it gates both, and is +discharged by repair of every in-set Blocker and Major followed by a pass that finds none. +The **hold** a surfaced finding places on closure is part of the ordering: it gates closing +while it stands, and is discharged by the answers that finding requires, below. +**No-clean-credit** — no pass that surfaced a scope stop is credited as clean — is part of +the ordering and a fact about that pass, discharged by nothing: a later pass is judged on its +own findings, so a decline never closes the cycle on the pass that surfaced the finding. **Second, only a pass that is not a clean completion can suspend** — that order is what makes "clean completion outranks the two-tell stop" executable rather than asserted. Three @@ -84,61 +106,78 @@ filter and the clean-final-pass rule stand while it does. Any non-empty set of t to one pass: **one surface, every reason reported, every question asked**, because a reason left out is a decision made by omission. A finding surfaced solely by the stuck or two-tell reading is not a scope-stop finding; one that is also outside the set, or also opens a -question, takes the scope stop's answers at that same surface — it is not asked twice. +question, takes the scope stop's answers at that same surface — it is not asked twice. A +re-raised finding that a decline of this cycle matches on all five fields raises no hold and +no stop; a difference in any field, or genuine uncertainty, is a new finding and a new stop. **What a suspension asks, and what ends it.** Only a scope stop raises a **hold**, on its -finding, and a hold ends with **any** answer, in either direction — a hold only accepting could -end would be the resolve duty under another name. At a membership stop the answer is -**accept** (the finding joins the fix set and resolves by its effective severity: Blocker or -Major before any pass can be clean, Minor or Nit collected and never iterated) or **decline** -(the finding stays outside, recorded, binding for the rest of this cycle — Mechanics, the -decline record). At a question stop the answer is the user's decision on the question, and -membership does not change: an in-set finding then resolves under that decision or is -dismissed with its one-line why; an out-of-set finding that opened the question is a -membership stop as well and takes accept or decline. **Decline is available only at a -membership stop.** The stuck and two-tell readings raise no hold: each asks one question, -**continue or stop**. Continue resumes the loop on the artifact as revised and the fix set as -the governing artifacts now assign it — where several plans or stories govern one cycle, the -union of the scopes they assign plus the obligations already accepted — and a finding that a -narrowing has put outside the set surfaces at the next pass as a membership stop. **Stop -leaves the suspension standing**: the cycle stays open under its nonce, resumable by a later -continue, nothing it wrote is a closing commit, and an open cycle nobody resumes is a human's -to resolve, exactly as the nonce rules already say of open cycles. +finding, and the hold ends when every answer that finding requires has been given — one for a +single-trigger finding, both the question decision and the membership answer for one carrying +both triggers — and in either direction: a hold only accepting could end would be the resolve +duty under another name. At a membership stop the answer is **accept** (the finding joins the +fix set and resolves by its effective severity: Blocker or Major before any pass can be clean, +Minor or Nit collected and never iterated) or **decline** (the finding stays outside, +recorded, binding for the rest of this cycle — Mechanics, the decline record). At a question +stop the answer is the user's decision on the question, and membership does not change: an +in-set finding then resolves under that decision or is dismissed with its one-line why; an +out-of-set finding that opened the question is a membership stop as well and takes accept or +decline. **Decline is available only at a membership stop.** The stuck and two-tell readings +raise no hold: each asks one question, **continue or stop**. Continue resumes the loop on the +artifact as revised and the fix set as the governing artifacts now assign it — where several +plans or stories govern one cycle, the union of the scopes they assign plus the obligations +already accepted — and a finding that a narrowing has put outside the set surfaces at the +next pass as a membership stop. **Stop leaves the suspension standing**: the cycle stays open +under its nonce, resumable by a later continue, nothing it wrote is a closing commit, and an +open cycle nobody resumes is a human's to resolve, exactly as the nonce rules already say of +open cycles. **Composition, and what cannot happen.** Every question is answered on its own, and the loop resumes only when every answer resumes it — accept or decline at a membership stop, a decision at a question stop, continue at the stuck or two-tell reading; one stop answer leaves the whole suspension standing, because a loop resumed over an unanswered question decides it by -running. **No pass that surfaced a scope stop is credited as clean** — a fact about that pass -no later answer rewrites — so a decline never closes the cycle on the surfacing pass; the -next pass is clean or not on its own findings. Two states cannot co-occur, and no rule ranks -them: clean completion and a scope stop, since a pass that surfaced one is not clean; and a -zero-finding pass and any suspension, since it has nothing to surface, nothing regenerating, -no cluster and no require↔withdraw pair. The Gate-B triviality skip is outside this ordering: -a skipped cycle runs no passes and ends by its own rule. +running. Two states cannot co-occur, and no rule ranks them: clean completion and a scope +stop, since a pass that surfaced one is not clean; and a zero-finding pass and any +suspension, since it has nothing to surface, nothing regenerating, no cluster and no +require↔withdraw pair. The Gate-B triviality skip is outside this ordering: a skipped cycle +runs no passes and ends by its own rule. **This ordering and the decline record in Mechanics +are one contract**: the hold's declined direction is released by that record and by nothing +else, so either present without the other is a partial adoption that stops under the +one-contract rule there. ``` Why each part is there, briefly. The evaluation order is the answer to the objection that a ranking between clean completion and the two-tell stop cannot fire if the stop can make the pass unclean: clean completion is read first, so a suspension is only ever evaluated on a pass -that did not close. Reading cleanliness at effective severity keeps the resolve duty and the -clean-pass test on one field, and leaves the health measures — passage (g) — on the other; -a demoted finding cannot then keep a pass unclean while owing no repair. The first paragraph -is **D1**, **D2** and **D3**, the three suspensions named as their paragraphs name them -(`b11`/`b13`, `c4`, `e7`); "nothing else closes a cycle" exists for story AC 4. The two -triggers of the scope stop are `b11` (membership) and `b13` (question) read separately, -because an in-set finding that opens a question can neither "join" nor "stay outside" the set -and needs its own answer. The composition paragraph is AC 1's demand that every reachable -conflict be covered and every unreachable one named with its reason (the `fic2` record -documents the instrument that legislating built). The stop answer is the one new state this -spec supplies beyond the table — a standing suspension, resumable, never a close — and it -uses only mechanisms already shipped: resumption on the revised artifact (`b18`), the -membership stop (`b11`), and the nonce rules' existing treatment of open cycles ("they stay -open, keep their own nonces, and are a human's to resolve", C:423–424). **D3** keeps the -clearly-stuck paragraph's own precedence sentence verbatim (§5, passage (c)); the block cites -it rather than restating its consequence clause. The "further pass" reading mirrors the -profile-change rule — "it is one more pass, not a licence to close on the next one" (C:763–764) -— so the two read alike. +that did not close. Reading every predicate over the union of the logical pass's branch files +keeps a `full` Gate-B pass from closing on one clean branch while the other carries an in-set +Blocker; reading them at effective severity keeps the resolve duty and the clean-pass test on +one field and leaves the health measures — passage (g) — on the other, so a demoted finding +cannot keep a pass unclean while owing no repair. The clean predicate names the fix set and +nothing else, so a decline is only ever a fact about membership: it cannot be read as +excusing an in-set finding, which is the gate-off route a fabricated record would otherwise +open. The close sentence points at the floor section's final-acceptance preconditions +("again before a clean pass is accepted as the cycle's final pass", C:116–118; "Any profile +change costs at least one further pass", C:760–764) instead of restating them, so a closing +pass under a changed header stays not-final as that section already says. The duties +paragraph is story AC 2 — each duty labelled precondition or ordering participant, with what +it gates and what discharges it. The first paragraph is **D1**, **D2** and **D3**, the three +suspensions named as their paragraphs name them (`b11`/`b13`, `c4`, `e7`); "nothing else +closes a cycle" exists for story AC 4. The two triggers of the scope stop are `b11` +(membership) and `b13` (question) read separately, because an in-set finding that opens a +question can neither "join" nor "stay outside" the set and needs its own answer; a finding +carrying both triggers requires both answers before its hold ends, so one answer cannot +release a hold the other question still holds. The declined re-raise raises no stop because +a decline that re-asked its own question every pass would bind for nothing (**D7**). The +composition paragraph is AC 1's demand that every reachable conflict be covered and every +unreachable one named with its reason (the `fic2` record documents the instrument that +legislating built). The stop answer is the one new state this spec supplies beyond the table +— a standing suspension, resumable, never a close — and it uses only mechanisms already +shipped: resumption on the revised artifact (`b18`), the membership stop (`b11`), and the +nonce rules' existing treatment of open cycles ("they stay open, keep their own nonces, and +are a human's to resolve", C:423–424). **D3** keeps the clearly-stuck paragraph's own +precedence sentence verbatim (§5, passage (c)); the block cites it rather than restating its +consequence clause. The closing contract sentence is reciprocal with one in the decline +block, so a copy that adopts one block without the other carries its own stop. What the block deliberately does not do: it does not define the floor, the tells, or the stuck reading — those stay in their paragraphs ("exactly one definition of the floor must be @@ -157,18 +196,40 @@ effect in any later one (**D7**); it must be an explicit, attributable decision finding — never silence, a general remark about scope, or an inference (**D8**). A later pass's finding is **the same finding** when all five of location, defect, severity, consequence and suggested fix match (**D9b**); the match is read on meaning, not on bytes, because a reviewer -rewrites its sentences between passes — and any difference, severity included, or any genuine -uncertainty, makes it a new finding and the hold applies. A decline releases the hold and never -qualifies the Blocker/Major-resolve duty, which the finding never reached (**D5**). - -**When it is recorded.** When made, into the cycle's next commit body on the branch — a spec or -plan revision commit for a Gate-A cycle, the `WIP:` amend for a Gate-B cycle — and restated in -every later body of that cycle, the closing one included. The advisory working record may -carry it too. A cycle that resumes with no commit body carrying a decline treats it as absent -and the hold applies again — the reading the unknown-start fallback gives, and the safe -direction: a lost decline costs a repeated question, an invented one releases a hold nobody -chose to release. **D9**'s "closing commit body" is therefore the last of the bodies that carry -it, not the first. +rewrites its sentences between passes. A matching finding raises no hold and no membership +stop when re-raised — it is not "surfaced", so the pass stays eligible for clean and the +decline is idempotent for its cycle (**D7**); any difference, severity included, or any genuine +uncertainty, makes it a new finding and a new stop, and the hold applies. A decline releases the +hold and never qualifies the Blocker/Major-resolve duty, which the finding never reached +(**D5**); it keeps a finding out and never excuses one that is in — a declined finding the fix +set later comes to include, through an accepted broadening or a governing artifact +re-assigning scope, owes resolution like any other, which is why the clean predicate in §3 +names the set and not the decline. + +**When it is recorded, and that it is recorded before the loop resumes.** When made, into the +cycle's next commit body on the branch — a spec or plan revision commit for a Gate-A cycle, the +`WIP:` amend for a Gate-B cycle — and that commit is made **before the next pass runs**: no pass +runs on a decline that no commit body on the branch carries, because a decline held only in a +session is one compaction away from a re-asked question, and **D7**'s remainder-of-cycle binding +is only as durable as its transport. The advisory working record may carry it meanwhile and +does not bind. It is restated in every later body of that cycle, the closing one included. A +cycle that resumes with no commit body carrying a decline treats it as absent and the hold +applies again — the reading the unknown-start fallback gives, and the safe direction: a lost +decline costs a repeated question, an invented one releases a hold nobody chose to release. +**D9**'s "closing commit body" is therefore the last of the bodies that carry it, not the first. + +**What the body does not record, and what a cycle does when it loses it.** An accepted +finding, a standing stuck or two-tell suspension and its continue-or-stop answer have no record +of their own; they live in the running session and, optionally, in the working record. No new +mandatory record is added for them — **D10** refused one for the same class of state, and the +existing rules already answer the loss: a cycle that cannot recover them has no identity and +starts a new cycle under the no-identity rule ("**No candidate, disagreeing sources, or more +than one candidate → no identity: start a new cycle**", C:419–420), which costs passes and +closes nothing. The reviewer re-reads the whole artifact each pass, so an unrepaired accepted +finding is re-raised, and one the new cycle's assigned scope does not include is re-asked as a +membership stop — a repeated question rather than a silent close, the safe direction. The +residual, stated: a finding the reviewer fails to re-raise is lost exactly as any missed +finding is, which this change neither creates nor removes. **The block that ships**, in Mechanics, placed **immediately after the whole human-exception passage** — after its last paragraph "**What the record is worth.**" (C:1028–1033 / @@ -189,8 +250,10 @@ two-space indent. Verbatim, byte-identical in both copies: Blocker/Major-resolve duty, which the finding never reached. A later pass raises **the same finding** when all five of location, defect, severity, consequence and suggested fix match, read on meaning rather than bytes, since a reviewer rewrites its sentences between passes; + a matching finding raises no hold and no stop, so a decline is not re-asked every pass; **any difference — severity included — or any genuine uncertainty makes it a new finding, - and the hold applies.** + and the hold applies.** A decline keeps a finding out and never excuses one that is in: a + declined finding the fix set later comes to include owes resolution like any other. ``` Declined: · · cycle @@ -202,29 +265,46 @@ two-space indent. Verbatim, byte-identical in both copies: escaped as `\|` exactly as there. **It carries the cycle nonce**, because it binds to one cycle and a record that cannot be attributed to its cycle cannot bind to it. **When and where:** written when made, into the cycle's next commit body on the branch — a spec or plan - revision commit for a Gate-A cycle, the `WIP:` amend for Gate B — and restated in every later - body of that cycle, the closing one included, because a body that drops it releases nothing - and re-asks the question; the advisory working record may carry it as well. A cycle that - resumes with no commit body carrying a decline **treats it as absent and the hold applies - again** — the reading the unknown-start fallback gives, and the safe direction. It is copied - on squash-merge with the other records (the carry rule above). **What it is worth:** an - unverified assertion of the same kind as the human exception — nothing checks that the - handle belongs to whoever decided, that a human was asked, or that the reason is honest. - **How it differs in force:** it releases a hold, which the human-exception form never does; - narrowness bounds what a false one can do — one fully identified finding, one cycle — and - that is not the same as making it safe. **What the body does not record, said here rather - than discovered:** an accepted finding, a stuck or two-tell surface and its continue-or-stop - answer leave no record of their own — from a closing body a reader can infer the close and - any declines, and nothing else about which suspensions the cycle passed through. + revision commit for a Gate-A cycle, the `WIP:` amend for Gate B — **and that commit is made + before the next pass runs**: no pass runs on a decline that no commit body on the branch + carries, because a decline held only in a session is one compaction away from a re-asked + question; the advisory working record may carry it meanwhile and does not bind. It is + restated in every later body of that cycle, the closing one included, because a body that + drops it releases nothing and re-asks the question. A cycle that resumes with no commit body + carrying a decline **treats it as absent and the hold applies again** — the reading the + unknown-start fallback gives, and the safe direction. It is copied on squash-merge with the + other records (the carry rule above). **What it is worth:** an unverified assertion of the + same kind as the human exception — nothing checks that the handle belongs to whoever + decided, that a human was asked, or that the reason is honest. **How it differs in force:** + it releases a hold, which the human-exception form never does; narrowness bounds what a + false one can do — one fully identified finding, one cycle, **as far as distinct nonces + allow**: two cycles sharing or redrawing a nonce are indistinguishable to this record as to + every other, and a replayed decline can then bind to the wrong cycle — and a bound is not + the same as safety. **This record and the closure ordering are one contract**: it releases + the hold that ordering defines, and either present without the other is a partial adoption + that stops under the one-contract rule above. **What the body does not record, said here + rather than discovered:** an accepted finding, a stuck or two-tell surface and its + continue-or-stop answer leave no record of their own — from a closing body a reader can + infer the close and any declines, and nothing else about which suspensions the cycle passed + through. They live in the running session and, optionally, the working record; a cycle that + cannot recover them has no identity and starts a new cycle under the recovery rule above, + which costs passes and closes nothing — the reviewer re-reads the whole artifact each pass, + so an unrepaired accepted finding is re-raised and one outside the new cycle's scope is + re-asked as a membership stop, a repeated question rather than a silent close. ```` **Which rules the answer modifies** (story AC 3): the scope stop's resume sentence (`b12`, which gains "accepting or declining"), and the hold and the clean-pass definition (both in the -§3 block, where the answer's two directions are one rule). It modifies **no other rule**, and -none other carries the qualification: not the Blocker/Major-resolve duty (Mechanics, Severity: -"both must resolve", C:784 — **D5**: it stays unqualified because a declined finding never -enters its scope), not the floor, not the tells, not the stuck reading. A reader finding -"decline" at any rule outside those three has found a defect. +§3 block, where the answer's two directions are one rule). Those are the rules whose **closure +behaviour** the answer qualifies, and it qualifies **no other closure rule**: not the +Blocker/Major-resolve duty (Mechanics, Severity: "both must resolve", C:784 — **D5**: it stays +unqualified because a declined finding never enters its scope), not the floor, not the tells, +not the stuck reading. A reader finding "decline" qualifying any *closure* rule outside those +three has found a defect. The decline is also **named**, without qualifying closure, at the +transport, attribution, activation and threat sites listed next — the squash carry, the nonce +set and exemption, the "Both shipped records" count, the unknown-start fallback, the gate-off +surface, the closing-message carry and the one-contract paragraph — and a reader finding it +absent from any of those has found the opposite defect. **Existing sentences that must name it**, both copies, old → new. Line numbers are C's; W's are in the site map and re-read at execution. @@ -256,8 +336,12 @@ are in the site map and re-read at execution. provenance-line duty owed, the curve duty owed, and the nonce duties at their strictest…". NEW: "…at minimum floor 3, severity classified without the demotion, the provenance-line duty owed, the curve duty owed, the nonce duties at their strictest, **every suspension - binding, and decline records treated as absent so that no hold is released**…". `i12`'s - sentence stays as written; this is the addition it invites. + binding, and decline records not attributable to the nonce this fallback minted treated as + absent, so that no inherited hold is released — declines recorded under that nonce are + honoured**…". `i12`'s sentence stays as written; this is the addition it invites. The + bound in time matters: the fallback mints a post-rule nonce, and a decline the cycle then + makes and records under it is its own, not an inherited one — treating every decline as + absent would leave that cycle unable to release any hold it raises, **D7** unmet. 6. **The human-exception "which commit" rule** (C:988–990 / W:1172–1174, `h3`–`h6`). Unchanged in wording. The decline block carries its own rule, which deliberately differs — written when made and restated in every later body, not written once into the closing body — @@ -294,7 +378,13 @@ are in the site map and re-read at execution. direction has no release rule, a decline record without the ordering is a release with no hold to release, and a decline record missing from the squash carry or the nonce set is a record that cannot survive a merge or be attributed." The stop sentence that follows is - unchanged and now covers these states. + unchanged and now covers these states. Because `/workflow-init` can merge this paragraph + and the two blocks independently, each block also carries a reciprocal one-contract + sentence of its own (§3, last sentence; §4, the decline block), so a copy holding one block + without the other stops on the block it has, not only on a paragraph it may not have. A + rollback while a cycle is open under the new text needs no new mechanism: "a cycle already + running finishes under the rules it started with" and "A revert is itself a shipping commit + for the old rules" (C:153–154, C:165–166) already govern it. The shorter "Copy every record into the squash body" sentence inside the human-exception block (C:1004–1007) is generic and already covers a decline record; it is not edited. The "records @@ -342,9 +432,15 @@ rule is not waived by surfacing, and the loop resumes on a continue." Accounting kept; `c9`, `c10`, `c11` kept verbatim; `c12` kept (pointer form); `c13`, `c14` moved (the zero-finding exception; the Minor-below-floor case is the block's "clean pass at or above the derived floor" read in the negative); `c15`, `c16`, `c17` kept in the pointer sentence; `c18` -moved (no pass that surfaced a scope stop is credited as clean — narrowed to the scope stop, -because a stuck or two-tell surface on a below-floor clean pass leaves that pass's cleanliness -as it was); `c19` **replaced** — old: the loop resumes on whatever the user decides, +**replaced, narrowed** — old: a clearly-stuck surface credits no pass as clean; new: no pass +that surfaced a *scope stop* is credited as clean. Deliberate, with its authority: the story +states the fourth duty as "no pass carrying a surfaced *finding* counts as clean" (story §1), +and the clearly-stuck exit requires regenerating Blocker/Major findings, so a pass that +surfaces it is unclean by the first step of the ordering without any separate rule; the +two-tell stop surfaces no finding, so a Blocker/Major-free pass below the floor that trips +two tells on the raw counts is clean *as a pass* and still cannot close, because below the +floor nothing closes, and at or above it clean completion outranks the tells by **D2** — so +the narrowing opens no closure bypass in either region; `c19` **replaced** — old: the loop resumes on whatever the user decides, unconditionally; new: it resumes on a resuming answer and a stop leaves the suspension standing — because an unconditional resume is the stop-with-no-transition path AC 4 forbids; `c20` **dropped** — it argued that reading the exit as "stop instead of fixing" would put it @@ -434,27 +530,41 @@ Appended to the "**From pass 4 onward**" paragraph, both copies, after "…deman earlier pass had removed.": ``` -**Where an earlier pass's findings file is unavailable** — its slot path absent from -`.context/codex-reviews/` in this workspace, as after a fresh checkout, a cleared `.context/` -or a cycle resumed elsewhere; a file that is present but fails validation is an incomplete -pass, already excluded, and a path resolved against the wrong root is the existing stop — the -report first reads the cycle's working record if one exists and says whether it used it, then -states which of the three lines it computed from the files it has, names the ones it could -not and why, and says that the two-tell threshold is being read on that reduced record. The -sensitivity is reduced concretely: the trend and the require↔withdraw comparison read only -the passes present, so a rising count or a pair that spans a missing pass cannot be seen. It -is not a stop of its own, and it does not make the working record mandatory: a report that -says what it could not see is the duty; a report that invents the trend, or omits the line -without saying so, is the failure. +**Where an earlier pass's findings file is unavailable**, the report says so before it reads +anything, and it establishes what "unavailable" means by two checks. First the root: the +slots live in `.context/codex-reviews/` under the top-level directory of the checkout this +cycle is running in — the same root every pass call passes as `workingDirectory` — and a +report reading any other directory would count files it never wrote or miss files it did, so +a mismatch is a **wrong-root state**, reported as such and never as absent history; this +check is what detects it, nothing upstream does. Then, per earlier pass, one of three states, +told apart by what the slot holds: the slot path **absent** — after a fresh checkout, a +cleared `.context/`, a cycle resumed elsewhere — is unavailable history; the slot **present, +accepted when it ran, and now failing validation** is a formerly valid pass whose artifact is +unusable, and its pass number stays counted while its series read `?` in every comparison +that uses them, as the curve already admits, since a corrupted record does not un-run a pass; +a pass **known incomplete when it ran** is excluded, as today, and a present-but-invalid slot +whose acceptance nothing records is read as that case — the direction that costs a pass +rather than credits one. The report then reads the cycle's working record if one exists and +says whether it used it, states which of the three lines it computed from the files it has, +names the ones it could not and why, and says that the two-tell threshold is being read on +that reduced record — reduced concretely: the trend and the require↔withdraw comparison read +only the passes present, so a rising count or a pair that spans a missing or unusable pass +cannot be seen. It is not a stop of its own, and it does not make the working record +mandatory: a report that says what it could not see is the duty; a report that invents the +trend, or omits the line without saying so, is the failure. ``` This is **D10**. The trend and the require↔withdraw pair are derivable from the mandated findings files alone (the `fic2` record verified that derivation reproduces the reported figures), so unavailability is a property of the workspace, not of the format, and the answer -is disclosure rather than a new stop or a new mandatory artifact. The three causes named are -partitioned by what the agent observes — path absent, file present but invalid, path resolved -against the wrong root — and only the first is this state; the other two already have their -own rules, which is why the sentence points at them rather than restating them. +is disclosure rather than a new stop or a new mandatory artifact. The state is partitioned by +what the agent observes, per `docs/prompt-standards.md` item 10: the root check first, +because every other observation is meaningless against the wrong directory; then absent, +present-but-unusable (once accepted), and known-incomplete, each with its own action. The +middle state exists because validation happens when a pass is accepted and a file can be +truncated or replaced afterwards; reading it as incomplete would retroactively uncount a +pass, reading it as usable would feed a broken record into the trend — `?` is the curve +grammar's word for exactly that. --- @@ -467,8 +577,10 @@ case it served no longer exists: every cycle started after the parent's rules bi nonce, and a cycle whose start cannot be established mints one rather than claiming `none (pre-rule)` (`i11`). A rule for a no-nonce cycle would legislate for an unreachable state, which is AC 1's prohibition. The `rle` naming that cycle used stays what its closing body -recorded it as — a plan-local exception under the old rules. This section and the story's §5 -are the record; no prompt text changes. +recorded it as — a plan-local exception under the old rules. The durable prior record is Plan +C's two drop notes (the Task 19 heading at line 970 and the Task 20 heading at line 1047 of +that plan, each followed by its note) and Task 23's second point (line 1298), which records +the `rle` exception; this section is the record of the dissolution; no prompt text changes. --- @@ -501,19 +613,28 @@ no decline record, and the (g) count is the one site where the parent is present removes it. No count against the parent is claimed as "contradictory". **The named verification of the risk path** (story AC 4). Walk every stop the shipped text -names — membership stop, question stop, clearly-stuck exit, two-tell stop, a hold awaiting its -answer, accept, decline, a stop answer, two or three suspensions at once, a below-floor clean -pass, a zero-finding pass, the unknown-start fallback — and write the **next-state table**: for -each, the input that ends it and the state the cycle is in afterwards, citing the shipped line -the row reads, in both copies. The table lives in the plan and is quoted by the closing commit -body, not here. What would be observed if the claim "no path leaves a cycle unable to close -and unable to suspend" were false: a row whose next state is the same stop with no input -consumed — the shape the parent cycle shipped once (a stop whose only answer resumed a cycle -that immediately stopped again). The wiring can produce it because every row is filled from -the shipped text, not from this spec, and the inputs the `fic2` instrument omitted — the user's -answer, and the decline — are input columns here; so is the stop answer, whose next state is a -standing suspension by design and must read as one in the table, not as the same stop -re-raised. +names — membership stop, question stop, a finding carrying both triggers, clearly-stuck exit, +two-tell stop, a hold awaiting its answer, accept, decline, a stop answer, two or three +suspensions at once, a below-floor clean pass, a zero-finding pass, the unknown-start +fallback — **and the stateful transitions** — a declined finding re-raised matching on all +five fields, and re-raised with one field changed; the fix set broadened to include a +declined finding; recovery after an accept and after a stop with the session lost (no +identity → new cycle); a `full` Gate-B pass with one branch clean and the other carrying an +in-set Blocker; a rollback with a cycle open under the new rules; a copy adopting the +ordering without the decline block, and the reverse — and write the **next-state table**: +for each row, the **record state** (which decline records the bodies carry, which working +record exists) and the **governing-scope state** (the fix set as currently assigned) as +explicit input columns beside the user's answer, then the input that ends the row and the +state the cycle is in afterwards, citing the shipped line the row reads, in both copies. The +table lives in the plan and is quoted by the closing commit body, not here. What would be +observed if the claim "no path leaves a cycle unable to close and unable to suspend" were +false: a row whose next state is the same stop with no input consumed — the shape the parent +cycle shipped once (a stop whose only answer resumed a cycle that immediately stopped again) +— or a row that closes with an in-set Blocker standing. The wiring can produce it because +every row is filled from the shipped text, not from this spec, and the inputs the `fic2` +instrument omitted — the user's answer, and the decline — are input columns here; so is the +stop answer, whose next state is a standing suspension by design and must read as one in the +table, not as the same stop re-raised. **What this is not.** It is not the `fic2` decision matrix: that instrument scored N states against old and new text with an expected output each, and Gate B found two defects in the From 0a5205f68f8798b7af607fb07a43d64e4ad9073a Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Thu, 10 Sep 2026 15:07:41 +0200 Subject: [PATCH 005/181] =?UTF-8?q?docs(specs):=20loop-rule=20consolidatio?= =?UTF-8?q?n=20=E2=80=94=20Gate-A=20pass=203=20revision?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Pass 3 returned 12 findings (4 Blocker, 6 Major, 2 Minor). Eleven fixed, one Minor collected, and the mandatory-record demand dismissed a second time with its residual stated instead. The ordering gains its third branch — a pass that neither closes nor suspends continues — plus a resolve duty that a validated dismissal can discharge, a decline that suppresses the membership trigger only, and both stuck conditions reading the pre-ceiling severity field. Q6 decides that a gap does not break the series. The severity split joins the one-contract coupling in both shipped blocks. Held at 691 lines: the §3 commentary and the §4 prose that duplicated the shipped blocks paid for the additions. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- ...26-09-10-loop-rule-consolidation-design.md | 415 +++++++++--------- 1 file changed, 207 insertions(+), 208 deletions(-) diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index c3b0b5d..cb4dfa0 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -1,6 +1,6 @@ # §5 loop-rule consolidation: one closure ordering — Design -**Date:** 2026-09-10 · **Status:** DRAFT — Gate A running (pass 2 revised) +**Date:** 2026-09-10 · **Status:** DRAFT — Gate A running (pass 3 revised) **Story:** `docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md` **Profile:** read from that header at every pass, never from here — the header is the only writable copy, and a value copied into this file would be a remembered value. @@ -29,9 +29,8 @@ directions; the answer to what a severity demotion does to the loop-health count to the pass-4 report when prior-pass history is unavailable; and the deferred slot discriminator, dissolved rather than shipped. -**What does not.** The pass floor and severity semantics (the parent shipped them and this spec -reads them as given); hook code; and the items the story's §2 parks — the pass-counter anomaly, -the CodeRabbit plan-metadata contradiction, and the fixture-per-predicate question. +**What does not.** The pass floor and severity semantics — the parent shipped them and this +spec reads them as given — and everything §11 lists as parked or out of scope. --- @@ -43,14 +42,12 @@ cycle established are read as given: a pass's cleanliness is a fact about what t and is **never rewritten** — an answer changes whether the *cycle* may close; and the findings files establish the **inventory** of findings, not their resolutions. -This spec adds what the table does not settle, each decided in the section that uses it: the -evaluation order and the definition of a clean pass under a decline, read at effective -severity over every branch file of the logical pass (§3); the classification of the four -duties (§3); the two triggers of the scope stop and what each answer does (§3); what the -user's answer to a stuck or two-tell surface produces, including a stop (§3); the wording, -the nonce status and the recording point of the decline record, and what happens to the -state the body does not record when a cycle loses it (§4); and the raw-severity rule for -the health measures (§5, passage (g)). +This spec adds what the table does not settle — the evaluation order and the field and file +set every predicate reads; the four duties' classification; the scope stop's two triggers and +what each answer does; what a stuck or two-tell answer produces, a stop included; the decline +record's wording, attribution and recording point, and what a cycle does when it loses the +state no record carries; and the raw-severity rule for the health measures — each decided in +the section that uses it (§3, §4, and §5 passage (g)). --- @@ -67,35 +64,26 @@ and a stated order is what stops them qualifying each other. Every predicate her validated findings file **or files** of the logical pass as one set — a `full` Gate-B pass has two, and one branch alone is already an incomplete pass — at **effective** severity, after the Mechanics severity ceiling, because cleanliness is about what the cycle must repair and -the ceiling is what decides that; only the named health measures — the per-pass counts, the -clusters and the tells — read the reviewer-written field before it (Mechanics, Severity). +the ceiling is what decides that; only the **health measures** read the reviewer-written field +before it (Mechanics, Severity): the per-pass counts, the clusters, the tells, and **both** +conditions of the stuck reading — its Blocker curve and its regenerating Blocker or Major +findings — since a predicate reading one field for half of itself could not be read at all. **First, clean completion.** A pass is clean when it carries no Blocker or Major at effective severity that is **in the assigned fix set**, and it surfaced no scope stop; a pass with **zero** findings is clean whatever the floor. A finding matched by a decline recorded in -this cycle is outside the set by that decision and is not surfaced when re-raised (Mechanics, -the decline record) — a decline keeps a finding out and never excuses one that is in, so a -declined finding the set later comes to include owes resolution like any other. **A clean -pass at or above the derived floor, or a zero-finding pass, closes the cycle**, under the -final-acceptance preconditions the floor section already states and this ordering does not -restate — the cited set and every profile re-read before the pass is accepted as final, a -header or profile that changed during it making the pass not final and costing the further -pass that section requires. A plateau or tells present on the closing pass go into the +this cycle is outside the set by that decision (Mechanics, the decline record) — a decline +keeps a finding out and never excuses one that is in, so a declined finding the set later +comes to include owes resolution like any other. **A clean pass at or above the derived floor, +or a zero-finding pass, closes the cycle**, under the final-acceptance preconditions the floor +section already states and this ordering does not restate — the cited set and every profile +re-read before the pass is accepted as final, a header or profile that changed during it +making the pass not final and costing the further pass that section requires. A plateau or +tells present on the closing pass go into the closing report and never block it, because reporting "will not converge" on a converged loop is a false report, and the clearly-stuck paragraph says the same of its own exit. **Nothing else closes a cycle.** -**The four standing duties, classified.** The **derived floor** is a precondition on -closure: it gates closing, and is discharged by the count of valid logical passes reaching -the floor with the last of them clean, or by the zero-finding exit. The **Blocker/Major-resolve -duty** is a precondition on closure and on any pass being clean: it gates both, and is -discharged by repair of every in-set Blocker and Major followed by a pass that finds none. -The **hold** a surfaced finding places on closure is part of the ordering: it gates closing -while it stands, and is discharged by the answers that finding requires, below. -**No-clean-credit** — no pass that surfaced a scope stop is credited as clean — is part of -the ordering and a fact about that pass, discharged by nothing: a later pass is judged on its -own findings, so a decline never closes the cycle on the pass that surfaced the finding. - **Second, only a pass that is not a clean completion can suspend** — that order is what makes "clean completion outranks the two-tell stop" executable rather than asserted. Three suspensions, by the names their paragraphs use: the **scope stop**, with two triggers — a @@ -107,8 +95,32 @@ to one pass: **one surface, every reason reported, every question asked**, becau left out is a decision made by omission. A finding surfaced solely by the stuck or two-tell reading is not a scope-stop finding; one that is also outside the set, or also opens a question, takes the scope stop's answers at that same surface — it is not asked twice. A -re-raised finding that a decline of this cycle matches on all five fields raises no hold and -no stop; a difference in any field, or genuine uncertainty, is a new finding and a new stop. +re-raised finding that a decline of this cycle matches on all five fields raises **no +membership stop and no membership hold**, and that alone: the five recorded fields record no +question, so a **question stop still fires** on it unless that same question has already been +answered in this cycle. A difference in any of the five fields, or genuine uncertainty, is a +new finding and a new stop of either kind. + +**Third, a pass that neither closes nor suspends continues** — the loop runs another pass on +the revised artifact. This is the ordinary case and not a residue: a clean pass below the +floor lands here, and so does a pass whose only findings are Minors and Nits, collected and +never iterated. It is a branch and not an inference, because "does not close" read alone says +nothing about whether to run again. + +**The four standing duties, classified.** The **derived floor** is a precondition on +closure: it gates closing, and is discharged by the count of valid logical passes reaching +the floor with the last of them clean, or by the zero-finding exit. The **Blocker/Major-resolve +duty** is a precondition on closure and on any pass being clean: it gates both, and is +discharged, for every in-set Blocker and Major, by a validated repair **or a validated +dismissal carrying its one-line why** — the advisory rule above, unchanged — followed by a +pass that finds none. **A dismissal is not a decline**: a dismissal is the author's judgement +that the finding is not true of the artifact, a decline is the user's decision that a true +finding stays outside the fix set, and only the second is an answer at a membership stop. +The **hold** a surfaced finding places on closure is part of the ordering: it gates closing +while it stands, and is discharged by the answers that finding requires, below. +**No-clean-credit** — no pass that surfaced a scope stop is credited as clean — is part of +the ordering and a fact about that pass, discharged by nothing: a later pass is judged on its +own findings, so a decline never closes the cycle on the pass that surfaced the finding. **What a suspension asks, and what ends it.** Only a scope stop raises a **hold**, on its finding, and the hold ends when every answer that finding requires has been given — one for a @@ -124,12 +136,14 @@ out-of-set finding that opened the question is a membership stop as well and tak decline. **Decline is available only at a membership stop.** The stuck and two-tell readings raise no hold: each asks one question, **continue or stop**. Continue resumes the loop on the artifact as revised and the fix set as the governing artifacts now assign it — where several -plans or stories govern one cycle, the union of the scopes they assign plus the obligations -already accepted — and a finding that a narrowing has put outside the set surfaces at the -next pass as a membership stop. **Stop leaves the suspension standing**: the cycle stays open -under its nonce, resumable by a later continue, nothing it wrote is a closing commit, and an -open cycle nobody resumes is a human's to resolve, exactly as the nonce rules already say of -open cycles. +plans or stories govern one cycle, the union of the scopes they assign — **plus the accepted +obligations this cycle can still recover**, from the running session or from its own records. +One it cannot recover is not in the set, and the reviewer raises its finding again on the next +pass like any other, which is the ordinary route and not a special one. A finding that a +narrowing has put outside the set surfaces at the next pass as a membership stop. **Stop +leaves the suspension standing**: the cycle stays open under its nonce, resumable by a later +continue, nothing it wrote is a closing commit, and an open cycle nobody resumes is a human's +to resolve, exactly as the nonce rules already say of open cycles. **Composition, and what cannot happen.** Every question is answered on its own, and the loop resumes only when every answer resumes it — accept or decline at a membership stop, a decision @@ -139,97 +153,53 @@ running. Two states cannot co-occur, and no rule ranks them: clean completion an stop, since a pass that surfaced one is not clean; and a zero-finding pass and any suspension, since it has nothing to surface, nothing regenerating, no cluster and no require↔withdraw pair. The Gate-B triviality skip is outside this ordering: a skipped cycle -runs no passes and ends by its own rule. **This ordering and the decline record in Mechanics -are one contract**: the hold's declined direction is released by that record and by nothing -else, so either present without the other is a partial adoption that stops under the -one-contract rule there. +runs no passes and ends by its own rule. **This ordering, the decline record in Mechanics and +the severity rule's raw-versus-effective split are one contract**: the hold's declined +direction is released by that record and by nothing else, and the split says which field each +predicate here reads, so any one of the three present without the others is a partial adoption +that stops under the one-contract rule there. ``` -Why each part is there, briefly. The evaluation order is the answer to the objection that a -ranking between clean completion and the two-tell stop cannot fire if the stop can make the -pass unclean: clean completion is read first, so a suspension is only ever evaluated on a pass -that did not close. Reading every predicate over the union of the logical pass's branch files -keeps a `full` Gate-B pass from closing on one clean branch while the other carries an in-set -Blocker; reading them at effective severity keeps the resolve duty and the clean-pass test on -one field and leaves the health measures — passage (g) — on the other, so a demoted finding -cannot keep a pass unclean while owing no repair. The clean predicate names the fix set and -nothing else, so a decline is only ever a fact about membership: it cannot be read as -excusing an in-set finding, which is the gate-off route a fabricated record would otherwise -open. The close sentence points at the floor section's final-acceptance preconditions -("again before a clean pass is accepted as the cycle's final pass", C:116–118; "Any profile -change costs at least one further pass", C:760–764) instead of restating them, so a closing -pass under a changed header stays not-final as that section already says. The duties -paragraph is story AC 2 — each duty labelled precondition or ordering participant, with what -it gates and what discharges it. The first paragraph is **D1**, **D2** and **D3**, the three -suspensions named as their paragraphs name them (`b11`/`b13`, `c4`, `e7`); "nothing else -closes a cycle" exists for story AC 4. The two triggers of the scope stop are `b11` -(membership) and `b13` (question) read separately, because an in-set finding that opens a -question can neither "join" nor "stay outside" the set and needs its own answer; a finding -carrying both triggers requires both answers before its hold ends, so one answer cannot -release a hold the other question still holds. The declined re-raise raises no stop because -a decline that re-asked its own question every pass would bind for nothing (**D7**). The -composition paragraph is AC 1's demand that every reachable conflict be covered and every -unreachable one named with its reason (the `fic2` record documents the instrument that -legislating built). The stop answer is the one new state this spec supplies beyond the table -— a standing suspension, resumable, never a close — and it uses only mechanisms already -shipped: resumption on the revised artifact (`b18`), the membership stop (`b11`), and the -nonce rules' existing treatment of open cycles ("they stay open, keep their own nonces, and -are a human's to resolve", C:423–424). **D3** keeps the clearly-stuck paragraph's own -precedence sentence verbatim (§5, passage (c)); the block cites it rather than restating its -consequence clause. The closing contract sentence is reciprocal with one in the decline -block, so a copy that adopts one block without the other carries its own stop. - -What the block deliberately does not do: it does not define the floor, the tells, or the -stuck reading — those stay in their paragraphs ("exactly one definition of the floor must be -present", C:176; "the other closure and stop predicates … not required to derive from the -floor", C:181–185). It does not restate any rule it does not own (`a13`). +Why the shape, where the block's own sentences do not already carry it. The evaluation order +answers the objection that a ranking between clean completion and the two-tell stop cannot +fire if the stop can make the pass unclean: clean completion is read first, so a suspension is +only ever evaluated on a pass that did not close. The three branches are story AC 4 — every +pass has a next state, and the third exists because "does not close" is not an instruction. +The duties paragraph is AC 2; the composition and cannot-co-occur sentences are AC 1's demand +that every reachable conflict be covered and every unreachable one named with its reason. The +scope stop's two triggers are `b11` (membership) and `b13` (question) read separately, because +an in-set finding that opens a question can neither join nor stay outside the set and needs +its own answer — which is also why a matching decline suppresses the membership trigger only. +The preconditions the close sentence points at are at C:116–118 ("again before a clean pass is +accepted as the cycle's final pass") and C:760–764 ("Any profile change costs at least one +further pass"). The stop answer is the one new +state this spec supplies beyond the table, and it reuses only shipped mechanisms: resumption +on the revised artifact (`b18`), the membership stop (`b11`), and the nonce rules' treatment +of open cycles ("they stay open, keep their own nonces, and are a human's to resolve", +C:422–424). **D3** keeps the clearly-stuck paragraph's own precedence sentence verbatim (§5, +passage (c)), so the block cites it instead of restating its consequence clause. The block +defines neither the floor nor the tells nor the stuck reading — those stay in their paragraphs +("exactly one definition of the floor must be present", C:176) — and restates no rule it does +not own (`a13`). --- ## 4. The decline rule and record -**The rule.** A decline is the user's recorded answer at a membership stop — that the surfaced -finding stays outside the assigned fix set — and it is available for no other finding -(**D6**): not for an in-set Blocker or Major, and not as the answer to a question stop, a stuck -or two-tell surface, or any obligation. It binds for the remainder of its cycle and has no -effect in any later one (**D7**); it must be an explicit, attributable decision on that specific -finding — never silence, a general remark about scope, or an inference (**D8**). A later pass's -finding is **the same finding** when all five of location, defect, severity, consequence and -suggested fix match (**D9b**); the match is read on meaning, not on bytes, because a reviewer -rewrites its sentences between passes. A matching finding raises no hold and no membership -stop when re-raised — it is not "surfaced", so the pass stays eligible for clean and the -decline is idempotent for its cycle (**D7**); any difference, severity included, or any genuine -uncertainty, makes it a new finding and a new stop, and the hold applies. A decline releases the -hold and never qualifies the Blocker/Major-resolve duty, which the finding never reached -(**D5**); it keeps a finding out and never excuses one that is in — a declined finding the fix -set later comes to include, through an accepted broadening or a governing artifact -re-assigning scope, owes resolution like any other, which is why the clean predicate in §3 -names the set and not the decline. - -**When it is recorded, and that it is recorded before the loop resumes.** When made, into the -cycle's next commit body on the branch — a spec or plan revision commit for a Gate-A cycle, the -`WIP:` amend for a Gate-B cycle — and that commit is made **before the next pass runs**: no pass -runs on a decline that no commit body on the branch carries, because a decline held only in a -session is one compaction away from a re-asked question, and **D7**'s remainder-of-cycle binding -is only as durable as its transport. The advisory working record may carry it meanwhile and -does not bind. It is restated in every later body of that cycle, the closing one included. A -cycle that resumes with no commit body carrying a decline treats it as absent and the hold -applies again — the reading the unknown-start fallback gives, and the safe direction: a lost -decline costs a repeated question, an invented one releases a hold nobody chose to release. -**D9**'s "closing commit body" is therefore the last of the bodies that carry it, not the first. - -**What the body does not record, and what a cycle does when it loses it.** An accepted -finding, a standing stuck or two-tell suspension and its continue-or-stop answer have no record -of their own; they live in the running session and, optionally, in the working record. No new -mandatory record is added for them — **D10** refused one for the same class of state, and the -existing rules already answer the loss: a cycle that cannot recover them has no identity and -starts a new cycle under the no-identity rule ("**No candidate, disagreeing sources, or more -than one candidate → no identity: start a new cycle**", C:419–420), which costs passes and -closes nothing. The reviewer re-reads the whole artifact each pass, so an unrepaired accepted -finding is re-raised, and one the new cycle's assigned scope does not include is re-asked as a -membership stop — a repeated question rather than a silent close, the safe direction. The -residual, stated: a finding the reviewer fails to re-raise is lost exactly as any missed -finding is, which this change neither creates nor removes. +**The rule ships as the block below**, which is the text and not a summary of it; this section +adds only what the block does not carry. The decisions it implements: availability at a +membership stop and nowhere else is **D6**; the remainder-of-cycle binding **D7**; the +explicit attributable decision **D8**; the five-field sameness test **D9b**; the +unverified-assertion reading **D9c**; the transport **D9**, whose "closing commit body" is +therefore the *last* of the bodies that carry the record, not the first. That a decline never +qualifies the Blocker/Major-resolve duty is **D5**, and it is why the clean predicate in §3 +names the fix set and not the decline: a decline is only ever a fact about membership, so it +cannot be read as excusing an in-set finding — the gate-off route a fabricated record would +otherwise open. That accepted obligations and standing suspensions get **no new mandatory +record** is deliberate: **D10** refused one for the same class of state, and the existing +no-identity rule ("**No candidate, disagreeing sources, or more than one candidate → no +identity: start a new cycle**", C:419–420 / W:613–614) already answers the loss in the safe +direction, at the cost of passes. **The block that ships**, in Mechanics, placed **immediately after the whole human-exception passage** — after its last paragraph "**What the record is worth.**" (C:1028–1033 / @@ -249,11 +219,14 @@ two-space indent. Verbatim, byte-identical in both copies: remainder of this cycle**, with no effect in any later one; it never qualifies the Blocker/Major-resolve duty, which the finding never reached. A later pass raises **the same finding** when all five of location, defect, severity, consequence and suggested fix match, - read on meaning rather than bytes, since a reviewer rewrites its sentences between passes; - a matching finding raises no hold and no stop, so a decline is not re-asked every pass; - **any difference — severity included — or any genuine uncertainty makes it a new finding, - and the hold applies.** A decline keeps a finding out and never excuses one that is in: a - declined finding the fix set later comes to include owes resolution like any other. + read on meaning rather than bytes, since a reviewer rewrites its sentences between passes. + A matching finding raises **no membership stop and no membership hold**, so a decline is not + re-asked every pass — **and that alone**: the five fields record no question, so a **question + stop** still fires on it unless that same question has already been answered in this cycle + (the closure ordering above). **Any difference — severity included — or any genuine + uncertainty makes it a new finding, and the hold applies.** A decline keeps a finding out and + never excuses one that is in: a declined finding the fix set later comes to include owes + resolution like any other. ``` Declined: · · cycle @@ -268,7 +241,11 @@ two-space indent. Verbatim, byte-identical in both copies: revision commit for a Gate-A cycle, the `WIP:` amend for Gate B — **and that commit is made before the next pass runs**: no pass runs on a decline that no commit body on the branch carries, because a decline held only in a session is one compaction away from a re-asked - question; the advisory working record may carry it meanwhile and does not bind. It is + question; the advisory working record may carry it meanwhile and does not bind. Where a + Gate-A cycle has no artifact revision to carry it — the declined finding was that pass's + only one — the record goes in an **empty commit of its own**, the destination the + human-exception rule above already blesses: "An empty commit carrying only the record is a + legitimate destination". It is restated in every later body of that cycle, the closing one included, because a body that drops it releases nothing and re-asks the question. A cycle that resumes with no commit body carrying a decline **treats it as absent and the hold applies again** — the reading the @@ -280,9 +257,11 @@ two-space indent. Verbatim, byte-identical in both copies: false one can do — one fully identified finding, one cycle, **as far as distinct nonces allow**: two cycles sharing or redrawing a nonce are indistinguishable to this record as to every other, and a replayed decline can then bind to the wrong cycle — and a bound is not - the same as safety. **This record and the closure ordering are one contract**: it releases - the hold that ordering defines, and either present without the other is a partial adoption - that stops under the one-contract rule above. **What the body does not record, said here + the same as safety. **This record, the closure ordering and the severity rule's + raw-versus-effective split are one contract**: this record releases the hold that ordering + defines, and that split says which severity field each of its predicates reads, so any one + of the three present without the others is a partial adoption that stops under the + one-contract rule above. **What the body does not record, said here rather than discovered:** an accepted finding, a stuck or two-tell surface and its continue-or-stop answer leave no record of their own — from a closing body a reader can infer the close and any declines, and nothing else about which suspensions the cycle passed @@ -343,10 +322,9 @@ are in the site map and re-read at execution. makes and records under it is its own, not an inherited one — treating every decline as absent would leave that cycle unable to release any hold it raises, **D7** unmet. 6. **The human-exception "which commit" rule** (C:988–990 / W:1172–1174, `h3`–`h6`). Unchanged - in wording. The decline block carries its own rule, which deliberately differs — written - when made and restated in every later body, not written once into the closing body — - because a decline must survive to the next pass while the human exception only has to - survive to history. + in wording. The decline block carries its own rule, deliberately different — written when + made and restated in every later body — because a decline must survive to the next pass + while the human exception only has to survive to history. 7. **The gate-off surface** (C:143–151 / W:350–358). OLD tail: "…silencing reminders; or not running a pass and reporting that it ran." NEW tail: "…silencing reminders; not running a pass and reporting that it ran; **or recording a decline nobody made, or one on an in-set @@ -371,20 +349,36 @@ are in the site map and re-read at execution. that say what a cycle owes when its starting rules cannot be established** depend on one another," NEW opening: "**These records are one contract, and a partial adoption breaks it.** The nonce, the slot naming, the provenance line, the curve, this carry rule, **the - closure ordering and the decline record**, and **the unknown-start activation semantics - that say what a cycle owes when its starting rules cannot be established** depend on one - another," and after "and a carry rule naming records a project does not produce is - inert." (C:885) add: "an ordering without the decline record is a hold whose declined - direction has no release rule, a decline record without the ordering is a release with no - hold to release, and a decline record missing from the squash carry or the nonce set is a - record that cannot survive a merge or be attributed." The stop sentence that follows is - unchanged and now covers these states. Because `/workflow-init` can merge this paragraph - and the two blocks independently, each block also carries a reciprocal one-contract - sentence of its own (§3, last sentence; §4, the decline block), so a copy holding one block - without the other stops on the block it has, not only on a paragraph it may not have. A - rollback while a cycle is open under the new text needs no new mechanism: "a cycle already - running finishes under the rules it started with" and "A revert is itself a shipping commit - for the old rules" (C:153–154, C:165–166) already govern it. + closure ordering, the decline record and the severity rule's raw-versus-effective split**, + and **the unknown-start activation semantics that say what a cycle owes when its starting + rules cannot be established** depend on one another," and after "and a carry rule naming + records a project does not produce is inert." (C:885) add: "an ordering without the decline + record is a hold whose declined direction has no release rule; a decline record without the + ordering is a release with no hold to release; an ordering whose raw-versus-effective split + has no counterpart in the severity rule, or a severity rule still calling that question + unsettled and mandating a stop beside an ordering that decides it, is two answers to one + question; and a decline record missing from the squash carry or the nonce set is a record + that cannot survive a merge or be attributed." The stop sentence that follows is unchanged + and now covers these states. Because `/workflow-init` can merge this paragraph and the + pieces independently, the two blocks each carry a reciprocal one-contract sentence of their + own (§3, last sentence; §4, the decline block), so a copy holding one without the others + stops on the text it has, not only on a paragraph it may not have. A rollback while a cycle + is open under the new text needs no new mechanism: "a cycle already running finishes under + the rules it started with" and "A revert is itself a shipping commit for the old rules" + (C:153–154, C:165–166) already govern it — and where a rollback leaves that cycle unable to + establish the rules it started with, the unknown-start fallback is the answer, extended by + item 5 with the loop rules. The residual, stated rather than papered over: the rule text a + cycle ran under is recorded nowhere, so "finishes under the rules it started with" is + recoverable only while that text is still present, and the fallback is what covers the rest. +10. **The no-identity rule's aftermath** (C:422–424 / W:616–618). OLD: "**Starting a new cycle + does not close, adopt or retire the cycles those candidates belong to** — they stay open, + keep their own nonces, and are a human's to resolve; the new cycle simply does not claim + them." NEW: "**Starting a new cycle does not close, adopt or retire the cycles those + candidates belong to** — they stay open, keep their own nonces, and are a human's to + resolve; the new cycle simply does not claim them, **and names them in its first pass + report**, so the human this rule makes responsible learns they exist." It adds no record: + the report names state the workspace already holds, and a cycle left open that nobody is + told about is the one shape this rule's "a human's to resolve" cannot reach. The shorter "Copy every record into the squash body" sentence inside the human-exception block (C:1004–1007) is generic and already covers a decline record; it is not edited. The "records @@ -395,9 +389,8 @@ conditional, like the human exception, and is not added. ## 5. Edits to the existing passages, with the old-conditions accounting -Ids are the inventory's. "Kept" means the sentence stays in its passage; "moved" means it now -lives in the §3 block and leaves the passage; "replaced" means the condition changes and says -how; "dropped" carries its reason. +Ids are the inventory's. "Kept" = the sentence stays; "moved" = it now lives in the §3 block; +"replaced" = the condition changes, and says how; "dropped" carries its reason. **(a) The floor paragraphs** (C:72–136 / W:279–343) — **trim** `a17`–`a19` and point at the block. OLD: "Your final pass must be clean — if the pass at the floor still finds @@ -410,8 +403,9 @@ prohibition on restating is what the pointer form obeys. **(b) What a loop absorbs** (C:195–223 / W:402–426) — **two sentence edits**, the triggers stay. `b12` OLD: "…and it resumes the moment the user says whether the set now includes it." NEW: -"…and it resumes the moment the user says whether the set now includes it — accepting or -declining it, under the closure ordering above." `b17`–`b18` OLD: "Stopping this way is **not +"…and it resumes once the user has said whether the set now includes it — accepting or +declining it — together with every other answer that pass's suspensions require, any decline +among them recorded as Mechanics requires, under the closure ordering above." `b17`–`b18` OLD: "Stopping this way is **not an exit from the gate**: the floor, the Blocker/Major filter and the clean-final-pass rule all stand, and the loop resumes on the revised artifact once the question is answered — what the stop prevents…" NEW: "Stopping this way is a **suspension** under the closure ordering above, @@ -428,19 +422,22 @@ satisfied the clean-final-pass rule — collect the Minors and Nits and close nothing closes**" to the end of "…and then nothing could satisfy both.": replaced by "**Below the floor nothing closes**, exactly as the closure ordering above says. This exit is a **suspension** under that ordering: you surface with the findings still open, the resolve -rule is not waived by surfacing, and the loop resumes on a continue." Accounting: `c1`–`c8` -kept; `c9`, `c10`, `c11` kept verbatim; `c12` kept (pointer form); `c13`, `c14` moved (the -zero-finding exception; the Minor-below-floor case is the block's "clean pass at or above the -derived floor" read in the negative); `c15`, `c16`, `c17` kept in the pointer sentence; `c18` +rule is not waived by surfacing, and the loop resumes when every answer that pass's +suspensions require has been given." Accounting: `c1`–`c8` +kept; `c9`, `c10`, `c11` kept verbatim; `c12` kept (pointer form); `c13` moved (the +zero-finding exception); `c14` moved **to the ordering's continue branch** — the +Minor-below-floor pass that "keeps looping" is precisely a pass that neither closes nor +suspends, and the branch states it rather than leaving it to be read out of a negation, which +is what finding 1 of pass 3 caught; `c15`, `c16`, `c17` kept in the pointer sentence; `c18` **replaced, narrowed** — old: a clearly-stuck surface credits no pass as clean; new: no pass that surfaced a *scope stop* is credited as clean. Deliberate, with its authority: the story states the fourth duty as "no pass carrying a surfaced *finding* counts as clean" (story §1), -and the clearly-stuck exit requires regenerating Blocker/Major findings, so a pass that -surfaces it is unclean by the first step of the ordering without any separate rule; the -two-tell stop surfaces no finding, so a Blocker/Major-free pass below the floor that trips -two tells on the raw counts is clean *as a pass* and still cannot close, because below the -floor nothing closes, and at or above it clean completion outranks the tells by **D2** — so -the narrowing opens no closure bypass in either region; `c19` **replaced** — old: the loop resumes on whatever the user decides, +and the clearly-stuck exit requires regenerating Blocker/Major findings, so a pass surfacing +it is already unclean by the ordering's first step without a separate rule. The two-tell stop +surfaces no finding, so a Blocker/Major-free pass below the floor tripping two tells is clean +*as a pass* and still cannot close — below the floor nothing closes, and at or above it clean +completion outranks the tells by **D2** — so the narrowing opens no bypass in either region; +`c19` **replaced** — old: the loop resumes on whatever the user decides, unconditionally; new: it resumes on a resuming answer and a stop leaves the suspension standing — because an unconditional resume is the stop-with-no-transition path AC 4 forbids; `c20` **dropped** — it argued that reading the exit as "stop instead of fixing" would put it @@ -453,7 +450,8 @@ applies. kept, unchanged. **(e) The five tells** (C:263–268 / W:467–472) — **add** one pointer sentence after `e10`: -"This stop is a **suspension** under the closure ordering above." `e1`–`e11` kept. +"This stop is a **suspension** under the closure ordering above, and the loop resumes when +every answer that pass's suspensions require has been given." `e1`–`e11` kept. **(f) The two rules above do not compete** (C:275–287 / W:473–483) — **unchanged**. "The two rules above" still names the absorb rule and the stuck reading; the block sits before both and @@ -493,9 +491,9 @@ and the decline block says which one item of that list it *is* the answer to. **(j) The squash carry** (C:892 / W:1076) — **extend** (§4 item 1). `j1`–`j4` kept. -Also touched, outside the inventoried passages: the nonce set, the nonce exemption, the "Both -shipped records" count, the gate-off list, the closing-message sentence and the one-contract -paragraph (§4 items 2, 3, 4, 7, 8, 9). +Also touched, outside the inventoried passages (§4 items 2, 3, 4, 7, 8, 9, 10): the nonce set +and exemption, the "Both shipped records" count, the gate-off list, the closing-message +sentence, the one-contract paragraph, the no-identity rule's aftermath. --- @@ -517,10 +515,8 @@ and W. The pre-existing divergences the inventory found are handled as follows: The parity check at execution: extract each edited passage from both files by its lead phrase and `diff` them; the only differences permitted are the rows marked as-is above, and the (g) row is gone. Any other difference is a defect, not a wording choice. The check's result is part -of the evidence entry (§9). - -The decline block and the human-exception block it follows are byte-identical between copies -today (the site map confirmed C:840–1033 = W:1024–1217), and stay so. +of the evidence entry (§9). The Mechanics region the decline block joins is byte-identical +between the copies today (the site map confirmed C:840–1033 = W:1024–1217) and stays so. --- @@ -549,7 +545,11 @@ says whether it used it, states which of the three lines it computed from the fi names the ones it could not and why, and says that the two-tell threshold is being read on that reduced record — reduced concretely: the trend and the require↔withdraw comparison read only the passes present, so a rising count or a pair that spans a missing or unusable pass -cannot be seen. It is not a stop of its own, and it does not make the working record +cannot be seen. **A gap does not break the series**: the consecutive *available* pass numbers +compare, and the report **names the missing pass numbers** so a reader can see which +comparison spans one. Refusing to compare across a gap would silence tell detection exactly +where the record is thinnest, and the tells read the direction of the loop rather than any one +adjacent pair. It is not a stop of its own, and it does not make the working record mandatory: a report that says what it could not see is the duty; a report that invents the trend, or omits the line without saying so, is the failure. ``` @@ -557,14 +557,11 @@ trend, or omits the line without saying so, is the failure. This is **D10**. The trend and the require↔withdraw pair are derivable from the mandated findings files alone (the `fic2` record verified that derivation reproduces the reported figures), so unavailability is a property of the workspace, not of the format, and the answer -is disclosure rather than a new stop or a new mandatory artifact. The state is partitioned by -what the agent observes, per `docs/prompt-standards.md` item 10: the root check first, -because every other observation is meaningless against the wrong directory; then absent, -present-but-unusable (once accepted), and known-incomplete, each with its own action. The -middle state exists because validation happens when a pass is accepted and a file can be -truncated or replaced afterwards; reading it as incomplete would retroactively uncount a -pass, reading it as usable would feed a broken record into the trend — `?` is the curve -grammar's word for exactly that. +is disclosure rather than a new stop or a new mandatory artifact. The partition is by what the +agent observes, per `docs/prompt-standards.md` item 10, and the middle state exists because a +file validated when its pass was accepted can be truncated or replaced afterwards: reading it +as incomplete would retroactively uncount a pass, reading it as usable would feed a broken +record into the trend, and `?` is the curve grammar's word for exactly that. --- @@ -598,6 +595,11 @@ where it does not: `git show 7c0d475:plugins/dev-workflow/commands/workflow-init.md` count **0**; - the (g) sentence "That question is owned by the loop-rule consolidation work" count **0** in both; against the parent tree, **1** in C and **0** in W; +- the (g) replacement's lead phrase `**The demotion changes what a cycle must resolve, never + what it counts.**` count **1** in C and **1** in W; **0** in the parent tree. This pair is + what the removal check above cannot do alone: deleting the old sentence and installing + nothing satisfies the removal count and the parity diff, and reports the central + demotion/loop-health outcome as verified while both copies carry no answer at all; - lead phrase `**Recording a decline.**` count **1** in each; **0** in the parent tree; - the six-member squash-carry sentence: `every decline record` count **1** in each; **0** in the parent tree; @@ -629,22 +631,21 @@ state the cycle is in afterwards, citing the shipped line the row reads, in both table lives in the plan and is quoted by the closing commit body, not here. What would be observed if the claim "no path leaves a cycle unable to close and unable to suspend" were false: a row whose next state is the same stop with no input consumed — the shape the parent -cycle shipped once (a stop whose only answer resumed a cycle that immediately stopped again) -— or a row that closes with an in-set Blocker standing. The wiring can produce it because -every row is filled from the shipped text, not from this spec, and the inputs the `fic2` -instrument omitted — the user's answer, and the decline — are input columns here; so is the -stop answer, whose next state is a standing suspension by design and must read as one in the -table, not as the same stop re-raised. - -**What this is not.** It is not the `fic2` decision matrix: that instrument scored N states -against old and new text with an expected output each, and Gate B found two defects in the -technique — a state's inputs must include every input the rule reads (the user's answer was -never a column), and a counterfactual must distinguish ABSENT from CONTRADICTORY. The check -above is a presence test whose parent state is absent by inspection; the walk above takes the -answer as an input and produces a next state rather than a scored expected output. **No fixture -per predicate is built** — that question is parked in the story's §2 and is not reopened. - -**Evidence entry**, in the closing commit body, names: the battery run; the five assert pairs +cycle shipped once — or a row that closes with an in-set Blocker standing. The wiring can +produce it because every row is filled from the shipped text rather than from this spec, and +the inputs the `fic2` instrument omitted — the user's answer, and the decline — are columns +here; so is the stop answer, whose next state is a standing suspension by design and must +read as one, not as the same stop re-raised. + +**What this is not.** Not the `fic2` decision matrix: that instrument scored N states against +old and new text with an expected output each, and Gate B found two defects in the technique — +a state's inputs must include every input the rule reads (the user's answer was never a +column), and a counterfactual must distinguish ABSENT from CONTRADICTORY. The check above is a +presence test whose parent state is absent by inspection; the walk takes the answer as an +input and produces a next state rather than a scored output. **No fixture per predicate is +built** — parked in the story's §2, not reopened. + +**Evidence entry**, in the closing commit body, names: the battery run; the six assert pairs with their working-tree and parent-tree counts; the §6 parity diff — the passages extracted, the differences observed, and that each is one of the permitted rows; and the next-state table's location in the plan plus its row count. It is revalidated before every Gate-B @@ -654,14 +655,12 @@ re-review and before the closing amend, as §5 requires. ## 10. AGENTS.md invariants touched -- **Invariant 11 — `docs/prompt-standards.md`, all 12 items.** Most at risk here: item 6 - (rules carry their why — each constraint in the §3 block and the decline block carries its - reason in the same sentence, and the (g) replacement gives two; the plan's review reads each - sentence for one); item 8 (token-lean — the block replaces closure sentences rather than - adding beside them, and `a13` is the reason the old sentences leave); item 10 (diagnostic - states name their causes — the Q6 sentence partitions the unavailable state from its two - neighbours and points each at its rule); and item 3 (stop conditions defined — the block's - stop answer is a named, resumable state). +- **Invariant 11 — `docs/prompt-standards.md`, all 12 items.** Most at risk: item 6 (rules + carry their why — each constraint in both shipped blocks carries its reason in the same + sentence, and the plan's review reads each sentence for one); item 8 (token-lean — the + blocks replace closure sentences rather than adding beside them, `a13` being why the old + ones leave); item 10 (the Q6 partition and its root check); and item 3 (stop conditions + defined — the stop answer is a named, resumable state). - **Don't: "Never replace a decision procedure without accounting for its old conditions."** Satisfied by §5: every inventory id for every edited passage is marked kept, moved, replaced or dropped with a reason. From b8d76734acca5f7daf24209c4c42388696616c4d Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Thu, 10 Sep 2026 15:55:31 +0200 Subject: [PATCH 006/181] =?UTF-8?q?docs(specs):=20loop-rule=20consolidatio?= =?UTF-8?q?n=20=E2=80=94=20Gate-A=20pass=204=20revision?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Applies the 18 findings of Gate-A spec pass 4 (cycle awsf1ec771). Authorised scope growth, decided by Daniel at this cycle's pass-4 scope stop: the change now ships one commit-body record form with two labels, Accepted: and Declined:, instead of a decline record alone. Three passes found the same hole — an acceptance put a repair obligation in the fix set and no artifact carried it, so a lost session left a cycle able to close over work it had agreed to do. The second label reuses the form, transport, nonce and carry rules the first already needed. Also: the Q6 partition rebuilt on two observables and its root-detection claim withdrawn; the rollback and slot-discriminator claims narrowed to what they can support; b12 and b18 marked replaced rather than kept; the severity paragraph given its own reciprocal one-contract sentence; the W/C "Mechanics" cross-reference aligned after the divergence's stated reason proved false. The 135-condition inventory the accounting cites is committed beside the spec, so a reader can check the accounting rather than take it. Docs-only; the Gate-B reminder is the CLAUDE.md §5 prose exemption. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- ...-rule-consolidation-condition-inventory.md | 481 ++++++++++++ ...26-09-10-loop-rule-consolidation-design.md | 711 +++++++++--------- 2 files changed, 851 insertions(+), 341 deletions(-) create mode 100644 docs/superpowers/specs/2026-09-10-loop-rule-consolidation-condition-inventory.md diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-condition-inventory.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-condition-inventory.md new file mode 100644 index 0000000..ace4047 --- /dev/null +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-condition-inventory.md @@ -0,0 +1,481 @@ +# Conditions inventory — §5 review-loop closure rules + +**What this is.** The id definition behind the old-conditions accounting in +`docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md` §5. That accounting marks +every condition `a1`…`j4` kept, moved, replaced or dropped; without this file the ids name +nothing a reader can audit, and an omitted condition is indistinguishable from an undefined id. +It is committed for that reason and for no other: it is a **snapshot**, not a live document. +Taken 2026-09-10 against the tree at `7c0d475` (main, plugin 0.11.0), before any edit of this +change landed. Line numbers therefore cite that tree and will drift; the spec's own line +citations are re-read at execution, and this file is not. A condition here is a claim about the +text as it stood at `7c0d475`, so where the two disagree the tree wins and the accounting is +what needs correcting. + +Read-only inventory taken on branch `loop-rule-consolidation` at +the repository root. No repo file was modified. + +Two copies of §5 exist and are compared throughout: + +- **C** = `CLAUDE.md` (canonical, 1068 lines) +- **W** = `plugins/dev-workflow/commands/workflow-init.md` + (scaffolded template, 1925 lines) + +Parity method: the corresponding line ranges were extracted and `diff -u`'d, so +"identical" below means byte-identical over the stated range, not merely +similar-looking. Line-wrapping differences are reported because they are real +byte differences even where the words match; they are flagged as such. + +Condition numbering is stable: the spec may cite `a1`, `c9`, `h17` etc. Every +condition is quoted with its operative phrase in backticks. Negations and +exceptions are numbered as conditions in their own right. + +**Total conditions listed: 135** (a 22 · b 18 · c 20 · d 7 · e 11 · f 7 · g 4 · +h 26 · i 16 · j 4). + +--- + +## Passage (a) — the HARD FLOOR paragraph + "The derived floor is the pass count a cycle owes" + +**Line ranges** + +| File | Range | Notes | +|---|---|---| +| C | **72–136** | para 1 = 72–121 (`**Both gates are a LOOP with a HARD FLOOR…`), blank 122, para 2 = 123–136 (`**The derived floor is the pass count a cycle owes…`) | +| W | **279–343** | para 1 = 279–328, blank 329, para 2 = 330–343 | + +**Scope note.** Per the task, the floor-derivation *arithmetic* in para 1 +(C 74–86, 96–116 — the max(risk, security) mapping, the `Story:` header +authority, the entry grammar, the per-cycle-kind governing header) is **not** +enumerated here. It is normative and will still need its own inventory if the +successor touches it. The conditions below are the ones that state how the loop +**closes, exits, or is stopped**, plus the sentences in para 1 that decide +whether a pass counts as the cycle's *final* pass — those are closure rules +living inside the arithmetic paragraph and would otherwise be lost. + +**Conditions** + +- **a1** — `Both gates are a LOOP with a HARD FLOOR: a minimum number of passes per run` — the floor is a *minimum pass count per run*, not a target. +- **a2** — `(Blocker/Major only)` — the floor's pass-counting filter is Blocker/Major only. +- **a3** — `A change to a profile or to the set therefore binds every open and future cycle — a raise costs an affected open cycle a further pass under the current profile, as the profile-change rule below requires`. +- **a4** — `a cycle that has already closed stands, its close having been valid under the profile current when it closed, which is the cycle-level form of passes already run keeping their count` — a closed cycle is never reopened by a later profile/set change. +- **a5** — `Some states derive **no** value rather than a second one, and each stops rather than defaulting: governing headers that disagree; a cited profile that is present but unresolvable; and a `Story:` header that cannot be read.` — three named stop states. +- **a6** — `Where they name different sets the premise of a single value has failed: **stop and surface the disagreement** rather than deriving from either, exactly as an unresolvable profile stops rather than defaulting.` +- **a7** — `Before each pass the deriving agent compares every such set, and again before a clean pass is accepted as the cycle's final pass.` — two distinct comparison points, the second being a closure precondition. +- **a8** — `A header or profile that changed during that pass means the pass is not final — the same answer a change gets at every other read point.` +- **a9** — `The derived floor is the pass count a cycle owes`. +- **a10** — `the hook's ratio is a reminder threshold that controls nothing`. +- **a11** — `a satisfied count is not a clean review`. +- **a12** — `a below-threshold reminder is noted in the pass report and disregarded where the cycle's own closure rules are satisfied` — two duties: note it, *and* disregard it under that condition. +- **a13** — `This replaces the pass-count number and nothing else. Every other rule stated here about how a cycle closes stands as written, and none of them is restated — a summary is where their conditions would get dropped.` — a standing prohibition on restating/summarising the closure rules. **Directly relevant to the successor: a single consolidated ordering must not read as a restatement that drops conditions.** +- **a14** — `Nothing here writes the floor knob: it stays the user's, never written, never removed, never read for this derivation.` +- **a15** — `Open a TodoWrite "Codex pass N" per pass`. +- **a16** — `fix Blocker/Major after each`. +- **a17** — `Your final pass must be clean`. +- **a18** — `if the pass at the floor still finds Blocker/Major, keep going until clean or clearly stuck → then STOP and surface to the user`. +- **a19** — `The only early exit below the floor is a pass with **zero** findings`. +- **a20** — `don't manufacture findings to pad`. +- **a21** — `Codex is advisory — validate before applying`. +- **a22** — `dismissed finding → one-line why`. + +**Parity differences** + +**None.** C 72–136 and W 279–343 are byte-identical. + +--- + +## Passage (b) — "**What a loop absorbs, and what stops it**" + +**Line ranges** + +| File | Range | +|---|---| +| C | **195–223** | +| W | **402–426** | + +**Conditions** + +- **b1** — `A finding that corrects the correction you just made **and stays inside the assigned fix set** is **inside this loop's scope**` — two conjuncts; ancestry alone is not enough. +- **b2** — `keep it here rather than handing it back`. +- **b3** — `then act on it by its severity exactly as Mechanics already says — Blocker/Major resolve, Minor/Nit collect and never iterate` (W says `exactly as the severity rule already says` — see parity). +- **b4** — `Ancestry decides where a finding belongs; it never decides what you do with it`. +- **b5** — `it grants no Minor or Nit a repair round it would not otherwise get`. +- **b6** — `The assigned fix set is fixed before the pass you are answering`. +- **b7** — the set `is the scope the approved story or plan assigns to this cycle, plus repair obligations you already accepted in earlier passes` — two components. +- **b8** — `A finding is in-set when repairing it stays inside that scope`. +- **b9** — `never merely because it arrived in the current pass, which would put every new finding in the set by definition and leave the boundary deciding nothing`. +- **b10** — `Where membership is genuinely unclear treat the finding as **outside**, which costs a question and never a silent expansion.` +- **b11** — `A correction that leaves that set stops the loop like any other out-of-scope finding`, **`even when it opens no new question at all`** — the qualifier is load-bearing. +- **b12** — `it resumes the moment the user says whether the set now includes it`. +- **b13** — `A finding that opens a **new structural or contract question** stops the loop and goes to the user`. +- **b14** — `size is not the test, novelty of the question is`, so `a structural finding that is genuinely small still stops it`. +- **b15** — `a long correction still aimed at the last correction does not [stop the loop] — provided that correction, too, stays inside the set, which its ancestry never supplies on its own`. +- **b16** — `When a finding is both … **the new question wins and the loop stops**: novelty overrides correction ancestry`. +- **b17** — `Stopping this way is **not an exit from the gate**: the floor, the Blocker/Major filter and the clean-final-pass rule all stand`. +- **b18** — `the loop resumes on the revised artifact once the question is answered`. + +**Parity differences** (three; C first, W second) + +1. **Cross-reference target.** + C: `then act on it by its severity exactly as **Mechanics** already says — Blocker/Major resolve, Minor/Nit collect and never iterate.` + W: `then act on it by its severity exactly as **the severity rule** already says — Blocker/Major resolve, Minor/Nit collect and never iterate.` + (Deliberate: the scaffolded template does not name this repo's `### Mechanics (reference)` heading.) +2. **Intensifier dropped.** + C: `because absorbing on ancestry is **exactly** how a contract decision gets made without anyone choosing it.` + W: `because absorbing on ancestry is how a contract decision gets made without anyone choosing it.` +3. **Closing rationale reworded, and the field-mint parenthetical exists in C only.** + C: `the loop resumes on the revised artifact once the question is answered — **what the stop prevents is a loop committing you to a design you never chose, which is a different failure from an unfinished review.** (Field-minted in `infinite-portfolio-canvas` and carried here because the alternative was observed there: handing back a three-line repair-of-a-repair wastes a session, and absorbing a contract question spends a decision that was not the loop's to make.)` + W: `the loop resumes on the revised artifact once the question is answered. **What it prevents is a loop committing you to a design nobody chose — a different failure from an unfinished review.**` — **the whole `(Field-minted in infinite-portfolio-canvas …)` sentence is absent from W.** + +--- + +## Passage (c) — "**Recognizing \"clearly stuck\"**" incl. "**Surfacing does not close the cycle**" + +**Line ranges** + +| File | Range | +|---|---| +| C | **225–246** (stuck-reading 225–241; `**Surfacing does not close the cycle…` 242–246) | +| W | **428–449** (stuck-reading 428–444; surfacing 445–449) | + +**Conditions** + +- **c1** — `Read the **Blocker curve across passes**, not any single pass's total`. +- **c2** — `it is the better of the two signals, the total says less than it looks like, and one low count is a snapshot rather than a plateau`. +- **c3** — `**Neither curve measures coverage:** a low Blocker count can sit beside an entirely unreviewed subsystem.` +- **c4** — `this exit needs three things **together**, and a missing one means keep going`. +- **c5** — condition 1: `a plateau visible across passes (six or more is where the field saw one)`. +- **c6** — condition 2: `an **affirmative judgement that coverage is sufficient**, stated` — it must be *stated*, not merely held. +- **c7** — `a known materially unreviewed area forbids this exit outright, and disclosing it does not license it`. +- **c8** — condition 3: `**Blocker or Major findings that keep regenerating across genuine repair attempts**, each round's fix producing the next`. +- **c9** — `**a clean completion takes precedence over this exit**`. +- **c10** — `a Blocker/Major-free pass **at or above the floor** has satisfied the clean-final-pass rule — collect the Minors and Nits and close`. +- **c11** — `reporting "will not converge" on a converged loop is a false report`. +- **c12** — `**Below the floor nothing closes**`. +- **c13** — `a zero-finding pass remains the only exception, exactly as above` (cross-references **a19**). +- **c14** — `a Blocker/Major-free pass below the floor carrying a Minor keeps looping`. +- **c15** — `**Surfacing does not close the cycle, and that is what makes this reachable.**` +- **c16** — `You surface *with the finding still open*`. +- **c17** — `the resolve rule is not waived`. +- **c18** — `no pass is credited as clean`. +- **c19** — `the loop resumes on whatever the user decides`. +- **c20** — `Reading it as "stop instead of fixing" would put the exit in competition with the rule that every Blocker and Major resolves, and then nothing could satisfy both.` — an explicit prohibition on the "stop instead of fixing" reading. + +**Parity differences** + +**None in the passage text** — C 225–246 and W 428–449 are byte-identical. + +**One structural difference immediately after the passage:** +C has **no blank line** between C:246 (`…nothing could satisfy both.`) and C:247 +(`**Every pass report states three things about the floor**…`) — the two run +together as one markdown paragraph. W has a **blank line at W:450**, so in the +template the surfacing block and the floor-reporting duty are separate +paragraphs. (Confirmed by blank-line map: C blanks at 254/262/269/274; W blanks +at 450/458/466/484.) + +--- + +## Passage (d) — "**From pass 4 onward every pass report carries three lines.**" + +**Line ranges** + +| File | Range | +|---|---| +| C | **255–261** | +| W | **459–465** | + +**Conditions** + +- **d1** — `From pass 4 onward every pass report carries three lines.` +- **d2** — `The carrier is **your own status report to the user**`. +- **d3** — `never the Codex reply, which stays exactly one line per branch`. +- **d4** — `never the findings file, which admits no line that is not a finding or the terminator`. +- **d5** — line (1): `the **trend** — findings and Blocker counts across the passes so far`. +- **d6** — line (2): `where this pass's findings **cluster** — product behaviour, the test instrument, or prose about either`. +- **d7** — line (3): `any **require↔withdraw pair** against earlier passes, meaning a pass demanding what an earlier pass had removed`. + +**Parity differences** + +**Wording: none.** The two copies differ only in where lines wrap +(C `…across the passes` / `so far;` vs W `…across the passes so` / `far;`, and +likewise at `the test instrument, / or prose` and `meaning a / pass demanding`). +Same words, same order. + +--- + +## Passage (e) — "Those three lines expose **five tells**" + +**Line ranges** + +| File | Range | +|---|---| +| C | **263–268**, plus the C-only rationale paragraph at **270–273** | +| W | **467–472** (no rationale paragraph) | + +**Conditions** + +- **e1** — `Those three lines expose **five tells**`. +- **e2** — tell 1: `the finding count rising rather than falling`. +- **e3** — tell 2: `the Blocker count failing to fall`. +- **e4** — tell 3: `findings clustering on the **instrument** rather than on product behaviour`. +- **e5** — tell 4: `findings clustering on **prose about** either`. +- **e6** — tell 5: `a require↔withdraw pair`. +- **e7** — `**Any two present makes stop-and-surface mandatory, not discretionary**`. +- **e8** — `you report the tells and hand the decision to the user` (W: `report the tells…`). +- **e9** — `the "clearly stuck" reading above is not a precondition for it`. +- **e10** — `A loop can be worth stopping long before it plateaus.` +- **e11** — (**C only**, 270–273) `That is why this is a reporting obligation with a mandatory threshold and not another heuristic to weigh.` — with its stated provenance `Recorded rationale, from the maintainer rather than from a measurement of this repo: in the Bricks consumer all five signals were measurable by **day two** of a week-long loop, and the cost was never detection — it was the absence of a duty to say so.` + +**Parity differences** (two) + +1. **Pronoun dropped.** + C: `— **you report** the tells and hand the decision to the user` + W: `— **report** the tells and hand the decision to the user` +2. **Whole paragraph exists in C only** (C 270–273): + `Recorded rationale, from the maintainer rather than from a measurement of this repo: in the Bricks consumer all five signals were measurable by **day two** of a week-long loop, and the cost was never detection — it was the absence of a duty to say so. That is why this is a reporting obligation with a mandatory threshold and not another heuristic to weigh.` + W has no counterpart, and **no blank line** between W:472 and W:473 — in the + template the five-tells paragraph and passage (f) are one continuous block. + +--- + +## Passage (f) — "**The two rules above do not compete**" + +**Line ranges** + +| File | Range | +|---|---| +| C | **275–287** | +| W | **473–483** | + +**Conditions** + +- **f1** — `**The two rules above do not compete**, and neither overrides the other`. +- **f2** — `the absorb rule decides whether *a finding* is inside this loop's scope`. +- **f3** — `this reading decides whether *the loop* can still converge`. +- **f4** — `A small correction-of-a-correction that stays inside the assigned fix set is absorbed and is not by itself evidence of a plateau.` +- **f5** — `the late Blockers were semantic contradictions rather than wording, which is why a low count is a signal to read and not a clearance`. +- **f6** — `Hence the sizing guidance: prefer **smaller specs with named interfaces** and let the plan carry the detail — **guidance, not a threshold**, because where the plateau starts is unmeasured.` +- **f7** — (**C only**) `That a round regenerates roughly half the findings it closes is a **hypothesis** in that record rather than a measurement; one lineage was established (the last pass's Blocker came from the previous pass's fix).` — an evidential qualifier that constrains how the curve may be cited. + +**Parity differences** (three) + +1. **Evidence framing.** + C: `The field measurement behind it, quoted at the precision its own record keeps: nineteen Gate-A passes…` + W: `Measured once, at the precision the record keeps: nineteen Gate-A passes…` +2. **The hypothesis qualifier is C-only.** C: `That a round regenerates roughly half the findings it closes is a **hypothesis** in that record rather than a measurement; one lineage was established (the last pass's Blocker came from the previous pass's fix).` — absent from W entirely. +3. **Punctuation/clause boundary around the late-Blockers claim.** + C: `Blockers from 11 to 0–1 from pass 7 on **— and** the late Blockers were semantic contradictions rather than wording**,** which is why a low count is a signal to read and not a clearance.` + W: `Blockers from 11 to 0–1 from pass 7 on**,** and the late Blockers were semantic contradictions rather than wording **—** which is why a low count is a signal to read and not a clearance.` + +--- + +## Passage (g) — Mechanics · Severity bullet · "**How this demotion bears on the loop-health measures**" + +**Line ranges** + +| File | Range | +|---|---| +| C | **810–815** | +| W | **996–999** | + +**Conditions** + +- **g1** — `How this demotion bears on the loop-health measures — the per-pass counts, the finding clusters and the stop thresholds — is not settled here, and this change does not settle it.` — three named unsettled surfaces. +- **g2** — `Until it is, a pass whose outcome would turn on that question reports the question and stops rather than deciding it` — two duties: report, and stop. +- **g3** — `the same answer any unresolved gate question gets`. +- **g4** — (**C only**) `That question is owned by the loop-rule consolidation work in `docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md`.` + +**Parity differences** (one — the expected deliberate one) + +C 814–815 carry the ownership sentence naming +`docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md`; W ends +at `— the same answer any unresolved gate question gets.` and names no story +path. This is the known deliberate difference; **g4 is the condition the +successor discharges**, so removing it in C without removing the corresponding +duty in W would desynchronise the two copies in the opposite direction. + +--- + +## Passage (h) — Mechanics · "**Recording a human exception.**" + +**Line ranges** + +| File | Range | +|---|---| +| C | **977–1033** | +| W | **1161–1217** | + +**Conditions** + +- **h1** — trigger: `Where a human decides that something **no applicable rule required** was nonetheless worth skipping — an optional check this environment cannot run, a review someone asked for and then stood down, a courtesy step — that decision goes in the closing commit body`. +- **h2** — the fixed three-line form: `Human exception: · ` / `Not done: ` / `Accepted because: `. +- **h3** — destination: `an ungated change records it in that commit`. +- **h4** — destination: `a Gate-A cycle in the spec or plan commit`. +- **h5** — destination: `a Gate-B cycle in the WIP commit, restated by the closing amend`. +- **h6** — `Several records accumulate; order means nothing.` +- **h7** — late decision: `**A decision made after its commit closed** — during PR review, say — goes in whichever of these exists: the next commit on the branch, the squash body, or a follow-up commit after the merge.` +- **h8** — `If none does — the branch is closed, unmerged, and heading for an ordinary or rebase merge — **add a commit for it.**` +- **h9** — `An empty commit carrying only the record is a legitimate destination: it changes no content, so it raises no review obligation.` +- **h10** — `**Do not expect silence from the gate hook, and do not read a reminder as a gate reopening.**` +- **h11** — `It is advisory, so it never blocks the commit attempt.` +- **h12** — `What is exempt is the **empty diff**, which `git show --stat` confirms — never a reminder that merely looks the same on a commit carrying content.` +- **h13** — `Copy every record into the squash body alongside the evidence entry (Mechanics, squash-merge carry).` — links to **j1**. +- **h14** — `**Nothing performs that carry and nothing checks afterwards that it happened** — it is on whoever prepares the merge.` +- **h15** — `If two copies of one record disagree, that is a copying error: stop and fix it rather than picking one.` +- **h16** — `**Scope, and it is narrow. This form supplies no permission.** It records a decision that was already the human's to make about something genuinely optional.` +- **h17** — `It is **never** the answer to a below-floor pass, an unclean final pass, a `STOP and surface`, a Gate-A or Gate-B obligation, or a profile-derived evidence requirement` — five named non-uses. **The successor's new "decline" record type reuses this transport and must be checked against each of the five.** +- **h18** — `**it authorizes nothing that any mandatory rule in this file or in `AGENTS.md` requires.**` +- **h19** — `Those have their own terminal actions and this paragraph changes none of them: on a STOP you still stop, and neither a human's assent nor this record lets an agent close or continue a cycle.` +- **h20** — `**"Mandatory" is not limited to this file.** A rule in `AGENTS.md`, a project doc, CI, a branch policy or the platform is equally out of reach`. +- **h21** — `under **Wait for**, `docs/pr-review-bots.md` requires a bot review unless an explicit recorded human decision permits proceeding without it, and this form is not that decision`. +- **h22** — `If you are reaching for it to get past something mandatory, the answer is no — take the operational route or stop.` +- **h23** — `**Nor is it for things that were simply never owed.** An absent review from a bot routed **opportunistically** blocks nothing and needs no exception and no record`. +- **h24** — `Record a decision, not a non-event.` +- **h25** — `**What the record is worth.** It is an **unverified assertion**, and reads as one: nothing checks that the handle belongs to whoever decided, that a human was asked, or that the reason is honest.` +- **h26** — `It supports no claim of authorization or review, and satisfies no evidence obligation. It exists because an exception nobody wrote down is invisible, not because writing it down makes it sound.` + +**Parity differences** + +**None.** C 977–1033 and W 1161–1217 are byte-identical. + +--- + +## Passage (i) — "**When these rules bind.**" (unknown-start fallback) + +**Line ranges** + +| File | Range | +|---|---| +| C | **153–167** | +| W | **360–374** | + +**Conditions** + +- **i1** — `From the commit that ships them`. +- **i2** — `a cycle already running finishes under the rules it started with`. +- **i3** — `Where a cycle's starting rules cannot be established it takes the stricter reading of every part this change touches`. +- **i4** — the strict-reading list, item 1: `at minimum floor 3`. +- **i5** — item 2: `severity classified without the demotion`. +- **i6** — item 3: `the provenance-line duty owed`. +- **i7** — item 4: `the curve duty owed`. +- **i8** — item 5: `the nonce duties at their strictest`. +- **i9** — `the cycle is treated as post-rule, so it owes a nonce, owes its provenance line and its curve or skip record, and uses that nonce in every cycle record it does write`. +- **i10** — `which changes what a record is named, never whether one is owed, so the working record stays optional and a skipped cycle still writes no findings slots`. +- **i11** — `Where it cannot recover a nonce it starts a new cycle rather than claiming `none (pre-rule)`, that reserved field being unavailable to a cycle whose start cannot be established`. +- **i12** — `**Each further rule this change ships adds its own strict reading to this list.**` — **this is the extension point the successor uses; the list at i4–i8 is open by construction.** +- **i13** — `Not a re-derivation, which could hand a level-0 cycle a floor of 1 and skip passes on the strength of not knowing when it started.` +- **i14** — `A user knob set above 3 is not lowered by this fallback.` +- **i15** — `A revert is itself a shipping commit for the old rules`. +- **i16** — `the activation rule wins wherever the start is determinable; the fallback covers only where it is not`. + +**Parity differences** + +**None.** C 153–167 and W 360–374 are byte-identical. + +--- + +## Passage (j) — Mechanics · the squash-merge carry sentence + +**Line ranges** + +| File | Range | +|---|---| +| C | **892** (single line) | +| W | **1076** (single line) | + +**Conditions** + +- **j1** — `On squash-merge, copy every evidence entry, every human-exception record, the provenance lines, the curves and any skipped cycle's skip record … in the squash range into the squash body` — five enumerated record kinds, scoped to *the squash range*. +- **j2** — `TOGETHER WITH THE SKIP REASON IT POINTS AT` (capitalised in source). +- **j3** — `a skip record carried without its reason is a pointer into a body the squash has made unreachable`. +- **j4** — `the squash commit is the only body the merge carries into `main`'s history, so anything left behind is unreachable from it`. + +**Parity differences** + +**None.** C:892 and W:1076 are byte-identical. + +--- + +## Other sites + +Every other place in either file that states which exit closes or suspends a +cycle, or that ranks one rule over another. Grepped for: `clean completion`, +`outranks`, `takes precedence`, `close the cycle`, `closes a cycle`, +`cycle-closing`, `suspend`, `stop and surface` / `STOP and surface` / +`stop-and-surface`, plus `closure`, `final clean pass`, `early exit`, +`toward the floor`, `at or above the floor`, `exit the loop`, `licence to close`. +(`outranks` and `suspend` occur **nowhere** in either file.) Hits already inside +passages (a)–(j) are omitted. + +### Closure / stop statements + +| # | C | W | Sentence | +|---|---|---|---| +| o1 | **63** | **262** | `The work loop includes the review gates: **spec ready → Gate A (spec) → plan ready → Gate A (plan) → execute → tests green → Gate B → commit** (see §5).` — the only place stating the gate *sequence* as an ordering. | +| o2 | **181–185** | **388–392** | `Everything else likewise keeps its own footing and is **not** required to derive from the floor: **the other closure and stop predicates** — assigned-fix-set membership, a new structural question, an accepted Blocker or Major, the tell thresholds; **independent reporting and diagnostic ordinals**, such as a duty owed from a given pass onward; and **the hook's reminder threshold together with any descriptive or historical pass number**…` — **names four closure/stop predicates as a set and asserts they are independent of the floor. The successor's single ordering must not collapse them into the floor.** | +| o3 | **189–192** | **396–399** | `A merge can produce any of them: the Gate-A loop description, **the pass-1 closure rule** and the re-review rationale each carry a fixed-three claim… In any such state nothing here resolves which rule governs: **stop, and have a human complete or revert the adoption, before running a gate under it.**` | +| o4 | **566** | **757** | `Ask for one line per finding and a literal `NO FINDINGS` when a pass is clean — **the explicit clean signal is what lets you exit the loop**` (Gate-A bullet). | +| o5 | **573** | **764** | `Each pass: validate, revise, re-run.` (Gate-A bullet — the loop's per-pass cadence.) | +| o6 | **583–584** | **774–775** | `Re-review after every fix — a fix changes the artifact, so the prior review no longer covers it. The hook merely notices, at commit time.` (Gate-B bullet.) | +| o7 | **484–486** | **676–678** | `Anything else … — is an **INCOMPLETE pass**, which is not a review: **don't act on the partial list, don't count it toward the floor, and don't read "no Blocker/Major visible" as clean.**` | +| o8 | **514–515** | **706–707** | `Spent and still incomplete → **STOP and surface**, naming which check failed.` (recovery budget.) | +| o9 | **410–411** | **604–605** | `a working record left by a closed cycle is not one either, which is why that record is **retired at closure** rather than left to be found later.` | +| o10 | **654–655** | **840–841** | `The Blocker/Major filter, the file-first findings protocol and the clean-final-pass rule are unchanged. The floor is not among them: it is no longer a fixed number but derives from the profile and the cited set.` | +| o11 | **664–670** | **850–856** | `The cited path **does not yield a readable story file** → **stop and surface which of these it was**…` (profile-reading case 3.) | +| o12 | **717–719** | **903–905** | `An **unobservable counterfactual is a blocking evidence gap**, not a free pass — **stop and surface**; the human may then lower the mode as a logged override.` | +| o13 | **727–730** | **913–916** | `…**revalidated before every Gate-B re-review and before the cycle-closing amend** … If revalidation changes the entry, **the clean pass no longer covers what is being committed: fix, re-review, close on the entry that pass validated.**` | +| o14 | **749–752** | **935–938** | `Passes already run under the lower profile **keep counting** toward the floor; only the **final clean pass** must run under the current profile. Inside an active Gate-B cycle, fold the edit into the active `WIP:` snapshot by amend — **a non-`WIP` commit reads to the hook as the cycle closing and would discard the accumulated passes.**` | +| o15 | **755–758** | **941–944** | `**While a gate is running, the floor derives from the current profile at each pass.** Passes already run keep counting; **closing requires the floor as currently derived.**` | +| o16 | **760–764** | **946–950** | `**Any profile change costs at least one further pass** … because the final clean pass must run under the current profile — so no already-banked pass can be it. **That further pass must itself be clean and every other closure duty must be satisfied; it is one more pass, not a licence to close on the next one.**` | +| o17 | **769–773** | **955–959** | `**The cited set is re-read at each pass, and the final clean pass runs against the current set** — whenever its membership changes, not only when the floor number moves… but **it never discharges an accepted in-set Blocker or Major: the acceptance put that finding in the fix set, not the citation.**` | +| o18 | **692–698** | **878–884** | `The **skip reason is recorded in the commit body** … **A skip removes the review, never the evidence** … **Neither is excused the records every cycle owes** — the provenance line, and a skip record in place of the curve.` (Gate-B triviality skip — the fourth way a cycle ends without passes.) | +| o19 | **783–784** | **969–970** | `**Severity:** Blocker (wrong/unsafe/breaks invariant) · Major (design flaw → rework) → **both must resolve**. Minor · Nit → **collect, never iterate**.` — the standing duty **b3**, **c17** and **c20** point at. | +| o20 | **788–789** | **974–975** | `If you cannot name both, the finding is Minor or below: collect, never iterate.` (severity demotion — the rule whose loop-health effect **g1** leaves unsettled.) | +| o21 | **825–826** | **1009–1010** | `A pre-review snapshot named anything else reads as a real commit and **closes the cycle, discarding the passes you just accumulated.**` | +| o22 | **827–829** | **1011–1013** | `**Finishing the cycle:** after the final clean pass, close it with `git commit --amend -m ""` — that replaces the WIP commit, and **the hook reads the amend as the real cycle-closing commit.**` | +| o23 | **885–890** | **1069–1074** | `**A project whose text carries some of them and not others, or carries all of them in versions that disagree, stops and has a human complete, revert or reconcile the adoption before running a gate under it**` (records-contract partial-adoption stop). | +| o24 | **174–178, 191–192** | **381–385, 398–399** | `**A partial adoption can leave a project's floor undefined or self-contradictory.** … **exactly one definition of the floor must be present, and every statement that defines or constrains the floor, or makes closing depend on it, must resolve to that one definition.**` — **the coherence requirement the successor's single ordering must itself satisfy.** | +| o25 | **178–180** | **385–387** | `**The unknown-start fallback is not a second definition**: it is explicitly conditional on a cycle's starting rules being undeterminable and governs only that state, so it coexists with the predicate rather than competing with it.` | +| o26 | **1011** (in h) | **1195** (in h) | `It is **never** the answer to a below-floor pass, an unclean final pass, a `STOP and surface`…` — listed here too because it is the only cross-reference from Mechanics back to the loop's four exits. | + +### "Takes precedence" / ranking statements + +Only **two** ranking claims exist across both files, and both are inside the +passages above: + +- **C:235 / W:438** — `a clean completion takes precedence over this exit` (= **c9**). +- **C:275 / W:473** — `**The two rules above do not compete**, and neither overrides the other` (= **f1**). + +`outranks` appears nowhere in either file. The phrase `takes precedence` also +appears once as a back-reference at **C:140 / W:347** — `what makes that +tolerable is **the precedence rule above** plus the hook exiting 0 on every +branch, not the reminder being harmless` (the "Named residual" paragraph, +C 138–141 / W 345–348, pointing at **a12**). That paragraph is byte-identical +between the copies. + +### Adjacent duty not inside (a)–(j), flagged because the successor will sit next to it + +**C 247–253 / W 451–457** — `**Every pass report states three things about the +floor**, from pass 1 onward: the derived floor, the risk and security values +read, and the cited stories they were read from. … This is owed by every pass; +the three lines below are owed from pass 4 and are a different obligation.` +Byte-identical between the copies. It is a *reporting* duty, not a closure +predicate, but it is the paragraph that separates (c) from (d) in W and is fused +onto the end of (c) in C (see the (c) parity note), so any re-paragraphing of +(c) or (d) touches it. + +--- + +## Whole-file parity summary for the inventoried ranges + +| Passage | Byte-identical? | Difference kind | +|---|---|---| +| (a) | yes | — | +| (b) | no | cross-reference target; intensifier; closing rationale reworded + C-only field-mint parenthetical | +| (c) | yes | (paragraph break *after* the passage differs) | +| (d) | words identical | line-wrapping only | +| (e) | no | `you report` → `report`; C-only "Recorded rationale" paragraph; no blank line before (f) in W | +| (f) | no | evidence framing reworded; C-only hypothesis qualifier; punctuation of the late-Blockers clause | +| (g) | no | C-only story-path ownership sentence (expected deliberate difference) | +| (h) | yes | — | +| (i) | yes | — | +| (j) | yes | — | diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index cb4dfa0..cb3657e 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -1,20 +1,20 @@ # §5 loop-rule consolidation: one closure ordering — Design -**Date:** 2026-09-10 · **Status:** DRAFT — Gate A running (pass 3 revised) +**Date:** 2026-09-10 · **Status:** DRAFT — Gate A running (pass 4 revised) **Story:** `docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md` -**Profile:** read from that header at every pass, never from here — the header is the only -writable copy, and a value copied into this file would be a remembered value. +**Profile:** read from that header at every pass, never from here — it is the only writable +copy, and a value copied here would be a remembered value. Prompt-only, in two mirrored copies: `CLAUDE.md` §5 (**C** below) and the inline template in `plugins/dev-workflow/commands/workflow-init.md` (**W** below). No file under `plugins/dev-workflow/hooks/` changes. Line numbers cite the tree at `7c0d475` (main) and are re-read at execution; the plan carries the `grep -n` sites. -Two research inventories were taken before this spec and the plan re-reads them: a -135-condition inventory of the passages this spec rewrites (ids `a1`…`j4`, cited below; -22 + 18 + 20 + 7 + 11 + 7 + 4 + 26 + 16 + 4), and a site map of every sentence that -enumerates the records a commit body carries. Where a sentence outside those passages is cited, -it is quoted by its lead phrase and C line. +The old-conditions accounting in §5 cites ids `a1`…`j4`, defined in +`docs/superpowers/specs/2026-09-10-loop-rule-consolidation-condition-inventory.md` beside this +file: 135 conditions (22 + 18 + 20 + 7 + 11 + 7 + 4 + 26 + 16 + 4) quoted from `7c0d475`, so a +reader can check the accounting rather than take it. A sentence outside those passages is +cited by its lead phrase and C line. --- @@ -23,14 +23,15 @@ it is quoted by its lead phrase and C line. **What ships.** One closure ordering, stated once in each copy, that says which exit *closes* a cycle and which merely *suspend* it, in what order a pass is read so the ranking is executable rather than asserted, what any set of suspensions at once does, and which of the four standing -duties participate in that ordering versus gate it as preconditions. On top of it: a decline -rule with a commit-body record, so a user's answer on a surfaced finding has one rule in both -directions; the answer to what a severity demotion does to the loop-health counts; the answer -to the pass-4 report when prior-pass history is unavailable; and the deferred slot -discriminator, dissolved rather than shipped. +duties participate in that ordering versus gate it as preconditions. On top of it: one +commit-body record form with two labels, `Accepted:` and `Declined:`, so a user's answer on a +surfaced finding has one rule in both directions and survives the session that made it; the +answer to what a severity demotion does to the loop-health counts; the answer to the pass-4 +report when prior-pass history is unavailable; and the deferred slot discriminator, dissolved +rather than shipped. -**What does not.** The pass floor and severity semantics — the parent shipped them and this -spec reads them as given — and everything §11 lists as parked or out of scope. +**What does not.** The pass floor and severity semantics, which the parent shipped and this +spec reads as given, and everything §11 lists as parked. --- @@ -44,10 +45,9 @@ files establish the **inventory** of findings, not their resolutions. This spec adds what the table does not settle — the evaluation order and the field and file set every predicate reads; the four duties' classification; the scope stop's two triggers and -what each answer does; what a stuck or two-tell answer produces, a stop included; the decline -record's wording, attribution and recording point, and what a cycle does when it loses the -state no record carries; and the raw-severity rule for the health measures — each decided in -the section that uses it (§3, §4, and §5 passage (g)). +what each answer does; what a stuck or two-tell answer produces, a stop included; the records' +wording, attribution and recording point, and the `Accepted:` label itself (§4); and the +raw-severity rule for the health measures — each decided in the section that uses it. --- @@ -72,7 +72,7 @@ findings — since a predicate reading one field for half of itself could not be **First, clean completion.** A pass is clean when it carries no Blocker or Major at effective severity that is **in the assigned fix set**, and it surfaced no scope stop; a pass with **zero** findings is clean whatever the floor. A finding matched by a decline recorded in -this cycle is outside the set by that decision (Mechanics, the decline record) — a decline +this cycle is outside the set by that decision (Mechanics, the answer records) — a decline keeps a finding out and never excuses one that is in, so a declined finding the set later comes to include owes resolution like any other. **A clean pass at or above the derived floor, or a zero-finding pass, closes the cycle**, under the final-acceptance preconditions the floor @@ -102,10 +102,10 @@ answered in this cycle. A difference in any of the five fields, or genuine uncer new finding and a new stop of either kind. **Third, a pass that neither closes nor suspends continues** — the loop runs another pass on -the revised artifact. This is the ordinary case and not a residue: a clean pass below the -floor lands here, and so does a pass whose only findings are Minors and Nits, collected and -never iterated. It is a branch and not an inference, because "does not close" read alone says -nothing about whether to run again. +the **current** artifact, revised or not. This is the ordinary case and not a residue: below +the floor a clean pass lands here, and so does a pass whose only findings are Minors and Nits, +which are collected and never iterated and so may leave nothing to revise. It is a branch and +not an inference, because "does not close" read alone says nothing about whether to run again. **The four standing duties, classified.** The **derived floor** is a precondition on closure: it gates closing, and is discharged by the count of valid logical passes reaching @@ -113,7 +113,9 @@ the floor with the last of them clean, or by the zero-finding exit. The **Blocke duty** is a precondition on closure and on any pass being clean: it gates both, and is discharged, for every in-set Blocker and Major, by a validated repair **or a validated dismissal carrying its one-line why** — the advisory rule above, unchanged — followed by a -pass that finds none. **A dismissal is not a decline**: a dismissal is the author's judgement +validated pass that finds **no in-set Blocker or Major at effective severity**, which is the +clean predicate's own wording so that the duty and the predicate cannot drift apart. +**A dismissal is not a decline**: a dismissal is the author's judgement that the finding is not true of the artifact, a decline is the user's decision that a true finding stays outside the fix set, and only the second is an answer at a membership stop. The **hold** a surfaced finding places on closure is part of the ordering: it gates closing @@ -128,18 +130,19 @@ single-trigger finding, both the question decision and the membership answer for both triggers — and in either direction: a hold only accepting could end would be the resolve duty under another name. At a membership stop the answer is **accept** (the finding joins the fix set and resolves by its effective severity: Blocker or Major before any pass can be clean, -Minor or Nit collected and never iterated) or **decline** (the finding stays outside, -recorded, binding for the rest of this cycle — Mechanics, the decline record). At a question +Minor or Nit collected and never iterated) or **decline** (the finding stays outside, binding +for the rest of this cycle). Either answer is **recorded** — Mechanics, the answer records — +and that record is where a later pass, or a cycle that lost its session, reads it. At a question stop the answer is the user's decision on the question, and membership does not change: an in-set finding then resolves under that decision or is dismissed with its one-line why; an out-of-set finding that opened the question is a membership stop as well and takes accept or decline. **Decline is available only at a membership stop.** The stuck and two-tell readings raise no hold: each asks one question, **continue or stop**. Continue resumes the loop on the artifact as revised and the fix set as the governing artifacts now assign it — where several -plans or stories govern one cycle, the union of the scopes they assign — **plus the accepted -obligations this cycle can still recover**, from the running session or from its own records. -One it cannot recover is not in the set, and the reviewer raises its finding again on the next -pass like any other, which is the ordinary route and not a special one. A finding that a +plans or stories govern one cycle, the union of the scopes they assign — **plus the obligations +this cycle accepted**, which its commit bodies carry (Mechanics, the answer records). One no +body carries is not in the set, and the reviewer raises its finding again on the next pass like +any other, which is the ordinary route and not a special one. A finding that a narrowing has put outside the set surfaces at the next pass as a membership stop. **Stop leaves the suspension standing**: the cycle stays open under its nonce, resumable by a later continue, nothing it wrote is a closing commit, and an open cycle nobody resumes is a human's @@ -153,183 +156,189 @@ running. Two states cannot co-occur, and no rule ranks them: clean completion an stop, since a pass that surfaced one is not clean; and a zero-finding pass and any suspension, since it has nothing to surface, nothing regenerating, no cluster and no require↔withdraw pair. The Gate-B triviality skip is outside this ordering: a skipped cycle -runs no passes and ends by its own rule. **This ordering, the decline record in Mechanics and -the severity rule's raw-versus-effective split are one contract**: the hold's declined -direction is released by that record and by nothing else, and the split says which field each +runs no passes and ends by its own rule. **This ordering, the answer records in Mechanics and +the severity rule's raw-versus-effective split are one contract**: a membership stop's two +answers are carried by those records and by nothing else, and the split says which field each predicate here reads, so any one of the three present without the others is a partial adoption that stops under the one-contract rule there. ``` -Why the shape, where the block's own sentences do not already carry it. The evaluation order -answers the objection that a ranking between clean completion and the two-tell stop cannot -fire if the stop can make the pass unclean: clean completion is read first, so a suspension is -only ever evaluated on a pass that did not close. The three branches are story AC 4 — every -pass has a next state, and the third exists because "does not close" is not an instruction. -The duties paragraph is AC 2; the composition and cannot-co-occur sentences are AC 1's demand -that every reachable conflict be covered and every unreachable one named with its reason. The -scope stop's two triggers are `b11` (membership) and `b13` (question) read separately, because -an in-set finding that opens a question can neither join nor stay outside the set and needs -its own answer — which is also why a matching decline suppresses the membership trigger only. -The preconditions the close sentence points at are at C:116–118 ("again before a clean pass is -accepted as the cycle's final pass") and C:760–764 ("Any profile change costs at least one -further pass"). The stop answer is the one new -state this spec supplies beyond the table, and it reuses only shipped mechanisms: resumption -on the revised artifact (`b18`), the membership stop (`b11`), and the nonce rules' treatment -of open cycles ("they stay open, keep their own nonces, and are a human's to resolve", -C:422–424). **D3** keeps the clearly-stuck paragraph's own precedence sentence verbatim (§5, -passage (c)), so the block cites it instead of restating its consequence clause. The block -defines neither the floor nor the tells nor the stuck reading — those stay in their paragraphs -("exactly one definition of the floor must be present", C:176) — and restates no rule it does -not own (`a13`). +Why the shape, where the block's own sentences do not carry it. The evaluation order answers +the objection that a ranking between clean completion and the two-tell stop cannot fire if the +stop can make the pass unclean: clean completion is read first, so a suspension is only ever +evaluated on a pass that did not close. The three branches are AC 4, the duties paragraph AC 2, +the composition and cannot-co-occur sentences AC 1. The scope stop's two triggers are `b11` +(membership) and `b13` (question) read separately, because an in-set finding that opens a +question can neither join nor stay outside the set and needs its own answer — which is also why +a matching decline suppresses the membership trigger only. Sentences the block points at rather +than restating (`a13`): C:116–118, C:760–764, C:422–424, C:176, and the clearly-stuck +paragraph's own precedence sentence, kept verbatim there under **D3**. --- -## 4. The decline rule and record +## 4. The answer records — one form, two labels **The rule ships as the block below**, which is the text and not a summary of it; this section -adds only what the block does not carry. The decisions it implements: availability at a -membership stop and nowhere else is **D6**; the remainder-of-cycle binding **D7**; the -explicit attributable decision **D8**; the five-field sameness test **D9b**; the -unverified-assertion reading **D9c**; the transport **D9**, whose "closing commit body" is -therefore the *last* of the bodies that carry the record, not the first. That a decline never -qualifies the Blocker/Major-resolve duty is **D5**, and it is why the clean predicate in §3 -names the fix set and not the decline: a decline is only ever a fact about membership, so it -cannot be read as excusing an in-set finding — the gate-off route a fabricated record would -otherwise open. That accepted obligations and standing suspensions get **no new mandatory -record** is deliberate: **D10** refused one for the same class of state, and the existing -no-identity rule ("**No candidate, disagreeing sources, or more than one candidate → no -identity: start a new cycle**", C:419–420 / W:613–614) already answers the loss in the safe -direction, at the cost of passes. +adds only what the block does not carry. Decisions implemented: availability at a membership +stop and nowhere else **D6**; the remainder-of-cycle binding **D7**; the explicit attributable +decision **D8**; the five-field sameness test **D9b**; the unverified-assertion reading +**D9c**; the transport **D9**, whose "closing commit body" is therefore the *last* body to +carry a record, not the first. That a decline never qualifies the Blocker/Major-resolve duty +is **D5**, and it is why §3's clean predicate names the fix set rather than the decline: a +decline is only ever a fact about membership, so it cannot be read as excusing an in-set +finding — the gate-off route a fabricated record would otherwise open. + +**The `Accepted:` label is authorised scope growth beyond the story**, decided by Daniel at +this cycle's Gate-A pass-4 scope stop and recorded in the pass-4 dispositions. The story +shipped a decline record only, and three successive passes found the same hole: an acceptance +put a repair obligation in the fix set and no artifact carried it, so a lost session left a +cycle able to close over work it had agreed to do. The cheapest close was a second label on a +record whose form, transport, nonce and carry rules already existed. This Gate-A cycle runs +under the pre-change rules and writes no such record for its own acceptance; the dispositions +file and the working record are what today's rules provide. **The block that ships**, in Mechanics, placed **immediately after the whole human-exception -passage** — after its last paragraph "**What the record is worth.**" (C:1028–1033 / -W:1212–1217) and before the next bullet "- **Timeout / abort:**" (C:1034 / W:1218), at the same -two-space indent. Verbatim, byte-identical in both copies: +passage** — after "**What the record is worth.**" (C:1028–1033 / W:1212–1217) and before +"- **Timeout / abort:**" (C:1034 / W:1218), at the same two-space indent. Verbatim, +byte-identical in both copies: ```` - **Recording a decline.** A decline is the user's answer at a **membership stop** (the - closure ordering above) that the surfaced finding stays outside the assigned fix set. It is - available there and nowhere else: not for an in-set Blocker or Major, which owes resolution, - and not as the answer to a question stop, a stuck or two-tell surface, a below-floor pass, an - unclean final pass, or any Gate-A, Gate-B or evidence obligation — of the list the - human-exception form is never the answer to, the membership stop is the one item this - record answers. It must be an **explicit, attributable decision on that specific finding** — - never silence, never a general remark about scope, never inferred — because a hold released - by inference is a hold nobody chose to release. It **releases the hold** and **binds for the - remainder of this cycle**, with no effect in any later one; it never qualifies the - Blocker/Major-resolve duty, which the finding never reached. A later pass raises **the same - finding** when all five of location, defect, severity, consequence and suggested fix match, - read on meaning rather than bytes, since a reviewer rewrites its sentences between passes. - A matching finding raises **no membership stop and no membership hold**, so a decline is not - re-asked every pass — **and that alone**: the five fields record no question, so a **question - stop** still fires on it unless that same question has already been answered in this cycle - (the closure ordering above). **Any difference — severity included — or any genuine - uncertainty makes it a new finding, and the hold applies.** A decline keeps a finding out and - never excuses one that is in: a declined finding the fix set later comes to include owes - resolution like any other. + **Recording an answer at a membership stop.** A membership stop asks whether a surfaced + finding joins the assigned fix set, and the user's answer is recorded under one of two + labels sharing one form. **Accepted** puts the finding in the set, where the + Blocker/Major-resolve duty governs it from then on. **Declined** keeps it out and releases + its hold. Both are available at a membership stop and nowhere else: not for an in-set + Blocker or Major, which owes resolution already, and neither is the answer to a question + stop, a stuck or two-tell surface, a below-floor pass, an unclean final pass, or any Gate-A, + Gate-B or evidence obligation — of the list the human-exception form is never the answer to, + the membership stop is the one item these records answer. An **acceptance also records a + question stop's decision where that decision put work in the set**, since that is the same + fact under another name: work the cycle now owes. Each must be an **explicit, attributable + decision on that specific finding** — never silence, never a general remark about scope, + never inferred — because a fix set changed by inference is a fix set nobody chose. ``` + Accepted: · · cycle + Finding: | | | | + Declined: · · cycle Finding: | | | | ``` The `Finding:` line is the finding line from the pass's findings file with its confidence field removed — the five fields the sameness test reads, in the file's order, a literal pipe - escaped as `\|` exactly as there. **It carries the cycle nonce**, because it binds to one - cycle and a record that cannot be attributed to its cycle cannot bind to it. **When and - where:** written when made, into the cycle's next commit body on the branch — a spec or plan - revision commit for a Gate-A cycle, the `WIP:` amend for Gate B — **and that commit is made - before the next pass runs**: no pass runs on a decline that no commit body on the branch - carries, because a decline held only in a session is one compaction away from a re-asked - question; the advisory working record may carry it meanwhile and does not bind. Where a - Gate-A cycle has no artifact revision to carry it — the declined finding was that pass's - only one — the record goes in an **empty commit of its own**, the destination the - human-exception rule above already blesses: "An empty commit carrying only the record is a - legitimate destination". It is - restated in every later body of that cycle, the closing one included, because a body that - drops it releases nothing and re-asks the question. A cycle that resumes with no commit body - carrying a decline **treats it as absent and the hold applies again** — the reading the - unknown-start fallback gives, and the safe direction. It is copied on squash-merge with the - other records (the carry rule above). **What it is worth:** an unverified assertion of the - same kind as the human exception — nothing checks that the handle belongs to whoever - decided, that a human was asked, or that the reason is honest. **How it differs in force:** - it releases a hold, which the human-exception form never does; narrowness bounds what a - false one can do — one fully identified finding, one cycle, **as far as distinct nonces - allow**: two cycles sharing or redrawing a nonce are indistinguishable to this record as to - every other, and a replayed decline can then bind to the wrong cycle — and a bound is not - the same as safety. **This record, the closure ordering and the severity rule's - raw-versus-effective split are one contract**: this record releases the hold that ordering - defines, and that split says which severity field each of its predicates reads, so any one - of the three present without the others is a partial adoption that stops under the - one-contract rule above. **What the body does not record, said here - rather than discovered:** an accepted finding, a stuck or two-tell surface and its - continue-or-stop answer leave no record of their own — from a closing body a reader can - infer the close and any declines, and nothing else about which suspensions the cycle passed - through. They live in the running session and, optionally, the working record; a cycle that - cannot recover them has no identity and starts a new cycle under the recovery rule above, - which costs passes and closes nothing — the reviewer re-reads the whole artifact each pass, - so an unrepaired accepted finding is re-raised and one outside the new cycle's scope is - re-asked as a membership stop, a repeated question rather than a silent close. + escaped as `\|` exactly as there. + + **Rules both labels share.** Each carries the **cycle nonce**, because a record that cannot + be attributed to its cycle cannot bind to it. Each is **written when made**, into the cycle's + next commit body on the branch — a spec or plan revision commit for a Gate-A cycle, the + `WIP:` amend for Gate B — **and that commit is made before the next pass runs**: an answer + held only in a session is one compaction away from being lost; the advisory working record + may carry it meanwhile and does not bind. Where the cycle has no artifact revision to carry + it — the answered finding was that pass's only one — the record goes in an **empty commit of + its own**, the destination the human-exception rule above already blesses: "An empty commit + carrying only the record is a legitimate destination". Each is **restated in every later body + of that cycle**, the closing one included, because a body that drops one loses the fact it + records, and each is **copied on squash-merge** (the carry rule above). A cycle that resumes + with no commit body carrying an answer **treats it as absent** — the unknown-start fallback's + reading, and the safe direction under both labels: an unrecorded decline means the hold + applies again, an unrecorded acceptance means the finding is raised afresh. Each is an + **unverified assertion** of the same kind as the human exception — nothing checks that the + handle belongs to whoever decided, that a human was asked, or that the reason is honest. + + **Binding, and the sameness test.** A decline **binds for the remainder of its cycle**, with + no effect in any later one, and never qualifies the Blocker/Major-resolve duty, which the + declined finding never reached. A later pass raises **the same finding** when all five of + location, defect, severity, consequence and suggested fix match, read on meaning rather than + bytes, since a reviewer rewrites its sentences between passes. A finding matching a decline + of this cycle raises **no membership stop and no membership hold**, so a decline is not + re-asked every pass — **and that alone**: the five fields record no question, so a + **question stop** still fires on it unless that same question has already been answered in + this cycle (the closure ordering above). **Any difference — severity included — or any + genuine uncertainty makes it a new finding**, and the stop applies. A decline keeps a + finding out and never excuses one that is in: a declined finding the fix set later comes to + include owes resolution like any other. An acceptance does not expire with its pass — the + obligation it records stands until the resolve duty discharges it. + + **What each is worth.** A decline **releases a hold** and an acceptance **creates an + obligation**, neither of which the human-exception form ever does. Narrowness bounds what a + false record can do — one fully identified finding, one cycle, **as far as distinct nonces + allow**: two cycles sharing or redrawing a nonce are indistinguishable to these records as + to every other, so a replayed answer can bind to the wrong cycle, and a bound is not safety. + **What the acceptance record buys, and what it does not:** a cycle that lost its session and + started fresh can read the branch's commit bodies and find what was accepted. Nothing makes + it read them and nothing checks that it did, so a replacement cycle can still review and + close the same artifact while an older one stays open; the record makes that discoverable + rather than invisible, which is less than preventing it. **The one thing no record carries:** + a stuck or two-tell surface and its continue-or-stop answer — from a closing body a reader + can infer the close, the acceptances and the declines, and nothing about those two. + + **These records, the closure ordering and the severity rule's raw-versus-effective split are + one contract**: these records carry the two answers that ordering's membership stop asks + for, and that split says which severity field each of its predicates reads, so any one of + the three present without the others is a partial adoption that stops under the one-contract + rule above. ```` -**Which rules the answer modifies** (story AC 3): the scope stop's resume sentence (`b12`, -which gains "accepting or declining"), and the hold and the clean-pass definition (both in the -§3 block, where the answer's two directions are one rule). Those are the rules whose **closure -behaviour** the answer qualifies, and it qualifies **no other closure rule**: not the -Blocker/Major-resolve duty (Mechanics, Severity: "both must resolve", C:784 — **D5**: it stays -unqualified because a declined finding never enters its scope), not the floor, not the tells, -not the stuck reading. A reader finding "decline" qualifying any *closure* rule outside those -three has found a defect. The decline is also **named**, without qualifying closure, at the -transport, attribution, activation and threat sites listed next — the squash carry, the nonce -set and exemption, the "Both shipped records" count, the unknown-start fallback, the gate-off -surface, the closing-message carry and the one-contract paragraph — and a reader finding it -absent from any of those has found the opposite defect. - -**Existing sentences that must name it**, both copies, old → new. Line numbers are C's; W's +**Which rules the answer modifies** (story AC 3): the scope stop's resume sentence (`b12`), +and the hold and the clean-pass definition (both in the §3 block, where the answer's two +directions are one rule). Those three are the rules whose **closure behaviour** the answer +qualifies, and it qualifies no other — not the Blocker/Major-resolve duty (Mechanics, +Severity: "both must resolve", C:784 — **D5**: a declined finding never enters its scope, and +an accepted one enters it without changing what the duty demands), not the floor, not the +tells, not the stuck reading. A reader finding either label qualifying a *closure* rule +outside those three has found a defect. The records are separately **named**, without +qualifying closure, at the transport, attribution, activation and threat sites listed next, +and a reader finding them absent from any of those has found the opposite defect. The two +lists answer different questions: what the answer *changes*, and what must *carry* it. + +**Existing sentences that must name them**, both copies, old → new. Line numbers are C's; W's are in the site map and re-read at execution. 1. **Squash carry** (C:892 / W:1076, `j1`). OLD: "…copy every evidence entry, every human-exception record, the provenance lines, the curves and any skipped cycle's skip record TOGETHER WITH THE SKIP REASON IT POINTS AT…". NEW: "…copy every evidence entry, - every human-exception record, **every decline record**, the provenance lines, the curves - and any skipped cycle's skip record TOGETHER WITH THE SKIP REASON IT POINTS AT…". Six - members; the rest of the sentence unchanged. + every human-exception record, **every answer record, one copy per cycle nonce and + five-field finding key — byte-identical repeats collapse, and copies that disagree stop + under the rule below** —, the provenance lines, the curves and any skipped cycle's skip + record TOGETHER WITH THE SKIP REASON IT POINTS AT…". Six members; the dedup clause exists + because a record restated in every body of its cycle reaches the squash range many times. 2. **The named nonce set** (C:385–387 / W:579–581). OLD: "**and that set is named rather than left open**: the provenance line, the per-pass curve (including a skip record standing in for one), the cycle's findings slots, and its advisory working record." NEW: "**and that set is named rather than left open**: the provenance line, the per-pass curve (including a skip - record standing in for one), **any decline record (Mechanics)**, the cycle's findings + record standing in for one), **any answer record (Mechanics)**, the cycle's findings slots, and its advisory working record." 3. **The nonce exemption** (C:397–399 / W:591–593). OLD: "The nonce is not required in records this change neither introduces nor keys to a cycle — the evidence entry and a human-exception record among them." NEW: "The nonce is not required in records that are not - keyed to a cycle — the evidence entry and a human-exception record among them; **a decline + keyed to a cycle — the evidence entry and a human-exception record among them; **an answer record is keyed to its cycle and carries it**." (The old sentence's "this change" dated it to the parent; the new one states the criterion.) 4. **"Both shipped records below"** (C:367 / W:561). OLD: "Both shipped records below carry a **cycle field**, because…". NEW: "Both shipped records below carry a **cycle field** — and - so does the decline record in Mechanics — because…". A load-bearing count that a third - cycle-attributed record would otherwise falsify. + so do the answer records in Mechanics — because…". A load-bearing count that further + cycle-attributed records would otherwise falsify. 5. **The unknown-start fallback** (C:153–167 / W:360–374, `i4`–`i8`, extended at `i12`'s invitation). OLD: "…at minimum floor 3, severity classified without the demotion, the provenance-line duty owed, the curve duty owed, and the nonce duties at their strictest…". NEW: "…at minimum floor 3, severity classified without the demotion, the provenance-line duty owed, the curve duty owed, the nonce duties at their strictest, **every suspension - binding, and decline records not attributable to the nonce this fallback minted treated as - absent, so that no inherited hold is released — declines recorded under that nonce are - honoured**…". `i12`'s sentence stays as written; this is the addition it invites. The - bound in time matters: the fallback mints a post-rule nonce, and a decline the cycle then - makes and records under it is its own, not an inherited one — treating every decline as - absent would leave that cycle unable to release any hold it raises, **D7** unmet. -6. **The human-exception "which commit" rule** (C:988–990 / W:1172–1174, `h3`–`h6`). Unchanged - in wording. The decline block carries its own rule, deliberately different — written when - made and restated in every later body — because a decline must survive to the next pass - while the human exception only has to survive to history. + binding, and answer records not attributable to the nonce this fallback minted treated as + absent, so that no inherited hold is released and no inherited obligation is claimed — + records made and recorded under that nonce are the cycle's own and are honoured**…". + `i12`'s sentence stays as written; this is the addition it invites. The bound in time + matters: treating *every* answer as absent would leave a fallback cycle unable to release + any hold it raises, **D7** unmet. +6. **The human-exception "which commit" rule** (C:988–990 / W:1172–1174, `h3`–`h6`). Unchanged. + The answer records carry their own, deliberately different — written when made and restated + in every later body — because an answer must survive to the next pass, while the human + exception only has to survive to history. 7. **The gate-off surface** (C:143–151 / W:350–358). OLD tail: "…silencing reminders; or not running a pass and reporting that it ran." NEW tail: "…silencing reminders; not running a - pass and reporting that it ran; **or recording a decline nobody made, or one on an in-set - finding**." The list says it is not complete; this change creates a route and names it, - as the parent did for the stated floor. + pass and reporting that it ran; **recording a decline nobody made, or one on an in-set + finding; or deleting an acceptance the cycle owes**." The list says it is not complete; + this change creates those routes and names them, as the parent did for the stated floor. 8. **The closing message and the soft-reset path** (C:834–838 / W:1018–1022, with the `git reset --soft` sentence at C:829–830 / W:1013–1014). OLD: "**The closing message carries the validated evidence entry for every cited profiled story** — one each, and none @@ -337,60 +346,57 @@ are in the site map and re-read at execution. wholesale, so an entry written only into the WIP body is destroyed exactly when the cycle closes." NEW: "**The closing message carries the validated evidence entry for every cited profiled story** — one each, and none for a cited unprofiled story, which owes no entry — - **and every decline record the cycle made, including any held only in WIP bodies a + **and every answer record the cycle made, including any held only in WIP bodies a `git reset --soft` collapsed**: the single commit after the reset carries all of them, because a body the reset discards is unreachable from the commit that replaces it. The amend replaces the WIP message wholesale, so an entry or record written only into the WIP - body is destroyed exactly when the cycle closes." The soft-reset sentence itself is - unchanged; this sentence is where the carry duty already lives. -9. **The one-contract paragraph** (C:879–890 / W:1063–1074). OLD opening: "**These records are - one contract, and a partial adoption breaks it.** The nonce, the slot naming, the - provenance line, the curve, this carry rule **and the unknown-start activation semantics - that say what a cycle owes when its starting rules cannot be established** depend on one - another," NEW opening: "**These records are one contract, and a partial adoption breaks - it.** The nonce, the slot naming, the provenance line, the curve, this carry rule, **the - closure ordering, the decline record and the severity rule's raw-versus-effective split**, + body is destroyed exactly when the cycle closes." The soft-reset sentence is unchanged. +9. **The one-contract paragraph** (C:879–890 / W:1063–1074). OLD opening: "…this carry rule + **and the unknown-start activation semantics that say what a cycle owes when its starting + rules cannot be established** depend on one another," NEW opening: "…this carry rule, **the + closure ordering, the answer records and the severity rule's raw-versus-effective split**, and **the unknown-start activation semantics that say what a cycle owes when its starting rules cannot be established** depend on one another," and after "and a carry rule naming - records a project does not produce is inert." (C:885) add: "an ordering without the decline - record is a hold whose declined direction has no release rule; a decline record without the - ordering is a release with no hold to release; an ordering whose raw-versus-effective split - has no counterpart in the severity rule, or a severity rule still calling that question - unsettled and mandating a stop beside an ordering that decides it, is two answers to one - question; and a decline record missing from the squash carry or the nonce set is a record - that cannot survive a merge or be attributed." The stop sentence that follows is unchanged + records a project does not produce is inert." (C:885) add: "an ordering without the answer + records is a membership stop whose two answers nothing carries; answer records without the + ordering are a release and an obligation with no stop that asks for them; an ordering whose + raw-versus-effective split has no counterpart in the severity rule, or a severity rule still + calling that question unsettled and mandating a stop beside an ordering that decides it, is + two answers to one question; and an answer record missing from the squash carry or the nonce + set cannot survive a merge or be attributed." The stop sentence that follows is unchanged and now covers these states. Because `/workflow-init` can merge this paragraph and the - pieces independently, the two blocks each carry a reciprocal one-contract sentence of their - own (§3, last sentence; §4, the decline block), so a copy holding one without the others - stops on the text it has, not only on a paragraph it may not have. A rollback while a cycle - is open under the new text needs no new mechanism: "a cycle already running finishes under - the rules it started with" and "A revert is itself a shipping commit for the old rules" - (C:153–154, C:165–166) already govern it — and where a rollback leaves that cycle unable to - establish the rules it started with, the unknown-start fallback is the answer, extended by - item 5 with the loop rules. The residual, stated rather than papered over: the rule text a - cycle ran under is recorded nowhere, so "finishes under the rules it started with" is - recoverable only while that text is still present, and the fallback is what covers the rest. + pieces independently, **each of the three pieces carries a reciprocal one-contract sentence + naming the other two** — §3's last sentence, the answer-record block's last paragraph, the + (g) replacement's last sentence — so a copy holding one without the others stops on the text + it has, not only on a paragraph it may not have. A rollback while a + cycle is open is governed by "a cycle already running finishes under the rules it started + with" and "A revert is itself a shipping commit for the old rules" (C:153–154, C:165–166) + where the cycle can still establish those rules, and by the unknown-start fallback where it + cannot. **Where the rollback removes the fallback text itself, neither is available** and + that cycle **stops and is a human's to resolve — it does not close**. The residual under all + of it: no record identifies the rule revision a cycle started under, so "finishes under the + rules it started with" is recoverable only while that text is present. Named, not fixed. 10. **The no-identity rule's aftermath** (C:422–424 / W:616–618). OLD: "**Starting a new cycle does not close, adopt or retire the cycles those candidates belong to** — they stay open, keep their own nonces, and are a human's to resolve; the new cycle simply does not claim them." NEW: "**Starting a new cycle does not close, adopt or retire the cycles those candidates belong to** — they stay open, keep their own nonces, and are a human's to resolve; the new cycle simply does not claim them, **and names them in its first pass - report**, so the human this rule makes responsible learns they exist." It adds no record: - the report names state the workspace already holds, and a cycle left open that nobody is - told about is the one shape this rule's "a human's to resolve" cannot reach. + report**, so the human this rule makes responsible learns they exist." It adds no record — + the report names state the workspace already holds — and closes the one shape "a human's to + resolve" cannot reach: an open cycle nobody is told about. The shorter "Copy every record into the squash body" sentence inside the human-exception block -(C:1004–1007) is generic and already covers a decline record; it is not edited. The "records -every cycle owes" list (C:698) enumerates unconditional records only; a decline record is +(C:1004–1007) is generic and already covers an answer record; it is not edited. The "records +every cycle owes" list (C:698) enumerates unconditional records only; an answer record is conditional, like the human exception, and is not added. --- ## 5. Edits to the existing passages, with the old-conditions accounting -Ids are the inventory's. "Kept" = the sentence stays; "moved" = it now lives in the §3 block; -"replaced" = the condition changes, and says how; "dropped" carries its reason. +Ids are the committed inventory's (§2). "Kept" = the sentence stays; "moved" = it now lives in +the §3 block; "replaced" = the condition changes, and says how; "dropped" carries its reason. **(a) The floor paragraphs** (C:72–136 / W:279–343) — **trim** `a17`–`a19` and point at the block. OLD: "Your final pass must be clean — if the pass at the floor still finds @@ -398,20 +404,31 @@ Blocker/Major, keep going until clean or clearly stuck → then STOP and surface only early exit below the floor is a pass with **zero** findings; don't manufacture findings to pad." NEW: "Your final pass must be clean; how a cycle closes, and what stops it short of closing, is the closure ordering below — don't manufacture findings to pad." Accounting: -`a1`–`a16`, `a20`–`a22` kept; `a17` kept (the clause survives); `a18`, `a19` moved. `a13`'s -prohibition on restating is what the pointer form obeys. - -**(b) What a loop absorbs** (C:195–223 / W:402–426) — **two sentence edits**, the triggers stay. -`b12` OLD: "…and it resumes the moment the user says whether the set now includes it." NEW: -"…and it resumes once the user has said whether the set now includes it — accepting or -declining it — together with every other answer that pass's suspensions require, any decline -among them recorded as Mechanics requires, under the closure ordering above." `b17`–`b18` OLD: "Stopping this way is **not -an exit from the gate**: the floor, the Blocker/Major filter and the clean-final-pass rule all -stand, and the loop resumes on the revised artifact once the question is answered — what the -stop prevents…" NEW: "Stopping this way is a **suspension** under the closure ordering above, -and the loop resumes on the revised artifact once the question is answered — what the stop -prevents…". Accounting: `b1`–`b16`, `b18` kept; `b17` moved (the block's "a suspension waives -nothing" sentence). W's three wording differences in this passage are untouched (§6). +`a1`–`a17`, `a20`–`a22` kept; `a18`, `a19` moved. The pointer form is what `a13`'s prohibition +on restating requires. + +**(b) What a loop absorbs** (C:195–223 / W:402–426) — **three sentence edits**, the triggers +stay. `b12` OLD: "…and it resumes the moment the user says whether the set now includes it." +NEW: "…and it resumes once the user has said whether the set now includes it — accepting or +declining it — **together with every other answer that pass's suspensions require, each +recorded as Mechanics requires**, under the closure ordering above." `b17`–`b18` OLD: +"Stopping this way is **not an exit from the gate**: the floor, the Blocker/Major filter and +the clean-final-pass rule all stand, and the loop resumes on the revised artifact once the +question is answered — what the stop prevents…" NEW: "Stopping this way is a **suspension** +under the closure ordering above, and the loop resumes on the revised artifact **once every +answer that pass's suspensions require has been given** — what the stop prevents…". Third, in +**W only**, the cross-reference "…by its severity exactly as the severity rule already says…" +becomes "…exactly as Mechanics already says…", matching C: W does carry a +`### Mechanics (reference)` heading (W:968), so the divergence rested on a false premise (§6). + +Accounting: `b1`–`b11`, `b13`–`b16` kept; `b12` **replaced** — old: an *immediate* resume on +the membership answer alone. Kept from it: that the membership answer is what the stop asks +for and that either direction ends it. Changed: the resume waits for every answer the pass's +suspensions require, each recorded, because a loop resumed over an unanswered question decides +it by running. `b17` moved (the block's "a suspension waives nothing" sentence); `b18` +**replaced** the same way and for the same reason — old: resume once *the* question is +answered; new: once *every* required answer is given. W's remaining wording differences here +are untouched (§6). **(c) Recognizing clearly stuck** (C:225–246 / W:428–449) — **keep** the precedence sentence in full and **trim** what follows it. The sentence kept byte-for-byte (**D3**): "That third @@ -431,20 +448,16 @@ suspends, and the branch states it rather than leaving it to be read out of a ne is what finding 1 of pass 3 caught; `c15`, `c16`, `c17` kept in the pointer sentence; `c18` **replaced, narrowed** — old: a clearly-stuck surface credits no pass as clean; new: no pass that surfaced a *scope stop* is credited as clean. Deliberate, with its authority: the story -states the fourth duty as "no pass carrying a surfaced *finding* counts as clean" (story §1), -and the clearly-stuck exit requires regenerating Blocker/Major findings, so a pass surfacing -it is already unclean by the ordering's first step without a separate rule. The two-tell stop -surfaces no finding, so a Blocker/Major-free pass below the floor tripping two tells is clean -*as a pass* and still cannot close — below the floor nothing closes, and at or above it clean -completion outranks the tells by **D2** — so the narrowing opens no bypass in either region; -`c19` **replaced** — old: the loop resumes on whatever the user decides, -unconditionally; new: it resumes on a resuming answer and a stop leaves the suspension -standing — because an unconditional resume is the stop-with-no-transition path AC 4 forbids; -`c20` **dropped** — it argued that reading the exit as "stop instead of fixing" would put it -in competition with the resolve duty, and the block classifies the exit as a suspension that -waives nothing, which is that argument's conclusion stated as a rule; a rationale for a -competition the ordering no longer permits would be prose about a rule that no longer -applies. +states the fourth duty as "no pass carrying a surfaced *finding* counts as clean" (story §1); +the clearly-stuck exit requires regenerating Blocker/Major findings, so a pass surfacing it is +already unclean by the ordering's first step, and the two-tell stop surfaces no finding, so a +Blocker/Major-free pass below the floor tripping two tells is clean *as a pass* and still +cannot close. No bypass opens in either region; `c19` **replaced** — old: the loop resumes on +whatever the user decides, unconditionally; new: it resumes on a resuming answer and a stop +leaves the suspension standing, because an unconditional resume is the stop-with-no-transition +path AC 4 forbids; `c20` **dropped** — it argued that reading the exit as "stop instead of +fixing" would compete with the resolve duty, and the block states that argument's conclusion +as a rule instead. **(d) From pass 4 onward** (C:255–261 / W:459–465) — **add** the Q6 sentence (§7). `d1`–`d7` kept, unchanged. @@ -470,42 +483,42 @@ the whole paragraph, both copies, with the answer. NEW: closure ordering above). Two reasons for the split. The curve must stay derivable from the findings files alone — counting finding lines and leading `BLOCKER` fields per pass reproduces it, which is the only thing that makes a self-reported curve checkable. And the - demotion is the author's judgement about the fix set; a loop spending passes on findings the - author keeps demoting is exactly what the prose-cluster tell exists to surface, and lowering - the counts by that same judgement would hide it. + demotion is the author's judgement about **the finding's repair severity**, never about + which findings the fix set contains — a separate predicate the closure ordering defines, and + one this must not be read as touching; a loop spending passes on findings the author keeps + demoting is exactly what the prose-cluster tell exists to surface, and lowering the counts by + that same judgement would hide it. **This split, the closure ordering and the answer records + are one contract**: the ordering's predicates read the field this sentence assigns them, and + its membership stop is answered by those records, so any one of the three present without + the others is a partial adoption that stops under the one-contract rule above. ``` Accounting: `g1` **dropped** (the unsettled statement, now settled); `g2`, `g3` **dropped** -(the interim report-and-stop duty existed only until the question was settled, and its trigger -no longer exists); `g4` **dropped** in C (the ownership sentence, discharged by this change), -and W, which never carried it, gets the same replacement paragraph — so the one deliberate -story-path difference between the copies is removed. +(the interim report-and-stop duty existed only until the question was settled); `g4` +**dropped** in C (the ownership sentence, discharged by this change), and W, which never +carried it, gets the same replacement — removing the one deliberate story-path difference. **(h) Recording a human exception** (C:977–1033 / W:1161–1217) — **unchanged**, and the -"**Recording a decline.**" block (§4, verbatim) is added immediately after its last paragraph, -before "- **Timeout / abort:**". `h1`–`h26` kept; `h17` stays true of the human-exception form, -and the decline block says which one item of that list it *is* the answer to. +"**Recording an answer at a membership stop.**" block (§4, verbatim) is added immediately after +its last paragraph, before "- **Timeout / abort:**". `h1`–`h26` kept; `h17` stays true of the +human-exception form, and the answer-record block says which item of that list it *is* the +answer to. **(i) When these rules bind** (C:153–167 / W:360–374) — **extend** the strict-reading list -(§4 item 5). `i1`–`i16` kept. - -**(j) The squash carry** (C:892 / W:1076) — **extend** (§4 item 1). `j1`–`j4` kept. - -Also touched, outside the inventoried passages (§4 items 2, 3, 4, 7, 8, 9, 10): the nonce set -and exemption, the "Both shipped records" count, the gate-off list, the closing-message -sentence, the one-contract paragraph, the no-identity rule's aftermath. +(§4 item 5); `i1`–`i16` kept. **(j) The squash carry** (C:892 / W:1076) — **extend** (§4 item +1); `j1`–`j4` kept. Also touched, outside the inventoried passages: §4 items 2, 3, 4, 7–10. --- ## 6. Parity -The two copies must agree on every rule this spec changes. The block (§3), the decline block -(§4), the (g) replacement, the Q6 sentence and every list extension ship byte-identical in C +The two copies must agree on every rule this spec changes. The block (§3), the answer-record +block (§4), the (g) replacement, the Q6 text and every list extension ship byte-identical in C and W. The pre-existing divergences the inventory found are handled as follows: | Divergence | Kind | This change | |---|---|---| -| (b) cross-reference "Mechanics" (C) vs "the severity rule" (W) | deliberate: W has no section by that name | **as-is**, stated | +| (b) cross-reference "Mechanics" (C) vs "the severity rule" (W) | **not deliberate**: the inventory's reason — that W has no section by that name — is false; `### Mechanics (reference)` is at W:968 | **aligned**: W takes C's wording (§5(b)) | | (b) "exactly how" (C) vs "is how" (W); closing rationale reworded; C-only `infinite-portfolio-canvas` parenthetical | rationale and field citation | **as-is**, stated | | (e) "you report" (C) vs "report" (W); C-only "Recorded rationale" Bricks paragraph | rationale | **as-is**, stated | | (f) evidence framing; C-only hypothesis qualifier; punctuation of the late-Blockers clause | rationale | **as-is**, stated; (f) is not edited | @@ -514,9 +527,9 @@ and W. The pre-existing divergences the inventory found are handled as follows: The parity check at execution: extract each edited passage from both files by its lead phrase and `diff` them; the only differences permitted are the rows marked as-is above, and the (g) -row is gone. Any other difference is a defect, not a wording choice. The check's result is part -of the evidence entry (§9). The Mechanics region the decline block joins is byte-identical -between the copies today (the site map confirmed C:840–1033 = W:1024–1217) and stays so. +row is gone. Any other difference is a defect, not a wording choice, and the result is part of +the evidence entry (§9). The Mechanics region the answer-record block joins is byte-identical +between the copies today (C:840–1033 = W:1024–1217) and stays so. --- @@ -527,41 +540,49 @@ earlier pass had removed.": ``` **Where an earlier pass's findings file is unavailable**, the report says so before it reads -anything, and it establishes what "unavailable" means by two checks. First the root: the -slots live in `.context/codex-reviews/` under the top-level directory of the checkout this -cycle is running in — the same root every pass call passes as `workingDirectory` — and a -report reading any other directory would count files it never wrote or miss files it did, so -a mismatch is a **wrong-root state**, reported as such and never as absent history; this -check is what detects it, nothing upstream does. Then, per earlier pass, one of three states, -told apart by what the slot holds: the slot path **absent** — after a fresh checkout, a -cleared `.context/`, a cycle resumed elsewhere — is unavailable history; the slot **present, -accepted when it ran, and now failing validation** is a formerly valid pass whose artifact is -unusable, and its pass number stays counted while its series read `?` in every comparison -that uses them, as the curve already admits, since a corrupted record does not un-run a pass; -a pass **known incomplete when it ran** is excluded, as today, and a present-but-invalid slot -whose acceptance nothing records is read as that case — the direction that costs a pass -rather than credits one. The report then reads the cycle's working record if one exists and -says whether it used it, states which of the three lines it computed from the files it has, -names the ones it could not and why, and says that the two-tell threshold is being read on -that reduced record — reduced concretely: the trend and the require↔withdraw comparison read -only the passes present, so a rising count or a pair that spans a missing or unusable pass -cannot be seen. **A gap does not break the series**: the consecutive *available* pass numbers -compare, and the report **names the missing pass numbers** so a reader can see which -comparison spans one. Refusing to compare across a gap would silence tell detection exactly -where the record is thinnest, and the tells read the direction of the loop rather than any one -adjacent pair. It is not a stop of its own, and it does not make the working record -mandatory: a report that says what it could not see is the duty; a report that invents the -trend, or omits the line without saying so, is the failure. +anything. **First the root.** The slots live in `.context/codex-reviews/` under the top-level +directory of the checkout this cycle is running in — the same root every pass call passes as +`workingDirectory`. A slot written under a different root is **indistinguishable from an +absent slot**: nothing records where a past call ran, so no check can tell them apart. Where +the current root cannot be established, the cycle **stops** and says so rather than reporting +benign unavailable history, and the recovery is to re-run the pass from the correct root. +**Then, per earlier pass, two questions.** Is the slot present and valid? And is the pass +**known accepted** — validated by this session, or recorded as accepted by a cycle record? +Present and valid is the ordinary case. Known accepted but absent, or present and now failing +validation, is a real pass whose artifact is unusable: its number **stays counted** and its +series read `?`, as the curve grammar already admits, since a corrupted or missing record does +not un-run a pass. **Acceptance unknown**, whether the slot is absent or present-but-invalid, +is **not counted toward the floor** and its series read `?` — crediting an unvalidated pass is +the dangerous direction and this is the other one. A pass **known incomplete when it ran** is +excluded, exactly as today. +**Then the report.** It reads the cycle's working record if one exists and says whether it +used it; states which of the three lines it computed and from which passes; **names the pass +numbers it could not read**, and why; and says the two-tell threshold is being read on that +reduced record. **A gap does not break the series**: the consecutive *available* passes +compare across a missing one, and a comparison that spans a gap is **visible but weaker +evidence**, named as such — while a comparison needing the missing pass as one of its two +endpoints is simply unavailable. Refusing to compare across a gap would silence tell detection +exactly where the record is thinnest, and the tells read the direction of the loop rather than +any one adjacent pair. This is not a stop of its own, and it does not make the working record +mandatory: a report that says what it could not see is the duty; one that invents the trend, +or omits a line without saying so, is the failure. Filled, over passes 1, 2 and 4 with 3 +unreadable: + + Root: `.context/codex-reviews/` under this checkout — confirmed. (Unconfirmable: + STOP, root not established; re-run from the correct root.) + Trend: findings 24, 17, —, 18; Blockers 5, 2, —, 1 — pass 3 not counted (slot + absent, acceptance unknown; series `?`); 2→4 spans that gap: rising, weaker + evidence. Cluster: product behaviour 12 of 18. require↔withdraw: none visible; + a pair with pass 3 as an endpoint cannot be read. Threshold read on 3 of 4. ``` -This is **D10**. The trend and the require↔withdraw pair are derivable from the mandated -findings files alone (the `fic2` record verified that derivation reproduces the reported -figures), so unavailability is a property of the workspace, not of the format, and the answer -is disclosure rather than a new stop or a new mandatory artifact. The partition is by what the -agent observes, per `docs/prompt-standards.md` item 10, and the middle state exists because a -file validated when its pass was accepted can be truncated or replaced afterwards: reading it -as incomplete would retroactively uncount a pass, reading it as usable would feed a broken -record into the trend, and `?` is the curve grammar's word for exactly that. +This is **D10**. Both historical lines are derivable from the mandated findings files alone +(the `fic2` record verified that), so unavailability is a property of the workspace, not of the +format, and the answer is disclosure rather than a new stop or mandatory artifact. The +partition is by what the agent can observe (`docs/prompt-standards.md` item 10): presence and +validity from the slot, acceptance from a record or this session, and where acceptance cannot +be established the count moves in the direction that costs a pass. The example is there for +item 4. --- @@ -569,15 +590,21 @@ record into the trend, and `?` is the curve grammar's word for exactly that. Plan C's Tasks 19 and 20 (`docs/superpowers/plans/2026-08-30-review-loop-economics-plan-c-rollout.md`, the drop note under Task 19) deferred "a general production" for a short deterministic -discriminator in the nonce's slot position to this story. It ships nothing here, because the -case it served no longer exists: every cycle started after the parent's rules bind holds a -nonce, and a cycle whose start cannot be established mints one rather than claiming `none -(pre-rule)` (`i11`). A rule for a no-nonce cycle would legislate for an unreachable state, -which is AC 1's prohibition. The `rle` naming that cycle used stays what its closing body -recorded it as — a plan-local exception under the old rules. The durable prior record is Plan -C's two drop notes (the Task 19 heading at line 970 and the Task 20 heading at line 1047 of -that plan, each followed by its note) and Task 23's second point (line 1298), which records -the `rle` exception; this section is the record of the dissolution; no prompt text changes. +discriminator in the nonce's slot position to this story. It ships nothing here, because for +**post-rule and unknown-start cycles** the case it served does not arise: every cycle started +after the parent's rules bind holds a nonce, and one whose start cannot be established mints +one rather than claiming `none (pre-rule)` (`i11`). A production for those would legislate for +an unreachable state, which is AC 1's prohibition. **The claim reaches no further, and one set +of cycles remains.** A cycle that began before the parent's rules shipped has no nonce, cannot +acquire one, and writes `cycle none (pre-rule)`; a rollback can make old-rule cycles reachable +again. That set is bounded and self-terminating — no later cycle can enter it — but while two +such cycles are observably live they compute the same bare slot paths, and **serializing them +is a human's job, not a shipped production**: one runs, then the other. Shipping the rejected +discriminator to serve a closing set would add a permanent rule for a temporary state. The +`rle` naming stays what its closing body recorded it as — a plan-local exception under the old +rules. The durable prior record is that plan's two drop notes (Task 19, line 970; Task 20, +line 1047) and Task 23's second point (line 1298); this section records the dissolution, and +no prompt text changes. --- @@ -586,9 +613,9 @@ the `rle` exception; this section is the record of the dissolution; no prompt te **Battery.** The quality command in `AGENTS.md` § Commands runs once, green, at the Gate-B WIP commit (`check-version-bump.sh main` needs the committed bump, §10). -**The check.** Inline shell asserts in the plan, run against the working tree and against the -parent tree, so the same assertion is observed passing where the change exists and failing -where it does not: +**The check.** Inline shell asserts in the plan, run against the working tree and the parent +tree, so each assertion is observed passing where the change exists and failing where it does +not: - lead phrase of the §3 block, `**How a cycle ends — one ordering`, count **1** in C and **1** in W; the same greps against `git show 7c0d475:CLAUDE.md` and @@ -600,19 +627,21 @@ where it does not: what the removal check above cannot do alone: deleting the old sentence and installing nothing satisfies the removal count and the parity diff, and reports the central demotion/loop-health outcome as verified while both copies carry no answer at all; -- lead phrase `**Recording a decline.**` count **1** in each; **0** in the parent tree; -- the six-member squash-carry sentence: `every decline record` count **1** in each; **0** in +- lead phrase `**Recording an answer at a membership stop.**` count **1** in each; **0** in the + parent tree, and with it both labels: `Accepted: ` and `Declined: ` count + **1** in each and **0** there, since one label shipping without the other is the shape the + block's whole form exists to prevent; +- the six-member squash-carry sentence: `every answer record` count **1** in each; **0** in the parent tree; - the Q6 lead phrase `**Where an earlier pass's findings file is unavailable**` count **1** in each; **0** in the parent tree. -If the claim "the ordering and the record ship in both copies" were false, one of the -working-tree counts would be **0** or the parent-tree counts would not differ from it. The -wiring can produce that observation: each grep reads the file bytes at the named revision, -nothing supplies its own input, and the parent tree is the actual prior state. **The -counterfactual is ABSENT, and is claimed as absent**: the parent carries no ordering block and -no decline record, and the (g) count is the one site where the parent is present and the change -removes it. No count against the parent is claimed as "contradictory". +If the claim "the ordering and the records ship in both copies" were false, one working-tree +count would be **0** or the parent-tree counts would not differ from it. The wiring can produce +that observation: each grep reads the file bytes at the named revision and nothing supplies its +own input. **The counterfactual is ABSENT, and is claimed as absent** — the parent carries no +ordering block and no answer records, and the (g) count is the one site where the parent is +present and the change removes it. Nothing is claimed as "contradictory". **The named verification of the risk path** (story AC 4). Walk every stop the shipped text names — membership stop, question stop, a finding carrying both triggers, clearly-stuck exit, @@ -621,60 +650,60 @@ suspensions at once, a below-floor clean pass, a zero-finding pass, the unknown- fallback — **and the stateful transitions** — a declined finding re-raised matching on all five fields, and re-raised with one field changed; the fix set broadened to include a declined finding; recovery after an accept and after a stop with the session lost (no -identity → new cycle); a `full` Gate-B pass with one branch clean and the other carrying an -in-set Blocker; a rollback with a cycle open under the new rules; a copy adopting the -ordering without the decline block, and the reverse — and write the **next-state table**: -for each row, the **record state** (which decline records the bodies carry, which working -record exists) and the **governing-scope state** (the fix set as currently assigned) as -explicit input columns beside the user's answer, then the input that ends the row and the -state the cycle is in afterwards, citing the shipped line the row reads, in both copies. The +identity → new cycle), **both with the answer records present in the branch's bodies and with +them absent**; a `full` Gate-B pass with one branch clean and the other carrying an in-set +Blocker; a rollback with a cycle open under the new rules, and a rollback that removes the +fallback text; a copy adopting the ordering without the answer records, the reverse, and +either without the severity split — and write the **next-state table**: for each row, the +**record state** (which answer records the bodies carry, which working record exists) and the +**governing-scope state** (the fix set as currently assigned) as explicit input columns beside +the user's answer, then the input that ends the row and the state the cycle is in afterwards, +citing the shipped line the row reads, in both copies. The table lives in the plan and is quoted by the closing commit body, not here. What would be observed if the claim "no path leaves a cycle unable to close and unable to suspend" were -false: a row whose next state is the same stop with no input consumed — the shape the parent -cycle shipped once — or a row that closes with an in-set Blocker standing. The wiring can -produce it because every row is filled from the shipped text rather than from this spec, and -the inputs the `fic2` instrument omitted — the user's answer, and the decline — are columns -here; so is the stop answer, whose next state is a standing suspension by design and must -read as one, not as the same stop re-raised. - -**What this is not.** Not the `fic2` decision matrix: that instrument scored N states against -old and new text with an expected output each, and Gate B found two defects in the technique — -a state's inputs must include every input the rule reads (the user's answer was never a -column), and a counterfactual must distinguish ABSENT from CONTRADICTORY. The check above is a -presence test whose parent state is absent by inspection; the walk takes the answer as an -input and produces a next state rather than a scored output. **No fixture per predicate is +false: a row whose next state is the same stop with no input consumed, or one that closes with +an in-set Blocker standing. The wiring can produce it because every row is filled from the +shipped text rather than from this spec, and the inputs the `fic2` instrument omitted — the +user's answer and the record state — are columns here. + +**What this is not.** Not the `fic2` decision matrix, whose two defects Gate B found in the +technique itself: a state's inputs must include every input the rule reads, and a +counterfactual must distinguish ABSENT from CONTRADICTORY. The check above is a presence test +whose parent state is absent by inspection; the walk takes the answer and the record state as +inputs and produces a next state rather than a scored output. **No fixture per predicate is built** — parked in the story's §2, not reopened. -**Evidence entry**, in the closing commit body, names: the battery run; the six assert pairs -with their working-tree and parent-tree counts; the §6 parity diff — the passages extracted, -the differences observed, and that each is one of the permitted rows; and the next-state -table's location in the plan plus its row count. It is revalidated before every Gate-B -re-review and before the closing amend, as §5 requires. +**Evidence entry**, in the closing commit body, names: the battery run; every assert pair +above with its working-tree and parent-tree counts; the §6 parity diff — the passages +extracted, the differences observed, and that each is one of the permitted rows; and the +next-state table's location in the plan plus its row count. It is revalidated before every +Gate-B re-review and before the closing amend, as §5 requires. --- ## 10. AGENTS.md invariants touched - **Invariant 11 — `docs/prompt-standards.md`, all 12 items.** Most at risk: item 6 (rules - carry their why — each constraint in both shipped blocks carries its reason in the same - sentence, and the plan's review reads each sentence for one); item 8 (token-lean — the + carry their why — every constraint in the shipped blocks carries its reason in the same + sentence **except the ordering's three settled axioms**: only clean completion closes, a + zero-finding pass is clean whatever the floor, and either answer is available only at a + membership stop. Those are asserted deliberately — their reasons are **D1**, the existing + zero-finding exit and **D6**, settled in the story and not re-argued in a prompt — and the + plan's review reads every *other* sentence for an inline why); item 8 (token-lean — the blocks replace closure sentences rather than adding beside them, `a13` being why the old - ones leave); item 10 (the Q6 partition and its root check); and item 3 (stop conditions - defined — the stop answer is a named, resumable state). + ones leave); item 4 (output structure shown — the Q6 example); item 10 (the Q6 partition and + its root check); item 3 (the stop answer is a named, resumable state). - **Don't: "Never replace a decision procedure without accounting for its old conditions."** - Satisfied by §5: every inventory id for every edited passage is marked kept, moved, replaced - or dropped with a reason. + Satisfied by §5, against the committed inventory. - **Don't: "Never rename or delete a doc section without grepping for references first."** The - (g) sentence names the story path; the grep finds it at `CLAUDE.md:815` (the site itself), - `docs/superpowers/plans/2026-08-29-review-loop-economics-plan-a-rules.md:878`, - `docs/superpowers/specs/2026-08-28-review-loop-economics-design.md:33` and - `docs/superpowers/stories/2026-08-28-review-loop-economics-pass-floor-story.md:78`. All - four cite the story file, which continues to exist; none cites the sentence. Nothing breaks. + (g) sentence names the story path; the grep finds it at `CLAUDE.md:815` (the site itself) + and in three artifacts of the parent cycle (`…plan-a-rules.md:878`, + `…review-loop-economics-design.md:33`, `…pass-floor-story.md:78`). All cite the story file, + which continues to exist; none cites the sentence. Nothing breaks. - **Invariant 12 — a plugin change requires a version bump.** `workflow-init.md` is under `plugins/`, so `plugins/dev-workflow/.claude-plugin/plugin.json` goes `0.11.0 → 0.12.0` with a - `CHANGELOG.md` entry: a minor bump, because the template gains a record type and a rule. - The entry also notes that the squash-carry sentence now lists six record kinds (the 0.9.0 - entry described two, and history entries are not edited). + `CHANGELOG.md` entry: a minor bump, because the template gains a record form and a rule. + The entry also notes that the squash-carry sentence now lists six record kinds. - **Invariant 4 / the hook.** Untouched: `codex-gate.sh` is not edited, and the §5 heading it greps (`Cross-Model Review`) does not move. @@ -686,5 +715,5 @@ re-review and before the closing amend, as §5 requires. - The CodeRabbit plan-metadata contradiction (`fic2` record). - The fixture-per-predicate question (`fic2` record; story §2). - Hook code under `plugins/dev-workflow/hooks/`. -- `todos.md`: both-branches-misread-each-other; self-consuming-deletion (prompt-standards - item 11); the three bot findings in resolved plans. +- `todos.md`: both-branches-misread-each-other; self-consuming-deletion (prompt-standards item + 11); the three bot findings in resolved plans. From 46256798fd57d3d78c1833d4a1f60e5b0111ad1e Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Thu, 10 Sep 2026 15:56:18 +0200 Subject: [PATCH 007/181] docs(stories): record the accepted scope expansion and its criterion MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Gate-A spec pass 4 raised the loss of accepted repair obligations as a scope stop. Daniel accepted it into scope on 2026-09-10: the change ships an `Accepted:` record label beside the decline record, sharing its form, transport, nonce and carry rules. §2 records the decision and why the reliance is older than this story; §3 gains criterion 7. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../2026-08-29-loop-rule-consolidation-story.md | 16 ++++++++++++++++ 1 file changed, 16 insertions(+) diff --git a/docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md b/docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md index f9e826b..e2648f6 100644 --- a/docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md +++ b/docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md @@ -47,6 +47,15 @@ which are preconditions; and what a user's answer on a surfaced finding does in - **The pass-counter anomaly**, the CodeRabbit plan-metadata contradiction, and the fixture-per-predicate question — all still parked. +**One expansion, authorised 2026-09-10 rather than absorbed.** Gate-A spec pass 4 raised, as a +scope stop, that an accepted repair obligation lives only in the running session: a cycle that +loses it can be replaced by one that closes the same artifact with the repair never made. The +reliance is older than this story — §5 already puts "repair obligations you already accepted in +earlier passes" in the fix set with nothing recording them — but this story is the one writing +the closure rules, so the gap became its question. **Daniel accepted it into scope**: the change +ships a second record label, `Accepted:`, sharing the decline record's form, transport, cycle +nonce and carry rules. Criterion 7 below is what that adds; nothing else in this section moves. + ## 3. Acceptance criteria - [ ] **The ordering is stated once, in both copies, and covers every reachable conflict.** A @@ -74,6 +83,13 @@ which are preconditions; and what a user's answer on a surfaced finding does in Don't. A requirement neither kept nor explicitly dropped is a dropped condition. - [ ] **The two copies stay in parity** on every rule this story changes, deliberate wording differences stated as such. +- [ ] **A user's answer that puts work into the fix set leaves a record, on the same terms as one + that keeps work out.** Both labels share one form, one transport, one attribution rule and + one carry rule, each stated once; recording an answer is a precondition to running the next + pass; and what the record does *not* buy is stated — it makes a lost obligation + recoverable, not automatically recovered, and nothing checks that a replacement cycle + looked. Checkable by reading: two labels with one set of rules, and a residual paragraph + that does not overclaim. *(Added 2026-09-10 with the expansion recorded in §2.)* ## 4. Settled inputs — decided, paid for, and not to be reopened From 726a5ded9664a0b909230e95652a93319234978e Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Thu, 10 Sep 2026 16:34:21 +0200 Subject: [PATCH 008/181] =?UTF-8?q?docs(specs):=20loop-rule=20consolidatio?= =?UTF-8?q?n=20=E2=80=94=20Gate-A=20pass=205=20revision?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Sixteen findings fixed, one Minor collected. Three concepts the design did not need are removed rather than patched, which is what most of pass 5's Majors were saying in different places: - the separate "obligation" a record was said to create, replaced by the shipped severity rules governing an in-set finding; - the question-stop decision the acceptance record was overloaded to carry, which its five fields cannot identify; - the new wrong-root stop, replaced by the INCOMPLETE pass state §5 already defines, so D10's "not a new stop condition" stands. One unsupportable claim goes with them: that the unknown-start fallback covers a rollback which removes the fallback text. The spec now states the residual instead, and §9 drops the walk row that could not cite shipped text. Also: the two senses of "clean" named; the fourth duty restored over its whole domain (c18 kept, not narrowed); the five-field key reading reviewer-written severity; a changed-field re-raise classified afresh; the closure-record contract given one name and one membership list; and the shipped recovery sentence edited so a Gate-A cycle can actually find the records it wrote. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- ...26-09-10-loop-rule-consolidation-design.md | 541 +++++++++--------- 1 file changed, 277 insertions(+), 264 deletions(-) diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index cb3657e..346bf8c 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -1,6 +1,6 @@ # §5 loop-rule consolidation: one closure ordering — Design -**Date:** 2026-09-10 · **Status:** DRAFT — Gate A running (pass 4 revised) +**Date:** 2026-09-10 · **Status:** DRAFT — Gate A running (pass 5 revised) **Story:** `docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md` **Profile:** read from that header at every pass, never from here — it is the only writable copy, and a value copied here would be a remembered value. @@ -8,9 +8,8 @@ copy, and a value copied here would be a remembered value. Prompt-only, in two mirrored copies: `CLAUDE.md` §5 (**C** below) and the inline template in `plugins/dev-workflow/commands/workflow-init.md` (**W** below). No file under `plugins/dev-workflow/hooks/` changes. Line numbers cite the tree at `7c0d475` (main) and are -re-read at execution; the plan carries the `grep -n` sites. - -The old-conditions accounting in §5 cites ids `a1`…`j4`, defined in +re-read at execution; the plan carries the `grep -n` sites. The old-conditions accounting in §5 +cites ids `a1`…`j4`, defined in `docs/superpowers/specs/2026-09-10-loop-rule-consolidation-condition-inventory.md` beside this file: 135 conditions (22 + 18 + 20 + 7 + 11 + 7 + 4 + 26 + 16 + 4) quoted from `7c0d475`, so a reader can check the accounting rather than take it. A sentence outside those passages is @@ -28,26 +27,22 @@ commit-body record form with two labels, `Accepted:` and `Declined:`, so a user' surfaced finding has one rule in both directions and survives the session that made it; the answer to what a severity demotion does to the loop-health counts; the answer to the pass-4 report when prior-pass history is unavailable; and the deferred slot discriminator, dissolved -rather than shipped. - -**What does not.** The pass floor and severity semantics, which the parent shipped and this -spec reads as given, and everything §11 lists as parked. +rather than shipped. **What does not:** the pass floor and severity semantics, which the parent +shipped and this spec reads as given, and everything §11 lists as parked. --- ## 2. Settled inputs The story's §4 table, decisions 1–10, is the design's starting point and is not restated here; -each is cited below as **D1**…**D10** (with **D9b**, **D9c**). Two implementation facts the parent -cycle established are read as given: a pass's cleanliness is a fact about what that pass found -and is **never rewritten** — an answer changes whether the *cycle* may close; and the findings -files establish the **inventory** of findings, not their resolutions. - -This spec adds what the table does not settle — the evaluation order and the field and file -set every predicate reads; the four duties' classification; the scope stop's two triggers and -what each answer does; what a stuck or two-tell answer produces, a stop included; the records' -wording, attribution and recording point, and the `Accepted:` label itself (§4); and the -raw-severity rule for the health measures — each decided in the section that uses it. +each is cited below as **D1**…**D10** (with **D9b**, **D9c**). Two implementation facts the +parent cycle established are read as given: a pass's cleanliness is a fact about what that pass +found and is **never rewritten** — an answer changes whether the *cycle* may close; and the +findings files establish the **inventory** of findings, not their resolutions. What the table +does not settle, this spec decides in the section that uses it: the evaluation order and the +field and file set each predicate reads; the duties' classification; the scope stop's two +triggers and what each answer does; what a stuck or two-tell answer produces; the records' +wording, attribution and recording point and the `Accepted:` label; and the raw-severity rule. --- @@ -64,17 +59,24 @@ and a stated order is what stops them qualifying each other. Every predicate her validated findings file **or files** of the logical pass as one set — a `full` Gate-B pass has two, and one branch alone is already an incomplete pass — at **effective** severity, after the Mechanics severity ceiling, because cleanliness is about what the cycle must repair and -the ceiling is what decides that; only the **health measures** read the reviewer-written field -before it (Mechanics, Severity): the per-pass counts, the clusters, the tells, and **both** -conditions of the stuck reading — its Blocker curve and its regenerating Blocker or Major -findings — since a predicate reading one field for half of itself could not be read at all. - -**First, clean completion.** A pass is clean when it carries no Blocker or Major at effective -severity that is **in the assigned fix set**, and it surfaced no scope stop; a pass with +the ceiling is what decides that. Two things read the reviewer-written field from before the +ceiling: the **health measures** (Mechanics, Severity) — the per-pass counts, the clusters, +the tells, and **both** conditions of the stuck reading, its Blocker curve and its +regenerating Blocker or Major findings, since a predicate reading one field for half of itself +could not be read at all — and the **five-field key** the answer records match on, whose +severity field is the one the reviewer wrote, so that a ceiling cannot make two findings the +same one. + +**First, clean completion.** §5 uses *clean* in two senses and always has. A **clean findings +file** is the `NO FINDINGS` signal the protocol defines. A **clean pass** is the predicate +below: a clean findings file always gives one, and a clean pass need not have one, because a +pass carrying only Minors, Nits or findings outside the fix set is clean without being empty. +**A pass is clean** when it carries no Blocker or Major at effective +severity that is **in the assigned fix set**, and it **surfaced no finding** — the scope stop +and the clearly-stuck exit both surface findings, while the two-tell stop surfaces tells and +not a finding, so it leaves cleanliness alone. A pass with **zero** findings is clean whatever the floor. A finding matched by a decline recorded in -this cycle is outside the set by that decision (Mechanics, the answer records) — a decline -keeps a finding out and never excuses one that is in, so a declined finding the set later -comes to include owes resolution like any other. **A clean pass at or above the derived floor, +this cycle is outside the set by that decision (Mechanics, the answer records). **A clean pass at or above the derived floor, or a zero-finding pass, closes the cycle**, under the final-acceptance preconditions the floor section already states and this ordering does not restate — the cited set and every profile re-read before the pass is accepted as final, a header or profile that changed during it @@ -95,17 +97,15 @@ to one pass: **one surface, every reason reported, every question asked**, becau left out is a decision made by omission. A finding surfaced solely by the stuck or two-tell reading is not a scope-stop finding; one that is also outside the set, or also opens a question, takes the scope stop's answers at that same surface — it is not asked twice. A -re-raised finding that a decline of this cycle matches on all five fields raises **no -membership stop and no membership hold**, and that alone: the five recorded fields record no -question, so a **question stop still fires** on it unless that same question has already been -answered in this cycle. A difference in any of the five fields, or genuine uncertainty, is a -new finding and a new stop of either kind. +re-raised finding that a decline of this cycle matches raises **no membership stop**, and one +differing from it in any of the five fields is a **new finding classified afresh**; Mechanics, +the answer records, states both, and this ordering does not restate them. **Third, a pass that neither closes nor suspends continues** — the loop runs another pass on -the **current** artifact, revised or not. This is the ordinary case and not a residue: below -the floor a clean pass lands here, and so does a pass whose only findings are Minors and Nits, -which are collected and never iterated and so may leave nothing to revise. It is a branch and -not an inference, because "does not close" read alone says nothing about whether to run again. +the **current** artifact, revised or not. Below the floor a clean pass lands here, and so does +a pass whose only findings are Minors and Nits, which are collected and never iterated and so +may leave nothing to revise. It is a branch and not an inference, because "does not close" +read alone says nothing about whether to run again. **The four standing duties, classified.** The **derived floor** is a precondition on closure: it gates closing, and is discharged by the count of valid logical passes reaching @@ -120,9 +120,9 @@ that the finding is not true of the artifact, a decline is the user's decision t finding stays outside the fix set, and only the second is an answer at a membership stop. The **hold** a surfaced finding places on closure is part of the ordering: it gates closing while it stands, and is discharged by the answers that finding requires, below. -**No-clean-credit** — no pass that surfaced a scope stop is credited as clean — is part of +**No-clean-credit** — no pass that **surfaced a finding** is credited as clean — is part of the ordering and a fact about that pass, discharged by nothing: a later pass is judged on its -own findings, so a decline never closes the cycle on the pass that surfaced the finding. +own findings, so an answer never closes the cycle on the pass that surfaced the finding. **What a suspension asks, and what ends it.** Only a scope stop raises a **hold**, on its finding, and the hold ends when every answer that finding requires has been given — one for a @@ -139,9 +139,9 @@ out-of-set finding that opened the question is a membership stop as well and tak decline. **Decline is available only at a membership stop.** The stuck and two-tell readings raise no hold: each asks one question, **continue or stop**. Continue resumes the loop on the artifact as revised and the fix set as the governing artifacts now assign it — where several -plans or stories govern one cycle, the union of the scopes they assign — **plus the obligations -this cycle accepted**, which its commit bodies carry (Mechanics, the answer records). One no -body carries is not in the set, and the reviewer raises its finding again on the next pass like +plans or stories govern one cycle, the union of the scopes they assign — **plus the findings +this cycle accepted into it**, which its commit bodies carry (Mechanics, the answer records). +One no body carries is not in the set, and the reviewer raises it again on the next pass like any other, which is the ordinary route and not a special one. A finding that a narrowing has put outside the set surfaces at the next pass as a membership stop. **Stop leaves the suspension standing**: the cycle stays open under its nonce, resumable by a later @@ -152,50 +152,47 @@ to resolve, exactly as the nonce rules already say of open cycles. resumes only when every answer resumes it — accept or decline at a membership stop, a decision at a question stop, continue at the stuck or two-tell reading; one stop answer leaves the whole suspension standing, because a loop resumed over an unanswered question decides it by -running. Two states cannot co-occur, and no rule ranks them: clean completion and a scope -stop, since a pass that surfaced one is not clean; and a zero-finding pass and any +running. Two states cannot co-occur, and no rule ranks them: clean completion and any +suspension that surfaces a finding — the scope stop, the clearly-stuck exit — since a pass +that surfaced one is not clean; and a zero-finding pass and any suspension, since it has nothing to surface, nothing regenerating, no cluster and no require↔withdraw pair. The Gate-B triviality skip is outside this ordering: a skipped cycle -runs no passes and ends by its own rule. **This ordering, the answer records in Mechanics and -the severity rule's raw-versus-effective split are one contract**: a membership stop's two -answers are carried by those records and by nothing else, and the split says which field each -predicate here reads, so any one of the three present without the others is a partial adoption -that stops under the one-contract rule there. +runs no passes and ends by its own rule. **This ordering is one component of the +closure-record contract** the one-contract rule names, and a copy carrying it without the +rest of that list is a partial adoption that stops there. ``` -Why the shape, where the block's own sentences do not carry it. The evaluation order answers -the objection that a ranking between clean completion and the two-tell stop cannot fire if the -stop can make the pass unclean: clean completion is read first, so a suspension is only ever +Why the shape, where the block's sentences do not carry it. The evaluation order answers the +objection that a ranking between clean completion and the two-tell stop cannot fire if the stop +can make the pass unclean: clean completion is read first, so a suspension is only ever evaluated on a pass that did not close. The three branches are AC 4, the duties paragraph AC 2, the composition and cannot-co-occur sentences AC 1. The scope stop's two triggers are `b11` (membership) and `b13` (question) read separately, because an in-set finding that opens a question can neither join nor stay outside the set and needs its own answer — which is also why a matching decline suppresses the membership trigger only. Sentences the block points at rather -than restating (`a13`): C:116–118, C:760–764, C:422–424, C:176, and the clearly-stuck -paragraph's own precedence sentence, kept verbatim there under **D3**. +than restating (`a13` as §5(a) replaces it): C:116–118, C:760–764, C:422–424, C:176, and the +clearly-stuck precedence sentence, kept verbatim under **D3**. --- ## 4. The answer records — one form, two labels **The rule ships as the block below**, which is the text and not a summary of it; this section -adds only what the block does not carry. Decisions implemented: availability at a membership -stop and nowhere else **D6**; the remainder-of-cycle binding **D7**; the explicit attributable -decision **D8**; the five-field sameness test **D9b**; the unverified-assertion reading -**D9c**; the transport **D9**, whose "closing commit body" is therefore the *last* body to -carry a record, not the first. That a decline never qualifies the Blocker/Major-resolve duty -is **D5**, and it is why §3's clean predicate names the fix set rather than the decline: a -decline is only ever a fact about membership, so it cannot be read as excusing an in-set -finding — the gate-off route a fabricated record would otherwise open. - -**The `Accepted:` label is authorised scope growth beyond the story**, decided by Daniel at -this cycle's Gate-A pass-4 scope stop and recorded in the pass-4 dispositions. The story -shipped a decline record only, and three successive passes found the same hole: an acceptance -put a repair obligation in the fix set and no artifact carried it, so a lost session left a -cycle able to close over work it had agreed to do. The cheapest close was a second label on a -record whose form, transport, nonce and carry rules already existed. This Gate-A cycle runs -under the pre-change rules and writes no such record for its own acceptance; the dispositions -file and the working record are what today's rules provide. +adds only what the block does not carry. Decisions implemented: **D6** availability, **D7** +binding, **D8** the explicit decision, **D9b** the sameness test, **D9c** the unverified +reading, **D9** the transport — whose "closing commit body" is therefore the *last* body to +carry a record, not the first. **D5** is why §3's clean predicate names the fix set rather than +the decline: a decline is only ever a fact about membership, so it cannot be read as excusing +an in-set finding, which is the gate-off route a fabricated record would otherwise open. + +**The `Accepted:` label was beyond the story's *original* scope and is now authorised by it**: +Daniel decided it at this cycle's Gate-A pass-4 scope stop, and the story records the expansion +in its §2 and adds acceptance criterion 7 (at `4625679`). Three passes had found the same hole +— an acceptance put a finding in the fix set and no artifact carried it, so a lost session left +a cycle able to close over work it had agreed to do. The cheapest close was a second label on a +record whose form, transport, nonce and carry rules already existed. This cycle runs under the +pre-change rules and writes no such record for its own acceptance; the dispositions file and +the working record are what today's rules provide. **The block that ships**, in Mechanics, placed **immediately after the whole human-exception passage** — after "**What the record is worth.**" (C:1028–1033 / W:1212–1217) and before @@ -205,17 +202,17 @@ byte-identical in both copies: ```` **Recording an answer at a membership stop.** A membership stop asks whether a surfaced finding joins the assigned fix set, and the user's answer is recorded under one of two - labels sharing one form. **Accepted** puts the finding in the set, where the - Blocker/Major-resolve duty governs it from then on. **Declined** keeps it out and releases - its hold. Both are available at a membership stop and nowhere else: not for an in-set + labels sharing one form. **Accepted** puts the finding in the set, from where the severity + rules already govern it: a Blocker or Major owes resolution, a Minor or Nit is collected and + never iterated. **Declined** keeps it out and releases its hold. Both are available at a + membership stop and nowhere else: not for an in-set Blocker or Major, which owes resolution already, and neither is the answer to a question stop, a stuck or two-tell surface, a below-floor pass, an unclean final pass, or any Gate-A, Gate-B or evidence obligation — of the list the human-exception form is never the answer to, - the membership stop is the one item these records answer. An **acceptance also records a - question stop's decision where that decision put work in the set**, since that is the same - fact under another name: work the cycle now owes. Each must be an **explicit, attributable - decision on that specific finding** — never silence, never a general remark about scope, - never inferred — because a fix set changed by inference is a fix set nobody chose. + the membership stop is the one item these records answer. Each must be an **explicit, + attributable decision on that specific finding** — never silence, never a general remark + about scope, never inferred — because a fix set changed by inference is a fix set nobody + chose. ``` Accepted: · · cycle @@ -250,47 +247,49 @@ byte-identical in both copies: no effect in any later one, and never qualifies the Blocker/Major-resolve duty, which the declined finding never reached. A later pass raises **the same finding** when all five of location, defect, severity, consequence and suggested fix match, read on meaning rather than - bytes, since a reviewer rewrites its sentences between passes. A finding matching a decline + bytes, since a reviewer rewrites its sentences between passes; the severity read is the one + the reviewer wrote, per the ordering's field rule. A finding matching a decline of this cycle raises **no membership stop and no membership hold**, so a decline is not re-asked every pass — **and that alone**: the five fields record no question, so a **question stop** still fires on it unless that same question has already been answered in this cycle (the closure ordering above). **Any difference — severity included — or any - genuine uncertainty makes it a new finding**, and the stop applies. A decline keeps a - finding out and never excuses one that is in: a declined finding the fix set later comes to - include owes resolution like any other. An acceptance does not expire with its pass — the - obligation it records stands until the resolve duty discharges it. - - **What each is worth.** A decline **releases a hold** and an acceptance **creates an - obligation**, neither of which the human-exception form ever does. Narrowness bounds what a + genuine uncertainty makes it a new finding**, classified afresh against the current fix set + and the question predicate rather than inheriting a stop from the finding it resembles. A + decline keeps a finding out and never excuses one that is in: a declined finding the fix set + later comes to include owes resolution like any other. An acceptance does not expire with + its pass: the finding is in the set until the cycle closes. + + **What each is worth.** A decline **releases a hold** and an acceptance **puts a finding in + the fix set**, neither of which the human-exception form ever does. Narrowness bounds what a false record can do — one fully identified finding, one cycle, **as far as distinct nonces allow**: two cycles sharing or redrawing a nonce are indistinguishable to these records as to every other, so a replayed answer can bind to the wrong cycle, and a bound is not safety. **What the acceptance record buys, and what it does not:** a cycle that lost its session and - started fresh can read the branch's commit bodies and find what was accepted. Nothing makes - it read them and nothing checks that it did, so a replacement cycle can still review and - close the same artifact while an older one stays open; the record makes that discoverable - rather than invisible, which is less than preventing it. **The one thing no record carries:** - a stuck or two-tell surface and its continue-or-stop answer — from a closing body a reader - can infer the close, the acceptances and the declines, and nothing about those two. - - **These records, the closure ordering and the severity rule's raw-versus-effective split are - one contract**: these records carry the two answers that ordering's membership stop asks - for, and that split says which severity field each of its predicates reads, so any one of - the three present without the others is a partial adoption that stops under the one-contract - rule above. + started fresh reads the branch's commit bodies among its recovery sources and finds what was + accepted. Nothing makes it act on them and nothing checks that it did, so a replacement + cycle can still review and close the same artifact while an older one stays open; the record + makes that discoverable rather than invisible, which is less than preventing it. **What no + record carries:** a question stop's decision, where it changed no membership — the five + fields identify a finding and not a question, and that decision's durable form is the + artifact revision it produces — and a stuck or two-tell surface with its continue-or-stop + answer. From a closing body a reader can infer the close, the acceptances and the declines, + and after a lost session neither of those two is recoverable from it. + + **These records are one component of the closure-record contract** the one-contract rule + above names, and a copy carrying them without the rest of that list is a partial adoption + that stops there. ```` **Which rules the answer modifies** (story AC 3): the scope stop's resume sentence (`b12`), -and the hold and the clean-pass definition (both in the §3 block, where the answer's two -directions are one rule). Those three are the rules whose **closure behaviour** the answer -qualifies, and it qualifies no other — not the Blocker/Major-resolve duty (Mechanics, -Severity: "both must resolve", C:784 — **D5**: a declined finding never enters its scope, and -an accepted one enters it without changing what the duty demands), not the floor, not the -tells, not the stuck reading. A reader finding either label qualifying a *closure* rule -outside those three has found a defect. The records are separately **named**, without -qualifying closure, at the transport, attribution, activation and threat sites listed next, -and a reader finding them absent from any of those has found the opposite defect. The two -lists answer different questions: what the answer *changes*, and what must *carry* it. +and the hold and the clean-pass definition, both in the §3 block. Those three are the rules +whose **closure behaviour** the answer qualifies, and it qualifies no other — not the +Blocker/Major-resolve duty (Mechanics, Severity: "both must resolve", C:784 — **D5**: a +declined finding never enters its scope, an accepted one enters it without changing what the +duty demands), not the floor, the tells or the stuck reading. A reader finding either label +qualifying a *closure* rule outside those three has found a defect; a reader finding the +records absent from any transport, attribution, activation or threat site listed next has +found the opposite one. The two lists answer different questions: what the answer *changes*, +and what must *carry* it. **Existing sentences that must name them**, both copies, old → new. Line numbers are C's; W's are in the site map and re-read at execution. @@ -301,8 +300,8 @@ are in the site map and re-read at execution. every human-exception record, **every answer record, one copy per cycle nonce and five-field finding key — byte-identical repeats collapse, and copies that disagree stop under the rule below** —, the provenance lines, the curves and any skipped cycle's skip - record TOGETHER WITH THE SKIP REASON IT POINTS AT…". Six members; the dedup clause exists - because a record restated in every body of its cycle reaches the squash range many times. + record TOGETHER WITH THE SKIP REASON IT POINTS AT…". Six members; the dedup clause is there + because a record restated in every body reaches the squash range many times. 2. **The named nonce set** (C:385–387 / W:579–581). OLD: "**and that set is named rather than left open**: the provenance line, the per-pass curve (including a skip record standing in for one), the cycle's findings slots, and its advisory working record." NEW: "**and that set @@ -313,69 +312,68 @@ are in the site map and re-read at execution. this change neither introduces nor keys to a cycle — the evidence entry and a human-exception record among them." NEW: "The nonce is not required in records that are not keyed to a cycle — the evidence entry and a human-exception record among them; **an answer - record is keyed to its cycle and carries it**." (The old sentence's "this change" dated it - to the parent; the new one states the criterion.) + record is keyed to its cycle and carries it**." The old "this change" dated the sentence to + the parent; the new one states the criterion. 4. **"Both shipped records below"** (C:367 / W:561). OLD: "Both shipped records below carry a **cycle field**, because…". NEW: "Both shipped records below carry a **cycle field** — and - so do the answer records in Mechanics — because…". A load-bearing count that further - cycle-attributed records would otherwise falsify. + so do the answer records in Mechanics — because…". A load-bearing count a third + cycle-attributed record would otherwise falsify. 5. **The unknown-start fallback** (C:153–167 / W:360–374, `i4`–`i8`, extended at `i12`'s invitation). OLD: "…at minimum floor 3, severity classified without the demotion, the provenance-line duty owed, the curve duty owed, and the nonce duties at their strictest…". NEW: "…at minimum floor 3, severity classified without the demotion, the provenance-line duty owed, the curve duty owed, the nonce duties at their strictest, **every suspension binding, and answer records not attributable to the nonce this fallback minted treated as - absent, so that no inherited hold is released and no inherited obligation is claimed — + absent, so that no inherited hold is released and no inherited acceptance is claimed — records made and recorded under that nonce are the cycle's own and are honoured**…". - `i12`'s sentence stays as written; this is the addition it invites. The bound in time - matters: treating *every* answer as absent would leave a fallback cycle unable to release - any hold it raises, **D7** unmet. + `i12`'s sentence stays as written; this is the addition it invites. The time bound matters: + treating *every* answer as absent would leave a fallback cycle unable to release a hold it + raised itself, **D7** unmet. 6. **The human-exception "which commit" rule** (C:988–990 / W:1172–1174, `h3`–`h6`). Unchanged. - The answer records carry their own, deliberately different — written when made and restated - in every later body — because an answer must survive to the next pass, while the human - exception only has to survive to history. + The answer records carry their own, deliberately different, because an answer must survive + to the next pass while the human exception only has to survive to history. 7. **The gate-off surface** (C:143–151 / W:350–358). OLD tail: "…silencing reminders; or not running a pass and reporting that it ran." NEW tail: "…silencing reminders; not running a pass and reporting that it ran; **recording a decline nobody made, or one on an in-set finding; or deleting an acceptance the cycle owes**." The list says it is not complete; - this change creates those routes and names them, as the parent did for the stated floor. -8. **The closing message and the soft-reset path** (C:834–838 / W:1018–1022, with the - `git reset --soft` sentence at C:829–830 / W:1013–1014). OLD: "**The closing message - carries the validated evidence entry for every cited profiled story** — one each, and none - for a cited unprofiled story, which owes no entry. The amend replaces the WIP message - wholesale, so an entry written only into the WIP body is destroyed exactly when the cycle - closes." NEW: "**The closing message carries the validated evidence entry for every cited - profiled story** — one each, and none for a cited unprofiled story, which owes no entry — + this change opens those routes and names them, as the parent did for the stated floor. +8. **The closing message and the soft-reset path** (C:834–838 / W:1018–1022). Two insertions + into a sentence whose other clauses are unchanged. After "…which owes no entry" add: "— **and every answer record the cycle made, including any held only in WIP bodies a `git reset --soft` collapsed**: the single commit after the reset carries all of them, - because a body the reset discards is unreachable from the commit that replaces it. The - amend replaces the WIP message wholesale, so an entry or record written only into the WIP - body is destroyed exactly when the cycle closes." The soft-reset sentence is unchanged. -9. **The one-contract paragraph** (C:879–890 / W:1063–1074). OLD opening: "…this carry rule + because a body the reset discards is unreachable from the commit that replaces it". In the + sentence after it, "so an entry written only into the WIP body" becomes "so an entry **or + record** written only into the WIP body". The `git reset --soft` sentence itself + (C:829–830 / W:1013–1014) is unchanged. +9. **The one-contract paragraph** (C:879–890 / W:1063–1074), which is where the contract gets + its **one name and one membership list**, so that each mergeable piece cites the name + instead of naming every peer. OLD opening: "…this carry rule **and the unknown-start activation semantics that say what a cycle owes when its starting rules cannot be established** depend on one another," NEW opening: "…this carry rule, **the - closure ordering, the answer records and the severity rule's raw-versus-effective split**, - and **the unknown-start activation semantics that say what a cycle owes when its starting - rules cannot be established** depend on one another," and after "and a carry rule naming - records a project does not produce is inert." (C:885) add: "an ordering without the answer - records is a membership stop whose two answers nothing carries; answer records without the - ordering are a release and an obligation with no stop that asks for them; an ordering whose - raw-versus-effective split has no counterpart in the severity rule, or a severity rule still - calling that question unsettled and mandating a stop beside an ordering that decides it, is - two answers to one question; and an answer record missing from the squash carry or the nonce - set cannot survive a merge or be attributed." The stop sentence that follows is unchanged - and now covers these states. Because `/workflow-init` can merge this paragraph and the - pieces independently, **each of the three pieces carries a reciprocal one-contract sentence - naming the other two** — §3's last sentence, the answer-record block's last paragraph, the - (g) replacement's last sentence — so a copy holding one without the others stops on the text - it has, not only on a paragraph it may not have. A rollback while a - cycle is open is governed by "a cycle already running finishes under the rules it started - with" and "A revert is itself a shipping commit for the old rules" (C:153–154, C:165–166) - where the cycle can still establish those rules, and by the unknown-start fallback where it - cannot. **Where the rollback removes the fallback text itself, neither is available** and - that cycle **stops and is a human's to resolve — it does not close**. The residual under all - of it: no record identifies the rule revision a cycle started under, so "finishes under the - rules it started with" is recoverable only while that text is present. Named, not fixed. + closure-record contract — the closure ordering, the answer records, the severity rule's + raw-versus-effective split, the answer records' membership in the named nonce set, the + unknown-start item covering them, the closing-message carry, the squash carry and the + recovery sources** — and **the unknown-start activation semantics that say what a cycle owes + when its starting rules cannot be established** depend on one another," and after "and a + carry rule naming records a project does not produce is inert." (C:885) add: "an ordering + without the answer records is a membership stop whose two answers nothing carries; answer + records without the ordering are a release and a set change with no stop that asks for them; + an ordering whose raw-versus-effective split has no counterpart in the severity rule, or a + severity rule still calling that question unsettled and mandating a stop beside an ordering + that decides it, is two answers to one question; and an answer record missing from the + squash carry, the nonce set or the recovery sources cannot survive a merge, be attributed, + or be found again." The stop sentence that follows is unchanged and now covers these states. + Because `/workflow-init` can merge this paragraph and the pieces independently, **each + separately mergeable piece names the contract by that one name** — §3's last sentence, the + answer-record block's last paragraph, the (g) replacement's last sentence — so a copy holding + one without the rest stops on the text it has. **Rollback.** A cycle open when the text is + reverted is governed by "a cycle already running finishes under the rules it started with" + and "A revert is itself a shipping commit for the old rules" (C:153–154, C:165–166) where it + can still establish those rules, and by the unknown-start fallback where it cannot. **Where + the rollback removes the fallback too, neither is available and nothing here replaces + them**: no record identifies the rule revision a cycle started under, so that cycle has no + text describing what it owes. A residual, named and left to a human — not a stop this change + can claim, since a stop needs shipped text and the rollback removed it. 10. **The no-identity rule's aftermath** (C:422–424 / W:616–618). OLD: "**Starting a new cycle does not close, adopt or retire the cycles those candidates belong to** — they stay open, keep their own nonces, and are a human's to resolve; the new cycle simply does not claim @@ -385,6 +383,14 @@ are in the site map and re-read at execution. report**, so the human this rule makes responsible learns they exist." It adds no record — the report names state the workspace already holds — and closes the one shape "a human's to resolve" cannot reach: an open cycle nobody is told about. +11. **The recovery sources** (C:416–417 / W:610–611), the sentence that would otherwise make an + answer record unreadable by the procedure meant to read it. OLD: "A Gate-A cycle + mid-run has no such commit and therefore has only the working record." NEW: "**A Gate-A cycle + mid-run has no closing commit, but it does have the commits it has made on the branch since + it started** — its artifact revisions, and any empty commit carrying an answer record — and + those bodies are a source for the records they carry, scoped by kind, artifact and nonce + exactly as the working record's candidates are." Left alone, the record would exist and the + procedure would never look at it, and story criterion 7 would fail on one unchanged sentence. The shorter "Copy every record into the squash body" sentence inside the human-exception block (C:1004–1007) is generic and already covers an answer record; it is not edited. The "records @@ -398,14 +404,22 @@ conditional, like the human exception, and is not added. Ids are the committed inventory's (§2). "Kept" = the sentence stays; "moved" = it now lives in the §3 block; "replaced" = the condition changes, and says how; "dropped" carries its reason. -**(a) The floor paragraphs** (C:72–136 / W:279–343) — **trim** `a17`–`a19` and point at the -block. OLD: "Your final pass must be clean — if the pass at the floor still finds -Blocker/Major, keep going until clean or clearly stuck → then STOP and surface to the user. The -only early exit below the floor is a pass with **zero** findings; don't manufacture findings to -pad." NEW: "Your final pass must be clean; how a cycle closes, and what stops it short of -closing, is the closure ordering below — don't manufacture findings to pad." Accounting: -`a1`–`a17`, `a20`–`a22` kept; `a18`, `a19` moved. The pointer form is what `a13`'s prohibition -on restating requires. +**(a) The floor paragraphs** (C:72–136 / W:279–343) — **two edits**. First, **trim** +`a17`–`a19` and point at the block. OLD: "Your final pass must be clean — if the pass at the +floor still finds Blocker/Major, keep going until clean or clearly stuck → then STOP and +surface to the user. The only early exit below the floor is a pass with **zero** findings; +don't manufacture findings to pad." NEW: "Your final pass must be clean; how a cycle closes, +and what stops it short of closing, is the closure ordering below — don't manufacture findings +to pad." Second, `a13`. OLD: "Every other rule stated here about how a +cycle closes stands as written, and none of them is restated — a summary is where their +conditions would get dropped." NEW: "**This paragraph** restates none of them — a summary is +where their conditions would get dropped — and the closure ordering below is where they are +stated once and in order." Accounting: `a1`–`a12`, `a14`–`a17`, `a20`–`a22` kept; `a18`, `a19` +moved; **`a13` replaced**. Old condition: *no* rule about how a cycle closes is restated +anywhere. Kept: the prohibition and its reason, scoped to the paragraph it was written to +police. Changed: the ordering does restate closure rules, deliberately and as the one +authority, so a categorical reading would leave the shipped text contradicting itself — the +failure `a13` exists to prevent, arriving from the other direction. **(b) What a loop absorbs** (C:195–223 / W:402–426) — **three sentence edits**, the triggers stay. `b12` OLD: "…and it resumes the moment the user says whether the set now includes it." @@ -417,13 +431,12 @@ the clean-final-pass rule all stand, and the loop resumes on the revised artifac question is answered — what the stop prevents…" NEW: "Stopping this way is a **suspension** under the closure ordering above, and the loop resumes on the revised artifact **once every answer that pass's suspensions require has been given** — what the stop prevents…". Third, in -**W only**, the cross-reference "…by its severity exactly as the severity rule already says…" -becomes "…exactly as Mechanics already says…", matching C: W does carry a -`### Mechanics (reference)` heading (W:968), so the divergence rested on a false premise (§6). +**W only**, "…by its severity exactly as the severity rule already says…" becomes "…exactly as +Mechanics already says…", matching C for the reason §6's first row gives. Accounting: `b1`–`b11`, `b13`–`b16` kept; `b12` **replaced** — old: an *immediate* resume on -the membership answer alone. Kept from it: that the membership answer is what the stop asks -for and that either direction ends it. Changed: the resume waits for every answer the pass's +the membership answer alone. Kept: that the membership answer is what the stop asks for and +that either direction ends it. Changed: the resume waits for every answer the pass's suspensions require, each recorded, because a loop resumed over an unanswered question decides it by running. `b17` moved (the block's "a suspension waives nothing" sentence); `b18` **replaced** the same way and for the same reason — old: resume once *the* question is @@ -443,16 +456,14 @@ rule is not waived by surfacing, and the loop resumes when every answer that pas suspensions require has been given." Accounting: `c1`–`c8` kept; `c9`, `c10`, `c11` kept verbatim; `c12` kept (pointer form); `c13` moved (the zero-finding exception); `c14` moved **to the ordering's continue branch** — the -Minor-below-floor pass that "keeps looping" is precisely a pass that neither closes nor -suspends, and the branch states it rather than leaving it to be read out of a negation, which -is what finding 1 of pass 3 caught; `c15`, `c16`, `c17` kept in the pointer sentence; `c18` -**replaced, narrowed** — old: a clearly-stuck surface credits no pass as clean; new: no pass -that surfaced a *scope stop* is credited as clean. Deliberate, with its authority: the story -states the fourth duty as "no pass carrying a surfaced *finding* counts as clean" (story §1); -the clearly-stuck exit requires regenerating Blocker/Major findings, so a pass surfacing it is -already unclean by the ordering's first step, and the two-tell stop surfaces no finding, so a -Blocker/Major-free pass below the floor tripping two tells is clean *as a pass* and still -cannot close. No bypass opens in either region; `c19` **replaced** — old: the loop resumes on +Minor-below-floor pass that "keeps looping" is precisely one that neither closes nor suspends, +and the branch states it rather than leaving it to be read out of a negation; `c15`, `c16`, +`c17` kept in the pointer sentence; `c18` +**kept** — a clearly-stuck surface credits no pass as clean, and the ordering carries that +condition over its whole domain and no narrower: no pass that **surfaced a finding** is +credited as clean, which the scope stop and the clearly-stuck exit both trigger. The two-tell +stop was never in its domain, surfacing tells and not a finding, and the ordering says so +rather than leaving it inferred; `c19` **replaced** — old: the loop resumes on whatever the user decides, unconditionally; new: it resumes on a resuming answer and a stop leaves the suspension standing, because an unconditional resume is the stop-with-no-transition path AC 4 forbids; `c20` **dropped** — it argued that reading the exit as "stop instead of @@ -487,10 +498,9 @@ the whole paragraph, both copies, with the answer. NEW: which findings the fix set contains — a separate predicate the closure ordering defines, and one this must not be read as touching; a loop spending passes on findings the author keeps demoting is exactly what the prose-cluster tell exists to surface, and lowering the counts by - that same judgement would hide it. **This split, the closure ordering and the answer records - are one contract**: the ordering's predicates read the field this sentence assigns them, and - its membership stop is answered by those records, so any one of the three present without - the others is a partial adoption that stops under the one-contract rule above. + that same judgement would hide it. **This split is one component of the closure-record + contract** the one-contract rule above names, and a copy carrying it without the rest of + that list is a partial adoption that stops there. ``` Accounting: `g1` **dropped** (the unsettled statement, now settled); `g2`, `g3` **dropped** @@ -506,7 +516,7 @@ answer to. **(i) When these rules bind** (C:153–167 / W:360–374) — **extend** the strict-reading list (§4 item 5); `i1`–`i16` kept. **(j) The squash carry** (C:892 / W:1076) — **extend** (§4 item -1); `j1`–`j4` kept. Also touched, outside the inventoried passages: §4 items 2, 3, 4, 7–10. +1); `j1`–`j4` kept. Also touched, outside the inventoried passages: §4 items 2, 3, 4, 7–11. --- @@ -518,7 +528,7 @@ and W. The pre-existing divergences the inventory found are handled as follows: | Divergence | Kind | This change | |---|---|---| -| (b) cross-reference "Mechanics" (C) vs "the severity rule" (W) | **not deliberate**: the inventory's reason — that W has no section by that name — is false; `### Mechanics (reference)` is at W:968 | **aligned**: W takes C's wording (§5(b)) | +| (b) cross-reference "Mechanics" (C) vs "the severity rule" (W) | **not deliberate**: the inventory's reason — that W has no such section — is false (W:968) | **aligned**: W takes C's wording (§5(b)) | | (b) "exactly how" (C) vs "is how" (W); closing rationale reworded; C-only `infinite-portfolio-canvas` parenthetical | rationale and field citation | **as-is**, stated | | (e) "you report" (C) vs "report" (W); C-only "Recorded rationale" Bricks paragraph | rationale | **as-is**, stated | | (f) evidence framing; C-only hypothesis qualifier; punctuation of the late-Blockers clause | rationale | **as-is**, stated; (f) is not edited | @@ -526,10 +536,10 @@ and W. The pre-existing divergences the inventory found are handled as follows: | (e)/(f) paragraph break: W runs the five-tells paragraph into "The two rules above" (no blank line at W:472/473); C has the Bricks paragraph between | structural | **aligned**: W gains a blank line before "The two rules above"; the Bricks paragraph stays C-only | The parity check at execution: extract each edited passage from both files by its lead phrase -and `diff` them; the only differences permitted are the rows marked as-is above, and the (g) -row is gone. Any other difference is a defect, not a wording choice, and the result is part of -the evidence entry (§9). The Mechanics region the answer-record block joins is byte-identical -between the copies today (C:840–1033 = W:1024–1217) and stays so. +and `diff` them; the only differences permitted are the rows marked as-is, and the result is +part of the evidence entry (§9). Any other difference is a defect, not a wording choice. The +Mechanics region the answer-record block joins is byte-identical between the copies today +(C:840–1033 = W:1024–1217) and stays so. --- @@ -540,20 +550,27 @@ earlier pass had removed.": ``` **Where an earlier pass's findings file is unavailable**, the report says so before it reads -anything. **First the root.** The slots live in `.context/codex-reviews/` under the top-level -directory of the checkout this cycle is running in — the same root every pass call passes as -`workingDirectory`. A slot written under a different root is **indistinguishable from an -absent slot**: nothing records where a past call ran, so no check can tell them apart. Where -the current root cannot be established, the cycle **stops** and says so rather than reporting -benign unavailable history, and the recovery is to re-run the pass from the correct root. +anything. **First the root**, which is a condition on the current pass being valid at all and +not a stop of its own. The slots live in `.context/codex-reviews/` under the top-level +directory of the checkout this cycle is running in — the same root the pass call is given as +`workingDirectory`. Establishing it fails in two observable ways, each with its own fix: the +working directory is **not inside a git repository** (run the pass from the checkout), or the +resolved top level **differs from the directory the call was given** (re-issue the call with +the resolved root). A pass whose root cannot be established is an **INCOMPLETE pass** — the +state this section already defines, already uncounted toward the floor and already excluded +from the curve — so nothing new is ranked in the closure ordering. What no check reaches: a +slot written under a different root **in the past** is **indistinguishable from an absent +slot**, because nothing records where a past call ran. Unobservable, not detected. **Then, per earlier pass, two questions.** Is the slot present and valid? And is the pass -**known accepted** — validated by this session, or recorded as accepted by a cycle record? -Present and valid is the ordinary case. Known accepted but absent, or present and now failing +**known accepted** — meaning **this session validated it**? Present and valid is the ordinary +case. Known accepted but absent, or present and now failing validation, is a real pass whose artifact is unusable: its number **stays counted** and its series read `?`, as the curve grammar already admits, since a corrupted or missing record does -not un-run a pass. **Acceptance unknown**, whether the slot is absent or present-but-invalid, -is **not counted toward the floor** and its series read `?` — crediting an unvalidated pass is -the dangerous direction and this is the other one. A pass **known incomplete when it ran** is +not un-run a pass. **After a lost session nothing is known accepted**, so every absent or +present-but-invalid slot is **acceptance unknown**: **not counted toward the floor**, series +`?`, disclosed. Crediting an unvalidated pass is the dangerous direction and this is the other +one — and no pass-acceptance record is invented to escape the cost, since the cost is one pass +and the record would be a permanent duty. A pass **known incomplete when it ran** is excluded, exactly as today. **Then the report.** It reads the cycle's working record if one exists and says whether it used it; states which of the three lines it computed and from which passes; **names the pass @@ -569,7 +586,7 @@ or omits a line without saying so, is the failure. Filled, over passes 1, 2 and unreadable: Root: `.context/codex-reviews/` under this checkout — confirmed. (Unconfirmable: - STOP, root not established; re-run from the correct root.) + this pass is INCOMPLETE, root not established; re-run from the checkout.) Trend: findings 24, 17, —, 18; Blockers 5, 2, —, 1 — pass 3 not counted (slot absent, acceptance unknown; series `?`); 2→4 spans that gap: rising, weaker evidence. Cluster: product behaviour 12 of 18. require↔withdraw: none visible; @@ -577,12 +594,13 @@ unreadable: ``` This is **D10**. Both historical lines are derivable from the mandated findings files alone -(the `fic2` record verified that), so unavailability is a property of the workspace, not of the -format, and the answer is disclosure rather than a new stop or mandatory artifact. The -partition is by what the agent can observe (`docs/prompt-standards.md` item 10): presence and -validity from the slot, acceptance from a record or this session, and where acceptance cannot -be established the count moves in the direction that costs a pass. The example is there for -item 4. +(the `fic2` record verified that), so unavailability is a property of the workspace and the +answer is disclosure, not a new stop or a mandatory artifact — the root condition reuses the +INCOMPLETE state rather than adding one, so **D10**'s "not a new stop condition" stands as +written. The partition is by what the agent can observe (`docs/prompt-standards.md` item 10): +presence and validity from the slot, acceptance from this session alone, and where acceptance +cannot be established the count moves in the direction that costs a pass. The example is +there for item 4. --- @@ -594,17 +612,17 @@ discriminator in the nonce's slot position to this story. It ships nothing here, **post-rule and unknown-start cycles** the case it served does not arise: every cycle started after the parent's rules bind holds a nonce, and one whose start cannot be established mints one rather than claiming `none (pre-rule)` (`i11`). A production for those would legislate for -an unreachable state, which is AC 1's prohibition. **The claim reaches no further, and one set -of cycles remains.** A cycle that began before the parent's rules shipped has no nonce, cannot -acquire one, and writes `cycle none (pre-rule)`; a rollback can make old-rule cycles reachable -again. That set is bounded and self-terminating — no later cycle can enter it — but while two -such cycles are observably live they compute the same bare slot paths, and **serializing them -is a human's job, not a shipped production**: one runs, then the other. Shipping the rejected -discriminator to serve a closing set would add a permanent rule for a temporary state. The -`rle` naming stays what its closing body recorded it as — a plan-local exception under the old -rules. The durable prior record is that plan's two drop notes (Task 19, line 970; Task 20, -line 1047) and Task 23's second point (line 1298); this section records the dissolution, and -no prompt text changes. +an unreachable state, AC 1's prohibition. **The claim reaches no further.** A cycle that began +before the parent's rules shipped has no nonce, cannot acquire one, and writes +`cycle none (pre-rule)`; a rollback can make old-rule cycles reachable again. That set is +bounded and self-terminating, but while two such cycles are observably live they compute the +same bare slot paths and can delete each other's findings files — the incident the parent +records. **Nothing shipped here prevents that**, and no text in this change reaches a cycle +running under the old rules that never reads this document. The trade, stated rather than +dressed as protection: a permanent production for a closing set costs more than the exposure +it removes, and the exposure is real meanwhile. The `rle` naming stays what its closing body +recorded — a plan-local exception under the old rules. The durable prior record is that plan's +drop notes (Task 19, line 970; Task 20, line 1047) and Task 23's second point (line 1298). --- @@ -623,10 +641,10 @@ not: - the (g) sentence "That question is owned by the loop-rule consolidation work" count **0** in both; against the parent tree, **1** in C and **0** in W; - the (g) replacement's lead phrase `**The demotion changes what a cycle must resolve, never - what it counts.**` count **1** in C and **1** in W; **0** in the parent tree. This pair is - what the removal check above cannot do alone: deleting the old sentence and installing - nothing satisfies the removal count and the parity diff, and reports the central - demotion/loop-health outcome as verified while both copies carry no answer at all; + what it counts.**` count **1** in C and **1** in W; **0** in the parent tree. The removal + check above cannot do this alone: deleting the old sentence and installing nothing satisfies + it and the parity diff, reporting the demotion/loop-health outcome as verified while both + copies carry no answer at all; - lead phrase `**Recording an answer at a membership stop.**` count **1** in each; **0** in the parent tree, and with it both labels: `Accepted: ` and `Declined: ` count **1** in each and **0** there, since one label shipping without the other is the shape the @@ -634,7 +652,9 @@ not: - the six-member squash-carry sentence: `every answer record` count **1** in each; **0** in the parent tree; - the Q6 lead phrase `**Where an earlier pass's findings file is unavailable**` count **1** in - each; **0** in the parent tree. + each; **0** in the parent tree; +- the recovery-source edit (§4 item 11): `has no closing commit, but it does have the commits` + count **1** in each; **0** in the parent tree, where the old sentence stands instead. If the claim "the ordering and the records ship in both copies" were false, one working-tree count would be **0** or the parent-tree counts would not differ from it. The wiring can produce @@ -643,34 +663,28 @@ own input. **The counterfactual is ABSENT, and is claimed as absent** — the pa ordering block and no answer records, and the (g) count is the one site where the parent is present and the change removes it. Nothing is claimed as "contradictory". -**The named verification of the risk path** (story AC 4). Walk every stop the shipped text -names — membership stop, question stop, a finding carrying both triggers, clearly-stuck exit, -two-tell stop, a hold awaiting its answer, accept, decline, a stop answer, two or three -suspensions at once, a below-floor clean pass, a zero-finding pass, the unknown-start -fallback — **and the stateful transitions** — a declined finding re-raised matching on all -five fields, and re-raised with one field changed; the fix set broadened to include a -declined finding; recovery after an accept and after a stop with the session lost (no -identity → new cycle), **both with the answer records present in the branch's bodies and with -them absent**; a `full` Gate-B pass with one branch clean and the other carrying an in-set -Blocker; a rollback with a cycle open under the new rules, and a rollback that removes the -fallback text; a copy adopting the ordering without the answer records, the reverse, and -either without the severity split — and write the **next-state table**: for each row, the -**record state** (which answer records the bodies carry, which working record exists) and the -**governing-scope state** (the fix set as currently assigned) as explicit input columns beside -the user's answer, then the input that ends the row and the state the cycle is in afterwards, -citing the shipped line the row reads, in both copies. The -table lives in the plan and is quoted by the closing commit body, not here. What would be -observed if the claim "no path leaves a cycle unable to close and unable to suspend" were -false: a row whose next state is the same stop with no input consumed, or one that closes with -an in-set Blocker standing. The wiring can produce it because every row is filled from the -shipped text rather than from this spec, and the inputs the `fic2` instrument omitted — the -user's answer and the record state — are columns here. - -**What this is not.** Not the `fic2` decision matrix, whose two defects Gate B found in the -technique itself: a state's inputs must include every input the rule reads, and a -counterfactual must distinguish ABSENT from CONTRADICTORY. The check above is a presence test -whose parent state is absent by inspection; the walk takes the answer and the record state as -inputs and produces a next state rather than a scored output. **No fixture per predicate is +**The named verification of the risk path** (story AC 4) is a **next-state table**, in the +plan and quoted by the closing commit body. **Rows** — every stop the shipped text names: +membership stop, question stop, a finding carrying both triggers, clearly-stuck exit, two-tell +stop, a hold awaiting its answer, accept, decline, a stop answer, two or three suspensions at +once, a below-floor clean pass, a zero-finding pass, the unknown-start fallback; **and the +stateful transitions**: a declined finding re-raised matching on all five fields and re-raised +with one changed, the fix set broadened to include a declined finding, recovery after an accept +and after a stop with the session lost (no identity → new cycle) **both with the answer records +present in the branch's bodies and with them absent**, a `full` Gate-B pass with one branch +clean and the other carrying an in-set Blocker, a rollback with a cycle open under the new +rules, and a copy adopting the ordering without the answer records, the reverse, or either +without the severity split. **Columns** — the **record state** (which answer records the bodies +carry, which working record exists) and the **governing-scope state** (the fix set as currently +assigned) beside the user's answer, then the input that ends the row, the state afterwards, and +the shipped line the row reads, in both copies. What would be observed if "no path leaves a +cycle unable to close and unable to suspend" were false: a row whose next state is the same +stop with no input consumed, or one that closes with an in-set Blocker standing. The wiring can +produce it because every row is filled from the shipped text rather than from this spec, and +the inputs the `fic2` instrument omitted — the user's answer and the record state — are columns +here. That is what makes this **not the `fic2` decision matrix**, whose two defects Gate B found +in the technique itself: a state's inputs must include every input the rule reads, and a +counterfactual must distinguish ABSENT from CONTRADICTORY. **No fixture per predicate is built** — parked in the story's §2, not reopened. **Evidence entry**, in the closing commit body, names: the battery run; every assert pair @@ -683,16 +697,15 @@ Gate-B re-review and before the closing amend, as §5 requires. ## 10. AGENTS.md invariants touched -- **Invariant 11 — `docs/prompt-standards.md`, all 12 items.** Most at risk: item 6 (rules - carry their why — every constraint in the shipped blocks carries its reason in the same - sentence **except the ordering's three settled axioms**: only clean completion closes, a - zero-finding pass is clean whatever the floor, and either answer is available only at a - membership stop. Those are asserted deliberately — their reasons are **D1**, the existing - zero-finding exit and **D6**, settled in the story and not re-argued in a prompt — and the - plan's review reads every *other* sentence for an inline why); item 8 (token-lean — the - blocks replace closure sentences rather than adding beside them, `a13` being why the old - ones leave); item 4 (output structure shown — the Q6 example); item 10 (the Q6 partition and - its root check); item 3 (the stop answer is a named, resumable state). +- **Invariant 11 — `docs/prompt-standards.md`, all 12 items.** Most at risk: item 6 (every + constraint in the shipped blocks carries its reason in the same sentence **except the + ordering's three settled axioms** — only clean completion closes, a zero-finding pass is + clean whatever the floor, either answer is available only at a membership stop. Those are + asserted deliberately, their reasons being **D1**, the existing zero-finding exit and + **D6**, settled in the story and not re-argued in a prompt; the plan's review reads every + *other* sentence for an inline why); item 8 (token-lean — the blocks replace closure + sentences rather than adding beside them); item 4 (the Q6 example); item 10 (the Q6 + partition and its root check); item 3 (the stop answer is a named, resumable state). - **Don't: "Never replace a decision procedure without accounting for its old conditions."** Satisfied by §5, against the committed inventory. - **Don't: "Never rename or delete a doc section without grepping for references first."** The @@ -701,9 +714,9 @@ Gate-B re-review and before the closing amend, as §5 requires. `…review-loop-economics-design.md:33`, `…pass-floor-story.md:78`). All cite the story file, which continues to exist; none cites the sentence. Nothing breaks. - **Invariant 12 — a plugin change requires a version bump.** `workflow-init.md` is under - `plugins/`, so `plugins/dev-workflow/.claude-plugin/plugin.json` goes `0.11.0 → 0.12.0` with a - `CHANGELOG.md` entry: a minor bump, because the template gains a record form and a rule. - The entry also notes that the squash-carry sentence now lists six record kinds. + `plugins/`, so `plugins/dev-workflow/.claude-plugin/plugin.json` goes `0.11.0 → 0.12.0` with + a `CHANGELOG.md` entry: a minor bump, the template gaining a record form and a rule. The + entry also notes that the squash-carry sentence now lists six record kinds. - **Invariant 4 / the hook.** Untouched: `codex-gate.sh` is not edited, and the §5 heading it greps (`Cross-Model Review`) does not move. From de66e00897717319ce4232893b586ea50509984a Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Thu, 10 Sep 2026 17:07:11 +0200 Subject: [PATCH 009/181] =?UTF-8?q?docs(specs):=20loop-rule=20consolidatio?= =?UTF-8?q?n=20=E2=80=94=20Gate-A=20pass=206=20revision?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Sixteen findings, all fixed. The centre is a structural repair: clean candidacy now reads the scope-stop TRIGGERS — a finding outside the assigned fix set, one opening a new structural or contract question — rather than the act of surfacing, which a lower branch performs. That makes the fixed evaluation order executable instead of asserted, and it dissolves four findings at once (the circular predicate, the collision with the NO FINDINGS vocabulary, and the contradiction with the D3 sentence kept verbatim). Four standing sentences are edited at their source rather than worked around: the two that use "clean" in the findings-file sense, the resolve duty that never stated its scope, and the recovery passage that would not have looked where the answer records live. The three axioms that had been exempted from prompt-standards item 6 now carry inline reasons, and the exemption paragraph is deleted; item 6 admits no "settled elsewhere" clause, and a scaffolded copy cannot reach the story the reasons were said to live in. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- ...26-09-10-loop-rule-consolidation-design.md | 276 +++++++++++++----- 1 file changed, 195 insertions(+), 81 deletions(-) diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index 346bf8c..20b9e4e 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -1,6 +1,6 @@ # §5 loop-rule consolidation: one closure ordering — Design -**Date:** 2026-09-10 · **Status:** DRAFT — Gate A running (pass 5 revised) +**Date:** 2026-09-10 · **Status:** DRAFT — Gate A running (pass 6 revised) **Story:** `docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md` **Profile:** read from that header at every pass, never from here — it is the only writable copy, and a value copied here would be a remembered value. @@ -43,6 +43,9 @@ does not settle, this spec decides in the section that uses it: the evaluation o field and file set each predicate reads; the duties' classification; the scope stop's two triggers and what each answer does; what a stuck or two-tell answer produces; the records' wording, attribution and recording point and the `Accepted:` label; and the raw-severity rule. +Four standing sentences are edited at their source rather than worked around — the two that +use *clean* in the file's sense, the resolve duty that never stated its scope, and the recovery +passage that would not have looked where the records live (§4 items 11–14). --- @@ -56,10 +59,13 @@ triggers and point at it. Verbatim as it will ship: **How a cycle ends — one ordering, stated here and referenced everywhere else.** A pass is read in a fixed order, because every rule below bears on one decision — may this cycle close — and a stated order is what stops them qualifying each other. Every predicate here reads the -validated findings file **or files** of the logical pass as one set — a `full` Gate-B pass -has two, and one branch alone is already an incomplete pass — at **effective** severity, after -the Mechanics severity ceiling, because cleanliness is about what the cycle must repair and -the ceiling is what decides that. Two things read the reviewer-written field from before the +validated findings file **or files** of the logical pass as their **concatenation** — a `full` +Gate-B pass has two, and one branch alone is already an incomplete pass — at **effective** +severity, after the Mechanics severity ceiling, because cleanliness is about what the cycle +must repair and the ceiling is what decides that. **Both branches' lines count** for the health +measures, the curve rule already summing the branches into one entry; a finding whose five +fields match one in the other branch is **one** finding for holds and answers, so it is +answered once. Two things read the reviewer-written field from before the ceiling: the **health measures** (Mechanics, Severity) — the per-pass counts, the clusters, the tells, and **both** conditions of the stuck reading, its Blocker curve and its regenerating Blocker or Major findings, since a predicate reading one field for half of itself @@ -67,16 +73,20 @@ could not be read at all — and the **five-field key** the answer records match severity field is the one the reviewer wrote, so that a ceiling cannot make two findings the same one. -**First, clean completion.** §5 uses *clean* in two senses and always has. A **clean findings -file** is the `NO FINDINGS` signal the protocol defines. A **clean pass** is the predicate -below: a clean findings file always gives one, and a clean pass need not have one, because a -pass carrying only Minors, Nits or findings outside the fix set is clean without being empty. -**A pass is clean** when it carries no Blocker or Major at effective -severity that is **in the assigned fix set**, and it **surfaced no finding** — the scope stop -and the clearly-stuck exit both surface findings, while the two-tell stop surfaces tells and -not a finding, so it leaves cleanliness alone. A pass with -**zero** findings is clean whatever the floor. A finding matched by a decline recorded in -this cycle is outside the set by that decision (Mechanics, the answer records). **A clean pass at or above the derived floor, +**First, clean completion.** §5 uses *clean* in two senses and now says which is which. A +**clean findings file** is the `NO FINDINGS` signal the protocol defines. A **clean pass** is +the predicate below: a clean findings file always gives one, and a clean pass need not have +one, because a pass carrying only Minors and Nits, or findings a decline of this cycle +matches, is clean without being empty. **A pass is clean** when its findings carry no Blocker +or Major at effective severity that is **in the assigned fix set**, and **no scope-stop +trigger** — no finding outside that set, and none opening a new structural or contract +question. Both halves are properties of the findings and the set, read before any branch below +runs, which is what makes this order executable rather than asserted: no predicate here waits +on an act a lower branch performs. A finding matched by a decline recorded in this cycle is +outside the set by that decision and triggers nothing (Mechanics, the answer records). A pass +with **zero** findings is clean whatever the floor, because a floor buys further looks at an +artifact that keeps yielding findings, and one yielding none has already given what those +looks were for. **A clean pass at or above the derived floor, or a zero-finding pass, closes the cycle**, under the final-acceptance preconditions the floor section already states and this ordering does not restate — the cited set and every profile re-read before the pass is accepted as final, a header or profile that changed during it @@ -84,7 +94,8 @@ making the pass not final and costing the further pass that section requires. A tells present on the closing pass go into the closing report and never block it, because reporting "will not converge" on a converged loop is a false report, and the clearly-stuck paragraph says the same of its own exit. **Nothing -else closes a cycle.** +else closes a cycle**, because every other way out of a pass leaves a finding or a question +open, and closing over one is the failure this ordering exists to prevent. **Second, only a pass that is not a clean completion can suspend** — that order is what makes "clean completion outranks the two-tell stop" executable rather than asserted. Three @@ -120,9 +131,14 @@ that the finding is not true of the artifact, a decline is the user's decision t finding stays outside the fix set, and only the second is an answer at a membership stop. The **hold** a surfaced finding places on closure is part of the ordering: it gates closing while it stands, and is discharged by the answers that finding requires, below. -**No-clean-credit** — no pass that **surfaced a finding** is credited as clean — is part of -the ordering and a fact about that pass, discharged by nothing: a later pass is judged on its -own findings, so an answer never closes the cycle on the pass that surfaced the finding. +**No-clean-credit** — no pass carrying a finding that gets surfaced is credited as clean — is +part of the ordering and a fact about that pass, discharged by nothing: a later pass is judged +on its own findings, so an answer never closes the cycle on the pass that surfaced the +finding. **The clean predicate carries it whole**, without reading the act: a scope-stop +trigger makes a pass unclean directly, and the clearly-stuck exit needs regenerating Blocker +or Major findings, which are either in the set — failing the predicate's first half — or +outside it, failing its second, so no pass either exit can surface on is clean. The two-tell +stop surfaces tells and not a finding, and was never in this duty's domain. **What a suspension asks, and what ends it.** Only a scope stop raises a **hold**, on its finding, and the hold ends when every answer that finding requires has been given — one for a @@ -132,17 +148,25 @@ duty under another name. At a membership stop the answer is **accept** (the find fix set and resolves by its effective severity: Blocker or Major before any pass can be clean, Minor or Nit collected and never iterated) or **decline** (the finding stays outside, binding for the rest of this cycle). Either answer is **recorded** — Mechanics, the answer records — -and that record is where a later pass, or a cycle that lost its session, reads it. At a question +and that record is where a later pass, or a cycle that lost its session, reads it. +**Membership is read when the answer is given, not frozen at the surface**: where the governing +artifacts have broadened the set so that the held finding is now inside it, that broadening +discharges the hold by itself — the finding is in-set, decline is unavailable to it, and it +resolves by its effective severity like any other in-set finding. At a question stop the answer is the user's decision on the question, and membership does not change: an -in-set finding then resolves under that decision or is dismissed with its one-line why; an +in-set finding then routes through its effective severity like any other — a Blocker or Major +resolves under that decision or is dismissed with its one-line why, a Minor or Nit is collected +and never iterated; an out-of-set finding that opened the question is a membership stop as well and takes accept or -decline. **Decline is available only at a membership stop.** The stuck and two-tell readings +decline. **Decline is available only at a membership stop**, because that is the only stop +whose question is whether a finding belongs to the set, and a decline anywhere else would +waive work the cycle owes. The stuck and two-tell readings raise no hold: each asks one question, **continue or stop**. Continue resumes the loop on the artifact as revised and the fix set as the governing artifacts now assign it — where several plans or stories govern one cycle, the union of the scopes they assign — **plus the findings -this cycle accepted into it**, which its commit bodies carry (Mechanics, the answer records). -One no body carries is not in the set, and the reviewer raises it again on the next pass like -any other, which is the ordinary route and not a special one. A finding that a +this cycle accepted into it**, which its **latest** commit body carries (Mechanics, the answer +records). One that body does not carry is not in the set, and the reviewer raises it again on +the next pass like any other, which is the ordinary route and not a special one. A finding that a narrowing has put outside the set surfaces at the next pass as a membership stop. **Stop leaves the suspension standing**: the cycle stays open under its nonce, resumable by a later continue, nothing it wrote is a closing commit, and an open cycle nobody resumes is a human's @@ -152,9 +176,12 @@ to resolve, exactly as the nonce rules already say of open cycles. resumes only when every answer resumes it — accept or decline at a membership stop, a decision at a question stop, continue at the stuck or two-tell reading; one stop answer leaves the whole suspension standing, because a loop resumed over an unanswered question decides it by -running. Two states cannot co-occur, and no rule ranks them: clean completion and any -suspension that surfaces a finding — the scope stop, the clearly-stuck exit — since a pass -that surfaced one is not clean; and a zero-finding pass and any +running. **Simultaneous health suspensions are one question, not two**: the stuck and two-tell +readings both ask continue or stop, so one answer carrying every reason ends both, and asking +twice would invite two answers to a question that has one. Two states cannot co-occur, and no +rule ranks them: clean completion and any suspension that surfaces a finding — the scope stop, +the clearly-stuck exit — since a pass carrying either's findings fails the clean predicate +above; and a zero-finding pass and any suspension, since it has nothing to surface, nothing regenerating, no cluster and no require↔withdraw pair. The Gate-B triviality skip is outside this ordering: a skipped cycle runs no passes and ends by its own rule. **This ordering is one component of the @@ -165,7 +192,9 @@ rest of that list is a partial adoption that stops there. Why the shape, where the block's sentences do not carry it. The evaluation order answers the objection that a ranking between clean completion and the two-tell stop cannot fire if the stop can make the pass unclean: clean completion is read first, so a suspension is only ever -evaluated on a pass that did not close. The three branches are AC 4, the duties paragraph AC 2, +evaluated on a pass that did not close — which works only because clean candidacy reads the +**triggers** and never the act of surfacing, as the block's own sentence says. +The three branches are AC 4, the duties paragraph AC 2, the composition and cannot-co-occur sentences AC 1. The scope stop's two triggers are `b11` (membership) and `b13` (question) read separately, because an in-set finding that opens a question can neither join nor stay outside the set and needs its own answer — which is also why @@ -215,16 +244,20 @@ byte-identical in both copies: chose. ``` - Accepted: · · cycle + Accepted: · · cycle · · Finding: | | | | - Declined: · · cycle + Declined: · · cycle · · Finding: | | | | ``` The `Finding:` line is the finding line from the pass's findings file with its confidence field removed — the five fields the sameness test reads, in the file's order, a literal pipe - escaped as `\|` exactly as there. + escaped as `\|` exactly as there. `` is one of the three cycle kinds and `` + the same key the working record uses — a Gate-A cycle's reviewed document path, a Gate-B + cycle's base commit as a full 40-character object name — because a record read back out of a + branch body must be attributable on its own, and an empty commit carrying only a record + supplies nothing else to attribute it by. **Rules both labels share.** Each carries the **cycle nonce**, because a record that cannot be attributed to its cycle cannot bind to it. Each is **written when made**, into the cycle's @@ -235,8 +268,11 @@ byte-identical in both copies: it — the answered finding was that pass's only one — the record goes in an **empty commit of its own**, the destination the human-exception rule above already blesses: "An empty commit carrying only the record is a legitimate destination". Each is **restated in every later body - of that cycle**, the closing one included, because a body that drops one loses the fact it - records, and each is **copied on squash-merge** (the carry rule above). A cycle that resumes + of that cycle**, the closing one included, so that the cycle's **latest** body carries the + complete answer set: **that body is the authoritative one**, and an answer missing from it is + lost whatever an earlier body says, because a rule that let any reachable body revive an + answer would make the set depend on how far back a reader looked. Each is **copied on + squash-merge** (the carry rule above). A cycle that resumes with no commit body carrying an answer **treats it as absent** — the unknown-start fallback's reading, and the safe direction under both labels: an unrecorded decline means the hold applies again, an unrecorded acceptance means the finding is raised afresh. Each is an @@ -271,9 +307,10 @@ byte-identical in both copies: makes that discoverable rather than invisible, which is less than preventing it. **What no record carries:** a question stop's decision, where it changed no membership — the five fields identify a finding and not a question, and that decision's durable form is the - artifact revision it produces — and a stuck or two-tell surface with its continue-or-stop - answer. From a closing body a reader can infer the close, the acceptances and the declines, - and after a lost session neither of those two is recoverable from it. + artifact revision it produces, **where it produces one**; a decision that changes nothing + stays in the session that made it — and a stuck or two-tell surface with its + continue-or-stop answer. From a closing body a reader can infer the close, the acceptances + and the declines, and after a lost session neither of those two is recoverable from it. **These records are one component of the closure-record contract** the one-contract rule above names, and a copy carrying them without the rest of that list is a partial adoption @@ -283,9 +320,11 @@ byte-identical in both copies: **Which rules the answer modifies** (story AC 3): the scope stop's resume sentence (`b12`), and the hold and the clean-pass definition, both in the §3 block. Those three are the rules whose **closure behaviour** the answer qualifies, and it qualifies no other — not the -Blocker/Major-resolve duty (Mechanics, Severity: "both must resolve", C:784 — **D5**: a -declined finding never enters its scope, an accepted one enters it without changing what the -duty demands), not the floor, the tells or the stuck reading. A reader finding either label +Blocker/Major-resolve duty, not the floor, the tells or the stuck reading. The Severity +bullet's edit (item 14) is not a counter-example and is the distinction worth holding: it +states **the duty's own scope**, the assigned fix set, which **D5** always implied. The answer +moves a finding into or out of that set and never changes what the duty demands of what is in +it, so no rule there mentions a decline. A reader finding either label qualifying a *closure* rule outside those three has found a defect; a reader finding the records absent from any transport, attribution, activation or threat site listed next has found the opposite one. The two lists answer different questions: what the answer *changes*, @@ -297,9 +336,9 @@ are in the site map and re-read at execution. 1. **Squash carry** (C:892 / W:1076, `j1`). OLD: "…copy every evidence entry, every human-exception record, the provenance lines, the curves and any skipped cycle's skip record TOGETHER WITH THE SKIP REASON IT POINTS AT…". NEW: "…copy every evidence entry, - every human-exception record, **every answer record, one copy per cycle nonce and - five-field finding key — byte-identical repeats collapse, and copies that disagree stop - under the rule below** —, the provenance lines, the curves and any skipped cycle's skip + every human-exception record, **every answer record (one copy per cycle nonce and + five-field finding key: byte-identical repeats collapse, and copies that disagree stop + under the rule below)**, the provenance lines, the curves and any skipped cycle's skip record TOGETHER WITH THE SKIP REASON IT POINTS AT…". Six members; the dedup clause is there because a record restated in every body reaches the squash range many times. 2. **The named nonce set** (C:385–387 / W:579–581). OLD: "**and that set is named rather than @@ -351,9 +390,10 @@ are in the site map and re-read at execution. **and the unknown-start activation semantics that say what a cycle owes when its starting rules cannot be established** depend on one another," NEW opening: "…this carry rule, **the closure-record contract — the closure ordering, the answer records, the severity rule's - raw-versus-effective split, the answer records' membership in the named nonce set, the - unknown-start item covering them, the closing-message carry, the squash carry and the - recovery sources** — and **the unknown-start activation semantics that say what a cycle owes + raw-versus-effective split and its assigned-fix-set boundary, the two clean-vocabulary + edits, the answer records' membership in the named nonce set, the unknown-start item + covering them, the closing-message carry, the squash carry, the recovery sources and the + no-identity report** — and **the unknown-start activation semantics that say what a cycle owes when its starting rules cannot be established** depend on one another," and after "and a carry rule naming records a project does not produce is inert." (C:885) add: "an ordering without the answer records is a membership stop whose two answers nothing carries; answer @@ -363,10 +403,14 @@ are in the site map and re-read at execution. that decides it, is two answers to one question; and an answer record missing from the squash carry, the nonce set or the recovery sources cannot survive a merge, be attributed, or be found again." The stop sentence that follows is unchanged and now covers these states. - Because `/workflow-init` can merge this paragraph and the pieces independently, **each - separately mergeable piece names the contract by that one name** — §3's last sentence, the - answer-record block's last paragraph, the (g) replacement's last sentence — so a copy holding - one without the rest stops on the text it has. **Rollback.** A cycle open when the text is + Then, before it: "**A project carrying any component of this contract owes all of them.** + The pieces are separately mergeable and are not separately adoptable, so a copy holding one + without the rest is an incomplete adoption and stops here." That is deliberately not the + finding's other option, a reciprocal marker on every hunk: one membership list is one thing + to keep in step, and eleven cross-references are eleven. Three pieces still name the + contract where a reader meets them first — §3's last sentence, the answer-record block's + last paragraph, the (g) replacement's last sentence — so a copy holding only those stops on + the text it has, but the obligation is the list's, not the marker's. **Rollback.** A cycle open when the text is reverted is governed by "a cycle already running finishes under the rules it started with" and "A revert is itself a shipping commit for the old rules" (C:153–154, C:165–166) where it can still establish those rules, and by the unknown-start fallback where it cannot. **Where @@ -383,14 +427,61 @@ are in the site map and re-read at execution. report**, so the human this rule makes responsible learns they exist." It adds no record — the report names state the workspace already holds — and closes the one shape "a human's to resolve" cannot reach: an open cycle nobody is told about. -11. **The recovery sources** (C:416–417 / W:610–611), the sentence that would otherwise make an - answer record unreadable by the procedure meant to read it. OLD: "A Gate-A cycle - mid-run has no such commit and therefore has only the working record." NEW: "**A Gate-A cycle - mid-run has no closing commit, but it does have the commits it has made on the branch since - it started** — its artifact revisions, and any empty commit carrying an answer record — and - those bodies are a source for the records they carry, scoped by kind, artifact and nonce - exactly as the working record's candidates are." Left alone, the record would exist and the - procedure would never look at it, and story criterion 7 would fail on one unchanged sentence. +11. **The recovery sources** (C:411–417 / W:605–611) — the **whole passage** is replaced, not a + sentence inside it, because the sentence that must change ("no search there") is the same one + that makes the rest coherent. OLD: "**Recovery has two sources, and they answer different + questions.** The **working record** is the source while the cycle runs, and it is the one the + candidate rules above apply to — several files may be present and the run must decide which, + if any, is its own. **History is the source once the cycle's own commit exists**, and there is + no search there: the cycle is reading **its own commit body**, so kind and artifact are settled + by which commit is being read, and the nonce is taken from the provenance line and the curve, + which must agree. A Gate-A cycle mid-run has no such commit and therefore has only the working + record." NEW: "**Recovery has three sources, and they answer different questions.** The + **working record** is the source while the cycle runs, and it is the one the candidate rules + above apply to — several files may be present and the run must decide which, if any, is its + own. **The cycle's own commits on the current branch** are the source for the records those + bodies carry — its artifact revisions, and any empty commit carrying an answer record — + **searched newest first and no further back than the branch point**, each candidate validated + by cycle field, kind and artifact exactly as a working record is; a bounded search, because an + unbounded one would reach other cycles' commits. **History is the source once the cycle's own + closing commit exists**, and there is no search there: the cycle is reading **its own commit + body**, so kind and artifact are settled by which commit is being read, and the nonce is taken + from the provenance line and the curve, which must agree." The next sentence's "Recovering a + single candidate from **either**" becomes "from **any of the three**" — a one-word repair the + passage's arithmetic forces, and the failure rule after it ("No candidate, disagreeing sources, + or more than one candidate → no identity") is unchanged and already governs all three, which is + why the NEW does not restate it. Left alone, an answer record would exist and the + procedure meant to read it would never look, and story criterion 7 would fail on one unchanged + sentence. + +12. **The findings-file protocol's clean sentence** (C:329 / W:523), inside the gate-prompt + template both gates paste. OLD: "A clean pass is the single body line `NO FINDINGS` with + `END OF FINDINGS (0 total)`." NEW: "A **clean findings file** is the single body line + `NO FINDINGS` with `END OF FINDINGS (0 total)`." It describes a file and always did; the + word *pass* in it is what makes §3's predicate look like a redefinition instead of the + other sense. Its condition — what a reviewer writes when it finds nothing — is unchanged. +13. **The Gate-A clean-signal sentence** (C:565–566 / W:756–757; the two copies wrap it + differently, so the shared fragment is what is quoted). OLD: "…when a pass is clean — the + explicit clean signal is what lets you exit the loop:". NEW: "…when a pass finds nothing — + the explicit signal is what lets a pass be read as clean without inspecting it further:". + Old condition: a `NO FINDINGS` file is what permits loop exit. Kept: the signal and why it + is demanded. Changed: it makes a pass readable as clean rather than being the only way to + be clean, since a pass carrying Minors alone is clean under the ordering and could never + produce this file. The other five uses of "clean pass" in each copy (C:117, 728, 750, 761, + 769, 827) are the closure sense the ordering defines and are correct as they stand — + checked, not assumed. +14. **The Severity bullet's resolve duty** (C:783–784 / W:969–970), which is the one place the + duty is stated and the only one without a scope. OLD: "- **Severity:** Blocker + (wrong/unsafe/breaks invariant) · Major (design flaw → rework) → both must resolve. Minor · + Nit → collect, never iterate." NEW: "- **Severity:** Blocker (wrong/unsafe/breaks + invariant) · Major (design flaw → rework) → both must resolve, **for every finding in the + assigned fix set**; one the user declined at a membership stop is outside that set and owes + nothing (the closure ordering above). Minor · Nit → collect, never iterate." Old condition: + every Blocker and Major resolves, unbounded. **Replaced**: the boundary **D5** always + implied and no sentence carried. Without it a declined finding must stay outside the set + and still bars closure, which is a pass that can neither close nor suspend. The (c) + pointer that calls this "the resolve rule" (`c17`) needs no edit: it names the rule, and + the rule now carries its own scope. The shorter "Copy every record into the squash body" sentence inside the human-exception block (C:1004–1007) is generic and already covers an answer record; it is not edited. The "records @@ -450,18 +541,27 @@ takes precedence over this exit**: a Blocker/Major-free pass **at or above the f satisfied the clean-final-pass rule — collect the Minors and Nits and close — and reporting "will not converge" on a converged loop is a false report." OLD, from "**Below the floor nothing closes**" to the end of "…and then nothing could satisfy both.": replaced by -"**Below the floor nothing closes**, exactly as the closure ordering above says. This exit is -a **suspension** under that ordering: you surface with the findings still open, the resolve +"That example is a clean-completion candidate only where the pass carries **no scope-stop +trigger** — a finding outside the assigned fix set, or one opening a new structural or +contract question — since a pass carrying either is not clean and the ordering above says +why. **Below the floor nothing closes**, exactly as that ordering says. This exit is +a **suspension** under it: you surface with the findings still open, the resolve rule is not waived by surfacing, and the loop resumes when every answer that pass's -suspensions require has been given." Accounting: `c1`–`c8` -kept; `c9`, `c10`, `c11` kept verbatim; `c12` kept (pointer form); `c13` moved (the +suspensions require has been given." The kept sentence is not touched; the qualification is +adjacent to it, which is how **D3**'s verbatim requirement and the ordering can both hold — +edited, the sentence would say a Blocker/Major-free pass carrying a new out-of-set Minor +closes, and the ordering says it suspends. Accounting: `c1`–`c8` +kept; `c9`, `c10`, `c11` **kept verbatim and qualified** — the sentence is unchanged and a new +sentence beside it names the condition its example assumed and never stated, so the conditions +`c10` carries are narrowed by adjacency rather than by edit; `c12` kept (pointer form); `c13` moved (the zero-finding exception); `c14` moved **to the ordering's continue branch** — the Minor-below-floor pass that "keeps looping" is precisely one that neither closes nor suspends, and the branch states it rather than leaving it to be read out of a negation; `c15`, `c16`, -`c17` kept in the pointer sentence; `c18` +`c17` kept in the pointer sentence, `c17`'s "resolve rule" now reading with the +assigned-fix-set boundary §4 item 14 gives it, so the pointer needs no edit of its own; `c18` **kept** — a clearly-stuck surface credits no pass as clean, and the ordering carries that -condition over its whole domain and no narrower: no pass that **surfaced a finding** is -credited as clean, which the scope stop and the clearly-stuck exit both trigger. The two-tell +condition over its whole domain and no narrower, reaching it through the pass's **findings** +rather than through the act, for the reason the duties paragraph states. The two-tell stop was never in its domain, surfacing tells and not a finding, and the ordering says so rather than leaving it inferred; `c19` **replaced** — old: the loop resumes on whatever the user decides, unconditionally; new: it resumes on a resuming answer and a stop @@ -516,7 +616,7 @@ answer to. **(i) When these rules bind** (C:153–167 / W:360–374) — **extend** the strict-reading list (§4 item 5); `i1`–`i16` kept. **(j) The squash carry** (C:892 / W:1076) — **extend** (§4 item -1); `j1`–`j4` kept. Also touched, outside the inventoried passages: §4 items 2, 3, 4, 7–11. +1); `j1`–`j4` kept. Also touched, outside the inventoried passages: §4 items 2, 3, 4, 7–14. --- @@ -565,12 +665,16 @@ slot**, because nothing records where a past call ran. Unobservable, not detecte **known accepted** — meaning **this session validated it**? Present and valid is the ordinary case. Known accepted but absent, or present and now failing validation, is a real pass whose artifact is unusable: its number **stays counted** and its -series read `?`, as the curve grammar already admits, since a corrupted or missing record does -not un-run a pass. **After a lost session nothing is known accepted**, so every absent or -present-but-invalid slot is **acceptance unknown**: **not counted toward the floor**, series -`?`, disclosed. Crediting an unvalidated pass is the dangerous direction and this is the other -one — and no pass-acceptance record is invented to escape the cost, since the cost is one pass -and the record would be a permanent duty. A pass **known incomplete when it ran** is +series read `?`, which the curve grammar already admits for a **valid** pass whose counts +cannot be recovered, since a corrupted or missing record does not un-run a pass. **After a +lost session nothing is known accepted**, so every absent or present-but-invalid slot is +**acceptance unknown**: **not counted toward the floor** and **omitted from the durable +curve**, which takes one entry per *valid* pass and has no way to say "may not have been +one" — `?` is a missing count, not a missing pass. The omission is named in the report beside +the reduced-sensitivity line, so it is disclosed where it is decided rather than inferred from +a gap in the curve. Crediting an unvalidated pass is the dangerous direction and this is the +other one — and no pass-acceptance record is invented to escape the cost, since the cost is +one pass and the record would be a permanent duty. A pass **known incomplete when it ran** is excluded, exactly as today. **Then the report.** It reads the cycle's working record if one exists and says whether it used it; states which of the three lines it computed and from which passes; **names the pass @@ -653,8 +757,15 @@ not: the parent tree; - the Q6 lead phrase `**Where an earlier pass's findings file is unavailable**` count **1** in each; **0** in the parent tree; -- the recovery-source edit (§4 item 11): `has no closing commit, but it does have the commits` - count **1** in each; **0** in the parent tree, where the old sentence stands instead. +- the recovery-source edit (§4 item 11): `Recovery has three sources` count **1** in each and + **0** in the parent tree, with `Recovery has two sources` at **0** in each and **1** there, + since the passage is replaced and a NEW installed beside a surviving OLD is the failure; +- the four source edits, each as a pair, because each is a standing sentence changing meaning + rather than new text appearing: `A **clean findings file** is the single body line` at **1** + in each and **0** in the parent, with `A\n> clean pass is the single body line` at **0** and + **1**; `when a pass finds nothing` at **1** and **0**, with `when a pass is clean — the + explicit clean signal` at **0** and **1**; `for every finding in the assigned fix set` at + **1** and **0**. A one-sided presence check would pass on a copy carrying both wordings. If the claim "the ordering and the records ship in both copies" were false, one working-tree count would be **0** or the parent-tree counts would not differ from it. The wiring can produce @@ -669,7 +780,8 @@ membership stop, question stop, a finding carrying both triggers, clearly-stuck stop, a hold awaiting its answer, accept, decline, a stop answer, two or three suspensions at once, a below-floor clean pass, a zero-finding pass, the unknown-start fallback; **and the stateful transitions**: a declined finding re-raised matching on all five fields and re-raised -with one changed, the fix set broadened to include a declined finding, recovery after an accept +with one changed, the fix set broadened to include a declined finding, **the fix set broadened +to include a finding whose hold is still awaiting its answer**, recovery after an accept and after a stop with the session lost (no identity → new cycle) **both with the answer records present in the branch's bodies and with them absent**, a `full` Gate-B pass with one branch clean and the other carrying an in-set Blocker, a rollback with a cycle open under the new @@ -698,12 +810,14 @@ Gate-B re-review and before the closing amend, as §5 requires. ## 10. AGENTS.md invariants touched - **Invariant 11 — `docs/prompt-standards.md`, all 12 items.** Most at risk: item 6 (every - constraint in the shipped blocks carries its reason in the same sentence **except the - ordering's three settled axioms** — only clean completion closes, a zero-finding pass is - clean whatever the floor, either answer is available only at a membership stop. Those are - asserted deliberately, their reasons being **D1**, the existing zero-finding exit and - **D6**, settled in the story and not re-argued in a prompt; the plan's review reads every - *other* sentence for an inline why); item 8 (token-lean — the blocks replace closure + constraint in the shipped blocks carries its reason in the same sentence — **including the + three that were exempted until pass 6**: nothing else closes a cycle *because every other way + out leaves a finding or a question open*; a zero-finding pass is clean whatever the floor + *because a floor buys further looks at an artifact that keeps yielding findings*; decline is + available only at a membership stop *because that is the only stop whose question is + whether a finding belongs to the set*. The exemption was wrong twice over: item 6 admits no + "settled elsewhere" clause, and a scaffolded copy cannot reach the story the reasons were + said to live in); item 8 (token-lean — the blocks replace closure sentences rather than adding beside them); item 4 (the Q6 example); item 10 (the Q6 partition and its root check); item 3 (the stop answer is a named, resumable state). - **Don't: "Never replace a decision procedure without accounting for its old conditions."** From 5aef816a3969d51cd007f1a8a40c815ba6653119 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Thu, 10 Sep 2026 17:40:45 +0200 Subject: [PATCH 010/181] =?UTF-8?q?docs(specs):=20loop-rule=20consolidatio?= =?UTF-8?q?n=20=E2=80=94=20Gate-A=20pass=207=20revision?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit All fourteen findings fixed. Two clusters carried the round. The decline/question/scope cluster (1-4) is one set of qualifications, not four patches: a matching decline suppresses the membership trigger only, a scope broadening discharges only the membership half of a dual-trigger hold, every decline-based exclusion reads "while the current set still excludes it", and the reachable clean/clearly-stuck overlap is admitted and left to the evaluation order rather than argued away. The snapshot transport (5, 6) simplifies rather than adds: the cycle's latest body is the authoritative answer set, squash carry copies that snapshot instead of every record in the range, an empty set is written explicitly, and a recovery candidate is a cycle identity rather than a body. Also: the standing curve paragraph now separates proof that a pass was valid from knowing it ran (8), per-hunk contract markers ship alongside the membership list, reversing this cycle's pass-6 choice (7), and seven sentence pairs that a fresh reader could take as two rules were found by reading both blocks end to end and repaired. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- ...26-09-10-loop-rule-consolidation-design.md | 288 ++++++++++++------ 1 file changed, 191 insertions(+), 97 deletions(-) diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index 20b9e4e..53d0ab0 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -43,9 +43,10 @@ does not settle, this spec decides in the section that uses it: the evaluation o field and file set each predicate reads; the duties' classification; the scope stop's two triggers and what each answer does; what a stuck or two-tell answer produces; the records' wording, attribution and recording point and the `Accepted:` label; and the raw-severity rule. -Four standing sentences are edited at their source rather than worked around — the two that -use *clean* in the file's sense, the resolve duty that never stated its scope, and the recovery -passage that would not have looked where the records live (§4 items 11–14). +Five standing sentences are edited at their source rather than worked around — the two that +use *clean* in the file's sense, the resolve duty that never stated its scope, the recovery +passage that would not have looked where the records live, and the curve's `?` rationale, +which answered the unavailable-history question the other way (§4 items 11–15). --- @@ -67,11 +68,12 @@ measures, the curve rule already summing the branches into one entry; a finding fields match one in the other branch is **one** finding for holds and answers, so it is answered once. Two things read the reviewer-written field from before the ceiling: the **health measures** (Mechanics, Severity) — the per-pass counts, the clusters, -the tells, and **both** conditions of the stuck reading, its Blocker curve and its -regenerating Blocker or Major findings, since a predicate reading one field for half of itself -could not be read at all — and the **five-field key** the answer records match on, whose -severity field is the one the reviewer wrote, so that a ceiling cannot make two findings the -same one. +the tells, and **both severity-bearing conditions of the three-condition stuck reading**, its +Blocker curve and its regenerating Blocker or Major findings, since a predicate reading one +field for half of itself could not be read at all, while its third condition, the stated +coverage-sufficiency judgement, reads no severity field and the ceiling does not touch it — +and the **five-field key** the answer records match on, whose severity field is the one the +reviewer wrote, so that a ceiling cannot make two findings the same one. **First, clean completion.** §5 uses *clean* in two senses and now says which is which. A **clean findings file** is the `NO FINDINGS` signal the protocol defines. A **clean pass** is @@ -83,7 +85,10 @@ trigger** — no finding outside that set, and none opening a new structural or question. Both halves are properties of the findings and the set, read before any branch below runs, which is what makes this order executable rather than asserted: no predicate here waits on an act a lower branch performs. A finding matched by a decline recorded in this cycle is -outside the set by that decision and triggers nothing (Mechanics, the answer records). A pass +outside the set by that decision **while the current assigned fix set still excludes it**, and +raises **no membership trigger and no membership hold** — and that alone: a **question stop** +still fires on it unless that same question has already been answered in this cycle (Mechanics, +the answer records). A pass with **zero** findings is clean whatever the floor, because a floor buys further looks at an artifact that keeps yielding findings, and one yielding none has already given what those looks were for. **A clean pass at or above the derived floor, @@ -107,10 +112,10 @@ filter and the clean-final-pass rule stand while it does. Any non-empty set of t to one pass: **one surface, every reason reported, every question asked**, because a reason left out is a decision made by omission. A finding surfaced solely by the stuck or two-tell reading is not a scope-stop finding; one that is also outside the set, or also opens a -question, takes the scope stop's answers at that same surface — it is not asked twice. A -re-raised finding that a decline of this cycle matches raises **no membership stop**, and one -differing from it in any of the five fields is a **new finding classified afresh**; Mechanics, -the answer records, states both, and this ordering does not restate them. +question, takes the scope stop's answers at that same surface — it is not asked twice. What a +decline of this cycle does to a re-raised finding, and what a changed field does, are stated +once each in Mechanics, the answer records, and in the first branch above; this branch adds +nothing to them and a reader who finds a rule here that is not there has found a defect. **Third, a pass that neither closes nor suspends continues** — the loop runs another pass on the **current** artifact, revised or not. Below the floor a clean pass lands here, and so does @@ -131,28 +136,36 @@ that the finding is not true of the artifact, a decline is the user's decision t finding stays outside the fix set, and only the second is an answer at a membership stop. The **hold** a surfaced finding places on closure is part of the ordering: it gates closing while it stands, and is discharged by the answers that finding requires, below. -**No-clean-credit** — no pass carrying a finding that gets surfaced is credited as clean — is -part of the ordering and a fact about that pass, discharged by nothing: a later pass is judged -on its own findings, so an answer never closes the cycle on the pass that surfaced the -finding. **The clean predicate carries it whole**, without reading the act: a scope-stop -trigger makes a pass unclean directly, and the clearly-stuck exit needs regenerating Blocker -or Major findings, which are either in the set — failing the predicate's first half — or -outside it, failing its second, so no pass either exit can surface on is clean. The two-tell -stop surfaces tells and not a finding, and was never in this duty's domain. +**No-clean-credit** — no pass that a scope stop surfaces on is credited as clean — is part of +the ordering and a fact about that pass, discharged by nothing: a later pass is judged on its +own findings, so an answer never closes the cycle on the pass that surfaced the finding. +**It is not a second test beside the clean predicate; it is that predicate's second half**, +which is why it is stated in the same words: a scope-stop trigger makes the pass unclean +directly, so no pass a scope stop can surface on is clean and nothing has to read the act. +**The other two exits are outside this duty and for different reasons**, said here so no +reader supplies a rule for them. The two-tell stop surfaces tells and not a finding, and was +never in the duty's domain. The clearly-stuck exit surfaces findings but is reached only on a +pass that did not close, and the fields differ — that exit reads reviewer-written severity +while cleanliness reads effective, so a demoted in-set Blocker leaves the pass clean, and a +clean pass is decided at the first branch, closing at or above the floor and continuing below +it. The order keeps the two apart; the duty never needed to. **What a suspension asks, and what ends it.** Only a scope stop raises a **hold**, on its -finding, and the hold ends when every answer that finding requires has been given — one for a -single-trigger finding, both the question decision and the membership answer for one carrying -both triggers — and in either direction: a hold only accepting could end would be the resolve -duty under another name. At a membership stop the answer is **accept** (the finding joins the +finding, and the hold ends when **every** answer that finding requires has been given — one for +a single-trigger finding, both the question decision and the membership answer for one carrying +both triggers — **and no direction is the wrong answer**, since a hold only accepting could end +would be the resolve duty under another name. At a membership stop the answer is **accept** (the finding joins the fix set and resolves by its effective severity: Blocker or Major before any pass can be clean, Minor or Nit collected and never iterated) or **decline** (the finding stays outside, binding for the rest of this cycle). Either answer is **recorded** — Mechanics, the answer records — and that record is where a later pass, or a cycle that lost its session, reads it. **Membership is read when the answer is given, not frozen at the surface**: where the governing artifacts have broadened the set so that the held finding is now inside it, that broadening -discharges the hold by itself — the finding is in-set, decline is unavailable to it, and it -resolves by its effective severity like any other in-set finding. At a question +discharges the hold's **membership component** by itself — the finding is in-set, decline is +unavailable to it, and it resolves by its effective severity like any other in-set finding. +Where that same finding also opened a question, **the question component stands until its +decision is given**, because a broadening answers who owns the work and never what the question +asked. At a question stop the answer is the user's decision on the question, and membership does not change: an in-set finding then routes through its effective severity like any other — a Blocker or Major resolves under that decision or is dismissed with its one-line why, a Minor or Nit is collected @@ -161,7 +174,8 @@ out-of-set finding that opened the question is a membership stop as well and tak decline. **Decline is available only at a membership stop**, because that is the only stop whose question is whether a finding belongs to the set, and a decline anywhere else would waive work the cycle owes. The stuck and two-tell readings -raise no hold: each asks one question, **continue or stop**. Continue resumes the loop on the +raise no hold: each asks one question, **continue or stop**. Continue is that suspension's +resuming answer, and where it is the last one outstanding the loop resumes on the artifact as revised and the fix set as the governing artifacts now assign it — where several plans or stories govern one cycle, the union of the scopes they assign — **plus the findings this cycle accepted into it**, which its **latest** commit body carries (Mechanics, the answer @@ -178,12 +192,16 @@ at a question stop, continue at the stuck or two-tell reading; one stop answer l whole suspension standing, because a loop resumed over an unanswered question decides it by running. **Simultaneous health suspensions are one question, not two**: the stuck and two-tell readings both ask continue or stop, so one answer carrying every reason ends both, and asking -twice would invite two answers to a question that has one. Two states cannot co-occur, and no -rule ranks them: clean completion and any suspension that surfaces a finding — the scope stop, -the clearly-stuck exit — since a pass carrying either's findings fails the clean predicate -above; and a zero-finding pass and any -suspension, since it has nothing to surface, nothing regenerating, no cluster and no -require↔withdraw pair. The Gate-B triviality skip is outside this ordering: a skipped cycle +twice would invite two answers to a question that has one. **Two pairings cannot occur**, and +no rule ranks them: clean completion and a **scope stop**, since that stop's triggers are the +clean predicate's own second half, so a pass raising one is not clean; and a zero-finding pass +and any suspension, since it has nothing to surface, nothing regenerating, no cluster and no +require↔withdraw pair. **Clean completion and the clearly-stuck exit can**, and the overlap is +admitted rather than argued away: that exit reads reviewer-written severity while cleanliness +reads effective, so an in-set Blocker the ceiling demotes can regenerate across passes on a +pass that is clean. **The order decides it and no new rule is needed** — the pass closes at or +above the floor, ranking the exit exactly as the clearly-stuck paragraph's own precedence +sentence says, and continues below it, where nothing closes anyway. The Gate-B triviality skip is outside this ordering: a skipped cycle runs no passes and ends by its own rule. **This ordering is one component of the closure-record contract** the one-contract rule names, and a copy carrying it without the rest of that list is a partial adoption that stops there. @@ -195,7 +213,10 @@ can make the pass unclean: clean completion is read first, so a suspension is on evaluated on a pass that did not close — which works only because clean candidacy reads the **triggers** and never the act of surfacing, as the block's own sentence says. The three branches are AC 4, the duties paragraph AC 2, -the composition and cannot-co-occur sentences AC 1. The scope stop's two triggers are `b11` +the composition sentences AC 1 — which asks that a conflict unable to co-occur be named with +its reason rather than legislated, and is met on both sides: the scope stop cannot co-occur +with clean completion and says why, the clearly-stuck exit can and is resolved by the order +already stated rather than by a rule invented for it. The scope stop's two triggers are `b11` (membership) and `b13` (question) read separately, because an in-set finding that opens a question can neither join nor stay outside the set and needs its own answer — which is also why a matching decline suppresses the membership trigger only. Sentences the block points at rather @@ -233,7 +254,8 @@ byte-identical in both copies: finding joins the assigned fix set, and the user's answer is recorded under one of two labels sharing one form. **Accepted** puts the finding in the set, from where the severity rules already govern it: a Blocker or Major owes resolution, a Minor or Nit is collected and - never iterated. **Declined** keeps it out and releases its hold. Both are available at a + never iterated. **Declined** keeps it out and ends the membership half of its hold — the + whole of it where membership was all that finding raised. Both are available at a membership stop and nowhere else: not for an in-set Blocker or Major, which owes resolution already, and neither is the answer to a question stop, a stuck or two-tell surface, a below-floor pass, an unclean final pass, or any Gate-A, @@ -269,19 +291,26 @@ byte-identical in both copies: its own**, the destination the human-exception rule above already blesses: "An empty commit carrying only the record is a legitimate destination". Each is **restated in every later body of that cycle**, the closing one included, so that the cycle's **latest** body carries the - complete answer set: **that body is the authoritative one**, and an answer missing from it is - lost whatever an earlier body says, because a rule that let any reachable body revive an - answer would make the set depend on how far back a reader looked. Each is **copied on - squash-merge** (the carry rule above). A cycle that resumes - with no commit body carrying an answer **treats it as absent** — the unknown-start fallback's + complete answer set: **that body is the authoritative snapshot**, and an answer missing from + it is lost whatever an earlier body says, because a rule that let any reachable body revive + an answer would make the set depend on how far back a reader looked. **A cycle with no + answers writes that too** — one line, `Cycle answers: none · cycle · · + `, in the same fields the record header carries — so a body stating an empty set is + distinguishable from one that dropped its records, which is the difference every reader of + these bodies turns on. Everything that reads these records reads **that snapshot and no + earlier body**: recovery, the squash carry, and the fix set the loop resumes with. Each is + **copied on squash-merge** (the carry rule above). A cycle that resumes and finds no answer + in that snapshot **treats it as absent** — the unknown-start fallback's reading, and the safe direction under both labels: an unrecorded decline means the hold - applies again, an unrecorded acceptance means the finding is raised afresh. Each is an + applies again, an unrecorded acceptance means the finding is raised afresh. There is no + second place to look, which is the point of naming one body authoritative. Each is an **unverified assertion** of the same kind as the human exception — nothing checks that the handle belongs to whoever decided, that a human was asked, or that the reason is honest. **Binding, and the sameness test.** A decline **binds for the remainder of its cycle**, with - no effect in any later one, and never qualifies the Blocker/Major-resolve duty, which the - declined finding never reached. A later pass raises **the same finding** when all five of + no effect in any later one, and never qualifies the Blocker/Major-resolve duty — which the + declined finding does not reach **while it stays outside the set**, the duty being scoped to + what is in it. A later pass raises **the same finding** when all five of location, defect, severity, consequence and suggested fix match, read on meaning rather than bytes, since a reviewer rewrites its sentences between passes; the severity read is the one the reviewer wrote, per the ordering's field rule. A finding matching a decline @@ -292,7 +321,12 @@ byte-identical in both copies: genuine uncertainty makes it a new finding**, classified afresh against the current fix set and the question predicate rather than inheriting a stop from the finding it resembles. A decline keeps a finding out and never excuses one that is in: a declined finding the fix set - later comes to include owes resolution like any other. An acceptance does not expire with + later comes to include owes resolution like any other. **The binding and the set are + different things**, which is how both hold at once: the decline binds the *decision* for the + rest of the cycle, so that question is never re-asked, while membership is owned by the + governing artifacts — a broadening puts the finding in the set without the decline having + expired, and every exclusion this record grants reads "while the current set still excludes + it". An acceptance does not expire with its pass: the finding is in the set until the cycle closes. **What each is worth.** A decline **releases a hold** and an acceptance **puts a finding in @@ -301,16 +335,18 @@ byte-identical in both copies: allow**: two cycles sharing or redrawing a nonce are indistinguishable to these records as to every other, so a replayed answer can bind to the wrong cycle, and a bound is not safety. **What the acceptance record buys, and what it does not:** a cycle that lost its session and - started fresh reads the branch's commit bodies among its recovery sources and finds what was - accepted. Nothing makes it act on them and nothing checks that it did, so a replacement - cycle can still review and close the same artifact while an older one stays open; the record - makes that discoverable rather than invisible, which is less than preventing it. **What no + **recovers its own identity** reads its latest body among its recovery sources and finds what + it had accepted. One that cannot recover it starts a new cycle, which **inherits nothing** — + it names the cycles it did not adopt, marks their answers unknown, and obtains its own. So a + replacement can still review and close the same artifact while an older one stays open; the + record makes that discoverable rather than invisible, which is less than preventing it. **What no record carries:** a question stop's decision, where it changed no membership — the five fields identify a finding and not a question, and that decision's durable form is the artifact revision it produces, **where it produces one**; a decision that changes nothing stays in the session that made it — and a stuck or two-tell surface with its continue-or-stop answer. From a closing body a reader can infer the close, the acceptances - and the declines, and after a lost session neither of those two is recoverable from it. + and the declines; after a lost session neither the question decision nor the + continue-or-stop answer is recoverable from it. **These records are one component of the closure-record contract** the one-contract rule above names, and a copy carrying them without the rest of that list is a partial adoption @@ -336,26 +372,33 @@ are in the site map and re-read at execution. 1. **Squash carry** (C:892 / W:1076, `j1`). OLD: "…copy every evidence entry, every human-exception record, the provenance lines, the curves and any skipped cycle's skip record TOGETHER WITH THE SKIP REASON IT POINTS AT…". NEW: "…copy every evidence entry, - every human-exception record, **every answer record (one copy per cycle nonce and - five-field finding key: byte-identical repeats collapse, and copies that disagree stop - under the rule below)**, the provenance lines, the curves and any skipped cycle's skip - record TOGETHER WITH THE SKIP REASON IT POINTS AT…". Six members; the dedup clause is there - because a record restated in every body reaches the squash range many times. + every human-exception record, **every cycle's latest answer-record snapshot — the complete + set as that cycle's newest body in the range states it, earlier restatements being + superseded rather than merged (part of the closure-record contract above; a copy carrying + this without the rest stops there)**, the provenance lines, the curves and any skipped + cycle's skip record TOGETHER WITH THE SKIP REASON IT POINTS AT…". Six members. Copying the + snapshot rather than every record in the range is what makes the carry agree with the + authoritative-body rule: a record restated in every body reaches the range many times, and + summing them would let a body that dropped an answer be overruled by an older one that + still carried it — the revival the snapshot rule exists to forbid. 2. **The named nonce set** (C:385–387 / W:579–581). OLD: "**and that set is named rather than left open**: the provenance line, the per-pass curve (including a skip record standing in for one), the cycle's findings slots, and its advisory working record." NEW: "**and that set is named rather than left open**: the provenance line, the per-pass curve (including a skip - record standing in for one), **any answer record (Mechanics)**, the cycle's findings - slots, and its advisory working record." + record standing in for one), **any answer record (Mechanics; part of the closure-record + contract below, and a copy carrying this without the rest stops there)**, the cycle's + findings slots, and its advisory working record." 3. **The nonce exemption** (C:397–399 / W:591–593). OLD: "The nonce is not required in records this change neither introduces nor keys to a cycle — the evidence entry and a human-exception record among them." NEW: "The nonce is not required in records that are not keyed to a cycle — the evidence entry and a human-exception record among them; **an answer - record is keyed to its cycle and carries it**." The old "this change" dated the sentence to - the parent; the new one states the criterion. + record is keyed to its cycle and carries it (part of the closure-record contract below; a + copy carrying this without the rest stops there)**." The old "this change" dated the + sentence to the parent; the new one states the criterion. 4. **"Both shipped records below"** (C:367 / W:561). OLD: "Both shipped records below carry a **cycle field**, because…". NEW: "Both shipped records below carry a **cycle field** — and - so do the answer records in Mechanics — because…". A load-bearing count a third + so do the answer records in Mechanics, part of the closure-record contract below, a copy + carrying this without the rest stopping there — because…". A load-bearing count a third cycle-attributed record would otherwise falsify. 5. **The unknown-start fallback** (C:153–167 / W:360–374, `i4`–`i8`, extended at `i12`'s invitation). OLD: "…at minimum floor 3, severity classified without the demotion, the @@ -364,7 +407,8 @@ are in the site map and re-read at execution. duty owed, the curve duty owed, the nonce duties at their strictest, **every suspension binding, and answer records not attributable to the nonce this fallback minted treated as absent, so that no inherited hold is released and no inherited acceptance is claimed — - records made and recorded under that nonce are the cycle's own and are honoured**…". + records made and recorded under that nonce are the cycle's own and are honoured (part of the + closure-record contract below; a copy carrying this without the rest stops there)**…". `i12`'s sentence stays as written; this is the addition it invites. The time bound matters: treating *every* answer as absent would leave a fallback cycle unable to release a hold it raised itself, **D7** unmet. @@ -374,13 +418,18 @@ are in the site map and re-read at execution. 7. **The gate-off surface** (C:143–151 / W:350–358). OLD tail: "…silencing reminders; or not running a pass and reporting that it ran." NEW tail: "…silencing reminders; not running a pass and reporting that it ran; **recording a decline nobody made, or one on an in-set - finding; or deleting an acceptance the cycle owes**." The list says it is not complete; - this change opens those routes and names them, as the parent did for the stated floor. + finding; or dropping an acceptance the cycle owes from the body that would carry it (part + of the closure-record contract below; a copy carrying this without the rest stops there)**." + The list says it is not complete; this change opens those routes and names them, as the + parent did for the stated floor. "Dropping from the body" and not "deleting", because under + the snapshot rule an omission is the whole of the act. 8. **The closing message and the soft-reset path** (C:834–838 / W:1018–1022). Two insertions into a sentence whose other clauses are unchanged. After "…which owes no entry" add: "— - **and every answer record the cycle made, including any held only in WIP bodies a - `git reset --soft` collapsed**: the single commit after the reset carries all of them, - because a body the reset discards is unreachable from the commit that replaces it". In the + **and the cycle's complete answer-record snapshot, including any answer held only in WIP + bodies a `git reset --soft` collapsed**: the single commit after the reset carries the whole + set, because a body the reset discards is unreachable from the commit that replaces it + (part of the closure-record contract below; a copy carrying this without the rest stops + there)". In the sentence after it, "so an entry written only into the WIP body" becomes "so an entry **or record** written only into the WIP body". The `git reset --soft` sentence itself (C:829–830 / W:1013–1014) is unchanged. @@ -391,9 +440,10 @@ are in the site map and re-read at execution. rules cannot be established** depend on one another," NEW opening: "…this carry rule, **the closure-record contract — the closure ordering, the answer records, the severity rule's raw-versus-effective split and its assigned-fix-set boundary, the two clean-vocabulary - edits, the answer records' membership in the named nonce set, the unknown-start item - covering them, the closing-message carry, the squash carry, the recovery sources and the - no-identity report** — and **the unknown-start activation semantics that say what a cycle owes + edits, the answer records' membership in the named nonce set and the nonce exemption's + criterion, the cycle-field count they falsify, the unknown-start item covering them, the + closing-message carry, the squash carry, the curve's validity rule, the recovery sources, + the no-identity report and the gate-off routes they open** — and **the unknown-start activation semantics that say what a cycle owes when its starting rules cannot be established** depend on one another," and after "and a carry rule naming records a project does not produce is inert." (C:885) add: "an ordering without the answer records is a membership stop whose two answers nothing carries; answer @@ -405,12 +455,13 @@ are in the site map and re-read at execution. or be found again." The stop sentence that follows is unchanged and now covers these states. Then, before it: "**A project carrying any component of this contract owes all of them.** The pieces are separately mergeable and are not separately adoptable, so a copy holding one - without the rest is an incomplete adoption and stops here." That is deliberately not the - finding's other option, a reciprocal marker on every hunk: one membership list is one thing - to keep in step, and eleven cross-references are eleven. Three pieces still name the - contract where a reader meets them first — §3's last sentence, the answer-record block's - last paragraph, the (g) replacement's last sentence — so a copy holding only those stops on - the text it has, but the obligation is the list's, not the marker's. **Rollback.** A cycle open when the text is + without the rest is an incomplete adoption and stops here." **Both halves of the guard ship, + and that reverses this cycle's earlier choice.** Pass 6 offered a central list or a marker + on every mergeable hunk; the list was taken alone, and pass 7 showed why that is not enough + — a list cannot police a merge that omits the list. So every hunk in this section also + carries a short marker naming the contract, and the three blocks keep the longer sentence + they already had. The list is what defines membership; the markers are what a partial merge + still sees. **Rollback.** A cycle open when the text is reverted is governed by "a cycle already running finishes under the rules it started with" and "A revert is itself a shipping commit for the old rules" (C:153–154, C:165–166) where it can still establish those rules, and by the unknown-start fallback where it cannot. **Where @@ -424,9 +475,12 @@ are in the site map and re-read at execution. them." NEW: "**Starting a new cycle does not close, adopt or retire the cycles those candidates belong to** — they stay open, keep their own nonces, and are a human's to resolve; the new cycle simply does not claim them, **and names them in its first pass - report**, so the human this rule makes responsible learns they exist." It adds no record — - the report names state the workspace already holds — and closes the one shape "a human's to - resolve" cannot reach: an open cycle nobody is told about. + report, marking their exit and their answers unknown**, so the human this rule makes + responsible learns they exist and nobody reads the new cycle as continuing them. **It + inherits no answer**: whatever those cycles accepted or declined, this one asks again (part + of the closure-record contract above; a copy carrying this without the rest stops there)." + It adds no record — the report names state the workspace already holds — and closes the one + shape "a human's to resolve" cannot reach: an open cycle nobody is told about. 11. **The recovery sources** (C:411–417 / W:605–611) — the **whole passage** is replaced, not a sentence inside it, because the sentence that must change ("no search there") is the same one that makes the rest coherent. OLD: "**Recovery has two sources, and they answer different @@ -443,14 +497,19 @@ are in the site map and re-read at execution. bodies carry — its artifact revisions, and any empty commit carrying an answer record — **searched newest first and no further back than the branch point**, each candidate validated by cycle field, kind and artifact exactly as a working record is; a bounded search, because an - unbounded one would reach other cycles' commits. **History is the source once the cycle's own + unbounded one would reach other cycles' commits. **A candidate is a cycle identity — the + cycle field, the kind and the artifact together — and not a body**, so one cycle's own + restatements across several bodies are one candidate and never trip the more-than-one rule; + **its newest body is the answer set and the search stops there**, since an older body is + superseded rather than merged. **History is the source once the cycle's own closing commit exists**, and there is no search there: the cycle is reading **its own commit body**, so kind and artifact are settled by which commit is being read, and the nonce is taken from the provenance line and the curve, which must agree." The next sentence's "Recovering a single candidate from **either**" becomes "from **any of the three**" — a one-word repair the passage's arithmetic forces, and the failure rule after it ("No candidate, disagreeing sources, or more than one candidate → no identity") is unchanged and already governs all three, which is - why the NEW does not restate it. Left alone, an answer record would exist and the + why the NEW does not restate it. Part of the closure-record contract above; a copy carrying + this without the rest stops there. Left alone, an answer record would exist and the procedure meant to read it would never look, and story criterion 7 would fail on one unchanged sentence. @@ -460,6 +519,7 @@ are in the site map and re-read at execution. `NO FINDINGS` with `END OF FINDINGS (0 total)`." It describes a file and always did; the word *pass* in it is what makes §3's predicate look like a redefinition instead of the other sense. Its condition — what a reviewer writes when it finds nothing — is unchanged. + Part of the closure-record contract below; a copy carrying this without the rest stops there. 13. **The Gate-A clean-signal sentence** (C:565–566 / W:756–757; the two copies wrap it differently, so the shared fragment is what is quoted). OLD: "…when a pass is clean — the explicit clean signal is what lets you exit the loop:". NEW: "…when a pass finds nothing — @@ -467,22 +527,42 @@ are in the site map and re-read at execution. Old condition: a `NO FINDINGS` file is what permits loop exit. Kept: the signal and why it is demanded. Changed: it makes a pass readable as clean rather than being the only way to be clean, since a pass carrying Minors alone is clean under the ordering and could never - produce this file. The other five uses of "clean pass" in each copy (C:117, 728, 750, 761, - 769, 827) are the closure sense the ordering defines and are correct as they stand — - checked, not assumed. + produce this file. Part of the closure-record contract below; a copy carrying this without + the rest stops there. The other **six** uses of "clean pass" in each copy (C:117, 728, 750, + 761, 769, 827; W:324, 914, 936, 947, 955, 1011) are the closure sense the ordering defines + and are correct as they stand — counted and checked, not assumed. 14. **The Severity bullet's resolve duty** (C:783–784 / W:969–970), which is the one place the duty is stated and the only one without a scope. OLD: "- **Severity:** Blocker (wrong/unsafe/breaks invariant) · Major (design flaw → rework) → both must resolve. Minor · Nit → collect, never iterate." NEW: "- **Severity:** Blocker (wrong/unsafe/breaks invariant) · Major (design flaw → rework) → both must resolve, **for every finding in the - assigned fix set**; one the user declined at a membership stop is outside that set and owes - nothing (the closure ordering above). Minor · Nit → collect, never iterate." Old condition: + assigned fix set**; one the user declined at a membership stop is outside that set, and owes + nothing **while the set still excludes it** (the closure ordering above; part of the + closure-record contract below, and a copy carrying this without the rest stops there). + Minor · Nit → collect, never iterate." Old condition: every Blocker and Major resolves, unbounded. **Replaced**: the boundary **D5** always implied and no sentence carried. Without it a declined finding must stay outside the set and still bars closure, which is a pass that can neither close nor suspend. The (c) pointer that calls this "the resolve rule" (`c17`) needs no edit: it names the rule, and the rule now carries its own scope. +15. **The curve's one-entry-per-valid-pass paragraph** (C:943–950 / W:1127–1134), which is + where `?` gets its rationale and where that rationale currently answers Q6 the other way. + OLD, the clause: "so a resumed cycle may know a pass happened and not what it found, and + zero and unknown are different facts." NEW: "so a resumed cycle may hold durable proof that + a pass was **valid** and no longer hold what it found, and zero and unknown are different + facts. **Knowing that a pass ran is not that proof**: where validity cannot be established + — the slot gone or unreadable, and nothing recording that it was accepted — the pass is + **omitted from the pass specification** rather than entered with `?`, because this grammar + takes one entry per *valid* pass and has no way to say "may not have been one", and the + report names the numbers it omitted (part of the closure-record contract above; a copy + carrying this without the rest stops there)." Old condition: a resumed cycle's knowledge + that a pass happened is enough to keep its entry, with `?` for the counts. **Replaced**: + knowledge that it ran is separated from proof that it was valid, because §7 makes the + second unavailable after a lost session and the two answers cannot both govern one slot. + Kept: `?` itself, per series, for a pass whose validity is established and whose counts are + not. Without this edit the same missing slot both keeps an entry and is omitted. + The shorter "Copy every record into the squash body" sentence inside the human-exception block (C:1004–1007) is generic and already covers an answer record; it is not edited. The "records every cycle owes" list (C:698) enumerates unconditional records only; an answer record is @@ -599,7 +679,7 @@ the whole paragraph, both copies, with the answer. NEW: one this must not be read as touching; a loop spending passes on findings the author keeps demoting is exactly what the prose-cluster tell exists to surface, and lowering the counts by that same judgement would hide it. **This split is one component of the closure-record - contract** the one-contract rule above names, and a copy carrying it without the rest of + contract** the one-contract rule below names, and a copy carrying it without the rest of that list is a partial adoption that stops there. ``` @@ -616,7 +696,7 @@ answer to. **(i) When these rules bind** (C:153–167 / W:360–374) — **extend** the strict-reading list (§4 item 5); `i1`–`i16` kept. **(j) The squash carry** (C:892 / W:1076) — **extend** (§4 item -1); `j1`–`j4` kept. Also touched, outside the inventoried passages: §4 items 2, 3, 4, 7–14. +1); `j1`–`j4` kept. Also touched, outside the inventoried passages: §4 items 2, 3, 4, 7–15. --- @@ -653,10 +733,14 @@ earlier pass had removed.": anything. **First the root**, which is a condition on the current pass being valid at all and not a stop of its own. The slots live in `.context/codex-reviews/` under the top-level directory of the checkout this cycle is running in — the same root the pass call is given as -`workingDirectory`. Establishing it fails in two observable ways, each with its own fix: the -working directory is **not inside a git repository** (run the pass from the checkout), or the -resolved top level **differs from the directory the call was given** (re-issue the call with -the resolved root). A pass whose root cannot be established is an **INCOMPLETE pass** — the +`workingDirectory`. Establishing it fails in three observable ways, each with its own fix: the +working directory is **not inside a git repository** (run the pass from the checkout); the +resolved top level **differs from the directory the call was given**, compared after both are +**canonicalized**, so that a symlink or a trailing slash is not a mismatch (re-issue the call +with the resolved root); or **the query itself fails** — git unavailable, repository ownership +rejected, metadata unreadable — which is reported with the error it returned and not as a +missing repository, since the fix is to make git usable in that checkout rather than to move. +A pass whose root cannot be established is an **INCOMPLETE pass** — the state this section already defines, already uncounted toward the floor and already excluded from the curve — so nothing new is ranked in the closure ordering. What no check reaches: a slot written under a different root **in the past** is **indistinguishable from an absent @@ -691,10 +775,11 @@ unreadable: Root: `.context/codex-reviews/` under this checkout — confirmed. (Unconfirmable: this pass is INCOMPLETE, root not established; re-run from the checkout.) - Trend: findings 24, 17, —, 18; Blockers 5, 2, —, 1 — pass 3 not counted (slot - absent, acceptance unknown; series `?`); 2→4 spans that gap: rising, weaker - evidence. Cluster: product behaviour 12 of 18. require↔withdraw: none visible; - a pair with pass 3 as an endpoint cannot be read. Threshold read on 3 of 4. + Trend: findings 24, 17, —, 18; Blockers 5, 2, —, 1 — pass 3 omitted: slot absent, + acceptance unknown, so it is neither counted toward the floor nor entered in the + curve. 2→4 spans that gap: rising, weaker evidence. Cluster: product behaviour + 12 of 18. require↔withdraw: none visible; a pair with pass 3 as an endpoint + cannot be read. Threshold read on 3 of 4. ``` This is **D10**. Both historical lines are derivable from the mandated findings files alone @@ -760,12 +845,21 @@ not: - the recovery-source edit (§4 item 11): `Recovery has three sources` count **1** in each and **0** in the parent tree, with `Recovery has two sources` at **0** in each and **1** there, since the passage is replaced and a NEW installed beside a surviving OLD is the failure; -- the four source edits, each as a pair, because each is a standing sentence changing meaning +- the five source edits, each as a pair, because each is a standing sentence changing meaning rather than new text appearing: `A **clean findings file** is the single body line` at **1** in each and **0** in the parent, with `A\n> clean pass is the single body line` at **0** and **1**; `when a pass finds nothing` at **1** and **0**, with `when a pass is clean — the explicit clean signal` at **0** and **1**; `for every finding in the assigned fix set` at - **1** and **0**. A one-sided presence check would pass on a copy carrying both wordings. + **1** and **0**; `Knowing that a pass ran is not that proof` at **1** and **0**, with + `may know a pass happened and not what it found` at **0** and **1**. A one-sided presence + check would pass on a copy carrying both wordings; +- the contract markers, since finding 7 turned on their absence: `part of the closure-record + contract` **case-insensitively** at **13** in each and **0** in the parent — one per source + edit in §4 except items 6 and 9, which are respectively unchanged and the contract itself, + and three of the thirteen open a sentence and so capitalise it — plus the three block-level + sentences (`This ordering is one component`, `These records are one component`, `This split + is one component`) at **1** each and **0** there. Sixteen markers in each copy; a + case-sensitive count would find ten and read as a failure. If the claim "the ordering and the records ship in both copies" were false, one working-tree count would be **0** or the parent-tree counts would not differ from it. The wiring can produce From 2feab0c7565b40c46fc41b3cf6ae384b49f1fc2b Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Thu, 10 Sep 2026 18:08:14 +0200 Subject: [PATCH 011/181] =?UTF-8?q?docs(specs):=20loop-rule=20consolidatio?= =?UTF-8?q?n=20=E2=80=94=20Gate-A=20pass=208=20revision?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit All twelve pass-8 findings fixed. Two changes carry the round. The scope stop's two triggers are now defined exactly once, in the ordering's first branch, so that cleanliness and suspension read the same words and a matching decline suppresses the membership trigger without a second sentence saying so. The four duties are reconciled rather than narrowed: the hold belongs to the scope stop, the only stop whose question is about a finding; no-clean-credit attaches to any pass carrying a scope-stop trigger; and where the clearly-stuck exit's regenerating findings are demoted below Major the pass is clean at effective severity and clean completion wins by D3. c16 and c18 are marked replaced, with their old conditions enumerated and their authority named. Also: every cycle body carries its answer snapshot and is found by its identity line; at most one effective label per cycle and finding key; four named root failures; and a false-red assert corrected, after which every assert and every stated count was re-checked against what it detects or enumerates. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- ...26-09-10-loop-rule-consolidation-design.md | 229 +++++++++++------- 1 file changed, 139 insertions(+), 90 deletions(-) diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index 53d0ab0..f46c490 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -1,6 +1,6 @@ # §5 loop-rule consolidation: one closure ordering — Design -**Date:** 2026-09-10 · **Status:** DRAFT — Gate A running (pass 6 revised) +**Date:** 2026-09-10 · **Status:** DRAFT — in Gate A, not yet approved **Story:** `docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md` **Profile:** read from that header at every pass, never from here — it is the only writable copy, and a value copied here would be a remembered value. @@ -81,14 +81,15 @@ the predicate below: a clean findings file always gives one, and a clean pass ne one, because a pass carrying only Minors and Nits, or findings a decline of this cycle matches, is clean without being empty. **A pass is clean** when its findings carry no Blocker or Major at effective severity that is **in the assigned fix set**, and **no scope-stop -trigger** — no finding outside that set, and none opening a new structural or contract -question. Both halves are properties of the findings and the set, read before any branch below +trigger**. There are two triggers, defined here and nowhere else, so that cleanliness and +suspension read the same words: a **membership trigger** is a finding outside the current +assigned fix set **and not matched by a binding decline of this cycle** — such a decline having +put it outside by the user's own answer, for as long as the current set still excludes it — and +a **question trigger** is a finding opening a new structural or contract question, in-set or +not, which a decline never suppresses, the five fields it records identifying a finding and +never a question. Both are properties of the findings and the set, read before any branch below runs, which is what makes this order executable rather than asserted: no predicate here waits -on an act a lower branch performs. A finding matched by a decline recorded in this cycle is -outside the set by that decision **while the current assigned fix set still excludes it**, and -raises **no membership trigger and no membership hold** — and that alone: a **question stop** -still fires on it unless that same question has already been answered in this cycle (Mechanics, -the answer records). A pass +on an act a lower branch performs. A pass with **zero** findings is clean whatever the floor, because a floor buys further looks at an artifact that keeps yielding findings, and one yielding none has already given what those looks were for. **A clean pass at or above the derived floor, @@ -98,24 +99,25 @@ re-read before the pass is accepted as final, a header or profile that changed d making the pass not final and costing the further pass that section requires. A plateau or tells present on the closing pass go into the closing report and never block it, because reporting "will not converge" on a converged loop -is a false report, and the clearly-stuck paragraph says the same of its own exit. **Nothing -else closes a cycle**, because every other way out of a pass leaves a finding or a question -open, and closing over one is the failure this ordering exists to prevent. +is a false report, and the clearly-stuck paragraph says the same of its own exit. **No other +pass outcome closes a cycle**, because every other way out of a pass leaves a finding or a +question open, and closing over one is the failure this ordering exists to prevent. The one +termination that is not a pass outcome is the Gate-B triviality skip, which ends a cycle +having run no passes and is outside this ordering entirely. **Second, only a pass that is not a clean completion can suspend** — that order is what makes "clean completion outranks the two-tell stop" executable rather than asserted. Three -suspensions, by the names their paragraphs use: the **scope stop**, with two triggers — a -**membership stop** (a finding outside the assigned fix set) and a **question stop** (a -finding, in-set or not, opening a new structural or contract question); the **clearly-stuck -exit**; and the **two-tell stop**. A suspension waives nothing — the floor, the Blocker/Major -filter and the clean-final-pass rule stand while it does. Any non-empty set of them can apply -to one pass: **one surface, every reason reported, every question asked**, because a reason -left out is a decision made by omission. A finding surfaced solely by the stuck or two-tell -reading is not a scope-stop finding; one that is also outside the set, or also opens a -question, takes the scope stop's answers at that same surface — it is not asked twice. What a -decline of this cycle does to a re-raised finding, and what a changed field does, are stated -once each in Mechanics, the answer records, and in the first branch above; this branch adds -nothing to them and a reader who finds a rule here that is not there has found a defect. +suspensions, by the names their paragraphs use: the **scope stop**, raised by either trigger +defined above — a **membership stop** by the first, a **question stop** by the second; the +**clearly-stuck exit**; and the **two-tell stop**. A suspension waives nothing — the floor, the +Blocker/Major filter and the clean-final-pass rule stand while it does. Any non-empty set of +them can apply to one pass: **one surface, every reason reported, every question asked**, +because a reason left out is a decision made by omission. A finding surfaced solely by the +stuck or two-tell reading is not a scope-stop finding; one that also carries either trigger +takes the scope stop's answers at that same surface — it is not asked twice. This branch +defines nothing: the triggers are the first branch's, and what a changed field does to a +re-raised finding is Mechanics', in the answer records. A reader who finds a rule here that is +not in one of those two has found a defect. **Third, a pass that neither closes nor suspends continues** — the loop runs another pass on the **current** artifact, revised or not. Below the floor a clean pass lands here, and so does @@ -135,20 +137,23 @@ clean predicate's own wording so that the duty and the predicate cannot drift ap that the finding is not true of the artifact, a decline is the user's decision that a true finding stays outside the fix set, and only the second is an answer at a membership stop. The **hold** a surfaced finding places on closure is part of the ordering: it gates closing -while it stands, and is discharged by the answers that finding requires, below. -**No-clean-credit** — no pass that a scope stop surfaces on is credited as clean — is part of -the ordering and a fact about that pass, discharged by nothing: a later pass is judged on its -own findings, so an answer never closes the cycle on the pass that surfaced the finding. -**It is not a second test beside the clean predicate; it is that predicate's second half**, -which is why it is stated in the same words: a scope-stop trigger makes the pass unclean -directly, so no pass a scope stop can surface on is clean and nothing has to read the act. -**The other two exits are outside this duty and for different reasons**, said here so no -reader supplies a rule for them. The two-tell stop surfaces tells and not a finding, and was -never in the duty's domain. The clearly-stuck exit surfaces findings but is reached only on a -pass that did not close, and the fields differ — that exit reads reviewer-written severity -while cleanliness reads effective, so a demoted in-set Blocker leaves the pass clean, and a -clean pass is decided at the first branch, closing at or above the floor and continuing below -it. The order keeps the two apart; the duty never needed to. +while it stands, and is discharged by the answers that finding requires, below. **A scope stop +is the only stop that holds a finding**, because it is the only one whose question is about a +finding; the health exits hold nothing, which is why their answer is continue or stop rather +than a direction on anything. **No-clean-credit** — no pass carrying a scope-stop trigger is +credited as clean — is part of the ordering and a fact about that pass, discharged by nothing: +a later pass is judged on its own findings, so an answer never closes the cycle on the pass +that surfaced the finding. **It is not a second test beside the clean predicate; it is that +predicate's second half**, which is why it is stated in the same words: the trigger makes the +pass unclean directly, so nothing has to read the act of surfacing. +**Where the other two exits stand in this duty**, said here so no reader supplies a rule for +them. The two-tell stop surfaces tells and not a finding, and was never in its domain. The +clearly-stuck exit does surface findings, and where those carry a scope-stop trigger the duty +reaches them exactly as it reaches any other. Where they do not — its regenerating findings +are in-set, reviewer-written Blocker or Major, and the ceiling demotes them below Major — the +pass is clean at effective severity, the two predicates reading different fields on purpose, +and **clean completion wins by D3**: the pass closes at or above the floor and continues below +it. The order decides that; the duty never needed to. **What a suspension asks, and what ends it.** Only a scope stop raises a **hold**, on its finding, and the hold ends when **every** answer that finding requires has been given — one for @@ -201,8 +206,7 @@ admitted rather than argued away: that exit reads reviewer-written severity whil reads effective, so an in-set Blocker the ceiling demotes can regenerate across passes on a pass that is clean. **The order decides it and no new rule is needed** — the pass closes at or above the floor, ranking the exit exactly as the clearly-stuck paragraph's own precedence -sentence says, and continues below it, where nothing closes anyway. The Gate-B triviality skip is outside this ordering: a skipped cycle -runs no passes and ends by its own rule. **This ordering is one component of the +sentence says, and continues below it, where nothing closes anyway. **This ordering is one component of the closure-record contract** the one-contract rule names, and a copy carrying it without the rest of that list is a partial adoption that stops there. ``` @@ -293,11 +297,17 @@ byte-identical in both copies: of that cycle**, the closing one included, so that the cycle's **latest** body carries the complete answer set: **that body is the authoritative snapshot**, and an answer missing from it is lost whatever an earlier body says, because a rule that let any reachable body revive - an answer would make the set depend on how far back a reader looked. **A cycle with no - answers writes that too** — one line, `Cycle answers: none · cycle · · - `, in the same fields the record header carries — so a body stating an empty set is - distinguishable from one that dropped its records, which is the difference every reader of - these bodies turns on. Everything that reads these records reads **that snapshot and no + an answer would make the set depend on how far back a reader looked. **Every body of the + cycle carries that snapshot, and a cycle with no answers carries it too** — one line, + `Cycle answers: none · cycle · · `, in the same fields the record + header carries. Two things follow, and both are why it is every body and not only the ones + with something to say. A body stating an empty set is distinguishable from one that dropped + its records, which is the difference every reader of these bodies turns on. And **the newest + body is found by that line and never by whether it happens to carry an answer**: a body + identified only by the records in it is invisible on the day it drops them, which is the + day it matters. **A body of this cycle whose snapshot is missing or malformed stops** the + cycle rather than being skipped, since skipping it is exactly what revives the superseded + body behind it. Everything that reads these records reads **that snapshot and no earlier body**: recovery, the squash carry, and the fix set the loop resumes with. Each is **copied on squash-merge** (the carry rule above). A cycle that resumes and finds no answer in that snapshot **treats it as absent** — the unknown-start fallback's @@ -313,11 +323,14 @@ byte-identical in both copies: what is in it. A later pass raises **the same finding** when all five of location, defect, severity, consequence and suggested fix match, read on meaning rather than bytes, since a reviewer rewrites its sentences between passes; the severity read is the one - the reviewer wrote, per the ordering's field rule. A finding matching a decline - of this cycle raises **no membership stop and no membership hold**, so a decline is not - re-asked every pass — **and that alone**: the five fields record no question, so a - **question stop** still fires on it unless that same question has already been answered in - this cycle (the closure ordering above). **Any difference — severity included — or any + the reviewer wrote, per the ordering's field rule. **At most one effective label per cycle + and five-field key.** Exactly equivalent same-label records collapse to one, which is what + makes restating the snapshot idempotent; two records disagreeing on that key — opposite + labels, or one label with disagreeing attribution — **stop**, because a finding both in the + set and kept out of it has no reading, and picking either would let a fabricated or replayed + record decide which. What a matching decline does to a re-raised finding is the closure + ordering's membership trigger, which is defined there and not restated here. **Any + difference — severity included — or any genuine uncertainty makes it a new finding**, classified afresh against the current fix set and the question predicate rather than inheriting a stop from the finding it resembles. A decline keeps a finding out and never excuses one that is in: a declined finding the fix set @@ -346,7 +359,12 @@ byte-identical in both copies: stays in the session that made it — and a stuck or two-tell surface with its continue-or-stop answer. From a closing body a reader can infer the close, the acceptances and the declines; after a lost session neither the question decision nor the - continue-or-stop answer is recoverable from it. + continue-or-stop answer is recoverable from it. **A cycle that recovers its own identity but + not those answers re-surfaces their questions and takes fresh ones before another pass runs** + — the same answer the no-identity case gets and for the same reason, that an answer nobody + can read is an answer nobody gave. Running on instead would bypass a standing stop precisely + because no record carries it, which is the one way a suspension can be lost without anyone + deciding to lift it. **These records are one component of the closure-record contract** the one-contract rule above names, and a copy carrying them without the rest of that list is a partial adoption @@ -443,7 +461,10 @@ are in the site map and re-read at execution. edits, the answer records' membership in the named nonce set and the nonce exemption's criterion, the cycle-field count they falsify, the unknown-start item covering them, the closing-message carry, the squash carry, the curve's validity rule, the recovery sources, - the no-identity report and the gate-off routes they open** — and **the unknown-start activation semantics that say what a cycle owes + the no-identity report, the gate-off routes they open, the four pointer sentences that send + a reader to the ordering — in the floor paragraph, the absorb rule, the clearly-stuck + paragraph and the tells paragraph — and the pass-4 report's unavailable-history block with + its root condition** — and **the unknown-start activation semantics that say what a cycle owes when its starting rules cannot be established** depend on one another," and after "and a carry rule naming records a project does not produce is inert." (C:885) add: "an ordering without the answer records is a membership stop whose two answers nothing carries; answer @@ -585,7 +606,8 @@ to pad." Second, `a13`. OLD: "Every other rule stated here about how a cycle closes stands as written, and none of them is restated — a summary is where their conditions would get dropped." NEW: "**This paragraph** restates none of them — a summary is where their conditions would get dropped — and the closure ordering below is where they are -stated once and in order." Accounting: `a1`–`a12`, `a14`–`a17`, `a20`–`a22` kept; `a18`, `a19` +stated once and in order (part of the closure-record contract below; a copy carrying this +without the rest stops there)." Accounting: `a1`–`a12`, `a14`–`a17`, `a20`–`a22` kept; `a18`, `a19` moved; **`a13` replaced**. Old condition: *no* rule about how a cycle closes is restated anywhere. Kept: the prohibition and its reason, scoped to the paragraph it was written to police. Changed: the ordering does restate closure rules, deliberately and as the one @@ -600,7 +622,8 @@ recorded as Mechanics requires**, under the closure ordering above." `b17`–`b1 "Stopping this way is **not an exit from the gate**: the floor, the Blocker/Major filter and the clean-final-pass rule all stand, and the loop resumes on the revised artifact once the question is answered — what the stop prevents…" NEW: "Stopping this way is a **suspension** -under the closure ordering above, and the loop resumes on the revised artifact **once every +under the closure ordering above (part of the closure-record contract below; a copy carrying +this without the rest stops there), and the loop resumes on the revised artifact **once every answer that pass's suspensions require has been given** — what the stop prevents…". Third, in **W only**, "…by its severity exactly as the severity rule already says…" becomes "…exactly as Mechanics already says…", matching C for the reason §6's first row gives. @@ -625,7 +648,8 @@ nothing closes**" to the end of "…and then nothing could satisfy both.": repla trigger** — a finding outside the assigned fix set, or one opening a new structural or contract question — since a pass carrying either is not clean and the ordering above says why. **Below the floor nothing closes**, exactly as that ordering says. This exit is -a **suspension** under it: you surface with the findings still open, the resolve +a **suspension** under it (part of the closure-record contract below; a copy carrying this +without the rest stops there): you surface with the findings still open, the resolve rule is not waived by surfacing, and the loop resumes when every answer that pass's suspensions require has been given." The kept sentence is not touched; the qualification is adjacent to it, which is how **D3**'s verbatim requirement and the ordering can both hold — @@ -636,14 +660,23 @@ sentence beside it names the condition its example assumed and never stated, so `c10` carries are narrowed by adjacency rather than by edit; `c12` kept (pointer form); `c13` moved (the zero-finding exception); `c14` moved **to the ordering's continue branch** — the Minor-below-floor pass that "keeps looping" is precisely one that neither closes nor suspends, -and the branch states it rather than leaving it to be read out of a negation; `c15`, `c16`, -`c17` kept in the pointer sentence, `c17`'s "resolve rule" now reading with the -assigned-fix-set boundary §4 item 14 gives it, so the pointer needs no edit of its own; `c18` -**kept** — a clearly-stuck surface credits no pass as clean, and the ordering carries that -condition over its whole domain and no narrower, reaching it through the pass's **findings** -rather than through the act, for the reason the duties paragraph states. The two-tell -stop was never in its domain, surfacing tells and not a finding, and the ordering says so -rather than leaving it inferred; `c19` **replaced** — old: the loop resumes on +and the branch states it rather than leaving it to be read out of a negation; `c15`, `c17` +kept in the pointer sentence, `c17`'s "resolve rule" now reading with the assigned-fix-set +boundary §4 item 14 gives it, so the pointer needs no edit of its own; `c16` **replaced** — +old: a clearly-stuck surface leaves its finding open, which read alone gives this exit a hold. +Kept: that the findings are still open at the surface, and that the resolve duty is what keeps +them so. Changed: the **hold** belongs to the scope stop, the only stop whose question is about +a finding, while this exit asks continue or stop and holds nothing. Authority **D1**, which +makes all three exits suspensions and gives a suspension its own transitions; `c18` +**replaced** — old: no pass is credited as clean on a clearly-stuck surface, unconditionally. +Kept: the rule wherever it decides anything — a pass carrying a scope-stop trigger is unclean, +and this exit is reached only on a pass that did not close. Changed: where the exit's +regenerating findings are in-set and the ceiling demotes them below Major, the pass is clean at +effective severity and closes at or above the floor. Authority **D3**, which ranks clean +completion above this exit — under the old reading D3's own preserved sentence and `c18` would +decide that pass in opposite directions. The two-tell +stop was never in either condition's domain, surfacing tells and not a finding, and the +ordering says so rather than leaving it inferred; `c19` **replaced** — old: the loop resumes on whatever the user decides, unconditionally; new: it resumes on a resuming answer and a stop leaves the suspension standing, because an unconditional resume is the stop-with-no-transition path AC 4 forbids; `c20` **dropped** — it argued that reading the exit as "stop instead of @@ -654,7 +687,8 @@ as a rule instead. kept, unchanged. **(e) The five tells** (C:263–268 / W:467–472) — **add** one pointer sentence after `e10`: -"This stop is a **suspension** under the closure ordering above, and the loop resumes when +"This stop is a **suspension** under the closure ordering above (part of the closure-record +contract below; a copy carrying this without the rest stops there), and the loop resumes when every answer that pass's suspensions require has been given." `e1`–`e11` kept. **(f) The two rules above do not compete** (C:275–287 / W:473–483) — **unchanged**. "The two @@ -733,13 +767,16 @@ earlier pass had removed.": anything. **First the root**, which is a condition on the current pass being valid at all and not a stop of its own. The slots live in `.context/codex-reviews/` under the top-level directory of the checkout this cycle is running in — the same root the pass call is given as -`workingDirectory`. Establishing it fails in three observable ways, each with its own fix: the -working directory is **not inside a git repository** (run the pass from the checkout); the +`workingDirectory`. Establishing it fails in **four** observable ways, each with its own fix: +the working directory is **not inside a git repository** (run the pass from the checkout); the resolved top level **differs from the directory the call was given**, compared after both are **canonicalized**, so that a symlink or a trailing slash is not a mismatch (re-issue the call -with the resolved root); or **the query itself fails** — git unavailable, repository ownership -rejected, metadata unreadable — which is reported with the error it returned and not as a -missing repository, since the fix is to make git usable in that checkout rather than to move. +with the resolved root); **git is not there at all**, the command not existing (install it, or +put it on the PATH); or **git ran and failed**, reported with the message it returned and never +as a missing repository — that message is the discriminator, and prose cannot partition better +than the tool does; the common shape is a refused repository ownership, whose fix is to trust +the checkout. An error none of the four explains is a residual: report it as it came and stop, +rather than filing it under the nearest shape. A pass whose root cannot be established is an **INCOMPLETE pass** — the state this section already defines, already uncounted toward the floor and already excluded from the curve — so nothing new is ranked in the closure ordering. What no check reaches: a @@ -747,10 +784,10 @@ slot written under a different root **in the past** is **indistinguishable from slot**, because nothing records where a past call ran. Unobservable, not detected. **Then, per earlier pass, two questions.** Is the slot present and valid? And is the pass **known accepted** — meaning **this session validated it**? Present and valid is the ordinary -case. Known accepted but absent, or present and now failing -validation, is a real pass whose artifact is unusable: its number **stays counted** and its -series read `?`, which the curve grammar already admits for a **valid** pass whose counts -cannot be recovered, since a corrupted or missing record does not un-run a pass. **After a +case. **A known-accepted pass whose slot is now absent or invalid** is a real pass whose +artifact is unusable: its number **stays counted** and its series read `?`, which the curve +grammar already admits for a **valid** pass whose counts cannot be recovered, since a corrupted +or missing record does not un-run a pass. **After a lost session nothing is known accepted**, so every absent or present-but-invalid slot is **acceptance unknown**: **not counted toward the floor** and **omitted from the durable curve**, which takes one entry per *valid* pass and has no way to say "may not have been @@ -770,7 +807,9 @@ endpoints is simply unavailable. Refusing to compare across a gap would silence exactly where the record is thinnest, and the tells read the direction of the loop rather than any one adjacent pair. This is not a stop of its own, and it does not make the working record mandatory: a report that says what it could not see is the duty; one that invents the trend, -or omits a line without saying so, is the failure. Filled, over passes 1, 2 and 4 with 3 +or omits a line without saying so, is the failure. This report and its root condition are one +component of the closure-record contract (Mechanics); a copy carrying them without the rest of +that list is a partial adoption that stops there. Filled, over passes 1, 2 and 4 with 3 unreadable: Root: `.context/codex-reviews/` under this checkout — confirmed. (Unconfirmable: @@ -838,28 +877,38 @@ not: parent tree, and with it both labels: `Accepted: ` and `Declined: ` count **1** in each and **0** there, since one label shipping without the other is the shape the block's whole form exists to prevent; -- the six-member squash-carry sentence: `every answer record` count **1** in each; **0** in - the parent tree; +- the six-member squash-carry sentence: `every cycle's latest answer-record snapshot` count + **1** in each; **0** in the parent tree. The phrase asserted is the one §4 item 1's NEW text + actually contains — an earlier revision asserted `every answer record`, which appears nowhere + in it, so that check could not have passed against a conforming implementation and would have + read as a false red; - the Q6 lead phrase `**Where an earlier pass's findings file is unavailable**` count **1** in each; **0** in the parent tree; - the recovery-source edit (§4 item 11): `Recovery has three sources` count **1** in each and **0** in the parent tree, with `Recovery has two sources` at **0** in each and **1** there, since the passage is replaced and a NEW installed beside a surviving OLD is the failure; -- the five source edits, each as a pair, because each is a standing sentence changing meaning - rather than new text appearing: `A **clean findings file** is the single body line` at **1** - in each and **0** in the parent, with `A\n> clean pass is the single body line` at **0** and - **1**; `when a pass finds nothing` at **1** and **0**, with `when a pass is clean — the - explicit clean signal` at **0** and **1**; `for every finding in the assigned fix set` at - **1** and **0**; `Knowing that a pass ran is not that proof` at **1** and **0**, with - `may know a pass happened and not what it found` at **0** and **1**. A one-sided presence - check would pass on a copy carrying both wordings; -- the contract markers, since finding 7 turned on their absence: `part of the closure-record - contract` **case-insensitively** at **13** in each and **0** in the parent — one per source - edit in §4 except items 6 and 9, which are respectively unchanged and the contract itself, - and three of the thirteen open a sentence and so capitalise it — plus the three block-level - sentences (`This ordering is one component`, `These records are one component`, `This split - is one component`) at **1** each and **0** there. Sixteen markers in each copy; a - case-sensitive count would find ten and read as a failure. +- the **other four** source edits, each as a pair, because each is a standing sentence changing + meaning rather than new text appearing, and a one-sided presence check would pass on a copy + carrying both wordings. **Every fragment here is a single line in the file it is grepped + from**, each OLD verified at 1 in both copies of the parent tree, because a fragment spanning + a line break makes `grep -F` count 0 and read as a failure — the earlier revision quoted + three of them across their wraps: `A **clean findings file** is the single body line` at + **1** in each and **0** in the parent, with `clean pass is the single body line` at **0** and + **1**; `when a pass finds nothing` at **1** and **0**, with `when a pass is clean` at **0** + and **1**; `for every finding in the assigned fix set` at **1** and **0**, with `both must + resolve. Minor` at **0** and **1**; `Knowing that a pass ran is not that proof` at **1** and + **0**, with `may know a pass happened` at **0** and **1**. The NEW fragments must be installed + unwrapped at those points — a constraint on how the plan writes the edits, not on what they + mean; +- the contract markers, since a partial merge sees them and not the membership list: `part of + the closure-record contract` **case-insensitively** at **17** in each and **0** in the parent + — one per source edit in §4 except items 6 and 9, which are respectively unchanged and the + contract itself, plus one on each of the four pointer sentences §5(a), (b), (c) and (e) + install; several open a sentence and capitalise, which is why the count is case-insensitive + and why a case-sensitive one would under-count and read as a failure — plus the four + block-level sentences (`This ordering is one component`, `These records are one component`, + `This split is one component`, `This report and its root condition are one component`) at + **1** each and **0** there. **Twenty-one** markers in each copy. If the claim "the ordering and the records ship in both copies" were false, one working-tree count would be **0** or the parent-tree counts would not differ from it. The wiring can produce From 18c742f9e549b4df03ef7a7baf38ff2db40f3a8c Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Thu, 10 Sep 2026 18:59:40 +0200 Subject: [PATCH 012/181] =?UTF-8?q?docs(specs):=20loop-rule=20consolidatio?= =?UTF-8?q?n=20=E2=80=94=20Gate-A=20pass=209=20revision=20(slim=20the=20an?= =?UTF-8?q?swer=20record)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Pass 9 tripped the two-tell threshold and the loop stopped. Daniel's answer: slim the answer record back to the pass-4 choice. Deleted with every consequence, marker, assert and next-state row that served them: the full-snapshot rule and its `Cycle answers` marker, the kind/artifact header fields, the authoritative-latest-body and omission-is-loss rules, the newest-first branch search and the whole recovery-passage rewrite, the malformed-snapshot stop, the durable-validity language, and the identity-preserving re-surface duty. Five of pass 9's fourteen findings dissolve with them. Kept: two labels, one form, the cycle nonce, written before the next pass, restated in the closing body, carried on squash, the sameness test, the cycle binding. In their place one residual paragraph — the record is legible in history to a human reading it and buys no automatic recovery — which is what story criterion 7 asks it to state. Finding 3 reverses the pass-6 choice back to the frozen pass-time fix set, as D4 and b6 both require: a hold discharged by a scope change is a hold nobody answered. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- ...26-09-10-loop-rule-consolidation-design.md | 450 +++++++++--------- 1 file changed, 214 insertions(+), 236 deletions(-) diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index f46c490..07f14f0 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -24,8 +24,8 @@ cycle and which merely *suspend* it, in what order a pass is read so the ranking rather than asserted, what any set of suspensions at once does, and which of the four standing duties participate in that ordering versus gate it as preconditions. On top of it: one commit-body record form with two labels, `Accepted:` and `Declined:`, so a user's answer on a -surfaced finding has one rule in both directions and survives the session that made it; the -answer to what a severity demotion does to the loop-health counts; the answer to the pass-4 +surfaced finding has one rule in both directions and is legible in history to whoever reads it; +the answer to what a severity demotion does to the loop-health counts; the answer to the pass-4 report when prior-pass history is unavailable; and the deferred slot discriminator, dissolved rather than shipped. **What does not:** the pass floor and severity semantics, which the parent shipped and this spec reads as given, and everything §11 lists as parked. @@ -44,9 +44,9 @@ field and file set each predicate reads; the duties' classification; the scope s triggers and what each answer does; what a stuck or two-tell answer produces; the records' wording, attribution and recording point and the `Accepted:` label; and the raw-severity rule. Five standing sentences are edited at their source rather than worked around — the two that -use *clean* in the file's sense, the resolve duty that never stated its scope, the recovery -passage that would not have looked where the records live, and the curve's `?` rationale, -which answered the unavailable-history question the other way (§4 items 11–15). +use *clean* in the file's sense, the resolve duty that never stated its scope, the curve's `?` +rationale, which answered the unavailable-history question the other way, and the Gate-A +cadence, which made a revision unconditional (§4 items 11–15). --- @@ -77,8 +77,10 @@ reviewer wrote, so that a ceiling cannot make two findings the same one. **First, clean completion.** §5 uses *clean* in two senses and now says which is which. A **clean findings file** is the `NO FINDINGS` signal the protocol defines. A **clean pass** is -the predicate below: a clean findings file always gives one, and a clean pass need not have -one, because a pass carrying only Minors and Nits, or findings a decline of this cycle +the predicate below, read on the **logical pass with every required branch file combined**: +one branch's clean findings file never establishes a clean pass, the other being free to carry +an in-set Blocker. Every required file carrying that signal gives one, and a clean pass need +not have it, because a pass carrying only Minors and Nits, or findings a decline of this cycle matches, is clean without being empty. **A pass is clean** when its findings carry no Blocker or Major at effective severity that is **in the assigned fix set**, and **no scope-stop trigger**. There are two triggers, defined here and nowhere else, so that cleanliness and @@ -100,8 +102,9 @@ making the pass not final and costing the further pass that section requires. A tells present on the closing pass go into the closing report and never block it, because reporting "will not converge" on a converged loop is a false report, and the clearly-stuck paragraph says the same of its own exit. **No other -pass outcome closes a cycle**, because every other way out of a pass leaves a finding or a -question open, and closing over one is the failure this ordering exists to prevent. The one +pass outcome closes a cycle**, because every other pass either leaves a required repair, a +finding hold or a question outstanding, or has an unmet closure precondition such as the +floor — and closing over any of those is the failure this ordering exists to prevent. The one termination that is not a pass outcome is the Gate-B triviality skip, which ends a cycle having run no passes and is outside this ordering entirely. @@ -149,81 +152,76 @@ pass unclean directly, so nothing has to read the act of surfacing. **Where the other two exits stand in this duty**, said here so no reader supplies a rule for them. The two-tell stop surfaces tells and not a finding, and was never in its domain. The clearly-stuck exit does surface findings, and where those carry a scope-stop trigger the duty -reaches them exactly as it reaches any other. Where they do not — its regenerating findings -are in-set, reviewer-written Blocker or Major, and the ceiling demotes them below Major — the +reaches them exactly as it reaches any other. Where they do not — the two shapes the +composition paragraph names, a demoted in-set finding and a declined out-of-set one — the pass is clean at effective severity, the two predicates reading different fields on purpose, and **clean completion wins by D3**: the pass closes at or above the floor and continues below it. The order decides that; the duty never needed to. **What a suspension asks, and what ends it.** Only a scope stop raises a **hold**, on its -finding, and the hold ends when **every** answer that finding requires has been given — one for -a single-trigger finding, both the question decision and the membership answer for one carrying -both triggers — **and no direction is the wrong answer**, since a hold only accepting could end -would be the resolve duty under another name. At a membership stop the answer is **accept** (the finding joins the -fix set and resolves by its effective severity: Blocker or Major before any pass can be clean, -Minor or Nit collected and never iterated) or **decline** (the finding stays outside, binding -for the rest of this cycle). Either answer is **recorded** — Mechanics, the answer records — -and that record is where a later pass, or a cycle that lost its session, reads it. -**Membership is read when the answer is given, not frozen at the surface**: where the governing -artifacts have broadened the set so that the held finding is now inside it, that broadening -discharges the hold's **membership component** by itself — the finding is in-set, decline is -unavailable to it, and it resolves by its effective severity like any other in-set finding. -Where that same finding also opened a question, **the question component stands until its -decision is given**, because a broadening answers who owns the work and never what the question -asked. At a question +finding. **The hold ends when every answer that finding requires has been given, in whichever +direction each is given** — one answer for a single-trigger finding, both the question decision +and the membership answer for one carrying both. That is **one rule with two parts**, how many +answers and which way each may go, and neither part is a test the other has to pass: a hold +only accepting could end would be the resolve duty under another name. +At a membership stop the answer is **accept** (the finding joins the +fix set, where its effective severity governs it) or **decline** (the finding stays outside, +binding for the rest of this cycle). Either answer is **recorded** — Mechanics, the answer +records — and that record is where a later pass reads it. +**Membership is answered against the set as it stood at the pass that raised the question**, +not against the set as it is when the answer arrives, and an explicit attributable answer is +required whatever the governing artifacts do meanwhile — a hold discharged by a scope change +is a hold nobody answered. A later broadening is a new fact the **next** pass reads; it never +discharges a standing hold. At a question stop the answer is the user's decision on the question, and membership does not change: an -in-set finding then routes through its effective severity like any other — a Blocker or Major -resolves under that decision or is dismissed with its one-line why, a Minor or Nit is collected -and never iterated; an -out-of-set finding that opened the question is a membership stop as well and takes accept or -decline. **Decline is available only at a membership stop**, because that is the only stop +in-set finding then routes through its effective severity like any other, under that decision; +an out-of-set finding that opened the question is a membership stop as well and takes accept or +decline. **Effective severity routes an in-set finding in one place only** — the resolve duty +above, which says what a Blocker or Major owes and that a Minor or Nit is collected and never +iterated — and every branch here that puts a finding in the set hands it to that duty rather +than restating it. **Decline is available only at a membership stop**, because that is the only stop whose question is whether a finding belongs to the set, and a decline anywhere else would waive work the cycle owes. The stuck and two-tell readings raise no hold: each asks one question, **continue or stop**. Continue is that suspension's resuming answer, and where it is the last one outstanding the loop resumes on the artifact as revised and the fix set as the governing artifacts now assign it — where several plans or stories govern one cycle, the union of the scopes they assign — **plus the findings -this cycle accepted into it**, which its **latest** commit body carries (Mechanics, the answer -records). One that body does not carry is not in the set, and the reviewer raises it again on -the next pass like any other, which is the ordinary route and not a special one. A finding that a +this cycle accepted into it**, which its commit bodies record (Mechanics, the answer records). +One no body records is not in the set, and the reviewer raises it again on the next pass like +any other, which is the ordinary route and not a special one. A finding that a narrowing has put outside the set surfaces at the next pass as a membership stop. **Stop leaves the suspension standing**: the cycle stays open under its nonce, resumable by a later continue, nothing it wrote is a closing commit, and an open cycle nobody resumes is a human's to resolve, exactly as the nonce rules already say of open cycles. -**Composition, and what cannot happen.** Every question is answered on its own, and the loop -resumes only when every answer resumes it — accept or decline at a membership stop, a decision -at a question stop, continue at the stuck or two-tell reading; one stop answer leaves the +**Composition, and what cannot happen.** Every **question** is answered on its own, and the +loop resumes only when every answer resumes it — accept or decline at a membership stop, a +decision at a question stop, continue at the health readings; one stop answer leaves the whole suspension standing, because a loop resumed over an unanswered question decides it by -running. **Simultaneous health suspensions are one question, not two**: the stuck and two-tell -readings both ask continue or stop, so one answer carrying every reason ends both, and asking -twice would invite two answers to a question that has one. **Two pairings cannot occur**, and +running. **The stuck and two-tell readings raise one question between them, not two**, both +asking continue or stop, so one answer carrying every reason ends both — which is not an +exception to the sentence before it but an instance of it, there being one question there to +answer. **Two pairings cannot occur**, and no rule ranks them: clean completion and a **scope stop**, since that stop's triggers are the clean predicate's own second half, so a pass raising one is not clean; and a zero-finding pass and any suspension, since it has nothing to surface, nothing regenerating, no cluster and no require↔withdraw pair. **Clean completion and the clearly-stuck exit can**, and the overlap is -admitted rather than argued away: that exit reads reviewer-written severity while cleanliness -reads effective, so an in-set Blocker the ceiling demotes can regenerate across passes on a -pass that is clean. **The order decides it and no new rule is needed** — the pass closes at or +admitted rather than argued away. **Two shapes reach it**, since that exit reads +reviewer-written severity and membership while cleanliness reads effective severity and the +set: an **in-set** Blocker or Major the ceiling demotes below Major, and an **out-of-set** one +matched by a binding decline of this cycle, which raises no membership trigger and sits outside +the set the clean predicate reads. Either can regenerate across passes on a pass that is clean. +**The order decides both and no new rule is needed** — the pass closes at or above the floor, ranking the exit exactly as the clearly-stuck paragraph's own precedence sentence says, and continues below it, where nothing closes anyway. **This ordering is one component of the closure-record contract** the one-contract rule names, and a copy carrying it without the rest of that list is a partial adoption that stops there. ``` -Why the shape, where the block's sentences do not carry it. The evaluation order answers the -objection that a ranking between clean completion and the two-tell stop cannot fire if the stop -can make the pass unclean: clean completion is read first, so a suspension is only ever -evaluated on a pass that did not close — which works only because clean candidacy reads the -**triggers** and never the act of surfacing, as the block's own sentence says. -The three branches are AC 4, the duties paragraph AC 2, -the composition sentences AC 1 — which asks that a conflict unable to co-occur be named with -its reason rather than legislated, and is met on both sides: the scope stop cannot co-occur -with clean completion and says why, the clearly-stuck exit can and is resolved by the order -already stated rather than by a rule invented for it. The scope stop's two triggers are `b11` -(membership) and `b13` (question) read separately, because an in-set finding that opens a -question can neither join nor stay outside the set and needs its own answer — which is also why -a matching decline suppresses the membership trigger only. Sentences the block points at rather +Where the block maps onto the criteria: the three branches are AC 4, the duties paragraph +AC 2, the composition sentences AC 1. The scope stop's two triggers are `b11` (membership) and +`b13` (question) read separately, because an in-set finding that opens a question can neither +join nor stay outside the set and needs its own answer. Sentences the block points at rather than restating (`a13` as §5(a) replaces it): C:116–118, C:760–764, C:422–424, C:176, and the clearly-stuck precedence sentence, kept verbatim under **D3**. @@ -242,9 +240,15 @@ an in-set finding, which is the gate-off route a fabricated record would otherwi **The `Accepted:` label was beyond the story's *original* scope and is now authorised by it**: Daniel decided it at this cycle's Gate-A pass-4 scope stop, and the story records the expansion in its §2 and adds acceptance criterion 7 (at `4625679`). Three passes had found the same hole -— an acceptance put a finding in the fix set and no artifact carried it, so a lost session left -a cycle able to close over work it had agreed to do. The cheapest close was a second label on a -record whose form, transport, nonce and carry rules already existed. This cycle runs under the +— an acceptance put a finding in the fix set and no artifact carried it, so a cycle could close +over work it had agreed to do with nothing left saying so. The cheapest close was a second +label on a record whose form, transport, nonce and carry rules already existed. **What the +label closes and what it leaves open** is criterion 7's second half and the block states it: +the decision becomes legible to a human reading the history, and no more — a later revision of +this spec had it drive an automatic recovery procedure, which grew an identity line, a +snapshot rule, an authoritative-body rule and a branch search, produced Blockers on each of +three successive passes, and was cut back to this on Daniel's decision at the pass-9 two-tell +stop. This cycle runs under the pre-change rules and writes no such record for its own acceptance; the dispositions file and the working record are what today's rules provide. @@ -256,9 +260,9 @@ byte-identical in both copies: ```` **Recording an answer at a membership stop.** A membership stop asks whether a surfaced finding joins the assigned fix set, and the user's answer is recorded under one of two - labels sharing one form. **Accepted** puts the finding in the set, from where the severity - rules already govern it: a Blocker or Major owes resolution, a Minor or Nit is collected and - never iterated. **Declined** keeps it out and ends the membership half of its hold — the + labels sharing one form. **Accepted** puts the finding in the set, where the resolve duty + governs it — this block adds nothing to what that duty says. **Declined** keeps it out and + ends the membership half of its hold — the whole of it where membership was all that finding raised. Both are available at a membership stop and nowhere else: not for an in-set Blocker or Major, which owes resolution already, and neither is the answer to a question @@ -270,20 +274,16 @@ byte-identical in both copies: chose. ``` - Accepted: · · cycle · · + Accepted: · · cycle Finding: | | | | - Declined: · · cycle · · + Declined: · · cycle Finding: | | | | ``` The `Finding:` line is the finding line from the pass's findings file with its confidence field removed — the five fields the sameness test reads, in the file's order, a literal pipe - escaped as `\|` exactly as there. `` is one of the three cycle kinds and `` - the same key the working record uses — a Gate-A cycle's reviewed document path, a Gate-B - cycle's base commit as a full 40-character object name — because a record read back out of a - branch body must be attributable on its own, and an empty commit carrying only a record - supplies nothing else to attribute it by. + escaped as `\|` exactly as there. **Rules both labels share.** Each carries the **cycle nonce**, because a record that cannot be attributed to its cycle cannot bind to it. Each is **written when made**, into the cycle's @@ -293,27 +293,9 @@ byte-identical in both copies: may carry it meanwhile and does not bind. Where the cycle has no artifact revision to carry it — the answered finding was that pass's only one — the record goes in an **empty commit of its own**, the destination the human-exception rule above already blesses: "An empty commit - carrying only the record is a legitimate destination". Each is **restated in every later body - of that cycle**, the closing one included, so that the cycle's **latest** body carries the - complete answer set: **that body is the authoritative snapshot**, and an answer missing from - it is lost whatever an earlier body says, because a rule that let any reachable body revive - an answer would make the set depend on how far back a reader looked. **Every body of the - cycle carries that snapshot, and a cycle with no answers carries it too** — one line, - `Cycle answers: none · cycle · · `, in the same fields the record - header carries. Two things follow, and both are why it is every body and not only the ones - with something to say. A body stating an empty set is distinguishable from one that dropped - its records, which is the difference every reader of these bodies turns on. And **the newest - body is found by that line and never by whether it happens to carry an answer**: a body - identified only by the records in it is invisible on the day it drops them, which is the - day it matters. **A body of this cycle whose snapshot is missing or malformed stops** the - cycle rather than being skipped, since skipping it is exactly what revives the superseded - body behind it. Everything that reads these records reads **that snapshot and no - earlier body**: recovery, the squash carry, and the fix set the loop resumes with. Each is - **copied on squash-merge** (the carry rule above). A cycle that resumes and finds no answer - in that snapshot **treats it as absent** — the unknown-start fallback's - reading, and the safe direction under both labels: an unrecorded decline means the hold - applies again, an unrecorded acceptance means the finding is raised afresh. There is no - second place to look, which is the point of naming one body authoritative. Each is an + carrying only the record is a legitimate destination". Each is **restated in the cycle's + closing body**, so that the record a reader of history meets is not buried in an intermediate + commit, and each is **copied on squash-merge** (the carry rule above). Each is an **unverified assertion** of the same kind as the human exception — nothing checks that the handle belongs to whoever decided, that a human was asked, or that the reason is honest. @@ -325,10 +307,12 @@ byte-identical in both copies: bytes, since a reviewer rewrites its sentences between passes; the severity read is the one the reviewer wrote, per the ordering's field rule. **At most one effective label per cycle and five-field key.** Exactly equivalent same-label records collapse to one, which is what - makes restating the snapshot idempotent; two records disagreeing on that key — opposite - labels, or one label with disagreeing attribution — **stop**, because a finding both in the - set and kept out of it has no reading, and picking either would let a fabricated or replayed - record decide which. What a matching decline does to a re-raised finding is the closure + makes a restatement idempotent; two records disagreeing on that key — opposite labels, or one + label with disagreeing attribution — are **a question for the user**, answered and recorded + like any other answer, because a finding both in the set and kept out of it has no reading, + and picking either would let a fabricated or replayed record decide which. It is a question + and not a cycle state: the loop already knows how to stop on one and resume on its answer. + What a matching decline does to a re-raised finding is the closure ordering's membership trigger, which is defined there and not restated here. **Any difference — severity included — or any genuine uncertainty makes it a new finding**, classified afresh against the current fix set @@ -347,24 +331,18 @@ byte-identical in both copies: false record can do — one fully identified finding, one cycle, **as far as distinct nonces allow**: two cycles sharing or redrawing a nonce are indistinguishable to these records as to every other, so a replayed answer can bind to the wrong cycle, and a bound is not safety. - **What the acceptance record buys, and what it does not:** a cycle that lost its session and - **recovers its own identity** reads its latest body among its recovery sources and finds what - it had accepted. One that cannot recover it starts a new cycle, which **inherits nothing** — - it names the cycles it did not adopt, marks their answers unknown, and obtains its own. So a - replacement can still review and close the same artifact while an older one stays open; the - record makes that discoverable rather than invisible, which is less than preventing it. **What no - record carries:** a question stop's decision, where it changed no membership — the five - fields identify a finding and not a question, and that decision's durable form is the - artifact revision it produces, **where it produces one**; a decision that changes nothing - stays in the session that made it — and a stuck or two-tell surface with its - continue-or-stop answer. From a closing body a reader can infer the close, the acceptances - and the declines; after a lost session neither the question decision nor the - continue-or-stop answer is recoverable from it. **A cycle that recovers its own identity but - not those answers re-surfaces their questions and takes fresh ones before another pass runs** - — the same answer the no-identity case gets and for the same reason, that an answer nobody - can read is an answer nobody gave. Running on instead would bypass a standing stop precisely - because no record carries it, which is the one way a suspension can be lost without anyone - deciding to lift it. + **What a record buys, and what it does not.** It makes the decision **legible in history to + a human reading it**, and that is the whole of it. It does not make an agent's recovery + automatic: nothing checks that a later cycle looked, no procedure is obliged to search for + it, and a cycle that lost its session is not promised its own answers back. Nor does any + record carry a question stop's decision where that changed no membership — the five fields + identify a finding and never a question — or a stuck or two-tell surface with its + continue-or-stop answer; after a lost session neither is reconstructable at all. A cycle that + resumes **continues on what it can read**, and the reviewer re-raises whatever is still true + of the artifact, which is the ordinary route and not a special one. So a replacement cycle + can review and close the same artifact while an older one stays open. The record makes that + discoverable to whoever reads the history, which is less than preventing it and is what this + record is for. **These records are one component of the closure-record contract** the one-contract rule above names, and a copy carrying them without the rest of that list is a partial adoption @@ -372,13 +350,10 @@ byte-identical in both copies: ```` **Which rules the answer modifies** (story AC 3): the scope stop's resume sentence (`b12`), -and the hold and the clean-pass definition, both in the §3 block. Those three are the rules -whose **closure behaviour** the answer qualifies, and it qualifies no other — not the -Blocker/Major-resolve duty, not the floor, the tells or the stuck reading. The Severity -bullet's edit (item 14) is not a counter-example and is the distinction worth holding: it -states **the duty's own scope**, the assigned fix set, which **D5** always implied. The answer -moves a finding into or out of that set and never changes what the duty demands of what is in -it, so no rule there mentions a decline. A reader finding either label +and the hold and the clean-pass definition, both in the §3 block — and no other, not the +Blocker/Major-resolve duty, the floor, the tells or the stuck reading. The Severity bullet's +edit (item 13) is not a counter-example: it states **the duty's own scope**, which **D5** +always implied, and names no decline. A reader finding either label qualifying a *closure* rule outside those three has found a defect; a reader finding the records absent from any transport, attribution, activation or threat site listed next has found the opposite one. The two lists answer different questions: what the answer *changes*, @@ -390,15 +365,11 @@ are in the site map and re-read at execution. 1. **Squash carry** (C:892 / W:1076, `j1`). OLD: "…copy every evidence entry, every human-exception record, the provenance lines, the curves and any skipped cycle's skip record TOGETHER WITH THE SKIP REASON IT POINTS AT…". NEW: "…copy every evidence entry, - every human-exception record, **every cycle's latest answer-record snapshot — the complete - set as that cycle's newest body in the range states it, earlier restatements being - superseded rather than merged (part of the closure-record contract above; a copy carrying - this without the rest stops there)**, the provenance lines, the curves and any skipped - cycle's skip record TOGETHER WITH THE SKIP REASON IT POINTS AT…". Six members. Copying the - snapshot rather than every record in the range is what makes the carry agree with the - authoritative-body rule: a record restated in every body reaches the range many times, and - summing them would let a body that dropped an answer be overruled by an older one that - still carried it — the revival the snapshot rule exists to forbid. + every human-exception record, **every answer record (part of the closure-record contract + above; a copy carrying this without the rest stops there)**, the provenance lines, the + curves and any skipped cycle's skip record TOGETHER WITH THE SKIP REASON IT POINTS AT…". + Six members. A record restated in the closing body and copied here reaches the squash body + once, because exactly equivalent same-label records collapse under the sameness rule. 2. **The named nonce set** (C:385–387 / W:579–581). OLD: "**and that set is named rather than left open**: the provenance line, the per-pass curve (including a skip record standing in for one), the cycle's findings slots, and its advisory working record." NEW: "**and that set @@ -436,16 +407,15 @@ are in the site map and re-read at execution. 7. **The gate-off surface** (C:143–151 / W:350–358). OLD tail: "…silencing reminders; or not running a pass and reporting that it ran." NEW tail: "…silencing reminders; not running a pass and reporting that it ran; **recording a decline nobody made, or one on an in-set - finding; or dropping an acceptance the cycle owes from the body that would carry it (part + finding; or omitting an acceptance the cycle owes from the body that would carry it (part of the closure-record contract below; a copy carrying this without the rest stops there)**." The list says it is not complete; this change opens those routes and names them, as the - parent did for the stated floor. "Dropping from the body" and not "deleting", because under - the snapshot rule an omission is the whole of the act. + parent did for the stated floor. 8. **The closing message and the soft-reset path** (C:834–838 / W:1018–1022). Two insertions into a sentence whose other clauses are unchanged. After "…which owes no entry" add: "— - **and the cycle's complete answer-record snapshot, including any answer held only in WIP - bodies a `git reset --soft` collapsed**: the single commit after the reset carries the whole - set, because a body the reset discards is unreachable from the commit that replaces it + **and the cycle's answer records, including any held only in WIP bodies a + `git reset --soft` collapsed**: the single commit after the reset carries them all, because + a body the reset discards is unreachable from the commit that replaces it (part of the closure-record contract below; a copy carrying this without the rest stops there)". In the sentence after it, "so an entry written only into the WIP body" becomes "so an entry **or @@ -460,7 +430,7 @@ are in the site map and re-read at execution. raw-versus-effective split and its assigned-fix-set boundary, the two clean-vocabulary edits, the answer records' membership in the named nonce set and the nonce exemption's criterion, the cycle-field count they falsify, the unknown-start item covering them, the - closing-message carry, the squash carry, the curve's validity rule, the recovery sources, + closing-message carry, the squash carry, the curve's validity rule, the Gate-A cadence, the no-identity report, the gate-off routes they open, the four pointer sentences that send a reader to the ordering — in the floor paragraph, the absorb rule, the clearly-stuck paragraph and the tells paragraph — and the pass-4 report's unavailable-history block with @@ -472,8 +442,8 @@ are in the site map and re-read at execution. an ordering whose raw-versus-effective split has no counterpart in the severity rule, or a severity rule still calling that question unsettled and mandating a stop beside an ordering that decides it, is two answers to one question; and an answer record missing from the - squash carry, the nonce set or the recovery sources cannot survive a merge, be attributed, - or be found again." The stop sentence that follows is unchanged and now covers these states. + squash carry or the nonce set cannot survive a merge or be attributed." The stop sentence + that follows is unchanged and now covers these states. Then, before it: "**A project carrying any component of this contract owes all of them.** The pieces are separately mergeable and are not separately adoptable, so a copy holding one without the rest is an incomplete adoption and stops here." **Both halves of the guard ship, @@ -502,46 +472,14 @@ are in the site map and re-read at execution. of the closure-record contract above; a copy carrying this without the rest stops there)." It adds no record — the report names state the workspace already holds — and closes the one shape "a human's to resolve" cannot reach: an open cycle nobody is told about. -11. **The recovery sources** (C:411–417 / W:605–611) — the **whole passage** is replaced, not a - sentence inside it, because the sentence that must change ("no search there") is the same one - that makes the rest coherent. OLD: "**Recovery has two sources, and they answer different - questions.** The **working record** is the source while the cycle runs, and it is the one the - candidate rules above apply to — several files may be present and the run must decide which, - if any, is its own. **History is the source once the cycle's own commit exists**, and there is - no search there: the cycle is reading **its own commit body**, so kind and artifact are settled - by which commit is being read, and the nonce is taken from the provenance line and the curve, - which must agree. A Gate-A cycle mid-run has no such commit and therefore has only the working - record." NEW: "**Recovery has three sources, and they answer different questions.** The - **working record** is the source while the cycle runs, and it is the one the candidate rules - above apply to — several files may be present and the run must decide which, if any, is its - own. **The cycle's own commits on the current branch** are the source for the records those - bodies carry — its artifact revisions, and any empty commit carrying an answer record — - **searched newest first and no further back than the branch point**, each candidate validated - by cycle field, kind and artifact exactly as a working record is; a bounded search, because an - unbounded one would reach other cycles' commits. **A candidate is a cycle identity — the - cycle field, the kind and the artifact together — and not a body**, so one cycle's own - restatements across several bodies are one candidate and never trip the more-than-one rule; - **its newest body is the answer set and the search stops there**, since an older body is - superseded rather than merged. **History is the source once the cycle's own - closing commit exists**, and there is no search there: the cycle is reading **its own commit - body**, so kind and artifact are settled by which commit is being read, and the nonce is taken - from the provenance line and the curve, which must agree." The next sentence's "Recovering a - single candidate from **either**" becomes "from **any of the three**" — a one-word repair the - passage's arithmetic forces, and the failure rule after it ("No candidate, disagreeing sources, - or more than one candidate → no identity") is unchanged and already governs all three, which is - why the NEW does not restate it. Part of the closure-record contract above; a copy carrying - this without the rest stops there. Left alone, an answer record would exist and the - procedure meant to read it would never look, and story criterion 7 would fail on one unchanged - sentence. - -12. **The findings-file protocol's clean sentence** (C:329 / W:523), inside the gate-prompt +11. **The findings-file protocol's clean sentence** (C:329 / W:523), inside the gate-prompt template both gates paste. OLD: "A clean pass is the single body line `NO FINDINGS` with `END OF FINDINGS (0 total)`." NEW: "A **clean findings file** is the single body line `NO FINDINGS` with `END OF FINDINGS (0 total)`." It describes a file and always did; the word *pass* in it is what makes §3's predicate look like a redefinition instead of the other sense. Its condition — what a reviewer writes when it finds nothing — is unchanged. Part of the closure-record contract below; a copy carrying this without the rest stops there. -13. **The Gate-A clean-signal sentence** (C:565–566 / W:756–757; the two copies wrap it +12. **The Gate-A clean-signal sentence** (C:565–566 / W:756–757; the two copies wrap it differently, so the shared fragment is what is quoted). OLD: "…when a pass is clean — the explicit clean signal is what lets you exit the loop:". NEW: "…when a pass finds nothing — the explicit signal is what lets a pass be read as clean without inspecting it further:". @@ -552,37 +490,51 @@ are in the site map and re-read at execution. the rest stops there. The other **six** uses of "clean pass" in each copy (C:117, 728, 750, 761, 769, 827; W:324, 914, 936, 947, 955, 1011) are the closure sense the ordering defines and are correct as they stand — counted and checked, not assumed. -14. **The Severity bullet's resolve duty** (C:783–784 / W:969–970), which is the one place the +13. **The Severity bullet's resolve duty** (C:783–784 / W:969–970), which is the one place the duty is stated and the only one without a scope. OLD: "- **Severity:** Blocker (wrong/unsafe/breaks invariant) · Major (design flaw → rework) → both must resolve. Minor · Nit → collect, never iterate." NEW: "- **Severity:** Blocker (wrong/unsafe/breaks invariant) · Major (design flaw → rework) → both must resolve, **for every finding in the - assigned fix set**; one the user declined at a membership stop is outside that set, and owes - nothing **while the set still excludes it** (the closure ordering above; part of the - closure-record contract below, and a copy carrying this without the rest stops there). - Minor · Nit → collect, never iterate." Old condition: + assigned fix set** (the closure ordering above; part of the closure-record contract below, + and a copy carrying this without the rest stops there). Minor · Nit → collect, never + iterate." Old condition: every Blocker and Major resolves, unbounded. **Replaced**: the boundary **D5** always implied and no sentence carried. Without it a declined finding must stay outside the set - and still bars closure, which is a pass that can neither close nor suspend. The (c) + and still bars closure, which is a pass that can neither close nor suspend. **The edit names + the set and never a decline**, which is AC 3's second half: the answer moves a finding into + or out of the set, and this duty says only what it demands of what is in it — a mention of + the decline here would read as a direct waiver of Blocker/Major resolution rather than the + membership decision it is. The (c) pointer that calls this "the resolve rule" (`c17`) needs no edit: it names the rule, and the rule now carries its own scope. -15. **The curve's one-entry-per-valid-pass paragraph** (C:943–950 / W:1127–1134), which is +14. **The curve's one-entry-per-valid-pass paragraph** (C:943–950 / W:1127–1134), which is where `?` gets its rationale and where that rationale currently answers Q6 the other way. OLD, the clause: "so a resumed cycle may know a pass happened and not what it found, and - zero and unknown are different facts." NEW: "so a resumed cycle may hold durable proof that - a pass was **valid** and no longer hold what it found, and zero and unknown are different - facts. **Knowing that a pass ran is not that proof**: where validity cannot be established - — the slot gone or unreadable, and nothing recording that it was accepted — the pass is + zero and unknown are different facts." NEW: "so a cycle may have **validated a pass this + session** and no longer hold what it found, and zero and unknown are different facts. + **Knowing that a pass ran is not knowing it was valid**: where this session did not validate + it and the slot is gone or unreadable, the pass is **omitted from the pass specification** rather than entered with `?`, because this grammar takes one entry per *valid* pass and has no way to say "may not have been one", and the report names the numbers it omitted (part of the closure-record contract above; a copy carrying this without the rest stops there)." Old condition: a resumed cycle's knowledge that a pass happened is enough to keep its entry, with `?` for the counts. **Replaced**: - knowledge that it ran is separated from proof that it was valid, because §7 makes the - second unavailable after a lost session and the two answers cannot both govern one slot. - Kept: `?` itself, per series, for a pass whose validity is established and whose counts are - not. Without this edit the same missing slot both keeps an entry and is omitted. + knowledge that it ran is separated from validation, using **§7's own predicate — this + session validated it** — rather than a durability test no artifact satisfies, since this + change introduces no pass-validity record. Kept: `?` itself, per series, for a validated + pass whose counts are not recoverable. Without this edit the same missing slot both keeps an + entry and is omitted. + +15. **The Gate-A cadence** (C:573 / W:764), which makes a revision unconditional between + passes. OLD: "Each pass: validate, revise, re-run." NEW: "Each pass: validate, revise + **where the severity and scope rules require a repair**, re-run (part of the closure-record + contract below; a copy carrying this without the rest stops there)." Old condition: every + pass is followed by a revision before the next. Kept: the cadence and its order — validate + first, re-run last. **Replaced**: the revision is conditional, because the ordering's + continue branch reaches a below-floor pass whose only findings are Minors and Nits, which + are collected and never iterated, and an unconditional "revise" tells that pass to + manufacture the repair the severity rule forbids. The shorter "Copy every record into the squash body" sentence inside the human-exception block (C:1004–1007) is generic and already covers an answer record; it is not edited. The "records @@ -628,7 +580,12 @@ answer that pass's suspensions require has been given** — what the stop preven **W only**, "…by its severity exactly as the severity rule already says…" becomes "…exactly as Mechanics already says…", matching C for the reason §6's first row gives. -Accounting: `b1`–`b11`, `b13`–`b16` kept; `b12` **replaced** — old: an *immediate* resume on +Accounting: `b1`–`b11`, `b13`–`b16` kept — `b6` among them, and it is worth naming: it fixes +the assigned set **before the pass being answered**, and the ordering's membership answer is +read against that same set rather than against a later one, so the two agree and `b6` needs no +edit. An earlier revision of this spec had membership re-read when the answer arrived, which +`b6` and **D4** both forbid — **D4** requiring a user answer for the hold to end at all — and +the reversal is recorded here rather than left as a silent narrowing. `b12` **replaced** — old: an *immediate* resume on the membership answer alone. Kept: that the membership answer is what the stop asks for and that either direction ends it. Changed: the resume waits for every answer the pass's suspensions require, each recorded, because a loop resumed over an unanswered question decides @@ -662,7 +619,7 @@ zero-finding exception); `c14` moved **to the ordering's continue branch** — t Minor-below-floor pass that "keeps looping" is precisely one that neither closes nor suspends, and the branch states it rather than leaving it to be read out of a negation; `c15`, `c17` kept in the pointer sentence, `c17`'s "resolve rule" now reading with the assigned-fix-set -boundary §4 item 14 gives it, so the pointer needs no edit of its own; `c16` **replaced** — +boundary §4 item 13 gives it, so the pointer needs no edit of its own; `c16` **replaced** — old: a clearly-stuck surface leaves its finding open, which read alone gives this exit a hold. Kept: that the findings are still open at the surface, and that the resolve duty is what keeps them so. Changed: the **hold** belongs to the scope stop, the only stop whose question is about @@ -686,10 +643,20 @@ as a rule instead. **(d) From pass 4 onward** (C:255–261 / W:459–465) — **add** the Q6 sentence (§7). `d1`–`d7` kept, unchanged. -**(e) The five tells** (C:263–268 / W:467–472) — **add** one pointer sentence after `e10`: -"This stop is a **suspension** under the closure ordering above (part of the closure-record -contract below; a copy carrying this without the rest stops there), and the loop resumes when -every answer that pass's suspensions require has been given." `e1`–`e11` kept. +**(e) The five tells** (C:263–268 / W:467–472) — **one edit and one added sentence**. The edit +is `e7`, the shared fragment both copies carry (C:266 / W:470). OLD: "**Any two present makes +stop-and-surface mandatory, not discretionary**". NEW: "**Any two present makes stop-and-surface +mandatory, not discretionary, where the clean-completion branch did not close the pass**". Then +the pointer, after `e10`: "This stop is a **suspension** under the closure ordering above (part +of the closure-record contract below; a copy carrying this without the rest stops there), and +the loop resumes when every answer that pass's suspensions require has been given." + +Accounting: `e1`–`e6`, `e8`–`e11` kept; `e7` **qualified** — old: two tells make the stop +mandatory, unconditionally. Kept: the threshold, its mandatory force, and that the stuck reading +is not a precondition for it. Changed: it is read after clean completion, not beside it. +Authority **D2**, which ranks clean completion above this stop — unqualified, `e7` and the +ordering decide a clean two-tell pass at or above the floor in opposite directions, which is +AC 1's reachable conflict and the one thing this passage must not leave standing. **(f) The two rules above do not compete** (C:275–287 / W:473–483) — **unchanged**. "The two rules above" still names the absorb rule and the stuck reading; the block sits before both and @@ -767,19 +734,16 @@ earlier pass had removed.": anything. **First the root**, which is a condition on the current pass being valid at all and not a stop of its own. The slots live in `.context/codex-reviews/` under the top-level directory of the checkout this cycle is running in — the same root the pass call is given as -`workingDirectory`. Establishing it fails in **four** observable ways, each with its own fix: -the working directory is **not inside a git repository** (run the pass from the checkout); the -resolved top level **differs from the directory the call was given**, compared after both are -**canonicalized**, so that a symlink or a trailing slash is not a mismatch (re-issue the call -with the resolved root); **git is not there at all**, the command not existing (install it, or -put it on the PATH); or **git ran and failed**, reported with the message it returned and never -as a missing repository — that message is the discriminator, and prose cannot partition better -than the tool does; the common shape is a refused repository ownership, whose fix is to trust -the checkout. An error none of the four explains is a residual: report it as it came and stop, -rather than filing it under the nearest shape. +`workingDirectory`, the two compared after both are **canonicalized**, so that a symlink or a +trailing slash is not a mismatch. That is **one condition**: either the slots being read are +under that checkout or they are not. A pass whose root cannot be established is an **INCOMPLETE pass** — the state this section already defines, already uncounted toward the floor and already excluded -from the curve — so nothing new is ranked in the closure ordering. What no check reaches: a +from the curve — so nothing new is ranked in the closure ordering. The report carries **git's +own message verbatim** and files it under no shape of its own: that message is the +discriminator, and prose cannot partition git's failures better than git does — a partition +written here would send a caller to the wrong repair on every case it guessed wrong, which is +the whole of what such a list would add. What no check reaches: a slot written under a different root **in the past** is **indistinguishable from an absent slot**, because nothing records where a past call ran. Unobservable, not detected. **Then, per earlier pass, two questions.** Is the slot present and valid? And is the pass @@ -813,7 +777,7 @@ that list is a partial adoption that stops there. Filled, over passes 1, 2 and 4 unreadable: Root: `.context/codex-reviews/` under this checkout — confirmed. (Unconfirmable: - this pass is INCOMPLETE, root not established; re-run from the checkout.) + this pass is INCOMPLETE, root not established — git said: "".) Trend: findings 24, 17, —, 18; Blockers 5, 2, —, 1 — pass 3 omitted: slot absent, acceptance unknown, so it is neither counted toward the floor nor entered in the curve. 2→4 spans that gap: rising, weaker evidence. Cluster: product behaviour @@ -825,9 +789,12 @@ This is **D10**. Both historical lines are derivable from the mandated findings (the `fic2` record verified that), so unavailability is a property of the workspace and the answer is disclosure, not a new stop or a mandatory artifact — the root condition reuses the INCOMPLETE state rather than adding one, so **D10**'s "not a new stop condition" stands as -written. The partition is by what the agent can observe (`docs/prompt-standards.md` item 10): -presence and validity from the slot, acceptance from this session alone, and where acceptance -cannot be established the count moves in the direction that costs a pass. The example is +written. The per-pass partition is by what the agent can observe +(`docs/prompt-standards.md` item 10): presence and validity from the slot, acceptance from this +session alone, and where acceptance cannot be established the count moves in the direction that +costs a pass. The root is deliberately **not** partitioned — item 10 asks that a diagnostic +state name its cause and its fix, and here git's message is the cause, named verbatim, while +any prose list would be this spec guessing at causes it cannot observe. The example is there for item 4. --- @@ -854,7 +821,12 @@ drop notes (Task 19, line 970; Task 20, line 1047) and Task 23's second point (l --- -## 9. Verification — mode `battery+check+verification` +## 9. Verification + +**The mode is read from the story's header at execution**, never from here — the same rule the +spec's own header states, and the reason no value is named in this heading. What follows is +what each level of that mode obliges, so that whichever it carries has its evidence described: +the battery, the check, and the named verification of the risk path. **Battery.** The quality command in `AGENTS.md` § Commands runs once, green, at the Gate-B WIP commit (`check-version-bump.sh main` needs the committed bump, §10). @@ -877,38 +849,40 @@ not: parent tree, and with it both labels: `Accepted: ` and `Declined: ` count **1** in each and **0** there, since one label shipping without the other is the shape the block's whole form exists to prevent; -- the six-member squash-carry sentence: `every cycle's latest answer-record snapshot` count - **1** in each; **0** in the parent tree. The phrase asserted is the one §4 item 1's NEW text - actually contains — an earlier revision asserted `every answer record`, which appears nowhere - in it, so that check could not have passed against a conforming implementation and would have - read as a false red; +- the six-member squash-carry sentence: `every answer record (part of the closure-record` count + **1** in each; **0** in the parent tree — the phrase checked against §4 item 1's NEW text + rather than assumed to match it; - the Q6 lead phrase `**Where an earlier pass's findings file is unavailable**` count **1** in each; **0** in the parent tree; -- the recovery-source edit (§4 item 11): `Recovery has three sources` count **1** in each and - **0** in the parent tree, with `Recovery has two sources` at **0** in each and **1** there, - since the passage is replaced and a NEW installed beside a surviving OLD is the failure; -- the **other four** source edits, each as a pair, because each is a standing sentence changing +- the **five** source edits, each as a pair, because each is a standing sentence changing meaning rather than new text appearing, and a one-sided presence check would pass on a copy carrying both wordings. **Every fragment here is a single line in the file it is grepped from**, each OLD verified at 1 in both copies of the parent tree, because a fragment spanning - a line break makes `grep -F` count 0 and read as a failure — the earlier revision quoted + a line break makes `grep -F` count 0 and read as a failure — an earlier revision quoted three of them across their wraps: `A **clean findings file** is the single body line` at **1** in each and **0** in the parent, with `clean pass is the single body line` at **0** and **1**; `when a pass finds nothing` at **1** and **0**, with `when a pass is clean` at **0** and **1**; `for every finding in the assigned fix set` at **1** and **0**, with `both must - resolve. Minor` at **0** and **1**; `Knowing that a pass ran is not that proof` at **1** and - **0**, with `may know a pass happened` at **0** and **1**. The NEW fragments must be installed + resolve. Minor` at **0** and **1**; `Knowing that a pass ran is not knowing it was valid` at + **1** and **0**, with `may know a pass happened` at **0** and **1**; and + `revise **where the severity and scope rules require a repair**` at **1** and **0**, with + `Each pass: validate, revise, re-run` at **0** and **1**. The NEW fragments must be installed unwrapped at those points — a constraint on how the plan writes the edits, not on what they mean; - the contract markers, since a partial merge sees them and not the membership list: `part of the closure-record contract` **case-insensitively** at **17** in each and **0** in the parent — one per source edit in §4 except items 6 and 9, which are respectively unchanged and the - contract itself, plus one on each of the four pointer sentences §5(a), (b), (c) and (e) - install; several open a sentence and capitalise, which is why the count is case-insensitive - and why a case-sensitive one would under-count and read as a failure — plus the four - block-level sentences (`This ordering is one component`, `These records are one component`, - `This split is one component`, `This report and its root condition are one component`) at - **1** each and **0** there. **Twenty-one** markers in each copy. + contract itself (thirteen), plus one on each of the four pointer sentences §5(a), (b), (c) + and (e) install; several open a sentence and capitalise, which is why the count is + case-insensitive and why a case-sensitive one would under-count and read as a failure — plus + the four block-level sentences (`This ordering is one component`, `These records are one + component`, `This split is one component`, `This report and its root condition are one + component`) at **1** each and **0** there. **Twenty-one** markers in each copy. **Every one + of the twenty-one must be installed on a single line**, the same constraint the source-edit + fragments carry and for the same reason: this spec itself wraps the marker phrase at several + of the sites that quote it, and a wrapped marker in the shipped copy makes `grep -F` count it + as absent. The constraint is on how the plan writes the line, not on the sentence's meaning, + which is why the marker sentences are short enough to fit one. If the claim "the ordering and the records ship in both copies" were false, one working-tree count would be **0** or the parent-tree counts would not differ from it. The wiring can produce @@ -924,9 +898,10 @@ stop, a hold awaiting its answer, accept, decline, a stop answer, two or three s once, a below-floor clean pass, a zero-finding pass, the unknown-start fallback; **and the stateful transitions**: a declined finding re-raised matching on all five fields and re-raised with one changed, the fix set broadened to include a declined finding, **the fix set broadened -to include a finding whose hold is still awaiting its answer**, recovery after an accept -and after a stop with the session lost (no identity → new cycle) **both with the answer records -present in the branch's bodies and with them absent**, a `full` Gate-B pass with one branch +while a hold is still awaiting its answer** — whose next state is the hold still standing, the +row that tests the frozen reading — two records disagreeing on one five-field key, recovery +after an accept and after a stop with the session lost (no identity → new cycle), +a `full` Gate-B pass with one branch clean and the other carrying an in-set Blocker, a rollback with a cycle open under the new rules, and a copy adopting the ordering without the answer records, the reverse, or either without the severity split. **Columns** — the **record state** (which answer records the bodies @@ -954,15 +929,18 @@ Gate-B re-review and before the closing amend, as §5 requires. - **Invariant 11 — `docs/prompt-standards.md`, all 12 items.** Most at risk: item 6 (every constraint in the shipped blocks carries its reason in the same sentence — **including the - three that were exempted until pass 6**: nothing else closes a cycle *because every other way - out leaves a finding or a question open*; a zero-finding pass is clean whatever the floor - *because a floor buys further looks at an artifact that keeps yielding findings*; decline is - available only at a membership stop *because that is the only stop whose question is - whether a finding belongs to the set*. The exemption was wrong twice over: item 6 admits no + three that were exempted until pass 6**: no other pass outcome closes a cycle *because every + other pass either leaves a required repair, a finding hold or a question outstanding, or has + an unmet closure precondition such as the floor*; a zero-finding pass is clean whatever the + floor *because a floor buys further looks at an artifact that keeps yielding findings*; + decline is available only at a membership stop *because that is the only stop whose question + is whether a finding belongs to the set*. The exemption was wrong twice over: item 6 admits no "settled elsewhere" clause, and a scaffolded copy cannot reach the story the reasons were said to live in); item 8 (token-lean — the blocks replace closure - sentences rather than adding beside them); item 4 (the Q6 example); item 10 (the Q6 - partition and its root check); item 3 (the stop answer is a named, resumable state). + sentences rather than adding beside them); item 4 (the Q6 example); item 10 (the Q6 per-pass + partition, and the root condition, whose cause is git's own message rather than a partition + written here — item 10 asks for a cause and a fix, and a guessed cause is neither); + item 3 (the stop answer is a named, resumable state). - **Don't: "Never replace a decision procedure without accounting for its old conditions."** Satisfied by §5, against the committed inventory. - **Don't: "Never rename or delete a doc section without grepping for references first."** The From 6cbc174ae3b9fba96350705efba3b2b4968ade76 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Thu, 10 Sep 2026 19:25:46 +0200 Subject: [PATCH 013/181] docs(stories): split the record-durability subject out of this story Gate-A spec pass 10 tripped the two-tell threshold and satisfied all three clearly-stuck conditions: ten passes, Blocker+Major never below 8, findings regenerating from the previous round's repairs. Nine of the twenty findings belonged to one subject this story never set out to answer, whether a record survives a session, an amend, a squash, a rollback or a moved checkout. The closure ordering had converged. Daniel's decision, 2026-09-10: split. The record and its transport (decisions 9, 9b, 9c), the unavailable-history report (decision 10), the checkout-root condition, the rollback reading and the discriminator dissolution move to a successor story. Those decisions stay settled and are implemented there. Criterion 7 moves with them. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- ...026-08-29-loop-rule-consolidation-story.md | 38 +++++++++++-------- 1 file changed, 23 insertions(+), 15 deletions(-) diff --git a/docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md b/docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md index e2648f6..03731fd 100644 --- a/docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md +++ b/docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md @@ -47,14 +47,25 @@ which are preconditions; and what a user's answer on a surfaced finding does in - **The pass-counter anomaly**, the CodeRabbit plan-metadata contradiction, and the fixture-per-predicate question — all still parked. -**One expansion, authorised 2026-09-10 rather than absorbed.** Gate-A spec pass 4 raised, as a -scope stop, that an accepted repair obligation lives only in the running session: a cycle that -loses it can be replaced by one that closes the same artifact with the repair never made. The -reliance is older than this story — §5 already puts "repair obligations you already accepted in -earlier passes" in the fix set with nothing recording them — but this story is the one writing -the closure rules, so the gap became its question. **Daniel accepted it into scope**: the change -ships a second record label, `Accepted:`, sharing the decline record's form, transport, cycle -nonce and carry rules. Criterion 7 below is what that adds; nothing else in this section moves. +**One expansion, authorised 2026-09-10 and then split out again on the same day.** Gate-A spec +pass 4 raised, as a scope stop, that an accepted repair obligation lives only in the running +session. Daniel accepted it into scope, and the change grew a second record label, `Accepted:`, +beside the decline record. Six passes later the loop had not converged and the evidence said why: +of pass 10's twenty findings, nine belonged to **one subject this story never set out to +answer** — whether a record survives a session, a commit amend, a squash, a rollback or a moved +checkout. The closure ordering itself had converged, its remaining findings small. + +**Split 2026-09-10 on Daniel's decision. Deferred to +`docs/superpowers/stories/2026-09-10-record-durability-story.md`**, whose subject is exactly that +one: the `Accepted:`/`Declined:` record and its transport (settled decisions 9, 9b and 9c), the +unavailable-history report (settled decision 10), the checkout-root condition that report grew, +the rollback reading, and the slot-discriminator dissolution. **Those decisions stay settled** — +they are not reopened, they are implemented there. Criterion 7 moves with them. + +**What this story keeps** is what its problem statement asked for: the ordering, the duty +classification, and the severity/loop-health answer handed to it by the parent. The rule covering +both directions of a user's answer stays here, because it is a rule about the loop; only the +record that transports it moves. ## 3. Acceptance criteria @@ -83,13 +94,10 @@ nonce and carry rules. Criterion 7 below is what that adds; nothing else in this Don't. A requirement neither kept nor explicitly dropped is a dropped condition. - [ ] **The two copies stay in parity** on every rule this story changes, deliberate wording differences stated as such. -- [ ] **A user's answer that puts work into the fix set leaves a record, on the same terms as one - that keeps work out.** Both labels share one form, one transport, one attribution rule and - one carry rule, each stated once; recording an answer is a precondition to running the next - pass; and what the record does *not* buy is stated — it makes a lost obligation - recoverable, not automatically recovered, and nothing checks that a replacement cycle - looked. Checkable by reading: two labels with one set of rules, and a residual paragraph - that does not overclaim. *(Added 2026-09-10 with the expansion recorded in §2.)* +*(Criterion 7 — that a user's answer putting work into the fix set leaves a record on the same +terms as one keeping work out — was added 2026-09-10 and moved the same day to +`docs/superpowers/stories/2026-09-10-record-durability-story.md` with the split recorded in §2. +It is not withdrawn, and it is not this story's to satisfy.)* ## 4. Settled inputs — decided, paid for, and not to be reopened From f3cb98addb6c5c9298dc2a8ae7f019e9752ef892 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Thu, 10 Sep 2026 19:28:38 +0200 Subject: [PATCH 014/181] docs(stories): record durability across sessions, commits and rollbacks Successor to the loop-rule consolidation story, created by Daniel's 2026-09-10 split. Ten Gate-A spec passes on that story's design never reached a clean pass, and nine of pass 10's twenty findings belonged to one subject the story never set out to answer: whether a record survives a lost session, a WIP amend, a soft reset, a squash, a rollback of the rules, or a checkout root that is not the one the pass ran in. Carries settled decisions 9, 9b, 9c and 10 across unreopened, plus the parent's measured lesson that a stated residual beats a built mechanism. Profile proposed, not confirmed: not executable until a human confirms it. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../2026-09-10-record-durability-story.md | 161 ++++++++++++++++++ 1 file changed, 161 insertions(+) create mode 100644 docs/superpowers/stories/2026-09-10-record-durability-story.md diff --git a/docs/superpowers/stories/2026-09-10-record-durability-story.md b/docs/superpowers/stories/2026-09-10-record-durability-story.md new file mode 100644 index 0000000..548e679 --- /dev/null +++ b/docs/superpowers/stories/2026-09-10-record-durability-story.md @@ -0,0 +1,161 @@ +# §5 record durability: sessions, commits, squashes and rollbacks — Story + +**Date:** 2026-09-10 · **Size:** story +**Risk:** *(proposed)* high · **Security:** *(proposed)* none · **Validation:** *(proposed)* battery+check+verification + +> **DRAFT — the profile above is proposed, not confirmed.** Per §5 a profile is confirmed by the +> human, and until it is this story is not executable. Nothing depends on it yet: the work it +> describes was split out of a cycle that is still running, and the successor starts when someone +> picks it up. The proposal's reasons are in §5. + +## 1. Problem statement + +**One question runs under every finding this story inherits: does a record survive?** §5 asks an +agent to write decisions into commit bodies and then to act on them a pass, a session or a merge +later. It never says what happens to a record across the six boundaries it will actually meet: a +**lost session**, an ordinary **`WIP:` amend**, a **`git reset --soft`** collapse of several WIP +snapshots into one, a **squash merge**, a **rollback** of the very rules that define the record, +and a **checkout whose root is not the one the pass ran in**. + +**The reliance is older than any of the records.** §5 already puts "repair obligations you already +accepted in earlier passes" inside the assigned fix set, and nothing anywhere holds them: they +live in the running session and vanish with it. Every record this story inherits inherits that +problem too. + +**The evidence is a cycle that could not converge while carrying this subject.** The +`Accepted:`/`Declined:` record was authorised into +`docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md` on 2026-09-10 and split +back out the same day. Ten Gate-A spec passes ran on that story's design under cycle nonce +`awsf1ec771`. Blocker+Major went **20, 15, 10, 11, 15, 13, 9, 8, 12, 17** and never reached zero; +findings went **24, 17, 12, 18, 17, 16, 14, 12, 14, 20**. + +**What the last pass's distribution showed is why this is its own story.** Of pass 10's twenty +findings, **nine belonged to this subject** — four on whether the record survives commits, three +on the checkout-root condition the unavailable-history report had grown, two on rollback and on +legacy cycles that hold no nonce. The closure ordering's own five findings were small. One +artifact was carrying two subjects, and only one of them was converging. + +**Where the pass-by-pass evidence lives:** `.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md`, +with the per-pass findings files beside it. That directory is gitignored, so the figures above are +quoted here rather than only cited — a reader on another machine has no way to open them. + +**Parent:** `docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md`, whose §2 +records the split and whose §4 holds the decisions this story implements. + +## 2. Desired outcome + +**One coherent answer to durability, covering every inherited decision, with each mechanism's +residual stated rather than mechanised.** A reader of either §5 copy can answer, without +judgement: what a record is worth after a session ends; which commit carries it through an amend, +a soft reset and a squash; what a cycle reads when the rules it started under are gone; and what +an unavailable prior-pass record does to the report and to the floor. + +**Prefer a stated residual to a built mechanism.** This is the parent's most expensive lesson and +it is an input here, not a discovery to be repeated — see §4. + +**Out of scope**, named so nothing absorbs them: +- **The closure ordering, the duty classification and the severity/loop-health answer.** The + parent story ships all three; this story treats them as given. +- **Reopening any decision in §4.** They are settled and paid for; the design starts from them. +- **Hook code** (anything under `plugins/dev-workflow/hooks/`). +- **The pass-counter anomaly**, the CodeRabbit plan-metadata contradiction, and the + fixture-per-predicate question — all still parked, as they were for the parent. + +## 3. Acceptance criteria + +- [ ] **A user's answer that puts work into the fix set leaves a record on the same terms as one + that keeps work out.** Both labels share one form, one transport, one attribution rule and + one carry rule, **each stated once**; recording an answer is a precondition to running the + next pass; and what the record does *not* buy is stated — it makes a decision legible, not + automatically recovered, and nothing checks that a later cycle looked. Falsifiable by + reading: two labels with one set of rules, and a residual paragraph that does not overclaim. + *(Carried from the parent's criterion 7, which moved here with the split.)* +- [ ] **A record made under a cycle survives that cycle's own git housekeeping**, and both copies + say how. Checkable by walking the three operations §5 already prescribes — a `WIP:` amend, a + `git reset --soft` onto the parent of the first WIP, and the closing `--amend` — and asking + of each which body holds the record afterwards. An operation whose answer is "none" is a + defect, not an accepted cost. +- [ ] **The squash-carry rule names every record this story ships**, and the enumeration matches + what the shipped text actually produces. The parent's members — evidence entries, human + exception records, provenance lines, curves, skip records with their reasons — are + **extended, never replaced**: a rewrite that drops one silently unships it. +- [ ] **A cycle whose starting rules cannot be established has a stated reading for every record + this story ships**, and that reading is the conservative one. Checkable against the + unknown-start fallback's existing list, which this story extends rather than rewrites. +- [ ] **An unavailable prior-pass record has one stated effect on the report and one on the + floor**, and they do not disagree. A reader can determine, for a slot that is absent, present + but invalid, or present under a root that cannot be established, whether the pass counts + toward the floor, what its curve entry is, and what the report must disclose. +- [ ] **The checkout root is answered one way and stated once** — a pass-validity condition or a + disclosure, not both and not a third thing in a third place. The parent reworked it three + times across passes 4, 8 and 9 and pass 10 still found it self-contradictory, which is what + this criterion exists to prevent. +- [ ] **Every condition of the replaced prose is accounted for** — for each rewritten passage, what + it required, each requirement marked kept, moved or deliberately dropped, per the AGENTS.md + Don't. A requirement neither kept nor explicitly dropped is a dropped condition. +- [ ] **The two copies stay in parity** on every rule this story changes, deliberate wording + differences stated as such. + +## 4. Settled inputs — inherited, not reopened + +Four decisions come across from `docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md` +§4, confirmed by Daniel during the parent cycle and **implemented here rather than re-decided**. +Their substance is reproduced so a reader need not hold both files open; the parent remains the +place they were settled. + +| # | Decision, as inherited | +|---|---| +| 9 | **The answer is recorded in the commit body**, reusing the human-exception transport as a **distinct record type** — the two differ in force, since the human-exception form authorizes nothing. | +| 9b | **The record stores exactly what the sameness test reads: location, defect, severity, consequence and suggested fix.** A test reading fields the record lacks is a wiring failure. Sameness requires all five to match; **any difference — including severity — makes it a new finding and the hold applies**, as does any genuine uncertainty. | +| 9c | **The record reads as an unverified assertion**, like the human-exception record beside it: nothing checks that the handle belongs to whoever decided. **Narrowness bounds what a false record can do — one fully-identified finding, one cycle — and that is not the same as making it safe.** | +| 10 | **When prior-pass history is unavailable**, the pass report states what is computable, names what is not and why, and discloses the reduced sensitivity. **Not a new stop condition**, not a mandatory resume note. | + +**A fifth input, from the parent's ten passes rather than from a decision: prefer a stated +residual to a built mechanism.** The parent grew a snapshot rule, an identity line, an +authoritative-body rule, a newest-first branch search, a malformed-snapshot stop and a +durable-validity proof, all to make records recover automatically. Pass 9 deleted roughly ninety +lines of that machinery on Daniel's decision, and pass 10 still returned Majors at **15**, the +pass-1 figure, with Blocker+Major up from 8 to 17. **So the machinery was not the whole cost — +but it was the part that regenerated**, each round's repair producing the next round's findings. +A design here that answers a durability question by building recovery, rather than by saying what +survives and what does not, is repeating a measured failure. + +**Two implementation facts the parent cycle established, carried so they are not rediscovered:** a +pass's cleanliness is a fact about what that pass found and is **never rewritten**; and the +findings files establish the **inventory** of findings, not their resolutions, which they do not +contain. + +**Passages shared with the parent.** The squash-carry rule and the unknown-start fallback are +rewritten by both stories. **This story extends what the parent leaves behind rather than +replacing it** — the parent's accounting covers its own half, and a replacement drops it silently. + +## 5. Open questions + +- **The profile.** Proposed `high / none / battery+check+verification`, on the same reasoning its + sibling used: the surface is the review gate itself, and a wrong rule mis-steers every future + cycle. **No named `high` trigger matches literally**, so this is a judgement call under intake's + "surfaces, not words", and the human decides it. The parent's experience is evidence for rather + than against: ten passes, two mandatory two-tell stops and a split. +- **Is the checkout root a pass-validity condition or a disclosure?** Settled decision 10 says + unavailable history is "not a new stop condition", and the parent's pass 10 read the root gate as + contradicting it. Answering it a validity condition means a pass can end INCOMPLETE for a reason + unrelated to its findings; answering it a disclosure means a wrong checkout can be reported as + benign missing history. The parent tried both and shipped neither cleanly. +- **Does the sameness test read the five fields byte-exact or on meaning?** This **interprets** + decision 9b rather than reopening it: 9b settles that any difference makes a new finding and is + silent on what counts as a difference. The parent read it as meaning, because a reviewer + rewrites its sentences between passes and byte equality would re-ask every declined finding; + pass 10 read it as bytes, citing 9b's own wording. Both readings satisfy the decision as + written, which is why a human answers it — it is the difference between a record that binds too + little and one that binds too much. +- **Does a rollback that removes these rules owe an open cycle any transition?** §5's activation + rule says a cycle finishes under the rules it started with, and nothing records which rules + those were. A rollback can therefore leave a cycle open with no text describing what it owes. + Whether that is a state to define or a residual to disclose is a requirement question. + +## 6. Suggested size + +`story` — one coherent subject, durability of records across boundaries, in two mirrored copies. +Not `chore`: it answers four open requirement questions and rewrites standing carry rules. Not +`epic-needs-splitting`: the boundaries are six named ones, and the decisions behind the records +are already made. From c8f96b8e3166e7419b6d198a169544b06d69de27 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Thu, 10 Sep 2026 19:36:31 +0200 Subject: [PATCH 015/181] docs(specs): narrow to the closure ordering; defer record durability Gate-A spec pass 10 tripped the two-tell threshold and satisfied all three clearly-stuck conditions: ten passes, Blocker+Major never below 8, findings regenerating from the previous round's own repairs. Nine of the twenty findings belonged to one subject the story never set out to answer -- whether a record survives a session, an amend, a squash, a rollback or a moved checkout -- while the ordering's own five were small. Daniel's decision, 2026-09-10: split. Removed from this spec and moved to docs/superpowers/stories/2026-09-10-record-durability-story.md: the Accepted/Declined record and its whole transport, the pass-4 report's unavailable-history block, the checkout-root condition, the rollback reading, the slot-discriminator dissolution, and the closure-record contract with its twenty-one per-hunk markers. Settled decisions 9, 9b, 9c and 10 are not reopened; they are implemented there. What remains is the ordering, the four-duty classification, the severity/loop-health answer, and six standing sentences the ordering cannot be adopted without. Twelve of the twenty findings are deferred with the material they are about; eight are fixed. 967 -> 647 lines. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- ...26-09-10-loop-rule-consolidation-design.md | 906 ++++++------------ 1 file changed, 293 insertions(+), 613 deletions(-) diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index 07f14f0..c299b87 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -15,6 +15,9 @@ file: 135 conditions (22 + 18 + 20 + 7 + 11 + 7 + 4 + 26 + 16 + 4) quoted from ` reader can check the accounting rather than take it. A sentence outside those passages is cited by its lead phrase and C line. +**Narrowed 2026-09-10: §9 lists what moved out and where.** Read it before reading anything +here as missing. + --- ## 1. Intent @@ -22,31 +25,36 @@ cited by its lead phrase and C line. **What ships.** One closure ordering, stated once in each copy, that says which exit *closes* a cycle and which merely *suspend* it, in what order a pass is read so the ranking is executable rather than asserted, what any set of suspensions at once does, and which of the four standing -duties participate in that ordering versus gate it as preconditions. On top of it: one -commit-body record form with two labels, `Accepted:` and `Declined:`, so a user's answer on a -surfaced finding has one rule in both directions and is legible in history to whoever reads it; -the answer to what a severity demotion does to the loop-health counts; the answer to the pass-4 -report when prior-pass history is unavailable; and the deferred slot discriminator, dissolved -rather than shipped. **What does not:** the pass floor and severity semantics, which the parent -shipped and this spec reads as given, and everything §11 lists as parked. +duties participate in that ordering versus gate it as preconditions. With it: the answer to what +a severity demotion does to the loop-health counts, and the six standing sentences the ordering +falsifies or leaves ambiguous if they are not edited at their source. + +**What does not:** the pass floor and severity semantics, which the parent shipped and this spec +reads as given, and everything §9 lists as moved or parked. + +**Why the split, since a reader of the ordering will look for the record.** Gate-A spec pass 10 +tripped the two-tell threshold and satisfied all three conditions of the clearly-stuck reading. +Nine of that pass's twenty findings belonged to **one subject this story never set out to +answer** — whether a record survives a session, a commit amend, a squash, a rollback or a moved +checkout — while the ordering's own five were small. Daniel split that subject out on +2026-09-10; §9 says what went and where. --- ## 2. Settled inputs -The story's §4 table, decisions 1–10, is the design's starting point and is not restated here; -each is cited below as **D1**…**D10** (with **D9b**, **D9c**). Two implementation facts the -parent cycle established are read as given: a pass's cleanliness is a fact about what that pass -found and is **never rewritten** — an answer changes whether the *cycle* may close; and the -findings files establish the **inventory** of findings, not their resolutions. What the table -does not settle, this spec decides in the section that uses it: the evaluation order and the -field and file set each predicate reads; the duties' classification; the scope stop's two -triggers and what each answer does; what a stuck or two-tell answer produces; the records' -wording, attribution and recording point and the `Accepted:` label; and the raw-severity rule. -Five standing sentences are edited at their source rather than worked around — the two that -use *clean* in the file's sense, the resolve duty that never stated its scope, the curve's `?` -rationale, which answered the unavailable-history question the other way, and the Gate-A -cadence, which made a revision unconditional (§4 items 11–15). +The story's §4 table is the design's starting point and is not restated here; each decision is +cited below as **D1**…**D8**. **D9**, **D9b**, **D9c** and **D10** stay settled and move to the +successor with the material they govern (§9). Two implementation facts the parent cycle +established are read as given: a pass's cleanliness is a fact about what that pass found and is +**never rewritten** — an answer changes whether the *cycle* may close; and the findings files +establish the **inventory** of findings, not their resolutions. + +What the table does not settle, this spec decides in the section that uses it: the evaluation +order and the field and file set each predicate reads; the duties' classification; the scope +stop's two triggers and what each answer does; what a stuck or two-tell answer produces; and the +raw-severity rule for the health measures. Six standing sentences are edited at their source +rather than worked around, each one the ordering falsifies or leaves ambiguous (§4). --- @@ -59,39 +67,41 @@ triggers and point at it. Verbatim as it will ship: ``` **How a cycle ends — one ordering, stated here and referenced everywhere else.** A pass is read in a fixed order, because every rule below bears on one decision — may this cycle close — -and a stated order is what stops them qualifying each other. Every predicate here reads the -validated findings file **or files** of the logical pass as their **concatenation** — a `full` -Gate-B pass has two, and one branch alone is already an incomplete pass — at **effective** -severity, after the Mechanics severity ceiling, because cleanliness is about what the cycle -must repair and the ceiling is what decides that. **Both branches' lines count** for the health -measures, the curve rule already summing the branches into one entry; a finding whose five -fields match one in the other branch is **one** finding for holds and answers, so it is -answered once. Two things read the reviewer-written field from before the -ceiling: the **health measures** (Mechanics, Severity) — the per-pass counts, the clusters, -the tells, and **both severity-bearing conditions of the three-condition stuck reading**, its -Blocker curve and its regenerating Blocker or Major findings, since a predicate reading one -field for half of itself could not be read at all, while its third condition, the stated -coverage-sufficiency judgement, reads no severity field and the ceiling does not touch it — -and the **five-field key** the answer records match on, whose severity field is the one the -reviewer wrote, so that a ceiling cannot make two findings the same one. +and a stated order is what stops them qualifying each other. **Every finding-derived predicate +here** reads the validated findings file **or files** of the logical pass as their +**concatenation** — a `full` Gate-B pass has two, and one branch alone is already an incomplete +pass — at **effective** severity, after the Mechanics severity ceiling, because cleanliness is +about what the cycle must repair and the ceiling is what decides that. **Both branches' lines +count** for the health measures, the curve rule already summing the branches into one entry; a +finding matching one in the other branch is **one** finding for holds and answers, so it is +answered once. The **health measures** (Mechanics, Severity) read the reviewer-written field +from before the ceiling — the per-pass counts, the clusters, the tells, and **both +severity-bearing conditions of the three-condition stuck reading**, its Blocker curve and its +regenerating Blocker or Major findings, since a predicate reading one field for half of itself +could not be read at all, while its third condition, the stated coverage-sufficiency judgement, +reads no severity field and the ceiling does not touch it. **The branches read more than the +findings**, named here so nobody applies the order to the findings file alone: closure reads the +**derived floor**, the **cited set and every profile** as re-read for final acceptance, and any +**hold still standing**; the membership trigger reads the **current assigned fix set** and the +answers this cycle has already given; the stuck reading adds its coverage judgement. **First, clean completion.** §5 uses *clean* in two senses and now says which is which. A **clean findings file** is the `NO FINDINGS` signal the protocol defines. A **clean pass** is the predicate below, read on the **logical pass with every required branch file combined**: one branch's clean findings file never establishes a clean pass, the other being free to carry an in-set Blocker. Every required file carrying that signal gives one, and a clean pass need -not have it, because a pass carrying only Minors and Nits, or findings a decline of this cycle -matches, is clean without being empty. **A pass is clean** when its findings carry no Blocker +not have it, because a pass carrying only Minors and Nits, or findings this cycle has declined, +is clean without being empty. **A pass is clean** when its findings carry no Blocker or Major at effective severity that is **in the assigned fix set**, and **no scope-stop -trigger**. There are two triggers, defined here and nowhere else, so that cleanliness and -suspension read the same words: a **membership trigger** is a finding outside the current -assigned fix set **and not matched by a binding decline of this cycle** — such a decline having -put it outside by the user's own answer, for as long as the current set still excludes it — and -a **question trigger** is a finding opening a new structural or contract question, in-set or -not, which a decline never suppresses, the five fields it records identifying a finding and -never a question. Both are properties of the findings and the set, read before any branch below -runs, which is what makes this order executable rather than asserted: no predicate here waits -on an act a lower branch performs. A pass +trigger**. There are two triggers, and this branch is where cleanliness and suspension both read +them, in the same words so they cannot drift: a **membership trigger** is a finding outside the +current assigned fix set **and not one this cycle has already declined** — such a decline having +put it outside by the user's own answer, for as long as the current set still excludes it, so +that an answered question is not asked again — and a **question trigger** is a finding opening a +new structural or contract question, in-set or not, which a decline never suppresses, an answer +about membership being no answer to a question. Both are properties of the findings and the set, +read before any branch below runs, which is what makes this order executable rather than +asserted: no predicate here waits on an act a lower branch performs. A pass with **zero** findings is clean whatever the floor, because a floor buys further looks at an artifact that keeps yielding findings, and one yielding none has already given what those looks were for. **A clean pass at or above the derived floor, @@ -116,16 +126,18 @@ defined above — a **membership stop** by the first, a **question stop** by the Blocker/Major filter and the clean-final-pass rule stand while it does. Any non-empty set of them can apply to one pass: **one surface, every reason reported, every question asked**, because a reason left out is a decision made by omission. A finding surfaced solely by the -stuck or two-tell reading is not a scope-stop finding; one that also carries either trigger -takes the scope stop's answers at that same surface — it is not asked twice. This branch -defines nothing: the triggers are the first branch's, and what a changed field does to a -re-raised finding is Mechanics', in the answer records. A reader who finds a rule here that is -not in one of those two has found a defect. +clearly-stuck reading is not a scope-stop finding; one that also carries either trigger takes +the scope stop's answers at that same surface — it is not asked twice. The two-tell stop +surfaces tells and not a finding, so no finding is surfaced by it at all. This branch defines +nothing: the triggers are the first branch's, and a reader who finds a rule here that is not +there has found a defect. **Third, a pass that neither closes nor suspends continues** — the loop runs another pass on -the **current** artifact, revised or not. Below the floor a clean pass lands here, and so does -a pass whose only findings are Minors and Nits, which are collected and never iterated and so -may leave nothing to revise. It is a branch and not an inference, because "does not close" +the **current** artifact, revised or not. A below-floor clean pass lands here **only where no +suspension applies to it**; where one does, the second branch has already taken it, because +clean completion did not close the pass and only closing outranks a suspension. So does a pass +whose only findings are Minors and Nits, which are collected and never iterated and so may +leave nothing to revise. It is a branch and not an inference, because "does not close" read alone says nothing about whether to run again. **The four standing duties, classified.** The **derived floor** is a precondition on @@ -166,29 +178,30 @@ answers and which way each may go, and neither part is a test the other has to p only accepting could end would be the resolve duty under another name. At a membership stop the answer is **accept** (the finding joins the fix set, where its effective severity governs it) or **decline** (the finding stays outside, -binding for the rest of this cycle). Either answer is **recorded** — Mechanics, the answer -records — and that record is where a later pass reads it. +binding for the rest of this cycle). Either is an **explicit, attributable decision on that +specific finding** — never silence, never a general remark about scope, never inferred, because +a fix set changed by inference is a fix set nobody chose. **Membership is answered against the set as it stood at the pass that raised the question**, -not against the set as it is when the answer arrives, and an explicit attributable answer is +not against the set as it is when the answer arrives, and an explicit answer is required whatever the governing artifacts do meanwhile — a hold discharged by a scope change is a hold nobody answered. A later broadening is a new fact the **next** pass reads; it never -discharges a standing hold. At a question +discharges a standing hold. A decline keeps a finding out and never excuses one that is in: a +declined finding the fix set later comes to include owes resolution like any other. At a question stop the answer is the user's decision on the question, and membership does not change: an in-set finding then routes through its effective severity like any other, under that decision; an out-of-set finding that opened the question is a membership stop as well and takes accept or -decline. **Effective severity routes an in-set finding in one place only** — the resolve duty -above, which says what a Blocker or Major owes and that a Minor or Nit is collected and never -iterated — and every branch here that puts a finding in the set hands it to that duty rather -than restating it. **Decline is available only at a membership stop**, because that is the only stop +decline. **Effective severity routes an in-set finding in one place only** — the Mechanics +severity rule, where a Blocker or Major's obligation to resolve and a Minor or Nit's collection +without iteration are both stated, and whose discharge the resolve duty above defines — and +every branch here that puts a finding in the set hands it there rather than restating it. +**Decline is available only at a membership stop**, because that is the only stop whose question is whether a finding belongs to the set, and a decline anywhere else would waive work the cycle owes. The stuck and two-tell readings raise no hold: each asks one question, **continue or stop**. Continue is that suspension's resuming answer, and where it is the last one outstanding the loop resumes on the artifact as revised and the fix set as the governing artifacts now assign it — where several plans or stories govern one cycle, the union of the scopes they assign — **plus the findings -this cycle accepted into it**, which its commit bodies record (Mechanics, the answer records). -One no body records is not in the set, and the reviewer raises it again on the next pass like -any other, which is the ordinary route and not a special one. A finding that a +this cycle accepted into it**. A finding that a narrowing has put outside the set surfaces at the next pass as a membership stop. **Stop leaves the suspension standing**: the cycle stays open under its nonce, resumable by a later continue, nothing it wrote is a closing commit, and an open cycle nobody resumes is a human's @@ -209,344 +222,110 @@ require↔withdraw pair. **Clean completion and the clearly-stuck exit can**, an admitted rather than argued away. **Two shapes reach it**, since that exit reads reviewer-written severity and membership while cleanliness reads effective severity and the set: an **in-set** Blocker or Major the ceiling demotes below Major, and an **out-of-set** one -matched by a binding decline of this cycle, which raises no membership trigger and sits outside +this cycle has declined, which raises no membership trigger and sits outside the set the clean predicate reads. Either can regenerate across passes on a pass that is clean. **The order decides both and no new rule is needed** — the pass closes at or above the floor, ranking the exit exactly as the clearly-stuck paragraph's own precedence -sentence says, and continues below it, where nothing closes anyway. **This ordering is one component of the -closure-record contract** the one-contract rule names, and a copy carrying it without the -rest of that list is a partial adoption that stops there. +sentence says, and continues below it, where nothing closes anyway. ``` Where the block maps onto the criteria: the three branches are AC 4, the duties paragraph AC 2, the composition sentences AC 1. The scope stop's two triggers are `b11` (membership) and `b13` (question) read separately, because an in-set finding that opens a question can neither join nor stay outside the set and needs its own answer. Sentences the block points at rather -than restating (`a13` as §5(a) replaces it): C:116–118, C:760–764, C:422–424, C:176, and the +than restating (`a13` as §5(a) replaces it): C:116–118, C:760–764, C:176, and the clearly-stuck precedence sentence, kept verbatim under **D3**. +**Two things the ordering names and does not define, both the successor's** (§9): what makes a +later finding *the same one* this cycle declined, and what form carries an answer into the +commit body. The ordering says what an answer does — a decline keeps a finding out of the set +for the cycle, an acceptance puts one in until the cycle closes — and **D9b** and **D9** say how +it is recognised and written down. Until the successor ships, an answer binds within the session +that made it and the reviewer re-raises what it cannot read, which is the ordinary route and not +a special one. + --- -## 4. The answer records — one form, two labels - -**The rule ships as the block below**, which is the text and not a summary of it; this section -adds only what the block does not carry. Decisions implemented: **D6** availability, **D7** -binding, **D8** the explicit decision, **D9b** the sameness test, **D9c** the unverified -reading, **D9** the transport — whose "closing commit body" is therefore the *last* body to -carry a record, not the first. **D5** is why §3's clean predicate names the fix set rather than -the decline: a decline is only ever a fact about membership, so it cannot be read as excusing -an in-set finding, which is the gate-off route a fabricated record would otherwise open. - -**The `Accepted:` label was beyond the story's *original* scope and is now authorised by it**: -Daniel decided it at this cycle's Gate-A pass-4 scope stop, and the story records the expansion -in its §2 and adds acceptance criterion 7 (at `4625679`). Three passes had found the same hole -— an acceptance put a finding in the fix set and no artifact carried it, so a cycle could close -over work it had agreed to do with nothing left saying so. The cheapest close was a second -label on a record whose form, transport, nonce and carry rules already existed. **What the -label closes and what it leaves open** is criterion 7's second half and the block states it: -the decision becomes legible to a human reading the history, and no more — a later revision of -this spec had it drive an automatic recovery procedure, which grew an identity line, a -snapshot rule, an authoritative-body rule and a branch search, produced Blockers on each of -three successive passes, and was cut back to this on Daniel's decision at the pass-9 two-tell -stop. This cycle runs under the -pre-change rules and writes no such record for its own acceptance; the dispositions file and -the working record are what today's rules provide. - -**The block that ships**, in Mechanics, placed **immediately after the whole human-exception -passage** — after "**What the record is worth.**" (C:1028–1033 / W:1212–1217) and before -"- **Timeout / abort:**" (C:1034 / W:1218), at the same two-space indent. Verbatim, -byte-identical in both copies: - -```` - **Recording an answer at a membership stop.** A membership stop asks whether a surfaced - finding joins the assigned fix set, and the user's answer is recorded under one of two - labels sharing one form. **Accepted** puts the finding in the set, where the resolve duty - governs it — this block adds nothing to what that duty says. **Declined** keeps it out and - ends the membership half of its hold — the - whole of it where membership was all that finding raised. Both are available at a - membership stop and nowhere else: not for an in-set - Blocker or Major, which owes resolution already, and neither is the answer to a question - stop, a stuck or two-tell surface, a below-floor pass, an unclean final pass, or any Gate-A, - Gate-B or evidence obligation — of the list the human-exception form is never the answer to, - the membership stop is the one item these records answer. Each must be an **explicit, - attributable decision on that specific finding** — never silence, never a general remark - about scope, never inferred — because a fix set changed by inference is a fix set nobody - chose. - - ``` - Accepted: · · cycle - Finding: | | | | - - Declined: · · cycle - Finding: | | | | - ``` - - The `Finding:` line is the finding line from the pass's findings file with its confidence - field removed — the five fields the sameness test reads, in the file's order, a literal pipe - escaped as `\|` exactly as there. - - **Rules both labels share.** Each carries the **cycle nonce**, because a record that cannot - be attributed to its cycle cannot bind to it. Each is **written when made**, into the cycle's - next commit body on the branch — a spec or plan revision commit for a Gate-A cycle, the - `WIP:` amend for Gate B — **and that commit is made before the next pass runs**: an answer - held only in a session is one compaction away from being lost; the advisory working record - may carry it meanwhile and does not bind. Where the cycle has no artifact revision to carry - it — the answered finding was that pass's only one — the record goes in an **empty commit of - its own**, the destination the human-exception rule above already blesses: "An empty commit - carrying only the record is a legitimate destination". Each is **restated in the cycle's - closing body**, so that the record a reader of history meets is not buried in an intermediate - commit, and each is **copied on squash-merge** (the carry rule above). Each is an - **unverified assertion** of the same kind as the human exception — nothing checks that the - handle belongs to whoever decided, that a human was asked, or that the reason is honest. - - **Binding, and the sameness test.** A decline **binds for the remainder of its cycle**, with - no effect in any later one, and never qualifies the Blocker/Major-resolve duty — which the - declined finding does not reach **while it stays outside the set**, the duty being scoped to - what is in it. A later pass raises **the same finding** when all five of - location, defect, severity, consequence and suggested fix match, read on meaning rather than - bytes, since a reviewer rewrites its sentences between passes; the severity read is the one - the reviewer wrote, per the ordering's field rule. **At most one effective label per cycle - and five-field key.** Exactly equivalent same-label records collapse to one, which is what - makes a restatement idempotent; two records disagreeing on that key — opposite labels, or one - label with disagreeing attribution — are **a question for the user**, answered and recorded - like any other answer, because a finding both in the set and kept out of it has no reading, - and picking either would let a fabricated or replayed record decide which. It is a question - and not a cycle state: the loop already knows how to stop on one and resume on its answer. - What a matching decline does to a re-raised finding is the closure - ordering's membership trigger, which is defined there and not restated here. **Any - difference — severity included — or any - genuine uncertainty makes it a new finding**, classified afresh against the current fix set - and the question predicate rather than inheriting a stop from the finding it resembles. A - decline keeps a finding out and never excuses one that is in: a declined finding the fix set - later comes to include owes resolution like any other. **The binding and the set are - different things**, which is how both hold at once: the decline binds the *decision* for the - rest of the cycle, so that question is never re-asked, while membership is owned by the - governing artifacts — a broadening puts the finding in the set without the decline having - expired, and every exclusion this record grants reads "while the current set still excludes - it". An acceptance does not expire with - its pass: the finding is in the set until the cycle closes. - - **What each is worth.** A decline **releases a hold** and an acceptance **puts a finding in - the fix set**, neither of which the human-exception form ever does. Narrowness bounds what a - false record can do — one fully identified finding, one cycle, **as far as distinct nonces - allow**: two cycles sharing or redrawing a nonce are indistinguishable to these records as - to every other, so a replayed answer can bind to the wrong cycle, and a bound is not safety. - **What a record buys, and what it does not.** It makes the decision **legible in history to - a human reading it**, and that is the whole of it. It does not make an agent's recovery - automatic: nothing checks that a later cycle looked, no procedure is obliged to search for - it, and a cycle that lost its session is not promised its own answers back. Nor does any - record carry a question stop's decision where that changed no membership — the five fields - identify a finding and never a question — or a stuck or two-tell surface with its - continue-or-stop answer; after a lost session neither is reconstructable at all. A cycle that - resumes **continues on what it can read**, and the reviewer re-raises whatever is still true - of the artifact, which is the ordinary route and not a special one. So a replacement cycle - can review and close the same artifact while an older one stays open. The record makes that - discoverable to whoever reads the history, which is less than preventing it and is what this - record is for. - - **These records are one component of the closure-record contract** the one-contract rule - above names, and a copy carrying them without the rest of that list is a partial adoption - that stops there. -```` - -**Which rules the answer modifies** (story AC 3): the scope stop's resume sentence (`b12`), -and the hold and the clean-pass definition, both in the §3 block — and no other, not the -Blocker/Major-resolve duty, the floor, the tells or the stuck reading. The Severity bullet's -edit (item 13) is not a counter-example: it states **the duty's own scope**, which **D5** -always implied, and names no decline. A reader finding either label -qualifying a *closure* rule outside those three has found a defect; a reader finding the -records absent from any transport, attribution, activation or threat site listed next has -found the opposite one. The two lists answer different questions: what the answer *changes*, -and what must *carry* it. - -**Existing sentences that must name them**, both copies, old → new. Line numbers are C's; W's -are in the site map and re-read at execution. - -1. **Squash carry** (C:892 / W:1076, `j1`). OLD: "…copy every evidence entry, every - human-exception record, the provenance lines, the curves and any skipped cycle's skip - record TOGETHER WITH THE SKIP REASON IT POINTS AT…". NEW: "…copy every evidence entry, - every human-exception record, **every answer record (part of the closure-record contract - above; a copy carrying this without the rest stops there)**, the provenance lines, the - curves and any skipped cycle's skip record TOGETHER WITH THE SKIP REASON IT POINTS AT…". - Six members. A record restated in the closing body and copied here reaches the squash body - once, because exactly equivalent same-label records collapse under the sameness rule. -2. **The named nonce set** (C:385–387 / W:579–581). OLD: "**and that set is named rather than - left open**: the provenance line, the per-pass curve (including a skip record standing in - for one), the cycle's findings slots, and its advisory working record." NEW: "**and that set - is named rather than left open**: the provenance line, the per-pass curve (including a skip - record standing in for one), **any answer record (Mechanics; part of the closure-record - contract below, and a copy carrying this without the rest stops there)**, the cycle's - findings slots, and its advisory working record." -3. **The nonce exemption** (C:397–399 / W:591–593). OLD: "The nonce is not required in records - this change neither introduces nor keys to a cycle — the evidence entry and a - human-exception record among them." NEW: "The nonce is not required in records that are not - keyed to a cycle — the evidence entry and a human-exception record among them; **an answer - record is keyed to its cycle and carries it (part of the closure-record contract below; a - copy carrying this without the rest stops there)**." The old "this change" dated the - sentence to the parent; the new one states the criterion. -4. **"Both shipped records below"** (C:367 / W:561). OLD: "Both shipped records below carry a - **cycle field**, because…". NEW: "Both shipped records below carry a **cycle field** — and - so do the answer records in Mechanics, part of the closure-record contract below, a copy - carrying this without the rest stopping there — because…". A load-bearing count a third - cycle-attributed record would otherwise falsify. -5. **The unknown-start fallback** (C:153–167 / W:360–374, `i4`–`i8`, extended at `i12`'s - invitation). OLD: "…at minimum floor 3, severity classified without the demotion, the - provenance-line duty owed, the curve duty owed, and the nonce duties at their strictest…". - NEW: "…at minimum floor 3, severity classified without the demotion, the provenance-line - duty owed, the curve duty owed, the nonce duties at their strictest, **every suspension - binding, and answer records not attributable to the nonce this fallback minted treated as - absent, so that no inherited hold is released and no inherited acceptance is claimed — - records made and recorded under that nonce are the cycle's own and are honoured (part of the - closure-record contract below; a copy carrying this without the rest stops there)**…". - `i12`'s sentence stays as written; this is the addition it invites. The time bound matters: - treating *every* answer as absent would leave a fallback cycle unable to release a hold it - raised itself, **D7** unmet. -6. **The human-exception "which commit" rule** (C:988–990 / W:1172–1174, `h3`–`h6`). Unchanged. - The answer records carry their own, deliberately different, because an answer must survive - to the next pass while the human exception only has to survive to history. -7. **The gate-off surface** (C:143–151 / W:350–358). OLD tail: "…silencing reminders; or not - running a pass and reporting that it ran." NEW tail: "…silencing reminders; not running a - pass and reporting that it ran; **recording a decline nobody made, or one on an in-set - finding; or omitting an acceptance the cycle owes from the body that would carry it (part - of the closure-record contract below; a copy carrying this without the rest stops there)**." - The list says it is not complete; this change opens those routes and names them, as the - parent did for the stated floor. -8. **The closing message and the soft-reset path** (C:834–838 / W:1018–1022). Two insertions - into a sentence whose other clauses are unchanged. After "…which owes no entry" add: "— - **and the cycle's answer records, including any held only in WIP bodies a - `git reset --soft` collapsed**: the single commit after the reset carries them all, because - a body the reset discards is unreachable from the commit that replaces it - (part of the closure-record contract below; a copy carrying this without the rest stops - there)". In the - sentence after it, "so an entry written only into the WIP body" becomes "so an entry **or - record** written only into the WIP body". The `git reset --soft` sentence itself - (C:829–830 / W:1013–1014) is unchanged. -9. **The one-contract paragraph** (C:879–890 / W:1063–1074), which is where the contract gets - its **one name and one membership list**, so that each mergeable piece cites the name - instead of naming every peer. OLD opening: "…this carry rule - **and the unknown-start activation semantics that say what a cycle owes when its starting - rules cannot be established** depend on one another," NEW opening: "…this carry rule, **the - closure-record contract — the closure ordering, the answer records, the severity rule's - raw-versus-effective split and its assigned-fix-set boundary, the two clean-vocabulary - edits, the answer records' membership in the named nonce set and the nonce exemption's - criterion, the cycle-field count they falsify, the unknown-start item covering them, the - closing-message carry, the squash carry, the curve's validity rule, the Gate-A cadence, - the no-identity report, the gate-off routes they open, the four pointer sentences that send - a reader to the ordering — in the floor paragraph, the absorb rule, the clearly-stuck - paragraph and the tells paragraph — and the pass-4 report's unavailable-history block with - its root condition** — and **the unknown-start activation semantics that say what a cycle owes - when its starting rules cannot be established** depend on one another," and after "and a - carry rule naming records a project does not produce is inert." (C:885) add: "an ordering - without the answer records is a membership stop whose two answers nothing carries; answer - records without the ordering are a release and a set change with no stop that asks for them; - an ordering whose raw-versus-effective split has no counterpart in the severity rule, or a - severity rule still calling that question unsettled and mandating a stop beside an ordering - that decides it, is two answers to one question; and an answer record missing from the - squash carry or the nonce set cannot survive a merge or be attributed." The stop sentence - that follows is unchanged and now covers these states. - Then, before it: "**A project carrying any component of this contract owes all of them.** - The pieces are separately mergeable and are not separately adoptable, so a copy holding one - without the rest is an incomplete adoption and stops here." **Both halves of the guard ship, - and that reverses this cycle's earlier choice.** Pass 6 offered a central list or a marker - on every mergeable hunk; the list was taken alone, and pass 7 showed why that is not enough - — a list cannot police a merge that omits the list. So every hunk in this section also - carries a short marker naming the contract, and the three blocks keep the longer sentence - they already had. The list is what defines membership; the markers are what a partial merge - still sees. **Rollback.** A cycle open when the text is - reverted is governed by "a cycle already running finishes under the rules it started with" - and "A revert is itself a shipping commit for the old rules" (C:153–154, C:165–166) where it - can still establish those rules, and by the unknown-start fallback where it cannot. **Where - the rollback removes the fallback too, neither is available and nothing here replaces - them**: no record identifies the rule revision a cycle started under, so that cycle has no - text describing what it owes. A residual, named and left to a human — not a stop this change - can claim, since a stop needs shipped text and the rollback removed it. -10. **The no-identity rule's aftermath** (C:422–424 / W:616–618). OLD: "**Starting a new cycle - does not close, adopt or retire the cycles those candidates belong to** — they stay open, - keep their own nonces, and are a human's to resolve; the new cycle simply does not claim - them." NEW: "**Starting a new cycle does not close, adopt or retire the cycles those - candidates belong to** — they stay open, keep their own nonces, and are a human's to - resolve; the new cycle simply does not claim them, **and names them in its first pass - report, marking their exit and their answers unknown**, so the human this rule makes - responsible learns they exist and nobody reads the new cycle as continuing them. **It - inherits no answer**: whatever those cycles accepted or declined, this one asks again (part - of the closure-record contract above; a copy carrying this without the rest stops there)." - It adds no record — the report names state the workspace already holds — and closes the one - shape "a human's to resolve" cannot reach: an open cycle nobody is told about. -11. **The findings-file protocol's clean sentence** (C:329 / W:523), inside the gate-prompt +## 4. The standing sentences edited at their source + +Six, and no more — working around any of them would ship two instructions that disagree. Line +numbers are C's; W's are re-read at execution. **Every OLD fragment quoted here is a single line +in the file it is grepped from**, verified at 1 in both copies (§7). + +1. **The findings-file protocol's clean sentence** (C:329 / W:523), inside the gate-prompt template both gates paste. OLD: "A clean pass is the single body line `NO FINDINGS` with `END OF FINDINGS (0 total)`." NEW: "A **clean findings file** is the single body line `NO FINDINGS` with `END OF FINDINGS (0 total)`." It describes a file and always did; the word *pass* in it is what makes §3's predicate look like a redefinition instead of the - other sense. Its condition — what a reviewer writes when it finds nothing — is unchanged. - Part of the closure-record contract below; a copy carrying this without the rest stops there. -12. **The Gate-A clean-signal sentence** (C:565–566 / W:756–757; the two copies wrap it + other sense. Old condition — what a reviewer writes when it finds nothing — unchanged. + +2. **The Gate-A clean-signal sentence** (C:565–566 / W:756–757; the two copies wrap it differently, so the shared fragment is what is quoted). OLD: "…when a pass is clean — the explicit clean signal is what lets you exit the loop:". NEW: "…when a pass finds nothing — the explicit signal is what lets a pass be read as clean without inspecting it further:". Old condition: a `NO FINDINGS` file is what permits loop exit. Kept: the signal and why it - is demanded. Changed: it makes a pass readable as clean rather than being the only way to - be clean, since a pass carrying Minors alone is clean under the ordering and could never - produce this file. Part of the closure-record contract below; a copy carrying this without - the rest stops there. The other **six** uses of "clean pass" in each copy (C:117, 728, 750, + is demanded. **Replaced**: it makes a pass readable as clean rather than being the only way + to be clean, since a pass carrying Minors alone is clean under the ordering and could never + produce this file. The other **six** uses of "clean pass" in each copy (C:117, 728, 750, 761, 769, 827; W:324, 914, 936, 947, 955, 1011) are the closure sense the ordering defines and are correct as they stand — counted and checked, not assumed. -13. **The Severity bullet's resolve duty** (C:783–784 / W:969–970), which is the one place the + +3. **The Severity bullet's resolve duty** (C:783–784 / W:969–970), which is the one place the duty is stated and the only one without a scope. OLD: "- **Severity:** Blocker (wrong/unsafe/breaks invariant) · Major (design flaw → rework) → both must resolve. Minor · Nit → collect, never iterate." NEW: "- **Severity:** Blocker (wrong/unsafe/breaks invariant) · Major (design flaw → rework) → both must resolve, **for every finding in the - assigned fix set** (the closure ordering above; part of the closure-record contract below, - and a copy carrying this without the rest stops there). Minor · Nit → collect, never - iterate." Old condition: - every Blocker and Major resolves, unbounded. **Replaced**: the boundary **D5** always - implied and no sentence carried. Without it a declined finding must stay outside the set - and still bars closure, which is a pass that can neither close nor suspend. **The edit names - the set and never a decline**, which is AC 3's second half: the answer moves a finding into - or out of the set, and this duty says only what it demands of what is in it — a mention of - the decline here would read as a direct waiver of Blocker/Major resolution rather than the - membership decision it is. The (c) - pointer that calls this "the resolve rule" (`c17`) needs no edit: it names the rule, and - the rule now carries its own scope. - -14. **The curve's one-entry-per-valid-pass paragraph** (C:943–950 / W:1127–1134), which is - where `?` gets its rationale and where that rationale currently answers Q6 the other way. - OLD, the clause: "so a resumed cycle may know a pass happened and not what it found, and - zero and unknown are different facts." NEW: "so a cycle may have **validated a pass this - session** and no longer hold what it found, and zero and unknown are different facts. - **Knowing that a pass ran is not knowing it was valid**: where this session did not validate - it and the slot is gone or unreadable, the pass is - **omitted from the pass specification** rather than entered with `?`, because this grammar - takes one entry per *valid* pass and has no way to say "may not have been one", and the - report names the numbers it omitted (part of the closure-record contract above; a copy - carrying this without the rest stops there)." Old condition: a resumed cycle's knowledge - that a pass happened is enough to keep its entry, with `?` for the counts. **Replaced**: - knowledge that it ran is separated from validation, using **§7's own predicate — this - session validated it** — rather than a durability test no artifact satisfies, since this - change introduces no pass-validity record. Kept: `?` itself, per series, for a validated - pass whose counts are not recoverable. Without this edit the same missing slot both keeps an - entry and is omitted. - -15. **The Gate-A cadence** (C:573 / W:764), which makes a revision unconditional between + assigned fix set** (the closure ordering above). Minor · Nit → collect, never iterate." + Old condition: every Blocker and Major resolves, unbounded. **Replaced**: the boundary + **D5** always implied and no sentence carried. Without it a declined finding must stay + outside the set and still bars closure, which is a pass that can neither close nor suspend. + **The edit names the set and never a decline**, which is AC 3's second half: the answer + moves a finding into or out of the set, and this duty says only what it demands of what is + in it — a mention of the decline here would read as a direct waiver of Blocker/Major + resolution rather than the membership decision it is. The (c) pointer that calls this "the + resolve rule" (`c17`) needs no edit: it names the rule, and the rule now carries its own + scope. + +4. **The lens paragraph's unchanged-list** (C:653–655 / W:839–841), which asserts of two rules + this change alters that they are unchanged. OLD, the single-line fragment: "clean-final-pass + rule are unchanged." NEW: "clean-final-pass rule are unchanged **by the lens sets**, which is + what this paragraph is about — the closure ordering above does change both, scoping the + filter to the assigned fix set and defining a clean pass at effective severity, and says so + there." Old conditions: lenses change what a pass asks and never how many passes a cycle + owes (**kept**); the Blocker/Major filter, the file-first findings protocol and the + clean-final-pass rule are unchanged (**replaced** — scoped to the lens sets, because two of + the three are changed by the ordering and an unscoped claim leaves both copies denying an + edit they carry); the floor is not among them (**kept**). + +5. **The Gate-A cadence** (C:573 / W:764), which makes a revision unconditional between passes. OLD: "Each pass: validate, revise, re-run." NEW: "Each pass: validate, revise - **where the severity and scope rules require a repair**, re-run (part of the closure-record - contract below; a copy carrying this without the rest stops there)." Old condition: every + **where the severity and scope rules require a repair**, re-run." Old condition: every pass is followed by a revision before the next. Kept: the cadence and its order — validate first, re-run last. **Replaced**: the revision is conditional, because the ordering's continue branch reaches a below-floor pass whose only findings are Minors and Nits, which are collected and never iterated, and an unconditional "revise" tells that pass to manufacture the repair the severity rule forbids. +6. **The unknown-start fallback** (C:153–167 / W:360–374, `i4`–`i8`, extended at `i12`'s + invitation) — an addition to a list rather than a change of meaning, so it is checked by + presence and not as a pair. After the single-line anchor "duty owed, and the nonce duties at + their strictest" the list gains "**every suspension binding**". `i12`'s sentence stays as + written; this is the addition it invites, and it is what this change owes that list: a cycle + that cannot establish its starting rules treats a suspension as binding rather than as + advisory, which is the strict reading of the rules §3 ships. + The shorter "Copy every record into the squash body" sentence inside the human-exception block -(C:1004–1007) is generic and already covers an answer record; it is not edited. The "records -every cycle owes" list (C:698) enumerates unconditional records only; an answer record is -conditional, like the human exception, and is not added. +(C:1004–1007) is not edited, and neither is the "records every cycle owes" list (C:698): this +change ships no record. --- ## 5. Edits to the existing passages, with the old-conditions accounting -Ids are the committed inventory's (§2). "Kept" = the sentence stays; "moved" = it now lives in -the §3 block; "replaced" = the condition changes, and says how; "dropped" carries its reason. +Ids are the committed inventory's (§ header). "Kept" = the sentence stays; "moved" = it now +lives in the §3 block; "replaced" = the condition changes, and says how; "dropped" carries its +reason. **Every inventoried passage has an entry**, including the four this narrowing no longer +edits, so that the accounting stays a complete map of the ten rather than a shorter list. **(a) The floor paragraphs** (C:72–136 / W:279–343) — **two edits**. First, **trim** `a17`–`a19` and point at the block. OLD: "Your final pass must be clean — if the pass at the @@ -558,8 +337,7 @@ to pad." Second, `a13`. OLD: "Every other rule stated here about how a cycle closes stands as written, and none of them is restated — a summary is where their conditions would get dropped." NEW: "**This paragraph** restates none of them — a summary is where their conditions would get dropped — and the closure ordering below is where they are -stated once and in order (part of the closure-record contract below; a copy carrying this -without the rest stops there)." Accounting: `a1`–`a12`, `a14`–`a17`, `a20`–`a22` kept; `a18`, `a19` +stated once and in order." Accounting: `a1`–`a12`, `a14`–`a17`, `a20`–`a22` kept; `a18`, `a19` moved; **`a13` replaced**. Old condition: *no* rule about how a cycle closes is restated anywhere. Kept: the prohibition and its reason, scoped to the paragraph it was written to police. Changed: the ordering does restate closure rules, deliberately and as the one @@ -569,27 +347,29 @@ failure `a13` exists to prevent, arriving from the other direction. **(b) What a loop absorbs** (C:195–223 / W:402–426) — **three sentence edits**, the triggers stay. `b12` OLD: "…and it resumes the moment the user says whether the set now includes it." NEW: "…and it resumes once the user has said whether the set now includes it — accepting or -declining it — **together with every other answer that pass's suspensions require, each -recorded as Mechanics requires**, under the closure ordering above." `b17`–`b18` OLD: +declining it — **together with every other answer that pass's suspensions require**, under the +closure ordering above." `b17`–`b18` OLD: "Stopping this way is **not an exit from the gate**: the floor, the Blocker/Major filter and the clean-final-pass rule all stand, and the loop resumes on the revised artifact once the question is answered — what the stop prevents…" NEW: "Stopping this way is a **suspension** -under the closure ordering above (part of the closure-record contract below; a copy carrying -this without the rest stops there), and the loop resumes on the revised artifact **once every +under the closure ordering above, and the loop resumes on the revised artifact **once every answer that pass's suspensions require has been given** — what the stop prevents…". Third, in **W only**, "…by its severity exactly as the severity rule already says…" becomes "…exactly as Mechanics already says…", matching C for the reason §6's first row gives. -Accounting: `b1`–`b11`, `b13`–`b16` kept — `b6` among them, and it is worth naming: it fixes +Accounting: `b1`–`b11`, `b13`–`b16` kept. Two are worth naming. `b6` fixes the assigned set **before the pass being answered**, and the ordering's membership answer is read against that same set rather than against a later one, so the two agree and `b6` needs no edit. An earlier revision of this spec had membership re-read when the answer arrived, which `b6` and **D4** both forbid — **D4** requiring a user answer for the hold to end at all — and -the reversal is recorded here rather than left as a silent narrowing. `b12` **replaced** — old: an *immediate* resume on -the membership answer alone. Kept: that the membership answer is what the stop asks for and -that either direction ends it. Changed: the resume waits for every answer the pass's -suspensions require, each recorded, because a loop resumed over an unanswered question decides -it by running. `b17` moved (the block's "a suspension waives nothing" sentence); `b18` +the reversal is recorded here rather than left as a silent narrowing. `b11` and `b13` are the +two trigger descriptions and **stay in their own words**, checked equivalent to the block's +definitions rather than merged into them, because editing them would rewrite the passage this +change deliberately leaves standing; §6 is where that check runs. `b12` **replaced** — old: an *immediate* +resume on the membership answer alone. Kept: that the membership answer is what the stop asks +for and that either direction ends it. Changed: the resume waits for every answer the pass's +suspensions require, because a loop resumed over an unanswered question decides it by running. +`b17` moved (the block's "a suspension waives nothing" sentence); `b18` **replaced** the same way and for the same reason — old: resume once *the* question is answered; new: once *every* required answer is given. W's remaining wording differences here are untouched (§6). @@ -605,8 +385,7 @@ nothing closes**" to the end of "…and then nothing could satisfy both.": repla trigger** — a finding outside the assigned fix set, or one opening a new structural or contract question — since a pass carrying either is not clean and the ordering above says why. **Below the floor nothing closes**, exactly as that ordering says. This exit is -a **suspension** under it (part of the closure-record contract below; a copy carrying this -without the rest stops there): you surface with the findings still open, the resolve +a **suspension** under it: you surface with the findings still open, the resolve rule is not waived by surfacing, and the loop resumes when every answer that pass's suspensions require has been given." The kept sentence is not touched; the qualification is adjacent to it, which is how **D3**'s verbatim requirement and the ordering can both hold — @@ -619,36 +398,33 @@ zero-finding exception); `c14` moved **to the ordering's continue branch** — t Minor-below-floor pass that "keeps looping" is precisely one that neither closes nor suspends, and the branch states it rather than leaving it to be read out of a negation; `c15`, `c17` kept in the pointer sentence, `c17`'s "resolve rule" now reading with the assigned-fix-set -boundary §4 item 13 gives it, so the pointer needs no edit of its own; `c16` **replaced** — +boundary §4 item 3 gives it, so the pointer needs no edit of its own; `c16` **replaced** — old: a clearly-stuck surface leaves its finding open, which read alone gives this exit a hold. -Kept: that the findings are still open at the surface, and that the resolve duty is what keeps -them so. Changed: the **hold** belongs to the scope stop, the only stop whose question is about -a finding, while this exit asks continue or stop and holds nothing. Authority **D1**, which -makes all three exits suspensions and gives a suspension its own transitions; `c18` -**replaced** — old: no pass is credited as clean on a clearly-stuck surface, unconditionally. -Kept: the rule wherever it decides anything — a pass carrying a scope-stop trigger is unclean, -and this exit is reached only on a pass that did not close. Changed: where the exit's -regenerating findings are in-set and the ceiling demotes them below Major, the pass is clean at -effective severity and closes at or above the floor. Authority **D3**, which ranks clean -completion above this exit — under the old reading D3's own preserved sentence and `c18` would -decide that pass in opposite directions. The two-tell -stop was never in either condition's domain, surfacing tells and not a finding, and the -ordering says so rather than leaving it inferred; `c19` **replaced** — old: the loop resumes on -whatever the user decides, unconditionally; new: it resumes on a resuming answer and a stop -leaves the suspension standing, because an unconditional resume is the stop-with-no-transition -path AC 4 forbids; `c20` **dropped** — it argued that reading the exit as "stop instead of -fixing" would compete with the resolve duty, and the block states that argument's conclusion -as a rule instead. - -**(d) From pass 4 onward** (C:255–261 / W:459–465) — **add** the Q6 sentence (§7). `d1`–`d7` -kept, unchanged. +Kept: that the findings are still open and that the resolve duty is what keeps them so. +Changed: the **hold** belongs to the scope stop, the only stop whose question is about a +finding, while this exit asks continue or stop. Authority **D1**, which makes all three exits +suspensions with their own transitions; `c18` **replaced** — old: no pass is credited as clean +on a clearly-stuck surface, unconditionally. Kept: the rule wherever it decides anything — a +pass carrying a scope-stop trigger is unclean, and this exit is reached only on a pass that did +not close. Changed: where the exit's regenerating findings are in-set and the ceiling demotes +them below Major, the pass is clean at effective severity and closes at or above the floor. +Authority **D3** — under the old reading D3's own preserved sentence and `c18` decide that pass +in opposite directions. The two-tell stop was never in either condition's domain, surfacing +tells and not a finding; `c19` **replaced** — old: the loop resumes on whatever the user +decides, unconditionally; new: it resumes on a resuming answer and a stop leaves the suspension +standing, because an unconditional resume is the stop-with-no-transition path AC 4 forbids; +`c20` **dropped** — it argued that reading the exit as "stop instead of fixing" would compete +with the resolve duty, and the block states that argument's conclusion as a rule instead. + +**(d) From pass 4 onward** (C:255–261 / W:459–465) — **no longer edited.** The Q6 +unavailable-history block this passage was to receive moved to the successor with **D10**, and +nothing the narrowed change ships touches the three-line duty. `d1`–`d7` kept, unedited. **(e) The five tells** (C:263–268 / W:467–472) — **one edit and one added sentence**. The edit is `e7`, the shared fragment both copies carry (C:266 / W:470). OLD: "**Any two present makes stop-and-surface mandatory, not discretionary**". NEW: "**Any two present makes stop-and-surface mandatory, not discretionary, where the clean-completion branch did not close the pass**". Then -the pointer, after `e10`: "This stop is a **suspension** under the closure ordering above (part -of the closure-record contract below; a copy carrying this without the rest stops there), and +the pointer, after `e10`: "This stop is a **suspension** under the closure ordering above, and the loop resumes when every answer that pass's suspensions require has been given." Accounting: `e1`–`e6`, `e8`–`e11` kept; `e7` **qualified** — old: two tells make the stop @@ -679,9 +455,7 @@ the whole paragraph, both copies, with the answer. NEW: which findings the fix set contains — a separate predicate the closure ordering defines, and one this must not be read as touching; a loop spending passes on findings the author keeps demoting is exactly what the prose-cluster tell exists to surface, and lowering the counts by - that same judgement would hide it. **This split is one component of the closure-record - contract** the one-contract rule below names, and a copy carrying it without the rest of - that list is a partial adoption that stops there. + that same judgement would hide it. ``` Accounting: `g1` **dropped** (the unsettled statement, now settled); `g2`, `g3` **dropped** @@ -689,23 +463,27 @@ Accounting: `g1` **dropped** (the unsettled statement, now settled); `g2`, `g3` **dropped** in C (the ownership sentence, discharged by this change), and W, which never carried it, gets the same replacement — removing the one deliberate story-path difference. -**(h) Recording a human exception** (C:977–1033 / W:1161–1217) — **unchanged**, and the -"**Recording an answer at a membership stop.**" block (§4, verbatim) is added immediately after -its last paragraph, before "- **Timeout / abort:**". `h1`–`h26` kept; `h17` stays true of the -human-exception form, and the answer-record block says which item of that list it *is* the -answer to. +**(h) Recording a human exception** (C:977–1033 / W:1161–1217) — **no longer edited.** The +answer-record block that was to follow it moved to the successor with **D9**. `h1`–`h26` kept, +unedited. + +**(i) When these rules bind** (C:153–167 / W:360–374) — **extend** the strict-reading list with +one item (§4 item 6); `i1`–`i16` kept, none replaced, the list being added to rather than +rewritten. + +**(j) The squash carry** (C:892 / W:1076) — **no longer edited.** It was extended to name the +answer record, which moved to the successor; this change ships no record for it to carry. +`j1`–`j4` kept, unedited. -**(i) When these rules bind** (C:153–167 / W:360–374) — **extend** the strict-reading list -(§4 item 5); `i1`–`i16` kept. **(j) The squash carry** (C:892 / W:1076) — **extend** (§4 item -1); `j1`–`j4` kept. Also touched, outside the inventoried passages: §4 items 2, 3, 4, 7–15. +Also touched, outside the inventoried passages: §4 items 1–5. --- ## 6. Parity -The two copies must agree on every rule this spec changes. The block (§3), the answer-record -block (§4), the (g) replacement, the Q6 text and every list extension ship byte-identical in C -and W. The pre-existing divergences the inventory found are handled as follows: +The two copies must agree on every rule this spec changes. The block (§3), the (g) replacement +and every source edit ship byte-identical in C and W. The pre-existing divergences the inventory +found are handled as follows: | Divergence | Kind | This change | |---|---|---| @@ -718,110 +496,16 @@ and W. The pre-existing divergences the inventory found are handled as follows: The parity check at execution: extract each edited passage from both files by its lead phrase and `diff` them; the only differences permitted are the rows marked as-is, and the result is -part of the evidence entry (§9). Any other difference is a defect, not a wording choice. The -Mechanics region the answer-record block joins is byte-identical between the copies today -(C:840–1033 = W:1024–1217) and stays so. +part of the evidence entry (§7). Any other difference is a defect, not a wording choice. The +same extraction also compares `b11` and `b13` against the block's two trigger definitions, per +§5(b) — a parity check between copies and an equivalence check within each. The correspondence +it confirms: `b11`'s "a correction that leaves that set" and "out-of-scope finding" is the +block's *outside the current assigned fix set*, and `b13`'s "opens a **new structural or +contract question**" is the block's question trigger word for word. --- -## 7. Q6 — the pass-4 report without prior-pass history - -Appended to the "**From pass 4 onward**" paragraph, both copies, after "…demanding what an -earlier pass had removed.": - -``` -**Where an earlier pass's findings file is unavailable**, the report says so before it reads -anything. **First the root**, which is a condition on the current pass being valid at all and -not a stop of its own. The slots live in `.context/codex-reviews/` under the top-level -directory of the checkout this cycle is running in — the same root the pass call is given as -`workingDirectory`, the two compared after both are **canonicalized**, so that a symlink or a -trailing slash is not a mismatch. That is **one condition**: either the slots being read are -under that checkout or they are not. -A pass whose root cannot be established is an **INCOMPLETE pass** — the -state this section already defines, already uncounted toward the floor and already excluded -from the curve — so nothing new is ranked in the closure ordering. The report carries **git's -own message verbatim** and files it under no shape of its own: that message is the -discriminator, and prose cannot partition git's failures better than git does — a partition -written here would send a caller to the wrong repair on every case it guessed wrong, which is -the whole of what such a list would add. What no check reaches: a -slot written under a different root **in the past** is **indistinguishable from an absent -slot**, because nothing records where a past call ran. Unobservable, not detected. -**Then, per earlier pass, two questions.** Is the slot present and valid? And is the pass -**known accepted** — meaning **this session validated it**? Present and valid is the ordinary -case. **A known-accepted pass whose slot is now absent or invalid** is a real pass whose -artifact is unusable: its number **stays counted** and its series read `?`, which the curve -grammar already admits for a **valid** pass whose counts cannot be recovered, since a corrupted -or missing record does not un-run a pass. **After a -lost session nothing is known accepted**, so every absent or present-but-invalid slot is -**acceptance unknown**: **not counted toward the floor** and **omitted from the durable -curve**, which takes one entry per *valid* pass and has no way to say "may not have been -one" — `?` is a missing count, not a missing pass. The omission is named in the report beside -the reduced-sensitivity line, so it is disclosed where it is decided rather than inferred from -a gap in the curve. Crediting an unvalidated pass is the dangerous direction and this is the -other one — and no pass-acceptance record is invented to escape the cost, since the cost is -one pass and the record would be a permanent duty. A pass **known incomplete when it ran** is -excluded, exactly as today. -**Then the report.** It reads the cycle's working record if one exists and says whether it -used it; states which of the three lines it computed and from which passes; **names the pass -numbers it could not read**, and why; and says the two-tell threshold is being read on that -reduced record. **A gap does not break the series**: the consecutive *available* passes -compare across a missing one, and a comparison that spans a gap is **visible but weaker -evidence**, named as such — while a comparison needing the missing pass as one of its two -endpoints is simply unavailable. Refusing to compare across a gap would silence tell detection -exactly where the record is thinnest, and the tells read the direction of the loop rather than -any one adjacent pair. This is not a stop of its own, and it does not make the working record -mandatory: a report that says what it could not see is the duty; one that invents the trend, -or omits a line without saying so, is the failure. This report and its root condition are one -component of the closure-record contract (Mechanics); a copy carrying them without the rest of -that list is a partial adoption that stops there. Filled, over passes 1, 2 and 4 with 3 -unreadable: - - Root: `.context/codex-reviews/` under this checkout — confirmed. (Unconfirmable: - this pass is INCOMPLETE, root not established — git said: "".) - Trend: findings 24, 17, —, 18; Blockers 5, 2, —, 1 — pass 3 omitted: slot absent, - acceptance unknown, so it is neither counted toward the floor nor entered in the - curve. 2→4 spans that gap: rising, weaker evidence. Cluster: product behaviour - 12 of 18. require↔withdraw: none visible; a pair with pass 3 as an endpoint - cannot be read. Threshold read on 3 of 4. -``` - -This is **D10**. Both historical lines are derivable from the mandated findings files alone -(the `fic2` record verified that), so unavailability is a property of the workspace and the -answer is disclosure, not a new stop or a mandatory artifact — the root condition reuses the -INCOMPLETE state rather than adding one, so **D10**'s "not a new stop condition" stands as -written. The per-pass partition is by what the agent can observe -(`docs/prompt-standards.md` item 10): presence and validity from the slot, acceptance from this -session alone, and where acceptance cannot be established the count moves in the direction that -costs a pass. The root is deliberately **not** partitioned — item 10 asks that a diagnostic -state name its cause and its fix, and here git's message is the cause, named verbatim, while -any prose list would be this spec guessing at causes it cannot observe. The example is -there for item 4. - ---- - -## 8. The slot discriminator — dissolved - -Plan C's Tasks 19 and 20 (`docs/superpowers/plans/2026-08-30-review-loop-economics-plan-c-rollout.md`, -the drop note under Task 19) deferred "a general production" for a short deterministic -discriminator in the nonce's slot position to this story. It ships nothing here, because for -**post-rule and unknown-start cycles** the case it served does not arise: every cycle started -after the parent's rules bind holds a nonce, and one whose start cannot be established mints -one rather than claiming `none (pre-rule)` (`i11`). A production for those would legislate for -an unreachable state, AC 1's prohibition. **The claim reaches no further.** A cycle that began -before the parent's rules shipped has no nonce, cannot acquire one, and writes -`cycle none (pre-rule)`; a rollback can make old-rule cycles reachable again. That set is -bounded and self-terminating, but while two such cycles are observably live they compute the -same bare slot paths and can delete each other's findings files — the incident the parent -records. **Nothing shipped here prevents that**, and no text in this change reaches a cycle -running under the old rules that never reads this document. The trade, stated rather than -dressed as protection: a permanent production for a closing set costs more than the exposure -it removes, and the exposure is real meanwhile. The `rle` naming stays what its closing body -recorded — a plan-local exception under the old rules. The durable prior record is that plan's -drop notes (Task 19, line 970; Task 20, line 1047) and Task 23's second point (line 1298). - ---- - -## 9. Verification +## 7. Verification **The mode is read from the story's header at execution**, never from here — the same rule the spec's own header states, and the reason no value is named in this heading. What follows is @@ -829,7 +513,7 @@ what each level of that mode obliges, so that whichever it carries has its evide the battery, the check, and the named verification of the risk path. **Battery.** The quality command in `AGENTS.md` § Commands runs once, green, at the Gate-B WIP -commit (`check-version-bump.sh main` needs the committed bump, §10). +commit (`check-version-bump.sh main` needs the committed bump, §8). **The check.** Inline shell asserts in the plan, run against the working tree and the parent tree, so each assertion is observed passing where the change exists and failing where it does @@ -845,104 +529,78 @@ not: check above cannot do this alone: deleting the old sentence and installing nothing satisfies it and the parity diff, reporting the demotion/loop-health outcome as verified while both copies carry no answer at all; -- lead phrase `**Recording an answer at a membership stop.**` count **1** in each; **0** in the - parent tree, and with it both labels: `Accepted: ` and `Declined: ` count - **1** in each and **0** there, since one label shipping without the other is the shape the - block's whole form exists to prevent; -- the six-member squash-carry sentence: `every answer record (part of the closure-record` count - **1** in each; **0** in the parent tree — the phrase checked against §4 item 1's NEW text - rather than assumed to match it; -- the Q6 lead phrase `**Where an earlier pass's findings file is unavailable**` count **1** in - each; **0** in the parent tree; -- the **five** source edits, each as a pair, because each is a standing sentence changing - meaning rather than new text appearing, and a one-sided presence check would pass on a copy - carrying both wordings. **Every fragment here is a single line in the file it is grepped - from**, each OLD verified at 1 in both copies of the parent tree, because a fragment spanning - a line break makes `grep -F` count 0 and read as a failure — an earlier revision quoted - three of them across their wraps: `A **clean findings file** is the single body line` at - **1** in each and **0** in the parent, with `clean pass is the single body line` at **0** and - **1**; `when a pass finds nothing` at **1** and **0**, with `when a pass is clean` at **0** - and **1**; `for every finding in the assigned fix set` at **1** and **0**, with `both must - resolve. Minor` at **0** and **1**; `Knowing that a pass ran is not knowing it was valid` at - **1** and **0**, with `may know a pass happened` at **0** and **1**; and - `revise **where the severity and scope rules require a repair**` at **1** and **0**, with - `Each pass: validate, revise, re-run` at **0** and **1**. The NEW fragments must be installed - unwrapped at those points — a constraint on how the plan writes the edits, not on what they - mean; -- the contract markers, since a partial merge sees them and not the membership list: `part of - the closure-record contract` **case-insensitively** at **17** in each and **0** in the parent - — one per source edit in §4 except items 6 and 9, which are respectively unchanged and the - contract itself (thirteen), plus one on each of the four pointer sentences §5(a), (b), (c) - and (e) install; several open a sentence and capitalise, which is why the count is - case-insensitive and why a case-sensitive one would under-count and read as a failure — plus - the four block-level sentences (`This ordering is one component`, `These records are one - component`, `This split is one component`, `This report and its root condition are one - component`) at **1** each and **0** there. **Twenty-one** markers in each copy. **Every one - of the twenty-one must be installed on a single line**, the same constraint the source-edit - fragments carry and for the same reason: this spec itself wraps the marker phrase at several - of the sites that quote it, and a wrapped marker in the shipped copy makes `grep -F` count it - as absent. The constraint is on how the plan writes the line, not on the sentence's meaning, - which is why the marker sentences are short enough to fit one. - -If the claim "the ordering and the records ship in both copies" were false, one working-tree -count would be **0** or the parent-tree counts would not differ from it. The wiring can produce -that observation: each grep reads the file bytes at the named revision and nothing supplies its -own input. **The counterfactual is ABSENT, and is claimed as absent** — the parent carries no -ordering block and no answer records, and the (g) count is the one site where the parent is -present and the change removes it. Nothing is claimed as "contradictory". - -**The named verification of the risk path** (story AC 4) is a **next-state table**, in the -plan and quoted by the closing commit body. **Rows** — every stop the shipped text names: -membership stop, question stop, a finding carrying both triggers, clearly-stuck exit, two-tell -stop, a hold awaiting its answer, accept, decline, a stop answer, two or three suspensions at -once, a below-floor clean pass, a zero-finding pass, the unknown-start fallback; **and the -stateful transitions**: a declined finding re-raised matching on all five fields and re-raised -with one changed, the fix set broadened to include a declined finding, **the fix set broadened +- the **five paired source edits** (§4 items 1–5), each as an OLD/NEW pair, because each is a + standing sentence changing meaning rather than new text appearing, and a one-sided presence + check would pass on a copy carrying both wordings. Each NEW counts **1** in each copy and + **0** in the parent tree, each OLD the reverse: `A **clean findings file** is the single body + line` against `clean pass is the single body line`; `when a pass finds nothing` against `when + a pass is clean`; `for every finding in the assigned fix set` against `both must resolve. + Minor`; `clean-final-pass rule are unchanged **by the lens sets**` against `clean-final-pass + rule are unchanged.`; and `revise **where the severity and scope rules require a repair**` + against `Each pass: validate, revise, re-run`. **Every fragment is a single line in the file + it is grepped from**, and the NEW ones must be installed unwrapped, because a fragment + spanning a line break makes `grep -F` count 0 and read as a failure — an earlier revision + quoted three of these across their wraps. That is a constraint on how the plan writes the + edits, not on what they mean; +- §4 item 6, the unknown-start addition, by presence alone since it adds to a list rather than + changing a meaning: `every suspension binding` count **1** in each and **0** in the parent. + +If the claim "the ordering ships in both copies" were false, one working-tree count would be +**0** or the parent-tree counts would not differ from it. The wiring can produce that +observation: each grep reads the file bytes at the named revision and nothing supplies its own +input. **The counterfactual is ABSENT, and is claimed as absent** — the parent carries no +ordering block, and the (g) count is the one site where the parent is present and the change +removes it. Nothing is claimed as "contradictory". + +**The named verification of the risk path** (story AC 4) is a **next-state table**, in the plan +and quoted by the closing commit body. **Rows** — every stop the shipped text names: membership +stop, question stop, a finding carrying both triggers, clearly-stuck exit, two-tell stop, a hold +awaiting its answer, accept, decline, a stop answer, two or three suspensions at once, a +below-floor clean pass with no suspension, **a below-floor clean pass carrying a health +reading** — whose next state is the suspension, the row that tests the third branch's +qualification — a zero-finding pass, and the unknown-start fallback; **and the stateful +transitions**: a finding this cycle declined re-raised on a later pass, **the fix set broadened while a hold is still awaiting its answer** — whose next state is the hold still standing, the -row that tests the frozen reading — two records disagreeing on one five-field key, recovery -after an accept and after a stop with the session lost (no identity → new cycle), -a `full` Gate-B pass with one branch -clean and the other carrying an in-set Blocker, a rollback with a cycle open under the new -rules, and a copy adopting the ordering without the answer records, the reverse, or either -without the severity split. **Columns** — the **record state** (which answer records the bodies -carry, which working record exists) and the **governing-scope state** (the fix set as currently -assigned) beside the user's answer, then the input that ends the row, the state afterwards, and -the shipped line the row reads, in both copies. What would be observed if "no path leaves a -cycle unable to close and unable to suspend" were false: a row whose next state is the same -stop with no input consumed, or one that closes with an in-set Blocker standing. The wiring can -produce it because every row is filled from the shipped text rather than from this spec, and -the inputs the `fic2` instrument omitted — the user's answer and the record state — are columns -here. That is what makes this **not the `fic2` decision matrix**, whose two defects Gate B found -in the technique itself: a state's inputs must include every input the rule reads, and a -counterfactual must distinguish ABSENT from CONTRADICTORY. **No fixture per predicate is -built** — parked in the story's §2, not reopened. - -**Evidence entry**, in the closing commit body, names: the battery run; every assert pair -above with its working-tree and parent-tree counts; the §6 parity diff — the passages -extracted, the differences observed, and that each is one of the permitted rows; and the -next-state table's location in the plan plus its row count. It is revalidated before every -Gate-B re-review and before the closing amend, as §5 requires. +row that tests the frozen reading — a fix set narrowed so an in-set finding falls outside it, +and a `full` Gate-B pass with one branch clean and the other carrying an in-set Blocker. +**Columns** — the **governing-scope state** (the fix set as currently assigned) and the +**answers already given in this cycle**, beside the user's answer for this row, then the input +that ends the row, the state afterwards, and the shipped line the row reads, in both copies. + +**The oracle.** A row **fails** when its required answer does not produce a **distinct** +resumable or closed state — the same stop returning, whether or not an input was consumed — +or when it closes with **any** closure precondition unmet: an in-set Blocker or Major, a +standing hold, an unanswered question, the derived floor, or a cited-set or profile change +during the pass. Naming only the first two would pass the exact no-progress defect AC 4 cites +from the parent cycle. The wiring can produce that observation because every row is filled from +the shipped text rather than from this spec, and the inputs the `fic2` instrument omitted — the +user's answer and the governing-scope state — are columns here. That is what makes this **not +the `fic2` decision matrix**, whose two defects Gate B found in the technique itself: a state's +inputs must include every input the rule reads, and a counterfactual must distinguish ABSENT +from CONTRADICTORY. **No fixture per predicate is built** — parked in the story's §2, not +reopened. + +**Evidence entry**, in the closing commit body, names: the battery run; every assert pair above +with its working-tree and parent-tree counts; the §6 parity diff — the passages extracted, the +differences observed, that each is one of the permitted rows, and the `b11`/`b13` equivalence +result; and the next-state table's location in the plan plus its row count. It is revalidated +before every Gate-B re-review and before the closing amend, as §5 requires. --- -## 10. AGENTS.md invariants touched - -- **Invariant 11 — `docs/prompt-standards.md`, all 12 items.** Most at risk: item 6 (every - constraint in the shipped blocks carries its reason in the same sentence — **including the - three that were exempted until pass 6**: no other pass outcome closes a cycle *because every - other pass either leaves a required repair, a finding hold or a question outstanding, or has - an unmet closure precondition such as the floor*; a zero-finding pass is clean whatever the - floor *because a floor buys further looks at an artifact that keeps yielding findings*; - decline is available only at a membership stop *because that is the only stop whose question - is whether a finding belongs to the set*. The exemption was wrong twice over: item 6 admits no - "settled elsewhere" clause, and a scaffolded copy cannot reach the story the reasons were - said to live in); item 8 (token-lean — the blocks replace closure - sentences rather than adding beside them); item 4 (the Q6 example); item 10 (the Q6 per-pass - partition, and the root condition, whose cause is git's own message rather than a partition - written here — item 10 asks for a cause and a fix, and a guessed cause is neither); - item 3 (the stop answer is a named, resumable state). +## 8. AGENTS.md invariants touched + +- **Invariant 11 — `docs/prompt-standards.md`, all 12 items.** Most at risk: item 6, every + constraint in the shipped block carrying its reason in the same sentence — **including the + three exempted until pass 6**, which now carry theirs inline in §3: that no other pass outcome + closes, that a zero-finding pass is clean whatever the floor, and that decline is available + only at a membership stop. The exemption was wrong twice over: item 6 admits no "settled + elsewhere" clause, and a scaffolded copy cannot reach the story the reasons were said to live + in. Then item 8 (token-lean — the block replaces closure sentences rather than adding beside + them) and item 3 (the stop answer is a named, resumable state). - **Don't: "Never replace a decision procedure without accounting for its old conditions."** - Satisfied by §5, against the committed inventory. + Satisfied by §5, against the committed inventory, with an entry for every inventoried passage + including the four this narrowing no longer edits. - **Don't: "Never rename or delete a doc section without grepping for references first."** The (g) sentence names the story path; the grep finds it at `CLAUDE.md:815` (the site itself) and in three artifacts of the parent cycle (`…plan-a-rules.md:878`, @@ -950,14 +608,36 @@ Gate-B re-review and before the closing amend, as §5 requires. which continues to exist; none cites the sentence. Nothing breaks. - **Invariant 12 — a plugin change requires a version bump.** `workflow-init.md` is under `plugins/`, so `plugins/dev-workflow/.claude-plugin/plugin.json` goes `0.11.0 → 0.12.0` with - a `CHANGELOG.md` entry: a minor bump, the template gaining a record form and a rule. The - entry also notes that the squash-carry sentence now lists six record kinds. + a `CHANGELOG.md` entry: a minor bump, the template gaining a closure ordering and six edited + sentences. - **Invariant 4 / the hook.** Untouched: `codex-gate.sh` is not edited, and the §5 heading it greps (`Cross-Model Review`) does not move. --- -## 11. Out of scope / parked +## 9. Moved, out of scope, and parked + +**Moved to `docs/superpowers/stories/2026-09-10-record-durability-story.md`** on Daniel's +decision of 2026-09-10, with the evidence in §1: + +- the `Accepted:`/`Declined:` answer record, its form, transport, attribution, recording point + and squash carry (**D9**), the sameness test by which a later pass recognises the same + finding (**D9b**), and the unverified-assertion reading (**D9c**); +- the pass-4 report's unavailable-history block (**D10**) and the checkout-root condition that + report grew; +- the rollback reading — what an open cycle owes when a revert removes the text it started + under; +- the slot-discriminator dissolution deferred here by Plan C's Tasks 19 and 20; +- **the partial-adoption guard for the whole set.** It was built as a "closure-record contract" + naming the record, with a marker on every mergeable hunk, and it leaves with the record it was + named for. The narrowed set is still mutually dependent — the ordering's clean predicate needs + §4 item 3's boundary, and its two senses of *clean* need items 1 and 2 — and **nothing here + guards that**. It belongs with the successor because the two sets are adopted together in + practice and one guard should name both; until then a partial `/workflow-init` merge of these + edits is caught only by the existing one-contract paragraph's semantic coherence rule, which + does not name them. Stated as a residual, not as protection. + +**Out of scope and parked**, unchanged: - The pass-counter anomaly (`fic2` record). - The CodeRabbit plan-metadata contradiction (`fic2` record). From 11b0e4739340dfca555361f89f702833a4e4ac3e Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Thu, 10 Sep 2026 20:05:27 +0200 Subject: [PATCH 016/181] =?UTF-8?q?docs(specs):=20loop-rule=20consolidatio?= =?UTF-8?q?n=20=E2=80=94=20Gate-A=20pass=2011=20revision=20(one=20authorit?= =?UTF-8?q?y=20per=20rule)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Pass 11's finding 14 was the centrepiece and findings 3, 4 and 6 were its symptoms: the closure-ordering block restated triggers, duties, preconditions and the severity answer that their own paragraphs still defined, so each prompt copy held two authorities and they had drifted. The repair inverts what the block does. It is authoritative for the evaluation order and for closure, and cites every other rule where that rule already lives; where a cited rule had to change to agree, it changed at its source rather than being restated. The block falls 162 -> 117 lines and §3 gains a table naming all eight cited rules and their single definitions. Also: the hold returns to every surfaced finding, as the story's third standing duty says; a below-floor clean pass carrying a health reading suspends; continue consumes the reading and stop parks the cycle, so both answers produce a distinct state; b7 and b11 are edited at their source for the multi-artifact union and the declined-finding exception; the partial-adoption overclaim is corrected to an admitted unsafe state; and every quoted OLD fragment is re-cited with its real line range in both copies plus the single-line substring the check counts. All 20 findings fixed. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- ...26-09-10-loop-rule-consolidation-design.md | 565 ++++++++++-------- 1 file changed, 328 insertions(+), 237 deletions(-) diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index c299b87..9b65828 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -51,10 +51,13 @@ established are read as given: a pass's cleanliness is a fact about what that pa establish the **inventory** of findings, not their resolutions. What the table does not settle, this spec decides in the section that uses it: the evaluation -order and the field and file set each predicate reads; the duties' classification; the scope -stop's two triggers and what each answer does; what a stuck or two-tell answer produces; and the -raw-severity rule for the health measures. Six standing sentences are edited at their source -rather than worked around, each one the ordering falsifies or leaves ambiguous (§4). +order and the file set each predicate reads; the duties' classification; which stop each of the +scope stop's two triggers raises and what each answer does; what a clearly-stuck or two-tell +answer produces; and the raw-severity rule for the health measures. **It defines no trigger, +duty or severity rule of its own** — each keeps its one definition where that definition already +lives, and where one had to change to agree with the ordering it changed **at its source**: six +standing sentences outside the inventoried passages (§4), and seven fragments inside them (§5), +each one the ordering falsifies or leaves ambiguous. --- @@ -65,204 +68,186 @@ in both copies, byte-identical. It states the ordering once; the paragraphs afte triggers and point at it. Verbatim as it will ship: ``` -**How a cycle ends — one ordering, stated here and referenced everywhere else.** A pass is -read in a fixed order, because every rule below bears on one decision — may this cycle close — -and a stated order is what stops them qualifying each other. **Every finding-derived predicate -here** reads the validated findings file **or files** of the logical pass as their -**concatenation** — a `full` Gate-B pass has two, and one branch alone is already an incomplete -pass — at **effective** severity, after the Mechanics severity ceiling, because cleanliness is -about what the cycle must repair and the ceiling is what decides that. **Both branches' lines -count** for the health measures, the curve rule already summing the branches into one entry; a -finding matching one in the other branch is **one** finding for holds and answers, so it is -answered once. The **health measures** (Mechanics, Severity) read the reviewer-written field -from before the ceiling — the per-pass counts, the clusters, the tells, and **both -severity-bearing conditions of the three-condition stuck reading**, its Blocker curve and its -regenerating Blocker or Major findings, since a predicate reading one field for half of itself -could not be read at all, while its third condition, the stated coverage-sufficiency judgement, -reads no severity field and the ceiling does not touch it. **The branches read more than the -findings**, named here so nobody applies the order to the findings file alone: closure reads the -**derived floor**, the **cited set and every profile** as re-read for final acceptance, and any -**hold still standing**; the membership trigger reads the **current assigned fix set** and the -answers this cycle has already given; the stuck reading adds its coverage judgement. - -**First, clean completion.** §5 uses *clean* in two senses and now says which is which. A -**clean findings file** is the `NO FINDINGS` signal the protocol defines. A **clean pass** is -the predicate below, read on the **logical pass with every required branch file combined**: -one branch's clean findings file never establishes a clean pass, the other being free to carry -an in-set Blocker. Every required file carrying that signal gives one, and a clean pass need -not have it, because a pass carrying only Minors and Nits, or findings this cycle has declined, -is clean without being empty. **A pass is clean** when its findings carry no Blocker -or Major at effective severity that is **in the assigned fix set**, and **no scope-stop -trigger**. There are two triggers, and this branch is where cleanliness and suspension both read -them, in the same words so they cannot drift: a **membership trigger** is a finding outside the -current assigned fix set **and not one this cycle has already declined** — such a decline having -put it outside by the user's own answer, for as long as the current set still excludes it, so -that an answered question is not asked again — and a **question trigger** is a finding opening a -new structural or contract question, in-set or not, which a decline never suppresses, an answer -about membership being no answer to a question. Both are properties of the findings and the set, -read before any branch below runs, which is what makes this order executable rather than -asserted: no predicate here waits on an act a lower branch performs. A pass -with **zero** findings is clean whatever the floor, because a floor buys further looks at an -artifact that keeps yielding findings, and one yielding none has already given what those -looks were for. **A clean pass at or above the derived floor, -or a zero-finding pass, closes the cycle**, under the final-acceptance preconditions the floor -section already states and this ordering does not restate — the cited set and every profile -re-read before the pass is accepted as final, a header or profile that changed during it -making the pass not final and costing the further pass that section requires. A plateau or -tells present on the closing pass go into the -closing report and never block it, because reporting "will not converge" on a converged loop -is a false report, and the clearly-stuck paragraph says the same of its own exit. **No other -pass outcome closes a cycle**, because every other pass either leaves a required repair, a -finding hold or a question outstanding, or has an unmet closure precondition such as the -floor — and closing over any of those is the failure this ordering exists to prevent. The one -termination that is not a pass outcome is the Gate-B triviality skip, which ends a cycle -having run no passes and is outside this ordering entirely. +**How a cycle ends — one ordering, stated here and referenced everywhere else.** A pass is read +in a fixed order, because every rule bearing on one decision — may this cycle close — otherwise +qualifies the others and the ranking survives only in a reader's head. **This paragraph is +authoritative for that order and for closure, and for nothing else.** Every trigger, duty and +severity rule it names keeps its one definition where that definition already lives; a reader +who finds a rule defined here rather than cited has found a defect. + +**What a pass is read from.** Every finding-derived predicate reads the validated findings file +**or files** of the logical pass as their **concatenation** — a `full` Gate-B pass has two, and +one branch alone is already an incomplete pass. **Which severity field each reads is settled in +Mechanics, Severity**: the closure and repair predicates read the effective severity, the health +measures the reviewer-written one. Beyond the findings, closure reads the **derived floor** and +the final-acceptance preconditions the floor section states, and any **hold still standing**; the +scope triggers read the **current assigned fix set** — what the absorb paragraph defines, plus +what this cycle has accepted, minus what it has declined — and the answers already given; the +clearly-stuck reading adds its own coverage judgement. **A line in one branch file and a line in +the other are distinct findings for holds and answers**, so a `full` pass asks twice rather than +risk resuming over one it never asked about. + +**First, clean completion.** §5 uses *clean* in two senses and now says which is which: a **clean +findings file** is the `NO FINDINGS` signal the protocol defines, and a **clean pass** is the +predicate here, read on the logical pass with every required branch file combined, so one +branch's clean file never establishes a clean pass. **A pass is clean** when its findings carry +no in-set Blocker or Major at effective severity and **no scope-stop trigger** — the two the +absorb paragraph defines, a finding outside the assigned fix set and a finding opening a new +structural or contract question, read here in its terms and not redefined. Two things this +branch adds, because they are about answers rather than about scope: a finding **this cycle has +declined** raises no membership trigger while the set still excludes it, and a structural or +contract question **this cycle has answered** is no longer new — otherwise an answered question +would be asked again on every pass. Both predicates are properties of the findings and the set, +readable before any branch below runs, which is what makes this order executable rather than +asserted. A pass with **zero** findings is clean whatever the floor, because a floor buys further +looks at an artifact that keeps yielding findings, and one yielding none has already given what +those looks were for. **A clean pass at or above the derived floor, or a zero-finding pass, +closes the cycle**, under those final-acceptance preconditions. A plateau or tells on the closing +pass go into the closing report and never block it, because reporting "will not converge" on a +converged loop is a false report. **No other pass outcome closes a cycle**, because every other +pass leaves a required repair, a hold or a question outstanding, or has an unmet closure +precondition — and closing over any of those is the failure this ordering exists to prevent. The +one termination that is not a pass outcome is the Gate-B triviality skip, which runs no passes +and is outside this ordering. **Second, only a pass that is not a clean completion can suspend** — that order is what makes -"clean completion outranks the two-tell stop" executable rather than asserted. Three -suspensions, by the names their paragraphs use: the **scope stop**, raised by either trigger -defined above — a **membership stop** by the first, a **question stop** by the second; the -**clearly-stuck exit**; and the **two-tell stop**. A suspension waives nothing — the floor, the -Blocker/Major filter and the clean-final-pass rule stand while it does. Any non-empty set of -them can apply to one pass: **one surface, every reason reported, every question asked**, -because a reason left out is a decision made by omission. A finding surfaced solely by the -clearly-stuck reading is not a scope-stop finding; one that also carries either trigger takes -the scope stop's answers at that same surface — it is not asked twice. The two-tell stop -surfaces tells and not a finding, so no finding is surfaced by it at all. This branch defines -nothing: the triggers are the first branch's, and a reader who finds a rule here that is not -there has found a defect. +"clean completion outranks the two-tell stop" executable rather than asserted. Three suspensions, +by the names their paragraphs use and read by those paragraphs: the **scope stop**, raised by +either trigger above — a **membership stop** by the first, a **question stop** by the second; the +**clearly-stuck exit**; and the **two-tell stop**. A suspension waives nothing. Any non-empty set +of them can apply to one pass: **one surface, every reason reported, every question asked**, +because a reason left out is a decision made by omission. A finding the clearly-stuck reading +surfaces that also carries either trigger takes the scope stop's answers at that same surface, so +it is not asked twice; the two-tell stop surfaces tells and not a finding. **Third, a pass that neither closes nor suspends continues** — the loop runs another pass on the **current** artifact, revised or not. A below-floor clean pass lands here **only where no suspension applies to it**; where one does, the second branch has already taken it, because clean completion did not close the pass and only closing outranks a suspension. So does a pass -whose only findings are Minors and Nits, which are collected and never iterated and so may -leave nothing to revise. It is a branch and not an inference, because "does not close" -read alone says nothing about whether to run again. - -**The four standing duties, classified.** The **derived floor** is a precondition on -closure: it gates closing, and is discharged by the count of valid logical passes reaching -the floor with the last of them clean, or by the zero-finding exit. The **Blocker/Major-resolve -duty** is a precondition on closure and on any pass being clean: it gates both, and is -discharged, for every in-set Blocker and Major, by a validated repair **or a validated -dismissal carrying its one-line why** — the advisory rule above, unchanged — followed by a -validated pass that finds **no in-set Blocker or Major at effective severity**, which is the -clean predicate's own wording so that the duty and the predicate cannot drift apart. -**A dismissal is not a decline**: a dismissal is the author's judgement -that the finding is not true of the artifact, a decline is the user's decision that a true -finding stays outside the fix set, and only the second is an answer at a membership stop. -The **hold** a surfaced finding places on closure is part of the ordering: it gates closing -while it stands, and is discharged by the answers that finding requires, below. **A scope stop -is the only stop that holds a finding**, because it is the only one whose question is about a -finding; the health exits hold nothing, which is why their answer is continue or stop rather -than a direction on anything. **No-clean-credit** — no pass carrying a scope-stop trigger is -credited as clean — is part of the ordering and a fact about that pass, discharged by nothing: -a later pass is judged on its own findings, so an answer never closes the cycle on the pass -that surfaced the finding. **It is not a second test beside the clean predicate; it is that -predicate's second half**, which is why it is stated in the same words: the trigger makes the -pass unclean directly, so nothing has to read the act of surfacing. -**Where the other two exits stand in this duty**, said here so no reader supplies a rule for -them. The two-tell stop surfaces tells and not a finding, and was never in its domain. The -clearly-stuck exit does surface findings, and where those carry a scope-stop trigger the duty -reaches them exactly as it reaches any other. Where they do not — the two shapes the -composition paragraph names, a demoted in-set finding and a declined out-of-set one — the -pass is clean at effective severity, the two predicates reading different fields on purpose, -and **clean completion wins by D3**: the pass closes at or above the floor and continues below -it. The order decides that; the duty never needed to. - -**What a suspension asks, and what ends it.** Only a scope stop raises a **hold**, on its -finding. **The hold ends when every answer that finding requires has been given, in whichever -direction each is given** — one answer for a single-trigger finding, both the question decision -and the membership answer for one carrying both. That is **one rule with two parts**, how many -answers and which way each may go, and neither part is a test the other has to pass: a hold -only accepting could end would be the resolve duty under another name. -At a membership stop the answer is **accept** (the finding joins the -fix set, where its effective severity governs it) or **decline** (the finding stays outside, -binding for the rest of this cycle). Either is an **explicit, attributable decision on that -specific finding** — never silence, never a general remark about scope, never inferred, because -a fix set changed by inference is a fix set nobody chose. -**Membership is answered against the set as it stood at the pass that raised the question**, -not against the set as it is when the answer arrives, and an explicit answer is -required whatever the governing artifacts do meanwhile — a hold discharged by a scope change -is a hold nobody answered. A later broadening is a new fact the **next** pass reads; it never -discharges a standing hold. A decline keeps a finding out and never excuses one that is in: a -declined finding the fix set later comes to include owes resolution like any other. At a question -stop the answer is the user's decision on the question, and membership does not change: an -in-set finding then routes through its effective severity like any other, under that decision; -an out-of-set finding that opened the question is a membership stop as well and takes accept or -decline. **Effective severity routes an in-set finding in one place only** — the Mechanics -severity rule, where a Blocker or Major's obligation to resolve and a Minor or Nit's collection -without iteration are both stated, and whose discharge the resolve duty above defines — and -every branch here that puts a finding in the set hands it there rather than restating it. -**Decline is available only at a membership stop**, because that is the only stop -whose question is whether a finding belongs to the set, and a decline anywhere else would -waive work the cycle owes. The stuck and two-tell readings -raise no hold: each asks one question, **continue or stop**. Continue is that suspension's -resuming answer, and where it is the last one outstanding the loop resumes on the -artifact as revised and the fix set as the governing artifacts now assign it — where several -plans or stories govern one cycle, the union of the scopes they assign — **plus the findings -this cycle accepted into it**. A finding that a -narrowing has put outside the set surfaces at the next pass as a membership stop. **Stop -leaves the suspension standing**: the cycle stays open under its nonce, resumable by a later -continue, nothing it wrote is a closing commit, and an open cycle nobody resumes is a human's -to resolve, exactly as the nonce rules already say of open cycles. - -**Composition, and what cannot happen.** Every **question** is answered on its own, and the -loop resumes only when every answer resumes it — accept or decline at a membership stop, a -decision at a question stop, continue at the health readings; one stop answer leaves the -whole suspension standing, because a loop resumed over an unanswered question decides it by -running. **The stuck and two-tell readings raise one question between them, not two**, both -asking continue or stop, so one answer carrying every reason ends both — which is not an -exception to the sentence before it but an instance of it, there being one question there to -answer. **Two pairings cannot occur**, and -no rule ranks them: clean completion and a **scope stop**, since that stop's triggers are the -clean predicate's own second half, so a pass raising one is not clean; and a zero-finding pass -and any suspension, since it has nothing to surface, nothing regenerating, no cluster and no -require↔withdraw pair. **Clean completion and the clearly-stuck exit can**, and the overlap is -admitted rather than argued away. **Two shapes reach it**, since that exit reads -reviewer-written severity and membership while cleanliness reads effective severity and the -set: an **in-set** Blocker or Major the ceiling demotes below Major, and an **out-of-set** one -this cycle has declined, which raises no membership trigger and sits outside -the set the clean predicate reads. Either can regenerate across passes on a pass that is clean. -**The order decides both and no new rule is needed** — the pass closes at or -above the floor, ranking the exit exactly as the clearly-stuck paragraph's own precedence -sentence says, and continues below it, where nothing closes anyway. +whose only findings are Minors and Nits, which may leave nothing to revise. It is a branch and +not an inference, because "does not close" read alone says nothing about whether to run again. + +**The four standing duties, classified.** The **derived floor** is a **precondition on closure**: +it gates closing, discharged by the count of valid logical passes reaching it with the last of +them clean, or by the zero-finding exit. The **Blocker/Major-resolve duty** is a **precondition +on closure and on any pass being clean**: Mechanics Severity, scoped to the assigned fix set, +states what it demands and what discharges it, and a validated pass finding **no in-set Blocker +or Major at effective severity** is what shows it discharged — the clean predicate's own wording, +so the two cannot drift. **A dismissal is not a decline**: a dismissal is the author's judgement +that the finding is not true of the artifact, a decline the user's decision that a true finding +stays outside the set, and only the second is an answer at a membership stop. The **hold** a +surfaced finding places on closure **participates in the ordering**: it gates closing while it +stands, and is discharged by the answers that surface requires. **It attaches to every surfaced +finding, whichever suspension surfaced it** — clean completion creates none, because it wins +before anything is surfaced. **No-clean-credit** — no pass carrying a scope-stop trigger is +credited as clean — also participates, and is a fact about that pass that nothing discharges, a +later pass being judged on its own findings. It is not a second test beside the clean predicate +but that predicate's second half, which is why it is stated in its words. + +**What a suspension asks, and what ends it.** **A hold ends when every answer its finding +requires has been given, in whichever direction each is given** — one answer for a +single-trigger finding, both for one carrying both. That is **one rule with two parts**, how many +answers and which way each may go, and neither is a test the other has to pass. At a **membership +stop** the answer is **accept**, the finding joining the fix set where Mechanics Severity governs +it, or **decline**, the finding staying outside and binding so for the rest of this cycle unless +the user explicitly reverses that decision. Either is an **explicit, attributable decision on +that specific finding** — never silence, never a general remark about scope, never inferred, +because a fix set changed by inference is a fix set nobody chose. **Membership is answered +against the set as it stood at the pass that raised the question**: a later broadening is a new +fact the **next** pass reads and never discharges a standing hold, a hold discharged by a scope +change being a hold nobody answered. At a **question stop** the answer is the user's decision on +the question and membership does not change; an out-of-set finding that opened one is a +membership stop as well. **Decline is available only at a membership stop**, that being the only +stop whose question is whether a finding belongs to the set. The **clearly-stuck and two-tell +readings** ask **continue or stop**. **Continue consumes the reading that raised the +suspension**: a further health suspension needs that reading recomputed over a pass run after the +answer, which is new data — so continue produces a distinct next state even on an unrevised +artifact, and the same reading cannot return the same stop unanswered. **Stop parks the cycle**: +open, not running, spending no passes, restarted only by an explicit later continue — a distinct +state from the suspended-awaiting-answer one it was in before the answer. Nothing a parked cycle +wrote is a closing commit, and a parked cycle nobody restarts is a human's to resolve, exactly as +the nonce rules already say of open cycles. + +**Composition, and what cannot happen.** Every **question** is answered on its own and the loop +resumes only when every answer resumes it — accept or decline at a membership stop, a decision at +a question stop, continue at the health readings; one stop answer parks the whole suspension, +because a loop resumed over an unanswered question decides it by running. **The clearly-stuck and +two-tell readings raise one question between them, not two**, both asking continue or stop, so +one answer carrying every reason ends both — an instance of the sentence before it, not an +exception. **Two pairings cannot occur**, and no rule ranks them: clean completion and a **scope +stop**, since that stop's triggers are the clean predicate's own second half, so a pass raising +one is not clean; and a zero-finding pass and any suspension, since it has nothing to surface, +nothing regenerating, no cluster and no require↔withdraw pair. **Clean completion and the +clearly-stuck exit can**, and the overlap is admitted rather than argued away: that exit reads +the reviewer-written field while cleanliness reads the effective one, so an **in-set** Blocker or +Major the ceiling demotes below Major, or an **out-of-set** one this cycle has declined, can +regenerate across passes on a pass that is clean. **The order decides it and no new rule is +needed** — at or above the floor the pass closes, ranking the exit exactly as the clearly-stuck +paragraph's own precedence sentence says; below it the pass **suspends**, clean completion having +not closed it. ``` Where the block maps onto the criteria: the three branches are AC 4, the duties paragraph -AC 2, the composition sentences AC 1. The scope stop's two triggers are `b11` (membership) and -`b13` (question) read separately, because an in-set finding that opens a question can neither -join nor stay outside the set and needs its own answer. Sentences the block points at rather -than restating (`a13` as §5(a) replaces it): C:116–118, C:760–764, C:176, and the -clearly-stuck precedence sentence, kept verbatim under **D3**. +AC 2, the composition sentences AC 1. + +**One authority per rule — the check finding 14 of pass 11 asked for.** The block cites eight +rules and defines none of them. Each has exactly one definition in the shipped text, and where +that definition had to change to agree with the ordering, it changed **at its source** (§4, §5) +rather than being restated here: + +| Rule the block cites | Its one definition | +|---|---| +| membership trigger | `b11`, the absorb paragraph (C:206 / W:413), qualified by §5(b) | +| question trigger | `b13`, the absorb paragraph (C:209–213 / W:416–420) | +| the assigned fix set | `b7`, the absorb paragraph (C:200–202 / W:407–409), edited by §5(b) | +| effective vs reviewer-written severity | the (g) replacement (§5(g)) | +| what a Blocker, Major, Minor or Nit demands | Mechanics · Severity (C:783–784 / W:969–970), scoped by §4 item 3 | +| the derived floor and final-acceptance preconditions | the floor section (C:116–118, C:760–764) | +| the clearly-stuck reading and its precedence | the clearly-stuck paragraph (C:225–246 / W:428–449), `c9` kept verbatim under **D3** | +| the two-tell threshold | the five-tells paragraph (C:263–268 / W:467–472), qualified by §5(e) | + +The scope stop's two triggers are `b11` and `b13` read separately, because an in-set finding +that opens a question can neither join nor stay outside the set and needs its own answer. **Two things the ordering names and does not define, both the successor's** (§9): what makes a later finding *the same one* this cycle declined, and what form carries an answer into the commit body. The ordering says what an answer does — a decline keeps a finding out of the set for the cycle, an acceptance puts one in until the cycle closes — and **D9b** and **D9** say how -it is recognised and written down. Until the successor ships, an answer binds within the session -that made it and the reviewer re-raises what it cannot read, which is the ordinary route and not -a special one. +it is recognised and written down. **Two consequences are stated in the block rather than left +to be discovered**: until a sameness rule ships, a line in one branch file and a line in the +other are distinct findings for holds and answers, so a `full` Gate-B pass may ask twice rather +than risk resuming over an unanswered one; and an answer binds within the session that made it, +the reviewer re-raising what it cannot read, which is the ordinary route and not a special one. --- ## 4. The standing sentences edited at their source -Six, and no more — working around any of them would ship two instructions that disagree. Line -numbers are C's; W's are re-read at execution. **Every OLD fragment quoted here is a single line -in the file it is grepped from**, verified at 1 in both copies (§7). - -1. **The findings-file protocol's clean sentence** (C:329 / W:523), inside the gate-prompt - template both gates paste. OLD: "A clean pass is the single body line `NO FINDINGS` with - `END OF FINDINGS (0 total)`." NEW: "A **clean findings file** is the single body line - `NO FINDINGS` with `END OF FINDINGS (0 total)`." It describes a file and always did; the - word *pass* in it is what makes §3's predicate look like a redefinition instead of the - other sense. Old condition — what a reviewer writes when it finds nothing — unchanged. - -2. **The Gate-A clean-signal sentence** (C:565–566 / W:756–757; the two copies wrap it - differently, so the shared fragment is what is quoted). OLD: "…when a pass is clean — the - explicit clean signal is what lets you exit the loop:". NEW: "…when a pass finds nothing — - the explicit signal is what lets a pass be read as clean without inspecting it further:". +Six, and no more — working around any of them would ship two instructions that disagree. + +**How each OLD is cited, since an earlier revision got this wrong three times.** A sentence +quoted here is quoted whole, and several of them **wrap across two lines** in one or both +copies, so each item gives the sentence's real line range in **both** copies and names the +**single-line substring the check actually counts**. A displayed sentence spanning a wrap +cannot be counted with `grep -F`, and quoting one as though it could is what made three +earlier checks read as false reds. Every counted substring below is verified at 1 in both +copies (§7). + +1. **The findings-file protocol's clean sentence**, inside the gate-prompt template both gates + paste. Sentence at **C:328–329 / W:522–523**, the word "A" ending the first line; **counted + substring `clean pass is the single body line`, C:329 / W:523**. OLD: "A clean pass is the + single body line `NO FINDINGS` with `END OF FINDINGS (0 total)`." NEW: "A **clean findings + file** is the single body line `NO FINDINGS` with `END OF FINDINGS (0 total)`." It describes + a file and always did; the word *pass* in it is what makes §3's predicate look like a + redefinition instead of the other sense. Old condition — what a reviewer writes when it + finds nothing — unchanged. + +2. **The Gate-A clean-signal sentence**. Sentence at **C:565–566**, wrapped, and **W:757**, + whole on one line; **counted substring `when a pass is clean`, C:565 / W:757** — the part + that is single-line in both, which is why the check uses it and not the displayed sentence. + OLD: "…when a pass is clean — the explicit clean signal is what lets you exit the loop:". + NEW: "…when a pass finds nothing — the explicit signal is what lets a pass be read as clean + without inspecting it further:". Old condition: a `NO FINDINGS` file is what permits loop exit. Kept: the signal and why it is demanded. **Replaced**: it makes a pass readable as clean rather than being the only way to be clean, since a pass carrying Minors alone is clean under the ordering and could never @@ -270,8 +255,9 @@ in the file it is grepped from**, verified at 1 in both copies (§7). 761, 769, 827; W:324, 914, 936, 947, 955, 1011) are the closure sense the ordering defines and are correct as they stand — counted and checked, not assumed. -3. **The Severity bullet's resolve duty** (C:783–784 / W:969–970), which is the one place the - duty is stated and the only one without a scope. OLD: "- **Severity:** Blocker +3. **The Severity bullet's resolve duty**, the one place the duty is stated and the only one + without a scope. Sentence at **C:783–784 / W:969–970**, wrapped in both; **counted substring + `both must resolve. Minor`, C:784 / W:970**. OLD: "- **Severity:** Blocker (wrong/unsafe/breaks invariant) · Major (design flaw → rework) → both must resolve. Minor · Nit → collect, never iterate." NEW: "- **Severity:** Blocker (wrong/unsafe/breaks invariant) · Major (design flaw → rework) → both must resolve, **for every finding in the @@ -286,9 +272,10 @@ in the file it is grepped from**, verified at 1 in both copies (§7). resolve rule" (`c17`) needs no edit: it names the rule, and the rule now carries its own scope. -4. **The lens paragraph's unchanged-list** (C:653–655 / W:839–841), which asserts of two rules - this change alters that they are unchanged. OLD, the single-line fragment: "clean-final-pass - rule are unchanged." NEW: "clean-final-pass rule are unchanged **by the lens sets**, which is +4. **The lens paragraph's unchanged-list**, which asserts of two rules this change alters that + they are unchanged. Sentence at **C:653–654 / W:839–840**; **counted substring + `clean-final-pass rule are unchanged.`, single-line at C:654 / W:840**. OLD: + "clean-final-pass rule are unchanged." NEW: "clean-final-pass rule are unchanged **by the lens sets**, which is what this paragraph is about — the closure ordering above does change both, scoping the filter to the assigned fix set and defining a clean pass at effective severity, and says so there." Old conditions: lenses change what a pass asks and never how many passes a cycle @@ -297,8 +284,9 @@ in the file it is grepped from**, verified at 1 in both copies (§7). the three are changed by the ordering and an unscoped claim leaves both copies denying an edit they carry); the floor is not among them (**kept**). -5. **The Gate-A cadence** (C:573 / W:764), which makes a revision unconditional between - passes. OLD: "Each pass: validate, revise, re-run." NEW: "Each pass: validate, revise +5. **The Gate-A cadence**, which makes a revision unconditional between passes. **Counted + substring `Each pass: validate, revise, re-run.`, single-line at C:573 / W:764.** + OLD: "Each pass: validate, revise, re-run." NEW: "Each pass: validate, revise **where the severity and scope rules require a repair**, re-run." Old condition: every pass is followed by a revision before the next. Kept: the cadence and its order — validate first, re-run last. **Replaced**: the revision is conditional, because the ordering's @@ -308,8 +296,9 @@ in the file it is grepped from**, verified at 1 in both copies (§7). 6. **The unknown-start fallback** (C:153–167 / W:360–374, `i4`–`i8`, extended at `i12`'s invitation) — an addition to a list rather than a change of meaning, so it is checked by - presence and not as a pair. After the single-line anchor "duty owed, and the nonce duties at - their strictest" the list gains "**every suspension binding**". `i12`'s sentence stays as + presence and not as a pair. After the anchor **`duty owed, and the nonce duties at their + strictest`, single-line at C:157 / W:364**, the list gains "**every suspension binding**". + `i12`'s sentence stays as written; this is the addition it invites, and it is what this change owes that list: a cycle that cannot establish its starting rules treats a suspension as binding rather than as advisory, which is the strict reading of the rules §3 ships. @@ -344,8 +333,25 @@ police. Changed: the ordering does restate closure rules, deliberately and as th authority, so a categorical reading would leave the shipped text contradicting itself — the failure `a13` exists to prevent, arriving from the other direction. -**(b) What a loop absorbs** (C:195–223 / W:402–426) — **three sentence edits**, the triggers -stay. `b12` OLD: "…and it resumes the moment the user says whether the set now includes it." +**(b) What a loop absorbs** (C:195–223 / W:402–426) — **five sentence edits**. This passage owns +the two triggers and the fix set, so where the ordering needs them qualified, they are qualified +**here**, which is what keeps one definition per rule. + +`b7`, the fix-set definition, **counted substring `scope the approved story or plan assigns to +this cycle`, single-line at C:201 / W:408**. OLD: "…it is the scope the approved story or plan +assigns to this cycle, plus repair obligations you already accepted in earlier passes." NEW: +"…it is the scope **every governing story or plan** assigns to this cycle — their union where +several do — plus repair obligations you already accepted in earlier passes, **minus any finding +this cycle has declined**." + +`b11`, the membership trigger, **counted substring `loop like any other out-of-scope finding**, +even when it opens no new question at all`, single-line at C:206 / W:413**. OLD: "**A correction +that leaves that set stops the loop like any other out-of-scope finding**, even when it opens no +new question at all…" NEW: "**A correction that leaves that set stops the loop like any other +out-of-scope finding**, even when it opens no new question at all — **unless this cycle has +already declined that same finding, which put it outside by the user's own answer**…" + +`b12` OLD: "…and it resumes the moment the user says whether the set now includes it." NEW: "…and it resumes once the user has said whether the set now includes it — accepting or declining it — **together with every other answer that pass's suspensions require**, under the closure ordering above." `b17`–`b18` OLD: @@ -353,19 +359,31 @@ closure ordering above." `b17`–`b18` OLD: the clean-final-pass rule all stand, and the loop resumes on the revised artifact once the question is answered — what the stop prevents…" NEW: "Stopping this way is a **suspension** under the closure ordering above, and the loop resumes on the revised artifact **once every -answer that pass's suspensions require has been given** — what the stop prevents…". Third, in -**W only**, "…by its severity exactly as the severity rule already says…" becomes "…exactly as -Mechanics already says…", matching C for the reason §6's first row gives. - -Accounting: `b1`–`b11`, `b13`–`b16` kept. Two are worth naming. `b6` fixes +answer that pass's suspensions require has been given** — what the stop prevents…". Fifth, in +**W only**, `b3`: "…by its severity exactly as the severity rule already says…" becomes +"…exactly as Mechanics already says…", matching C for the reason §6's first row gives. + +Accounting: `b1`, `b2`, `b4`–`b6`, `b8`–`b10`, `b13`–`b16` kept. Five are worth naming. +`b3` is an **edited cross-reference whose operative condition is kept** — W's pointer changes +target while what it points at is unchanged, so it is not in the kept range above and it is +checked in §7 like any other edit. `b6` fixes the assigned set **before the pass being answered**, and the ordering's membership answer is read against that same set rather than against a later one, so the two agree and `b6` needs no edit. An earlier revision of this spec had membership re-read when the answer arrived, which `b6` and **D4** both forbid — **D4** requiring a user answer for the hold to end at all — and -the reversal is recorded here rather than left as a silent narrowing. `b11` and `b13` are the -two trigger descriptions and **stay in their own words**, checked equivalent to the block's -definitions rather than merged into them, because editing them would rewrite the passage this -change deliberately leaves standing; §6 is where that check runs. `b12` **replaced** — old: an *immediate* +the reversal is recorded here rather than left as a silent narrowing. `b7` **replaced** — old: +the scope *the* approved story or plan assigns, singular, with no subtraction. Kept: that the +set is fixed before the pass being answered and carries the obligations already accepted. +Changed: it is the union where several artifacts govern one cycle, and it subtracts this cycle's +declines — the singular reading left a Gate-B cycle fed by several plans, or a Gate-A artifact +citing several stories, with no defined set on its **first** pass, before any suspension could +have raised the question. `b11` **replaced** — old: any out-of-set finding stops the loop, +unconditionally. Kept: the trigger and its reason. Changed: a finding this cycle has already +declined does not raise it again, because unqualified `b11` and the ordering's clean predicate +decide a re-raised declined finding in opposite directions, and re-asking an answered question +on every pass is the non-idempotent path AC 4 forbids. `b13` is the question trigger and +**stays in its own words**, checked equivalent to the block's reference rather than merged into +it; §6 is where that check runs. `b12` **replaced** — old: an *immediate* resume on the membership answer alone. Kept: that the membership answer is what the stop asks for and that either direction ends it. Changed: the resume waits for every answer the pass's suspensions require, because a loop resumed over an unanswered question decides it by running. @@ -398,21 +416,25 @@ zero-finding exception); `c14` moved **to the ordering's continue branch** — t Minor-below-floor pass that "keeps looping" is precisely one that neither closes nor suspends, and the branch states it rather than leaving it to be read out of a negation; `c15`, `c17` kept in the pointer sentence, `c17`'s "resolve rule" now reading with the assigned-fix-set -boundary §4 item 3 gives it, so the pointer needs no edit of its own; `c16` **replaced** — -old: a clearly-stuck surface leaves its finding open, which read alone gives this exit a hold. -Kept: that the findings are still open and that the resolve duty is what keeps them so. -Changed: the **hold** belongs to the scope stop, the only stop whose question is about a -finding, while this exit asks continue or stop. Authority **D1**, which makes all three exits -suspensions with their own transitions; `c18` **replaced** — old: no pass is credited as clean +boundary §4 item 3 gives it, so the pointer needs no edit of its own; `c16` **kept** — a +clearly-stuck surface leaves its finding open, and the ordering's hold attaches to **every** +surfaced finding, whichever suspension surfaced it, discharged by the answers that surface +requires. An earlier revision of this spec narrowed the hold to scope stops, which contradicted +both `c16` and the story's own third standing duty; the reversal is recorded here rather than +left as a silent narrowing, and it is why the duties paragraph now says "every surfaced +finding" instead of naming one stop; `c18` **replaced** — old: no pass is credited as clean on a clearly-stuck surface, unconditionally. Kept: the rule wherever it decides anything — a pass carrying a scope-stop trigger is unclean, and this exit is reached only on a pass that did not close. Changed: where the exit's regenerating findings are in-set and the ceiling demotes -them below Major, the pass is clean at effective severity and closes at or above the floor. -Authority **D3** — under the old reading D3's own preserved sentence and `c18` decide that pass -in opposite directions. The two-tell stop was never in either condition's domain, surfacing +them below Major, the pass is clean at effective severity — it closes at or above the floor and +**suspends below it**, clean completion having not closed it. Authority **D3** — under the old +reading D3's own preserved sentence and `c18` decide that pass in opposite directions. The +two-tell stop was never in this condition's domain, surfacing tells and not a finding; `c19` **replaced** — old: the loop resumes on whatever the user -decides, unconditionally; new: it resumes on a resuming answer and a stop leaves the suspension -standing, because an unconditional resume is the stop-with-no-transition path AC 4 forbids; +decides, unconditionally; new: it resumes on a resuming answer, and a stop **parks** the cycle — +open, not running, restarted only by an explicit later continue — because an unconditional +resume is the stop-with-no-transition path AC 4 forbids and a post-stop state identical to the +pre-answer one is that same path wearing a different name; `c20` **dropped** — it argued that reading the exit as "stop instead of fixing" would compete with the resolve duty, and the block states that argument's conclusion as a rule instead. @@ -496,12 +518,24 @@ found are handled as follows: The parity check at execution: extract each edited passage from both files by its lead phrase and `diff` them; the only differences permitted are the rows marked as-is, and the result is -part of the evidence entry (§7). Any other difference is a defect, not a wording choice. The -same extraction also compares `b11` and `b13` against the block's two trigger definitions, per -§5(b) — a parity check between copies and an equivalence check within each. The correspondence -it confirms: `b11`'s "a correction that leaves that set" and "out-of-scope finding" is the -block's *outside the current assigned fix set*, and `b13`'s "opens a **new structural or -contract question**" is the block's question trigger word for word. +part of the evidence entry (§7). Any other difference is a defect, not a wording choice. + +The same extraction runs a second, different check within each copy: that `b11` and `b13` as +edited say what the block cites them as saying. **It compares the complete predicates, not a +shared phrase** — comparing only the phrase both carry is what let an earlier revision call two +wordings equivalent while one of them lacked the decline exception the other had. Each side is +read whole: + +| Block cites | `b11`/`b13` as edited must carry | +|---|---| +| a finding **outside the assigned fix set** | `b11`'s "a correction that leaves that set… like any other out-of-scope finding" | +| …**and not one this cycle has declined** | `b11`'s "unless this cycle has already declined that same finding" (§5(b)) | +| a finding **opening a new structural or contract question** | `b13`'s "opens a **new structural or contract question**", word for word | +| in-set or not | `b13` carries no membership condition, and adding one would narrow it | + +A difference in either direction is a defect: a condition in the block and not in `b11`/`b13` +ships two triggers that disagree, and one in `b11`/`b13` and not in the block means the block +cites a rule it has not read. --- @@ -543,7 +577,24 @@ not: quoted three of these across their wraps. That is a constraint on how the plan writes the edits, not on what they mean; - §4 item 6, the unknown-start addition, by presence alone since it adds to a list rather than - changing a meaning: `every suspension binding` count **1** in each and **0** in the parent. + changing a meaning: `every suspension binding` count **1** in each and **0** in the parent; +- **the passage edits in §5, each as an OLD/NEW pair on the same terms.** They change standing + meanings exactly as §4's do, and checking only §4 would leave both copies able to carry the + new ordering beside stale resume, trigger, fix-set, surfacing and two-tell rules while every + stated count and the parity diff still passed. Seven pairs, each OLD counting **1** in the + parent tree and **0** in the working tree and each NEW the reverse: `final pass must be + clean — if the pass at the floor still finds Blocker/Major, keep going until` (`a17`–`a19`, + C:133 / W:340); `written, and none of them is restated` (`a13`, C:130 / W:337); `scope the + approved story or plan assigns to this cycle` (`b7`, C:201 / W:408); `loop like any other + out-of-scope finding**, even when it opens no new question at all` (`b11`, C:206 / W:413); + `the moment the user says whether the set now includes it.` (`b12`, C:208 / W:415); + `**Below the floor nothing closes**, and a zero-finding pass remains` (the (c) trim, C:238 / + W:441); and `**Any two present makes stop-and-surface mandatory, not discretionary**` (`e7`, + C:266 / W:470). **`b3` is W-only** and is checked in W alone — `severity exactly as the + severity rule already says` counting **1** in the parent W and **0** in the working W, against + `severity exactly as Mechanics already says` counting **0** then **1**; C already carries the + target wording at C:198 and must be unchanged, which is what makes this an alignment rather + than an edit to both. If the claim "the ordering ships in both copies" were false, one working-tree count would be **0** or the parent-tree counts would not differ from it. The wiring can produce that @@ -555,17 +606,29 @@ removes it. Nothing is claimed as "contradictory". **The named verification of the risk path** (story AC 4) is a **next-state table**, in the plan and quoted by the closing commit body. **Rows** — every stop the shipped text names: membership stop, question stop, a finding carrying both triggers, clearly-stuck exit, two-tell stop, a hold -awaiting its answer, accept, decline, a stop answer, two or three suspensions at once, a +awaiting its answer, accept, decline, **a continue answer on an artifact left unrevised** — +whose next state is a further pass, the row that tests that continue consumes the reading — +**a stop answer** — whose next state is the **parked** cycle, distinct from the +suspended-awaiting-answer state it was in before, the row that tests AC 4 at the one transition +that produces no pass — two or three suspensions at once, a below-floor clean pass with no suspension, **a below-floor clean pass carrying a health -reading** — whose next state is the suspension, the row that tests the third branch's +reading** — whose next state is the **suspension**, the row that tests the third branch's qualification — a zero-finding pass, and the unknown-start fallback; **and the stateful -transitions**: a finding this cycle declined re-raised on a later pass, **the fix set broadened +transitions**: **a finding this cycle declined re-raised on a later pass** — whose next state is +no membership trigger and a pass still eligible for clean, the row that tests `b11`'s +qualification — **a structural question this cycle answered, re-raised** — whose next state is +no question stop — an **accepted** finding re-raised before repair, **the fix set broadened while a hold is still awaiting its answer** — whose next state is the hold still standing, the -row that tests the frozen reading — a fix set narrowed so an in-set finding falls outside it, -and a `full` Gate-B pass with one branch clean and the other carrying an in-set Blocker. -**Columns** — the **governing-scope state** (the fix set as currently assigned) and the -**answers already given in this cycle**, beside the user's answer for this row, then the input -that ends the row, the state afterwards, and the shipped line the row reads, in both copies. +row that tests the frozen reading — **a confirmed profile change while a hold stands** — whose +next state is the hold still standing *and* the further pass under the current profile that a +profile change costs, run against both answer directions, since final acceptance re-reads every +profile and a hold surviving one is not the same claim as a hold surviving a scope change — a +fix set narrowed so an in-set finding falls outside it, and a `full` Gate-B pass with one branch +clean and the other carrying an in-set Blocker. +**Columns** — the **governing-scope state** (the fix set as currently assigned), the +**answers already given in this cycle**, and the **profile as currently read**, beside the +user's answer for this row, then the input that ends the row, the state afterwards, and the +shipped line the row reads, in both copies. **The oracle.** A row **fails** when its required answer does not produce a **distinct** resumable or closed state — the same stop returning, whether or not an input was consumed — @@ -586,6 +649,16 @@ differences observed, that each is one of the permitted rows, and the `b11`/`b13 result; and the next-state table's location in the plan plus its row count. It is revalidated before every Gate-B re-review and before the closing amend, as §5 requires. +**One observability residual, stated because the lens set asks for it and nothing here answers +it.** A closing commit body records that a cycle closed and what its curve was; it records +**nothing about which exit the cycle took** — whether a scope stop, a clearly-stuck surface or +two tells ever suspended it, what was asked, or how it was answered. A reader of history +therefore cannot audit that every suspension was answered before the cycle closed. Neither the +evidence entry above nor the per-pass curve supplies this, and saying otherwise would be the +overclaim `AGENTS.md` names as this repo's most persistent defect. The transport that could +carry it left with the record (§9), so **this belongs to the successor**; until then it is an +admitted gap, not a guarded one. + --- ## 8. AGENTS.md invariants touched @@ -596,8 +669,15 @@ before every Gate-B re-review and before the closing amend, as §5 requires. closes, that a zero-finding pass is clean whatever the floor, and that decline is available only at a membership stop. The exemption was wrong twice over: item 6 admits no "settled elsewhere" clause, and a scaffolded copy cannot reach the story the reasons were said to live - in. Then item 8 (token-lean — the block replaces closure sentences rather than adding beside - them) and item 3 (the stop answer is a named, resumable state). + in. Then **item 8 (token-lean), which an earlier revision claimed on the wrong ground** — it + said the block replaces closure sentences rather than adding beside them, while the block was + in fact restating triggers, duties, preconditions and the severity answer that their own + paragraphs still defined, which is two authorities per copy and the drift had already begun. + The claim now rests on what the block does: it is authoritative for the evaluation order and + for closure and **cites** every other rule where that rule is defined, so each has one + definition in the shipped text and §3's table is the check; where a cited rule had to change + to agree, it changed at its source (§4, §5) rather than being restated. Then item 3 (the stop + answer produces a named state, **parked**, with its own restart transition). - **Don't: "Never replace a decision procedure without accounting for its old conditions."** Satisfied by §5, against the committed inventory, with an entry for every inventoried passage including the four this narrowing no longer edits. @@ -625,17 +705,28 @@ decision of 2026-09-10, with the evidence in §1: finding (**D9b**), and the unverified-assertion reading (**D9c**); - the pass-4 report's unavailable-history block (**D10**) and the checkout-root condition that report grew; -- the rollback reading — what an open cycle owes when a revert removes the text it started - under; +- the rollback reading — **whether what an open cycle already has is enough** when a revert + removes the text it started under. It is not a gap here, and an earlier revision of this spec + wrongly implied one: the activation paragraph this change ships and leaves unchanged already + gives that cycle a defined transition — a cycle whose starting rules cannot be established + takes the **stricter reading**, which §4 item 6 extends with "every suspension binding", so it + neither closes nor is left without a next state. What is open is whether the stricter reading + is *sufficient* for a rollback specifically, since no record identifies the rule revision a + cycle started under. That question is the successor's; the transition exists meanwhile; - the slot-discriminator dissolution deferred here by Plan C's Tasks 19 and 20; - **the partial-adoption guard for the whole set.** It was built as a "closure-record contract" naming the record, with a marker on every mergeable hunk, and it leaves with the record it was named for. The narrowed set is still mutually dependent — the ordering's clean predicate needs - §4 item 3's boundary, and its two senses of *clean* need items 1 and 2 — and **nothing here - guards that**. It belongs with the successor because the two sets are adopted together in - practice and one guard should name both; until then a partial `/workflow-init` merge of these - edits is caught only by the existing one-contract paragraph's semantic coherence rule, which - does not name them. Stated as a residual, not as protection. + §4 item 3's boundary, its two senses of *clean* need items 1 and 2, and its triggers need + §5(b)'s `b7` and `b11` — and **nothing catches a partial adoption of it**. Said exactly: + the live one-contract paragraph (C:880–890 / W:1063–1074) names the nonce, the slots, the + provenance line, the curve, the carry rule and the unknown-start semantics, and **it does not + name this block or any of its coupled edits**, so its coherence rule does not reach them. An + earlier revision of this spec said that rule caught them; it does not, and a downstream merge + can take the clean predicate without the scoped resolve duty, or either sense of *clean* + without the other, and run. **That is an admitted unsafe state, not a guarded one.** The guard + belongs with the successor because the two sets are adopted together in practice and one guard + should name both. **Out of scope and parked**, unchanged: From d5a4f14ece0fce3ccca18e5d90187adbb5d42c6c Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Thu, 10 Sep 2026 20:24:30 +0200 Subject: [PATCH 017/181] docs(stories): terminate repeated harness findings by criterion, not by agent MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Daniel brought a criterion back from a sibling project, where the question had been whether the workflow needs a triage agent holding a whole-project overview. The answer he endorses is that the missing thing is a criterion, not an agent: a product Major always blocks; a harness Major blocks only where it makes a claim about the committed code vacuous; and the second finding of the same shape against the same mechanism ends that mechanism's rounds. The kit already carries most of the first two in the Mechanics severity carve-out, and detects harness drift through the instrument tell. The third has no counterpart, and the severity ceiling never says what a non-vacuous harness Major becomes — "collect" names no destination. Evidence is this repository's own Gate-A cycle awsf1ec771: twelve passes, no clean pass, and two mechanisms each producing one shape of finding for four rounds while every individual repair was cheap enough to absorb. Intake only. The profile is proposed and unconfirmed. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- ...09-10-harness-finding-termination-story.md | 160 ++++++++++++++++++ 1 file changed, 160 insertions(+) create mode 100644 docs/superpowers/stories/2026-09-10-harness-finding-termination-story.md diff --git a/docs/superpowers/stories/2026-09-10-harness-finding-termination-story.md b/docs/superpowers/stories/2026-09-10-harness-finding-termination-story.md new file mode 100644 index 0000000..8af1242 --- /dev/null +++ b/docs/superpowers/stories/2026-09-10-harness-finding-termination-story.md @@ -0,0 +1,160 @@ +# Harness findings: what blocks a commit, and when one mechanism's loop ends — Story + +**Date:** 2026-09-10 · **Size:** story +**Risk:** *(proposed)* high · **Security:** *(proposed)* none · **Validation:** *(proposed)* battery+check+verification + +> **DRAFT — the profile above is proposed, not confirmed.** Per §5 a profile is confirmed by the +> human, and until it is this story is not executable. Nothing depends on it yet. The proposal's +> reasons are in §5. + +## 1. Problem statement + +**Daniel's question, from another project, was whether the workflow needs an agent that keeps a +whole-project overview and makes triage calls so review loops stop spinning on edge-case +definitions.** The answer he brought back and endorses is that the missing thing is **not an agent +but a criterion** — a second agent would need the first one's entire context, hold no authority +because the human decides, and have to read the code to triage honestly; anything triaging on +plausibility alone dismisses the real defect as an unrealistic edge case. + +**The criterion he proposes is three lines:** + +1. A Major **in the product** blocks the commit, always. +2. A Major **in the harness** blocks only if it makes a claim about the code being committed + vacuous; otherwise it becomes a ticket with a tripwire. +3. **The second finding of the same shape against the same mechanism ends the loop for that + mechanism**: narrow the claim, print the residual, escalate by ticket rather than by another + round. + +**Two of the three already have most of a counterpart, and a design that ignores that would +rebuild them.** `CLAUDE.md` §5's severity procedure carries an **instrument carve-out** — "an +instrument finding keeps its severity whenever it shows the instrument changes what a gate +concludes about product behaviour — a false green, and equally a false red or a check blocking a +valid change" — which is close to lines 1 and 2. The five tells include "findings clustering on the +**instrument** rather than on product behaviour", so the kit **detects** harness drift and makes +stop-and-surface mandatory at two tells. And `dev-workflow:harden-finding` with the fingerprinted +ledger already handles recurrence, which AGENTS.md describes as "a recurring finding escalates one +rung harder (prose → lint → type → test)". + +**What none of them does is the part that costs rounds.** + +- **The severity procedure decides severity and never says what a non-vacuous harness Major + *becomes*.** It sets a ceiling — "the finding is Minor or below: collect, never iterate" — and + "collect" names no destination. Nothing is ticketed, nothing is tripwired, and the finding leaves + no trace outside that pass's report. +- **The tells stop the whole loop, not one mechanism.** Two tells hand the entire cycle to the + human. There is no way to end one mechanism's rounds while the loop keeps working on everything + else. +- **Line 3 has no counterpart at all.** `harden-finding` escalates a *fix* after a finding is + closed; it is not a move a running loop can make to stop re-examining a mechanism. + +**The concrete miss is measurable in this repository's own cycle.** Gate-A spec cycle nonce +`awsf1ec771` on the loop-rule consolidation design ran twelve passes without a clean pass. +Blocker+Major by pass: **20, 15, 10, 11, 15, 13, 9, 8, 12, 17, 14, 14.** The pass-by-pass record is +`.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md`, which is gitignored — hence the figures +quoted rather than only cited. + +**Two mechanisms inside that cycle each produced the same shape of finding for four rounds, and +each was repaired one instance at a time.** The spec's own verification assert list produced +"an assertion that does not detect what it claims" at passes 8, 11 and 12 — one found by the +reviewer, three more by the revision's own audit, then four more, then three more. Separately, the +one-authority principle produced four "the block still restates what it cites" findings across +passes 11 and 12. Under line 3 each would have ended after its second instance, with the claim +narrowed and the residual printed. + +**Why it kept happening, in Daniel's words: cheap to fix is not the criterion, and treating it as +one is exactly the drift a written rule prevents.** Every individual repair was a few lines, so +every one was absorbed rather than ending the round. + +## 2. Desired outcome + +**A reviewer and an author can tell, without inferring it, whether a given Major blocks this +commit or becomes tracked work** — and when it becomes tracked work, where that work is recorded +and what would catch its return. + +**One mechanism's repeated findings can be ended without ending the loop.** A running cycle gains a +move between "absorb another round of this" and "hand the whole cycle to the human", so a mechanism +that keeps producing the same shape of finding stops consuming rounds while the loop continues on +everything else. + +**Out of scope**, named so nothing absorbs them: +- **The severity definitions.** Blocker, Major, Minor and Nit keep their current meanings; this + story changes what follows from a severity, not what earns one. +- **The loop-health counts and the two-tell threshold.** `CLAUDE.md` §5 already defers how a + *demoted* finding bears on "the per-pass counts, the finding clusters and the stop thresholds" to + `docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md`. That story owns counting; + this one owns blocking and terminating. Neither reopens the other. +- **Any triage agent.** Considered and rejected, with the reason recorded above: it would need the + reviewing agent's whole context, hold no authority, and triage on plausibility unless it read the + code. +- **Hook code** (anything under `plugins/dev-workflow/hooks/`). + +## 3. Acceptance criteria + +- [ ] **Every Major has exactly one stated disposition, determinable from the finding and the diff + rather than from judgement about intent.** A reader can decide whether a given Major blocks + this commit or becomes tracked work without consulting a person. +- [ ] **A harness Major that is not vacuum-making has a named destination.** "Collect" resolves to + a specific artifact a later reader can open, and the shipped text says which one. Falsifiable + by reading: a disposition that ends at "collect" with no destination fails. +- [ ] **A tracked harness finding carries something that would catch its return**, and the shipped + text says what that is and who writes it. A ticket with no tripwire is the state this + criterion exists to prevent. +- [ ] **"The same shape against the same mechanism" is defined precisely enough that two readers, + given the same two findings, agree on whether they match.** Checkable by applying the + definition to the two instances in §1 — the assert-list findings and the one-authority + findings — and to a deliberate near-miss. +- [ ] **"Narrow the claim, print the residual" obliges specific written output**, and the shipped + text names it. A reader can tell whether an author who ended a mechanism's loop actually did + so or merely stopped looking. +- [ ] **Per-mechanism termination and the existing five tells do not duplicate or contradict each + other.** Both copies say which applies when both would fire, and neither weakens the two-tell + mandatory stop. +- [ ] **Every condition of the replaced prose is accounted for**, each marked kept, moved or + deliberately dropped, per the AGENTS.md Don't — and **the two prompt copies stay in parity** + on every rule this story changes, deliberate wording differences stated as such. + +## 4. Affected AGENTS.md invariants + +- `## What this project is` — "**The product is prompts.** Skills, slash commands, agent + definitions, hook reminder messages and every template `/workflow-init` scaffolds are the + deliverable — plus one POSIX-shell hook." +- `## What this project is` — "There is no application code, so there is no typechecker to catch a + defect; review and `docs/prompt-standards.md` are the only gates a prompt passes through." +- `## What this project is` — "a fingerprinted hardening ledger where a recurring finding escalates + one rung harder (prose → lint → type → test), and one repo-enforced quality command." +- `### Prompts and scaffolding` — "11. **Prompt changes pass `docs/prompt-standards.md`** — all 12 + checklist items, for any skill, command, agent definition, hook message, or scaffolded template. + The prompts are the product, and **no comprehensive mechanical checker exists for them**: review + is the gate." +- `## Don'ts` — "**Never describe what a gate proves without checking what it actually compares.**" +- `## Don'ts` — "**Never replace a decision procedure without accounting for its old conditions.**" + +## 5. Open questions + +- **The profile.** Proposed `high / none / battery+check+verification`. The surface is the review + gate's own severity ladder, and a wrong rule there mis-steers every future cycle in every project + that adopts the kit. **No named `high` trigger matches literally**, so this is a judgement call + under intake's "surfaces, not words", and the human decides it. +- **Is "harness" definable without a per-project list?** In this repository the product is prompts + and the harness is shell test suites and CI checkers, so the line is unusually clean. Elsewhere a + test helper can be both. Whether one definition transfers, or each project names its own the way + `docs/hardening-taxonomy.md` holds project vocabulary, is a requirement question — and invariant + 10 says project vocabulary must not leak into the shipped skill. +- **Where does a tracked harness finding live — `docs/hardening-log.md` or `todos.md`?** The ledger + is append-only, fingerprinted and recurrence-grepped, which fits a tripwire; `todos.md` holds the + backlog that `pending` ledger rows already point to by `ref`. Both are plausible and they behave + differently on recurrence. +- **Does the rule bind Gate A, Gate B, or both?** Gate A reviews text with no test harness in the + usual sense, yet this cycle's four repeated findings were all Gate A, against the spec's own + verification section. Whether "harness" at Gate A means the spec's instrument, and whether that + reading holds downstream, is unsettled. +- **How does per-mechanism termination interact with the existing escalation ladder?** Hardening a + harness finding escalates toward a test — which is the harness. Whether the ladder's rungs mean + something different for a harness finding, or whether termination bypasses the ladder, is open. + +## 6. Suggested size + +`story` — one coherent change to the severity ladder's consequences in two mirrored copies, plus +one new loop move. Not `chore`: it answers a structural question and changes what a standing duty +produces. Not `epic-needs-splitting`: blocking and terminating are one decision procedure, and the +counting half is already owned elsewhere. From 78e5f9761cb89cb0f62758dec5b68b0522d65309 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Thu, 10 Sep 2026 20:32:35 +0200 Subject: [PATCH 018/181] =?UTF-8?q?docs(specs):=20loop-rule=20consolidatio?= =?UTF-8?q?n=20=E2=80=94=20Gate-A=20pass=2012=20revision=20(end=20two=20me?= =?UTF-8?q?chanisms)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Applies Daniel's repeat criterion, brought in mid-loop from a sibling project: the second finding of the same shape against the same mechanism ends that mechanism's rounds — narrow the claim, print the residual, move the work. Two mechanisms here were on their fourth round of one shape. The §7 assert list moves to the plan. A spec cannot build an exact substring check for text that does not exist yet, so each of four revisions guessed and each list carried at least one fragment that could not do what it claimed. §7 now states what a check must establish and which edits owe a discriminating pair; the plan builds the fragments against the real files. The completeness claim did not move with it, because nothing supports it. The block stops restating what it cites, checkably: a cited rule now contributes zero predicate words to it. The two triggers return to b11 and b13, the fix set to b7, the severity split to the (g) replacement. c9's precedence sentence moves the other way, into the block, because precedence is evaluation order and stating it in both places is the drift the one-authority check exists to catch. Also: the decline reversal exception removed (D7 admits none), a16 scoped by pointer, four source pointers deferring to the ordering on what an answer does, the profile-change verification row split in two, and the partial-adoption and observability residuals owned here rather than assigned to a successor whose scope excludes them. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- ...26-09-10-loop-rule-consolidation-design.md | 437 ++++++++++-------- 1 file changed, 242 insertions(+), 195 deletions(-) diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index 9b65828..052d45f 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -55,9 +55,13 @@ order and the file set each predicate reads; the duties' classification; which s scope stop's two triggers raises and what each answer does; what a clearly-stuck or two-tell answer produces; and the raw-severity rule for the health measures. **It defines no trigger, duty or severity rule of its own** — each keeps its one definition where that definition already -lives, and where one had to change to agree with the ordering it changed **at its source**: six -standing sentences outside the inventoried passages (§4), and seven fragments inside them (§5), -each one the ordering falsifies or leaves ambiguous. +lives, and where one had to change to agree with the ordering it changed **at its source**. + +**Nineteen edits in all**, counting **one edit per contiguous replacement or addition at one +site**, which is the unit because a single sentence can be replaced once and a list added to once +without either being two: **five** sentences outside the inventoried passages (§4 items 1–5), and +**fourteen** inside them (§5). One of the fourteen — §4 item 6, passage (i)'s list addition — is +described in §4 for locality and counted here in §5(i), never twice. --- @@ -77,12 +81,11 @@ who finds a rule defined here rather than cited has found a defect. **What a pass is read from.** Every finding-derived predicate reads the validated findings file **or files** of the logical pass as their **concatenation** — a `full` Gate-B pass has two, and -one branch alone is already an incomplete pass. **Which severity field each reads is settled in -Mechanics, Severity**: the closure and repair predicates read the effective severity, the health -measures the reviewer-written one. Beyond the findings, closure reads the **derived floor** and -the final-acceptance preconditions the floor section states, and any **hold still standing**; the -scope triggers read the **current assigned fix set** — what the absorb paragraph defines, plus -what this cycle has accepted, minus what it has declined — and the answers already given; the +one branch alone is already an incomplete pass. **Which severity field each of them reads is +settled in Mechanics, Severity**, which is where that split lives and is not repeated here. +Beyond the findings, closure reads the **derived floor** and the final-acceptance preconditions +the floor section states, and any **hold still standing**; the scope triggers read the **current +assigned fix set** as the absorb paragraph defines it, and the answers already given; the clearly-stuck reading adds its own coverage judgement. **A line in one branch file and a line in the other are distinct findings for holds and answers**, so a `full` pass asks twice rather than risk resuming over one it never asked about. @@ -92,14 +95,10 @@ findings file** is the `NO FINDINGS` signal the protocol defines, and a **clean predicate here, read on the logical pass with every required branch file combined, so one branch's clean file never establishes a clean pass. **A pass is clean** when its findings carry no in-set Blocker or Major at effective severity and **no scope-stop trigger** — the two the -absorb paragraph defines, a finding outside the assigned fix set and a finding opening a new -structural or contract question, read here in its terms and not redefined. Two things this -branch adds, because they are about answers rather than about scope: a finding **this cycle has -declined** raises no membership trigger while the set still excludes it, and a structural or -contract question **this cycle has answered** is no longer new — otherwise an answered question -would be asked again on every pass. Both predicates are properties of the findings and the set, -readable before any branch below runs, which is what makes this order executable rather than -asserted. A pass with **zero** findings is clean whatever the floor, because a floor buys further +absorb paragraph defines, read there and not redefined here, each already carrying the +qualification an answer given in this cycle puts on it. Both are properties of the findings and +the set as that paragraph reads them, settled before any branch below runs, which is what makes +this order executable rather than asserted. A pass with **zero** findings is clean whatever the floor, because a floor buys further looks at an artifact that keeps yielding findings, and one yielding none has already given what those looks were for. **A clean pass at or above the derived floor, or a zero-finding pass, closes the cycle**, under those final-acceptance preconditions. A plateau or tells on the closing @@ -149,11 +148,14 @@ requires has been given, in whichever direction each is given** — one answer f single-trigger finding, both for one carrying both. That is **one rule with two parts**, how many answers and which way each may go, and neither is a test the other has to pass. At a **membership stop** the answer is **accept**, the finding joining the fix set where Mechanics Severity governs -it, or **decline**, the finding staying outside and binding so for the rest of this cycle unless -the user explicitly reverses that decision. Either is an **explicit, attributable decision on -that specific finding** — never silence, never a general remark about scope, never inferred, +it, or **decline**, the finding staying outside and binding so for the rest of this cycle. +A later answer that contradicts a decline is a **contradiction to surface**, not a reversal these +rules permit: **D7** admits no exception, and allowing one would let a finding be moved out of the +set and back into it to escape what it owes inside it. Either answer is an **explicit, +attributable decision on that specific finding** — never silence, never a general remark about +scope, never inferred, because a fix set changed by inference is a fix set nobody chose. **Membership is answered -against the set as it stood at the pass that raised the question**: a later broadening is a new +against the set as the absorb paragraph fixes it for the pass that raised the question**: a later broadening is a new fact the **next** pass reads and never discharges a standing hold, a hold discharged by a scope change being a hold nobody answered. At a **question stop** the answer is the user's decision on the question and membership does not change; an out-of-set finding that opened one is a @@ -178,13 +180,18 @@ exception. **Two pairings cannot occur**, and no rule ranks them: clean completi stop**, since that stop's triggers are the clean predicate's own second half, so a pass raising one is not clean; and a zero-finding pass and any suspension, since it has nothing to surface, nothing regenerating, no cluster and no require↔withdraw pair. **Clean completion and the -clearly-stuck exit can**, and the overlap is admitted rather than argued away: that exit reads -the reviewer-written field while cleanliness reads the effective one, so an **in-set** Blocker or -Major the ceiling demotes below Major, or an **out-of-set** one this cycle has declined, can -regenerate across passes on a pass that is clean. **The order decides it and no new rule is -needed** — at or above the floor the pass closes, ranking the exit exactly as the clearly-stuck -paragraph's own precedence sentence says; below it the pass **suspends**, clean completion having -not closed it. +clearly-stuck exit can**, and the overlap is admitted rather than argued away: the two read +different severity fields, as Mechanics · Severity sets out, so an **in-set** Blocker or Major the +ceiling demotes below Major, or an **out-of-set** one this cycle has declined, can regenerate +across passes on a pass that is clean. **The order decides it and no new rule is needed.** The +clearly-stuck paragraph's own precedence sentence is stated here rather than there, because +precedence is evaluation order and this paragraph is where evaluation order is stated once; its +opening words point back to that paragraph, which is where the reading itself lives. That third +condition is what makes a plateau rather than a finish, and it is why **a clean completion takes +precedence over this exit**: a Blocker/Major-free pass **at or above the floor** has satisfied the +clean-final-pass rule — collect the Minors and Nits and close — and reporting "will not converge" +on a converged loop is a false report. Below the floor the pass **suspends**, clean completion +having not closed it. ``` Where the block maps onto the criteria: the three branches are AC 4, the duties paragraph @@ -200,15 +207,25 @@ rather than being restated here: | membership trigger | `b11`, the absorb paragraph (C:206 / W:413), qualified by §5(b) | | question trigger | `b13`, the absorb paragraph (C:209–213 / W:416–420) | | the assigned fix set | `b7`, the absorb paragraph (C:200–202 / W:407–409), edited by §5(b) | -| effective vs reviewer-written severity | the (g) replacement (§5(g)) | +| effective vs reviewer-written severity | the (g) replacement, in Mechanics · Severity (§5(g)) | | what a Blocker, Major, Minor or Nit demands | Mechanics · Severity (C:783–784 / W:969–970), scoped by §4 item 3 | | the derived floor and final-acceptance preconditions | the floor section (C:116–118, C:760–764) | -| the clearly-stuck reading and its precedence | the clearly-stuck paragraph (C:225–246 / W:428–449), `c9` kept verbatim under **D3** | +| the clearly-stuck reading | the clearly-stuck paragraph (C:225–246 / W:428–449) | | the two-tell threshold | the five-tells paragraph (C:263–268 / W:467–472), qualified by §5(e) | The scope stop's two triggers are `b11` and `b13` read separately, because an in-set finding that opens a question can neither join nor stay outside the set and needs its own answer. +**One rule runs the other way, and it is the one exception to the table.** The clearly-stuck +exit's **precedence** against a clean completion is not part of that exit's reading; it is +evaluation order, which is the block's own subject. So `c9`'s sentence — the one **D3** requires +preserved verbatim — **moves into the block** rather than being cited from where it stood, and +the clearly-stuck paragraph keeps only its reading and points forward (§5(c)). Stated in both +places, it would be exactly the drift this table exists to prevent, and stated only in the +clearly-stuck paragraph it would put evaluation order somewhere the ordering does not govern. +The sentence's words are untouched by the move; its opening clause refers back to that +paragraph's third condition, and the block says so where it quotes it. + **Two things the ordering names and does not define, both the successor's** (§9): what makes a later finding *the same one* this cycle declined, and what form carries an answer into the commit body. The ordering says what an answer does — a decline keeps a finding out of the set @@ -223,19 +240,18 @@ the reviewer re-raising what it cannot read, which is the ordinary route and not ## 4. The standing sentences edited at their source -Six, and no more — working around any of them would ship two instructions that disagree. +Six items, five of them outside the passages §5 accounts for; the sixth is passage (i)'s list +addition, described here because it belongs with these and counted in §5(i). Working around any +of them would ship two instructions that disagree. -**How each OLD is cited, since an earlier revision got this wrong three times.** A sentence -quoted here is quoted whole, and several of them **wrap across two lines** in one or both -copies, so each item gives the sentence's real line range in **both** copies and names the -**single-line substring the check actually counts**. A displayed sentence spanning a wrap -cannot be counted with `grep -F`, and quoting one as though it could is what made three -earlier checks read as false reds. Every counted substring below is verified at 1 in both -copies (§7). +**How each OLD is located, and what this section does not do.** A sentence quoted here is quoted +whole, and several **wrap across two lines** in one or both copies, so each item gives the +sentence's real line range in **both**. This section identifies the sentence and states what +changes; it does **not** name the substrings a check would count. Those are built in the plan, +against the real files, for the reason §7 gives. 1. **The findings-file protocol's clean sentence**, inside the gate-prompt template both gates - paste. Sentence at **C:328–329 / W:522–523**, the word "A" ending the first line; **counted - substring `clean pass is the single body line`, C:329 / W:523**. OLD: "A clean pass is the + paste. Sentence at **C:328–329 / W:522–523**, the word "A" ending the first line. OLD: "A clean pass is the single body line `NO FINDINGS` with `END OF FINDINGS (0 total)`." NEW: "A **clean findings file** is the single body line `NO FINDINGS` with `END OF FINDINGS (0 total)`." It describes a file and always did; the word *pass* in it is what makes §3's predicate look like a @@ -243,8 +259,7 @@ copies (§7). finds nothing — unchanged. 2. **The Gate-A clean-signal sentence**. Sentence at **C:565–566**, wrapped, and **W:757**, - whole on one line; **counted substring `when a pass is clean`, C:565 / W:757** — the part - that is single-line in both, which is why the check uses it and not the displayed sentence. + whole on one line. OLD: "…when a pass is clean — the explicit clean signal is what lets you exit the loop:". NEW: "…when a pass finds nothing — the explicit signal is what lets a pass be read as clean without inspecting it further:". @@ -256,8 +271,7 @@ copies (§7). and are correct as they stand — counted and checked, not assumed. 3. **The Severity bullet's resolve duty**, the one place the duty is stated and the only one - without a scope. Sentence at **C:783–784 / W:969–970**, wrapped in both; **counted substring - `both must resolve. Minor`, C:784 / W:970**. OLD: "- **Severity:** Blocker + without a scope. Sentence at **C:783–784 / W:969–970**, wrapped in both. OLD: "- **Severity:** Blocker (wrong/unsafe/breaks invariant) · Major (design flaw → rework) → both must resolve. Minor · Nit → collect, never iterate." NEW: "- **Severity:** Blocker (wrong/unsafe/breaks invariant) · Major (design flaw → rework) → both must resolve, **for every finding in the @@ -273,8 +287,7 @@ copies (§7). scope. 4. **The lens paragraph's unchanged-list**, which asserts of two rules this change alters that - they are unchanged. Sentence at **C:653–654 / W:839–840**; **counted substring - `clean-final-pass rule are unchanged.`, single-line at C:654 / W:840**. OLD: + they are unchanged. Sentence at **C:653–654 / W:839–840**. OLD: "clean-final-pass rule are unchanged." NEW: "clean-final-pass rule are unchanged **by the lens sets**, which is what this paragraph is about — the closure ordering above does change both, scoping the filter to the assigned fix set and defining a clean pass at effective severity, and says so @@ -284,8 +297,8 @@ copies (§7). the three are changed by the ordering and an unscoped claim leaves both copies denying an edit they carry); the floor is not among them (**kept**). -5. **The Gate-A cadence**, which makes a revision unconditional between passes. **Counted - substring `Each pass: validate, revise, re-run.`, single-line at C:573 / W:764.** +5. **The Gate-A cadence**, which makes a revision unconditional between passes. Sentence at + **C:573 / W:764**, single-line in both. OLD: "Each pass: validate, revise, re-run." NEW: "Each pass: validate, revise **where the severity and scope rules require a repair**, re-run." Old condition: every pass is followed by a revision before the next. Kept: the cadence and its order — validate @@ -296,9 +309,8 @@ copies (§7). 6. **The unknown-start fallback** (C:153–167 / W:360–374, `i4`–`i8`, extended at `i12`'s invitation) — an addition to a list rather than a change of meaning, so it is checked by - presence and not as a pair. After the anchor **`duty owed, and the nonce duties at their - strictest`, single-line at C:157 / W:364**, the list gains "**every suspension binding**". - `i12`'s sentence stays as + presence and not as a pair. The strict-reading list, at **C:157 / W:364**, gains "**every + suspension binding**". `i12`'s sentence stays as written; this is the addition it invites, and it is what this change owes that list: a cycle that cannot establish its starting rules treats a suspension as binding rather than as advisory, which is the strict reading of the rules §3 ships. @@ -316,7 +328,7 @@ lives in the §3 block; "replaced" = the condition changes, and says how; "dropp reason. **Every inventoried passage has an entry**, including the four this narrowing no longer edits, so that the accounting stays a complete map of the ten rather than a shorter list. -**(a) The floor paragraphs** (C:72–136 / W:279–343) — **two edits**. First, **trim** +**(a) The floor paragraphs** (C:72–136 / W:279–343) — **three edits**. First, **trim** `a17`–`a19` and point at the block. OLD: "Your final pass must be clean — if the pass at the floor still finds Blocker/Major, keep going until clean or clearly stuck → then STOP and surface to the user. The only early exit below the floor is a pass with **zero** findings; @@ -326,44 +338,60 @@ to pad." Second, `a13`. OLD: "Every other rule stated here about how a cycle closes stands as written, and none of them is restated — a summary is where their conditions would get dropped." NEW: "**This paragraph** restates none of them — a summary is where their conditions would get dropped — and the closure ordering below is where they are -stated once and in order." Accounting: `a1`–`a12`, `a14`–`a17`, `a20`–`a22` kept; `a18`, `a19` -moved; **`a13` replaced**. Old condition: *no* rule about how a cycle closes is restated +stated once and in order." Third, `a16`, single-line at **C:132 / W:339**. OLD: "Open a +TodoWrite "Codex pass N" per pass; fix Blocker/Major after each." NEW: "Open a TodoWrite "Codex +pass N" per pass; fix Blocker/Major after each **as Mechanics · Severity requires**." +Accounting: `a1`–`a12`, `a14`, `a15`, `a17`, `a20`–`a22` kept; `a18`, `a19` +moved; **`a13` replaced**; **`a16` replaced**. Old condition: *no* rule about how a cycle closes is restated anywhere. Kept: the prohibition and its reason, scoped to the paragraph it was written to police. Changed: the ordering does restate closure rules, deliberately and as the one authority, so a categorical reading would leave the shipped text contradicting itself — the -failure `a13` exists to prevent, arriving from the other direction. - -**(b) What a loop absorbs** (C:195–223 / W:402–426) — **five sentence edits**. This passage owns +failure `a13` exists to prevent, arriving from the other direction. `a16`'s old condition: +**every** Blocker and Major is fixed after each pass, unbounded, standing beside a Severity +bullet §4 item 3 now scopes to the assigned fix set. Kept: the per-pass cadence and the duty +itself. Changed: it points at the rule that carries the scope instead of restating an unscoped +version of it — a declined finding is out of the set and owes nothing, and an unscoped `a16` +commands its repair anyway, which is the same both-must-resolve contradiction §4 item 3 exists +to close, one paragraph earlier. + +**(b) What a loop absorbs** (C:195–223 / W:402–426) — **six sentence edits**. This passage owns the two triggers and the fix set, so where the ordering needs them qualified, they are qualified -**here**, which is what keeps one definition per rule. - -`b7`, the fix-set definition, **counted substring `scope the approved story or plan assigns to -this cycle`, single-line at C:201 / W:408**. OLD: "…it is the scope the approved story or plan -assigns to this cycle, plus repair obligations you already accepted in earlier passes." NEW: -"…it is the scope **every governing story or plan** assigns to this cycle — their union where -several do — plus repair obligations you already accepted in earlier passes, **minus any finding -this cycle has declined**." - -`b11`, the membership trigger, **counted substring `loop like any other out-of-scope finding**, -even when it opens no new question at all`, single-line at C:206 / W:413**. OLD: "**A correction +**here**, which is what keeps one definition per rule. **None of these edits states what an +answer then does** — that is the ordering's, and each pointer below defers to it rather than +repeating a version of it that can drift. + +`b7`, the fix-set definition (C:200–202 / W:407–409). OLD: "…it is the scope the approved story +or plan assigns to this cycle, plus repair obligations you already accepted in earlier passes." +NEW: "…it is the scope **every governing story or plan** assigns to this cycle — their union +where several do, since a cycle a single artifact does not govern has no set at all under the +singular reading — plus repair obligations you already accepted in earlier passes, **minus any +finding this cycle has declined**, a decline being the user's answer that it stays out." + +`b11`, the membership trigger (C:206 / W:413). OLD: "**A correction that leaves that set stops the loop like any other out-of-scope finding**, even when it opens no new question at all…" NEW: "**A correction that leaves that set stops the loop like any other out-of-scope finding**, even when it opens no new question at all — **unless this cycle has already declined that same finding, which put it outside by the user's own answer**…" +`b13`, the question trigger (C:209–213 / W:416–420). OLD: "…that opens a **new structural or +contract question** stops the loop and goes to the user…" NEW: "…that opens a **new structural +or contract question** — new meaning **not already answered in this cycle**, since an answered +question re-raised is the same question and asking it again on every pass is a loop that cannot +end — stops the loop and goes to the user…" + `b12` OLD: "…and it resumes the moment the user says whether the set now includes it." -NEW: "…and it resumes once the user has said whether the set now includes it — accepting or -declining it — **together with every other answer that pass's suspensions require**, under the -closure ordering above." `b17`–`b18` OLD: +NEW: "…and the user's answer, accepting or declining it, is what ends the hold; **what the pass +does then, once every answer its suspensions require has been given, is the closure ordering +above**." `b17`–`b18` OLD: "Stopping this way is **not an exit from the gate**: the floor, the Blocker/Major filter and the clean-final-pass rule all stand, and the loop resumes on the revised artifact once the question is answered — what the stop prevents…" NEW: "Stopping this way is a **suspension** -under the closure ordering above, and the loop resumes on the revised artifact **once every -answer that pass's suspensions require has been given** — what the stop prevents…". Fifth, in +under the closure ordering above, and **what ends it, and what the loop does next, are stated +there** — what the stop prevents…". Sixth, in **W only**, `b3`: "…by its severity exactly as the severity rule already says…" becomes "…exactly as Mechanics already says…", matching C for the reason §6's first row gives. -Accounting: `b1`, `b2`, `b4`–`b6`, `b8`–`b10`, `b13`–`b16` kept. Five are worth naming. +Accounting: `b1`, `b2`, `b4`–`b6`, `b8`–`b10`, `b14`–`b16` kept. Six are worth naming. `b3` is an **edited cross-reference whose operative condition is kept** — W's pointer changes target while what it points at is unchanged, so it is not in the kept range above and it is checked in §7 like any other edit. `b6` fixes @@ -381,37 +409,44 @@ have raised the question. `b11` **replaced** — old: any out-of-set finding sto unconditionally. Kept: the trigger and its reason. Changed: a finding this cycle has already declined does not raise it again, because unqualified `b11` and the ordering's clean predicate decide a re-raised declined finding in opposite directions, and re-asking an answered question -on every pass is the non-idempotent path AC 4 forbids. `b13` is the question trigger and -**stays in its own words**, checked equivalent to the block's reference rather than merged into -it; §6 is where that check runs. `b12` **replaced** — old: an *immediate* +on every pass is the non-idempotent path AC 4 forbids. `b13` **replaced** — old: any finding +opening a new structural or contract question stops the loop, with *new* undefined. Kept: the +trigger, the novelty test and its reason, and that size is not the test. Changed: *new* now +excludes a question this cycle has already answered, because unqualified it and the ordering's +clean predicate decide a re-raised answered question in opposite directions — the same +non-idempotent path as `b11`'s, arriving through the other trigger. The qualification is added +**here rather than in the block**, which is what keeps both triggers to one definition each; +§6 checks that the block's reference and this sentence say the same thing. `b12` **replaced** — old: an *immediate* resume on the membership answer alone. Kept: that the membership answer is what the stop asks -for and that either direction ends it. Changed: the resume waits for every answer the pass's -suspensions require, because a loop resumed over an unanswered question decides it by running. +for and that either direction ends the hold. Changed: it says the answer ends the hold and +**defers to the ordering** for what the pass does next, because a loop resumed over an unanswered +question decides it by running, and because a second statement of what an answer does is a second +authority — an earlier revision of this spec had four of them, each saying the loop resumes once +every answer is given and none saying a stop answer parks the cycle. `b17` moved (the block's "a suspension waives nothing" sentence); `b18` **replaced** the same way and for the same reason — old: resume once *the* question is -answered; new: once *every* required answer is given. W's remaining wording differences here -are untouched (§6). - -**(c) Recognizing clearly stuck** (C:225–246 / W:428–449) — **keep** the precedence sentence -in full and **trim** what follows it. The sentence kept byte-for-byte (**D3**): "That third -condition is what makes a plateau rather than a finish, and it is why **a clean completion -takes precedence over this exit**: a Blocker/Major-free pass **at or above the floor** has -satisfied the clean-final-pass rule — collect the Minors and Nits and close — and reporting -"will not converge" on a converged loop is a false report." OLD, from "**Below the floor -nothing closes**" to the end of "…and then nothing could satisfy both.": replaced by -"That example is a clean-completion candidate only where the pass carries **no scope-stop -trigger** — a finding outside the assigned fix set, or one opening a new structural or -contract question — since a pass carrying either is not clean and the ordering above says -why. **Below the floor nothing closes**, exactly as that ordering says. This exit is -a **suspension** under it: you surface with the findings still open, the resolve -rule is not waived by surfacing, and the loop resumes when every answer that pass's -suspensions require has been given." The kept sentence is not touched; the qualification is -adjacent to it, which is how **D3**'s verbatim requirement and the ordering can both hold — -edited, the sentence would say a Blocker/Major-free pass carrying a new out-of-set Minor -closes, and the ordering says it suspends. Accounting: `c1`–`c8` -kept; `c9`, `c10`, `c11` **kept verbatim and qualified** — the sentence is unchanged and a new -sentence beside it names the condition its example assumed and never stated, so the conditions -`c10` carries are narrowed by adjacency rather than by edit; `c12` kept (pointer form); `c13` moved (the +answered; new: the ordering says what ends the suspension and what follows. W's remaining wording +differences here are untouched (§6). + +**(c) Recognizing clearly stuck** (C:225–246 / W:428–449) — **one replacement**, from "That +third condition is what makes a plateau rather than a finish" to the end of "…and then nothing +could satisfy both." The paragraph keeps its reading — the three conditions and why each is +needed — and stops carrying evaluation order. NEW: "**What that third condition means for +closure — how this exit ranks against a clean completion, and what happens below the floor — +is the closure ordering above**, which carries this paragraph's precedence sentence word for +word so evaluation order is stated in one place. This exit is a **suspension** under that +ordering: you surface with the findings still open, the resolve rule is not waived by +surfacing, and **what its answer does is stated there**." + +**The precedence sentence is moved, not rewritten, and that is what satisfies D3.** Its words +are untouched; only the paragraph it sits in changes, because precedence is evaluation order +and the block is where evaluation order is stated once (§3). Left here it would be a second +authority for the one thing the block exists to own — and an earlier revision of this spec kept +it here *and* stated the same ranking in the block, which is the drift the one-authority check +was added to catch. Accounting: `c1`–`c8` +kept; `c9`, `c10`, `c11` **moved** — the sentence in full, into the block's composition +paragraph, where the block names it as this paragraph's and says why its opening clause points +back here; `c12` kept (pointer form); `c13` moved (the zero-finding exception); `c14` moved **to the ordering's continue branch** — the Minor-below-floor pass that "keeps looping" is precisely one that neither closes nor suspends, and the branch states it rather than leaving it to be read out of a negation; `c15`, `c17` @@ -431,10 +466,11 @@ them below Major, the pass is clean at effective severity — it closes at or ab reading D3's own preserved sentence and `c18` decide that pass in opposite directions. The two-tell stop was never in this condition's domain, surfacing tells and not a finding; `c19` **replaced** — old: the loop resumes on whatever the user -decides, unconditionally; new: it resumes on a resuming answer, and a stop **parks** the cycle — -open, not running, restarted only by an explicit later continue — because an unconditional -resume is the stop-with-no-transition path AC 4 forbids and a post-stop state identical to the -pre-answer one is that same path wearing a different name; +decides, unconditionally. Kept: that an answer is what moves the cycle. Changed: this passage no +longer says what the answer does and defers to the ordering, which states it once — a resuming +answer resumes, a stop **parks**. An unconditional resume is the stop-with-no-transition path +AC 4 forbids, and a passage carrying its own version of the transition is how four sites came to +disagree with the block about a stop answer; `c20` **dropped** — it argued that reading the exit as "stop instead of fixing" would compete with the resolve duty, and the block states that argument's conclusion as a rule instead. @@ -447,7 +483,7 @@ is `e7`, the shared fragment both copies carry (C:266 / W:470). OLD: "**Any two stop-and-surface mandatory, not discretionary**". NEW: "**Any two present makes stop-and-surface mandatory, not discretionary, where the clean-completion branch did not close the pass**". Then the pointer, after `e10`: "This stop is a **suspension** under the closure ordering above, and -the loop resumes when every answer that pass's suspensions require has been given." +**what its answer does is stated there**." Accounting: `e1`–`e6`, `e8`–`e11` kept; `e7` **qualified** — old: two tells make the stop mandatory, unconditionally. Kept: the threshold, its mandatory force, and that the stuck reading @@ -547,60 +583,48 @@ what each level of that mode obliges, so that whichever it carries has its evide the battery, the check, and the named verification of the risk path. **Battery.** The quality command in `AGENTS.md` § Commands runs once, green, at the Gate-B WIP -commit (`check-version-bump.sh main` needs the committed bump, §8). - -**The check.** Inline shell asserts in the plan, run against the working tree and the parent -tree, so each assertion is observed passing where the change exists and failing where it does -not: - -- lead phrase of the §3 block, `**How a cycle ends — one ordering`, count **1** in C and **1** - in W; the same greps against `git show 7c0d475:CLAUDE.md` and - `git show 7c0d475:plugins/dev-workflow/commands/workflow-init.md` count **0**; -- the (g) sentence "That question is owned by the loop-rule consolidation work" count **0** in - both; against the parent tree, **1** in C and **0** in W; -- the (g) replacement's lead phrase `**The demotion changes what a cycle must resolve, never - what it counts.**` count **1** in C and **1** in W; **0** in the parent tree. The removal - check above cannot do this alone: deleting the old sentence and installing nothing satisfies - it and the parity diff, reporting the demotion/loop-health outcome as verified while both - copies carry no answer at all; -- the **five paired source edits** (§4 items 1–5), each as an OLD/NEW pair, because each is a - standing sentence changing meaning rather than new text appearing, and a one-sided presence - check would pass on a copy carrying both wordings. Each NEW counts **1** in each copy and - **0** in the parent tree, each OLD the reverse: `A **clean findings file** is the single body - line` against `clean pass is the single body line`; `when a pass finds nothing` against `when - a pass is clean`; `for every finding in the assigned fix set` against `both must resolve. - Minor`; `clean-final-pass rule are unchanged **by the lens sets**` against `clean-final-pass - rule are unchanged.`; and `revise **where the severity and scope rules require a repair**` - against `Each pass: validate, revise, re-run`. **Every fragment is a single line in the file - it is grepped from**, and the NEW ones must be installed unwrapped, because a fragment - spanning a line break makes `grep -F` count 0 and read as a failure — an earlier revision - quoted three of these across their wraps. That is a constraint on how the plan writes the - edits, not on what they mean; -- §4 item 6, the unknown-start addition, by presence alone since it adds to a list rather than - changing a meaning: `every suspension binding` count **1** in each and **0** in the parent; -- **the passage edits in §5, each as an OLD/NEW pair on the same terms.** They change standing - meanings exactly as §4's do, and checking only §4 would leave both copies able to carry the - new ordering beside stale resume, trigger, fix-set, surfacing and two-tell rules while every - stated count and the parity diff still passed. Seven pairs, each OLD counting **1** in the - parent tree and **0** in the working tree and each NEW the reverse: `final pass must be - clean — if the pass at the floor still finds Blocker/Major, keep going until` (`a17`–`a19`, - C:133 / W:340); `written, and none of them is restated` (`a13`, C:130 / W:337); `scope the - approved story or plan assigns to this cycle` (`b7`, C:201 / W:408); `loop like any other - out-of-scope finding**, even when it opens no new question at all` (`b11`, C:206 / W:413); - `the moment the user says whether the set now includes it.` (`b12`, C:208 / W:415); - `**Below the floor nothing closes**, and a zero-finding pass remains` (the (c) trim, C:238 / - W:441); and `**Any two present makes stop-and-surface mandatory, not discretionary**` (`e7`, - C:266 / W:470). **`b3` is W-only** and is checked in W alone — `severity exactly as the - severity rule already says` counting **1** in the parent W and **0** in the working W, against - `severity exactly as Mechanics already says` counting **0** then **1**; C already carries the - target wording at C:198 and must be unchanged, which is what makes this an alignment rather - than an edit to both. +commit (`scripts/check-version-bump.sh main` needs the committed bump, §8). + +**The check — what it must establish, and where it is built.** Every edit that changes a +standing meaning owes a **discriminating pair of counts**: one showing the new wording present, +one showing the old wording gone. Each half is run in **both copies** and against **both** the +working tree and the parent tree (`git show 7c0d475:CLAUDE.md` and +`git show 7c0d475:plugins/dev-workflow/commands/workflow-init.md`), so every assertion is +observed passing where the change exists and failing where it does not. A one-sided presence +check is not enough: a copy carrying the new wording **and** the old one satisfies it, which is +exactly the two-instructions-that-disagree failure §4 exists to prevent. An edit that only adds +to a list is checked by **presence alone**, because nothing is being replaced. + +**Which edits owe a pair**, named here because this spec knows which sentences it changes: +§4 items 1–5; and in §5, `a17`–`a19` and `a13` and `a16` in (a); `b7`, `b11`, `b12`, `b13`, +`b17`–`b18` in (b), plus `b3` in **W alone**, since C already carries that target wording and +must be unchanged, which is what makes it an alignment rather than an edit to both; the +replacement in (c); `e7` in (e); and the paragraph replacement in (g), whose old sentence is the +one site where the parent is present and the change removes it. **Presence alone:** the §3 +block's lead phrase, the (e) pointer sentence, and §4 item 6. + +**The plan builds each pair against the real files and runs both directions there.** It carries +two constraints: a counted fragment must be **single-line in the file it is grepped from**, since +one spanning a line break makes `grep -F` count 0 and read as a failure; and the new wording must +therefore be **installed unwrapped**. Both are constraints on how an edit is written, not on what +it means. + +**Why the fragments are not enumerated here, stated as a residual rather than repaired again.** +Four consecutive revisions of this spec listed them, and each list contained at least one +fragment that could not do what it claimed — a substring preserved inside its own replacement, so +the old-wording-gone count could never reach zero; three quoted across their line wraps, so they +counted zero in a correct tree; and a meaning-changing passage with no pair at all. The cause is +structural: an exact substring check for text that does not yet exist can only be guessed, and a +guess that is wrong reads as a failed check rather than as a wrong check. **Nothing here verifies +that the list of edits above is complete, or that the plan's chosen fragments discriminate.** The +enumeration moved to where the text exists; the completeness claim did not move with it, because +nothing supports it. If the claim "the ordering ships in both copies" were false, one working-tree count would be **0** or the parent-tree counts would not differ from it. The wiring can produce that observation: each grep reads the file bytes at the named revision and nothing supplies its own input. **The counterfactual is ABSENT, and is claimed as absent** — the parent carries no -ordering block, and the (g) count is the one site where the parent is present and the change +ordering block, and the (g) sentence is the one site where the parent is present and the change removes it. Nothing is claimed as "contradictory". **The named verification of the risk path** (story AC 4) is a **next-state table**, in the plan @@ -617,12 +641,19 @@ qualification — a zero-finding pass, and the unknown-start fallback; **and the transitions**: **a finding this cycle declined re-raised on a later pass** — whose next state is no membership trigger and a pass still eligible for clean, the row that tests `b11`'s qualification — **a structural question this cycle answered, re-raised** — whose next state is -no question stop — an **accepted** finding re-raised before repair, **the fix set broadened +no question stop, the row that tests `b13`'s — **a decline that leaves nothing to revise** — +whose next state is a further pass on the **unrevised** artifact, the row that tests the continue +branch against a valid path with no repair in it — an **accepted** finding re-raised before +repair, **the fix set broadened while a hold is still awaiting its answer** — whose next state is the hold still standing, the -row that tests the frozen reading — **a confirmed profile change while a hold stands** — whose -next state is the hold still standing *and* the further pass under the current profile that a -profile change costs, run against both answer directions, since final acceptance re-reads every -profile and a hold surviving one is not the same claim as a hold surviving a scope change — a +row that tests the frozen reading — and **a confirmed profile change while a hold is still +awaiting its answer**, which is **two rows and not one**: the profile change itself, whose next +state is the hold still standing and **no pass run**, since a standing hold forbids one; and then +the answer, **in each direction**, whose next state is the parked cycle or a resume, with only a +resuming answer starting the further pass under the current profile that a profile change costs. +Split because final acceptance re-reads every profile, and a hold surviving a profile change is +not the same claim as an answer paying for one — a single row asserting both would have to run a +pass through an unanswered hold to be filled in. Then a fix set narrowed so an in-set finding falls outside it, and a `full` Gate-B pass with one branch clean and the other carrying an in-set Blocker. **Columns** — the **governing-scope state** (the fix set as currently assigned), the @@ -643,8 +674,9 @@ inputs must include every input the rule reads, and a counterfactual must distin from CONTRADICTORY. **No fixture per predicate is built** — parked in the story's §2, not reopened. -**Evidence entry**, in the closing commit body, names: the battery run; every assert pair above -with its working-tree and parent-tree counts; the §6 parity diff — the passages extracted, the +**Evidence entry**, in the closing commit body, names: the battery run; every pair the plan +built, with the fragment it counted and its working-tree and parent-tree counts in each copy, +and every presence check beside them; the §6 parity diff — the passages extracted, the differences observed, that each is one of the permitted rows, and the `b11`/`b13` equivalence result; and the next-state table's location in the plan plus its row count. It is revalidated before every Gate-B re-review and before the closing amend, as §5 requires. @@ -656,8 +688,9 @@ two tells ever suspended it, what was asked, or how it was answered. A reader of therefore cannot audit that every suspension was answered before the cycle closed. Neither the evidence entry above nor the per-pass curve supplies this, and saying otherwise would be the overclaim `AGENTS.md` names as this repo's most persistent defect. The transport that could -carry it left with the record (§9), so **this belongs to the successor**; until then it is an -admitted gap, not a guarded one. +carry it left with the record (§9). **No story has taken it**, and naming one that has not is +the same defect in a smaller place, so it is recorded here as an **admitted gap of this change** +— unowned, unguarded, and available to whoever picks it up. --- @@ -683,15 +716,22 @@ admitted gap, not a guarded one. including the four this narrowing no longer edits. - **Don't: "Never rename or delete a doc section without grepping for references first."** The (g) sentence names the story path; the grep finds it at `CLAUDE.md:815` (the site itself) - and in three artifacts of the parent cycle (`…plan-a-rules.md:878`, - `…review-loop-economics-design.md:33`, `…pass-floor-story.md:78`). All cite the story file, - which continues to exist; none cites the sentence. Nothing breaks. -- **Invariant 12 — a plugin change requires a version bump.** `workflow-init.md` is under - `plugins/`, so `plugins/dev-workflow/.claude-plugin/plugin.json` goes `0.11.0 → 0.12.0` with - a `CHANGELOG.md` entry: a minor bump, the template gaining a closure ordering and six edited - sentences. -- **Invariant 4 / the hook.** Untouched: `codex-gate.sh` is not edited, and the §5 heading it - greps (`Cross-Model Review`) does not move. + and in three artifacts of the parent cycle + (`docs/superpowers/plans/2026-08-29-review-loop-economics-plan-a-rules.md:878`, + `docs/superpowers/specs/2026-08-28-review-loop-economics-design.md:33`, + `docs/superpowers/stories/2026-08-28-review-loop-economics-pass-floor-story.md:78`). All cite + the story file, which continues to exist; none cites the sentence. Nothing breaks. +- **Invariant 12 — a plugin change requires a version bump.** + `plugins/dev-workflow/commands/workflow-init.md` is under `plugins/`, so + `plugins/dev-workflow/.claude-plugin/plugin.json` goes `0.11.0 → 0.12.0` with a + `plugins/dev-workflow/CHANGELOG.md` entry: a minor bump, the template gaining a closure + ordering and nineteen edited or extended sentences. +- **Invariant 4 / the hook.** Untouched: `plugins/dev-workflow/hooks/codex-gate.sh` is not + edited, and the §5 heading it greps (`Cross-Model Review`) does not move. + +**Every path in this spec is written repository-relative and in full** — no ellipsis shorthand +and no bare basename. An abbreviated citation fails the path-existence check a review pass runs +mechanically, and five of them did. --- @@ -705,28 +745,35 @@ decision of 2026-09-10, with the evidence in §1: finding (**D9b**), and the unverified-assertion reading (**D9c**); - the pass-4 report's unavailable-history block (**D10**) and the checkout-root condition that report grew; -- the rollback reading — **whether what an open cycle already has is enough** when a revert - removes the text it started under. It is not a gap here, and an earlier revision of this spec - wrongly implied one: the activation paragraph this change ships and leaves unchanged already - gives that cycle a defined transition — a cycle whose starting rules cannot be established - takes the **stricter reading**, which §4 item 6 extends with "every suspension binding", so it - neither closes nor is left without a next state. What is open is whether the stricter reading - is *sufficient* for a rollback specifically, since no record identifies the rule revision a - cycle started under. That question is the successor's; the transition exists meanwhile; +- the rollback reading — **what an open cycle owes when a revert removes the rules it started + under**. This change does not answer it and no longer claims to. An earlier revision said the + activation paragraph already supplies the transition; that is wrong in a specific way. A revert + of this change removes the ordering, the source edits **and §4 item 6's "every suspension + binding" extension together**, so the stricter reading such a cycle would fall back to is + itself part of what the revert takes away, and it no longer mentions the suspensions that cycle + is holding. What remains is the pre-change §5 — the loose ordering this story exists to + replace — read by a cycle that started under a different one. **Whether that is enough is + unanswered here and nothing is shipped for it**, since no record identifies the rule revision a + cycle started under. An admitted residual, and the successor's question; - the slot-discriminator dissolution deferred here by Plan C's Tasks 19 and 20; -- **the partial-adoption guard for the whole set.** It was built as a "closure-record contract" - naming the record, with a marker on every mergeable hunk, and it leaves with the record it was - named for. The narrowed set is still mutually dependent — the ordering's clean predicate needs - §4 item 3's boundary, its two senses of *clean* need items 1 and 2, and its triggers need - §5(b)'s `b7` and `b11` — and **nothing catches a partial adoption of it**. Said exactly: - the live one-contract paragraph (C:880–890 / W:1063–1074) names the nonce, the slots, the - provenance line, the curve, the carry rule and the unknown-start semantics, and **it does not - name this block or any of its coupled edits**, so its coherence rule does not reach them. An - earlier revision of this spec said that rule caught them; it does not, and a downstream merge - can take the clean predicate without the scoped resolve duty, or either sense of *clean* - without the other, and run. **That is an admitted unsafe state, not a guarded one.** The guard - belongs with the successor because the two sets are adopted together in practice and one guard - should name both. +- **the partial-adoption guard as it applied to the record.** It was built as a "closure-record + contract" naming the record, with a marker on every mergeable hunk, and it leaves with the + record it was named for. + +**One residual this change owns rather than moves: partial adoption of the narrowed set.** The +set is mutually dependent — the ordering's clean predicate needs §4 item 3's boundary, its two +senses of *clean* need items 1 and 2, and its triggers need §5(b)'s `b7`, `b11` and `b13` — and +**nothing catches a downstream merge that takes some of it**. Said exactly: the live one-contract +paragraph (C:880–890 / W:1063–1074) names the nonce, the slots, the provenance line, the curve, +the carry rule and the unknown-start semantics, and **it does not name this block or any of its +coupled edits**, so its coherence rule does not reach them. An earlier revision of this spec said +that rule caught them; it does not, and a project can take the clean predicate without the scoped +resolve duty, or either sense of *clean* without the other, and run. **That is an admitted unsafe +state, not a guarded one, and it is this change's own.** An earlier revision assigned the guard to +the successor; that was wrong on the face of the successor's own scope, which is record +durability and **excludes the closure ordering by name**, so the assignment named an owner that +had not taken it — the same defect as claiming a mechanism that does not exist. Nothing here +builds one, and a downstream project gets no coverage for this set. **Out of scope and parked**, unchanged: From 8e11f3783a40e80667046d8fe463fd63afdca094 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Thu, 10 Sep 2026 21:00:48 +0200 Subject: [PATCH 019/181] docs(stories): add the two missing loop bounds to the termination story MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Daniel asked how a project that develops this kit avoids blocking itself while using it. The diagnosis from cycle awsf1ec771 is that the cost was not self-application: the loop had no ceiling and the artifact had no size limit, and self-application only multiplied the readings each round had to weigh. Both bounds are mechanical rather than a reading, which is why they belong beside per-mechanism termination rather than in a story of their own. A pass ceiling turns an unbounded grind into a mandatory stop-and-surface, the shape the two-tell rule already has. A size limit checked before pass 1 is a line count. Evidence: thirteen Gate-A spec passes against a floor of 3 with no clean pass, on a spec that reached 989 lines of which 173 were the design. Three acceptance criteria added. A third bound needed no new rule and is recorded as a lapse instead: §5 already requires settling mechanically what a parser decides, and the precheck was written and then not run. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- ...09-10-harness-finding-termination-story.md | 35 +++++++++++++++++++ 1 file changed, 35 insertions(+) diff --git a/docs/superpowers/stories/2026-09-10-harness-finding-termination-story.md b/docs/superpowers/stories/2026-09-10-harness-finding-termination-story.md index 8af1242..738605f 100644 --- a/docs/superpowers/stories/2026-09-10-harness-finding-termination-story.md +++ b/docs/superpowers/stories/2026-09-10-harness-finding-termination-story.md @@ -76,6 +76,30 @@ move between "absorb another round of this" and "hand the whole cycle to the hum that keeps producing the same shape of finding stops consuming rounds while the loop continues on everything else. +**Two bounds exist that today do not, and both are mechanical rather than a reading.** Added +2026-09-10 from the same session that produced the evidence above, on Daniel's question of how a +project developing this kit avoids blocking itself while using it. The diagnosis that prompted them +is that the cycle was not slow because of self-application; it was slow because **the loop had no +ceiling and the artifact had no size limit**, and self-application only multiplied the readings each +round had to consider. + +- **A pass ceiling.** §5 fixes a floor and no maximum. The only two ways up and out — the + clearly-stuck exit and the two-tell threshold — both require a *reading*, so a cycle that trips + neither grinds without anyone being obliged to decide. A ceiling turns that into a mandatory + stop-and-surface, the same shape the two-tell rule already has, where the human chooses to split, + to accept with stated residuals, or to continue for a named reason. Cycle `awsf1ec771` ran + **thirteen** Gate-A spec passes against a floor of 3 and reached no clean pass. +- **An artifact size limit before the cycle starts.** §5's sizing guidance — "prefer smaller specs + with named interfaces and let the plan carry the detail" — is advice with no number, and it was + read and not followed. The same spec reached **989 lines**, of which the design was **173**; the + remaining **58%** was bookkeeping about the change, and it took roughly half the findings of every + pass. A limit checked before the first pass is a `wc -l`, not a judgement. + +**A third bound already exists in §5 and was simply not honoured**, which is worth recording because +it needed no new rule: the instruction to settle mechanically what a parser can decide before +spending a read pass on it. A precheck script was written before pass 1 of that cycle and then never +run again, and roughly a quarter of the later passes' findings were things it decides in seconds. + **Out of scope**, named so nothing absorbs them: - **The severity definitions.** Blocker, Major, Minor and Nit keep their current meanings; this story changes what follows from a severity, not what earns one. @@ -109,6 +133,17 @@ everything else. - [ ] **Per-mechanism termination and the existing five tells do not duplicate or contradict each other.** Both copies say which applies when both would fire, and neither weakens the two-tell mandatory stop. +- [ ] **A cycle cannot run unbounded without a human deciding.** A ceiling exists, it is stated as a + number derived the way the floor is, and reaching it without a clean pass is a mandatory + stop-and-surface naming the options. Checkable by reading: a pass count alone decides it, with + no reading of a curve or a cluster. +- [ ] **An oversized artifact is stopped before the cycle starts, not diagnosed after it.** A stated + limit applies before pass 1, it is checkable by counting lines, and the shipped text says what + an over-limit artifact does instead — split, or move detail to the plan behind a named + interface. +- [ ] **The two bounds and the existing exits compose without a fourth reading.** For every state + where a bound and an exit could both apply, the shipped text says which governs and why, or + names the pair as unable to co-occur. - [ ] **Every condition of the replaced prose is accounted for**, each marked kept, moved or deliberately dropped, per the AGENTS.md Don't — and **the two prompt copies stay in parity** on every rule this story changes, deliberate wording differences stated as such. From 0168f880307497af325a157c491e7f47ea9ec375 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Thu, 10 Sep 2026 21:05:33 +0200 Subject: [PATCH 020/181] docs(specs): cut the spec to the design; the plan carries the bookkeeping MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Measured before deciding: the spec was 785 lines, of which the design it exists to state — the closure ordering — was 173, or 22%. The bookkeeping about the change was 456 lines, or 58%: quoted OLD and NEW text, the per-condition dispositions, the parity divergence list, the verification substrings. That 58% had been taking roughly half the findings of every Gate-A pass while describing work the plan performs against real files. Pass 12 had already run the experiment at small scale. Moving the verification substrings to the plan killed three Majors at once and regenerated nothing, because a spec cannot build an exact check for text that does not yet exist. The same holds for every OLD/NEW pair and every kept/moved/dropped disposition. Daniel's decision, 2026-09-10: cut now. All four are moved to the plan, not dropped, and each section names the plan as carrier. Story acceptance criterion 5 is satisfied by the plan's disposition list against the committed 135-condition inventory; §5 keeps the ten-passage map so no passage falls out of view. 785 -> 532 lines. Gate-A pass 13 applied in the same commit. Fixed: the validated dismissal as a resolution route, with a re-raised dismissal counting as regeneration so a repeat false positive reaches a suspension instead of continuing forever; a state for an answer contradicting a binding decline, preserving D7; clean completion as eligibility, the closing amend as the closure itself; b8 tested against the set b7 computes; b3 as a pure pointer; the evidence entry's revalidation rule cited among the closure preconditions; continue permitting an unrevised artifact only where no repair is owed; the hold's answer count scoped to scope-stop answers; the next-state table's claim narrowed to answer-state transitions with separate named checks beside it. Dissolved with the cut: two count claims and two OLD quotations. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- ...26-09-10-loop-rule-consolidation-design.md | 829 ++++++------------ 1 file changed, 288 insertions(+), 541 deletions(-) diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index 052d45f..1b16d7f 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -7,16 +7,17 @@ copy, and a value copied here would be a remembered value. Prompt-only, in two mirrored copies: `CLAUDE.md` §5 (**C** below) and the inline template in `plugins/dev-workflow/commands/workflow-init.md` (**W** below). No file under -`plugins/dev-workflow/hooks/` changes. Line numbers cite the tree at `7c0d475` (main) and are -re-read at execution; the plan carries the `grep -n` sites. The old-conditions accounting in §5 -cites ids `a1`…`j4`, defined in +`plugins/dev-workflow/hooks/` changes. Condition ids `a1`…`j4` are defined in `docs/superpowers/specs/2026-09-10-loop-rule-consolidation-condition-inventory.md` beside this -file: 135 conditions (22 + 18 + 20 + 7 + 11 + 7 + 4 + 26 + 16 + 4) quoted from `7c0d475`, so a -reader can check the accounting rather than take it. A sentence outside those passages is -cited by its lead phrase and C line. +file: 135 conditions quoted from `7c0d475`, so a reader can check an accounting rather than +take it. -**Narrowed 2026-09-10: §9 lists what moved out and where.** Read it before reading anything -here as missing. +**Narrowed twice on 2026-09-10, and §9 lists what moved.** First the record-durability subject +went to a successor story. Then the bookkeeping went to the plan: this spec was 785 lines of +which the design was 173, and the remaining 58% — quoted OLD and NEW text, per-condition +dispositions, the parity divergence list, the verification substrings — took roughly half the +findings of every Gate-A pass while describing work the plan performs against real files. **It +is moved, not dropped.** The plan is the carrier and each section below names what it owes. --- @@ -26,18 +27,20 @@ here as missing. cycle and which merely *suspend* it, in what order a pass is read so the ranking is executable rather than asserted, what any set of suspensions at once does, and which of the four standing duties participate in that ordering versus gate it as preconditions. With it: the answer to what -a severity demotion does to the loop-health counts, and the six standing sentences the ordering -falsifies or leaves ambiguous if they are not edited at their source. +a severity demotion does to the loop-health counts, and the standing sentences the ordering +falsifies or leaves ambiguous if they are not edited at their source — **five outside the +inventoried passages and fifteen inside them, eighteen replacing a sentence and two adding +one**. **What does not:** the pass floor and severity semantics, which the parent shipped and this spec reads as given, and everything §9 lists as moved or parked. -**Why the split, since a reader of the ordering will look for the record.** Gate-A spec pass 10 -tripped the two-tell threshold and satisfied all three conditions of the clearly-stuck reading. -Nine of that pass's twenty findings belonged to **one subject this story never set out to -answer** — whether a record survives a session, a commit amend, a squash, a rollback or a moved -checkout — while the ordering's own five were small. Daniel split that subject out on -2026-09-10; §9 says what went and where. +**Why the record split, since a reader of the ordering will look for the record.** Gate-A spec +pass 10 tripped the two-tell threshold and satisfied all three conditions of the clearly-stuck +reading. Nine of that pass's twenty findings belonged to **one subject this story never set out +to answer** — whether a record survives a session, a commit amend, a squash, a rollback or a +moved checkout — while the ordering's own five were small. Daniel split that subject out; §9 +says what went and where. --- @@ -57,11 +60,11 @@ answer produces; and the raw-severity rule for the health measures. **It defines duty or severity rule of its own** — each keeps its one definition where that definition already lives, and where one had to change to agree with the ordering it changed **at its source**. -**Nineteen edits in all**, counting **one edit per contiguous replacement or addition at one +**Twenty source edits**, counting **one edit per contiguous replacement or addition at one site**, which is the unit because a single sentence can be replaced once and a list added to once -without either being two: **five** sentences outside the inventoried passages (§4 items 1–5), and -**fourteen** inside them (§5). One of the fourteen — §4 item 6, passage (i)'s list addition — is -described in §4 for locality and counted here in §5(i), never twice. +without either being two. §4 lists all twenty. The **closure-ordering block is an addition +beside them**, not one of the twenty, so a reader counting changes to the two copies counts +twenty-one. --- @@ -84,11 +87,15 @@ who finds a rule defined here rather than cited has found a defect. one branch alone is already an incomplete pass. **Which severity field each of them reads is settled in Mechanics, Severity**, which is where that split lives and is not repeated here. Beyond the findings, closure reads the **derived floor** and the final-acceptance preconditions -the floor section states, and any **hold still standing**; the scope triggers read the **current -assigned fix set** as the absorb paragraph defines it, and the answers already given; the -clearly-stuck reading adds its own coverage judgement. **A line in one branch file and a line in -the other are distinct findings for holds and answers**, so a `full` pass asks twice rather than -risk resuming over one it never asked about. +the floor section states, the **evidence entry's revalidation rule** — a changed entry meaning +the clean pass no longer covers what is being committed — and any **hold still standing**; the +scope triggers read the **current assigned fix set** as the absorb paragraph defines it, and the +answers already given; the clearly-stuck reading adds its own coverage judgement. **A line in +one branch file and a line in the other are distinct findings for holds and answers**, so a +`full` pass asks twice rather than risk resuming over one it never asked about. **Within one +running cycle an answer binds to the finding or question as the pass that raised it recorded +them** — which is what an agent running the cycle can do with nothing written down. Recognising +the same finding or question across a lost session needs a record these rules do not ship. **First, clean completion.** §5 uses *clean* in two senses and now says which is which: a **clean findings file** is the `NO FINDINGS` signal the protocol defines, and a **clean pass** is the @@ -100,14 +107,16 @@ qualification an answer given in this cycle puts on it. Both are properties of t the set as that paragraph reads them, settled before any branch below runs, which is what makes this order executable rather than asserted. A pass with **zero** findings is clean whatever the floor, because a floor buys further looks at an artifact that keeps yielding findings, and one yielding none has already given what -those looks were for. **A clean pass at or above the derived floor, or a zero-finding pass, -closes the cycle**, under those final-acceptance preconditions. A plateau or tells on the closing -pass go into the closing report and never block it, because reporting "will not converge" on a -converged loop is a false report. **No other pass outcome closes a cycle**, because every other -pass leaves a required repair, a hold or a question outstanding, or has an unmet closure -precondition — and closing over any of those is the failure this ordering exists to prevent. The -one termination that is not a pass outcome is the Gate-B triviality skip, which runs no passes -and is outside this ordering. +those looks were for. **A clean pass at or above the derived floor, or a zero-finding pass, makes +the cycle eligible to close**, under those final-acceptance preconditions; **the closing amend +Mechanics · Finishing the cycle describes is the closure itself**, so nothing is closed before +it, and a profile, cited set or evidence entry that changes in between still gates it. A plateau +or tells on that pass go into the closing report and never block it, because reporting "will not +converge" on a converged loop is a false report. **No other pass outcome makes a cycle eligible +to close**, because every other pass leaves a required repair, a hold or a question outstanding, +or has an unmet closure precondition — and closing over any of those is the failure this ordering +exists to prevent. The one termination that is not a pass outcome is the Gate-B triviality skip, +which runs no passes and is outside this ordering. **Second, only a pass that is not a clean completion can suspend** — that order is what makes "clean completion outranks the two-tell stop" executable rather than asserted. Three suspensions, @@ -120,11 +129,13 @@ surfaces that also carries either trigger takes the scope stop's answers at that it is not asked twice; the two-tell stop surfaces tells and not a finding. **Third, a pass that neither closes nor suspends continues** — the loop runs another pass on -the **current** artifact, revised or not. A below-floor clean pass lands here **only where no -suspension applies to it**; where one does, the second branch has already taken it, because -clean completion did not close the pass and only closing outranks a suspension. So does a pass -whose only findings are Minors and Nits, which may leave nothing to revise. It is a branch and -not an inference, because "does not close" read alone says nothing about whether to run again. +the **current** artifact, revised where the severity and scope rules require a repair and +unrevised where they do not. A below-floor clean pass lands here **only where no suspension +applies to it**; where one does, the second branch has already taken it, because clean completion +did not close the pass and only closing outranks a suspension. So does a pass whose only findings +are Minors and Nits, which are collected and never iterated and may leave nothing to revise. It +is a branch and not an inference, because "does not close" read alone says nothing about whether +to run again. **The four standing duties, classified.** The **derived floor** is a **precondition on closure**: it gates closing, discharged by the count of valid logical passes reaching it with the last of @@ -132,43 +143,57 @@ them clean, or by the zero-finding exit. The **Blocker/Major-resolve duty** is a on closure and on any pass being clean**: Mechanics Severity, scoped to the assigned fix set, states what it demands and what discharges it, and a validated pass finding **no in-set Blocker or Major at effective severity** is what shows it discharged — the clean predicate's own wording, -so the two cannot drift. **A dismissal is not a decline**: a dismissal is the author's judgement -that the finding is not true of the artifact, a decline the user's decision that a true finding -stays outside the set, and only the second is an answer at a membership stop. The **hold** a -surfaced finding places on closure **participates in the ordering**: it gates closing while it -stands, and is discharged by the answers that surface requires. **It attaches to every surfaced -finding, whichever suspension surfaced it** — clean completion creates none, because it wins -before anything is surfaced. **No-clean-credit** — no pass carrying a scope-stop trigger is -credited as clean — also participates, and is a fact about that pass that nothing discharges, a -later pass being judged on its own findings. It is not a second test beside the clean predicate -but that predicate's second half, which is why it is stated in its words. +so the two cannot drift. It is discharged by a **repair** or by a **validated dismissal**: the +author's judgement, carrying the one-line why this section already requires of a dismissed +finding, that the finding is not true of the artifact. A dismissal does not rewrite the pass that +found it, and the later clean pass is still owed. **A dismissal is not a decline**: a dismissal +says the finding is false, a decline is the user's decision that a **true** finding stays outside +the set, and only the second is an answer at a membership stop. **A dismissal the reviewer keeps +re-raising across passes is regeneration**, counting toward the clearly-stuck reading's third +condition — so a false positive that returns every pass reaches that suspension instead of +continuing forever. The **hold** a surfaced finding places on closure **participates in the +ordering**: it gates closing while it stands, and is discharged by the answers that surface +requires. **It attaches to every surfaced finding, whichever suspension surfaced it** — clean +completion creates none, because it wins before anything is surfaced. **No-clean-credit** — no +pass carrying a scope-stop trigger is credited as clean — also participates, and is a fact about +that pass that nothing discharges, a later pass being judged on its own findings. It is not a +second test beside the clean predicate but that predicate's second half, which is why it is +stated in its words. **What a suspension asks, and what ends it.** **A hold ends when every answer its finding -requires has been given, in whichever direction each is given** — one answer for a +requires has been given, in whichever direction each is given** — one **scope-stop** answer for a single-trigger finding, both for one carrying both. That is **one rule with two parts**, how many -answers and which way each may go, and neither is a test the other has to pass. At a **membership -stop** the answer is **accept**, the finding joining the fix set where Mechanics Severity governs -it, or **decline**, the finding staying outside and binding so for the rest of this cycle. -A later answer that contradicts a decline is a **contradiction to surface**, not a reversal these -rules permit: **D7** admits no exception, and allowing one would let a finding be moved out of the -set and back into it to escape what it owes inside it. Either answer is an **explicit, +answers and which way each may go, and neither is a test the other has to pass. Where a health +suspension applies to the same pass, its shared continue-or-stop answer is **additional** to +those and not counted among them, the health readings asking about the loop rather than about +this finding. At a **membership stop** the answer is **accept**, the finding joining the fix set +where Mechanics Severity governs it, or **decline**, the finding staying outside and binding so +for the rest of this cycle. A later answer that contradicts a decline **does not reverse it**: +the decline **remains binding**, the contradiction is **surfaced to the user as information**, +and the **cycle continues** — nothing here turns one answer into another, since that would let a +finding be moved out of the set and back into it to escape what it owes inside it, and **D7** +admits no exception. What the user may do is **withdraw** the decline, which is not a reversal by +these rules but a fresh explicit decision by the same authority that declined; the finding then +takes the ordinary route at the next pass that raises it. Either answer is an **explicit, attributable decision on that specific finding** — never silence, never a general remark about -scope, never inferred, -because a fix set changed by inference is a fix set nobody chose. **Membership is answered -against the set as the absorb paragraph fixes it for the pass that raised the question**: a later broadening is a new -fact the **next** pass reads and never discharges a standing hold, a hold discharged by a scope -change being a hold nobody answered. At a **question stop** the answer is the user's decision on -the question and membership does not change; an out-of-set finding that opened one is a -membership stop as well. **Decline is available only at a membership stop**, that being the only -stop whose question is whether a finding belongs to the set. The **clearly-stuck and two-tell -readings** ask **continue or stop**. **Continue consumes the reading that raised the -suspension**: a further health suspension needs that reading recomputed over a pass run after the -answer, which is new data — so continue produces a distinct next state even on an unrevised -artifact, and the same reading cannot return the same stop unanswered. **Stop parks the cycle**: -open, not running, spending no passes, restarted only by an explicit later continue — a distinct -state from the suspended-awaiting-answer one it was in before the answer. Nothing a parked cycle -wrote is a closing commit, and a parked cycle nobody restarts is a human's to resolve, exactly as -the nonce rules already say of open cycles. +scope, never inferred, because a fix set changed by inference is a fix set nobody chose. +**Membership is answered against the set as the absorb paragraph fixes it for the pass that +raised the question**: a later broadening is a new fact the **next** pass reads and never +discharges a standing hold, a hold discharged by a scope change being a hold nobody answered. At +a **question stop** the answer is the user's decision on the question and membership does not +change; an out-of-set finding that opened one is a membership stop as well. **Decline is +available only at a membership stop**, that being the only stop whose question is whether a +finding belongs to the set. The **clearly-stuck and two-tell readings** ask **continue or stop**. +**Continue consumes the reading that raised the suspension**: a further health suspension needs +that reading recomputed over a pass run after the answer, which is new data — so continue +produces a distinct next state, and the same reading cannot return the same stop unanswered. It +permits an **unrevised** artifact **only where no repair is owed**; where effective severity or +scope requires one, that repair comes before the post-answer pass, since a pass run over an +unrepaired in-set Blocker or Major spends a look on text the rules already say must change. +**Stop parks the cycle**: open, not running, spending no passes, restarted only by an +explicit later continue — a distinct state from the suspended-awaiting-answer one it was in +before the answer. Nothing a parked cycle wrote is a closing commit, and a parked cycle nobody +restarts is a human's to resolve, exactly as the nonce rules already say of open cycles. **Composition, and what cannot happen.** Every **question** is answered on its own and the loop resumes only when every answer resumes it — accept or decline at a membership stop, a decision at @@ -197,21 +222,21 @@ having not closed it. Where the block maps onto the criteria: the three branches are AC 4, the duties paragraph AC 2, the composition sentences AC 1. -**One authority per rule — the check finding 14 of pass 11 asked for.** The block cites eight -rules and defines none of them. Each has exactly one definition in the shipped text, and where -that definition had to change to agree with the ordering, it changed **at its source** (§4, §5) -rather than being restated here: +**One authority per rule.** The block cites nine rules and defines none of them. Each has exactly +one definition in the shipped text, and where that definition had to change to agree with the +ordering, it changed **at its source** (§4) rather than being restated here: | Rule the block cites | Its one definition | |---|---| -| membership trigger | `b11`, the absorb paragraph (C:206 / W:413), qualified by §5(b) | -| question trigger | `b13`, the absorb paragraph (C:209–213 / W:416–420) | -| the assigned fix set | `b7`, the absorb paragraph (C:200–202 / W:407–409), edited by §5(b) | -| effective vs reviewer-written severity | the (g) replacement, in Mechanics · Severity (§5(g)) | -| what a Blocker, Major, Minor or Nit demands | Mechanics · Severity (C:783–784 / W:969–970), scoped by §4 item 3 | -| the derived floor and final-acceptance preconditions | the floor section (C:116–118, C:760–764) | -| the clearly-stuck reading | the clearly-stuck paragraph (C:225–246 / W:428–449) | -| the two-tell threshold | the five-tells paragraph (C:263–268 / W:467–472), qualified by §5(e) | +| membership trigger | `b11`, the absorb paragraph, qualified by §4 | +| question trigger | `b13`, the absorb paragraph, qualified by §4 | +| the assigned fix set, and what is in it | `b7` and `b8`, the absorb paragraph, both edited by §4 | +| effective vs reviewer-written severity | the (g) replacement, in Mechanics · Severity | +| what a Blocker, Major, Minor or Nit demands | Mechanics · Severity, scoped by §4 | +| the derived floor and final-acceptance preconditions | the floor section | +| the evidence entry's revalidation rule | the profiles section, unedited | +| the clearly-stuck reading | the clearly-stuck paragraph | +| the two-tell threshold | the five-tells paragraph, qualified by §4 | The scope stop's two triggers are `b11` and `b13` read separately, because an in-set finding that opens a question can neither join nor stay outside the set and needs its own answer. @@ -220,285 +245,52 @@ that opens a question can neither join nor stay outside the set and needs its ow exit's **precedence** against a clean completion is not part of that exit's reading; it is evaluation order, which is the block's own subject. So `c9`'s sentence — the one **D3** requires preserved verbatim — **moves into the block** rather than being cited from where it stood, and -the clearly-stuck paragraph keeps only its reading and points forward (§5(c)). Stated in both -places, it would be exactly the drift this table exists to prevent, and stated only in the -clearly-stuck paragraph it would put evaluation order somewhere the ordering does not govern. -The sentence's words are untouched by the move; its opening clause refers back to that -paragraph's third condition, and the block says so where it quotes it. +the clearly-stuck paragraph keeps only its reading and points forward. Stated in both places, it +would be exactly the drift this table exists to prevent; stated only in the clearly-stuck +paragraph, it would put evaluation order somewhere the ordering does not govern. The sentence's +words are untouched by the move, and the block says whose sentence it is where it quotes it. **Two things the ordering names and does not define, both the successor's** (§9): what makes a -later finding *the same one* this cycle declined, and what form carries an answer into the -commit body. The ordering says what an answer does — a decline keeps a finding out of the set -for the cycle, an acceptance puts one in until the cycle closes — and **D9b** and **D9** say how -it is recognised and written down. **Two consequences are stated in the block rather than left -to be discovered**: until a sameness rule ships, a line in one branch file and a line in the -other are distinct findings for holds and answers, so a `full` Gate-B pass may ask twice rather -than risk resuming over an unanswered one; and an answer binds within the session that made it, -the reviewer re-raising what it cannot read, which is the ordinary route and not a special one. +later finding *the same one* this cycle declined, and what form carries an answer into the commit +body. The ordering says what an answer does; **D9b** and **D9** say how it is recognised across +sessions and written down. **Within one cycle the block supplies what it needs and says so**: a +line in each branch file is a distinct finding for holds and answers, and an answer binds to the +finding or question as the pass that raised it recorded them. Both sentences are in the block +above — the ordering claims no in-session consequence it does not state there. --- ## 4. The standing sentences edited at their source -Six items, five of them outside the passages §5 accounts for; the sixth is passage (i)'s list -addition, described here because it belongs with these and counted in §5(i). Working around any -of them would ship two instructions that disagree. - -**How each OLD is located, and what this section does not do.** A sentence quoted here is quoted -whole, and several **wrap across two lines** in one or both copies, so each item gives the -sentence's real line range in **both**. This section identifies the sentence and states what -changes; it does **not** name the substrings a check would count. Those are built in the plan, -against the real files, for the reason §7 gives. - -1. **The findings-file protocol's clean sentence**, inside the gate-prompt template both gates - paste. Sentence at **C:328–329 / W:522–523**, the word "A" ending the first line. OLD: "A clean pass is the - single body line `NO FINDINGS` with `END OF FINDINGS (0 total)`." NEW: "A **clean findings - file** is the single body line `NO FINDINGS` with `END OF FINDINGS (0 total)`." It describes - a file and always did; the word *pass* in it is what makes §3's predicate look like a - redefinition instead of the other sense. Old condition — what a reviewer writes when it - finds nothing — unchanged. - -2. **The Gate-A clean-signal sentence**. Sentence at **C:565–566**, wrapped, and **W:757**, - whole on one line. - OLD: "…when a pass is clean — the explicit clean signal is what lets you exit the loop:". - NEW: "…when a pass finds nothing — the explicit signal is what lets a pass be read as clean - without inspecting it further:". - Old condition: a `NO FINDINGS` file is what permits loop exit. Kept: the signal and why it - is demanded. **Replaced**: it makes a pass readable as clean rather than being the only way - to be clean, since a pass carrying Minors alone is clean under the ordering and could never - produce this file. The other **six** uses of "clean pass" in each copy (C:117, 728, 750, - 761, 769, 827; W:324, 914, 936, 947, 955, 1011) are the closure sense the ordering defines - and are correct as they stand — counted and checked, not assumed. - -3. **The Severity bullet's resolve duty**, the one place the duty is stated and the only one - without a scope. Sentence at **C:783–784 / W:969–970**, wrapped in both. OLD: "- **Severity:** Blocker - (wrong/unsafe/breaks invariant) · Major (design flaw → rework) → both must resolve. Minor · - Nit → collect, never iterate." NEW: "- **Severity:** Blocker (wrong/unsafe/breaks - invariant) · Major (design flaw → rework) → both must resolve, **for every finding in the - assigned fix set** (the closure ordering above). Minor · Nit → collect, never iterate." - Old condition: every Blocker and Major resolves, unbounded. **Replaced**: the boundary - **D5** always implied and no sentence carried. Without it a declined finding must stay - outside the set and still bars closure, which is a pass that can neither close nor suspend. - **The edit names the set and never a decline**, which is AC 3's second half: the answer - moves a finding into or out of the set, and this duty says only what it demands of what is - in it — a mention of the decline here would read as a direct waiver of Blocker/Major - resolution rather than the membership decision it is. The (c) pointer that calls this "the - resolve rule" (`c17`) needs no edit: it names the rule, and the rule now carries its own - scope. - -4. **The lens paragraph's unchanged-list**, which asserts of two rules this change alters that - they are unchanged. Sentence at **C:653–654 / W:839–840**. OLD: - "clean-final-pass rule are unchanged." NEW: "clean-final-pass rule are unchanged **by the lens sets**, which is - what this paragraph is about — the closure ordering above does change both, scoping the - filter to the assigned fix set and defining a clean pass at effective severity, and says so - there." Old conditions: lenses change what a pass asks and never how many passes a cycle - owes (**kept**); the Blocker/Major filter, the file-first findings protocol and the - clean-final-pass rule are unchanged (**replaced** — scoped to the lens sets, because two of - the three are changed by the ordering and an unscoped claim leaves both copies denying an - edit they carry); the floor is not among them (**kept**). - -5. **The Gate-A cadence**, which makes a revision unconditional between passes. Sentence at - **C:573 / W:764**, single-line in both. - OLD: "Each pass: validate, revise, re-run." NEW: "Each pass: validate, revise - **where the severity and scope rules require a repair**, re-run." Old condition: every - pass is followed by a revision before the next. Kept: the cadence and its order — validate - first, re-run last. **Replaced**: the revision is conditional, because the ordering's - continue branch reaches a below-floor pass whose only findings are Minors and Nits, which - are collected and never iterated, and an unconditional "revise" tells that pass to - manufacture the repair the severity rule forbids. - -6. **The unknown-start fallback** (C:153–167 / W:360–374, `i4`–`i8`, extended at `i12`'s - invitation) — an addition to a list rather than a change of meaning, so it is checked by - presence and not as a pair. The strict-reading list, at **C:157 / W:364**, gains "**every - suspension binding**". `i12`'s sentence stays as - written; this is the addition it invites, and it is what this change owes that list: a cycle - that cannot establish its starting rules treats a suspension as binding rather than as - advisory, which is the strict reading of the rules §3 ships. - -The shorter "Copy every record into the squash body" sentence inside the human-exception block -(C:1004–1007) is not edited, and neither is the "records every cycle owes" list (C:698): this -change ships no record. - ---- - -## 5. Edits to the existing passages, with the old-conditions accounting - -Ids are the committed inventory's (§ header). "Kept" = the sentence stays; "moved" = it now -lives in the §3 block; "replaced" = the condition changes, and says how; "dropped" carries its -reason. **Every inventoried passage has an entry**, including the four this narrowing no longer -edits, so that the accounting stays a complete map of the ten rather than a shorter list. - -**(a) The floor paragraphs** (C:72–136 / W:279–343) — **three edits**. First, **trim** -`a17`–`a19` and point at the block. OLD: "Your final pass must be clean — if the pass at the -floor still finds Blocker/Major, keep going until clean or clearly stuck → then STOP and -surface to the user. The only early exit below the floor is a pass with **zero** findings; -don't manufacture findings to pad." NEW: "Your final pass must be clean; how a cycle closes, -and what stops it short of closing, is the closure ordering below — don't manufacture findings -to pad." Second, `a13`. OLD: "Every other rule stated here about how a -cycle closes stands as written, and none of them is restated — a summary is where their -conditions would get dropped." NEW: "**This paragraph** restates none of them — a summary is -where their conditions would get dropped — and the closure ordering below is where they are -stated once and in order." Third, `a16`, single-line at **C:132 / W:339**. OLD: "Open a -TodoWrite "Codex pass N" per pass; fix Blocker/Major after each." NEW: "Open a TodoWrite "Codex -pass N" per pass; fix Blocker/Major after each **as Mechanics · Severity requires**." -Accounting: `a1`–`a12`, `a14`, `a15`, `a17`, `a20`–`a22` kept; `a18`, `a19` -moved; **`a13` replaced**; **`a16` replaced**. Old condition: *no* rule about how a cycle closes is restated -anywhere. Kept: the prohibition and its reason, scoped to the paragraph it was written to -police. Changed: the ordering does restate closure rules, deliberately and as the one -authority, so a categorical reading would leave the shipped text contradicting itself — the -failure `a13` exists to prevent, arriving from the other direction. `a16`'s old condition: -**every** Blocker and Major is fixed after each pass, unbounded, standing beside a Severity -bullet §4 item 3 now scopes to the assigned fix set. Kept: the per-pass cadence and the duty -itself. Changed: it points at the rule that carries the scope instead of restating an unscoped -version of it — a declined finding is out of the set and owes nothing, and an unscoped `a16` -commands its repair anyway, which is the same both-must-resolve contradiction §4 item 3 exists -to close, one paragraph earlier. - -**(b) What a loop absorbs** (C:195–223 / W:402–426) — **six sentence edits**. This passage owns -the two triggers and the fix set, so where the ordering needs them qualified, they are qualified -**here**, which is what keeps one definition per rule. **None of these edits states what an -answer then does** — that is the ordering's, and each pointer below defers to it rather than -repeating a version of it that can drift. - -`b7`, the fix-set definition (C:200–202 / W:407–409). OLD: "…it is the scope the approved story -or plan assigns to this cycle, plus repair obligations you already accepted in earlier passes." -NEW: "…it is the scope **every governing story or plan** assigns to this cycle — their union -where several do, since a cycle a single artifact does not govern has no set at all under the -singular reading — plus repair obligations you already accepted in earlier passes, **minus any -finding this cycle has declined**, a decline being the user's answer that it stays out." - -`b11`, the membership trigger (C:206 / W:413). OLD: "**A correction -that leaves that set stops the loop like any other out-of-scope finding**, even when it opens no -new question at all…" NEW: "**A correction that leaves that set stops the loop like any other -out-of-scope finding**, even when it opens no new question at all — **unless this cycle has -already declined that same finding, which put it outside by the user's own answer**…" - -`b13`, the question trigger (C:209–213 / W:416–420). OLD: "…that opens a **new structural or -contract question** stops the loop and goes to the user…" NEW: "…that opens a **new structural -or contract question** — new meaning **not already answered in this cycle**, since an answered -question re-raised is the same question and asking it again on every pass is a loop that cannot -end — stops the loop and goes to the user…" - -`b12` OLD: "…and it resumes the moment the user says whether the set now includes it." -NEW: "…and the user's answer, accepting or declining it, is what ends the hold; **what the pass -does then, once every answer its suspensions require has been given, is the closure ordering -above**." `b17`–`b18` OLD: -"Stopping this way is **not an exit from the gate**: the floor, the Blocker/Major filter and -the clean-final-pass rule all stand, and the loop resumes on the revised artifact once the -question is answered — what the stop prevents…" NEW: "Stopping this way is a **suspension** -under the closure ordering above, and **what ends it, and what the loop does next, are stated -there** — what the stop prevents…". Sixth, in -**W only**, `b3`: "…by its severity exactly as the severity rule already says…" becomes -"…exactly as Mechanics already says…", matching C for the reason §6's first row gives. - -Accounting: `b1`, `b2`, `b4`–`b6`, `b8`–`b10`, `b14`–`b16` kept. Six are worth naming. -`b3` is an **edited cross-reference whose operative condition is kept** — W's pointer changes -target while what it points at is unchanged, so it is not in the kept range above and it is -checked in §7 like any other edit. `b6` fixes -the assigned set **before the pass being answered**, and the ordering's membership answer is -read against that same set rather than against a later one, so the two agree and `b6` needs no -edit. An earlier revision of this spec had membership re-read when the answer arrived, which -`b6` and **D4** both forbid — **D4** requiring a user answer for the hold to end at all — and -the reversal is recorded here rather than left as a silent narrowing. `b7` **replaced** — old: -the scope *the* approved story or plan assigns, singular, with no subtraction. Kept: that the -set is fixed before the pass being answered and carries the obligations already accepted. -Changed: it is the union where several artifacts govern one cycle, and it subtracts this cycle's -declines — the singular reading left a Gate-B cycle fed by several plans, or a Gate-A artifact -citing several stories, with no defined set on its **first** pass, before any suspension could -have raised the question. `b11` **replaced** — old: any out-of-set finding stops the loop, -unconditionally. Kept: the trigger and its reason. Changed: a finding this cycle has already -declined does not raise it again, because unqualified `b11` and the ordering's clean predicate -decide a re-raised declined finding in opposite directions, and re-asking an answered question -on every pass is the non-idempotent path AC 4 forbids. `b13` **replaced** — old: any finding -opening a new structural or contract question stops the loop, with *new* undefined. Kept: the -trigger, the novelty test and its reason, and that size is not the test. Changed: *new* now -excludes a question this cycle has already answered, because unqualified it and the ordering's -clean predicate decide a re-raised answered question in opposite directions — the same -non-idempotent path as `b11`'s, arriving through the other trigger. The qualification is added -**here rather than in the block**, which is what keeps both triggers to one definition each; -§6 checks that the block's reference and this sentence say the same thing. `b12` **replaced** — old: an *immediate* -resume on the membership answer alone. Kept: that the membership answer is what the stop asks -for and that either direction ends the hold. Changed: it says the answer ends the hold and -**defers to the ordering** for what the pass does next, because a loop resumed over an unanswered -question decides it by running, and because a second statement of what an answer does is a second -authority — an earlier revision of this spec had four of them, each saying the loop resumes once -every answer is given and none saying a stop answer parks the cycle. -`b17` moved (the block's "a suspension waives nothing" sentence); `b18` -**replaced** the same way and for the same reason — old: resume once *the* question is -answered; new: the ordering says what ends the suspension and what follows. W's remaining wording -differences here are untouched (§6). - -**(c) Recognizing clearly stuck** (C:225–246 / W:428–449) — **one replacement**, from "That -third condition is what makes a plateau rather than a finish" to the end of "…and then nothing -could satisfy both." The paragraph keeps its reading — the three conditions and why each is -needed — and stops carrying evaluation order. NEW: "**What that third condition means for -closure — how this exit ranks against a clean completion, and what happens below the floor — -is the closure ordering above**, which carries this paragraph's precedence sentence word for -word so evaluation order is stated in one place. This exit is a **suspension** under that -ordering: you surface with the findings still open, the resolve rule is not waived by -surfacing, and **what its answer does is stated there**." - -**The precedence sentence is moved, not rewritten, and that is what satisfies D3.** Its words -are untouched; only the paragraph it sits in changes, because precedence is evaluation order -and the block is where evaluation order is stated once (§3). Left here it would be a second -authority for the one thing the block exists to own — and an earlier revision of this spec kept -it here *and* stated the same ranking in the block, which is the drift the one-authority check -was added to catch. Accounting: `c1`–`c8` -kept; `c9`, `c10`, `c11` **moved** — the sentence in full, into the block's composition -paragraph, where the block names it as this paragraph's and says why its opening clause points -back here; `c12` kept (pointer form); `c13` moved (the -zero-finding exception); `c14` moved **to the ordering's continue branch** — the -Minor-below-floor pass that "keeps looping" is precisely one that neither closes nor suspends, -and the branch states it rather than leaving it to be read out of a negation; `c15`, `c17` -kept in the pointer sentence, `c17`'s "resolve rule" now reading with the assigned-fix-set -boundary §4 item 3 gives it, so the pointer needs no edit of its own; `c16` **kept** — a -clearly-stuck surface leaves its finding open, and the ordering's hold attaches to **every** -surfaced finding, whichever suspension surfaced it, discharged by the answers that surface -requires. An earlier revision of this spec narrowed the hold to scope stops, which contradicted -both `c16` and the story's own third standing duty; the reversal is recorded here rather than -left as a silent narrowing, and it is why the duties paragraph now says "every surfaced -finding" instead of naming one stop; `c18` **replaced** — old: no pass is credited as clean -on a clearly-stuck surface, unconditionally. Kept: the rule wherever it decides anything — a -pass carrying a scope-stop trigger is unclean, and this exit is reached only on a pass that did -not close. Changed: where the exit's regenerating findings are in-set and the ceiling demotes -them below Major, the pass is clean at effective severity — it closes at or above the floor and -**suspends below it**, clean completion having not closed it. Authority **D3** — under the old -reading D3's own preserved sentence and `c18` decide that pass in opposite directions. The -two-tell stop was never in this condition's domain, surfacing -tells and not a finding; `c19` **replaced** — old: the loop resumes on whatever the user -decides, unconditionally. Kept: that an answer is what moves the cycle. Changed: this passage no -longer says what the answer does and defers to the ordering, which states it once — a resuming -answer resumes, a stop **parks**. An unconditional resume is the stop-with-no-transition path -AC 4 forbids, and a passage carrying its own version of the transition is how four sites came to -disagree with the block about a stop answer; -`c20` **dropped** — it argued that reading the exit as "stop instead of fixing" would compete -with the resolve duty, and the block states that argument's conclusion as a rule instead. - -**(d) From pass 4 onward** (C:255–261 / W:459–465) — **no longer edited.** The Q6 -unavailable-history block this passage was to receive moved to the successor with **D10**, and -nothing the narrowed change ships touches the three-line duty. `d1`–`d7` kept, unedited. - -**(e) The five tells** (C:263–268 / W:467–472) — **one edit and one added sentence**. The edit -is `e7`, the shared fragment both copies carry (C:266 / W:470). OLD: "**Any two present makes -stop-and-surface mandatory, not discretionary**". NEW: "**Any two present makes stop-and-surface -mandatory, not discretionary, where the clean-completion branch did not close the pass**". Then -the pointer, after `e10`: "This stop is a **suspension** under the closure ordering above, and -**what its answer does is stated there**." - -Accounting: `e1`–`e6`, `e8`–`e11` kept; `e7` **qualified** — old: two tells make the stop -mandatory, unconditionally. Kept: the threshold, its mandatory force, and that the stuck reading -is not a precondition for it. Changed: it is read after clean completion, not beside it. -Authority **D2**, which ranks clean completion above this stop — unqualified, `e7` and the -ordering decide a clean two-tell pass at or above the floor in opposite directions, which is -AC 1's reachable conflict and the one thing this passage must not leave standing. - -**(f) The two rules above do not compete** (C:275–287 / W:473–483) — **unchanged**. "The two -rules above" still names the absorb rule and the stuck reading; the block sits before both and -adds no third rule between them. Accounting: `f1`–`f7` kept, unchanged; `f1` remains true -because the block composes the suspensions and ranks none over another. - -**(g) Mechanics · Severity · the handed-over question** (C:810–815 / W:996–999) — **replace** -the whole paragraph, both copies, with the answer. NEW: +Twenty edits. Each row names the sentence, where it lives, what changes and why, and whether it +is a replacement or an addition. **No OLD or NEW text and no line numbers appear here**: the plan +quotes each sentence from the real file, writes the replacement, and re-greps the site, for the +reason §7 gives. Working around any of these would ship two instructions that disagree. + +| # | Sentence | Where | Change, and why | Kind | +|---|---|---|---|---| +| 1 | the findings-file protocol's clean sentence | the gate-prompt template both gates paste, C and W | it calls a `NO FINDINGS` file a *clean pass*; renamed a **clean findings file**, which is what it always described, so §3's predicate is not read as redefining it | replacement | +| 2 | the Gate-A clean-signal sentence | the Gate-A section, C and W | it makes the `NO FINDINGS` signal the only route to clean; it becomes what lets a pass be read as clean without inspecting it, since a pass carrying Minors alone is clean and could never produce that file | replacement | +| 3 | the Severity bullet's resolve duty | Mechanics · Severity, C and W | the only statement of the duty and the only one without a scope; scoped to **the assigned fix set**, the boundary **D5** implied and no sentence carried. It names the set and never a decline — AC 3's second half, since naming one here would read as a waiver of resolution rather than the membership decision it is | replacement | +| 4 | the lens paragraph's unchanged-list | the profiles section, C and W | it asserts the Blocker/Major filter and the clean-final-pass rule are unchanged, which the ordering falsifies; scoped to **the lens sets**, which is what that paragraph is about | replacement | +| 5 | the Gate-A cadence | the Gate-A section, C and W | "validate, revise, re-run" makes a revision unconditional, telling a Minor-only pass to manufacture the repair the severity rule forbids; the revision becomes conditional on a repair being required | replacement | +| 6 | `a17`–`a19`, the floor paragraph's closure sentences | passage (a), C and W | they state closure and the early exit in the paragraph that owns the floor; trimmed to point at the ordering, which states them once | replacement | +| 7 | `a13` | passage (a) | "no rule about how a cycle closes is restated" becomes categorically false once the ordering exists; scoped to that paragraph, which is what it was written to police | replacement | +| 8 | `a16` | passage (a) | "fix Blocker/Major after each" stands unscoped beside a Severity bullet item 3 now scopes; it points at that rule instead of restating an unscoped version | replacement | +| 9 | `b3` | passage (b), C and W | it points at Mechanics and then **restates the four severity actions**, giving them two definitions; it becomes a pure pointer. W's pointer also changes target, matching C (§6) | replacement | +| 10 | `b7`, the fix-set definition | passage (b) | singular "the approved story or plan" leaves a cycle governed by several with no set at all; it becomes their **union**, plus obligations already accepted, **minus this cycle's declines** | replacement | +| 11 | `b8`, the membership test | passage (b) | it tests a finding against "that scope" independently, so a broadened scope puts back a finding `b7` keeps out; it tests against **the assigned fix set as `b7` computes it** | replacement | +| 12 | `b11`, the membership trigger | passage (b) | unqualified, it and the clean predicate decide a re-raised declined finding in opposite directions; it excludes a finding this cycle has already declined | replacement | +| 13 | `b12` | passage (b) | it resumes on the membership answer alone; it says the answer ends the hold and **defers to the ordering** for what the pass does next | replacement | +| 14 | `b13`, the question trigger | passage (b) | *new* is undefined, so an answered question re-raised stops the loop again on every pass; *new* excludes a question already answered in this cycle | replacement | +| 15 | `b17`–`b18` | passage (b) | they state their own version of what ends a suspension; they name it a suspension and defer to the ordering | replacement | +| 16 | the clearly-stuck closure sentences | passage (c), from the third condition to the end | they carry evaluation order in the paragraph that owns the reading; the paragraph keeps its reading, and the precedence sentence moves into the block **word for word**, which is what satisfies **D3** | replacement | +| 17 | `e7`, the two-tell threshold | passage (e) | unqualified, it and the ordering decide a clean two-tell pass at or above the floor in opposite directions; it is read **after** the clean-completion branch. Authority **D2** | replacement | +| 18 | the five-tells pointer | passage (e) | the passage says nothing about what its answer does; one sentence names the stop a suspension and defers to the ordering | addition | +| 19 | the handed-over severity question | Mechanics · Severity, passage (g) | the paragraph says the demotion/loop-health question is unsettled and mandates a stop; replaced by the answer, which is the (g) text below | replacement | +| 20 | the unknown-start strict-reading list | passage (i) | the list of what a cycle takes at its strictest omits the suspensions this change ships; it gains **every suspension binding, since unknown starting rules cannot waive an open hold** — the reason carried inline, as prompt-standards item 6 requires of the clause itself | addition | + +**Passage (g)'s replacement is design rather than bookkeeping, so it is stated here in full:** ``` **The demotion changes what a cycle must resolve, never what it counts.** The per-pass @@ -516,181 +308,131 @@ the whole paragraph, both copies, with the answer. NEW: that same judgement would hide it. ``` -Accounting: `g1` **dropped** (the unsettled statement, now settled); `g2`, `g3` **dropped** -(the interim report-and-stop duty existed only until the question was settled); `g4` -**dropped** in C (the ownership sentence, discharged by this change), and W, which never -carried it, gets the same replacement — removing the one deliberate story-path difference. +Two sentences are **not** edited and are named so nobody looks for them: the "Copy every record +into the squash body" sentence inside the human-exception block, and the "records every cycle +owes" list. This change ships no record. -**(h) Recording a human exception** (C:977–1033 / W:1161–1217) — **no longer edited.** The -answer-record block that was to follow it moved to the successor with **D9**. `h1`–`h26` kept, -unedited. +--- -**(i) When these rules bind** (C:153–167 / W:360–374) — **extend** the strict-reading list with -one item (§4 item 6); `i1`–`i16` kept, none replaced, the list being added to rather than -rewritten. +## 5. The passage map, and who carries the accounting -**(j) The squash carry** (C:892 / W:1076) — **no longer edited.** It was extended to name the -answer record, which moved to the successor; this change ships no record for it to carry. -`j1`–`j4` kept, unedited. +**The plan carries the kept / moved / replaced / dropped disposition for every one of the 135 +conditions in the committed inventory**, per passage, **beside the edit it belongs to** — which +is the only place it can be checked against the file being changed. **Story acceptance criterion +5 is satisfied by that list, not by this section**, and nothing here is dropped by being absent +here. This table says what happens to each inventoried passage, so the map stays complete at ten. -Also touched, outside the inventoried passages: §4 items 1–5. +| Passage | This change | Edits (§4) | +|---|---|---| +| (a) the floor paragraphs | edited — closure sentences trimmed to a pointer, the no-restating prohibition scoped, the per-pass fix command pointed at Mechanics | 6, 7, 8 | +| (b) what a loop absorbs | edited — owns both triggers and the fix set, so every qualification the ordering needs is made **here**, which is what keeps one definition per rule | 9, 10, 11, 12, 13, 14, 15 | +| (c) recognizing clearly stuck | edited — keeps its three-condition reading, stops carrying evaluation order; its precedence sentence moves into the block unchanged | 16 | +| (d) from pass 4 onward | **no longer edited.** The unavailable-history block moved to the successor with **D10** | — | +| (e) the five tells | edited — the threshold is read after clean completion, and a pointer says what its answer does | 17, 18 | +| (f) the two rules above do not compete | **unchanged.** "The two rules above" still names the absorb rule and the stuck reading; the block sits before both and adds no third rule between them | — | +| (g) Mechanics · Severity, the handed-over question | edited — the unsettled statement and its interim report-and-stop duty are replaced by the answer, in both copies, removing the one deliberate story-path divergence | 19 | +| (h) recording a human exception | **no longer edited.** The answer-record block that was to follow it moved to the successor with **D9** | — | +| (i) when these rules bind | edited — the strict-reading list is added to, not rewritten | 20 | +| (j) the squash carry | **no longer edited.** It was to name the answer record, which moved; this change ships no record for it to carry | — | + +**Three reversals are recorded here rather than left as silent narrowings**, because each was +made in an earlier revision of this spec and each contradicted something settled. Membership was +briefly re-read when the answer arrived, which `b6` and **D4** both forbid. The hold was briefly +narrowed to scope stops, which contradicted `c16` and the story's own third standing duty; it +attaches to **every** surfaced finding. And `c18` — no pass credited as clean on a clearly-stuck +surface — is **replaced** rather than kept: where that exit's regenerating findings are in-set +and the ceiling demotes them below Major, the pass is clean at effective severity, closes at or +above the floor and suspends below it. Authority **D3**, under which the old reading and D3's own +preserved sentence decide that pass in opposite directions. --- ## 6. Parity -The two copies must agree on every rule this spec changes. The block (§3), the (g) replacement -and every source edit ship byte-identical in C and W. The pre-existing divergences the inventory -found are handled as follows: +The two copies must agree on every rule this change ships: the block, the (g) replacement and +every edit in §4 are byte-identical in C and W. **The plan carries the divergence list** — which +pre-existing wording differences are deliberate and stay, which are not and are aligned — and +**performs the extraction and diff**, passage by passage, against the real files. One divergence +is decided here because it is a correctness call rather than a wording one: W's `b3` pointer +names "the severity rule" on the inventory's reasoning that W has no Mechanics section, which is +false, so W takes C's wording (§4 item 9). -| Divergence | Kind | This change | -|---|---|---| -| (b) cross-reference "Mechanics" (C) vs "the severity rule" (W) | **not deliberate**: the inventory's reason — that W has no such section — is false (W:968) | **aligned**: W takes C's wording (§5(b)) | -| (b) "exactly how" (C) vs "is how" (W); closing rationale reworded; C-only `infinite-portfolio-canvas` parenthetical | rationale and field citation | **as-is**, stated | -| (e) "you report" (C) vs "report" (W); C-only "Recorded rationale" Bricks paragraph | rationale | **as-is**, stated | -| (f) evidence framing; C-only hypothesis qualifier; punctuation of the late-Blockers clause | rationale | **as-is**, stated; (f) is not edited | -| (c)/(d) paragraph break: C fuses the surfacing block onto "Every pass report states three things" (no blank line at C:246/247); W separates them | structural | **aligned**: the (c) edit re-paragraphs the surfacing block, and C gains the blank line, so the floor-report paragraph stands alone in both | -| (e)/(f) paragraph break: W runs the five-tells paragraph into "The two rules above" (no blank line at W:472/473); C has the Bricks paragraph between | structural | **aligned**: W gains a blank line before "The two rules above"; the Bricks paragraph stays C-only | - -The parity check at execution: extract each edited passage from both files by its lead phrase -and `diff` them; the only differences permitted are the rows marked as-is, and the result is -part of the evidence entry (§7). Any other difference is a defect, not a wording choice. - -The same extraction runs a second, different check within each copy: that `b11` and `b13` as -edited say what the block cites them as saying. **It compares the complete predicates, not a -shared phrase** — comparing only the phrase both carry is what let an earlier revision call two -wordings equivalent while one of them lacked the decline exception the other had. Each side is -read whole: - -| Block cites | `b11`/`b13` as edited must carry | -|---|---| -| a finding **outside the assigned fix set** | `b11`'s "a correction that leaves that set… like any other out-of-scope finding" | -| …**and not one this cycle has declined** | `b11`'s "unless this cycle has already declined that same finding" (§5(b)) | -| a finding **opening a new structural or contract question** | `b13`'s "opens a **new structural or contract question**", word for word | -| in-set or not | `b13` carries no membership condition, and adding one would narrow it | - -A difference in either direction is a defect: a condition in the block and not in `b11`/`b13` -ships two triggers that disagree, and one in `b11`/`b13` and not in the block means the block -cites a rule it has not read. +The same extraction runs a second check within each copy: that `b11` and `b13` as edited say what +the block cites them as saying, **comparing the complete predicates and not a shared phrase** — +including `b11`'s already-declined exception and `b13`'s already-answered qualification, in both +directions. A condition in the block and not in the source ships two triggers that disagree; one +in the source and not in the block means the block cites a rule it has not read. --- ## 7. Verification -**The mode is read from the story's header at execution**, never from here — the same rule the -spec's own header states, and the reason no value is named in this heading. What follows is -what each level of that mode obliges, so that whichever it carries has its evidence described: -the battery, the check, and the named verification of the risk path. +**The mode is read from the story's header at execution**, never from here — the same rule this +spec's header states. What follows is what each level of that mode obliges. **Battery.** The quality command in `AGENTS.md` § Commands runs once, green, at the Gate-B WIP commit (`scripts/check-version-bump.sh main` needs the committed bump, §8). -**The check — what it must establish, and where it is built.** Every edit that changes a -standing meaning owes a **discriminating pair of counts**: one showing the new wording present, -one showing the old wording gone. Each half is run in **both copies** and against **both** the -working tree and the parent tree (`git show 7c0d475:CLAUDE.md` and -`git show 7c0d475:plugins/dev-workflow/commands/workflow-init.md`), so every assertion is -observed passing where the change exists and failing where it does not. A one-sided presence -check is not enough: a copy carrying the new wording **and** the old one satisfies it, which is -exactly the two-instructions-that-disagree failure §4 exists to prevent. An edit that only adds -to a list is checked by **presence alone**, because nothing is being replaced. - -**Which edits owe a pair**, named here because this spec knows which sentences it changes: -§4 items 1–5; and in §5, `a17`–`a19` and `a13` and `a16` in (a); `b7`, `b11`, `b12`, `b13`, -`b17`–`b18` in (b), plus `b3` in **W alone**, since C already carries that target wording and -must be unchanged, which is what makes it an alignment rather than an edit to both; the -replacement in (c); `e7` in (e); and the paragraph replacement in (g), whose old sentence is the -one site where the parent is present and the change removes it. **Presence alone:** the §3 -block's lead phrase, the (e) pointer sentence, and §4 item 6. - -**The plan builds each pair against the real files and runs both directions there.** It carries -two constraints: a counted fragment must be **single-line in the file it is grepped from**, since -one spanning a line break makes `grep -F` count 0 and read as a failure; and the new wording must -therefore be **installed unwrapped**. Both are constraints on how an edit is written, not on what -it means. - -**Why the fragments are not enumerated here, stated as a residual rather than repaired again.** -Four consecutive revisions of this spec listed them, and each list contained at least one -fragment that could not do what it claimed — a substring preserved inside its own replacement, so -the old-wording-gone count could never reach zero; three quoted across their line wraps, so they -counted zero in a correct tree; and a meaning-changing passage with no pair at all. The cause is -structural: an exact substring check for text that does not yet exist can only be guessed, and a -guess that is wrong reads as a failed check rather than as a wrong check. **Nothing here verifies -that the list of edits above is complete, or that the plan's chosen fragments discriminate.** The -enumeration moved to where the text exists; the completeness claim did not move with it, because -nothing supports it. - -If the claim "the ordering ships in both copies" were false, one working-tree count would be -**0** or the parent-tree counts would not differ from it. The wiring can produce that -observation: each grep reads the file bytes at the named revision and nothing supplies its own -input. **The counterfactual is ABSENT, and is claimed as absent** — the parent carries no -ordering block, and the (g) sentence is the one site where the parent is present and the change -removes it. Nothing is claimed as "contradictory". - -**The named verification of the risk path** (story AC 4) is a **next-state table**, in the plan -and quoted by the closing commit body. **Rows** — every stop the shipped text names: membership -stop, question stop, a finding carrying both triggers, clearly-stuck exit, two-tell stop, a hold -awaiting its answer, accept, decline, **a continue answer on an artifact left unrevised** — -whose next state is a further pass, the row that tests that continue consumes the reading — -**a stop answer** — whose next state is the **parked** cycle, distinct from the -suspended-awaiting-answer state it was in before, the row that tests AC 4 at the one transition -that produces no pass — two or three suspensions at once, a -below-floor clean pass with no suspension, **a below-floor clean pass carrying a health -reading** — whose next state is the **suspension**, the row that tests the third branch's -qualification — a zero-finding pass, and the unknown-start fallback; **and the stateful -transitions**: **a finding this cycle declined re-raised on a later pass** — whose next state is -no membership trigger and a pass still eligible for clean, the row that tests `b11`'s -qualification — **a structural question this cycle answered, re-raised** — whose next state is -no question stop, the row that tests `b13`'s — **a decline that leaves nothing to revise** — -whose next state is a further pass on the **unrevised** artifact, the row that tests the continue -branch against a valid path with no repair in it — an **accepted** finding re-raised before -repair, **the fix set broadened -while a hold is still awaiting its answer** — whose next state is the hold still standing, the -row that tests the frozen reading — and **a confirmed profile change while a hold is still -awaiting its answer**, which is **two rows and not one**: the profile change itself, whose next -state is the hold still standing and **no pass run**, since a standing hold forbids one; and then -the answer, **in each direction**, whose next state is the parked cycle or a resume, with only a -resuming answer starting the further pass under the current profile that a profile change costs. -Split because final acceptance re-reads every profile, and a hold surviving a profile change is -not the same claim as an answer paying for one — a single row asserting both would have to run a -pass through an unanswered hold to be filled in. Then a -fix set narrowed so an in-set finding falls outside it, and a `full` Gate-B pass with one branch -clean and the other carrying an in-set Blocker. -**Columns** — the **governing-scope state** (the fix set as currently assigned), the -**answers already given in this cycle**, and the **profile as currently read**, beside the -user's answer for this row, then the input that ends the row, the state afterwards, and the -shipped line the row reads, in both copies. - -**The oracle.** A row **fails** when its required answer does not produce a **distinct** -resumable or closed state — the same stop returning, whether or not an input was consumed — -or when it closes with **any** closure precondition unmet: an in-set Blocker or Major, a -standing hold, an unanswered question, the derived floor, or a cited-set or profile change -during the pass. Naming only the first two would pass the exact no-progress defect AC 4 cites -from the parent cycle. The wiring can produce that observation because every row is filled from -the shipped text rather than from this spec, and the inputs the `fic2` instrument omitted — the -user's answer and the governing-scope state — are columns here. That is what makes this **not -the `fic2` decision matrix**, whose two defects Gate B found in the technique itself: a state's -inputs must include every input the rule reads, and a counterfactual must distinguish ABSENT -from CONTRADICTORY. **No fixture per predicate is built** — parked in the story's §2, not -reopened. - -**Evidence entry**, in the closing commit body, names: the battery run; every pair the plan -built, with the fragment it counted and its working-tree and parent-tree counts in each copy, -and every presence check beside them; the §6 parity diff — the passages extracted, the -differences observed, that each is one of the permitted rows, and the `b11`/`b13` equivalence -result; and the next-state table's location in the plan plus its row count. It is revalidated -before every Gate-B re-review and before the closing amend, as §5 requires. +**The check — what it must establish, and where it is built.** Every edit that changes a standing +meaning owes a **discriminating pair of counts**: one showing the new wording present, one +showing the old wording gone. Each half runs in **both copies** and against **both** the working +tree and the parent tree, so every assertion is observed passing where the change exists and +failing where it does not. A one-sided presence check is not enough: a copy carrying the new +wording **and** the old one satisfies it, which is the two-instructions-that-disagree failure §4 +exists to prevent. An edit that only adds to a list — §4 items 18 and 20, and the block itself — +is checked by **presence alone**, because nothing is being replaced. **The plan builds each pair +against the real files and runs both directions there**, under two constraints: a counted +fragment must be **single-line in the file it is grepped from**, since one spanning a line break +makes `grep -F` count zero and read as a failure; and the new wording must therefore be +**installed unwrapped**. + +**Why no fragment is named here, stated as a residual rather than repaired again.** Four +consecutive revisions of this spec listed them, and each list held at least one fragment that +could not do what it claimed: one preserved inside its own replacement, so its +old-wording-gone count could never reach zero; three quoted across their line wraps, so they +counted zero in a correct tree; one meaning-changing passage with no pair at all. The cause is +structural — an exact check for text that does not yet exist can only be guessed, and a wrong +guess reads as a failed check rather than a wrong one. **Nothing verifies that §4's list is +complete or that the plan's fragments discriminate.** The enumeration moved to where the text +exists; the completeness claim did not, because nothing supports it. + +The counterfactual is **ABSENT and is claimed as absent**: the parent carries no ordering block, +and passage (g)'s old sentence is the one site where the parent is present and the change removes +it. Nothing is claimed as "contradictory" — the second of the two defects Gate B found in the +`fic2` instrument. + +**The named verification of the risk path** (story AC 4) is a **next-state table**, written in the +plan and quoted by the closing commit body. **Its claim is narrow and stated as such: it covers +answer-state transitions once the predicates producing them are established**, which is the first +`fic2` defect — a state's inputs must include every input the rule reads — answered by +restricting the claim rather than by widening the table. So the plan writes **separate named +checks** for what the table therefore does not establish: that a logical pass was validated +across every required branch file, and that each final-acceptance precondition the block cites +held — the floor, the cited set and profile, and the evidence entry's revalidation. **No fixture +per predicate is built**; that question is parked in the story's §2 and is not reopened. + +**The oracle.** A row **fails** when its required answer does not produce a **distinct** resumable +or closed state — the same stop returning, whether or not an input was consumed — or when it +closes with **any** closure precondition unmet: an in-set Blocker or Major, a standing hold, an +unanswered question, the derived floor, a changed evidence entry, or a cited-set or profile change +during the pass. Naming only the first would pass the exact no-progress defect AC 4 cites from +the parent cycle. + +**Evidence entry**, in the closing commit body, names: the battery run; every pair the plan built +with its counts in each copy and each tree, and every presence check beside them; the §6 parity +diff and the `b11`/`b13` equivalence result; and the next-state table's location plus its row +count. It is revalidated before every Gate-B re-review and before the closing amend, as §5 +requires. **One observability residual, stated because the lens set asks for it and nothing here answers it.** A closing commit body records that a cycle closed and what its curve was; it records **nothing about which exit the cycle took** — whether a scope stop, a clearly-stuck surface or -two tells ever suspended it, what was asked, or how it was answered. A reader of history -therefore cannot audit that every suspension was answered before the cycle closed. Neither the -evidence entry above nor the per-pass curve supplies this, and saying otherwise would be the -overclaim `AGENTS.md` names as this repo's most persistent defect. The transport that could -carry it left with the record (§9). **No story has taken it**, and naming one that has not is -the same defect in a smaller place, so it is recorded here as an **admitted gap of this change** -— unowned, unguarded, and available to whoever picks it up. +two tells ever suspended it, what was asked, or how it was answered. A reader of history cannot +audit that every suspension was answered before the cycle closed. Neither the evidence entry nor +the curve supplies this, and saying otherwise would be the overclaim `AGENTS.md` calls this +repo's most persistent defect. The transport that could carry it left with the record (§9), and +**no story has taken it** — naming one that has not is the same defect in a smaller place. An +**admitted gap of this change**: unowned, unguarded, open to whoever picks it up. --- @@ -699,21 +441,21 @@ the same defect in a smaller place, so it is recorded here as an **admitted gap - **Invariant 11 — `docs/prompt-standards.md`, all 12 items.** Most at risk: item 6, every constraint in the shipped block carrying its reason in the same sentence — **including the three exempted until pass 6**, which now carry theirs inline in §3: that no other pass outcome - closes, that a zero-finding pass is clean whatever the floor, and that decline is available - only at a membership stop. The exemption was wrong twice over: item 6 admits no "settled - elsewhere" clause, and a scaffolded copy cannot reach the story the reasons were said to live - in. Then **item 8 (token-lean), which an earlier revision claimed on the wrong ground** — it - said the block replaces closure sentences rather than adding beside them, while the block was - in fact restating triggers, duties, preconditions and the severity answer that their own - paragraphs still defined, which is two authorities per copy and the drift had already begun. - The claim now rests on what the block does: it is authoritative for the evaluation order and - for closure and **cites** every other rule where that rule is defined, so each has one - definition in the shipped text and §3's table is the check; where a cited rule had to change - to agree, it changed at its source (§4, §5) rather than being restated. Then item 3 (the stop - answer produces a named state, **parked**, with its own restart transition). + makes a cycle eligible to close, that a zero-finding pass is clean whatever the floor, and that + decline is available only at a membership stop. The exemption was wrong twice over: item 6 + admits no "settled elsewhere" clause, and a scaffolded copy cannot reach the story the reasons + were said to live in. §4 item 20 carries its reason in the shipped clause for the same reason. + Then **item 8 (token-lean), which an earlier revision claimed on the wrong ground** — it said + the block replaces closure sentences rather than adding beside them, while the block was in + fact restating triggers, duties, preconditions and the severity answer that their own + paragraphs still defined, which is two authorities per copy. The claim now rests on what the + block does: it is authoritative for the evaluation order and for closure and **cites** every + other rule where that rule is defined, so each has one definition in the shipped text and §3's + table is the check. Then item 3 (the stop answer produces a named state, **parked**, with its + own restart transition). - **Don't: "Never replace a decision procedure without accounting for its old conditions."** - Satisfied by §5, against the committed inventory, with an entry for every inventoried passage - including the four this narrowing no longer edits. + Satisfied by the plan's per-condition disposition list against the committed inventory (§5), + with this spec's passage map keeping the ten-passage set complete. - **Don't: "Never rename or delete a doc section without grepping for references first."** The (g) sentence names the story path; the grep finds it at `CLAUDE.md:815` (the site itself) and in three artifacts of the parent cycle @@ -725,7 +467,7 @@ the same defect in a smaller place, so it is recorded here as an **admitted gap `plugins/dev-workflow/commands/workflow-init.md` is under `plugins/`, so `plugins/dev-workflow/.claude-plugin/plugin.json` goes `0.11.0 → 0.12.0` with a `plugins/dev-workflow/CHANGELOG.md` entry: a minor bump, the template gaining a closure - ordering and nineteen edited or extended sentences. + ordering and twenty edited or extended sentences. - **Invariant 4 / the hook.** Untouched: `plugins/dev-workflow/hooks/codex-gate.sh` is not edited, and the §5 heading it greps (`Cross-Model Review`) does not move. @@ -748,32 +490,37 @@ decision of 2026-09-10, with the evidence in §1: - the rollback reading — **what an open cycle owes when a revert removes the rules it started under**. This change does not answer it and no longer claims to. An earlier revision said the activation paragraph already supplies the transition; that is wrong in a specific way. A revert - of this change removes the ordering, the source edits **and §4 item 6's "every suspension - binding" extension together**, so the stricter reading such a cycle would fall back to is - itself part of what the revert takes away, and it no longer mentions the suspensions that cycle - is holding. What remains is the pre-change §5 — the loose ordering this story exists to - replace — read by a cycle that started under a different one. **Whether that is enough is - unanswered here and nothing is shipped for it**, since no record identifies the rule revision a - cycle started under. An admitted residual, and the successor's question; + of this change removes the ordering, the source edits **and §4 item 20's suspension-binding + extension together**, so the stricter reading such a cycle would fall back to is itself part of + what the revert takes away, and it no longer mentions the suspensions that cycle is holding. + What remains is the pre-change §5 — the loose ordering this story exists to replace — read by a + cycle that started under a different one. **Whether that is enough is unanswered here and + nothing is shipped for it**, since no record identifies the rule revision a cycle started + under. An admitted residual, and the successor's question; - the slot-discriminator dissolution deferred here by Plan C's Tasks 19 and 20; - **the partial-adoption guard as it applied to the record.** It was built as a "closure-record contract" naming the record, with a marker on every mergeable hunk, and it leaves with the record it was named for. +**Moved to the plan** — the implementation artifact, not a story: the per-condition disposition +list for all 135 inventoried conditions (§5); the OLD and NEW text of all twenty edits with their +line ranges (§4); the parity divergence list and the extraction-and-diff (§6); and every +verification fragment with its counts (§7). Each is work this change still owes; none of it is +work a spec can do correctly, because all four are checked against files the plan edits. + **One residual this change owns rather than moves: partial adoption of the narrowed set.** The set is mutually dependent — the ordering's clean predicate needs §4 item 3's boundary, its two -senses of *clean* need items 1 and 2, and its triggers need §5(b)'s `b7`, `b11` and `b13` — and -**nothing catches a downstream merge that takes some of it**. Said exactly: the live one-contract -paragraph (C:880–890 / W:1063–1074) names the nonce, the slots, the provenance line, the curve, -the carry rule and the unknown-start semantics, and **it does not name this block or any of its -coupled edits**, so its coherence rule does not reach them. An earlier revision of this spec said -that rule caught them; it does not, and a project can take the clean predicate without the scoped -resolve duty, or either sense of *clean* without the other, and run. **That is an admitted unsafe -state, not a guarded one, and it is this change's own.** An earlier revision assigned the guard to -the successor; that was wrong on the face of the successor's own scope, which is record -durability and **excludes the closure ordering by name**, so the assignment named an owner that -had not taken it — the same defect as claiming a mechanism that does not exist. Nothing here -builds one, and a downstream project gets no coverage for this set. +senses of *clean* need items 1 and 2, and its triggers need items 10 through 14 — and **nothing +catches a downstream merge that takes some of it**. Said exactly: the live one-contract paragraph +names the nonce, the slots, the provenance line, the curve, the carry rule and the unknown-start +semantics, and **it does not name this block or any of its coupled edits**, so its coherence rule +does not reach them. An earlier revision of this spec said that rule caught them; it does not, and +a project can take the clean predicate without the scoped resolve duty, or either sense of *clean* +without the other, and run. **That is an admitted unsafe state, not a guarded one, and it is this +change's own.** An earlier revision assigned the guard to the successor; that was wrong on the +face of the successor's own scope, which is record durability and **excludes the closure ordering +by name**, so the assignment named an owner that had not taken it — the same defect as claiming a +mechanism that does not exist. Nothing here builds one. **Out of scope and parked**, unchanged: From c7c140331b500effec156774746d786c341793a6 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Fri, 11 Sep 2026 09:31:07 +0200 Subject: [PATCH 021/181] docs(stories): cut the two bounds to what the field record supports An independent assessment (.context/assessment-brief.md, read-only) checked the two bounds proposed on 2026-09-10 against the committed field reports. Both are cut back; neither is withdrawn. PASS CEILING -> a report line, not a fourth stop. Kept: the gap is real and measured. Pass 12 of awsf1ec771 sat at Blocker+Major 14 with zero of five tells, invisible to every rule the kit has. Dropped: the mandatory stop-and-surface. Section 5 already has three exits and they fire -- awsf1ec771 at passes 9 and 10, Plan C at passes 4 and 5 -- and the answer was "continue" every time, because a stop hands a question over and ends nothing. And it costs where the loop works: the rle spec cycle ran 34 passes, closed clean, and still returned Blockers at passes 30-32, so a floor+3 ceiling would have interrupted it roughly nine times for no new information. SIZE LIMIT -> a restatement prohibition, not a line count. Kept: an oversized artifact is dealt with before pass 1, and both measurements. Dropped: the line count as the test. C1 falsifies it -- nine sentences in eight replacements, and its curve rose anyway (8, 8, 10) until it stopped on two tells. The field record draws the conclusion itself: "the cost is not carried by the plan's size". What ships instead is what both successful cuts actually removed: an artifact quotes no text it does not change and restates no rule its own repo already states. That is the one generator every field report names independently. Acceptance criteria 7, 8 and 9 follow. Criterion 9 shrinks on purpose: neither bound adds a reading to the loop, so no composition rule against the three existing exits is owed. RECORD DURABILITY -> deferred, trigger-gated, not re-scoped. The two records section 5 requires in a closing commit body exist in exactly one commit out of 110 -- 7c0d475, the commit that shipped the rule requiring them. There is no second data point on whether the discipline holds, so durability rules would be built on a record produced once by its own inventor. Trigger: two further cycles close and write a conformant provenance line and curve. Checkable with one command, named in the banner. Per the AGENTS.md Don't on replacing a decision procedure, each proposal's old conditions are marked kept or dropped in the story text rather than rewritten away. No gate: every staged path is docs/**.md, which section 5 exempts as prose -- these describe the product rather than being it, and neither story is executable (both profiles are still unconfirmed). Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- ...09-10-harness-finding-termination-story.md | 88 ++++++++++++------- .../2026-09-10-record-durability-story.md | 19 +++- 2 files changed, 74 insertions(+), 33 deletions(-) diff --git a/docs/superpowers/stories/2026-09-10-harness-finding-termination-story.md b/docs/superpowers/stories/2026-09-10-harness-finding-termination-story.md index 738605f..803c723 100644 --- a/docs/superpowers/stories/2026-09-10-harness-finding-termination-story.md +++ b/docs/superpowers/stories/2026-09-10-harness-finding-termination-story.md @@ -50,8 +50,9 @@ rung harder (prose → lint → type → test)". **The concrete miss is measurable in this repository's own cycle.** Gate-A spec cycle nonce `awsf1ec771` on the loop-rule consolidation design ran twelve passes without a clean pass. Blocker+Major by pass: **20, 15, 10, 11, 15, 13, 9, 8, 12, 17, 14, 14.** The pass-by-pass record is -`.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md`, which is gitignored — hence the figures -quoted rather than only cited. +`.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md`, tracked since 2026-09-10 — the figures +are quoted here because they were quoted before that path was tracked, and they now cite something +a reader can open. **Two mechanisms inside that cycle each produced the same shape of finding for four rounds, and each was repaired one instance at a time.** The spec's own verification assert list produced @@ -76,24 +77,46 @@ move between "absorb another round of this" and "hand the whole cycle to the hum that keeps producing the same shape of finding stops consuming rounds while the loop continues on everything else. -**Two bounds exist that today do not, and both are mechanical rather than a reading.** Added -2026-09-10 from the same session that produced the evidence above, on Daniel's question of how a -project developing this kit avoids blocking itself while using it. The diagnosis that prompted them -is that the cycle was not slow because of self-application; it was slow because **the loop had no -ceiling and the artifact had no size limit**, and self-application only multiplied the readings each -round had to consider. - -- **A pass ceiling.** §5 fixes a floor and no maximum. The only two ways up and out — the - clearly-stuck exit and the two-tell threshold — both require a *reading*, so a cycle that trips - neither grinds without anyone being obliged to decide. A ceiling turns that into a mandatory - stop-and-surface, the same shape the two-tell rule already has, where the human chooses to split, - to accept with stated residuals, or to continue for a named reason. Cycle `awsf1ec771` ran - **thirteen** Gate-A spec passes against a floor of 3 and reached no clean pass. -- **An artifact size limit before the cycle starts.** §5's sizing guidance — "prefer smaller specs - with named interfaces and let the plan carry the detail" — is advice with no number, and it was - read and not followed. The same spec reached **989 lines**, of which the design was **173**; the - remaining **58%** was bookkeeping about the change, and it took roughly half the findings of every - pass. A limit checked before the first pass is a `wc -l`, not a judgement. +**Two bounds exist that today do not. Both were proposed on 2026-09-10 as mandatory mechanisms and +both were cut back on 2026-09-10 after an independent assessment read the field record against +them** (`.context/assessment-brief.md`; the reading is in this section). The diagnosis that prompted +them was that the cycle was slow because **the loop had no ceiling and the artifact had no size +limit**. The first half survives as a *visibility* gap; the second half does not survive contact +with the record. + +Per the AGENTS.md Don't on replacing a decision procedure, what each proposal required and what +became of it: + +- **A pass ceiling — kept as a report line, dropped as a stop.** *Kept:* the gap it named is real + and measured. §5 fixes a floor and no maximum, and its two ways out both need a *reading*, so a + flat loop can trip nothing: pass 12 of `awsf1ec771` sat at Blocker+Major 14 with **zero of five + tells**, invisible to every rule the kit has. Cycle `awsf1ec771` ran **thirteen** Gate-A spec + passes against a floor of 3 and reached no clean pass. *Dropped:* the mandatory stop-and-surface, + on two grounds. **§5 already has three exits and they fire** — `awsf1ec771` at passes 9 and 10, + Plan C at passes 4 and 5 — and the answer was "continue" every time, because a stop hands a + question over and ends nothing; a fourth stop asks the same question a fourth time. And **it + costs where the loop works**: the `rle` spec cycle ran 34 passes, closed clean, and still returned + Blockers at passes 30–32 (`docs/field-reports/2026-08-29-gate-a-rle-cycle-evidence.md`) — a + floor+3 ceiling would have interrupted it roughly nine times for no new information. *What ships + instead:* the pass report already carries the floor and the trend; it also carries **passes run + against the derived floor**, as a number. That makes the flat loop visible, obliges no new + reading, and composes with nothing. +- **An artifact size limit — replaced by a restatement prohibition.** *Kept:* the observation that + an oversized artifact must be dealt with before pass 1 rather than diagnosed after it, and both + measurements behind it. The `awsf1ec771` spec reached **989 lines** of which the design was + **173**; cutting it to 532 (design 192) took Blocker+Major from 14 to 10, and the `rle` spec cut + from 604 to 332 lines took it from 28 to 12. *Dropped:* the line count as the test. **C1 + falsifies it.** That plan was nine sentences in eight replacements — about as small as a plan of + this kind gets — and its curve rose anyway (Blocker+Major 8, 8, 10) until it stopped on two + tells; the field record draws the conclusion itself: "the cost is not carried by the plan's + **size**" (`docs/field-reports/2026-08-30-gate-a-rle-plan-cycles.md`). A line limit would pass a + 200-line artifact made entirely of restatement and split a 600-line one made entirely of design. + *What ships instead:* the rule states what both successful cuts actually removed — **an artifact + quotes no text it does not change, and restates no rule its own repo already states.** That is + the one generator every field report names independently: "the remedy is never a better summary"; + "a plan that restates a protocol its own repo already governs creates a second copy that drifts, + and Gate A will review the copy instead of the work"; "the block restates rules their own + paragraphs still own". Checkable by reading, before pass 1. **A third bound already exists in §5 and was simply not honoured**, which is worth recording because it needed no new rule: the instruction to settle mechanically what a parser can decide before @@ -133,17 +156,20 @@ run again, and roughly a quarter of the later passes' findings were things it de - [ ] **Per-mechanism termination and the existing five tells do not duplicate or contradict each other.** Both copies say which applies when both would fire, and neither weakens the two-tell mandatory stop. -- [ ] **A cycle cannot run unbounded without a human deciding.** A ceiling exists, it is stated as a - number derived the way the floor is, and reaching it without a clean pass is a mandatory - stop-and-surface naming the options. Checkable by reading: a pass count alone decides it, with - no reading of a curve or a cluster. -- [ ] **An oversized artifact is stopped before the cycle starts, not diagnosed after it.** A stated - limit applies before pass 1, it is checkable by counting lines, and the shipped text says what - an over-limit artifact does instead — split, or move detail to the plan behind a named - interface. -- [ ] **The two bounds and the existing exits compose without a fourth reading.** For every state - where a bound and an exit could both apply, the shipped text says which governs and why, or - names the pair as unable to co-occur. +- [ ] **A loop that is going nowhere is visible without anyone reading a curve.** Every pass report + states the passes run against the derived floor, as a number, from pass 1 onward. Falsifiable + by reading a report: the number is there or it is not. **It creates no stop and no exit**, so + the shipped text adds no rule about how it ranks against the three that exist — a criterion + that demanded one would rebuild the lattice this change exists to avoid. +- [ ] **An oversized artifact is stopped before the cycle starts, not diagnosed after it** — and the + test is restatement, not length. The shipped text says an artifact quotes no text it does not + change and restates no rule its own repo already states, and says what an offending artifact + does instead: delete the restatement and cite, or move the detail to the plan behind a named + interface. Falsifiable by reading one artifact against the rules it cites. **No line count + appears**, because C1 converged worse at nine sentences than the `rle` spec did at 332 lines. +- [ ] **Neither bound adds a reading to the loop.** The visibility line is reported and never + evaluated; the restatement rule applies before pass 1 and is spent by then. The shipped text + states both facts, so no reader looks for a composition rule against the existing exits. - [ ] **Every condition of the replaced prose is accounted for**, each marked kept, moved or deliberately dropped, per the AGENTS.md Don't — and **the two prompt copies stay in parity** on every rule this story changes, deliberate wording differences stated as such. diff --git a/docs/superpowers/stories/2026-09-10-record-durability-story.md b/docs/superpowers/stories/2026-09-10-record-durability-story.md index 548e679..a1325f4 100644 --- a/docs/superpowers/stories/2026-09-10-record-durability-story.md +++ b/docs/superpowers/stories/2026-09-10-record-durability-story.md @@ -8,6 +8,20 @@ > describes was split out of a cycle that is still running, and the successor starts when someone > picks it up. The proposal's reasons are in §5. +> **DEFERRED — trigger-gated, Daniel 2026-09-10.** Not withdrawn and not re-scoped; the problem +> below is real and the story stands as written. What is unbuilt is the premise it rests on. +> **The two records §5 requires in a closing commit body exist in exactly one commit out of 110** +> — `7c0d475`, the commit that shipped the rule requiring them. Every earlier commit predates the +> rule and every later one is mid-cycle. So there is **no second data point on whether the +> discipline holds at all**, and rules for how a record survives a session, an amend, a squash or +> a rollback would be built on a record that has been produced once, by its own inventor. +> **Trigger: two further cycles close and write a conformant provenance line and curve.** If the +> discipline holds, this story starts with its premise established. If it does not, the answer is +> a different transport rather than durability rules for a commit body nobody fills in — and that +> is a cheaper thing to learn before the eight criteria below are spent on it. +> Checking the trigger is one command: +> `git log --all --format='%b' | grep -cE 'floor [0-9]+ per '`. + ## 1. Problem statement **One question runs under every finding this story inherits: does a record survive?** §5 asks an @@ -36,8 +50,9 @@ legacy cycles that hold no nonce. The closure ordering's own five findings were artifact was carrying two subjects, and only one of them was converging. **Where the pass-by-pass evidence lives:** `.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md`, -with the per-pass findings files beside it. That directory is gitignored, so the figures above are -quoted here rather than only cited — a reader on another machine has no way to open them. +with the per-pass findings files beside it. That directory was gitignored when these figures were +quoted, which is why they are quoted rather than only cited; it is tracked from 2026-09-10, so a +reader on another machine can now open them. **Parent:** `docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md`, whose §2 records the split and whose §4 holds the decisions this story implements. From c06dd694a15230dc88b238ad00e615f1676d75d5 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Fri, 11 Sep 2026 09:31:25 +0200 Subject: [PATCH 022/181] chore: track .context/codex-reviews/ so pass curves survive on their own Three field reports exist only because someone hand-carried per-pass counts out of this directory before a .context/ clear destroyed them (2026-08-26-fic2, 2026-08-29-gate-a-rle, 2026-08-30-gate-a-rle-plan-cycles). Tracked, those counts survive without the rescue, and any curve is one grep: grep -cE '^(BLOCKER|MAJOR|MINOR|NIT) \|' Checked before changing anything, because AGENTS.md forbids describing a gate without reading what it compares: codex-gate.sh excludes .context/ from ALL THREE fingerprint components independently of this file -- a :(exclude) pathspec on the tracked diff, git rm --cached on the throwaway index, the same pathspec on add -A -- and its own comment records that ".context/ is committed in some projects". So Gate B does not break. Verified with git check-ignore: the mutable counters, spec-precheck.py and the assessment files stay ignored; codex-gate.on stays visible. KNOWN DIVERGENCE, deliberate and local. CLAUDE.md section 5 ("Why .context/codex-reviews/") still says to ignore that path, and the /workflow-init template still writes that instruction into target projects. Both are unchanged on purpose: editing section 5 is a two-copy prompt change owing a full Gate-B cycle and a plugin version bump, which is a separate decision. This repo deviates; nothing shipped does. The divergence is stated in the .gitignore comment so it is met as a decision rather than as an oversight. Two todos.md rows said nothing here is durable because .context/ is git-ignored. That is now false for this repo and still true for target projects; both rows are corrected rather than deleted, and both keep their triggers. RECORDS cycle 4win97lk9j; floor 3 per none; hook reminder threshold absent cycle 4win97lk9j; Gate B: skipped (see skip reason) Skip reason: behaviourally trivial. One .gitignore negation, two backlog corrections, and review artifacts this repo already produced. No prompt, hook, plugin manifest or scaffolded template is touched, so nothing the product ships changes and no plugin version bump is due. No story is cited, so the skip keeps the existing judgement-based form and owes no mode-derived evidence entry. Battery, run 2026-09-10 over this content, exit 0: shellcheck --shell=sh over both hook files and all four checker files; the hook suite under sh and under dash; check-invariants.test.sh (148 assertions) then check-invariants.sh; check-version-bump.test.sh (36 assertions) then check-version-bump.sh main; claude plugin validate . --strict. All green. That is the battery, not a review, and this body says so. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- ...-CLOSURE.2026-08-12-ledger-supersession.md | 108 ++ .context/codex-reviews/gate-a-plan-CLOSURE.md | 69 ++ ...-dispositions.2026-07-26-profiles-cycle.md | 16 + ...ispositions.2026-07-30-classifier-cycle.md | 54 + ...dispositions.2026-08-04-hardening-round.md | 78 ++ ...ositions.2026-08-12-ledger-supersession.md | 57 + ...plan-pass-1.2026-07-30-classifier-cycle.md | 41 + ...-plan-pass-1.2026-08-04-hardening-round.md | 20 + ...n-pass-1.2026-08-12-ledger-supersession.md | 19 + .context/codex-reviews/gate-a-plan-pass-1.md | 19 + ...-dispositions.2026-07-26-profiles-cycle.md | 11 + ...ispositions.2026-07-30-classifier-cycle.md | 38 + ...dispositions.2026-08-04-hardening-round.md | 61 + ...ositions.2026-08-12-ledger-supersession.md | 50 + ...plan-pass-2.2026-07-30-classifier-cycle.md | 33 + ...-plan-pass-2.2026-08-04-hardening-round.md | 12 + ...n-pass-2.2026-08-12-ledger-supersession.md | 12 + .context/codex-reviews/gate-a-plan-pass-2.md | 21 + ...ispositions.2026-07-30-classifier-cycle.md | 48 + ...ositions.2026-08-12-ledger-supersession.md | 68 ++ ...plan-pass-3.2026-07-30-classifier-cycle.md | 36 + ...-plan-pass-3.2026-08-04-hardening-round.md | 20 + ...n-pass-3.2026-08-12-ledger-supersession.md | 10 + .context/codex-reviews/gate-a-plan-pass-3.md | 18 + ...ispositions.2026-07-30-classifier-cycle.md | 45 + ...dispositions.2026-08-04-hardening-round.md | 69 ++ ...ositions.2026-08-12-ledger-supersession.md | 64 ++ .../gate-a-plan-pass-4-dispositions.md | 66 ++ ...plan-pass-4.2026-07-30-classifier-cycle.md | 30 + ...-plan-pass-4.2026-08-04-hardening-round.md | 12 + ...n-pass-4.2026-08-12-ledger-supersession.md | 9 + .context/codex-reviews/gate-a-plan-pass-4.md | 16 + ...ispositions.2026-07-30-classifier-cycle.md | 34 + ...dispositions.2026-08-04-hardening-round.md | 59 + ...ositions.2026-08-12-ledger-supersession.md | 58 + ...plan-pass-5.2026-07-30-classifier-cycle.md | 20 + ...-plan-pass-5.2026-08-04-hardening-round.md | 10 + ...n-pass-5.2026-08-12-ledger-supersession.md | 7 + .context/codex-reviews/gate-a-plan-pass-5.md | 10 + ...ispositions.2026-07-30-classifier-cycle.md | 60 + .../gate-a-plan-pass-6-dispositions.md | 69 ++ ...plan-pass-6.2026-07-30-classifier-cycle.md | 38 + ...n-pass-6.2026-08-12-ledger-supersession.md | 3 + .context/codex-reviews/gate-a-plan-pass-6.md | 10 + ...ispositions.2026-07-30-classifier-cycle.md | 56 + ...ositions.2026-08-12-ledger-supersession.md | 28 + ...plan-pass-7.2026-07-30-classifier-cycle.md | 27 + ...n-pass-7.2026-08-12-ledger-supersession.md | 3 + ...plan-pass-8.2026-07-30-classifier-cycle.md | 34 + ...plan-pass-9.2026-07-30-classifier-cycle.md | 38 + .../gate-a-plan-plana-CLOSURE.md | 102 ++ .../gate-a-plan-plana-pass-1-dispositions.md | 97 ++ .../codex-reviews/gate-a-plan-plana-pass-1.md | 22 + .../gate-a-plan-plana-pass-10.md | 7 + .../gate-a-plan-plana-pass-11-dispositions.md | 105 ++ .../gate-a-plan-plana-pass-11.md | 9 + .../gate-a-plan-plana-pass-12.md | 2 + .../gate-a-plan-plana-pass-2-dispositions.md | 91 ++ .../codex-reviews/gate-a-plan-plana-pass-2.md | 17 + .../gate-a-plan-plana-pass-3-dispositions.md | 111 ++ .../codex-reviews/gate-a-plan-plana-pass-3.md | 17 + .../codex-reviews/gate-a-plan-plana-pass-4.md | 8 + .../gate-a-plan-plana-pass-5-dispositions.md | 132 +++ .../codex-reviews/gate-a-plan-plana-pass-5.md | 6 + .../codex-reviews/gate-a-plan-plana-pass-6.md | 3 + .../codex-reviews/gate-a-plan-plana-pass-7.md | 4 + .../codex-reviews/gate-a-plan-plana-pass-8.md | 5 + .../gate-a-plan-plana-pass-9-dispositions.md | 93 ++ .../codex-reviews/gate-a-plan-plana-pass-9.md | 7 + ...te-a-plan-plana-passes-6-7-dispositions.md | 91 ++ .../gate-a-plan-planb-CLOSURE.md | 98 ++ .../codex-reviews/gate-a-plan-planb-pass-1.md | 12 + .../codex-reviews/gate-a-plan-planb-pass-2.md | 4 + .../codex-reviews/gate-a-plan-planb-pass-3.md | 7 + .../codex-reviews/gate-a-plan-planb-pass-4.md | 11 + .../codex-reviews/gate-a-plan-planb-pass-5.md | 10 + .../codex-reviews/gate-a-plan-planb-pass-6.md | 8 + .../codex-reviews/gate-a-plan-planb-pass-7.md | 2 + ...te-a-plan-planb-passes-1-4-dispositions.md | 76 ++ .../gate-a-plan-planc-CLOSURE.md | 86 ++ .../codex-reviews/gate-a-plan-planc-pass-1.md | 19 + .../codex-reviews/gate-a-plan-planc-pass-2.md | 21 + .../codex-reviews/gate-a-plan-planc-pass-3.md | 21 + .../gate-a-plan-planc-pass-4-dispositions.md | 93 ++ .../codex-reviews/gate-a-plan-planc-pass-4.md | 24 + .../codex-reviews/gate-a-plan-planc-pass-5.md | 23 + .../codex-reviews/gate-a-plan-planc-pass-6.md | 20 + .../codex-reviews/gate-a-plan-planc-pass-7.md | 30 + ...te-a-plan-planc-passes-1-2-dispositions.md | 101 ++ .../gate-a-plan-planc1-pass-1-dispositions.md | 66 ++ .../gate-a-plan-planc1-pass-1.md | 15 + .../gate-a-plan-planc1-pass-2.md | 13 + .../gate-a-plan-planc1-pass-3.md | 18 + .../codex-reviews/gate-a-plan-pr1-pass-1.md | 8 + .../codex-reviews/gate-a-plan-pr1-pass-2.md | 8 + .../codex-reviews/gate-a-plan-pr1-pass-3.md | 6 + .../codex-reviews/gate-a-plan-pr1-pass-4.md | 7 + .../codex-reviews/gate-a-plan-pr1-pass-5.md | 7 + .../codex-reviews/gate-a-plan-pr1-pass-6.md | 5 + .../codex-reviews/gate-a-plan-pr1-pass-7.md | 6 + .../codex-reviews/gate-a-plan-pr1-pass-8.md | 6 + .../codex-reviews/gate-a-plan-pr1-pass-9.md | 6 + .../codex-reviews/gate-a-plan-pr2-pass-1.md | 9 + .../codex-reviews/gate-a-plan-pr2-pass-2.md | 9 + .../codex-reviews/gate-a-plan-pr2-pass-3.md | 8 + .../codex-reviews/gate-a-plan-pr2-pass-4.md | 10 + .../codex-reviews/gate-a-plan-pr2-pass-5.md | 6 + .../codex-reviews/gate-a-plan-pr2-pass-6.md | 6 + .../codex-reviews/gate-a-plan-pr2-pass-7.md | 4 + .../codex-reviews/gate-a-plan-pr2-pass-8.md | 2 + .context/codex-reviews/gate-a-plan-resume.md | 57 + .../gate-a-plan-rle-pass-1-dispositions.md | 70 ++ .../codex-reviews/gate-a-plan-rle-pass-1.md | 34 + ...-plan-rle-pass-1.superseded-pre-rewrite.md | 22 + ...-2026-08-02.2026-07-30-classifier-cycle.md | 19 + ...pec-CLOSURE.2026-07-30-classifier-cycle.md | 80 ++ .context/codex-reviews/gate-a-spec-CLOSURE.md | 166 +++ ...e-a-spec-awsf1ec771-pass-1-dispositions.md | 26 + .../gate-a-spec-awsf1ec771-pass-1.md | 25 + ...-a-spec-awsf1ec771-pass-10-dispositions.md | 27 + .../gate-a-spec-awsf1ec771-pass-10.md | 21 + ...-a-spec-awsf1ec771-pass-11-dispositions.md | 45 + .../gate-a-spec-awsf1ec771-pass-11.md | 21 + ...-a-spec-awsf1ec771-pass-12-dispositions.md | 63 + .../gate-a-spec-awsf1ec771-pass-12.md | 22 + ...-a-spec-awsf1ec771-pass-13-dispositions.md | 56 + .../gate-a-spec-awsf1ec771-pass-13.md | 17 + ...e-a-spec-awsf1ec771-pass-2-dispositions.md | 19 + .../gate-a-spec-awsf1ec771-pass-2.md | 18 + ...e-a-spec-awsf1ec771-pass-3-dispositions.md | 12 + .../gate-a-spec-awsf1ec771-pass-3.md | 13 + ...e-a-spec-awsf1ec771-pass-4-dispositions.md | 43 + .../gate-a-spec-awsf1ec771-pass-4.md | 19 + ...e-a-spec-awsf1ec771-pass-5-dispositions.md | 17 + .../gate-a-spec-awsf1ec771-pass-5.md | 18 + ...e-a-spec-awsf1ec771-pass-6-dispositions.md | 16 + .../gate-a-spec-awsf1ec771-pass-6.md | 17 + ...e-a-spec-awsf1ec771-pass-7-dispositions.md | 32 + .../gate-a-spec-awsf1ec771-pass-7.md | 15 + ...e-a-spec-awsf1ec771-pass-8-dispositions.md | 12 + .../gate-a-spec-awsf1ec771-pass-8.md | 13 + ...e-a-spec-awsf1ec771-pass-9-dispositions.md | 16 + .../gate-a-spec-awsf1ec771-pass-9.md | 15 + .../gate-a-spec-awsf1ec771-resume.md | 387 +++++++ .../gate-a-spec-pass-1-dispositions.md | 138 +++ ...-pass-1-dispositions.stopped-2tier-debt.md | 103 ++ ...-spec-pass-1-dispositions.stopped-3tier.md | 113 ++ ...-pass-1-dispositions.stopped-tier3-core.md | 107 ++ .../gate-a-spec-pass-1.stopped-2tier-debt.md | 31 + .../gate-a-spec-pass-1.stopped-3tier.md | 38 + .../gate-a-spec-pass-1.stopped-tier3-core.md | 29 + ...pec-pass-10-dispositions.pre-2026-08-14.md | 33 + .../gate-a-spec-pass-10.pre-2026-08-14.md | 6 + ...pec-pass-11-dispositions.pre-2026-08-14.md | 41 + .../gate-a-spec-pass-11.pre-2026-08-14.md | 13 + ...pec-pass-12-dispositions.pre-2026-08-14.md | 37 + .../gate-a-spec-pass-12.pre-2026-08-14.md | 9 + ...pec-pass-13-dispositions.pre-2026-08-14.md | 54 + .../gate-a-spec-pass-13.pre-2026-08-14.md | 7 + ...pec-pass-14-dispositions.pre-2026-08-14.md | 55 + .../gate-a-spec-pass-14.pre-2026-08-14.md | 4 + ...pec-pass-15-dispositions.pre-2026-08-14.md | 35 + .../gate-a-spec-pass-15.pre-2026-08-14.md | 5 + ...pec-pass-16-dispositions.pre-2026-08-14.md | 46 + .../gate-a-spec-pass-16.pre-2026-08-14.md | 8 + ...pec-pass-17-dispositions.pre-2026-08-14.md | 43 + .../gate-a-spec-pass-17.pre-2026-08-14.md | 7 + ...pec-pass-18-dispositions.pre-2026-08-14.md | 37 + .../gate-a-spec-pass-18.pre-2026-08-14.md | 7 + ...pec-pass-19-dispositions.pre-2026-08-14.md | 52 + .../gate-a-spec-pass-19.pre-2026-08-14.md | 6 + ...spec-pass-2-dispositions.pre-2026-08-14.md | 18 + ...-pass-2-dispositions.stopped-2tier-debt.md | 115 ++ ...-spec-pass-2-dispositions.stopped-3tier.md | 125 ++ ...-pass-2-dispositions.stopped-tier3-core.md | 141 +++ .context/codex-reviews/gate-a-spec-pass-2.md | 26 + .../gate-a-spec-pass-2.pre-2026-08-14.md | 12 + .../gate-a-spec-pass-2.stopped-2tier-debt.md | 35 + .../gate-a-spec-pass-2.stopped-3tier.md | 40 + .../gate-a-spec-pass-2.stopped-tier3-core.md | 30 + ...pec-pass-20-dispositions.pre-2026-08-14.md | 42 + .../gate-a-spec-pass-20.pre-2026-08-14.md | 7 + ...pec-pass-21-dispositions.pre-2026-08-14.md | 55 + .../gate-a-spec-pass-21.pre-2026-08-14.md | 8 + ...pec-pass-22-dispositions.pre-2026-08-14.md | 51 + .../gate-a-spec-pass-22.pre-2026-08-14.md | 6 + ...pec-pass-23-dispositions.pre-2026-08-14.md | 33 + .../gate-a-spec-pass-23.pre-2026-08-14.md | 5 + ...spec-pass-3-dispositions.pre-2026-08-14.md | 15 + ...-pass-3-dispositions.stopped-2tier-debt.md | 96 ++ ...-spec-pass-3-dispositions.stopped-3tier.md | 78 ++ ...-pass-3-dispositions.stopped-tier3-core.md | 102 ++ .context/codex-reviews/gate-a-spec-pass-3.md | 24 + .../gate-a-spec-pass-3.pre-2026-08-14.md | 9 + .../gate-a-spec-pass-3.stopped-2tier-debt.md | 33 + .../gate-a-spec-pass-3.stopped-3tier.md | 41 + .../gate-a-spec-pass-3.stopped-tier3-core.md | 35 + .../gate-a-spec-pass-4-dispositions.md | 97 ++ ...spec-pass-4-dispositions.pre-2026-08-14.md | 28 + .context/codex-reviews/gate-a-spec-pass-4.md | 21 + .../gate-a-spec-pass-4.pre-2026-08-14.md | 12 + ...spec-pass-5-dispositions.pre-2026-08-14.md | 45 + .context/codex-reviews/gate-a-spec-pass-5.md | 17 + .../gate-a-spec-pass-5.pre-2026-08-14.md | 15 + ...spec-pass-6-dispositions.pre-2026-08-14.md | 61 + .context/codex-reviews/gate-a-spec-pass-6.md | 13 + .../gate-a-spec-pass-6.pre-2026-08-14.md | 17 + ...spec-pass-7-dispositions.pre-2026-08-14.md | 56 + .context/codex-reviews/gate-a-spec-pass-7.md | 17 + .../gate-a-spec-pass-7.pre-2026-08-14.md | 16 + .../gate-a-spec-pass-8-dispositions.md | 85 ++ ...spec-pass-8-dispositions.pre-2026-08-14.md | 51 + .context/codex-reviews/gate-a-spec-pass-8.md | 18 + .../gate-a-spec-pass-8.pre-2026-08-14.md | 10 + .../gate-a-spec-pass-9-dispositions.md | 72 ++ ...spec-pass-9-dispositions.pre-2026-08-14.md | 52 + .context/codex-reviews/gate-a-spec-pass-9.md | 10 + .../gate-a-spec-pass-9.pre-2026-08-14.md | 11 + ...spec-rejected-design.stopped-tier3-core.md | 1016 +++++++++++++++++ ...-spec-resume.2026-08-03-hardening-round.md | 75 ++ .context/codex-reviews/gate-a-spec-resume.md | 90 ++ .../gate-a-spec-resume.stopped-2tier-debt.md | 68 ++ .../gate-a-spec-resume.stopped-3tier.md | 45 + .../gate-a-spec-resume.stopped-tier3-core.md | 103 ++ .../codex-reviews/gate-a-spec-rle-pass-1.md | 28 + .../codex-reviews/gate-a-spec-rle-pass-10.md | 28 + .../codex-reviews/gate-a-spec-rle-pass-11.md | 39 + .../codex-reviews/gate-a-spec-rle-pass-12.md | 25 + .../codex-reviews/gate-a-spec-rle-pass-13.md | 33 + .../codex-reviews/gate-a-spec-rle-pass-14.md | 19 + .../codex-reviews/gate-a-spec-rle-pass-15.md | 14 + .../codex-reviews/gate-a-spec-rle-pass-16.md | 5 + .../codex-reviews/gate-a-spec-rle-pass-17.md | 7 + .../codex-reviews/gate-a-spec-rle-pass-18.md | 3 + .../codex-reviews/gate-a-spec-rle-pass-19.md | 8 + .../codex-reviews/gate-a-spec-rle-pass-2.md | 31 + .../codex-reviews/gate-a-spec-rle-pass-20.md | 7 + .../codex-reviews/gate-a-spec-rle-pass-21.md | 10 + .../codex-reviews/gate-a-spec-rle-pass-22.md | 5 + .../codex-reviews/gate-a-spec-rle-pass-23.md | 3 + .../codex-reviews/gate-a-spec-rle-pass-24.md | 8 + .../codex-reviews/gate-a-spec-rle-pass-25.md | 6 + .../codex-reviews/gate-a-spec-rle-pass-26.md | 5 + .../codex-reviews/gate-a-spec-rle-pass-27.md | 7 + .../codex-reviews/gate-a-spec-rle-pass-28.md | 7 + .../codex-reviews/gate-a-spec-rle-pass-29.md | 4 + .../codex-reviews/gate-a-spec-rle-pass-3.md | 55 + .../codex-reviews/gate-a-spec-rle-pass-30.md | 4 + .../codex-reviews/gate-a-spec-rle-pass-31.md | 2 + .../codex-reviews/gate-a-spec-rle-pass-32.md | 3 + .../codex-reviews/gate-a-spec-rle-pass-33.md | 2 + .../codex-reviews/gate-a-spec-rle-pass-34.md | 2 + .../codex-reviews/gate-a-spec-rle-pass-4.md | 41 + .../codex-reviews/gate-a-spec-rle-pass-5.md | 34 + .../codex-reviews/gate-a-spec-rle-pass-6.md | 35 + .../codex-reviews/gate-a-spec-rle-pass-7.md | 33 + .../codex-reviews/gate-a-spec-rle-pass-8.md | 34 + .../codex-reviews/gate-a-spec-rle-pass-9.md | 29 + .../gate-a-spec-vision-pass-1-dispositions.md | 29 + .../gate-a-spec-vision-pass-1.md | 53 + .../gate-a-spec-vision-pass-2-dispositions.md | 41 + .../gate-a-spec-vision-pass-2.md | 35 + .../gate-a-spec-vision-pass-3-dispositions.md | 51 + .../gate-a-spec-vision-pass-3.md | 33 + .../gate-a-spec-vision-pass-4-dispositions.md | 85 ++ .../gate-a-spec-vision-pass-4.md | 27 + .../gate-a-spec-vision-pass-5-dispositions.md | 55 + .../gate-a-spec-vision-pass-5.md | 16 + .../gate-b-0.8.0-dispositions.md | 214 ++++ .../gate-b-0.8.0-pass-1-dispositions.md | 86 ++ .../gate-b-fic2-parked-review-economics.md | 228 ++++ .../gate-b-pass-1-dispositions.md | 15 + .../gate-b-passes-2-4-dispositions.md | 34 + .../gate-b-quality-fic2-pass-1.md | 10 + .../gate-b-quality-fic2-pass-2.md | 12 + .../gate-b-quality-fic2-pass-3.md | 11 + .../gate-b-quality-fic2-pass-4.md | 4 + .../gate-b-quality-fic2-pass-5.md | 5 + .../gate-b-quality-fic2-pass-6.md | 5 + .../gate-b-quality-fic2-pass-7.md | 2 + .../gate-b-quality-harden-pass-1.md | 5 + .../gate-b-quality-harden-pass-2.md | 2 + .../gate-b-quality-harden-pass-3.md | 2 + .../gate-b-quality-harden-pass-4.md | 2 + .../gate-b-quality-pass-1-dispositions.md | 6 + .../gate-b-quality-pass-1.pre-2026-08-16.md | 10 + .../gate-b-quality-pass-10-dispositions.md | 10 + .../codex-reviews/gate-b-quality-pass-10.md | 2 + .../gate-b-quality-pass-11-dispositions.md | 6 + .../codex-reviews/gate-b-quality-pass-11.md | 2 + .../gate-b-quality-pass-12-dispositions.md | 7 + .../codex-reviews/gate-b-quality-pass-12.md | 2 + .../gate-b-quality-pass-13-dispositions.md | 5 + .../codex-reviews/gate-b-quality-pass-13.md | 2 + .../codex-reviews/gate-b-quality-pass-14.md | 2 + .../gate-b-quality-pass-15-dispositions.md | 7 + .../codex-reviews/gate-b-quality-pass-15.md | 2 + .../codex-reviews/gate-b-quality-pass-16.md | 2 + .../gate-b-quality-pass-2-dispositions.md | 6 + ...lity-pass-2.STALE-other-cycle-preserved.md | 4 + .../gate-b-quality-pass-3-dispositions.md | 6 + .../codex-reviews/gate-b-quality-pass-3.md | 2 + .../gate-b-quality-pass-4-dispositions.md | 6 + .../codex-reviews/gate-b-quality-pass-4.md | 4 + .../gate-b-quality-pass-5-dispositions.md | 40 + .../codex-reviews/gate-b-quality-pass-5.md | 5 + .../gate-b-quality-pass-6-dispositions.md | 6 + .../codex-reviews/gate-b-quality-pass-6.md | 2 + .../gate-b-quality-pass-7-dispositions.md | 6 + .../codex-reviews/gate-b-quality-pass-7.md | 5 + .../gate-b-quality-pass-8-dispositions.md | 5 + .../codex-reviews/gate-b-quality-pass-8.md | 5 + .../gate-b-quality-pass-9-dispositions.md | 6 + .../codex-reviews/gate-b-quality-pass-9.md | 2 + .../gate-b-quality-pr15-pass-1.md | 3 + .../gate-b-quality-pr15-pass-2.md | 2 + .../gate-b-quality-pr15-pass-3.md | 6 + .../gate-b-quality-pr15-pass-4.md | 3 + .../gate-b-quality-pr15-pass-5.md | 2 + .../gate-b-quality-pr15-pass-6.md | 3 + .../gate-b-quality-pr15-pass-7.md | 4 + .../gate-b-quality-pr15-pass-8.md | 2 + .../gate-b-quality-pr15-pass-9.md | 2 + .../gate-b-quality-pr15b-pass-1.md | 2 + .../gate-b-quality-rle-pass-1.md | 3 + .../gate-b-quality-rle-pass-2.md | 14 + .../gate-b-quality-rle-pass-3.md | 10 + .../gate-b-quality-rle-pass-4.md | 10 + .../gate-b-quality-rle-pass-5.md | 11 + .../gate-b-rle-pass-5-dispositions.md | 57 + .context/codex-reviews/gate-b-rle-resume.md | 66 ++ .../codex-reviews/gate-b-spec-fic2-pass-1.md | 6 + .../codex-reviews/gate-b-spec-fic2-pass-2.md | 14 + .../codex-reviews/gate-b-spec-fic2-pass-3.md | 3 + .../codex-reviews/gate-b-spec-fic2-pass-4.md | 2 + .../codex-reviews/gate-b-spec-fic2-pass-5.md | 3 + .../codex-reviews/gate-b-spec-fic2-pass-6.md | 3 + .../codex-reviews/gate-b-spec-fic2-pass-7.md | 3 + .../gate-b-spec-harden-pass-1.md | 4 + .../gate-b-spec-harden-pass-2.md | 2 + .../gate-b-spec-harden-pass-3.md | 2 + .../gate-b-spec-harden-pass-4.md | 2 + .../gate-b-spec-pass-1-dispositions.md | 97 ++ .context/codex-reviews/gate-b-spec-pass-1.md | 19 + .../gate-b-spec-pass-1.pre-2026-08-16.md | 8 + .../gate-b-spec-pass-10-dispositions.md | 8 + .context/codex-reviews/gate-b-spec-pass-10.md | 2 + .../gate-b-spec-pass-11-dispositions.md | 7 + .context/codex-reviews/gate-b-spec-pass-11.md | 2 + .context/codex-reviews/gate-b-spec-pass-12.md | 2 + .../gate-b-spec-pass-13-dispositions.md | 7 + .context/codex-reviews/gate-b-spec-pass-13.md | 2 + .../gate-b-spec-pass-14-dispositions.md | 8 + .context/codex-reviews/gate-b-spec-pass-14.md | 2 + .context/codex-reviews/gate-b-spec-pass-15.md | 2 + .context/codex-reviews/gate-b-spec-pass-16.md | 2 + .../gate-b-spec-pass-2-dispositions.md | 60 + .../gate-b-spec-pass-2.RACED-discarded.md | 16 + .context/codex-reviews/gate-b-spec-pass-2.md | 15 + .../gate-b-spec-pass-3-dispositions.md | 62 + .context/codex-reviews/gate-b-spec-pass-3.md | 9 + .../gate-b-spec-pass-4-dispositions.md | 56 + .context/codex-reviews/gate-b-spec-pass-4.md | 8 + .../gate-b-spec-pass-5-dispositions.md | 7 + .context/codex-reviews/gate-b-spec-pass-5.md | 14 + .../gate-b-spec-pass-6-dispositions.md | 76 ++ .context/codex-reviews/gate-b-spec-pass-6.md | 10 + .../gate-b-spec-pass-7-dispositions.md | 49 + .context/codex-reviews/gate-b-spec-pass-7.md | 5 + .../gate-b-spec-pass-8-dispositions.md | 7 + .context/codex-reviews/gate-b-spec-pass-8.md | 5 + .../gate-b-spec-pass-9-dispositions.md | 6 + .context/codex-reviews/gate-b-spec-pass-9.md | 3 + .../codex-reviews/gate-b-spec-pr15-pass-1.md | 3 + .../codex-reviews/gate-b-spec-pr15-pass-2.md | 2 + .../codex-reviews/gate-b-spec-pr15-pass-3.md | 4 + .../codex-reviews/gate-b-spec-pr15-pass-4.md | 2 + .../codex-reviews/gate-b-spec-pr15-pass-5.md | 3 + .../codex-reviews/gate-b-spec-pr15-pass-6.md | 3 + .../codex-reviews/gate-b-spec-pr15-pass-7.md | 2 + .../codex-reviews/gate-b-spec-pr15-pass-8.md | 2 + .../codex-reviews/gate-b-spec-pr15-pass-9.md | 2 + .../codex-reviews/gate-b-spec-pr15b-pass-1.md | 2 + .../codex-reviews/gate-b-spec-rle-pass-1.md | 15 + .../codex-reviews/gate-b-spec-rle-pass-2.md | 17 + .../codex-reviews/gate-b-spec-rle-pass-3.md | 17 + .../codex-reviews/gate-b-spec-rle-pass-4.md | 17 + .../codex-reviews/gate-b-spec-rle-pass-5.md | 14 + .gitignore | 15 + todos.md | 13 +- 390 files changed, 12196 insertions(+), 4 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-CLOSURE.2026-08-12-ledger-supersession.md create mode 100644 .context/codex-reviews/gate-a-plan-CLOSURE.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-1-dispositions.2026-07-26-profiles-cycle.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-1-dispositions.2026-07-30-classifier-cycle.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-1-dispositions.2026-08-04-hardening-round.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-1-dispositions.2026-08-12-ledger-supersession.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-1.2026-07-30-classifier-cycle.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-1.2026-08-04-hardening-round.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-1.2026-08-12-ledger-supersession.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-1.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-2-dispositions.2026-07-26-profiles-cycle.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-2-dispositions.2026-07-30-classifier-cycle.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-2-dispositions.2026-08-04-hardening-round.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-2-dispositions.2026-08-12-ledger-supersession.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-2.2026-07-30-classifier-cycle.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-2.2026-08-04-hardening-round.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-2.2026-08-12-ledger-supersession.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-2.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-3-dispositions.2026-07-30-classifier-cycle.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-3-dispositions.2026-08-12-ledger-supersession.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-3.2026-07-30-classifier-cycle.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-3.2026-08-04-hardening-round.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-3.2026-08-12-ledger-supersession.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-3.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-4-dispositions.2026-07-30-classifier-cycle.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-4-dispositions.2026-08-04-hardening-round.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-4-dispositions.2026-08-12-ledger-supersession.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-4-dispositions.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-4.2026-07-30-classifier-cycle.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-4.2026-08-04-hardening-round.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-4.2026-08-12-ledger-supersession.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-4.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-5-dispositions.2026-07-30-classifier-cycle.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-5-dispositions.2026-08-04-hardening-round.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-5-dispositions.2026-08-12-ledger-supersession.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-5.2026-07-30-classifier-cycle.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-5.2026-08-04-hardening-round.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-5.2026-08-12-ledger-supersession.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-5.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-6-dispositions.2026-07-30-classifier-cycle.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-6-dispositions.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-6.2026-07-30-classifier-cycle.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-6.2026-08-12-ledger-supersession.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-6.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-7-dispositions.2026-07-30-classifier-cycle.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-7-dispositions.2026-08-12-ledger-supersession.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-7.2026-07-30-classifier-cycle.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-7.2026-08-12-ledger-supersession.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-8.2026-07-30-classifier-cycle.md create mode 100644 .context/codex-reviews/gate-a-plan-pass-9.2026-07-30-classifier-cycle.md create mode 100644 .context/codex-reviews/gate-a-plan-plana-CLOSURE.md create mode 100644 .context/codex-reviews/gate-a-plan-plana-pass-1-dispositions.md create mode 100644 .context/codex-reviews/gate-a-plan-plana-pass-1.md create mode 100644 .context/codex-reviews/gate-a-plan-plana-pass-10.md create mode 100644 .context/codex-reviews/gate-a-plan-plana-pass-11-dispositions.md create mode 100644 .context/codex-reviews/gate-a-plan-plana-pass-11.md create mode 100644 .context/codex-reviews/gate-a-plan-plana-pass-12.md create mode 100644 .context/codex-reviews/gate-a-plan-plana-pass-2-dispositions.md create mode 100644 .context/codex-reviews/gate-a-plan-plana-pass-2.md create mode 100644 .context/codex-reviews/gate-a-plan-plana-pass-3-dispositions.md create mode 100644 .context/codex-reviews/gate-a-plan-plana-pass-3.md create mode 100644 .context/codex-reviews/gate-a-plan-plana-pass-4.md create mode 100644 .context/codex-reviews/gate-a-plan-plana-pass-5-dispositions.md create mode 100644 .context/codex-reviews/gate-a-plan-plana-pass-5.md create mode 100644 .context/codex-reviews/gate-a-plan-plana-pass-6.md create mode 100644 .context/codex-reviews/gate-a-plan-plana-pass-7.md create mode 100644 .context/codex-reviews/gate-a-plan-plana-pass-8.md create mode 100644 .context/codex-reviews/gate-a-plan-plana-pass-9-dispositions.md create mode 100644 .context/codex-reviews/gate-a-plan-plana-pass-9.md create mode 100644 .context/codex-reviews/gate-a-plan-plana-passes-6-7-dispositions.md create mode 100644 .context/codex-reviews/gate-a-plan-planb-CLOSURE.md create mode 100644 .context/codex-reviews/gate-a-plan-planb-pass-1.md create mode 100644 .context/codex-reviews/gate-a-plan-planb-pass-2.md create mode 100644 .context/codex-reviews/gate-a-plan-planb-pass-3.md create mode 100644 .context/codex-reviews/gate-a-plan-planb-pass-4.md create mode 100644 .context/codex-reviews/gate-a-plan-planb-pass-5.md create mode 100644 .context/codex-reviews/gate-a-plan-planb-pass-6.md create mode 100644 .context/codex-reviews/gate-a-plan-planb-pass-7.md create mode 100644 .context/codex-reviews/gate-a-plan-planb-passes-1-4-dispositions.md create mode 100644 .context/codex-reviews/gate-a-plan-planc-CLOSURE.md create mode 100644 .context/codex-reviews/gate-a-plan-planc-pass-1.md create mode 100644 .context/codex-reviews/gate-a-plan-planc-pass-2.md create mode 100644 .context/codex-reviews/gate-a-plan-planc-pass-3.md create mode 100644 .context/codex-reviews/gate-a-plan-planc-pass-4-dispositions.md create mode 100644 .context/codex-reviews/gate-a-plan-planc-pass-4.md create mode 100644 .context/codex-reviews/gate-a-plan-planc-pass-5.md create mode 100644 .context/codex-reviews/gate-a-plan-planc-pass-6.md create mode 100644 .context/codex-reviews/gate-a-plan-planc-pass-7.md create mode 100644 .context/codex-reviews/gate-a-plan-planc-passes-1-2-dispositions.md create mode 100644 .context/codex-reviews/gate-a-plan-planc1-pass-1-dispositions.md create mode 100644 .context/codex-reviews/gate-a-plan-planc1-pass-1.md create mode 100644 .context/codex-reviews/gate-a-plan-planc1-pass-2.md create mode 100644 .context/codex-reviews/gate-a-plan-planc1-pass-3.md create mode 100644 .context/codex-reviews/gate-a-plan-pr1-pass-1.md create mode 100644 .context/codex-reviews/gate-a-plan-pr1-pass-2.md create mode 100644 .context/codex-reviews/gate-a-plan-pr1-pass-3.md create mode 100644 .context/codex-reviews/gate-a-plan-pr1-pass-4.md create mode 100644 .context/codex-reviews/gate-a-plan-pr1-pass-5.md create mode 100644 .context/codex-reviews/gate-a-plan-pr1-pass-6.md create mode 100644 .context/codex-reviews/gate-a-plan-pr1-pass-7.md create mode 100644 .context/codex-reviews/gate-a-plan-pr1-pass-8.md create mode 100644 .context/codex-reviews/gate-a-plan-pr1-pass-9.md create mode 100644 .context/codex-reviews/gate-a-plan-pr2-pass-1.md create mode 100644 .context/codex-reviews/gate-a-plan-pr2-pass-2.md create mode 100644 .context/codex-reviews/gate-a-plan-pr2-pass-3.md create mode 100644 .context/codex-reviews/gate-a-plan-pr2-pass-4.md create mode 100644 .context/codex-reviews/gate-a-plan-pr2-pass-5.md create mode 100644 .context/codex-reviews/gate-a-plan-pr2-pass-6.md create mode 100644 .context/codex-reviews/gate-a-plan-pr2-pass-7.md create mode 100644 .context/codex-reviews/gate-a-plan-pr2-pass-8.md create mode 100644 .context/codex-reviews/gate-a-plan-resume.md create mode 100644 .context/codex-reviews/gate-a-plan-rle-pass-1-dispositions.md create mode 100644 .context/codex-reviews/gate-a-plan-rle-pass-1.md create mode 100644 .context/codex-reviews/gate-a-plan-rle-pass-1.superseded-pre-rewrite.md create mode 100644 .context/codex-reviews/gate-a-plan-sweep-2026-08-02.2026-07-30-classifier-cycle.md create mode 100644 .context/codex-reviews/gate-a-spec-CLOSURE.2026-07-30-classifier-cycle.md create mode 100644 .context/codex-reviews/gate-a-spec-CLOSURE.md create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-1-dispositions.md create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-1.md create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-10-dispositions.md create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-10.md create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-11-dispositions.md create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-11.md create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-12-dispositions.md create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-12.md create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-13-dispositions.md create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-13.md create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-2-dispositions.md create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-2.md create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-3-dispositions.md create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-3.md create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-4-dispositions.md create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-4.md create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-5-dispositions.md create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-5.md create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-6-dispositions.md create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-6.md create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-7-dispositions.md create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-7.md create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-8-dispositions.md create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-8.md create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-9-dispositions.md create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-9.md create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-resume.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-1-dispositions.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-1-dispositions.stopped-2tier-debt.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-1-dispositions.stopped-3tier.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-1-dispositions.stopped-tier3-core.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-1.stopped-2tier-debt.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-1.stopped-3tier.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-1.stopped-tier3-core.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-10-dispositions.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-10.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-11-dispositions.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-11.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-12-dispositions.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-12.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-13-dispositions.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-13.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-14-dispositions.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-14.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-15-dispositions.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-15.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-16-dispositions.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-16.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-17-dispositions.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-17.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-18-dispositions.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-18.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-19-dispositions.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-19.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-2-dispositions.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-2-dispositions.stopped-2tier-debt.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-2-dispositions.stopped-3tier.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-2-dispositions.stopped-tier3-core.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-2.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-2.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-2.stopped-2tier-debt.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-2.stopped-3tier.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-2.stopped-tier3-core.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-20-dispositions.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-20.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-21-dispositions.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-21.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-22-dispositions.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-22.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-23-dispositions.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-23.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-3-dispositions.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-3-dispositions.stopped-2tier-debt.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-3-dispositions.stopped-3tier.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-3-dispositions.stopped-tier3-core.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-3.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-3.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-3.stopped-2tier-debt.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-3.stopped-3tier.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-3.stopped-tier3-core.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-4-dispositions.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-4-dispositions.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-4.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-4.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-5-dispositions.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-5.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-5.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-6-dispositions.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-6.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-6.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-7-dispositions.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-7.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-7.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-8-dispositions.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-8-dispositions.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-8.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-8.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-9-dispositions.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-9-dispositions.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-9.md create mode 100644 .context/codex-reviews/gate-a-spec-pass-9.pre-2026-08-14.md create mode 100644 .context/codex-reviews/gate-a-spec-rejected-design.stopped-tier3-core.md create mode 100644 .context/codex-reviews/gate-a-spec-resume.2026-08-03-hardening-round.md create mode 100644 .context/codex-reviews/gate-a-spec-resume.md create mode 100644 .context/codex-reviews/gate-a-spec-resume.stopped-2tier-debt.md create mode 100644 .context/codex-reviews/gate-a-spec-resume.stopped-3tier.md create mode 100644 .context/codex-reviews/gate-a-spec-resume.stopped-tier3-core.md create mode 100644 .context/codex-reviews/gate-a-spec-rle-pass-1.md create mode 100644 .context/codex-reviews/gate-a-spec-rle-pass-10.md create mode 100644 .context/codex-reviews/gate-a-spec-rle-pass-11.md create mode 100644 .context/codex-reviews/gate-a-spec-rle-pass-12.md create mode 100644 .context/codex-reviews/gate-a-spec-rle-pass-13.md create mode 100644 .context/codex-reviews/gate-a-spec-rle-pass-14.md create mode 100644 .context/codex-reviews/gate-a-spec-rle-pass-15.md create mode 100644 .context/codex-reviews/gate-a-spec-rle-pass-16.md create mode 100644 .context/codex-reviews/gate-a-spec-rle-pass-17.md create mode 100644 .context/codex-reviews/gate-a-spec-rle-pass-18.md create mode 100644 .context/codex-reviews/gate-a-spec-rle-pass-19.md create mode 100644 .context/codex-reviews/gate-a-spec-rle-pass-2.md create mode 100644 .context/codex-reviews/gate-a-spec-rle-pass-20.md create mode 100644 .context/codex-reviews/gate-a-spec-rle-pass-21.md create mode 100644 .context/codex-reviews/gate-a-spec-rle-pass-22.md create mode 100644 .context/codex-reviews/gate-a-spec-rle-pass-23.md create mode 100644 .context/codex-reviews/gate-a-spec-rle-pass-24.md create mode 100644 .context/codex-reviews/gate-a-spec-rle-pass-25.md create mode 100644 .context/codex-reviews/gate-a-spec-rle-pass-26.md create mode 100644 .context/codex-reviews/gate-a-spec-rle-pass-27.md create mode 100644 .context/codex-reviews/gate-a-spec-rle-pass-28.md create mode 100644 .context/codex-reviews/gate-a-spec-rle-pass-29.md create mode 100644 .context/codex-reviews/gate-a-spec-rle-pass-3.md create mode 100644 .context/codex-reviews/gate-a-spec-rle-pass-30.md create mode 100644 .context/codex-reviews/gate-a-spec-rle-pass-31.md create mode 100644 .context/codex-reviews/gate-a-spec-rle-pass-32.md create mode 100644 .context/codex-reviews/gate-a-spec-rle-pass-33.md create mode 100644 .context/codex-reviews/gate-a-spec-rle-pass-34.md create mode 100644 .context/codex-reviews/gate-a-spec-rle-pass-4.md create mode 100644 .context/codex-reviews/gate-a-spec-rle-pass-5.md create mode 100644 .context/codex-reviews/gate-a-spec-rle-pass-6.md create mode 100644 .context/codex-reviews/gate-a-spec-rle-pass-7.md create mode 100644 .context/codex-reviews/gate-a-spec-rle-pass-8.md create mode 100644 .context/codex-reviews/gate-a-spec-rle-pass-9.md create mode 100644 .context/codex-reviews/gate-a-spec-vision-pass-1-dispositions.md create mode 100644 .context/codex-reviews/gate-a-spec-vision-pass-1.md create mode 100644 .context/codex-reviews/gate-a-spec-vision-pass-2-dispositions.md create mode 100644 .context/codex-reviews/gate-a-spec-vision-pass-2.md create mode 100644 .context/codex-reviews/gate-a-spec-vision-pass-3-dispositions.md create mode 100644 .context/codex-reviews/gate-a-spec-vision-pass-3.md create mode 100644 .context/codex-reviews/gate-a-spec-vision-pass-4-dispositions.md create mode 100644 .context/codex-reviews/gate-a-spec-vision-pass-4.md create mode 100644 .context/codex-reviews/gate-a-spec-vision-pass-5-dispositions.md create mode 100644 .context/codex-reviews/gate-a-spec-vision-pass-5.md create mode 100644 .context/codex-reviews/gate-b-0.8.0-dispositions.md create mode 100644 .context/codex-reviews/gate-b-0.8.0-pass-1-dispositions.md create mode 100644 .context/codex-reviews/gate-b-fic2-parked-review-economics.md create mode 100644 .context/codex-reviews/gate-b-pass-1-dispositions.md create mode 100644 .context/codex-reviews/gate-b-passes-2-4-dispositions.md create mode 100644 .context/codex-reviews/gate-b-quality-fic2-pass-1.md create mode 100644 .context/codex-reviews/gate-b-quality-fic2-pass-2.md create mode 100644 .context/codex-reviews/gate-b-quality-fic2-pass-3.md create mode 100644 .context/codex-reviews/gate-b-quality-fic2-pass-4.md create mode 100644 .context/codex-reviews/gate-b-quality-fic2-pass-5.md create mode 100644 .context/codex-reviews/gate-b-quality-fic2-pass-6.md create mode 100644 .context/codex-reviews/gate-b-quality-fic2-pass-7.md create mode 100644 .context/codex-reviews/gate-b-quality-harden-pass-1.md create mode 100644 .context/codex-reviews/gate-b-quality-harden-pass-2.md create mode 100644 .context/codex-reviews/gate-b-quality-harden-pass-3.md create mode 100644 .context/codex-reviews/gate-b-quality-harden-pass-4.md create mode 100644 .context/codex-reviews/gate-b-quality-pass-1-dispositions.md create mode 100644 .context/codex-reviews/gate-b-quality-pass-1.pre-2026-08-16.md create mode 100644 .context/codex-reviews/gate-b-quality-pass-10-dispositions.md create mode 100644 .context/codex-reviews/gate-b-quality-pass-10.md create mode 100644 .context/codex-reviews/gate-b-quality-pass-11-dispositions.md create mode 100644 .context/codex-reviews/gate-b-quality-pass-11.md create mode 100644 .context/codex-reviews/gate-b-quality-pass-12-dispositions.md create mode 100644 .context/codex-reviews/gate-b-quality-pass-12.md create mode 100644 .context/codex-reviews/gate-b-quality-pass-13-dispositions.md create mode 100644 .context/codex-reviews/gate-b-quality-pass-13.md create mode 100644 .context/codex-reviews/gate-b-quality-pass-14.md create mode 100644 .context/codex-reviews/gate-b-quality-pass-15-dispositions.md create mode 100644 .context/codex-reviews/gate-b-quality-pass-15.md create mode 100644 .context/codex-reviews/gate-b-quality-pass-16.md create mode 100644 .context/codex-reviews/gate-b-quality-pass-2-dispositions.md create mode 100644 .context/codex-reviews/gate-b-quality-pass-2.STALE-other-cycle-preserved.md create mode 100644 .context/codex-reviews/gate-b-quality-pass-3-dispositions.md create mode 100644 .context/codex-reviews/gate-b-quality-pass-3.md create mode 100644 .context/codex-reviews/gate-b-quality-pass-4-dispositions.md create mode 100644 .context/codex-reviews/gate-b-quality-pass-4.md create mode 100644 .context/codex-reviews/gate-b-quality-pass-5-dispositions.md create mode 100644 .context/codex-reviews/gate-b-quality-pass-5.md create mode 100644 .context/codex-reviews/gate-b-quality-pass-6-dispositions.md create mode 100644 .context/codex-reviews/gate-b-quality-pass-6.md create mode 100644 .context/codex-reviews/gate-b-quality-pass-7-dispositions.md create mode 100644 .context/codex-reviews/gate-b-quality-pass-7.md create mode 100644 .context/codex-reviews/gate-b-quality-pass-8-dispositions.md create mode 100644 .context/codex-reviews/gate-b-quality-pass-8.md create mode 100644 .context/codex-reviews/gate-b-quality-pass-9-dispositions.md create mode 100644 .context/codex-reviews/gate-b-quality-pass-9.md create mode 100644 .context/codex-reviews/gate-b-quality-pr15-pass-1.md create mode 100644 .context/codex-reviews/gate-b-quality-pr15-pass-2.md create mode 100644 .context/codex-reviews/gate-b-quality-pr15-pass-3.md create mode 100644 .context/codex-reviews/gate-b-quality-pr15-pass-4.md create mode 100644 .context/codex-reviews/gate-b-quality-pr15-pass-5.md create mode 100644 .context/codex-reviews/gate-b-quality-pr15-pass-6.md create mode 100644 .context/codex-reviews/gate-b-quality-pr15-pass-7.md create mode 100644 .context/codex-reviews/gate-b-quality-pr15-pass-8.md create mode 100644 .context/codex-reviews/gate-b-quality-pr15-pass-9.md create mode 100644 .context/codex-reviews/gate-b-quality-pr15b-pass-1.md create mode 100644 .context/codex-reviews/gate-b-quality-rle-pass-1.md create mode 100644 .context/codex-reviews/gate-b-quality-rle-pass-2.md create mode 100644 .context/codex-reviews/gate-b-quality-rle-pass-3.md create mode 100644 .context/codex-reviews/gate-b-quality-rle-pass-4.md create mode 100644 .context/codex-reviews/gate-b-quality-rle-pass-5.md create mode 100644 .context/codex-reviews/gate-b-rle-pass-5-dispositions.md create mode 100644 .context/codex-reviews/gate-b-rle-resume.md create mode 100644 .context/codex-reviews/gate-b-spec-fic2-pass-1.md create mode 100644 .context/codex-reviews/gate-b-spec-fic2-pass-2.md create mode 100644 .context/codex-reviews/gate-b-spec-fic2-pass-3.md create mode 100644 .context/codex-reviews/gate-b-spec-fic2-pass-4.md create mode 100644 .context/codex-reviews/gate-b-spec-fic2-pass-5.md create mode 100644 .context/codex-reviews/gate-b-spec-fic2-pass-6.md create mode 100644 .context/codex-reviews/gate-b-spec-fic2-pass-7.md create mode 100644 .context/codex-reviews/gate-b-spec-harden-pass-1.md create mode 100644 .context/codex-reviews/gate-b-spec-harden-pass-2.md create mode 100644 .context/codex-reviews/gate-b-spec-harden-pass-3.md create mode 100644 .context/codex-reviews/gate-b-spec-harden-pass-4.md create mode 100644 .context/codex-reviews/gate-b-spec-pass-1-dispositions.md create mode 100644 .context/codex-reviews/gate-b-spec-pass-1.md create mode 100644 .context/codex-reviews/gate-b-spec-pass-1.pre-2026-08-16.md create mode 100644 .context/codex-reviews/gate-b-spec-pass-10-dispositions.md create mode 100644 .context/codex-reviews/gate-b-spec-pass-10.md create mode 100644 .context/codex-reviews/gate-b-spec-pass-11-dispositions.md create mode 100644 .context/codex-reviews/gate-b-spec-pass-11.md create mode 100644 .context/codex-reviews/gate-b-spec-pass-12.md create mode 100644 .context/codex-reviews/gate-b-spec-pass-13-dispositions.md create mode 100644 .context/codex-reviews/gate-b-spec-pass-13.md create mode 100644 .context/codex-reviews/gate-b-spec-pass-14-dispositions.md create mode 100644 .context/codex-reviews/gate-b-spec-pass-14.md create mode 100644 .context/codex-reviews/gate-b-spec-pass-15.md create mode 100644 .context/codex-reviews/gate-b-spec-pass-16.md create mode 100644 .context/codex-reviews/gate-b-spec-pass-2-dispositions.md create mode 100644 .context/codex-reviews/gate-b-spec-pass-2.RACED-discarded.md create mode 100644 .context/codex-reviews/gate-b-spec-pass-2.md create mode 100644 .context/codex-reviews/gate-b-spec-pass-3-dispositions.md create mode 100644 .context/codex-reviews/gate-b-spec-pass-3.md create mode 100644 .context/codex-reviews/gate-b-spec-pass-4-dispositions.md create mode 100644 .context/codex-reviews/gate-b-spec-pass-4.md create mode 100644 .context/codex-reviews/gate-b-spec-pass-5-dispositions.md create mode 100644 .context/codex-reviews/gate-b-spec-pass-5.md create mode 100644 .context/codex-reviews/gate-b-spec-pass-6-dispositions.md create mode 100644 .context/codex-reviews/gate-b-spec-pass-6.md create mode 100644 .context/codex-reviews/gate-b-spec-pass-7-dispositions.md create mode 100644 .context/codex-reviews/gate-b-spec-pass-7.md create mode 100644 .context/codex-reviews/gate-b-spec-pass-8-dispositions.md create mode 100644 .context/codex-reviews/gate-b-spec-pass-8.md create mode 100644 .context/codex-reviews/gate-b-spec-pass-9-dispositions.md create mode 100644 .context/codex-reviews/gate-b-spec-pass-9.md create mode 100644 .context/codex-reviews/gate-b-spec-pr15-pass-1.md create mode 100644 .context/codex-reviews/gate-b-spec-pr15-pass-2.md create mode 100644 .context/codex-reviews/gate-b-spec-pr15-pass-3.md create mode 100644 .context/codex-reviews/gate-b-spec-pr15-pass-4.md create mode 100644 .context/codex-reviews/gate-b-spec-pr15-pass-5.md create mode 100644 .context/codex-reviews/gate-b-spec-pr15-pass-6.md create mode 100644 .context/codex-reviews/gate-b-spec-pr15-pass-7.md create mode 100644 .context/codex-reviews/gate-b-spec-pr15-pass-8.md create mode 100644 .context/codex-reviews/gate-b-spec-pr15-pass-9.md create mode 100644 .context/codex-reviews/gate-b-spec-pr15b-pass-1.md create mode 100644 .context/codex-reviews/gate-b-spec-rle-pass-1.md create mode 100644 .context/codex-reviews/gate-b-spec-rle-pass-2.md create mode 100644 .context/codex-reviews/gate-b-spec-rle-pass-3.md create mode 100644 .context/codex-reviews/gate-b-spec-rle-pass-4.md create mode 100644 .context/codex-reviews/gate-b-spec-rle-pass-5.md diff --git a/.context/codex-reviews/gate-a-plan-CLOSURE.2026-08-12-ledger-supersession.md b/.context/codex-reviews/gate-a-plan-CLOSURE.2026-08-12-ledger-supersession.md new file mode 100644 index 0000000..869438a --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-CLOSURE.2026-08-12-ledger-supersession.md @@ -0,0 +1,108 @@ +# Gate A — plan — CLOSURE RECORD + +**Cycle:** a supersession convention for the hardening ledger. +**Plan:** `docs/superpowers/plans/2026-08-12-hardening-ledger-supersession.md` +**Spec:** `docs/superpowers/specs/2026-08-05-hardening-ledger-supersession-design.md` (its own Gate A +closed at pass 23 — `gate-a-spec-CLOSURE.md`). +**Branch:** `ledger-supersession`. **Closed 2026-08-12 on plan pass 7.** + +**Gate A (plan) is CLOSED.** `superpowers:executing-plans` is open — sequential, mirror edits in the +same commit. Passes 1–7 applied in full; nothing is held. + +Reviewer for all seven passes: `mcp-codex-dev@1.0.1`, `codex-cli 0.144.5`, model `gpt-5.6-sol`. +Every pass satisfied the file-first protocol: terminator exact, count matching, no non-finding body +lines. No pass was INCOMPLETE; the recovery budget was never spent. + +## Why it closed here + +Pass 7 returned **two findings, both Minor, zero Blocker and zero Major**. §5's loop is a +Blocker/Major loop — Minor and Nit are collected, never iterated — so a pass with none is its exit +condition, and both Minors were applied rather than collected because each was a one-line +cross-reference fix. + +## Trend + +| Pass | 1 | 2 | 3 | 4 | 5 | 6 | 7 | +|---|---|---|---|---|---|---|---| +| Findings | 18 | 11 | 9 | 8 | 6 | 2 | 2 | +| Blocker | 2 | 1 | 4 | 3 | 1 | – | – | +| Major | 12 | 8 | 2 | 4 | 3 | 2 | – | +| Minor | 4 | 2 | 3 | 1 | 2 | – | 2 | + +Monotone in the count from pass 2. The Blocker spike at pass 3 was not a regression: those four had +been latent since pass 1 and surfaced only once the surrounding defects cleared. Both severity +columns reach zero together. + +## Two cuts applied before pass 1, by the sweep + +- **428 lines of shell** across 31 fences, against spec §6's own "the executable form is written at + execution time, carried in the plan with one label per check" → cut to labels, properties and + oracles. Zero lines of shell remain. +- **Check 1's candidates were never bounded** to the label-to-`Columns:` interval, contradicting + §6's own 1d oracle — which is also §8 held item 4's trace. + +## What the passes actually caught + +**Two Blockers about the Gate-B mechanics** (pass 1): the consolidation reset to *the parent of the +first WIP* and then reviewed an empty range, with the closing `--amend` set to rewrite `BASE` +itself; and a Gate-B fix was to be "re-reviewed" without being amended into the WIP commit, so +`mcp__codex__review` would have re-read the **stale** committed range and reported a clean pass on +code it never saw. + +**Two Blockers where the validation violated the convention it validates** (pass 3): `C1b`'s +counter-check mutated the real 2026-07-20 row and `C3`'s moved the real block and added and removed +a complete entry — operations §2.1 and §2.2 forbid the moment the text exists, committed or not, in +the very file the convention is being added to. Every such mutation now runs on a scratch copy, with +the boundary stated: *prose* mutations are covered by neither rule and stay in place. + +**Two staging Blockers** (passes 2 and 4): `git add -A` would have swept the untracked +`docs/research/` tree and the executor's fixtures into a commit or an amend; and `git reset --soft` +cannot be repaired by adding paths, because an explicit `add` does not *unstage* an extra one — so +the equality assertion would have halted execution with no route forward. Now `--mixed`, five named +paths, and an asserted cached set. + +**One Blocker per pass, 3 through 5, of one shape: a fix that landed in one site and not in its +mirror.** §8's held-not-fixed intro still said Gate B reviews the checks after §6 had been corrected +to say the opposite; the plan promised the reviewer a paste of the check source that +`additionalContext` did not carry; `BASE` contained the plan while §8 claimed the plan was inside the +reviewed range. Naming that failure mode in the pass-5 and pass-6 prompts is what drove it to zero. + +## Settled decisions + +| Decision | Settled at | +|---|---| +| **The plan carries labels, properties and oracles; no executable form** | pre-pass sweep | +| **Eight labels** — `C1a`–`C1d`, `C1f`, `C2a`, `C2b`, `C3`; **spec 1e is not implemented**, recorded in §8 | pass 1 | +| **Gate B compares the implementation range only** — neither the plan nor the checks are inside it; both reach the reviewer as context | passes 3–5 | +| **Spec §6 governs, except on the four §8 held-not-fixed items**, where the plan's remedies are authoritative | pass 1 | +| **No counter-check may mutate a row or a complete entry in the real ledger**; prose mutations may stay in place | pass 3 | +| **Never a pathspec broader than the step's file list**, with the staged set asserted before every commit and amend | passes 2, 4 | +| **The evidence entry is one artifact**, composed before the first Gate-B call, carrying the story path inside it | passes 4, 6 | +| **One baseline, one advertised mutation per fixture**; seventeen rows with expected outcomes, PASS rows as well as FAIL | passes 2, 6 | +| **Five labels fail on the untouched tree** — `C1a`, `C1d`, `C2a`, `C2b`, `C3`; `C1b` and `C1c` are protective | pass 1 | + +## The four §8 held-not-fixed items, traced + +1. **row-date equality** → `C1d`.4 plus the `entry-wrong-rowdate` fixture — executable. +2. **the pre-narrowing "prose-only" claim** → the Global Constraint stating §2.2 governs — wording, + which is all this one admits. +3. **plural scope** → `C1d`.1 plus `entry-good`'s second inert entry — executable, as of pass 2. +4. **the format-example guard** → `C3` plus Task 5's below-the-table counter-check — executable. Not + `C1d`.2, which *ignores* a line outside the interval rather than detecting one; that + mis-attribution was itself a pass-4 finding. + +## Spec edits made during the plan cycle + +Five, each from a plan-pass finding: §6's "reviewed by Gate B against the real diff" corrected to +*supplied to the reviewer alongside* it; §8's held-not-fixed intro brought into line; a new §8 +residual recording that the checks are never fingerprinted and that `battery+check` evidence is +produced by unfingerprinted code; the §8 record that check 1e is not implemented; and the §6 anchor +oracle's false "two begin with `-`" corrected to one. + +## Where everything lives + +- **Uncommitted at closure:** the spec, the story and this plan. Task 1 of the plan commits all + three and pins `BASE`. +- **On disk only, gitignored:** `.context/codex-reviews/` — seven plan pass files, their + dispositions, this record, and the spec cycle's twenty-three. `.gitignore:13` ignores `.context/*`; + a push does not back these up. diff --git a/.context/codex-reviews/gate-a-plan-CLOSURE.md b/.context/codex-reviews/gate-a-plan-CLOSURE.md new file mode 100644 index 0000000..ec6fdae --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-CLOSURE.md @@ -0,0 +1,69 @@ +# Gate A — plan — **CLOSED on dispositions, with four post-pass fixes** + +Closed 2026-08-16 by Daniel's decision, after pass 6. + +| Pass | 1 | 2 | 3 | 4 | 5 | 6 | +|---|---|---|---|---|---|---| +| Findings | 18 | 20 | 17 | 15 | 9 | 9 | +| Blockers | 1 | 0 | 1 | 1 | 1 | 1 | + +**88 findings, none dismissed.** Six passes; the floor was met at pass 3 and the loop continued +because no pass was clean. + +## Why this closes on dispositions rather than a clean pass + +**The checker's verification channel is execution, not review.** Every defect check 4c ever +had was found by running it, and none by reading it: + +| Defect | Found by | +|---|---| +| Fence-mode parser returned `MULTIEND` on the real file — the check could never pass | running the `awk` against `workflow-init.md` | +| `checklist parser failure fires` had stopped testing 4b (its grep matched 4c's diagnostic too) | the mutation run — 21 flips against a recorded 22 | +| Whole-file counting silently dropped the placement guarantee | planting the line in the command file's prose and watching the battery stay green | +| Placement failed open when its terminator moved | renaming `### 2.2` and planting a line in the widened gap | + +Four review passes read the first as prose without finding it. The last two were found in +minutes by executing the thing. + +**So the remaining assurance comes from the tools and from Gate B, not from a seventh pass.** +The scripts are in the reviewed range: `baseSha` = `c0a6ed2`, and the WIP snapshot carries +`scripts/check-invariants.{sh,test.sh}` along with the spec and this plan. Gate B reads them +**as code**, which is the review they should have had all along. + +## The four post-pass fixes + +1. **Terminator validation.** The placement range's end must be `### 2.2`; anything else fails + loudly, naming what it found. Three fixtures, including the verified exploit — rename the + terminator, plant the line in the widened gap. Confirmed rejected. +2. **Mutation re-measured.** 4c flips **19** (18 rejects + the parser case), no accept case + moved. It read 13 before the placement fixtures and 16 before the terminator ones; each + superseded number was replaced by re-running, never extrapolated. The block now says so. +3. **Task 0 deleted.** It committed the spec, plan and scripts immediately before the WIP — + making that commit the WIP *parent*, which `baseSha..HEAD` excludes, so Gate B would have + excluded exactly what it existed to include. It was also a non-WIP commit of executable + code with a red battery and no Gate-B loop. The one-line fix it should have been: the WIP + snapshot carries those four files, `baseSha` = `c0a6ed2`. +4. **Self-review and inventory rewritten from the code as built** — `severity_rule_scan ` + returning four integers, not the withdrawn `severity_rule_count(file, mode)` sentinel API; + 24 `sev_case` cases, not 15. And Task 4 Step 4 now follows §6's settled **TRIGGER FIRED** + rather than reopening it. + +## State at closure + +- **148 assertions** green under `sh` and `dash` (123 before), shellcheck clean on both scripts. +- **The counterfactual is observed, not asserted:** with 4c present and the canonical line not + yet in either prompt copy, `sh scripts/check-invariants.sh` exits 1 naming only 4c on both + files. The battery is **red by design** until Task 1's prose edits land. +- Uncommitted: the twice-amended spec, this plan, both scripts. `c0a6ed2` is the only commit. + +## What Gate A did not settle + +The plan's prose has been the slow part, not the design: passes 5 and 6 were dominated by +descriptions going stale as the code changed under them. Gate B reviews the code and the +shipped prompt text; it does **not** re-review this plan. A reader following the plan should +treat the built scripts as authoritative where the two disagree. + +## Next + +Execution: test-first, battery green before the Gate-B loop, §5 cited rather than restated, +evidence per `battery+check` with the observed counterfactual. diff --git a/.context/codex-reviews/gate-a-plan-pass-1-dispositions.2026-07-26-profiles-cycle.md b/.context/codex-reviews/gate-a-plan-pass-1-dispositions.2026-07-26-profiles-cycle.md new file mode 100644 index 0000000..772a2b6 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pass-1-dispositions.2026-07-26-profiles-cycle.md @@ -0,0 +1,16 @@ +# Gate A — plan — pass 1 dispositions + +12 findings, 3 Blockers. All accepted. + +1 BLOCKER closing commit body echoes risk/security/mode — ACCEPT. My own Global Constraint, violated in the plan's own example. The entry carries the story PATH and the named evidence, nothing else. +2 BLOCKER new Flow step 6 asks inside step 3's round, which has already closed — ACCEPT. Unimplementable as written. The profile assessment moves BEFORE the question round: read -> detect ambiguity -> assess profile -> one round carrying clarifications AND the proposal -> pause. Whole flow renumbered to 11 steps. +3 BLOCKER battery runs before the version bump is committed, and `check-version-bump.sh main` compares COMMITS — ACCEPT. It would fail on a tree whose plugin edits are committed with no bump. The WIP snapshot now includes the manifest and CHANGELOG, and the battery runs at that snapshot's SHA. +4 MAJOR Gate-B call omits the evidence entry it is required to quote — ACCEPT. The entry is written into the WIP body before the first pass, and every pass's context quotes that exact entry. +5 MAJOR no handling for a header whose Validation contradicts the axes — ACCEPT, and it is genuinely new: the spec's case 3 covers a value outside the enums, not an in-enum value inconsistent with `max(risk, security)` or a `+abuse-path` suffix disagreeing with security. Added to the plan's §5 text AND to the spec in the same change, so the two do not disagree. +6 MAJOR profile-log block has no shown placement — ACCEPT. A conditional example goes inside the story-template code block directly beneath the header, marked absent until the first event (prompt standard 4: show the output format). +7 MAJOR mode-override events have no direction in the log grammar — ACCEPT. Direction applies to mode overrides too; one example line per event kind now ships in the template guidance. +8 MAJOR the agreement diff covers only the Profiles subsection, not the Mechanics sentence — ACCEPT, and it is the sharper half: the paired edit was half-verified while claiming to be checked. A second diff covers the Finishing-the-cycle bullet. +9 MAJOR battery command drops `claude plugin validate . --strict` — ACCEPT. AGENTS.md's quality row includes it and the plan claimed plugin validation ran. "Never document a command that wasn't run." +10 MAJOR evidence says "run at " for a battery run on an uncommitted tree — ACCEPT. Same fix as 3: snapshot first, run at that SHA, re-run after every fix, quote the SHA actually reviewed. +11 MAJOR Task 4's replacement text starts mid-sentence, so literal execution corrupts the paragraph — ACCEPT. Anchor widened to the full sentence beginning "The `intake` skill". +12 MINOR audit verdicts live only in WIP messages that the soft reset discards — ACCEPT. They move into the final commit body (no profile values), where a reviewer can still tell "read and accurate" from "never checked". diff --git a/.context/codex-reviews/gate-a-plan-pass-1-dispositions.2026-07-30-classifier-cycle.md b/.context/codex-reviews/gate-a-plan-pass-1-dispositions.2026-07-30-classifier-cycle.md new file mode 100644 index 0000000..fb50043 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pass-1-dispositions.2026-07-30-classifier-cycle.md @@ -0,0 +1,54 @@ +# Gate A (plan) — pass 1 dispositions + +Advisory companion to `gate-a-plan-pass-1.md` (7 BLOCKER, 29 MAJOR, 4 MINOR = 40). One line +per finding, in file order. **Six findings dissolved rather than being fixed**: the human +confirmed collapsing the spec's two-locator design to one, so the machinery that existed to +reconcile the two paths is deleted (spec §3.1 amended inline, 2026-08-01). + +1. BLOCKER `jq -r … | head -n1` — DISSOLVED. The `jq` locator is deleted. +2. MAJOR span check unimplemented — DISSOLVED. `jq` reports no byte offsets, which is what made the check unbuildable and the second path pointless. +3. MAJOR duplicate-key divergence — DISSOLVED. One locator; a repeated depth-1 key exits 2 → `unrecognized`, tested. +4. MAJOR `jq` status masked / newline stripping — DISSOLVED with the path. The awk locator's status is read directly; command substitution's newline stripping is harmless because a JSON string cannot hold a raw newline (stated in Task 5 Step 3). +5. BLOCKER `locate_scan` mis-parses the real fixture — FIXED. Replaced by `locate.awk`, green 35/35 against the real captures under `sh` and `dash`; `locate_result` is now defined, not prose. +6. MAJOR malformed-JSON contract — PARTLY ACCEPTED. The scan validates exactly what it must walk to reach the block; trailing garbage and post-value balance are **not** checked, and A2 says so rather than claiming full validation. Routable-but-malformed → `unrecognized` (terminal default), which spec §3.3's unroutable-payload rule never covered. +7. MAJOR `RS="\0"` portability — FIXED. Records are accumulated and processed in `END`, with the separator restored; verified under two shells. +8. BLOCKER `case` glob admits a reordered envelope — FIXED. Brace stripped, whitespace consumed by a bounded loop, then a fixed prefix; `{"status":…,"success":false}` → `unrecognized`, tested both collision directions. +9. MAJOR `[[:space:]]*` is not zero-or-more — FIXED. Three explicit grammar points plus a token-end check (`truely` → `unrecognized`); compact, tab, CRLF and space-before-colon all tested. +10. BLOCKER anchor expects a literal quote — FIXED. Anchor is `MCP tool \"` in escaped form; the real `shape3` capture classifies `backgrounded`. +11. MAJOR A4 constraints unvalidated — PARTLY DISMISSED. Duration format, unit and task-id are **deliberately** outside the anchor (spec §4): narrowing to shapes observed once would widen C1, not close it. Implemented and tested: start anchor, segment before any `\n`, and all four named near-misses. +12. MAJOR unicode test supplies an ASCII space — FIXED. Literal ` ` in the table; the canonical-form ordering rule it tested is deleted with the second path (spec §3.3 amended). +13. MAJOR §7.3 matrix gaps — FIXED. The 35-row table from `verify.sh` is ported across Tasks 5–7, including every `no-result` shape, both collisions and the polarity encodings. +14. MAJOR the driver itself needs `jq` — FIXED. `payload`/`resp` are `printf`-only and `payload_from` retargets a fixture with `sed`; no classification test depends on `jq` existing. +15. MAJOR undefined helpers and state variables — FIXED. Task 2 defines all of them before first use. `HOOK_LOCATE`/`HOOK_CLASSIFY` are **deleted**: the five classes are externally distinguishable, so no debug entry point is added to the product. +16. MAJOR writer-failure test asserts nothing — FIXED. `run_closed` runs the hook with stdout closed and captures its status separately; the marker consequences are asserted where markers exist (Tasks 7–8). +17. MAJOR `note_discarded` used before defined — FIXED. Old Tasks 6 and 7 merged, so every committed state is functional. +18. MAJOR preservation coverage incomplete — FIXED. `seeded_preserves` runs five discarded shapes (incl. two `no-result` variants) × both gates × default and mapped names, byte-comparing all four state files. +19. MINOR review fixture unused — FIXED. Task 7 Step 2 asserts the success class and Gate-B state through `shape0-success-review.json`. +20. BLOCKER `note_unverified` never invoked — FIXED. Called from both gate branches; delivery, marker writes and pending live in `flush_notes`, which is the only place that knows whether anything was written. +21. MAJOR shown+pending unreachable behind the early return — FIXED. Coexistence is resolved at the **top** of `flush_notes`, before any decision; the row is tested. +22. MAJOR A5 incomplete, no reset cleanup — FIXED. Thirteen rows including the `bgAdvice` lifecycle and every write/delete failure; `reset_all` extended to all three markers (Task 2 Step 3). +23. MAJOR `out=$(rev)` cannot observe output — FIXED. `rev` (silent) and `revout` (capturing) split, named for what they do. +24. BLOCKER composition is one sentence — FIXED. `note`/`flush_notes` are real code across Tasks 4 and 8: buffer, single flush, status propagation, marker transitions, `Earlier:` prefix, and the removal of the early `exit 0` that would have skipped the flush. +25. MAJOR composition assertions pass on zero output — FIXED. `= 1`, never `-le 1`; the silent case asserts `0`; plus disclosure-first ordering, separator, pending cleared, and `jq -e` parse, across 13 scenarios. +26. MAJOR clause greps instead of goldens — FIXED. `field_of` plus per-branch exact-equality assertions on both output fields (Task 7 Step 3, Task 9 Step 5). +27. MAJOR B3 omits `codex-gate.sh` — FIXED. Added to Task 9's file list and its `git add`. +28. MAJOR spec §9 setup documentation missing — FIXED. Task 9 Step 6, a step of its own, with every settled clause enumerated. +29. MAJOR carried-scope contradiction — FIXED. The "none before Task 9" rule is narrowed rather than the edit moved: B2's hook-side occurrence is a comment on the function Task 3 rewrites, and the checklist says so. +30. MAJOR census greps match nothing — FIXED. Claim-oriented patterns, run 2026-08-01; the real hits are recorded as file:line in the B items. +31. BLOCKER version bump after the first plugin commit — FIXED. The bump is Task 1, before any other plugin file is committed, with the reason stated in the ordering rationale. +32. MAJOR `git stash` counterfactual is empty — FIXED. The pre-change hook is materialized with `git show` into a temp dir beside the new suite; the working tree is not touched. +33. MAJOR `jq` reserialization defeats byte-exactness — FIXED. Fixtures are hand-sanitized; verification compares the **located block** byte-for-byte and a human reads the `diff`. The bound of that claim is stated rather than overclaimed. +34. MAJOR fixtures leak machine and prompt data — FIXED. `tool_input.workingDirectory` and `.instruction` added to the sanitize table; every fixture is read end to end, not only the review one. +35. MAJOR named verification not reproducible — FIXED. Replaced with a procedure that never touches the installed plugin cache: captured payloads through the repo hook in a disposable repo, with exact expected readings. The one reading it cannot produce is named and sourced to the existing `foreground-env-*` captures. +36. MINOR argv-sized grep pattern — DISSOLVED with the `jq` path. Separately, `strip_ws` is bounded at 64 units and the blank test is one `sed` pass, so no shell loop scales with external input. +37. MINOR B and C items name no task — FIXED. All now do, including C4. +38. MINOR version target undecided — FIXED. `0.8.0`, minor, with the reason named where the checker cannot judge it. +39. MAJOR marker concurrency residual unrecorded — FIXED. Added as **C4** with both observable directions; the A5 table states it describes sequential behaviour only. +40. MAJOR locator tests grep fragments — FIXED. Every locator row asserts the resulting class through `class_of`, with an exact expectation; no fragment greps remain in that section. + +--- + +**Housekeeping.** The prior (profiles) cycle's dispositions for slots 1 and 2 were archived to +`*-dispositions.2026-07-26-profiles-cycle.md` rather than overwritten — the slot names carry +no cycle component, which is the defect already filed in `todos.md`. Their findings files were +overwritten by this cycle before that was noticed. diff --git a/.context/codex-reviews/gate-a-plan-pass-1-dispositions.2026-08-04-hardening-round.md b/.context/codex-reviews/gate-a-plan-pass-1-dispositions.2026-08-04-hardening-round.md new file mode 100644 index 0000000..832c71c --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pass-1-dispositions.2026-08-04-hardening-round.md @@ -0,0 +1,78 @@ +# Gate A — plan — pass 1 dispositions + +19 findings, 13 Major, 6 Minor. All 19 accepted; none dismissed. One reached back into a +committed artifact (the spec's story), which was corrected rather than worked around. + +## The finding that mattered most + +**1 (MAJOR)** — `${lit:0:60}` is bash-only substring expansion. My pre-commit sweep ran `sh -n` +over every fenced block and passed it, because `sh -n` is a **parse** check and this construct +is syntactically valid; it fails at *runtime* under `dash` with `Bad substitution`. Verified +directly: `dash -c 'lit=abcdef; echo "${lit:0:3}"'` → `dash: 1: Bad substitution`. + +The plan now uses `printf '%.Ns'`, and Global Constraints names the class of constructs `sh -n` +cannot catch. The sweep itself was extended: parse under **both** `sh -n` and `dash -n`, then +grep for `${var:o:l}`, `[[ ]]`, `<(...)`, `local`, arrays. + +## Fixed — mechanical correctness + +- **3 (MAJOR)** — `grep -c` exits 1 on a zero count, so every expected-zero check would abort on + the correct result. All are now captured with `|| true` and compared numerically. +- **2 (MAJOR)** — Task 2's "only AGENTS.md" grep could never succeed: the spec and the plan both + quote the phrase, so the check always reached its stop branch. Narrowed to `AGENTS.md` alone, + with the reason stated — those other occurrences are quotations of the rule, not copies of it. +- **12 (MAJOR)** — the battery includes `check-version-bump.sh main`, which fails on a committed + `plugins/` change with no bump. With the bump at Task 8, Tasks 2–7 would each have asserted + "expected exit 0" against a battery that cannot pass. The manifest bump moved into **Task 1**, + alongside the first `plugins/` change. +- **11 (MINOR)** — six vs seven `todos.md` updates. Corrected everywhere to seven, one new plus + six existing. +- **13 (MINOR)** — the ledger precondition stopped only on a count that was too high; too low + equally invalidates the premise. Now requires exact equality on all four and prints the + matching rows so the stop report can name them. +- **19 (MINOR)** — `FIRST_WIP` could match several commits after a retry. Now asserts exactly one + anchor and prints the range before any `reset --soft`. + +## Fixed — structure and procedure + +- **6 (MAJOR)** — the three trigger stories had seven `##` sections; the template allows six. The + inheritance inventory became a `###` subsection of §1. Verified: all four embedded stories now + have exactly six. +- **7, 8 (MAJOR)** — the collision procedure classified once up front and left the write + unspecified. Now each path is classified immediately before its own write, with `set -C` so + creation fails if an entry appeared meanwhile, and `[ -L ]` tested separately because `[ -e ]` + is false for a dangling symlink. +- **14 (MAJOR)** — the cross-finding conflict check ran at validation time; the spec requires it + before any edit. It is now **Task 0**. +- **16 (MAJOR)** — mirror parity ran once before the first commit. Now also in Task 9, after every + Gate-B fix, and immediately before the closing amend. +- **18 (MINOR)** — `todos.md` mutations were unconditional. Each now has an idempotency guard: + absent → apply, identical → skip, different → stop. + +## Fixed — claims the plan could not support + +- **5 (MAJOR)** — `/path/to/final-message.txt` and `/path/to/pr-body.md` were placeholders in a + plan asserting it had none. Replaced by `.context/evidence-0.8.1.md` with its full template, + defined in Task 9 Step 1 before any step reads it, and a `&-`; the only assertion is the already-required hook rc 0, so after removing the undefined variable the test passes before the emit change | Emit status 2, pending creation/retention, and shown-marker non-burning remain untested in both parser modes | Invoke `sh "$HOOK"` directly with stdout closed or a failing writer, capture its process status separately, and assert shown absent plus pending created or retained with and without jq +MAJOR | high | Task 6 lines 533-551, task ordering | Task 6 calls `note_discarded` but Task 7 does not define it until the next commit | The intermediate implementation prints `note_discarded: not found`; review paths hide it behind `exit 0`, so the suite can look green while a shipped commit is broken | Define a complete no-op-safe diagnostic interface before wiring it, or combine Tasks 6-7 so every committed state is functional +MAJOR | high | Task 6 lines 498-515, state-preservation coverage | The seeded loop omits `no-result` and all mapped exec/review names, despite the spec requiring every discarded class, both gates, and mapped failures | Clearing or incrementing earned state on unreadable or mapped-tool results can escape tests | Seed and byte-compare all four pass-state files for failure, timeout, backgrounded, and every no-result shape through default and mapped exec/review tools +MINOR | high | Tasks 1-2 and Task 5 success coverage | The costly real `shape0-success-review.json` capture is not used by `rev()` or any class assertion; `rev()` reuses the exec response | Review-envelope drift can be shipped even though the plan claims real coverage of both gate tools | Use the review fixture for `rev()` or add an explicit success-class and Gate-B state assertion for it +BLOCKER | high | Tasks 6 and 8, unrecognized disclosure wiring | Task 6 deliberately lets `unrecognized` fall through to counting, but Task 8 only defines `note_unverified()` and never invokes it | Settled decision 2 is only half implemented: unrecognized calls count with no once-per-workspace disclosure or pending debt | Call the disclosure lifecycle on every unrecognized default or mapped gate call while preserving count/store behavior, and test gate-on, gate-off, and failed-emit paths +MAJOR | high | Task 8 lines 645-687, shown-plus-pending state | The table requires `unverified=present` plus `pending=present` to delete pending, but `note_unverified()` returns immediately when shown exists | A failed earlier delete leaves permanent contradictory state and every-event cleanup promised by A5 never happens | Resolve coexistence before the early return, retry deletion best-effort, and assert deletion success and failure behavior +MAJOR | high | Task 8 lines 643-671, A5 completeness | The claimed complete table contains no bgAdvice lifecycle row, omits absent/present write and delete failures, and its tests exercise only a fraction of rows; no task updates `reset_all()` to clear the three new markers | Tests become order-dependent and marker-write/delete/retry failures, coexistence precedence, and one-shot non-burning can violate the intended delivery semantics unnoticed | Provide the full sequential table across all three markers, define reset cleanup, and add one test per transition and failure edge in both parser modes +MAJOR | high | Task 8 lines 663-670, pending flush test | `out=$(rev)` can never observe a disclosure because Task 2 defines `rev()` with an internal `>/dev/null` | The stated off-to-on test necessarily fails even if pending composition works | Use `run "$(payload ...)"` for output-bearing cases or split silent and capturing helpers +BLOCKER | high | Task 9 lines 701-740, A6-A7 implementation | The implementation step is only “Route every emit through one function”; it supplies no raw writer/composer boundary, pending flush point, status propagation, marker updates, or handling of Task 6's early exit | A6 and A7 are not implementation-grade, and a naïve wrapper can recurse through `note_unverified`, double-prefix disclosure, clear debt on failed writes, or miss silent events | Provide complete POSIX shell pseudocode/code for one buffered message plus one final flush, with explicit status and marker transitions, before this plan can execute +MAJOR | high | Task 9 lines 726-732, composition assertions | `grep -c '^{` with `-le 1` passes zero output and arbitrary non-JSON output, while the scenario list omits Gate-B stale/below-floor/satisfied variants and Gate-A satisfied; it does not assert disclosure-first ordering, separators, both fields, or pending clearing | Broken or dropped composition can pass the advertised branch matrix | Assert exactly one parseable JSON document for every emitting collision, zero only for genuinely silent/no-debt cases, exact two-field content/order/separators, and post-write marker state for every existing branch +MAJOR | high | Tasks 7-10, prompt golden coverage | The spec requires exact golden assertions for both output fields of every failure, no-result, long/short backgrounded, unrecognized, pending, and composed message, but the plan uses clause greps and later mentions only unknown-tool plus scaffolded CLAUDE text | Negation, cause/remedy drift, or dropped clauses in shipped prompts can pass, risking invariant 11 | Add exact expected `additionalContext` and `systemMessage` fixtures for every new and composed branch, reviewed against all 12 prompt standards +MAJOR | high | Task 10 lines 744-784, B3 file scope | B3 requires changing the unknown-tool message in `codex-gate.sh`, but Task 10's file list and exact `git add` omit that file and no earlier task makes the namespace/remedy edit | A shipped prompt continues teaching incomplete mapping semantics while the checklist is claimed discharged | Include `codex-gate.sh` in Task 10's edits and staging, update its golden output, or assign that exact edit to an earlier named task +MAJOR | high | Task 10, spec §9 setup requirement | No step adds `CLAUDE_CODE_MCP_AUTO_BACKGROUND_MS` to README Setup with Claude Code ≥2.1.212, launch-environment/restart semantics, `0`, a sufficiently long positive value, and the failure consequence | A story acceptance criterion and the primary defense against accepted residual C1 are unimplemented | Add a dedicated documentation step and exact assertions or review checklist covering every settled setup clause +MAJOR | high | Carried obligations B lines 31-37 versus Task 3 lines 268-270 | The checklist says no B item is edited before Task 10, but Task 3 explicitly edits and claims the hook-side B2 occurrence | The plan contradicts its own sequencing contract and makes checklist status/audit history unreliable | Either move the B2 prose edit to Task 10 or revise the carried scope rule and task mapping explicitly before execution +MAJOR | high | Task 10 lines 751-758, docs census | The grep looks for “keys on tool name” and “never inspects”, while the actual settled sites say “keyed on tool name” and “never sees the file”; the command returns no B1 hit | The claimed census can miss this repo's CLAUDE.md and the inline template, repeating the documented docs-drift class | Use claim-oriented patterns covering keyed/keys and sees/inspects plus manually read all hits; add exact known-site assertions after edits +BLOCKER | high | Global line 61, Tasks 1-11 commit/verification order | Plugin commits begin in Task 1, but the manifest bump is left uncommitted until Task 11; the required full quality command's commit-based version checker therefore fails after the first plugin commit, and Task 11 runs it before the bump can enter a commit | The plan cannot satisfy “battery before every commit”, cannot leave each task green, and cannot reach its final verification in the stated order | Put the manifest/changelog bump in the first plugin-touching commit or make one WIP commit containing all plugin changes before running the commit-based checker, then structure later commits/checks consistently +MAJOR | high | Task 11 lines 798-808, +check counterfactual | `git stash` only hides uncommitted Task-11 edits; Tasks 3-9 already committed the changed hook, so the test still runs against the post-change hook and will not demonstrate pre-change failure | Required `+check` evidence is false or empty, and stashing unrelated tracked work adds conflict/data-loss risk | Run the new tests against a hook materialized from the actual pre-change commit in a disposable worktree/path, without mutating the user's working tree +MAJOR | high | Task 1 lines 89-127, fixture byte-exactness | Rewriting the whole payload with jq reserializes `tool_response`, and comparing `jq -c '.tool_response'` proves semantic equality only, not raw-byte equality | Canonicalization can erase precisely the escape/whitespace variants the raw classifier and span check must test while the README falsely claims byte-exact provenance | Preserve the raw response slice when sanitizing and compare hashes or exact extracted raw bytes before and after +MAJOR | high | Task 1 lines 91-148, security/assets | The sanitizer leaves `tool_input.workingDirectory` and instructions untouched; the real captures contain Daniel's private temp/user path, while only the synthetic review fixture gets a whole-file privacy inspection | Shipped fixtures disclose machine identity/path data and can later expose prompt content across the plugin trust boundary | Neutralize all non-response machine/prompt fields and manually inspect every fixture end-to-end before staging, not only the review fixture +MAJOR | high | Task 11 lines 810-816, named verification | The referenced INDEX records payload dumping but not a runnable setup for launching/restarting Claude Code with each process-start environment, configuring the 135-second probe, restoring the installed cached hook, or recording exact expected readings | The high-risk `battery+check+verification` criterion is not reproducible and can leave a globally installed hook modified if interrupted | Specify disposable repo/server/session commands, before/after cache hashes and trap/restore steps, exact counter/fingerprint expectations, and the durable evidence destination +MINOR | medium | Task 4 lines 372-375, risk+security abuse path | The full externally supplied result text becomes a command-line grep pattern; a mapped server can return enough data to exceed argv limits or make hook classification disproportionately expensive | A malicious or pathological external result can delay PostToolUse or force uncertainty/no-result behavior, harming availability and observability | Avoid argv-sized patterns and rescan payload/text in streaming awk with bounded linear work; add a large-payload compatibility test +MINOR | high | Carried obligations lines 31-43 | Unlike A1-A7, B1-B3 and C1-C3 do not name the task that discharges or preserves them, despite the checklist's explicit “every item names the task” contract | Reviewers cannot mechanically audit those six mappings and C residual preservation can silently drift | Add `→ Task 10` to B items and explicit preservation/evidence task references to C items +MINOR | high | Task 11 lines 789-796, version choice | “Bump the manifest version” leaves the target version and semantic rationale undecided even though the current version is 0.7.1 and invariant 12's checker deliberately cannot judge correctness or direction | Implementers can choose inconsistent patch/minor/decreasing values and still pass CI | Name the exact next version and why it matches this behavior change before execution +MAJOR | high | Task 8 A5 and spec §6 concurrency contract | The plan carries unserialized counters and trusted `.context/` as C3 but omits the spec's separate accepted concurrent check-emit-write behavior for diagnostic markers, where disclosures may duplicate or be lost | A “complete” lifecycle review can accidentally overclaim sequential guarantees or engineer partial locking that changes only one state family | Record concurrent marker races as an unchanged accepted residual, state their observable directions, and keep sequential tests from claiming stronger guarantees +MAJOR | high | Task 4 lines 296-317, locator test assertions | The positive locator tests grep for a marker fragment rather than compare exact extracted bytes, so the embedded awk implementation's entire-tail output satisfies them | The concrete real-fixture mis-parse can survive the tests that claim to pin A1 | Assert exact output and status for each locator case, including that no container suffix or later payload field is present +END OF FINDINGS (40 total) diff --git a/.context/codex-reviews/gate-a-plan-pass-1.2026-08-04-hardening-round.md b/.context/codex-reviews/gate-a-plan-pass-1.2026-08-04-hardening-round.md new file mode 100644 index 0000000..7d168c3 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pass-1.2026-08-04-hardening-round.md @@ -0,0 +1,20 @@ +MAJOR | high | Mechanical sweep: Task 1 Step 5, shell block at line 133 | `${lit:0:60}` is bash-only substring expansion even though all 24 shell fences parse under `sh -n` | the parity command aborts with `Bad substitution` under POSIX `sh` or `dash`, violating invariant 4 and preventing its expected output | truncate portably with `printf '%.60s' "$lit"` or print the full literal +MAJOR | high | Mechanical sweep: Task 2 Step 3 normalized-anchor audit | the quoted grep cannot produce "only AGENTS.md" because the committed plan contains the phrase twice and the committed spec contains it twice | the task always reaches its stop-and-surface branch even when there is no second operative copy | restrict the canonical check to `AGENTS.md` and give any repository-wide search an explicit allowlist for the spec and plan quotations +MAJOR | high | Mechanical sweep: Task 3 Step 2 and Task 7 Step 1 | both expected-zero `grep -c` checks return status 1; Task 3's standalone block and Task 7's loop ending on the new-class zero therefore finish nonzero | an executor can stop on a mechanically correct result and never reach the write or collision decision | count with `awk`, or capture `grep -c ... \|\| true`, then assert the numeric value explicitly +MAJOR | high | Mechanical sweep: `b9111a2` five-fix premise and source story AC 7 | the commit updates the spec and amendment log to "before design resumes" with a named step, but the story's three-trigger-stories criterion still says "before design begins" | the settled source artifact contradicts the plan's story text and preserves the pass-9 unsatisfiable wording | add a plan edit that replaces story AC 7 with the full named-step "before design resumes" criterion and verify both affected story criteria +MAJOR | high | Mechanical sweep: cited-path existence and Task 9 Steps 7-8 | `/path/to/final-message.txt` and `/path/to/pr-body.md` do not exist, and the plan supplies neither artifact's full text despite claiming no placeholders and exact artifact text | an engineer must invent the durable evidence-bearing commit message and PR description, so Gate-B context and closure can diverge from the settled requirements | provide exact templates, named workspace-relative paths, required fields, creation steps, and expected validation for both artifacts +MAJOR | high | Read audit: Task 5 Steps 2-4 | each of the three trigger-story bodies has seven `##` sections because `Inherited conditions` was inserted as section 5 | all three violate the settled six-section story-template requirement while Task 5 Step 5 never checks section count | keep the six canonical sections and place each inheritance inventory under an existing section or a `###` subsection, then assert six `##` headings for every story +MAJOR | high | Read audit: Tasks 4-5 collision procedure | the plan classifies with `[ -e ]` and then leaves the write operation unspecified; it neither uses an exclusive create nor rechecks bytes immediately before staging, and `[ -e ]` also misses a dangling symlink | a path can appear or change after classification and be overwritten or silently committed, despite the spec's fail-if-appeared requirement and invariant 9 | classify symlinks too, create missing files with an operation that fails if any directory entry now exists, and byte-verify existing or newly created content immediately before staging +MAJOR | high | Read audit: Task 5 Step 1 | all three paths are classified once before any story is written instead of each path being classified immediately before its own Step 2, 3, or 4 write | the second and third checks are stale by construction and cannot detect collisions arising while earlier stories are written | move one classification check directly in front of each corresponding write and retain stop-on-different behavior +MAJOR | high | Read audit: Task 4 Step 2, unprofiled explanation | the story says a fabricated but confirmed-looking profile is classified by CLAUDE.md §5 as unresolvable and stopped, but §5 cannot detect missing human confirmation when the values are syntactically and semantically valid | the new artifact overstates what the gate proves, violating the gate-claims Don't and prompt-standard item 11 in the round hardening that class | say that such a profile would look valid and §5 would not reveal that nobody confirmed it, which is why the debt stays explicit +MINOR | high | Read audit: Task 5 Step 2, Affected AGENTS.md invariants | the ledger-supersession story considers changing every scaffolded ledger but lists invariants 8 and 9 only, omitting prompt conformance and the plugin version bump | choosing the downstream-template branch later would begin from a story that fails to name invariants 11 and 12 | add invariants 11 and 12 as conditional constraints when the design touches the inline template or another plugin path +MINOR | high | Read audit: architecture, deliverables, File Structure, and Task 6 | the plan says six `todos.md` row updates and one new plus five existing, but Steps 1-7 modify one new row and six existing rows | scope totals and completion accounting are internally inconsistent even though the individual edits enumerate seven rows | change every total to seven and every breakdown to one new plus six existing +MAJOR | high | Read audit: Task 2 Step 4 through Task 8 Step 4 | the recommended execution skill requires an isolated branch, but after Task 1 commits a plugin change the full battery's version-bump check fails until the manifest bump is committed; Task 8 runs that battery before committing the bump as well | the stated exit-0 expectation is impossible for every intervening full-battery step on the normal execution path, and invariant 12 cannot be validated where the plan places it | commit the manifest bump with the first plugin edit or distinguish an interim battery that omits the commit-range checker, then run the actual full battery only after the bump commit +MINOR | high | Mechanical sweep and read audit: Task 7 Step 1 | the procedure says to stop only when a fingerprint count is higher than 5, 5, 1, 0, although any lower count also invalidates the reviewed ledger premise, and its command prints counts without the row the stop report must name | a deleted or rewritten row leaves the engineer guessing whether to proceed, while a higher count requires an unplanned second command to identify it | require exact equality for all four counts and print the anchored matching rows whenever any value differs +MAJOR | high | Stop-path audit: Task 9 Step 5 | the cross-finding conflict check runs only after Tasks 1-8 have applied and committed every edit, while the spec requires it before applying any edit | the named stop condition is detectable but too late to prevent the plan from choosing and encoding conflicting directions | move the check ahead of Task 1 and repeat it if a Gate-B fix changes a hardening's site or direction +MAJOR | high | Executability audit: Task 9 Steps 2-5 | the plan repeatedly says to record battery numbers, five named-verification verdicts, prompt-conformance results, story reviews, and the conflict verdict but declares no new file and names no destination or record schema | the engineer must guess where evidence lives, and most of the validation can disappear with session context instead of being auditable at closure | name one durable evidence artifact or exact commit-body sections and provide the complete field structure plus expected entries +MAJOR | high | State-path audit: Task 1 Steps 5-6 and Task 9 Gate-B loop | mirror parity is checked only before the first WIP commit and is not rerun before Gate B, after Gate-B fixes, or before the closing amend | a review fix can change only `CLAUDE.md` or only the inline template and still reach the final commit without the plan's parity oracle running again | make parity part of Task 9 pre-call evidence and revalidate it after every fix and immediately before the closing amend +MINOR | high | Read audit: Task 1 Steps 4-6 | the deliverable requires byte-identical sentence text, but both verification commands normalize whitespace and therefore accept changed wrapping, indentation, or internal spacing | the checks prove normalized prose equivalence, not the byte identity the plan claims | either define normalized equivalence as the settled requirement or extract the inserted ranges and compare their exact bytes after removing only explicitly permitted surrounding indentation +MINOR | medium | Retry-state audit: Task 6 Steps 1-7 | every backlog mutation is phrased as an unconditional add or append with no precondition for a partially completed task being resumed | an interruption after some edits can duplicate status blocks or the new parked row when the task is retried | before each mutation assert absent, byte-identical-already-applied, or conflicting; reuse the identical state and stop on conflict +MINOR | medium | Executability audit: Task 9 Step 1 | `FIRST_WIP` collects every historical subject match and the plan neither asserts exactly one hash nor states the expected three-log-line shape | a resumed or repeated Task 1 makes the quoted revision invalid and blocks the squash at a risky history-rewrite step | select an anchored unique commit in the intended range, assert one hash and its parent, and state the expected post-squash log result before resetting +END OF FINDINGS (19 total) diff --git a/.context/codex-reviews/gate-a-plan-pass-1.2026-08-12-ledger-supersession.md b/.context/codex-reviews/gate-a-plan-pass-1.2026-08-12-ledger-supersession.md new file mode 100644 index 0000000..83ae618 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pass-1.2026-08-12-ledger-supersession.md @@ -0,0 +1,19 @@ +MAJOR | high | mechanical sweep — Task 1 Step 1 | the expected change set omits the untracked plan itself, which is present alongside the two modified artifacts and `docs/research/`, and no later task commits the plan | the step cannot observe its stated expected state and the approved plan would remain outside the repository and outside `BASE` | add the plan path to Task 1's files, status expectation, prose-only classification, and docs-only baseline commit after Gate A closes +MAJOR | high | mechanical sweep — Task 4 Step 5 | the quoted end sentinel is ``…and nothing checks the difference.`` but the spec explicitly requires the literal bytes ``and nothing checks the difference.`` and forbids an editorial ellipsis in that code span | the quoted locator is not byte-findable, so an executor following it cannot locate the insertion boundary by the stated oracle | replace the ellipsis form with the literal sentinel from spec §4 +MINOR | high | mechanical sweep — fixtures table `entry-future-dated` | the fixture is assigned to `C1d`.6, but distinction 6 is zero-versus-at-least-one while the future-date failure exercises the eligibility bound in `C1d`.4 | the fixture-to-oracle map and the self-review's label-consistency claim are false | relabel the fixture to the locator-and-eligibility distinction and update every trace that cites it +BLOCKER | high | read pass — Task 8 Steps 3–5 | Step 3 soft-resets HEAD to the parent of the first WIP and then Step 4 runs Gate B before creating a replacement WIP commit, so the review range is empty; Step 5 then amends `BASE` itself rather than a WIP commit | Gate B reviews none of the implementation and the close folds the Task 1 baseline artifacts into an unreviewed final commit, contradicting the cited §5 mechanics | soft-reset to `BASE`, immediately create one consolidated `WIP:` commit, review `BASE..WIP`, then amend that WIP to the real commit only after the final clean pass +BLOCKER | high | read pass — Task 8 Step 4 | after a Gate-B fix the plan says only to re-review; it never amends the fix into the WIP snapshot before the next call | `mcp__codex__review` reads the committed git range, so a worktree-only fix produces a re-review of the stale diff while appearing to satisfy the loop | after every fix, revalidate and amend it into the consolidated WIP commit while retaining the `WIP:` prefix, then call Gate B again against `BASE` +MAJOR | high | read pass — opening contract, Task 2, and Task 8 | the plan says the execution-time shell is reviewed by Gate B, but the harness and fixtures are scratch artifacts that are never committed and Task 8 passes neither their paths nor contents into a reviewable range | the named checks can be buggy while Gate B has no artifact from which to review them, so the plan's enforcement claim is impossible as written | keep the plan oracles-only; either define a reviewable executable artifact and update the spec's change surface, or revise the spec and plan to state that the scratch checks are self-tested and shellchecked but not Gate-B-reviewed and record that residual in §8 +MAJOR | high | read pass — Global Constraints lines 38–39 and Self-Review item 2 | the item-2 remedy says the syntax is a standing convention, but the next bullet says “Nothing standing is added” | the remedy merely mentions the held residual and immediately reinstates the same pre-narrowing claim, so held-not-fixed item 2 is not closed | replace it with “No standing check is added”; keep the distinction between a standing authoring convention and no standing machine consumer +MAJOR | high | read pass — check authority paragraph and `C1d`.1–`.4 | the plan says spec §6 governs every disagreement even though the held remedies intentionally correct §6's missing row-date equality and plural scope wording | an executor can follow the declared authority and discard the very corrections this pass is required to land; `C1d`.1 and `C1d`.2 are substantive closures and `C1d`.4 plus `entry-wrong-rowdate` is substantive locally, but the precedence sentence makes those closures non-authoritative | state that §6 governs except where the four explicit §8 held remedies supplement or correct it, and make the plan's mapped remedies authoritative for those four residuals +MAJOR | high | read pass — Global Constraints | invariant 9 is absent even though the story, spec, and review brief all identify the inline-template change as touching `/workflow-init`'s no-silent-overwrite contract | the plan never requires the executor to preserve or verify the existing present-and-different → diff-and-ask path for an older scaffolded ledger | add the spec §4 invariant-9 decision as a constraint and include its verdict in Task 8's prompt-conformance read +MAJOR | high | read pass — `C1e`, `rows-backdated`, and Tasks 4–5 | the plan builds a one-time chronology check for a backdated row even though this change appends no table row and spec §8 itself says this change cannot enter that state | this is a check for an unreachable change state and adds shell and fixtures that cannot validate the implementation under review | delete `C1e` and `rows-backdated` from this plan and record the absent standing chronology validation in spec §8, as required by the known cut +MAJOR | high | read pass — fixtures table and Task 4 Steps 1–3 | the fixture matrix omits spec §6's required many-match case, the on-date boundary, a calendar-invalid entry date, a longer malformed row, and the empty-`finding` versus parse-failure distinction | implementations that require exactly one match, use a strict-before bound, accept impossible dates, accept overlong rows, or conflate an empty field with parse failure can pass every named fixture | add explicit positive and negative fixtures for each omitted distinction and state the expected result per label under both `sh` and `dash` +MAJOR | high | read pass — fixtures heading and Task 4 Step 1 | `rows-escapes` is a valid adversarial row that a correct parser must accept, yet the plan requires every fixture to fail its target label before it can be trusted | the task is impossible for the correct checker and encourages treating valid escaped content as malformed | replace “each must be able to fail” with an expected-outcome matrix; require `rows-escapes` to pass the correct check and demonstrate that it kills the specific naive parser mutant +MAJOR | high | read pass — Tasks 3–5 discrimination steps | no mutation proves `C1b` rejects a changed protected row, `C1c` independently guards all four current/base inputs and unique delimiters, or `C2b` rejects a one-surface parity change | those checks can be implemented as no-ops or weaker three-input/presence-only checks and still pass every prescribed run | add reversible counter-checks for the protected row, each header input and delimiter guard, and a one-surface region mismatch; also self-test stderr-as-failure as required by the global oracle +MINOR | high | read pass — Task 4 Step 5 and `C3` | the plan mandates one blank line on each side of the entry list but no label asserts that layout and no mutation removes either blank | all mechanical checks can pass while CommonMark renders the following `Columns:` paragraph into the list, violating the stated output | add blank-line cardinality and adjacency to `C3` and prove each missing-blank mutation fails +MAJOR | high | read pass — Task 3 Step 7 and Task 8 Steps 3–5 | generic §5 protocol is restated in the plan through WIP behavior, reset mechanics, `baseSha`, branch-file deletion, re-review, spec-update, evidence, and amend rules instead of being cited | the duplicate has already drifted into the empty-range and wrong-amend blockers above, exactly the failure the citation-only rider prevents | strip the protocol copy, cite the precise CLAUDE.md §5 prose-exemption and Mechanics paragraphs, and retain only change-specific resolved values such as `BASE`, the story path, target findings paths, and the current evidence entry +MINOR | high | read pass — plan-wide narration | the plan records cycle history and war stories in the `$D` explanation, pass ranges in Task 1, “past harness got wrong” in Task 2, and “sites that have drifted before” in Task 8 | the standing rider requires an executable artifact of decisions and oracles, not narration of how the cycle arrived there | strip those historical clauses and state only the current date rule, current artifact state, required harness properties, and concrete falsification targets +MINOR | high | read pass — opening contract and check table | the plan says every label names an observation that fails before the change and titles the third column “Fails without the change because”, but `C1b`, `C1c`, and `C1e` pass on the clean pre-change tree and fail only under mutations | the plan overstates the counterfactual coverage and can make passing baseline checks look defective | rename the column to “Falsifying observation”, state exactly which labels fail on the actual pre-change tree, and reserve the mode evidence for a check that genuinely does +MINOR | high | read pass — Task 6 interface and Task 8 Step 5 | Task 6 promises that the closing commit body names `0.8.2`, while Task 8 specifies only the story path and evidence entry for that body | the promised deliverable is not independently required by the closing step | either require `0.8.2` in the final message or remove that promise from Task 6's interface +END OF FINDINGS (18 total) diff --git a/.context/codex-reviews/gate-a-plan-pass-1.md b/.context/codex-reviews/gate-a-plan-pass-1.md new file mode 100644 index 0000000..22989d0 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pass-1.md @@ -0,0 +1,19 @@ +BLOCKER | high | Global Constraints lines 34-35, Task 1 Step 11, Tasks 2-5 commit steps, and Gate B lines 612-624 | The plan promises one Gate-B cycle after all five tasks but creates ordinary non-WIP commits at the end of every task before that review. | CLAUDE.md §5 requires Gate B before a real commit and says every non-WIP commit closes the cycle and discards its accumulated passes; the proposed history therefore creates several unreviewed closed cycles and cannot produce the claimed single-cycle evidence. | Keep the whole change in one Gate-B WIP snapshot, use WIP amends for task checkpoints, run the battery and counterfactual plus the Gate-B loop on the complete diff, then close once with the reviewed evidence entry; make any desired task history only after preserving that reviewed closing state. +MAJOR | high | Task 1 Step 8, Task 2 Step 4, Task 3 Step 2, and Task 5 Steps 3-4 | Deferring the plugin version bump until an uncommitted Task 5 edit makes the intermediate full batteries red on a feature branch, and Task 5 runs the commit-based version checker before the bump is committed. | After Task 1's ordinary commit, check-version-bump.sh main sees committed plugin changes with no committed version change; before Task 5's commit it still cannot see the working-tree bump. On main the same commands pass only because the merge-base is HEAD, which the plan itself calls trivial and useless. | Put all plugin edits and the 0.9.0 bump into the Gate-B WIP commit before running the quality battery, then run check-version-bump.sh against the actual current base ref; do not claim each task is green while committed plugin changes precede the bump. +MAJOR | high | Task 1 Step 4 lines 259-269 and Step 6 lines 304-316 | Check 4c counts index($0, canon), so any longer line containing the canonical sentence, including a prefix, suffix, or negating wrapper, satisfies the purported byte-for-byte rule. | Spec §5.2 requires the canonical line itself byte-for-byte and exactly once; the proposed check establishes only one line containing that substring, and a direct awk probe accepted a longer negating line. | Compare the complete record with $0 == canon and add reject fixtures for leading text, trailing text, and two canonical substrings on one physical line. +MAJOR | high | Task 1 Step 4 lines 246-270 | Fence mode does not identify the CLAUDE.md template block or its opening fence; it takes the sole ## 5. anywhere in the command and the first later four-backtick line, and it cannot detect a premature or duplicate candidate closing fence. | A misplaced severity rule in another scaffold block can satisfy 4c while the actual inline CLAUDE.md §5 lacks it, contradicting spec §5.2's region definition and its fail-closed duplicate-boundary requirement. | Parse and uniquely identify the CLAUDE.md template opening before accepting its ## 5., pair it with that block's closing fence, reject extra or premature candidate boundaries, and exercise those states with full-shaped workflow-init fixtures. +MAJOR | high | Task 1 Step 4 lines 273-279 | The caller silently continues when either required prompt copy is absent because of [ -f "$sev_file" ] \|\| continue. | The assertion is defined over both prompt copies and is required to fail closed; deleting one entire file currently turns that half of the check off and can leave the checker green. | Treat a missing or non-regular required path as a named failure, then pass every existing file to the parser so unreadability remains a separate fail-closed result. +MAJOR | high | Task 1 Step 2 lines 96-220 | The fixture matrix omits several cases required by spec §5.2 and by the review brief: an empty overwrite is impossible because empty $3 or $4 retains the valid initializer, and there are no template-side before/after-boundary, absent-start, absent-end, duplicate-end, unreadable-input, or 4c-specific parser-failure cases. | The untested branches include the parser's weakest file-specific logic, so the suite can pass while the implementation violates the approved path table and fail-closed contract. | Make overwrite intent separate from body content so an empty file is representable, then add symmetric boundary cases for both file shapes plus missing, unreadable, and selectively injected awk-failure fixtures whose diagnostics name 4c. +MAJOR | high | Task 1 Step 4 lines 280-295 | The plan says both fail messages contain closed severity set, but the awk-failure diagnostic contains neither that substring nor a specific closed-set label; BAD also collapses absent start, duplicate start, absent end, duplicate end, and unreadability into one generic diagnosis. | A parser-failure fixture would fail the SEV diagnostic-isolation assertion, and operators cannot tell which distinct repair applies, contrary to the plan's isolation claim and prompt-standard item 10's cause-specific failure guidance. | Give parser failure and each BAD class stable cause codes and cause-specific messages containing the 4c diagnostic key, and test every code independently. +MAJOR | high | Task 1 Step 2 lines 151-224 and Self-review lines 639-647 | The supposedly executable test patch still contains ${TPL_OK/$CANON/nothing here}, a non-POSIX substitution, and delegates its replacement to an unspecified manual rewrite while also claiming every edit carries exact shell. | Applying the shown patch verbatim makes dash exit 2 with Bad substitution before the red TDD oracle can run; a warning about later rewriting is not an executable, decidable plan step. | Replace every non-POSIX expansion in the plan itself with complete literal fixture bodies or a POSIX helper, then run the red phase under both sh and dash before presenting the code as exact. +MAJOR | high | Task 1 Step 10 lines 361-377 | The mutation recipe says to run for 4a, 4b, and 4c but hard-codes only the 4c markers and omits the header procedure's baseline-status, mutant-status, nonempty-flip-set, and accept-case checks. | It cannot execute the stated three mutations as written and can record a syntax-broken or otherwise invalid mutant as evidence, reintroducing the verification-masks-failure class the repository's canonical procedure prevents. | Parameterize the marker per check and reproduce the complete guarded procedure from check-invariants.sh for each run, including status captures and an explicit comparison of the flipped names against reject and accept inventories. +MINOR | high | Task 1 Step 10 lines 375-377 and Self-review lines 645-648 | The plan asserts the existing 20 and 22 flip counts are now wrong while also saying the new counts cannot be predicted; the proposed 4c fixtures do not exercise 4a or 4b, so those two counts may remain exactly unchanged. | This is an unverified numerical claim in the step meant to repair stale mutation evidence and conflicts with the plan's own oracle discipline. | Say the recorded counts are invalidated and must be remeasured, without asserting that their values change; replace them only with observed results. +MAJOR | high | Task 2 Step 1 lines 408-461 | The block labelled verbatim from spec §2.1 is not verbatim: it removes the condition that the no-destination case is a closed, unmerged branch heading for ordinary or rebase merge, drops the explicit does-not-reopen-any-gate sentence, and changes the squash-carry and mandatory-scope wording. | Those edits alter the approved applicability and gate-calibration text in the most safety-sensitive shipped paragraph, despite the plan promising the settled block rather than a new paraphrase. | Copy §2.1's shipped block byte-for-byte into the plan and both prompt copies, preserving paragraph boundaries and every limiting clause. +MINOR | high | Task 3 Step 1 lines 509-517 | The sentence labelled verbatim from spec §3 is line-wrapped across three source lines, while spec §3 explicitly pins the shipped sentence byte-for-byte on one line. | The rendered prose is similar, but the plan contradicts its stated oracle and cannot demonstrate byte equality to the approved source. | Copy the exact one-line sentence from spec §3, then use that same source line in both prompt copies. +MAJOR | high | Task 4 Step 3 lines 559-563 | The backlog task claims to cover spec §6 but lists only four new rows and omits tracked re-review debt and tier-2 counting/containment; it merely says the gate-cycle slot-collision row stays open even though no such todos.md row currently exists. | Required parked work and its trigger would disappear from the durable backlog, making the Self-review's §6 coverage claim false. | Enumerate every §6 backlog row as an explicit todos.md edit, including tracked re-review debt, creation of the fired-but-open slot-collision row, and tier-2 counting/containment pointing to its story, with each approved trigger copied exactly. +MAJOR | high | Task 4 Files lines 539-542 and Steps 1-4 lines 546-570 | docs/hardening-log.md is declared modified and staged, but no step specifies any row to append, its fields, or an oracle for the append-only edit. | The task is not executable or decidable, and either stages no ledger change or invites an implementer to invent one outside the approved plan. | Either remove hardening-log.md from the task if the spec requires no ledger append, or provide the exact append-only row and verification that resolves prompt-vague-criteria at check 4c's actual guarded spelling. +MAJOR | high | Global Constraints and Tasks 1-3 verification steps | The plan never performs the manual all-12-item prompt-standards review required by AGENTS.md invariant 11 and spec §5.3; it relies on the battery even while stating that only one task has a mechanical oracle. | scripts/check-invariants.sh is explicitly only a narrow floor, so a green battery cannot establish conformance of the new reader rule, exception form, or squash instruction. | Add a named, recorded review of all 12 checklist items for each changed prompt copy after the final prose is assembled and before Gate B, without misrepresenting it as mechanical enforcement. +MAJOR | high | Task 1 Step 9 lines 341-354 | The quoted replacement target begins at Two narrow checks and excludes the surrounding review-is-the-gate and Every-other-item sentences, but the replacement text includes both of those sentences. | Following the instruction literally duplicates adjacent invariant text and leaves AGENTS.md internally malformed; the edit is not anchored to the exact text it says it replaces. | Either replace the full existing paragraph including both surrounding sentences, or make the replacement begin at Three narrow checks and end at a floor, not coverage without repeating untouched text. +MINOR | high | Self-review lines 634-637 | The plan says all ten spec §4 paths appear across Tasks 1-5, but neither story path appears in any task; those two were already changed in c0a6ed2. | The implementation scope may still be correct, but the coverage proof is factually false and hides the distinction between already-landed closure work and remaining salvage work. | Account for each §4 row individually: mark the two story files satisfied by c0a6ed2 with verification, and map only the remaining eight paths to Tasks 1-5. +MINOR | high | Task 2 commit message lines 481-494 | The second task's commit body says it closes the story's remaining scope even though rider (c), backlog work, release metadata, the final evidence run, and Gate B are still pending. | The durable history would claim completion before the plan's own completion conditions hold, undermining the record form's emphasis on calibrated commit evidence. | Describe the commit as the record-form portion only, or make the completion statement solely in the final reviewed closing commit after every task and evidence obligation is complete. +END OF FINDINGS (18 total) diff --git a/.context/codex-reviews/gate-a-plan-pass-2-dispositions.2026-07-26-profiles-cycle.md b/.context/codex-reviews/gate-a-plan-pass-2-dispositions.2026-07-26-profiles-cycle.md new file mode 100644 index 0000000..70f35d7 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pass-2-dispositions.2026-07-26-profiles-cycle.md @@ -0,0 +1,11 @@ +# Gate A — plan — pass 2 dispositions + +7 findings, 2 Blockers. All accepted. 12 -> 7. + +1 BLOCKER the evidence entry cites a SHA that the amend then invalidates — ACCEPT. A commit body cannot name its own SHA; every amend produces a new one, so "battery green at " is stale by construction. Fixed by removing the SHA entirely: the entry names the battery and its result, and the tested state IS the tree this commit records, because revalidation is required before the close. Non-circular and checkable. +2 BLOCKER fixes are not amended into the WIP snapshot before the battery and re-review — ACCEPT. Uncommitted fixes sit outside `baseSha..HEAD`, so Gate B could return clean on the pre-fix diff while the closing commit carries unreviewed changes. Every fix now stages and amends the active `WIP:` snapshot (preserving the evidence body) BEFORE the battery re-run and the re-review. +3 MAJOR the agreement check guards only one of four extractions — ACCEPT. Missing anchors on both sides diff clean. All four captures are asserted non-empty before the two diffs. +4 MAJOR shipped §5 text omits the spec's multi-story aggregation — ACCEPT. The Profiles subsection only said every cited story must be skip-eligible; it now carries the whole rule (battery once per cycle, per-story mode and suffix with its own evidence entry, lens sets unioned) in both copies and inside the agreement check. +5 MAJOR any historical override excuses a mode/axes mismatch — ACCEPT, and sharp: an axis change VOIDS prior overrides, so a stale one must not resolve a mismatch. Resolution is now ordered — only the latest `mode override` after the latest `axis change`, with a compatible direction, explains a mismatch; anything else stops. Applied to the spec too, so the two agree. +6 MAJOR the plan states this story's concrete profile values — ACCEPT. Third violation of my own single-copy constraint in this artifact. Values removed; the plan cites the story path and instructs a fresh header read, noting that this story predates profiles and therefore runs unprofiled unless adopted through §6's procedure. +7 MINOR getting-started teaches only half the feature — ACCEPT. The paragraph now says the axes select review lenses AND the derived mode sets the evidence required before Gate B. diff --git a/.context/codex-reviews/gate-a-plan-pass-2-dispositions.2026-07-30-classifier-cycle.md b/.context/codex-reviews/gate-a-plan-pass-2-dispositions.2026-07-30-classifier-cycle.md new file mode 100644 index 0000000..484e52c --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pass-2-dispositions.2026-07-30-classifier-cycle.md @@ -0,0 +1,38 @@ +# Gate A (plan) — pass 2 dispositions + +Advisory companion to `gate-a-plan-pass-2.md` (5 BLOCKER, 23 MAJOR, 4 MINOR = 32). One line +per finding, in file order. All accepted; two partly, with reasons. Three were fixed in the +verified drafts rather than in prose, and `verify.sh` grew 35 → 50 rows to pin them. + +1. BLOCKER Tasks 5–7 cannot be green at their commits — ACCEPT. Locator, classifier, wiring, messages and disclosure are now **one task and one commit** (Task 5). The classifier is unobservable until it is wired and the wiring calls the diagnostic interface, so every split produced a commit shipping either an uncalled function or a call to a missing one. Steps stay ordered and separately verifiable; only the commit is single. +2. MAJOR `class_of` / `LAST_DISCARDED` reads cross-test state — ACCEPT. One observer, defined in Task 2, that resets and classifies **one** invocation from that invocation's own counter, marker and exact message. `LAST_DISCARDED` is gone. +3. MAJOR harness still incomplete — ACCEPT. Task 2 now defines `off_file`, `resp`, `resp_from`, `resp_success`, `unrec`, `payload_from`, `run_scenario`, `class_of`, `golden`, `field_of` and every restricted-PATH runner, with exact shell. The either-or on response fixtures is decided: sibling `*.response.json` slices, with a permanent test pinning each slice to its payload. +4. BLOCKER Task 7 calls `note_unverified` and `bg_advice_file` before they exist — ACCEPT. Dissolved by the merge in finding 1. +5. MAJOR Task 3's writer test drives a branch that emits nothing — ACCEPT, and it was worse than stated: a successful review `PostToolUse` never calls `emit`, so both assertions held before the change. Now driven through the unknown-tool note, whose one-shot marker is the observable consequence of status 2, in both emitter modes. +6. MAJOR jq-free `emit` swallows encoder failure — ACCEPT. Both `sed` substitutions must succeed or return 2, plus a test with `sed` removed from `PATH`. +7. BLOCKER the scan is not the conservative walk claimed — ACCEPT, fixed in code: containers keep a **closer stack** (`[1}` is refused, not balanced); primitives must be a complete JSON token; member grammar rejects stray and trailing commas; the outer object is walked in full and garbage after its `}` is refused. Seven malformed-but-routable shapes, **each carrying a success envelope**, are now tested. +8. MAJOR duplicate members and escaped key spellings — ACCEPT for duplicates: a repeated `type` or `text` inside the selected element now refuses instead of last-wins, tested in both orders. PARTLY DISMISSED for escaped key spellings: an escaped key simply fails to match, so the payload takes the fail-closed path or an unambiguous other candidate — never a wrong verdict — and decoding `\uXXXX` would put a unicode decoder inside a hook whose safety argument is that it decodes nothing. Stated in A2 as out of scope by decision. +9. MAJOR ambiguity fixtures cannot fail — ACCEPT, and it applied to more rows than named. Every refusal case now carries a **success** envelope in at least one candidate, so a locator that picks a block instead of refusing changes the observable class. +10. MAJOR two spec cases missing from the ported table — ACCEPT. Added: a result quoting `tool_response` **after** the real field, and a genuine Unicode-escaped marker key, distinct from the escaped-space row. +11. MAJOR the ` ` spelling was erased — ACCEPT, and it was my own pass-1 edit that did it. Restored as literal code text in spec §3.3 and in the plan's A3. +12. MAJOR byte-exact proof ignores locator status — ACCEPT. Both extractions must exit 0 before their outputs are compared; two failures would otherwise compare equal and print `OK`. +13. MAJOR preservation tests miss the clean-state case — ACCEPT. Every discarded class now asserts **both**: nothing created from clean state, and byte-identical from seeded state. An implementation that rewrites the seeded fingerprint with an identical current hash passes the second and fails the first. +14. MAJOR `reset_all` leaves the opt-out marker — ACCEPT. It now clears it, so a test that wants the gate off sets it explicitly after — intent rather than inheritance. This was making the closed-writer assertion pass on suppression, not on write failure. +15. MAJOR the C2 test destroys adoption — ACCEPT, and the assertion was fully vacuous. `chmod 500 .context` instead, keeping `codex-gate.on` readable, with a `skip -` guard for platforms that ignore it. +16. MAJOR A5 rows untested; no jq-free closed writer — ACCEPT. One assertion block per row, including shown-write failure (a **directory** at the marker path), pending retention and every `bgAdvice` transition, each repeated through `nojq_run` / `nojq_run_closed`. +17. MAJOR composed assertions compare fragments — ACCEPT. Both fields compared whole against an expected value built by the A7 rule. Fragment greps passed on truncation, reordering, duplication and negation alike. +18. MAJOR the `silent` scenario is unreachable — ACCEPT, good catch: `hooks.json` matches `^(Bash|Skill|mcp__codex__.*)$`, so a `PostToolUse` for `Edit` never reaches the hook. Replaced with a non-commit Bash `PostToolUse`, which is reachable and matches no branch. +19. MAJOR the jq-free PATH lacks `mktemp` and `cp` — ACCEPT. The full command list is linked, and the guard is itself asserted: under that PATH a success fixture must record a usable, self-matching fingerprint, or the jq-free Gate-B rows compare two different code paths. +20. MAJOR `codextool` carries text `x` — ACCEPT, my own regression. It now carries a real success envelope; a separately named helper exists for the intentionally unrecognized case. +21. BLOCKER message placeholders — ACCEPT. All six pairs are literal in Task 5 Step 1, declared unconditionally so `set -u` cannot abort, and written against `docs/prompt-standards.md` with the binding items named. +22. MAJOR layout trees drift when a shipped directory is added — ACCEPT. `AGENTS.md` and `docs/architecture.md` are in Task 6's file list with their own step and a grep-first check. +23. MAJOR an intermediate commit publishes 0.8.0 with a changelog for absent behaviour — ACCEPT, and it is resolvable without weakening pass-1 finding 31: the **manifest number** is what the commit-based checker needs and stays in Task 1; the **entry** is what a reader needs and moves to Task 7, landing with the behaviour it describes. +24. MAJOR "run the battery" is not what the tasks run — ACCEPT. Each commit step now carries the exact pre-commit command list and the post-commit `check-version-bump.sh`, with the reason the latter runs after. +25. BLOCKER Gate B would review one task — ACCEPT. `baseSha` is the **merge-base with main** (the SHA recorded at Task 1), not a WIP commit's parent; the cycle closes by amend or `reset --soft` to that base. +26. MAJOR the named verification substitutes pre-change captures — ACCEPT. The two `foreground-env-*` captures now run **through the changed hook**, which is exactly the prevented reading criterion 10 asks for. The entry states what they establish (the classification) and what they do not (that the variable still keeps a call in the foreground on the current build — that was established when they were captured, and the date is cited). +27. MAJOR the `+check` grep succeeds on any failure — ACCEPT. Four exact test labels must be present against the old hook and absent against the new one. +28. MINOR `HEAD~9` is fragile — ACCEPT. The base SHA is recorded at Task 1's commit and the materialized blob is verified against it. +29. MAJOR the missing-`awk` contract is promised, not tested — ACCEPT. A `noawk` PATH, a **failure** envelope through both gate tools asserting fail-open disclosure, and an exit-0 assertion. The failure envelope is deliberate: it is the case where fail-open costs a real count. +30. MINOR two shells is not two `awk`s — ACCEPT, my overclaim. The provenance note now says what 50/50 covers, and CI's `mawk` is named as the second implementation, with "unverified until CI runs" stated rather than implied. +31. MINOR the locator is unbounded — ACCEPT. A payload past **1 MiB** is refused rather than accumulated, routed to `unrecognized`; tested. +32. MINOR the sanitization block exits nonzero on success — ACCEPT. `diff`'s status is accepted at 0 or 1 and only a real error fails the step. diff --git a/.context/codex-reviews/gate-a-plan-pass-2-dispositions.2026-08-04-hardening-round.md b/.context/codex-reviews/gate-a-plan-pass-2-dispositions.2026-08-04-hardening-round.md new file mode 100644 index 0000000..f5f8be0 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pass-2-dispositions.2026-08-04-hardening-round.md @@ -0,0 +1,61 @@ +# Gate A — plan — pass 2 dispositions + +11 findings, 10 Major, 1 Minor. Down from 19/13. All 11 accepted; none dismissed. + +**Ten of the eleven are one class:** a verification block that reports a failure and then exits +0, or that cannot see the failure at all. In a plan whose deliverable includes the +`verification-masks-failure` hardening, that is the round committing its own subject matter — +found by a reviewer, not by the sweep I ran before committing. + +## The cluster + +- **2 (MAJOR)** — every verification block printed its verdict and ended on an `echo`, so shell + status was 0 on every mismatch. Nine blocks affected. All now accumulate into `rc` and end with + `exit "$rc"`, or use an explicit failing branch. A new Global Constraint states the rule. +- **3 (MAJOR)** — the cited-story loop ran in a pipeline subshell, so `rc=1` inside it was + discarded; the check could print `MISS` and still report `todos exit=0`. Rewritten to read from + a redirected file, and it now also asserts exactly four distinct paths and rejects symlinks. +- **5, 6 (MAJOR)** — `applied()` and `FIRST_WIP` were defined in one fenced block and consumed in + later ones. Agentic workers run each block in a fresh shell, so neither was in scope. The guard + is now written out per step, and the squash recomputes and re-validates its anchor inside the + same block that resets. +- **4 (MAJOR)** — marker counts cannot distinguish "already applied, identical" from "partially + applied or independently edited". The guard now compares the whole intended block, + whitespace-normalized, with three explicit branches. +- **8 (MAJOR)** — mirror parity extracted matches from whole files without checking cardinality, + so duplicates or two empty extractions would report equivalence. Cardinality is now asserted + first. Placement under the corresponding heading is confirmed by reading, and the plan says so + rather than implying the grep establishes it. + +## Path handling + +- **9 (MAJOR)** — a symlink whose target held identical text was eligible for reuse; `git add` + would stage the link, and every shape check would follow the target and pass. A symlink is now + an unconditional stop, tested before `[ -e ]`, which is false for a dangling link and true for + a live one. +- **10 (MAJOR, medium)** — reserving an empty placeholder with `set -C` closed the + check-to-create race and opened a create-to-write one. Replaced with write-to-temp then `ln`, + which fails if the target exists, so content is installed atomically. +- **11 (MINOR)** — every creation failure was reported as a concurrent writer. The failure branch + now re-stats the path and names a collision only when an entry exists, otherwise surfacing the + real cause — the `prompt-diagnostic-cause-unnamed` class in this repo's own taxonomy. + +## Claims the plan could not support + +- **7 (MAJOR)** — the plan told the engineer to record counts "the run actually printed" for + shellcheck and the hook suite, which print none. Established by running them: the hook suite + prints `all passed` with no total (467 `ok` lines when counted); the invariants and version-bump + suites print `all passed (123 assertions)` and `(36 assertions)`; shellcheck prints nothing. + The evidence template now asks for each suite's terminal line verbatim and labels the + hook-suite count as **derived**, with the deriving command — which was run before being + written down, per "Never document a command that wasn't run". +- **1 (MAJOR)** — the no-PR story's inheritance inventory marked one trigger alternative moved and + omitted the other ("or a project reporting an empty ledger across cycles that fixed findings"). + A dropped condition in the inventory whose whole purpose is to drop none. Added and marked + **kept**, since it did not fire. + +## Sweep after fixing + +38 fenced blocks: all parse under `sh -n` and `dash -n`; none uses `${var:o:l}`, `[[ ]]`, +`<(...)`, `local` or `declare`; none accumulates `rc` inside a pipeline subshell; every block +that can fail exits nonzero when it does. diff --git a/.context/codex-reviews/gate-a-plan-pass-2-dispositions.2026-08-12-ledger-supersession.md b/.context/codex-reviews/gate-a-plan-pass-2-dispositions.2026-08-12-ledger-supersession.md new file mode 100644 index 0000000..8dcc538 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pass-2-dispositions.2026-08-12-ledger-supersession.md @@ -0,0 +1,50 @@ +# Gate A — plan — pass 2 dispositions + +Artifact: `docs/superpowers/plans/2026-08-12-hardening-ledger-supersession.md`. +11 findings: **1 Blocker**, 8 Major, 2 Minor. All eleven valid. **All eleven applied.** + +Reviewer: `mcp-codex-dev@1.0.1`, `codex-cli 0.144.5`, model `gpt-5.6-sol`. Pass file valid: +12 lines, 11 finding lines, terminator exact. + +Plan passes: **18, 11.** + +## The Blocker + +**Task 4 Step 9 through Task 8 Step 5 — `git add -A` stages what it must not.** The working tree +carries `docs/research/`, untracked and not ours, and the executor creates fixtures. Either could be +swept into a commit or an amend — changing the Gate-B range, and publishing an unrelated tree. +Applied as three constraints: never a pathspec broader than the step's own file list; **assert the +staged set equals the intended paths before every commit and amend**, via +`git diff --cached --name-only`; and the harness lives **outside the repository working tree**, so +nothing the executor creates appears in `git status` at all. Task 8 Step 5 and Task 1 Step 3 were +rewritten to stage explicitly. + +## The eight Majors + +| Finding | Verdict | Applied as | +|---|---|---| +| "Counter-check every pattern against the pre-edit tree: it must return 0 there" | **Valid, and self-contradictory as written.** Carried over from the spec's fence rider, where it meant *the pattern must not match*; sitting beside an exit-status convention where 0 is success and a table saying five labels must **fail** pre-change, it instructs the executor to record the required counterfactual as a harness defect. | Replaced with an explicit clean-tree outcome table: five labels must fail, two must pass, and each direction's failure mode named. | +| `rows-truncated` / `rows-overlong` expected only `C1a` to fail | **Valid.** `C1d`.7 independently requires `C1d` to reject incomplete and overlong rows; a matcher accepting them passed every stated outcome. | Both cells now **FAIL `C1a`, `C1d`**. | +| No fixture pairs a good mandated entry with an extra inert one | **Valid — this is what closes §8 item 3 with evidence rather than words.** A checker still quantifying over every added entry passed the whole matrix and the real tree. | `entry-good` now carries the mandated entry **plus a second inert entry**, expected PASS. | +| `entry-bad-date` invalidated only the entry date | **Valid.** `C1d`.5 covers the locator's row date too; a checker validating one satisfied every listed outcome. | Split into `entry-bad-date-entry` and `entry-bad-date-locator`, both FAIL. | +| `entry-empty-field` said "one prose field empty", unspecified | **Valid.** The oracle requires both fields independently. | Split into `entry-empty-first-field` and `entry-empty-second-field`. | +| Matrix rows naming `C3` were to be run in Task 4, before `C3` exists | **Valid — an ordering defect, not a wording one.** The step could not be completed as written, and Task 5 never re-ran them. | Task 4 Step 3 now runs only rows naming `C1a`/`C1d`; Task 5 gained a step running the deferred `C3` rows. | +| Task 5's counter-check moved the whole block | **Valid.** A moved block still has every entry inside its own interval, so the interval oracle stayed unverified. | Added: append a second entry-shaped line **below the table while the valid block stays in place**. | +| `C1c`'s counter-check mutated "one input" | **Valid, and precise.** A three-input implementation reads both current surfaces and one base; mutating any of those three still fails, so the mutation did not identify the omitted input. | Now names it: corrupt the **base version of the template** specifically, confirm the guard fires before comparison. | +| Task 8 restated §5's target-file deletion rule as "delete both before each call" | **Valid on both counts.** It violates the citation-only rider **and** contradicts §5, which requires a single-branch resume to delete only the failed branch — deleting both makes the both-files check fail by construction and can spend the sole recovery attempt on a path that cannot succeed. | Reduced to `reviewType: full` plus a citation; the lifecycle and recovery rules stay §5's. | + +## The two Minors + +"Two prose paragraphs" → **"Two convention blocks"**, since §2.2's block is an intro paragraph, an +indented example and a closing paragraph. Narration stripped at three sites: the landed-files list +reframed as *outside this execution's scope, verify and do not rewrite*; Task 1's commit message +described by artifact content rather than by cycle closure; Self-Review's §7 line reworded to what +§7 is rather than to the state of a finished review. + +## Sweep after applying + +Step numbering contiguous in all eight tasks (Task 5's insertion renumbered); 17 fixture rows with +no stale count claim anywhere; eight labels, `C1e` appearing only where its absence is recorded; 35 +anchors byte-identical to the spec; 12 fence lines, balanced; every stated count — eight labels, +five falsifying labels, four inputs, seven distinctions, thirty-five anchors, twelve items — +verified against what it counts; `add -A` appears once, in the constraint forbidding it. diff --git a/.context/codex-reviews/gate-a-plan-pass-2.2026-07-30-classifier-cycle.md b/.context/codex-reviews/gate-a-plan-pass-2.2026-07-30-classifier-cycle.md new file mode 100644 index 0000000..18e4e6f --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pass-2.2026-07-30-classifier-cycle.md @@ -0,0 +1,33 @@ +BLOCKER | high | Tasks 5–7, test and wiring order | Task 5 tests call `class_of` before Task 6 defines it, and neither the locator nor classifier is invoked by the hook until Task 7, so implementing Tasks 5 or 6 cannot make their advertised end-to-end tests pass | The Task-5 and Task-6 commits cannot be green or prove the code they add, violating the plan's commit-by-commit contract | Combine locator, classifier, and state wiring into one functional task, or add a concrete test-only harness before Task 5 that directly exercises each layer +MAJOR | high | Tasks 5–7, `class_of` and `LAST_DISCARDED` | `class_of` claims discarded classes are recovered from `LAST_DISCARDED`, but locator cases discard hook output, a hook subprocess cannot set a test-shell variable, and no shown code assigns that variable; success versus unrecognized also remains indistinguishable until Task 8 creates markers | The five-class assertions can read stale or unset state and either abort under `set -u` or report a class unrelated to the invocation under test | Make one assertion helper classify the same invocation from its counter, fingerprint, marker, and exact emitted message, without cross-test global state +MAJOR | high | Task 2 interfaces and later uses | The supposedly complete harness still omits concrete definitions for `off_file`, `resp_from`, `resp_success`, `run_scenario`, and the `LAST_DISCARDED` mechanism; `resp_from` is retroactively assigned to Task 2 in Task 7 with an unresolved either-or implementation choice, while `run_scenario` is only promised in prose | The suite aborts under `set -u` or leaves implementers inventing load-bearing test code after the task that was meant to define it | Put exact POSIX-shell implementations and every path variable in Task 2 before first use, choosing one response-fixture representation now +BLOCKER | high | Task 7 Step 4–5 | Both gate branches call `note_unverified`, but Task 8 defines it later, and `note_backgrounded` reads `bg_advice_file` before the plan says to define that path; these are live discarded/unrecognized paths under hook `set -u` | The Task-7 commit emits command-not-found errors or exits nonzero, breaking the explicit fix for pass-1 finding 17 and invariant 1 | Define the complete diagnostic interface and all marker paths in Task 7, or merge Tasks 7 and 8 so no committed hook calls undefined names +MAJOR | high | Task 3 Step 1, writer-failure test | The test closes stdout on a successful review `PostToolUse`, but that branch emits no reminder, so `emit` is never called; both asserted outcomes already hold before the proposed `emit` change | Task 3's claimed failing test is behaviorally unrelated to writer-status propagation, leaving the new 0/1/2 contract unverified at its commit | Drive an existing emitting branch such as the unknown-tool note or a commit reminder, and assert its one-shot marker is not written when stdout is closed +MAJOR | high | Task 3 Step 3, jq-free `emit` | The two `sed` command substitutions do not propagate encoder failure; if either fails it yields an empty field, `printf` can still succeed, and `emit` returns 0 | A malformed or empty reminder can burn `toolNote`, `bgAdvice`, or `unverified`, losing one-shot disclosure while the code claims a complete document was written | Require both encoder substitutions to succeed or return 2, and add a selective encoder-failure test that checks marker retention and hook exit 0 +BLOCKER | high | Task 5 A2 and `.context/plan-drafts/locate.awk:17-48` | The scanner is not the conservative structural walk the plan claims: `skipval` uses one undifferentiated depth for braces and brackets, accepts arbitrary primitive tokens and stray commas, and the top loop ignores trailing garbage; a grep-routable payload with a mismatched prior container or a valid outer object followed by garbage can return a `success` block | Malformed payloads can become trusted success and mutate pass state in jq-free routing while jq routing ignores them, creating a false checkmark and contradicting invariant 4 plus the amended spec's safe-mislocation claim | Validate matching delimiters, member/value grammar, comma placement, and end-of-document after the outer object, then add malformed-but-routable success-prefix cases that must be `unrecognized` +MAJOR | high | Task 5 A2 and `.context/plan-drafts/locate.awk:43-69` | Ambiguity detection covers only the literal spelling of a depth-1 `tool_response`; an equivalent escaped key and duplicate `type` or `text` members are accepted with last-value behavior | A payload with semantically duplicate members can be classified success or failure instead of taking the settled uncertainty path, and an external MCP result sits across the stated trust boundary | Define duplicate semantics for every classification-relevant member, decode or conservatively reject escaped key spellings, and test conflicting duplicates in both orders +MAJOR | high | Task 5 Step 1 and `.context/plan-drafts/verify.sh:71-74` | The repeated-key fixture uses unmatched texts `a` and `b`, and the truncated-array fixture uses unmatched text `x`, so a buggy locator that selects either block instead of refusing still produces final class `unrecognized` | The tests cannot establish the locator-refusal behavior their labels and pass-1 disposition claim to pin | Put a success envelope in one ambiguous candidate and a failure envelope in the other, and put a success envelope in every truncated/unwalkable case so an incorrect location changes the observable class +MAJOR | high | Spec §7.3 and Tasks 5–6 coverage | The ported 35-row table still omits two explicit spec cases: a result that quotes `tool_response` after the real field and a Unicode-escaped `success` marker; the existing Unicode row is only an escaped space | Regressions to a last-anchor heuristic or Unicode marker handling can ship despite the plan claiming the full boundary/refusal matrix is discharged | Add end-to-end cases for a quoted response-key decoy after the selected block and for an actual Unicode-escaped marker key, with exact expected classes +MAJOR | high | Spec §3.3 line 161 and Task 6 A3 | The pass-1 amendment erased the literal `\u0020` spelling, leaving `A Unicode-escaped space (` `)` in both the normative explanation and plan while `verify.sh` still tests six literal escape bytes | The blank grammar is no longer self-contained or unambiguous, so A3 cannot be reviewed or implemented from the approved artifacts | Restore the literal escaped spelling as code text in the spec and plan and keep separate tests for ASCII space versus the six-byte Unicode escape +MAJOR | high | Task 1 Step 3, byte-exact proof | Both locator calls are captured without checking their exit statuses, so two failures produce equal empty variables and print `OK` | Fixture sanitization can destroy the response under test while the provenance README falsely claims byte-exact preservation | Capture and require status 0 for both extractions before comparing, then fail the step on either locator error +MAJOR | high | Task 7 Step 1, discarded-state coverage | Seeded byte comparisons do not include the complementary empty-state assertion that a discarded Gate-B call creates no fingerprint; an erroneous store of the current hash rewrites the seeded fingerprint to identical bytes and passes | The story's no-fingerprint acceptance criterion can regress while every preservation assertion stays green | For every discarded class and default/mapped review name, assert both absence from a clean state and byte preservation from seeded state +MAJOR | high | Task 7 Step 6, failed-write one-shot test | `reset_all` deliberately does not remove the opt-out marker, so after the suppressed backgrounding case the next `run_closed` remains suppressed and never attempts a write | The test passes because `emit` returns suppression status 1, not because writer status 2 preserves the one-shot | Remove the off marker explicitly before `run_closed`, or introduce separate reset helpers whose opt-out behavior is explicit +MAJOR | high | Task 8 Step 1, marker-write-failure test | Replacing `.context` with a regular file also removes `.context/codex-gate.on`; with no gate-bearing `CLAUDE.md` in the fixture, `is_adopted` exits before classification or marker writing | The only advertised C2 write-failure assertion is vacuous and cannot support the disposition that the accepted silent-count residual is tested | Keep adoption through a temporary gate heading in `CLAUDE.md` or fail only the individual marker path while leaving the project adopted, then assert the count, missing disclosure markers, and exit status +MAJOR | high | Task 8 A5 tests versus marker table | The shown tests omit the promised one-per-row cases for failed shown-marker writes, failed pending deletion, pending retention after suppressed/failed flush, and failed `bgAdvice` writes; saying they run under `nojq_run` also cannot test closed stdout because no jq-free closed-writer helper exists | Sequential disclosure debt can be lost or duplicated on precisely the failure edges A5 was deferred to the plan to settle | Add explicit portable fault injection and post-state assertions for every table row in both emitter modes, including a `nojq_run_closed` helper +MAJOR | high | Task 8 Step 4, composed-message assertions | The matrix checks exact document count and only prefix/separator fragments in `additionalContext`; it never compares either composed field to a golden or verifies disclosure-first ordering and separation in `systemMessage` | A branch message or operator remedy can be truncated, reordered, duplicated, or negated while all assertions pass, so pass-1 findings 25–26 and invariant 11 remain incompletely dispositioned | Build exact expected `additionalContext` and `systemMessage` values for every scenario and compare both fields, updating the unknown-tool composition golden again in Task 9 +MAJOR | high | Task 8 Step 4, `silent` scenario | The load-bearing silent event is `PostToolUse` for `Edit`, but `hooks.json` invokes this hook after only Bash, Skill, or `mcp__codex__*` tools | The test proves debt flushing on an event production can never deliver, not on a reachable silent invocation | Use a matched silent event such as a non-commit Bash `PostToolUse` and assert it flushes exactly one disclosure document +MAJOR | high | Task 2 jq-free PATH setup | The inherited symlink list lacks `mktemp` and `cp`, both required by `tree_hash`; merely adding `awk` makes jq-free Gate-B scenarios compute `unavailable` and enter different branches | The Task-4/8 branch matrix cannot establish behavior parity for stale, floor, or satisfied Gate-B output without jq | Include every external command used by the hook, especially `mktemp` and `cp`, and assert the jq-free success fixture records a usable, self-matching fingerprint +MAJOR | high | Task 2 `codextool` helper | Mapped-tool tests are changed to carry text `x`, which classifies `unrecognized`, not a real success envelope | Existing mapped success assertions can pass by fail-open counting while also creating disclosure state, masking a regression in decision 1 and contaminating later one-shot tests | Make counting mapped calls use a captured success response and reserve a separately named helper for intentionally unrecognized text +BLOCKER | high | Tasks 7–8 message implementation and Self-Review | Runnable code still contains literal long/short-form placeholders, and `UNVERIFIED_CTX` plus `UNVERIFIED_MSG` are described but never assigned, even though exact goldens and all 12 prompt standards depend on their final text | The plan is not implementation-complete, hook `set -u` can abort, and Gate A cannot review the actual shipped prompts required by invariant 11 | Put the final exact strings in the plan and matching exact test fixtures, then review those strings against all 12 checklist items before execution +MAJOR | high | Task 1 file additions and Task 9 file scope | Adding shipped `hooks/fixtures/` changes the meaningful repository surface, but neither `AGENTS.md` nor `docs/architecture.md` is in the plan's update census and both layout trees still enumerate only three hook files | The plan creates the docs-drift condition AGENTS.md explicitly says to prevent when adding files | Update both layout trees to include the fixture directory and include those edits in the appropriate task and review scope +MAJOR | medium | Task 1 version/changelog ordering | The first commit publishes version 0.8.0 and a changelog entry claiming the new classifier behavior, while Tasks 3–8 have not implemented that behavior yet | Any checkout or install at that intermediate commit has a release version whose shipped hook contradicts its changelog, so the workaround for invariant 12 makes intermediate history misleading | Keep implementation in one WIP release commit, or make the bump/changelog and all behavior changes atomic before any non-WIP release point +MAJOR | high | Global Constraints versus Tasks 1–9 commit steps | The plan says the full AGENTS.md quality row runs before every commit, but Task 1 runs only the hook suite and one invariant script, Task 2 runs only the hook suite, and later commit steps likewise omit the full lint, invariant suites, version suite/checker, and strict plugin validation | The stated per-commit green-tree guarantee is not executable from the task commands and can miss packaging or POSIX failures until the end | Give each task the exact applicable full battery, explain the commit-based checker timing explicitly, and record post-commit validation where the current commit must exist to be checked +BLOCKER | high | Task 10 Step 4, Gate B range and commit history | Tasks 1–9 already create nine commits, leaving no change for the requested WIP commit, and reviewing a new WIP only against its parent would cover at most Task 9 rather than the full implementation | Gate B can declare clean without reviewing the classifier, state machine, or most tests, violating CLAUDE.md §5 and the high-risk profile | Record the pre-Task-1 base, review the full base-to-WIP range, and squash or amend the WIP sequence into the cycle-closing commit only after the clean pass +MAJOR | high | Story acceptance criterion 10 and Task 10 Step 2 | The story requires the probe methodology to be re-run against the changed hook with backgrounding both absent and prevented, but the plan explicitly does not produce the prevented reading and substitutes captures made before the change | The named high-risk verification does not satisfy the approved acceptance criterion or demonstrate the changed hook on that path | Run the foreground captures through the changed hook and perform the required end-to-end launch-environment probe, or obtain a human-confirmed story amendment before implementation +MAJOR | high | Task 10 Step 1, `+check` counterfactual | Piping the old suite to `grep '^FAIL'` succeeds when any unrelated new assertion fails and does not require the two failure-class regression tests named by the story to fail | The evidence entry can claim the required counterfactual while the actual defect tests are vacuous | Match and record the exact failure-test labels, assert both are present against the old hook, and assert those same labels pass against the new hook +MINOR | high | Task 10 Step 1, base selection | `HEAD~9` assumes exactly one commit per numbered task and no review-fix or split commit, despite the plan itself inviting fixes and manual capture work | One extra or squashed commit silently materializes the wrong supposedly pre-change hook and corrupts the counterfactual | Record the exact base SHA before Task 1 or derive it from a stable merge-base, then verify the selected hook blob matches that base +MAJOR | high | Global awk failure contract and Tasks 5–8 coverage | The plan promises missing or failing `awk` maps to counted-and-disclosed `unrecognized`, but no test removes or stubs `awk` and checks class, state, disclosure, and hook exit 0 | A load-bearing dependency failure can instead become `no-result`, abort under `set -u`, or violate invariant 1 without detection | Add missing-command and nonzero/partial-output awk cases through both gate tools and assert fail-open state plus best-effort disclosure +MINOR | high | Plan provenance claim and POSIX awk compatibility | Running `verify.sh` under `sh` and `dash` invokes the same `/usr/bin/awk`, so 35/35 under two shells does not test the newly load-bearing code across awk implementations | The evidence overstates portability on uncontrolled machines and leaves implementation-specific awk behavior unexamined | State the actual shell-only bound and add at least one materially different POSIX awk implementation in CI or a named compatibility verification +MINOR | medium | Risk high plus security standard abuse-path lens, `.context/plan-drafts/locate.awk` | The locator accumulates the entire external payload and repeatedly appends every scanned string without any input-size or work bound; only the later shell whitespace loop is bounded | A malicious or pathological mapped MCP result can make an advisory PostToolUse hook consume disproportionate CPU or memory and delay the workflow | Define an accepted payload ceiling or demonstrate bounded linear behavior at the maximum supported result size, routing oversize or resource failure to `unrecognized` +MINOR | high | Task 1 Step 3, human diff command | Each expected sanitized `diff` returns status 1, so the displayed verification block finishes nonzero even when every edit is correct and aborts immediately if run in a fail-fast shell | Implementers can mistake successful sanitization for failure or skip the remaining files, weakening the manual privacy check | Wrap the expected diff status explicitly while still failing on operational errors, and require a recorded human inspection for all four files +END OF FINDINGS (32 total) diff --git a/.context/codex-reviews/gate-a-plan-pass-2.2026-08-04-hardening-round.md b/.context/codex-reviews/gate-a-plan-pass-2.2026-08-04-hardening-round.md new file mode 100644 index 0000000..cce635f --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pass-2.2026-08-04-hardening-round.md @@ -0,0 +1,12 @@ +MAJOR | high | Task 5 Step 2, Conditions inherited from the source row | the no-PR story marks only the "next round that touches §5" trigger as moved and omits the source row's alternative trigger, "or a project reporting an empty ledger across cycles that fixed findings" | the replacement story silently drops one settled condition, violating the Don't that every old condition be marked kept, moved, or deliberately dropped | add the alternative trigger to the inheritance inventory and mark whether it was kept, moved, or deliberately dropped +MAJOR | high | Task 1 Steps 6-7; Task 2 Step 3; Task 3 Step 2; Task 4 Step 3; Task 5 Step 4; Task 6 Step 8; Task 7 Steps 1 and 6; Task 8 Step 2 | the verification blocks print a failure verdict but finish with a successful echo, so every mismatch tested in the mechanical sweep returned shell status 0 | an agent or wrapper that trusts command status can proceed through a promised stop path after parity, version, story-shape, ledger-count, or evidence prerequisites fail, recreating the verification-masks-failure class this round hardens | end each block with a nonzero status when its accumulated check fails, for example `exit "$rc"`, and make the single-condition blocks use an explicit failing branch +MAJOR | high | Task 6 Step 8 | the cited-story loop prints `MISS` but never sets `rc`, and the pipeline subshell could not propagate such an assignment even if one were added inside the loop | a nonexistent story can still produce `todos exit=0`, so the check does not enforce the docs-drift condition it claims to verify | avoid the pipeline subshell by feeding a temporary or redirected list into the loop, set `rc=1` on every missing path, assert exactly four unique paths, and exit with `rc` +MAJOR | high | Task 6 preamble and Steps 1-7 | marker counts cannot implement the promised absent versus present-identical versus present-different decision; after a marker is found, the plan gives no comparison against the full intended block | a partial, stale, or independently edited status block can be treated as an identical completed mutation and skipped, leaving `todos.md` semantically wrong while the final marker-count check passes | give each mutation a full normalized-content or bounded-row comparison and explicit branches for absent, identical, and different before editing +MAJOR | high | Task 6 preamble | `applied()` is defined in one fenced shell invocation but every later guard consumes it outside that invocation | agentic workers normally execute fenced blocks in fresh shells, so the guards are not executable from the text as written and require the engineer to guess that the function must be redefined or that an interactive shell must be retained | inline the count command in every guard or provide each guard as a self-contained fenced block that defines and calls `applied` +MAJOR | high | Task 9 Step 2 | `FIRST_WIP` is assigned inside the inspection block and then consumed by the separately fenced `git reset --soft "$FIRST_WIP"^` command | in a fresh shell the variable is empty, so the squash cannot execute as specified and the engineer must reconstruct a history-rewriting target by guesswork | recompute and revalidate `FIRST_WIP` in the same self-contained block that performs the reset, or pass the validated hash as an explicit literal after the inspection +MAJOR | high | Task 9 Step 3 | the plan requires recording counts for shellcheck files and hook-suite assertions "the run actually printed", but the exact quality command prints neither a shellcheck file count nor hook assertion totals; the mechanical run emitted only uncounted `ok` lines and `all passed` for each hook shell | the evidence template cannot be completed exactly as instructed without inventing numbers or adding an unstated counting procedure, conflicting with the rule never to record a number not read from output | either require only the exit status and totals the suites really emit, or supply a tee-and-count wrapper with explicit expected labels and verify that wrapper before using its counts +MAJOR | high | Task 1 Step 6 and Task 9 Step 6 | mirror verification extracts matching sentence prefixes from each entire file but checks neither cardinality nor the required placement under the corresponding heading | duplicate copies or equally misplaced copies can report three `EQUIVALENT` lines even though the spec requires one copy at the same semantic site in both files | scope each extraction to its anchored paragraph or section and assert exactly one normalized match in each file before comparing the full sentence +MAJOR | high | Task 4 Step 1 and Task 5 path-classification procedure | a non-dangling symlink whose target currently has byte-identical story text is eligible for "reuse", and the shape checks follow that target | git would stage a symlink rather than the required durable story bytes, so another checkout can receive a dangling or externally retargeted path while every planned verification passes | treat any `-L "$p"` result as an unconditional stop; only a regular file with identical bytes may be reused +MAJOR | medium | Task 4 Steps 1-2 and Task 5 Steps 1-3 | noclobber atomically reserves only an empty placeholder; the later content write has no ownership check or atomic compare-and-swap, so another writer can modify that placeholder before Step 2 and be silently overwritten | the procedure closes the initial check-to-create race but leaves a create-to-content-write race that violates the stated no-silent-overwrite collision rule | write complete content to a private temporary file, then install it atomically with no-replace semantics, or verify the reserved inode/content immediately before an atomic replacement and stop if it changed +MINOR | high | Task 4 Step 1 and Task 5 path-classification procedure | every noclobber creation failure is reported as "APPEARED between the test and the create" even when the cause is permissions, a missing parent, filesystem failure, or another I/O error | the stop is conservative but its diagnostic makes the engineer investigate the wrong state and overclaims what the check proved | after failure, re-stat the path; report a collision only when an entry now exists, otherwise surface the real create failure without assigning a concurrent-writer cause +END OF FINDINGS (11 total) diff --git a/.context/codex-reviews/gate-a-plan-pass-2.2026-08-12-ledger-supersession.md b/.context/codex-reviews/gate-a-plan-pass-2.2026-08-12-ledger-supersession.md new file mode 100644 index 0000000..93747e9 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pass-2.2026-08-12-ledger-supersession.md @@ -0,0 +1,12 @@ +MINOR | high | Mechanical sweep — Architecture | the shared addition is called "Two prose paragraphs", but the copied §2.1 and §2.2 blocks contain three prose paragraphs separated by the indented format example | the opening count contradicts the block structure even though Task 3 later names the two fenced blocks correctly | replace "Two prose paragraphs" with "Two convention blocks" or explicitly say three prose paragraphs plus one indented example +MAJOR | high | Fixture matrix closing sentence | "Counter-check every pattern against the pre-edit tree: it must return 0 there" contradicts the check table, which says five labels must fail on the untouched tree, and the plan's exit-status convention makes 0 mean success | an executor cannot satisfy both instructions and can record the required failing baseline as a harness defect | replace the sentence with the explicit clean-tree outcome per label: five failures and three passes or mutation-only checks +MAJOR | high | Fixture matrix `rows-truncated` and `rows-overlong` | both rows require only `C1a` to fail even though `C1d`.7 independently requires `C1d` to reject incomplete and overlong table rows | a `C1d` matcher that accepts malformed rows passes every stated `C1d` outcome while violating its row-side oracle | change both expected cells to FAIL `C1a`, `C1d` and run both labels for those fixtures +MAJOR | high | `C1d`.1 and fixture `entry-good` | the mandated-entry-only remedy says other added entries are ignored, but no fixture combines a good mandated entry with an additional inert entry | a checker that still quantifies over every added entry passes the matrix and real tree, so held-not-fixed item 3 is mentioned but not closed by executable evidence | make `entry-good` include the correct mandated entry plus an extra inert entry and keep its expected result PASS `C1d` +MAJOR | high | `C1d`.5 and fixture `entry-bad-date` | the only bad-date fixture invalidates the entry date; none isolates the requirement that the locator row date is calendar-valid | a checker that validates only the entry date satisfies every listed outcome while accepting an impossible locator date | make this fixture cover two variants, including a valid entry date with a matching table row and locator both dated `2026-02-30`, and require both to fail +MAJOR | medium | `C1d`.3 and fixture `entry-empty-field` | "one prose field empty" supplies only one unspecified case although the oracle requires both prose fields independently to be non-empty | an implementation can validate the exercised field only and pass the full matrix | define two variants under this fixture, first field empty and second field empty, and require both to fail +MAJOR | high | Task 4 Step 3 and Task 5 | the matrix expects `entry-outside-block` and `ledger-unreadable` to fail `C3`, but `C3` is not implemented until the next task and Task 5 never reruns those fixtures; its mutations also move the whole block rather than leave a valid block in place with an extra entry-shaped line outside it | the full-matrix step cannot be completed as written and the oracle that every entry-shaped line lies inside the interval remains independently unverified | after implementing `C3`, run every matrix outcome naming `C3` and add a counter-check that keeps the valid block intact while appending a second entry-shaped line below the table +MAJOR | high | Task 4 Step 7 `C1c` counter-check | removing a delimiter from "one input" does not identify the base-template input that a three-input implementation omits | mutating any of the three inputs such an implementation already reads still fails and therefore does not prove the required fourth input is guarded | explicitly mutate the extracted base version of the template, confirm failure before comparison, restore it, and rerun +BLOCKER | high | Task 4 Step 9 through Task 8 Step 5 | the scratch harness and fixtures exist untracked when later commits are made, and Step 5 explicitly runs `git add -A`; the known unrelated `docs/research/` tree is also untracked | the amend necessarily stages scratch and user-owned files, contradicting the opening contract, changing the Gate-B range, and risking an unrelated commit | stage only the named implementation paths for every post-fixture commit, use deletion-aware explicit pathspecs, and assert that neither scratch nor unrelated paths entered the index before each commit or amend +MAJOR | high | Task 8 Step 4 second bullet | the plan copies §5's target-file deletion protocol back into the plan as "delete both ... before each call", while §5 requires a single-branch resume to delete only the failed branch because deleting both makes the both-files check fail by construction | the copied rule both violates the citation-only rider and can spend the sole recovery attempt on an impossible resume | retain only the change-specific choice `reviewType: full` and cite §5 for file lifecycle and recovery; do not restate its deletion rule +MINOR | high | File Structure closing paragraph, Task 1 Step 3, and Self-Review | the plan narrates what already landed, instructs the commit message to name Gate-A closure and held-review history, and records that §7 is closed | the plan is supposed to carry current scope, decisions, and oracles rather than cycle history, so these copies become stale review narration | strip the history; express the three landed paths only as excluded from this execution, make the commit message describe artifact content, and remove the closed-pass note from Self-Review +END OF FINDINGS (11 total) diff --git a/.context/codex-reviews/gate-a-plan-pass-2.md b/.context/codex-reviews/gate-a-plan-pass-2.md new file mode 100644 index 0000000..e140838 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pass-2.md @@ -0,0 +1,21 @@ +MAJOR | high | Task 1 Step 11, Global Constraints lines 33–49, Tasks 2–3 battery steps | Task 1 commits plugin changes before the manifest bump, despite the plan claiming every plugin edit and the bump are in the same snapshot | On a feature branch the full batteries in Tasks 2 and 3 run the commit-based version checker against `main` and fail on the committed unbumped plugin diff; on `main` they pass only because the comparison is empty, so the sequence cannot be both green and meaningful | Put the version bump into the first WIP snapshot, create the WIP only after all five tasks are complete, or defer intermediate full batteries while explicitly replacing them with checks that are valid before the deciding committed snapshot +MAJOR | high | Task 5 Step 3 lines 672–683 | The step named “verify the bump mechanically” runs `scripts/check-version-bump.sh main` while the stated current branch is `main`, and the plan itself admits that this comparison passes trivially | The required invariant-12 check never decides whether the WIP snapshot contains a bump, so the battery evidence can be green without testing the plugin change it is cited for | Capture the WIP parent and run the checker against that exact commit, or require a feature branch whose current base ref is verified before execution; treat an empty HEAD-to-base range as VOID rather than verification +MAJOR | high | Gate B lines 700–712 | The fix loop says only to fix, revalidate the evidence entry, and re-review; it never stages and amends each fix into the WIP snapshot that `mcp__codex__review` reads | A re-review can inspect the previous committed range while the purported fix remains only in the worktree, producing a clean pass over stale content | After every Gate-B fix, stage the complete intended diff, amend the still-WIP commit, rerun the battery and counterfactual, rerun parity and the 12-item prompt review when prompt text changed, then call the re-review with the unchanged WIP parent as `baseSha` +MAJOR | high | Task 1 Step 4, `severity_rule_count` fence mode lines 266–274 | Fence mode accepts `## 5.` inside any four-backtick fenced block, not specifically inside the `CLAUDE.md` template block | If the intended template loses its heading while a later fenced scaffold gains a matching heading and canonical line, the checker returns one occurrence and passes, contradicting the spec’s path-specific region and the plan’s claim that a stray heading elsewhere cannot be taken as the mirror | Anchor fence mode to the `### 2.1 CLAUDE.md` template and its immediately associated outer fence, and add a fixture with the intended heading absent but a valid-looking section 5 in a later fenced template +MAJOR | high | Task 1 Steps 2 and 4, spec §5.2 fixture row for duplicate end boundaries | The plan neither tests nor implements rejection of two candidate end boundaries; the boolean fence toggle simply accepts the first fence that ends the region, and heading mode accepts the first later level-2 heading | A malformed or duplicated boundary can still resolve as valid even though the approved spec requires an ambiguous end boundary to fail closed | Define exactly what constitutes multiple candidate ends for each file shape, make the parser return a distinct ambiguity code, and add reject fixtures for both heading and fence modes +MAJOR | high | Task 1 Step 2 lines 218–222 and Step 4 default case | The plan declares the parser-failure branch untestable, but the existing suite already has a selective PATH-injection seam that identifies an `awk` invocation by a unique program marker | The default `case` branch remains unreachable in the fixture matrix despite the approved spec explicitly requiring parser-stage failure coverage, so it can regress to fail-open without moving any test | Add a unique marker comment to the 4c awk program and use `inject_case` to fail only that invocation, asserting the `closed severity set parser failed` diagnostic +MAJOR | high | Task 1 Steps 2–3, unterminated-template and missing-template fixtures | These two reject fixtures already fail check 4b because the raw unterminated body has no checklist and the missing command file is itself a 4b parser error, so before 4c they do not fail with the promised `(exited 0)` result | The fixtures are not isolated to their named 4c causes and the red-phase expectation is false; after 4c they pass the harness while still carrying unrelated diagnostics | Keep a valid checklist reachable in the unterminated-fence fixture, and isolate missing-template behavior from 4b with a dedicated 4c harness or a selective checker mode; assert that reject output contains the expected 4c diagnostic and no other checker diagnostic +MINOR | high | Task 1 Step 2 unreadable fixture | `@LOCK@` uses mode 000 without the existing suite’s root guard, while root can still satisfy `-r` and read the file | The suite fails in root-run containers even though the checker is correct for ordinary users; the existing scan-error case already documents and handles this environment difference | Skip the permission fixture when `id -u` is 0 or run the checker as a known unprivileged user, and add the symmetric unreadable-template case if per-copy behavior is intended to be covered +MAJOR | high | Task 1 Step 9, proposed edit to `scripts/check-invariants.sh:261` | Changing “the two prompt-conformance checks below” to “three” makes the scan-domain comment false because `PROMPT_EXCL` and the Markdown recursive scan govern only 4a and 4b; 4c reads two fixed paths directly | This introduces a new enforcement-scope overclaim in the exact change meant to calibrate what the checker proves | Rewrite the comment to say the scan domain applies to checks 4a and 4b and that 4c has its own fixed two-file domain; change only the true total checker count in the header and AGENTS.md +MAJOR | high | Task 1 Step 10 lines 402–415 and Gate-B counterfactual lines 720–728 | Both temporary-copy recipes drop the checked header procedure’s `mktemp` failure guard and cleanup trap; the mutation loop also uses `continue` for a red baseline or passing mutant and has no nonempty-flip guard | A failed `mktemp` can redirect copies toward `/repo`, temporary trees leak, and a partially VOID mutation run can still finish successfully and be recorded as evidence despite the header explicitly forbidding that outcome | Restore `mktemp -d || exit 1`, validate the nonempty directory, install a cleanup trap, fail the whole run on either guard, require a nonempty flipped set, and preserve a nonzero final status for every VOID condition in both recipes +MINOR | high | Task 1 Steps 9–10 | The checked-in mutation-procedure edit is not executable as written: Step 9 merely says to add a 4c line “mirroring 4a”, while only the separate one-off run in Step 10 is parameterized across 4a, 4b, and 4c | An implementer can run the loop once yet leave the checker header hard-coded or incomplete, contradicting spec §4 and the plan’s claim that every edit carries exact shell | Provide the exact replacement header procedure, parameterized over all three markers and containing the full guards, rather than an informal line-level instruction +MINOR | high | Global Constraints line 27 and per-task parity steps | The plan requires each prose task to record parity in “that task’s commit body”, but Tasks 2–5 only amend a WIP commit whose message deliberately remains unchanged | The named record has no destination at the time the plan says to write it, so compliance depends on silently reinterpreting the rule as a final-close obligation | State that each task records its result in working notes and that all parity results are written once into the final closing amend, or explicitly amend the WIP body while preserving its WIP subject +MINOR | high | Task 1 Step 2 line 213 | The fixture summary says there are five accepting cases, but only three cases have expected status 0: both copies valid, blockquoted valid, and outside-plus-inside valid | The plan’s numerical self-review is already stale and can mislead the required mutation inventory check | Correct the count to three or add the two missing accepting near-neighbor cases and update the 22-case total accordingly +MINOR | high | Task 1 Step 1 and existing test-suite comments at lines 15–24 and 325 | Adding `sev_case` creates another fixture builder, but the plan does not update the existing “all four builders” claim or the “checks 4a and 4b” section heading | The implementation would leave exact count and inventory comments false, recreating the repository’s standing docs-drift class | Update the initializer comments to the new builder count and rename the prompt-conformance section and mutation record to cover 4a through 4c +MINOR | high | Task 1 Step 6 lines 351–352 | The plan says 4c matches the canonical sentence “as a substring of one line”, while the proposed implementation and the pass-1 correction require normalized equality | This contradicts the load-bearing reason leading and trailing text are rejected and can prompt an implementer to weaken the matcher | Replace “substring” with “for equality after the specified blockquote and indentation stripping” +MINOR | high | Gate B standing-falsification lens lines 745–750 | The plan says 4c depends on `## 6.` following section 5, but the parser actually ends the region at any following line beginning `## ` | This overstates the exact comparison and can cause irrelevant coupling to the section number while missing changes to the actual generic boundary rule | Describe the dependency as exactly one section-5 start and a later level-2 heading; mention that `## 6.` is only the current concrete boundary +MINOR | high | Task 4 Files lines 592–595 versus Steps 3–4 | The file summary says six new todo rows, but Step 3 adds six parked rows and Step 4 separately creates the absent slot-collision row, for seven new rows | The task inventory and its executable steps disagree, weakening the plan’s claim that all §6 backlog changes were counted individually | Change the summary to seven new rows and enumerate them there, or explain which Step-3 item updates an existing row rather than creating one +MINOR | high | Global Constraints lines 43–45 and Gate B close lines 711–712 | The plan says per-task closing-message shares are given for all five tasks, but Tasks 4 and 5 define no closing-message share | The final amend cannot mechanically satisfy the instruction to carry “the five tasks’ message shares”, and backlog or release evidence can be omitted without violating any task-local checklist | Add explicit closing-message shares to Tasks 4 and 5, including the nine backlog dispositions, version bump, changelog, deciding version-check result, battery, counterfactual, parity, and prompt-standards review as applicable +NIT | high | Task 1 Interfaces lines 67–70 | The interface says `CANON` is quoted by Task 5’s changelog entry, but Task 5 only asks the changelog to cover the enum and never quotes the canonical line | This is a small false dependency in the task graph | Remove the claimed reuse or require the changelog to include the exact canonical sentence +NIT | high | Task 1 Step 4 comments lines 237–242 versus Global Constraints lines 22–25 | The checker comment says the line ships indented inside the template fence, while the plan correctly states and the real template shows that the inline §5 mirror is flush-left | The comment gives a false reason for normalization and obscures which whitespace forms are intentionally accepted versus merely tolerated | Say that indentation is tolerated for formatting variation while the current template is flush-left, or narrow stripping to the formatting forms the settled design actually requires +END OF FINDINGS (20 total) diff --git a/.context/codex-reviews/gate-a-plan-pass-3-dispositions.2026-07-30-classifier-cycle.md b/.context/codex-reviews/gate-a-plan-pass-3-dispositions.2026-07-30-classifier-cycle.md new file mode 100644 index 0000000..1dc2f98 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pass-3-dispositions.2026-07-30-classifier-cycle.md @@ -0,0 +1,48 @@ +# Gate A (plan) — pass 3 dispositions + +Advisory companion to `gate-a-plan-pass-3.md` (5 BLOCKER, 22 MAJOR, 8 MINOR = 35). One line +per finding, in file order. + +**The cycle stopped and surfaced after this pass** (B+M 36 → 28 → 27, not converging), and +the human decided the one blocking question: **narrow the locator's claim rather than grow +the parser**, with three riders — state that walk integrity rests on string-boundary +tracking alone; freeze the three counterexamples as fixtures asserting today's behaviour; +and treat a fourth claim-vs-code finding on this same spot as the stuck signal. Findings 2, +3, 4 and 5 are dispositioned under that decision. Everything else is ordinary work, all +accepted. + +1. BLOCKER locator rows cannot route through the hook — ACCEPT, and it was worse than stated: bare `{"tool_response":…}` fragments carry no `hook_event_name` at all. Task 5 Step 2 now requires every ported row to be a routable payload, and names the **two groups that stay in the drafts** — deliberately unroutable documents, and documents whose routability depends on `jq` — with the reason, so they are excluded rather than silently dropped. +2. MAJOR `skipval` is not a recursive parser; `readstr` accepts invalid escapes — ACCEPT as fact, **claim narrowed rather than code grown** (human decision). Verified in-session: `[1,]`, `{"a" 1}` and `"\q"` in a sibling value each return the block with status 0. A2 now says what actually carries safety — string-boundary tracking, which cannot be redirected by malformation *outside* a string — and the three examples are frozen as fixtures asserting exactly that behaviour, so the claim-code correspondence is pinned mechanically instead of argued a fourth time. +3. MAJOR escaped key spellings bypass duplicate detection — ACCEPT as fact, scope stated. A2 records it as the one case where the scan differs from a real parser's duplicate handling, and why it is out of scope: the payload's producer is Claude Code, whose serializer emits valid JSON, so a document carrying both spellings is synthetic. The pass-2 "never a wrong verdict" rationale is **withdrawn** as stated and replaced with the producer argument, which is the true one. +4. MINOR `text` is not recognized as `type: text` — ACCEPT, contract stated explicitly: `type` is compared as raw bytes, so an exotic-but-conforming tool takes the fail-closed `no-result` path. Decoding `\uXXXX` would put a unicode decoder in a hook whose safety argument is that it decodes nothing. +5. MAJOR the ceiling does not bound memory — ACCEPT, my overclaim. `payload=$(cat)` has already stored the whole input before `awk` starts. The plan and the `awk` comment now both say the ceiling caps the **scan** and nothing else, and name what bounding the input would require. +6. MAJOR the encoder-failure test is vacuous without `sed` — ACCEPT, good catch: `field()` needs `sed` to route at all, so the hook never reached `emit`. Replaced with a shim that fails only the escaper's own script and passes everything else through, plus a `skip -` guard. +7. BLOCKER `run_scenario` / `expected_ctx` / `expected_msg` are prose — ACCEPT. All three are now complete POSIX shell in Task 2 Steps 6–7, including per-scenario setup, cleanup for the two scenarios that dirty the worktree, and the A7 composition rule applied once rather than restated per assertion. +8. MAJOR both provenance checks fail open — ACCEPT. Both now `exit 1`. Step 7's claim is also narrowed to what it compares: the **located block**, not the whole array. +9. MAJOR `seeded_preserves` never runs jq-free — ACCEPT. Parameterized over `run` and `nojq_run`, both gates, default and mapped names, all five discarded shapes. +10. MAJOR the missing-`awk` test asserts only disclosure — ACCEPT. It now asserts the **count and a usable fingerprint** for Gate B and `passCountA` for Gate A, across three fault shapes: absent, nonzero exit, and partial output then failure. +11. MAJOR A5 rows still unpinned — ACCEPT. Added: pending retention under suppression and under a failed flush, shown-write failure while flushing an existing pending, failed pending delete plus its cleanup on the next event, and `bgAdvice` write failure after a successful emit. +12. MAJOR the C2 test does not produce the residual — ACCEPT, and `chmod 500 .context` observed neither half. Replaced with surgical injection: pass-state writable, emit suppressed, only the pending path unwritable — then assert counted, no shown marker, no pending marker. +13. MINOR C2 is overstated in the A5 table — ACCEPT. Split into its own row, with the other marker failures listed and their different outcomes named. +14. MINOR "writes nothing at all" contradicts `bgAdvice` — ACCEPT. Relabelled to **gate-pass** state, with the diagnostic-marker effects asserted separately as their own family. +15. MAJOR the composed document is never tested jq-free — ACCEPT, and the plan explicitly claimed that path is stressed hardest by exactly this. One real two-message branch now runs through `nojq_run` with whole-field goldens and an independent `jq -e` parse afterwards. +16. MAJOR the field split is inverted — ACCEPT, the most consequential prompt finding: operator checks and remedies sat in the model's field while the operator got "see the note". Every operator-performable action moved to `systemMessage`; `additionalContext` keeps the gate consequence and what Claude does next. +17. MAJOR items 1, 3, 5 unmet — ACCEPT. Target model named once for all ten strings (item 1), a uniform *state → consequence → next → stop* structure (item 5), and a bounded stop tied to §5's one-attempt recovery budget on every retry-capable state (item 3). +18. MAJOR `FAILURE_CTX` names no causes — ACCEPT. It now names the error codes, what each means, the matching remedy, and where the retry budget ends. +19. MAJOR `.mcp.json` is not authoritative for the effective server — ACCEPT. The message now says to check with `claude mcp list` and why (scope precedence), and gives the two causes separate, complete remedies. +20. MAJOR `UNVERIFIED_CTX` omits known causes — ACCEPT. Oversize refusal, missing/failing `awk` and locator refusal added; the backgrounding remedy is **inlined** rather than pointing at a message a reworded notice never produces. +21. MAJOR the short form drops the variable name — ACCEPT. Settled decision 4 requires it precisely because `bgAdvice` outlives the session that saw the long form. +22. MAJOR the backgrounded call can still write its findings slot — ACCEPT, and this is the finding I would have missed: the hook's new diagnosis would otherwise *create* a race against §5's own file protocol, where a late writer leaves a correctly terminated file from the wrong run. Both backgrounding messages now tell Claude to stop or await the task id before re-running. +23. MINOR "has counted" precedes the best-effort write — ACCEPT. Reworded to "classified as countable and attempted to record it", which is prompt-standards item 11 applied to our own message. +24. MAJOR Task 6's prompts are not literal — ACCEPT. The B1 sentences, the two narrower phrasings, and the unknown-tool message's inserted paragraph are all written out. +25. MAJOR the census misses three sites — ACCEPT, verified: `workflow-init.md:250`, `README.md:59`, `AGENTS.md:70`. The grep now carries three alternations and the plan lists all five hits with a disposition each, including one marked *read and decide* rather than pre-judged. +26. MINOR the AGENTS.md declaration census is skipped — ACCEPT. Both greps plus a same-change read of `plugin.json` are now in Task 6 Step 5. +27. BLOCKER the counterfactual deletes its own evidence — ACCEPT. `$EVID` lives outside the temp directory, a trap does the cleanup, and the assertions run before it. +28. MAJOR the base-blob guard continues on failure — ACCEPT. Every guard exits nonzero; an empty blob is checked too. +29. BLOCKER Task 7's own edit is outside the reviewed range — ACCEPT. The CHANGELOG edit is staged in a named `WIP:` snapshot before Gate B, so the review covers it and the counters are not reset. +30. BLOCKER merge-base resolves to HEAD on this checkout — ACCEPT, verified: the work is on `main`, so `git merge-base main HEAD` is HEAD and Gate B would have received an **empty range**. `baseSha` is now the recorded pre-Task-1 SHA, passed literally, with the branch case as a check rather than a substitute. +31. MINOR the per-commit battery claim is false — ACCEPT. The block is written out once and later tasks require *that exact block, run and green*; the claim that it is copied everywhere is removed. +32. MAJOR the capture mutates the globally installed hook — ACCEPT. Backup, hash before, trap on EXIT/INT/TERM/HUP, restore, hash after and a hard stop on mismatch; plus deleting payloads dumped for unrelated calls, and a preference for an isolated install where one exists. +33. MINOR no rollback — ACCEPT. Task 7 Step 5, with the point that `codex-gate.off` is **not** a rollback (it suppresses messages while classification keeps discarding), the version-keyed cache path is, and the diagnostic markers survive it. +34. MAJOR the load-bearing code lives in ignored scratch — ACCEPT. Both bodies are embedded in Task 5 Steps 5–6. `.context/plan-drafts/` remains where the test corpus lives and where a change is re-verified; Step 2 says the two copies are kept in step by hand. +35. MINOR "six pairs" — ACCEPT, there are five (ten strings). Corrected, with the derived composed and `Earlier:`-prefixed forms named as derived rather than missing. diff --git a/.context/codex-reviews/gate-a-plan-pass-3-dispositions.2026-08-12-ledger-supersession.md b/.context/codex-reviews/gate-a-plan-pass-3-dispositions.2026-08-12-ledger-supersession.md new file mode 100644 index 0000000..981aaa5 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pass-3-dispositions.2026-08-12-ledger-supersession.md @@ -0,0 +1,68 @@ +# Gate A — plan — pass 3 dispositions + +Artifact: `docs/superpowers/plans/2026-08-12-hardening-ledger-supersession.md` +(three findings landed in the spec instead — see below). +9 findings: **4 Blocker**, 2 Major, 3 Minor. All nine valid. **All nine applied.** + +Reviewer: `mcp-codex-dev@1.0.1`, `codex-cli 0.144.5`, model `gpt-5.6-sol`. Pass file valid: +10 lines, 9 finding lines, terminator exact. + +Plan passes: **18, 11, 9.** Blockers: **2, 1, 4.** The count falls while the severity does not — the +three-pass floor is met and the pass is not clean, so the loop continues. + +## The four Blockers + +**1 — the plan claimed a spec §8 record that did not exist, and contradicted §6.** Pass 1's +finding 6 offered two exits; the plan half of one was applied and the spec half was not. §6 said the +executable form is *"reviewed by Gate B against the real diff"*, the plan said the opposite, and the +plan's own precedence rule makes §6 govern outside the four held items — so the plan was +self-defeating. Applied in the spec: §6 now says the executable form is **supplied to the Gate-B +reviewer alongside the diff** and is not itself inside the reviewed range, and §8 gains the residual +in full — the checks are never fingerprinted, the `battery+check` evidence is produced by +unfingerprinted code, and a later reader cannot recover which version of a check produced a result. + +**2 and 3 — the validation procedure executed the operation the convention forbids, in the file the +convention is being added to.** `C1b`'s counter-check changed a character of the real 2026-07-20 +row; `C3`'s moved the real block, and appended and removed a complete entry. §2.1 protects a row the +moment it exists, committed or not, and §2.2 protects a complete entry the same way — neither waits +for a commit. Applied as a plan-wide rule with a stated boundary: **no counter-check may mutate a +row or a complete entry in the real ledger**; every such mutation runs against a scratch copy with +the check pointed at it. Mutations to *prose* — a header paragraph, an anchored clause — are covered +by neither rule and stay in place, which is why `C1c` and `C2a` are unaffected. + +**4 — the consolidation reopened the pass-2 staging blocker.** `git reset --soft $BASE` leaves +whatever accumulated since `BASE` staged, which is not the same as the intended set, and Task 8 +declares no modified files, so the mandated exact-set assertion had nothing to compare against. +Applied: the five tracked paths are named in the step — `docs/hardening-log.md`, +`plugins/dev-workflow/commands/workflow-init.md`, +`plugins/dev-workflow/.claude-plugin/plugin.json`, `plugins/dev-workflow/CHANGELOG.md`, `todos.md` — +re-staged explicitly after the reset, with the cached set asserted equal to that list. + +## The two Majors + +**`ledger-unreadable` was routed nowhere.** Its expected cell said "every label that reads it, `C3` +included", so Task 4's `C1a`/`C1d` run excluded it and Task 5 checked only `C3` — leaving the +read-failure oracle unexercised for six labels. Now the cell names all seven, Task 3 gained a step +running it for `C2a`/`C2b`, Task 4's run widened to `C1a`–`C1d`, and Task 5 keeps `C3`. + +**The matrix declared outcomes "under both `sh` and `dash`" and no step required both.** A +shell-specific parser defect would have survived behind a happy path that passes twice. All three +matrix-running steps now require both shells. + +## The three Minors + +The `grep -F` rationale was **factually wrong in the direction that matters**: fixed-string grep +treats each line of a multi-line pattern as a separate pattern, so it never establishes the +contiguous sequence and can exit 0 on a component line — a false *positive*, not the "clean absence" +the plan claimed, which would have had the executor self-test the wrong failure mode. The anchor +data claim said **two** anchors begin with `-`; exactly one does, and five contain backticks — the +error was in the spec's §6 oracle as well and both are corrected. `C1e`'s rationale carried a +totality claim ("can only be made to fail by mutating the table by hand") when an unreadable ledger +fails it too; narrowed to the spec's own wording. + +## Sweep after applying + +Step numbering contiguous in all eight tasks (Task 3's insertion renumbered to eight steps); the +false "two begin with `-`" claim gone from both artifacts; the §8 harness residual present; the +scratch-copy rule stated globally and at both counter-check sites; 35 anchors byte-identical; 12 +fence lines, balanced. diff --git a/.context/codex-reviews/gate-a-plan-pass-3.2026-07-30-classifier-cycle.md b/.context/codex-reviews/gate-a-plan-pass-3.2026-07-30-classifier-cycle.md new file mode 100644 index 0000000..3732fcc --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pass-3.2026-07-30-classifier-cycle.md @@ -0,0 +1,36 @@ +BLOCKER | high | Task 5 Step 2 and Task 2 `class_of` | the plan says every locator row is ported through the real hook, but malformed documents do not route when `jq` is present and most minimal `verify.sh` payloads have no `hook_event_name` or gate `tool_name` at all | those rows produce no class or `DISCARDED-UNIDENTIFIED`, so the merged Task 5 cannot make its advertised table pass and would test a contract the spec says unroutable JSON never reaches | split locator-level tests from hook-level tests; exercise raw locator status directly for unroutable documents, add complete routing fields to routable cases, and assert the documented jq versus jq-free routing difference +MAJOR | high | Task 5 A1-A2 and `.context/plan-drafts/locate.awk:25-47` | `skipval` validates only matching container closers and does not parse members or elements recursively, while `readstr` accepts invalid escapes and raw control bytes; for example balanced `[1,]`, `{"a" 1}`, or `"\q"` before a real success block still returns that block with status 0 | the plan's claims that `skipval` consumes exactly one JSON value, that the outer path is fully validated, and that malformed routable structures refuse are stronger than the code, allowing malformed input to become trusted `success` | implement recursive JSON value/member grammar and valid string-escape checks, or narrow the settled contract and claims and add success-bearing counterexamples for every accepted malformed form +MAJOR | high | Task 5 A2 and pass-2 disposition 8 | escaped spellings do not merely fail safely: a semantic duplicate such as one `\u0074ool_response` plus one literal `tool_response`, or escaped plus literal `type` or `text`, bypasses duplicate detection and can select a trusted success block | this refutes the disposition's “never a wrong verdict” rationale and reopens the ambiguity-to-success false-checkmark path at the external-payload boundary | conservatively refuse relevant escaped keys, decode the bounded ASCII key grammar, or treat any escaped key coexisting with a classification-relevant literal as uncertainty; test both orders for all three keys +MINOR | high | Task 5 A2 and `.context/plan-drafts/locate.awk:95-101` | the locator compares the raw encoded `type` bytes to `text`, so the valid JSON value `te\u0078t` is not recognized even though the spec selects the first element whose string value is exactly `text` | a conforming mapped tool can be misclassified as `no-result` and have a genuine pass discarded, contradicting the settled block-selection contract | either decode the narrowly bounded `type` value, explicitly amend the contract to raw literal spelling, and add a compatibility fixture for the chosen behavior +MAJOR | high | risk high plus security standard abuse path, Task 5 A2 ceiling and current hook line 17 | the 1 MiB ceiling does not prevent unbounded payload memory or input work: the shell has already stored all of `payload`, `awk` must still ingest the full input, and a normal one-line MCP payload places the whole oversized record in `$0`; the bound also measures implementation/locale-dependent `awk` length and rejects an exact no-newline 1 MiB input because it adds a synthetic newline | a malicious or pathological mapped server can still impose unbounded memory and read work while the plan claims the trust-boundary risk is answered, and legitimate large prompts can trigger a misleading unrecognized disclosure | describe it only as a state-machine scan ceiling, or enforce a byte limit before whole-payload storage under `LC_ALL=C` with exact boundary tests and an explicit oversize diagnosis +MAJOR | high | Task 3 Step 4 | removing `sed` from the jq-free PATH prevents fallback `field()` from routing the unknown-tool event before `emit` is reached, so no output and no marker occur for the wrong reason | the proposed encoder-failure test can pass vacuously while `emit` still converts a failed encoder into an empty successful document and burns one-shots | use a selective `sed` shim that succeeds for routing and fails only the encoder substitutions, or expose a test-only writer path that demonstrably invokes `emit`, then assert the emitted status consequence +BLOCKER | high | Task 2 Step 6, Task 5 Step 12, and Self-Review “Type consistency” | `run_scenario`, `expected_ctx`, and `expected_msg` are promised in prose but no executable definitions or branch cleanup are provided, despite the disposition claiming exact shell and Task 5 invoking them | implementation still requires inventing load-bearing state setup and golden text, and the suite aborts or stays red on undefined commands under the stated step order | provide complete POSIX-shell definitions for every scenario and expected pair, including deterministic setup and cleanup, before the first Task 4 use +MAJOR | high | Task 1 Steps 3 and 7 | both provenance checks fail open: Step 3 ends a nonzero `rc` with `echo STOP`, returning success, while Step 7 only prints `DRIFT`; moreover Step 7 compares only the selected text although Step 6 and the README claim the entire response array slice is byte-identical | destroyed or drifted fixtures can pass the task and battery while the shipped provenance claim remains categorical | make every mismatch or extraction error exit nonzero, compare the raw bounded array slice if array identity is claimed, and stage a permanent failing test in the task that owns the test-file edit +MAJOR | high | spec §7.3 versus Task 5 Steps 3 and 11 | `seeded_preserves` is hard-wired to the normal `run`, and the marker-row repetitions do not cover `failure` or `no-result` gate-pass effects through the jq-free field reader | invariant 4's required same behavior without `jq` can regress for two discarded classes while all listed jq-free checks stay green | parameterize the clean/seeded state matrix over normal and `nojq_run` for every discarded class and both default and mapped gate names +MAJOR | high | Task 5 Step 7 | the missing-`awk` test asserts only a disclosure marker and final exit 0; it never asserts Gate-A count, Gate-B count, or fingerprint storage, and it does not cover an `awk` that exists but exits nonzero or emits partial output | an implementation can disclose uncertainty yet fail closed on pass state, violating the explicit counted-and-disclosed contract without failing this test | assert count and usable fingerprint for review, `passCountA` for exec, plus disclosure and exit 0 for missing, nonzero, and partial-output `awk` faults +MAJOR | high | Task 5 A5 and Step 11 | the claimed one-test-per-row matrix omits pending retention under suppression and closed stdout, shown-marker failure while flushing existing pending, pending-delete failure and retry, and a successful emit followed by `bgAdvice` marker-write failure | the exact debt-loss and duplicate-delivery edges A5 deferred to the plan remain unpinned, so pass-2 disposition 16 is incomplete | add selective fault injection and post-state assertions for each omitted transition in both emitter modes, including the reachable shown-plus-pending retry state +MAJOR | high | Task 5 Step 11 row 10 and accepted residual C2 | `chmod 500 .context` makes the pass counter and fingerprint writes fail while stdout still writes successfully to `/dev/null`; it does not create a counted call whose disclosure is neither delivered nor persisted | the test cannot establish C2's accepted silent-count residual and may report coverage while no count was recorded and the disclosure writer succeeded | leave pass-state paths writable, suppress or fail the emit, make only `unverifiedPending` unwritable, then assert the count/fingerprint advanced, neither disclosure marker exists, and the hook exits 0 +MINOR | high | Task 5 A5 row “any marker write fails” | the table says C2 applies to any marker-write failure, but C2 applies only when disclosure delivery and pending persistence both fail; failed `bgAdvice`, a failed shown write followed by a successful pending write, and a failed pending delete have different outcomes | the residual is overstated and obscures which failure combination actually permits a silent counted pass | split the row by marker and delivery outcome and cite C2 only on the undelivered-plus-unpersisted disclosure transition +MINOR | high | Task 5 Step 3 `seeded_preserves` | the comment and test label say a discarded call creates “nothing at all” and “writes no state,” but a backgrounded call intentionally creates `codex-gate.bgAdvice`; the assertion checks only the four pass-state files | the test claims more than it asserts and contradicts diagnostic-state behavior settled in spec §5.2 | rename the comment and labels to “no gate-pass state” and separately assert the expected diagnostic marker effects +MAJOR | high | spec §7.3 and Task 5 Steps 11-12 | the repeated jq-free marker rows emit only one note at a time or an earlier disclosure alone, while the actual two-message composition matrix in Step 12 runs only through the normal emitter | the fallback escaper is never tested on the separator-bearing composed document that the plan says stresses it hardest | repeat at least one real two-message branch through `nojq_run`, compare both complete fields, and parse the captured result with an independent JSON parser after restoring normal PATH +MAJOR | high | Task 5 Step 1 message pairs versus spec §6 “Field split” | operator checks and remedies for `no-result` and `unrecognized`, including changing process-start environment and pinning a server, are placed in model-facing `additionalContext`, while their user-visible `systemMessage` fields contain no action beyond “see the note” | the operator who can perform those fixes may never receive them, directly contradicting the settled field ownership | keep gate consequences in context and move or duplicate every operator-only check and remedy into the corresponding `systemMessage`, then update exact goldens +MAJOR | high | Task 5 Step 1 and Step 12 prompt-conformance claim | the literal model-facing messages do not name Claude via Claude Code as their target, the long diagnostics are undelimited prose rather than context/task/rules sections, and retrying/backgrounded paths lack a bounded stop or escalation rule | the strings do not verifiably satisfy prompt-standards items 1, 3, and 5 despite the categorical all-12 claim, and Codex is mentioned prominently enough to make the executor ambiguous | rewrite the model-facing strings with an explicit target and compact labeled structure, and give each retry-capable state a stop/escalation condition tied to CLAUDE.md §5 +MAJOR | high | Task 5 `FAILURE_CTX` | the diagnostic says only that the call “reported failure” and to retry once “the cause is fixed,” but the class covers every immediate-first false envelope and supplies no cause list, distinguishing check, or cause-specific fix | prompt-standard item 10 is not met and the agent cannot know what “fixed” means or when the shared recovery attempt is spent | direct Claude to inspect the tool result's exact error code/message, distinguish known execution failure from timeout and unknown codes, apply the matching remedy, and stop per the one-attempt recovery rule if it persists +MAJOR | high | Task 5 `NORESULT_CTX` | `.mcp.json` is described as the place to check the effective MCP registration, even though scope precedence can make a different same-named server effective, and neither of the two named causes receives a complete distinct remediation | this repeats the diagnostic failure that motivated prompt-standard item 10 and can send an operator around a restart/config loop that never changes the active server | name the command or inspection that reveals Claude Code's effective registration, compare it with `.mcp.json` and `.context/codex-gate.tools`, and give separate hook-contract reporting and third-party-tool replacement/unmapping fixes +MAJOR | high | Task 5 `UNVERIFIED_CTX` | the supposedly known-cause list omits deliberate oversize refusal, missing/failing `awk`, and malformed-structure refusal, and its backgrounding remedy points to “the discarded-pass message” even though a reworded notice never produces that message | distinct known causes with different fixes collapse into one state, and the one residual that most needs the environment setting can withhold the actual setting instructions | enumerate these known causes with checks and fixes, and inline the launch-environment, restart, zero, positive-threshold, and minimum-version guidance for a suspected reworded notice +MAJOR | high | Task 5 `BG_SHORT_MSG` and settled decision 4 | the short user-visible form does not name `CLAUDE_CODE_MCP_AUTO_BACKGROUND_MS`; only model-facing context names it | because `bgAdvice` outlives the session that showed the long form, a later operator receives no actionable pointer, exactly the case decision 4 required the short form to cover | include the variable name and concise action in the short `systemMessage` and update its golden +MAJOR | high | risk high concurrency and observability lens, Task 5 `BG_LONG_CTX` | the message admits the backgrounded review may still be running but gives no instruction to stop or await it before retrying; that original call can later write the same findings slot after the retry deletes or rewrites it | two concurrent writers can leave a validly terminated stale findings file that passes the file protocol, defeating the sequential-pass assumption and making the hook's new diagnosis create a review-artifact race | instruct Claude to stop the task id from the tool result or wait for completion before deleting/retrying the slot, confirm the task is no longer active, and test/document the late-writer path +MINOR | high | Task 5 `UNVERIFIED_CTX` and Step 8 write order | “has counted at least one gate call” is emitted before best-effort counter and fingerprint writes are known to have succeeded | an unwritable pass-state path can produce a categorical recorded-state claim that is false, violating prompt-standard item 11 even though the conservative user behavior is harmless | say the hook classified the call as countable and attempted to record it, or make the wording conditional on observable state-write success +MAJOR | high | Task 6 Steps 2, 4, 7, 8 and Self-Review “Placeholders” | Task 6 changes the repo CLAUDE prompt, the inline CLAUDE template, and the unknown-tool hook prompt but provides no literal replacement strings or blocks for Gate A to review; exact goldens are deferred to implementation | invariant 11 cannot be checked against the actual shipped prompts, and implementation must invent wording after this plan's Gate A | include every final replacement prompt literally in the plan and review it against all 12 items before execution, while keeping authoritative prose references out of non-self-contained copies +MAJOR | medium | Task 6 census and B1 accounting | the patterns and explicit B1 list do not account for `commands/workflow-init.md` saying “counts passes by TOOL NAME” around line 250 or README's “counters key on those two tool names” setup claim | one large rewrite can leave another shipped prompt teaching the old name-only mechanism, repeating the exact dropped-condition and docs-drift class AGENTS.md warns about | broaden the claim-oriented census to include “by tool name” and “key on those tool names,” list every resulting site, and mark each kept, narrowed, or replaced +MINOR | high | Task 6 Step 5 and AGENTS.md manifest-claim Don't | Task 6 edits both AGENTS.md and `docs/architecture.md` but runs only a `codex-gate` grep; it omits the mandated broad declaration/convention-loading census and a same-change read of the plugin manifest | a nearby layout edit can preserve or introduce the duplicate-hooks belief that previously broke plugin loading, with no mechanical checker for prose accuracy | run the exact declaration/convention-loading grep from AGENTS.md, read `plugins/dev-workflow/.claude-plugin/plugin.json`, and disposition every hit before editing either tree +BLOCKER | high | Task 7 Step 2 counterfactual | the command deletes `$t` and therefore `old-fails.txt` before the next paragraph requires exact-label assertions and verbatim evidence from that file | the required `+check` counterfactual cannot be validated or recorded in the stated order | assert and copy the four named lines before cleanup, retain them in a durable evidence variable/file, and use a trap to remove the temp directory afterward +MAJOR | high | Task 7 Step 2 base-blob guard | a wrong or failed `git show` reaches `echo "STOP: wrong base blob"`, whose zero status lets the procedure continue against an empty or wrong hook | the evidence can attribute failures to the pre-change implementation without having materialized that implementation | make either `git show` or `cmp` failure terminate the step nonzero before running the suite +BLOCKER | high | Task 7 Steps 1 and 6 | Task 7 edits `CHANGELOG.md` but has no staging or WIP-commit step; Gate B reviews only `baseSha..HEAD`, and the closing amend is told neither to stage nor include the changelog | the final pass excludes a plugin file that the cycle then intends to ship, and the changelog may remain uncommitted or enter the amend without review | stage Task 7's edit in an explicitly named WIP snapshot, review the recorded pre-Task-1 base through that snapshot, then close by amend with the validated evidence body +BLOCKER | high | Task 7 Step 6 in the observed `main` checkout | “baseSha is the merge-base with main” resolves to current HEAD when the tasks are executed directly on `main`, not to the parent SHA recorded in Task 1 | Gate B receives an empty range and can return clean without seeing any implementation, so pass-2 disposition 25 is incomplete for the actual working state | define `BASE` as the recorded pre-Task-1 commit unconditionally and pass that exact SHA; use a live merge-base only after confirming work is on a distinct branch and equality with `BASE` +MINOR | high | Global “battery, per commit” versus Tasks 2-6 | the plan says the exact pre-commit command list is copied into every task because references were insufficient, but Tasks 2-6 say only “as in Task 1 Step 9” and show staging/commit commands | the audit claim is false and an executor following only each code block can omit the promised checks | either paste the full pre-commit block into every task as claimed or define one named reusable block and explicitly require its successful output before each commit +MAJOR | high | security standard assets/trust-boundary lens, Task 1 Step 4 | capturing the review fixture by the INDEX method mutates the globally installed cached hook to dump complete future payloads, but the task adds no backup hash, trap, interruption recovery, or proof of restoration | interruption can leave a machine-wide hook recording private prompts, paths, and review content beyond this synthetic capture | capture through an isolated plugin/session if possible; otherwise back up the exact installed file, install a trap before mutation, restore on every exit, and verify the final hash before staging the fixture +MINOR | medium | risk high rollback lens, Global Constraints and Task 7 | the plan gives no rollback procedure if the load-bearing classifier discards legitimate passes on a supported machine, and `.context/codex-gate.off` is not a rollback because classification and state tracking continue while messages are suppressed | operators can be left with permanently non-advancing counters and no documented safe path back while 0.8.0 is investigated | name and verify a recoverable rollback to the prior plugin version or hook, including how to preserve/clear only diagnostic state and how to confirm restored counting +MAJOR | medium | Plan provenance, Task 5 Steps 5-6, and Self-Review “Placeholders” | the actual locator and matcher are represented by placeholders that depend on ignored `.context/plan-drafts` files, even though the plan calls itself implementation-complete and may be executed from a fresh clone or later session where those files do not exist | the central load-bearing code and its provenance can disappear while the reviewed plan remains, blocking execution or inviting an unreviewed reimplementation | embed the verified bodies in the plan or move them to a tracked reviewable source before Gate-A closure, and verify their hashes when installing them +MINOR | high | Task 5 Step 1 and Self-Review | the plan repeatedly says “all six message pairs” are literal, but the code declares five pairs, ten strings: failure, no-result, background long, background short, and unverified | the prompt census and claimed all-12 review scope are numerically unreliable, making it unclear whether a required pending/composed form is missing or merely miscounted | correct the inventory to five base pairs plus derived pending/composed forms, or add and name the genuinely missing sixth pair before prompt review +END OF FINDINGS (35 total) diff --git a/.context/codex-reviews/gate-a-plan-pass-3.2026-08-04-hardening-round.md b/.context/codex-reviews/gate-a-plan-pass-3.2026-08-04-hardening-round.md new file mode 100644 index 0000000..ececd43 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pass-3.2026-08-04-hardening-round.md @@ -0,0 +1,20 @@ +BLOCKER | high | Task 9 Step 9 | the plan never creates or requires a feature branch, and the reviewed checkout is currently `main`, so `git push -u origin HEAD` pushes the seven-ahead local main directly to `origin/main` before `gh pr create` | the PR-only version-bump gate is bypassed and PR creation then has no head branch to open | add a pre-edit branch preflight that stops on the default branch or creates a named feature branch, and verify the PR head/base before pushing +BLOCKER | high | Task 9 Steps 8-9 | the amend and PR-create fences read `$EVIDENCE` in fresh shells without defining it, despite the global self-contained-block rule | `cat ""` fails inside command substitution while `git commit --amend` or `gh pr create` can still succeed, silently omitting the required evidence from the commit and PR | recompute and validate `EVIDENCE` inside each mutating fence, or combine each guard and consumer in one block whose final status includes the read +MAJOR | high | Task 9 Steps 7-8 | Gate-B fixes are told to be applied and re-reviewed, but no step stages them before the closing `git commit --amend` | reviewed fixes can remain only in the worktree while the final commit preserves the pre-fix index | after every fix stage only the intended paths, verify the staged and unstaged sets, and immediately before amend prove the index contains the reviewed snapshot and no intended fix remains unstaged +MAJOR | high | Task 9 Step 3 | the derived hook-count pipeline returns `grep`'s status and never captures the hook suite's status | a suite that exits nonzero after printing any `ok` lines is recorded as a valid partial assertion count, recreating `verification-masks-failure` | capture the suite output in a per-run temporary file, save its exit status, count only after completion, and fail unless both the suite status and count command succeed +MAJOR | high | Task 9 Step 2 inspection block | failures from `git rev-parse "$first_wip"^` or `git log "$first_wip"^..HEAD` are followed by unconditional `exit 0` | the safety inspection can report success without resolving the reset target or showing the commits a destructive reset will rewrite | resolve and validate the parent into a variable, make the range log failure fatal, and remove the unconditional success exit +MAJOR | high | Task 9 Step 2 squash | uniqueness of the first subject is the only range guard; the plan does not require the commits through `HEAD` to be exactly the eight expected WIP commits or require a clean index/worktree | unrelated later commits or staged changes are folded into the squash and recommitted | compare the exact ordered subject list and count to Tasks 1-8, require `HEAD` to be the expected last WIP, and stop on any staged, unstaged, or untracked in-scope change before reset +MAJOR | high | Tasks 1-8 commit steps | each task stages named paths and commits without first checking for pre-existing staged changes or verifying the final staged set | unrelated user work already in the index is swept into a WIP commit and later into the squash | before every `git add`, require an empty index or record the permitted staged set; after staging, compare `git diff --cached --name-only` to the task's exact path list +MAJOR | high | Task 4 Step 2 and Task 5 Steps 1-3 | the atomic-install fence contains only the comment `write the full story text ... here`; executing it verbatim never creates `$tmp`, and the prose below does not specify an executable hand-off into the live shell variable | absent targets fail with a misleading create-failed diagnosis, while identical existing files are misclassified as different | provide a complete quoted heredoc for each story inside its install block, then validate the temporary file before comparing or linking it +MAJOR | medium | Task 4 Step 2 and Task 5 Steps 1-3 | the temporary pathname `.tmp-$$-story.md` is neither freshly allocated nor checked for an existing file or symlink before the unspecified write | PID reuse or a planted/stale symlink can cause the story write to overwrite another path, and cleanup failures can leave misleading state | use a same-directory `mktemp` result with a cleanup trap, reject non-regular temporary state, and make write and cleanup failures explicit +MAJOR | high | Task 6 idempotency guard | after flattening all of `todos.md` to one line, `grep -c` can only return zero or one, so the documented `n != 1` duplicate branch is unreachable; a read error is also converted to zero by `|| true` and reported as ABSENT | duplicate or unreadable state is accepted as safe to mutate, defeating all three promised idempotency branches | first require a readable regular `todos.md`, then count matches with `grep -oE ...` plus a separate numeric count or an `awk` matcher that counts occurrences +MAJOR | high | Task 6 Steps 1-7 | every guard depends on an `INTENDED` file, but no step gives that file a path or a command that writes the supplied full block into it | on resume, the present-once branch cannot distinguish an identical completed edit from a partial or divergent edit as promised | make each guard self-contained with a uniquely allocated intended-text file populated by the exact step text and cleaned safely +MAJOR | high | Task 6 Step 8 | `/tmp/cited-stories.txt` is a fixed shared pathname opened with truncation and later deleted | concurrent runs can corrupt each other's verification, and a pre-existing symlink can redirect the write into another user-writable file | allocate the file with `mktemp`, install a trap, validate creation, and use that private path throughout +MAJOR | high | Task 2 Step 3 | `tr` flattens `AGENTS.md` to one line before `grep -c`, so two or more copies of the operative phrase still produce `n=1` | the check explicitly claiming exactly one operative copy passes duplicates | emit each match with `grep -oF` before counting, or count non-overlapping occurrences with a tool whose unit is matches rather than lines +MINOR | high | Task 4 Step 3 and Task 5 Step 4 | the criterion check also flattens each story before `grep -c`, and it checks only presence rather than that the matching criterion is first | duplicate copies pass as one, and moving the profile-confirmation criterion later still passes the shape check | count emitted matches and parse the first checklist item under `## 3. Acceptance criteria` explicitly +MAJOR | high | Task 9 Step 1 | `cat > "$EVIDENCE"` silently replaces any evidence from a resumed run, and the final `-s` test accepts a partial write as long as one byte landed | rerunning the task destroys durable validation work, while disk or write failure can leave a truncated template reported as created | classify the evidence path first, preserve or compare existing content, write a complete template to a private temporary file, validate its structure, and install atomically +MAJOR | high | Task 7 Step 6 versus spec §6.4 and §12 | the plan adds a post-append exact-count observation that detects a same-fingerprint row landing after Step 1 but before Step 6, while the approved spec says any row landing after the re-read is not detected and that no observation point exists past it | the plan and settled design make incompatible claims about the concurrency window, violating the rule against overstating what a gate or check compares | update the spec to name both count observations and the remaining race after Step 6, or remove the added observation and retain the approved limitation +MINOR | high | Global Constraints, Task 1 Step 8, and Self-Review | the plan says deferring the bump makes intervening local battery runs fail and that Task 1 passes because the bump landed, then correctly notes that on current `main` the merge-base is `HEAD` and the version check passes trivially | the battery is presented as enforcing an ordering it cannot observe in this checkout, an overclaim against the gate-claims Don't | state that ordering is required for PR or non-main comparisons, that the local main invocation proves nothing about the bump, and that Step 7 is the only current-tree version assertion +MINOR | medium | Task 9 Step 8 ` · supersedes ` does | this is a mechanically false count in the data-handling rationale and the Task 1 sweep does not catch it | change the claim to one anchor or add the missing second adversarial anchor if two are required +MINOR | high | The checks — labels, properties, oracles | the C1e rationale says the check can only be made to fail by mutating the table by hand, although an unreadable or missing ledger is another failure path for a correct chronology check | the totality claim overstates the reason C1e lacks change-specific evidence | use the spec's narrower wording that C1e's falsifying observation requires a table mutation and that this diff cannot produce it +END OF FINDINGS (9 total) diff --git a/.context/codex-reviews/gate-a-plan-pass-3.md b/.context/codex-reviews/gate-a-plan-pass-3.md new file mode 100644 index 0000000..bec1863 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pass-3.md @@ -0,0 +1,18 @@ +BLOCKER | high | Task 1 Steps 1–2, lines 98–104 and 173–187 | Both template fixtures make `# Command` the first content line in the four-backtick fence, while the proposed parser recognizes a template only when that line is exactly `# ` | The initializer makes every pre-existing fixture fail 4c with NOSTART, and even the three named 4c accept cases reject, so the planned suite cannot become green | Change both `init_prompt_fixtures` and `sev_tpl` to reproduce the real fence opening and first content line: ````markdown followed by `# ` +MAJOR | high | Task 1 Step 7, lines 403–419 versus approved spec §3 lines 322–337 | The proposed tolerant-reader paragraph is newly worded rather than the byte-for-byte rider (b) text settled in spec §3; it changes the opening rule, collapses the explicit structural-failure inventory, and adds record/incident prose from other spec paragraphs | The product is prompts, so semantically similar prose is still a changed interface and violates the stated requirement that §3 supplies the shipped text byte-for-byte | Paste the approved rider (b) shipped block exactly, or amend and re-approve the spec before using different wording +MAJOR | high | Task 1 Step 4, lines 301–370, and Step 2 fixture inventory | `MULTIEND` is unreachable: heading mode increments `ends` only for the first heading while `inregion`, and fence mode increments it only when closing an active region; after either event `inregion` is zero. There is also no `0:MULTIEND` case branch and no two-end fixture | Approved spec §5.2 explicitly requires rejection coverage for two candidate end boundaries, while the implementation silently routes the nominal code through the generic wrong-count diagnostic and the plan falsely says every abnormality has a named branch | Define a reachable ambiguity condition for each mode, add the required two-end fixture against the real file shapes, and handle `0:MULTIEND` explicitly; otherwise remove the code only after reconciling the approved spec +MAJOR | high | Task 1 Step 2, lines 217–223 and 255–269 | The literal unterminated-template fixture has neither the required `# ` anchor nor the checklist that 4b reads, although later prose instructs the implementer to add a reachable checklist without giving the corrected body | In the red phase it rejects for 4b or is classified NOSTART, not NOEND, so it can pass for a reason other than its name and contradicts Step 3's required `(exited 0)` failure | Replace the shown fixture with one complete literal using the real anchor and an external reachable checklist, then assert the NOEND-specific diagnostic +MAJOR | high | Task 1 Step 2, lines 225 and 255–269 | The missing-template case is left as an unresolved choice: keep a non-isolated test or drop it, while Step 3 simultaneously requires every reject fixture to fail only because 4c is absent | Before 4c, deleting the template already makes 4b fail, so this case cannot produce the promised `(exited 0)` red-phase result and the plan is not executable as written | Choose one settled route in the plan: omit this redundant fixture, or give it a specialized harness/red-phase assertion that proves the 4c diagnostic becomes newly present without claiming diagnostic isolation +MAJOR | high | Global Constraints lines 54–57, Task 1 Step 11 lines 507–514, and Task 5 Step 3 lines 758–767 | `WIP_PARENT` is assigned in one shell command block and consumed several tasks later, but shell variables do not persist across ordinary agent tool calls or resumed execution sessions | The version check will receive an empty or unset argument, and Gate B can lose the exact base this sequencing was introduced to preserve | Make every consumer self-contained, for example recompute and validate `WIP_PARENT=$(git rev-parse HEAD^)` from the still-WIP snapshot immediately before the version check and Gate-B call, or persist the SHA in an explicitly named durable working note +MAJOR | high | Gate B fix loop, lines 797–805 | After a Gate-B fix the plan reruns the full battery against `main` and the counterfactual, but does not rerun `scripts/check-version-bump.sh "$WIP_PARENT"`; on the current `main` branch the battery's `main` comparison is the known empty range | A fix that changes plugin content or accidentally loses the manifest bump can be re-reviewed and closed without any deciding invariant-12 check after the diff changed | Add the explicit WIP-parent version check, with a non-empty-range guard, to every fix/re-review iteration and once immediately before the closing amend +MAJOR | high | Global parity rule lines 22–30 and Task 2 Step 3 lines 600–608 | The global rule and story AC require the newly written regions to match, but Task 2 permits them to differ by unspecified “template list indentation”; the real inline §5 template is flush-left | This creates an undefined human normalization and can certify divergent shipped prompt bytes despite the settled parity decision | Require an exact diff of the inserted bytes in their current flush-left shapes; if structural indentation truly must differ, specify the exact extraction/normalization and reconcile that change with the approved spec +MAJOR | medium | Task 2 Step 1 lines 542–598 | The exact block is a top-level blockquote beginning with `>`, but the placement instruction says to add it “as a list item” inside Mechanics, whose existing entries begin with `-`; those are different Markdown structures | An implementer must choose between byte-for-byte text and the stated list placement, so the rendered prompt structure is not decidable from the plan | State explicitly that the approved blockquote is inserted between Mechanics list items, or provide the exact list-indented bytes and amend the spec; do not describe both shapes +MAJOR | medium | Gate-B counterfactual, lines 813–826 | The command sequence does not guard `mktemp`, the tracked-file copy, either `git show`, baseline exit 0, mutant exit 1, or diagnostic exclusivity, and it has no cleanup trap; semicolons allow later steps to run after setup failure | A void or multi-cause run can still reach the end and be recorded as the required discriminating evidence, contrary to the plan's own proof-calibration rules | Turn each requirement into a checked branch that exits nonzero on failure, validate exactly one 4c diagnostic and no others, guard the temp path, and install cleanup before copying +MINOR | high | Task 1 Step 10, lines 466–497 | `flips=$(diff ... \| grep -c '^[<>]')` counts both removed and added output lines rather than flipped assertions, and the pipeline masks a failing `diff`; this silently changes the unit behind the existing 20/22 assertion counts | The checked-in mutation record can acquire doubled or otherwise invalid numbers while being presented as comparable evidence | Capture and status-check `diff` separately, extract one canonical side such as `^> FAIL`, and record assertion names/counts in the same unit as the existing evidence +MINOR | high | Task 1 Steps 8–11, lines 421–515 | The only full battery in Task 1 runs before the AGENTS inventory edit, mutation-evidence/header rewrite, manifest bump, and WIP commit; Step 11 runs only two shellcheck commands before declaring the snapshot complete | The first committed plugin-changing intermediate state is not actually verified by the quality command the plan globally requires, so a validation or invariant regression introduced late in Task 1 is deferred to a later task | Run the complete AGENTS quality battery after all Task 1 edits and after creating the WIP snapshot, using the deciding WIP-parent version check where the battery's `main` argument is empty +MINOR | high | Self-review lines 868–875 | The self-review says the parser returns only count/NOSTART/MULTISTART/NOEND and that the awk-failure branch is untested, contradicting the proposed `MULTIEND` return and the `inject_case` parser-failure fixture earlier in the same plan | The plan's final consistency audit preserves stale pass-2 claims and can mislead the implementer into dropping coverage that was just added | Include `MULTIEND` in the return inventory and state that parser failure is exercised through the marker-based PATH injection, then describe only any genuinely untested branch +MINOR | high | Self-review path table lines 848–856 | The mapping still says `plugins/dev-workflow/.claude-plugin/plugin.json` belongs to Task 5 even though pass 2 moved the bump and Task 1's file list and Step 11 now perform it | The plan's own coverage ledger contradicts the sequencing correction whose purpose is to keep intermediate commits green | Change the manifest mapping to Task 1 and leave only CHANGELOG in Task 5 +MINOR | high | Self-review line 861 versus approved spec §6 lines 562–605 | The plan claims §6 has nine items, but the approved backlog section contains ten top-level bullets, including the version/CHANGELOG bullet handled across Tasks 1 and 5 | An exact coverage claim with the wrong cardinality weakens the audit and can hide a dropped row in a later revision | Enumerate all ten §6 bullets and map each to Task 4 or the version/CHANGELOG tasks instead of relying on the stale count +MINOR | medium | Task 2 Step 4 lines 610–614 | “The battery is confirming that the edit broke nothing else” outruns what the battery compares, immediately after admitting that nothing mechanical covers the paragraph | This is the gate-proof overclaim class AGENTS.md warns about: green fixed checks cannot establish arbitrary absence of prompt regressions | Say only that the named lint, suites, invariant checks, version check, and plugin validation passed; leave semantic prompt safety to the required 12-item review and Gate B +NIT | high | Task 1 Step 2 lines 128–137 | The fixture API comment says “Two sentinels” and then defines three: `@KEEP@`, `@GONE@`, and `@LOCK@` | Exact-count comments are a named docs-drift class in this repository and this one is wrong in the code the plan asks to paste | Change the comment to “Three sentinels” +END OF FINDINGS (17 total) diff --git a/.context/codex-reviews/gate-a-plan-pass-4-dispositions.2026-07-30-classifier-cycle.md b/.context/codex-reviews/gate-a-plan-pass-4-dispositions.2026-07-30-classifier-cycle.md new file mode 100644 index 0000000..63953a7 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pass-4-dispositions.2026-07-30-classifier-cycle.md @@ -0,0 +1,45 @@ +# Gate A (plan) — pass 4 dispositions + +Advisory companion to `gate-a-plan-pass-4.md` (3 BLOCKER, 21 MAJOR, 5 MINOR = 29). One line +per finding, in file order. All accepted; one partly dismissed with a reason. + +**Rider 3 checked, and my call recorded.** The human's third rider was: *"If pass 4 raises +claim-vs-code on this same spot again after the narrowing and fixtures, that is the stuck +signal — stop and surface rather than iterate a fifth wording."* Findings 3 and 5 are +claim-vs-code on the locator. **I judged the rider does not fire**, and the reasoning is +here so it can be overruled: both ask me to *weaken* a sentence I failed to narrow — the A1 +prose and the embedded `skipval` comment still said "consumes exactly one JSON value", and +the drafts README still carried the withdrawn wider claims — while the rider is aimed at +re-arguing that the code should *validate more*, which pass 4 does not do. The fix was +deleting a phrase in three places, not a fifth wording of a contested claim. The trend also +does not read as stuck: Blocker+Major 36 → 28 → 27 → **24**, blockers 7 → 5 → 5 → **3**. + +1. BLOCKER `branch_msg` is a literal ellipsis and no expected constant is defined — ACCEPT. Written out in full, with the rule that every `*_EXPECTED` is a literal copy (a golden reading the hook's own variable agrees with any text the hook emits), and a note that four of them interpolate `$passes`/`$floor`/`$policy` and must be written for the counts their scenario seeds. +2. MAJOR `run_scenario` inherits state across the Task 4 matrix — ACCEPT, and the example is exact: `gateb_stale` leaves three passes, so `gateb_floor` reaches the **satisfied** branch under its own below-floor label. Added `reset_gate_state` (the four pass-state files only) at the top of `run_scenario` — it still cannot call `reset_all`, because Task 5's matrix seeds pending before invoking it. +3. MAJOR A1 and the embedded comment still say "exactly one JSON value" — ACCEPT, and this is the narrowing left incomplete outside A2. Both now say `skipval` consumes one **span**, that container members are not parsed, and that `[1,]` is consumed rather than refused. See the rider note above. +4. MAJOR Task 1's checks still read the ignored `.context/plan-drafts/locate.awk` — ACCEPT. Step 3 writes the plan's own body to a temp file; the **permanent** slice check moves to Task 5, where the locator is in the hook and the suite can drive it with no untracked reference. Shipping a permanent test that reads an ignored path is green here and broken in every clone. +5. MAJOR the drafts README still says 50 and keeps the withdrawn claims — ACCEPT, my omission when the narrowing landed. Rewritten to 53 and to the same string-boundary-only, walkable-invalid and scan-not-memory bounds A2 uses, plus what the file is *for* now that the bodies live in the plan. +6. BLOCKER the cache backup/restore is unguarded and disarms its own traps — ACCEPT. Every setup step hard-stops; one `restore_cache` function restores, verifies and only then releases the backup; signal handlers **exit** after cleanup; and the normal path no longer runs `trap -` before restoring, which is what turned a failed restore into a silent one. +7. MINOR the sed shim hard-codes `/usr/bin/sed` and has no guard — ACCEPT. The real path comes from `command -v sed` and is baked in; the build is guarded with a `skip -`; and the shim itself is verified both ways before it is trusted. +8. MAJOR `sed` is load-bearing for classification and untested — ACCEPT, and it was the wrong-direction failure: a failing `sed` yields an empty substitution, which reads as blank and classifies a genuine **success** envelope as `no-result` — fail-closed. `classify_block` now checks `sed`'s status and maps failure to `unrecognized`; verified in-session (broken `sed` → `unrecognized`, not `no-result`); and the Global Constraints declare both `awk` and `sed`. +9. MINOR the ceiling is called 1 MiB but `awk`'s `length()` counts characters — ACCEPT. Restated as a 1 Mi-**unit** scan ceiling, with the reason for not setting a locale inside the hook: the exact cut-off is a backstop, not a contract. +10. MAJOR C2 carries only one of spec §5.2's two directions — ACCEPT. Both rows are in the table now: suppressed-emit plus failed pending write, and failed-emit plus failed pending write. +11. MAJOR a current-`unrecognized` event emits the carried debt without the `Earlier:` prefix — ACCEPT as a **table** gap rather than a code bug, and settled explicitly: the prefix marks that no statement is being made about the current call, so when the current call is itself `unrecognized` the statement *is* about it and there is no prefix. New A5 row, and a scenario that would have shown the disagreement. +12. MAJOR the marker rows never actually run jq-free — ACCEPT, claimed in prose while every block called `run` directly. Parameterized on a runner pair with the loop shown, so the substitution *is* the coverage. +13. MAJOR the pending-delete test asserts nothing — ACCEPT, and the diagnosis was right twice over: a directory at the pending path is not seen by `[ -f ]` as pending at all, and the test then removed it by hand before the event it claimed did the cleaning. Replaced with a real pending file plus a non-writable `.context/`, which fails the `rm` while an existing file can still be truncated — the one fault that separates the two operations. +14. MINOR the post-fix paragraph still describes the removed `chmod` test — ACCEPT. Rewritten to describe the selective injection actually used, and it now records **both** blunt approaches that failed and why, rather than only the older one. +15. MAJOR no delivered string names its target model — ACCEPT the observation, **partly dismiss the remedy**. The declaration is made once in a header comment above the strings, naming each field's consumer; repeating it inside five injected diagnostics restates what the API channel already fixes. If a reviewer holds item 1 requires it in the delivered text, that is a real disagreement about the checklist's scope for injected diagnostics and belongs in `docs/prompt-standards.md`, not in boilerplate — the plan says so rather than quietly claiming conformance. +16. MAJOR item 4 unmet for the contexts that request a report — ACCEPT. `FAILURE_CTX` and `NORESULT_CTX` now give the exact one-line reporting shape. +17. MAJOR `FAILURE_CTX`'s remedy is still vague — ACCEPT, and pass-3 disposition 18 overstated what it fixed. Each error code now carries a distinct check and action, and the unknown-code case says what to do with it. +18. MAJOR `UNVERIFIED_MSG` promises "each with its check" and lists three without one; oversize is misstated — ACCEPT. Causes are split into those with a user-side fix and those without, and the ceiling is named as the **whole-payload** one, which is what the code enforces. +19. MAJOR `UNVERIFIED_MSG` still says "was counted" — ACCEPT; pass 3 fixed one field and left the other. Both now say classified-as-countable-and-attempted. +20. MAJOR "will not repeat" contradicts the plan's own failure table and C4 — ACCEPT. Reworded to normally-once, with the two conditions under which it repeats named. +21. MINOR "treat the count as a floor" inverts the meaning — ACCEPT, and it is the sharpest wording finding of the pass: this count can *overstate* completed reviews, so "floor" invited the very false checkmark the disclosure exists to prevent. Now "a mechanical tally … it can overstate them". +22. MINOR `FAILURE_MSG` carries no operator action while the structure claim says all five do — ACCEPT. The claim is narrowed to four, with the asymmetry stated: nothing an operator can do fixes a Codex call that reported failure, and inventing an action to satisfy a uniformity claim is worse. +23. MAJOR the CLAUDE replacement restores the over-broad claim — ACCEPT, and this is the dropped-condition class AGENTS.md names. The replacement now says **first** property, **the current notice wording**, and that unrecognized shapes count — with the three precisions listed as load-bearing so the next editor cannot drop them silently. +24. MAJOR the README replacement excludes `unrecognized` — ACCEPT. It contradicted settled decision 2 in the primary setup doc. Reworded to name the three skipped classes and to say an unreadable result still counts. +25. MAJOR the setup text hides C1 — ACCEPT. It now states both outcomes and why the setting, not the anchor, is the primary defence. +26. MAJOR the unknown-tool insert repeats the field-split inversion — ACCEPT. Registering an MCP server is an operator action, so the namespace check and the register-as-`codex` remedy move into `systemMessage`. +27. MAJOR the Gate-B success row uses an `exec` capture that replied "ok" — ACCEPT. Switched to `shape0-success-review`, which is the distinction Task 1 spent a synthetic-repo capture to create; using the other one would claim Gate-B continuity for content no reviewer saw. +28. BLOCKER `git reset --soft` after the clean pass invalidates the fingerprint — ACCEPT, and it is invariant 3 turning on the plan: a reset changes `git diff HEAD` and the effective index, two of the three `tree_hash` inputs, so the pass just earned is spent and Gate B correctly fires again. Closeout is `--amend` on the WIP snapshot; any squashing happens **before** the reviewed snapshot exists. +29. MAJOR the rollback verification is two comments and an `ls` — ACCEPT. Guarded commands with expected state, a hard stop if the prior copy is absent or differs from the release, a behavioural probe (0.7.1 must *count* a failure envelope), and the version-switch command left to be documented from what was actually run rather than invented. diff --git a/.context/codex-reviews/gate-a-plan-pass-4-dispositions.2026-08-04-hardening-round.md b/.context/codex-reviews/gate-a-plan-pass-4-dispositions.2026-08-04-hardening-round.md new file mode 100644 index 0000000..bc0c4b0 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pass-4-dispositions.2026-08-04-hardening-round.md @@ -0,0 +1,69 @@ +# Gate A — plan — pass 4 dispositions + +11 findings, 1 Blocker, 10 Major. Plan trend 19 → 11 → 19 → 11. All accepted; none dismissed. +One went to Daniel because it required a structural decision. + +## The Blocker + +**1 (BLOCKER)** — Task 9 told the executor to *stage* each Gate-B fix and never to amend it into +the WIP commit. `mcp__codex__review` reads a **git range**, so a staged-but-uncommitted fix is +not in it: every post-fix pass would have re-reviewed the pre-fix commit and reported clean on +unreviewed content. The staged index also makes the clean-tree precondition before the amend +unsatisfiable, so any Gate-B finding either deadlocks or closes on content nobody reviewed. + +This is the parked "Gate-B fingerprints disk; the reviewer reads history" row, reproduced inside +the fix loop. + +## The structural decision (Daniel's call) + +Eight of the eleven findings sat in the git/Gate-B/PR procedure; the hardening content has been +stable for four passes. The cause: `CLAUDE.md` §5 already specifies that protocol +authoritatively, and Task 9 **restated** it — each restatement a fresh chance to drift, which is +what `prompt-standards` item 11 warns about by name. The Blocker is a restatement that dropped +§5's amend step. + +**Decision: cite §5, do not restate it.** Task 9 now keeps only what is specific to this round — +the evidence template, the four self-test cases, the prompt-conformance scope — and points at §5 +for the squash, Gate-B loop, amend and close. One clause names the dropped step explicitly +("§5 governs, including its amend-every-fix-into-the-WIP-before-re-review requirement"), because +naming a step is not re-specifying a procedure, and that step proved droppable. + +That deletion resolves findings 1, 2, 7 and most of 8 by construction. + +## Fixed directly + +- **10 (MAJOR)** — the plan claimed "five executable checks, each self-contained and exiting + nonzero on failure". The self-tests are prose readings with no block and no exit status, so + the claim was false and the rider's execution report impossible. Now stated honestly: **four + executable checks plus one recorded reading**, with the reason — applying a prompt sentence to + a case is a judgement, and a script asserting it would assert its author's opinion. An + enforcement claim with no mechanism, inside the round that hardens that class. +- **3 (MAJOR)** — the branch preflight validated only the name. It now requires a clean tree and + index first, and on a resumed run prints the ahead-count so the executor can confirm the range + is this plan's commits. A pre-existing edit would otherwise be swept into whichever task + commits the same path, indistinguishable in the diff from this round's work. +- **4 (MAJOR)** — the cut dropped the spec's requirement that an absent story path be created by + an operation that fails if the path appeared meanwhile. Restored as a must-be-true naming the + mechanism (write to a temp file, `ln` into place), with the reason: a truncating redirect's + clobber is invisible in the final diff. +- **5 (MAJOR, medium)** — byte-identical reuse was a success state, yet each task still committed + unconditionally and the squash expected a fixed count. Resume semantics stated: identical means + skip the commit, and the squash folds whatever commits exist. +- **6 (MAJOR)** — the hook-count derivation redirected to a literal `out` in the repo root, which + would collide with a user file and leave an untracked artifact for the clean-tree guard to trip + over. Now `mktemp`, with the suite's own exit status captured before counting. +- **8 (MAJOR)** — push accepted any non-`main` branch and the PR block had no refusal at all. + Both now require the exact branch name; a `!= main` test passes on any wrong branch. +- **9 (MAJOR)** — `--body "$(cat "$EVIDENCE")"` swallows a `cat` failure, so a read error would + have opened the PR with an empty body. Read into a variable with an explicit stop. +- **11 (MAJOR, medium)** — resumed evidence was reused on path existence alone. Now requires a + regular non-symlink file whose header names this story and this branch; a stale file from + another cycle would otherwise pass the ``. Introduced by *pass 2's* fence-anchoring fix. +- **Pass 4's blocker** — fence mode counts `ends` on every closing fence after the region + opens, so on the real `workflow-init.md` it returns `MULTIEND` with `ends=7` and **4c can + never pass**. Introduced by *pass 3's* MULTIEND-reachability fix. Confirmed independently: + 13 fence lines follow §5 in that file, so `ends > 1` is certain. + +That is fix-of-fix oscillation on a fixed surface, which is the stop condition. + +## Where pass 4's findings actually live + +| Area | Findings | Count | +|---|---|---| +| 4c's `awk` region parser and its fixtures | 1, 2, 3, 4, 5, 15 | 6 | +| The two temp-tree shell recipes (mutation, counterfactual) | 8, 9, 10, 11, 12 | 5 | +| `WIP_PARENT` plumbing | 6, 14 | 2 | +| Prose placement of the §2.1 block | 7 | 1 | +| `todos.md` row text | 13 | 1 | + +**Eleven of fifteen are in shell embedded in a plan document** — roughly eighty lines of `awk` +and `sh` that nobody has run, being reviewed by reading. Pass 4's findings are the kind a +tool answers instantly: `shellcheck` finds the unguarded pipeline (8) and the trap +reassignment (10); running the fixture suite finds the unreachable `case` branch (2), the +inventory that does not match the literals (3), and the parser returning `MULTIEND` on the real +file (1) — which is exactly how Codex found it. + +**This is the spec cycle's lesson recurring one level down.** Design §1.6 records it: *a spec +must not ask mechanical questions about artifacts only humans read.* The plan's inverse is +asking a **reader** to decide questions a **machine** settles — and the review is doing that +job badly and expensively, four passes running. + +## The simplification that removes the subject matter + +The whole region-bounding apparatus — fence nesting, template anchoring, `ends` counting, +`NOSTART`/`MULTISTART`/`NOEND`/`MULTIEND` — exists to stop an occurrence in surrounding command +prose from satisfying the check. **Requiring the canonical line to appear exactly once in each +whole file achieves the same thing with `grep -c`**: an occurrence in the wrong place makes the +count 2 and fails, and there is no parser to get wrong. Six lines instead of thirty, and +findings 1, 2, 3, 5 and 15 stop existing rather than being fixed. + +The cost is a real contract change, so it is Daniel's call, not mine: spec §5.2 currently says +*"occurrences outside the region are ignored"*, and whole-file counting makes them fatal. That +is **stricter**, not weaker — but it is different, and it needs a one-line spec amendment +rather than a quiet reinterpretation. + +## Status + +Stopped and surfaced. Nothing applied from pass 4. The plan is uncommitted in the working +tree at its post-pass-3 state. diff --git a/.context/codex-reviews/gate-a-plan-pass-4.2026-07-30-classifier-cycle.md b/.context/codex-reviews/gate-a-plan-pass-4.2026-07-30-classifier-cycle.md new file mode 100644 index 0000000..49885e1 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pass-4.2026-07-30-classifier-cycle.md @@ -0,0 +1,30 @@ +BLOCKER | high | Task 2 Step 7 and Self-Review lines 373-405, 1506-1510 | the expected-message builders are still placeholders: none of the expected constants is defined and `branch_msg() { case "$1" in ... same shape ... ;; esac; }` is not executable shell | Task 5 Step 12 will abort under `set -u`, so pass-3 disposition 7 and the self-review claim that all helpers are complete are false | include every literal expected constant and the full `branch_msg` case before execution +MAJOR | high | Task 2 Step 6 and Task 4 Step 4 lines 348-371, 582-587 | `run_scenario` deliberately does not reset state, but Task 4 runs all scenarios sequentially without resetting; after `gateb_stale` leaves three passes, `gateb_floor` seeds a fourth pass and reaches the satisfied branch instead of the below-floor branch named by its label | the one-document matrix can pass without exercising every emit branch it claims to cover, leaving the behavior-neutral conversion incompletely tested | call `reset_all` before each Task 4 scenario or make a separate isolated runner for this matrix, then assert the expected branch text as well as one document +MAJOR | high | Task 5 A1 and embedded `skipval` comment, lines 610 and 816-818 | the prose still says `skipval` consumes exactly one JSON value even though the settled frozen cases prove it accepts balanced containers such as `[1,]` and `{"a" 1}` that are not JSON values | this is an old wider-validation claim surviving outside A2, exactly the claim-code drift the human decision asked this pass to find | say it consumes one quote-aware string or primitive token, or one balanced container span, without calling the latter a JSON value +MAJOR | high | Plan provenance and Task 1 Steps 3 and 7, lines 17, 127-146, 188-204 | the central bodies were embedded to make a fresh checkout self-contained, but both early fixture proofs and the promised permanent shipped test still call ignored `.context/plan-drafts/locate.awk`, which does not exist in a fresh checkout or installed plugin | execution and permanent fixture-drift coverage still depend on the disappearing scratch artifact that pass-3 disposition 34 claimed to eliminate | provide a tracked or inline test helper before Task 1, and specify a self-contained permanent slice check that ships with the suite +MAJOR | high | `.context/plan-drafts/README.md` lines 11-19 and 34-48 | the verification README still reports 50 rather than 53 cases and retains the withdrawn claims that refusal prevents every wrong verdict, the selected bytes arrive through a fully validated path, and an oversized payload is not accumulated | the stated verification bounds contradict the narrowed A2 contract and the actual 53-case corpus, so the provenance record can reintroduce the rejected wider claim | update the count to 53 and restate the same string-boundary-only, walkable-invalid, and scan-not-memory bounds used by Task 5 A2 +BLOCKER | high | Task 1 Step 4 lines 159-174 | the globally installed hook backup/restore sequence does not guard `mktemp`, `cp`, or hash failures, signal traps restore but do not terminate, and the normal path disables every trap before attempting restoration | an interruption or failed restore can leave the machine-wide hook dumping unrelated projects' prompts and review content, or overwrite it from an invalid backup | require isolated capture; if global mutation is unavoidable, hard-stop on every setup failure and use one cleanup function that restores and verifies before traps are disarmed, with signal handlers that exit after cleanup +MINOR | high | Task 3 Step 4 lines 484-501 | the executable shim hard-codes `/usr/bin/sed` while the following prose says to pin `command -v sed`, and the promised build guard and skip branch are not present | on a platform with another sed path the test can fail in routing rather than in the encoder and pass or fail for the wrong reason, so pass-3 disposition 6 remains incomplete | provide the complete shim using a captured exported real-sed path and an explicit preflight that skips only when selective injection cannot be built +MAJOR | high | Global Constraints and Task 5 Step 6 lines 64-66, 941-976 | `sed` is now load-bearing for classification but only `awk` is declared and fault-tested; if `sed` is missing or exits nonzero, the blank-test command substitution becomes empty and a genuine success block classifies as `no-result` | a supported jq-present environment with broken sed can permanently stop counters and emit the wrong two-cause diagnosis instead of taking the settled fail-open uncertainty path | check the sed status and map failure or partial output to `unrecognized`, with missing, nonzero, and partial-output tests, or remove sed from classification +MINOR | medium | Task 5 A2 and embedded locator ceiling lines 622 and 837-840 | the plan calls the limit 1 MiB, but POSIX awk `length()` counts locale characters and implementations differ, while the code never sets a byte locale | BWK awk and CI mawk can enforce different byte ceilings on non-ASCII payloads, weakening the claimed cross-implementation bound and compatibility evidence | run the locator under `LC_ALL=C` with exact boundary tests or describe the ceiling in awk length units rather than MiB +MAJOR | high | Task 5 A5 table and Step 11 lines 638-653, 1146-1163 | C2 is reduced to suppression plus pending-write failure, but spec §5.2 explicitly includes the second direction where gate-on emit fails and the pending write also fails | half of the accepted silent-count residual is absent from both the complete table and its test, so pass-3 dispositions 12-13 are incomplete | add a C2 row and surgical test combining a failed writer with an unwritable pending path while pass-state remains writable +MAJOR | high | Task 5 `note_unverified` and `flush_notes` lines 1061-1089 | when an earlier pending disclosure exists and the current event is also `unrecognized`, `note_unverified` sets `want_disclosure=1`, preventing the pending check from changing it to 2, so the carried debt is emitted without the required `Earlier:` prefix | this contradicts A5's absent-plus-present `any unsuppressed event` row and A7, and the composition matrix has no current-unrecognized scenario to reveal it | settle whether the current event absorbs prior debt; if not, detect pending before current desire and add a pending-plus-unrecognized golden +MAJOR | high | Task 5 Step 11 lines 1109-1203 | the plan says every marker-table row runs again through `nojq_run` or `nojq_run_closed`, but every shown assertion block uses only `run` or `run_closed` and no parameterization or second pass exists | jq-free marker and failure transitions remain untested despite pass-2 disposition 16 and pass-3 disposition 11 claiming complete dual-emitter coverage | parameterize every row over normal and jq-free runners, including selective write and marker faults, and show the executable loop +MAJOR | high | Task 5 Step 11 pending-delete test lines 1196-1202 | the test creates a directory, which is not recognized by `[ -f "$pending_file" ]` as an existing pending marker, then manually removes it with `rm -rf` before the event labeled as cleaning coexistence | the final assertion passes even if `flush_notes` never retries or clears a real shown-plus-pending state, so it does not assert what its label claims | inject one failed `rm` against a real pending file, preserve shown-plus-pending, restore deletion, and let the next hook event perform and prove the cleanup +MINOR | high | Task 5 Step 11 line 1205 | the explanatory paragraph says row 10 uses `chmod` on `.context/`, while the rewritten row uses gate-off suppression and a directory at the pending path and explicitly says chmod is wrong | this stale post-fix instruction can make an implementer restore the vacuous test that pass 3 removed | delete the paragraph or rewrite it to describe the actual selective pending-path fault +MAJOR | high | Task 5 Step 1 lines 661-679 | the plan names Claude via Claude Code only in surrounding design prose; none of the delivered `additionalContext` strings states its target model | prompt-standard item 1 says the prompt itself must name the executing model, so the five rewritten model-facing prompts still do not verifiably satisfy invariant 11 | include one compact target declaration in each delivered context or in a shared prefix that `note` actually emits +MAJOR | medium | Task 5 Step 1 lines 665-679 | `FAILURE_CTX` and `NORESULT_CTX` tell Claude to surface or report an outcome, but no output structure or example is supplied | prompt-standard item 4 remains unmet for the retry/stop outputs despite the plan's all-12 conformance claim | add a compact exact reporting format and one example for the contexts that request a report or escalation +MAJOR | high | Task 5 `FAILURE_CTX` line 666 | the message names execution-failed and timeout codes but gives neither a distinguishing check beyond rereading the code nor a cause-specific fix; `addressing what that code names` is the same vague remedy the rewrite was meant to remove | prompt-standard item 10 is still unmet, and pass-3 disposition 18 incorrectly claims matching remedies were added | give concrete separate checks and actions for execution setup failure, executor timeout, and unknown codes before the one allowed retry +MAJOR | high | Task 5 `UNVERIFIED_MSG` line 679 | `Known causes, each with its check` is false: oversize, ambiguity, and missing or failing awk have no checks or fixes, and the oversize cause is misstated as a result over 1 MiB although the ceiling applies to the whole payload | the operator can enter an unbounded diagnosis loop and the shipped enforcement claim violates prompt-standard items 10 and 11 | name the whole-payload ceiling and pair every listed cause with a concrete discriminator and action, or remove causes that cannot be made actionable +MAJOR | high | Task 5 `UNVERIFIED_MSG` and its load-bearing rationale lines 679, 684 | the user-visible message still says a call `was counted` even though counter and fingerprint writes are best-effort and may both fail | the categorical recorded-state claim that pass 3 fixed in `UNVERIFIED_CTX` survives in the other field and is false on an unwritable `.context/` | use the same `classified as countable and attempted to record` wording in the system message +MAJOR | high | Task 5 `UNVERIFIED_CTX` line 678 versus A5 and C4 | `this note ... will not repeat` is false when the shown-marker write fails, pending deletion fails, or concurrent invocations race | a prompt-standard item-11 guarantee contradicts the plan's own sequential failure table and accepted concurrency residual | say the note is normally once per workspace and may repeat when diagnostic state cannot be persisted or concurrent hooks race +MINOR | high | Task 5 `UNVERIFIED_CTX` line 678 | telling Claude to treat a counter containing false passes `as a floor` reverses the usual meaning: such a count is not a lower bound on completed reviews and can overstate them | the wording can encourage the exact false checkmark the disclosure is meant to prevent | call it only a mechanical counter signal and require independently discounting every incomplete or unverified call +MINOR | high | Task 5 Step 1 line 663 and `FAILURE_MSG` line 667 | the plan claims every system message has state followed by operator action, but the failure system message contains only the state and no action | the asserted uniform prompt structure is not true, weakening the all-12 review audit | either add the operator action, state explicitly that this class needs none because Claude owns recovery, or narrow the uniformity claim +MAJOR | high | Task 6 Step 2 literal CLAUDE replacements lines 1300-1316 | the replacement again says any `success: false`, any harness backgrounding notice, and any payload with no readable result is not counted, omitting immediate-first recognition, C1's reworded-notice fallback, and locator-uncertainty fail-open behavior | it restores the exact broader claim the story criterion was amended to remove and teaches downstream agents incorrect counter semantics | describe only recognized immediate-first failure, the recognized notice anchor, and unambiguous no-result, with unrecognized shapes explicitly counted +MAJOR | high | Task 6 Step 2 README replacement line 1318 | `pass counters additionally require a result the hook can read as a completed review` excludes `unrecognized`, which is deliberately counted even though the hook cannot read it as completed | the main setup documentation contradicts settled decision 2 and can make users trust a satisfied counter as inspected | say that recognized failure, backgrounded, and no-result classes do not count while success and unrecognized do +MAJOR | high | Task 6 Step 6 line 1355 | the planned setup text says a backgrounded pass is discarded rather than counted without limiting that claim to the recognized notice wording | this hides accepted residual C1 in the primary-defense documentation and becomes false as soon as harness prose changes | state that the observed recognized notice is discarded and that a reworded notice falls to counted `unrecognized`, which is why the environment setting remains primary +MAJOR | high | Task 6 Step 4 unknown-tool prompt lines 1328-1334 | the new namespace diagnosis and register-as-codex remedy are placed in model-facing `additionalContext`, while the operator-visible field stays `see the note` and the plan claims the operator has no action | this recreates the field-split inversion fixed in Task 5: only the operator can change MCP registration | move the concrete registration action and namespace check into `systemMessage`, leaving model context with the gate consequence +MAJOR | high | Task 7 Step 3 lines 1433-1456 | the success verification retargets `shape0-success`, an exec call that only replied `ok`, to the review tool and then calls it a genuine pass with a stored Gate-B fingerprint even though Task 1 created `shape0-success-review` for exactly this distinction | the named high-risk verification does not satisfy the story's genuine-review reading and can claim Gate-B continuity for content no reviewer saw | use `shape0-success-review` for the Gate-B success row and retain the exec capture only for Gate-A success if desired +BLOCKER | high | Task 7 Step 7 line 1496 | the offered `git reset --soft ` closeout occurs after the clean Gate-B pass and changes HEAD-relative tracked diff plus the effective index, both inputs to the content fingerprint | the real commit immediately becomes stale under invariant 3, so this advertised closeout cannot preserve the clean pass and will correctly fire Gate B again | squash to the recorded base before creating the reviewed WIP snapshot, or keep the task history and amend the WIP; if resetting after review, rerun Gate B on the post-reset fingerprint before committing +MAJOR | high | Task 7 Step 5 lines 1462-1475 | the rollback verification is only two comments and `ls`; it gives no exact way to select the prior hook, run it in an adopted disposable repo, assert the old count, or handle a missing or modified 0.7.1 cache | the high-risk rollback obligation remains non-reproducible, so pass-3 disposition 33 overstates what was fixed and an operator can discover at incident time that no verified rollback exists | provide exact guarded commands and expected state, stop release if the prior copy is absent or differs, and document the actual version-switch command separately from the behavioral probe +END OF FINDINGS (29 total) diff --git a/.context/codex-reviews/gate-a-plan-pass-4.2026-08-04-hardening-round.md b/.context/codex-reviews/gate-a-plan-pass-4.2026-08-04-hardening-round.md new file mode 100644 index 0000000..38649b6 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pass-4.2026-08-04-hardening-round.md @@ -0,0 +1,12 @@ +BLOCKER | high | Task 9 Step 6 | fixes are staged but never amended into the WIP snapshot before re-review | `mcp__codex__review` reads the git range, so every post-fix pass reviews the old commit; the staged index also makes the required clean-state precondition impossible, so any Gate-B finding deadlocks or closes on unreviewed content | after every fix, stage it and amend the WIP commit while preserving its WIP subject, then re-run the battery and Gate B against the unchanged parent +MAJOR | high | Task 9 Step 2 | the squash block promises exactly eight plan commits and an exact first subject but checks neither | a prefix-matching anchor followed by any number of unrelated commits is soft-reset without a failing guard, so the wrong history can be rewritten | compare the eight subjects and their order exactly over the candidate range, require a count of eight, and only then run `git reset --soft` +MAJOR | high | Task 0 Step 1 | the branch preflight validates only the branch name, not a clean tree/index or a clean branch range based on the intended start commit | pre-existing edits can be folded into same-path task commits, and an already-named branch can carry earlier commits that Gate B excludes by using the WIP parent as its base; neither ownership defect is reliably visible to the battery or reviewer | before creating or accepting the feature branch, require a clean status and verify the expected base/range, with an explicit resumable-state branch for this plan's own commits +MAJOR | high | Tasks 4–5 path classification | the plan drops the spec's requirement that an absent story path be created by an operation that fails if the path appears meanwhile | another session can create or change the path after classification and have its bytes silently replaced; the final diff cannot reveal that clobber | require atomic no-clobber creation immediately after classification and stop on an exists/collision result +MAJOR | medium | Tasks 4–6 commit steps | byte-identical story reuse and identical `todos.md` skips are declared successful states, but each task still unconditionally runs a commit and Task 9 later requires exactly eight new WIP commits | a valid resumed or concurrent-completion path reaches an empty commit failure, or cannot satisfy the squash count even though the deliverable is already correct | define task-level resume semantics: skip the commit when the task is already present in HEAD, or deliberately create an allowed empty commit and make the squash validator account for it +MAJOR | high | Task 9 Steps 1 and 3 | the documented hook-count derivation redirects to the fixed repo-root path `out` without collision handling or cleanup | it can overwrite an existing user file and otherwise leaves an untracked file that makes the later clean-tree guard fail | write to a uniquely created temporary file with cleanup, or to a validated dedicated path under `.context/`, and capture the suite status before counting +MAJOR | high | Task 9 Step 7 | the `` template block, then stop considering later fences; add a fixture containing a valid target template followed by another real-shaped fenced scaffold block +MAJOR | high | Task 1 Step 4, lines 348–388; Self-review lines 919–923 | `severity_rule_count` returns `MULTIEND`, but the caller has no `0:MULTIEND` case and falls into the generic occurrence-count failure, producing the nonsensical detail `found MULTIEND occurrences`; the self-review nevertheless says every return code has a distinct branch | the two-end fixture can pass while reporting the wrong cause, and operators are told to repair the canonical-line count instead of the ambiguous boundary | add a dedicated `0:MULTIEND` diagnostic and make its fixture assert that unique diagnostic +MAJOR | high | Task 1 Step 2–3, lines 239–282; Self-review lines 929–931 | the literal fixture list still includes the missing-template case and an unconditional unreadable-file case, while later prose says to drop the former and wrap the latter; it also calls 23 `sev_case` cases “Twenty-four” before adding the parser case, while the claimed final dropped inventory would be 22 `sev_case` calls plus the parser case on non-root runs | the implementation is not decidable, the missing-template red phase fails on 4b rather than with the promised `(exited 0)`, root runs duplicate or mishandle the permission case, and the mutation assertion count has no stable source inventory | provide one final literal fixture block: remove the missing-template call, replace rather than supplement the unreadable call with the root guard, state the resulting count including the parser fixture, and emit an explicit skip assertion if counts must be root-stable +MAJOR | high | Task 1 Step 2, `sev_case`, lines 145–169 and 394–398 | every reject case accepts any nonzero result containing the shared substring `closed severity set`; it never checks the cause-specific diagnostic | a fixture named duplicate start, missing end, unreadable input, or wrong count can exercise a different 4c branch and still pass, so the requested proof that all fixtures fail for the named reason is absent | pass an expected cause token or exact diagnostic into `sev_case` and assert it per fixture, retaining the shared substring only as the isolation check +MAJOR | high | Task 1 Step 4, lines 342–346; Step 6 lines 411–418 | the equality normalizer removes trailing spaces and tabs even though the approved spec calls the canonical line byte-for-byte and the plan says only the leading blockquote marker and indentation are stripped | a line with trailing whitespace, including Markdown’s semantically significant two-space hard break, passes without being the pinned canonical line; no fixture covers this accepted drift | remove the trailing-whitespace substitution or amend the approved contract explicitly, and add a trailing-space fixture matching the chosen rule +MAJOR | high | Global Constraints lines 54–60; Task 1 interface lines 83–87; Task 5 lines 773–795; Gate B lines 826–835 | the plan says every consumer recomputes `WIP_PARENT`, but the task interfaces still say it is captured and carried from Task 1, and the executable Task 5 and Gate-B commands use the variable without first running the recomputation and WIP-subject guard | shell variables do not survive task tool calls, so the version check and review base can receive an unset or stale value despite the stated pass-3 fix | put the two-line recomputation and subject guard immediately before every Task 5 and Gate-B consumer, and change the interfaces to say the parent is derived locally rather than consumed from Task 1 +MAJOR | high | Task 2 Steps 1–3, lines 567–639 | the literal settled §2.1 text is a flush-left blockquote, but the placement instructions call it a list item, tell the mirror to match list indentation, and allow an indentation difference during parity | an implementer can follow the prose and alter the byte-for-byte shipped block or accept divergent copies, contradicting both the literal block and the stated settled decision | state that the blockquote is inserted flush-left as shown in both files, remove the list-item and indentation-exception language, and require a byte-identical diff +MAJOR | high | Mutation recipe line 497; counterfactual lines 853–855 | both temp-tree recipes pipe `git ls-files -z` directly into `xargs`; POSIX shell reports only the `xargs` status, so a failed or partial producer can be treated as a successful copy, and interpolating each pathname into shell source also breaks on quote-bearing tracked names | a partial tree can become the supposedly current green baseline, allowing evidence to pass without comparing the full tracked worktree | write the NUL list in a separately status-checked command, feed it to an `xargs` script that receives paths as positional arguments rather than source interpolation, check every copy status, and verify the copied tracked set before running either baseline +MAJOR | high | Mutation recipe lines 497–503 | mutant construction is not guarded: neither the copy nor the `sed` deletion is checked before the mutant suite runs | an empty, truncated, or otherwise malformed checker can make the mutant suite fail and create a nonempty flip set, so a construction failure can be recorded as load-bearing evidence despite the claim that every VOID condition exits nonzero | guard the copy and `sed` statuses, verify exactly one marked region was removed and the remaining checker matches the baseline outside it, then run the mutant +MINOR | high | Mutation recipe lines 493–509 | `trap` is reassigned inside the three-iteration loop, so only the last temporary directory is removed on normal exit; the HUP, INT, and TERM trap also cleans up without explicitly terminating | two temp trees leak on every successful run, and a trapped signal can continue against a removed tree | allocate one parent temp directory with one trap outside the loop, or clean each child before replacing the trap; make signal handlers clean up and exit nonzero +MINOR | high | Mutation recipe line 503; counterfactual lines 850–867 | the recipes are not self-contained with respect to errexit: an inherited `set -e` aborts on the expected `diff` status 1 before its status guard, while the counterfactual executes `set +e` and never restores it | behavior depends on caller shell state and the counterfactual can silently disable fail-fast behavior for subsequent commands in the same shell | place `diff` in an `if` or other tested context, capture the mutant status in a guarded construct, and restore the prior errexit state before returning +MAJOR | medium | Counterfactual lines 860–870 | the recipe accepts any nonzero mutant status although the evidence claims exit 1, and its “nothing else” comparison counts only lines beginning `Invariant` or `Prompt standards`; unrelated shell, git, awk, or copy diagnostics are invisible to `all` | the command can print its success sentence while the mutant exited operationally or emitted additional uncounted diagnostics, overstating what the evidence establishes | require `st` to equal 1 and validate the complete normalized output against the allowed 4c diagnostic blocks, rejecting every other nonblank line +MAJOR | medium | Task 4 Steps 3–4, lines 727–751; Self-review lines 909–917 | the plan says every edit is exact, but the seven new `todos.md` rows are only summaries; notably the external-authority summary drops the settled trusted-signer, role-policy, independent-availability-attestation, and “one direction, not the only possible one” constraints, and the tier-2 and slot-collision entries provide no complete final row text | an implementer following the plan can create materially weaker backlog records while still checking off all §6 bullets | include the exact final Markdown for all seven rows or explicitly require copying each full §6 bullet verbatim and then show the precise additions needed to fit the backlog format +MINOR | high | Task 5 Step 3, lines 791–800 | the prose says an empty comparison range is VOID, but the only executable command is `check-version-bump.sh`, which deliberately exits green on an empty range and does not report whether the range contained changes | the claimed guard is not executable and a stale or wrong parent can produce evidence indistinguishable from a deciding pass | add an explicit nonempty plugin-range assertion against the recomputed parent before invoking the checker +MINOR | medium | Task 1 Step 4, lines 334–341 | heading mode treats any column-zero `## 5.` line as a section heading without tracking Markdown fences | a fenced example can become a false start, duplicate start, or even the only accepted region, contrary to the contract’s reference to the actual section-5 heading | make heading mode fence-aware or document the line-oriented limitation and add accept/reject fixtures for fenced heading lookalikes +END OF FINDINGS (15 total) diff --git a/.context/codex-reviews/gate-a-plan-pass-5-dispositions.2026-07-30-classifier-cycle.md b/.context/codex-reviews/gate-a-plan-pass-5-dispositions.2026-07-30-classifier-cycle.md new file mode 100644 index 0000000..820df3e --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pass-5-dispositions.2026-07-30-classifier-cycle.md @@ -0,0 +1,34 @@ +# Gate A (plan) — pass 5 dispositions + +Advisory companion to `gate-a-plan-pass-5.md` (2 BLOCKER, 15 MAJOR, 2 MINOR = 19). One line +per finding, in file order. **All accepted, none dismissed** — the first pass in this cycle +with no partial dismissal. + +Trend: Blocker+Major 36 → 28 → 27 → 24 → **17**; blockers 7 → 5 → 5 → 3 → **2**. + +**Three of my own pass-4 dispositions were overstated, and pass 5 caught each.** Disposition +1 said the expected constants were "written out in full" — only the `branch_msg` case was. +Disposition 12 said the marker rows were parameterized over both emitters — the wrapper was +declared and every block still called `run`. Disposition 4 said the ignored-scratch +dependency was removed — Step 3's was; Step 7's was not. That is a pattern worth naming: I +wrote the disposition from the fix I intended rather than from the text I left behind. + +1. BLOCKER `{ : > "$marker"; }` exits `dash` on a redirection failure — ACCEPT, verified in-session and worse than reported. `:` is a POSIX **special builtin**: with a directory at the target, `{ : > m; } 2>/dev/null || true; echo REACHED` prints nothing and exits **2** under `dash`, while the `printf` form prints `REACHED` and exits 0. `/bin/sh` on Ubuntu — which CI runs — is `dash`. **This is a live defect in the shipped hook**, not only in the plan: `codex-gate.sh:394` already uses that form, so on Linux an unwritable `.context/` makes the hook exit 2 from the unknown-tool branch — a direct violation of invariant 1. All four writes converted to `printf '%s' ''`, the pre-existing line is fixed in Task 4, and the marker faults are asserted under `dash` as well as `sh`. The existing "special-builtin redirection regression" test only ever ran under macOS `sh`, where the form happens to survive, which is why this stood. +2. BLOCKER no `*_EXPECTED` constant is assigned — ACCEPT; my pass-4 disposition was wrong. All thirty are now assigned, split by provenance: the five new pairs copied from Task 5 Step 1, the ten pre-existing ones copied verbatim from `codex-gate.sh` (not re-quoted here — Gate A has no business re-reviewing prompts nobody is editing), plus a mechanical inventory check, because "I copied them all" is the claim that has now failed twice. +3. MAJOR the dual-emitter wrapper is declared and never used — ACCEPT; same class as 2. Replaced with a real `marker_rows()` function taking the runner pair, called once per mode, with the raw-`sh "$HOOK"` cases parameterized too. +4. MAJOR C2 direction 2 has a table row and no test — ACCEPT. Added: gate on, writer closed, pending path unwritable, pass-state writable — asserting exit 0, counted, and neither marker present. +5. MAJOR the current-`unrecognized` prefix row has no scenario — ACCEPT. Added, asserting one document, **no** `Earlier:` prefix, the exact `UNVERIFIED_CTX` golden, shown written, pending cleared, and the count still advancing. +6. MAJOR `sed` faults are promised and only `awk` is tested — ACCEPT. Three selective `sed` shims (absent, nonzero, partial output), asserting that a **success** envelope still counts — which is the assertion that moves if the fail-closed direction ever returns. +7. MAJOR the Unicode-escaped marker row tests the wrong thing — ACCEPT, verified: the row supplied raw quotes, making the outer payload malformed, so it reached `unrecognized` from the locator (rc=2) rather than from the matcher. Pass-2 disposition 10 was wrong. The row now spells the key with real `"` escapes and the payload is well-formed. +8. MAJOR jq-free goldens compare decoded expectations against escaped bytes — ACCEPT, and it would have failed the suite outright on a jq-free machine, since every new reporting format contains quotes. The expectation is now escaped the same way the emitter escapes, on the **known** side, which keeps the hook's escaper out of its own assertion. +9. MAJOR Task 1 Step 7 still executes the ignored path — ACCEPT; my pass-4 disposition covered Step 3 only. Both extractions now use Step 3's temp copy. +10. MAJOR item 1 is not satisfied by a source comment — ACCEPT, and the previous dismissal is **withdrawn**: item 1 says *the prompt* names the model, and a comment the model never receives is not the prompt. Every `additionalContext` now opens `Claude Code gate hook —`; four words that also give these injected strings the provenance they otherwise lack. The header comment stays for the "checked that model's prompting page" half, which a prefix cannot carry. +11. MAJOR the `CODEX_TIMEOUT` remedy narrows the review scope — ACCEPT, and it is the sharpest finding of the pass: "re-run with a smaller instruction or fewer files" would count a pass for less than the artifact or diff the gate requires, recreating the false checkmark from inside the fix. Now: re-run the same scope, and raise the executor timeout or reduce load **outside** the review. +12. MAJOR real remedies filed under "no user-side fix" — ACCEPT. Unmapping or replacing a third-party server, and repairing `awk`/`sed`, are actions an operator can take; only the scan ceiling, an ambiguous payload and a parser defect have none. Regrouped. +13. MAJOR `NORESULT_MSG` diagnoses before running its own check — ACCEPT: absence of a mapping file is not proof the pinned server is effective, since scope precedence can put a different server under the same `codex` name — which the same message elsewhere admits. Both checks now run first, and the pinned-server diagnosis is reached only when both confirm it. +14. MINOR the message still says 1 MiB — ACCEPT; the design prose was narrowed to Mi-units and the prompt kept the stronger byte claim. +15. MAJOR the CLAUDE and README replacements invert decision 3 — ACCEPT, and this one is mine from pass 4: fixing decision 2 ("unrecognized counts") I wrote "an unreadable result counts", which is decision 3's `no-result` and does **not**. Both replacements now turn on whether any result text was *obtained*, and the load-bearing precision list grew from three to four so the distinction cannot be dropped again. +16. MAJOR the rollback probe cannot fail — ACCEPT: the mismatch branch ended in `echo`, inside a subshell, so the release continued past a failed check. Now exits nonzero with an outer hard stop. +17. MAJOR `git show v0.7.1:…` cites a tag that does not exist — ACCEPT, verified: this repo has **no tags at all**, so the mandatory rollback check would stop before comparing anything. The release commit is resolved from the manifest history instead, with creating a tag noted as a prerequisite rather than assumed. +18. MAJOR six non-WIP commits precede the only Gate-B cycle — ACCEPT. Each closed a cycle and reset the counters without a review, and on a shared branch they are what someone else pulls. All six are `WIP:`-prefixed now, squashed into one reviewed snapshot **before** Gate B runs — before, because pass-4 finding 28 established that squashing after a clean pass invalidates the fingerprint it recorded. +19. MINOR `shasum | cut` reports `cut`'s status — ACCEPT: a checksum emitting a partial line and failing would be accepted as a verified hash, in the procedure whose whole job is proving a globally installed hook was restored. Now a checked helper, used identically before mutation and after restore. diff --git a/.context/codex-reviews/gate-a-plan-pass-5-dispositions.2026-08-04-hardening-round.md b/.context/codex-reviews/gate-a-plan-pass-5-dispositions.2026-08-04-hardening-round.md new file mode 100644 index 0000000..1a8d2f1 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pass-5-dispositions.2026-08-04-hardening-round.md @@ -0,0 +1,59 @@ +# Gate A — plan — pass 5 dispositions + +9 findings, 5 Major, 4 Minor, no Blockers. Plan trend 19 → 11 → 19 → 11 → 9. All 9 accepted; +none dismissed; **none required a decision from Daniel** — no new class, no reversal of a settled +decision. That is the pinned exit, so the plan's Gate A closes here. + +Both settled decisions the pass was asked to judge on their merits survived: the reviewer +reported no dropped check whose violation would reach the commit undetected, and no round-specific +content lost in the §5 citation. + +## Fixed + +- **1 (MAJOR)** — the plan appended evidence case 3 beneath the scope-blind row's own premise, + which this round disproved: the row still said the workaround "lives in ledger prose, which + agents do not read", while `harden-finding` step 3 does re-read the log. Task 6 Step 2 now + corrects the premise to the decision-branch diagnosis **in the same change**, before appending. + Leaving it would have been `docs-drift` created by the round that hardens it. +- **2 (MAJOR)** — the branch preflight's resumed path printed an ahead-*count* while the prose + claimed the executor could confirm the range holds only this plan's commits. A count cannot + establish that. It now prints the actual `main..feature` commit list with an explicit stop + instruction. +- **6 (MAJOR)** — evidence creation had the classify-then-create race the story paths had already + been fixed for. Same treatment: temporary file, `ln` into place, `trap` cleanup. This file is + quoted verbatim into the commit body and the PR, so foreign content reaching it is worse than a + stray story copy. +- **7 (MAJOR)** — the `` before `gh pr create`, + so a commit made between push and PR would open a PR whose remote head omits the reviewed close. + Compared explicitly, with a stop on divergence. + +## Sweep after fixing + +22 blocks. All parse under `sh -n` and `dash -n`; no `${var:o:l}`, `[[ ]]`, `<(...)`, `declare` or +`local`; every block reading `EVIDENCE` defines it; no unanchored `Branch:` grep remains; the +stated verification totals (four executable checks plus one recorded reading) match what exists. + +## Gate A closes for the plan + +Five valid passes under the file-first protocol, every finding dispositioned in writing. The +final pass returned no Blocker and nothing requiring a decision. The plan is ready to execute. diff --git a/.context/codex-reviews/gate-a-plan-pass-5-dispositions.2026-08-12-ledger-supersession.md b/.context/codex-reviews/gate-a-plan-pass-5-dispositions.2026-08-12-ledger-supersession.md new file mode 100644 index 0000000..810e7cf --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pass-5-dispositions.2026-08-12-ledger-supersession.md @@ -0,0 +1,58 @@ +# Gate A — plan — pass 5 dispositions + +Artifact: `docs/superpowers/plans/2026-08-12-hardening-ledger-supersession.md` (one finding landed +in the spec). 6 findings: **1 Blocker**, 3 Major, 2 Minor. All six valid. **All six applied.** + +Reviewer: `mcp-codex-dev@1.0.1`, `codex-cli 0.144.5`, model `gpt-5.6-sol`. Pass file valid: +7 lines, 6 finding lines, terminator exact. + +Plan passes: **18, 11, 9, 8, 6.** Blockers: **2, 1, 4, 3, 1.** + +## The Blocker + +**`BASE` contains the plan, so `$BASE..HEAD` excludes it — while spec §8, as pass 4 rewrote it, said +the committed plan falls inside the reviewed range.** The pass-4 correction had landed in the spec +and not in its plan mirror, which is the failure mode this pass was aimed at and found once more. + +Resolved toward the classification Task 1 already establishes rather than by moving the baseline: +the plan is a Gate-A artifact committed as prose, so it is **not** in the Gate-B range, and §8 now +says so — neither the plan nor the check source is inside it, both reach the reviewer as context. +Moving `BASE` to the Task-1 commit's parent was rejected: it would drag the spec and story into a +range they were classified out of, replacing one contradiction with a worse one. Task 8's +`additionalContext` now carries the plan's path alongside the story's. + +## The three Majors + +**Task 2 Step 5 called the self-test, `shellcheck` and the counter-checks the harness's "only +review"** — contradicting the opening contract, Task 8 and spec §8, all of which add the reviewer's +read of the pasted source. Two incompatible accounts of the same guard. Now: *outside the reviewed +range* is the true claim, *unreviewed* is not, and the reviewer read is listed with the rest. + +**`entry-outside-block` had no determinate expected result.** "An entry-shaped line below the table" +is two different fixtures: the mandated entry **moved** there (must fail `C1d`) or an extra line +added while the mandated entry stays (must **pass**, under `C1d`.1 and `.2`). Building the second +and expecting failure would have made the pass-2 scope remedy look broken. Defined as the move, with +the addition left to Task 5's separate counter-check, and both constructions named in the cell. + +**`C1c`'s counter-check never re-ran after the final restore.** The last green result predated the +mutations, so an incomplete restore could have been committed under a verdict that never saw it. A +green run on the restored real surfaces, both shells, now closes the step before the commit. + +## The two Minors + +Task 4 Step 5's "everything else is verbatim from §3.1" also covered the `**Superseded rows:**` +label, which is §2.2's — the verbatim claim is now scoped to the entry line, with the label +attributed separately. Task 8 Step 3's restatement of what the `WIP: ` prefix does was replaced by a +citation; the leftover duplicate sentence was removed in the same edit. + +## Sweep after applying + +Step numbering contiguous in all eight tasks; the WIP-prefix restatement gone; 35 anchors +byte-identical; 12 fence lines, balanced. + +## Trend + +Findings **18 → 11 → 9 → 8 → 6**; Blockers **2 → 1 → 4 → 3 → 1**. Monotone in the count since pass 2 +and falling in severity since pass 3. The class has narrowed to one shape — *a fix that landed in +one site and not in its mirror* — which is what passes 4 and 5 were aimed at, and pass 5 found one +instance where pass 4 found three. Not oscillating. diff --git a/.context/codex-reviews/gate-a-plan-pass-5.2026-07-30-classifier-cycle.md b/.context/codex-reviews/gate-a-plan-pass-5.2026-07-30-classifier-cycle.md new file mode 100644 index 0000000..42a8e5f --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pass-5.2026-07-30-classifier-cycle.md @@ -0,0 +1,20 @@ +BLOCKER | high | Task 5 Steps 10-11 marker writes and directory-at-marker fault injections | the planned `{ : > "$marker"; }` writes use the POSIX special builtin `:`, and a redirection failure on a directory exits `dash` immediately even inside the shown `if` or `... || true` forms; the same probe survives under macOS `sh` but exits 2 under `dash` | the row-4, C2, and bgAdvice directory faults do not reach their asserted marker transitions on Ubuntu, and the shipped hook violates invariants 1 and 4 on those paths | use a non-special writer such as `printf '%s' '' > "$marker"`, then assert every directory-at-marker path exits 0 under both `sh` and `dash` +BLOCKER | high | Task 2 Step 7 expected-message builders | no `*_EXPECTED` constant is assigned anywhere in the plan; the code still contains only a comment instructing the implementer to copy literals, despite pass-4 disposition 1 saying they were written out in full | Task 5 Step 12 aborts under `set -u` before any composition golden runs, so the plan is not executable and the disposition is false | include literal assignments for every expected context/message before the builders, including the interpolated fixed-count variants, and verify the constant inventory mechanically +MAJOR | high | Task 5 Step 11 dual-emitter wrapper | the wrapper declares `$R` and `$RC`, but every shown assertion block still calls `run`, `run_closed`, or raw `sh "$HOOK"` directly | the promised jq-free replay never occurs, so pass-4 disposition 12 remains incomplete and fallback marker-state regressions can ship | replace every direct runner in the wrapped blocks with `$R` or `$RC`, provide a runner for the selective raw-hook cases, and assert the loop executed both modes +MAJOR | high | Task 5 A5 C2 direction 2 and Step 11 | the table now contains the failed-emit plus failed-pending-write C2 row, but the only C2 test suppresses emit with the off marker; no test combines a failing writer with an unwritable pending path | half of the accepted silent-count residual remains unpinned despite pass-4 disposition 10 claiming both the row and surgical test were added | add a closed/failing-stdout case with a directory at the pending path while pass-state remains writable, then assert exit 0, count/store occurred, no output was delivered, and neither disclosure marker is a file +MAJOR | high | Task 5 A5 current-unrecognized prefix row and Steps 11-12 | the new row settling pending plus a current `unrecognized` event to one unprefixed disclosure has no scenario or golden; the composition matrix seeds pending but never uses an unrecognized gate call as the current scenario | the precedence that motivated pass-4 finding 11 can regress without any test moving, so that disposition is incomplete | add a pending-plus-current-unrecognized scenario asserting one unprefixed disclosure, successful shown-marker creation, pending deletion, and unchanged count/store semantics +MAJOR | high | Global Constraints and Task 5 Steps 6-7 | the plan promises missing, nonzero, and partial-output fault tests for both load-bearing tools, but Step 7 implements those three shapes only for `awk`; no classification-path `sed` fault test exists | pass-4 disposition 8 verifies the implementation branch only in prose, leaving a future unchecked sed status able to turn a real success into fail-closed `no-result` | add selective classification-sed shims for absent, nonzero, and partial-output failure and assert both gate tools count/store and disclose as `unrecognized` +MAJOR | high | Task 5 Step 2 and `.context/plan-drafts/verify.sh` Unicode-escaped marker row | the row labeled `unicode-escaped marker key` supplies `{"success": true}` without any Unicode escape; through `tcls` that is malformed outer JSON with raw inner quotes, so it reaches `unrecognized` for the wrong reason | pass-2 disposition 10's claim that a genuine Unicode-escaped key was added is wrong, and the matcher behavior for the intended valid exotic spelling remains untested | construct a valid routed payload whose result text genuinely spells the marker key with Unicode escapes, separately validate the outer payload, and assert `unrecognized` +MAJOR | high | Task 5 Step 4 `field_of` fallback and jq-free goldens | without jq, the sed fallback extracts JSON-escaped field bytes but compares them with decoded literal expectations; the new reporting formats contain quotes, so correct jq-free output returns `\"...\"` and fails equality against `"..."` | the suite itself fails on a supported jq-free machine and cannot establish invariant 4 there | compare against independently escaped expected JSON bytes in fallback mode or add a portable decoder whose behavior is tested independently from the hook emitter +MAJOR | high | Task 1 Step 7 response-slice proof | the command still invokes `.context/plan-drafts/locate.awk` even though the surrounding prose says it uses Step 3's temp copy and pass-4 disposition 4 says the ignored dependency was removed | Task 1 fails in a fresh checkout, exactly the disappearing-scratch failure the disposition claimed to close | invoke the Step-3 temp locator `$L` in both extractions, or place the proof after the tracked hook locator exists; leave no executable Task-1 reference to `.context/plan-drafts` +MAJOR | high | Task 5 Step 1 prompt-standard item 1 disposition | a source comment naming Claude as the consumer is not part of any delivered `additionalContext` prompt, while checklist item 1 explicitly requires the prompt to state which model executes it | all five model-facing shipped prompts still fail invariant 11, so pass-4 disposition 15 dismisses a binding checklist requirement without amending that requirement | emit one compact shared target-model prefix in every context, or obtain a human-approved amendment to `docs/prompt-standards.md` before claiming conformance +MAJOR | high | Task 5 Step 1 `FAILURE_CTX` CODEX_TIMEOUT action | telling Claude to retry with a smaller instruction or fewer files can narrow Gate A's required one broad prompt or Gate B's reviewed range | a retry can count while reviewing less than the artifact or diff, contradicting CLAUDE.md section 5 and recreating the false-checkmark direction | keep the required review scope fixed; recommend a timeout/configuration remedy that preserves it, and stop after the shared single retry if that full-scope call still times out +MAJOR | high | Task 5 Step 1 `UNVERIFIED_MSG` cause grouping | the message labels mapped third-party envelopes and missing/failing awk or sed as causes with no user-side fix, although the operator can unmap or replace the server and can install or repair the required tools/PATH | the rewritten diagnostic still violates prompt-standard item 10 by assigning real remedies to the no-remedy bucket, sending operators away from actionable recovery | group every cause by its actual check and action, reserving “no user-side fix” only for parser defects or genuinely unsupported envelopes after replacement/unmapping is considered +MAJOR | high | Task 5 Step 1 `NORESULT_MSG` provenance decision | “If nothing is mapped, this is a payload-contract change the pinned server cannot produce” treats absence of `.context/codex-gate.tools` as proof that the pinned server is effective, even though the same message acknowledges MCP scope precedence can select a different server under the same `codex` name | the diagnostic can assert the wrong cause before running its own distinguishing check, repeating the failure class prompt-standard item 10 exists to prevent | first inspect both the mapping file and `claude mcp list`; only diagnose a pinned-server contract change when the effective registration and version are verified, otherwise diagnose the effective third-party server +MINOR | high | Task 5 Step 1 `UNVERIFIED_MSG` oversize cause | the shipped message still says “1 MiB whole-payload scan ceiling” after A2 was narrowed to a 1 Mi-unit limit because POSIX awk length is locale/implementation character-counted | pass-4 disposition 9 fixed the design prose but left a stronger byte claim in the prompt | call it the 1 Mi-unit whole-payload scan ceiling, or run under a verified byte locale before using MiB +MAJOR | high | Task 6 Step 2 literal CLAUDE and README replacements | the replacements say an “unreadable third-party result” or “a result it cannot read at all” counts with disclosure, but the normative `no-result` class is precisely an unambiguously unobtainable result and does not count; only located text whose envelope cannot be interpreted is counted as `unrecognized` | the primary shipped explanations now invert settled decision 3 while trying to preserve decision 2, so pass-4 dispositions 23-24 introduced a new contradiction | say that no obtainable text is discarded, while a located nonblank text block whose envelope/anchor is unrecognized counts with disclosure +MAJOR | high | Task 7 Step 5 rollback behavioral probe | the old-hook mismatch branch ends with `echo "STOP: ..."`, so the subshell returns 0 and the plan continues after a failed rollback behavior check | release can proceed with an unverified or nonfunctional rollback despite the preceding hard-stop contract | make the mismatch branch exit nonzero and guard the subshell with an outer hard stop before cleanup or release proceeds +MAJOR | high | Task 7 Step 5 rollback release-byte check | the procedure requires `git show v0.7.1:...`, but the reviewed repository currently has no `v0.7.1` tag, so the mandatory rollback check deterministically stops before comparison | the high-risk plan has no runnable source of truth for the prior released bytes | record and verify the actual commit that introduced manifest version 0.7.1, or establish a verified release tag as an explicit prerequisite before using it in the rollback command +MAJOR | high | Tasks 1-6 commit steps versus Task 7 Gate B | six real non-WIP commits are created before the only Gate-B cycle, although CLAUDE.md section 5 requires tests green then Gate B then commit and reserves WIP-prefixed commits for pre-review snapshots | every intermediate commit closes and resets the hook cycle without review, so an interrupted or shared branch contains plugin changes that bypassed the mandatory gate even though the eventual range review may cover them | make intermediate snapshots WIP-prefixed, squash them before the reviewed WIP snapshot, then run Gate B and close by amend; alternatively keep changes uncommitted until the single reviewed snapshot +MINOR | medium | Task 1 Step 4 global-cache backup hash | both `BEFORE=$(shasum ... \| cut ...)` and the restore check read the pipeline status from `cut`, so a shasum that emits partial output and fails can still be accepted as a verified hash | the interruption-safety procedure can claim a globally installed payload-dumping hook was restored without having a successful checksum operation | capture shasum output and status directly, validate the parsed token separately, and use the same checked helper before mutation and after restoration +END OF FINDINGS (19 total) diff --git a/.context/codex-reviews/gate-a-plan-pass-5.2026-08-04-hardening-round.md b/.context/codex-reviews/gate-a-plan-pass-5.2026-08-04-hardening-round.md new file mode 100644 index 0000000..bee7bf5 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pass-5.2026-08-04-hardening-round.md @@ -0,0 +1,10 @@ +MAJOR | high | Task 6 Step 2 | the append leaves the source row's known-false premise that agents do not read ledger rows, although the story amendment and current skill establish that step 3 rereads the log and the defect is the decision branch ignoring guards | the backlog remains internally contradictory and the new evidence points at a diagnosis this round already disproved, recreating docs-drift | in the same row change, replace the stale premise with the decision-branch diagnosis before appending evidence; this still counts as one of the six existing-row changes +MAJOR | high | Task 0 Step 1 | the resumed-run path prints only an ahead count, yet the prose says the executor can confirm the range contains this plan's commits and nothing else | unrelated commits on an existing feature branch can be swept into the squash and PR with no detectable stop | print and inspect the actual main..feature commit list, and require an explicit stop when any commit is not part of this run +MINOR | high | Task 1 Steps 5-6 | the first battery runs before the plugin edit and version bump are committed, so check-version-bump.sh main compares commits that contain neither and its version-bump result is vacuous despite Step 5 calling the branch invocation a real check | the first kept battery can report green without testing invariant 12 at the point the plan presents it as the CI battery | state that this invocation's version-bump component is not evidence, and add or identify the first post-commit battery as the invariant-12 observation +MINOR | high | Task 6 Steps 3-7 | the snippets name a row but no quoted before-or-after anchor within it, while the task requires whole-block idempotency and the audit requires every insertion point to be locatable | executors can place fired, not-fired, or observation text on different sides of the existing trigger or mistake a partial independent edit for the completed mutation | give each step an exact surrounding anchor and say whether the new block is inserted before or after the existing trigger sentence +MINOR | medium | Tasks 4-5 creation protocol | the temp-file-plus-ln procedure never specifies cleanup of the temporary name on either success or ln failure | a successful hard link leaves an extra untracked story copy unless removed, and a collision stop can leave scratch content that later batteries scan or clean-tree checks reject | provide a POSIX trap-backed creation recipe that removes only the resolved temporary file on every exit after classification +MAJOR | high | Task 9 Step 1 | evidence creation has a classify-then-create race but, unlike the story paths, specifies no fail-if-it-appeared operation | another session can create foreign evidence after classification and an executor using a redirect can silently overwrite it, violating this step's own no-overwrite rule and contaminating durable commit evidence | create through a same-directory temporary file and atomic no-clobber link or another explicitly exclusive operation, with safe cleanup +MAJOR | high | Task 9 Step 6 | the guard's branch binding is only an unanchored grep inside the evidence file and never checks the currently checked-out branch or the exact Story line | text such as OldBranch: harden-0-8-0-and-pr-21 can pass, and the amend can mutate an unexpected branch before Step 7 finally stops | require exact fixed-line matches for Story and Branch and verify git rev-parse --abbrev-ref HEAD equals harden-0-8-0-and-pr-21 inside the guard-and-amend block +MAJOR | high | Task 9 Step 7 PR block | the PR body rereads a mutable ignored evidence path after the Step 6 validation window but checks only that cat succeeds and the result is nonempty | a concurrent replacement, symlink, foreign-cycle file, or reintroduced `, ``, ``, `` — inside the single-line encoding A7 requires. +- **22** — `NORESULT_MSG` turned provenance into causality, which spec §6 says the tool name cannot establish. It now says the checks **narrow but do not prove**. +- **23** — the unknown-tool prompt lacked the target-model prefix the other four carry. Added, and its `systemMessage` gains the operator action (see 26). +- **24** — the scaffolded prose promised a once-per-workspace disclosure categorically, which C2 and C4 forbid. Now "attempted and normally once … can be lost or repeated". +- **25** — "Everything else counts" dropped the routing precondition. Now "every other **routed** gate call". +- **26** — "absent" narrowed `no-result`, which also covers null, empty, non-array, non-object, non-text and blank. Now "no usable text". Both this and 25 are added to the load-bearing precision list, which grew from four to six. +- **27** — the counterfactual checked only for `FAIL` lines, so a crash or truncated run read as green. Status **and** completion marker are now required on both runs. +- **28** — the named verification replayed old captures and said it was not a live session, silently weakening story criterion 10. Now: run the live probe in an isolated profile for both variable states, **or obtain and record a human amendment to the criterion before Gate B**. +- **29** — the replay rows asserted nothing about fixture-read status, `sed` status, routing or class, so a missing fixture would print the expected result. Every row is now guarded, with an explicit expected value and an exact row count. +- **30** — side-by-side cache bytes are not an operable rollback. The activation path must be recorded, verified, and the active hook's bytes confirmed — or the release stops. +- **31** — no commit command carried the evidence entry. It is now in the reviewed snapshot's body from the start, revalidated after every fix, and closed with the same validated entry. +- **32** — **verified**: `is_wip_commit` greps for `-m … wip`, so `--amend --no-edit` reads as a real cycle-closing commit and **resets the counters mid-cycle**. Every fix now amends with an explicit `-m "WIP: …"`. +- **33** — `reset --soft "$BASE"` folded every commit since the base, including concurrent ones. Now guarded by an ancestry check and two reads that must be verified before the reset. +- **34** — no task checked the index for pre-staged files. Now a global constraint: `git diff --cached --name-only` must contain nothing outside the task's list. +- **35** — the final `check-version-bump.sh` had no argument, and `main` on this checkout compares HEAD with itself. Now `"$BASE"`, after the squash, with the exact command recorded. +- **36** — "after pushing" named no route and could have meant pushing WIP history to `main`; and CI evidence from before the Gate-B fixes would not cover the final bytes. Both addressed. +- **37** — the counterfactual's trap did not exit on a signal and left the evidence files outside cleanup. Now a cleanup function whose handlers exit, covering both files. diff --git a/.context/codex-reviews/gate-a-plan-pass-6-dispositions.md b/.context/codex-reviews/gate-a-plan-pass-6-dispositions.md new file mode 100644 index 0000000..0546361 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pass-6-dispositions.md @@ -0,0 +1,69 @@ +# Gate A — plan — pass 6 dispositions — **not clean; surfaced per the pinned exit** + +9 findings (1 BLOCKER, 5 MAJOR, 3 MINOR). **None dismissed. None applied.** + +The pinned exit was *clean or dispositions-only closes and opens execution; anything else comes +here.* These are real fixes, not dispositions, so it comes here. + +## Trend + +| Pass | 1 | 2 | 3 | 4 | 5 | 6 | +|---|---|---|---|---|---|---| +| Findings | 18 | 20 | 17 | 15 | 9 | **9** | +| Blockers | 1 | 0 | 1 | 1 | 1 | **1** | + +Flat at 9. But the **composition** changed, and that is the useful signal. + +## The code has one defect. The document has the rest. + +**One finding is about the checker, and it is real — confirmed by running it.** + +**F3 — the placement rule fails open when its terminator moves.** The scan finds the unique +`### 2.1` anchor but never validates the numbered heading that ends the range. Rename `### 2.2` +to something unnumbered and `intpl` stays true until `### 2.3`, so a canonical line sitting in +the *2.2* section counts as inside the template and the checker exits 0. **Verified:** renaming +that heading and planting the line in the resulting gap left the battery green. This is exactly +the anchor-drift class the amendment promised would fail loudly, and the second amendment's own +"holds no state across lines" (F9) is a false description of an `awk` that carries `intpl` +between lines. + +**One is a measured number I let go stale.** F4: deleting the 4c block now flips **16**, not +the recorded 13 — 15 reject fixtures plus the parser case. I measured 13, then added six +placement fixtures (three of them rejects) and never re-measured. That is documenting an +unverified result, the Don't this repository names, in the block whose entire purpose is to +carry measured evidence. + +**Two are Task 0, which I invented at pass 5 and got wrong twice over.** + +- **F1 (blocker):** Task 0 commits the spec, plan and scripts immediately before the WIP + snapshot — so that commit becomes the WIP parent, and a range `baseSha..HEAD` **excludes + baseSha itself**. Gate B would exclude the very artifacts Task 0 exists to include. +- **F2:** it is a **non-WIP commit of executable code** with the battery deliberately red and + no Gate-B loop. §5 requires tests green and Gate B before a non-prose commit, and the hook + reads a non-WIP commit as a cycle close. + +**The correct fix for pass 5's finding was one line, not a task:** put the spec, plan and both +scripts **into the WIP snapshot**, with `baseSha` = `c0a6ed2`. Task 0 should not exist. + +**The remaining four are documentation drift** — the self-review still describes the +pre-amendment checker (`severity_rule_count(file, mode)`, the `NOSTART`/`MULTIEND` sentinels, +the `sev-region-parser` marker, `CANON`), the file inventory says 15 `sev_case` cases where the +tree has 21, Task 4 Step 4 reopens a settled backlog decision with a rationale that is also +wrong (Task 1 *does* touch the acceptance rules), and Task 0's commit command omits the body it +requires. + +## What this says + +The **checker converged**: built, run, 145 assertions under `sh` and `dash`, one defect left, +and that defect was found by executing it rather than reading it — as were the two before it. + +The **plan document has not**, because it narrates code that keeps changing, and every change +leaves stale sentences behind. Six passes have not fixed that, and a seventh describing a +seventh version of the checker is not obviously different from the sixth. + +## Status + +Surfaced, not closed. Nothing applied from pass 6. The plan, the twice-amended spec, and the +built checker and suite are all uncommitted in the working tree; `c0a6ed2` is the only commit. +The battery is red by design — 4c fires because the canonical line is not yet in either prompt +copy. diff --git a/.context/codex-reviews/gate-a-plan-pass-6.2026-07-30-classifier-cycle.md b/.context/codex-reviews/gate-a-plan-pass-6.2026-07-30-classifier-cycle.md new file mode 100644 index 0000000..d69ccd7 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pass-6.2026-07-30-classifier-cycle.md @@ -0,0 +1,38 @@ +BLOCKER | high | Task 2 Step 7 expected-message builders | the plan still provides no literal assignment for any `*_EXPECTED` constant, and its inventory is internally inconsistent because the listed five new pairs plus nine pre-existing pairs total 28 constants, not the claimed thirty | `set -u` aborts the suite before the composition goldens, while the proposed inventory loop can only check names already referenced and cannot detect an omitted pair | write every literal assignment into the plan, correct the pair and constant counts, and compare an explicit required-name inventory with the assigned-name inventory +BLOCKER | high | Task 5 Step 11 `marker_rows` | the promised dual-emitter implementation is still a skeleton containing `# ... every remaining row ...`, followed by the old blocks that call `run`, `run_closed`, and raw `sh` directly | the pass-5 disposition claims parameterization that is not in the plan, so most marker-table rows still run only through the jq emitter | write the entire function body with every row using its runner parameters, remove the duplicate direct blocks, and make the two invocations the sole execution path +MAJOR | high | Task 5 Step 11 C2 direction 2 | `$SH_RUN` is neither passed to `marker_rows` nor defined by executable code, and the proposed jq-free scalar value beginning with `PATH=...` cannot become an assignment prefix when expanded as `$SH_RUN` | the test aborts under `set -u` or tries to execute a command literally named `PATH=...`, leaving one accepted silent-count residual untested in the fallback emitter | use runner functions that perform the redirection and return the hook status, then pass those function names explicitly to `marker_rows` +BLOCKER | high | Task 5 Step 7 sed-failure tests | `nosed`, `sedfail`, and `sedpart` are consumed by `eval` but no shim definitions are present anywhere in the plan | the first iteration aborts under `set -u`, so the claimed three sed fault shapes and their fail-open behavior are not tested | provide complete selective shim construction and self-checks for all three modes before the loop +BLOCKER | high | Task 2 Step 3 restricted-PATH construction | the exact snippets `mk_path nojq $HOOK_CMDS` and `mk_path noawk $(...)` trigger SC2086 and SC2046 under the required ShellCheck 0.11.0 policy | Task 2 cannot pass the battery the plan requires before its WIP commit | construct the command list as positional arguments without unquoted expansion or command substitution splitting, then run the exact repository ShellCheck command on the resulting suite +MAJOR | high | Global Constraints and Task 5 Step 11 dash coverage | the plan says marker-write faults are asserted under both `dash` and `sh`, but no task invokes the hook suite with `dash`; only the ignored locator/matcher verifier was run under both shells | the pass-5 disposition again claims work absent from the text, and the special-builtin regression that motivated this requirement is specifically shell-dependent | add an explicit `dash plugins/dev-workflow/hooks/codex-gate.test.sh` run locally and in the named battery or narrow the claim to the actual CI-on-Ubuntu evidence +MAJOR | high | Task 5 Step 12 composition matrix | `run_scenario` seeds `gateb_stale`, `gateb_floor`, `gateb_ok`, `gatea_floor`, and `gatea_ok` with successful gate calls after the pending disclosure is created and the off-switch removed | the first setup call flushes and clears the debt with output redirected away, so the target branch cannot emit the expected `Earlier:` composition and five matrix rows fail for a harness reason | keep suppression active during scenario setup or split setup from the single observed event, then remove suppression only immediately before that event +MAJOR | high | Task 5 Steps 4 and 12 jq-free field assertions | `field_of` returns JSON-escaped bytes without jq, but the composition matrix and the dedicated jq-free composed assertions compare them directly with decoded constants; only `golden()` applies `esc_like_emitter` | quote-bearing failure and disclosure messages fail exact comparison on a machine that genuinely lacks jq, obscuring whether the hook output is correct | route every exact field comparison through one helper that decodes with jq or escapes the expected value consistently in fallback mode +MAJOR | high | Task 5 Step 3 real review capture assertion | `shape0-success-review` is checked only for count and fingerprint state, which are identical for `success` and `unrecognized` | a malformed or drifted real review fixture can pass while being counted through the fail-open terminal class, contrary to spec §7.3 and the pass-1 disposition | also assert `class_of` is `success` and that neither disclosure marker is created +MAJOR | high | Task 3 Step 4 selective sed shim | the `else` branch prints `skip`, but the three assertions remain after the `fi` and still execute against an absent or unverified shim | a platform unable to construct the fault injector records a skip and then produces misleading failures or exercises routing failure instead of encoder failure | put shim verification and all dependent assertions inside the successful branch, and make the skip branch bypass them completely +MAJOR | high | Task 1 Step 4 cache restoration | the normal path calls `restore_cache` explicitly, deletes `$BACKUP` on success, and then the still-armed EXIT trap calls it again | every successful capture emits `STOP: cache NOT restored`, and signal exits also run restoration twice, making the safety result non-idempotent and untrustworthy | track a restored flag or disarm only the EXIT trap after a successful verified restore while retaining the backup on any failure +MAJOR | medium | Task 1 Steps 3 and 7 locator provenance | Step 7 assumes `$L` and its temp file from Step 3 still exist across separate checklist steps and shell code blocks | a normal implementation session can run the steps in distinct shells or after the Step-3 EXIT trap, causing an unset-variable abort or a missing locator rather than provenance verification | recreate the locator temp file in Step 7 or combine both checks into one explicitly single-shell procedure with one cleanup boundary +MAJOR | high | Task 1 Step 4 review fixture capture | the globally mutating procedure still replaces its critical operation with an ellipsis and merely prefers isolation when available, while the dump hook records every concurrent MCP payload on the machine | the procedure is not reproducible and can capture unrelated prompts, paths, or review content across the external Claude Code boundary, which is a high-risk security and privacy failure | make an isolated profile or scratch install mandatory, provide the exact insert, single-call correlation, collection, and restore commands, and stop if isolation cannot be established +MAJOR | high | Task 5 A2 scan ceiling [risk high — abuse; security standard — abuse paths, trust boundary, external system] | the external MCP server controls `tool_response`, while `skipval` grows and truncates a string stack and the plan admits quadratic-ish work up to 1 Mi-unit after the whole payload is already buffered | a valid deeply nested response can consume extreme CPU and wedge the advisory commit hook; the ceiling does not bound the work soon enough | replace the string stack with indexed depth storage or impose a conservative nesting or operation cap that returns `unrecognized`, and add a terminating adversarial-depth test without expanding JSON validation +MAJOR | high | Task 5 A2 duplicate-member contract | the prose says repeated `type` or `text` is refused inside the selected element, but the embedded locator exits 2 for duplicates in every object it examines before selection, including a non-text block preceding a valid text block | the implementation rejects payloads outside the stated ambiguity boundary and contradicts the first-text-element selection rule | either narrow the code to reject duplicates only in the element that would be selected or explicitly amend the contract and add a non-selected-duplicate fixture for the stricter behavior +MAJOR | high | Task 5 Step 1 `FAILURE_CTX` | `CODEX_EXECUTION_FAILED` is stated to mean the call never started, but the pinned server uses that generic code for abort-before-start, abort-after-start, and child execution errors | Claude can apply the wrong remedy and overwrite or collide with artifacts from a call that did start | describe the code generically, distinguish causes using the available error message or other evidence, and preserve the one-retry stop without claiming start state from the code alone +MAJOR | high | Task 5 Step 1 failure message pair [risk high — observability; security standard — roles] | `FAILURE_CTX` tells Claude to raise the executor timeout while the plan says operator-only configuration belongs in `systemMessage`, and `FAILURE_MSG` incorrectly claims there is no operator action | the model may lack authority to change the server timeout while the operator receives no actionable configuration diagnosis | place the verified timeout-setting check and remedy in `systemMessage`, and tell Claude only the retry or escalation action it can actually perform +MAJOR | high | Task 5 Step 1 `UNVERIFIED_MSG` versus approved spec §6 | the plan calls unmapping or replacing an unreadable mapped third-party tool a user-side fix, while the approved spec explicitly classifies that cause as having no user-side fix | the shipped prompt and implementation plan no longer implement the settled design | reconcile the spec and message before implementation; absent a human amendment, keep the spec's no-user-fix disposition verbatim +MAJOR | high | Task 5 Step 1 `UNVERIFIED_MSG` POSIX diagnostics [risk high — compatibility] | `awk --version` and `sed --version` are non-POSIX checks, and healthy BSD `sed` exits nonzero for `--version` | a supported macOS installation is falsely diagnosed as broken by the recovery prompt for invariant 4 | use `command -v` plus small portable functional probes that exercise the exact awk locator and sed substitution behavior +MAJOR | high | Task 5 Step 1 `UNVERIFIED_MSG` diagnostic grouping [risk high — observability] | the text promises a check for the ceiling and ambiguity causes but supplies no check that distinguishes either condition; saying jq changes neither is not a diagnostic | prompt-standard item 10 remains unmet and the pass-5 disposition overstates the rewrite | give a bounded size check and a reproducible locator-status probe with the exact payload capture/report path, or stop claiming those causes have distinguishing checks +MAJOR | high | Task 5 Step 1 five model-facing messages | the five `additionalContext` values are single flowing shell-string paragraphs with inline labels, not context, task, rules, and output sections separated by headings or XML tags | they do not satisfy prompt-standard item 5 despite the plan's categorical all-12 conformance claim | structure each model-facing prompt with compact tagged sections while preserving the no-literal-newline encoding constraint +MAJOR | high | Task 5 Step 1 `NORESULT_MSG` [security standard — trust boundary, external system] | the message says that finding a mapping or third-party server proves that tool returned empty or non-text content, although the approved spec says tool identity cannot distinguish that cause from a hooks-API payload contract change | the diagnostic converts provenance evidence into a causal conclusion it cannot support and can send the operator down the wrong repair path | report both causes when provenance is mapped, state that the checks narrow but do not prove causality, and avoid the categorical `that tool returned` wording +MAJOR | high | Task 6 Step 4 unknown-tool prompt | the literal insertion changes the existing `additionalContext` but leaves it opening with the current `Codex tool ...` text rather than the target-model prefix applied to the other changed hook prompts | this shipped prompt still fails prompt-standard item 1 while Task 6 claims all edited prompts receive golden and all-12 review | add the same Claude Code target prefix to the unknown-tool `additionalContext` and update both branch and composition goldens +MAJOR | high | Task 6 Step 2 scaffolded disclosure prose | `with a once-per-workspace disclosure` and README's `says so once` are categorical, but spec C2 and C4 allow the disclosure to be lost, repeated, or counted silently | the scaffolded workflow overclaims a best-effort diagnostic as enforcement, violating prompt-standard item 11 | say disclosure is attempted and normally once in sequential no-failure operation, with the accepted loss and duplication residuals scoped explicitly +MINOR | medium | Task 6 Step 2 literal CLAUDE replacement | `Everything else counts` is broader than the classifier state table because an unroutable malformed payload exits without touching pass state | the rewritten decision procedure drops the routing precondition and teaches a false totality claim | scope the sentence to routed gate invocations whose result reaches classification +MAJOR | high | Task 6 Step 2 README replacement | `failed, backgrounded, or absent` narrows `no-result` to absence, but that class also includes null, empty, non-array, non-object, non-text, and blank result content | users can infer that present-but-unreadable results count even though the hook discards them | replace `absent` with `no usable result text could be obtained` and keep the unrecognized-text distinction explicit +MAJOR | high | Task 7 Step 2 counterfactual | both suite commands are piped through `grep`, and the procedure checks only whether `FAIL` lines were captured, never the suite's exit status or completion marker | a crash, `set -u` abort, missing dependency, or truncated run with no matching line is accepted as a green new suite, while a partial old run can masquerade as valid counterfactual evidence | capture full output and command status separately, require the new suite to exit 0 and print its completion signal, and require the old suite to terminate normally with the exact expected defect failures +MAJOR | high | Task 7 Step 3 named verification versus story acceptance criterion 10 | the story requires the probe methodology to be re-run with shape 3 under the variable absent and set, but the plan only replays old captured payloads and expressly says it is not a fresh end-to-end session | the required `+verification` evidence is not produced, and the plan silently weakens a live acceptance criterion after five rewrites | run the live probe in an isolated profile against the changed hook for both variable states, or obtain and record a human amendment to the story criterion before Gate B +MAJOR | high | Task 7 Step 3 named verification | each negative row feeds a `sed` pipeline into an always-zero advisory hook and makes no assertion about fixture-read status, substitution status, routing, exact class, or expected row count | a missing fixture or failed `sed` supplies an empty payload and still prints the expected absent counter and fingerprint, so most verification evidence can be false while looking correct | guard every input transform, assert the intended tool and class, compare every row with an explicit expected value, require the exact row count, and fail the step on any mismatch +MAJOR | medium | Task 7 Step 5 rollback [risk high — rollback] | the executable check verifies old bytes and old behavior but never verifies that an operator can activate 0.7.1; the actual version switch is deferred to unspecified environment prose | side-by-side cache bytes are not an operable rollback, especially because cache paths are an implementation detail and the marketplace command has no version syntax | make a tested activation or downgrade procedure and active-hook byte check mandatory evidence, and stop the release if the environment has no verified switch path +MAJOR | high | Task 7 Step 7 Gate-B snapshot | neither WIP commit command nor the closing amend command includes the story's validated evidence entry in the commit body | this contradicts CLAUDE.md §5, and the final `-m` replacement would erase any body added informally before Gate B | create the reviewed WIP snapshot with the current evidence body, revalidate and update that body after fixes, and close with a final subject plus the same validated entry +BLOCKER | high | Task 7 Step 7 fix loop | fixes use `git commit --amend --no-edit`, but the shipped `is_wip_commit` recognizes WIP only from a `-m` argument containing `wip` | the first fix amend is treated as a real cycle-closing commit and PostToolUse resets the counters, so the stated three passes cannot keep accumulating | amend every fix with an explicit `-m "WIP: ..."` subject or first extend and test WIP recognition for `--no-edit` by reading the current commit message safely +MAJOR | high | Task 7 Step 7 soft squash [risk high — data loss, idempotency, concurrent change] | `git reset --soft "$BASE"` unconditionally folds every commit since the recorded base into the feature snapshot without checking ancestry or authorship | concurrent or unrelated commits can be rewritten into the squashed change, and rerunning the step can widen the range again | assert `$BASE` is an ancestor, enumerate and approve every commit and changed path in the range, and stop if the history differs from the expected WIP set +MAJOR | high | Tasks 1–7 staging [risk high — data loss, concurrent change] | exact `git add` path lists do not check for files that were already staged before a task begins | unrelated user changes can enter each WIP commit and the later squash even though the plan never named them | inspect the effective index before each commit, require an allowlisted staged-path set, and stop for pre-existing staged content rather than committing it +MAJOR | medium | Task 7 Step 6 invariant-12 check | the final battery says `check-version-bump.sh` uses the branch's real base ref but gives no exact argument even though the work is on `main`, where passing `main` is explicitly known to succeed trivially | the release evidence can claim invariant 12 without comparing the reviewed feature range | run `sh scripts/check-version-bump.sh "$BASE"` after the squash commit and record that exact successful command +MAJOR | high | Task 7 Step 4 awk portability evidence [risk high — compatibility, observability] | `after pushing` names no branch or PR workflow even though the work is on `main`, and CI evidence obtained before Gate-B fixes would not necessarily cover the final amended awk bytes | the plan either encourages pushing unreviewed WIP history to main or cites mawk evidence from a different snapshot | define a safe branch or pinned local mawk route, bind the CI result to the exact reviewed commit tree, and rerun it after any classifier or locator fix before closing +MINOR | high | Task 7 Step 2 cleanup trap | `trap 'rm -rf "$t"' EXIT INT TERM HUP` does not exit after a caught signal and leaves `$EVID` and `$EVID.new` outside the cleanup target | an interrupted evidence run may continue with a deleted fixture tree and leave stale temp evidence that can be mistaken for the current run | use a cleanup function with signal handlers that exit nonzero, remove both evidence files, and disarm or make cleanup idempotent after normal completion +END OF FINDINGS (37 total) diff --git a/.context/codex-reviews/gate-a-plan-pass-6.2026-08-12-ledger-supersession.md b/.context/codex-reviews/gate-a-plan-pass-6.2026-08-12-ledger-supersession.md new file mode 100644 index 0000000..88010dd --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pass-6.2026-08-12-ledger-supersession.md @@ -0,0 +1,3 @@ +MAJOR | high | Read pass — Task 8 Steps 4–7 | the plan restates §5's evidence, re-review/amend, and closing protocol despite the standing cite-only rider, and its copied evidence-entry recipe omits the story path that §5 says the entry itself carries while Step 7 lists that path beside the entry | the durable commit can contain both strings without satisfying the required story-to-evidence association, and this is another protocol mirror already drifting from its authority | replace the generic §5 instructions with precise citations and retain only this change's resolved values, including the story path as a value supplied to the cited evidence-entry procedure +MAJOR | high | Read pass — fixture matrix and Task 4 Steps 1–3 | the fixtures have no rule that each negative case starts from a named green baseline and changes only its advertised dimension; in particular `rows-two-matching` says the rows share "the entry's date" even though an entry has its own `$D` date and a distinct locator row date | a fixture can receive its expected failure from an unrelated locator or parse defect, and constructing the two rows with `$D` would not match the mandated `2026-07-20` locator, so the run would not prove the oracle it claims | define a common valid post-change fixture baseline, require one advertised mutation per negative fixture, and state that `rows-two-matching` duplicates complete rows matching the mandated locator's row date and fingerprint while the entry date remains `$D` +END OF FINDINGS (2 total) diff --git a/.context/codex-reviews/gate-a-plan-pass-6.md b/.context/codex-reviews/gate-a-plan-pass-6.md new file mode 100644 index 0000000..77204c2 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pass-6.md @@ -0,0 +1,10 @@ +BLOCKER | high | plan Task 0 lines 70-103; Global Constraints lines 42-66; Gate B lines 591-606 | Task 0 commits the amended spec, plan, checker, and suite immediately before Task 1 creates the WIP, so that Task 0 commit becomes the WIP commit's parent; a review range whose baseSha is that parent excludes the parent commit itself, exactly opposite the stated reason for Task 0 | Gate B would review the later prompt/plugin diff while excluding the authoritative amendments and executable 4c changes the plan says must be included, so a clean pass cannot cover the claimed implementation | Put the Task 0 artifacts into the WIP snapshot and use their pre-change commit as baseSha, or explicitly use Task 0's parent as baseSha and revise every WIP-parent derivation, guard, version-range check, and evidence statement to match that wider range +MAJOR | high | plan Task 0 Step 2 lines 91-103 | The prescribed non-WIP commit contains executable changes to both scripts while the repository battery is deliberately red and no Gate-B loop precedes the commit | CLAUDE.md section 5 requires tests green and Gate B before committing a nontrivial non-prose diff, and a non-WIP commit is also interpreted by the hook as a real cycle close; Task 0 therefore violates the procedure it is meant to preserve even if the later review range is widened | Do not make a real Task 0 commit containing the scripts; make the first snapshot a WIP containing all Gate-A-approved artifacts, or otherwise run and close a valid Gate-B cycle for the executable Task 0 commit before proceeding +MAJOR | high | scripts/check-invariants.sh lines 497-505 and 539-545; spec section 5.2 lines 469-474; plan lines 150-156 | The placement scan counts the 2.1 anchor but never counts or validates its terminating numbered heading; if the immediate 2.2 heading is renamed to an unnumbered heading, intpl remains true until 2.3, so a canonical line placed in the 2.2 section is misclassified as inside 2.1 and the checker exits 0 | The asserted placement guarantee can fail open when the boundary moves, allowing workflow-init to pass while the CLAUDE.md template it scaffolds still lacks the rule; this is the same anchor-drift class the amendment says must fail loudly | Track the first numbered terminator after the unique 2.1 anchor, require it to exist in the expected structural position, and add reject fixtures for a missing or renamed immediate terminator and for a canonical line in the resulting expanded gap +MAJOR | high | plan lines 249-269 and 690-692; scripts/check-invariants.test.sh lines 341-370; spec section 5.2 lines 518-524 | The recorded 4c mutation result is stale: deleting the marked 4c block on this non-root checkout flips 16 assertions, comprising 15 reject fixtures plus the parser-failure fixture, not 13 comprising 12 rejects plus the parser; the three new placement rejects are precisely the missing delta | The plan and approved spec cite a run result that the current suite does not produce, undermining the load-bearing-evidence claim and violating the rule against documenting an unverified command result | Re-run the documented mutation procedure after the placement fixtures, record 16 with its 15-plus-1 composition and root-dependent unreadable-case caveat, and update the spec, plan, and test evidence block together +MINOR | high | plan Task 1 file inventory lines 116-123 | The plan says the suite has 15 sev_case cases and one inject_case for 4c, but the tree has 21 sev_case invocations for 4c plus its inject_case, totaling the correctly cited 22 assertions | The contradictory inventory can cause an implementer or reviewer to omit the six cases added by the second amendment while still believing the task's stated inventory was verified | Change the inventory to 21 sev_case cases and one 4c inject_case, ideally naming the six post-amendment cases that raised the count +MAJOR | high | plan Self-review lines 684-692 | The self-review describes an obsolete checker: the suite variable is SEV_LINE rather than CANON, the function is severity_rule_scan(file) rather than severity_rule_count(file, mode), it emits three integers rather than count or NOSTART/MULTISTART/NOEND/MULTIEND sentinels, and the injection marker is sev-canon-count rather than sev-region-parser | This section claims type and branch coverage for symbols and branches that do not exist, directly contradicting the real code and the plan's earlier accurate description, so it cannot serve as the claimed final self-review | Rewrite the self-review from the current implementation and enumerate the actual missing-file, unreadable-file, awk-status, whole-count, anchor-count, and in-template-count branches and their fixtures +MAJOR | high | plan Task 4 Step 4 lines 495-508; approved spec section 6 lines 586-589 | The approved spec settles the slot-collision item as trigger fired and row stays open, but the plan reopens that decision as a judgement and supplies a not-fired rationale; that rationale also says acceptance rules are unchanged even though Task 1 adds the severity reader directly to pass acceptance | Execution can record another NOT FIRED disposition contrary to the settled decision and lose the approved meaning of the second observed destruction of findings and dispositions | Make Step 4 deterministic: amend the existing row with the second occurrence, state TRIGGER FIRED and row stays open exactly as spec section 6 requires, and remove the contrary no-acceptance-rule-change rationale +MINOR | high | plan Task 0 Step 2 lines 91-103 | The command uses a single git commit -m subject, but the following paragraph requires the deliberately red-battery explanation to be in the commit body | Following the command literally omits evidence the plan explicitly requires, leaving history unable to distinguish an intentional red transitional state from an accidental broken commit | Add an explicit second message paragraph to the command or provide the complete subject-and-body commit message +MINOR | high | approved spec section 5.2 lines 491-500 | The amendment says the heading range holds no state across lines, but the implemented awk necessarily carries the intpl flag from the 2.1 line through later lines until a numbered heading | This is a false description of the exact comparison and state machine in the authoritative spec, violating the gate-proof calibration rule and obscuring the missing-terminator failure mode | Replace the no-state claim with the exact state transition the awk performs and state how absent, renamed, and duplicated boundaries are handled +END OF FINDINGS (9 total) diff --git a/.context/codex-reviews/gate-a-plan-pass-7-dispositions.2026-07-30-classifier-cycle.md b/.context/codex-reviews/gate-a-plan-pass-7-dispositions.2026-07-30-classifier-cycle.md new file mode 100644 index 0000000..2934b04 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pass-7-dispositions.2026-07-30-classifier-cycle.md @@ -0,0 +1,56 @@ +# Gate A (plan) — pass 7 dispositions + +Advisory companion to `gate-a-plan-pass-7.md` (4 BLOCKER, 18 MAJOR, 4 MINOR = 26). **All +accepted, none dismissed.** First pass on the narrowed plan (531 lines, was 1732). + +Trend: 40 → 32 → 35 → 29 → 19 → 37 → **26**; B+M 36 → 28 → 27 → 24 → 17 → 35 → **22**. + +**The narrowing worked as intended, and rider 1 earned its keep.** Six findings (3, 4, 6, 7, +10) are *dropped-coverage* findings — assertions earlier passes established that no oracle +still required. That is exactly what "every test keeps its label and its oracle" was for, +and it caught what the rewrite lost. No finding said "the harness shell is missing", which +was the scope risk. + +**Four verified in-session before acting:** + +- **`locate.awk` contained two ASCII apostrophes** (`caller's`, `machine's`) — in comments I + added at pass 5 when narrowing the ceiling claim. The plan says to embed it in a + **single-quoted** shell variable; either one would terminate the quote and leave the hook + unparseable. Removed; the file is now apostrophe-free and the constraint is recorded in + the drafts README as a trap. +- **`git commit -m … -F …` is rejected outright** — *"options '-m' and '-F' cannot be used + together"*. The squash command could not have created the reviewed snapshot at all. +- **`printf x | sed s/x/y/` emits `y` with no newline**, so my own health probe fails its own + "must print a line" requirement — it would report a healthy POSIX `sed` as broken. +- **The battery is four script runs, not three.** Miscounting is how + `check-invariants.test.sh` or `check-version-bump.test.sh` gets dropped while the task + still claims the full battery. The block is now written out verbatim. + +## Dispositions + +1. BLOCKER apostrophe in the embedded locator — ACCEPT, verified. Removed, plus a mechanical check after embedding (`sh -n`, `shellcheck`, one classification test) and the constraint recorded where a reconstruction would read it. +2. MAJOR `locate_result`'s contract vanished in the narrowing — ACCEPT. Restated: feed `$payload` unchanged to `LOCATE_AWK`, return `awk`'s status untouched, never consult `jq`, never collapse "anything else" — collapsing it is how a missing `awk` (127) becomes `no-result` instead of fail-open. +3. MAJOR the preceding-non-text duplicate was claimed verified and had no row — ACCEPT; my pass-6 disposition overstated again. Row added to `verify.sh` (green), plus its own oracle. +4. MAJOR block selection had no oracle — ACCEPT. An `element [0]` implementation passed every listed group while violating spec §3.1's settled rule. Oracle added. +5. MAJOR the three frozen walkable-invalid rows cannot route with `jq` — ACCEPT, and the contradiction was real: the plan required them ported *and* excluded jq-routing-dependent documents. They are now **locator-level contract tests, exempt by label**, which is the level their claim is about. +6. MAJOR the permanent slice test lost its home — ACCEPT. Task 1 deferred it to Task 5 and Task 5 had no such oracle. Added, one per fixture, requiring both locator statuses before comparing. +7. MAJOR the ported set is named by group, not by label — ACCEPT. The plan now requires enumerating every `verify.sh` label as ported or excluded-with-reason; a group heading cannot show that a case quietly stopped being required. +8. MAJOR thirteen scenarios named, fourteen sources listed — ACCEPT. Nine existing branches plus five new emissions. Counted explicitly, because an omitted branch is the dropped-output failure the oracle exists to catch. +9. MAJOR the Unicode row's wording recreates pass-5's wrong test — ACCEPT. It now says six-byte `"`, no raw quotes, and asserts the locator succeeded before the matcher's verdict. +10. MINOR the depth cap had no verifier row — ACCEPT. Added at 201 openers, plus an ordinary-nesting row so the cap cannot be tightened into refusing real payloads. +11. MAJOR retries omit CLAUDE.md §5's target-file cleanup — ACCEPT, and this is the sharpest finding of the pass: all four retry-capable messages allowed a re-run without deleting the target findings file and confirming it gone, which is how a died-part-way call's valid-looking artifact survives into the next pass. Added to each ``, after the stop-or-await for the backgrounded pair — deleting a slot the original call can still write is the race those messages exist to prevent. +12. MINOR the `sed` probe prints no line — ACCEPT, verified. Both probes now use `printf 'x\n'` with the exact expected output and status. +13. MAJOR oversize and ambiguous are not distinguishable — ACCEPT. The message now says so, gives the one check that *is* available (measure against the ceiling), and treats the rest as unresolved rather than claiming a discriminator it lacks. +14. MAJOR the message told the operator to report a raw payload — ACCEPT, and Task 1 establishes exactly why: those payloads carry prompts, absolute paths, review content, session identifiers and unrelated concurrent call data. Now: keep it local and access-restricted, sanitize before showing anyone, never attach it unsanitized. +15. BLOCKER `-m` with `-F` — ACCEPT, verified. Two `-m` arguments: the first carries the `WIP:` subject `is_wip_commit` matches, the second the evidence body. +16. BLOCKER the CHANGELOG never reaches the snapshot — ACCEPT. `reset --soft` preserves the unstaged edit and the commit omits it, so invariant 12's release note would sit outside the reviewed range and the commit-based path audit could not see it. Staged explicitly before the reset, with a post-commit `show --stat` read. +17. MAJOR Step 6 ran the post-squash battery before Step 7 created the squash — ACCEPT. Resequenced: squash (6), battery (7), Gate B (8). +18. MAJOR no amend form carries subject and evidence together — ACCEPT. `-m` alone erases the body; `--no-edit` is not recognized as WIP and resets the counters. Only the two-`-m` form satisfies both, and it is now written out for the fix loop and the close. +19. BLOCKER the rollback drill leaves 0.7.1 active — ACCEPT, and the consequence is severe: every Gate-B pass would then run under the hook that counts failed calls, so the review's pass accounting is the accounting this change exists to fix. Either an isolated profile, or reactivate 0.8.0 and **byte-verify** before the first pass. +20. MAJOR the rollback hides what it costs — ACCEPT. Reverting restores the original false ✓ **and** the `dash` invariant-1 exit. Both recorded, with the narrow condition under which rollback is the right call. +21. MAJOR `$BASE` used before it is recorded, and not durable — ACCEPT. Moved to a new **Task 0**, recorded as a literal SHA in a durable note rather than a shell variable that lived in one terminal, and ancestor-validated before each use. +22. MAJOR the battery block is referenced and never written — ACCEPT, and it undercounted. Written out verbatim, with the count corrected and a note that the `AGENTS.md` row remains the source of truth. +23. MAJOR the four counterfactual labels are never named — ACCEPT. Named, with only the mechanical runner suffix left to execution, and with why those four: each fails for the change's own reason rather than a harness difference. +24. MAJOR the working tree is not clean at execution start — ACCEPT, and it is true right now: the spec amendment and this plan are unstaged. **Task 0** commits the approved Gate-A artifacts separately (prose, so Gate B is N/A) and stops on anything unaccounted, before `$BASE` is fixed. +25. MINOR the drafts README still says the bodies live in the plan — ACCEPT; it predated the narrowing by one pass. Rewritten: these files are now the **only** copy, with the reconstruction fallback stated. +26. MINOR the README's A2 contract is weaker than the code — ACCEPT. Duplicate refusal in **any** examined element and the depth cap both added, so a reconstruction guided by it cannot drop two settled conditions. diff --git a/.context/codex-reviews/gate-a-plan-pass-7-dispositions.2026-08-12-ledger-supersession.md b/.context/codex-reviews/gate-a-plan-pass-7-dispositions.2026-08-12-ledger-supersession.md new file mode 100644 index 0000000..bf0d1fb --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pass-7-dispositions.2026-08-12-ledger-supersession.md @@ -0,0 +1,28 @@ +# Gate A — plan — pass 7 dispositions (CLOSING PASS) + +Artifact: `docs/superpowers/plans/2026-08-12-hardening-ledger-supersession.md`. +2 findings: **2 Minor. Zero Blocker, zero Major.** Both valid. **Both applied.** + +Reviewer: `mcp-codex-dev@1.0.1`, `codex-cli 0.144.5`, model `gpt-5.6-sol`. Pass file valid: +3 lines, 2 finding lines, terminator exact. + +Plan passes: **18, 11, 9, 8, 6, 2, 2.** Blockers: **2, 1, 4, 3, 1, 0, 0.** + +| # | Verdict | Applied as | +|---|---|---| +| 1 | **Valid.** Task 6's Interfaces pointed the `0.8.2` closing-body deliverable at Task 8 **Step 6**, which amends fixes into the WIP; the real closing body is Step 7. The reference sent an auditor to a step that cannot establish the claim. | Reference corrected to Step 7. | +| 2 | **Valid.** Self-Review mapped spec §4 to Tasks 3, 4, 6 and 7 and named the three pre-landed files, omitting the **governing story's amendments** — which §4 requires and Task 1 commits. The plan's own coverage summary disagreed with its task mapping, making a required change look unaccounted for. | Task 1 added to the §4 mapping, with the three outside-scope files named explicitly. | + +## Why this closes the gate + +§5's loop is a **Blocker/Major** loop — Minor and Nit are collected, never iterated — so a pass +returning none is its exit condition. Both Minors were applied rather than collected because each is +a one-line cross-reference fix, not a judgement call worth carrying forward. + +Neither is the cycle's dominant failure mode. Both are internal cross-references that were correct +when written and went stale when a step was inserted (Task 8 gained a step at pass 4) or a task's +scope widened (Task 1 gained the plan at pass 1) — drift *behind* a fix rather than a fix landing in +one site only. That class produced a Blocker in each of passes 3, 4 and 5 and has produced none +since it was named directly in the pass-5 and pass-6 prompts. + +Closure record: `gate-a-plan-CLOSURE.md`. diff --git a/.context/codex-reviews/gate-a-plan-pass-7.2026-07-30-classifier-cycle.md b/.context/codex-reviews/gate-a-plan-pass-7.2026-07-30-classifier-cycle.md new file mode 100644 index 0000000..25d9a5f --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pass-7.2026-07-30-classifier-cycle.md @@ -0,0 +1,27 @@ +BLOCKER | high | Task 5 Step 2, `LOCATE_AWK` installation | the plan says the locator can be embedded in a single-quoted shell variable because it contains no single quote, but `.context/plan-drafts/locate.awk:59` contains the ASCII apostrophe in `caller's` | copying the verified seed as directed produces an unterminated shell quote, so Task 5 cannot parse or pass ShellCheck | remove or escape the apostrophe, or specify a safe materialization mechanism, and require `sh -n` plus ShellCheck on the embedded result +MAJOR | high | Task 5 Step 2, `classify()` | `classify()` calls `locate_result`, but the narrowed plan neither defines that function's contract nor points to a seed that contains it | an executor can leave the command undefined, making every located result fall through status 127 to `unrecognized` and count with disclosure; pass-1's disposition claim that this was defined has silently disappeared | specify that `locate_result` feeds the unchanged payload to the fixed awk program, preserves statuses 0, 1, and all other statuses, never consults jq, and add an end-to-end success/failure oracle +MAJOR | high | Task 5 A2 and classification oracles | A2 claims a duplicate `type` or `text` in a non-text element preceding a valid text block was verified, but `verify.sh` has no such row and the generic duplicate oracle can be satisfied by its existing selected-element cases | the stricter any-examined-element behavior can regress while all named coverage remains green, repeating a pass-6 disposition that claimed verification not present | add the exact preceding-non-text-duplicate payload and oracle, followed by a valid success text block, expecting `unrecognized` +MAJOR | high | Task 5 classification oracle table versus spec §7.3 Block selection | the narrowed table has no oracle for a non-text first response element followed by a real text element | an element-zero implementation can pass the listed no-result and capture groups while violating the settled first-text-element rule | add a named routed test whose first element is non-text and second is a failure or success text block, with the corresponding class as its oracle +MAJOR | high | Task 5 classification lines 338-353 | the three frozen walkable-invalid documents are required to be ported through the hook, yet they are malformed outer JSON whose routability depends on jq, while the next paragraph says every jq-dependent-routing document stays in the drafts | with jq they cannot reach classification and without jq they can, so the current requirements cannot produce one non-vacuous portable hook test for the settled fixtures | explicitly run these three as locator-level contract tests or through a named jq-free routed runner, and exempt them by label from the general hook-routing rule +MAJOR | high | Task 1 Step 7 and Task 5 oracle tables | Task 1 says the permanent payload-to-response-slice byte-identity test belongs in Task 5, but Task 5 contains no such label or oracle | later fixture or slice drift can be class-equivalent and leave every classification test green while the claimed byte-exact duplicated representation is false | add one permanent per-fixture oracle that checks both locator statuses and fails on any located-block byte difference +MAJOR | high | narrowed test-coverage inventory | the plan promises every test's label and oracle, but the classification section only says 53 rows are ported and then names broad groups; it never enumerates which `verify.sh` labels are retained versus the two excluded groups | cases such as non-object-element skipping and array-walk boundaries can disappear without any reviewable indication, violating AGENTS.md's decision-procedure replacement rule | list every retained verifier label with its expected class and list every excluded label with its explicit disposition +MAJOR | high | Task 5 Messages and composition, `all thirteen` | Task 4 establishes nine existing emitting branches and Task 5 adds five standalone emissions: failure, no-result, backgrounded long, backgrounded short, and unverified disclosure, totaling fourteen rather than thirteen | one branch can be silently omitted from the exactly-one-document matrix, the exact dropped-output failure this oracle exists to catch | enumerate all scenario labels and require fourteen, or name and justify the one scenario covered by an equivalent independent exact-document oracle +MAJOR | high | Task 5 classification oracle, Unicode-escaped marker | the row says the Unicode-escaped key is spelled with real quote bytes, but the verified case must contain the six-byte `\u0022` spelling and no raw quote around `success` | following the wording recreates pass-5's malformed-payload test and reaches `unrecognized` for the wrong reason | say explicitly that the valid located text contains `\u0022success\u0022`, not raw quote bytes, and assert locator success before the matcher verdict +MINOR | high | Task 5 A2 evidence claim | A2 says every listed refusal is tested and the plan cites 53/53 as seed evidence, but the newly added depth-200 refusal has no row in `verify.sh` | the depth cap currently works, but the evidence claimed for the seed does not protect it before execution writes the future hook-suite test | add a depth-201 verifier row or narrow the current-evidence claim until the Task 5 test exists +MAJOR | high | Task 5 `FAILURE_CTX`, `NORESULT_CTX`, `BG_LONG_CTX`, and `BG_SHORT_CTX` | all four paths permit or direct a retry without requiring the review target file or files to be deleted and confirmed gone after the failed or backgrounded call has stopped | CLAUDE.md §5 requires that cleanup because these calls may leave a complete-looking stale artifact; accepting it on retry recreates a false clean pass outside the counter fix | add the policy's exact-target cleanup and confirmation before the one allowed retry, preserving the two-file versus one-branch distinction for Gate B +MINOR | high | Task 5 `UNVERIFIED_MSG` awk and sed probes | `printf x \| sed s/x/y/` emits byte `y` without a terminating newline, so it does not satisfy the message's own requirement that both probes print a line | a healthy POSIX sed can be reported unhealthy by the shipped diagnostic, violating prompt-standard item 10 | use `printf 'x\n'` for both probes and state the exact expected outputs and statuses +MAJOR | high | Task 5 `UNVERIFIED_MSG`, oversized versus ambiguous diagnosis | the proposed distinguishing check is to save the payload and run the hook against it, but both an oversized payload and an ambiguous payload produce the same `unrecognized` message and state effects | the check cannot tell the named causes apart, so the prompt fails item 10 while claiming it does | provide separate observable checks, such as a size check plus a locator diagnostic, or stop claiming they can be distinguished and direct the capture to maintainers as unresolved +MAJOR | high | security standard, Task 5 `UNVERIFIED_MSG` | the operator is told to save and report the raw hook payload even though Task 1 correctly establishes that it can contain prompts, absolute paths, review content, session identifiers, and unrelated concurrent call data | following the shipped prompt can disclose sensitive project and session assets across the external-MCP trust boundary | require local restricted storage, explicit sanitization of metadata and sensitive result content, and prohibit attaching an unsanitized payload to a report +BLOCKER | high | Task 7 Step 7 squash command | `git commit -m "WIP: gate-pass result classification (0.8.0)" -F ` is invalid because Git rejects `-m` and `-F` together | the reviewed snapshot cannot be created, so Gate B has no valid range and the workflow stops before review | use two `-m` arguments, one for the WIP subject and one for the evidence body, or generate one complete message file while also updating the WIP detector to recognize that safe form +BLOCKER | high | Task 7 Steps 1 and 7 | the CHANGELOG is edited in Step 1 but no Task-7 commit or staging step includes it before `git reset --soft`; the reset preserves the unstaged edit and the subsequent commit omits it | the release note required by invariant 12 and the plan's own file structure remains outside the reviewed snapshot, while the pre-reset path audit cannot see it | check the worktree, stage the exact CHANGELOG path before the squash or create the promised Task-7 WIP snapshot, then verify the post-squash index contains every named path +MAJOR | high | Task 7 Steps 6-7 ordering | Step 6 requires the full battery and version-bump check after the squash commit, but the squash commit is not created until Step 7 | the numbered procedure is impossible to execute in order and can record pre-squash evidence as though it covered the reviewed tree | move the squash to the start of the Gate-B preparation, then run the post-squash battery and version check before the first review call +MAJOR | high | Task 7 Gate-B amendment loop | every fix must amend with an explicit WIP `-m`, while the evidence body must survive and be updated, but no amend form carries both; a normal `git commit --amend -m "WIP: ..."` replaces the whole message and erases the evidence | later passes can review a commit lacking the profile-required evidence, or the closing commit can preserve stale evidence | specify the exact two-message amend form for every fix and the exact closing amend form, each rebuilding subject plus the revalidated evidence body +BLOCKER | high | risk high rollback, Task 7 Step 5 before Step 7 | the rollback drill ends by activating and confirming 0.7.1, with no step that restores and confirms the 0.8.0 candidate before Gate B | Gate-B calls can run under the old hook that counts failed calls, defeating the change's validation and pass accounting | run the drill in an isolated profile, or explicitly reactivate and byte-verify the candidate plus reload the session before any Gate-B pass +MAJOR | high | risk high rollback, Task 7 Step 5 | the rollback notes only persistent diagnostic markers; it does not disclose that 0.7.1 deliberately restores the original failed-call false checkmark and the known dash special-builtin invariant-1 violation | an operator can choose rollback believing it only removes classification, then trust counters that again count failures or encounter a nonzero advisory hook | record both lost protections as rollback consequences and define the narrow conditions under which that rollback is acceptable +MAJOR | high | Task 1 Step 9 and `$BASE` lifecycle | the version checker uses `$BASE` before the next sentence records it, and no durable location is specified for a value needed across later task shells | the first check can receive an empty base and later Gate B can reconstruct the wrong SHA, producing a trivial or widened review range | capture the pre-Task-1 HEAD before the commit, persist the literal SHA in the execution evidence or another named durable note, validate it as an ancestor, and only then use it +MAJOR | high | Global Constraints battery and Task 1 Step 9 | the plan says the exact pre-commit block is written once in Task 1, but Step 9 contains only `see The battery`, while The battery incorrectly calls the remaining commands three test/checker scripts even though the AGENTS.md quality row minus only `check-version-bump.sh` has four script invocations | later tasks refer to a nonexistent exact block and can omit `check-version-bump.test.sh` or another required command while claiming the full battery | copy the exact verified AGENTS.md command minus only the commit-comparison checker and reference that literal block from later tasks +MAJOR | high | Task 7 Step 2 counterfactual | the procedure demands four exact `seeded_preserves` labels but the narrowed plan never names those four labels, only a parameterized family whose eventual runner suffix is left to implementation | the counterfactual can select the wrong failures or be impossible to audit against the plan, contrary to the human-confirmed label-and-oracle scope | enumerate the four semantic labels and their expected old/new statuses now, and allow only the mechanical runner suffix to be filled in by the suite +MAJOR | medium | execution precondition before Task 1 | the working tree currently contains the human-confirmed spec amendment and this reviewed plan as unstaged edits, while the plan checks only the cached index and never requires Gate-A artifacts to be committed or the remaining worktree to be clean before `$BASE` is fixed | execution can leave the settled amendment outside every WIP snapshot and Gate-B range, or accidentally fold unrelated unstaged work into the fingerprint without including it in the reviewed diff | add a start guard that commits the approved Gate-A artifacts separately or stops on any unaccounted worktree/index path before recording BASE +MINOR | high | `.context/plan-drafts/README.md` versus plan scoping section | the README still says the locator and matcher bodies live in the plan and that the two copies are kept in step, while the narrowed plan explicitly removed those bodies and makes the ignored drafts the only seed | an executor following the required evidence README can search for nonexistent in-plan bodies or assume a synchronization guarantee that no longer exists | update the README to describe the pass-6 narrowing and the actual reconstruction fallback +MINOR | high | `.context/plan-drafts/README.md` A2 contract | the README still limits duplicate `type` or `text` refusal to the selected element and omits the depth-200 refusal, while both the seed and plan now specify refusal in every examined element and the depth cap | the evidence index is weaker than the source it documents and can guide a reconstruction that drops two settled conditions | align the README's refusal list and verification claims with current A2, including which new cases are and are not present in `verify.sh` +END OF FINDINGS (26 total) diff --git a/.context/codex-reviews/gate-a-plan-pass-7.2026-08-12-ledger-supersession.md b/.context/codex-reviews/gate-a-plan-pass-7.2026-08-12-ledger-supersession.md new file mode 100644 index 0000000..8b5aef2 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pass-7.2026-08-12-ledger-supersession.md @@ -0,0 +1,3 @@ +MINOR | high | Task 6 Interfaces | the `0.8.2` closing-body deliverable points to Task 8 Step 6, but Step 6 only amends fixes into the WIP and the real closing body is created in Task 8 Step 7 | the cross-reference sends an executor or auditor to a step that cannot establish the claimed deliverable | change the reference to Task 8 Step 7 +MINOR | high | Self-Review, Spec coverage | the claim that spec §4 maps to Tasks 3, 4, 6, and 7 omits the governing-story change that §4 requires and Task 1 commits, while describing only the three other pre-landed files | the plan's own coverage summary disagrees with its executable task mapping and makes a required change appear unaccounted for | include Task 1 in the §4 mapping and name the three already-landed files explicitly +END OF FINDINGS (2 total) diff --git a/.context/codex-reviews/gate-a-plan-pass-8.2026-07-30-classifier-cycle.md b/.context/codex-reviews/gate-a-plan-pass-8.2026-07-30-classifier-cycle.md new file mode 100644 index 0000000..fbd8356 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pass-8.2026-07-30-classifier-cycle.md @@ -0,0 +1,34 @@ +BLOCKER | high | Task 5 Step 1, literal message assignments | `UNVERIFIED_MSG` is single-quoted but contains `\'` around both `printf` samples; POSIX shell does not escape an apostrophe inside single quotes, and piping the ten assignments to `sh -n` fails with an unmatched quote/backtick. The mandated cleanup sentence also contains `pass's`, so inserting it verbatim would break four more assignments | copying the plan's product strings makes the hook unparsable and violates invariant 1 before any classifier path can run | rephrase inner examples to use double quotes and say `target findings file for the pass`, or encode apostrophes with a POSIX-safe quote break, then require `sh -n` on the exact ten assignments as an oracle +MAJOR | high | Task 5 Step 1 lines 337-353 and pass-7 disposition 11 | the disposition says the target-findings-file cleanup was added to each retry-capable ``, but none of the four literal `FAILURE_CTX`, `NORESULT_CTX`, `BG_LONG_CTX`, or `BG_SHORT_CTX` strings contains it; the later prose is not part of the declared string | an executor copying the promised full message pairs ships retries that can accept a stale, valid-looking findings file and recreate a false clean pass | insert the cleanup text into all four literal assignments, after stop-or-await in both background variants, and update their exact goldens +MAJOR | high | Task 5 Step 1, `BG_LONG_CTX` and `BG_SHORT_CTX` | both background messages stop retries only while the original task is active; after it stops they allow an unlimited sequence of re-runs, unlike the one-attempt budget in CLAUDE.md §5 and unlike the other two retry-capable messages | repeated backgrounding can loop indefinitely and prompt-standard item 3 is not satisfied | add the shared one-retry budget and a stop-and-surface outcome after a second backgrounding to both background variants +MAJOR | high | Task 5 Step 1, `FAILURE_CTX` | every envelope beginning with `success: false` is class `failure`, but the diagnostic names only `CODEX_EXECUTION_FAILED` and `CODEX_TIMEOUT`; there is no cause/check/action for a missing or future error code, despite pass-4 disposition 17 claiming the unknown-code case was added | a valid failure envelope outside the two captures reaches a diagnostic state that violates prompt-standard item 10 and gives no bounded recovery instruction | add an explicit missing-or-unrecognized-code branch that preserves same-scope retry, records the exact code/message and pinned/effective server evidence, and stops after the shared attempt +MAJOR | medium | Task 5 Step 1, item-1 conformance claim | `Claude Code gate hook —` identifies the source of the injection but does not explicitly say that Claude via Claude Code is the executing model, and the claimed header comment recording the prompting-page check does not exist in the current hook or as literal text in the plan | the plan declares prompt-standard item 1 satisfied without supplying either half of that item's evidence, risking invariant 11 | use an explicit delivered prefix such as `For Claude via Claude Code — gate hook`, and add the concrete source comment with the checked model page and revalidation date +MAJOR | high | Task 6 Step 4 and Step 8, unknown-tool prompt | the plan only prefixes and inserts paragraphs into the current flowing unknown-tool `additionalContext`; it never converts that changed prompt to the structured context/task/rules/output or XML-tag form required by prompt-standard item 5 | Step 8's claim that every changed string passes all 12 items is false for a shipped prompt | specify the complete final unknown-tool pair in tagged sections, review that full result against all 12 items, and pin it in both goldens +MAJOR | medium | Task 5 Step 1, `UNVERIFIED_MSG` security-standard external-system diagnosis | the third-party-envelope cause uses `.context/codex-gate.tools` as its provenance check, but a third-party server can be registered under the name `codex` and expose the default tool names with no mapping file, exactly the indistinguishable provenance path the no-result message already acknowledges | an operator can wrongly rule out the external MCP trust boundary and apply the pinned-server remedy to the wrong server | make the effective MCP registration from `claude mcp list` or `claude mcp get codex` part of this cause's check too, and avoid treating absence of the mapping file as provenance +MAJOR | high | `What this plan specifies` line 23 and Task 5 | the plan says it specifies the class table, but there is no five-row class/counter/fingerprint table anywhere in the plan; only prose, seed code, and a pointer to the spec collectively reconstruct it | the human-confirmed scope claims a reviewable table that can silently drift or disappear while self-review still reports it present | include the exact five-row table or explicitly make the spec §3.3 table the sole table and remove the false in-plan claim +MAJOR | high | Task 5 A2 refusal list | A2 says the locator refuses `a stray comma` without qualification, while the verified seed deliberately returns success for `trailing comma AFTER the selected element is not walked` | the locator contract is stronger than `locate.awk` and its frozen verifier, repeating the claim/code class the settled narrowing was meant to end | qualify refusal as a stray comma on the path walked before selection and list the post-selection trailing-comma label among the non-refusals +MAJOR | high | Task 5 A2 and classification oracle for preceding non-text duplicates | the prose promises repeated `type` or `text` refusal in every examined element, but the restored preceding-non-text fixture repeats only `type`; selected-element tests do not protect preceding-element `text` handling | an implementation can stop refusing repeated `text` in a preceding non-text element while every named row stays green, so pass-7 disposition 3 still overstates the coverage | add a second routed locator case with repeated `text` in a preceding non-text element followed by a valid text block, expecting `unrecognized` +MINOR | high | Task 5 A3 and polarity oracle group | the 64-unit bound applies to all four `strip_ws` call sites, but the only past-bound row places the run after `{`; there is no oracle for past-bound whitespace after the key, after the colon, or after the value token | a later implementation can preserve the first bound but make another externally controlled shell loop unbounded without moving the stated coverage | add parameterized past-64 labels for the other three call sites, each expecting `unrecognized` +MAJOR | high | Scoping section line 30 and Task 5 classification inventory | the exact 56 test labels and the only executable verifier live in ignored `.context/plan-drafts/verify.sh`, yet the fallback says that if the directory is missing the code is reconstructed and re-verified against that same now-missing file; the plan itself contains only group summaries | this is within the human-confirmed label-and-oracle scope, not missing harness shell, and a fresh clone cannot recover or audit the promised decision procedure | track the verifier corpus or inline a non-shell table of all 56 labels and expected classes, plus a reconstructable verification route that does not depend on the missing directory +MINOR | high | `.context/plan-drafts/harness.sh:3-4` | the explicitly unrun seed still says the locator and matcher are green at 53/53 while the plan, README, and observed runs are 56/56 | execution starts from a stale evidence header and can inventory only 53 ports while believing the seed is current | update the header to 56/56 and name the three added labels or refer to `verify.sh` as the count authority +MAJOR | high | Task 1 Steps 2 and 5 | the plan never enumerates the seven destination fixture filenames, and Step 5 tells the executor to copy `shape1-fast-fail.json`, a name absent from `.context/probe-payloads/`, whose source is `shape1-fast-fail-execution-failed.json` | the fixture task is not executable deterministically and later slice, provenance, replay, and counterfactual inventories can choose incompatible names | add a seven-row source-to-destination table, including exact collision and review names, and use those names in every later label and row-count oracle +MAJOR | medium | Task 1 Steps 3-7 sequencing | Step 3 must run in the same shell and locator lifetime as Step 7, but Steps 4-6 that create the review fixture, collisions, and slices sit between them; executing Step 3 when reached either leaves a temporary process/resource alive across manual work or requires deferring the numbered step until later | the advertised start-to-finish order is internally misleading and the byte proof can be run over an incomplete fixture set | move both source-copy and slice checks into one verification step after Step 6, materialize the locator there once, and make the earlier sanitation step a manual edit only +MAJOR | high | Task 5 A5 row `absent/present, pending delete fails` and marker-fault note | the proposed non-writable-`.context` fault cannot let an absent `unverified` marker be created and simultaneously deny deletion of `pending`; directory write permission controls both creation and unlink, so the shown write fails before the delete is reached | the test claimed by pass-4 disposition 13 cannot exercise the pending-delete branch and can pass on the wrong transition | use a self-verified selective `rm` shim or another operation-specific fault that allows marker creation but fails only the target delete, then assert coexistence and next-event cleanup +MAJOR | high | Task 5 A5 row `absent/present, any unsuppressed event` | the event column says `unsuppressed`, but its result column includes emit status 1, which is precisely suppression; there is no separate row for pending debt while the gate remains off | the table is neither mutually intelligible nor complete for an accepted state, so implementations can consume or duplicate debt under opt-out | change the event to `any event` and state gate-on composition versus gate-off retention explicitly, or split suppressed and unsuppressed rows with their exact effects +MAJOR | high | Task 5 A5 independence claim and marker oracles | `bgAdvice` is declared independent, but no row or oracle covers a pending disclosure composed with a background message when one marker write succeeds and the other fails | an early return after one marker failure can lose the other family's transition while all isolated marker tests and the happy composed golden pass | add cross-family rows and labels for pending plus background success, shown-write failure with bgAdvice success, and bgAdvice-write failure with shown success +MAJOR | high | Task 4 unconditional `flush_notes` and spec §3.3 malformed routing | the plan flushes immediately before the final exit on every adopted-repo invocation; if `field()` cannot parse `hook_event_name`, an existing pending disclosure can therefore emit with an empty event name and clear or mutate markers even though the spec says an unroutable payload emits nothing and touches no state | malformed input can burn the one-shot on an invalid hook document and violates the settled routing boundary | gate flushing and marker cleanup on a successfully parsed recognized hook event, and add a pending-plus-unroutable-payload oracle for zero output and byte-identical marker state +MAJOR | high | Task 3 fallback-encoder oracle | Step 1 requires both `sed` substitutions to be status-checked, but the single `encoder failure` label can fail the first substitution every time and never prove that failure of the second, `systemMessage`, substitution returns 2 without output or marker burn | an implementation checking only the first encoder still satisfies the stated test | add separate, self-verified first-encoder and second-encoder failure labels with identical exit/output/marker oracles +MAJOR | high | Task 5 unrecognized state-effect coverage | wiring requires `unrecognized` to count and disclose for both default and mapped exec and review tools, but the state table covers discarded classes, and the fault/marker rows do not explicitly require disclosure and correct pass-state effects for all four gate/tool-name combinations | forgetting `note_unverified` in the exec branch or only on mapped names can pass while violating settled decision 2 | add a four-combination unrecognized matrix asserting Gate-A count versus Gate-B count/fingerprint plus exactly one shown-or-pending disclosure +MINOR | medium | Task 5 sed-fault oracle | the sed-fault row for review requires only that the success envelope counts and discloses, not that Gate B stores a usable, self-matching fingerprint as the equivalent awk-fault row does | a restricted-PATH or wiring defect can leave `unavailable` state while this fault-tolerance row reports coverage | require a usable fingerprint and fresh-state effect for the review half of every sed fault +MAJOR | high | Task 4 and Task 5 message/composition labels | ``, ``, `the nine existing emitting branches`, and `one composed pair` are templates or counts, not the per-test labels promised by the human-confirmed scope; the nine branches and the composed pair are never enumerated | pass 7 corrected thirteen to fourteen but still allows a branch or composition collision to disappear without an auditable label | list all nine existing branch labels, all five new labels, every pending-composition label, and identify the exact pair used for the standalone composed golden +MAJOR | high | Task 6 Step 1 census commands | the recursive greps do not exclude ignored `.mcp/`; in the current working directory they already return stale `.mcp/pass5-revert/codex-gate.sh` and test hits, despite AGENTS.md defining `.mcp/` as generated state | `the greps decide` can send execution into generated cache copies or make the claimed five-site census false | use `git grep` over tracked files or add `.mcp/` to the exclusion and record dispositions only for tracked hits +MAJOR | high | B2 carried obligation and Task 6 accurate census | after excluding generated state, the `accurate` grep still finds the suite comment `Full state machine keeps running while off (so re-enable is accurate)` and the label `re-enable sees accurate state`, but B2 and its edit steps disposition only README, workflow-init, and the hook comment | tests continue asserting the exact overclaim the change retires, and the plan's own census has undispositioned tracked hits | rename and reword both suite sites to `same counting semantics as gate-on, not evidence of review`, and include them in the Task 6 staged-path and golden/label inventory +MAJOR | high | Task 6 Step 7 | `Golden assertions for the two edited prompts` neither identifies the two prompts nor gives test labels or failure oracles; the same step then mentions two distinct unknown-tool golden locations, while Task 6 changes root CLAUDE, its inline template, workflow-init instructions, and the hook prompt | this directly violates the settled requirement that every planned test has a label and oracle, and prompt drift can be declared covered by an ambiguous count | enumerate each golden by artifact and label, state its independent expected source, and say whether it fails on root/template drift, branch-message drift, or composition drift +BLOCKER | high | Task 7 Steps 2, 6, and 8, `$EVIDENCE` | no step assigns an evidence path or creates the complete evidence entry, yet the exact snapshot and amend commands run `cat "$EVIDENCE"`; Step 2 discusses temporary counterfactual evidence files without connecting them to this variable | the reviewed snapshot can be created with an empty/missing body or fail outright, so the profile-required evidence never reaches Gate B or the final commit | define one durable ignored path and `EVIDENCE=` before first use, specify how Steps 2-5 atomically build it, validate it is readable/non-empty, and preserve it from temporary cleanup +MAJOR | high | Task 0 `$BASE` durability versus Tasks 1 and 7 commands | Task 0 explicitly says the SHA is not merely a shell variable, but every later exact command assumes `$BASE` is populated in a new shell and no step reloads it from the durable note | an empty or unset base makes the version check fail, compare the wrong range, or under `set -u` abort; manually remembering the SHA defeats the durability fix | define the durable note format and a checked command that reads the literal SHA into `BASE` at each task boundary, then run the ancestor check immediately +BLOCKER | high | Task 7 Step 8 Gate-B fix loop | after a review fix the plan immediately runs `git commit --amend` but never stages the fixed paths; plain amend does not include unstaged changes | the next reviewer reads the old committed range while the fingerprint sees the changed worktree, and the final commit can omit the accepted fixes entirely, breaking the Gate-B range and invariant-3 reasoning | inspect status, stage only the exact fix paths, audit `git diff --cached`, amend with both messages, then verify a clean worktree and the committed diff before every re-review +MAJOR | high | Global per-commit battery rule versus Task 7 Steps 6-7 | the global rule requires the full pre-commit block before every commit, but Task 7 deliberately creates the squashed snapshot in Step 6 and runs the battery only afterward in Step 7 | the plan is internally inconsistent and the snapshot commit is the one commit whose CHANGELOG staging and full combined tree have not passed the claimed pre-commit gate | run the full block after staging CHANGELOG and before the squash commit as well as the required post-squash battery, or state and justify a narrow snapshot exception instead of claiming per-commit coverage +MAJOR | high | Task 7 Step 4 CI portability evidence versus Step 6 squash | Step 4 demands CI bound to the exact reviewed tree before Step 6 has committed Task 7's CHANGELOG or created the squashed candidate; CI can only run a commit, so at that point it necessarily covers a different tree | the evidence entry can cite mawk/CI for bytes other than the snapshot handed to Gate B, contradicting its own exact-tree claim | move the CI run after the squash and post-squash battery, and after every relevant amend before re-review, or narrow and verify the claim to the exact locator blob hash rather than the whole tree +MAJOR | high | Task 7 Step 3 named verification | the prose requires an explicit expected value per row and an exact row count but supplies neither the row inventory nor the expected counter/fingerprint/class table for replay and live absent/set probes | a procedure can omit a required shape or record a plausible wrong reading while still claiming it followed the task, repeating the disposition-overstatement pattern | add a named table for every replay and live probe row with source, variable state, expected class, pass-count delta, and fingerprint effect, plus the literal total +MAJOR | high | Task 7 Step 5 rollback drill, risk-high concurrency and external system | the non-isolated alternative temporarily activates 0.7.1 in the shared plugin environment without requiring other Claude Code sessions or projects to be quiescent; restoring 0.8.0 later does not undo gate calls that fired during the window | concurrent work can receive the known false-checkmark and dash-exit behavior the drill is meant to document, causing cross-project state corruption that byte verification cannot detect | require an isolated profile/install for the rollback drill, or mandate a verified global quiescence guard and state audit before permitting the shared activation path +END OF FINDINGS (33 total) diff --git a/.context/codex-reviews/gate-a-plan-pass-9.2026-07-30-classifier-cycle.md b/.context/codex-reviews/gate-a-plan-pass-9.2026-07-30-classifier-cycle.md new file mode 100644 index 0000000..1e4c403 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pass-9.2026-07-30-classifier-cycle.md @@ -0,0 +1,38 @@ +MAJOR | high | §What this plan specifies line 23 and §Self-Review | the plan claims it contains the class table, but no five-row class, counter, and fingerprint table exists anywhere in the plan | the human-confirmed review scope and self-review both report a reviewable contract that has silently disappeared, leaving the core effects reconstructable only from the spec and scattered prose | either include the exact five-row table or state that spec §3.3 is the sole class table and remove both in-plan existence claims +MAJOR | high | Task 5 A3 line 311 versus `.context/plan-drafts/match.sh:23-25` | A3 says whitespace is accepted at exactly three points, but `_token_ends` calls `strip_ws` after the value token, making post-value whitespace a fourth accepted point | the stated encoding contract is not exactly as strong as the verified matcher and can guide a different implementation | enumerate all four `strip_ws` sites and their shared alphabet and 64-unit bound +MINOR | high | Task 5 polarity oracle group line 410 | the sole past-64 fixture exercises whitespace after the opening brace, with no label or oracle for the independently coded bounds after the key, after the colon, or after the value | removing or unbounding any of three externally controlled loops leaves the named coverage green | add one past-64 label and `unrecognized` oracle for each remaining `strip_ws` call site +MAJOR | high | Task 5 A2 lines 301-303 | the refusal list says the locator rejects “a stray comma” without scoping that claim to the path walked before selection, while `verify.sh` freezes a trailing comma after the selected element as `success` | the contract is stronger than the seed and repeats the exact claim-versus-code class the settled narrowing was meant to close | qualify the refusal and name `trailing comma AFTER the selected element is not walked` among the non-refusals +MAJOR | high | Task 5 classification oracle line 415 | the preceding non-text duplicate fixture repeats only `type`; there is no counterpart for the separately implemented repeated-`text` branch | an implementation can stop refusing duplicate `text` in a preceding non-text element while every required label remains green, so the pass-7 disposition still overstates coverage | add a routed preceding-element duplicate-`text` label followed by a valid text block, expecting `unrecognized` +MINOR | high | Task 5 A2 line 303 versus spec §3.1 | the prior disposition said the plan would state that the `type` value is compared as raw bytes, but the current text mentions only escaped key spellings; `te\u0078t` is semantically `text` under the spec yet the seed skips it and can return `no-result` | the current plan silently drops the accepted compatibility boundary and disagrees with the approved first-`text`-element contract | explicitly state the raw-value behavior, amend the spec contract to match, and add a fixture, or implement the semantic comparison +MAJOR | high | §What this plan specifies line 30 and Task 5 lines 398-421 | when the ignored `.context/plan-drafts/` directory is missing, the plan says to reconstruct and re-verify against `verify.sh`, but that verifier and the only complete label inventory are in the same missing directory | a fresh clone cannot reconstruct or audit the human-required per-label coverage, and the plan itself contains only group summaries | track the verifier corpus or inline a non-shell table of every label and expected class plus a verification route independent of ignored state +MAJOR | high | Task 1 Steps 3-7 lines 180-200 | Step 3 and Step 7 must run in one shell with one temporary locator, but Steps 4-6 that create the review fixture, collision fixtures, and response slices intervene between them | executing in numbered order either holds an ephemeral shell and temp resource across external manual capture work or reaches Step 7 without the promised lifetime, so the task is not cleanly executable start to finish | move both byte checks into one verification step after Step 6 and materialize the locator there once +MAJOR | high | Task 5 A5 row 326 and marker-fault method line 434 | making `.context` non-writable cannot allow creation of an absent `unverified` marker while denying deletion of `pending`; the same directory permission controls both operations | the stated surgical fault never reaches the pending-delete transition, so its label can pass on the wrong marker failure | use a self-verified selective `rm` shim or another operation-specific fault, then assert coexistence and next-event cleanup +MAJOR | high | Task 5 A5 row 323 | the event is named “any unsuppressed event” but its result includes flush status 1, which is suppression, and no separate row states what happens when pending debt remains while the gate stays off | the supposedly complete state table is internally contradictory and leaves an accepted opt-out state ambiguous | rename the event to any event and spell out gate-on composition versus gate-off retention, or split it into suppressed and unsuppressed rows +MAJOR | high | Task 5 A5 independence claim and marker oracles lines 315-335 and 436-446 | `bgAdvice` is called independent, but no row or label covers pending disclosure composed with background advice when one family's marker write succeeds and the other's fails | an early return after one marker failure can lose the other state transition while every isolated row and happy composition test passes | add cross-family cases for both writes succeeding, shown-write failure with `bgAdvice` success, and `bgAdvice` failure with shown success +MAJOR | high | Task 4 unconditional `flush_notes` line 272 and spec §3.3 | after Task 5 adds pending debt, the unconditional final flush can emit that debt and mutate markers on an adopted-repo payload whose event or tool could not be routed, even though the spec requires unroutable payloads to emit nothing and touch no state | malformed input can burn a one-shot and produce a hook document with an empty event name | flush diagnostic debt only after successful event routing and add a pending-plus-unroutable oracle for zero output and byte-identical marker state +MAJOR | high | Task 3 encoder-failure oracle line 258 | one generic selective-`sed` failure label can always fail the first context encoder and never prove that failure of the second system-message encoder returns status 2 without output or marker burn | an implementation checking only the first substitution satisfies the stated oracle | require separate self-verified first-encoder and second-encoder labels with the same exit, output, and marker expectations +MAJOR | high | Task 5 state-effect oracles lines 423-432 | ordinary `unrecognized` behavior is not required across the four default-or-mapped exec-or-review combinations; discarded classes get that matrix, while success and tool-fault rows do not cover it | omitting disclosure or the correct Gate-A versus Gate-B state effect in one branch can pass despite settled decision 2 | add a four-combination `unrecognized` matrix asserting the exact counter, usable review fingerprint where applicable, and shown-or-pending disclosure +MINOR | high | Task 5 sed-fault oracle line 453 | the review half requires only that a success envelope counts and discloses, not that Gate B stores a usable self-matching fingerprint and fresh-state effect | a restricted-path or wiring defect can store `unavailable` while the fault-tolerance test reports success | give every review-side sed-fault label the same usable-fingerprint oracle as the awk-fault row +MAJOR | high | Task 4 and Task 5 message/composition oracle labels lines 276-285 and 455-466 | ``, ``, “all nine,” “fourteen,” and “one composed pair” are templates or counts rather than the per-test labels promised by the settled scope, and the composed pair is never identified | an emitting branch or collision can disappear without an auditable label, exactly the dropped-coverage failure earlier passes established | enumerate the nine existing branch labels, five new labels, every pending-composition label, and the exact pair used for the standalone and jq-free composed goldens +MAJOR | high | Task 6 census commands lines 480-499 and sweep check 6 | the displayed commands remain recursive `grep` calls that include generated `.mcp/` copies even though the following prose says they use `git grep`; running them verbatim currently returns `.mcp/pass5-revert` hits | the sweep's claim that the census defects were fixed is false, and execution can edit or count generated cache state instead of shipped files | replace the commands themselves with tracked-file `git grep` equivalents or explicitly exclude `.mcp/`, then rerun and record the tracked hit set +MAJOR | high | Task 6 Step 7 | “golden assertions for the two edited prompts” identifies neither prompts nor test labels or independent expected sources, while the task changes the scaffolded CLAUDE prompt and the unknown-tool pair and mentions two separate unknown-tool golden sites | the step violates the settled label-and-oracle scope and cannot show which drift each test catches | enumerate each golden by artifact and exact label, state whether it catches root/template drift, branch-message drift, or composition drift, and name its independently copied expectation +BLOCKER | high | Task 7 Steps 2-8 and sweep check 3 | `$EVIDENCE` is never assigned, no durable evidence path or complete entry-construction step exists, and the commit commands accept `cat "$EVIDENCE"`; the scratch dry run could still create a commit with an empty second message, so its clean result did not test the profile evidence | the reviewed snapshot can lack the required battery, counterfactual, verification, portability, and rollback evidence while appearing executable | define one ignored durable path before first use, specify how Steps 2-5 build it, require readable non-empty validated content, and make the dry run assert that exact body +MAJOR | high | Task 0 `$BASE` durability versus Tasks 1 and 7 | Task 0 says the SHA must outlive different shells, but every later exact command assumes `$BASE` is already populated and no checked reload command or durable-note format is defined | a new shell can abort, compare the wrong range, or tempt the executor to reconstruct the value from current history | define the note format and a checked command that reads the literal SHA into `BASE` and validates ancestry at every task boundary +BLOCKER | high | Task 7 Gate-B fix loop lines 599-605 and sweep check 4 | the amend commands never inspect or stage the files changed by a review fix; a dry run with no representative unstaged fix cannot expose this | the next reviewer can receive the old committed range while the worktree fingerprint changed, and the final commit can omit accepted fixes entirely | before every amend inspect status, stage only approved fix paths, audit the index, amend with both messages, and verify the committed diff and clean worktree +MAJOR | high | Task 7 Step 4 before Step 6 | portability evidence must be bound to the exact reviewed tree, but Step 4 precedes staging the CHANGELOG and creating the squashed candidate commit that CI can actually run | the numbered procedure cannot satisfy its own exact-tree claim start to finish | move the initial CI run after the squash and post-squash battery, and rerun it after each locator or classifier amend before re-review +MAJOR | high | Task 7 Step 3 named verification lines 552-557 | the plan demands an explicit expectation per row and an exact row count but supplies neither the replay and live-probe row inventory nor their class, counter delta, and fingerprint expectations | required profile evidence can omit a story shape or record plausible but wrong readings while claiming compliance | add a literal table for every captured replay and both live variable states, with expected class and state effects plus the exact total +MAJOR | high | Task 7 rollback drill lines 560-568 | the non-isolated alternative temporarily activates 0.7.1 in the shared plugin environment without requiring other Claude Code sessions and projects to be quiescent | concurrent work can receive the known false-checkmark and dash-exit behavior during the drill, corrupting gate state in projects outside this task | require an isolated profile or install, or require a verified global quiescence guard and post-window state audit before shared activation +MAJOR | high | Task 7 rollback release resolution line 562 and sweep check 4 | “the newest commit whose manifest still contains 0.7.1” identifies the last repository bytes under that version, not bytes proven released or cached; AGENTS.md explicitly says the version checker cannot prove release or tagging, and this repository has a history of unbumped plugin commits | a legitimate cached 0.7.1 can fail the byte check against unreleased later bytes, leaving the plan with no rollback despite the sweep proving only `git -S` semantics | establish an authoritative release artifact, tag, or recorded hash for 0.7.1, or stop calling the derived commit released bytes and define how its candidate is installed and verified +MAJOR | medium | Task 7 Step 6 squash commands lines 574-587 | manual log, path, and status reads do not bind the subsequently reset tip; another process can create a commit after the reads and before `git reset --soft`, and that commit is silently folded into the snapshot | the high-risk squash still has a time-of-check/time-of-use concurrent-data path after the earlier guard fix | prohibit concurrent writers for the operation and verify an unchanged captured tip immediately before reset, or use an atomic compare-and-swap ref move with equivalent soft-reset semantics +MAJOR | high | Task 1 sanitization lines 168-192 and `.context/probe-payloads/shape0-success.json` | the plan forbids changing any byte inside captured `tool_response`, but that block contains a real Codex `sessionId` and the background capture contains a real task id; the plan's own unverified warning classifies session identifiers as sensitive data to strip before sharing | tracked shipped fixtures can disclose live provenance identifiers despite claiming sanitized captures | explicitly redact and document these inner identifiers without reserialization and narrow the byte-exact claim to the untouched bytes, or recapture disposable identifiers and prove they are inert +MAJOR | medium | Task 5 Step 1 item-1 claim lines 346-348 and Task 6 unknown-tool prefix | `Claude Code gate hook —` names the source surface but does not state that Claude via Claude Code is the executing model, and the claimed source comment recording the prompting-page check is not supplied as literal planned text | prompt-standard item 1 is asserted rather than demonstrated for every changed model-facing hook prompt | use an explicit delivered target such as `For Claude via Claude Code`, and specify the source comment with the checked page and revalidation date +MAJOR | medium | `BG_LONG_CTX` and `BG_SHORT_CTX` lines 357-360 | both prompts require a second backgrounding to be surfaced but specify no output structure or example for that escalation | prompt-standard item 4 remains unmet on two retry-capable agent prompts | add a compact exact one-line escalation format and example to both variants +MAJOR | high | four retry-capable context strings and line 367 | each delivered cleanup rule says to delete the target findings file and confirm it is gone, but none carries the reason that a failed call may have left a valid-looking stale artifact; the rationale exists only in surrounding plan prose | prompt-standard item 6 is not met inside the prompts, weakening the instruction that closes a false-clean-pass path | append the stale-artifact reason in one concise clause to each cleanup instruction +MAJOR | high | `FAILURE_CTX` line 351 | “every envelope whose first property is success false reaches this state” outruns the classifier, which requires a routed call, a located first text block, raw key spelling, and the accepted encoding grammar; reordered or exotic valid spellings become `unrecognized` | the diagnostic makes an item-11 enforcement claim broader than the mechanism and hides cases that count instead of discard | say every envelope classified as `failure` reaches this state, or state the exact recognized prefix contract +MAJOR | high | `UNVERIFIED_MSG` line 364 versus A2 | the cause tree reduces locator refusals to “oversized or ambiguous,” omitting unwalkable structure and the depth cap, and “measure against the 1 Mi-unit ceiling” is not a complete discriminator or command | prompt-standard item 10 is not met for known terminal-class causes, so operators can exhaust checks without identifying the actual state | enumerate the known refusal families, give the checks that exist, and label structurally indistinguishable cases unresolved without promising a discriminator +MAJOR | medium | `NORESULT_CTX` line 354 versus `NORESULT_MSG` line 355 | the model-facing field says the only external cause is a mapped third-party tool, while the operator field correctly allows a third-party server to be effective under the default `codex` registration with no mapping file | the two audiences receive contradictory trust-boundary diagnoses and Claude can report the wrong provenance | describe the cause as an effective third-party server, then use both the mapping file and `claude mcp list` only as narrowing evidence +MAJOR | high | Task 6 Step 4 unknown-tool prompt lines 515-521 | the plan adds a prefix and paragraphs to the existing long flowing `additionalContext` but never gives the changed prompt the tagged context, task, rules, and output structure required by prompt-standard item 5 | Task 6 Step 8's all-12 conformance claim is false for a shipped product prompt | specify the complete final unknown-tool pair in structured sections and review and golden-test that complete pair +MAJOR | high | Global Constraints line 80 and Task 5 tool-fault contract | the claim that any `sed` failure maps to `unrecognized` ignores jq-free routing, where the existing `field()` also needs `sed`; with jq absent and sed wholly missing, the hook cannot establish the event or tool and exits silently before classification | the compatibility promise is stronger than the executable architecture and the broken-tool disclosure can disappear on the exact fallback environment invariant 4 covers | scope the guarantee to classification-specific sed failure after routing, document the jq-free routing dependency, and add the corresponding no-route oracle or redesign the field fallback +MAJOR | high | Task 5 A4 and both background context strings | the settled notice anchor does not require a task id, but both prompts unconditionally instruct Claude to stop or await the original call by the task id in the result | a broad-but-recognized notice without an id reaches a recovery instruction that cannot be followed | make the task-id action conditional and provide the no-id stop-and-surface path without narrowing the frozen anchor +MINOR | high | `UNVERIFIED_CTX` line 363 | “It repeats only when its marker cannot be persisted or two hook runs race” omits manual or automated marker deletion, which also makes the note repeat | the product makes another categorical item-11 claim beyond marker mechanics | replace “only” with calibrated examples or include marker removal as a cause +END OF FINDINGS (37 total) diff --git a/.context/codex-reviews/gate-a-plan-plana-CLOSURE.md b/.context/codex-reviews/gate-a-plan-plana-CLOSURE.md new file mode 100644 index 0000000..be2bf29 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-plana-CLOSURE.md @@ -0,0 +1,102 @@ +# Gate A — Plan A cycle — CLOSED CLEAN at pass 12 + +Advisory human note. Not a findings file; participates in no pass validation. + +## Closure + +**Artifact:** `docs/superpowers/plans/2026-08-29-review-loop-economics-plan-a-rules.md` +at commit `89056e6` (revision 14). + +**Pass 12: CLEAN.** File validated structurally, not assumed: + +``` +NO FINDINGS␊ +END OF FINDINGS (0 total)␊ +``` + +Exactly two lines; line 1 exactly `NO FINDINGS`; line 2 exactly the terminator. Valid clean +pass under the file-first protocol. + +**Floor satisfied.** The cited story is +`docs/superpowers/stories/2026-08-28-review-loop-economics-pass-floor-story.md`, read fresh at +close: `**Risk:** high · **Security:** none`. Under the rules in force — the constant 3, since +this cycle runs under the OLD rules by its own activation constraint — the floor is 3. Twelve +passes were run. **Final pass clean. The cycle closes.** + +## Curve + +| pass | findings | Blockers | Majors | B+M | +|---|---|---|---|---| +| 1 | 21 | 7 | 10 | 17 | +| 2 | 16 | 6 | 8 | 14 | +| 3 | 16 | 6 | 7 | 13 | +| 4 | 7 | 0 | 3 | 3 | +| 5 | 5 | 0 | 4 | 4 | +| 6 | 2 | 0 | 1 | 1 | +| 7 | 3 | 0 | 2 | 2 | +| 8 | 4 | 0 | 2 | 2 | +| 9 | 6 | 0 | 5 | 5 | +| 10 | 6 | 1 | 5 | 6 | +| 11 | 8 | 1 | 4 | 5 | +| **12** | **0** | **0** | **0** | **0** | + +## Coverage statement + +Stated affirmatively, as the closure requires, and bounded honestly. + +**Reviewed across the twelve passes:** the shipped replacement text of all fourteen edits, read +as final §5 rather than as a diff; both prompt copies, including the two passages where they +already diverge and the one where this plan makes them diverge deliberately; spec §2, §2.1, +§2.2, §2.4, §3 and §10 against the text that implements them; **spec §9's exclusion list item +by item**, which was decisive twice; the old-conditions accounting against the actual old text +of all eleven passages; the downstream `/workflow-init` template for dependencies on this +repo's layout, which found one Blocker; the successor story's reciprocal obligation; the commit +protocol against the hook's actual source; and the fourteen assert-new patterns executed +against a simulated post-edit tree. + +**Not covered, by construction rather than omission:** Plan A's interaction with Plans B and C, +which do not exist yet. That is scope, not a gap in this artifact, and Plan C's five inherited +obligations are named in the plan so they cannot be lost between documents. + +**What this closure does not claim.** Gate A reviewed the plan. It did not review the edits — +the plan describes fourteen prose replacements and Gate B will read the real diff. Each task's +single check establishes that its edit landed at its site, and nothing more; correctness of +what landed is Gate B's. + +## The traced regeneration chain — preserved verbatim for the field record + +The best single specimen this cycle produced of fix-generates-finding in a prose artifact. +Each step is a repair that created the next round's finding: + +- **pass 9** → ship §3 only, mark the deferral by naming the successor story. +- **pass 10 BLOCKER** → that named path ships into the `/workflow-init` template, which + scaffolds `CLAUDE.md` into *other people's* repositories, where the path does not exist. +- **revision 13** → make the pointer `CLAUDE.md`-only, a deliberate divergence. +- **pass 11 BLOCKER** → the divergence is *declared in prose* but never *implemented as a step*: + the architecture still called all thirteen edits mirrored, the constraint permitted only + pre-existing divergences, the accounting still named two diverging passages, and Task 12 still + supplied one identical both-copies replacement. **Executing the plan exactly produced + identical copies and no pointer at all.** +- **revision 14** → emit it as Task 14, an actual step with an asymmetric check (1 in + `CLAUDE.md`, 0 in the template), and **delete the declaration layer that kept disagreeing with + the shipped text** rather than reconciling it again. +- **pass 12** → clean. + +**What broke the chain was deletion, not reconciliation** — the third time in this cycle that +deletion was the cure. Pass 4 deleted the verification machinery after it drew 92% of findings. +Pass 10's decision deleted the out-of-scope reach into the loop rules. Pass 12 followed the +deletion of the plan's self-description. Every attempt to *reconcile* a describing layer with +the thing it described produced another finding; every *deletion* of one produced a drop. + +## Two mandatory stops, and what they cost + +Passes 9 and 11 were mandatory stop-and-surfaces under the five-tells rule — three tells each +time. Both were correct to take: pass 9's stop produced the §3-only reduction, and pass 11's +produced the deletion of the declaration layer. **Neither would have been reached by iterating** +— both times the available repair was a smaller reconciliation, and both times the right answer +was a structural cut that only a human could authorize. + +## Residue + +**None open.** All findings from passes 1-11 are resolved, routed and answered, or explicitly +collected as Minor/Nit per §5. Pass 12 found nothing. diff --git a/.context/codex-reviews/gate-a-plan-plana-pass-1-dispositions.md b/.context/codex-reviews/gate-a-plan-plana-pass-1-dispositions.md new file mode 100644 index 0000000..616205e --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-plana-pass-1-dispositions.md @@ -0,0 +1,97 @@ +# Gate A — Plan A cycle — pass 1 dispositions + +Advisory human note. Not a findings file; participates in no pass validation. + +## Result + +Artifact 3b94dbf. VALID pass: terminator exact, 21 finding lines, 0 non-finding lines. +**7 BLOCKER · 10 MAJOR · 4 MINOR = 17 Blocker/Major.** + +Routed contract: each plan opens <= 12 B/M; **above ~15 routes back before pass 2**. +17 is above it. **Routed, not iterated.** + +## The split worked, directionally + +| | single plan (5c00f8c) | Plan A (3b94dbf) | +|---|---|---| +| findings | 33 | 21 | +| Blocker+Major | 31 | 17 | +| **Blockers** | **21** | **7** | + +Blockers fell by two thirds. That is the same signal the spec cycle gave when it was +split and when it was slimmed, and it is the third time those two interventions are the +only ones that moved a curve. Scope is smaller here, so this is not like-for-like — but +the Blocker share fell from 64% of B+M to 41%, which scope alone does not explain. + +## Two Blockers are consequences of the split itself + +These are the routing question. Neither existed before the split and neither is a +drafting error. + +- **B4 — Plan A cannot produce a green battery.** Tasks 2-7 change + `plugins/dev-workflow/commands/workflow-init.md`; the manifest bump is Plan C's. + `scripts/check-version-bump.sh main` therefore fails at Plan A's WIP HEAD, so the + battery cannot go green, so the story's `battery+check+verification` evidence cannot + be produced, so Plan A cannot close its own Gate-B cycle. +- **B5 — Plan A's findings slots are illegal under the rules in force.** Three plans + running three Gate-B cycles need a per-cycle discriminator; I wrote + `gate-b--plana-pass-

.md`. The infix slot grammar is **Plan B's + deliverable**. Under §5 as it stands today that path is an INCOMPLETE pass, so every + Plan-A Gate-B pass would be unusable. + +Both say the same thing: **splitting the Gate-B cycle creates a bootstrap problem that +splitting the Gate-A cycles does not.** Each plan's own cycle needs infrastructure a +later plan ships. + +## The other five Blockers are mine + +- **B1** — I copied the profile values into the plan (`high`, three 3-pass floors, + `battery+check+verification`) in the same document that says it never copies them. + Predecessor finding M2, unfixed, one paragraph after promising the fix. +- **B2** — the accounting has ten passages; Task 3 Step 5 edits an eleventh + (`pass 1 carrying a Minor`, 126/322). Task 7 Step 3 would then reject the plan's own + required edit as an out-of-inventory hunk. +- **B3 — the sequence defect one level deeper.** I fixed the expected *values* for their + point in the sequence and left the *line numbers* at their untouched-tree positions. + Task 2 inserts two large blocks at 72-80, so by Task 3 the `carrying a Minor` line is + no longer at 126/322 and Task 4's `sed -n '133,140p'` and Task 5's `sed -n '480,493p'` + read unrelated text. Same class as last round's B8, caught at the value level and + missed at the address level. +- **B6** — evidence not revalidated after each accepted fix nor before the closing amend, + which §5 requires explicitly. +- **B7** — Plan A's own Gate-B cycle opens **before** the commit that ships the new + severity rule, so by the activation rule Plan A ships, that cycle must finish under the + **old** severity semantics. Nothing in the plan says so. Left as written, the cycle + reviewing this change could apply the new Minor-or-below ceiling to itself and + under-iterate on exactly the change that introduces it. + +## Findings verified independently before routing + +- **M1 — half confirmed, half refuted, and the confirmed half is a real defect.** My + identity proof compares single anchor lines and concludes whole passages match. + Checked properly with `diff`: the **pass-report passage genuinely differs** between the + copies — line wrapping, and one substantive difference (`you report the tells` in + `CLAUDE.md`, `report the tells` in the template). So that passage needs two accounting + rows, not one. M1's further claim that the **Gate-A passage differs in content is + wrong**: it is identical across all thirteen lines, as is the floor paragraph. +- **M6 — confirmed, and it is the pattern `AGENTS.md` warns about happening live.** I + replaced a false causal claim (`which is where the 3 come from`) with another one: my + replacement makes the hook's advisory fingerprint the *cause* of the re-review + obligation. The actual reason a fix costs another pass is that the prior review no + longer covers the changed artifact; the hook only compares a fingerprint at + commit/review events. Each correction introducing a subtler version of the same claim + is the four-round failure the Don'ts section records, reproduced in one round. +- **M13 — confirmed.** The global constraint `no hook state file is written` is + unqualified and false as written: Gate-B calls and commit events write pass counters, + fingerprints and disclosure markers under `.context/`. Only + `.context/codex-gate.floor` is deliberately untouched. + +## Disposition + +All 17 Blocker/Major carried open to the routing decision. No fixes applied: B4 and B5 +change what the plans *are*, and repairing the other fifteen against a topology that may +not survive would be work spent twice. + +The four Minor are collected, not iterated: Task 7 Step 3's incomplete proof, the fixed +`/tmp` path, the unqualified hook-state claim, and the Self-Review's blanket +sequence-expected-value claim being false in two places. diff --git a/.context/codex-reviews/gate-a-plan-plana-pass-1.md b/.context/codex-reviews/gate-a-plan-plana-pass-1.md new file mode 100644 index 0000000..2b42bcc --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-plana-pass-1.md @@ -0,0 +1,22 @@ +BLOCKER | high | Header lines 21-24, Why line 46, and Task 8 Step 2 | the plan says it never copies profile values, but it records the story as `high`, derives three 3-pass floors from that copy, and hard-codes `battery+check+verification` immediately before telling the executor to read the mode from the story | a later confirmed profile change can leave Plan A running stale passes, lenses, or evidence, recreating the gate-off risk this change exists to remove and leaving predecessor finding M2 unfixed | remove every copied profile value and derived consequence; have each affected step read and validate the story header at execution time +BLOCKER | high | Old-conditions accounting and Task 3 Step 5 | the accounting's ten-passage inventory omits the `Blocker/Major-free pass 1 carrying a Minor keeps looping` passage that Task 3 rewrites at current lines 126/322 | a changed decision procedure has no kept/moved/dropped accounting, and Task 7's instruction to reject every hunk outside the ten will reject the plan's own required edit, leaving predecessor finding B2 unfixed | add this as an eleventh passage for both copies, enumerate every condition in its surrounding closure rule, and update the inventory and diff assertion +BLOCKER | high | Task 3 Step 2, Task 4 Step 2, and Task 5 Step 2 | the sequence-sensitive line checks use untouched-tree locations after Task 2 has inserted the much larger floor, residual, and gate-off blocks: `carrying a Minor keeps` can no longer be at 126/322, and the later `sed -n 133,140p` and `sed -n 480,493p` commands read unrelated text | the stated expected output is wrong at its actual point in the sequence and executors can inspect or edit the wrong passage, which is a blocking anchor/check failure under the review brief | resolve each live location with a unique content anchor after prior tasks, or perform and record all fixed-line reads before any line-adding edit +BLOCKER | high | Task 8 Step 1 | the full battery includes `scripts/check-version-bump.sh main`, but Tasks 2-7 commit a change under `plugins/dev-workflow/` while Plan A deliberately leaves the manifest bump to Plan C | the checker necessarily fails at Plan A's WIP HEAD, so `BATTERY-GREEN` and the required high-risk evidence cannot be produced and Plan A depends on unlanded Plan-C content to close | choose a topology that makes the bump present before this battery, or keep the prompt edits unclosed until the packaging plan is folded into the same reviewed snapshot +BLOCKER | high | Task 8 Step 4 findings paths | the prescribed `gate-b--plana-pass-

.md` names are outside current §5's exact `gate-b--pass-

.md` slot grammar; Plan B's nonce-bearing slot rules do not exist yet | current §5 classifies a wrong path as an INCOMPLETE pass, so every Plan-A Gate-B pass can be unusable and cannot satisfy the floor | use the current legacy slots for this pre-rule cycle, or ship the new slot grammar before relying on different names +BLOCKER | high | Task 8 Steps 3-5 | evidence is produced once before the review loop, but after an accepted fix the plan immediately amends and re-reviews without revalidating the battery/check/verification entry, and it performs no final revalidation immediately before the closing amend | a clean pass and closing body can cover stale evidence for an earlier diff, directly violating current §5 and spec §8 | build the exact evidence entry before the first call, revalidate and update it after every fix before amending/re-reviewing, and revalidate it again against the closing snapshot +BLOCKER | high | Task 6 Step 3 activation rule and Task 8 Step 4 | Plan A's Gate-B cycle opens before the commit that ships the new severity rule, so the inserted activation text says this cycle must finish under the old severity semantics, yet the review-loop instructions never carry or verify that starting-rule choice | the cycle reviewing this change can prematurely apply the new Minor-or-below ceiling, silently reducing Blocker/Major iterations in the very pre-rule cycle the activation rule excludes | state explicitly that this cycle uses the pre-rule severity definitions, include that in every Gate-B call, and verify the pass reports did not apply the new demotion +MAJOR | high | Task 1 Step 2 and Old-conditions accounting identity claim | the claimed proof compares only each anchor line, not each passage; independent block comparison already shows the pass-report paragraph differs in wrapping and the Gate-A passage differs in content, while the table contains ten shared rows rather than the claimed twenty per-copy rows | a condition unique to one mirror can still hide behind a shared disposition and predecessor finding M1 is not actually resolved | compare the complete exact rewrite ranges byte-for-byte and use shared rows only for ranges that match, with separate rows for every differing range +MAJOR | high | File Structure, Old-conditions accounting, and Task 1 Step 3 | spec §6 requires the accounting in one artifact written once, but Plan A writes it in the plan, copies it to a second conditions artifact after Gate A, and later mutates that second copy with result tables | the two records can diverge and the final conditions artifact is not the single artifact accepted alongside the plan, despite the table claiming predecessor B1 is dissolved | make the reviewed conditions artifact the sole accounting source before execution, or amend the approved spec explicitly before using a duplicated extraction model +MAJOR | high | Tasks 1-7 preflights | the preflights do not define safe complete, partial, asymmetric, or duplicate states: Task 1 trusts any existing file then tries another ordinary commit, Task 2's completed path still reaches a no-change commit, Tasks 4-6 duplicate insertions/appends on rerun, and Task 7 has no preflight | interruption between mirror edits or rerunning a completed task can leave divergent copies, duplicate shipped rules, or a failed commit, so predecessor finding M9 remains | give every task an exact 0/0 proceed, 1/1 byte-verify-and-skip, asymmetric stop/repair, and duplicate stop state, including the correct commit/amend action for resumed work +MAJOR | high | Task 5 Steps 1 and 4 | Step 4 says to compare against a pre-edit token count captured in Step 1, but Step 1 captures only the new-phrase count; moreover the appended text itself adds `keep counting` and several `final clean pass` matches, so the post-edit nondecrease test can pass after old conditions were dropped | the check advertised as confirming all nine old conditions cannot establish that claim and has no sequence-specific expected number | capture the untouched old paragraph or its digest before editing and require that exact block to survive, while checking the appended block separately +MAJOR | high | Task 3 Step 3 lenses replacement | `passes a cycle owes, which the profile alone decides` contradicts Task 2's own predicate because cited-set emptiness, membership, unprofiled members, and unresolvable members also determine whether the result is 3 or a stop | readers can derive a floor from one profile while ignoring another cited story or the no-story/unprofiled branches, creating a new unlicensed-floor path | say the floor is decided by the derived-floor rule over the current cited-story set and its profiles +MAJOR | high | Task 3 Step 3 Gate-B rationale | `the hook invalidates the prior pass — which is why a fix costs another pass` makes an advisory fingerprint mechanism the cause of the review obligation; the hook only compares its recorded fingerprint at commit/review events, while the prior review stops covering the artifact because the revision changed and §5 requires re-review | the shipped rationale overclaims what the hook proves and can make the rule appear contingent on hook behavior, violating the gate-mechanism invariant | state that the changed revision is no longer covered by the prior review, then separately state the hook's exact fingerprint comparison and the profile-derived floor +MAJOR | high | Task 7 Steps 1-2 | the promised one-row-per-rule parity and conformance gates have no canonical row inventory or asserted row count, and the supplied parity script compares only equal marker occurrence counts, printing `PARITY` even when the surrounding shipped texts differ | an executor can omit Task 3 sites or divergent rule text and still produce the expected eight `PARITY 1` lines, so predecessor finding M7 is only cosmetically addressed | enumerate every changed rule ID and all twelve item/artifact pairs, compare exact extracted rule text, and fail on a missing, extra, duplicate, or non-identical row +MAJOR | high | Task 8 Step 4 recovery behavior | `delete both branch targets and confirm them gone before each call` conflicts with current §5's single-branch recovery rule, which requires deleting only the failed branch because recreating one branch after deleting both makes the both-files validation fail by construction | a recoverable one-branch write failure can spend the sole retry on a call that cannot yield a valid full pass | distinguish full reruns from branch resumes and preserve the successful branch exactly as §5 requires +MAJOR | high | Task 8 Step 3 floor-file verification | the only knob check occurs at close and asks `was it present at start?` without recording start-time existence or bytes | removal or modification of a user's knob during execution can be reported as clean, so the plan does not verify its own no-write/no-remove and data-loss claims | record existence and a byte digest before Task 1, compare at every revalidation and close, and record a reasoned N/A only when it was absent at both points +MAJOR | high | Task 8 Step 4 accepted-fix amend | `git add -u` stages every tracked modification in the repository without a clean-worktree preflight or path review | unrelated user changes can be folded into the WIP, reviewed under the wrong story, and closed in Plan A's commit | stage only an explicitly inspected path set belonging to the accepted fix and stop on unrelated tracked changes +MINOR | high | Task 7 Step 3 | the command shows only a stat for both prompt files and counts changed lines only in `CLAUDE.md`; it never displays the template hunks it claims the executor must verify, and the count has no stated expected value | an out-of-inventory template edit can escape the advertised proof even after the missing passage is fixed | display and inspect both full diffs against an explicit eleven-passage inventory and state the expected hunk or changed-line set +MINOR | medium | Task 8 Step 5 | the closing body uses the fixed global path `/tmp/plan-a-close.txt` | a concurrent Plan-A execution or unrelated process can overwrite the body between validation and amend, producing the wrong evidence or commit message | create a private file with `mktemp`, validate that exact path, and remove it after the amend +MINOR | high | Header Architecture, Global Constraints, and Task 8 | `no hook state file is written` is unqualified even though the planned Gate-B calls and commit events normally write pass counters, fingerprints, and disclosure markers under `.context/`; only `.context/codex-gate.floor` is intentionally untouched | an executor can read normal hook state as a plan violation or delete it, disrupting the active cycle | narrow the claim to implementation-owned writes and explicitly exempt the hook's normal operational state while retaining the floor-file prohibition +MINOR | high | Self-Review "Placeholders" and Task 7 Step 3 | the blanket claim that every check states its sequence-specific expected value is false: Task 5 omits its promised baseline and Task 7's changed-line count has no expected result | the self-review can report mechanically complete checks even when an executor cannot distinguish the intended state from drift | list a concrete expected value or accepted closed set for every output-producing check and remove the blanket claim until that is true +END OF FINDINGS (21 total) diff --git a/.context/codex-reviews/gate-a-plan-plana-pass-10.md b/.context/codex-reviews/gate-a-plan-plana-pass-10.md new file mode 100644 index 0000000..90c2c74 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-plana-pass-10.md @@ -0,0 +1,7 @@ +BLOCKER | high | Task 12 NEW lines 789-791 in the inline `/workflow-init` copy | The shipped scaffold cites `docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md`, a repo-internal file that `/workflow-init` does not scaffold and that a downstream project is not promised to have, contrary to the architecture boundary and prompt-standards item 11's self-contained-template exception | Every initialized downstream `CLAUDE.md` contains an unresolvable authority for part of its gate semantics, so the installed plugin depends on this repo's internal layout and leaves the loop-health interaction unusable outside this checkout | Keep the named reciprocal handoff in the plan and successor story, but make the shipped sentence self-contained without a repo-local path, or deliberately give the two copies different self-contained wording and account for that divergence +MAJOR | high | Task 12 NEW lines 789-791 and successor story §4 lines 127-138 | The shipped sentence says the demotion interaction "is settled by" the successor story, while that story explicitly says the parent does not settle it and that the successor design still owes an answer; the referenced draft contains an obligation, not the settlement | A reader of final §5 is told that operative answers already exist, follows the reference, and still cannot determine how demoted findings affect counts, clusters, or stop thresholds | Say the interaction is deferred to or will be settled by the successor design, or put the actual settled answer in the referenced artifact before claiming it is settled +MAJOR | high | Task 1 NEW lines 212-225 | The governing-header model still has no closed result for an expected artifact with no `Story:` header or for repeated `Story:` headers: the comparison covers only headers that "exist", so an absent plan contributes no empty set, and neither identical nor conflicting repeats are normalized or rejected | In the multi-plan Gate-B case an omitted or ambiguous header can disappear from the union and from the disagreement check, allowing a lower floor and omitting that story's lenses, evidence duties, and review scope | Make every expected spec or contributing plan contribute a set, with absence contributing the empty set, and define repeated headers as either one explicitly deduplicated set or a malformed state that stops and surfaces +MAJOR | high | Task 1 NEW lines 205-226 | The text claims that "at any moment there is exactly one value" before defining reachable states in which governing headers disagree or a present profile is unresolvable and therefore no value may be derived | A deriving agent can follow the unconditional one-value assertion instead of the later stop, and the shipped explanation overclaims the mechanism in the exact class AGENTS.md forbids | Qualify the one-value claim to moments when all governing headers agree and every cited present profile resolves; state that the disagreement and unresolvable states produce no derived value +MAJOR | high | Task 1 NEW lines 212-216 and Task 12 NEW lines 789-791 | Task 1 says a story path in an artifact body "governs nothing", but Task 12 deliberately uses such a body path to govern where the unresolved demotion interaction is handed off; being non-citing for floor derivation does not make a reference non-governing for every purpose | The final §5 gives an agent mutually exclusive instructions to ignore the body path and to treat it as the authority for the deferred interaction, undermining both the citation rule and the handoff | Narrow Task 1 to say body paths do not contribute to the cited-story set or floor derivation, without claiming they govern nothing for other purposes +MAJOR | high | Task 13 NEW lines 856-863 against Task 10 and untouched §5's pass-report paragraph | The coherence rule still over-triggers because it covers every normative specific pass ordinal but exempts only closure and stop predicates, hook thresholds, and descriptive or historical numbers; the fully adopted text itself contains independent normative ordinals such as Task 10's "from pass 1 onward" and the untouched "From pass 4 onward" reporting duty, neither of which derives from the floor | A complete correct adoption can be classified as incoherent and forced to stop even though its floor definition and all floor-dependent closure statements agree | Restrict the rule to normative pass counts that define or constrain the floor and floor-based closure, or explicitly exempt independent reporting and diagnostic thresholds as well as the already named closure predicates +END OF FINDINGS (6 total) diff --git a/.context/codex-reviews/gate-a-plan-plana-pass-11-dispositions.md b/.context/codex-reviews/gate-a-plan-plana-pass-11-dispositions.md new file mode 100644 index 0000000..2d846c9 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-plana-pass-11-dispositions.md @@ -0,0 +1,105 @@ +# Gate A — Plan A cycle — pass 11 dispositions + +Advisory human note. Not a findings file; participates in no pass validation. + +## Result + +Artifact 7208756 (revision 13). VALID pass: terminator exact, 8 finding lines, 0 non-finding +lines. **1 BLOCKER · 4 MAJOR · 2 MINOR · 1 NIT = 5 Blocker/Major.** + +**MANDATORY STOP-AND-SURFACE — three tells. Second stop in three passes. Nothing repaired.** + +## Trend — the whole cycle + +| pass | findings | Blockers | Majors | B+M | +|---|---|---|---|---| +| 1 | 21 | 7 | 10 | 17 | +| 2 | 16 | 6 | 8 | 14 | +| 3 | 16 | 6 | 7 | 13 | +| 4 | 7 | 0 | 3 | 3 | +| 5 | 5 | 0 | 4 | 4 | +| **6** | **2** | **0** | **1** | **1** ← the bottom | +| 7 | 3 | 0 | 2 | 2 | +| 8 | 4 | 0 | 2 | 2 | +| 9 | 6 | 0 | 5 | 5 | +| 10 | 6 | 1 | 5 | 6 | +| 11 | 8 | 1 | 4 | 5 | + +**The loop bottomed at pass 6 and has not returned in five passes.** + +## Tells + +| # | tell | present? | +|---|---|---| +| 1 | finding count rising | **YES** — 2, 3, 4, 6, 6, 8 | +| 2 | Blocker count failing to fall | **YES** — zero for six passes, then 1, 1 | +| 3 | clustering on the instrument | no — the instrument was stripped at pass 4 | +| 4 | clustering on **prose about** either | **YES** — 4 of 8 (findings 1, 5, 7, 8) are the plan's own declarations, accounting and rationale disagreeing with its own shipped text | +| 5 | require↔withdraw pair | not as a pair, but see the regeneration chain below | + +**Three present. Stop is mandatory, not discretionary.** + +## The regeneration chain, traced + +Each round's repair produced the next round's finding, and it is traceable rather than +impressionistic: + +- **pass 9** → ship §3 only, mark the deferral by naming the successor story. +- **pass 10 BLOCKER** → that named path ships into the `/workflow-init` template, which + scaffolds `CLAUDE.md` into *other people's* repositories where the path does not exist. +- **revision 13** → make the pointer `CLAUDE.md`-only, a deliberate divergence. +- **pass 11 BLOCKER** → the divergence is *declared in prose* but never *implemented as a + step*: the architecture still calls all thirteen edits mirrored, the constraint permits only + pre-existing divergences, the accounting still names two diverging passages, and Task 12 + still supplies one identical both-copies replacement. **Executing the plan exactly produces + identical text in both copies and never creates the pointer at all.** + +That last one is mine and it is exact: the `CLAUDE.md`-only edit exists in the simulation +source but was never emitted into the plan as a task step, because the generator emits only +the thirteen tasks the notes file drives. + +## Two findings worth naming individually + +- **NIT 8 — a mis-cited invariant, and the conclusion was still right.** I justified the + no-downstream-path rule with invariant 7. Invariant 7 governs `examples/` as read-only + reference and says nothing about story paths in scaffolded templates. What actually supports + it is the architecture dependency boundary and prompt-standards item 11. The rule I shipped + is correct; the reason I gave for it was not, which teaches a future maintainer to lean on a + rule that does not cover the case. +- **MINOR 7 — the accounting overclaims twice.** Row 1 says conditions d–h are "kept verbatim" + when Task 1 rewrites their wording; row 8 says the hook-invalidation condition is kept when + Task 7 deliberately *drops* that causal claim and replaces it. The accounting is the plan's + sole stated guard against dropped conditions, so an inaccurate disposition is the guard + overclaiming itself. + +## The "clearly stuck" reading — offered, not taken + +§5 requires three things **together**, and I can state two affirmatively: + +1. **A plateau visible across passes** — yes. Six passes since the pass-6 bottom, none + returning to it. +2. **Blocker/Major regenerating across genuine repair attempts** — yes, and the chain above + traces it rather than asserting it. +3. **An affirmative judgement that coverage is sufficient** — **I can state this, with one + caveat named.** Eleven passes have walked the shipped text of both copies, spec §§2–3 and + §10, §9's exclusion list item by item, the old-conditions accounting, the downstream + template, and the successor story's reciprocal. I know of no materially unreviewed area + *within Plan A*. The caveat is that Plans B and C do not exist, so Plan A's interaction + with them is unreviewed by construction — that is scope, not a coverage gap in this + artifact. + +**But the exit is not mine to take, and I am not taking it.** A clean completion takes +precedence over this exit and we do not have one; reporting "will not converge" would also be +premature while every finding is individually small and fixable. What I am reporting is that +the *shape* now matches the one §5 describes, and that the decision belongs to the human. + +## Disposition + +All five Blocker/Major carried open. Nothing repaired — the tells rule hands the decision over +rather than permitting another round. + +**Cost of resuming, so the choice is informed:** every one of the eight is individually small +and mechanical. The Blocker is a missing task step plus four prose sites to bring into +agreement. None requires a new decision. The question is not whether this round can be fixed — +it can — but whether a loop that has produced a fix-generates-finding chain for five +consecutive passes should keep running in this form. diff --git a/.context/codex-reviews/gate-a-plan-plana-pass-11.md b/.context/codex-reviews/gate-a-plan-plana-pass-11.md new file mode 100644 index 0000000..11db69b --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-plana-pass-11.md @@ -0,0 +1,9 @@ +BLOCKER | high | Architecture and Global Constraints lines 11 and 80-81, old-conditions accounting lines 136-154, Task 12 lines 762-817 | The plan declares one new CLAUDE.md-only successor pointer, but its architecture still calls all thirteen edits mirrored, the global constraint permits only pre-existing divergences, the accounting still says only two passages diverge, and Task 12 supplies only one identical BOTH-copies replacement with no exact second sentence or CLAUDE.md-only edit | Executing the tasks exactly produces identical new severity text in both copies and never creates the claimed reciprocal handoff, while an executor who tries to honor the prose must invent unreviewed wording | State the new divergence consistently in the architecture, parity paragraph, constraint, and accounting, and give Task 12 an exact CLAUDE.md-only insertion while retaining its single settled assert-new check +MAJOR | high | Task 1 NEW lines 205-209 and 225-234 | The text says exactly two states derive no value, but it later defines a third: a Story header that cannot be read as a list of paths is malformed and stops before any cited set or floor can be derived | The exhaustive two-state claim overstates the procedure and can make an agent force a malformed header into disagreement or profile-unresolvability instead of reporting the actual cause | Add malformed or unreadable Story headers to the no-value states, or make the list explicitly non-exhaustive +MAJOR | high | Task 13 NEW lines 848-878 | The coherence rule requires exactly one floor definition and treats a fixed-number obligation beside the derived predicate as broken, while the same final text deliberately contains the unknown-start fallback floor of 3 beside that predicate and does not exempt or distinguish this disjoint activation state | A reader can classify a complete correct adoption as self-contradictory and stop every gate whose starting rules are unknown, defeating the fallback that is supposed to govern that state | Define coherence per activation state, or say the ordinary derived predicate must be unique while the explicitly conditional unknown-start fallback is not a competing definition +MAJOR | high | Task 12 NEW lines 799-800 in the scaffolded template, against the untouched pass-report and two-tell rules | The downstream copy says the demotion interaction with counts, clusters, and stop thresholds is not settled, but provides neither an authority it can reach nor a terminal action for the first pass containing a demoted finding | Different downstream agents can include or exclude the same finding from the Blocker curve, clusters, and mandatory two-tell stop, so the shipped gate becomes nondeterministic at a reachable state | Keep the substantive interaction deferred, but make the template self-contained by specifying a safe unresolved-state action such as stop and surface until the project's loop-rule contract settles it, or sequence shipment with the rule that supplies the answer +MAJOR | high | Successor story section 4 lines 127-138, as relied on by Task 12 line 762 | The claimed reciprocal still says the parent ships a sentence naming this story as where the interaction is settled, while revision 13's shipped sentence neither names the story nor says the interaction is settled and the successor paragraph itself says its future design still owes the answer | The handoff artifacts disagree about whether an operative answer exists and about what the parent actually ships, recreating the misleading settlement claim pass 10 removed from Plan A | Rewrite the reciprocal to say the parent defers the interaction and, in this repo's CLAUDE.md-only wording, points to this story as the owner of the unresolved obligation rather than as an existing settlement +MINOR | medium | Task 1 NEW lines 214-230 | The new multi-story authority has no accepted syntax for a Story header that carries a list of paths, even though validity now depends on distinguishing a readable list from a malformed one and normalizing repeated entries | Legitimate multi-story artifacts can be parsed differently by different agents and fail closed for availability reasons, with no diagnostic cause or distinct fix as prompt-standards item 10 requires | Define a minimal header representation and normalization rule, for example one Story field per path with exact duplicate paths deduplicated, and name the malformed cases and fixes +MINOR | high | Old-conditions accounting rows 1 and 8, lines 142 and 151 | Row 1 says conditions d-h are kept verbatim although Task 1 rewrites their wording, and row 8 says the old hook-invalidation condition is kept even though Task 7 explicitly removes that causal claim and replaces it with artifact-coverage reasoning plus a narrower hook observation without marking the old claim deliberately dropped | The accounting is the plan's sole stated guard against dropped conditions, so inaccurate dispositions make that guard overclaim what it records | Distinguish semantically kept from verbatim text in row 1, and split row 8 into the re-review duty kept and the hook-causes-invalidation rationale deliberately dropped and replaced +NIT | high | Task 12 rationale line 762 | The rationale says invariant 7 forbids the downstream repo-local story pointer, but invariant 7 governs examples as read-only reference and says nothing about story paths in scaffolded templates; only the architecture dependency boundary and the actual workflow-init target list support this conclusion | The conclusion remains correct, but the false invariant citation teaches a future maintainer to rely on a rule that does not cover the case | Remove the invariant 7 citation or replace it with the self-contained-template exception in prompt-standards item 11 plus the architecture boundary +END OF FINDINGS (8 total) diff --git a/.context/codex-reviews/gate-a-plan-plana-pass-12.md b/.context/codex-reviews/gate-a-plan-plana-pass-12.md new file mode 100644 index 0000000..8d781f4 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-plana-pass-12.md @@ -0,0 +1,2 @@ +NO FINDINGS +END OF FINDINGS (0 total) diff --git a/.context/codex-reviews/gate-a-plan-plana-pass-2-dispositions.md b/.context/codex-reviews/gate-a-plan-plana-pass-2-dispositions.md new file mode 100644 index 0000000..69f03b3 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-plana-pass-2-dispositions.md @@ -0,0 +1,91 @@ +# Gate A — Plan A cycle — pass 2 dispositions + +Advisory human note. Not a findings file; participates in no pass validation. + +## Result + +Artifact 83b17e0 (revision 2). VALID pass: terminator exact, 16 finding lines, 0 +non-finding lines. **6 BLOCKER · 8 MAJOR · 2 MINOR = 14 Blocker/Major.** + +Routed contract: pass 2 <= 8 B/M, converging <= 4; **above ~12 routes back**. 14 is above +it. **Routed, not iterated** — and see the tells below, which make it mandatory rather +than discretionary. + +## Curve, Plan A cycle + +| pass | findings | Blockers | Majors | B+M | +|---|---|---|---|---| +| 1 | 21 | 7 | 10 | 17 | +| 2 | 16 | 6 | 8 | 14 | + +## The tells — two present, so stopping is mandatory + +§5 owes these lines from pass 4; reading them at pass 2 is free and they already fire. + +1. **Finding count rising** — NO. 21 → 16, falling. +2. **Blocker count failing to fall** — 7 → 6. A fall of one across a full repair round. + **Ambiguous, leaning present.** +3. **Findings clustering on the INSTRUMENT rather than product behaviour** — **PRESENT, + overwhelmingly.** Five of six Blockers are broken checks, not wrong rules: B1 quotes a + replacement boundary mid-line; B3's three fixed-string greps address text that does not + exist on disk as contiguous bytes; B4's marker is split across lines and its `grep -cF` + pattern begins `- `, which grep reads as an option and exits 2; B5's end anchors do not + occur on any line. Majors 4, 5, 7 and Minor 10 are the same class. +4. **Findings clustering on PROSE ABOUT either** — **present.** M2, M3, M6 and MINOR 10 + are about the plan's own accounting and its own self-review claims. +5. **A require↔withdraw pair** — none identified. + +**Tells 3 and 4 are present, so stop-and-surface is mandatory, not discretionary.** The +routed >12 threshold points the same way. Both are reported to the sparring session. + +## Root cause — one error, not fourteen + +Every instrument finding is the same mistake: **I wrote checks against Markdown prose and +asserted they would work instead of executing them against the actual post-edit bytes.** +The four shapes it took: + +- a phrase containing `**emphasis**` matched as if it were plain text; +- a phrase split across a line break matched as one contiguous string; +- a `grep -cF` pattern beginning `- ` passed without `--` or `-e`, so grep parses it as an + option and exits 2; +- an `awk` range whose end anchor does not exist on any line, so the range runs to EOF. + +**Codex found these by building a sandbox at `.context/plan-a-pass2-sim/`, applying Tasks +1-5 in sequence, and running the checks.** I did not. That asymmetry is the whole finding: +the reviewer executed the plan and I only read it. + +## Findings verified against source before routing + +- **M1 — CONFIRMED, and it is the third failure of one sentence.** The text on disk is + `Re-review after every fix — a fix changes the diff and the hook` / `invalidates the + prior pass, which is where the 3 come from.` My replacement swaps only the second line, + so the shipped sentence reads **"a fix changes the diff and the hook no longer covers the + artifact"** — the hook is *still* the grammatical subject and the causal claim is still + wrong. Round 1 claimed the number came from invalidation; round 2 made the hook the cause; + round 3 leaves the hook as the subject. The `AGENTS.md` entry says this took four Gate-B + rounds because each correction searched for the previous **phrase** rather than the + **claim**. Three rounds here, same mechanism. **The fix is to replace both lines.** +- **M2 — CONFIRMED, and it refutes my own pass-1 verification.** I reported the Gate-A + passage as byte-identical across copies. It is not: over its **full** extent (30 lines vs + 29) `CLAUDE.md` carries `` (`docs/prompt-standards.md`, "coverage first, filter later") `` + which the template **drops entirely** — substantive, not wrapping. My check compared + `C 300-312` against `T 485-497`, a 13-line window that **stops before the divergence**, + and I reported IDENTICAL. A check wired so it cannot observe the thing it claims to + check — while verifying a finding *about* identity claims. The accounting's shared-row + structure is therefore unsound for this passage too, not only for the pass-report one. +- **B3 — CONFIRMED mechanically.** `grep -c 'you report the tells' CLAUDE.md` returns **0**; + the phrase breaks after `you report` at line 144. + +## Disposition + +All 14 Blocker/Major carried open. No fixes applied. + +The repair is mechanical and bounded, and it is one action rather than fourteen: **build the +simulated post-edit tree, run every check in the plan against it, and ship only checks that +demonstrably produce their stated output.** That is the validation pass I skipped, and +`.context/plan-a-pass2-sim/` already exists as a starting point. B6 is the one finding +outside that class — it is a real gap in the newly-decided one-WIP topology (no immutable +pre-cycle base SHA, so stacked WIPs pass `CYCLE OPEN` while `git diff HEAD~1` silently omits +an earlier WIP from the only Gate-B review) and needs a topology answer, not a better check. + +The two Minor are collected, not iterated. diff --git a/.context/codex-reviews/gate-a-plan-plana-pass-2.md b/.context/codex-reviews/gate-a-plan-plana-pass-2.md new file mode 100644 index 0000000..503be04 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-plana-pass-2.md @@ -0,0 +1,17 @@ +BLOCKER | high | Task 1 Step 3, lines 268-276 | the block labelled “Replace exactly” is not present byte-for-byte in either target: the real line is `review. Open a TodoWrite ...`, while the quoted replacement target ends `review.` followed by a newline | the first edit cannot be performed as specified; an executor must guess the boundary and can either stop or swallow the closure-rule tail that the plan promises to preserve | quote the actual bytes including `review. Open ...`, or replace the complete paragraph with one complete resulting paragraph and verify the preserved tail explicitly +BLOCKER | high | Task 1 Step 1 and Steps 3-4 | `old=0 new=1` is called complete as soon as the floor replacement lands, even if the residual and gate-off blocks have not been inserted, and the same pair is reported when only one mirror has received Step 4; the state table also has no `new=2` duplicate state | an interruption after Step 3 or between the Step-4 mirror edits is misclassified as complete, skips directly to the WIP commit, and can ship without the named residual or gate-off disclosure that the spec requires | count independent markers for the replacement, residual, gate-off block, and both mirrors; define a closed state matrix for partial, duplicate, asymmetric, pre-commit, and already-committed states before choosing Step 4, Step 6, or STOP +BLOCKER | high | Task 3 Steps 1 and 3, lines 476-517 | all three load-bearing fixed-string checks are addressed to text that does not exist: the inserted text contains `the **cited stories they were read from**`, and the pre-existing `you report the tells` and `report the tells` phrases are split across lines; sequence simulation returned 0 for each instead of the promised 1 | Task 3 can never prove completion, its preflight continues to read a completed insertion as not started, and a rerun can duplicate the every-pass rule | use contiguous byte anchors that include the Markdown actually on disk, and compare the exact pre-existing paragraphs or their digests rather than grepping prose across formatting boundaries +BLOCKER | high | Task 5 Step 4, lines 686-699 | the third check searches for `adds its own strict reading to this list`, but the proposed text breaks that phrase across two lines, and the fourth invokes `grep -cF` with a pattern beginning `- ` without `--` or `-e`, which exits 2 | the stated four `1/1` results are impossible even after the correct edits, so the severity and activation task cannot satisfy its own verification | use a contiguous marker such as `adds its own strict reading`, and pass the severity pattern with `grep -cF -- ...` or `grep -cF -e ...`; run the repaired block under `sh` and `dash` +BLOCKER | high | Task 6 Step 2 ranges 3, 11, and 13 | the end anchors `not the reminder being harmless`, `that is the product`, and `what it cannot is said` do not occur on any line of the proposed text because of line wrapping or intervening Markdown emphasis; after applying Tasks 1-5 in sequence, those ranges ran to EOF and the loop produced only 10 of 13 parity rows | the required 13/13 gate always fails on the intended implementation, and the script’s `-s` test cannot diagnose a missing end anchor because a start-only extraction is non-empty | choose byte-present single-line end anchors, assert exactly one start and one end match per file before extraction, and reject a range unless the extracted final line matches its end anchor +BLOCKER | high | Commit protocol, Task 6 Step 3, and Task 7 Step 2 | the one-WIP topology records no immutable pre-cycle base and checks only that the tip subject begins `WIP: review-loop economics`; a resumed Task 1 or accidental second WIP can stack commits while still passing `CYCLE OPEN`, after which `git diff HEAD~1` and a Plan-C review based on the tip’s parent omit the earlier WIP | part of the combined A+B+C diff can escape the only Gate-B review, the closing amend can leave an older WIP in history, and the topology no longer guarantees one reviewed diff | record the pre-cycle base SHA before Task 1, require the tip to be the sole WIP child of that base at every handoff, use the recorded base for every diff and Gate-B call, and stop or deliberately collapse stacked WIPs before continuing +MAJOR | high | Task 2 Step 3(e), lines 400-421 | replacing only `invalidates the prior pass ...` leaves the preceding words `a fix changes the diff and the hook`, so the shipped sentence reads `the hook no longer covers the artifact`; the third wording still makes the hook the subject instead of the prior review | the prompt carries another false gate-mechanism rationale and leaves pass-1 M6 unfixed, inviting readers to treat re-review as contingent on advisory hook behaviour | replace the complete sentence from `Re-review after every fix` so it says the changed artifact is no longer covered by the prior pass, then state the hook’s fingerprint/reminder role in a separate sentence +MAJOR | high | Old-conditions accounting, lines 158-205 | the assertion that ten whole passages are byte-identical is false: the Gate-A passage has substantive root-only text and different wrapping, while rows 5a/5b attribute `you report the tells` versus `report the tells` to the pass-report paragraph even though those words live in the following paragraph | the per-copy accounting can hide requirements unique to one mirror and pass-1 M1 is only cosmetically repaired | define an exact start and end for every accounted passage, diff each complete range, and create separate per-copy rows for every non-identical range, including Gate A and any multi-paragraph range intentionally containing the tells rule +MAJOR | high | Old-conditions accounting rows 1 and 5a/5b | row 1 omits the hook limitations that it cannot read findings or distinguish the spec run from the plan run; if row 5 is intended to include the next tells paragraph, it omits the five tells, the any-two threshold, and the mandatory stop-and-surface consequence, while if it is not, condition h is outside the passage | the eleven-passage guard does not account for every condition of the prose it claims to cover, so a later rewrite can drop behaviour while every listed disposition still reads “kept” | inventory every normative condition inside each newly explicit exact range and mark each kept, moved, or deliberately dropped; keep the tells paragraph wholly inside or wholly outside row 5 rather than straddling the boundary +MAJOR | high | Task 4 Steps 1 and 3, lines 535-603 | the baseline pipeline reports only eight of the promised nine conditions because `append one profile-log line` is split across two physical lines, and matching marker phrases cannot prove the paragraph stayed byte-identical; the produced `BASELINE ...` and `AFTER ...` lines also cannot literally be identical because their labels differ | the advertised preservation check can pass after the log-append condition or other surrounding semantics are changed, leaving pass-1 M4 incompletely repaired | capture and compare an exact digest of the original `Changing a profile` paragraph, verify the appended block separately, and state the expected nine-condition diagnostic only as supplemental evidence +MAJOR | high | Tasks 2-6 preflights and Self-Review lines 904-907 | contrary to the claimed four-state repair, Task 2 defines only counts 6 and 0, Tasks 3-5 define only marker counts 0 and 1, and Task 6 merely says “if the tables already exist”; partial counts, duplicates, damaged text, cross-copy asymmetry, and whether the WIP already exists have no prescribed terminal action | interruption or rerun can duplicate blocks, continue from divergent mirrors, fail a commit, or apply an initial commit where an amend is required, so pass-1 M3 remains broadly unresolved | give every mutating task a complete state table over both mirrors and git state, with exact verify-and-skip, reconcile, amend, and STOP actions; make the final self-review describe those actual states rather than referring back only to Task 1 +MAJOR | high | Finding ownership table lines 75-78 and Plan-C handoff | pass-1 M8 recovery, M9 floor-knob byte preservation, M10 broad staging, and the fixed `/tmp/plan-a-close.txt` Minor belonged to the removed Gate-B close, but the table marks them repaired by Plan A and points to unrelated Plan-A checks; B6 and B7 are merely promised to Plan C, and there is no receiving Plan-B or Plan-C artifact in which any of these inherited conditions can be verified | deleting the Gate-B section has relocated rather than repaired the findings, so the eventual combined close can repeat the wrong branch-recovery deletion, lose knob-byte evidence, stage unrelated changes, use an unsafe close-body file, omit revalidation, or apply the new severity rule to the pre-rule cycle | assign each moved finding to Plan C explicitly, copy its full old conditions and acceptance test into a mandatory handoff contract, and require that receiving plan to exist and pass Gate A before the shared WIP can advance to the step that depends on it +MAJOR | high | Task 6 Steps 1, 2, and 4 | the thirteen-row parity table and twenty-four-row conformance table have no exact table template, unique marker, row-count command, duplicate check, or stop rule for `DEFECT` and `FAIL`; moreover the whole-artifact conformance result is recorded before Plans B and C make further edits to the same command prompt | the executor can append malformed or incomplete result records, proceed with failed prompt-standard items, and leave a stale “resulting artifact” assessment in the combined commit while pass-1 M7 appears closed | provide exact table schemas and mechanical cardinality/status validation that exits nonzero on any defect, label Plan-A-only snapshots as non-final, and defer or repeat the whole-artifact conformance gate after the final Plan-C text +MAJOR | medium | Task 1 Step 6 and every later amend | the plan has no clean-worktree and inspected-hunk preflight before `git add` stages the complete two prompt files, and the shared-WIP protocol does not forbid concurrent executors | pre-existing or concurrently arriving edits inside either target file can be folded into the combined WIP, reviewed under the wrong story, or mixed across plan boundaries even though the cached name list still contains exactly the two expected paths | require a clean starting worktree and index, serialize A→B→C to one executor at a time, inspect the complete staged diff before every amend, and stop on any hunk not produced by the active task +MINOR | medium | Task 3 replacement, lines 490-495 | every pass must report the risk and security values read, but the shipped text gives no representation for the no-story and unprofiled branches where no axis values exist, nor an explicit per-story mapping for a multi-story set | agents can omit the fields, invent axis values, or emit an ambiguous list, weakening the derivation visibility this paragraph exists to add | state explicit report values for `no story` and `unprofiled`, and require each cited story to be paired with the risk/security values read from its header +MINOR | high | Self-Review lines 890-907 | the plan says only two checks intentionally lack numeric expectations and that every task distinguishes not-started, complete, duplicate, damaged, and asymmetric states, but Task 4 has no expected nine-marker baseline, the result tables are only declared “asserted”, and Tasks 2-6 lack those state machines | the repaired artifact’s own audit trail overstates what was checked and obscures the remaining pass-1 defects | after repairing the executable checks, enumerate every genuinely non-numeric judgement and every task’s actual state matrix; remove the blanket claims until they are mechanically true +END OF FINDINGS (16 total) diff --git a/.context/codex-reviews/gate-a-plan-plana-pass-3-dispositions.md b/.context/codex-reviews/gate-a-plan-plana-pass-3-dispositions.md new file mode 100644 index 0000000..1cffbb5 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-plana-pass-3-dispositions.md @@ -0,0 +1,111 @@ +# Gate A — Plan A cycle — pass 3 dispositions + +Advisory human note. Not a findings file; participates in no pass validation. + +## Result + +Artifact 3799ac9 (revision 3). VALID pass: terminator exact, 16 finding lines, 0 +non-finding lines. **6 BLOCKER · 7 MAJOR · 2 MINOR · 1 NIT = 13 Blocker/Major.** + +Routed contract: pass 3 <= 6 B/M; **above ~10, or instrument-dominated again, routes back.** +Both conditions are met. **The standing escalation fires** — this goes to Daniel, not into +another repair round. + +## Curve + +| pass | findings | Blockers | Majors | B+M | +|---|---|---|---|---| +| 1 | 21 | 7 | 10 | 17 | +| 2 | 16 | 6 | 8 | 14 | +| 3 | 16 | 6 | 7 | 13 | + +**Blockers have not fallen in two rounds: 7 → 6 → 6.** Findings flat at 16. + +## The cluster, counted rather than asserted + +| category | findings | B+M | +|---|---|---| +| **instrument** — checks, the base-SHA mechanism, commit staging | B1-B6, M1-M6, MINOR 9 = **13** | **12** | +| **shipped rules** | M7 = **1** | **1** | +| **prose about the plan** | MINOR 8, NIT 10 = **2** | 0 | + +**Twelve of thirteen Blocker/Major are instrument — 92%.** One finding in the entire pass +concerns the rules being shipped. + +## Tells + +1. finding count rising — **no**, flat at 16. +2. **Blocker count failing to fall — PRESENT.** 7 → 6 → 6. +3. **clustering on the instrument — PRESENT, 92%.** +4. clustering on prose about either — marginal (MINOR 8, NIT 10). +5. require↔withdraw pair — **none.** What is happening is escalation, not reversal: each + round's fix is accepted and then found insufficient at a deeper level. Pass 2 said bound + the extractor; revision 3 bounded it; pass 3 says the bound must also assert uniqueness + and line counts. + +Tells 2 and 3 present → mandatory stop, independently of the routed threshold. + +## What the three rounds actually show + +**The rules are converging and the instrument is diverging.** Across pass 3 the shipped +replacement text drew exactly one Blocker/Major (M7). Everything else is about the +machinery that checks the edits, and that machinery has grown for three rounds while the +reviewer keeps finding new ways it can report success with the thing it checks absent: + +- **B5** — Task 2's verification proves the six OLD texts are gone and never that the NEW + texts arrived. **Deleting the six sentences outright produces the plan's exact pasted + observations.** The check cannot fail in the direction that matters. +- **B4** — the `awk` scoping I introduced to stop a check passing before its edit is itself + unclosed: if the end marker is absent the range runs to EOF and matches the pre-existing + Profiles-section occurrence, so all four Task-1 counts can read `1` with the predicate + never inserted. +- **B1** — Task 6's preflight greps for strings that occur in **its own prose and its own + command**. Executed on the committed plan it returns **4**, not the pasted `0`. +- **B6** — the parity loop asserts 13 iterations and 13 equal extractions but never the + documented line counts, so the extractor failure those counts exist to catch passes. +- **B2, B3** — the base-SHA repair does not hold: the record is rewritten unconditionally on + a resumed Task 1, and `rev-list --count base..HEAD = 1` does not prove `base` is HEAD's + parent. The reviewer **demonstrated** the second by constructing a sibling commit with the + same tree: `n=1`, correct WIP subject, zero-file diff, `parent_equal=no`. + +Each of those is a real defect. The pattern across three rounds is that a plan carrying its +own executable verification in prose keeps generating a larger instrument than the change it +verifies — which is what the escalation was set up to catch. + +## Verified before routing + +- **B1 — CONFIRMED by execution.** `grep -c 'PARITY RESULTS — Plan A\|CONFORMANCE RESULTS — Plan A'` + on the committed plan returns **4**; the strings occur at lines 771, 775, 852, 878 — the + preflight's own prose, its own command, and both table-recording instructions. +- **MINOR 8 — CONFIRMED, and it is mine twice over.** The plan says the simulation applied + "all sixteen edits"; there are **13** OLD→NEW pairs. I noticed this while reporting and + corrected it in conversation, then left the wrong number standing in both the artifact and + the commit message. Noticing an error is not fixing it. +- **NIT 10 — CONFIRMED.** The Plan-C inheritance table attributes evidence revalidation to + pass-2 B6; it is pass-**1** B6. Pass-2 B6 is the base-SHA defect, which Plan A owns. + +## Disposition + +All 13 Blocker/Major carried open. **No fixes applied — the escalation is explicit that this +is not to be absorbed.** + +Two findings would need fixing wherever the checks end up living, because they are not +instrument: **M7** — the Lenses paragraph still ends "The floor, the Blocker/Major filter, +the file-first findings protocol and the clean-final-pass rule are unchanged", so the +shipped sentence still tells a reader the floor is unchanged in the change that makes it +profile-dependent; and **MINOR 8 / NIT 10**, both plain factual errors in the plan's prose. + +## Lesson logged: noticing an error is not fixing it + +MINOR 8 — the plan claimed the simulation applied "all sixteen edits" when there are +thirteen — was **caught by me, in conversation, one turn before the pass that found it.** +The generator script's summary line hardcoded `16`; I saw it, said so in the status report, +and moved on. The wrong number stayed in the artifact and in the commit message, and the +reviewer had to spend a finding on it. + +The failure is not the arithmetic. It is treating a spoken correction as a completed one. +A correction that lives only in the conversation is invisible to every later reader of the +artifact, and the conversation is exactly where it feels most like the work is done. + +**Rule taken from it:** when an error is noticed in a written artifact, fix the artifact in +the same turn or record it as an open item. Saying it aloud counts as neither. diff --git a/.context/codex-reviews/gate-a-plan-plana-pass-3.md b/.context/codex-reviews/gate-a-plan-plana-pass-3.md new file mode 100644 index 0000000..34b32b0 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-plana-pass-3.md @@ -0,0 +1,17 @@ +BLOCKER | high | Task 6 Step 1 and Self-Review “Every expected value is an observation” | the exact preflight command returns `4` on revision 3 before Task 6, not the stated `0`, because the two marker strings already occur in the preflight prose, its command, and both table-recording instructions; after appending the two tables it would return `6`, not `2` | Task 6 always misclassifies the untouched plan as already run, cannot distinguish complete from partial output, and directly falsifies the central pasted-observation claim | anchor the grep to exact table-heading lines or use dedicated sentinel lines absent from the instructions, then execute and paste the real `0`/`2` observations +BLOCKER | high | Commit protocol, Task 1 Step 1, and Task 1 rerun path | `git rev-parse HEAD > .context/plan-a-base-sha` unconditionally overwrites the value called immutable; after the first WIP, a resumed Task 1 records that WIP as the new base, and an accidental second WIP is then one child above it and passes the handoff count while the first WIP is excluded; the redirection also follows an existing ignored symlink and can truncate its target | the stacked-WIP hole remains open on the exact resume path the repair claims to close, and hostile or stale per-clone state can cause data loss | create the record atomically only when absent, refuse symlinks and non-regular files, validate an existing value without rewriting it, and recover the original pre-cycle base from a durable authenticated record +BLOCKER | high | Commit protocol handoff check and every later diff/Gate-B base | `git rev-list --count "$base"..HEAD = 1` does not prove that `base` is HEAD’s parent or even an ancestor; in an independent replay, pointing the file at a sibling commit with the WIP’s same tree produced `n=1`, the required WIP subject, and a zero-file diff while `parent_equal=no` | a stale or modified base record can make every topology check pass and make the only Gate-B review see an empty diff, omitting all Plan-A changes | require a single-parent WIP whose exact parent SHA equals the recorded base, verify ancestry explicitly, and reject any base that is not that parent before any diff, reset, or Gate-B call +BLOCKER | high | Task 1 Step 5 predicate verification | the `awk '/HARD FLOOR/,/never read for this derivation/'` range is not closed when the new end marker is absent; if Step 4’s residual and gate-off blocks land while Step 3’s old HARD FLOOR remains, awk runs to EOF and finds the pre-existing Profiles-section `max(risk, security)`, so all four advertised verification counts can still be `1` | Task 1 can green and commit without the derived-floor predicate, silently leaving the fixed three-pass rule in force | assert exactly one new start and end marker and zero old marker before extracting, make a missing end a nonzero failure, and compare the exact resulting floor blocks rather than one token inside a potentially unbounded range +BLOCKER | high | Task 2 Step 5 | verification proves only that the six old regex fragments disappeared; except for the separate Minor sentence, it never asserts that replacements (a)–(f) exist, so deleting those old sentences outright yields the stated `0/0`, `1/1`, `0/0`, `1/1` observations | Gate-A wording, incomplete-pass accounting, the re-review rationale, and the lenses rule can all be absent while Task 2 reports success and amends the WIP | check each of the six OLD texts is zero and each complete NEW text is exactly one in each copy, with a failing status for missing, duplicate, or asymmetric results +BLOCKER | high | Task 6 Step 2 parity loop | the final predicate checks only 13 iterations and 13 pairwise-equal extractions; it never asserts the documented expected line count for any anchor and awk selects only the first occurrence, so an identically overrun, truncated, or duplicated block can still print `PARITY-COMPLETE 13/13` | the check can return success for the extractor failure its line-count expectations are supposed to detect, or for duplicate shipped rules left by a rerun | encode the 13 expected line counts in the loop, assert every anchor occurs exactly once in both files, and increment `ok` only after uniqueness, bounds, line count, and byte parity all pass +MAJOR | high | Tasks 3 and 5 verification blocks | these checks count one marker from each inserted rule plus a few untouched anchors, not the complete replacement blocks; most of the every-pass reporting rule, severity exclusions, activation fallback, or downstream-adoption warning can be deleted or changed identically in both copies while every stated count remains `1` | truncated or semantically wrong product prompts can be committed and handed to Plan B with all per-task checks green | compare each complete resulting block against the plan’s exact NEW bytes and assert unique boundaries, leaving the later parity check to prove mirror equality rather than using marker counts as content proof +MAJOR | high | Task 4 Steps 1 and 3 and old-conditions row 10 | the nine-marker pipeline is not a preservation check: it can still return `9` after deleting conditions such as “in both directions” or “raised or lowered”, or after changing the surrounding decision while retaining each marker phrase | a profile-change condition can be dropped while the plan claims all nine stayed verbatim, repeating the decision-procedure loss invariant this check exists to prevent | extract the original paragraph at baseline, save a POSIX digest or exact temporary copy, require byte identity after the append, and keep the nine markers only as supplemental diagnostics +MAJOR | high | Tasks 2–5 preflights and Self-Review “Rerun and interruption” | only Task 1 has a state matrix; Tasks 2–5 state not-started and complete counts but prescribe no action for partial replacements, duplicate blocks, cross-copy asymmetry, or whether the WIP commit already exists, while the Self-Review claims those preflights distinguish the relevant states | abandonment between mirrors or a rerun can fail on a missing OLD string, amend divergent copies, duplicate insertions, or attempt the wrong commit operation before parity notices much later | give every mutating task an explicit per-copy OLD/NEW count matrix plus git state, with verify-and-skip, reconcile, amend, and STOP actions, and remove the blanket Self-Review claim until those states exist +MAJOR | high | Task 1 Step 6 and Tasks 2–6 amend steps | staging the complete two prompt files plus `git status --porcelain` cannot distinguish plan-owned hunks from pre-existing or concurrently arriving edits inside those same files, and later amends do not inspect the staged diff before committing | unrelated user work can be folded into the shared WIP and reviewed or closed under the wrong story; an edit inside an accounted passage can be mistaken for plan output | require clean target blobs and index at the recorded base, serialize writers, inspect the full cached diff before every commit/amend, and stop on any hunk not produced by the active task +MAJOR | high | Commit protocol stacked-state recovery message | every failed topology check tells the executor to run `git reset --soft $base, then one WIP commit` without first classifying the commits above base | legitimate or concurrent commits can lose their branch-visible boundaries, messages, and authorship and be folded into this story; a corrupt base record makes the reset target itself unsafe | STOP and inspect the full ancestry/diffs first, preserve a recovery ref, and permit a collapse only for commits positively identified as this cycle’s WIPs with explicit human approval when unrelated commits exist +MAJOR | high | Commit protocol handoff to Plans B and C | the sole base record is under gitignored, per-clone `.context/`, but no missing-file or fresh-clone recovery procedure exists; a new session in the same clone usually works, while a cleared context or another clone cannot establish the required base | a legitimate cross-session or cross-machine handoff is stranded, or an executor is tempted to guess `HEAD~1`, reopening the omission the protocol was added to close | put the original base SHA in the WIP commit body or another durable cycle record and define a fail-closed recovery that cross-checks it against the WIP’s exact parent +MAJOR | high | Task 2 replacement (f), resulting Profiles “Lenses” paragraph, and old-conditions row 9 | the repair still produces `The floor, the Blocker/Major filter, the file-first findings protocol and the clean-final-pass rule are unchanged`, so the same paragraph continues to call the floor unchanged in the change that replaces a fixed floor with a derived one; row 9’s claim that this overstatement was removed is therefore false | readers receive mutually inconsistent instructions and can retain the fixed-floor interpretation this change is meant to eliminate | exclude the floor from the unchanged list and say explicitly that lenses do not alter the derived-floor rule stated above, while the filter, file protocol, and clean-final-pass rule remain unchanged +MINOR | high | Revision-3 central simulation claim, lines 35–41 | the plan says the simulation applied “all sixteen edits” to both copies, but the executable plan contains 13 OLD→NEW replacement pairs: 2 in Task 1, 7 in Task 2, 1 in Task 3, 1 in Task 4, and 2 in Task 5 | the completeness claim and generated-source audit trail disagree with the artifact they are meant to authenticate, making omissions harder to diagnose | state 13 replacements, enumerate stable replacement IDs, and generate the count from the same replacement manifest used by the replay +MINOR | medium | Task 3 Step 1 and Task 5 Step 1 preflights | the standalone `grep -cF` commands correctly print the expected zero counts but exit status 1 when no match exists; the independent replay observed exit 1 for both nominal preflight states | an executor or wrapper that treats a nonzero check as failure can abort before valid work or misreport the expected untouched state as an error | normalize the expected-no-match status explicitly while preserving real read errors, or use an awk count that exits zero after successfully reading both files +NIT | high | “Findings inherited by Plan C” table | `pass-2 B6` is attributed to evidence revalidation, but the sixth Blocker in the pass-2 findings file is the base-SHA/topology defect; evidence revalidation is pass-1 Blocker 6 | the handoff requirement itself is present, but its provenance is wrong and future repair audits cannot reliably map it back to the originating finding | relabel the row as pass-1 B6 and record the pass-2 base-SHA finding as repaired here only after the immutable-base and exact-parent defects are closed +END OF FINDINGS (16 total) diff --git a/.context/codex-reviews/gate-a-plan-plana-pass-4.md b/.context/codex-reviews/gate-a-plan-plana-pass-4.md new file mode 100644 index 0000000..5404c57 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-plana-pass-4.md @@ -0,0 +1,8 @@ +MAJOR | high | Resulting §5: Task 12 severity rule versus the untouched “Any two present makes stop-and-surface mandatory” paragraph | The new procedure demotes consequence-free instrument and rationale findings to Minor or below and says to collect them without iterating, while the existing five-tells rule makes instrument clustering plus prose clustering sufficient for a mandatory stop even on a Blocker/Major-free pass at or above the floor; the earlier clean-completion precedence is stated only against the separate “clearly stuck” exit, and the tells rule expressly says that reading is not its precondition | The same pass is simultaneously eligible to close and required to stop, so the new demotion can fail to reduce loop cost on exactly the finding classes it targets | State whether clean Blocker/Major-free completion at or above the floor also overrides the two-tell stop, or restrict the tell threshold to findings that remain gating under the severity procedure +MAJOR | high | Task 13 “Downstream has no shipping commit” against Tasks 1 and 3–9 | The partial-adoption warning names only the floor-without-severity coupling, but a merge that takes Task 1 without all of the coordinated numeric replacements leaves the derived-floor rule beside explicit `pass 3`, `below 3`, and `3-pass floor` obligations; the text supplies no precedence for that direct contradiction and then overclaims that what prompt text can do has been done | A downstream project can acquire two incompatible floors, including a newly stated floor its cited profiles do not license, and the contradiction may persist under invariant 9 | State that a partial floor adoption can leave contradictory pass obligations and must stop for resolution before a gate runs, or otherwise define a strict precedence for that state; remove the categorical “what prompt text can do is done” claim +MAJOR | medium | Task 1 “One derived value governs all three cycles” versus Task 11 “these pass-count rules apply while a gate is running and are silent otherwise” | No rule defines a work-wide cited set or handles a profile or cited-set change after one cycle closes but before the next opens; a level change in that interval can therefore leave the completed Gate-A cycle at one numeric floor and the later cycles at another even though Task 1 promises one value for all three | Pass debt and provenance become ambiguous at an ordinary between-cycle edit, and an agent must either reopen a closed cycle without authority or violate the one-value claim | Define the shared cited-set/profile snapshot and the effect of between-cycle changes on already closed cycles, or narrow the promise to one derivation predicate applied independently to each cycle’s current set +MINOR | medium | Task 13 “A user knob set above 3 is not lowered by this fallback” | In the activation paragraph this reads as though an unknown-start cycle may owe the higher knob value, contradicting Task 1’s precedence rule that the knob is only a reminder threshold and controls no pass obligation | A cautious reader can recreate an unlicensed floor above 3 and spend extra passes, the central class this change is meant to remove | Say explicitly that the fallback obligation is 3 while the hook’s independently configured reminder threshold remains unchanged and non-binding even when it is higher +MINOR | high | Task 11 “a pass run before it joined did not cover them” | The categorical antecedent covers the new story’s lens set, evidence obligations, and review scope, but adding a high-risk story to a set that already contains another high-risk story adds no new risk lenses, and review scope can overlap even though the new story’s per-story evidence was not owed | The rationale overstates what the earlier pass failed to cover and repeats the gate-coverage overclaim class called out in AGENTS.md | Ground the consequence in accountability: the earlier pass was not run against the current cited set and did not owe the new story’s per-story evidence, regardless of coincidental lens or scope overlap +MINOR | high | Task 12 kinship sentence and untouched §5 “What counts as prose” exemption | Calling the two rules “one principle at two granularities” hides a real disagreement: the path rule exempts every `docs/**.md` change as explanatory, while the new finding-level rule treats operational consumption as decisive and `docs/hardening-log.md` is explicitly consumed to choose hardening escalation | A reader can infer equivalence where the same operational document is non-product under the path rule but product-affecting under the severity rule | Describe the kinship as a shared intent but state that the path exemption is coarse and can diverge from the operational-reader test, or remove the equivalence claim +NIT | high | Task 10 pass-report rationale | A number-only report does not make the reader unable to check the derivation, and the running pass is not literally the only time checking is cheap: the reader can reconstruct it from the artifact and story headers, and Plan B’s provenance record exists for later reconstruction | The absolute wording overstates the benefit of the reporting fields and weakens trust in the surrounding calibrated mechanism claims | Say that a number alone forces external reconstruction and that the running pass is the cheapest time to check, rather than the only possible cheap check +END OF FINDINGS (7 total) diff --git a/.context/codex-reviews/gate-a-plan-plana-pass-5-dispositions.md b/.context/codex-reviews/gate-a-plan-plana-pass-5-dispositions.md new file mode 100644 index 0000000..35b359e --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-plana-pass-5-dispositions.md @@ -0,0 +1,132 @@ +# Gate A — Plan A cycle — pass 5 dispositions + +Advisory human note. Not a findings file; participates in no pass validation. + +## Result + +Artifact 895a94d (revision 5). VALID pass: terminator exact, 5 finding lines, 0 non-finding +lines. **0 BLOCKER · 4 MAJOR · 1 NIT = 4 Blocker/Major.** + +Three Majors repaired in revision 6 (26e7885). **One routed as a contract question.** + +## The three pass-report lines (owed from pass 4) + +**1 — Trend.** + +| pass | findings | Blockers | Majors | B+M | +|---|---|---|---|---| +| 1 | 21 | 7 | 10 | 17 | +| 2 | 16 | 6 | 8 | 14 | +| 3 | 16 | 6 | 7 | 13 | +| 4 | 7 | 0 | 3 | 3 | +| 5 | 5 | 0 | 4 | 4 | + +Findings falling steadily. **Blockers at zero for two consecutive passes.** B+M ticked 3 → 4, +which is noise at this size and not a plateau: the four are a different set from pass 4's +three, and three of them are consequences of pass 4's repairs rather than survivals. + +**2 — Cluster.** Product behaviour, entirely. **Zero instrument findings for the second pass +running.** Every finding concerns the prose that will ship into `CLAUDE.md` §5 and the +template. That is the strip working: passes 1-3 spent 92% of their Blocker/Major budget on +the plan's own checks; passes 4-5 spent none. + +**3 — require↔withdraw.** **None.** Nothing in pass 5 demands what an earlier pass removed. +The closest candidate is not one: pass 4 required a precedence rule between demotion and the +tells, and pass 5 narrowed it — that is refinement of an accepted requirement, not a reversal. + +**Tells present: none.** Findings falling, Blockers zero, clustering on product behaviour, +no require↔withdraw pair. This is a converging loop. + +## Repaired — absorbed under §5's absorb rule + +All three correct a pass-4 fix and stay inside the assigned fix set, so they belong in this +loop rather than being handed back. + +- **M2** — the clean-completion precedence was broader than the collision it repaired. "A + Blocker/Major-free pass at or above the floor closes" would have let the new rule bypass + §5's untouched scope stops, which fire when a finding leaves the assigned fix set or opens + a new structural question **at any severity**. Now limited explicitly to the two-tell stop, + and it says outright that it overrides nothing else. +- **M3** — "demotion does not change what the tells observe" was wrong in one direction, and + the reviewer found the direction. Demoting a Blocker to Minor **does** remove it from the + Blocker curve; that is what demoting is for. Stated per tell now: total, clusters and + require↔withdraw see every reported finding; the Blocker curve reads severity after the + ceiling. +- **M4** — the partial-adoption trigger named three spellings and missed the ones a partial + merge happens to leave. Three further sites carry a fixed-three claim — the Gate-A loop + description, the pass-1 closure rule, the re-review rationale — and a merge can take some + tasks and not others. The trigger is semantic now, not a list. + +## Routed — a contract question, and it stops the loop + +**M1.** Pass-4 M3 found that spec §2's promise is undefined across an ordinary profile edit +between cycles. Revision 5 answered by weakening it to a per-cycle derivation. Pass 5 shows +that contradicts the approved spec, which reads: + +> **One derived value governs the Gate-A spec loop, the Gate-A plan loop and the Gate-B +> cycle.** Not because they are one cycle — §5 is explicit that they are **three separate +> cycles** — but because they derive from **the same cited-story set**. + +Verified against revision 36 directly. The spec says one **value**; revision 5 shipped one +**derivation**. Those are different cross-cycle contracts, and the story's criterion tracks +the spec's wording. + +**Why this stops the loop rather than being absorbed:** §5 — "When a finding is both — it +corrects the last correction *and* opens a new structural or contract question — the new +question wins and the loop stops." Absorbing it would settle a contract by drafting, which +is the failure mode that rule exists to prevent. Two answers are available and neither is an +agent's to pick: define the shared snapshot the spec promises, or revise the spec and the +story criterion to approve per-cycle derivation. + +Task 1 now carries a note that its between-cycle sentence is awaiting that decision and must +not be executed until resolved. + +## Collected, not iterated + +**NIT** — the plan's stated test for the re-review sentence, "delete the hook and the +sentence is still true", is literally false of the sentence's trailing clause ("The hook +merely notices, at commit time"), which describes the hook and cannot outlive it. The +load-bearing half does survive the test. The task note now says which half the test applies +to. Collected under §5; a Nit never earns a repair round, and this one got a wording +adjustment only because the note was being rewritten anyway. + +## M1 resolved as (C), current-header-governs — verified before encoding + +Neither of the two options I routed. A third reading was returned and I was asked to verify +it textually before encoding it. **It verifies**, and the verification changed what I think +the original finding was. + +**Quoted from revision 36 and from `CLAUDE.md`, checked one at a time:** + +1. §2 — *"**One derived value governs the Gate-A spec loop, the Gate-A plan loop and the + Gate-B cycle.** Not because they are one cycle — §5 is explicit that they are **three + separate cycles** — but because they derive from **the same cited-story set**."* + This argues from the **source**, never from time. "One value at any moment because one + source" satisfies it; nothing in it freezes a number. +2. §2.4 — *"The floor derives from the **current** profile at each pass."* Directly supports + current-header-governs and directly contradicts a snapshot. +3. §2.4 — *"**These are pass-count rules, so they apply while §5 says a gate is running** and + are silent otherwise. What §5 says about when a gate runs — **including how a moving + profile or cited set bears on that** — is §5's, unchanged and deliberately not summarised + here."* The spec **delegates** the between-cycle case rather than leaving it open. +4. `CLAUDE.md:387-389` — *"The **story header is the single writable copy** … read the values + fresh at each pass, **never a remembered or copied value**."* Option (A)'s snapshot + violates this in those words. (C) satisfies it. + +**No sentence contradicts (C).** Encoded in Task 1, plan-only, citing §2's sentence. + +**What the verification changed.** Pass-4 M3's premise — that no rule handles a between-cycle +change — was **partly wrong**, and so was my acceptance of it. §2.4 quote 3 shows the spec +deliberately routed that question to §5, which answers it by reading fresh. I treated a +delegation as a gap and invented an answer, and the invented answer then contradicted §2. +The lesson is narrower than "verify findings": **a finding that says "no rule covers X" needs +the same check as a finding that says "rule Y is wrong" — that no rule covers it.** + +**One correction to the routed reading:** it cited "§2.4/§4.2". **There is no §4.2.** §4 is +"The per-pass curve — required properties" and has no subsections. All the supporting text is +in §2.4. The substance is unaffected; the citation is not. + +**Why (A) was wrong in a way worth keeping:** it fails invariant 2's firing direction. Under a +snapshot, a human-confirmed **raise** between cycles would leave in-flight work reviewed under +the weaker profile — the under-review direction, which is the one this repo treats as +dangerous. I recommended (A) and missed that. diff --git a/.context/codex-reviews/gate-a-plan-plana-pass-5.md b/.context/codex-reviews/gate-a-plan-plana-pass-5.md new file mode 100644 index 0000000..544c13e --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-plana-pass-5.md @@ -0,0 +1,6 @@ +MAJOR | high | Task 1 lines 198–204 against design §2 and story criterion 1 | Revision 5 changes the settled requirement from “one derived value governs all three cycles” to independent per-cycle derivation and explicitly allows a between-cycle profile or cited-set change to give later cycles a different value; that is not merely a clarification, and “the same predicate reads the same cited-story set” is no longer a reason for one value when each cycle reads a different snapshot | Plan A under-delivers the approved spec and acceptance criterion, so the shipped copies can pass Gate B while implementing a different cross-cycle contract | Either define and retain the shared snapshot/value the spec requires, or first revise the spec and story criterion to approve per-cycle snapshots and their potentially different floors; do not claim the latter implements the former +MAJOR | high | Resulting §5 Task 12 clean-completion paragraph versus the untouched assigned-fix-set and new-question stop rules | “Where a pass is Blocker/Major-free at or above the floor ... the cycle closes” is broader than the collision it repairs: the untouched rules independently require stopping when even a non-gating finding leaves the assigned fix set or opens a new structural or contract question | A clean pass carrying an out-of-scope Minor or a new contract question is simultaneously required to stop for a human decision and told to close, allowing the new precedence rule to bypass a mandatory scope stop | Limit the precedence explicitly to the two-tell condition, for example: if no other §5 stop applies, the two-tell threshold alone does not bar clean completion at or above the floor +MAJOR | medium | Task 12 lines 767–770 against the untouched five-tells definition | The blanket claim that demotion “does not change what the pass-report tells observe” leaves the Blocker-count tell contradictory: demoting a Blocker to Minor necessarily removes it from the Blocker curve, while the following sentence preserves only instrument and prose clustering and never says that total findings and require↔withdraw still count but the Blocker series uses post-demotion severity | Agents can either keep demoted findings in the Blocker count, defeating the new severity semantics and triggering false stop-and-surfaces, or remove them and violate the blanket instruction | State the per-tell rule precisely: all reported findings remain in the total, subject clusters, and require↔withdraw analysis, while the Blocker curve reflects severity after the reachability ceiling +MAJOR | high | Task 13 lines 835–839 against Tasks 3–9 and partial merges allowed by invariant 9 | The partial-adoption stop names only `pass 3`, `below 3`, and `3-pass floor`, but coordinated legacy obligations also include `each its own 3-pass loop`, the pass-1-Minor rule that forbids closing at a derived floor of 1, and the hook rationale saying where “the 3” come from; a partial merge can take Tasks 1, 3, 4, 5, and 8 while leaving Tasks 6, 7, and 9, so none of the named obligations remains even though the resulting rules still contradict the derived floor | The new repair leaves a concrete partial-adoption state that can run without the mandated stop and state a floor the cited profiles do not license, preserving the central gate-off risk | Make the trigger semantic and exhaustive over this edit set—any surviving fixed-three or pass-1 closure claim beside the derived predicate stops—or name all coordinated Tasks 3–9 conditions and require completing or reverting the adoption before a gate runs +NIT | high | Task 7 lines 472–485 | The plan says the replacement passes the literal test “delete the hook and the sentence is still true,” but the replacement itself asserts “The hook merely notices, at commit time”; after deleting the hook that clause is false even though the re-review rationale remains valid | The third-attempt acceptance claim is mechanically untrue and keeps an unnecessary hook claim in the sentence whose truth was meant to be hook-independent | End the replacement after “the prior review no longer covers it,” or restate the test as independence of the re-review obligation rather than truth of the whole replacement +END OF FINDINGS (5 total) diff --git a/.context/codex-reviews/gate-a-plan-plana-pass-6.md b/.context/codex-reviews/gate-a-plan-plana-pass-6.md new file mode 100644 index 0000000..f71e76b --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-plana-pass-6.md @@ -0,0 +1,3 @@ +MAJOR | high | Task 1 NEW, especially “one source” and the open/future-cycle rule, against untouched §5 Profiles and Task 11 | The current-header decision gives each cited story one authoritative profile header, but nothing gives the cited-set membership one authoritative source: untouched §5 defines it from “the artifact” and the spec, plan, and Gate-B inputs can carry different citations. “Read fresh at each pass” therefore yields one value per set, not necessarily one value across all three cycles; with two cycles open over disagreeing artifact citations, one can derive 1 while the other derives 3 at the same moment, and no rule says which set is current or stops on the disagreement. §2.4 is explicitly silent outside a running gate, so it also cannot supply the claimed between-cycle answer | The shipped text can claim the spec’s one-value contract while permitting concurrent or successive cycles to use different values, making the principal revision-7 repair non-actionable in exactly the disagreement state it must settle | State one prompt-level authority for the current cited set across the change and require stop-and-surface when open-cycle or artifact citations disagree or change concurrently; if no such shared set is intended, revise the spec and story criterion to approve per-cycle sets rather than calling them one value +NIT | high | Old-conditions accounting row 1, disposition for condition b | The table says condition b, “Blocker/Major only,” is kept verbatim outside the replaced range, but it is inside Task 1’s OLD block and is re-emitted inside Task 1’s NEW block | The condition is preserved, but the sole accounting record misstates how it is preserved and weakens confidence that its range claims were checked against the actual old text | Split the disposition: b is kept verbatim in the replacement; i–m are kept verbatim outside the replaced range +END OF FINDINGS (2 total) diff --git a/.context/codex-reviews/gate-a-plan-plana-pass-7.md b/.context/codex-reviews/gate-a-plan-plana-pass-7.md new file mode 100644 index 0000000..f23a284 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-plana-pass-7.md @@ -0,0 +1,4 @@ +MAJOR | high | Task 1 NEW cited-set authority and the preceding §2.4 note | Absorbing a conservative disagreement stop is compatible with spec §2's same-set contract, but “the citations of the artifact under review” is not one authority across the change: each cycle reviews a different artifact, a Gate-B diff need not cite any story, incidental story references are not distinguished from governing citations, and no cycle is told to compare its set with an open or already-closed cycle's artifact. The current spec already cites the successor loop-rule story while this plan cites only the assigned pass-floor story, so the ambiguity is present in the artifacts implementing the rule; “would run” also does not clearly reach a successive cycle after the first has run. The note's claim that §5's read-fresh rule answers the between-cycle question therefore remains false: it yields a fresh value per artifact-defined set, not one shared set | Concurrent or successive cycles can silently derive different floors, or Gate B can derive from an empty/incidental set, while the shipped text claims the one-value contract is satisfied; the new stop can be missed because the cycle lacks both a shared carrier and an instruction that discovers the disagreement | Define one change-level carrier for the governing cited set, distinguish it from incidental references, and require each open or new cycle to reconcile its artifact's set with that carrier and any prior-cycle record; on mismatch, surface both sets and resume only after they agree. If choosing that carrier is not already a settled contract, route that choice to the human rather than absorbing it +MAJOR | high | Task 13 downstream partial-adoption trigger, especially “sitting beside the derived predicate” | The trigger is one-directional: it stops when Task 1's predicate lands while an old fixed-number or specific-pass obligation survives, but it does not stop the inverse partial merge or partial revert. For example, Task 6 can leave “loop at the derived floor,” or Task 8 can say the floor derives from the profile and cited set, while Task 1 is absent and the old fixed-three definition remains; other dependent “the floor” edits can likewise land without any derived predicate. `/workflow-init` explicitly permits a user-selected merge, so these are reachable downstream states | A partially adopted project can have an undefined or contradictory floor and run or close a gate without triggering the promised stop, creating another way to state or act on a floor the cited profiles do not license | State the coherence rule bidirectionally and semantically: exactly one floor definition must exist and every pass-count and closure reference must resolve to it; stop both when fixed and derived definitions coexist and when any derived-floor claim or dependency exists without the predicate that defines it +NIT | high | Task 1 NEW, “a raise costs an affected open cycle a further pass under the current profile, as above” | No preceding text states the one-further-pass rule; the detailed rule is appended later by Task 11 in the Profiles section | The cross-reference sends a reader to the wrong place and makes an already dense lifecycle rule harder to verify | Replace “as above” with “as the Profiles section below requires” or remove the cross-reference +END OF FINDINGS (3 total) diff --git a/.context/codex-reviews/gate-a-plan-plana-pass-8.md b/.context/codex-reviews/gate-a-plan-plana-pass-8.md new file mode 100644 index 0000000..26694d8 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-plana-pass-8.md @@ -0,0 +1,5 @@ +MAJOR | high | Task 12 NEW lines 783-794 and old-conditions accounting rows 5a, 5b and 11 | The added two-tell precedence is not severity semantics from spec §3: it changes the untouched rule that any two tells make stop-and-surface mandatory into a rule that permits completion whenever the pass is Blocker/Major-free at or above the floor. The spec and story explicitly move loop-rule consolidation out of this change, while the accounting calls the two-tell condition untouched | A cycle can now close instead of handing two convergence warnings to the human, and the plan ships a contract decision outside its assigned spec while falsely recording the old condition as preserved | Remove this precedence from Plan A, or route the interaction for an explicit spec decision and then account for the two-tell condition as deliberately changed rather than untouched +MAJOR | high | Task 1 NEW lines 202-220, especially the Story-header authority and disagreement stop | The chosen header convention is not operationally complete across the three cycles. A spec and plan are files with Story headers, but Gate B reviews a diff and unchanged §5 only requires its additionalContext to carry every cited story path; it does not make that call a header-bearing artifact or say whether Gate B derives from the spec header, the plan header, a WIP or closing-body field, or a comparison of them. Likewise, saying each artifact header is authoritative creates replicated authorities, while the disagreement rule does not define the artifact set the deriving agent must enumerate or the point at which it compares them | Gate A and Gate B can still derive different sets or Gate B can treat a required path carried outside a Story header as non-citing, so the promised single value, unanimity rule and lens union are not reliably determined and an omitted higher-risk member remains a gate-off path | Keep the human-selected Story-header authority, but map each cycle to the exact governing header set, require the Gate-B call to carry that whole set in the required form, and require the deriving agent to compare all then-existing governing headers before each pass and stop on disagreement +MINOR | high | Task 13 NEW lines 857-866, the sentence Two states break it | The two-state list is not complete over its own exactly-one-definition requirement. A partial merge can leave zero floor definitions without an explicitly derived dependency, or can duplicate two derived definitions, and neither state is a fixed obligation beside one predicate nor a derived dependency with no predicate. These are reachable through the command's user-selected merge path, especially an append or conflict resolution on a previously modified CLAUDE.md | A reader treating the explicitly numbered pair as exhaustive can miss an undefined or multiply defined floor and run a gate under an incoherent partial adoption | Make the examples explicitly non-exhaustive, or add zero and multiple definitions plus any normative pass or closure reference that resolves to no unique definition +MINOR | high | Task 13 NEW lines 859-862, every statement about pass counts or about closing | The coherence requirement is broader than the defect it is meant to detect: final §5 deliberately contains a separate hook reminder count and descriptive historical pass numbers, neither of which should derive from the cycle floor. The following two-state explanation narrows the concern to obligations and dependencies, so the rule contradicts its own literal scope | A literal downstream agent can stop a fully coherent adoption merely because the hook threshold or a historical measurement does not resolve to the derived floor | Limit the requirement to normative statements that set how many passes a gate owes or when a cycle may close, and explicitly leave reminder thresholds and descriptive measurements outside it +END OF FINDINGS (4 total) diff --git a/.context/codex-reviews/gate-a-plan-plana-pass-9-dispositions.md b/.context/codex-reviews/gate-a-plan-plana-pass-9-dispositions.md new file mode 100644 index 0000000..2b4be40 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-plana-pass-9-dispositions.md @@ -0,0 +1,93 @@ +# Gate A — Plan A cycle — pass 9 dispositions + +Advisory human note. Not a findings file; participates in no pass validation. + +## Result + +Artifact 3b95e10 (revision 11). VALID pass: terminator exact, 6 finding lines, 0 non-finding +lines. **0 BLOCKER · 5 MAJOR · 1 MINOR = 5 Blocker/Major.** + +**MANDATORY STOP-AND-SURFACE. Three tells present. Nothing repaired this pass.** + +## The three pass-report lines + +**1 — Trend.** + +| pass | findings | Blockers | Majors | B+M | +|---|---|---|---|---| +| 1 | 21 | 7 | 10 | 17 | +| 2 | 16 | 6 | 8 | 14 | +| 3 | 16 | 6 | 7 | 13 | +| 4 | 7 | 0 | 3 | 3 | +| 5 | 5 | 0 | 4 | 4 | +| 6 | **2** | 0 | 1 | **1** | +| 7 | 3 | 0 | 2 | 2 | +| 8 | 4 | 0 | 2 | 2 | +| 9 | **6** | 0 | 5 | **5** | + +Findings bottomed at 2 on pass 6 and have risen every pass since: **2 → 3 → 4 → 6**. B+M +**1 → 2 → 2 → 5**. + +**2 — Cluster.** Split, and the split is the point. Three of six (findings 2, 3, 6) are the +plan's **own descriptive prose contradicting its own shipped text** — a stale task preamble +still describing a rule that was withdrawn, an accounting row whose claims no longer match, +and an expansion from two states to four that left "either" in the sentences below it. The +other three are product behaviour. **Half the pass is now spent on the plan describing +itself inconsistently**, which is where passes 1-3 went before the strip, in a different +costume. + +**3 — require↔withdraw.** **PRESENT, and it is a chain rather than a pair.** The +demotion/tells interaction has now been: + +- **pass 4 M1** — *required*: the collision between demotion and the five-tells rule must be + resolved; state the precedence. +- **pass 5 M2** — *required narrower*: the precedence is too broad, limit it to the two-tell + stop. +- **pass 8 MAJOR 1** — *required removed*: the precedence is out of scope under spec §9 and + must go. It went. +- **pass 9 finding 3** — *requires more removed*: the remaining per-tell text is also §4 and + successor-story material, not §3, and should follow it out. + +A second, shorter chain in the same pass: **pass 8 MINOR 4** required the coherence rule be +narrowed to normative statements; it was; **pass 9 finding 5** says that narrowing +over-triggers and must be narrowed again. + +## Tells + +| # | tell | present? | +|---|---|---| +| 1 | finding count rising | **YES** — 2 → 3 → 4 → 6 across four passes | +| 2 | Blocker count failing to fall | no — zero for six passes; it cannot fall further | +| 3 | clustering on the instrument | not as such; the instrument was stripped | +| 4 | clustering on **prose about** either | **YES** — 3 of 6 | +| 5 | a require↔withdraw pair | **YES** — two chains, above | + +**Three present. Any two make stop-and-surface mandatory, not discretionary.** Reported and +handed over; no fixes applied, per the rule. + +## What I think is happening, offered as a reading and not a decision + +**One task is generating most of this: Task 12.** Spec §3 is the severity procedure and that +is what Plan A is licensed to ship. Everything I have added around it — how demotion affects +the finding total, the clusters, the require↔withdraw comparison and the Blocker curve, and +what that means for the two-tell stop — is **spec §4 (the per-pass curve) and successor-story +material**, and §9 excludes both. Each pass has removed one layer of that reach and found +another underneath. + +The withdrawal at revision 11 removed the *conclusion* I had drawn. Finding 3 says the +*premises* are equally out of scope. On the text, it is right: "the Blocker curve reads +severity after the ceiling" is a statement about the curve, and the curve is §4, which is +Plan B's. + +**The available move is to ship §3 and only §3** — the procedure, the exclusions, the +symmetric instrument carve-out, the rationale rule, coverage-first and the kinship sentence — +and keep every statement about what demotion does to the curve and the tells in **non-shipped +plan prose**, or hand it to the successor story. That is a reduction in what Plan A delivers, +which is why it is not mine to take. + +**Findings 1 and 4 are separate and would survive that move**: the governing-header mapping +is still not closed for the multi-plan case this very change executes (Gate B reviews a +combined A+B+C diff and is assigned "the plan's" header, singular), and no rule re-reads the +profile or the header *after* a pass and before accepting it as final, so a concurrent change +during the last pass is invisible. Both are real. Neither was repaired, because the stop +applies to the pass, not to a selection of it. diff --git a/.context/codex-reviews/gate-a-plan-plana-pass-9.md b/.context/codex-reviews/gate-a-plan-plana-pass-9.md new file mode 100644 index 0000000..ffeed90 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-plana-pass-9.md @@ -0,0 +1,7 @@ +MAJOR | high | Task 1 NEW lines 202-224 and plan topology lines 114-116 | The governing-header mapping is still not closed over this change or the general multi-plan case: Gate B is assigned singular "the plan's" header although it reviews the combined A+B+C diff, "every governing header that exists" names no exact artifact set, and an expected artifact with no `Story:` header drops out of the comparison instead of contributing the empty cited set; repeated `Story:` entries likewise have no normalization or ambiguity rule | Different Gate-A plan cycles and Gate B can derive different floors or omit a higher-risk story while the text still claims one source and one value, so the new gate-off route remains available in exactly the split-plan case this plan executes | Define the comparison set explicitly as the current spec plus every implementation plan contributing to the reviewed diff, make every expected artifact contribute a set even when its `Story:` header is absent, define identical repeats as deduplicated or malformed, require Gate B to carry every contributing header, and stop when any resulting set differs +MAJOR | high | Task 12 preamble line 750 | The retained instruction says the last two paragraphs "settle" the five-tells interaction and describes precedence against the two-tell stop, while the replacement at lines 794-798 says the interaction is not settled and that two tells still make stopping mandatory | An executor or reviewer can follow the stale task instruction and reintroduce the withdrawn out-of-scope completion rule even though the replacement and accounting say the opposite | Replace the preamble with the revision-11 decision: no precedence is chosen here, loop-rule ownership is deferred, and Task 12 changes severity semantics only +MAJOR | high | Task 12 NEW lines 787-798, old-conditions accounting row 11, and Global Constraints | The shipped severity insertion still specifies how the finding total, clusters, require↔withdraw comparison and Blocker curve consume demoted findings, then restates the exact two-tell conclusion; those are per-pass curve and loop-rule semantics from spec §4 and the successor story, not severity semantics from §3, and the restatement contradicts the plan's own rule that other closure rules are only referred to | Plan A still reaches past spec §9 after withdrawing the explicit precedence, duplicates an authoritative loop rule where it can drift, and makes the accounting's claims that the pass-report passage is untouched and Task 12 implements only §3 materially incomplete | Limit the shipped Task 12 text to §3's procedure, exclusions, symmetric instrument carve-out, rationale rule, coverage rule and kinship; keep the unresolved interaction only in non-shipped plan accounting or route it to the successor story, then correct row 11 +MAJOR | high | Task 1 NEW lines 207-224 and Task 11 NEW lines 693-708 | Header comparison and profile/set reads happen only before or "at" each pass, but a clean pass can be accepted and the cycle closed without a second read; a concurrent profile or `Story:` change during that pass is therefore invisible, and untouched §5 explicitly admits that nothing checks whether a header changed mid-call | A final pass can run under the old lower profile or cited set and close after the source has changed, contradicting the promises that every change binds an open cycle and costs a further pass under the current source | Re-read every governing cited set and profile after the call and immediately before accepting a final pass; if either differs from the pre-pass values, do not accept that pass as final and run the required further pass, or narrow the guarantee and disclose this concurrency residual explicitly +MAJOR | high | Task 13 NEW lines 861-866 | The coherence rule over-triggers by requiring every normative statement about when a cycle may close to resolve to the floor; untouched §5 deliberately has independent closure and stop predicates such as assigned-fix-set membership, new structural questions, accepted Blocker/Major findings and the two-tell stop, none of which derives from a pass-count definition | A fully and correctly adopted final §5 can satisfy the new floor predicate yet be declared incoherent and stopped because its other mandatory closure rules do not "resolve to" that predicate | Restrict the coherence requirement to normative statements that state or assume a numeric floor, a specific pass ordinal, or a dependency on the derived floor; say explicitly that independent closure predicates coexist with the floor and need not derive from it +MINOR | high | Task 13 NEW lines 867-874 | The breaking-state list was expanded from two states to four and made non-exhaustive, but the operational sentences still say "a merge can produce either" and "in either state" | The newly added no-definition and multiple-definition states are listed yet are not unambiguously included in the stop-and-human-repair branch, recreating the under-trigger the expansion was meant to remove | Replace both uses of "either" with "any such state" and, if examples remain, make clear they illustrate rather than limit the states routed to the stop +END OF FINDINGS (6 total) diff --git a/.context/codex-reviews/gate-a-plan-plana-passes-6-7-dispositions.md b/.context/codex-reviews/gate-a-plan-plana-passes-6-7-dispositions.md new file mode 100644 index 0000000..ac79141 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-plana-passes-6-7-dispositions.md @@ -0,0 +1,91 @@ +# Gate A — Plan A cycle — passes 6 and 7 dispositions + +Advisory human note. Not a findings file; participates in no pass validation. + +## The three pass-report lines + +**1 — Trend.** + +| pass | findings | Blockers | Majors | B+M | instrument share | +|---|---|---|---|---|---| +| 1 | 21 | 7 | 10 | 17 | most | +| 2 | 16 | 6 | 8 | 14 | 5 of 6 Blockers | +| 3 | 16 | 6 | 7 | 13 | 12 of 13 B+M (92%) | +| 4 | 7 | 0 | 3 | 3 | **0** | +| 5 | 5 | 0 | 4 | 4 | **0** | +| 6 | 2 | 0 | 1 | 1 | **0** | +| 7 | 3 | 0 | 2 | 2 | **0** | + +**2 — Cluster.** Product behaviour, four passes running. Zero instrument since the strip. + +**3 — require↔withdraw.** None across the cycle. The repeated findings on the cited-set axis +(pass 5 M1 → pass 6 MAJOR → pass 7 MAJOR 1) are **successive refinements of one unsettled +question**, not a pass demanding what an earlier pass removed. + +**Tells:** one marginal — findings rose 2 → 3. Blockers have been zero for four passes, the +cluster is on product behaviour, and there is no require↔withdraw pair. **One tell is not +two**, so no mandatory stop from that rule. The loop stops on pass 7 MAJOR 1's novelty +instead. + +## Pass 6 + +- **MAJOR — absorbed.** No authority for cited-set membership. Absorbed on the reasoning that + §2 already contracts one set, so supplying a stop for when reality disagrees implements the + assertion rather than changing it. Pass 7 shows that reasoning was **too generous** — see + below. +- **NIT — fixed while the row was being edited.** Accounting row 1 said condition b + ("Blocker/Major only") was kept outside the replaced range; it is inside Task 1's OLD block + and re-emitted in its NEW block. Disposition split. A Nit earns no repair round of its own + and did not get one. + +## Pass 7 + +- **MAJOR 2 — absorbed.** The downstream adoption trigger ran one direction only. It stopped + when the predicate landed beside a surviving fixed-number obligation, but not the inverse: + a merge can take Task 6's or Task 8's derived-floor language while Task 1 is absent, leaving + a project claiming a derived floor with nothing defining it. `/workflow-init` explicitly + permits a user-selected partial merge, so both states are reachable. Restated as a + **bidirectional coherence requirement**: exactly one floor definition present, and every + pass-count and closure statement resolving to it. +- **NIT — fixed, verified mechanically.** Task 1 said a raise costs a further pass "as above". + `grep` for `further pass` before Task 1 (line 160): **0 occurrences**; it first appears at + line 689, inside Task 11. The cross-reference pointed backward at nothing and now points + forward. + +## MAJOR 1 — routed, and my pass-6 absorption reconsidered + +This is the third finding on the cited-set axis. It asks for **one change-level carrier for +the governing cited set**, with each cycle reconciling against it, and says outright: *"If +choosing that carrier is not already a settled contract, route that choice to the human +rather than absorbing it."* + +**It is not settled.** §2 asserts the three cycles derive from "the same cited-story set" and +never says **where that set lives**. A carrier is therefore a new mechanism, not an +implementation of an approved sentence. §5's rule for a genuinely unclear boundary is to +treat the finding as **outside** — "which costs a question and never a silent expansion." + +**Its supporting example is wrong, and I checked before routing.** The finding says the spec +cites the successor loop-rule story while the plan cites only the pass-floor story. Verified: + +``` +spec :4 **Story:** …/2026-08-28-review-loop-economics-pass-floor-story.md +spec :32 …/2026-08-29-loop-rule-consolidation-story.md <- a scope DISCLAIMER +plan :19 **Story:** …/2026-08-28-review-loop-economics-pass-floor-story.md +``` + +Both governing `**Story:**` headers cite the **same** story. Line 32 states the consolidation +is *not* in the spec. **The artifacts do not disagree.** + +**The gap survives the bad example**, which is why it routes rather than being dismissed: +nothing in the shipped text distinguishes a **governing citation** from an **incidental +mention**, and an agent grepping story paths finds two in the spec and one in the plan. Two +fixes exist and they are not the same size — naming the `**Story:**` header as the governing +citation, or the carrier-plus-reconciliation the finding asks for. Choosing between them is +the routed question. + +**On my pass-6 absorption.** I absorbed the set-authority finding on the ground that a fix +implementing the approved sentence existed. That was defensible and it was also the more +generous of two readings, and the question came straight back one pass later in sharper +form. The rule I should have applied is the one §5 states for exactly this case: where +membership is genuinely unclear, treat it as outside. Two rounds on one axis is the signal +that it was unclear. diff --git a/.context/codex-reviews/gate-a-plan-planb-CLOSURE.md b/.context/codex-reviews/gate-a-plan-planb-CLOSURE.md new file mode 100644 index 0000000..fd2a021 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-planb-CLOSURE.md @@ -0,0 +1,98 @@ +# Gate A — Plan B cycle — CLOSED CLEAN at pass 7 + +Advisory human note. Not a findings file; participates in no pass validation. + +## Closure + +**Artifact:** `docs/superpowers/plans/2026-08-30-review-loop-economics-plan-b-records.md` +at commit `f132c57` (revision 8). + +**Pass 7: CLEAN.** Validated structurally: exactly two lines, line 1 exactly `NO FINDINGS`, +line 2 exactly `END OF FINDINGS (0 total)`. + +**Floor satisfied.** Story read fresh at close: risk `high`, security `none`. Under the rules in +force — the constant 3, since this cycle runs under the OLD rules by its own activation +constraint — the floor is 3. Seven passes run, final pass clean. **The cycle closes.** + +## Curve, and the comparison with Plan A + +| pass | findings | Blockers | Majors | B+M | +|---|---|---|---|---| +| 1 | 11 | 2 | 7 | 9 | +| 2 | 3 | 0 | 3 | 3 | +| 3 | 6 | 0 | 4 | 4 | +| 4 | 10 | 0 | 7 | 7 | +| 5 | 9 | 0 | 5 | 5 | +| 6 | 7 | 0 | 6 | 6 | +| **7** | **0** | **0** | **0** | **0** | + +**Seven passes against Plan A's twelve, for comparable content in a leaner artifact.** The whole +saving came at the start: Plan A opened at 17 B+M and spent three passes shedding a +self-description layer; Plan B was written without one and opened at 9. + +**Instrument and prose-about never reached a third of any pass** — 18%, 0, 17%, 20%, 22%, 14%. +The immediate-routeback condition never fired. Every pass was product churn on a genuinely +intricate mechanism, which is what the converged shape was supposed to leave room for. + +## Coverage statement + +**Reviewed across the seven passes:** both pinned grammars character-for-character against spec +§2.3 and §4 in every pass after the first; the nonce properties of §5 one at a time; the curve +and provenance properties of §4 and §2.3 one at a time; §6's accounting method applied to Plan +B's own passages; the §10 extension; **spec §9's exclusion list item by item**; the sequencing — +Plan A's fourteen edits then Plan B's six against the result, each OLD matching exactly once when +its task runs; all six assert-new patterns before and after in both copies; and the downstream +template for dependencies on this repo's layout. + +**Not covered, by construction:** Plan C, which does not exist. Its five inherited obligations +are named in Plan A and carried forward. + +**What this closure does not claim.** Gate A reviewed the plan, not the edits. Each task's single +check establishes that its edit landed at its site and nothing more. Correctness of what lands is +Gate B's, reading the real diff. *(Same boundary Plan A closed on, and the right sentence to carry +into every remaining closure.)* + +## Three results worth keeping + +**1 — A review finding that did not survive checking.** Pass 4 said the nonce property +contradicted the approved spec and story. Both already said what was shipped; what they also did, +and the plan had collapsed, was **separate the requirement from what a check can establish**. +That sharpened the absorb/route line: **a finding is absorbable when a fix implementing the +approved sentence exists** — and the spec's own §5-requirement/§8-checkability structure +demonstrated the shape of that fix. Plan A's pass-5 M1 routed because no such fix existed. Third +round on one axis is a signal to look for the route, not a rule that forces it. + +**2 — A rule that described a situation which cannot arise.** Pass 5 asked for an observable +procedure for the slot refusal; the pass-5 fix gave one that was unreachable — resolving your own +path never yields a different nonce, and scanning the family would flag legitimate siblings. +Pass 6 caught the same defect in new words. The fix was to bind the rule to **the deletion step** +§5 already requires, which is where the damage actually happens and where the check is a question +about a path the step itself computed. **The incident that motivated the rule — a nonce-holding +cycle computing a legacy bare path and deleting another cycle's file — is exactly the case the +reachable version covers, and the unreachable version did not.** + +**3 — A transcription that ate the thing it was transcribing.** Embedding the provenance grammar +through a string step interpreted `\"` as `"` and `\\` as `\`, silently rewriting **the +production that specifies which escapes are legal**. Both grammars are now inserted from the +spec's own block at build time and never re-keyed, and every later pass re-verified them +character-for-character. + +## Tooling built during the cycle + +The generator refuses to build unless **every assert pattern occurs exactly once in its task's +NEW text and zero times in its OLD**. It caught Task 1's pattern drifting three times as the slot +rule was rewritten — each time the plan's own check would have failed at execution. + +**Its first falsification test did not fire**, because the sabotage truncated a word and the +truncation was still a prefix of the real text: a test that could not fail. Re-tested with a +genuinely absent pattern and with one that also appears in the OLD text; it fires on both, naming +the pattern and both counts. **§8's wiring rule, practised on my own tooling.** + +## Residue + +**None open.** All findings from passes 1-6 are resolved, or routed and answered, or collected as +Minor/Nit per §5. Pass 7 found nothing. + +**Carried to the final-round story batch:** the criterion's worked example gains +"one per cycle; a change running N cycles holds N (e.g. five for this three-plan change)", so the +confirmed reading survives to the next reader rather than living only in gitignored notes. diff --git a/.context/codex-reviews/gate-a-plan-planb-pass-1.md b/.context/codex-reviews/gate-a-plan-planb-pass-1.md new file mode 100644 index 0000000..16a87ad --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-planb-pass-1.md @@ -0,0 +1,12 @@ +BLOCKER | high | Task 3 NEW, lines 251-267 | The provenance grammar is not copied character-for-character from spec §2.3: production alignment and side-condition line breaks were changed throughout, despite the plan's explicit verbatim claim | The shipped prompt creates a second textual form for a machine-parsed record, so the implementation cannot satisfy the pinned-grammar contract | Replace the whole grammar payload with the exact §2.3 block, preserving every character after only the uniform Markdown container indent is removed +BLOCKER | high | Task 4 NEW, lines 327-342 | The curve grammar materially diverges from spec §4: it changes the unknown-count gloss, replaces the pinned `[!-~]{1,}` bare-model production with prose, drops the `undetermined` meaning and real-model quoting side-condition, and removes the control-character and rejected-byte side-conditions from the grammar block | A parser written to the pinned grammar and a writer following the shipped grammar can accept and emit different languages, defeating the one-form requirement | Paste the exact §4 grammar block into Task 4 without re-alignment, rewording, or moving grammar side-conditions into surrounding prose +MAJOR | high | Task 2 NEW, lines 178-183, against the re-emitted optional-companion block | The rule says the advisory working record carries the nonce and follows the slot rules, but the only working-record spellings left in the resulting text are the exact bare names `gate-a-spec-resume.md`, `gate-a-plan-resume.md`, and `gate-b-resume.md`; no nonce-bearing working-record grammar is defined | Concurrent post-rule cycles can overwrite the same resume note or choose incompatible inferred spellings, so the principal Gate-A recovery source can be lost or undiscoverable | Define the exact nonce-bearing resume-note names in both copies, reserve the three bare names for the pre-rule case, and add that changed passage to the old-conditions accounting +MAJOR | medium | Plan topology lines 30-31 and Task 2 lines 164-169 plus Task 3 closing-line rule | The plan binds three Gate-A cycles plus one Gate-B cycle for the A/B/C sequence, while the shipped text says this section defines three cycles and that three cycles produce three lines and three nonce fields | An executor cannot tell whether every actual plan cycle gets its own provenance line and curve or whether one of the four plan-sequence cycles is meant to disappear from the record | Distinguish the three cycle types from cycle instances and state the exact record count owed by the A/B/C execution topology, including how the already-run spec cycle is treated +MAJOR | high | Task 2 NEW, lines 164-198 | The text promises three distinct post-rule nonces and says later readers need no contextual attribution, but uniqueness is checked only against cycles open when a nonce is generated; sequential cycles may legally reuse a closed cycle's nonce | Aggregated or squash-carried records can share a cycle field and become indistinguishable, so the stated attribution property outruns the rule that is supposed to establish it | Either require uniqueness against the durable scope in which records are later read, or weaken the distinctness and context-free-attribution claims to the settled open-cycle guarantee and state the remaining ambiguity +MAJOR | medium | Task 2 NEW, lines 195-198 | The unique-nonce start rule has a check-then-start race: two concurrent cycles can generate the same value, each observe no open collision before either publishes ownership, and both then start | The two cycles use identical findings and working-record slots, recreating the overwrite and mixed-provenance failure the nonce is meant to prevent | Bind nonce creation to an atomic ownership reservation before the cycle starts, treating a failed reservation as the named collision cause within the same bounded retry policy +MAJOR | high | Task 2 recovery rule, lines 185-193 | On no candidate, disagreeing sources, or multiple candidates the rule starts a new cycle but never closes, abandons, or retires the prior open candidates and their working records | The next recovery still sees the old candidates, now possibly plus the new one, so ambiguity can persist indefinitely and old writers can race the replacement cycle | Define a durable abandonment transition that makes every rejected candidate no longer open and retires its working record before the replacement cycle starts +MAJOR | high | Task 6 NEW, lines 469-472, against Task 2 lines 171 and 200-203 | `the nonce duties at their strictest` gives no executable outcome when the starting rules are unknown and conflicts with both nonce creation only at cycle start and the ban on a pre-rule cycle acquiring one | An agent can either mint late provenance and falsely attribute an existing cycle or use `none (pre-rule)` despite the strict fallback, with no rule choosing between them | State the terminal branch explicitly: recover one valid existing nonce if uniquely provable; otherwise abandon the indeterminate cycle and start a new post-rule cycle, never attaching a newly generated nonce to the old cycle +MAJOR | high | Tasks 2, 3, and 4 replacement operations | Each insertion re-emits its OLD text byte-for-byte, so that OLD still occurs once after the task; if execution is interrupted after the edit or amend and the task is resumed, the same replacement inserts the nonce or pinned grammar block a second time | Duplicate rule blocks create competing definitions and invalidate the one-form grammar promise even though the task's presence pattern still matches | Use the existing single assert-new check as the task state gate: count zero means apply, one means the edit already landed and must not be applied again, and any other count stops and surfaces +MINOR | high | Task 2 retry policy, lines 149 and 195-198 | The plan fixes `at most three attempts`, but the shipped rule says `retry at most three times`, which normally permits an initial attempt plus three retries | Executors can make either three or four generation attempts, so the supposedly fixed bound is not actually fixed; the three failure causes themselves are named correctly | Say `make at most three generation attempts total, then stop and surface` in both the plan instruction and shipped text +NIT | high | Old-conditions accounting row 1, line 75 | The accounting calls the old slot grammar exactly three names, but `gate-b--pass-

` expands to two concrete branch names, making three syntactic forms but four concrete slot families | The disposition's cardinality is factually wrong even though Task 1 preserves all of the actual forms | Change the row to say `three syntactic forms, yielding four concrete slot families` +END OF FINDINGS (11 total) diff --git a/.context/codex-reviews/gate-a-plan-planb-pass-2.md b/.context/codex-reviews/gate-a-plan-planb-pass-2.md new file mode 100644 index 0000000..55ee930 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-planb-pass-2.md @@ -0,0 +1,4 @@ +MAJOR | high | Global Constraints lines 54-58 and Tasks 2-4 | The assert-new preflight defines only the all-absent `0/0` and all-present `1/1` states; an interruption after editing one copy leaves `1/0` or `0/1`, and a prior duplicate leaves a count above one, with no prescribed branch | A resumed executor can apply the replacement to both copies, duplicating a nonce or grammar block in one while repairing the other, or proceed with divergent prompt copies | Using the same single assert-new check, define `0/0` as apply both, `1/1` as skip, and every mixed or above-one result as STOP and surface without replacing either copy +MAJOR | high | Task 2 NEW lines 177-180 and 214-224 | The text still claims that attribution means a later reader needs no cycle context and that a cycle cannot start without a `unique` nonce, while the residual immediately admits an unseen-open collision or check-then-start race can let two cycles start with the same value | A later reader can merge two cycles' records under one field even though the earlier categorical wording says that ambiguity is eliminated, repeating the gate-mechanism overclaim class in AGENTS.md | Qualify the benefit as probabilistic and replace `unique` with the exact property established: valid and not equal to any nonce observed on a known-open cycle at check time, subject to the two stated residuals +MAJOR | medium | Task 2 NEW lines 214-218 against `docs/prompt-standards.md` item 10 | The nonce-start failure names randomness unavailable, invalid output, and known-open collision and explicitly says they need different fixes, but the shipped prompt supplies neither the checks that distinguish them nor a cause-to-fix mapping | The surfaced diagnostic leaves the executor or human to invent remediation and the changed prompt does not pass invariant 11's cause-specific diagnostic requirement | Add a concise check and fix for each named cause while retaining the same three-attempt total and terminal STOP +END OF FINDINGS (3 total) diff --git a/.context/codex-reviews/gate-a-plan-planb-pass-3.md b/.context/codex-reviews/gate-a-plan-planb-pass-3.md new file mode 100644 index 0000000..1763f10 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-planb-pass-3.md @@ -0,0 +1,7 @@ +MAJOR | high | Global Constraints preflight table lines 54-64 | The `anything above 1` recovery is not executable for asymmetric states such as `2/0`: removing the one `extra copy` leaves `1/0`, whose row immediately STOPs, so the task cannot be retried as instructed | An executor resuming after a duplicate in only one copy remains permanently stuck or improvises a repair that can duplicate or corrupt one mirror | Keep STOP-and-surface, but require the hand repair to normalize the pair to an explicitly valid state—either remove every inserted block to `0/0` or restore exactly one complete block in each copy to `1/1`—before retry or skip +MAJOR | high | Plan header lines 11-15 and Task 2 lines 221-242 against spec §5 lines 338-370, spec §8 lines 529-539, and story lines 231-250 | The plan now intentionally guarantees only inequality with nonces observed on known-open cycles at check time, but the governing spec and story still require three distinct or valid unique nonces and difference from every open cycle at generation; no task updates them | Execution would knowingly ship behavior that fails the approved acceptance criterion and leaves the future evidence checkpoint evaluating an impossible stronger property | Add a task to update the spec, story, and future-discharge language to the weaker human-settled property, then rerun the required Gate-A review for the changed governing artifact before implementation +MAJOR | high | Goal lines 7-9; Task 1 NEW lines 122-130; Task 2 NEW lines 209-242; Task 4 NEW lines 410-415 | Despite admitting that two cycles can receive the same nonce, the plan still says slot naming keeps cycles from overwriting, a single recovered candidate keeps identity, starting anew prevents cross-reading, and the nonce keeps curve identity; same-nonce cycles use identical slot and working-record paths, and the `another nonce` refusal cannot distinguish them | A concurrent or unseen collision can overwrite findings or working records and misattribute a curve while the text assures the reader that identity is preserved | Qualify every ownership, recovery, and identity claim as probabilistic and explicitly state that a same-nonce collision can share slots and defeat attribution; retain the disclosure as a limitation, not a guard +MAJOR | high | Tasks 1-6 against the post-Plan-A downstream partial-adoption paragraph and workflow-init Rules 1-3 | Plan B adds a coupled nonce, slot, provenance, curve, and carry contract but no binding rule for a partial downstream merge; `/workflow-init` explicitly permits merge or skip, while the existing coherence STOP covers only the floor, so the curve grammar can land without `` or `` and nonce duties can land without infixed slots or squash carry | A downstream `CLAUDE.md` can contain an undefined or contradictory record contract, producing malformed or unattributable records or retaining the overwrite path this change is meant to close | Extend the downstream-adoption paragraph with a self-contained semantic coherence rule for Plan B's record set and STOP for a human to complete or revert a partial adoption +MINOR | high | Task 2 NEW lines 227-234 | The cause-specific diagnostics overclaim: a randomness source error or empty result can be transient despite `Retrying does not help`, and two collisions from a uniform source remain possible despite `means the source is not behaving randomly` | The terminal report can falsely blame the generator and send the human to the wrong remediation, so the mapping does not satisfy prompt-standards item 10 accurately | Separate transient source failure from persistent unavailability, and describe repeated collision as evidence or suspicion rather than proof while retaining the bounded STOP +MINOR | medium | Task 2 NEW lines 221-234 | The three-attempt terminal rule says to name which one cause occurred, but separate attempts can fail for different causes and it never says whether to report the last cause, every cause, or an aggregate | An executor can surface an incomplete or misleading diagnosis and omit a required cause-specific fix after a mixed failure sequence | Require the terminal report to name every distinct cause observed, by attempt, and pair each with its fix +END OF FINDINGS (6 total) diff --git a/.context/codex-reviews/gate-a-plan-planb-pass-4.md b/.context/codex-reviews/gate-a-plan-planb-pass-4.md new file mode 100644 index 0000000..4281ce6 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-planb-pass-4.md @@ -0,0 +1,11 @@ +MAJOR | high | Plan B header and Tasks 2-3 against spec §4-§5 and story criteria 2, 5, and 10 | The plan says there are three Gate-A cycles plus one Gate-B cycle, while its shipped text correctly says a three-plan change has five cycles: the Gate-A spec cycle, three distinct Gate-A plan cycles, and Gate B; the governing spec and story still describe three branch bodies and one plan cycle | Plan C can omit a provenance line or curve, or reconstruct an already-closed cycle while the spec says no reconstruction is needed, so the durable record set has no single accepted cardinality | Enumerate the exact five cycles and the commit/body disposition of each, then align the plan header, spec evidence section, story criteria, and Plan-C handoff before execution +MAJOR | high | Task 2 nonce start rule against spec §5 and story criterion 5 | The plan deliberately guarantees only inequality with nonces observed on known-open cycles at check time, but the governing spec still requires a valid unique nonce and three distinct nonces and the story requires uniqueness among all cycles open when generated | Execution would knowingly ship a weaker contract than the approved acceptance criteria and leave the future nonce checkpoint judging a property the shipped procedure cannot establish | Record the settled collision-resistant property in the spec and story and re-gate those governing artifacts, or restore a procedure that actually satisfies their stronger property +MAJOR | high | Goal; Task 2 recovery paragraph; Task 4 unknown-count paragraph | The later collision disclaimer does not repair the categorical claims that slot naming keeps cycles from overwriting, one recovered candidate keeps identity, a new cycle prevents cross-reading, and a nonce keeps curve identity; every claim is false when two cycles share the same field | A same-nonce collision can delete findings and merge provenance while nearby operational text assures an executor or later reader that identity was preserved | Qualify each mechanism sentence at its own site with the exact different-nonce or probabilistic bound instead of relying on one blanket caveat elsewhere +MAJOR | high | Task 2 opening attribution paragraph and “Two residuals” paragraph | The supposedly complete two-cause account omits a third reachable collision: generation intentionally ignores closed cycles, so a new cycle can redraw a visible closed cycle's nonce even without an unseen-open cycle or a simultaneous publication race | The new cycle reuses the closed cycle's findings and resume paths, deletes those surviving targets under the ordinary same-nonce retry rule, and makes durable records read as one cycle's while the residual list says the mechanism has only two gaps | Add collisions with closed cycles to the disclosed paths and make the enumeration explicitly non-exhaustive; keep the same probabilistic consequence statement +MAJOR | high | Task 2 recovery candidate and working-record paragraphs | Recovery requires a candidate to be keyed to this cycle's artifact, but the only working-record form supplied is a kind-plus-nonce filename; it contains no required artifact key, no Gate-A versus Gate-B binding rule, and no encoding for an artifact path that needs quoting | A resumed executor must guess that two records describe the same artifact, either adopting another cycle's identity or conservatively restarting every time the filename alone is insufficient | Define the exact artifact identity used for each cycle kind and require the working record to carry it in an unambiguous form, reusing a quoted-path production where a path is the key +MAJOR | medium | Task 5 partial-adoption coherence stop and Task 6 activation extension | The coherence set names nonce, slots, provenance, curve, and squash carry but not the unknown-start activation rule added by Task 6, so a downstream project can carry every named member, miss Task 6, and still appear complete | A cycle whose start cannot be established can then claim the pre-rule field or leave nonce duties undefined instead of taking the required strict post-rule path | Include the activation and unknown-start semantics in the coupled adoption set and require the adopted definitions to agree, not merely all five named blocks to be present +MAJOR | high | Task 6 NEW strict-fallback sentence against Task 2 optional-companions text and spec §5 | “owes a nonce and every record that carries one” makes every nonce-bearing record sound mandatory, including the explicitly optional working record and findings slots for a skipped cycle | The unknown-start fallback contradicts “Nothing depends on it existing” and can manufacture new closure duties or make a legitimately skipped cycle impossible to represent | State that the cycle owes the provenance line and curve or skip record, and that every cycle record it does write uses the nonce; expressly preserve the working record's optional status and the absence of findings slots for a skip +MINOR | high | Task 4 logical-pass paragraph | The text requires the commit body to say when hook call count differs from logical-pass count, but neither the pinned curve grammar nor an adjacent record form says what values to write or how to bind the disclosure to its cycle | Authors can satisfy the sentence with incompatible prose or omit one of the two counts, leaving the promised observability unusable across cycles | Define one small adjacent disclosure form naming the cycle field, logical-pass count, and hook-call count, and include it in the squash-carry contract +MINOR | medium | Task 4 quoted-model control-character branch | A control-bearing model identifier becomes `undetermined` and the body must instead record its source and rejected bytes, but no safe representation, binding to the affected pass, or squash-carry treatment is defined | The diagnostic can be lost, attached to the wrong pass, or accidentally reproduce the raw control value the rule forbids | Specify a safe escaped or hexadecimal adjacent diagnostic keyed to cycle and pass, and state that it travels with the curve on squash +MINOR | high | Old-conditions accounting rows 2-3 and Tasks 5-6 rationale | The plan says both changes append without rewriting and keep the old text verbatim, but Task 5 changes “entry and every” to “entry, every” and Task 6 removes “and”, replaces the semicolon, and changes the following clause's capitalization and placement | The accounting's sole guard against dropped conditions contains assertions disproved by its own OLD and NEW blocks, so a reviewer or successor can trust a preservation claim that was never true | Mark the requirements as kept semantically and extended, list their exact unchanged meanings, and reserve “verbatim” for text that is byte-for-byte retained +END OF FINDINGS (10 total) diff --git a/.context/codex-reviews/gate-a-plan-planb-pass-5.md b/.context/codex-reviews/gate-a-plan-planb-pass-5.md new file mode 100644 index 0000000..f3087da --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-planb-pass-5.md @@ -0,0 +1,10 @@ +MAJOR | high | Goal lines 7-9 and Task 4 lines 448-450 | The pass-4 same-nonce qualification is still absent at two categorical sites: the Goal says slot naming keeps two cycles from overwriting, and the curve rationale says a nonce keeps cycle identity, although Task 2 correctly admits that same-nonce cycles overwrite and collapse into one identity | An executor or later reader can rely on the unqualified claims at exactly the sites that explain the mechanisms, understating the accepted data-loss and misattribution path | Qualify the Goal with the different-nonce or probabilistic bound and change the curve sentence to say identity is retained only as far as the field distinguishes cycles +MAJOR | high | Task 4 lines 455-461 against spec §4 lines 309-313 | The spec says the plan fixes how the tracked reviewed commit is encoded, but the plan only names that concept and supplies no source, representation, capture point, or exact comparison for the two branches | A split or resumed Gate-B pass cannot reliably decide whether its branches reviewed one revision, so it can sum two artifacts or conservatively discard a valid pass depending on an executor's guess | State the exact full commit identifier each branch uses and the byte-for-byte equality comparison before branches are summed, as a procedure rather than a new body-record layer +MAJOR | high | Task 2 lines 217-239 | Recovery is keyed by cycle kind plus artifact, but the only artifact encoding added is “its path”; that works for Gate-A spec and plan files but Gate B reviews a diff or tracked commit and has no artifact path, and the plan never defines a key per kind | A resumed Gate-B executor cannot select recovery candidates by the stated predicate and may either adopt another cycle or restart every time | Define the stable artifact key for each kind, including a Gate-B key that survives WIP amends, and state its unambiguous representation in the working record +MAJOR | medium | Task 1 lines 135-145 | “A target owned by a DIFFERENT nonce” has no observable decision procedure: different nonces produce different target paths, while the findings-file grammar permits only findings plus the terminator and carries no owner metadata | The refusal rule is either vacuous for exact target paths or can be misread as scanning and refusing harmless sibling nonce paths, so the claimed protection cannot be applied consistently | Define the exact colliding target set and an existing observable ownership signal, or remove the refusal claim from the plan and spec and describe the protection as path disjointness only +MAJOR | medium | Task 5 lines 512-521 | The coupled-adoption rule now requires all adopted definitions to agree, but its only terminal action covers the missing-member state “some of them and not others”; it gives no action when every member is present but two versions disagree | A downstream project can satisfy presence, retain contradictory nonce, curve, carry, or activation semantics, and proceed to run a gate despite violating the new agreement requirement | Extend the stop sentence to cover either a missing member or any disagreement among present definitions, with the same complete-or-revert action +MINOR | high | Task 4 lines 426-433 | A control-bearing model identifier becomes `undetermined` and the body must describe its source and rejected bytes, but no safe byte representation or binding to the affected cycle and pass is defined | Authors can attach the diagnostic to the wrong pass or accidentally reproduce the forbidden raw control value while trying to identify it | In adjacent procedural prose, require the cycle and pass to be named and the rejected bytes to use a control-safe representation such as hexadecimal, never the raw value +MINOR | high | Task 2 lines 241-257 | The three nonce-generation causes are not mutually exclusive as written: an empty source result is “returns nothing” under randomness-unavailable and is also a drawn value that is not 8 to 16 characters under invalid-value | The required ordered cause report and per-cause fix can diagnose the same attempt two ways, violating prompt-standards item 10's distinguishable-cause rule | State evaluation precedence and make invalid-value apply only after a non-empty successful source result has been obtained and fails the alphabet or length check +NIT | high | Task 2 rationale line 181 and shipped text lines 202-203 and 265-271 | The revision added a third listed collision path and made the residual list non-exhaustive, but nearby text still says “Two residuals,” “two reasons,” “both,” and “Neither” | The risk disclosure is internally miscounted and makes the closed-cycle redraw path look outside the probability and no-guard conclusions that follow it | Replace the numeric and paired wording with “the listed residuals,” “these paths,” and “none is a guard” +NIT | medium | Task 2 lines 208-213 | The rationale says names, timestamps, and commits “each” collide exactly where sibling cycles do, but artifact names can differ, timestamps depend on resolution, and distinct plan commits can differ | The random-only rule remains clear, but its categorical why is false and can mislead later maintenance of the nonce design | Keep the random-source rule while changing the rationale to the narrower claim that deterministic candidates can be shared, unavailable at cycle start, or otherwise fail to provide the required independent draw +END OF FINDINGS (9 total) diff --git a/.context/codex-reviews/gate-a-plan-planb-pass-6.md b/.context/codex-reviews/gate-a-plan-planb-pass-6.md new file mode 100644 index 0000000..8ce26ca --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-planb-pass-6.md @@ -0,0 +1,8 @@ +MAJOR | high | Goal lines 7-9 and Task 4 lines 461-463 | Pass-5 Major 1 remains at the two claim sites it named: the Goal says slot naming keeps two cycles from overwriting each other, and the curve prose says a cycle keeps its identity through the nonce, with neither claim bounded to distinct nonces; Task 2's caveats elsewhere do not qualify these claims where they are made | The same plan acknowledges that same-nonce cycles delete and write the same paths and collapse into one identity, so an executor or downstream adopter can rely on protections the design explicitly does not provide | Qualify both sites in place with the distinct-nonce or probabilistic bound and describe the Goal as collision-resistant rather than as preventing overwrite categorically +MAJOR | high | Task 1 lines 138-147 | Pass-5 Major 4 remains Major: the revision says a different-nonce path is not one this cycle would write, then requires refusal when such a path is found while resolving a slot, but it defines neither a scan that could find it nor the exact operation and terminal state being refused; resolving the current nonce's exact path can never yield a different nonce, while scanning the slot family would encounter legitimate concurrent siblings | The rule is still either unreachable or an instruction to block the concurrency that nonce-infixed paths are meant to permit, so it is not an observable executable implementation of spec §5's refusal property | Define the exact candidate path set, observation, and stopped operation without treating harmless sibling paths as collisions, or remove the unreachable refusal claim and revise spec §5 to state that distinct-nonce paths coexist by construction +MAJOR | high | Task 2 lines 224-247 | The recovery artifact keys are stable but not selective: two Gate-A cycles can review the same document path and two Gate-B cycles commonly share the same base commit; if this cycle loses its own working record while exactly one such sibling remains observable, the kind-plus-artifact-plus-open predicate returns one candidate and adopts the sibling rather than taking the more-than-one ambiguity path | Recovery can merge two distinct cycles under one nonce, misattribute their curves and provenance, and make them delete or overwrite one another's cycle records even though their generated nonces differed | Do not treat a sole kind-plus-artifact match as identity; add a non-circular discriminator that separates concurrent cycles or declare shared-key recovery unresolvable and start a new cycle unless the candidate is positively linked to the current run +MAJOR | high | Task 2 lines 224-244 | The stated two-source recovery comparison is not executable from the records the plan defines: only the optional working record is told to carry the artifact key, no exact field or line identifies that key inside its free-form contents, and the provenance line and curve available in history carry no reviewed document path or Gate-B base commit; moreover an open Gate-A cycle's closing commit does not yet exist while closed cycles are expressly ineligible | An executor cannot enumerate or compare history candidates by the required kind, artifact, and open predicates, so disagreement between the two sources cannot be detected reliably and recovery will restart unnecessarily or adopt by guess | Define the exact recoverable record and fields for nonce, kind, artifact key, and open status in each source, using an existing call field where available, or remove history as a claimed source and narrow the recovery guarantees to what the working record can establish +MAJOR | high | Task 4 line 448 and Task 5 line 540 | The skipped-cycle line is literally `skipped (see skip reason)`, but the squash rule copies only the skip record and does not copy the separate skip reason that existing §5 records in the original commit body | A squash makes the original body unreachable from main while preserving only a dangling pointer, so history no longer says why the cycle was legitimately skipped and the promised skip evidence is lost | Copy each skip record together with its referenced skip reason, or make the pinned skip record self-contained so the one carried line contains the reason +MAJOR | medium | Task 4 lines 468-476 | Pass-5 Major 2 is only partially addressed: the representation is now a full 40-character object name, but the procedure still does not name the existing source field, when each branch's value is captured, or the exact equality check before summing; after an intervening WIP amend, current HEAD no longer identifies the first branch's reviewed commit | A split or resumed pass can still combine different revisions or discard a valid pair because the executor has to invent how to recover and compare the first branch's value | Name each call's existing tracked head commit field as the source, capture its full object name with that branch result, and require exact equality of the two captured names before the branches form one logical pass, without adding a commit-body self-description field +NIT | high | Old-conditions accounting row 3 line 105 and Task 6 lines 580-593 | The accounting says conditions a-h are all kept verbatim and calls Task 6 an append, but the replacement does not re-emit its OLD verbatim: it removes the conjunction before severity and changes the retained `each further rule` clause's capitalization while extending the same sentence | No old condition is semantically lost, but the literal-preservation claim is false in the artifact whose purpose is to prove that rewrites did not silently drop conditions | Change the disposition to `kept in substance` and state that Task 6 extends the decision list, or structure NEW so the OLD bytes remain verbatim and the additional duties are appended separately +END OF FINDINGS (7 total) diff --git a/.context/codex-reviews/gate-a-plan-planb-pass-7.md b/.context/codex-reviews/gate-a-plan-planb-pass-7.md new file mode 100644 index 0000000..8d781f4 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-planb-pass-7.md @@ -0,0 +1,2 @@ +NO FINDINGS +END OF FINDINGS (0 total) diff --git a/.context/codex-reviews/gate-a-plan-planb-passes-1-4-dispositions.md b/.context/codex-reviews/gate-a-plan-planb-passes-1-4-dispositions.md new file mode 100644 index 0000000..d69e77c --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-planb-passes-1-4-dispositions.md @@ -0,0 +1,76 @@ +# Gate A — Plan B cycle — passes 1-4 dispositions + +Advisory human note. Not a findings file; participates in no pass validation. + +## Curve + +| pass | findings | Blockers | Majors | B+M | instrument + prose-about | +|---|---|---|---|---|---| +| 1 | 11 | 2 | 7 | 9 | 2 of 11 (18%) | +| 2 | 3 | 0 | 3 | 3 | 0 | +| 3 | 6 | 0 | 4 | 4 | 1 of 6 (17%) | +| 4 | 10 | 0 | 7 | 7 | 2 of 10 (20%) | + +Written lean from the start, so there was no declaration layer to shed: pass 1 opened at 9 B+M +against Plan A's 17. **Instrument and prose-about stayed below a third in every pass**, so the +immediate-routeback condition never fired. This is product churn, not Plan A's disease. + +## The cycle-cardinality reading — CONFIRMED, and recorded where it survives + +The story's criterion reads *"one per cycle, so a run of all three holds three."* This change +runs **five** cycles: one Gate-A spec, three Gate-A plan, one Gate B. + +**Confirmed reading: "one per cycle" is the rule; the clause after "so" is a worked example** +written before the three-plan split existed, and a stale example does not outrank the rule it +illustrates. The five-cycle shape is a consequence of two decisions already taken — the plan +split, and the single Gate-B topology — not a new claim by this plan. + +**Recorded in three places on purpose, because two of them do not survive.** This file and the +closure record are under `.context/`, which is gitignored and per-clone: a reading recorded only +here reaches nobody. So it also went into **the plan's binding constraints**, which are +committed, and into the commit body. The example clause itself gets a one-line touch-up in the +final-round story-edit batch, so the next reader meets the corrected example rather than +re-deriving this. **That gitignored records do not reach the next reader is this cycle's own +availability lesson, applied to itself.** + +## Absorb-versus-route, and the line this cycle sharpened + +Pass 4 said my weakened nonce property contradicted the approved spec and story. **Checking it +changed the fix, and the finding was half wrong** — the first review finding this cycle that did +not survive verification. + +``` +spec §8: "it must differ from that of any cycle open at the time it was generated + — which is the uniqueness the rule actually requires, and all it can check." +story: "unique among cycles open when it was generated" +``` + +The governing artifacts already say what I shipped. What they also do, and I had collapsed, is +**separate the requirement from what a check can establish** — §5 states the rule, §8 explains +the limit. Restoring that separation implements the contract rather than amending it. + +**That is the line between absorb and route**, stated more precisely than before: **a finding is +absorbable when a fix implementing the approved sentence exists** — and here the spec's own +requirement-versus-check structure demonstrated the shape of that fix. Plan A's pass-5 M1 routed +because no such fix existed: every option there changed what the spec promised. Third round on +one axis is a signal to look for the route, not a rule that forces it. + +## What my own tooling caught that the review did not + +Editing the slot sentence broke Task 1's assert pattern, which still grepped the old wording — +**the plan's own check would have failed at execution.** The generator now refuses to build +unless every assert pattern occurs exactly once in its task's NEW text and zero times in its OLD. + +**The first test of that guard did not fire**, because I sabotaged a pattern by truncating a word +and the truncation was still a prefix of the real text — a falsification test that could not +falsify. Re-tested with a genuinely absent pattern and with one that also appears in the OLD +text; the guard fires on both, with the pattern and both counts named. This is §8's wiring rule — +*name the observation that would exist if the claim were false, and confirm the wiring could have +produced it* — practised on my own tooling. + +## Collected, not iterated + +Three Minors from pass 4: the curve grammar carries no field for the logical-pass versus +hook-call discrepancy its own prose requires be stated; the rejected-bytes record for a +control-bearing model identifier has no defined form; and the accounting calls Tasks 5 and 6 pure +appends when each also rewords a clause. diff --git a/.context/codex-reviews/gate-a-plan-planc-CLOSURE.md b/.context/codex-reviews/gate-a-plan-planc-CLOSURE.md new file mode 100644 index 0000000..4531cee --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-planc-CLOSURE.md @@ -0,0 +1,86 @@ +# Gate A — Plan C cycle — CLOSED AS NOT CONVERGED at pass 7 + +Advisory human note. Not a findings file; participates in no pass validation. + +**Artifact:** `docs/superpowers/plans/2026-08-30-review-loop-economics-plan-c-rollout.md` +at `bcd45dc` (revision 7). **No clean pass was reached, and none is claimed anywhere.** + +**Decision (Daniel, 2026-08-30):** close the cycle as not-converged and split Plan C by +statement site into C1, C2 and C3, each its own Gate-A cycle. The split is the fallback he +named at the pass-6 tripwire and chose against then; it is chosen now on the data below. + +## The curve + +``` +pass 1 2 3 4 5 6 7 +Findings 18 20 20 23 22 19 29 +Blockers 5 4 2 4 5 2 2 +Majors 11 13 13 16 13 16 19 +B+M 16 17 15 20 18 18 21 +``` + +Seven passes, seven revisions. **Blocker/Major never left the 15–21 band**, and the highest +total and highest B+M of the cycle are both pass 7 — the last one. For contrast, from the same +change: Plan A cleared its Blockers by pass 4 and closed clean at 12; Plan B carried one Blocker +in its whole cycle and closed clean at 7. + +## Two mandatory stops and one tripwire + +| Pass | What fired | What it produced | +|---|---|---| +| 4 | two-tell stop (4 of 5 tells) | revision 5 deleted the plan's copy of `CLAUDE.md` §5 | +| 5 | two-tell stop + both standing route-immediately conditions | revision 6 moved the pre-rule records out of the commit body | +| 6 | Daniel's numeric tripwire (≤8 B+M, instrument share collapsing) | revision 7's claim sweep | +| 7 | bounded contract expired, not clean | this closure and the split | + +## The claim sweep, and its honest outcome + +Revision 7's method was Daniel's: enumerate the contested **claims**, grep every statement site +of each, fix every occurrence in one pass — `AGENTS.md`'s own recipe (*search for the claim, not +the phrase*) applied as a method rather than as a warning. + +**Three of six held. Two regenerated inside the sweep itself.** + +| Claim | Outcome at pass 7 | +|---|---| +| 1 — what the hook stores and counts | **fixed**, not re-contested | +| 5 — the WIP preflight's desired-state exit code | **fixed** | +| 4 — what a `grep -E` match can decide | **improved, incomplete** — per-pass model-key duplication and ordering uncovered; cross-record nonce agreement has only a positive case | +| 6 — completing the field report | **new task, two new findings** — updates the curve but not the analysis prose a clean close falsifies; its assert checks removal without checking insertion | +| 2 — whether a parser exists | **still open** — the sweep found the spec site it had missed and **missed two more**: both shipped prompt copies still carry Plan B's *"because the deferred metrics work parses it"* | +| 3 — which slot forms are valid | **worse; both remaining Blockers** — shipping the discriminator production is out of Plan C's §7/§8 scope **and unreachable**: this cycle runs under the old rules by the plan's own activation constraint, so a production shipping in the same commit cannot bind it | + +**The finding that matters more than any individual one:** a sweep whose entire premise was +*find every site* missed two sites of the very claim it was sweeping, and the fix written for a +contradiction was itself contradictory. **That is a width result, not a care result** — Plan C +carried a rollout across seven files, a spec, a CHANGELOG and a live Gate-B cycle in one +artifact, and at that width a single agent's sweep does not converge. + +**Cross-reference for the field record:** this is the strongest evidence so far for the +machine-readable architecture projection in `docs/superpowers/specs/2026-08-30-dark-factory-vision.md` +§10 — *generated from* `AGENTS.md`, never a second source. What defeated the sweep is exactly +what a projection would make mechanical: knowing every site a claim occupies. + +## The six standing structural findings, unrepaired + +Carried into C2 and C3 rather than fixed here: + +1. No task instantiates the preflight; the only one is a placeholder template. +2. No task constructs or stores the evidence entry every Gate-B call must quote verbatim. +3. The knob before-state has no storage location or resumption rule (inherited M9). +4. The grammar check has no actual EREs, fixtures or fail-capable assertions. +5. The `mktemp` closing-body obligation (inherited MINOR 12) is declared moot while a multiline + body still has to be composed with no named mechanism. +6. The risk-path verification reads one `Story:` header where the shipped rule requires the + union across the spec and all three plans. + +## What replaces Plan C + +`C1` user-facing docs · `C2` packaging and descriptions · `C3` the close. Sequential, each its +own Gate-A cycle at floor 3, slot infixes `planc1` / `planc2` / `planc3`. One combined Gate-B +cycle at the end, unchanged. + +**Bootstrap fix carried into C3:** the discriminator production ships **for future cycles**; +this cycle's slot naming rests on the `fic2`/`pr15` field-practice precedent. The plan claims no +production binds the cycle that ships it — which dissolves both standing Blockers by removing +the unreachable claim rather than by arguing with it. diff --git a/.context/codex-reviews/gate-a-plan-planc-pass-1.md b/.context/codex-reviews/gate-a-plan-planc-pass-1.md new file mode 100644 index 0000000..de1f207 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-planc-pass-1.md @@ -0,0 +1,19 @@ +BLOCKER | high | Task 13 Step 2 and Task 14 Step 1, lines 669 and 716 | Both commands read `.context/plan-a-base-sha`, but Plan A Task 1 only prints `git rev-parse HEAD~1`; no task writes that file, and it is absent in the current workspace | The differential check and the Gate-B cycle-intact check fail before they can establish a base, so the combined Gate-B cycle cannot run as written | Persist the base SHA when Plan A opens the WIP, or add a Plan-C precondition that recovers the exact WIP parent, validates it as the recorded cycle base, and writes the file before either consumer runs +BLOCKER | high | Task 11 assert-new command, line 544 | Shell parsing removes the intended JSON quotes, so the command searches for `version: 0.11.0`; that substring is absent even after the manifest contains `"version": "0.11.0"` | The task's one permitted check remains 0 after a correct edit, preventing the manifest-bump task from completing and invalidating the claimed A-to-B-to-C execution verification | Quote the fixed string once with shell-safe outer quotes, for example a single-quoted `"version": "0.11.0"` pattern +MAJOR | high | Task 6 NEW, lines 318-320 | The hook does not emit one value N in all three positions: it reports total cycle calls over the hook threshold, then a separate fresh-fingerprint count, so a satisfied state can be `5/3 cycle, 1 on current fingerprint` | The replacement remains a false user-facing description of the hook and can make readers treat the threshold, total calls, and fresh calls as equal | Use distinct placeholders such as passes, threshold, and fresh, or keep the literal 3/3/3 as a clearly bounded example rather than a general form +MAJOR | high | Old-conditions accounting and Task 9; `docs/coding-workflow.md:261-281` | The falsified-sentence set misses the methodology's instruction to record the pass model in an evidence entry or dispositions file; Plan B instead pins the model inside each cycle's per-pass curve | Users following the surviving sentence can omit the model from the mandatory curve and write it into records that Plan B explicitly does not key for this purpose | Add a Plan-C replacement at this site that points to the pinned per-pass curve and preserves the health-probe procedure and unchecked-bookkeeping caveat +BLOCKER | high | Old-conditions accounting and Tasks 1-12; `plugins/dev-workflow/commands/process-pr-review.md:159-162` | The shipped PR-processing prompt still says a skipped cycle records only the skip reason, battery result, and profile evidence, while Plan B newly requires every skipped cycle to carry a provenance line and a cycle-field skip record in place of its curve | The two prompt surfaces give incompatible closure instructions, so a normal trivial-fix skip can close without mandatory records; this violates invariant 11's no-contradictions requirement | Add this prompt to Plan C's falsified-site accounting and extend its skip instructions with the provenance line and pinned skip record, preserving the existing reason, battery, and evidence duties +MINOR | high | Task 1 NEW, line 118 | The README says the obliged floor is derived from the cited story's profile, but the mechanism also assigns floor 3 when no story is cited or a cited story is unprofiled, where no such profile exists | The knob row still teaches a narrower mechanism than the one shipped and is false for two ordinary states | Describe the floor as derived under §5 from the governing cited-story set, explicitly including its no-story and unprofiled defaults +MINOR | high | Task 12 CHANGELOG entry, lines 596-599 | The entry says the two forms are pinned because a program parses them, but this repository has no parser yet; the design routes parsing to deferred P8 work and separately requires a constructed-string parse check for this branch | The changelog presents a future consumer as an existing mechanism, the same enforcement-overclaim class the repository warns against | Say the forms are pinned for the deferred program to parse, and distinguish that future consumer from this change's one-time parse verification +MAJOR | high | Task 13 Step 2, lines 668-679 | Both grep commands search the common tail `carrying a Minor keeps` and return success whether or not Plan A changed the governing prefix; if the change is absent, both commands still exit zero | The purported check that fails without the change can pass mechanically on an unchanged tree, so it does not satisfy the story's `battery+check+verification` counterfactual | Make the comparison observe the deciding text: require the base to contain `pass 1 carrying`, require HEAD to contain `pass below the floor`, and make either reversed or missing observation fail +MAJOR | high | Task 13 Step 3, lines 681-686 | The spec's named risk-path verification recomputes every provenance line from its cited headers, but this step checks pass reports instead and runs before the Gate-B reports and closing provenance lines exist | The evidence pack neither verifies the durable provenance mechanism the story names nor has a reliable artifact set on which its stated pass-report check can execute | Move the risk-path verification to the completed closing-body draft, recompute each of the five provenance floors from the authoritative headers, and separately check any available pass-report duty without substituting it for provenance verification +MAJOR | high | Task 13 evidence pack | Spec §8 requires a parse check over constructed valid and invalid strings for both pinned grammars, including quoted paths, unusable-knob causes, gapped ranges, split models, skips, and count cardinality; Plan C contains no such step | The new machine-readable contracts can ship with an unexercised grammar, and the branch cannot claim the validation mode's named verification | Add the specified constructed-string parse check and record which grammar feature each accepted and rejected case exercises +MAJOR | high | Task 13 evidence pack | Spec §8 requires parity verification for every changed rule across `CLAUDE.md` and the scaffolded copy, but Plan C performs only per-task presence checks and the battery's narrow invariant checks | A one-sided or semantically divergent Plan-A or Plan-B edit can survive while every Plan-C check passes, violating the story's parity criterion | Add the required rule-by-rule parity pass over the resulting two prompt copies, retaining the one deliberate successor-pointer divergence +BLOCKER | high | Task 13 evidence pack and invariant 11 | Spec §8 and the story require a fresh twelve-item prompt-standards pass over both the outer `workflow-init` command and the resulting scaffolded `CLAUDE.md`; Plan C only inserts the item-1 n/a note and runs the narrow mechanical checker | Eleven template items and all twelve outer-command items can remain unreviewed while the plan declares invariant 11 discharged | Add an explicit complete checklist pass for both prompt artifacts, recording only the approved item-1 n/a for the scaffolded template and concrete results for every other item +MAJOR | high | Task 13 Step 4, lines 693-706 | The knob check treats every `-e` path as a readable file and nests unchecked `wc` and `shasum` failures inside `printf`; directories, unreadable files, broken symlinks, or a missing `shasum` can yield a normal-looking present record, and it never classifies the numeric or unusable value needed by the provenance grammar | A changed or unusable knob can be misreported as preserved, and the closing provenance lines lack a sound source for their threshold clause | Distinguish absent, readable regular, and unusable states with checked commands; capture byte count and digest for preservation and independently classify the value as numeric or one of the pinned unusable causes +MAJOR | high | Task 13 Step 1, lines 652-657 | The version-bump checker is run against local `main`, while “confirm main is current” is only prose with no stop condition; against a stale main that predates an earlier bump, an unchanged 0.10.0 manifest can differ from the stale base and pass | The battery can produce a false green for invariant 12 precisely on the omitted-bump path it is meant to catch | Use the recorded pre-WIP base SHA for this local comparison, or make freshness against the actual PR base an executable prerequisite that stops on mismatch +BLOCKER | high | Task 14 Step 2, lines 733-734 versus Task 15 lines 782-785 | The Gate-B cycle is explicitly pre-rule and has no nonce, yet Task 14 requires nonce-inflected findings paths; Plan B reserves those paths for nonce-holding cycles and says a pre-rule cycle cannot mint one | The executor must either fabricate late provenance, violate the slot grammar, or be unable to name the review files; this also contradicts the later claim that the cycle identifier is undemonstrable on this branch | Use the old bare Gate-B branch slots for this pre-rule cycle and reserve nonce-inflected slots for cycles that actually start after the rules ship +MAJOR | high | Task 15 Step 2, lines 775-785 | The plan places all five provenance lines and curves in the Gate-B closing commit, but the pinned contract and story require each cycle's records in that cycle's own closing commit body; the four Gate-A cycles have already closed, and their actual commits contain no pinned records | The branch does not demonstrate native per-cycle recording and necessarily reconstructs four curves later despite spec §8 saying no reconstruction is needed or claimed | Reconcile the topology with the contract before implementation: either amend the four owning commits with validated native records or obtain an explicit spec and story decision permitting one later aggregate body and documenting reconstruction +MAJOR | high | Task 15 Step 2, lines 770-790 | The only check on the dynamic closing body is absence of angle brackets followed by visual output; it does not validate that there are exactly five provenance lines and five curves, that each parses, that counts match pass ranges, that models and cycle fields agree, or that the two undemonstrable clauses are recorded | A malformed, missing, duplicated, or internally inconsistent evidence pack can close while the plan claims every demonstrable field was verified | Validate the completed body before amend against both pinned grammars and explicit cardinality and cross-record checks, then record the two genuinely undemonstrable fields as such +MAJOR | high | Task 15 Step 3, lines 795-799 | The temporary body is deleted even if `git commit --amend -F` fails, and the final subject check only prints the old subject rather than failing when it still begins with WIP | A transient commit failure destroys the only assembled evidence body and leaves the review cycle open while the step continues | Delete the temporary file only after a successful amend and an explicit non-WIP assertion; on failure preserve the path, report it, and stop +END OF FINDINGS (18 total) diff --git a/.context/codex-reviews/gate-a-plan-planc-pass-2.md b/.context/codex-reviews/gate-a-plan-planc-pass-2.md new file mode 100644 index 0000000..c897104 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-planc-pass-2.md @@ -0,0 +1,21 @@ +BLOCKER | high | Task 15, lines 730-751 | Base recovery is circular: it defines the base as the current tip's parent and then proves that exactly one commit is above that parent; the line-738 rev-list guard is a no-op, and a stack of two identically titled WIP commits therefore recovers the newer WIP as the base and passes with count 1. It also overwrites any previously persisted base before comparison | Gate B can exclude the first WIP and review only the last commit, while destroying the one local datum that could have exposed the wrong base | If a persisted base exists, validate it before any write; otherwise recover from an independent anchor, reject a parent whose subject is the WIP, validate the complete intended path and commit range, and write only after all checks pass +BLOCKER | high | Task 14 versus `CLAUDE.md:424-433` and `plugins/dev-workflow/commands/workflow-init.md:603-612` | Task 14 repairs `process-pr-review.md`, but both governing §5 copies still say a skipped profiled cycle owes only battery, reason, and evidence entry and that an unprofiled one owes nothing more; Plan B requires every skipped cycle's provenance line and skip record | The shipped prompts remain contradictory, so an agent following either governing copy can close a skipped cycle without the two mandatory records, violating invariant 11 | Add both sites to the falsified-sentence accounting and make the same duty-preserving replacement in root `CLAUDE.md` and the inline scaffolded copy +BLOCKER | high | Tasks 16 Step 7 and 18 Step 2, lines 832-854 and 919-930 | Four Gate-A cycles already closed before the knob snapshot procedure existed, and the plan has no contemporaneous evidence of whether `.context/codex-gate.floor` existed, what bytes it held, or how it classified during those cycles; reading it immediately before Gate B cannot establish their historical knob clauses, and the grammar has no unknown form | Whichever record-placement answer the human chooses, four provenance lines cannot be populated truthfully without guessing historical state | Extend the blocking contract decision to specify an honest historical-knob representation or supply authoritative contemporaneous evidence; do not infer the four values from the Gate-B-time snapshot +BLOCKER | high | Task 16 Step 6, lines 825-830, after Task 14 changes `process-pr-review.md` | The fresh prompt-standards pass still covers only the scaffolded template and outer `workflow-init.md`; revision 2 newly changes `process-pr-review.md`, which is itself a command prompt covered by invariant 11 | Plan C can close without checking all twelve requirements against one of the two changed command files the user explicitly requires reviewed | Add a complete twelve-item status matrix for `process-pr-review.md`, in addition to the two existing artifacts, and stop on any FAIL +MAJOR | high | Task 16 Step 1, lines 762-778 | The only literal version-check command named is `sh scripts/check-version-bump.sh`, which exits 2 because the checker requires exactly one base-ref argument; the prose says to pass the recovered base but never gives an executable full battery or a check that the invoked argument is that base | The required battery cannot be run verbatim and an executor can omit or substitute the critical invariant-12 comparison while still claiming the prose was followed | Spell out the complete command ending with `sh scripts/check-version-bump.sh "$base"` and assert the argument equals the validated persisted base before invocation +MAJOR | high | Task 16 Step 2, lines 780-801 | The three discriminating greps print expected counts but never compare them with 1, 0, and 1; an unexpected first or second result is masked when the final grep succeeds | The counterfactual check can exit zero even when the base lacks the old rule or HEAD still contains it, so it still does not mechanically fail without the intended change | Capture or directly test every count with numeric assertions in a fail-fast compound command, then record the two revision answers +MAJOR | high | Task 16 Step 3, lines 803-809 | The risk-path step requires comparing five recomputed floors with five provenance lines before Task 18 has drafted those lines; four owning commits contain none and the Gate-B cycle is not yet closed | The named verification has no comparison artifacts when scheduled, so the evidence pack cannot be completed in order | Draft the records before this step, run the verification over that draft, and repeat it after the human-selected placement and immediately before closure +MAJOR | high | Task 16 Step 4, lines 811-817 | The parse check supplies no parser or executable accept/reject procedure, specifies only a curve-cardinality rejection, and says every feature is exercised for each grammar even though gapped specs, split models, skips, and count cardinality are curve-only while knob causes are provenance-only | An executor cannot demonstrate both pinned grammars or the required invalid cases consistently, and a malformed provenance grammar can ship untested | Provide an actual parser invocation and fixture table partitioned by grammar, including at least one rejected provenance instance and one rejected curve instance plus every applicable valid production +MAJOR | high | Task 16 Step 7, lines 834-854 | The state machine tests `! -e` before lexical path type, so a broken symlink is classified as absent; the promised separate numeric, empty, non-numeric, and out-of-range classifier is prose only and has no command or recorded result | An existing unusable knob can be recorded as absent or left unclassified, producing a false provenance clause | Detect a symlink before the dereferencing existence test and implement a checked classifier that emits exactly the pinned absent, numeric, or unusable-cause token +MAJOR | medium | Task 16 Step 7, lines 843-853 | Byte count, digest, and later value classification read the live path separately; a concurrent writer or symlink retarget can make those observations describe different file states | Preservation and provenance evidence can look internally valid while mixing bytes and a classification that never coexisted | Copy the knob once into a securely created snapshot after validating the path, derive byte count, digest, and class from that snapshot, and repeat the same atomic procedure at close +MAJOR | high | Task 17 Step 2, lines 862-882 | Single-branch recovery preserves the successful branch file but does not capture and compare the tracked reviewed commit for both branches, despite the pinned rule that two branches form one logical pass only when they reviewed the same artifact revision | A failed branch resumed after an amend or external change can be summed with a successful branch from another revision and recorded as one valid pass | Record the reviewed commit with each branch result and require exact equality; on mismatch mark the earlier branch incomplete and start a new logical pass +MAJOR | high | Tasks 17 Step 3 and 18 Step 1, lines 884-901 | The plan says evidence is revalidated after a fix and before close but never asserts that current HEAD is the exact commit reviewed by both branches of the final clean Gate-B pass, nor explicitly stops and re-reviews when it differs | A fix or amend landing after evidence production can reach the closing commit without Gate-B coverage even when the revalidated evidence itself is unchanged | Persist the final clean pass's reviewed commit, compare it with HEAD immediately before close, and require a full re-review against the same base on any mismatch +MAJOR | high | Global Constraints and Tasks 1-15, lines 64-68 and 728-754 | Explicit per-task `git add` commands do not detect unrelated content that was already staged when Plan A opened the WIP or already committed into the recovered range; `git status --porcelain` later cannot reveal unrelated content already inside HEAD | The combined closing commit can absorb out-of-scope user changes and present them as reviewed under this story | At base validation, compare the full `base..HEAD` path set and current index/worktree state with the exact declared A+B+C allowlist and stop on every extra path +MAJOR | high | Task 18 Step 2, lines 903-917 | The shown validation commands merely print counts: they do not assert zero placeholders or one record per cycle, and the promised parser, cardinality, and cross-record checks have no executable procedure; a body containing a placeholder still leaves all three greps successful | A malformed, duplicated, incomplete, or internally inconsistent body can be used for the closing amend while the plan claims it was validated | Implement a fail-fast validator that asserts exact counts, parses every record, checks curve cardinality and cycle-field equality, and only then permits the amend +MAJOR | high | Task 18 Steps 2-3, lines 903-940 | Without re-reporting the open ownership choice, the executable path is already hard-coded to aggregate every cycle's records into one temp body and amend only the tip; it has no branch for amending four owner commits, and rewriting those ancestors after Gate B would change the reviewed history and recovered base | One permitted human answer is not executable, and naïvely implementing it would invalidate the completed Gate-B review before closure | After the human decision, encode separate procedures: aggregate only when authorized; otherwise rewrite owner commits before Gate B, recover the resulting base and tip again, and rerun Gate B over the final history +MAJOR | high | Task 17 Step 2, lines 871-880, and existing `.context/codex-reviews/gate-b-spec-pass-1.md` / `gate-b-quality-pass-3.md` | Bare names are correct for this settled pre-rule cycle, but the required deletion collides with pre-existing bare findings slots from earlier cycles—the exact data-loss class Plan B's nonce names were introduced to prevent | Executing the plan destroys prior review evidence before the new calls begin | Preflight every bare target; if any exists, preserve it at a clearly archival non-slot path or stop for human disposition before deleting only the new cycle's known targets +MAJOR | medium | Tasks 10 and 12, lines 463-506 and 558-620 | Both insertion tasks re-emit their OLD anchor inside NEW, so after an interrupted or resumed run OLD and NEW coexist and the replacement remains applicable; their grep checks accept duplicate NEW text | Retrying either task can duplicate the prompt-standards note or the 0.11.0 changelog entry, making the rollout non-idempotent | Add a one-check preflight state machine that applies only in the OLD-only state, skips the exact OLD-plus-one-NEW state, and stops on duplicates or every mixed/unknown state +MINOR | high | Task 1 NEW, line 118 | The README says the obliged floor is derived from the cited story's profile, but §5 also assigns floor 3 when no story is cited or a cited story is unprofiled, where no profile supplies the value | The corrected knob row still gives a false account of two ordinary floor states | Say the floor derives under §5 from the governing cited-story set, explicitly including the no-story and unprofiled defaults +MINOR | high | Task 12 CHANGELOG entry, lines 596-599 | The entry says the two forms are pinned because a program parses them, but P8 is deferred and Task 16 is only a one-time verification, not a shipped programmatic consumer | The release history overclaims present enforcement and obscures the distinction between a future parser and rollout evidence | Use future-facing wording for the deferred parser and separately name this branch's one-time parse check +MINOR | high | Falsified-sentence catalogue versus `docs/coding-workflow.md:123-131` and `docs/getting-started.md:102-106` | Both live explanatory duty summaries still enumerate battery, reason, and profile evidence for a skipped cycle but omit Plan B's newly mandatory provenance line and skip record | User-facing documentation teaches an incomplete close even after the prompts are repaired | Add both sites to the catalogue and extend their existing skipped-cycle lists with the provenance line and skip record without dropping current duties +END OF FINDINGS (20 total) diff --git a/.context/codex-reviews/gate-a-plan-planc-pass-3.md b/.context/codex-reviews/gate-a-plan-planc-pass-3.md new file mode 100644 index 0000000..b6ede3e --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-planc-pass-3.md @@ -0,0 +1,21 @@ +BLOCKER | high | Task 20 Steps 2-3, lines 929-969, and Self-Review lines 989-990 | Step 2 never assigns `body`, yet Step 3 expands `$body`; stripping revision 2's validation block also stripped its `body=$(mktemp)` action while the Self-Review still claims inherited MINOR 12 is discharged at Task 20 | The closing command cannot be executed from the plan in a fresh shell, and an ambient `body` value could target an unrelated file, so the cycle cannot close safely | Start Step 2 with `body=$(mktemp)` followed by an immediate nonzero-exit failure branch, assemble the aggregate into that exact path, and preserve it on every failure path +BLOCKER | high | Task 17, lines 826-843 | Rejecting only a merge does not establish that HEAD is the intended sole WIP or that its parent is Plan A's true base; a stacked WIP, an out-of-order run, or a pre-existing `.context/plan-a-base-sha` is silently converted into a new parent-derived base and any prior value is overwritten | Gate B can review only the newest WIP while excluding earlier A/B/C changes from the combined diff, recreating pass-2 Blocker 1 as an unguarded action | Keep base recording as an action, but recover the parent of the first intended WIP from an independent execution anchor, require the complete intended A+B+C range, and preserve or validate any existing base file before writing +MAJOR | high | Task 18 item 1, lines 853-856 | The only literal version-check invocation is `sh scripts/check-version-bump.sh`, which exits 2 with a usage error because the checker requires exactly one base-ref argument; prose later saying to pass the recorded base does not make the documented command executable | An executor following the named full battery cannot obtain green evidence for invariant 12 and may omit the checker while believing the obligation was followed | State the executable invocation as `base=$(cat .context/plan-a-base-sha)` followed by `sh scripts/check-version-bump.sh "$base"` within the full AGENTS.md battery +MAJOR | high | Task 18 item 3 versus Task 20 Step 2, lines 874-876 and 929-959 | The risk-path verification is scheduled before any provenance lines are assembled; four reconstructed lines are first mentioned in Task 20 and the Gate-B line cannot exist until Task 19 has finished | The item cannot recompute and compare the five claimed lines at its scheduled point, so the named risk verification remains unperformable in plan order | Draft the four reconstructed records before their verification, then assemble the Gate-B record after Task 19 and repeat the five-line comparison immediately before closure +MAJOR | high | Task 18 item 5, lines 884-888 | "Construct and parse, for each" assigns every listed feature to both grammars even though quoted paths and `unusable(...)` are provenance-only while gapped specs, split models, skips, and unknown counts are curve-only | The evidence obligation is internally impossible as written, so an executor must silently reinterpret it and can leave one grammar's real productions uncovered | Partition the fixtures by grammar, require every applicable valid production and at least one invalid provenance instance plus one invalid curve instance, and retain the feature-to-instance record +MAJOR | high | Task 18 item 4, lines 878-882 | The stripped knob state machine's broken-symlink defect has no surviving state rule: "if ... exists" does not say whether lexical presence with a missing target is absent or `unusable(unreadable)` | A present unusable knob can again be recorded as absent, producing a false provenance clause | Define lexical symlink presence before dereferenced existence and route a broken or non-regular path to the pinned unreadable cause +MAJOR | medium | Task 18 item 4, lines 878-882 | Bytes, digest, and value class are separate observations with no requirement that they come from one immutable snapshot; a concurrent writer or symlink retarget can make them describe states that never coexisted | Preservation evidence and the provenance clause can agree individually while being jointly false | Require one securely captured snapshot for the before observation and one for the after observation, deriving bytes, digest, and class from each snapshot +MAJOR | high | Task 19 single-branch recovery, lines 913-916 | The recovery instructions preserve the successful branch file but do not capture the full reviewed commit with each branch or require equality, despite Plan B and the story defining two branches as one logical pass only on the same tracked revision | A resumed failed branch after an amend can be summed with a successful branch from an older commit and recorded as one valid Gate-B pass | Capture the reported 40-character reviewed commit with each branch result and end the earlier logical pass as incomplete on any mismatch +MAJOR | high | Tasks 19-20, lines 918-927 | Re-review after an accepted fix is instructed, but no final obligation compares HEAD with the commit reviewed by both branches of the final clean pass | An external edit, amend, or late fix after the last review can enter the closing commit without Gate-B coverage while all evidence text remains unchanged | Persist the final clean pass's reviewed commit, compare it with HEAD immediately before closure, and require a new full pass on any mismatch +MAJOR | high | Global Constraints and Tasks 1-17 | Explicit `git add` paths do not expose unrelated changes already staged or already inside the WIP range, and no step inspects the full recorded-base-to-HEAD path set against the declared A+B+C surface | Out-of-scope user work can be absorbed into the aggregate commit and closed under this story even though it was never assigned to the fix set | Before Gate B, inspect the complete range plus index/worktree state against an exact A+B+C allowlist and stop on every extra path +MAJOR | high | Task 19, lines 909-916 | The pre-rule cycle uses bare findings slots and unconditionally deletes them, but this workspace already contains bare `gate-b-spec-pass-1.md` and many bare `gate-b-quality-pass-

.md` files from earlier cycles | Execution destroys prior review evidence, the exact data-loss class Plan B's nonce rule exists to prevent | Preflight every bare target and preserve an occupied prior-cycle file at a clearly archival non-slot name or stop for human disposition before deletion +MAJOR | high | Tasks 10 and 12, lines 477-512 and 570-626 | Both insertion replacements re-emit their OLD anchor, while their checks only require the NEW marker to exist and are not declared as preflight state machines | Resuming or rerunning either task can duplicate the prompt-standards note or the 0.11.0 changelog entry, so the plan is not idempotent under interruption | Give each task OLD-only, exactly-one-NEW, mixed, and duplicate states; apply only in OLD-only, skip exactly-one-NEW, and stop on every other state +MAJOR | high | Task 20 Steps 1-2, lines 927-959 | Removing the closing-body validator also removed the only obligation to check the actual five provenance lines and five curves; Task 18 item 5 parses constructed fixtures, not the branch's aggregate instances | The final body can contain malformed records, wrong cardinality, mismatched cycle fields, or placeholders while all seven evidence obligations are still claimed complete | Add a pre-close author obligation over the assembled body itself: five exact grammar-valid lines of each form, matching cycle fields and curve cardinalities, with no placeholders +MAJOR | high | Task 20 Step 2, lines 934-939 | The four reconstructed records must be marked, but the plan does not say where that marker goes and the pinned provenance and curve grammars have no reconstruction field | An executor can append the marker to a pinned record line, making the aggregate body unparseable, or omit it to preserve grammar and silently claim contemporaneity | Specify a separate adjacent reconstruction-source line for each cycle, outside both pinned record lines, naming the validated pass files while leaving the record bytes grammar-valid +MAJOR | high | `docs/superpowers/plans/2026-08-29-review-loop-economics.md`, especially lines 45-63 | The superseded single-plan implementation artifact remains titled and written as an executable plan with no superseded marker; it directs a different commit topology, including a separate docs-only commit and its own Gate-B close | An executor discovering it can follow a stale path that contradicts the settled A→B→C sequence and single combined Gate-B cycle | Add an unequivocal top-of-file superseded notice pointing to Plans A, B, and C and stating that none of its tasks are executable +MINOR | high | Task 1 NEW, line 118 | The README replacement says the obliged floor is derived from the cited story's profile, but §5 assigns floor 3 when no story is cited or a cited story is unprofiled, where no profile supplies the value | The corrected user-facing knob row still teaches a false derivation for two ordinary states | Say the obliged floor is determined by §5 from the governing cited-story set, including its no-story and unprofiled defaults +MINOR | high | Falsified-sentence catalogue and Tasks 14-16 versus `docs/coding-workflow.md:123-131` and `docs/getting-started.md:102-106` | Task 14 says the skipped-cycle claim lives in three files, but both explanatory duty summaries also omit the newly mandatory provenance line and skip record | The sweep remains incomplete and users can still close a skipped cycle from documentation without the records Plan B requires | Add both sites to the catalogue and extend each summary with the per-cycle provenance line and skip record while preserving its existing battery, reason, and evidence duties +MINOR | high | Task 12 CHANGELOG replacement, lines 596-599 | The new release entry says the forms are pinned "because a program parses them", but the parser is deferred to P8 and Task 18 is only one-time rollout evidence, not a shipped programmatic consumer | The changelog overclaims current enforcement and blurs a future parser with branch-local evidence | Use future-facing wording for P8 and describe Task 18's parse exercise separately +MINOR | high | Task 20 Step 2, lines 934-946 | The rationale says all four Gate-A cycles closed before Plan B pinned the forms, but the Plan B and Plan C Gate-A cycles occurred after the Plan B artifact stated those forms; what is true is that all four closed before the rules shipped and without native records | The reconstruction annotation carries a false chronology, weakening the very provenance distinction it is meant to make honest | State the chronology per cycle: the spec cycle predates the forms, while the three plan cycles predate rule activation and closed without emitting the forms +NIT | high | Task 17 rationale, line 842 | `HEAD^` is not ambiguous on a merge; Git defines it as the first parent, while the real issue is that choosing that parent is not justified as the combined cycle base | The explanation gives the wrong reason for the otherwise conservative merge refusal and can mislead later maintenance of base recovery | Say a merge is refused because first-parent selection would not establish the intended review base +END OF FINDINGS (20 total) diff --git a/.context/codex-reviews/gate-a-plan-planc-pass-4-dispositions.md b/.context/codex-reviews/gate-a-plan-planc-pass-4-dispositions.md new file mode 100644 index 0000000..b78cad8 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-planc-pass-4-dispositions.md @@ -0,0 +1,93 @@ +# Plan C — Gate-A pass 4 dispositions (23 findings) + +Pass 4 triggered CLAUDE.md §5's mandatory stop-and-surface: four of the five tells present +(count rising 18/20/20/23; Blockers 5/4/2/4; 17 of 23 clustering on the instrument; +require<->withdraw pair on the superseded-artifact marker). Daniel ruled 2026-08-30, routed +through dev-workflow-kit-56. Revision 5 is f2cfbc7. + +## The governing ruling +Reduce Tasks 18-21 to "run the Gate-B cycle per CLAUDE.md §5" plus exactly three plan-specific +facts: record the base commit at cycle open; the `rle` slot discriminator; the cycle runs under +the OLD rules per the activation constraint. §5 stays the only source; no protocol +re-specification in the plan. + +## Per finding + +FIXED BY DELETION OF THE RESTATEMENT (not repaired individually, per the ruling — they were +parts of the copy, and repairing them would have kept it): +- BLOCKER 1010-1057 — unassigned `body` across fenced blocks, empty mktemp file. +- MAJOR 884-892 — `.context` never created before the redirect. +- MAJOR 886-892 — redirect follows a symlink and truncates its target. +- MINOR 887-890 — `-\?wip` matches "wip" anywhere in a subject. +- MAJOR 1049-1060 — no recovery rule for a failed closing amend. +- MINOR 1049-1057 — post-amend check accepts any non-`WIP:` subject. +- MAJOR 991-994, 1000-1002 — branch/reviewed-commit persistence across a resume. §5's own + single-branch recovery rule; the plan defers rather than restating. +- MAJOR 904-962, 981-982, 1020-1032 — evidence-entry construction and carry. §5 states the + revalidation obligation (inherited B6); Task 19 produces the entry. +- BLOCKER 927-931, 978-982, 1008-1018 — risk-path verification ordering. Task 19 item 3 now + states the ordering constraint without scheduling it against deleted task numbers. +- MAJOR 967-976 — `git diff --name-only` overclaimed. Kept as a check, reworded to the + comparison it actually performs: path names, not content. +- MAJOR 60-68, 97-875, 879-900 — preflight ordering across the amend chain. §5's cycle + procedure; the plan no longer sequences it. +- MAJOR 984-994 — no ownership guard before slot deletion. §5's delete-before-call step owns it. +- MAJOR 933-945 — knob snapshot procedure. Task 19 item 4 keeps the one-observation rule + (inherited M9); the storage/digest procedure was instrument and is gone. +- MAJOR 1020-1026 — reconstruction sources under ignored `.context`. Task 20 states the + sources; the inventory procedure was instrument. + +RESOLVED BY THE MARKER MOVE: +- BLOCKER 1023-1032 — the in-record `— reconstructed ...` suffix cannot coexist with the pinned + grammars. The marker moves BESIDE each record as its own adjacent line in the closing body. + Grammars stay pure; the adjacent line rides the same squash carry, since the carry copies the + body's records and the line sits in the same block. + +RESOLVED BY DELETION OF AN UNEXECUTABLE OBLIGATION: +- BLOCKER 947-952, 1028-1032 — the parse obligation names no parser and none exists. Scoped to + executable checks, consistent with the assert-new philosophy revision 3 applied to the six + checks it stripped: the records are written FROM the pinned grammars and checked by reading + (the Gate-B reviewer receives the closing body). Task 19 item 5 now also names the productions + this branch cannot demonstrate, with reasons, rather than implying enforcement that does not + exist. P8 ships the programmatic consumer. + +RESOLVED IN THE SPEC'S FAVOUR: +- MAJOR 826-875 — Task 17 vs spec scope. Task 17 removed; supersession remedies are out of + scope. The superseded single-plan artifact gets one line in the closure record. This is the + withdraw half of the require<->withdraw pair with pass 3's MAJOR, which had required the + marker. + +ANSWERED, NOT REVISED: +- MAJOR 1020-1041 — "aggregate reconstruction is a contract change made by the loop". Answered + by the standing (B)-transitional decision, grounded in spec §10's activation rule: the records + contract binds cycles that start under the new rules, and all five cycles here are pre-rule. + Cited here rather than revising the spec. If pass 5 contests it again as a contract change, it + routes to Daniel as a spec question and is not absorbed a second time. + +FIXED IN THE ROLLOUT (in scope — these are the deliverable, not the instrument): +- MAJOR 78-94, 1064-1067 — the skipped-cycle sweep was two files short. + `docs/getting-started.md:105` and `docs/coding-workflow.md:129` enumerate the same duties and + omit the provenance line and skip record. New Tasks 17 and 18; Tasks 14-16 renumbered "of 5"; + accounting rows 17 and 18 added. +- MAJOR 73-94, 631-875 — the old-conditions accounting stopped at Task 12. Extended through + Task 18, and its intro no longer claims none of these passages are §5 rules: rows 15 and 16 + are. +- MINOR 16-18, 631-672, 904-962 — the header pointed at Task 13 for the mode; now Task 19. + +FIXED ACROSS EVERY EDIT TASK (both are Majors, so both resolve — not carried): +- MAJOR 121-127, 862-868 — every `grep -cF` was labelled an assertion but succeeded at any + positive count. All 18 asserts are now equality tests that exit nonzero on any other value. + The two remaining bare `grep -c` invocations are deliberate: Task 19 item 2 is a + counterfactual observation across two revisions, which must print rather than assert. +- MAJOR 477-506, 570-620, 840-868 — no done-state preflight, so a rerun could duplicate an + insertion. One Global Constraint makes every task skip-if-applied: run its assert first, `1` + means skip, `0` means apply, anything else is a stop rather than a retry. A second constraint + names Tasks 10 and 12 as the two that re-emit their anchor by design, so only their NEW count + discriminates. One rule rather than twenty state tables. + +Every bash fence in the plan was checked with `sh -n`: 37 blocks, 0 syntax failures. + +## Inherited obligations after revision 5 +M9 at Task 19 item 4. M10 in Global Constraints. M8 and B6 are §5's own rules and are deferred +to, not restated. MINOR 12 is moot: it asked for `mktemp` over a fixed `/tmp` path, and the plan +no longer scripts the close. diff --git a/.context/codex-reviews/gate-a-plan-planc-pass-4.md b/.context/codex-reviews/gate-a-plan-planc-pass-4.md new file mode 100644 index 0000000..021adc2 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-planc-pass-4.md @@ -0,0 +1,24 @@ +BLOCKER | high | lines 927-931, 978-982, 1008-1018 | the named risk-path verification is scheduled only after Gate B and after the final evidence revalidation, even though the governing mode requires the named verification before Gate B and its result changes the evidence entry | the first review calls cannot carry a truthful current evidence entry, and the closing amend can carry evidence changed after the last revalidation | draft the provenance lines and run this verification before Task 20, carry that result in every Gate-B call, then rerun the verification and the final evidence revalidation after the final pass and before assembly +BLOCKER | high | lines 1010-1057 | `body` is assigned in one fenced shell block and expanded in a later block, so Step 3 has no `body` variable in a fresh shell; moreover the plan never writes a subject or any content to the empty `mktemp` file | the closing command is not executable exactly as written and will either fail on an unset or empty path or attempt an amend with an empty message, so the cycle cannot close | put allocation, population, validation, amend, and cleanup in one fenced script, or persist the temp path safely; include an explicit non-WIP subject and the complete body template +BLOCKER | high | lines 1023-1032 | the four reconstructed records gain a trailing `— reconstructed ...` suffix, while the pinned provenance and curve grammars admit no such suffix, yet the same step requires every record to parse against those grammars | four required records cannot simultaneously carry the marker and pass the mandatory grammar validation | obtain an approved grammar change that represents reconstruction, or keep each record byte-conformant and define a separately pinned, squash-carried companion record for reconstruction provenance +BLOCKER | high | lines 947-952, 1028-1032 | the plan requires positive and negative parse tests and later requires the assembled body to parse, but names no parser, implementation, or executable assertion, and the repository contains no parser for either new grammar | the required parse evidence and closing validation have no reproducible way to succeed or fail, so malformed machine-consumed records can close | provide a concrete checked-in or disposable parser plus exact commands and expected exit statuses for every valid and invalid fixture and for the final body +MAJOR | high | lines 1020-1041 | four already-closed Gate-A cycles are reconstructed into the combined implementation commit even though the spec and story require each cycle's records in that cycle's own closing commit body | an aggregate later body is not the historical location the settled contract requires, so the plan silently changes a machine-read record-placement rule after Plans A and B closed | either rewrite the four actual closing commits with their records before Gate B, or stop and obtain an explicit spec and story revision permitting marked aggregate reconstruction +MAJOR | high | lines 884-892 | Task 18 redirects into `.context/plan-a-base-sha` without creating `.context`, although ignored state is absent in a fresh checkout and may be cleared between Plans B and C | the first persistent base step fails on a valid resumption environment, and later commands then read a missing base | create and verify `.context` before the redirect, fail immediately if the write or read-back fails, and validate the stored value as the expected full commit id +MAJOR | medium | lines 886-892 | the base-record redirect follows an existing symlink and truncates any file it targets, with no type or ownership check on the ignored path | a hostile or stale workspace entry can turn a bookkeeping step into unrelated data loss | refuse a symlink or non-regular destination, write through a private temporary file in a verified directory, then atomically rename after validating the value +MAJOR | high | lines 1049-1060 | the failure branch preserves the message file but gives no recovery rule for the hook's actual behavior: any non-WIP commit command resets Gate-B state even when the commit fails | retrying the preserved body after a failed amend can close with all reviewed-pass state discarded and no fresh Gate-B cycle | state that any failed closing amend invalidates the cycle, repair the cause while HEAD remains WIP, rerun the required Gate-B passes and evidence revalidation, then retry with the preserved body +MAJOR | high | lines 967-976 | the step claims to prove the tree contains only this change, but `git diff --name-only` proves only that path names look expected and does not inspect the content the reviewer will see | unrelated or hostile edits inside an allowed path can be folded into the WIP and pass this check, contradicting the plan's observability claim | inspect and record the full `base..HEAD` diff, verify exact expected paths and hunks, and reword the claim to the precise comparisons actually made +MAJOR | high | lines 60-68, 97-875, 879-900 | Tasks 1-17 repeatedly amend before the first executable check that HEAD is the intended WIP and before any clean-index or staged-content preflight | an out-of-order or resumed run can rewrite an unrelated tip, and pre-staged content can be absorbed into every subsequent amend before Task 20 notices it | move the merge, exact WIP-subject, parent, clean-worktree, and staged-diff checks ahead of Task 1 and repeat the WIP identity check before each amend +MAJOR | high | lines 477-506, 570-620, 840-868 | Tasks 10, 12, and 17 re-emit their OLD anchor inside NEW but have no done-state preflight, so rerunning them can insert the note, changelog release, or supersession banner again | interruption and retry can create duplicated shipped prose and duplicate changelog versions; the later `grep -cF` still exits successfully | add the same zero-or-one state table Plan B uses, skip at exactly one, apply only at zero, and stop on asymmetric or duplicate counts +MAJOR | high | lines 121-127 and the repeated assert steps through lines 862-868 | every task labels `grep -cF` an assertion, but the command succeeds for any positive count and never checks that OLD disappeared | a duplicate insertion or an additive edit leaving contradictory OLD text can be reported as successful | use explicit equality tests for exactly one NEW occurrence and zero OLD occurrences, with a nonzero exit on every other state +MAJOR | high | lines 73-94, 631-875 | the old-conditions accounting stops at Task 12 and says the passages are not §5 rules, but Tasks 13-17 were added later and Tasks 15-16 directly rewrite the governing §5 decision procedure | required kept, moved, or deliberately-dropped conditions for five rewrites are absent, recreating the exact invariant-11 failure class the accounting is meant to prevent | extend the table through Tasks 13-17 and account separately for every condition in the model-recording rule, all three skip-duty copies, and the supersession edit +MAJOR | high | lines 78-94, 1064-1067 | the claimed complete rollout omits two user-facing skip descriptions: `docs/getting-started.md:102-107` and `docs/coding-workflow.md:123-132` still enumerate battery, reason, and evidence without the provenance line and skip record Plan B makes mandatory | readers can follow those documents and close a skipped cycle without the records the shipped prompts require | add both sites to the rollout catalogue, accounting, edit tasks, exact assertions, parity review, and prompt or documentation review scope as applicable +MAJOR | high | lines 826-875 | Task 17 adds a supersession remedy even though the approved spec explicitly places any remedy to the supersession convention out of scope | Plan C expands the settled change and the combined Gate-B diff without an approved decision, while also spending the versioned implementation cycle on unrelated repository hygiene | remove Task 17 or first revise and approve the spec and story to include this bounded supersession action and its acceptance evidence +MAJOR | high | lines 1020-1026 | the four reconstructed records depend on unspecified `validated pass files` under ignored `.context` state, with no inventory, presence check, committed source, or unknown-count fallback | a fresh checkout or interrupted handoff cannot reproduce the records and may force the executor to guess counts or falsely claim validation | enumerate every source path and expected digest before execution, stop if provenance is unavailable, and define honest `?` or `undetermined` handling where the pinned grammars permit it +MAJOR | medium | lines 984-994 | the deterministic `rle` discriminator avoids historical bare slots but has no ownership or concurrency guard before deleting its targets | two simultaneous or resumed Plan-C executors compute the same paths, so one can delete or overwrite the other's live result exactly like the incident the slot rule addresses | establish exclusive cycle ownership before deletion, or detect any live writer and stop; record the owner and invocation identity in a durable resume record without misrepresenting it as the new-rule nonce +MAJOR | high | lines 933-945 | the required one-observation knob verification specifies classifications but no executable snapshot, storage location, digest algorithm, race handling, or after-cycle comparison command | bytes, digest, path type, and value class can come from different filesystem states, and a later reader cannot audit whether the user's knob survived | define a safe snapshot procedure that records lstat state and captured bytes atomically enough for the stated bound, stores the pre-state outside the reviewed diff, and performs an explicit post-state equality assertion +MAJOR | high | lines 991-994, 1000-1002 | single-branch recovery says to capture the surviving branch's reviewed commit and final acceptance says to compare HEAD, but neither defines how to extract, persist, or validate the full commit id or the branch session ids | an interruption loses the only evidence needed to prove both branches reviewed one revision, making a resumed logical pass unverifiable | specify a durable resume record containing branch, full reviewed commit, session id, slot, validation status, and recovery budget, plus exact equality checks and retirement at close +MAJOR | high | lines 904-962, 981-982, 1020-1032 | Task 20 assumes a `current evidence entry` exists and Task 21 assumes the body carries it, but no prior step assembles, stores, or validates the exact entry containing the seven evidence results | Gate-B calls cannot quote a stable entry verbatim and the closing body can omit or mutate evidence without detection | add an explicit evidence-entry construction task with a fixed temporary location, exact story path and named results, counterfactual observations, checksum or byte comparison, and a byte-for-byte carry into every call and the closing body +MINOR | high | lines 16-18, 631-672, 904-962 | the header says Task 13 reads the story mode and produces its evidence, but Task 13 is now the model-recording documentation edit and the evidence pack is Task 19 | stale numbering can send an executor to the wrong task and undermines the plan's own execution map | change the header reference from Task 13 to Task 19 +MINOR | high | lines 887-890 | the parent-WIP check uses the basic regular expression `-\?wip`, which matches `wip` anywhere in the subject rather than the documented WIP prefix | ordinary subjects containing that substring can trigger a false stacked-WIP stop | match the exact accepted WIP subject or an anchored case-insensitive `^WIP:` prefix +MINOR | high | lines 1049-1057 | the post-amend success check rejects only a subject starting `WIP:` and accepts any other subject, including an evidence or provenance line accidentally placed first | the command can print CLOSED for a malformed commit message and gives weak observability of what actually shipped | specify the intended final subject and assert exact equality, then separately verify the expected body records +END OF FINDINGS (23 total) diff --git a/.context/codex-reviews/gate-a-plan-planc-pass-5.md b/.context/codex-reviews/gate-a-plan-planc-pass-5.md new file mode 100644 index 0000000..b59a5e1 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-planc-pass-5.md @@ -0,0 +1,23 @@ +BLOCKER | high | lines 985-995 | Task 19 replaces the approved spec §8 parse check with unaided reading of ten branch records and explicitly refuses to provide a parser | the settled evidence contract requires executable valid and invalid grammar cases, including rejection and cardinality behavior that these native records cannot exercise, so Plan C does not implement §8 as claimed | restore a concrete parser or executable checker and run constructed accepted and rejected fixtures plus the final records, or obtain an approved spec and story revision before execution +BLOCKER | high | lines 942-1004, 1008-1062 | Task 19 is ordered before Task 20 but item 3 waits for provenance lines that Task 20 only describes, item 4 cannot perform its after-cycle comparison until Task 20 finishes, and item 5 needs the Gate-B curve that does not exist until Task 20 runs | task-by-task subagent execution has a dependency cycle, so Task 19 cannot truthfully complete and Task 20 cannot start with the promised completed evidence pack | split Task 19 into an explicit pre-cycle evidence task and a post-final-pass revalidation and assembly task, with the latter preceding the closing amend +BLOCKER | high | lines 1010-1012, 1040-1062, 1080-1093 | deferral to §5 removes every action that allocates, populates, validates, supplies to the reviewer, and commits the Plan-C-specific closing body; §5 supplies only a generic closing amend with a placeholder message | an executor has no exact final subject, no safe body file, no command that writes the evidence and ten records, and no way for Gate B to read the uncommitted body, so the required close is not executable or reviewable | keep §5 authoritative for the loop but add Plan-C-specific body construction using mktemp, exact population and validation, verbatim inclusion in every relevant review context, and one checked amend command with cleanup only after success +BLOCKER | high | lines 947-960, 1017-1019, 1035-1038 | `$base` is expanded in Task 19 before any assignment in Plan C, while Task 20 says to record it at cycle open even though Plan A opened the cycle earlier and only printed it | the quoted `git show` command and the version-bump battery substitution fail in a fresh shell, and the later diff and every Gate-B call lack a durable verified base | before Task 19, verify the exact WIP tip and single non-WIP parent, assign and persist its full commit id safely, read it back into `base` in every fresh shell, and use that value throughout +MAJOR | high | lines 70-75, 142-155 | the global resume rule says to run each equality assertion first and treat count 0 as permission to apply, but every assertion exits 1 at count 0 | following the stated preflight procedure stops every unapplied task before its edit, while following the per-task checkbox order forfeits the promised skip-if-applied safety | replace each preflight with an explicit case over 0, 1, and all other states that returns an apply or skip decision without aborting on the valid 0 state, then run a separate post-edit assertion +MAJOR | high | lines 70-78, 521-527, 637-643 | a count of one short sentinel is treated as proof that a whole replacement or insertion completed, without checking the full NEW block or that OLD disappeared where it should | a half-applied edit containing the sentinel, or contradictory OLD and NEW text together, is misclassified as done and survives into the shipped prompts despite the plan claiming half-applied states stop | verify the complete expected block exactly once and the complete OLD block zero times for replacements; for insertions verify the full inserted block and its exact placement around the retained anchor +MAJOR | high | lines 70-75, 150-155, 887-938 | the done-state rule skips the whole task when text is present but does not verify that the path was included in the WIP amend | after an edit succeeds and `git commit --amend` fails or execution is interrupted, resumption skips the amend and can leave dirty or staged content for a later task to absorb accidentally or for Task 20 to stop on without a recovery path | make resume state include both content and commit state; if the exact intended path diff is still outside HEAD, inspect it, stage that path explicitly, and complete the WIP amend before declaring the task done +MAJOR | high | lines 60-68, 118-938 | Tasks 1-18 amend HEAD repeatedly without first proving that HEAD is the intended single WIP commit, that its parent is the cycle base, and that no unrelated staged content is present | out-of-order execution or a resumed run can rewrite an unrelated commit and fold pre-staged user changes into the reviewed story before the late Task-20 status check notices anything | add a preflight before Task 1 and before each amend that checks the exact WIP subject, non-merge shape, expected parent, clean index except for the inspected task paths, and absence of stacked WIPs +BLOCKER | high | lines 1042-1071 | the plan reconstructs four already-closed Gate-A cycles into the later combined implementation commit even though the approved spec and governing story require each cycle's record in that cycle's own closing commit body and explicitly say reconstruction is neither needed nor claimed | an adjacent disclaimer does not change the required historical destination, so the machine-read records would be stored in the wrong commit under a contract the plan has no authority to rewrite | amend the actual four closing commits before Gate B, or stop and obtain explicit approved revisions to the spec and story that define aggregate reconstruction, its destination, and its machine semantics +MAJOR | high | lines 985-990, 1042-1051 | the four reconstructed curves and provenance lines are said to come from validated pass files, but the plan inventories no source paths, digests, pass numbers, counts, models, or missing-data behavior | ignored `.context` state may be absent or stale after interruption or on another checkout, leaving the executor unable to reproduce the records without guessing and leaving later readers unable to audit them | enumerate every source artifact and expected digest before execution, validate it, and use the grammar's `?` and `undetermined` values honestly where recoverable evidence is missing +MAJOR | high | lines 1045-1051 | the plan claims an adjacent reconstruction marker will travel with records through squash carry, but the shipped carry rule manually enumerates provenance lines, curves, and skip records and does not include arbitrary adjacent prose | a squasher following the rule can copy the records and omit the marker, making reconstructed records appear contemporaneous | add an approved, explicitly squash-carried companion record or extend the carry rule and both prompt copies to name the marker; adjacency alone is not a mechanism +MAJOR | high | lines 1010-1012, 1040-1062 | no recovery is defined for a failed closing amend even though the actual hook resets all Gate-B state after any non-WIP commit command regardless of whether Git succeeds | retrying a preserved or rebuilt body after an amend failure can close with the reviewed-pass counters and fingerprint discarded | prevalidate everything before issuing the non-WIP amend and state that any failed attempt invalidates the cycle, requiring repair while HEAD remains WIP, a fresh full Gate-B floor, evidence revalidation, and only then another close attempt +MAJOR | high | lines 966-1004, 1032-1043 | Task 20 assumes a current evidence entry that can be quoted verbatim and the closing body assumes the same entry, but no step assembles, stores, validates, or byte-compares that exact artifact | different subagents or resumptions can send one claim to Gate B and commit another, and several Task-19 results can be silently omitted | define the evidence-entry format and durable temporary location, populate it from all required observations, validate it, and carry the same bytes into every call and the closing body with a final equality check +MAJOR | high | lines 971-983 | the inherited knob verification names states but supplies no executable one-observation snapshot, safe storage, digest algorithm, race bound, or post-cycle comparison | separate existence, type, byte, digest, and classification reads can describe different filesystem states, while a stale or hostile symlink can cause unrelated reads or writes and the executor cannot later prove the knob survived | specify a symlink-aware snapshot into a private temporary location, derive classification and digest from the captured bytes, record the bounded race limitation, and perform an explicit after-cycle type and byte equality assertion +MAJOR | high | lines 1035-1038 | the only combined-diff scope check reads path names and knowingly does not inspect hunks inside allowed files | unrelated or hostile edits already inside `CLAUDE.md`, `workflow-init.md`, or another legitimate path can be closed under this story even though explicit staging only protects against stray paths | require an executor to inspect and record the full base-to-HEAD diff against every task's exact OLD-to-NEW change, while keeping the path-set check as a separate coarse guard +MAJOR | medium | lines 1021-1026 | every concurrent or restarted Plan-C executor uses the same deterministic `rle` slots with no ownership or live-writer check before §5's destructive delete-before-call step | two actors can delete or overwrite each other's valid findings files, recreating the data-loss incident the slot work is meant to prevent | establish an exclusive cycle ownership record or lock before any deletion, validate ownership on resume, and stop on a live conflicting owner without mislabeling the pre-rule discriminator as a nonce +MAJOR | high | lines 985-995, 1059-1062 | the plan names only the real nonce and non-absent knob as productions this branch cannot demonstrate, but native records also cannot cover quoted paths, unprofiled and no-story sets, unusable-knob causes, skipped records, rejected duplicate paths, count-cardinality rejection, and other grammar alternatives | the claimed grammar coverage materially overstates what ten homogeneous records establish and conflicts with the spec's explicit reason for constructed fixtures | enumerate every production and constraint separately, mark native coverage honestly, and exercise the rest with valid and invalid parser fixtures +MINOR | high | lines 955-964 | the two counterfactual commands merely print independent match counts and assert neither the expected old and new states nor mutual exclusion of the contradictory phrases | a base or result containing both forms, or an unexpected count, can still be recorded without the named check failing even though the task labels it a check that fails without the change | use explicit expected equalities for old-present and new-absent at the base and old-absent and new-present at HEAD, in both prompt copies where the rule ships +MINOR | high | lines 329-346 | Task 6 describes the hook numerator as all calls the cycle made and the fresh value as how many ran against the fingerprint | the hook excludes recognized failure, no-result, and backgrounded calls and can count unrecognized result shapes, so the new user-facing explanation is mechanically false | say `calls the hook counted this cycle` and `counted calls whose stored fingerprint equals the current fingerprint`, while retaining the warning that neither proves a valid pass +MINOR | high | lines 627-629 | the CHANGELOG says findings slots take a per-cycle infix and deletion now removes only nonce-bearing paths without qualifying that the shipped rule preserves bare slots for pre-rule or legacy no-nonce cycles | the release note contradicts Plan B's compatibility exception and even the special pre-rule cycle Plan C itself runs | qualify both claims with `for a nonce-holding cycle` and state that pre-rule or legacy no-nonce cycles retain the bare form +MAJOR | medium | lines 1053-1057 | adding a closing-body declaration that the old single-plan artifact is superseded and non-executable is itself a supersession remedy, although the approved spec places any remedy to the supersession convention out of scope | the plan still expands the settled change after deleting the direct file edit, only moving the new authority into history where an executor may not see it | remove the line, or obtain an approved scope revision defining this closure-record remedy and how readers discover it +NIT | medium | lines 1021-1024 | the exact claim that the workspace already holds 61 bare-slot files is mutable execution-time state with no observation date or recheck | cleanup or another review cycle can make the plan's rationale factually stale even though the safer slot choice remains valid | state that 61 was observed when the plan was written and require a fresh non-destructive inventory at execution time without making correctness depend on the count +END OF FINDINGS (22 total) diff --git a/.context/codex-reviews/gate-a-plan-planc-pass-6.md b/.context/codex-reviews/gate-a-plan-planc-pass-6.md new file mode 100644 index 0000000..9ea319f --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-planc-pass-6.md @@ -0,0 +1,20 @@ +BLOCKER | high | lines 1117-1126 and 644-648 | Task 21 mandates deterministic `rle`-infixed Gate-B slots, but Plan B's shipped slot rule and Task 12's own CHANGELOG text say a pre-rule or other no-nonce cycle keeps the bare form; the `rle` form is admitted by neither the old exact slot list nor Plan B's nonce grammar | An executor cannot simultaneously follow the plan and the governing prompt, and the findings files may fail the protocol whose validation makes a pass count | Resolve the source-of-truth conflict before execution: either add an explicit approved pre-rule discriminator production to both prompt copies and align the CHANGELOG, or use the bare form with a separately specified preservation procedure +MAJOR | high | lines 1117-1126 | Re-inventorying before deletion has no terminal action when an exact `rle` target or an active writer already exists; because `rle` is deterministic, a resumed run or concurrent sibling can own the same path and §5's delete-before-call step will erase or race it | Findings files are gitignored, can be irreproducible, and a late writer can leave a valid-looking file from the wrong invocation, so this is both data loss and false-pass risk | Require an ownership and inactivity check for every exact target, stop rather than delete an unknown owner's file, and either preserve existing artifacts or choose a collision-resistant cycle-scoped target +MAJOR | high | lines 70-84 | The plan says every task is skip-if-applied, but Tasks 1-19 contain no instantiated preflight; the only preflight is a template containing literal `` and `` placeholders and is not executable as written | A fresh or resumed executor has no task-local decision procedure and can retry a replacement, duplicate an insertion, or fail on a missing OLD block | Put a concrete preflight before every edit, with the actual file, NEW sentinel, and expected OLD-state check for that task +MAJOR | high | lines 76-94 | Even if the generic preflight is instantiated, its `1` branch declares the task complete and exits before validating the site, while lines 85-90 explicitly admit that a half-applied edit can contain the sentinel exactly once alongside OLD text | The advertised resumption path accepts a corrupt partial state as successfully applied and can carry contradictory prompt text into the WIP commit | On a count of one, read and validate the entire expected NEW block and the required OLD-text absence, with the stated Tasks 10 and 12 exceptions, before declaring the task complete; otherwise stop for repair +MAJOR | high | lines 1032-1035 | The battery step names an adaptation of the `AGENTS.md` chain but gives no exact command or captured success criterion, especially for replacing the checker's `main` argument with `HEAD^` | An executor can run the documented `main` form, which can compare the wrong range or pass trivially, while still claiming the story's battery evidence | Paste the complete executable chain with `sh scripts/check-version-bump.sh HEAD^`, state the required final status and tool prerequisites, and record the output used by the evidence entry +MAJOR | high | lines 1023-1092 and 1132-1133 | Task 20 never constructs, formats, or stores the "current evidence entry" that every Gate-B call must quote verbatim and the close must carry | After an interruption there is no authoritative text to resume from, and different calls can review different evidence claims while all appearing to follow the plan | Define the exact evidence-entry text after items 1-7, write it to a cycle-local durable scratch artifact, validate it, and require every call and the closing body to read the identical bytes from that artifact +MAJOR | high | lines 1052-1065 and 1154-1157 | The before-state of `.context/codex-gate.floor` is said to include bytes, digest, type, and value class, but no storage location or atomic capture procedure is defined | Task 22 cannot perform a grounded after-cycle comparison after context loss or interruption, so inherited obligation M9 can degrade into recollection | Capture the input once into a private cycle-local snapshot, record a safe byte encoding plus digest, type, and class from that same capture, and define how resumption locates and validates the snapshot +MAJOR | medium | lines 1058-1063 and 1154-1157 | The knob comparison treats path type plus readable target bytes as identity but does not record symlink target bytes and gives no protection against the path changing between type, byte, digest, and classification reads | A broken symlink can be retargeted and still compare as the same unusable state, while a concurrently replaced regular target can produce internally inconsistent evidence, contradicting the claim that all fields describe one observation | Use `lstat`-style path classification, record `readlink` bytes for symlinks, snapshot a readable regular target once before hashing and classifying it, and reject a path that changes during capture +BLOCKER | high | lines 994-1003 and 1067-1079 | The plan names cardinality as the only rule a single `grep -E` match cannot decide, but the pinned forms also contain non-regular constraints: repeated-path rejection, strictly ascending non-overlapping ranges, exact per-pass model keys, and representability rules; Task 20 nevertheless requires a repeated path to fail the provenance grep | The proposed evidence can report both grammars checked while accepting malformed records, and the revised spec would overstate what the named mechanism proves | Keep `grep -E` for lexical productions, add explicit shell checks for uniqueness, ordering, range overlap, per-pass key equality, control characters, and cardinality, and revise Task 19 to name every supplemental comparison rather than calling cardinality the sole exception +MAJOR | high | lines 1071-1074 and 1180-1185 | Task 20's constructed cases omit a real nonce and a valid numeric knob value, even though Task 22 says the constructed strings cover everything the branch's own pre-rule and absent-knob records cannot demonstrate | The two fields expressly unavailable from the live cycle remain unexercised while the close records them as covered | Add valid and invalid nonce cases to both record forms, a valid numeric knob case, and assertions that the same nonce is carried where cross-record agreement is being checked; keep randomness and live attribution recorded as undemonstrable +MAJOR | high | lines 1085-1092 | The parity and prompt-standards obligations have no extraction boundaries, expected comparison output, status artifact, or command that can fail when a changed rule or checklist item is wrong | An executor can mark both items complete after an informal read, and a later reader cannot distinguish a performed 36-row checklist from an assertion that it happened | Specify the compared rule blocks and expected sole divergence, provide a fail-capable parity procedure, and write the twelve-item status matrix for each of the three prompt artifacts to a named evidence artifact consumed by the entry +MAJOR | high | lines 1108-1112 | The quoted WIP preflight exits nonzero in the desired state: the final `git log ... \| grep ... && { ...; }` returns grep's status 1 when the parent is correctly not WIP, and the first negative guard also aborts immediately in any shell running with `errexit` | The exact command cannot serve as a green precondition in a fresh shell and may stop Task 21 before any review call | Rewrite all three guards as explicit `if` statements, positively verify that `HEAD^` exists, and end the block with an explicit successful command +MAJOR | high | lines 356-358 | The satisfied-message replacement defines `` as how many counted calls carry a stored fingerprint equal to the current one, but the hook stores only the last fingerprint and a consecutive streak; an earlier matching fingerprint separated by another fingerprint is not counted | This is a new user-facing overclaim about the exact comparison the hook performs and violates the gate-description invariant | Describe `` as the consecutive counted calls on the current fingerprint since the last fingerprint change, and state that only the last fingerprint plus the streak is retained +MAJOR | high | lines 636-639 and 972-1003 | Task 12 says the two formats are pinned "because a program parses them" while Task 19's rationale says no parser for either grammar exists in this repo and changes the check for precisely that reason | The release notes and the revised spec give incompatible accounts of existing enforcement, causing readers to trust a consumer the plan says does not exist | Identify an actual parser if one exists; otherwise use future-tense wording about intended P8 consumption and describe the current evidence as lexical and supplemental shell comparisons +MAJOR | high | lines 1149-1152 | The risk-path verification recomputes the floor from the `Story:` header of "the artifact this cycle reviewed", but Gate B reviews a diff and has no single artifact header; Plan A requires comparing the spec and every contributing plan's set and derives Gate B from the union of the plans | Looking at only one header can miss an absent, malformed, or disagreeing header and still bless a provenance line with an unlicensed floor | Re-read the spec plus Plans A, B, and C as the expected artifacts, compare their cited sets exactly as the shipped rule requires, then derive and verify the Gate-B union from all contributing plan headers +MAJOR | high | lines 41-45, 1187-1191, and 1215-1218 | The inherited `mktemp` obligation is promised at the top and then declared moot, although Task 22 still has to compose a multi-line real commit body containing several conditional records and gives no executable construction command | Dropping an inherited recovery and safety requirement without accounting reintroduces fixed-path, quoting, truncation, and partial-body risks at the irreversible close | Build the complete body in a `mktemp` file with cleanup trapping, validate its exact lines, and close with `git commit --amend -F "$body_file"`; explain that WIP `-m` recognition is needed for internal amends, not for the deliberately cycle-closing real commit +MAJOR | high | lines 1163-1170 | The section labels its four-item list complete and forbids anything else, but "decline records" are undefined by the listed settled sources while §5 conditionally requires human-exception records when such a decision occurred | The executor has no way to decide what a decline record is and can be told to omit a real record that the governing protocol requires | Remove the undefined record unless its owning rule is already shipped and cited, include every conditional §5 record that this cycle actually produced, and avoid a closed list unless every source obligation is accounted for +MAJOR | high | lines 1171-1178 and 1195-1218 | No task updates `docs/field-reports/2026-08-30-gate-a-rle-plan-cycles.md` with Plan C's final Gate-A curve, although that file currently says the cycle is open at five passes and explicitly promises the closing figures in a later revision | The artifact selected as the durable record for all four pre-rule Gate-A cycles remains knowingly incomplete after execution, defeating the observability reason for moving records out of the commit body | Add a Gate-A-close or pre-execution step that mechanically extracts the final Plan C counts, replaces the provisional open-cycle passage, verifies the totals, and lands that report update without reconstructing records into the implementation body +MINOR | high | lines 51-56 and 1171-1185 | The plan consistently calls the Gate-B cycle old-rule and records `cycle none (pre-rule)`, but the settled field report says the implementation commit carries the one cycle "the new rules actually bind" | The durable narrative and the executable plan assign different activation states to the same cycle, obscuring whether its new-form records are mandatory by activation or are evidence required by this story | State in both places that the cycle is pre-rule and writes the new-form records natively because this branch's acceptance and evidence criteria require them, not because the new operational rules bind the cycle +END OF FINDINGS (19 total) diff --git a/.context/codex-reviews/gate-a-plan-planc-pass-7.md b/.context/codex-reviews/gate-a-plan-planc-pass-7.md new file mode 100644 index 0000000..ac3dd18 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-planc-pass-7.md @@ -0,0 +1,30 @@ +BLOCKER | high | Plan C lines 980-1017 and 1412-1428; approved spec lines 377-386 and 513-519 | Tasks 19 and 20 turn the plan-local deterministic discriminator mentioned in spec section 8 into a general shipped slot production, but the approved slot rules reserve bare names for no-nonce cycles and contain no such production, and Task 21 updates only sections 1 and 8 | This changes specified operational behaviour beyond Plan C's stated section 7 and 8 rollout scope while leaving the source-of-truth slot section stale, so the prompt and approved spec cannot both govern execution | Either keep the discriminator as an explicit Plan-C-local safety procedure or approve and update the spec's slot rules, story criterion, accounting, and both prompt copies in the same change +BLOCKER | high | Plan C lines 51-59, 980-1017, and 1282-1302 | The plan says the Gate-B cycle finishes under the old rules, yet justifies its `rle` slot names solely by the new discriminator production Tasks 19 and 20 will ship; that production cannot bind this pre-rule cycle, while every future cycle governed by it must have a nonce | The exception is unreachable under the activation rule and does not make the current non-bare slots valid, so Task 23 still cannot follow both the old file-first protocol and the plan | Define an approved plan-local exception under the old rules with its own preservation and ownership procedure, or change the settled activation semantics and re-review every affected source +MAJOR | high | Plan C lines 1081-1094 and 1134-1151; Plan B lines 360-364 | The parser-existence sweep updates the spec's two statements but misses the two shipped prompt copies Plan B leaves saying the provenance form has no informal variant "because the deferred metrics work parses it" | The finished product will still claim a parser exists while the CHANGELOG and revised spec say none ships, violating the goal to correct every falsified user-facing sentence and invariant 11 | Add paired Plan-C tasks replacing that sentence in root `CLAUDE.md` and the scaffolded template with accurate future-facing P8 wording, with old-condition accounting and parity checks +MAJOR | high | Plan C lines 1007-1017 | The new `` production defines neither an allowed character set, length bound, derivation algorithm, nor canonical mapping from a cycle to a discriminator; `rle` appears only as this plan's example | A hostile or careless discriminator can introduce separators or path components, and independent executors can choose different paths while each believes it followed the rule | Pin a path-safe discriminator grammar and deterministic derivation procedure, including validation and a stop for unrepresentable inputs +MAJOR | high | Plan C lines 1012-1017 and 1295-1298 | The text admits two cycles can compute the same deterministic discriminator, but deletion is authorized by matching that discriminator and the Task-23 ownership check is neither specified nor atomic | A sibling can pass the check and begin writing after it, or an existing same-discriminator target can be deleted as "own", causing irrecoverable findings loss or a valid-looking file from the wrong invocation | Require an atomic ownership mechanism for the exact discriminator namespace and refuse deletion without proven ownership, or use a collision-resistant cycle-scoped namespace that the settled identity rules permit +MINOR | high | Plan C lines 1007-1017 and 1047-1052; Plan B lines 226-239 | The discriminator addition names only Gate-A and Gate-B findings-file forms, while Plan B makes the advisory resume record a cycle slot governed by the same naming and collision rules | A pre-rule cycle encountering an existing bare resume record still has no defined safe filename or retirement rule, so the claimed slot repair is incomplete | Define discriminator forms and ownership behaviour for all three resume-record names as well as findings slots, or explicitly prove the existing "exactly as findings slots" wording supplies an unambiguous spelling +MAJOR | high | Plan C lines 120-124, 901-904, and 948-950 | Tasks 17 and 18 say a skipped cycle runs the battery but omit the battery result from the commit-body list, despite the accounting claiming the skip reason, battery result, and evidence duties are all kept | An unprofiled skipped cycle has no mode-derived evidence entry, so these two user-facing summaries still let it close without the battery result that both governing prompt copies require | Add the battery result explicitly to both documentation lists, preserving the provenance line, skip record, and per-profile evidence entries +MAJOR | high | Plan C lines 70-84 and Tasks 1-21 | The only advertised preflight is a template containing literal `` and `` placeholders; it is not executable as written, and no task instantiates it before editing | A fresh or resumed executor lacks the promised task-local decision procedure and can duplicate an insertion, retry a replacement against missing OLD text, or fail on shell redirection from a nonexistent `FILE` | Put a concrete executable preflight with the real file, OLD anchor, and NEW sentinel before every edit +MAJOR | high | Plan C lines 70-94 | Even if instantiated, the count-one branch exits before the required whole-block read and before determining whether the edit is only in the worktree, staged, or already in the WIP commit | A half-applied block is accepted, and interruption after editing but before amend causes a rerun to skip the commit step, leaving changes stranded until the final clean-status check | Make the preflight validate the complete block and OLD-text disposition, distinguish HEAD/index/worktree states, and complete the pending stage-and-amend when the edit is correct but uncommitted +MAJOR | high | Plan C lines 1179-1182 | The mandatory battery task does not provide the exact command, output capture, or success assertion; it only tells the executor to adapt the `AGENTS.md` chain by replacing its base argument | An executor can run the documented `main` form, compare the wrong range or pass trivially, yet still record battery evidence | Paste the full fresh-shell command with `sh scripts/check-version-bump.sh HEAD^`, require exit 0 for every component, and record the observed result used by the evidence entry +MAJOR | high | Plan C lines 1170-1250, 1304-1305, and 1331-1341 | Task 22 never constructs or stores the "current evidence entry", the grammar feature matrix, parity result, or checklist rows that Task 23 must quote verbatim and Task 24 must revalidate | Task-by-task subagents and interrupted sessions have no authoritative bytes to resume from, so calls and the closing commit can silently carry different evidence claims | Define an exact evidence-entry format and a private cycle-local durable artifact written atomically, validate it, and require every call and the close to read the same recorded bytes +MAJOR | high | Plan C lines 1199-1212 and 1326-1329 | The before-state of `.context/codex-gate.floor` has no storage location, integrity check, or resumption rule despite being needed after the whole Gate-B loop | After context loss or a different subagent takes over, the after-state comparison can only rely on recollection, so inherited obligation M9 is not demonstrable | Persist one cycle-local snapshot containing safe byte encoding, digest, lstat type, value class, and capture metadata, then make Task 24 locate and validate that snapshot before comparison +MAJOR | medium | Plan C lines 1203-1210 and 1326-1329 | The knob observation does not define an atomic capture and treats a symlink to a readable file as target bytes without recording the link target bytes; unusable non-regular paths are compared only by broad type | Retargeting a symlink, replacing a directory, or racing the type and byte reads can pass as unchanged or produce internally inconsistent evidence | Use lstat-style classification, record symlink target bytes and relevant non-regular identity, snapshot readable content once before hashing and classifying it, and reject any path that changes during capture +MAJOR | high | Plan C lines 1109-1123 and 1214-1241 | The grammar check supplies no actual EREs, constructed strings, extraction commands, or fail-capable assertions, while the pinned "grammars" contain prose productions and constraints that cannot be passed directly to `grep -E` | The required evidence step is not executable or observable, and an implementation can claim coverage with an unanchored or incomplete regex that never tested the pinned forms | Pin the exact anchored EREs and fixture table in the plan, with commands whose expected statuses distinguish every valid and invalid case +MAJOR | high | Plan C lines 1117-1123 and 1231-1238 | The four supplemental comparisons do not cover the curve rule that per-pass model keys occur exactly once and in ascending order; set difference against `` does not detect a duplicate or reordered key | A malformed curve can pass every named check while violating a pinned required property | Add explicit duplicate-key and ordering checks, each with passing and failing fixtures, in addition to set equality with the expanded pass specification +MAJOR | high | Plan C lines 1219-1229 and 1231-1238 | Using the same constructed nonce in the provenance and curve instances demonstrates only the positive case; no comparison or failing different-nonce pair verifies the cross-record agreement the plan says it exercises | Broken wiring that never compares the two cycle fields will still report the attribution requirement covered | Add an explicit equality comparison over extracted cycle fields plus a mismatched-nonce pair that must fail +MINOR | high | Plan C lines 1117-1123 | The revised spec first says "Every constraint" a match cannot decide is checked, then says the resulting list is not proven complete | The universal claim outruns the stated evidence and repeats the gate-overclaim class the edit is meant to correct | Say "Every identified constraint" or otherwise bound the claim to the listed axes and keep the incompleteness statement +MAJOR | high | Plan C lines 1243-1250 and 1444-1446 | Parity and the three twelve-item prompt-standard passes have no extraction boundaries, expected sole-difference assertion, named status artifact, or procedure that fails on a wrong row | Invariant 11 can be marked complete after an informal read, and later readers cannot distinguish a performed 36-row review from an assertion that it happened | Specify the compared rule blocks, a fail-capable parity check, and a persisted 12-by-3 status matrix consumed by the evidence entry +MAJOR | high | Plan C lines 1321-1324 | The risk-path verification derives from the singular `Story:` header of "the artifact" even though Gate B reviews the combined diff and Plan A requires reading every contributing plan header, comparing expected artifact sets, and deriving from their union | A missing, malformed, or disagreeing header can be ignored while the provenance line is still blessed; the fact all current paths intend to name one story does not verify the required wiring | Re-read the spec and all three plan headers, validate each expected set, compare the sets as the shipped rule requires, and derive the Gate-B union before checking the line +MAJOR | high | Plan C lines 34-45, 1359-1363, and 1448-1451 | The inherited `mktemp` closing-body obligation is declared moot, but the plan still must compose a multiline body and gives neither a safe body file nor an exact close command | Quoting, truncation, partial assembly, or an ad hoc fixed temporary path can corrupt the irreversible close, and the inherited obligation is dropped rather than discharged | Build the complete body in a `mktemp` file with a cleanup trap, validate every required record, close with `git commit --amend -F`, and preserve the file if the amend fails +MAJOR | high | Plan C lines 1335-1342 | The supposedly complete closing-body list includes an undefined successor-story "decline record" but omits current section 5 human-exception records when such a decision occurred | The executor cannot determine the undefined record's form and can omit a real governing record because the list explicitly says nothing else belongs | Remove the unshipped decline-record dependency and account for every conditional record the old rules can actually produce, including human exceptions, without claiming completeness until the sources are reconciled +MAJOR | high | Plan C lines 1367-1406; field report lines 68-86 | Task 25 replaces only the provisional curve passage, leaving the analysis that says Plan C never reached zero Blockers and stops its revision table and mandatory-stop narrative at pass 5 | A final clean pass necessarily falsifies that prose, so the durable report remains knowingly contradictory even after its curve is updated | Re-read and update every Plan-C analytical statement, revision row, stop count, and conclusion affected by passes 6 through closure, not just the three count rows +MAJOR | high | Plan C lines 1395-1403 | The assertion labelled as checking both removal and presence checks only that the provisional phrase has count zero; deleting the passage without inserting any closed record passes | The durable evidence file can be truncated or left without Plan C's curve while Task 25 reports success | Assert the exact closed heading, pass range, and all three extracted series once each, and cross-check their lengths and final clean values +MINOR | high | Plan C lines 1382-1393 | The extraction loop stops at the first missing pass file and counts severity prefixes without validating each file's terminator, declared total, or line grammar | A gap silently hides later passes and a stale or malformed file can seed authoritative-looking durable counts | Enumerate the complete matching set, reject gaps and unexpected suffixes, validate every file by the section 5 protocol, and compare the counted total to its terminator before extraction +MINOR | high | Plan C lines 70-94 and 1367-1406 | Task 25 is covered by the blanket skip-if-applied claim but has no applied-state sentinel: on rerun the provisional passage is absent, its sole assertion passes, and the requested replacement has no target | A resumed execution can attempt an empty commit or report the task complete without validating the existing closed record | Add a preflight that validates and skips an already-correct closed block, applies only from the unique provisional state, and stops on every mixed state +MINOR | high | Plan C lines 348-360 | The satisfied-message replacement lists three withheld result shapes without limiting the statement to successfully routed PostToolUse results or saying the list is incomplete; the hook also does not count unroutable payloads, unattributable tools, and failures delivered outside that route | The revised user-facing explanation still overstates exactly what the counter represents, the claim this revision says it swept | Qualify the three as the recognized discarded classes after successful routing and state that other unroutable or undelivered calls can also remain uncounted +MINOR | medium | Plan C lines 358-360 | The text states the hook keeps the last fingerprint and consecutive streak without the hook's best-effort persistence limit; independent state-file and counter-write failures can reset, preserve, or distort that streak | In an unwritable or partially writable `.context`, the displayed fresh value is not guaranteed to be the true consecutive call count | Describe `` as the hook's best-effort stored streak and point readers to state-write failures as a cause of divergence +MINOR | high | Plan C lines 11-12 and 1071-1166; spec line 3 | Task 21 edits two approved-spec sections but leaves the spec header at revision 36 and the plan continuing to identify revision 36 | The revision identifier no longer names stable bytes, weakening the provenance used by the closed Gate-A spec cycle | Either avoid editing the approved spec by implementing its existing parse requirement, or advance and consistently reference the spec revision with an explicit review/accounting decision +NIT | high | Plan C lines 57-59 | The global constraint says Task 13 reads `.context/codex-gate.floor`, but Task 13 edits the model-recording sentence and the actual knob observation is Task 22 item 4a | The wrong cross-reference sends an executor to a task that cannot discharge the stated evidence obligation | Change the reference from Task 13 to Task 22 item 4a and Task 24 item 4b +END OF FINDINGS (29 total) diff --git a/.context/codex-reviews/gate-a-plan-planc-passes-1-2-dispositions.md b/.context/codex-reviews/gate-a-plan-planc-passes-1-2-dispositions.md new file mode 100644 index 0000000..f78d18e --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-planc-passes-1-2-dispositions.md @@ -0,0 +1,101 @@ +# Gate A — Plan C cycle — passes 1-2 dispositions, and a stop + +Advisory human note. Not a findings file; participates in no pass validation. + +## Result + +| pass | findings | Blockers | Majors | B+M | +|---|---|---|---|---| +| 1 | 18 | 5 | 11 | 16 | +| 2 | 20 | 4 | 13 | **17** | + +**Fifteen findings were fixed between them and the count went up.** Three of pass 2's four +Blockers are consequences of those fixes. + +## Why this stops here + +**The routed classification question was never answered, and pass 2 settles it as a practical +matter.** After pass 1 I flagged that instrument-versus-product was ambiguous for Plan C — it is +largely a *procedure*, so its commands arguably are the product — and proceeded on the strict +reading, which put instrument at ~10%. On the loose reading it was 37%. + +Pass 2 puts it at roughly **40% on the loose reading** and, more to the point, supplies the +evidence the classification was standing in for: + +- **`Task 15`'s base recovery is circular.** It defines the base as the current tip's parent and + then proves exactly one commit sits above that parent — which is true by construction. **A check + that cannot fail, for the sixth time in this cycle.** I wrote it one round after fixing the + fifth. +- The differential check's three greps **print counts and never compare them**, so an unexpected + value passes silently. +- The parse check **supplies no parser**, so the thing it validates is a claim. +- The knob state machine tests `! -e` before path type, so **a broken symlink classifies as + absent** — the exact case it was added to catch. +- The closing-body validation **prints counts rather than asserting them**. + +That is not a classification artifact. **The instrument is where the errors are**, again, and the +standing rule was written to fire on sight for precisely this. + +## Two findings are entangled with the open contract question + +The routed question — which commit carries each cycle's provenance line and curve — is still +open, and pass 2 shows it reaches further than Task 18 Step 2: + +- **BLOCKER 3:** the four already-closed Gate-A cycles have **no contemporaneous knob evidence** + either, for the same reason they have no records — the procedure did not exist when they closed. + Whatever answers the records question answers this too. +- **MAJOR 11:** Task 18's executable path is **already hard-coded to aggregate**, so the plan + quietly presumes one answer while the note above it says the question is open. + +**Repairing around an open contract question is how the answer gets made by drafting.** That is +the failure §5's stop rule exists to prevent, and it applies twice over here. + +## The finding I would have missed + +**BLOCKER 2 is the one worth carrying forward regardless of what happens next.** Task 14 repairs +`process-pr-review.md`'s skipped-cycle duties — and **both §5 copies still say the old thing** +(`CLAUDE.md:424-433`, `workflow-init.md:603-612`). I fixed one of three surfaces and reported it +as the fix. The completeness sweep found the first surface at pass 1 and the remaining two at +pass 2, which is the same lesson arriving twice: **a statement's other homes are not found by +fixing the one you noticed.** + +## Disposition + +**Nothing repaired this round.** All 17 Blocker/Major carried open, plus 3 Minor. + +**What a resumption needs, in order:** the contract answer, which unblocks two findings and +settles Task 18's shape; then a decision on whether Plan C's executable steps carry checks at all, +given that six of them have now been written so they cannot fail; then the remaining fifteen, +which are individually small and specific. + +**What is not in doubt:** Plans A and B are closed clean and committed, the falsified-sentence +catalogue is now three sites larger than the one built before them, and every finding here is +recorded rather than lost. + +## Delivery record — corrected, because the label matters + +The sparring session reports that no routed question reached it and reads this as the +**announce-without-sending** pattern in a new spot. **That is not what happened, and the +distinction is only visible from this side**, so it is recorded here rather than argued about: + +- The routing message **was sent** — `SendMessage` returned success, `msg_id` + `8e0f4438-e521-44af-a7ff-e94933afa4e5`. +- A **system notice then reported it dropped at the recipient's inbox and NOT delivered** — "a + relay loop between sessions was cut" — and instructed: treat as unsent, **do not resend now**, + fold into one later message **after finishing other work**, never retry in a loop. +- I followed that: ran pass 2, recorded the stop, and folded everything into one message. + +So the failure was **delivery, not omission**. From the peer's side the two are indistinguishable, +which is exactly why it was reported as the known pattern — a reasonable reading of the evidence +available to it. + +**What the peer is right about, and it survives the correction:** the commit body at `b280099` +says "one contract question routed" as settled fact, written **before delivery was confirmed**. +The routing was real and the claim was true when written, but the body asserts an outcome whose +confirmation had not arrived and could not have. **A commit body should record what this side did +— "routed, delivery unconfirmed" — not what the other side received.** That is a genuine lesson +and it is not the same lesson as announce-without-sending. + +**Occurrence count is unchanged**: the four announce-then-idle occurrences stand at four. Adding a +fifth here would corrupt a record the sparring session is using to track a real failure of mine, +by putting a delivery fault in the same column as an omission. diff --git a/.context/codex-reviews/gate-a-plan-planc1-pass-1-dispositions.md b/.context/codex-reviews/gate-a-plan-planc1-pass-1-dispositions.md new file mode 100644 index 0000000..62f63b8 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-planc1-pass-1-dispositions.md @@ -0,0 +1,66 @@ +# Plan C1 — Gate-A pass 1 dispositions (14 findings: 0 B, 8 MAJOR, 3 MINOR, 3 NIT) + +Opening: 8 B+M, against Daniel's prediction of <=6 and a route-back threshold of 12. Below the +threshold, so the loop continued. For comparison, Plan C opened at 16 B+M and Plan B at 9. + +## MAJOR — six fixed in revision 2 + +- **Singular "the cited story's profile"** (lines 31-38, 121-125, 170-175, 371-375). VALID and + verified against Plan A: the floor derives from the cited **set** — 1 only if every cited story + is level 0, 3 if no story is cited or any is unprofiled, and an unresolvable profile **stops**. + The singular phrasing is false on three of four states. Every occurrence now uses Plan A's own + words, *the profile and the cited set*, and the plan states the four states once. +- **An eighth site was missed** (`docs/getting-started.md:35`, the `1/3` counter example). VALID + and verified live. It teaches a fixed denominator and would have sat one line from Task 2's + derived-floor sentence, contradicting it on the same screen. Now Task 3. Explicitly NOT C2's: + C2 owns the *satisfied* message at line 58; this is a below-floor message, the same claim as + Task 5. +- **Preflight checked only the new sentinel** (lines 64-68, 105-138, ...). VALID, and the resume + hole was the real cost: an interruption between Replace and Amend left the edit stranded in the + worktree while the task reported itself done. Every preflight is now a state machine over both + texts, and the already-replaced branch checks whether the edit reached the WIP before skipping. +- **No amend verified HEAD is the WIP** (lines 54-62 and every amend step). VALID. Each amend now + checks `git log -1 --pretty=%s` against the exact subject first. +- **Explicit paths do not scope content inside the file** (line 61 and every amend). VALID. Each + amend step reads `git diff -- ` before staging, and the plan says why explicit paths are + not a scope guard here. +- **Asserts counted substrings, not the replacement** (lines 127-132, ...). VALID, and the example + was concrete: Task 7 could have dropped half its replacement and still gone green. Six of the + eight replacements are whole lines and now use `grep -cxF` — exact whole-line equality. Tasks 2 + and 3 land mid-line, cannot use it, and each says so. + +## MAJOR — two routed, not absorbed + +- **`docs/coding-workflow.md` axes and model-recording sentences are unassigned.** +- **The skipped-cycle duty claim across five files is unassigned.** + +Both are correct, and both are correct about the **close**, not about C1: the story requires every +falsified shipped sentence corrected in the same change, so C3 cannot close while these have no +owner. They were already named in C1's scope table as UNASSIGNED; pass 1 confirms that naming them +is not the same as solving them. **Routed to Daniel.** Absorbing them into C1 would rebuild Plan C, +which is the failure the split exists to prevent. C1's scope table now says explicitly that they +block C3's close rather than this plan. + +## MINOR / NIT — fixed + +- **Accounting row 6/7 was written from a summary, not the live sentence.** VALID and the sharper + finding of the two Minors: the live sentence asserts profile-supplies-eligibility, battery-still- + owed, and floor-unchanged. The row claimed a baseline-questions clause the original never made — + and the draft replacement had *added* that clause. Row rewritten from the live text; the added + clause is gone, so the replacement now changes exactly the one false clause. +- **Task 9 claimed the knob is described "identically" in both files.** It is not, and should not + be — README names the source, getting-started does not. Reworded to the property actually meant. +- **No rollback checkpoint.** Now a Global Constraint: record the WIP SHA before Task 1. +- **NIT, assert contract wording** — the constraint said "anything but 1 fails" while every + negative half wants 0. Stated correctly now. +- **NIT, hook-reset overclaim** — "discarding the cycle's passes" was too broad. The Bash commit + branch resets **Gate-B** state, regardless of whether the commit succeeded, because it cannot + observe exit status; Gate-A state is reset by skill events. Corrected. +- **NIT, Task 9 does not always amend.** Stated: Tasks 1-8 always amend, Task 9 only after a fix. + +## Found while fixing, not from the pass + +The generic knob regex written for Task 9 **errored on this machine's `grep`** ("exceeds +complexity limits", ugrep). A command that cannot run reports nothing and reads as a pass, so it +was replaced with fixed-string checks and the reason recorded in the plan. Both Task 9 greps were +then executed live against the current files and exit as intended. diff --git a/.context/codex-reviews/gate-a-plan-planc1-pass-1.md b/.context/codex-reviews/gate-a-plan-planc1-pass-1.md new file mode 100644 index 0000000..5a3f9ce --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-planc1-pass-1.md @@ -0,0 +1,15 @@ +MAJOR | high | lines 31-38, 121-125, 170-175, 371-375 | The plan's defining claim and three NEW blocks reduce the source to a singular "cited story's profile", "its profile", or "the profile"; in Task 2, "its" grammatically points at the revised spec, which has no profile | Plan A derives one value from the current governing cited-story set, uses unanimity for several stories, defaults to 3 for no citation or an unprofiled member, and stops on a present-but-unresolvable profile, so the proposed user text is false or ambiguous on supported states | Say that the floor is derived under §5 from the current governing cited-story set and its profiles; where this summary needs the edge cases, name the no-story/unprofiled default and unresolvable stop explicitly +MAJOR | high | lines 36-38, 143-151, 445-461 | The claimed seven-site sweep omits the separate sentence at live `docs/getting-started.md:35`, which quotes `Gate A below floor (1/3)` and calls it "the counter" without saying that 3 is the hook threshold rather than the cycle's obligation; Task 8 does not match `1/3` | After Task 2 that sentence still teaches a fixed denominator beside a derived-floor sentence, and unlike the `3/3` satisfied example explicitly assigned to C2, this below-floor example has no owner | Add an eighth C1 replacement that calls it the hook's reminder-threshold counter, or assign this exact site to C2, and extend the completion sweep to cover numeric ratios with explicit exclusions for C2-owned sites +MAJOR | high | lines 45, 49-50, 489-491 | The `docs/coding-workflow.md` axes and model-recording corrections required by spec §7 are explicitly left unassigned even though C3 is still supposed to close the combined change | The story requires every falsified shipped sentence to be corrected in the same change, so the combined A/B/C1/C2/C3 close cannot satisfy acceptance while this row has no owner | Assign these exact sites to C2, C3, or a named additional sub-plan and make C3's close depend on that plan completing +MAJOR | high | lines 46, 49-50, 489-491 | The skipped-cycle duty claim across five shipped copies is explicitly left unassigned | The plan itself says splitting that claim by file recreates the defect, yet the announced combined close has no artifact responsible for correcting all copies, leaving a known cross-copy contradiction at closure | Create one named plan that owns the claim across all five sites and add it as a prerequisite of C3's evidence and closing cycle +MAJOR | high | lines 64-68, 105-138, 153-188, 407-440 | Every preflight inspects only the NEW sentinel and treats one occurrence as "ALREADY APPLIED — skip"; it neither proves the OLD block is still present exactly once nor distinguishes a committed task from an edit completed just before interruption | On resume after Replace but before Amend, the task exits successfully without staging or amending; Task 1 and Task 7 have no later stage of their file to rescue that state, while missing or drifted OLD anchors can be misclassified as ready to apply | Make each preflight a state machine over both exact OLD and NEW blocks: OLD once and NEW zero means apply, OLD zero and NEW once means verify/stage/amend if not already in the WIP, and every mixed, missing, or duplicate state stops +MAJOR | high | lines 54-62, 134-139, 184-189, 232-237, 284-289, 332-337, 384-389, 436-440 | None of the repeated `git commit --amend` steps verifies that HEAD is the intended `WIP: review-loop economics` commit or even a WIP commit at all | If C1 is started out of order, resumed in the wrong checkout, or the WIP was already closed, the first amend rewrites an unrelated commit and the hook treats the command according to a cycle state the plan never established | Add one mandatory global preflight that records and verifies the expected WIP commit identity and exact subject before any edit, then re-check that identity or an explicit successor identity before every amend +MAJOR | high | lines 61, 134-139, 184-189, 232-237, 284-289, 332-337, 384-389, 436-440 | Explicit path staging still stages every concurrent or pre-existing change inside that file, and the plan never inspects the scoped worktree or index diff before `git add` or before amend | A careless or concurrent actor can have unrelated README/getting-started edits silently absorbed into the shared WIP; the per-task sentinel can remain correct while extra content is committed | Before editing and again immediately before each amend, inspect the scoped worktree and index diffs and stop on content outside the exact planned replacement; after staging, verify the staged diff is the intended cumulative C1 diff +MAJOR | high | lines 127-132, 177-182, 225-230, 277-282, 325-330, 377-382, 429-434, 463-469, 481-483 | The asserts count short sentinel substrings rather than the exact NEW blocks, despite the self-review claiming the new text is present exactly once | A malformed or contradictory replacement can delete the OLD phrase, place the sentinel once, and pass; for example Task 6 can omit the baseline-questions clause entirely and still go green, so these checks do not falsify the edit they name | Assert the complete expected replacement block at each site, its exact count, the OLD block's absence, and the intended surrounding anchor; make Task 8 rerun these exact checks after any repair +MINOR | high | lines 76-84, 341-375 | Old-conditions row 6 says the replaced sentence asserted that the axes never subtract baseline questions, but the live sentence actually asserts that the profile supplies only eligibility, the battery remains owed, and Gate A's floor is fixed | The NEW text happens to retain the battery clause, but the accounting does not account for that real condition and attributes a new baseline-question claim to the old prose, defeating the invariant's audit trail | Rewrite row 6 from the full live sentence: keep profile-only eligibility and battery owed, deliberately drop the fixed-floor clause, and mark the baseline-question sentence as a new clarification rather than a kept condition +MINOR | high | lines 423-469 | Task 8 says the knob is described "identically" in both files, but the proposed rows are not identical: README explicitly says it does not change the obligated floor and names the source, while getting-started only says it moves the reminder threshold; the two commands merely count a shared substring and never test the claimed non-binding semantics | The check can pass if either line also says in different words that the knob changes the obligated floor, and its success cannot support the prose claim made above it | Change "identically" to the narrower semantic property actually intended and assert each complete expected line plus a site-specific check that no positive floor-binding claim remains +MINOR | medium | lines 54-68, 134-139, 471-473 | The plan supplies no rollback checkpoint or scoped undo procedure before repeatedly replacing the same shared WIP commit | A half-executed C1 can be recovered through reflog knowledge, but an executor has no recorded pre-C1 WIP identity or documented way to remove only C1 while preserving completed Plan A and Plan B work | Record the pre-C1 WIP SHA before Task 1 and document a non-destructive, scoped rollback or forward-repair procedure that preserves the earlier plans +NIT | high | lines 64-65 | The global constraint says the assert treats anything but 1 as failure, but every negative half deliberately succeeds at 0 and fails at any nonzero count | This contradicts the commands and makes the preflight/assert contract harder to reason about during recovery | State that the positive sentinel must occur once and the OLD sentinel must occur zero times +NIT | high | lines 56-60 | The warning says a non-`-m` amend makes the hook reset and discard "the cycle's passes", while the hook's Bash commit branch removes only Gate-B fingerprint/pass state; Gate-A state is reset by skill events, and the Bash branch resets even when the commit actually failed because it cannot observe exit status | The overbroad wording can lead an executor to infer that Gate-A records were discarded or that reset proves the amend succeeded | Say specifically that an amend not recognized as WIP is treated as a Gate-B cycle boundary and its Gate-B state is reset regardless of commit success +NIT | high | lines 56-57, 445-473 | The global constraint says every task amends the WIP, but Task 8 explicitly amends only if a check led to a fix and otherwise stages nothing | The internal mismatch is small but obscures the terminal state of a clean Task 8 run | Say Tasks 1-7 always amend and Task 8 amends only after a repaired failure, followed by a rerun of its checks +END OF FINDINGS (14 total) diff --git a/.context/codex-reviews/gate-a-plan-planc1-pass-2.md b/.context/codex-reviews/gate-a-plan-planc1-pass-2.md new file mode 100644 index 0000000..63de16e --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-planc1-pass-2.md @@ -0,0 +1,13 @@ +BLOCKER | high | lines 653-663 | Task 9's first command exits 1 for every match, but the plan says the C2-owned `3/3 cycle` line is expected to remain when C1 finishes | C1 is explicitly first and C2 owns that line, so the completeness task cannot reach its desired state when this plan is executed in the stated order | Make the check accept exactly the one full C2-owned line and fail on any additional hit, or explicitly defer this check until after C2 and change the stated C1 completion boundary +MAJOR | high | lines 36-40 and 130-133 | The claimed four-state derivation omits the resolved nonzero-profile case that yields floor 3 and also omits Plan A's no-value stops for disagreeing governing headers and an unreadable or malformed `Story:` header | The plan's own explanation is not a complete or faithful account of the byte-frozen rule, so an executor or reviewer can validate the replacements against a materially incomplete decision procedure | State the rule exactly: floor 1 iff the cited set is nonempty and every member is profiled, resolvable, and level 0; floor 3 for no story, any unprofiled member, or any resolved level-1/2 member; stop for present-unresolvable profiles and governing-header read, parse, or disagreement failures +MINOR | high | lines 31-34 | The headline says “Three passes is not the floor,” although 3 remains the derived floor for this high-risk story and for every other non-level-0, uncited, or unprofiled case | The central summary contradicts the settled predicate even though its intended point is only that 3 is no longer a universal constant | Change it to “Three passes is not a universal or constant floor” +NIT | high | lines 130-133 | The Task 1 rationale says “level 1” where the settled rule says “floor 1” | Risk/security levels and pass floors are different types, and this wording needlessly conflates them in the paragraph meant to teach the derivation | Replace “level 1” with “floor 1” and include the missing higher-profile floor-3 arm +MAJOR | high | lines 200-235, 276-311, and 705-708 | Tasks 2 and 3 treat short substrings as proof that their complete multi-line replacements landed; each preflight can classify a partial bad edit as complete, and each assert can pass after load-bearing clauses are dropped | Task 2 can lose the zero-finding early exit, and Task 3 can lose “hook's own reminder threshold” or “not against the floor §5 obliges,” recreating exactly the dropped-condition failure the plan says it prevents | Assert the complete replacement blocks across the line break, or fix task order and assert every resulting whole line plus every preserved clause exactly +MAJOR | medium | lines 42-44, 648-696, and live `docs/getting-started.md:6-8` | The claimed complete eight-site sweep omits the introductory statement that the hook “reminds ... when a gate isn't satisfied,” which still presents reminder state as gate-satisfaction state and is not matched by Task 9 | Under a floor-1 cycle the hook can emit a below-threshold reminder after the owed floor is met, while after enough counted calls it can report threshold satisfaction despite unresolved findings; leaving this sentence defeats the threshold-versus-obligation clarification | Add this as a ninth owned site and rewrite it to say the hook reports advisory counter and fingerprint state while §5 determines whether a gate is satisfied +MAJOR | high | lines 75-76 and each amend block, first at 171-177 | The scoped read uses `git diff -- `, which excludes staged changes, while `git commit --amend` commits the entire pre-existing index, including staged content in the target file and staged paths outside it | A partial, concurrent, or careless run can silently absorb unrelated work even though the plan claims explicit path staging keeps other files out | Before every amend inspect `git diff HEAD -- ` and the full cached path set, refuse unexpected staged paths or content, then stage the target and re-check the exact index that will be committed +MAJOR | medium | lines 72-74 and each HEAD check, first at 173-177 | Matching only the subject `WIP: review-loop economics` does not establish that HEAD is this cycle's WIP; an unrelated commit or checkout can use the same subject and pass the guard | The subsequent amend rewrites that commit, so the advertised wrong-checkout and out-of-order protection is weaker than stated | Record the original WIP's parent and repository/worktree identity, then require each amend candidate to have that same parent and expected branch or worktree in addition to the subject check +MAJOR | high | lines 86-88 | The rollback is neither instantiated nor correct as written: no task records the SHA, and an unspecified `git reset ` defaults to mixed reset and leaves the C1 worktree edits in place, while a hard reset could destroy concurrent or pre-existing work | A half-executed plan has no safe, observable undo path despite the high-risk rollback and data-loss lens | Add a mandatory checkpoint step that durably records the SHA and clean index/worktree assumptions, plus an exact path-scoped rollback procedure that first detects unexpected changes and verifies the two files and index after rollback +MAJOR | medium | each amend block, first at lines 173-177 | The amend commands do not fail fast and perform no post-amend verification; a failed `git diff` or `git add`, or a concurrent edit after the pre-assert, can still be followed by a successful amend of stale index content | The task can appear successful while its replacement is absent from HEAD or while later content was committed without being asserted | Guard every command with explicit failure handling and, after the amend, run the exact old/new assertions against the blobs in HEAD plus clean worktree/index checks +MINOR | medium | lines 135-149, 200-214, 276-290, 339-353, 404-418, 465-479, 532-546, and 597-611 | The preflights describe duplicate detection but `grep -c` counts matching lines, not occurrences, so two copies of a substring on one line are reported as one | A hostile or malformed partial edit can enter the “apply” or “already applied” branch instead of the duplicate stop state | Count occurrences rather than matching lines for substring sentinels, and prefer exact full-line or full-block state checks wherever possible +NIT | high | lines 263-270 | Task 3 says its below-floor message carries the same threshold-versus-floor claim as Task 4, but Task 4 is the plan-loop sentence; the corresponding below-floor threshold task is Task 5 | The incorrect cross-reference slows review and can send an executor to the wrong rationale | Replace “Task 4” with “Task 5” +END OF FINDINGS (12 total) diff --git a/.context/codex-reviews/gate-a-plan-planc1-pass-3.md b/.context/codex-reviews/gate-a-plan-planc1-pass-3.md new file mode 100644 index 0000000..fcdbe0a --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-planc1-pass-3.md @@ -0,0 +1,18 @@ +MAJOR | high | plan:16-18,38-39 | The execution-time directive says the plan states no pass count derived from the story, but the headline later states that three is this story's derived floor | If the story header changes before execution, the plan contains a stale derived value while telling the executor none exists | Either remove the story-specific pass count or revise the directive to say it is explanatory only and that execution must re-read and recompute from the story header +MAJOR | high | plan:44-52 | The purported exact rule gives floor 3 for any unprofiled member or any resolvable member above level 0 while separately saying any unresolvable member stops, so a mixed cited set containing both satisfies incompatible outcomes | Plan A and the spec make an unresolvable member stop the cycle rather than default, regardless of other members; ambiguous precedence can turn a required stop into floor 3 | State the stop cases first and qualify every floor-3 arm with "when no stop condition applies," matching Plan A's stop-before-default semantics +MINOR | high | plan:41-52 | The text calls itself a byte-frozen quotation of Plan A, but it is a three-bullet rewrite rather than Plan A's actual text at lines 140-183 | The claimed safeguard against another faulty summary is absent, so later readers cannot distinguish faithful wording from a new paraphrase | Paste the relevant Plan A passage verbatim or label this as a restatement and remove "quoted," "exactly," and "byte-frozen" +MAJOR | high | plan:67,70-73 | The two falsified `docs/coding-workflow.md` sentences are explicitly UNASSIGNED | The governing story requires every falsified shipped sentence to be corrected in the same combined change, so the A+B+C rollout cannot close as planned | Assign these sites to a named reviewed sub-plan before C1 execution, with exact anchors, replacements, assertions, and combined-close ownership +MAJOR | high | plan:68,70-73 | The skipped-cycle duty claim across five files is explicitly UNASSIGNED and the five sites are not enumerated | The combined close is knowingly missing a story requirement, and an unnamed site list cannot be checked for completeness or handed off safely | Enumerate all five exact sites and assign them to a reviewed sub-plan before implementation +MAJOR | high | plan:85-87,195-203 | Branch name, WIP subject, and a non-WIP parent do not establish commit identity statelessly; a different or concurrently amended WIP with the same subject passes all three checks | A stale executor can stage and amend into the wrong snapshot without any check noticing, especially because every amend rewrites the WIP SHA | Record the expected HEAD SHA for each task, re-check it immediately before staging and committing, update it after each amend, and require exclusive worktree ownership or an explicit lock +MAJOR | high | plan:121,130-133,199-202 | The clean-tree and empty-index checks mask command failure inside `test -z "$(...)"`, and the scoped `git diff HEAD` read has no checked status; empty stdout from a failed Git command is accepted as clean | A corrupt or unreadable index, failing diff driver, or hostile environment can bypass the safety preconditions and reach `git add`, amend, or hard reset | Capture each Git command's output only after checking its status, stop on failure, and check the scoped diff command before permitting staging +MINOR | high | plan:90-91,206-208 | The plan says `git diff HEAD -- ` covers the worktree and index, but it shows only the net HEAD-to-worktree result and can hide an index change canceled by the worktree; it also says nothing about staged paths outside the scope | The read is weaker than its description and cannot independently establish what the whole index would contribute to the amend | Inspect `git diff --cached HEAD` and `git diff` separately, while retaining a fail-closed whole-index emptiness check +MAJOR | high | plan:164-170,199-202 | If `git add` succeeds but `git commit --amend` fails or is interrupted, rerunning the preflight says "run the amend step only," but that step immediately rejects the task's now-nonempty index | The documented resume path dead-ends on a normal partial-execution state and gives no safe way to distinguish intended staging from foreign staging | Add an explicit staged-state recovery that verifies the cached path set and patch, then either permits the amend or unstages only the verified task path before rerunning +MAJOR | high | plan:125-139 | A clean worktree does not make `git reset --hard "$sha"` safe: unrelated commits made after the checkpoint are invisible to `git status`, and a concurrent write can occur between the check and reset | Rollback can remove another actor's clean committed work from the branch or destroy a racing tracked edit; the stated safety claim is false | Require exclusive execution, verify the current HEAD is the exact expected C1 WIP shape, create a backup ref before resetting, and re-check immediately before the destructive operation +MAJOR | high | plan:112-139,798-799 | The rollback needs a manually copied SHA and only handles the clean state after a completed amend; a half-executed replacement or staged edit merely gets "stop and inspect" with no recovery procedure | Interruption is one of the requested risk paths, and losing terminal output or stopping between replace, stage, and amend leaves the promised rollback unavailable | Persist the checkpoint as a named Git ref and specify a non-destructive recovery for verified C1-only unstaged and staged states before any reset +MINOR | high | plan:186-203,264-284 | Exact-line assertions prove the intended replacement exists and the old line is gone, but they do not reject unrelated edits elsewhere in the same file, and `git add ` stages the entire file | A careless replacement or concurrent edit can delete or alter unrelated documentation while every assertion and HEAD read-back still passes | Compare the complete patch against an expected patch or generated expected file and fail if any hunk lies outside the task's declared replacement +MAJOR | high | plan:56-58,723-747 | The plan says Task 9 re-runs the discovery grep and turns completeness into a check, but Task 9 checks only selected numeric spellings and two knob substrings; the three known nonnumeric claims are covered only by their own old-line assertions | A tenth nonnumeric sentence stating the same claim would pass Task 9, so the plan has not mechanized the completeness property it claims | Narrow the claim to the exact bounded checks actually run and add a reviewed explicit inventory or a separate semantic sweep for nonnumeric satisfaction and derivation claims +MINOR | high | plan:731-740 | The C2-owned line is described as filtered by exact text, but `grep -vF 'Codex Gate B satisfied (3/3 cycle'` removes any line containing that substring | A modified line can append another fixed-floor claim and still be hidden from the sweep | Filter the complete expected line with whole-line fixed matching, or assert that exact C2 line once before excluding it +MINOR | high | plan:728-735,743-747 | The heading claims no numeric floor claim survives, but the regex misses ordinary numeric spellings such as "floor of 3," "pass 3," and "three-pass" | An undiscovered or newly malformed fixed-floor sentence can survive while the named check exits successfully | Either cover the numeric forms the claim promises or rename the check as a bounded scan for the exact legacy spellings and leave completeness to an explicit inventory +NIT | high | plan:100-103,749-754,780-784 | The global constraints and self-review say every check uses whole-line `grep -cxF`, but Task 9 uses substring `grep -cF` checks and regex pipelines | The plan's own description of its verification method is factually inconsistent, making the review evidence look stronger than it is | Limit the whole-line claim to replacement preflights and assertions, or convert the Task 9 property checks to exact whole-line checks +MINOR | high | plan:769-770 | When Task 9 finds an unexpected site, the plan tells the executor to fix and amend "as the tasks above do" without an OLD block, NEW truth, preflight, exact assertion, instantiated identity checks, or HEAD read-back | The exceptional path bypasses the very safeguards the plan says every edit and amend carries and lets an unreviewed scope expansion enter the WIP | Make an unexpected hit a stop-and-revise-plan condition, or add a fully instantiated task for each newly discovered site and send it through Gate A before editing +END OF FINDINGS (17 total) diff --git a/.context/codex-reviews/gate-a-plan-pr1-pass-1.md b/.context/codex-reviews/gate-a-plan-pr1-pass-1.md new file mode 100644 index 0000000..9252b78 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pr1-pass-1.md @@ -0,0 +1,8 @@ +MAJOR | high | Task 3, lines 72-96 | The plan defines exclusions for 4a and 4b but never defines either scan's positive file domain or extensions, so "in-scope files" in Tasks 1 and 2 has no deterministic meaning | An implementer could reuse the existing YAML/Markdown/JSON/TOML domain, scan only Markdown, or scan every text file, producing materially different enforcement and false-positive behavior; the two pending ledger rows cannot honestly claim a bounded mechanical rung without naming what is scanned | Add a per-check inclusion domain alongside the exclusion table, including file extensions/path roots and whether symlinks/hidden files are included, then make the fixtures exercise a file at each relevant boundary +MAJOR | high | Task 4, lines 104-113 | Parameterizing the source `CHECKER` variable alone does not make a differently named scratch mutant executable by the suite: all four copies currently use `cp "$CHECKER" "$work/r/scripts/"`, while all four invocations hard-code `sh scripts/check-invariants.sh` | An override such as `/tmp/check-invariants-mutant.sh` is copied under its mutant basename, leaving `scripts/check-invariants.sh` absent; mutation runs then test a missing-file failure instead of the mutant, recreating the exact no-op/wrong-reason risk this prerequisite is meant to remove | Require every call site to copy the override to the fixed destination `"$work/r/scripts/check-invariants.sh"` (or require the mutation path to have the production basename) and add a sentinel mutation proving the invoked file is the override +MAJOR | high | Task 5 Verify, lines 151-152 | The plan requires re-running "the invariant-5/6 mutation," but the actual `scripts/check-invariants.test.sh` contains no mutation procedure for either existing check and the plan neither defines mutations nor enumerates the old fixtures expected to flip | A fresh implementer cannot deterministically perform this verification or distinguish preserved diagnostic isolation from a suite that remains green because the old checks were accidentally disabled | Specify the exact invariant-5 and invariant-6 mutations, the complete expected fixture names/status changes for each, the command/output assertions, and the restore procedure, or replace this requirement with concrete targeted pre-existing reject cases whose diagnostics are re-verified +MAJOR | high | Task 6, lines 160-169 | The known doc-site list omits `README.md:126`, which currently describes `scripts/check-invariants.sh` as enforcing only invariants 5 and 6 | The PR would knowingly leave a public contributor-facing description stale after adding checks 4a/4b, contradicting Task 6's "every sentence" requirement and reproducing the `docs-drift` class this work is meant to harden | Add `README.md` to the required files and update its checker description and surrounding check count consistently +MAJOR | high | Success criterion 1, lines 220-222, versus Task 6 line 164 | The plan says no prompt artifact is edited and therefore skips the invariant-11 review, but Task 6 explicitly edits `AGENTS.md`, an instruction file read directly by coding/review models and treated by `CLAUDE.md` §5 as prompt/product rather than prose | This false scope claim can let changed model instructions bypass the repository's only prompt-quality gate, directly risking invariant 11 | Treat the `AGENTS.md` edit as a prompt change and require a 12-item `docs/prompt-standards.md` review of the changed instruction text, or avoid editing `AGENTS.md` only if its checker description can truthfully remain unchanged +MINOR | high | Task 8, lines 200-207 | "Mark the checker half done; leave the template-sync half pointing at PR 2" does not map to the actual `todos.md` structure: the checker is one standalone unchecked item, while template sync is a separate preceding unchecked item, so there is no checker/template pair or "half" within one item | Different implementers may check off the checker item, rewrite its stale resolution-vehicle prose, split it, or partially edit it, creating unnecessary backlog drift | State exactly that the standalone "Prompt-standards conformance checker" item becomes checked and how its obsolete canvas-round/trigger text is rewritten, while the separate "ad-hoc task briefs" template-sync item remains byte-unchanged for PR 2 +MINOR | medium | Task 1 line 24 and Task 5 fixtures | The required token test is written with `\b`, but the plan does not state which grep mode/implementation supplies that boundary semantics even though this checker is run on both macOS/BSD userland and Linux CI | A nominally identical implementation can interpret the boundary differently or rely on a non-POSIX extension, causing platform-specific acceptance of values such as `ClaudeX` and undermining deterministic implementation | Specify a portable boundary such as `\([^[:alnum:]_]\|$\)` in the chosen BRE/ERE mode, name the exact grep mode, and add reject fixtures for token prefixes like `ClaudeX`, `Codex2`, and `GPTfoo` +END OF FINDINGS (7 total) diff --git a/.context/codex-reviews/gate-a-plan-pr1-pass-2.md b/.context/codex-reviews/gate-a-plan-pr1-pass-2.md new file mode 100644 index 0000000..32a1134 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pr1-pass-2.md @@ -0,0 +1,8 @@ +MAJOR | high | Task 5, lines 178-186 | Three of the four supposedly named pre-existing reject cases do not exist under the stated names: the suite has `major-only ref rejected`, not `unpinned action ref rejected`; `ubuntu-latest rejected`, not `ubuntu-latest runner rejected`; and no case named `re-declares a convention-loaded component` (that text is only the diagnostic substring shared by `hooks key rejected`, `skills key rejected`, and `key/colon split across lines rejected`). Only `unpinned npx in a shell script rejected` exists exactly as claimed. The four `cp "$CHECKER"` sites and four hard-coded invocations do exist. | An implementer following the required named-case verification literally cannot select or report three requested cases, so the plan is not ready to implement deterministically and its claim that Pass 1 F3 was addressed is false. | Replace each nonexistent name with one exact existing case name and diagnostic, for example `major-only ref rejected` + `$ACTION`, `ubuntu-latest rejected` + `$RUNNER`, and `hooks key rejected` + `$MANIFEST`, or explicitly require all three manifest reject cases if that is the intended coverage. +MAJOR | high | Invariants touched, lines 255-260 | The plan says preservation of invariants 5 and 6 is “proven by re-running their mutation,” directly contradicting Task 5’s corrected statement that no invariant-5/6 mutation procedure exists and that named reject cases must be re-verified instead. | This leaves two mutually exclusive acceptance procedures and resurrects the exact unexecutable instruction Pass 1 F3 was meant to remove; different implementers can reasonably do different work or falsely claim a nonexistent mutation was run. | Replace the mutation sentence with the Task 5 procedure: preservation is evidenced by re-running the exact selected pre-existing reject cases and confirming each case’s own diagnostic. +MINOR | high | Task 3, lines 108-110 | The claimed observed path behavior is wrong in the reviewed working directory: both `grep -rl … .` and `grep -rno … .` currently emit paths with a `./` prefix, not one without and one with. | The prefix-independent exclusion requirement remains prudent, but presenting an untrue reproduction as verified weakens the evidence for a mechanical guard and can mislead implementation or debugging across grep variants. | State that prefix form varies by grep implementation and must not be assumed; retain paired excluded/non-excluded controls, and remove the false claim about what the current checkout emits. +MINOR | high | Task 1, lines 29-33 | The environment claim that local `grep` resolves to ugrep 7.5.0 is false for the supplied working directory/session: `grep --version` reports BSD grep 2.6.0-FreeBSD. | The portable ERE boundary decision is still correct, but a plan intended to be independently executable should not ground that decision in stale or unverifiable tool identity. | Replace the version-specific assertion with the actual verified implementation/version, or make the rationale implementation-neutral: CI and developer environments may use different grep implementations, so only POSIX ERE constructs are allowed. +MINOR | high | Task 6, lines 198-206 | The plan calls the sweep “Five sites” but enumerates six distinct edit locations: the script header, two AGENTS.md locations, docs/architecture.md, README.md, and the CI workflow. All are present and stale as described. | A cardinality contradiction inside a plan for hardening count drift invites one of the enumerated locations to be skipped and makes the final sweep evidence harder to audit. | Say “six locations across five files” (if file count is the intended metric) and enumerate all six consistently. +MINOR | medium | Task 1, lines 18-29 | Rejecting every `Target model:` value containing the substring ` or ` is broader than the defect being hardened: it also rejects a single named model with alternative execution-surface prose such as `Claude via Claude Code or the API`, even though checklist item 1 requires a named model and gives execution surfaces as explanatory context. | This creates avoidable false positives outside the ambiguous-model state, and the ledger row would overstate how narrowly the mechanical rung guards the settled defect. | Reject alternation only when it introduces another recognized model token (for example `Claude or Codex`), or define and test a stricter single-model grammar that still permits ordinary prose about one model’s execution surface. +MINOR | medium | Success criteria 2, lines 273-275 | The summary says the exact quality command runs “both regression suites,” but the authoritative AGENTS.md command runs three test suites: the hook suite, the invariant-checker suite, and the version-bump-checker suite. | The command itself is referenced verbatim, so execution can still be correct, but the adjacent count is another stale prose count in the artifact specifically designed to prevent count drift. | Say “all three regression suites” or avoid the count and name the three test commands. +END OF FINDINGS (7 total) diff --git a/.context/codex-reviews/gate-a-plan-pr1-pass-3.md b/.context/codex-reviews/gate-a-plan-pr1-pass-3.md new file mode 100644 index 0000000..a383e6d --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pr1-pass-3.md @@ -0,0 +1,6 @@ +MAJOR | high | Task 2, lines 60-67 | the rule requires exactly one `## Checklist` heading in each definition, but `/usr/bin/grep` shows both actual headings are `## Checklist (each item must be verifiably true)`, so an exact implementation of the specified heading rejects the real repository before comparing any items | the plan is not implementable literally and cannot make the required real-repository invariant check green; loosening the match ad hoc would make different implementers choose different heading grammars | specify and fixture the actual accepted heading grammar, for example exactly `## Checklist (each item must be verifiably true)`, or explicitly define an anchored `^## Checklist([[:space:]].*)?$` contract with accepted and near-miss cases +MAJOR | high | Task 6, lines 229-239 | the measured inventory is seven stale claims, not six: besides the script header, two AGENTS.md sites, docs/architecture.md, and README.md, `.github/workflows/ci.yml` has two distinct stale sites at lines 75 and 83—the `Invariants 5 and 6 mechanically` comment and the `Invariant checks (pinning, manifest) + both checker suites` step name | trusting the claimed six-location inventory can leave one CI-facing description stale, directly repeating the docs-drift class this PR is meant to harden | change the count to seven locations across five files and enumerate both CI lines separately; retain the semantic final sweep +MAJOR | high | Tasks 4-5, lines 161-164 and 194-198 | the mutation verification still does not define how to disable each whole check, which exact source range or switch is changed, how the mutant is created and restored, or the command and assertions that establish the complete expected case-status delta; the sentinel proves only path injection and does not make the later “neutered copy” procedure deterministic | a fresh implementer can mutate different logic, accidentally disable shared initialization or diagnostics, or report a partial/irrelevant delta while believing the load-bearing requirement passed, so the plan does not yet provide reproducible evidence for honestly resolving the pending rows | specify a test-only per-check disable switch or exact marker-delimited mutation, then give the concrete commands and expected named FAIL/no-movement assertions for 4a and 4b, including separately captured suite status and restoration +MINOR | high | Task 1, lines 18-27 | “Count all `Target model:` declarations” has no declaration-recognition grammar even though validation is anchored to `^Target model:`; it is unspecified whether indented lines, blockquotes, fenced examples, prose mentions, or lines with leading whitespace count toward cardinality | two POSIX implementations can follow the plan and disagree on cardinality, and a future prompt containing an example label can be rejected or accepted inconsistently | define one exact anchored declaration regex used for both counting and value validation, and add near-miss fixtures for prose, indentation, blockquotes, and fenced examples according to the chosen contract +MINOR | medium | Task 5, lines 185-192 | the exclusions row says to accept violating content “inside each excluded path” without enumerating the per-check matrix, even though 4a and 4b have different exclusions and several entries denote directory prefixes while the self-exclusions denote two exact files | an implementer cannot derive a unique fixture set or know whether every listed exclusion has both its positive control and neighboring rejection control, leaving path-form and overbroad-exclusion regressions unevenly tested | enumerate every applicable check/path pair and the exact excluded and neighboring fixture filenames, explicitly omitting `docs/hardening-log.md` only from 4b +END OF FINDINGS (5 total) diff --git a/.context/codex-reviews/gate-a-plan-pr1-pass-4.md b/.context/codex-reviews/gate-a-plan-pr1-pass-4.md new file mode 100644 index 0000000..d62fba2 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pr1-pass-4.md @@ -0,0 +1,7 @@ +MAJOR | high | Task 1, Rule step 1 and near-miss fixtures | NOT READY TO IMPLEMENT: the sole declared recognition grammar `^Target model:` necessarily matches a column-zero `Target model:` line inside a fenced block, while the same paragraph requires fenced examples not to count; no fence-aware preprocessing or parser contract is specified | a conforming POSIX-grep implementation cannot satisfy the required fenced near-miss fixture, so implementers must invent materially different Markdown parsing behavior and check 4a is not deterministic | specify the complete fence-aware line-selection algorithm, including backtick and tilde fences, variable fence lengths, indentation, and unclosed fences, or remove the claim that fenced declarations are excluded and change the fixture contract accordingly +MAJOR | high | Task 4, sentinel verification | adding a marker to checker stdout cannot make that marker appear in the regression-suite output: all four checker invocations capture stdout/stderr into `out`, successful cases print only their assertion names, and correctly diagnosed reject cases also discard `out` | the mandated sentinel proof is impossible against the actual suite, so the override can remain unproven or force an implementer to change unrelated reporting behavior without authorization | define a sentinel that changes an asserted status/diagnostic in one named case, add an explicit test-only trace that identifies the invoked checker path, or run a separately specified scratch fixture directly through the override and assert its marker there +MAJOR | high | Task 5, mutation procedure step 4 | 4b's exact expected mutation delta is not named: “precisely its own set” points only to descriptive table entries, while the procedure requires comparing every named test status and gives no exact fixture-output names or parsing/assertion command for 4b | a failing mutant suite alone cannot prove that all and only the intended 4b cases flipped, and two implementers can choose different case names or comparison logic while claiming the same evidence | enumerate the exact 4b assertion names expected to flip, enumerate or mechanically derive the complement, and specify the command that compares captured mutant output/status with the unmodified baseline and fails on missing or extra movements +MAJOR | high | Tasks 2 and 7, `docs-drift` pending-row resolution | check 4b explicitly excludes word-form counts, but the actual pending 2026-07-25 occurrence it claims to resolve is the word-form spelling “all ten items”; therefore the proposed rung does not reject recurrence of the motivating defect even though Task 7 marks that pending hardening resolved | the appended ledger row would overstate what was hardened and would resolve the pending row dishonestly, undermining the fingerprint ladder and invariant 11's enforcement-claim requirement | either mechanically recognize the bounded word forms needed for the checklist count and add reject/accept fixtures, or narrow the new row to digit-form claims and leave the 2026-07-25 pending row unresolved with an explicit follow-up for its actual spelling +MINOR | high | Task 5, named pre-existing reject-case verification | the instruction to `grep -cF` each fixture name does not uniquely verify `ubuntu-latest rejected`: that substring occurs on three lines in the actual suite, including quoted and matrix cases, although the intended exact assertion exists only once | a fresh implementer following the prescribed verification can get a count of 3 and cannot deterministically identify which legacy case must retain its own diagnostic | verify the quoted assertion argument or anchored emitted test name, for example a fixed-string search including `"ubuntu-latest rejected"`, and state the expected count of one for every selected case +MINOR | medium | Task 1, Rule step 3 | the single-model rule rejects only the literal `Claude or Codex` shape, leaving other values that begin with an accepted token but name multiple executing models—such as `Claude or GPT`, `Codex or Claude`, `Claude and Codex`, or `Claude/Codex`—without a stated disposition or fixtures | check 4a can report conformance for an ambiguous multi-model declaration even though its rationale says the field must name one executing model, making the new ledger mechanism narrower than its prose claim | define the exact supported single-model value grammar or explicitly bound the alternation check to a tested separator and all distinct recognized-token pairs, then add accept/reject fixtures for the chosen boundary +END OF FINDINGS (6 total) diff --git a/.context/codex-reviews/gate-a-plan-pr1-pass-5.md b/.context/codex-reviews/gate-a-plan-pr1-pass-5.md new file mode 100644 index 0000000..851e429 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pr1-pass-5.md @@ -0,0 +1,7 @@ +BLOCKER | high | Tasks 2, 3, and 7; docs/hardening-log.md:27 | Check 4b scans docs/hardening-log.md and requires every recognized claim to equal N, but the append-only pending row that must remain byte-unchanged contains the deliberately historical claim `all ten items`; `/usr/bin/grep` therefore finds three in-domain claims, not the plan's claimed two claims both equal to 12 | The checker will reject the real repository forever, so success criterion 3 cannot pass and appending a resolution row cannot honestly resolve the pending row | Exclude docs/hardening-log.md from 4b as quoted historical evidence, add the same exact-path exclusion/neighbour fixtures required for 4a, correct the measured claim total, and state this exclusion and its instruction-backed risk in the new ledger row +MAJOR | high | Tasks 3 and 5, Markdown-only scan domain versus checker-file exclusion matrix | Both checks are specified to scan only `*.md`, yet the plan says the `.sh` checker and test need self-exclusion and requires `scripts/check-invariants-helper.sh` plus the two `.sh` exact-path fixtures to exercise rejection/acceptance; none of those files can enter a Markdown-only scan, and the stated self-match cannot occur | These required fixtures are impossible to satisfy and cannot prove either exact-path filtering or that real-repository fixture literals are harmless, so a fresh implementer has mutually contradictory requirements | Remove the `.sh` self-exclusions and their 4a/4b fixture rows because `--include='*.md'` already excludes them, or deliberately widen the scan domain and then specify why shell source is prompt text; retain a real-repository acceptance test for the actual Markdown domain +MAJOR | high | Task 5, mutation delta comparison | The plan requires a symmetric difference of sorted raw `ok -`/`FAIL -` output lines to equal the enumerated set of case names, but each status flip contributes two unequal raw lines—its baseline `ok - NAME` and mutant `FAIL - NAME (got green)`—and the FAIL suffix is not normalized away | The prescribed mechanical oracle cannot equal a one-name-per-case expected set, so mutation evidence is not deterministic as written | Specify an exact normalization that converts both forms to `NAMESTATUS`, compare the baseline and mutant maps by name, and assert the exact expected `ok→FAIL` transitions plus absence of added, removed, duplicate, or differently directed cases +MAJOR | high | Tasks 1 and 5, second-model-token fixtures | Rule 4a declares four tested separators and rejection of any second distinct recognized token, but the mandatory fixture table contains only `Claude or Codex`; it has no required cases for ` and `, `/`, comma, alternate first tokens, or another distinct pairing despite claiming the separator boundary is covered | An implementation can hard-code the one historical literal or implement only ` or ` and still pass every enumerated fixture, repeating the pass-4 defect while the plan calls the broader grammar mechanically hardened | Add named reject fixtures covering every separator and enough token orderings to prove distinct-token parsing, plus accept fixtures showing repeated prose uses of the same model token and the required `Claude via Claude Code or the API` boundary +MAJOR | high | Task 5, mutation commands and restoration | The commands write to `$TMP/mutant.sh` and `$TMP/out`, but the plan never creates or validates `$TMP`, and its “restore” step merely says to delete it without an exact cleanup command or trap | In an ordinary shell an unset TMP expands to the root-level paths `/mutant.sh` and `/out`, making the procedure unsafe or permission-dependent and therefore not executable as written | Prepend `TMP=$(mktemp -d)` with a failure check and a trap that removes that exact validated directory, then give the exact cleanup and both 4a/4b mutant commands +MINOR | high | Task 4, direct marker proof | The first required proof says to set the checker override and then run the marker-bearing checker directly; a direct invocation does not read `CHECK_INVARIANTS_BIN`, so it proves only that the copied checker prints a marker, not that the suite's override path is wired | The plan labels a non-proof as path-injection evidence, weakening the verification narrative even though the altered-diagnostic suite proof can establish injection behaviorally | Either drop the direct proof as redundant or invoke the suite through the override and add a narrowly scoped reporting path that makes the marker observable; describe the direct run only as a mutant sanity check +END OF FINDINGS (6 total) diff --git a/.context/codex-reviews/gate-a-plan-pr1-pass-6.md b/.context/codex-reviews/gate-a-plan-pr1-pass-6.md new file mode 100644 index 0000000..42b5ac2 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pr1-pass-6.md @@ -0,0 +1,5 @@ +MAJOR | high | Task 5 lines 287-294, mutation expected delta | the exact 4a delta omits five required reject fixtures (`Claude or GPT`, `Codex or Claude`, `Claude and Codex`, `Claude/Codex`, and `Claude, Codex`), and the exact 4b delta omits all three required near-miss-heading fixtures; deleting either marked check will make those cases transition ok→FAIL too, contradicting the requirement that precisely the enumerated names flip | the prescribed load-bearing mutation verification cannot pass for a conforming implementation, so the plan is NOT READY to implement deterministically and success criterion 4 is unreachable as written | enumerate every required reject fixture in each mutation delta, preferably derive the expected names from one authoritative per-check list, and assert that complete set +MAJOR | high | Task 1 rule 3 and Task 5 line 251 | the rule rejects any second distinct recognized token after the first for every listed separator, but the fixtures exercise only six hand-picked combinations and do not cover each separator in both token orders or any GPT-leading value; the phrase “one reject per tested separator and ordering” is therefore false, and an implementation specialized to those examples can pass while accepting cases such as `GPT/Codex` or `Codex, Claude` | check 4a would overstate its guarded spelling in the new ledger row, weakening the honest rung-2 resolution of the 2026-07-25 enforcement-claim row and risking invariant 11 | either define the guarded pair/order matrix narrowly and state that limitation, or generate fixtures for every ordered pair of distinct Claude/Codex/GPT tokens across every supported separator plus same-token accepts +MAJOR | high | Task 5 lines 298-302, result normalization | the only normalization example does not match the suite’s actual success prefix `ok - ` and replaces a prefix with tab-plus-status rather than producing the promised `NAMESTATUS`; the plan also does not give an exact parser for `FAIL - NAME (reason)` that preserves names while stripping only the diagnostic suffix | different implementers can produce empty or malformed maps and falsely accept or reject the mutation delta, recreating the verification-masks-failure class this procedure is meant to prevent | specify and fixture-test an exact POSIX normalization command against the actual `ok - NAME` and `FAIL - NAME (reason)` forms, require one unique record per assertion name, and show the exact joined record shape used for transition checks +MAJOR | medium | Task 4 lines 213-228, altered-diagnostic injection proof | the plan requires altering “the 4a diagnostic string” and confirming the named 4a rejects report `wrong diagnostic`, but it never specifies whether 4a has one shared diagnostic or several diagnostics for cardinality, invalid value, and multiple-model failures, nor which complete case set must flip in this proof | a fresh implementer can alter one message and legitimately observe only a subset flip, or collapse distinct diagnostics merely to satisfy the proof, so the required path-injection evidence is not deterministic | prescribe one exact diagnostic mutation and the exact fixture names expected to report `wrong diagnostic`, or add a dedicated sentinel reject fixture whose sole purpose is proving the override path +END OF FINDINGS (4 total) diff --git a/.context/codex-reviews/gate-a-plan-pr1-pass-7.md b/.context/codex-reviews/gate-a-plan-pr1-pass-7.md new file mode 100644 index 0000000..a3443c2 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pr1-pass-7.md @@ -0,0 +1,6 @@ +BLOCKER | high | Task 5, Mutation procedure step 4 | The plan says every 4a/4b fixture registers in `CASES_4A`/`CASES_4B`, then requires disabling a check to make exactly every registered name transition `ok→FAIL`; accept fixtures and the excluded side of each exclusion pair remain `ok` when enforcement is removed, so the required delta is impossible as written | The mutation proof cannot pass without either violating its authoritative-list rule or falsely reporting unchanged accepts as movers, so the plan is NOT READY to implement deterministically | Define authoritative expected-status maps, or register only fixtures expected to flip and separately derive/assert that every other normalized record is unchanged; state explicitly how unregistered new reject and accept fixtures are detected +MAJOR | high | Task 4, behavioural override proof | The dedicated `override sentinel` is required to have a diagnostic marker that no other fixture uses and whose mutation flips only that fixture, but the plan specifies neither a rule-valid sentinel input nor a production diagnostic branch that could emit such a unique marker; all specified 4a violations share cardinality, invalid-value, or multi-model diagnostics with other fixtures | An implementer must invent either an unplanned production special case solely for the test or a different mutation mechanism, so the supposedly deterministic path-injection proof is not implementable from the plan and could distort the checker merely to satisfy its test | Specify a concrete sentinel input and legitimate general diagnostic it exercises plus a mutation seam unique to its output, or prove injection by overriding with a wrapper/mutant whose observable behavior is unique without adding sentinel-only behavior to the production checker +MINOR | high | Task 1 Rule 3 and Task 5 generated 24-case matrix | The first model token has a portable token boundary, but the second “recognized token” grammar is not specified and the matrix tests only exact second tokens; an implementation that rejects `Claude or Codex2` or `GPT/ClaudeX` by prefix match passes all 24 generated rejects and the existing first-token near-misses | The checker can produce false positives outside the stated distinct-recognized-token rule, contrary to invariant 2's directional calibration and the plan's claim that the generated matrix pins the grammar | Require the same `([^[:alnum:]_]\|$)` boundary after the second token and generate acceptance fixtures with token-prefix near-misses in second position for every separator or otherwise derive them across the token/separator matrix +MINOR | high | Task 2 recognized claim grammar | `all (\|)( checklist)? items` has no left boundary before `all` or right boundary after `items`, so a direct ERE substring implementation also recognizes text such as `small ten items` and `all ten itemsized`; the plan does not say whether those are claims or near-misses | Different conforming implementations can disagree on scan results, and the obvious grep implementation can reject unrelated prose while the ledger overstates the bounded spelling actually guarded | Specify portable outer boundaries, for example `(^\|[^[:alnum:]_])all ... items([^[:alnum:]_]\|$)`, and add accept fixtures for embedded-prefix and embedded-suffix near-misses +MINOR | medium | Task 5, `norm()` | The FAIL-with-reason sed arm accepts only a final parenthesized reason containing no parentheses (`([^()]*)`), although suite failures interpolate checker output and fixture text that may themselves contain parentheses; the current 61 green records validate only the no-failure path and do not establish normalization of such a mutated failure | A legitimate mutation can disappear from the normalized map or fail the uniqueness/join checks for a parser artifact, making the required mutation evidence brittle and contradicting the claim that the `FAIL - NAME (reason)` form is generally parsed | Parse the status prefix first and strip only the known final reason suffix with a rule that tolerates nested literal parentheses in its contents, or constrain and test the failure-reason grammar explicitly with a representative nested-parenthesis line +END OF FINDINGS (5 total) diff --git a/.context/codex-reviews/gate-a-plan-pr1-pass-8.md b/.context/codex-reviews/gate-a-plan-pr1-pass-8.md new file mode 100644 index 0000000..0c274eb --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pr1-pass-8.md @@ -0,0 +1,6 @@ +MAJOR | high | Task 5, lines 299-313 | the narrative adopts distinct-token counting and says the fixture set is four separator-independent rejects plus three accepts, but the authoritative fixture table still requires the removed separator/order suite (`Claude or GPT`, `Codex or Claude`, `Claude/Codex`, etc.) and omits the settled `Claude or Codex2` accept | an implementer following the table will preserve residue of the deleted separator grammar and will not pin a specifically settled boundary, so the plan is internally contradictory and not deterministic | replace the 4a table cells with exactly the four settled two-token rejects and the three settled accepts, including `Claude or Codex2`, and remove “one reject per tested separator and ordering” +MAJOR | high | Task 4 procedure, lines 242-255 | the procedure captures mutant status in `st` but never prints or asserts it, never captures or asserts the baseline status, and does not assert that `diff` found the expected changes | a failing baseline, a mutant that aborts for an unrelated reason, or an empty/unexpected diff can still be recorded as mutation evidence, recreating the logged `verification-masks-failure` class despite claiming to avoid it | capture both statuses, require baseline status 0, require mutant status nonzero, derive and compare the exact flipped assertion names, and fail unless that mapping is exactly the expected mapping before recording it +MAJOR | high | Tasks 2 and 5, lines 92-129 and 310-313 | the admitted `[0-9]+` grammars do not define how leading-zero checklist labels or digit claims such as `01. **...` and `all 012 items` compare to canonical `1..N`, nor how arbitrarily large digit strings are handled without shell-arithmetic overflow | two conforming implementations can disagree, and a naïve POSIX-shell numeric comparison can error or misclassify malformed state, conflicting with invariant 2’s fire-on-uncertainty direction | require canonical decimal spellings (`0` or leading-zero labels rejected, claims `[1-9][0-9]*`) or specify safe string normalization and overflow-independent comparison, then add reject/accept fixtures for zero, leading zeros, and a very long digit claim +MINOR | high | Task 4 recording trigger, lines 257-275 | the mandated re-run trigger is only “when modifying the scan logic,” although changing fixtures, assertion names, marker placement, or the suite harness can invalidate the recorded mutant-to-assertion mapping without modifying scan logic | the permanent evidence comment can become stale while its stated trigger says no re-run is needed, weakening the honesty required by prompt-standards item 11 | broaden the trigger to changes to either marked check, its markers, its fixtures/assertion names, or the harness that runs them, and require updating both recorded mappings after such a run +MINOR | high | Task 4 procedure, lines 242-252 | scratch cleanup occurs only on the successful fall-through path and there is no trap, while several commands can terminate or be manually interrupted before `rm -rf "$TMP"` | failed development runs leave full repository copies, including `.git` and ignored review artifacts, in temporary storage and make repeated evidence runs unnecessarily stateful | install a validated, quoted EXIT/HUP/INT/TERM cleanup trap immediately after `mktemp -d`, and clear it only after explicit cleanup if desired +END OF FINDINGS (5 total) diff --git a/.context/codex-reviews/gate-a-plan-pr1-pass-9.md b/.context/codex-reviews/gate-a-plan-pr1-pass-9.md new file mode 100644 index 0000000..a4e32a6 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pr1-pass-9.md @@ -0,0 +1,6 @@ +BLOCKER | high | Task 2, lines 102-108 and Task 5 fixture requirements | the plan requires `all 012 items` to fire as malformed, but its only claim grammar accepts canonical `[1-9][0-9]*` digits, so `all 012 items` is not recognized and is silently ignored; unlike an invalid checklist label, the plan specifies no required claim cardinality or separate malformed-numeric-claim detector that could reject it | a conforming implementation can pass the required `all 012 items` reject fixture only by inventing behavior absent from the rule, and the stated canonical-decimal protection plus invariant-2 firing direction are false as specified, so the plan is NOT READY and the docs-drift row cannot yet be resolved honestly | add an explicit broader near-claim detector for `all [0-9]+( checklist)? items` that rejects noncanonical numeric spellings before canonical claim comparison, with boundaries and fixtures, or explicitly accept/ignore `all 012 items` and remove every claim that it fires +MAJOR | high | Task 4, lines 262-277 | the procedure does not compare the flipped set to an expected set despite claiming that this catches an unrelated mutant abort and that assertions flipped “exactly”: it asserts only mutant nonzero plus a non-empty list of newly printed FAIL lines | a syntax/runtime failure in the neutered checker can make accept fixtures fail and produce a non-empty flipped list, which this procedure records as load-bearing evidence even though the mutant failed for the wrong reason; the dry-run's three named flips were observed manually, not enforced by the written procedure | define the complete expected flipped assertion set for each marked check and compare sorted expected versus actual exactly, while also rejecting unexpected checker diagnostics/abort output; if exact-set validation is intentionally manual, delete the claim that the procedure catches unrelated aborts and specify the required human validation explicitly +MINOR | high | Task 5, lines 329-343 | the prose says there are “three accepts” and “all seven verdicts,” but the fixture table requires four accepts and therefore eight verdicts by additionally including `Claude in a chat interface, upstream of Claude Code` | this is residual count drift in the section meant to be the authoritative settled fixture inventory, so implementers and recorded mutation evidence can disagree about whether the fourth accept is mandatory | change “three accepts” to “four accepts,” include the chat-interface case in that sentence's enumeration, and change “seven” to “eight” +MINOR | high | Task 2, lines 109-112 | the checklist-label requirement contains the broken duplicate phrase “The item labels must also form the contiguous sequence 1..N, and whose item labels form the contiguous sequence 1..N” | the malformed sentence obscures the subject and makes an otherwise important malformed-state rule look like two partially merged requirements | replace it with one sentence stating that each checklist section's labels must form exactly the contiguous canonical sequence 1 through N +MINOR | medium | Success criterion 4 versus Task 4, lines 262-268 | “The mutation run is evidenced by captured output, not asserted” conflicts with Task 4's explicit requirement that baseline status, mutant status, and non-empty flips are asserted, and it does not say whether “not asserted” means not merely claimed in prose | a fresh implementer cannot tell whether the acceptance criterion requires raw output retention, executable assertions, or both, which weakens the evidence handoff | rewrite the criterion to require both the asserted validity checks and retained exact flipped-assertion output, and say that an unsupported prose attestation is insufficient +END OF FINDINGS (5 total) diff --git a/.context/codex-reviews/gate-a-plan-pr2-pass-1.md b/.context/codex-reviews/gate-a-plan-pr2-pass-1.md new file mode 100644 index 0000000..7585ea3 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pr2-pass-1.md @@ -0,0 +1,9 @@ +MAJOR | high | Task 1, lines 42-46 | “run the anchored recurrence grep” does not say how to derive the canonical fingerprint, whether to repeat the check for every fixed Blocker/Major, or how the grep result determines “anything worth keeping” | the observable anchor is deterministic, but the action at that anchor is not; an agent can run one arbitrary grep, miss other findings, or stop after a no-match without ever making the follow-up decision | require one intake/fingerprint/anchored-column-2 recurrence check per fixed Blocker/Major, and state the explicit outcomes: no match, prior real rung, and pending, while leaving application to a separate `harden-finding` cycle +MAJOR | high | Task 2, lines 67-72 | `-resume.md` does not uniquely name the independent Gate-A spec and plan cycles, and a single Gate-B name also leaves the two concurrently running full-review branches able to overwrite the same interruption note | stale spec context can be consumed during the plan run, and concurrent Gate-B writers can recreate the session-bound loss this companion is meant to prevent | define collision-free cycle identities such as `gate-a-spec-resume.md`, `gate-a-plan-resume.md`, and branch-specific Gate-B resume files, or explicitly make one coordinator the sole writer of a single Gate-B cycle note +MINOR | medium | Task 2, lines 69-76 | “overwritten on interruption, removed when it closes” does not define who writes it, what minimum resume state it carries, or whether “closes” means a valid pass, a clean gate run, or the cycle-closing commit | different agents can preserve incompatible state or delete the note before it has served its purpose, leaving the resume path nondeterministic | specify the owner, minimum fields (cycle identity, next valid pass, spent recovery budget, surviving artifact/session IDs), and the exact creation/replacement/removal transitions +MINOR | high | Task 3, lines 90-101 | “held to this checklist in spirit” is the only required semantic content, so an implementation can imply that every ad-hoc brief is formally reviewed against all 12 items even though the settled repo paragraph explicitly says it is not | that would turn a habit-level principle into a false process/enforcement claim in every scaffolded project | require the downstream-neutral paragraph to preserve both halves: briefs should embody the checklist’s habits, but no per-brief exhaustive checklist review is claimed +MAJOR | high | Task 6(c), lines 148-153 | the sketched scope-aware recurrence rule is reversed: it says a recurrence outside the prior row’s guard should escalate, which is precisely the scope-blind over-escalation the story identifies | a new spelling or surface that the prior mechanism never claimed to cover does not prove that rung failed, so automatic escalation can demand unjustified machinery | outside the stated guard, select the fitting rung independently without recurrence escalation; inside the guard, treat it as regression and repair or strengthen the existing mechanism based on the diagnosed failure +MAJOR | high | Task 7, lines 162-166 | the plan says a green CodeRabbit check does not prove the final head was reviewed, then retains that same check as “the completion signal”; the current document also explicitly says non-pending status is proof it finished | the resulting instructions can both block on and trust a signal the new evidence says is insufficient, allowing a rate-limited non-review to be treated as completed | distinguish “the check stopped pending” from “the final head was reviewed,” and update the table, Wait-for text, and completion-signal paragraphs together with a deterministic comment/head verification and a bounded stop/escalation path for rate limiting +MINOR | high | Task 5, lines 121-129 | only the A ledger row is required to admit its instruction-backed, unenforced nature; C’s optional companion convention is equally instruction-backed and can be skipped without any mechanism noticing | the C row can otherwise read as if `P std` made session context durable rather than merely recommending durable notes | require both rows to name their actual prompt mechanism and explicitly state that no checker or hook enforces compliance +MINOR | medium | Task 2, line 77 | the plan requires the field-practice credit in both copies, which ships `infinite-portfolio-canvas` provenance into every downstream CLAUDE.md without affecting agent behavior | this is repo-specific attribution in a token-sensitive inline template and conflicts with prompt-standard item 8’s token-lean requirement | keep the credit in repo documentation or the CHANGELOG, and omit it from the downstream scaffolded prompt unless a downstream reader needs it to follow the rule +END OF FINDINGS (8 total) diff --git a/.context/codex-reviews/gate-a-plan-pr2-pass-2.md b/.context/codex-reviews/gate-a-plan-pr2-pass-2.md new file mode 100644 index 0000000..ce69719 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pr2-pass-2.md @@ -0,0 +1,9 @@ +MAJOR | high | Task 1, “for each Blocker/Major fixed this cycle” plus the real-rung outcome | Two fixed findings assigned the same candidate/canonical class are each carried into a separate follow-up; after the first follow-up appends the hardening row, the second will be classified as a recurrence even though both findings predated that mechanism | This falsely concludes that the new rung failed and drives an unjustified escalation, corrupting the ledger semantics the nudge is meant to improve | Group same-class findings into one harden-finding follow-up or require the follow-up to distinguish occurrences that predate the latest hardening row before treating them as recurrence +MAJOR | high | Task 1 real-rung outcome versus Task 6(c) and plugins/dev-workflow/skills/harden-finding/SKILL.md step 3 | The nudge says the follow-up reads the prior row’s guard and selects independently outside it, but this PR only parks that behavior in todos.md; the actual harden-finding skill still mandates “prior rung didn’t hold” and one-rung-stronger for every real-rung fingerprint match | Following the plan’s own instruction to run harden-finding therefore produces the scope-blind over-escalation that Task 6 identifies as a bug | Either update harden-finding’s recurrence semantics in this PR and include the shipped-skill/version/changelog scope, or make the nudge defer without prescribing the unimplemented guard rule and explicitly block scope-ambiguous recurrence until the parked story lands +MAJOR | high | Task 2 resume-note lifecycle, “overwritten on each interruption” | An abrupt timeout, crash, context loss, or session termination leaves no agent able to write “on interruption,” and the stated lifecycle does not require checkpointing the note before a risky call or after each state transition | The exact session-bound state this feature is intended to preserve can still exist only in chat, so finding C remains unhandled on its primary error path | Define write-ahead transitions: create/replace the note before each gate call with current pass, branch/target slots, recovery budget and session data; update it after validation/disposition; remove it only after the cycle-closing commit succeeds +MAJOR | high | Task 2 minimum resume-note fields | Cycle identity, next pass number, one recovery flag, and “any session id worth resuming” are insufficient for a deterministic Gate-B full recovery, which may have two branch session IDs, one failed branch, distinct target files, a validation failure cause, and a recovery attempt scoped to the current pass | A resumed agent can delete the wrong targets, rerun both branches unnecessarily, omit reviewType, or spend/reuse the recovery budget incorrectly, contradicting the existing bounded-recovery protocol | Require current pass, per-branch status, exact target path(s), validation state/cause, specSessionId/qualitySessionId with reviewType, and whether the current pass’s shared recovery attempt is spent +MAJOR | high | Task 1 raw-finding nudge combined with Task 2 optional dispositions | The close-time nudge needs to know which Blocker/Major findings were actually fixed, but findings files contain findings rather than verdicts and dispositions remain optional; after context loss, neither the resume-note minimum nor another durable artifact preserves accepted/fixed versus rejected findings | A resumed cycle closer can silently skip eligible findings or harden rejected ones, recreating session-bound context loss at the new mandatory step | Make dispositions durable for every Blocker/Major acted on, or add the accepted/fixed finding list and candidate classes to the resume-note/checkpoint protocol +MAJOR | high | Task 7, “a review comment whose reviewed-commit matches the merge head” | The verification object and query are underspecified: issue comments have no reviewed commit, inline comments expose commit_id but do not prove a completed whole-head review, while pull-request reviews expose commit_id through `gh api .../pulls//reviews` and `gh pr view --json reviews`; “review comment” could therefore be implemented against the wrong channel or accept an incidental bot comment | The document could replace one false completion proof with another and still merge an unreviewed head | Specify an exact CodeRabbit-authored pull-request review record, the exact `gh`/API field and author login, comparison to the PR head OID immediately before merge, and rejection of rate-limited/non-review records; document the verified one-shot re-trigger command too +MINOR | high | Task 2 Credit placement and Task 8 CHANGELOG rule | Task 2 requires the infinite-portfolio-canvas credit in CHANGELOG but lists only CLAUDE.md and workflow-init.md as files, while Task 8 requires feature names without requiring that credit | An implementer can satisfy every per-task Files/Verify instruction yet omit a settled attribution requirement | Add plugins/dev-workflow/CHANGELOG.md to Task 2’s Files/Verify scope or explicitly require the credit in Task 8’s new 0.6.0 entry +MINOR | medium | Task 1 pending outcome, “record that,” versus “The look is read-only” | The plan does not name where the prerequisite-blocked result is recorded, and writing it to the ledger would violate both the read-only close and harden-finding’s rule not to append another pending row merely for rediscovery | Different implementers may write into the closing diff, leave only ephemeral chat, or create duplicate ledger state | State that the result is carried in the durable follow-up/resume/disposition artifact with the existing row’s ref, and that the close performs no ledger or product-file write +END OF FINDINGS (8 total) diff --git a/.context/codex-reviews/gate-a-plan-pr2-pass-3.md b/.context/codex-reviews/gate-a-plan-pr2-pass-3.md new file mode 100644 index 0000000..ca84580 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pr2-pass-3.md @@ -0,0 +1,8 @@ +MAJOR | high | Task 1 lines 42-62 | The close greps a merely candidate class before harden-finding performs canonical fingerprinting, so an alias or near-match can report no match even when the canonical fingerprint already has a real or pending ledger row | The nudge can misclassify a recurrence as new or bypass a prerequisite, defeating the observable anchor it is meant to add | Canonicalize against both taxonomies before the grep, or carry the raw finding to one follow-up whose first action is canonicalization plus the authoritative grep and make no close-time outcome claim +MAJOR | high | Task 1 lines 46-50 | Grouping is keyed by candidate class even though two different candidates can canonicalize to the same fingerprint | Separate follow-ups can still make the second finding look like a recurrence against the mechanism created for the first, so the stated fix for M1 is incomplete | Group only after canonical fingerprinting, or make the follow-up intake canonicalize and coalesce all carried findings before any row is appended +MAJOR | high | Task 1 lines 42-50 and Task 2 lines 120-126 | The plan requires one follow-up for a group of findings, but harden-finding's shipped intake, application, severity, and log contract are explicitly one finding and one row | An implementer has no defined rule for choosing the group's source, severity, mechanism, finding text, or single ledger row, so grouping can silently discard distinctions or invoke the skill repeatedly and recreate the false recurrence | Define a group handoff that preserves every raw finding and tell the follow-up to run one canonicalization/coalescing decision before one harden-finding invocation, including deterministic severity/source and row-text rules +MAJOR | high | Task 2 lines 95-130 | Resume notes are called optional, advisory, deletable, and unnecessary for a zero-finding pass while the same section requires creating or replacing one before every gate call and relies on it as the only durable accepted/fixed list for the close nudge | The lifecycle cannot be followed literally, and treating the note as optional permits the exact context loss this requirement claims to prevent | Separate the optional dispositions file from a required write-ahead recovery note, require the note before every call including calls that later return zero findings, and limit deletion/rebuild to an explicit recovery procedure +MAJOR | medium | Task 2 lines 100-119 | The plan treats Gate-B spec and quality as independent cycles needing separate resume files because the reviewers run concurrently, but the outer agent writes the notes and a full pass has one shared recovery budget and coupled two-file validation | Two notes can disagree about shared-attempt state, both session IDs, validation state, or which branch completed, leaving recovery with no authoritative record; the claimed findings-file writer race does not apply to agent-owned resume notes | Use one cycle-stable Gate-B resume note containing both branch records and shared state, or specify an atomic ownership and reconciliation protocol that makes one of the two files authoritative for shared fields +MAJOR | high | Task 7 lines 253-260 | The prescribed gh query emits only commit_id and state, yet the acceptance rule also rejects review records whose body says the review was rate limited | Following the exact command cannot evaluate one of the required rejection conditions, so it can bless a rate-limited record on the final head | Query body as well and filter it explicitly, for example by returning commit_id, state, and body and rejecting the rate-limit marker before accepting a head match +MINOR | high | Task 7 lines 253-260 | The plan requires equality with the PR head OID read immediately before merge but specifies no command or object for obtaining that OID, despite otherwise presenting the verification as exact | Different operators may compare against local HEAD, a stale checkout, or the merge commit and reproduce the false-proof class this task is meant to close | Name an exact API query for the live PR head OID immediately before merge and show the comparison against the selected pull-request review record +END OF FINDINGS (7 total) diff --git a/.context/codex-reviews/gate-a-plan-pr2-pass-4.md b/.context/codex-reviews/gate-a-plan-pr2-pass-4.md new file mode 100644 index 0000000..2b5c395 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pr2-pass-4.md @@ -0,0 +1,10 @@ +MAJOR | high | Task 1 lines 42-49 and 101-104 | the nudge says to hand any Blocker/Major finding to a follow-up cycle but never defines whether multiple findings are handed over individually, together, or as separate cycles, even though harden-finding requires one finding text, source, and severity per invocation | the close-time classifier is gone, but the prior grouped-handoff failure survives at the remaining boundary and an implementer cannot deterministically invoke the one-finding-one-row skill when a cycle fixed several findings | say “for each Blocker/Major finding fixed this cycle, start a separate harden-finding invocation after closing this cycle” and require passing that finding’s text, source, and severity +MAJOR | high | Task 1 lines 37-49 and 62-71 | “Before the cycle-closing commit” orders the handoff before the current cycle closes, while “follow-up cycle” and “hardening cannot invalidate the reviewed diff” require hardening to begin only after that commit; “hand it” does not distinguish scheduling the work from executing the skill | a literal implementer can run harden-finding before the closing commit, change product files, and invalidate the reviewed diff, directly defeating the claimed isolation | make the sequence explicit: retain the findings before the closing commit, close the reviewed cycle, then invoke harden-finding in a new cycle before ending the session +MAJOR | high | Task 1 lines 59-64 and Verify lines 70-71 | the plan claims “the close stays observable and terminating,” “the look terminates (read-only, once),” and requires the shipped text to say “read-only,” but the close-time look was deleted and the remaining handoff is neither a read-only look nor defined as an observable artifact or tool result | these are stale claims from the deleted design and would force implementation text that misdescribes the mechanism, violating prompt-standard item 11 | delete the look/read-only/observable claims and verify only the actual ordered handoff, or define a concrete non-classifying observation whose result proves the handoff occurred +MAJOR | high | Task 1 lines 33-49 and 101-104 | the replacement still depends on volatile same-session memory to know every fixed finding at close time and explicitly accepts losing that set; a normal context compaction, interruption, or agent handoff can therefore erase the only input before the nudge or between commit and follow-up | finding A is only partially served: projects without PRs still have no durable or enforced route to harden findings, so the ledger can remain empty despite cycles that fixed Blocker/Major findings | either require a minimal durable per-finding handoff queue that is independent of the optional resume note, or narrow the stated outcome to a best-effort same-session reminder and leave finding A unresolved/pending rather than logging it as hardened +MAJOR | high | Task 1 lines 42-49 versus process-pr-review step 5 | the new non-PR path covers only Blocker/Major findings, while the existing PR-only mandated hardening step covers every accepted actionable finding fixed there and harden-finding accepts minor/nit severities | the plan claims to address the defect that the only mandated ledger check is PR-only, but it leaves fixed lower-severity findings with exactly the same PR-only gap and can still produce an empty ledger in a non-PR project | apply the handoff to every finding actually fixed during the cycle, while retaining Blocker/Major only as the review-loop iteration threshold +MAJOR | high | Task 7 lines 235-247 | the jq filter tests the lowercase case-sensitive phrase `rate limited`, but the observed rejection text is quoted as `Review rate limited` with an uppercase R, so that record passes the filter; the #13 verification only exercised the mismatched commit_id path and did not verify this rejection branch | a rate-limited review on the live head would be accepted as proof of review, recreating the exact false-proof failure this task is meant to close | use a case-insensitive null-safe predicate such as `select((.body // \"\") \| test(\"review rate limited\"; \"i\") \| not)` and verify it against a fixture or real record whose commit_id equals the queried head +MINOR | high | Task 7 lines 231-247 | the query emits `.state` for any matching review record but defines no accepted state or success test for zero, one, or multiple output lines | an implementer can treat any output—including a dismissed or otherwise non-counting review state—as proof, and the documentation still does not give a deterministic boolean check | specify the allowed CodeRabbit review state(s), make the command exit nonzero when no qualifying record exists, and define how multiple records are reduced +MINOR | medium | Task 4 lines 151-161 | “neither duplicates an existing class” names no required comparison against both the project taxonomy and the base skill lists, even though harden-finding’s minting contract requires grepping both lists for the closest class and aliases before minting | an implementer following only this plan can add near-duplicate fingerprints and corrupt recurrence grouping | require the skill’s two-list closest-match grep before adding either class and record the compared near matches in verification +MINOR | high | Task 6 lines 185-213 | “four entries, edit only” does not identify the destination rows/sections for three new parked stories or require preservation of the trigger-gated backlog structure, and only the calibration item names an existing row | two implementers can place or phrase the stories differently, accidentally create active work under Now, or duplicate an existing parked concern while both satisfying the plan | name the exact todos.md heading and insertion/update target for each item, state that a/c/d are new unchecked trigger-gated rows and b mutates the existing P2 row, and add a no-duplicate verification +END OF FINDINGS (9 total) diff --git a/.context/codex-reviews/gate-a-plan-pr2-pass-5.md b/.context/codex-reviews/gate-a-plan-pr2-pass-5.md new file mode 100644 index 0000000..6aee46f --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pr2-pass-5.md @@ -0,0 +1,6 @@ +MAJOR | high | Task 1, Rule and "What this does and does not do" | Task 1 is not sound as written because it says to run harden-finding once for every accepted actionable finding and repeatedly claims this is exactly the set process-pr-review step 5 hardens, but step 5 invokes the skill only when a finding matches an existing class or a new class is clearly warranted | The narrowed reminder still expands the non-PR path beyond the settled PR-path contract, so an agent can be instructed to harden one-off findings that step 5 would only inspect and decline; the explicit exact-scope claim is therefore false | Mirror step 5's condition in the reminder: check every accepted actionable fixed finding, but invoke harden-finding once per finding only when it matches an existing class or a new class is clearly warranted, and revise the exact-match claims accordingly +MAJOR | high | Task 7, proposed pull-request-reviews query | The query reads only the first page of the reviews endpoint | A PR with more than the API page size of review records can have the qualifying final-head review on a later page and be falsely classified as unreviewed, sending the workflow into an unnecessary retrigger or human-override path | Make pagination explicit and aggregate all pages into one boolean, for example with gh api --paginate --slurp and a jq expression over the combined arrays; verify both a qualifying record after page one and an all-pages-no-match case +MAJOR | medium | Tasks 4-5, taxonomy minting and ledger append verification | The plan requires a closest-match grep and an anchored postcondition, but omits harden-finding's concurrency rule to re-read the ledger immediately before each append and append only if no matching fingerprint appeared meanwhile | Another branch or agent can add the class or row after the initial grep, producing a duplicate taxonomy class or duplicate first-occurrence ledger hardening and corrupting recurrence counts in the append-only ledger | Require a fresh read of both taxonomy lists before minting and a fresh read plus anchored column-2 grep immediately before each row append; if a match appeared, reuse or reconcile it instead of appending +MINOR | high | Task 7, paragraph claiming a rate-limited record carries a commit id | The stated reason for filtering review bodies overclaims the measured mechanism: on PRs #12 and #13 the rate-limit warning is in the CodeRabbit issue comment, while the pull-request review record is an earlier completed review; the plan provides no evidence that a rate-limited pull-request review record itself carries commit_id | This repeats the repository's recurring failure mode of turning a defensive check into a claim about what was observed, and future maintainers may believe the query directly validates the warning object when it does not | State the measured object split exactly: a qualifying completed review record must exist for the live head, while the rate-limit warning was observed in the issue comment; retain the case-insensitive body rejection only as bounded defensive filtering unless an actual rate-limited review record is measured +NIT | high | Task 1, "What this does and does not do" | The plan calls the proposed addition "three sentences", but the quoted shipped text contains seven sentences across three paragraphs | This is a small internal inconsistency in a plan explicitly narrowing the feature on the basis of its tiny prompt cost | Say "three short paragraphs" or remove the sentence-count claim +END OF FINDINGS (5 total) diff --git a/.context/codex-reviews/gate-a-plan-pr2-pass-6.md b/.context/codex-reviews/gate-a-plan-pr2-pass-6.md new file mode 100644 index 0000000..bc84ed0 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pr2-pass-6.md @@ -0,0 +1,6 @@ +MAJOR | high | Task 1 lines 78-85 | The optional-companion task still says “Task 1's nudge” hands fixed findings to `harden-finding` in the same session, even though Finding A and its nudge were cut entirely | This is a direct dangling behavioral requirement and makes C appear dependent on the removed same-session hardening path, contradicting the plan's claim that C ships independently | Delete the removed-nudge history and replacement-behavior sentences; retain only the relevant explanation that companions are optional, advisory, non-validating, and non-load-bearing +MINOR | high | Task 3 lines 132-153 | A one-class task still says “each with alias hints,” “either class,” and “Neither duplicates,” preserving the prior two-class shape | The singular scope and plural verification instructions disagree, so an implementer cannot tell whether a second class is omitted or the prose is stale | Change these phrases to singular and keep the two-list search wording only where it refers to the base and project taxonomy lists +MAJOR | high | Task 4 lines 157-170 | The task is titled “one appended row” and shows one C row, but its rule repeatedly requires “Both rows,” discusses “The A row,” and says “these rows” | Finding A's ledger row was supposed to be removed; following the rule would either reintroduce cut scope or leave the implementer inventing a nonexistent second row, breaking the claimed one-row release shape | Rewrite the rule entirely in the singular for the C row and remove every A-row reference +MAJOR | high | Task 6 lines 251-265 | The concrete command that the documentation is told to use is still the first-page-only `gh api ... --jq '[ .[] ... ]'` form, while the following paragraph requires `--paginate --slurp` and a combined-array query | The plan simultaneously specifies a known-bad executable example and its correction; an implementer can copy the fenced command and ship the exact false-negative path pass 5 was meant to close | Replace the fenced command itself with the full `gh api --paginate --slurp ... --jq '[ .[][] ... ] \| length'` form, then verify that exact command with the page-two-match and all-pages-no-match cases +MAJOR | high | Task 7 lines 303-308 | The release task still requires the CHANGELOG to name “A's nudge” even though the plan says A is cut and the PR ships only C, the template sync, and release housekeeping | This would publish a false user-visible claim that 0.6.0 includes behavior it deliberately does not ship, violating the plan's scope statement and the repository's rule against claims outrunning mechanisms | Remove A's nudge from the CHANGELOG requirement and require entries only for C, the downstream-neutral template sync, the taxonomy/ledger/todo/doc changes as appropriate, and the release bump +END OF FINDINGS (5 total) diff --git a/.context/codex-reviews/gate-a-plan-pr2-pass-7.md b/.context/codex-reviews/gate-a-plan-pr2-pass-7.md new file mode 100644 index 0000000..03420e5 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pr2-pass-7.md @@ -0,0 +1,4 @@ +MAJOR | high | Task 2 and Task 5 | the plan implements the deferred ad-hoc-brief template sync but never removes or resolves the existing unchecked todos.md row that says to perform that sync when workflow-init is next touched | after implementation the backlog will still advertise completed work as pending and retain a now-false “upcoming” resolution vehicle, creating exactly the dangling reference left by a cut or completed task | add a Task 5 edit that marks the existing template-sync row resolved or removes it, and update the stated edit counts accordingly +MAJOR | high | opening scope summary and Finding A closing paragraph | “This PR takes” and “This PR therefore ships” enumerate only C, the template sync, and the release, omitting the planned taxonomy class, ledger row, four todos edits, and pr-review-bots caveat | the plan contradicts both its own Tasks 3–6 and the settled seven-part scope, so an implementer or reviewer using the declared scope can incorrectly treat required work as incidental or out of scope | replace both summaries with the complete settled scope: C, template sync, one taxonomy class, one ledger row, four todos edits, the pr-review-bots caveat, and the 0.6.0 release +MINOR | high | Task 7 “CHANGELOG entry” | “omitting the third would under-report” is a dangling cardinality reference even though the sentence now names only C's companions and the template synchronization | it is stale cut-era text and makes the release requirement ambiguous about an unnamed third CHANGELOG item, inviting either accidental A inclusion or another invented entry | change it to “omitting either would under-report” or explicitly enumerate the two required CHANGELOG subjects +END OF FINDINGS (3 total) diff --git a/.context/codex-reviews/gate-a-plan-pr2-pass-8.md b/.context/codex-reviews/gate-a-plan-pr2-pass-8.md new file mode 100644 index 0000000..8d781f4 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-pr2-pass-8.md @@ -0,0 +1,2 @@ +NO FINDINGS +END OF FINDINGS (0 total) diff --git a/.context/codex-reviews/gate-a-plan-resume.md b/.context/codex-reviews/gate-a-plan-resume.md new file mode 100644 index 0000000..4a3a43f --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-resume.md @@ -0,0 +1,57 @@ +# Gate-A plan cycle — Plan C (rle) — CYCLE ENDED at pass 7, NOT CLEAN + +Artifact: docs/superpowers/plans/2026-08-30-review-loop-economics-plan-c-rollout.md +Revision 7, commit bcd45dc. Passes 1..7 present and valid. NOT closed clean. + +## Why it ended here +Daniel's bounded contract before pass 7: clean or Blocker/Major-free -> close per §5; +not clean -> no pass 8, full numbers to Daniel, cycle ends within the contract. +Pass 7 returned 21 B+M. The cycle ends. No further repair was made. + +## Curve — the whole cycle +pass | total | BLOCKER | MAJOR | B+M +1 | 18 | 5 | 11 | 16 +2 | 20 | 4 | 13 | 17 +3 | 20 | 2 | 13 | 15 +4 | 23 | 4 | 16 | 20 +5 | 22 | 5 | 13 | 18 +6 | 19 | 2 | 16 | 18 +7 | 29 | 2 | 19 | 21 + +Highest total and highest B+M of the run are both pass 7, the last one. +Blocker/Major never left the 15-21 band across seven passes and seven revisions. + +## How the six swept claims fared +1. hook stores/counts — FIXED, not re-contested. +2. parser existence — STILL OPEN. The sweep found the spec site and missed two more: + Plan B leaves a sentence in BOTH shipped prompt copies saying the provenance form + has no informal variant "because the deferred metrics work parses it". The sweep + for incompleteness was itself incomplete. +3. slot validity — WORSE. Both Blockers land here. Shipping the discriminator + production into the prompt is called out of Plan C's §7/§8 scope AND unreachable: + this cycle runs under the OLD rules by the plan's own activation constraint, so a + production shipping in this same commit cannot bind it. The resolution is + self-contradictory, and I introduced that while sweeping for self-contradictions. +4. grep limits — IMPROVED, still incomplete. Per-pass model-key duplication and + ordering are uncovered; cross-record nonce agreement has only a positive case. +5. preflight exit code — FIXED. The uninstantiated template is re-raised from pass 6. +6. field report completion — the NEW task generated two findings: it updates the + curve but not the analysis prose a clean close would falsify, and its assert + checks removal without checking insertion. + +## Standing structural findings, unrepaired +- No task instantiates the preflight; the only one is a placeholder template. +- No task constructs or stores the evidence entry every call must quote verbatim. +- The knob before-state has no storage location or resumption rule. +- The grammar check has no actual EREs, fixtures or fail-capable assertions. +- The mktemp closing-body obligation is declared moot while a multiline body still + has to be composed with no named mechanism. +- The risk-path verification reads one Story: header where the shipped rule requires + the union across the spec and all three plans. + +## Named fallback, chosen against for now +Split Plan C by statement site. + +## Resume procedure if the cycle is reopened +Next pass is 8: delete .context/codex-reviews/gate-a-plan-planc-pass-8.md, confirm +gone, re-run the same broad Gate-A prompt (risk-high lens set, security none). diff --git a/.context/codex-reviews/gate-a-plan-rle-pass-1-dispositions.md b/.context/codex-reviews/gate-a-plan-rle-pass-1-dispositions.md new file mode 100644 index 0000000..6b72f56 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-rle-pass-1-dispositions.md @@ -0,0 +1,70 @@ +# Gate A — plan cycle (rle) — pass 1 dispositions + +Advisory human note. Not a findings file; participates in no pass validation. + +## Count restart + +This is **pass 1 of a restarted count**. The previous pass 1 ran against the plan +committed at 1470094 and returned 21 findings, 20 of them Blocker/Major; its file is +preserved at `gate-a-plan-rle-pass-1.superseded-pre-rewrite.md` rather than deleted, +because it holds the evidence that motivated the rewrite. + +**Restart reason:** the artifact was rewritten, not repaired (5c00f8c). A carried count +would flatter the new artifact — passes spent reading a document that no longer exists +say nothing about the one that replaced it. + +## Pass 1 result — VALID, and it routes back + +- File: `gate-a-plan-rle-pass-1.md`, terminator `END OF FINDINGS (33 total)`, 33 finding + lines, 0 non-finding lines. **Valid pass.** +- Severity mix: **21 BLOCKER · 10 MAJOR · 1 MINOR · 1 NIT** → **31 Blocker/Major**. +- Routed contract for this rewrite: converge in 2-3 passes at <= 8 B/M; **above ~12 + routes back before pass 2**. 31 is nearly triple that. **Routed, not iterated.** + +This is not the "clearly stuck" exit — that needs a plateau across passes and this is +pass 1. It is the pre-agreed routing threshold firing on its first reading. The findings +stay open, no pass is credited as clean, and the loop resumes on whatever is decided. + +## The prediction was wrong, and how + +Predicted <= 8 B/M on the grounds that the four habits behind the predecessor's 20 were +structurally prevented. Three of the four recurred in new form: + +- **"checks written to look like TDD"** -> B8. Task 2's recorded `7/7` is a real + measurement of the tree *today*, but the plan runs it *after* Task 1 has already + removed one of the seven matches. In sequence it is 6, not 7. The number was honest + and the sequence was not, which is the same defect wearing a demonstrated output. +- **"a plan for a §5 change that does not follow §5"** -> B15, B16, B14. Fixed + "nine ordinary commits" and introduced an amend the hook cannot recognize. +- **"anchors typed from memory"** -> NIT 12. One line of one anchor block is the + editorial token `(identical)` rather than pasted output, so the blanket claim that + every anchor is pasted is not literally true. + +Only "references to spec contents that no longer exist" was actually prevented, and +B2 argues even the passage list is unsound in the other direction. + +## Findings verified before routing + +Two were checked against source rather than accepted on the reviewer's word. + +- **B15 — CONFIRMED, and decisive.** `plugins/dev-workflow/hooks/codex-gate.sh:763`: + `is_wip_commit() { printf '%s' "$1" | grep -Eiq -- "-m[[:space:]]*['\"]?[[:space:]]*wip"; }` + It matches the **Bash command string**, not git state or the commit message. So + `git commit --amend --no-edit` carries no `-m` and is NOT a WIP commit to the hook; + at `:886` `is_commit "$cmd" && ! is_wip_commit "$cmd"` is then true and the cycle + RESETS. Tasks 2-6 each run exactly that command. The plan's central structural fix + destroys the cycle it was written to protect. + Corroborated incidentally this session: a `cat` heredoc merely *containing* the text + `WIP:` fired the hook's WIP notice, because the match is on the command string. +- **MAJOR 10 — CONFIRMED as stated.** `check-version-bump.sh` does compare committed + state; the plan's ordering fix is right. The finding is that the plan never checks + `main` is current, which `AGENTS.md` names as a precondition. Correct and additive. + +## Disposition + +All 31 Blocker/Major carried open to the routing decision. No fixes applied in this +pass: at this density the artifact is being re-decided, not repaired, and applying 31 +repairs before that decision is how a rewrite becomes a patch pile. + +The two Minor/Nit are collected, not iterated: MINOR 11 (CHANGELOG entry content +unspecified), NIT 12 (the `(identical)` anchor line). diff --git a/.context/codex-reviews/gate-a-plan-rle-pass-1.md b/.context/codex-reviews/gate-a-plan-rle-pass-1.md new file mode 100644 index 0000000..5ea918b --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-rle-pass-1.md @@ -0,0 +1,34 @@ +BLOCKER | high | Task 0 opening and Steps 1-5 | the conditions artifact is absent in the reviewed tree and is scheduled to be created only after this Gate-A plan cycle closes, although spec §6 requires it to be reviewed alongside the plan before any replacement text is written | execution must either start without a required accepted input or pretend this pass reviewed a file that did not exist | create and populate the artifact now, include it in this Gate-A plan cycle, and begin executable tasks only after the cycle accepts both files +BLOCKER | high | Task 0 "The nineteen passages — provenance stated" | the claimed reconciliation is not sound: it imports revision 15's eighteen rows and adds only the pass-report row, but revision 36 also reverses row 14's premise by excluding human-exception records from nonce scope and adds substantial profile, recovery, activation, revert, and adoption rules; the exact-once grep proves only that the chosen lead-ins exist, not that the set is complete | the sole guard against dropped old conditions can certify a stale, overinclusive, or incomplete passage set | derive the set from the actual planned edits against current §5, remove obsolete rows, add every current passage actually rewritten, and fail a canonical set comparison for missing, extra, or duplicate entries +MAJOR | high | Task 0 Step 3 | one row per passage with one singular "existing prose" column does not implement spec §6's requirement to account for the two copies separately wherever their current text differs | a condition unique to one mirror can be silently dropped while the shared row still reads complete | give each copy its own requirement and disposition rows or prove and record byte identity passage by passage before using a shared row +BLOCKER | high | Task 0 passage 19 and Task 1 Step 3 | Task 0 says the change rewrites `From pass 4 onward every pass report carries three lines`, but no implementation step edits that paragraph; Task 1 instead adds new every-pass fields to the floor paragraph | the shipped prompt can simultaneously require new fields on every pass and still prescribe the old three-line pass-4 carrier, leaving the report shape contradictory or underspecified | add an explicit anchored edit for the pass-report paragraph that integrates the new fields while preserving and dispositioning every existing trend, cluster, and require-withdraw condition +MAJOR | high | plan header `Story:` | the plan copies `risk high`, `security none`, and the validation mode into itself even though current §5 makes the story header the single writable copy and says plans carry the path and read values fresh | a later confirmed profile change can leave the plan executing stale lenses, evidence, or a stale pass floor, which is the central gate-off risk of this change | keep only the story path in the plan and require each pass and execution step to read the current header +BLOCKER | high | Task 1 Step 2 | the instruction says to replace only `min 3 passes per run` and keep every other requirement verbatim, while the supplied block replaces the whole bold lead and Task 1 Step 3 requires the hook ratio to control nothing; preserving the adjacent `counted by the hook` clause contradicts the new rule, while replacing the whole lead silently drops a condition without saying so | two literal executors can ship opposite floor authorities, one of which lets the hook knob govern closure | delimit the exact old range being replaced, explicitly disposition `counted by the hook`, and provide the complete resulting paragraph including the preserved tail +BLOCKER | high | Task 1 Step 3 gate-off disclosure | the proposed text lists four routes as "all routes" but neither labels the list non-exhaustive nor carries the spec's other known routes such as falsified evidence, silenced reminders, or reporting an unrun pass as run | readers can treat an incomplete enumeration as complete, precisely the gate-overclaim forbidden by AGENTS.md, and overlook another way to state too few passes | use the spec's "routes known today, not a complete list" framing, include every currently named route, and retain the explicit no-guard statement +BLOCKER | high | Task 1 and Self-Review mapping for spec §2.4 | no step adds the settled current-profile and current-cited-set behavior: re-read membership each pass, require a further clean pass after any profile or membership change even when the floor is unchanged, recompute all derived duties, and never discharge an accepted in-set Blocker or Major by removing a citation | a mid-cycle profile or story-set edit can silently reuse an old clean pass and close under obligations that were never reviewed | add exact mirrored edits anchored at `Changing a profile:` and account every old condition there +BLOCKER | high | Self-Review mapping for spec §10 | the plan claims Task 1 ships §10, but no task adds activation at the shipping commit, the five-part unknown-start strict fallback, revert precedence, or downstream partial-adoption behavior | cycles that straddle shipping, reversion, or partial scaffolder adoption have no defined rule and can take the cheaper floor or demotion on uncertainty | add a production-ready mirrored activation block implementing every §10 branch and include it in the passage accounting and parity review +BLOCKER | high | Task 2 Steps 1-3 | the recorded pre-edit output cannot occur in sequence: after Task 1 removes the line-72 `min 3 passes` match the command returns 6 per copy, not 7, and the regex never matches the separate `pass 1 carrying a Minor` site at all | the demonstrated output is false and the check reaches zero even if the central floor-1 residual remains unchanged | record the actual 6-per-copy sequential output and add a separately asserted pre/post check for the `pass 1` sentence, or use a complete canonical inventory +MAJOR | high | Task 2 Step 2 replacement table | four replacements are written with literal ellipses rather than complete resulting text despite the Self-Review claim that every replacement is written out | an executor must invent what the ellipses retain and can drop neighboring closure conditions in the exact rewrite class AGENTS.md warns about | provide exact old and exact new strings, including line breaks, for all seven sites +MAJOR | high | Task 2 Step 2 Gate-B replacement | `a fix changes the diff and invalidates the prior pass, which is where the floor's lower bound comes from` is false under the new design: the lower bound comes from the story profile, while invalidation only explains why a fix requires another review | the prompt misstates what the gate comparison proves and supplies a causal rationale incompatible with floor 1 | say that a fix requires re-review because the reviewed revision changed, and state separately that the profile derives the floor +MAJOR | medium | Task 2 Step 2 Lenses replacement | changing `The 3-pass floor ... is unchanged` to `The floor ... is unchanged` still tells readers the floor is unchanged in the very change that makes it profile-dependent | readers can retain the fixed-three interpretation or conclude that axes may add lenses but never alter pass count | say lenses are not additional passes beyond the independently derived floor, then list only the rules that truly remain unchanged +BLOCKER | high | Task 4 Steps 2-4 | the largest shipped insertion has no destination anchor and no complete production text; `copy grammars` plus summary bullets leaves placement, section structure, and rationale to the executor | the two prompt copies can acquire different rule order, scope, or wording and still satisfy the token-count greps | provide unique current-file anchors and the exact complete resulting blocks for each mirror +BLOCKER | high | Task 4 Step 4 | the spec delegates slot spelling to the plan, but the plan says only "optional per-cycle infix" and never defines the filenames for nonce-bearing findings slots or advisory working records | concurrent cycles can invent incompatible names, collide, or fail to recover each other's records | pin the exact three findings-slot productions and three working-record names, including nonce placement and the reserved legacy forms +BLOCKER | high | Task 4 Steps 2-3 per-pass curve | copying the grammar and adding four summary properties omits settled operational rules: one record for each cycle type, logical-pass aggregation, incomplete-pass exclusion with consumed numbers, same-tracked-revision branching, split-model completeness, revision mismatch behavior, unknown-model handling, and the self-reported limitation | a syntactically valid curve can merge different revisions, omit contributors, or claim evidence the spec explicitly disallows | enumerate and ship every §4 required property around the pinned grammar, with dispositions for the existing commit-body rules it extends +BLOCKER | high | Task 4 Step 3 nonce | the nonce summary omits immutability and uniform generation, candidate scoping by cycle type and artifact, the two recovery sources, agreeing versus disagreeing sources, closed-cycle exclusion, working-record retirement, the cost of starting fresh, three-distinct-cycle behavior, and the bounded pre-rule exception | interruption and concurrency paths remain ambiguous and can reuse another cycle's identity or silently retain banked passes | translate every settled §5 nonce property into exact mirrored prompt text and specify the recovery state table +BLOCKER | high | Task 4 Step 5 parse check | no parser, command, acceptance algorithm, fixture strings, or demonstrated output is supplied for the promised accept-and-reject checks | the check cannot be run, so a malformed pinned grammar can be declared parsed by inspection and the predecessor's non-failing-test defect remains | include a POSIX-sh-invoked parser or deterministic validator plus explicit valid and invalid fixtures and asserted outputs +MAJOR | high | Task 4 Step 3 and Task 7 prompt-standards item 10 | nonce failures are collapsed into one retry-and-stop response without the cause-specific checks and fixes required for diagnostic states by prompt-standards item 10 | an unavailable randomness source, malformed generator output, and a live-cycle collision present the same stop with no actionable diagnosis, violating invariant 11 | require the shipped prompt to name how each cause is distinguished and the distinct corrective action before retry or surface +MAJOR | high | Task 7 Steps 1-4 | the parity and twelve-item results have no target section, row schema, required status vocabulary, or per-rule and per-artifact fields | an executor can append an unauditable paragraph or nothing material and still make the docs-only commit | define exact tables with one row per changed rule and one status-plus-reason row per checklist item for each artifact, then assert their row sets +BLOCKER | high | Global commit protocol and Self-Review | the plan gives three incompatible cycle memberships: the opening protocol includes Tasks 1-5, 8, and 9; Task 6 amends the WIP; Task 8 explicitly commits N/A; and the Self-Review instead names Tasks 1-6 and 9 | there is no single executable commit topology, so the executor cannot know which content the Gate-B cycle reviews or closes | choose one authoritative ordered topology and make every task header, commit step, and summary match it +BLOCKER | high | Tasks 2-6 amend commands | `git commit --amend --no-edit` preserves the resulting commit subject but the actual hook recognizes WIP only from a `-m ... WIP` token in the Bash command; the command was mechanically classified as NON-WIP | every amend is read as a real cycle boundary and resets Gate-B state despite the plan's claim that the prefix is carried | amend with an explicit `-m` value beginning `WIP:` while safely carrying the full body, or change the topology so no cycle-internal commit command is misclassified +BLOCKER | high | Task 7 Step 4, Task 8 Step 4, and Task 9 Step 3 | the two ordinary docs commits occur after the WIP snapshot; any non-WIP commit resets the hook regardless of staged paths, advances HEAD beyond the WIP, and makes Task 9's `--amend` amend the Task-8 docs commit rather than the WIP | the WIP remains stranded in history, packaging is folded into the wrong commit, and the closing amend cannot close the claimed cycle | place all ordinary commits before opening the WIP or include their content in the one WIP and close it before any later ordinary commit +BLOCKER | high | Task 8 Step 2 getting-started line 58 correction | `N/N cycle ... N being the derived floor` falsely identifies the hook's displayed denominator as the derived obligation even though the spec and Task 1 say it is the independent reminder threshold | a floor-1 cycle with the default hook threshold can wait for 3/3 and silently run extra passes, while a custom threshold can be mistaken for the obligation | describe the displayed ratio as the hook threshold and state that cycle closure follows the profile-derived floor even when the reminder remains below threshold +BLOCKER | high | Task 9 Step 5 fix loop | the plan says to re-review after each fix but never stages and amends the WIP before the next `mcp__codex__review`; that tool reviews the committed git range, not uncommitted worktree fixes | a re-review can inspect the old snapshot and produce a clean pass over content different from the closing amend | after every accepted fix and evidence update, stage and explicitly amend the WIP, then review that new commit against the same base +BLOCKER | high | Task 9 Steps 4-6 and Task 10 placement | Task 9 runs Gate B and closes the commit before the later Task 10 produces and revalidates the evidence pack that §5 requires in every Gate-B call and before closure | reviewers cannot receive the current evidence entry and the close cannot carry evidence verified against its final content | move Task 10's production and revalidation steps before the first Gate-B call, repeat them after each fix, and close only after the final revalidation +BLOCKER | high | Task 9 Step 6 | the executable closing command contains a literal `` placeholder even though the Self-Review claims no placeholders | literal execution destroys the WIP body and records neither the evidence entry, provenance line, nor curve | provide the fully instantiated body or an explicit build-and-validate step that substitutes real records before invoking the amend +BLOCKER | high | Task 10 Step 5 | the plan requires provenance and curves in the already-closed Gate-A spec and current Gate-A plan commits, but supplies no commit identities, history-rewrite procedure, rollback point, or verification; the actual revision-36 and plan commit bodies contain neither pinned form | the three-body acceptance criterion is impossible to satisfy safely and an ad hoc rebase can lose the long review history | name each owning commit, specify a recoverable rewrite sequence before descendant work, and verify those exact three bodies rather than an arbitrary recent range +MAJOR | high | Task 10 Step 5 future nonce checkpoint | the text says the implementation commit records the future-cycle discharge obligation, but neither Task 9's closing body template nor any other step adds that obligation | the branch can close while falsely claiming the undemonstrable nonce field has a durable checkpoint | pin the exact checkpoint record and require it in the validated closing body +MAJOR | high | Tasks 1-6 interruption and rerun behavior | manual append-and-replace steps have no missing, partial, complete, or duplicate preflight and parity is deferred until Task 7 | abandonment after editing one mirror leaves the shipped copies disagreeing, while rerunning append tasks can duplicate classifiers, grammars, nonce rules, or the item-1 note | add per-task state detection with exact counts in both files, make completed tasks no-ops, and stop on asymmetric or duplicate state before editing +MAJOR | medium | Task 9 Steps 3-4 | `check-version-bump.sh main` does compare committed HEAD and ignore worktree and index as claimed, but the plan never verifies that `main` is current, which AGENTS.md states as a precondition | a stale base can validate against the wrong manifest version and make the recorded BUMP-OK evidence misleading | resolve and update the PR base before the cycle, record the merge-base used, and pass that current ref to both standalone and battery checks +MINOR | high | Task 9 Step 2 | `Add the CHANGELOG entry` specifies neither required content nor a check beyond position | an empty or materially incomplete `0.11.0` section satisfies the literal step while failing to explain the shipped prompt behavior | require the entry to cover the floor predicate, severity test, record forms, activation and adoption limits, then assert one newest `0.11.0` heading +NIT | high | Task 5 Step 2 anchor | the second line of the claimed pasted `grep -nF` output is the editorial token `(identical)` rather than the actual template line | the rewrite's blanket claim that every anchor is pasted output is not literally true, weakening mechanical traceability despite the verified line being identical | paste the full template grep line or label the block as one pasted line plus a separately verified equality assertion +END OF FINDINGS (33 total) diff --git a/.context/codex-reviews/gate-a-plan-rle-pass-1.superseded-pre-rewrite.md b/.context/codex-reviews/gate-a-plan-rle-pass-1.superseded-pre-rewrite.md new file mode 100644 index 0000000..f934e53 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-rle-pass-1.superseded-pre-rewrite.md @@ -0,0 +1,22 @@ +BLOCKER | high | Task 0 Interfaces and Step 3 | the cited spec §5.1 site inventory and §5.2 eighteen-row passage list do not exist; the spec instead says the plan carries the list, but the plan carries none | the hard-gated conditions artifact has no authoritative passage set and cannot be completed | put the complete eighteen-passage, per-copy inventory in the plan and make Task 0 consume that exact list +BLOCKER | high | Task 0 Step 4 | the proposed source count returns 0 now because its sed range matches no table, while `grep -c '^\| '` on the future artifact would also count its header and separator and no command compares the two sets | the completeness gate can neither produce the expected source count nor prove that every passage was dispositioned | compare canonical passage identifiers or lead-ins from both artifacts with a set diff and fail on any missing, extra, or duplicate entry +BLOCKER | high | Task 0 Steps 5-6 | Task 0 is placed inside plan execution, but its artifact must be accepted by this already-running Gate-A plan cycle before execution may begin and the artifact is absent at pass 1 | the workflow must either close Gate A without a required input or start executing an unapproved plan, so the stated gate cannot be followed | make the conditions artifact a pre-implementation input created now and include it in subsequent Gate-A plan passes; start the executable task list only after that cycle closes +MAJOR | high | Task 1 Step 1 | the supposed failing check already prints 1 for both files because each existing Profiles section contains `max(risk, security)` | the check has exactly the same expected result before and after the edit and proves nothing about the new floor predicate | scope the check to the HARD FLOOR paragraph or grep for a new complete sentence that is absent now +MAJOR | high | Task 1 Step 2 | the replacement says `The floor is counted by the hook`, but the settled rule says the derived floor is text-bound and the hook ratio is a reminder threshold that controls nothing | an executor can make the hook's independently configured threshold govern closure, reversing the settled precedence | replace that clause with the spec §2.1 precedence wording and state that the hook counts calls against its own threshold, not the derived obligation +MAJOR | high | Plan self-review and Tasks 1-4 | no proposed shipped wording implements the required pass-report fields, knob preservation and gate-off residual, current-profile and current-cited-set re-evaluation, accepted-finding retention, activation fallback, revert precedence, or partial downstream adoption; the self-review nevertheless maps spec §§2.1, 2.2, 2.4 and 10 to Tasks 1 and 4 | story criteria 1, 4, and 7 remain unsatisfied in both prompt copies | add exact mirrored replacement blocks and anchors for every omitted obligation, then map each story clause to the task and check that implements it +MAJOR | high | Task 2 Step 2 second replacement | the quoted current anchor `a Blocker/Major-free pass 1 carrying a Minor keeps looping` does not exist verbatim because both files split it between `keeps` and `looping` | a literal executor cannot apply the replacement as written | quote the actual two-line anchor including its newline or use a shorter unique verbatim anchor +MAJOR | high | Task 4 Steps 2-4 | `Copy §2.3, §4 and §5 of the spec's grammars verbatim` is not an executable edit: §5 has no grammar, the instruction does not delimit grammar blocks from repo-specific rationale, and no insertion anchors or full replacement text are supplied | different executors can ship materially different rule sets or copy non-transferable spec narration into the scaffold | include the exact production-ready text for each copy and identify a unique verbatim anchor for every insertion +MAJOR | high | Task 4 Step 2 | the nonce instructions omit the bounded generation retry policy the spec explicitly delegates to the plan, candidate scoping by cycle type and artifact, the two recovery sources, disagreement and multiplicity handling, working-record retirement, and uniqueness against open cycles | story criterion 6's interruption, collision, and concurrent-cycle behavior remains undefined | specify the retry bound and exact generation, recovery, collision, retirement, and stop-and-surface procedure in both copies +MAJOR | high | Task 4 Step 5 | the spec contains no five provenance examples and no curve example to hand-parse, and the required spec §8 check instead calls for constructed valid strings plus invalid strings that must be rejected | the step cannot run against its named inputs and would not test rejection or cardinality even if examples existed | put explicit valid and invalid fixtures plus a parser or deterministic validation procedure in the plan and assert every expected accept and reject +MAJOR | high | Task 6 Steps 1-4 | the task says to record parity and twelve-item results but gives no target section, schema, required per-rule or per-item fields, or replacement text; only the later `git add` implies an edit | the task can be half-executed with nothing changed or with an unauditable narrative and then fail its commit | define the exact sections and tables appended to the conditions artifact, including one row per changed rule and one status-plus-reason row per checklist item and artifact +MAJOR | high | Task 6 before Task 8 | the conformance pass runs before Task 8 changes the outer `workflow-init.md` prompt by adding the item-1 note | the final changed command prompt is never assessed as a complete artifact against all twelve items | move Task 8 before Task 6 or rerun and record the full conformance pass after Task 8 +MAJOR | high | Global Constraints Gate-B classification | the plan says Tasks 1-6 all fire full Gate B, but Tasks 0 and 6 each stage only `docs/superpowers/specs/*.md`, which §5 classifies as explanatory documentation and therefore N/A when committed alone | the plan encodes false-red gate decisions and contradicts its own Task 7 classification rule | classify Tasks 0 and 6 as N/A when their staged sets remain docs-only and keep full Gate B for Tasks 1-5 and 8-9 +BLOCKER | high | Commit steps in Tasks 1-6 and 8-9 | each prompt or plugin task ends with an ordinary final `git commit`, while full Gate B is only mentioned after the fact; §5 requires a `WIP:` snapshot, review, re-review after fixes, and a closing amend, and every ordinary commit closes and resets the cycle | literal execution lands unreviewed prompt commits and creates multiple reset cycles instead of the one Gate-B cycle the spec requires | replace the per-task final commits with a gate-compatible WIP and amend strategy, or batch the implementation into the single reviewed Gate-B snapshot and close it only after a clean pass +MAJOR | high | Task 7 Step 2 anchors at getting-started lines 44-45 and coding-workflow lines 79-80 | both quoted Current strings are split by newlines and therefore do not exist verbatim in the named files | those literal replacements cannot be applied as written | quote each actual multiline anchor or provide a unique verbatim substring that exists on one line +MAJOR | high | Task 7 Step 2 at getting-started line 58 | `show the ratio as the derived floor, not a literal 3/3` is an instruction-shaped placeholder rather than replacement wording | an executor must invent the user-facing example and can produce a form inconsistent with the hook's independent reminder threshold | provide the exact corrected sentence and example, explicitly distinguishing the derived obligation from the hook ratio +MAJOR | high | Task 7 Steps 1 and 3 | the grep finds only five of the nine listed sites and does not match the stale `same 3-pass loop`, `If the Gate-A floor...`, `three passes, final clean`, or literal `3/3` text | the check reaches zero while four required corrections can remain undone | extend the before-and-after check to match every listed current phrase and assert each expected pre-edit count before requiring zero afterward +MINOR | high | Task 8 Step 1 | the step is titled `watch it fail` but the command correctly returns 1 now and must still return 1 after the note; no check asserts that the note itself was added | the advertised red-green evidence is false and omission of the note can survive the numeric check | treat the Target-model count as a preservation check and add a separate initially failing exact check for the outside-fence note +BLOCKER | high | Task 9 Steps 2-5 | `check-version-bump.sh` compares committed HEAD with the merge-base of main and ignores worktree and index changes, but Step 4 reruns it before Step 5 commits the manifest bump | Step 4 still sees version 0.10.0, reports the plugin changed without a bump, and the full battery cannot reach green, so Task 9 cannot complete as ordered | create the gate-compatible committed WIP containing the bump before running the checker, then run the battery and close by amend after review +MAJOR | high | Task 10 Steps 3-5 before Task 11 | Task 10 requires recomputing `each provenance line` and writing evidence from it before Task 11 writes any of the three closing-body lines | the risk verification has no records to inspect and can only be fabricated or deferred silently | draft all three body records before Task 10, verify those exact drafts against story headers, and only then write them into their owning closing commits +MAJOR | high | Task 11 Steps 1-4 | after the plan's many explicit task commits, `git log --format='%b' -3` cannot select the Gate-A spec, Gate-A plan, and Gate-B closing commits, and the task supplies no commands for locating or safely rewriting the already-closed spec and plan bodies | the expected count of three does not establish the three required records and retroactive history edits are left ambiguous and risky | name the three owning commit identities, specify the safe amend or rebase procedure and rollback point, and verify those exact bodies rather than the last three arbitrary commits +END OF FINDINGS (21 total) diff --git a/.context/codex-reviews/gate-a-plan-sweep-2026-08-02.2026-07-30-classifier-cycle.md b/.context/codex-reviews/gate-a-plan-sweep-2026-08-02.2026-07-30-classifier-cycle.md new file mode 100644 index 0000000..1e97733 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-sweep-2026-08-02.2026-07-30-classifier-cycle.md @@ -0,0 +1,19 @@ +## Mechanical sweep before pass 9 — 13 checks, 8 defects found and fixed + +| # | Check | Result | +|---|---|---| +| 1 | `sh -n` on all fenced `sh` blocks | clean | +| 2 | `sh -n` + content invariants on the 10 message assignments (no newline, no apostrophe, target prefix, four tags, §5 cleanup, retry budget) | **1 defect**: `NORESULT_CTX` bounded the retry but never tied it to §5's shared budget | +| 3 | Scratch-clone dry run, Task 0 → Task 7 Step 6 | clean; confirmed `CHANGELOG.md` shows `??` after `reset --soft` and must be staged | +| 4 | Scratch-clone dry run: Step 8 fix loop, Step 2 counterfactual, Step 5 rollback resolution | **1 defect**: `git log -S'"version": "0.7.1"'` resolves to the **0.8.0 bump that removed it**, so the rollback byte-check would compare the cached 0.7.1 hook against 0.8.0's bytes. "The commit that set it" is wrong too — a later docs commit can change the hook at the same version | +| 5 | `verify.sh` labels vs the plan's port inventory | **1 defect**: 53 call sites produce 56 assertions; the plan called both "56 rows" | +| 6 | Task 6 census greps, run verbatim | **2 defects**: grep 1 returns 7 lines for 5 sites (two sentences wrap), so a count comparison reports a phantom discrepancy; grep 2 has two undispositioned hits in `codex-gate.test.sh` asserting the retired "accurate" claim | +| 7 | Plan's battery block vs the `AGENTS.md` quality row | clean — 11 vs 12, sole delta `check-version-bump.sh main`, exactly as stated | +| 8 | Fixture source names vs `.context/probe-payloads/` | **1 defect**: seven destinations and one rename (`shape1-fast-fail-execution-failed.json` → `shape1-fast-fail.json`) were never named | +| 9 | All 14 `file:line` references say what the plan claims | clean | +| 10 | `is_wip_commit` against the plan's four exact commit commands | clean — three cycle-internal, the closing amend correctly a real commit | +| 11 | `shellcheck --shell=sh` on every extracted block | clean | +| 12 | Cross-artifact refs: 14 spec §§, 5 prompt-standards items, 6 invariants, `hooks.json` | **1 defect**: `hooks.json` has **two** matchers — `PostToolUse` is `^(Bash\|Skill\|mcp__codex__.*)$`, `PreToolUse` is the unanchored `Bash\|Skill` — and the plan cited one unqualified | +| 13 | `harness.sh` seed header | **1 defect**: stale 53/53 | + +All eight fixed. Post-fix state: 7 blocks + 10 assignments `sh -n` and `shellcheck` clean; drafts 56/56 under `sh` and `dash`. diff --git a/.context/codex-reviews/gate-a-spec-CLOSURE.2026-07-30-classifier-cycle.md b/.context/codex-reviews/gate-a-spec-CLOSURE.2026-07-30-classifier-cycle.md new file mode 100644 index 0000000..1492d07 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-CLOSURE.2026-07-30-classifier-cycle.md @@ -0,0 +1,80 @@ +# Gate A (spec) — closure record, result-classification cycle + +## How it closed + +EIGHT passes. MAJOR counts: 17, 13, 14, 4, 7, 7, 6, 6. Every pass was VALID (file present, +terminator exact, count matching, all body lines finding lines). + +**Gate A closed on JUDGEMENT, not on a NO FINDINGS pass.** That is recorded plainly so nobody +later reads this cycle as having ended clean. The final pass returned 6 MAJOR; all six were +fixed, and none challenged the design core — the five classes, the fail-open/fail-closed +split, the state effects and the message structure have been stable since pass 6. Passes 7-8 +returned peripheral findings: scope discoveries in shipped files, and encoding edges. + +Human decision, after the trend was surfaced: accept with residuals, carry the obligations +into the plan, and expect the PLAN's Gate A to reach a genuine NO FINDINGS — its material is +implementation-grade and therefore decidable. **If the plan's Gate A also ends on judgement, +STOP and surface: two judgement exits in one cycle is a pattern, not a coincidence.** + +## Carried obligations — a CHECKLIST for the top of the plan + +Each item is checked off or explicitly re-dispositioned during the plan's own Gate A. None +evaporates between artifacts. + +### A. Implementation contracts (spec §11) +- [ ] 1. The jq-free scanner as a state machine: quote state, backslash parity, value + boundaries, operational meaning of "depth 1". +- [ ] 2. Duplicate depth-1 keys and malformed-JSON recognition (the CLASS is settled; only + recognition is deferred). +- [ ] 3. Accepted raw encodings around every token, including the blank-byte grammar and its + ordering against canonical-form validation. +- [ ] 4. Full notice grammar: duration format, task-id boundary, quote representation, and + the near-misses that must NOT match. +- [ ] 5. Complete marker state table across both disclosure markers and bgAdvice, including + coexistence precedence and every write/delete/retry failure. +- [ ] 6. Composition against every existing emit branch (Gate A, Gate B, WIP, docs-only, + unknown-tool) and events that would otherwise emit nothing. +- [ ] 7. Separator and encoding rules for composed messages. + +### B. Shipped-doc scope (pass 8) — SCOPE, NOT WORK +Which files the change touches. **None is edited before the plan says so.** +- [ ] Inline CLAUDE template in commands/workflow-init.md, and this repo's CLAUDE.md §5 — + both say the hook keys on tool name and never inspects results. +- [ ] "Accurate counters" on opt-out: README.md, codex-gate.sh reminder text, + commands/workflow-init.md. +- [ ] Every mapping instruction in commands/workflow-init.md — including the preflight + remedy that renames the server AWAY from `codex` — plus the README knob row and the + unknown-tool hook message. + +### C. Accepted residuals — must survive into the plan unchanged +- [ ] §4: a reworded backgrounding notice on any runtime where auto-backgrounding is still + effective (unset variable, Claude Code < 2.1.212, or a positive value shorter than the + call) is counted again. +- [ ] §5.2: an `unrecognized` call whose disclosure is neither delivered nor persisted is + counted silently. +- [ ] §5.1: counter mutation is unserialized, and `.context/` is trusted. Both pre-existing, + both filed in todos.md with triggers. + +## Story amendments made during this Gate A +§2 at passes 2 and 5; §3 criteria at passes 4 and 8. Each is recorded inline in the story +with what it replaced. §2 is now marked a SUMMARY deferring to the spec, so a future +amendment has one target. + +## Watch-item for the plan's Gate A + +**If findings concentrate on contracts 5-7 — the marker state table, composition against +every emit branch, and separator/encoding rules — STOP and surface rather than elaborating.** + +Those three are the diagnostic-messaging machinery, and they are the part of this design that +grew by accretion: pending state was added at pass 2, the gate-on failure path at pass 4, and +one-emit composition at pass 3, each closing a real gap and each enlarging the surface the +next pass had to keep consistent. Findings clustering there would mean the composition +semantics are carrying more than they should, not that the spec needs finer prose. + +**The named pressure valve is a SIMPLER composition semantics** — for instance, dropping the +compose-into-one-emit rule in favour of a single deferred-disclosure flush, or accepting a +duplicated disclosure instead of tracking pending/shown separately. That trade changes what +the design promises about message delivery, so **it is decided upstream by the human, not +absorbed into the plan.** + +Recorded here rather than in the plan because the plan is what would be tempted to elaborate. diff --git a/.context/codex-reviews/gate-a-spec-CLOSURE.md b/.context/codex-reviews/gate-a-spec-CLOSURE.md new file mode 100644 index 0000000..aebfcc0 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-CLOSURE.md @@ -0,0 +1,166 @@ +# Gate A — spec — CLOSURE RECORD + +**Cycle:** a supersession convention for the hardening ledger. +**Spec:** `docs/superpowers/specs/2026-08-05-hardening-ledger-supersession-design.md` +**Story:** `docs/superpowers/stories/2026-08-04-hardening-ledger-supersession-story.md` +**Branch:** `ledger-supersession`. **Closed 2026-08-11 on pass 23.** + +**Gate A (spec) is CLOSED.** `superpowers:writing-plans` is open. Passes 1–22 applied; **pass 23's +four findings are held-not-fixed and recorded as §8 residuals** — Daniel took the stop decision +before pass 23 ran, so repairing them would have produced exactly the unreviewed revision the +closing pass existed to end. + +Reviewer for passes 4–23: `mcp-codex-dev@1.0.1`, `codex-cli 0.144.5`, model `gpt-5.6-sol`. +Passes 1–3's model was never recorded, so a change across the Codex outage cannot be ruled out. +Every pass satisfied the file-first protocol: terminator exact, count matching, no non-finding +body lines. No pass was INCOMPLETE; the recovery budget was never spent. + +--- + +## The two generators + +Nineteen of the cycle's twenty-three passes chased consequences of two constructs. Both were +**deleted rather than fixed**, and each deletion is what ended the finding stream around it. + +### 1. The landed-row boundary — deleted at pass 7 + +The story's append-only rule was narrowed at pass 3 to "never edit a **landed** row", to make +lawful what PR #22 had already accepted in practice. That required a definition of *landed*, and +three were designed and deleted in turn: **authorship-by-cycle**, **reachability from +`origin/main`**, and **content presence in the published ledger**. Passes 4 through 7 each found +the *replacement* for the previous definition unsound — three times inside the very sentence +written to fix its predecessor. The last was pass 7's blocker: a stale-but-readable ref returns a +confident *absent* for a row already published, licensing an edit to it. + +Each definition also carried its own old-conditions accounting, six-case self-test, concurrency +assumption and residual list, every one of which grew defects of its own. + +**Rationale for deletion:** with no amendable class there is no test to get wrong, no set to +enumerate, no verdict that can flip, and nothing to keep in step across two prompt surfaces. The +rule is bound to **rows**, not commits, so the drafting floor answers the "may I fix a typo before +I commit" question without rebuilding the boundary. Cost: **one entry, once** — row D's in-place +amend becomes an entry that was never written. Pass 7 restored the absolute wording; AC 3 holds +literally again. Recorded as **explored-and-deleted** in the story so it is not re-explored. + +### 2. Check 1d's second half — deleted at pass 21 + +Pass 19 found the original 1d unsatisfiable: *"no entry this change appends is inert"* fails on +§2.2's own sanctioned repair, since a mistyped locator is a complete entry the moment it exists, +entries are never removed, and the correction is an append. The replacement added a successor rule +— *no inert added entry stands uncorrected by a later added one*. + +That clause then produced four Majors across passes 20 and 21: it is satisfied by any later +non-inert entry including one about an unrelated row; it needed a duplicate-alignment oracle +because union merge can produce identical entry lines; it is **unsatisfiable** for an entry aimed +at a row that never existed, a case §2.2 explicitly admits; and the repair relation it rests on is +authorial intent the ledger does not encode, so 1f could not carry what pass 20 moved into it. + +**Rationale for deletion:** 1d now reads only *"the entry §3.1 mandates matches at least one row +dated on or before its own date."* It still passes the sanctioned typo-then-correct case — the +mandated entry **is** the corrected one — and still fails the mistyped-locator case it was written +for. The successor rule, the alignment oracle and 1f's fifth confirmation went with it. Pass 22, +the first pass after the deletion, returned **no mechanism defect, no unhandled path, no invariant +risk** — the only pass in the cycle to do so. + +**The lesson both share:** the machinery was built for states this change cannot enter, and each +repair enlarged the surface that generated the next finding. + +--- + +## Severity trend + +| Pass | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | 17 | 18 | 19 | 20 | 21 | 22 | 23 | +|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---| +| Findings | 11 | 14 | 16 | 15 | 9 | 10 | 5 | 12 | 8 | 6 | 3 | 4 | 7 | 6 | 6 | 5 | 6 | 7 | 5 | 4 | +| Blocker | – | – | – | 1 | 1 | 2 | – | – | – | 1 | 1* | – | – | – | – | – | – | – | – | – | +| Major | 7 | 11 | 11 | 9 | 6 | 6 | 3 | 7 | 5 | 4 | 1 | 3 | 5 | 1 | 4 | 4 | 6 | 6 | 3 | 2 | + +\* pass 14's "Blocker" is the only finding dismissed **as stated** in the whole cycle — its premise +was factually wrong (the `+` *was* escaped); its remedy was applied anyway as portability hardening. + +**Reading it.** The count never converged to zero and gives no signal on its own. The **class** did: +mechanism defects and shell disappeared after pass 16; the 20–21 spike was self-inflicted by +generator 2 and vanished when it was deleted; passes 22 and 23 returned claims-about-the-design and +one over-coverage finding. Pass 7's findings were superseded wholesale by the boundary cut. Rider 4 +(oracle coverage) came back empty at 17 and 18, produced real gaps at 19–21, and at 22 ran in the +**over**-coverage direction for the first time. + +--- + +## Settled decisions, with the passes that settled them + +| Decision | Settled | Passes | +|---|---|---| +| **No amendable class.** Every row protected; a correction is always an append | absolute, restored | 3 narrowed · 4–7 tested three boundaries · **7 restored** | +| **The floor is bound to rows, not commits** — drafting below it | stated in shared prose | 5, 8 | +| **Entries take the same floor, at COMPLETE entry shape** — partial lines are drafting | stated in shared prose, anchor 35 | 20 raised · **21 set the boundary at completeness** | +| **Entries are row-markers; latest-in-file governs** — the correcting-entry path is cut | cut | 10, 12 | +| **Match semantics:** an entry applies to every row it matches; uniqueness withdrawn | settled | 10 (zero-match) · 11 (temporal bound) | +| **Inert entries stand as history** — never removed, repaired by appending | settled | 10 | +| **Bounded backwards in time:** on-or-before the entry's own date; backdating forbidden | settled | 11 | +| **Inseparability is per ROW, not per pair** — no permitted distinguishing fragment | generalised | 17 narrowed · **18 generalised** · 19 example fixed | +| **`harden-finding` is out of the change surface** — the convention narrows nothing | settled | 4 | +| **Both surfaces**, resting on a parity claim check 2 validates | resolved | 4 | +| **§6 carries no shell** — properties, falsifying observations, oracles only | settled | 16 | +| **No standing machine consumer parses an entry**; §6's parser is validation-only | narrowed | **22** | +| **1d asks about the mandated entry only** | narrowed | 19 · 20 · **21 deleted the second half** | +| **Thirty-five anchors** (33 + fragment narrowing + entry floor) | extended | 19, 20 | +| **The story carries eight amendments**, one per prompting pass, each with old-condition accounting | — | 1, 3, 4, 5, 7, 9, 18, 19 | + +--- + +## Held-not-fixed — pass 23's four findings + +Recorded in the spec's §8 as a single bullet, and actionable by whoever writes the plan: + +1. **MAJOR — 1d's matching oracle does not name row-date equality.** A checker comparing only the + fingerprint passes every stated fixture and still reports non-inert an entry whose locator + matches no row. The executable check must validate exact row-date **and** fingerprint equality + before the eligibility bound, with a wrong-row-date fixture. +2. **MAJOR — two sites still carry the pre-narrowing "prose-only" claim** (§5, and §8's + discriminator-rejected bullet). The accurate claim is **no standing machine consumer**: the + syntax *is* a standing convention future authors must honour. §2.2 governs where they disagree. +3. **MINOR — 1d's scope oracle reads as plural** and can be misread as quantifying over every added + entry, recreating the unsatisfiable property the narrowing removed. +4. **MINOR — the format-example guard is understated as "shape alone".** For this change, location + guards it too (1d's interval, check 3's rejection outside it) — once. The claim is right about + *standing* enforcement. + +--- + +## What the plan must carry — standing riders into `superpowers:writing-plans` + +- **Per-check labels and oracles.** One label per §6 check; the plan carries the oracles, since §6 + states properties and the plan states how they are met. +- **Fences are written at execution time**, for Gate B against the real diff and for the harness — + not in the spec. Extract every `sh` fence, run it **standalone under both `sh` and `dash`** in a + clone so `git show BASE:` resolves; **assert exit status, never printed output** (a fence that + prints its failure and exits 0 is the defect this catches, and it caught exactly that twice); + treat any stderr shell error as failure; run against fixtures that can fail — good / inert / + edited-target-row / truncated-row / pre-change — and counter-check every pattern against the + pre-edit tree, where it must return 0. **Beware the harness itself:** three harness bugs produced + false results this cycle (a multi-line pattern under `grep -F` that could never match; + `echo "$(basename $f) exit=$?"`, where the command substitution resets `$?`; a `sed` delimiter + colliding with `\|` in a fixture). A harness that reports success is not evidence until it has + been shown able to fail. +- **Decisions without history.** The plan states what to do, not how the decision was revised. +- **Mechanical sweep before every plan pass** — cited paths, quoted passages byte-for-byte + including emphasis markers, stated counts, cross-references, fence balance. A read pass spends + expensive judgement on what a parser settles in seconds, and misses it anyway. +- **Per-fix landing check, two levels.** After applying any finding: one match proving the edit is + in the file, one proving the superseded text is gone — **and the same for any instruction + elsewhere that tells an implementer what to do to that file.** A pass-9 blocker hid at exactly + that second level. +- **The change surface is §4, not the scope paragraph.** §4 is authoritative and larger: version + bump, changelog, the backlog row and its two parked rows, two stories, `docs/coding-workflow.md`, + and the plan-snapshot note. + +## Where everything lives + +- **Committed and pushed** (`5c0dfc6`, `5e295f0`, `c3fc742`): the spec through pass 17, both + stories, `docs/coding-workflow.md`, the plan snapshot note. +- **Uncommitted at closure:** the spec and the story, carrying passes 18–23. Docs-only, so Gate B + is N/A by §5's prose exemption — **verify path by path before committing**, not by assumption. +- **On disk only, gitignored:** everything in `.context/codex-reviews/` — 23 pass files, their + dispositions, the resume note and this record. `.gitignore:13` ignores `.context/*`. **A push + does not back these up.** diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-1-dispositions.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-1-dispositions.md new file mode 100644 index 0000000..9a01406 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-1-dispositions.md @@ -0,0 +1,26 @@ +# Gate-A spec · cycle awsf1ec771 · pass 1 dispositions (advisory companion, not a findings file) + +1 | fixed | decline block moved after the whole human-exception passage (before "- **Timeout / abort:**"); §4 and §5(h) updated +2 | fixed | §3 block gains an evaluation order: clean completion read first, suspensions only on a non-closing pass +3 | fixed | cleanliness and resolve read effective (post-ceiling) severity; only counts/clusters/tells read the pre-ceiling field; (g) says the same +4 | fixed | stop answer = standing, resumable suspension under the cycle's nonce; no abandonment transition invented +5 | fixed | composition rule: resume only when every answer resumes; one stop answer leaves the whole suspension standing +6 | fixed | "any non-empty set" of suspensions, every reason reported, every question asked +7 | fixed | "surfaced solely by" stuck/two-tell; a finding also out-of-set or opening a question takes the scope answers at that surface +8 | fixed | scope stop split into membership stop and question stop; question answer leaves membership unchanged; decline only at membership stop +9 | fixed | decline recorded when made in the next commit body, restated in every later body; resume with no body carrying it = absent, hold applies +10 | fixed | §4 item 8: closing-message sentence carries every decline record incl. those in soft-reset-collapsed WIP bodies +11 | fixed | §4 item 9: one-contract paragraph names the ordering and the decline record and the three inconsistent states +12 | fixed | c19 marked replaced with reason +13 | fixed | (c) keeps the full precedence sentence byte-for-byte; c9–c11 kept verbatim; block no longer restates the consequence clause +14 | fixed | 135 = 22+18+20+7+11+7+4+26+16+4 stated in the header +15 | fixed | (f) accounting f1–f7 kept, unchanged +16 | fixed | Q6 states reduced sensitivity concretely (trend and pair read only present passes) +17 | fixed | Q6 defines "unavailable" (slot path absent) and partitions it from invalid-file and wrong-root, each pointed at its rule +18 | fixed | continue resumes on the union of governing artifacts' scopes plus accepted obligations +19 | fixed | complete paste-ready "Recording a decline." block in §4, verbatim, Mechanics indent +20 | fixed | every constraint in the block carries its why inline; §10 claim narrowed to "each constraint … in the same sentence" +21 | fixed | o-ids replaced by quoted lead phrase + C line +22 | fixed | copied profile values removed from the header +23 | fixed | parity diff and result added to the evidence entry +24 | fixed | disclosure sentence in the decline block: accepted stops and stuck/two-tell surfaces leave no record diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-1.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-1.md new file mode 100644 index 0000000..8fd5141 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-1.md @@ -0,0 +1,25 @@ +BLOCKER | high | §5(h) "Recording a human exception" | The decline block is placed before "Do not expect silence", which is inside the current human-exception passage, not after it; the untouched tail at C:999–1033 and W:1183–1217 then follows the decline heading and says "this form" is never an answer to a STOP, directly contradicting the decline rule | Both shipped copies can instruct that a decline both is and is not the answer to the scope stop, so the hold has no deterministic transition | Place the decline block after the full human-exception passage at C:1033 and W:1217, or explicitly relabel and re-account every retained tail paragraph so its antecedent remains the human-exception form +BLOCKER | high | §3 "When two apply at once" | Clean completion cannot actually outrank the two-tell stop: the two-tell suspension raises a hold, while the proposed clean-pass definition requires that the pass raised no hold | A Blocker/Major-free pass at or above the floor with two historical tells is made unclean by the lower-priority stop, so continuing can repeat forever and the D2 precedence never fires | Define a pre-hold clean-completion candidate and suppress the losing two-tell hold, or otherwise state an evaluation sequence that makes the promised precedence executable +BLOCKER | high | §3 "A pass is clean" / §5(g) | Cleanliness reads the raw severity carried by the findings file, while passage (g) says a ceiling-demoted Blocker changes what the cycle must resolve but remains a Blocker only for counts and curves | A raw Blocker demoted to Minor owes no repair and cannot be declined in-set, yet still prevents every later pass carrying it from being clean, leaving the cycle unable to close | State that resolve and cleanliness use effective post-ceiling severity, while only the explicitly named health measures use the reviewer-written pre-ceiling field +BLOCKER | high | §3 "The answer to a clearly-stuck or two-tell surface" | "Any answer ends" the hold, but a stop answer neither resumes nor closes and the spec names no suspended, abandoned, or later-resumable state after that answer | The cycle is left open and non-running after its hold has ended, with no input that can move it back toward a clean close; this is the prohibited unable-to-close and unable-to-suspend path | Define stop as preserving a resumable suspension with an explicit later input, or define a separate abandonment transition and its lifecycle without calling it a close +BLOCKER | high | §3 "When two apply at once" | For simultaneous suspensions, "resume when every question raised is answered" conflicts with the stuck/two-tell rule when one answer is continue and another is stop | The same fully answered scope-plus-two-tell state both resumes and does not resume, so the next-state table cannot assign one result and the cycle can strand | Specify composition of answer outcomes, including mixed accept or decline plus continue or stop, and give stop a deterministic state transition +MAJOR | high | §3 "When two apply at once" | The text handles exactly two suspensions with "both reasons", but all three are reachable together when an out-of-set or structural finding lands on a pass that is both clearly stuck and has two tells | AC1 promises every reachable conflict without inference, yet the three-way report, answer set, and resume condition are undefined | Generalize the rule to any non-empty set of suspensions and require every applicable reason and question to be carried and answered +MAJOR | high | §4 "The rule" | The blanket statement that a finding surfaced by stuck or two-tell is not a scope-stop finding until a later pass contradicts the rule that suspensions compose when that same finding is already outside-set or opens a new structural question | A simultaneously applicable scope stop is erased for one pass, so the user cannot decline the finding at the surface where D6 makes decline available and may be asked again unnecessarily | Say "surfaced solely by" stuck or two-tell, and explicitly preserve immediate decline eligibility when scope is one of the simultaneous reasons +MAJOR | high | §3 "A pass is clean" / §4 "The rule" | The accept-or-decline model assumes every scope-stop finding is outside the fix set, but the retained scope rule also stops on a new structural or contract question even when its finding is already in-set | An in-set structural finding cannot be "put" into the set or "stay outside" it, and the prose can either authorize an invalid decline or fail to say what answer releases its hold | Split membership stops from in-set structural-question stops; define the latter's answer and hold release without changing set membership or extending decline beyond D6 +MAJOR | high | §4 "The record" | A decline must bind for the rest of an active cycle, but its only required durable home is the closing commit body; a Gate-A cycle has no spec or plan commit until after it closes, and the existing resume record remains optional | Session loss, compaction, or resume elsewhere can erase a valid explicit decline before the next pass, violating D7 and causing replayed holds or repeated user questions | Require active declines in the cycle working record or another mandatory cycle-local store, define recovery and conflict handling, then copy them into the closing body +MAJOR | high | §4 "Which commit" | Gate-B declines are put in a WIP body and restated by a closing amend, but current Mechanics also permits several WIP commits followed by a soft reset and one new commit; the spec gives no carry rule for that local collapse | Decline records held only in earlier WIP bodies can be destroyed even though squash-merge carry is correct | Extend the multi-WIP soft-reset path to collect every decline record from the collapsed WIP range into the final body and define disagreement handling +MAJOR | high | §5(h) / partial-adoption contract | The spec adds a coupled ordering, decline block, nonce membership, squash carry, and strict unknown-start behavior but leaves the existing "These records are one contract" partial-adoption inventory at C:879–890 and W:1063–1074 unchanged | Downstream `/workflow-init` can merge only some of the new clauses without triggering the mandated coherence stop, leaving references to absent records or records that are not carried or attributable | Extend the semantic partial-adoption contract to the closure ordering and every decline-dependent component, with the exact inconsistent states that must stop +MAJOR | high | §5(c) old-conditions accounting | `c19` is marked moved, but the current text at C:242–246 and W:445–449 says the loop "resumes on whatever the user decides", whereas the replacement makes stop an answer that does not resume | The accounting hides a changed condition instead of marking it deliberately dropped or replaced, violating AC5 and the AGENTS.md decision-procedure rule | Mark the old unconditional-resume condition as replaced, state why the stop branch is deliberately different, and include its new state transition +MAJOR | high | §5(c) "Recognizing clearly stuck" | The spec and D3 claim the existing clean-precedence sentence is preserved verbatim, but the current sentence at C:234–238 and W:437–441 continues through the at-or-above-floor consequence and false-report rationale, while the proposed text ends it immediately after "this exit" | A settled decision is not met literally and the old-condition accounting mislabels moved text as a verbatim keep | Preserve the full current sentence byte-for-byte, or retain it as a clearly delimited unchanged sentence and move only separate surrounding sentences +MAJOR | high | §2 "Settled inputs" | The stated 135-condition inventory does not match the ranges enumerated in §5: 22 + 18 + 20 + 7 + 11 + 1 + 4 + 26 + 16 + 4 totals 129 | The load-bearing completeness claim is mechanically false, so the reviewer cannot distinguish a bad count from omitted old conditions | Correct the stated count or enumerate the missing six conditions and their accounting +MAJOR | high | §5(f) old-conditions accounting | Only `f1` is accounted for in the unchanged (f) passage; the six-condition gap between the claimed 135 and the 129 enumerated conditions is consistent with missing `f2`–`f7`, yet those conditions are neither kept, moved, nor dropped | Conditions in a passage that directly interprets the interacting scope and stuck rules escape AC5 review and can contradict the new ordering unnoticed | Publish the complete (f) inventory and mark every condition kept, moved, or deliberately dropped even if the prose remains byte-unchanged +MAJOR | high | §7 "Q6" | The shipped sentence says the threshold is read on a "reduced record" but never states D10's settled fact that sensitivity is reduced | A literal reader can disclose missing files without warning that missing history can hide a rising trend or require-withdraw pair, so the user may over-trust a non-stop result | Add an explicit statement that unavailable prior passes reduce tell detection sensitivity and name which comparisons are weakened +MAJOR | high | §7 "Q6" / prompt standard 10 | The new diagnostic state "earlier passes' findings files are unavailable" gives examples but no exhaustive cause partition, distinguishing checks, or cause-specific fixes as required by `docs/prompt-standards.md` item 10 | Missing paths, unreadable files, wrong-root resolution, and genuinely absent history can all produce the same report while requiring different recovery, so invariant 11 is not met | Enumerate the reachable causes, say how the agent distinguishes them, and pair each with its fix or explicitly justify why no fix is attempted before reduced-record reporting +MAJOR | medium | §3 final paragraph | Resume uses "the story or plan" and one assigned fix set, but current §5 permits a Gate-B cycle whose diff is contributed by several plans and whose header set cites several stories | In a multi-plan cycle, the agent has no rule for which artifact's scope wins or how accepted obligations combine, so scope-stop and decline decisions vary by reader | Refer to all governing artifacts and state the exact aggregation rule for their assigned scopes and previously accepted obligations +MAJOR | medium | §4 / §5(h) "Recording a decline" | Unlike the closure block, severity paragraph, and Q6 paragraph, the actual decline prompt block is not provided verbatim; §5(h) only tells the implementer to synthesize a block from design commentary, references, and a two-line form | The prompt is the product, so Gate A cannot review the exact rule wording or guarantee byte parity, antecedents, and prompt-standard compliance before implementation | Include the complete paste-ready decline block exactly as it must appear in both copies +MAJOR | medium | §10 "Invariant 11" | The spec claims every sentence in the new closure block carries its reason inline, but operational constraints such as stop closing nothing, any answer ending a hold, and suspension composition are asserted without an inline why; the explanatory rationale remains outside the shipped prompt | The changed prompt fails `docs/prompt-standards.md` item 6 and future readers cannot tell which rationale is load-bearing when editing the rules | Add concise causal clauses to each new constraint or narrow the §10 claim and revise the block until every constraint actually carries its reason +MINOR | high | §2 / inventory references | The spec declares inventory ids `a1` through `j4`, but later cites `o2`, `o16`, `o19`, and `o24`; no definition or source for any `o` id exists in the repository | Those rationales and cross-references cannot be mechanically traced to current prose, weakening the claimed old-condition audit | Replace each `o` id with the applicable defined inventory id or cite and commit the inventory that defines the `o` namespace +MINOR | high | Header "Profile" | The spec copies `high / none / battery+check+verification` even while saying the story header is the only writable copy; current §5 explicitly says specs carry the path and never a remembered or copied value | The copied profile can drift and can steer a reader despite the caveat, violating the single-source rule the review was told to enforce | Remove the copied values and retain only the story path plus the instruction to read its header fresh +MINOR | medium | §9 "Evidence entry" | §6 requires a passage-by-passage parity diff, but the closing evidence entry is specified to name only the battery, four presence asserts, and the next-state table | The durable record omits the check that supports AC6, and the four asserts do not cover the Q6 text or most list and passage edits | Add the §6 parity diff and its result to the named evidence entry +MINOR | medium | §4 / §9 observability | The commit-body design records declines but not accepted scope stops, clearly-stuck surfaces, two-tell surfaces, or their continue/stop answers | A reader of a later closing commit cannot observe which suspension exits the cycle took or verify that simultaneous reasons were reported; only the final clean close and any declines are inferable | Either add a compact suspension-event record to the closing body or explicitly disclose this observability limit and narrow claims about what history records +END OF FINDINGS (24 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-10-dispositions.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-10-dispositions.md new file mode 100644 index 0000000..2f9f525 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-10-dispositions.md @@ -0,0 +1,27 @@ +# Gate-A spec pass 10 — dispositions (cycle awsf1ec771) + +Pass 10 tripped the two-tell threshold and satisfied all three clearly-stuck conditions. Daniel's +answer, 2026-09-10: **split**. Twelve findings move with the material they are about to +`docs/superpowers/stories/2026-09-10-record-durability-story.md`; eight survive the narrowing and +are fixed in the pass-10 revision. + +1 | deferred | The rollback reading moves whole to docs/superpowers/stories/2026-09-10-record-durability-story.md; the spec no longer carries a rollback paragraph to be wrong about. +2 | deferred | Conflicting same-key records presuppose the record; both move to docs/superpowers/stories/2026-09-10-record-durability-story.md. +3 | fixed | The continue branch now applies only where no suspension applies; a below-floor pass with a health reading is taken by the second branch, since clean completion did not close it (D2, D3). +4 | deferred | Acceptance versus recording is a record rule; moves to docs/superpowers/stories/2026-09-10-record-durability-story.md. The ordering keeps the membership fact: an acceptance puts the finding in the set until the cycle closes. +5 | deferred | The empty-commit carrier and its Gate-B WIP interaction are record transport; move to docs/superpowers/stories/2026-09-10-record-durability-story.md. +6 | deferred | Record survival through amends and a soft reset is the successor's whole subject; moves to docs/superpowers/stories/2026-09-10-record-durability-story.md. +7 | deferred | The five-field sameness test is D9b and moves to docs/superpowers/stories/2026-09-10-record-durability-story.md; the ordering names "a finding this cycle has already declined" and says there that recognising it across passes is not its own rule. +8 | deferred | The unknown-start nonce-minting reading was needed only to classify records; moves to docs/superpowers/stories/2026-09-10-record-durability-story.md. The narrowed §4 item 6 adds only "every suspension binding", which needs no nonce. +9 | deferred | The checkout-root condition and its D10 tension move to docs/superpowers/stories/2026-09-10-record-durability-story.md with the unavailable-history block. +10 | deferred | The root-mismatch branch moves with the root condition to docs/superpowers/stories/2026-09-10-record-durability-story.md. +11 | deferred | The git-failure partition moves with the root condition to docs/superpowers/stories/2026-09-10-record-durability-story.md; raised at passes 7, 8 and 10 and unresolved in three attempts here. +12 | fixed | §4 item 4 edits the lens paragraph at its source: the filter and the clean-final-pass rule are unchanged **by the lens sets**, and the ordering changes both and says so. Old conditions enumerated in that item. +13 | fixed | The uniqueness claim is narrowed rather than the absorb paragraph rewritten: the block is where cleanliness and suspension both read the triggers, b11 and b13 stay in their own words, and §6's extraction checks the correspondence. +14 | deferred | The per-hunk contract marker moves with the contract itself to docs/superpowers/stories/2026-09-10-record-durability-story.md; §9 records the resulting unguarded partial adoption as a residual rather than as protection. +15 | deferred | The discriminator dissolution and its pre-rule exposure move to docs/superpowers/stories/2026-09-10-record-durability-story.md. +16 | fixed | The oracle now fails a row whose answer produces no distinct resumable or closed state, and any close with any unmet closure precondition, not only an in-set Blocker. +17 | fixed | The row set drops the departed material and gains the two rows the ordering's own transitions need: a below-floor clean pass carrying a health reading, and a fix set narrowed so an in-set finding falls outside it. +18 | fixed | "or two-tell" removed from the surfaced-finding sentence; the block says instead that the two-tell stop surfaces tells and not a finding, so no finding is surfaced by it at all. +19 | fixed | The effective-severity routing sentence now points at the Mechanics severity rule, which carries the Minor/Nit half, rather than at the duty paragraph, which does not. +20 | fixed | "Every predicate here" narrowed to every finding-derived predicate, and the block names the other state each branch reads: the floor, the cited set and profiles, a standing hold, the current fix set and the answers given, the coverage judgement. diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-10.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-10.md new file mode 100644 index 0000000..92952bd --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-10.md @@ -0,0 +1,21 @@ +BLOCKER | high | §4 item 9 "Rollback" | The spec explicitly permits a rollback that removes both the activation rule and the unknown-start fallback, then says an already-open new-rule cycle has no text describing what it owes and that this is not a stop | That reachable cycle has neither a rule by which it can close nor a shipped suspension with a resuming transition, matching the review contract's Blocker condition | Make safe rollback conditional on resolving or explicitly suspending every open affected cycle before the rules disappear, or retain a transition rule that gives those cycles a defined next state +BLOCKER | high | §4 "Binding, and the sameness test" | Conflicting records are said to become a user question "answered and recorded like any other answer", but question-stop decisions deliberately have no record form and another Accepted or Declined record cannot supersede the already-conflicting pair under "at most one effective label" | The user can answer yet the record state remains contradictory, so the resumed pass reaches the same stop again and the cycle cannot progress | Define a concrete state-changing resolution such as correcting the bad body before resumption or abandoning this nonce and starting a new cycle; do not claim the existing answer form records this question +MAJOR | high | §3 "Second" and "Third" | A below-floor clean pass carrying a clearly-stuck reading or two tells qualifies for suspension because it is not a clean completion, but the continue branch and the clearly-stuck-duty paragraph both say a below-floor clean pass continues | The same reachable pass has two next states, so the advertised fixed ordering is not executable | Say that a below-floor clean pass continues only when no suspension applies; under D2 and D3, a below-floor pass with a health stop suspends because clean completion did not close it +MAJOR | high | §3 "What a suspension asks, and what ends it" | The block first says the user's acceptance immediately joins the fix set, then says an acceptance with no commit-body record "is not in the set" | Recording becomes constitutive of membership even though story criterion 7 makes it a precondition to the next pass and says the user's answer is what puts work in the set | Keep the accepted finding in the fix set and make a missing record stop the next pass until the body is corrected; never turn an omitted record into an out-of-set finding +MAJOR | high | §4 "Rules both labels share" | The no-revision fallback orders an empty commit "of its own" without excluding Gate B, immediately after ordering Gate-B answers into a WIP amend | On Gate B, a separate non-WIP empty commit is read by the existing hook protocol as cycle-closing and discards the accumulated pass state even when the floor or other duties remain unmet | Scope the empty-commit fallback to cycles with no existing carrier; state that Gate B always amends the active WIP commit even when the tree does not change +MAJOR | high | §4 "Rules both labels share" and item 8 | Records are carried to the closing and squash bodies, but no rule preserves all earlier answer records through ordinary later WIP amends, and the unchanged soft-reset instruction runs before the new text says the replacement body carries records from discarded WIP bodies | A later amend or soft reset can erase the only durable copy before a compaction or handoff, losing an acceptance or decline that the record was introduced to preserve | Require every WIP amend to carry forward all answer records already made and require capturing those bodies before a soft reset, then write the captured records into the replacement commit +MAJOR | high | §4 "Binding, and the sameness test" | D9b says any changed one of the five stored fields makes a new finding, but the spec changes matching to meaning rather than bytes and therefore lets a textually changed field inherit the old decline when an agent judges it semantically equivalent | A changed defect, consequence, or suggested-fix field can suppress the new membership hold that the settled decision requires | Use exact field equality as D9b states, or obtain and record a new settled decision before introducing semantic equivalence +MAJOR | high | §4 item 5 and §8 "The slot discriminator" | The unknown-start rule tests records against "the nonce this fallback minted" and §8 says every unknown-start cycle mints one, while current condition i11 mints only when no nonce can be recovered | An unknown-start cycle that successfully recovers its nonce has no minted nonce against which records can be classified, leaving its holds and accepted set without a defined reading | Refer to the cycle's current nonce, whether recovered or newly minted, and correct the claim in §8 +MAJOR | high | §7 "First the root" | The spec makes failure to establish the checkout root invalidate the current pass as INCOMPLETE, even though settled D10 says unavailable prior history is disclosure and not a new stop condition | A pass that otherwise validated can consume recovery and end in the existing incomplete-pass stop solely because of the new Q6 root predicate | Remove this new validity gate and disclose root uncertainty as unavailable-history context, or obtain an explicit scope decision changing D10 +MAJOR | high | §7 "First the root" | The root predicate defines two outcomes, slots under the canonical checkout or not, but assigns INCOMPLETE only when the root cannot be established and gives no action for an established mismatch | On a mismatched-root path the agent cannot decide whether to use the slots, discount the pass, suspend, or continue | Give the false branch an explicit report and next state consistent with D10, including whether those files are ignored as unavailable history +MAJOR | high | §7 "First the root" and §10 | "Git's own message" is treated as the cause and discriminator without naming the git operation, while distinct canonicalization and root-establishment failures have no checks or fixes | The shipped diagnostic cannot be executed consistently and directly violates prompt-standards item 10, which requires distinct causes, distinguishing checks, and fixes | Either remove the root diagnostic or specify the exact operation and enumerate each observable failure with its check and fix; a raw unspecified error message is not that partition +MAJOR | high | §5 old-conditions accounting and current C:653-655/W:839-841 | The standing sentence says the Blocker/Major filter and clean-final-pass rule are unchanged, while this spec scopes the filter to the assigned fix set and changes cleanliness to include effective severity and the absence of scope triggers | Both shipped copies would retain a sentence denying two semantic changes made by the proposed text | Edit that standing sentence at its source to say only what remains unchanged, and account for every condition it currently carries +MAJOR | high | §3 "First, clean completion" | The block says the two scope triggers are defined "here and nowhere else" and that the suspension branch defines nothing, while §5(b) deliberately leaves b11 and b13 as standing trigger definitions in both copies | The prompt ships duplicate authorities despite a uniqueness claim, so later drift can make cleanliness and suspension read different triggers | Make the existing scope paragraph point to the authoritative definitions without restating them, or narrow the uniqueness claim and verify the two formulations are deliberately equivalent +MAJOR | medium | §4 item 9 and §5(b) | The spec claims every separately mergeable contract hunk carries a partial-adoption marker, but the b12 resume edit can be adopted without the later b17 marker and then requires answer records that the copy may not contain | A downstream partial adoption can resume on an undefined recording rule instead of taking the contract's mandatory stop | Put the contract marker on b12 itself, and audit every behavior-changing edit rather than relying on an adjacent sentence landing in the same merge hunk +MAJOR | high | §8 "The slot discriminator" | The nonce-less legacy set is called bounded and self-terminating even though an open cycle may remain open indefinitely and the same paragraph admits rollback can start additional old-rule, nonce-less cycles | The concurrent bare-slot collision and deletion exposure is neither guaranteed to terminate nor bounded across rollback, so the spec understates a real data-loss residual | State that the currently known set is finite only while the new rules remain installed and that termination is human-dependent; retain the admitted collision risk without the stronger claim +MAJOR | high | §9 "The named verification of the risk path" | The proposed oracle declares failure only for a row that repeats a stop with no input consumed or closes with an in-set Blocker, omitting an answer that is consumed but returns to the same stop and closure over an outstanding hold, question, floor, or profile precondition | The verification can pass the exact no-progress defect AC4 cites from the parent cycle and can accept several forbidden closes | Fail any row whose required answer does not produce a distinct resumable or closed state, and fail closure with any unmet closure precondition, not only an in-set Blocker +MAJOR | high | §9 "Rows" | The supposedly complete next-state verification omits the new root-unestablished and root-mismatch states and contains no squash or intermediate-amend transition that demonstrates answer-record survival | The high-risk verification never exercises the new Q6 terminal path or the data-loss path of criterion 7, while the presence greps prove only that prose exists | Add rows for both root outcomes and for record state before and after an ordinary WIP amend, soft reset, and squash merge +MINOR | high | §3 "Second" | The text discusses "a finding surfaced solely by the stuck or two-tell reading", then later correctly says the two-tell stop surfaces tells and not a finding | The ordering invents an unreachable two-tell finding state that AC1 explicitly says not to legislate for | Remove "or two-tell" from the finding sentence and keep finding-specific handling on the clearly-stuck and scope paths +MINOR | high | §3 "What a suspension asks, and what ends it" | The block says the resolve duty "above" states that Minor and Nit findings are collected and never iterated, but that duty paragraph states only the Blocker/Major discharge; the Minor/Nit rule remains in Mechanics below | The cross-reference sends the reader to text that does not contain the claimed rule | Point to the Mechanics severity rule or add the claimed Minor/Nit clause to the duty paragraph +MINOR | medium | §3 opening input rule | "Every predicate here" is said to read the findings files as a concatenation at effective severity, although the floor, profile and set finality checks, outstanding holds, answer records, and coverage-sufficiency judgement read external state instead | A reader can omit non-file inputs while applying closure because the block overstates a single data source | Narrow the sentence to finding-derived predicates and name the additional state each closure branch reads +END OF FINDINGS (20 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-11-dispositions.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-11-dispositions.md new file mode 100644 index 0000000..c56d101 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-11-dispositions.md @@ -0,0 +1,45 @@ +# Gate-A spec pass 11 — dispositions (cycle awsf1ec771) + +Advisory companion. One line per finding: verdict + reason. Spec revised at the commit whose +subject is `docs(specs): loop-rule consolidation — Gate-A pass 11 revision (one authority per rule)`. + +**The centrepiece was finding 14**, and findings 3, 4 and 6 were its symptoms. The block was +restating triggers, duties, preconditions and the severity answer that their own paragraphs still +defined, so each prompt copy held two authorities and they had already drifted. The repair +inverts it: the block is authoritative for the evaluation order and for closure and **cites** +every other rule where that rule lives; where a cited rule had to change to agree, it changed at +its source. §3 now carries a one-authority table naming all eight cited rules and their single +definitions, which is the check finding 14 was really asking for. + +1 | fixed | The rollback transition already exists — the activation paragraph's stricter reading, which §4 item 6 extends with "every suspension binding". §9 no longer implies a gap; what is open is whether that reading is *sufficient*, and that is the successor's. +2 | fixed | File aggregation and severity selection separated into two sentences; severity now points at the (g) replacement as the single authority instead of restating the partition. +3 | fixed | A below-floor clean pass carrying a health reading **suspends**. Corrected at all three sites — the third branch, the composition paragraph's closing sentence, and §5(c)'s `c18` entry — and §7's health-reading row already expects the suspension. +4 | fixed | The hold **attaches to every surfaced finding, whichever suspension surfaced it**, discharged by the answers that surface requires. This reverses the passes 6–7 narrowing to scope stops, which contradicted both `c16` and the story's own third standing duty; `c16` returns to **kept** and the reversal is recorded in §5(c) rather than left silent. +5 | fixed | A declined finding stays excluded for the cycle unless the user explicitly reverses that decision. The current set is now stated once, in `b7`: what every governing artifact assigns, plus what this cycle accepted, minus what it declined. +6 | fixed | `b11` qualified at both source sites with the already-declined exception, marked **replaced** in the accounting, and §6's equivalence check rebuilt to compare **complete predicates** in a four-row table — comparing a shared phrase is what let the earlier revision call two unequal wordings equivalent. +7 | fixed | A structural or contract question already answered in this cycle is no longer new, so re-raising it opens no question stop. Within-cycle only, no durable record. Two rows added to the next-state table. +8 | fixed | **Continue consumes the reading** that raised the suspension; a further health suspension needs it recomputed over a pass run after the answer, which is new data. That gives continue a distinct next state even on an unrevised artifact. Row added. +9 | fixed | Stop **parks** the cycle — open, not running, spending no passes, restarted only by an explicit later continue — a named state distinct from the suspended-awaiting-answer one it was in before the answer. `c19`'s accounting updated to match. +10 | fixed | Conservative, since sameness is deferred: a line in one branch file and a line in the other are **distinct** findings for holds and answers, so a `full` pass asks twice rather than risk resuming over one it never asked about. Stated as a rule in the shipped block, not as a temporal hedge. +11 | fixed | `b7` edited at its source in both copies with an OLD/NEW pair and its accounting: the union where several artifacts govern one cycle. The singular reading left a multi-plan Gate-B cycle with no defined fix set on its **first** pass, before any suspension could raise the question. +12 | fixed | Overclaim corrected. The live one-contract paragraph (C:880–890 / W:1063–1074) names the nonce, slots, provenance, curve, carry and unknown-start records and **not** this block or its coupled edits, so its coherence rule does not reach them. §9 now states an **admitted unsafe state**, not protection. +13 | fixed | Seven OLD/NEW assert pairs added for the §5 passage edits (`a17`, `a13`, `b7`, `b11`, `b12`, the (c) trim, `e7`), plus `b3` checked in W alone with C required unchanged, which is what makes it an alignment rather than a two-copy edit. +14 | fixed | See the note above. §8's item-8 claim also corrected: it rested on "replaces rather than adds beside", which was false while the block was restating; it now rests on what the block does, with §3's table as the check. +15 | fixed | Profile-change-while-a-hold-stands row added, run against both answer directions, with the further pass the change costs; the profile as currently read is now a column. +16 | fixed | Observability residual stated plainly: a closing body records that a cycle closed and its curve, and **nothing about which exit it took**, so a reader cannot audit that every suspension was answered. Assigned to the successor; the evidence entry and curve are explicitly not claimed to supply it. +17 | fixed | `b3` taken out of the blanket kept range and marked an **edited cross-reference whose operative condition is kept**; its edit is checked in §7 like any other. +18 | fixed | Real ranges cited (C:328–329 / W:522–523) and the counted single-line substring named. +19 | fixed | Real ranges cited (C:565–566 wrapped, W:757 whole) and the counted substring named — the part single-line in **both**, which is why the check uses it rather than the displayed sentence. +20 | fixed | Real ranges cited (C:783–784 / W:969–970) and the counted substring named. + +**18–20 also produced a general repair.** §4's preamble no longer claims every quoted OLD is a +single line. It now says a sentence is quoted whole, that several wrap, and that each item gives +its real range in both copies plus the single-line substring the check counts — because a +displayed sentence spanning a wrap cannot be counted with `grep -F`, and quoting one as though it +could is what made three earlier checks read as false reds. **Every OLD fragment was then +re-verified on that basis**, not only the three named: all fourteen shared substrings count 1/1. + +**One thing found while re-reading that no finding raised**, repaired: the shipped block carried +"Until the sameness rule ships…", a temporal hedge that means nothing in a scaffolded copy where +no successor is coming. It is now stated as the rule it is; the fact that a successor may replace +it is spec commentary in §3, not prompt text. diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-11.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-11.md new file mode 100644 index 0000000..abb1867 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-11.md @@ -0,0 +1,21 @@ +BLOCKER | high | §5(i) "When these rules bind" and §9 "Moved, out of scope, and parked" | The spec keeps i2, i15, and i16, so a determinable-start cycle must finish under the rules it started with, while §9 defers the rollback reading and the successor story says a rollback can remove the only text describing what that open cycle owes; the removal leaves a live activation rule with no closing or suspending transition | A cycle open when this ordering is reverted can neither close nor suspend, which is the review's stated Blocker condition | Make shipment depend on the successor supplying that transition, or otherwise keep this new ordering out of the unresolved rollback path without rebuilding the deferred transport here +MAJOR | high | §3 "The closure ordering" opening severity rules | Lines 70–82 say every finding-derived predicate reads effective severity and then say the health predicates, including two conditions of clearly stuck, read reviewer-written severity | Readers can apply the ceiling to the Blocker curve and regenerating-finding test or not, producing opposite closure decisions and failing the handed-over severity requirement | Separate file aggregation from severity selection: say all predicates read the combined validated files, then explicitly partition closure and repair predicates to effective severity and health predicates to reviewer-written severity +MAJOR | high | §3 "The closure ordering" below-floor overlap | The second branch and lines 135–138 say a below-floor clean pass carrying a health reading suspends, but the duties discussion at lines 164–171 and composition at lines 221–229 say the same pass continues below the floor; §7 lines 559–561 again expect suspension | The central ordering yields two next states for the exact overlap row it claims to settle, so a reader can bypass the required surface and answer | Change both below-floor overlap statements to suspension, matching the second branch and verification oracle, or change all three sites together if a different settled transition is intended +MAJOR | high | §3 "The four standing duties" | Settled D4 and the story's third standing duty cover a surfaced finding, including the finding surfaced by clearly stuck, but lines 154–168 redefine the hold as scope-stop-only and say health exits hold nothing | A clearly-stuck surface loses the classified hold the story requires; after continue, a demoted or out-of-set finding can cease gating closure even though the settled duty says the surfaced finding stays open until answered | Apply the hold to every surfaced finding and explain that clean completion creates no hold because it wins before surfacing; then let the relevant scope or health answer discharge it +MAJOR | high | §3 "What a suspension asks" decline lifetime | Lines 179–181 say decline keeps the finding outside and binds for the rest of the cycle, but lines 188–189 say a later fix-set broadening makes that same declined finding in-set and owing resolution | The later-scope rule silently overrides the specific user decision, contradicting settled D5 and D7 and making the membership answer's effect depend on an unstated precedence rule | Keep a declined finding excluded for the cycle unless the user explicitly reverses that specific decision; compute the current set as governing scopes plus accepted findings minus still-binding declines +MAJOR | high | §5(b) and §6 b11 equivalence | The block exempts a finding already declined this cycle from the membership trigger, but b11 is kept unchanged at C:205–208 and W:412–415 and still says an out-of-set finding stops; §6 nevertheless calls the two wordings equivalent | A declined finding re-raised on a later pass both must and must not suspend, violating AC 3's requirement that the qualification appear at every modified rule and defeating idempotent handling | Qualify b11 at both source sites with the already-declined exception, add b11 to the changed-condition accounting, and compare the complete predicates rather than only their shared outside-set phrase +MAJOR | high | §3 "First, clean completion" question trigger | The membership trigger consults answers already given, but the question trigger does not say that a decision already made on the same structural or contract question makes it no longer new; §7 tests only a re-raised decline | A reviewer can rephrase or re-raise the answered question on every pass, recreating the same hold and suspension after the required answer and violating AC 4's distinct-next-state rule | Define the within-cycle effect of an answered question without adding a durable record, and add re-raised accepted and question-answered rows to the next-state table +MAJOR | high | §3 "What a suspension asks" health continue | Continue merely runs another pass on an artifact that may be unrevised, while both the two-tell predicates and the clearly-stuck predicates can remain true on that next pass; no state consumes or changes the health signal | The same mandatory health stop can return immediately after its prescribed resuming answer, exactly the no-progress defect AC 4 and §7's oracle say must fail | State what new fact makes a repeated health reading a new suspension or make continue consume the current reading until such a fact exists, then verify the unchanged-artifact path +MAJOR | high | §3 "What a suspension asks" health stop | Stop is described as leaving the same suspension standing and resumable by the same later continue that was already available before the answer | The stop-answer row has no distinct post-answer state and therefore fails the spec's own oracle at lines 570–574 even though AC 4 requires every introduced terminal answer to change state | Name a distinct parked-open state and its later transition, or choose another transition consistent with only clean completion closing; do not describe the pre-answer and post-stop states as the same standing suspension +MAJOR | high | §3 full Gate-B branch composition | Lines 71–77 require a matching finding in the other branch to become one finding for holds and answers, but no matching predicate or conservative uncertainty behavior is defined; lines 239–245 defer only recognition of a later declined finding, and the successor is unconfirmed | A full pass with semantically similar but differently worded branch findings has no determinate number of holds or required answers, so it may resume over an unanswered finding or ask twice | Until the successor owns a usable sameness rule, treat branch lines as distinct for holds and answers or state a conservative fallback for uncertain matches; record this dangling use explicitly rather than importing the deferred discriminator +MAJOR | high | §3 current assigned fix set | The only standing source definition is singular at C:200–202 and W:407–409, while the proposed union of several plans or stories appears only in the sentence describing how a suspended cycle resumes | The initial pass of a Gate-B cycle with several contributing plans or a Gate-A artifact citing several stories has no defined assigned fix set, so membership, cleanliness, and the resolve duty can differ before any suspension occurs | Define the multi-artifact union once as the assigned fix set for every pass, edit and account for b7 in both copies, and have resume reuse that same definition +MAJOR | high | §9 "Moved, out of scope, and parked" partial adoption | Lines 631–638 say partial adoption is caught by an existing one-contract semantic-coherence rule, but the live contract at C:880–890 and W:1063–1074 names nonce, slots, provenance, curve, carry, and unknown-start records, not the closure block or its six coupled edits | The spec overstates compatibility protection: a downstream merge can take the clean predicate without the scoped severity duty, or vice versa, and no named rule catches the inconsistent prompt before it governs closure | Keep the guard deferred, but correct the residual to say no existing mechanism catches this set and identify partial adoption as an admitted unsafe state +MAJOR | high | §7 "The check" | The parent-versus-working-tree assertions cover the block, the severity replacement, and §4 items 1–6, but omit the meaning-changing OLD and NEW edits in §5(a), §5(b), §5(c), and §5(e), including the W-only b3 alignment | Both copies can retain the same stale resume, surfacing, and two-tell rules, pass every stated count and parity check, and still contradict the new ordering | Add paired parent and working-tree assertions for every changed passage fragment, with the parity diff remaining a separate cross-copy check +MAJOR | high | §3 and §8 prompt-standard compliance | The roughly 160-line block repeats the scope triggers, floor and profile preconditions, resolve semantics, health-exit behavior, and raw-severity answer while the original passages remain; §8 calls this token-lean merely because a few closure sentences are trimmed | The shipped prompt violates invariant 11 item 8, multiplies authorities inside each copy, and creates the very drift already visible in b11 and the below-floor transition | Reduce the ordering to one compact decision procedure and place each trigger and duty definition in one authoritative location that the other passages reference +MINOR | high | §7 risk-path verification | The stateful rows test a fix-set change while a hold stands but omit the separately requested profile-change-while-held path, even though profile changes impose a further-pass requirement and final acceptance rereads every profile | The verification can pass without demonstrating that the hold survives the profile change and that resumption still pays the new profile's pass and evidence obligations | Add a row with a standing hold, a confirmed profile change, each relevant answer direction, and the required subsequent pass under the new profile +MINOR | high | §7 evidence entry and commit-body observability | The closing body records the battery, assertions, parity evidence, table location, and curves, but no actual suspension reason or exit taken; a WIP or later closing commit cannot distinguish scope stop, clearly stuck, two tells, or their composition | A history reader cannot observe which exit the cycle took or audit whether every reported suspension was answered, despite observability being in the high-risk lens set | State this as an explicit observability residual and assign it to the successor if transport remains deferred; do not imply the existing evidence entry or curve supplies it +MINOR | high | §5(b) old-conditions accounting | The spec changes W b3 from "the severity rule" at W:405 to "Mechanics" but then includes b3 in the blanket statement that b1–b11 are kept, even though §5 defines kept as "the sentence stays" | The accounting is internally false and obscures a real source edit from the condition map used by the plan and reviewers | Mark b3 as an edited cross-reference whose operative condition is kept, and include that edit in verification +NIT | high | §4 item 1 mechanical OLD check | The quoted OLD sentence begins at C:328 and W:522, with the word "A" on the preceding line and the remainder at C:329 and W:523, contrary to the claim that every quoted OLD fragment is a single line at the cited site | The exact quoted sentence cannot be verified by the claimed one-line grep, so the mechanical evidence overstates what was checked | Cite C:328–329 and W:522–523 and either rewrap the source before asserting a one-line match or quote only the actual single-line fragment being counted +NIT | high | §4 item 2 mechanical OLD check | The quoted OLD fragment from "when a pass is clean" through the colon spans C:565–566, although §4 says every OLD fragment is a single line; W carries it on W:757 | The stated cross-copy one-line premise is false and a literal grep of the displayed fragment behaves differently between C and W | Quote the shorter single-line fragment actually used by the check or make the wrapping and multiline check explicit +NIT | high | §4 item 3 mechanical OLD check | The quoted OLD Severity sentence spans C:783–784 and W:969–970 rather than occurring on one line as §4 asserts | A literal one-line verification of the displayed OLD sentence cannot produce the claimed count in either copy | Cite both line ranges and specify the tested single-line substring, or rewrap before making the assertion +END OF FINDINGS (20 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-12-dispositions.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-12-dispositions.md new file mode 100644 index 0000000..c8c1d6b --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-12-dispositions.md @@ -0,0 +1,63 @@ +# Gate-A spec pass 12 — dispositions (cycle awsf1ec771) + +Advisory companion. Not a findings file; participates in no pass validation. + +## Two mechanisms ended this round, and the criterion that ended them + +Daniel brought a criterion mid-loop from a sibling project and endorsed it: **the second finding +of the same shape against the same mechanism ends that mechanism's rounds** — narrow the claim, +print the residual, move the work rather than running another round on it. (Its two companion +lines — a product Major always blocks, a harness Major blocks only where it makes a claim +vacuous — are already in `CLAUDE.md` Mechanics as the instrument carve-out. The third line has no +counterpart; captured as +`docs/superpowers/stories/2026-09-10-harness-finding-termination-story.md`.) + +Both mechanisms below were on their **fourth** round of one shape. + +**Mechanism 1 — §7's assert list. Moved to the plan.** Shape: "an assertion that does not detect +what it claims." Pass 8 found one and that revision's own audit three more; pass 11 found four, +three of them mis-cited line ranges; pass 12 found three (findings 10, 11, 16). The cause is +structural, not clerical: a spec cannot build an exact single-line substring check for text that +does not exist yet, so every round guesses and a wrong guess reads as a failed check rather than +a wrong one. §7 now states what a check must establish, names which edits owe a discriminating +pair and which are presence-only, hands fragment construction to the plan where the text exists, +and prints the residual — nothing verifies that the list of edits is complete or that the plan's +fragments discriminate. + +**Mechanism 2 — the block restating what it cites. Made mechanical.** Shape: "the block carries a +predicate its own source paragraph still defines." Pass 11's finding 14 named it; pass 12 found +four more (5, 6, 7, 8). The principle is no longer asserted but checkable: **a cited rule +contributes zero predicate words to the block.** Verified by probe — the eight predicate phrases +that were in the block now count 0 in it, while the citations count 1–3. + +## Findings + +1 | fixed | §9 no longer claims a rollback has a defined transition. A revert removes this change's own text, §4 item 6 included, so the stricter reading is itself part of what goes; stated as an admitted residual, and the successor's open question, which its §5 does carry. +2 | fixed | All four source pointers (`b12`, `b17`–`b18`, the (c) replacement, the (e) pointer) now defer to the ordering for what an answer does. Four sites each carrying their own version is how they came to disagree with the block about a stop answer. `c19`'s accounting corrected to match its own NEW text. +3 | fixed | `b17`–`b18` no longer says "the revised artifact"; it defers. §7 gains the row: a decline that leaves nothing to revise, next state a further pass on the unrevised artifact. +4 | fixed | `a16` ("fix Blocker/Major after each") now points at Mechanics · Severity, which carries the assigned-fix-set scope §4 item 3 gives it. Marked **replaced**, out of the kept range, old condition enumerated. Pointer chosen over restating the scope, per the directive. +5 | fixed | Block cites `b11` and `b13` and restates neither. The decline exception was already in `b11`; the answered-question qualification is added to `b13` at source as a new sixth (b) edit, with its own accounting. +6 | fixed | The fix-set formula lives only in `b7`. The block says "the current assigned fix set as the absorb paragraph defines it" and computes nothing. +7 | fixed | The effective/reviewer-written partition lives only in the (g) replacement, inside Mechanics · Severity. The block names which rule settles it and stops restating the split, in both places it had it. +8 | fixed | `c9`'s precedence sentence **moved** into the block, word for word — precedence is evaluation order, which is the block's subject. The clearly-stuck paragraph keeps only its reading and points forward. Accounting changed from "kept verbatim and qualified" to **moved**. D3 is satisfied by a move, which does not touch the words; the block names the sentence as that paragraph's so its opening clause resolves. +9 | fixed | "unless the user explicitly reverses that decision" deleted. D7 admits no exception; a contradicting later answer is a contradiction to surface, and permitting a reversal would let a finding be moved out of the set and back in to escape what it owes there. +10 | dissolved | The b11 OLD substring could never reach 0 because it survives inside its own replacement. Dissolves with Mechanism 1: no substring is named here any more. +11 | dissolved | The missing `b17`–`b18` pair. Dissolves with Mechanism 1; §7 now names `b17`–`b18` among the edits that owe a pair, and the plan builds it. +12 | fixed | The profile-change row is split in two: the profile change while the hold awaits its answer, next state the hold standing and no pass run; then the answer in each direction, only a resuming one starting the further pass. A single row would have to run a pass through an unanswered hold to be filled in. +13 | fixed | The partial-adoption guard is now an **owned residual of this change**, not assigned to the successor — whose scope is record durability and excludes the closure ordering by name, so the assignment named an owner that had not taken it. Same correction applied to §7's observability residual, which had the same defect. +14 | fixed | `b7`'s NEW text carries its reasons inline: the union because a cycle no single artifact governs has no set at all under the singular reading, the subtraction because a decline is the user's answer that the finding stays out. +15 | fixed | §2 recounts with the unit stated — one edit per contiguous replacement or addition at one site: **nineteen**, five outside the inventoried passages and fourteen inside, with §4 item 6 counted once, in §5(i). +16 | dissolved | The missing `c9` assertion. Dissolves with Mechanism 1 — and the underlying risk changed shape anyway, since `c9` is now **moved** rather than kept in place, so §7 lists the (c) replacement among the edits owing a pair. +17 | fixed | Full repository-relative path, and a rule added to §8 covering all of them. +18 | fixed | Same. +19 | fixed | Same. +20 | fixed | Same. +21 | fixed | Same. Whole-file sweep run: every backticked path now resolves, and each was checked to exist. + +## Notes + +- The moved `c9` sentence wraps across two lines in the spec, so `grep -F` on the whole phrase + counts 0. Verified by reading instead. That is the same false-red class Mechanism 1 documents, + reproduced in this session's own checking. +- `scripts/check-version-bump.sh main` was written as a bare basename and is now full, found by + the sweep rather than by a finding. diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-12.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-12.md new file mode 100644 index 0000000..18be264 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-12.md @@ -0,0 +1,22 @@ +BLOCKER | high | §9 "Moved, out of scope, and parked" | The spec says a rollback already has a defined stricter-reading transition, but the successor story explicitly says a rollback can remove the only text describing what the open cycle owes, and a revert also removes this spec's proposed "every suspension binding" extension | An open cycle whose ordering is reverted can have neither a valid close nor a defined suspension, exactly the review's Blocker condition, while the false "transition exists" claim hides the deferred gap | Delete the claim that the transition exists and state the unresolved rollback path accurately in the successor without rebuilding its mechanism here +MAJOR | high | §3 "What a suspension asks" and §5(b), (c), (e) | The block says any stop answer parks the whole suspension, but every proposed source pointer says the loop resumes once all answers have merely been given, and the §5(c) accounting even claims a "resuming answer" qualification and parked state that its quoted NEW text does not contain | On a pass combining a scope stop with a health stop, answering membership and choosing stop simultaneously requires resume and park, so the ordering has two next states | Make every source site defer to the block or require that every answer be resuming, and preserve the stop-parks exception in each affected replacement +MAJOR | high | §5(b) "What a loop absorbs" | The proposed b17-b18 replacement still says the loop resumes on the "revised artifact", although a membership decline can end the only hold without requiring or permitting a repair and the block explicitly allows the current artifact to be unrevised | A valid decline-only path has no compliant next pass unless the agent manufactures an unrelated edit, reproducing the no-progress terminal defect AC 4 forbids | Say current artifact, revised or not, and add an unrevised-decline transition to the verification table +MAJOR | high | §4 item 3 and §5(a) old-condition a16 | The spec calls Mechanics Severity the sole unscoped resolve duty and scopes it to the assigned fix set, but it keeps the separate standing command "fix Blocker/Major after each" at C:132 and W:339 unchanged and marks a16 kept | A declined out-of-set Blocker or Major is simultaneously exempt from resolution and still commanded to be fixed, leaving the two prompt copies internally contradictory | Scope a16 to the assigned fix set or replace it with a pointer to Mechanics, mark a16 replaced, and verify the OLD and NEW forms +MAJOR | high | §3 "First, clean completion" and §6 trigger equivalence | The block redefines both scope triggers by restating b11's decline exception and by defining an answered structural question as no longer new; b11 also defines its exception at source, while b13 remains unchanged and §6's equivalence check omits the answered-question condition | The claimed one-authority design has two definitions for membership and no source definition for re-raised answered questions, so either trigger can drift or be applied differently | Keep both complete trigger definitions in b11 and b13, including the within-cycle answered-question qualification, and have the block cite them without restating their predicates +MAJOR | high | §3 "What a pass is read from" and §3 one-authority table | The block computes the assigned fix set as the source definition plus accepted work minus declined work, while proposed b7 already defines the full governing-artifact union plus accepted work minus declined work | The same set has two authorities despite the table claiming b7 is its one definition, creating an immediate drift surface in the predicate every closure branch reads | Leave the complete set formula only in b7 and make the block cite the current assigned fix set from that paragraph +MAJOR | high | §3 "What a pass is read from" and §5(g) Severity | The block states the complete effective-versus-reviewer-written severity partition immediately after saying Mechanics Severity settles it, and §5(g) installs the same partition as that rule's source | The shipped prompt duplicates the severity rule while claiming the block only cites it, violating the one-authority principle and invariant 11's token-lean requirement | Put the partition only in Mechanics Severity and let the ordering identify each predicate by reference to the field Mechanics assigns it +MAJOR | high | §3 "Composition" and the retained c9 sentence | The block owns and states clean-completion precedence over clearly stuck, while the table assigns "the clearly-stuck reading and its precedence" to c9 and §5(c) preserves that separate precedence sentence verbatim | Evaluation order has two authorities even though the spec says it is stated once, so a later edit can change one precedence statement without changing the other | Move the verbatim c9 precedence sentence into the authoritative ordering and leave the clearly-stuck source paragraph defining only its reading, with a pointer back to the ordering +MAJOR | high | §3 "What a suspension asks" | The proposed decline binds for the rest of the cycle "unless the user explicitly reverses that decision", but settled D7 says it binds for the remainder without an exception, and no ordering resolves simultaneous or successive accept and decline answers | A true finding can be moved in and back out of the fix set to bypass its in-set resolve duty, and re-raised findings are no longer idempotent | Remove the reversal exception and treat a conflicting later answer as a surfaced contradiction unless a future settled decision explicitly changes D7 +MAJOR | high | §7 "The check" b11 pair | The asserted OLD b11 substring "loop like any other out-of-scope finding**, even when it opens no new question at all" is preserved verbatim inside the proposed NEW b11 sentence before the decline exception | The required working-tree OLD count can never become 0, so correct implementation fails the check and the claimed counterfactual is mechanically impossible | Choose mutually distinguishing single-line OLD and NEW substrings that include the old continuation versus the new decline exception +MAJOR | high | §7 "The check" passage coverage | The check claims paired assertions for every meaning-changing §5 passage edit but omits the b17-b18 replacement entirely, and the spec gives no exact single-line OLD substring for that quoted replacement | Both copies can retain the stale immediate-resume and revised-artifact rule while the block, all listed counts, and parity checks pass | Add a discriminating OLD and NEW assertion pair for b17-b18 and state the exact single-line substrings counted in each copy +MAJOR | high | §7 profile-change-with-hold transition | The named row says its next state contains both a still-standing hold and the further pass a profile change costs, then says to run it against both answer directions; before an answer the hold forbids that pass, while after either resuming membership answer the hold has ended | The verification cannot represent the requested hold-outlives-profile-change path without violating either the hold or the answer transition, so it can bless a pass run through an unanswered hold | Split this into profile change to still-suspended hold, then answer to park or resume, with only a resuming answer starting the required pass under the current profile +MAJOR | high | §9 partial-adoption deferral | The spec assigns a new guard covering the closure block and coupled edits to the successor, but that story neither scopes nor accepts such a guard and explicitly treats the closure ordering as out of scope | The admitted unsafe partial-adoption state has no owner, so downstream can indefinitely ship a clean predicate without its scoped duty or trigger qualifications | Either add the guard explicitly to the successor's desired outcome and acceptance criteria or leave it as an owned residual rather than claiming it belongs there +MINOR | high | §5(b) proposed b7 source text | The new governing-artifact union and declined-finding subtraction are constraints in shipped prompt text, but their reasons live only in this design spec rather than in the same clauses as required by prompt-standards item 6 | Downstream readers cannot judge why union and subtraction remain necessary, and the prompt change fails invariant 11 even if its mechanics are corrected | Add concise inline reasons to b7 for unioning every governing artifact and retaining a cycle's decline subtraction +MINOR | high | §2 "Settled inputs" edit counts | The spec says six standing sentences are outside the inventoried passages and seven fragments are inside, but §4 item 6 is explicitly passage i, while the inventoried edits comprise the two a edits, five b edits, the c trim, e edit and addition, g replacement, and i addition | The stated split does not match its own enumeration, undermining the mechanical accounting readers are told to trust | Recount with an explicit unit of counting and correct the outside-versus-inventory totals +MAJOR | high | §5(c) retained c9 verification | The spec quotes the settled c9 sentence as preserved byte-for-byte but supplies no exact single-line substring or parent-versus-working-tree assertion for it; parity only proves both copies agree with each other | Both copies can change the settled D3 sentence identically and all stated checks still pass | Name a single-line c9 substring and assert it exactly once in both parent and working copies, or compare the complete normalized sentence against the parent +NIT | high | §8 parent-artifact citation at line 686 | The cited path `…plan-a-rules.md:878` does not exist in the tree; the actual file is `docs/superpowers/plans/2026-08-29-review-loop-economics-plan-a-rules.md` | The mandated path-existence check fails and a reader cannot use the citation literally | Replace the ellipsis shorthand with the exact repository-relative path and line +NIT | high | §8 parent-artifact citation at line 687 | The cited path `…review-loop-economics-design.md:33` does not exist in the tree; the actual file is `docs/superpowers/specs/2026-08-28-review-loop-economics-design.md` | The mandated path-existence check fails and a reader cannot use the citation literally | Replace the ellipsis shorthand with the exact repository-relative path and line +NIT | high | §8 parent-artifact citation at line 687 | The cited path `…pass-floor-story.md:78` does not exist in the tree; the actual file is `docs/superpowers/stories/2026-08-28-review-loop-economics-pass-floor-story.md` | The mandated path-existence check fails and a reader cannot use the citation literally | Replace the ellipsis shorthand with the exact repository-relative path and line +NIT | high | §8 plugin file citation at line 689 | The cited relative path `workflow-init.md` does not exist at the repository root; the file the sentence means is `plugins/dev-workflow/commands/workflow-init.md` | A literal path-existence check fails even though W was defined earlier | Use the exact repository-relative path in the invariant-12 accounting +NIT | high | §8 plugin file citation at line 691 | The cited relative path `CHANGELOG.md` does not exist at the repository root; the actual changelog is `plugins/dev-workflow/CHANGELOG.md` | A literal path-existence check fails and the version-bump edit target is underspecified | Use the exact repository-relative changelog path +END OF FINDINGS (21 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-13-dispositions.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-13-dispositions.md new file mode 100644 index 0000000..107edd4 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-13-dispositions.md @@ -0,0 +1,56 @@ +# Gate-A spec pass 13 — dispositions (cycle awsf1ec771) + +Advisory companion, not a findings file. Findings: `gate-a-spec-awsf1ec771-pass-13.md`. + +**This round is a cut first and a repair second.** Measured before deciding: the spec was 785 +lines, of which §3's closure ordering — the design it exists to state — was **173, 22%**. The +bookkeeping about the change (§4 quoted OLD/NEW 83, §5 per-condition accounting 216, §6 parity +list 38, §7 verification substrings 119) was **456 lines, 58%**, and had been taking roughly half +the findings of every pass. Pass 12 had already run the experiment at small scale: moving the +verification substrings to the plan killed three Majors at once and regenerated nothing, because +a spec cannot build an exact check for text that does not yet exist. Daniel's decision: cut now. + +All four are **moved to the plan, not dropped** — the plan performs them against real files, +which is the only place they can be checked. Story acceptance criterion 5 is satisfied by the +plan's per-condition disposition list; §5 keeps the ten-passage map so nothing falls out of view. + +Result: **533 → 532 lines after trims, delta −253 from 785.** §3 unchanged in role and grown only +by the findings applied to it. + +1 | fixed | validated dismissal is a resolution route beside repair, discharging the duty for that finding without rewriting the pass that found it; a dismissal the reviewer keeps re-raising is regeneration and counts toward the clearly-stuck third condition, so the repeat false positive reaches a suspension instead of continuing forever +2 | fixed | the contradiction gets a state: the decline remains binding, the contradiction is surfaced as information, the cycle continues; withdrawal is a fresh explicit decision by the same authority, not a reversal by these rules, so D7 keeps its no-exception reading +3 | fixed | clean completion at or above the floor makes the cycle eligible to close; the closing amend Mechanics · Finishing the cycle describes is the closure itself, so nothing is closed before it and a profile, cited set or evidence entry changing in between still gates it +4 | fixed | `b8` tests membership against the assigned fix set as `b7` computes it, replaced rather than kept — §4 item 11 +5 | fixed | `b3` becomes a pure pointer to Mechanics and stops restating the four severity actions, in both copies — §4 item 9. W's pointer target also aligns to C (§6) +6 | fixed | §6's equivalence check names `b13`'s already-answered qualification alongside `b11`'s already-declined exception, compared in both directions +7 | fixed | the false claim is deleted; the block itself now carries both in-session consequences — a line per branch file is a distinct finding, and an answer binds to the finding or question as the pass that raised it recorded them. Cross-session recognition is named as the successor's +8 | fixed | the evidence entry's revalidation rule joins what closure reads and the one-authority table, cited and not restated, and joins the oracle's unmet-precondition list +9 | fixed | continue permits an unrevised artifact only where no repair is owed, with its reason inline; the third branch carries the same qualification +10 | fixed | §7's next-state table claim is narrowed to answer-state transitions once the predicates are established, with separate named checks for logical-pass completeness and each cited final-acceptance precondition. No fixture-per-predicate mechanism added +11 | fixed | the one-or-two answer count is scoped to scope-stop answers; a shared health continue-or-stop answer is additional and not counted among them +12 | fixed | §4 item 20 carries its reason in the shipped clause — unknown starting rules cannot waive an open hold — per prompt-standards item 6 +13 | dissolved | the count is recomputed against its own enumeration: twenty source edits, the block an addition beside them, twenty-one changes in all. §8's version-bump rationale moves with it +14 | dissolved | §1 now reads five outside the inventoried passages and fifteen inside, eighteen replacements and two additions — verified against §4's twenty rows +15 | dissolved | the OLD quotations left with the cut; §4 names each sentence and what changes about it, and the plan quotes it from the real file +16 | dissolved | same as 15 + +## Checks run before commit + +- **Precheck** `.context/spec-precheck.py`: exit 0. Fences balanced, 11 cited paths all exist, no + ellipsis shorthand, no OLD fenced blocks left to verify. +- **Counts against enumerations:** inventory 135 ✓; §4 twenty rows, 18 replacements + 2 additions + ✓; §3 "nine rules" against nine table rows ✓; §5 ten passage rows against the ten inventoried + passages ✓; §4 items 6–20 all referenced by the passage map, 1–5 correctly outside it ✓. +- **One-authority probe:** the block cites nine rules and defines none; each has one definition in + the shipped text, and the one exception — `c9`'s precedence sentence moving *into* the block — + is stated as such, because precedence is evaluation order. +- **Transition walk**, the check pass 13's finding 2 exists for. Every state named, its input, its + next state: clean at/above floor → eligible → amend → **closed**; clean below floor with no + suspension → continue; not clean, no suspension → continue with repair where owed; membership + stop → accept or decline → hold discharged → continue; question stop → decision → continue; + both triggers → both answers → continue; clearly-stuck or two-tell → continue (reading consumed) + or stop → **parked** → later continue → resumes; several suspensions → one stop answer parks the + whole; contradiction to a decline → surfaced, cycle continues; re-raised validated dismissal → + regeneration → clearly-stuck → continue or park; zero-finding pass → eligible → closed; + precondition changed before the amend → not closed, fix and re-review → continue. **Close or + park is reachable from every state; no state returns itself with its input consumed.** diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-13.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-13.md new file mode 100644 index 0000000..eee56c7 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-13.md @@ -0,0 +1,17 @@ +BLOCKER | high | §3 "The four standing duties, classified" | A dismissal is defined as the author's judgement that a finding is false, but the resolve duty is discharged only when a later validated pass carries no in-set Blocker or Major, and neither the ordering nor retained a22 gives a re-raised dismissed finding a suspension path; a repeat false-positive Blocker can be in-set, have no structural question, lack the genuine repair attempts clearly-stuck requires, and expose only one tell | Every pass takes the continue branch forever, so this reachable cycle can neither close nor suspend | Restore dismissal as a stated resolution route without rewriting the pass that found it, require the later clean pass, and give a repeatedly re-raised validated dismissal a defined suspension transition such as eligibility for the clearly-stuck reading +BLOCKER | high | §3 "What a suspension asks, and what ends it" | A later answer contradicting a binding decline is ordered to be surfaced, but this newly introduced surface is not one of the three suspensions, has no required answer, and has no resume, park, continue, or close transition | Once the contradictory answer exists the cycle has no executable next state, meeting the stated Blocker condition of being unable both to close and to suspend | Preserve D7 by stating what the contradiction does to state, for example that the original decline remains binding and the cycle either continues after an informational surface or parks until the contradictory answer is withdrawn, and cover that transition in §7 +MAJOR | high | §3 "First, clean completion" versus Mechanics "Finishing the cycle" | The block says the qualifying pass closes the cycle, while the standing Mechanics text at C:827-829 and W:1011-1013 says the cycle is closed only afterwards by the real amend | A profile, cited set, or evidence entry can change between pass and amend, and readers can treat the cycle as already closed under the block even though the standing final-acceptance rules still forbid closure | Define clean completion as eligibility to perform the required closing operation, or define it to include that operation, and make both source passages name one actual closure transition +MAJOR | high | §5(b) b7-b8 accounting | New b7 computes the assigned fix set from governing scopes plus acceptances minus declines, but retained b8 still says a finding is in-set when repairing it stays inside "that scope"; after a governing scope broadens to include a previously declined finding, b7 keeps it out while b8 can put it back in | The same re-raised finding can either avoid the membership trigger or re-enter the resolve duty, defeating D7 and AC 3's requirement to qualify every modified rule | Rewrite b8 to test membership against the fully computed assigned fix set defined by b7 rather than independently against the governing scope, mark b8 replaced, and add it to the paired verification list +MAJOR | high | §4 item 3 and §5(b) b3 | The spec calls Mechanics Severity the one definition of what each severity demands, yet retained b3 still spells out "Blocker/Major resolve, Minor/Nit collect and never iterate" after pointing to Mechanics | The severity action has two definitions that can drift, contrary to the mandated one-authority design and the spec's claim that each cited rule has exactly one definition | Make b3 a pure pointer to Mechanics without restating the four actions, account for the removed operative words, and add the source edit to §7 +MAJOR | high | §6 b11-b13 equivalence table | The table says it compares the complete edited predicates, but its b13 rows omit the new condition that "new" excludes a structural or contract question already answered in this cycle | Both copies can lose or alter the answer qualification and still receive the claimed b13 equivalence result, reintroducing the repeated question stop the qualification exists to prevent | Add the already-answered qualification to the semantic equivalence table and require the parity evidence to compare that condition in both directions +MAJOR | high | §3 "Two things the ordering names and does not define" | The live b11 and b13 qualifications require recognizing the same finding and the same structural question on a later pass, but the only shipped interim identity rule distinguishes lines across Gate-B branch files; the prose then falsely says both in-session binding consequences are stated in the fenced block even though only the cross-branch rule is there | A changed or paraphrased re-raise can be treated as already answered and close over a new duty, or treated as new and re-suspend indefinitely, while the successor owns the deferred finding-sameness rule and owns no question-sameness rule | Remove the false claim and either defer the dependent re-raise qualifications and verification rows with their identity dependency, or state a conservative session-local association rule without rebuilding the deferred durable record +MAJOR | high | §3 one-authority table "final-acceptance preconditions" | The block cites only C:116-118 and C:760-764 for final acceptance, omitting the standing evidence-entry rule at C:725-729 and W:911-915 that a changed revalidation invalidates the clean pass and requires fix and re-review | The ordering can be read to close on a clean at-floor pass whose evidence entry changed, directly contradicting a live closure precondition in both copies | Add the evidence-entry source to the one-authority citation and to the closure oracle, referring to it without restating its rule in the block +MAJOR | high | §3 health "Continue" transition | Continue after a clearly-stuck or two-tell suspension is said to produce a further pass even on an unrevised artifact, without the qualification used by the Gate-A cadence that revision is still required wherever effective severity and scope require repair | A user can continue past an in-set effective Blocker or Major and spend passes on unchanged text, violating the resolve duty and potentially obtaining a nondeterministic clean result without the required repair | Permit the unrevised transition only when no repair is owed; otherwise require the severity and scope repair before the post-answer pass +MAJOR | high | §7 next-state table and oracle | The verification is introduced as covering every stop with every input the rule reads, but its rows and columns omit validated logical-pass completeness, the current findings and severity fields, health coverage and history, and evidence-entry revalidation, and several rows begin from already-derived labels such as clean or health reading | The table can award the expected transition without demonstrating that the predicates or closure preconditions producing it were present, repeating the incomplete-input defect it says distinguishes this instrument from fic2 | Narrow the table's claim to answer-state transitions after predicates are established, and add separate named checks for logical-pass completeness and every cited final-acceptance precondition without adding the parked fixture-per-predicate mechanism +MINOR | high | §3 hold discharge count | The hold rule says one answer is required for a single-trigger finding and two for one carrying both, but a finding can carry both scope triggers and also be surfaced by the clearly-stuck reading, which adds the shared continue-or-stop question | In the explicit three-suspension case the text gives both two-answer and three-answer readings for when that finding's hold ends, even though the later composition sentence prevents an immediate resume | Scope the one-or-two count explicitly to scope-stop answers and state that any shared health answer is additional when a health suspension also applies +MINOR | high | §4 item 6 and invariant 11 | The shipped unknown-start addition is only the bare constraint "every suspension binding"; its reason appears in the design discussion but not in the prompt clause or surrounding source list | A downstream reader cannot tell why uncertainty must preserve a suspension, so this new prompt rule fails prompt-standards item 6 despite §8 claiming the changed prompts pass all 12 items | Add a concise inline reason in both prompt copies, such as that unknown starting rules cannot waive an open hold +NIT | high | §2 "Nineteen edits in all" | Under the section's own unit of a contiguous replacement or addition, the enumeration contains nineteen source edits plus the new §3 closure-block addition, so the total is at least twenty even when each mirrored C-W pair is counted once | The mechanical edit count and the repeated "nineteen" version-bump rationale do not match the spec's own verification inventory | Call the nineteen count source edits excluding the block, or change the total and the §8 rationale to include the block +NIT | high | §1 "six standing sentences" | The six-item §4 enumeration contains five sentence edits and one strict-reading list addition, not six standing sentences falsified or made ambiguous by the ordering | The stated count does not match its own enumeration | Say five standing sentences plus one list addition +NIT | high | §4 item 2 OLD quotation | The source sentence found at C:565-566 is "Ask for one line per finding and a literal `NO FINDINGS` when a pass is clean — the explicit clean signal is what lets you exit the loop:", and at W:756-757 it wraps differently, but the spec quotes only its tail with an ellipsis after claiming every sentence here is quoted whole | The mandatory OLD-sentence mechanical check cannot verify the displayed quotation as a whole source sentence | Quote the complete sentence and note both real wraps while leaving verification substrings to the plan +NIT | high | §4 item 4 OLD quotation | The source sentence found at C:653-655 and W:839-841 is "The Blocker/Major filter, the file-first findings protocol and the clean-final-pass rule are unchanged.", but the OLD field contains only "clean-final-pass rule are unchanged." after the section claims each sentence is quoted whole | The mandatory OLD-sentence mechanical check succeeds only by treating the field as an unlabelled fragment | Quote the complete sentence, while keeping the later plan responsible for choosing its discriminating substring +END OF FINDINGS (16 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-2-dispositions.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-2-dispositions.md new file mode 100644 index 0000000..2554e73 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-2-dispositions.md @@ -0,0 +1,19 @@ +# Gate-A spec · cycle awsf1ec771 · pass 2 dispositions (advisory companion, not a findings file) + +1 | fixed | clean predicate names the fix set only ("no in-set Blocker/Major at effective severity, no scope stop surfaced"); a decline keeps a finding out and never excuses one that is in; a declined finding the set later includes owes resolution — §3 block, §4 rule, decline block +2 | fixed | a re-raised finding matching a decline of this cycle on all five fields raises no hold and no membership stop; any field difference or uncertainty = new finding, new stop — §3 suspension paragraph, §4 rule, decline block +3 | fixed | a hold ends when every answer the finding requires is given: one for a single-trigger finding, both the decision and the membership answer for a two-trigger one — §3 +4 | fixed | four-duty classification paragraph added to the block (floor: precondition on closure; resolve: precondition on closure and on cleanliness; hold: ordering participant; no-clean-credit: ordering participant, discharged by nothing) — §3 +5 | fixed | every predicate reads the validated findings file or files of the logical pass as one set; a full Gate-B pass has two and one branch alone is already incomplete — §3 first paragraph +6 | fixed | close sentence is read under the floor section's final-acceptance preconditions (set and profiles re-read before final acceptance; a mid-pass change makes the pass not final and costs the further pass) — cited by lead phrase C:116–118, C:760–764, not restated — §3 +7 | fixed | c18 marked REPLACED, narrowed to the scope stop, with authority (story §1's "surfaced finding") and why it is no closure bypass (stuck surfaces are unclean by step one; two-tell below the floor cannot close; at/above it D2) — §5(c) +8 | fixed-partly | recovery reading added (§4 paragraph + decline-block disclosure): accepted obligations and standing suspensions live in the session and optionally the working record; a cycle that cannot recover them has no identity and starts a new cycle under C:419–420, which costs passes and closes nothing; the reviewer re-raises unrepaired findings and out-of-scope ones re-ask as membership stops; residual (a finding not re-raised is lost as any missed finding is) stated. DISMISSED: the demand for a new mandatory cycle-attributed record — it would add a mandatory artifact D10 refused for the same class of state, and the existing no-identity rule already gives the safe direction (repeat question, never silent close) +9 | fixed | unknown-start item bounded in time: only declines not attributable to the nonce the fallback minted are treated as absent; declines recorded under that nonce are honoured — §4 item 5 +10 | fixed | recording is a precondition to resuming: the commit carrying the decline is made before the next pass runs; the working record may carry it meanwhile and does not bind — §4 and decline block +11 | fixed | Q6 partition: slot absent = unavailable; slot present, accepted when it ran, now invalid = pass stays counted, series read `?`; known-incomplete = excluded; present-but-invalid with no record of acceptance reads as incomplete (costs a pass rather than credits one) — §7 +12 | fixed | Q6 first step is the root check (`.context/codex-reviews/` under the checkout's top level, the same root passed as workingDirectory); a mismatch is reported as wrong-root, never as absent history; the sentence says this check is what detects it — §7 +13 | fixed | narrowness bound qualified "as far as distinct nonces allow"; shared or redrawn nonce makes a replayed decline bind to the wrong cycle — decline block +14 | fixed | reciprocal one-contract sentence in both shipped blocks; rollback with an open cycle pointed at the existing activation rule (C:153–154, C:165–166), no new mechanism — §3 block last sentence, decline block, §4 item 9 +15 | fixed | next-state walk gains the stateful rows (matching vs changed-field re-raise, scope broadening after a decline, recovery after accept and after stop, full Gate-B branch union, rollback, partial adoption both ways) with record state and governing-scope state as explicit input columns — §9 +16 | fixed | "any decline elsewhere is a defect" narrowed to closure rules; the transport/attribution/activation/threat sites that must name it enumerated, with the opposite defect stated — §4 +17 | fixed | §8 cites Plan C Task 19 (line 970), Task 20 (line 1047) and Task 23's second point (line 1298) as the prior record; the story-§5 claim removed diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-2.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-2.md new file mode 100644 index 0000000..04c676c --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-2.md @@ -0,0 +1,18 @@ +BLOCKER | high | §3 "A pass is clean" / §4 "Recording a decline" | The clean predicate exempts an in-set effective Blocker or Major when it is "matched by a decline", even though D5 and the decline block say an in-set Blocker or Major can never be declined; because decline records are unverified, a fabricated in-set record—or a real record whose finding later enters the set after scope broadens—satisfies the exemption literally | An in-set Blocker or Major can be treated as clean without resolving, creating the exact gate-off route the proposed gate-off list admits exists | Remove the decline exception from the in-set Blocker/Major predicate; make a decline affect only an otherwise matching finding while it remains outside the assigned set, and require resolution if that finding later becomes in-set +MAJOR | high | §3 "Second, only a pass that is not a clean completion can suspend" | The membership-stop trigger applies to every out-of-set finding without exempting a finding matched by this cycle's existing decline, while cleanliness separately requires that the pass "surfaced no scope stop" | Re-raising the same genuinely declined finding triggers another scope stop and makes every such pass unclean, so D7 does not make the decline idempotent and a persistent reviewer finding can prevent clean closure indefinitely | State that a valid matching decline suppresses a new hold and scope-stop surface for that same finding, while any five-field difference or uncertainty creates a new membership stop +MAJOR | high | §3 "What a suspension asks, and what ends it" | For one out-of-set finding that also opens a structural or contract question, the text requires both a question decision and accept-or-decline, but says the single hold on that finding ends with "any" answer | One answer can end the hold while the other required question remains unanswered, contradicting the composition rule and leaving no deterministic next state for this explicitly reachable two-trigger path | Define either one hold per required answer or one aggregate hold that ends only after both the question decision and membership answer have been supplied +MAJOR | high | §3 "The closure ordering" | The shipped block never explicitly classifies each of the four standing duties as an ordering participant or a closure precondition, and it does not say for each precondition exactly what it gates and what discharges it | Story AC2 is unmet, so readers must still infer the role of the floor, Blocker/Major resolution, the open surfaced finding, and the no-clean-credit rule from separate sentences | Add one compact four-duty classification that labels each duty, names the closure state it gates, and names its discharge condition without changing the settled semantics +MAJOR | high | §3 "First, clean completion" | Cleanliness is read on "the pass's validated findings file" singular, but a full Gate-B logical pass has two independently written branch files and the proposed ordering never says its predicates read their union | A clean spec branch can be read as closing the cycle while the quality branch carries an in-set Blocker/Major or a scope-stop finding, leaving the concurrent full-review path ambiguous despite the existing both-files validation rule | Say "validated findings file or files" and require cleanliness, scope triggers, and every suspension reason to be evaluated over the union of all required branch findings for the logical pass +MAJOR | high | §3 "A clean pass ... closes the cycle" | The absolute close rule does not preserve the standing preconditions at CLAUDE.md:116-121 and 745-776 / workflow-init.md:323-328 and 931-962: the cited set and profiles must be re-read before final acceptance, a mid-pass change makes the pass non-final, and any profile change costs a further pass | A reader can close on a clean or zero-finding pass whose governing header or profile changed during the pass, directly contradicting text the spec leaves standing; the rationale's claimed "further pass" reading is not present in the shipped block | Make clean completion subject to the existing final-acceptance preconditions, explicitly including stable current headers/profiles and the required post-change further pass, or place those preconditions in the fixed evaluation order +MAJOR | high | §5(c) old-conditions accounting | Current CLAUDE.md:242-246 and workflow-init.md:445-449 say that a clearly-stuck surface leaves the finding open and credits no pass as clean, but the replacement moves the no-clean rule only to scope stops and nevertheless labels c18 "moved" rather than replaced or dropped | The spec silently narrows both the current clearly-stuck condition and the story's stated standing duty; a below-floor pass that is clean at effective severity but trips a raw-severity health stop can remain credited clean | Preserve the no-clean-credit condition for every surface to which it currently applies, or mark c18 replaced, state the deliberate semantic change and its authority, and reconcile it with the settled four-duty requirement +BLOCKER | high | §3 "Continue resumes" / §4 "What the body does not record" | The next fix set includes "obligations already accepted" and a stop answer leaves a resumable standing suspension, but the spec explicitly gives accepted findings and stuck/two-tell answers no record while the working record remains optional | After compaction, workspace loss, or recovery into a new cycle, an accepted in-set Blocker/Major or an outstanding stop can disappear from all recoverable state; findings files hold inventory rather than resolutions, so a later clean pass can close without the accepted repair or continuation decision | Persist accepted obligations and standing suspension answers in a mandatory cycle-attributed record, define recovery and conflict handling, and carry that state through closing and squash bodies; disclosure alone is insufficient for state that controls closure +MAJOR | high | §4 item 5 "The unknown-start fallback" | "Decline records treated as absent" is unbounded in time, so it also tells an unknown-start cycle to ignore a decline newly made and recorded after the fallback minted that cycle's post-rule nonce | The same out-of-set finding can raise a hold on every later pass even after repeated explicit declines, violating D7 and leaving the unknown-start path permanently unable to benefit from the only release record it can create | Treat only inherited, unattributable declines as absent at activation; once the fallback establishes a new nonce, honor declines explicitly made and recorded under that nonce +MAJOR | high | §4 "When it is recorded" | A decline is assigned to the cycle's "next commit body", but neither the block nor the surrounding Gate-A workflow requires that body to be written before the next review pass runs | A Gate-A cycle can resume and run several passes with the decline existing only in chat, so session loss re-applies the hold and D7's remainder-of-cycle binding is not actually provided by the specified transport | Make successful recording an atomic precondition to resume—amend or create the cycle commit before the next pass—or require the cycle working record while no commit body yet carries the decline +MAJOR | high | §7 "Where an earlier pass's findings file is unavailable" | A historical slot that is present but fails validation is called "an incomplete pass, already excluded", but validation occurs when a pass is accepted and the file can be corrupted, truncated, or replaced afterwards | A previously valid counted pass is retroactively misclassified, and the unavailable-history behavior remains undefined for exactly the later-corruption state that now prevents recomputing its trend and comparisons | Distinguish a pass known to have been incomplete from a formerly accepted pass whose historical artifact is now unusable; preserve the pass count, mark affected series unknown, and disclose reduced sensitivity for the latter +MAJOR | high | §7 "Where an earlier pass's findings file is unavailable" / prompt standard 10 | The new diagnostic bundles fresh checkout, cleared `.context`, and resume elsewhere under an indistinguishable absent-path observation, gives no cause-specific checks or fixes, and asserts that wrong-root resolution is "the existing stop" even though CLAUDE.md:298-320 and workflow-init.md:492-514 define working-directory/deletion rules, not a check that distinguishes a missing historical slot at the wrong root | A wrong checkout can be reported as benignly unavailable history and continue with reduced sensitivity, while invariant 11's diagnostic-state requirement remains unmet | Specify the root check first, then partition genuinely absent history from present-but-unusable history using observable tests and give each state its action; do not claim an existing stop unless the cited standing text actually detects this read path +MAJOR | high | §4 "What it is worth" / cycle nonce residual | The decline block says a false record is bounded to "one fully identified finding, one cycle" without the existing nonce qualification at CLAUDE.md:455-468 / workflow-init.md:649-662 that two cycles sharing or redrawing a nonce are indistinguishable and inherit each other's attribution bound | A replayed decline can release a hold in a later colliding cycle, so the proposed narrowness statement contradicts the standing residual and overstates the abuse boundary | Qualify the one-cycle claim with "as far as distinct nonces allow" and explicitly state that collision or reuse can make a replayed record bind to the wrong cycle +MAJOR | high | §4 item 9 "The one-contract paragraph" / compatibility and rollback | The coherence stop is added only to one of the clauses that `/workflow-init` may merge in part; if the closure block or decline block is adopted without the updated contract paragraph—or the new text is rolled back while a cycle that started under it remains open—the surviving text has no local check that its counterpart exists | Downstream partial adoption can run an ordering with no decline transport or a release record with no hold, and an open new-rule cycle cannot reliably finish under the rules it started with | Put reciprocal presence/version checks at the new ordering and decline blocks themselves or make their scaffolded adoption atomic, and define how an open cycle loads its starting rules across a rollback +MAJOR | medium | §9 "The named verification of the risk path" | The proposed next-state walk covers named stops but omits the stateful transitions most likely to fail: a matching versus changed-field re-raise, scope broadening after a decline, state recovery after accept or stop, full Gate-B branch union, and rollback or partial adoption | The high-risk named verification can pass while the closure bypasses and non-idempotent paths introduced by this design remain unexercised | Add rows for each temporal, recovery, concurrent, and compatibility transition, with the record state and governing-scope state as explicit inputs rather than precomputed outcomes +MINOR | high | §4 "Which rules the answer modifies" | The spec says any occurrence of "decline" at a rule outside three named sites is a defect, then deliberately adds decline behavior to the unknown-start fallback, gate-off surface, nonce set, closing-message carry, squash carry, and one-contract rules | The implementation and parity review are given a self-contradictory audit instruction and can reject required text or overlook which qualification AC3 actually constrains | Narrow the assertion to rules whose closure behavior the user's answer qualifies, and separately enumerate the transport, attribution, activation, and threat-model rules that must mention declines +MINOR | high | §8 "The slot discriminator — dissolved" | The statement that "this section and the story's §5 are the record" is mechanically false: story §5 at lines 141-154 contains the profile and ordering open questions but no slot-discriminator decision; the actual deferred-production and `rle` exception record is in plan C Task 19 at lines 970-980 and 1298-1314 | The dissolution points readers at a section that cannot substantiate it, weakening the audit trail for a deliberately unshipped production | Cite the two plan C drop-note locations as the durable prior record and remove the unsupported story-§5 claim, or add the settled dissolution to the story's designated record if that is required +END OF FINDINGS (17 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-3-dispositions.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-3-dispositions.md new file mode 100644 index 0000000..19a7c5c --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-3-dispositions.md @@ -0,0 +1,12 @@ +1 | fixed | Third branch added: a pass that neither closes nor suspends continues. c14 re-marked moved to that branch, not read out of a negation. +2 | fixed | Resolve duty discharged by validated repair OR validated dismissal with its one-line why (the standing advisory rule); dismissal-is-not-decline stated in one sentence. +3 | fixed | (b), (c) and (e) resume sentences now defer to the aggregate condition — every answer every standing suspension requires, any decline among them recorded first. No local "resumes the moment" reading survives. +4 | fixed | A matching decline suppresses the membership trigger only; a question stop still fires unless that same question was already answered this cycle. Stated in the ordering block and the decline block, cross-referenced. +5 | dismissed-in-part | No new mandatory record: D10 refused one for the same class of state, and the no-identity rule (C:419-420 / W:613-614) already answers the loss in the safe direction. Fixed instead: the resume fix set is narrowed to recoverable accepted obligations; a new cycle names in its first pass report the open cycles it found and did not adopt (item 10, no new record); the residual is stated — a replacement cycle can close the same artifact while an older one stays open, nothing detects it. +6 | fixed | A Gate-A decline with no artifact revision to carry it goes in an empty commit, the destination the human-exception rule already blesses. No new commit operation invented. +7 | fixed | Rollback leaving a cycle unable to establish its starting rules takes the unknown-start fallback, extended by item 5 with the loop rules. Residual stated: the rule text a cycle ran under is recorded nowhere. +8 | fixed | The (g) raw-versus-effective split joins the reciprocal one-contract coupling; both asymmetric partial-adoption states named in item 9 and in both shipped blocks. +9 | fixed | A gap does not break the series: consecutive available pass numbers compare and the report names the missing pass numbers. Reason stated — refusing to compare would silence detection where the record is thinnest. +10 | fixed | Both stuck conditions — the Blocker curve and the regenerating Blocker/Major findings — read the pre-ceiling field, with the reason that a predicate split across both fields could not be read. +11 | collected | Squash carry has no dedup identity for a decline restated in several bodies, so one decision can multiply into many identical records. Minor: collected for the plan, not iterated. Candidate key: cycle nonce plus the five finding fields, exact duplicates collapsed, conflicting copies stopping. +12 | fixed | Severity corrected MINOR -> MAJOR on the instrument carve-out: the check produced a false green, deleting the old (g) sentence without installing the new paragraph passed both the removal assert and the parity diff. Assert pair added on the replacement's lead phrase, 1 in C and 1 in W, 0 in the parent tree. diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-3.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-3.md new file mode 100644 index 0000000..c27102e --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-3.md @@ -0,0 +1,13 @@ +BLOCKER | high | §3 "Composition, and what cannot happen" / §5(c) | The ordering has no fall-through transition for a clean, non-zero-finding pass below the floor when no suspension predicate fires; the accounting says the old "carrying a Minor keeps looping" condition is preserved merely by reading the close predicate in the negative, but "does not close" does not say to continue | This reachable Minor-or-Nit-only path neither closes nor suspends and therefore meets the review's explicit Blocker condition | Add an explicit final branch that a non-closing pass with no suspension continues to the next pass, and mark the old keep-looping condition as moved to that branch +BLOCKER | high | §3 "The four standing duties, classified" | The Blocker/Major duty is discharged only by "repair of every" in-set finding, while retained §5 text permits a finding to be validated and dismissed with a one-line why and the new question-stop rule expressly permits an in-set finding to be dismissed | A false-positive in-set Blocker or Major can be dismissed and absent from the next pass yet can never satisfy the repair-only precondition, leaving the cycle unable to close and with no suspension to resolve | Define resolution to include either a validated repair or a justified dismissal, followed by the required clean pass, while keeping decline unavailable to in-set findings +MAJOR | high | §5(b)/(c) old-conditions replacements | The local resume sentences still say a membership stop resumes the moment its membership answer is given, a question stop resumes once its question is answered, and a clearly-stuck suspension resumes on a continue; the central composition rule instead requires every co-occurring answer to resume, and a decline additionally must reach a commit body before the next pass | A dual-trigger finding or simultaneous suspension can resume with another answer outstanding, and a declined finding can resume before its binding record exists, so the edited copies retain contradictory transition rules at the exact rules AC3 requires to be qualified | Change every local resume sentence to defer to the aggregate condition: all answers required by every simultaneous suspension must resume, and any decline must first be recorded as Mechanics requires +BLOCKER | high | §3 "Second" / §4 "Recording a decline" | A five-field decline match is said to raise "no hold and no stop", although D6 makes decline an answer only to the membership stop and the separate question-stop predicate is not one of the five recorded fields; the narrative version says no membership stop but still suppresses every hold | The same five-field finding can later open a new structural or contract question, especially after context or governing scope changes, and a prior membership decline can then suppress the question hold and allow a Minor-or-Nit pass to close without the user's required decision | Make a matching decline suppress only the membership-stop trigger and its share of the hold; preserve any independently applicable question stop unless that specific question has already been answered +BLOCKER | high | §4 "What the body does not record" | Accepted obligations and standing stuck or two-tell suspensions are deliberately session-only with an optional working record, and loss starts a new cycle while merely leaving the old cycle open; nothing prevents that fresh cycle from reviewing and closing the same artifact without recovering the accepted obligation or the user's later-continue requirement | Compaction or workspace loss can turn a mandatory accepted repair or an explicit stop into unreachable state and let a replacement cycle ship past it, so the claimed safe direction "costs passes and closes nothing" holds only at the instant the new cycle starts | Persist closure-controlling accepted obligations and standing suspensions in a mandatory cycle-attributed record, or prohibit any replacement cycle from closing the same artifact until the old cycle is explicitly resolved and define how that check is made +MAJOR | high | §4 "When it is recorded" | A Gate-A decline must be written into the next spec or plan revision commit before another pass, but the reachable case where the declined out-of-set finding is the only finding produces no artifact revision and the spec defines neither a message-only amend nor an allowed empty record commit | The hold is nominally released but the cycle cannot lawfully run its required later clean pass without inventing a commit operation or violating the durability precondition | Specify the exact zero-content Gate-A path, such as amending the current artifact commit's message or creating an explicitly permitted empty record commit, and state its hook/review consequences +MAJOR | medium | §4 item 9 "rollback" | The spec says an open cycle needs no rollback mechanism because it finishes under its starting rules, but neither the nonce, commit records nor optional working record identifies or preserves the rule version; after a full rollback the current prompt can be internally coherent while no longer containing the ordering the cycle is told to apply | A recovered open cycle cannot reconstruct its governing transitions and can silently run under the restored rules or remain indefinitely open, so the compatibility claim is not operationally supported | Record a recoverable starting-rule revision or define a history lookup and failure stop; treat inability to recover that exact rule text as an explicit suspended state rather than asserting the existing activation sentence is sufficient +MAJOR | high | §4 item 9 / §5(g) partial adoption | The one-contract expansion couples only the closure ordering and decline record, but the ordering also states raw-severity health semantics while the separately merged Mechanics paragraph currently says that question is unsettled and mandates a stop; adopting either the ordering or the new paragraph without the other yields a contradiction or a dangling "ordering above" reference that none of the reciprocal checks names | Downstream partial adoption can make identical demoted findings produce either a computed tell, a mandatory unresolved-question stop, or an undefined reference, defeating the compatibility stop the design claims to provide | Include the severity/health replacement in the reciprocal one-contract coherence rule and enumerate both asymmetric partial-adoption states, or keep the severity decision in one atomic block referenced from the other site +MAJOR | high | §7 "Q6 — the pass-4 report without prior-pass history" | The proposed text says trend and require-withdraw comparisons read the passes present, then concludes that a comparison spanning a missing pass cannot be seen, without defining whether non-adjacent available pass numbers may be compared or whether a gap partitions the series | Different readers can count different tells from the same reduced history, and because two tells mandate suspension the unavailable-history path has no deterministic next state | Define comparison adjacency explicitly: either gaps break the series and only consecutive available pass numbers compare, or non-adjacent observations compare and the reduced-sensitivity claim is narrowed accordingly +MAJOR | medium | §3 evaluation field / §5(g) handed-over severity question | The raw/effective split names counts, clusters and tell thresholds as raw and cleanliness/resolve as effective, but the clearly-stuck predicate combines a raw Blocker curve with "Blocker or Major findings that keep regenerating" and never says which severity that latter condition reads | A reviewer-written Blocker demoted to Minor can either satisfy or fail the third stuck condition depending on inference, changing whether the mandatory clearly-stuck suspension exists | State explicitly that both the Blocker-curve condition and the regenerating-finding condition read raw severity, or that the curve is raw while the regenerating findings are effective, and give the reason for the chosen split +MINOR | high | §4 decline carry / squash merge | Each decline is restated in every later cycle body, while the squash rule says to copy every decline record in the range and supplies no deduplication identity | One user decision can be multiplied into many identical records in the squash body, making the transport non-idempotent and obscuring genuine duplicate or conflicting decisions | Define squash carry as one copy of each distinct decline decision keyed by cycle nonce and the five finding fields, with exact duplicates collapsed and conflicting copies stopping +MINOR | high | §9 "The check" | The only assertion for the handed-over severity change checks removal of the C-only ownership sentence; parent W already has count zero, and deleting that sentence from C without installing the new raw-severity paragraph makes both that assertion and the parity diff pass | The evidence can report the central demotion/loop-health outcome as verified when it was omitted from both shipped copies | Assert the new replacement paragraph's lead phrase exactly once in C and W and zero times in the parent tree, in addition to checking removal of the old ownership sentence +END OF FINDINGS (12 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-4-dispositions.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-4-dispositions.md new file mode 100644 index 0000000..24d596e --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-4-dispositions.md @@ -0,0 +1,43 @@ +# Gate-A spec pass 4 — dispositions (cycle awsf1ec771) + +Advisory companion. Not a findings file; it participates in no pass validation. +Spec revised at 07c88a1 → this revision. Findings: `gate-a-spec-awsf1ec771-pass-4.md`. + +**Scope stop, surfaced and ANSWERED.** Finding 1 opened a new contract question — whether §5 +gains a record for accepted obligations — and was surfaced to Daniel with the finding still +open, per §5's scope-stop rule. **Daniel's answer, 2026-09-10: accept.** The change ships a +second label, `Accepted:`, sharing the decline record's form, transport, nonce and carry rules. +This is authorised scope growth beyond the story's §2. + +**This cycle writes no `Accepted:` record of its own.** It began before these rules ship and +finishes under the rules it started with, which have no such record. This file and +`gate-a-spec-awsf1ec771-resume.md` are what today's rules provide for that acceptance. + +1 | accepted-and-shipped | Third raise, and the D10 dismissal was wrong: D10 settles unavailable-pass-history reporting, not loss of closure state. Surfaced as a scope stop; Daniel accepted. Ships the `Accepted:` label plus the honest residual — a replacement cycle can still close the same artifact, and the record makes that discoverable, not impossible. +2 | fixed | The question-stop resume clause (`b18`) now takes the same aggregate precondition as `b12`. No local "resumes once that question is answered" reading survives. +3 | fixed | Q6 self-contradiction removed. One gap policy: consecutive available passes compare across a gap, visible but weaker evidence; a comparison needing the missing pass as an endpoint is unavailable. +4 | fixed | Q6 partition rebuilt on two observables — slot present/valid, and pass known-accepted. Acceptance-unknown is not counted toward the floor, series `?`, disclosed: crediting an unvalidated pass is the dangerous direction. +5 | fixed | The root-detection claim is withdrawn. A slot written under another root is stated as indistinguishable from an absent one; where the current root cannot be established the cycle stops, and the recovery is to re-run from the correct root. +6 | fixed | The (g) severity paragraph gains its own reciprocal one-contract sentence. All three independently mergeable pieces now name the other two. +7 | fixed | The rollback overclaim is withdrawn. A rollback that removes the fallback text leaves a cycle that stops and is a human's to resolve; it does not close. Residual named: no record identifies the rule revision a cycle started under. +8 | fixed | §8's dissolution claim is narrowed to post-rule and unknown-start cycles. Pre-rule cycles still exist, are bounded and self-terminating, and two observably live ones are serialized by a human rather than by a shipped production. +9 | fixed | `b12` marked **replaced**, with its old immediate-resume condition enumerated and the kept/changed halves stated. `b18` marked replaced for the same reason. +10 | fixed | The 135-condition inventory is committed as `docs/superpowers/specs/2026-09-10-loop-rule-consolidation-condition-inventory.md`, machine-local paths stripped, with a header stating it is a snapshot against 7c0d475 and the id definition the accounting cites. +11 | fixed | "the author's judgement about the fix set" → "about the finding's repair severity", with an explicit clause saying set membership is a separate predicate this must not be read as touching. +12 | fixed | Continue branch: below the floor, and on the current artifact whether revised or not, since Minors and Nits are collected and may leave nothing to revise. +13 | fixed | Discharge clause tied verbatim to the clean predicate — "no in-set Blocker or Major at effective severity" — so duty and predicate cannot drift apart. +14 | fixed | Second raise of a NIT, and the sentence was already being edited: squash carry takes one copy per cycle nonce plus five-field key, byte-identical repeats collapse, conflicting copies stop under the existing rule. +15 | fixed | The parity rationale was mechanically false — `### Mechanics (reference)` is at W:968. The row is re-marked "not deliberate" and W's cross-reference is **aligned** to C's wording as a third (b) edit. +16 | fixed | With the acceptance record shipping, only one thing stays unrecorded: a stuck or two-tell surface and its continue-or-stop answer. Disclosed in one sentence in the shipped block. +17 | fixed | One filled Q6 example added to the shipped text — the three lines over a gapped history plus the root line, six lines, per prompt-standards item 4. +18 | fixed by narrowing | §10 no longer claims every constraint carries an inline why. It names three settled axioms as deliberately unmotivated (D1, the existing zero-finding exit, D6) and says the plan's review reads every other sentence for one. + +**Severity correction carried forward:** pass 3's finding 12 was raised MINOR and treated as +MAJOR under the instrument carve-out (a false green in the evidence). Recorded in that pass's +dispositions; noted here because the assert it added is extended by finding 1's disposition. + +**Mechanical self-check, this revision.** Every OLD sentence the spec quotes counts 1/1 in both +copies (27 fragments, lines joined), with two deliberate exceptions: the (g) ownership sentence +is 1/0 because only C carries it, and the (b) cross-reference is 1/0 because W says "the +severity rule" — which is the divergence finding 15 aligns. New lead phrases count 0/0. Fences +balanced. No placeholders. Spec 719 lines, under the 720 ceiling. diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-4.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-4.md new file mode 100644 index 0000000..68cb872 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-4.md @@ -0,0 +1,19 @@ +BLOCKER | high | §3 "What a suspension asks" / §4 "What the body does not record" | Accepted obligations, answered structural questions, and standing stuck or two-tell suspensions remain session-only while the working record is optional; the fallback then drops any accepted obligation it cannot recover and merely assumes the reviewer will raise it again, despite the retained fix-set rules at C:200–202 and C:769–776 saying acceptance keeps the finding in-set and despite no mechanism guaranteeing a re-raise | After compaction, workspace loss, or replacement-cycle recovery, an unrepaired accepted Blocker or Major or an explicit stop can disappear and a fresh cycle can close the same artifact, so closure-controlling state is lost rather than conservatively suspended | Persist accepted obligations and standing suspension or question state in a mandatory cycle-attributed record, or prevent any replacement cycle from closing the same artifact until the old cycle is explicitly resolved; do not cite D10, which settles unavailable pass-history reporting rather than loss of closure state +MAJOR | high | §5(b) "What a loop absorbs" | The proposed membership-stop sentence defers resume until every simultaneous answer exists and every decline is recorded, but the proposed question-stop sentence still says the loop resumes once that question alone is answered | A finding carrying both membership and question triggers, or a question stop co-occurring with a stuck or two-tell stop, can resume prematurely under the local rule while the central composition rule says it remains suspended | Replace the question-stop resume clause with the same aggregate precondition used for membership stops: every answer required by the pass must resume, and any decline must already be recorded +MAJOR | high | §7 "Q6 — the pass-4 report without prior-pass history" | The text first says a rising count or require-withdraw pair spanning a missing pass cannot be seen, then says a gap does not break the series and consecutive available passes are compared across that missing pass | The same reduced history can yield different tell counts, and because two tells mandate suspension the closure ordering has no deterministic next state | Decide one gap policy and state it consistently; if passes 2 and 4 compare across missing pass 3, say that comparison is visible but less trustworthy, while comparisons needing the missing endpoint remain unavailable +MAJOR | high | §7 "Q6 — the pass-4 report without prior-pass history" | The claimed three-state partition is not observable from what the slot holds: a present invalid file does not reveal whether it was accepted when valid, and an absent formerly accepted pass is not classified at all; only the present-invalid accepted case is explicitly kept in the pass count | Readers can retroactively uncount a real pass, credit an unvalidated one, or disagree about whether the current floor is met and which curve positions receive question marks | Add an observable durable acceptance fact, or split known-accepted from acceptance-unknown for both absent and invalid slots and take the conservative count action for the unknown state +MAJOR | high | §7 "wrong-root state" / prompt standards 10–11 | The proposed prompt says its root check detects a historical pass written under a different working directory, but it only defines the correct current root and no recorded prior-call root or search mechanism can distinguish an elsewhere-written slot from an absent slot; it also gives no cause-specific fix | A wrong checkout can be reported as benign unavailable history and continue with reduced sensitivity, while the spec falsely claims a detecting mechanism and violates invariant 11's diagnostic-state standard | Define an observable comparison against recorded invocation-root data and its repair, or remove the detection claim and conservatively classify an unprovable root as a stop with the exact recovery action +MAJOR | high | §4 item 9 "one contract" | The ordering and decline blocks each contain a local partial-adoption stop, but the separately mergeable raw-versus-effective severity paragraph has no reciprocal check when it is adopted alone without those blocks | A downstream partial merge can install the new severity answer with a dangling "closure ordering above" reference and continue instead of taking the promised coherence stop | Add a local reciprocal one-contract instruction to the severity paragraph or make all three pieces an atomic adoption unit with a check that is present in every independently adoptable piece +MAJOR | medium | §4 item 9 "rollback" | The spec admits that no record identifies the rule revision an open cycle started under, yet claims the unknown-start fallback covers a rollback; a rollback can remove the new fallback extension and coherence text themselves, leaving only the older rules that cannot reconstruct the new ordering | A recovered open cycle can silently apply restored rules, start a replacement cycle that bypasses its state, or strand because "finish under the rules it started with" is not executable | Record or recover the starting rule revision from history and suspend explicitly when that revision cannot be loaded; do not claim a fallback removed by the rollback can cover its own absence +MAJOR | high | §8 "The slot discriminator — dissolved" | The claim that the no-nonce case no longer exists ignores determinable cycles that started before the parent rules, which current §5 expressly allows to finish under their starting rules and represent as `cycle none (pre-rule)`; rollback can also make old-rule no-nonce cycles reachable again | Concurrent legacy cycles can still collide on the bare slots, so dissolving the discriminator as legislation for an unreachable state leaves a known compatibility path unhandled | Limit the dissolution claim to post-rule and unknown-start cycles, then define safe serialization or explicit human handling for any observable known pre-rule cycle without shipping the rejected general discriminator +MAJOR | high | §5(b) old-conditions accounting | The accounting marks `b1`–`b16` kept even though `b12` is explicitly rewritten from immediate resume after the membership answer to aggregate resume after all answers and durable recording | The mechanically checkable accounting is false at the exact transition AC3 changes, violating story AC5 and the AGENTS.md decision-procedure rule | Mark `b12` replaced or extended, enumerate its old immediate-resume condition, and state which portion is kept and which new qualifications alter it +MAJOR | high | §2 / §5 "135-condition inventory" | The artifact gives only opaque ranges such as `a1`–`a22` and bulk dispositions, but never maps most ids to the conditions they denote or cites a durable inventory that does | The claimed complete old-condition accounting cannot be audited against current prose, so an omitted condition is indistinguishable from an undefined id despite AC5 requiring each requirement to be marked | Include the condition-to-id inventory in the spec or a cited committed appendix, then retain the disposition for every individually visible condition +MAJOR | medium | §5(g) "The demotion changes what a cycle must resolve" | The rationale calls demotion "the author's judgement about the fix set", but the retained severity procedure makes demotion a judgement about an in-system consumer and changed decision; assigned-fix-set membership is a separate scope predicate and a demoted finding may remain in-set | The shared technical term can make readers treat severity demotion as removal from scope, contradicting the closure block and changing whether membership stops or accepted obligations apply | Replace "judgement about the fix set" with "judgement about the finding's repair severity" or another term that cannot be read as assigned-set membership +MINOR | high | §3 "Third, a pass that neither closes nor suspends continues" | The example says a pass whose only findings are Minors and Nits lands in the continue branch without limiting it to below-floor passes, although the first branch closes that pass at or above the floor; it also requires running on "the revised artifact" even though those findings are collected and never iterated, so no revision may exist | A literal reader can continue after a valid clean completion or invent a prohibited Minor or Nit repair merely to produce a revised artifact | Say that a Minor-or-Nit-only pass continues only below the floor and reruns on the current artifact, whether revised or unchanged +MINOR | high | §3 "Blocker/Major-resolve duty" | The discharge clause ends with "followed by a pass that finds none" without saying whether "none" means no findings at all, none of the original findings, or no in-set Blocker or Major at effective severity | A pass carrying only declined out-of-set findings or Minors and Nits is clean under the first branch but can be rejected under this duty, reintroducing two closure definitions | Tie the discharge clause exactly to the clean predicate: a later validated pass finds no in-set Blocker or Major at effective severity +NIT | high | §4 "Recording a decline" / squash carry | Every decline is restated in every later cycle body while the squash rule copies every decline record in the range, but no identity or deduplication rule collapses exact repeats | A single decision can appear many times in the squash body, obscuring whether duplicates are harmless copies or distinct decisions even though conflicting copies already stop under the generic record rule | Carry one copy per cycle nonce plus the five-field finding key, collapse byte-identical repeats, and retain the existing stop for conflicts +MINOR | high | §6 "Parity" | The table preserves the C "Mechanics" versus W "the severity rule" divergence on the stated ground that W has no section named Mechanics, but the current scaffolded copy has `### Mechanics (reference)` at W:968 | The parity exception rests on a mechanically false claim and can preserve unnecessary drift between the two prompt copies | Correct the rationale after reading W:968, then either align the cross-reference or state the actual reason the wording must differ +MINOR | medium | §4 "What the body does not record" / observability | The spec explicitly records no accepted scope-stop result, question decision, stuck or two-tell surface, or continue-or-stop answer, so a WIP or later closing body exposes declines and final clean closure but not which suspension exits the cycle took | A reader cannot determine from commit history why a cycle remains open, whether a later continue was required, or whether simultaneous suspension reasons were surfaced, failing the requested observability risk test | Add a compact cycle-attributed suspension-event record, or explicitly accept this as an observability limitation and pair it with a rule that prevents unobserved open-cycle state from being bypassed +MINOR | high | §7 "Q6" / prompt standard 4 | The new unavailable-history report is described through many required statements but no concrete output example shows the three lines, wrong-root report, missing-pass list, question-mark values, and reduced-sensitivity disclosure together | Two implementations can satisfy the prose with materially different and incomplete reports, and the changed prompt fails invariant 11's requirement that output structure be shown rather than only described | Add one filled report example covering a gap and one compact wrong-root example in both prompt copies +MINOR | medium | §10 "Invariant 11" / prompt standard 6 | The spec claims every constraint in both shipped blocks carries its reason in the same sentence, but absolute rules such as "a pass with zero findings is clean whatever the floor", "nothing else closes a cycle", and "decline is available only at a membership stop" are asserted without an inline causal clause | The invariant-conformance claim is stronger than the proposed prompt text, leaving future editors unable to tell which rationale is load-bearing | Add concise inline reasons for the unmotivated constraints or narrow the §10 claim and document why those settled axioms are an intentional exception +END OF FINDINGS (18 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-5-dispositions.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-5-dispositions.md new file mode 100644 index 0000000..0e1382a --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-5-dispositions.md @@ -0,0 +1,17 @@ +1 | fixed | Five-field key and sameness test read the reviewer-written severity; effective severity limited to cleanliness, resolution and severity-driven action. Stated once, in the ordering's field paragraph. +2 | fixed | Named the two senses of "clean" §5 already carried — clean findings file (the NO FINDINGS signal) versus clean pass (the ordering's predicate) — rather than inventing a third term. +3 | fixed | The fourth duty is restored over its original domain: no pass that surfaced a FINDING is credited as clean, which the scope stop and the clearly-stuck exit both trigger; the two-tell stop surfaces tells, not a finding, and never was in that domain. c18 marked kept, not replaced. +4 | fixed | Both labels are exclusive to membership stops. The "an acceptance also records a question stop's decision" overload is deleted, so the block and §10 now say one thing. +5 | fixed | A question-stop decision that changes no membership leaves no record, with the reason stated (the five fields identify a finding, not a question) and the residual disclosed rather than a field invented. +6 | fixed | The separate "obligation" concept is deleted. An accepted finding is in the set and the shipped severity rules govern it — Blocker/Major owes resolution, Minor/Nit is collected and never iterated. No obligation a rule cannot discharge, no exception to never-iterate. +7 | fixed | A changed-field re-raise is a new finding classified afresh against the current fix set and question predicate; it never inherits a stop from the finding it resembles. +8 | fixed | The new root stop is deleted. Root establishment is a pre-pass validity condition and a pass failing it is an INCOMPLETE pass — the state §5 already defines, already uncounted and already excluded from the curve — so D10's "not a new stop condition" stands and nothing new is ranked in the ordering. +9 | fixed | "Known accepted" now means this session validated the pass. After a lost session every absent or invalid slot is acceptance-unknown: uncounted toward the floor, series `?`, disclosed. No pass-acceptance record invented. +10 | fixed | Two observable root-failure causes named with their own fixes (not inside a git repository; resolved top level differs from the directory the call was given). A past wrong-root write is stated as unobservable, not detected. +11 | fixed | The shipped recovery sentence is edited (§4 item 11): a Gate-A cycle mid-run also has the commits it made on the branch, scoped by kind, artifact and nonce. Without it the record would exist and the recovery procedure would never look at it. +12 | fixed | The contract gets one name and one membership list in the one-contract paragraph; each mergeable piece cites the name instead of enumerating its peers. One list to keep in step. +13 | fixed | The rollback stop claim is dropped rather than defended, and its row is removed from the §9 walk — a row that cannot cite shipped text is not evidence. What ships is the true statement: no record identifies the rule revision a cycle started under. +14 | fixed | §8 stops claiming a human serializing pre-rule cycles protects anything. It states the exposure: two live pre-rule cycles compute the same bare slots and can delete each other's findings files, and nothing shipped here prevents it. +15 | fixed | a13 marked REPLACED with its old condition enumerated. The categorical "none of them is restated" becomes scoped to its own paragraph, since the ordering restates closure rules deliberately and as the one authority. +16 | collected | Simultaneous clearly-stuck and two-tell suspensions ask the identical continue-or-stop question and the composition rule requires each answered on its own; one labelled answer covering both may be the better rule. For the plan. +17 | fixed | The label was beyond the story's ORIGINAL scope and is authorised by the current story's §2 and acceptance criterion 7, at 4625679. Corrected wherever the spec called it unauthorised growth. diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-5.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-5.md new file mode 100644 index 0000000..9d8dfce --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-5.md @@ -0,0 +1,18 @@ +MAJOR | high | §3 "Field selection" / §4 "Binding, and the sameness test" | The ordering says every predicate reads effective severity except the health measures, while the five-field decline matcher is a non-health predicate whose recorded severity comes verbatim from the reviewer-written findings line | A finding re-raised with a different raw severity but the same ceiling-demoted effective severity can match the old decline and suppress the new membership hold, contradicting settled D9b that any severity difference makes a new finding | State that the five-field key and sameness test read reviewer-written raw severity, and limit effective severity to cleanliness, resolution, and severity-governed action +MAJOR | high | §3 "First, clean completion" / standing Gate-A findings protocol | The new predicate calls a pass clean when it has no in-set Blocker or Major at effective severity, but the unchanged protocol at C:323–329 and W:517–523 says a clean pass has the single body line `NO FINDINGS`, and C:565–566 / W:756–757 says that explicit signal is what permits loop exit | A pass containing only Minors, Nits, or a previously declined out-of-set finding is simultaneously clean under the ordering and non-clean under the required file signal, so readers can either reject a valid close or close without the signal the standing rule requires | Introduce a distinct term and explicit file test for closure-clean passes, then update every standing clean-signal sentence and account for those conditions, or retain the zero-finding definition consistently +MAJOR | medium | §3 "The four standing duties" / §5(c) old-condition accounting | The fourth duty is narrowed to passes that surfaced a scope stop even though the story defines it as no pass carrying a surfaced finding counts as clean and current c18 applies to the clearly-stuck surface; calling c18 replaced does not supply authority to change that settled duty | Under the new raw-health/effective-cleanliness split, a below-floor clearly-stuck pass can surface reviewer-written Blocker or Major findings that demote below Major yet still receive clean credit, so one of the four required duties is no longer classified over its original domain | Preserve no-clean-credit for every pass that surfaces a finding, including clearly stuck, while allowing clean completion to win before any surface occurs +MAJOR | high | §4 "Recording an answer at a membership stop" / §10 invariant claim | The block says both `Accepted:` and `Declined:` are available at a membership stop and nowhere else, then says an acceptance also records a question-stop decision; §10 repeats the absolute claim that either answer is available only at a membership stop | An in-set question-only finding has contradictory instructions about whether an `Accepted:` record is permitted, while a dual-trigger finding leaves readers unsure whether one label answers one trigger or both | Keep both labels exclusive to membership stops and record question decisions separately, or state one precise dual-trigger exception consistently in the rule and §10 +MAJOR | high | §4 "Recording an answer at a membership stop" | Even where a question decision is meant to be carried by `Accepted:`, the form stores only handle, date, nonce, and the reviewer's five finding fields; it stores neither the question nor the user's chosen answer | After session loss, the record can prove membership acceptance but cannot reconstruct the structural or contract decision needed to discharge the question hold, so a reader may either release an unanswered hold or ask indefinitely | Add a cycle-attributed question-and-answer field or distinct question-decision record, and define its sameness and carry rules +MAJOR | high | §3 "What a suspension asks" / §4 "What each is worth" | The text repeatedly says every acceptance creates an obligation and records work the cycle owes, while the same ordering says an accepted Minor or Nit is collected and never iterated and the unchanged severity rule forbids a repair round for it | A reader can repair an accepted Minor or Nit contrary to the out-of-scope severity semantics, or carry a purported obligation that no rule can ever discharge | Define acceptance as a durable membership fact; state that it creates a repair obligation only when effective severity is Blocker or Major and that Minor or Nit acceptance is collection-only +MAJOR | high | §3 "Second" / §4 "Binding, and the sameness test" | A changed field is declared to create a new finding and a new stop automatically, without re-evaluating that new finding against the current assigned fix set and question predicate; elsewhere a formerly declined finding that the set later includes owes ordinary resolution instead of a membership stop | After scope broadens, a changed-field re-raise can acquire a membership hold even though it is now in-set, but decline is forbidden and the record block also forbids membership labels for an in-set Blocker or Major, leaving incompatible transition rules | Treat a changed-field re-raise as a new finding, then classify it afresh: in-set findings follow severity, out-of-set findings raise a membership stop, and new structural or contract questions raise a question stop +MAJOR | high | §7 "Q6 — the pass-4 report without prior-pass history" | The block creates a current-root failure that stops the cycle, yet D10 and the same block say unavailable history is not a new stop condition, and the closure ordering neither classifies this root stop nor ranks it against zero-finding completion | On pass 4 or later a zero-finding pass with an unestablished root can both close first under §3 and stop under §7, so the single ordering no longer covers every reachable terminal conflict | Make root establishment an explicit pre-pass validity precondition with a defined incomplete-pass result, or remove the stop and use D10's disclosure path; in either case add it to the ordering and conflict walk +MAJOR | high | §7 "known accepted" | No named cycle record records that an individual pass was validated and accepted: findings files record inventory, the optional working record has no acceptance schema, and the durable curve is author-written, unchecked, and normally written only when the cycle closes | The known-accepted-but-missing branch is not mechanically observable after session loss, while treating an unchecked note or curve as acceptance can fabricate floor credit for a pass whose findings cannot be validated | Name and define a validated pass-acceptance record, or restrict known acceptance to this session and classify every absent or invalid historical slot after loss as acceptance unknown and not counted +MAJOR | high | §7 "First the root" / prompt-standards item 10 | The diagnostic state "current root cannot be established" has no specified check, no enumeration of distinct causes, and only the blanket instruction to re-run from the correct root; the text also admits a prior wrong-root write is indistinguishable from absence | A caller already using the correct working directory can repeat the same failure without learning whether the cause is a non-repository checkout, an invocation mismatch, Git failure, or another condition, violating invariant 11's cause-specific diagnostic requirement | Define the exact root-establishment check, enumerate its observable failure causes, and pair each with its own recovery; keep historical wrong-root placement classified as unobservable rather than detected +MAJOR | high | §4 "Rules both labels share" / unchanged recovery source at C:411–417 and W:605–611 | The new transport creates spec or plan revision commits, including empty commits, during an open Gate-A cycle and says a lost session can recover answers from branch bodies, while the unchanged recovery prose says a Gate-A cycle mid-run has no commit of its own and that history recovery performs no search | A Gate-A answer record can exist but the authoritative recovery procedure will never locate or adopt it, so the accepted obligation remains only theoretically recoverable and story criterion 7 is unmet | Extend the recovery procedure in both copies to identify and validate open Gate-A answer-record commits by kind, artifact, and nonce, or move the record into a mandatory cycle-stable artifact the existing recovery path reads +MAJOR | high | §4 item 9 "one contract" / compatibility | The local reciprocal guards name only the ordering, answer-record block, and severity split, although independently mergeable nonce-set, unknown-start, closing-message, and squash-carry edits are also required for an answer record to remain attributable and durable | A downstream partial adoption can contain all three guarded core pieces while omitting one transport extension, pass every local three-piece check, and then lose an acceptance on recovery or squash | Make the entire closure-record contract one atomic replaceable block, or have every independently adoptable piece name and validate all required components and versions +MAJOR | high | §4 item 9 "rollback" / §9 risk-path verification | The rule that a cycle stops when rollback removes the fallback text appears only in design explanation, not in any quoted shipped replacement, yet §9 requires the next-state table to cite a shipped line for that row | Once the rollback has removed the fallback and coherence text, an open new-rule cycle has no executable instruction selecting the claimed stop, so implementation can satisfy the specified edits while the rollback row is impossible to evidence | Add the rollback recovery/stop rule to both shipped copies in text that survives or can be recovered across rollback, or record the starting rule revision and load it; otherwise remove the unsupported claim +MAJOR | high | §8 "The slot discriminator — dissolved" | The only protection for two observably live pre-rule cycles is "serializing them is a human's job", but §8 explicitly ships no prompt text and cycles finishing under old rules never read this design document | Concurrent legacy or rollback-created cycles can still compute and delete the same bare findings slots, recreating the documented data-loss incident with no user-visible instruction to serialize them | Put an adoption-time compatibility check and serialization instruction in a surface the human actually reads before either legacy cycle runs, or retain a bounded discriminator for this reachable state +MAJOR | medium | §5(a) `a13` accounting | The spec marks a13 kept, leaving C:128–130 and W:335–337 to say every other closure rule stands as written and none is restated, while the new central block restates and changes clean completion, suspensions, holds, duty discharge, and continuation | The shipped prompt will tell readers both that closure rules are not restated and to use a new restatement as the single authority, contrary to the old-condition accounting rule and prompt-standards consistency requirement | Replace a13 with a scoped statement that the floor paragraph does not summarize the ordering and points to the one authoritative block, and mark the old categorical condition replaced +MINOR | medium | §3 "Composition, and what cannot happen" | Clearly-stuck and two-tell suspensions each ask the identical `continue or stop` question, but the composition rule says every question is answered on its own and never says whether one explicit answer can cover both simultaneous health suspensions | A user who says "continue" once can reasonably believe the cycle resumed while a literal agent keeps one suspension standing, producing avoidable repeated escalation | Either collapse simultaneous health suspensions into one labeled continue-or-stop decision carrying every reason, or require and exemplify separate labeled answers +MINOR | high | §4 scope-expansion provenance | The spec says `Accepted:` is scope growth "beyond the story" and that the story shipped only a decline record, but the current story at 4625679 records the expansion in §2 and adds criterion 7 requiring the accepted-work record | Reviewers are directed to treat an input that now authorizes the change as if it were stale, obscuring the actual settled authority and making future scope audits disagree | Say the label was beyond the story's original scope and is now authorized by current story §2 and acceptance criterion 7 +END OF FINDINGS (17 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-6-dispositions.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-6-dispositions.md new file mode 100644 index 0000000..1780598 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-6-dispositions.md @@ -0,0 +1,16 @@ +1 | fixed | Fixed at the source, not by inventing a term: C:329/W:523 now says "clean findings file" (it describes a file and always did) and C:565/W:756 says "when a pass finds nothing". §4 items 12-13. The other five "clean pass" uses in each copy were checked and are the closure sense the ordering defines. +2 | fixed | Severity bullet C:783-784/W:969-970 gains "for every finding in the assigned fix set", the boundary D5 always implied and no sentence carried. §4 item 14, accounted as replaced. c17's pointer needs no edit: it names the rule, and the rule now carries its own scope. +3 | fixed | Membership is read when the answer is given, not frozen at the surface: a broadening that puts the held finding in-set discharges the hold by itself, decline becomes unavailable, effective severity governs. One sentence in the block plus a §9 walk row. +4 | fixed | Conformed rather than claimed an exception. All three axioms gained an inline causal clause in the shipped block; §10's exemption paragraph is deleted. The exemption was wrong twice: item 6 admits no "settled elsewhere" clause, and a scaffolded copy cannot reach the story. +5 | fixed | The structural repair. Clean candidacy now reads the triggers (a finding outside the set, one opening a question), never the act of surfacing, so no predicate waits on a lower branch. The outside-set example is narrowed to findings a decline of this cycle matches. +6 | fixed | The D3 sentence stays byte-identical; an adjacent sentence names the condition its example assumed — a clean-completion candidate only where no scope-stop trigger applies. c9/c10/c11 accounted as kept verbatim and qualified by adjacency. +7 | fixed | The branch aggregate is the concatenation: both branches' lines count for the health measures (the curve rule already sums them), and a finding matching one in the other branch on all five fields is one finding for holds and answers. +8 | fixed | After a question decision an in-set finding routes through its effective severity like any other: Blocker/Major resolves or is dismissed, Minor/Nit is collected and never iterated. No Minor gets a repair round. +9 | fixed | The whole recovery passage is replaced, not one sentence: three sources, the branch search bounded newest-first to the branch point, candidates validated by cycle field, kind and artifact. The following sentence's "from either" becomes "from any of the three"; the standing failure rule is unchanged and not restated. +10 | fixed | The record header becomes ` · · cycle · · `, so a record read out of a branch body — including an empty commit carrying only a record — is attributable on its own. +11 | fixed | Took the named-list option over a marker on every hunk: the one-contract paragraph names all eleven components and says a project carrying any of them owes all of them. One list to keep in step rather than eleven cross-references; the three reader-facing markers stay, but the obligation is the list's. +12 | fixed | An acceptance-unknown pass is omitted from the durable curve rather than given a `?` slot: the grammar takes one entry per valid pass and has no way to say "may not have been one". `?` stays a missing count. The omission is disclosed in the report. +13 | fixed | The cycle's latest body is authoritative and must carry the complete answer set; an answer missing from it is lost whatever an earlier body says, because a rule letting any reachable body revive an answer would make the set depend on how far back a reader looked. The ordering's fix-set sentence is aligned to it. +14 | fixed | Third raise, and the fix is smaller than the rule it replaces: simultaneous health suspensions are one question, so one continue-or-stop answer carrying every reason ends both. The per-question rule still governs holds and question stops. +15 | fixed | "the artifact revision it produces, where it produces one; a decision that changes nothing stays in the session that made it". +16 | fixed | The `—,` sequence is gone; the dedup clause is parenthesised instead. diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-6.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-6.md new file mode 100644 index 0000000..3499390 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-6.md @@ -0,0 +1,17 @@ +BLOCKER | high | §3 "First, clean completion" / standing findings protocol | The proposed distinction makes a pass carrying Minors, Nits, or a matched decline clean, but the unchanged protocol says "A clean pass is" the `NO FINDINGS` body at C:329 and W:523 and says that signal permits loop exit at C:565–566 and W:756–757; the spec neither edits nor accounts for those standing uses | Such a nonempty pass has no permitted clean signal and, when no suspension trigger remains, can neither close nor suspend | Give the closure predicate a distinct term and update every standing clean-signal use and inventory entry, or retain `NO FINDINGS` as the only clean pass +BLOCKER | high | §3 "Blocker/Major-resolve duty" / Mechanics Severity | The ordering limits resolution to in-set Blockers and Majors, but the unchanged Severity sentence at C:783–784 and W:969–970 still says Blockers and Majors "both must resolve" without an in-set boundary | After a matching decline suppresses the membership stop, the finding must stay outside the fix set yet the standing duty still bars closure; with no other suspension it can neither close nor suspend | Qualify the standing Severity duty and its c17 pointer in both copies with the assigned-fix-set boundary, and account for those current conditions explicitly +BLOCKER | high | §3 "What a suspension asks" / §4 "Recording an answer" | A membership hold is classified from the pass that raised it, but no transition handles the governing scope broadening before the answer so that the held finding is now in-set; the hold still requires accept or decline while the record rule forbids those labels for an in-set Blocker or Major and decline cannot override the governing artifact | The hold remains standing with no permitted answer, leaving the cycle unable to close or resume | Define whether pending membership is frozen at surface time or re-evaluated, and give scope broadening an explicit cancel, conversion, or recorded-accept transition exercised by the next-state table +BLOCKER | high | §10 "Invariant 11" | The spec claims all 12 prompt standards pass while expressly exempting three shipped constraints from item 6's requirement that every rule carry its why in the same clause; settled decisions in the repository story are not an exception in the standard and are unavailable to a downstream scaffolded copy | The design knowingly violates non-negotiable invariant 11, so the prompt cannot pass its required review gate as specified | Add concise inline reasons for all three axioms in both shipped blocks, or change the standard through separately authorized work before claiming conformance +MAJOR | high | §3 "First" and "Second" | Clean completion is evaluated before suspensions, yet its predicate depends on whether the pass "surfaced" a finding, an action produced only by the lower-priority scope or stuck branch; the preceding explanation also says an outside-set finding is clean even though a fresh outside-set finding necessarily triggers and surfaces a membership stop | A literal first-branch evaluation can close before the scope/question stop is evaluated, while the opposite reading requires inferring a look-ahead the promised executable ordering never states | Define clean candidacy using pre-surface predicates, explicitly excluding scope-stop triggers but not the lower-priority health exits, and limit the outside-set example to findings matched by a current-cycle decline +MAJOR | high | §5(c) `c9`–`c10` accounting | The spec preserves verbatim the sentence at C:235–237 and W:438–440 that any Blocker/Major-free pass at or above the floor has satisfied the clean-final-pass rule and closes | An at-floor pass carrying an in-set Minor that opens a new contract question or a new out-of-set Minor must suspend under the new ordering but must close under this retained sentence | Preserve the required precedence wording while adding an adjacent qualification that its Blocker/Major-free example is a clean-completion candidate only when no scope-stop trigger applies, and account for that qualification +MAJOR | high | §3 "Field selection" / Mechanics full Gate-B pass | Reading two branch files "as one set" does not say whether duplicate five-field findings are deduplicated, while the standing curve rule at C:952–965 and W:1136–1149 requires one summed logical-pass entry | When both reviewers raise the same finding, readers can count it once or twice, derive different clusters and tells, and disagree whether one or two answer records and suspensions exist | Define the branch aggregate as a concatenated multiset or a semantic union, state how conflicting severities and duplicate keys behave, and align the health counts and answer-record cardinality with the existing summed-entry rule +MAJOR | high | §3 "What a suspension asks" | The question-stop branch says an in-set finding resolves under the user's decision or is dismissed, without applying the effective-severity split | An in-set Minor or Nit that opens a structural question is assigned a repair/dismissal round even though the unchanged severity semantics require it to be collected and never iterated, changing an expressly out-of-scope rule | After the question answer, route the finding through ordinary effective severity: Blocker/Major resolves or is dismissed, while Minor/Nit is collected without iteration +MAJOR | high | §4 item 11 "Recovery sources" | The proposed replacement makes open Gate-A branch bodies a recovery source, but it leaves standing C:413–416 and W:607–610, which say history recovery performs no search and reads only the cycle's own commit body | A lost-session replacement cannot perform the branch-body lookup that criterion 7 and the answer-record residual rely on, so an accepted finding is nominally durable but not recoverable by the authoritative procedure | Replace the whole conflicting recovery passage, defining the bounded branch-history search, candidate validation, conflicts, and no-match result in both copies +MAJOR | high | §4 "Rules both labels share" / item 11 | The recovery edit says answer-record commits are scoped by kind, artifact, and nonce, but the record form carries only handle, date, nonce, and the five finding fields; an empty answer-only commit supplies neither cycle kind nor reviewed artifact, and no start boundary identifies "commits since it started" after session loss | Multiple open Gate-A cycles on one branch cannot reliably attribute or reject an empty record commit, enabling replay/misattribution or making the accepted work unrecoverable | Add kind and artifact identity plus a recoverable start/range key to the record or move answers to a mandatory cycle record whose existing identity fields make the proposed scoping decidable +MAJOR | high | §4 item 9 / §6 "Parity" | The spec says every independently mergeable contract piece names the closure-record contract, but it names only the ordering, answer-record block, and severity replacement; the nonce-set, unknown-start, closing-message, squash-carry, no-identity-report, and recovery edits are independently mergeable and carry no local contract marker | A downstream partial adoption can take one transport hunk without the updated one-contract paragraph, leaving an undefined or lossy answer-record rule without triggering the promised compatibility stop | Make the contract atomic for `/workflow-init`, or put a reciprocal presence/version check on every independently mergeable hunk and verify each asymmetric omission +MAJOR | high | §7 "Q6" / standing curve grammar | Acceptance-unknown prior passes are said to be uncounted but to occupy a `?` series position because the curve grammar admits it, while the unchanged grammar at C:943–950 and W:1127–1134 permits one entry only per valid pass and excludes incomplete passes | After session loss the pass is not known valid, so including it violates the curve grammar and omitting it violates Q6; floor accounting and tell comparisons can diverge | Define a distinct unknown-validity representation and update the curve grammar and counting rule, or omit unknown passes from the durable curve and state that `?` applies only to known-valid passes with unrecoverable counts +MAJOR | medium | §3 "Continue" / §4 "Rules both labels share" | The answer block says every later body must restate each answer because a body that drops it loses the fact, but the ordering keeps accepted work whenever the cycle's commit bodies carry it and excludes it only when no body carries it | If an earlier reachable body has an acceptance and the latest body omits it, one rule keeps the finding in-set while the other treats the omission as loss, with no defined stop or recovery action | Name the authoritative body set and omission transition, such as requiring the latest reachable cycle body to contain the complete answer set and stopping on any omission, then align recovery, close, and squash rules +MINOR | high | §3 "Composition" | Clearly-stuck and two-tell suspensions each ask the identical `continue or stop` question, while composition requires every question to be answered on its own and gives no labels or example distinguishing the two answers | One unqualified "continue" can be read as resuming both suspensions or only one, causing avoidable repeated surfaces and inconsistent state | Collapse simultaneous health exits into one continue-or-stop decision carrying every reason, or require explicit labeled answers for each health suspension +NIT | high | §4 "What each is worth" | The text says a question decision's durable form is the artifact revision it produces even though a valid no-change decision or dismissal can produce no revision, then later admits the decision is unrecoverable after session loss | The blanket durability sentence overstates observability and obscures the already admitted residual | Say "an artifact revision, when the decision produces one" and state that otherwise the decision remains session-only +NIT | high | §4 item 1 "Squash carry" | The exact NEW sentence places a comma immediately after the closing em dash of the inserted answer-record clause, producing the literal sequence `—, the provenance lines` | The malformed punctuation makes the six-member list harder to parse in the prompt's most load-bearing carry rule | Remove either the comma or the closing em dash and keep the six members grammatically unambiguous +END OF FINDINGS (16 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-7-dispositions.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-7-dispositions.md new file mode 100644 index 0000000..e8cff07 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-7-dispositions.md @@ -0,0 +1,32 @@ +1 | fixed | "triggers nothing" → "no membership trigger and no membership hold"; the question-stop test stays independent, in both the ordering and the record block. +2 | fixed | A broadening discharges only the hold's membership component; the question component stands until its decision is given. +3 | fixed | The overlap is admitted rather than argued away: the stuck reading is raw, cleanliness effective, so a demoted in-set Blocker satisfies both. The evaluation order decides it — close at or above the floor, continue below — and the c9 sentence stays untouched with the qualification adjacent. +4 | fixed | Every decline-based exclusion now reads "while the current assigned fix set still excludes it"; a new sentence says how that coexists with D7 — the decline binds the decision, the governing artifacts own membership. +5 | fixed | The latest body is the authoritative snapshot and the only thing anything reads; squash carry copies that snapshot, not every record in the range; an empty answer set is written explicitly as `Cycle answers: none`. +6 | fixed | A recovery candidate is a cycle identity tuple (cycle field, kind, artifact), not a body, so one cycle's restatements are one candidate; its newest body is the answer set and the search stops there. +7 | fixed | Both halves ship, reversing this cycle's pass-6 choice. Pass 6 offered a central list or per-hunk markers; the list alone was taken, and pass 7 showed a list cannot police a merge that omits the list. Thirteen source edits now carry a short marker, the three blocks keep their longer sentence, and the gate-off edit joins the membership list. +8 | fixed | New item 15 edits the standing curve paragraph: durable proof that a pass was valid is separated from knowing it ran; a pass whose validity cannot be established is omitted from the specification rather than entered with `?`. +9 | fixed | The filled example now shows pass 3 omitted from the curve, not given `?` — it was teaching the representation the rule forbids. +10 | fixed | A third root-failure branch for a git query that fails for any other reason, reported with its error; both paths canonicalized before comparison. +11 | fixed | Six, not five — counted in both copies (C:117, 728, 750, 761, 769, 827; W:324, 914, 936, 947, 955, 1011) and the W sites added. +12 | fixed | "both severity-bearing conditions of the three-condition stuck reading", with the coverage-sufficiency condition named as reading no severity field. +13 | fixed | "above" → "below" in the (g) replacement: the one-contract paragraph is at C:879-890 and (g) at C:810-815. Directions audited across all sixteen markers; two others were backwards and were corrected the same way. +14 | fixed | The no-identity report marks the unadopted cycles' exit and answers unknown, and the replacement inherits nothing and obtains its own. No suspension-state record added. + +Extra check (directive): read both shipped blocks end to end for sentence pairs a fresh reader +could take as two rules. Seven found and fixed, none of them raised by Codex: +- no-clean-credit read as a second test beside the clean predicate; now stated as that + predicate's second half, with the other two exits explicitly outside the duty. +- the Second branch half-restating the decline and changed-field rules it says it does not + restate; now points at the two places that state them. +- "Continue resumes the loop" read as resuming unilaterally against the composition rule; now + "is that suspension's resuming answer, and where it is the last one outstanding". +- "Two states cannot co-occur" → "Two pairings cannot occur" (it enumerates pairs). +- "Declined keeps it out and releases its hold" read as unconditional; now the membership half. +- "treats it as absent" scoped to the authoritative snapshot rather than to any body. +- the resolve duty "which the declined finding never reached" against a later broadening making + it owe resolution; now time-bound to while it stays outside the set. + +Mechanical self-check: 31 quoted OLD fragments, all 1/1 in both copies; one deliberate +asymmetry, the (g) ownership sentence at C=1 W=0; 13 new lead phrases all 0/0; fences balanced +(3 three-tick pairs, 1 four-tick pair, 2 indented inner pairs); no placeholders. diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-7.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-7.md new file mode 100644 index 0000000..4d518d7 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-7.md @@ -0,0 +1,15 @@ +MAJOR | high | §3 "First, clean completion" / §4 "Binding, and the sameness test" | The clean branch says a finding matched by a current-cycle decline "triggers nothing", but the answer-record rule says that match suppresses only the membership stop and a question stop still fires unless that question was answered | A declined finding that reopens an unanswered structural or contract question can be treated as clean and close the cycle over the open question | Change "triggers nothing" to "raises no membership trigger or hold" and preserve the independent question-stop test +MAJOR | high | §3 "What a suspension asks, and what ends it" | A dual-trigger finding requires both a membership answer and a question decision, yet scope broadening is said to discharge "the hold by itself" as soon as it makes that finding in-set | The loop can resume or close while the structural or contract question from the same finding is still unanswered | State that broadening discharges only the membership component; keep the question component of the hold until its decision is given +MAJOR | high | §3 "No-clean-credit" / "Composition, and what cannot happen" | The non-co-occurrence proof assumes every regenerating reviewer-written Blocker or Major fails the clean predicate, but the severity split makes the stuck reading use raw severity while cleanliness uses effective severity; an in-set raw Blocker demoted to Minor can therefore satisfy the stuck predicate and the clean predicate together | Readers can disagree whether D3 makes clean completion win, whether clearly stuck surfaces, and whether the pass receives clean credit | Admit this reachable overlap, say clean completion wins by D3, and clarify adjacent to the preserved c9 sentence that its clean-completion severity test is effective while stuck detection remains raw +MAJOR | high | §3 "First, clean completion" / §4 item 14 "Severity" | The clean text and proposed Severity sentence say a matching declined finding is outside the set and owes nothing, while the answer-record rule says a later governing-scope broadening can put that same declined finding in-set and make it owe ordinary resolution | A reader can use the stale decline to ignore an in-set Blocker or Major after a scope change, or can re-raise a membership stop that the current set no longer warrants | Qualify every decline-based exclusion with "while the current assigned fix set still excludes it" and state how that qualification coexists with D7's cycle binding +MAJOR | high | §4 "Rules both labels share" / item 1 "Squash carry" / item 11 "Recovery sources" | The latest cycle body is declared authoritative and an answer omitted from it is declared lost, but squash carry copies every answer record anywhere in the squash range and branch recovery searches older bodies; moreover a later Gate-A body that omits all records carries no nonce marker proving it belongs to that cycle | An earlier decline can be replayed to release a hold after the authoritative set dropped it, and an earlier acceptance can be revived or lost depending on which reader scans the history | Put an explicit complete answer-set snapshot or cycle marker in every cycle body, define how an empty set is represented, and make recovery and squash carry read only the latest validated snapshot +MAJOR | medium | §4 item 11 "Recovery sources" | Answer records are intentionally restated in every later body, but the new branch-history source does not say whether those repeated bodies collapse to one candidate before the unchanged "more than one candidate → no identity" rule runs | A normal answered cycle with two revision commits can be treated as ambiguous and forced into a new cycle, defeating the recovery path criterion 7 needs | Define a candidate as a distinct cycle identity tuple, select its newest body as authoritative, collapse byte-identical older restatements, and stop only on distinct or disagreeing candidates +MAJOR | high | §4 item 9 "The one-contract paragraph" | A central membership list cannot protect a partial downstream merge that omits that very paragraph, and only three of the separately mergeable pieces carry local contract markers; the list also omits the gate-off edit even though §4 says absence from any listed threat site is a defect | A downstream project can adopt a clean-vocabulary, carry, recovery, nonce, or gate-off hunk alone and continue running without ever seeing the promised partial-adoption stop | Make the update atomic in workflow-init or add a compact reciprocal contract/version marker to every independently mergeable hunk, and include every required source and threat edit in the contract membership +MAJOR | high | §7 "Q6 — the pass-4 report without prior-pass history" | The new rule says no pass is known accepted after a lost session and therefore absent or invalid prior passes are omitted, while the unchanged curve rule at C:943–950 and W:1127–1134 uses a resumed cycle that knows a pass happened but not its counts as the rationale for retaining that valid pass with question-mark counts | A resumed reader has two incompatible instructions for floor credit and the durable curve, so the same missing slot can either count or disappear | Edit and account for the standing curve paragraph so it distinguishes durable proof that a pass was valid from mere knowledge that it ran, and align its resumed-cycle example with Q6 +MAJOR | high | §7 "Q6" example | The normative text says an acceptance-unknown pass is omitted from the durable curve and that a question mark is only a missing count for a known-valid pass, but the filled example labels pass 3 "acceptance unknown" and still says "series `?`" | The required output example teaches the exact invalid representation the preceding rule prohibits, so agents can write malformed or misleading curves | Remove the question-mark claim for pass 3 and show it omitted from the curve's pass specification, or change the example to a separately identified known-accepted pass if a question mark is intended +MINOR | medium | §7 "First the root" | Root establishment is claimed to have exactly two observable failures, but a Git invocation can also fail because Git is unavailable, repository ownership is rejected, metadata is unreadable, or another command error occurs; the text also does not say that the supplied directory is canonicalized before comparison | Distinct causes can be mislabeled as "not inside a git repository" or a semantically identical symlink or trailing-slash path can make a valid pass INCOMPLETE with the wrong fix | Add a command-failure diagnostic branch and compare canonical paths, with a cause-specific report and fix for failures that are not Git's not-a-repository result +MINOR | high | §4 item 13 "The Gate-A clean-signal sentence" | The spec says there are "other five uses" of clean pass but immediately enumerates six current C sites: 117, 728, 750, 761, 769, and 827; W likewise has six counterparts | The required mechanical count is false and weakens the claim that every standing use was checked | Change the count to six and retain all six citations +MINOR | high | §3 "How a cycle ends" | The field-selection paragraph says it covers "both conditions of the stuck reading" and enumerates the Blocker curve and regenerating findings, but the preserved stuck procedure has three conjuncts, the third being the stated coverage-sufficiency judgement | The stated count is mechanically false and can be read as silently dropping the coverage condition from clearly-stuck evaluation | Say "both severity-bearing conditions of the three-condition stuck reading" and explicitly state that coverage sufficiency reads no severity field +NIT | high | §5(g) "Mechanics · Severity" | The replacement says the closure-record contract is named by "the one-contract rule above", but the one-contract paragraph remains later in Mechanics at C:879–890 and W:1063–1074 | The directional cross-reference sends readers away from the contract definition in both shipped copies | Change "above" to "below" in the byte-identical replacement +MINOR | medium | §4 "What each is worth" / observability risk | The design deliberately records neither question-stop decisions that produce no revision nor stuck or two-tell continue-or-stop answers, so a branch body can show answer records without showing which simultaneous suspension remained standing or why | After session loss, a reader cannot reconstruct which exit the cycle took or what input would resume it, making the human-owned open cycle difficult to distinguish from an abandoned or already-answered one | Add a durable suspension-state record or explicitly require the first replacement-cycle report to mark the prior exit and answer state as unknown and obtain fresh answers before relying on it +END OF FINDINGS (14 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-8-dispositions.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-8-dispositions.md new file mode 100644 index 0000000..5096fe8 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-8-dispositions.md @@ -0,0 +1,12 @@ +1 | fixed | Identity-preserving recovery that cannot recover a health-stop answer or a no-change question decision re-surfaces those questions and takes fresh ones before another pass runs; the finding's own second option, stated once in §4 beside the no-identity case rather than in two places. +2 | fixed | Duties reconciled instead of narrowed: the hold belongs to the scope stop, the only stop whose question is about a finding; no-clean-credit attaches to any pass carrying a scope-stop trigger; a clearly-stuck pass whose regenerating in-set findings the ceiling demotes below Major is clean at effective severity and clean completion wins by D3. c16 and c18 both marked REPLACED with their old conditions enumerated and authority named (D1, D3). +3 | fixed | The membership trigger and the question trigger are defined once in §3's first branch and nowhere else; the suspension list now says "raised by either trigger defined above" and Mechanics references the definition instead of restating it. +4 | fixed | Every cycle body carries the identity line and the complete snapshot, `Cycle answers: none` included; the newest body is found by that line independently of whether it carries an answer; a missing or malformed snapshot stops rather than being skipped. +5 | fixed | At most one effective label per cycle and five-field key: exactly equivalent same-label records collapse idempotently, conflicting labels or disagreeing attribution stop. Placed with the sameness test where the key is defined. +6 | fixed | §5(a), (b), (c) and (e) pointer edits and the §7 Q6 block added to the contract membership list, each carrying the same marker the other hunks carry. Markers now 17 "part of" + 4 block-level = 21 per copy; the marker assert re-counted. +7 | fixed | Root establishment split into four observable failures — not inside a repository, root mismatch after canonicalization, git not present, git ran and failed — each with its own fix, git's own message named as the discriminator, plus an explicit residual for an error none of the four explains. +8 | fixed | The squash-carry assert named `every answer record`, which appears nowhere in the NEW text; corrected to `every cycle's latest answer-record snapshot`. Every other §9 assert then re-checked against the NEW text it detects: three OLD fragments were quoted across their line wraps and would have counted 0 under `grep -F`, reading as false reds; all four are now single-line fragments verified at 1 in both copies of the parent tree, and the constraint that NEW fragments install unwrapped is stated. +9 | fixed | "A known-accepted pass whose slot is now absent or invalid…" as its own subject, with the acceptance-unknown branch following as a separate sentence. +10 | fixed | "No other pass outcome closes a cycle", with the Gate-B triviality skip named in the same breath as the one termination that is not a pass outcome; the duplicate skip sentence later in the block removed. +11 | fixed | "the five source edits, each as a pair" corrected to "the other four", the recovery pair being its own bullet, and the Severity edit given the pair it lacked (`both must resolve. Minor` at 0/1). Every other stated count then re-checked against its enumeration: the 135-condition sum, the six-member squash list, the six other "clean pass" uses, the three recovery sources, the twelve prompt-standard items, and the a/b/c accounting arithmetic (22, 18, 20) all match. +12 | fixed | Header status is "DRAFT — in Gate A, not yet approved"; the pass number is gone from durable metadata, since a number that must be updated on every revision is a number that will be stale. diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-8.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-8.md new file mode 100644 index 0000000..3a2b086 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-8.md @@ -0,0 +1,13 @@ +BLOCKER | high | §3 "What a suspension asks" / §4 "What no record carries" | After session loss, the design handles a health-stop answer only when identity is lost and a new cycle starts; when the nonce is successfully recovered, neither the stuck or two-tell continue-or-stop answer nor a no-change question decision is recoverable, and no fallback says whether the recovered cycle must remain suspended, ask again, or run another pass | A recovered cycle can bypass an explicit stop, or remain unable to determine the input that permits it to resume and eventually close | Persist suspension and question-answer state in the authoritative cycle snapshot, or require every identity-preserving recovery with missing state to re-surface all such questions and receive fresh answers before another pass +MAJOR | high | §3 "The four standing duties" / §5(c) old-condition accounting | The proposed ordering limits both the surfaced-finding hold and no-clean-credit to scope stops and expressly excludes clearly stuck, while current c16 and c18 at CLAUDE.md:242-244 and workflow-init.md:445-447 apply those duties to a clearly-stuck surface; §5 nevertheless marks c18 kept rather than replaced | The spec silently narrows two of the four duties named by the story, and a raw-severity clearly-stuck pass that is clean at effective severity can surface findings while retaining clean credit | Reconcile the reachable raw-versus-effective overlap with D3, preserve the duties for every surface still in their domain, and mark every narrowed current condition replaced with its settled authority rather than kept +MAJOR | high | §3 "First, clean completion" / "Second" | The clean predicate defines any finding outside the assigned fix set as a scope-stop trigger, and the suspension list repeats that definition, but the next sentences say an outside-set finding matched by a current-cycle decline raises no membership trigger or hold | The same re-raised finding can both make the pass unclean and be exempt from the only stop that would explain why, so readers can either close or re-suspend indefinitely | Define the membership trigger once as an outside-set finding not matched by a binding current-cycle decline while the current set excludes it, and use that exact predicate in cleanliness and suspension evaluation +MAJOR | high | §4 "Rules both labels share" / item 11 "Recovery sources" | The latest cycle body is declared authoritative and omission from it is declared loss, but branch recovery can recognize a cycle body only from cycle-attributed content it carries; a later revision body that accidentally drops the entire answer snapshot may carry no cycle identity, so the newest-first search skips it and revives an older answer | A stale decline can be replayed to release a hold, or a stale acceptance can remain in the fix set, contrary to the authoritative-snapshot rule and its stated omission behavior | Require a cycle identity and an explicit complete snapshot, including `Cycle answers: none`, in every cycle body, define how the newest body is identified independently of answer presence, and stop on a missing or malformed snapshot +MAJOR | high | §4 "Binding, and the sameness test" | The snapshot has no uniqueness or conflict rule for two records with the same five-field key, including one `Accepted:` and one `Declined:` record or repeated same-label records with divergent attribution | A fabricated, replayed, or duplicated record can simultaneously put a finding in-set and claim it was kept out, leaving hold release and resolution duties dependent on reader choice | Require at most one effective label per cycle and five-field key, collapse exactly equivalent same-label duplicates idempotently, and stop on conflicting labels or attribution +MAJOR | high | §4 item 9 "The one-contract paragraph" / §5 pointer edits / §7 Q6 | The closure-record contract list and local markers cover the central blocks and §4 transport edits, but not independently mergeable edits such as the floor and suspension pointers in §5(a)-(e) or the Q6 report block; those hunks can be adopted without the block or curve edit they reference and without carrying a local partial-adoption stop | A downstream partial merge can leave dangling ordering or Mechanics references, or install Q6's omission rule beside the old question-mark curve rule, while bypassing the compatibility guard | Include every independently mergeable pointer and Q6 hunk in the named contract with a local reciprocal marker, or make the scaffold update an atomic versioned replacement that detects and stops on partial adoption +MAJOR | high | §7 "First the root" / prompt standard 10 | The third root failure class groups Git being unavailable, ownership rejection, and unreadable metadata under one query-failed state and only prescribes making Git usable, although these are distinguishable causes with different repairs | The shipped diagnostic can send a caller through the wrong recovery and fails AGENTS.md invariant 11 via `docs/prompt-standards.md` item 10's cause-specific check-and-fix requirement | Split query failure into the observable causes the returned error distinguishes and pair each with its own fix, retaining an explicit residual for genuinely unclassified command errors +MAJOR | high | §9 "The check" | The squash-carry assertion requires the literal `every answer record` once in each shipped copy, but the specified NEW text at §4 item 1 contains `every cycle's latest answer-record snapshot` and never contains the asserted phrase | The mandatory check cannot pass against a conforming implementation, so the specified battery-plus-check evidence is mechanically impossible | Assert the actual NEW lead phrase, with the intended hyphenation and snapshot wording, and retain the parent-tree zero count +MINOR | high | §7 "Then, per earlier pass" | `Known accepted but absent, or present and now failing validation, is a real pass` grammatically applies known acceptance only to the absent case, while the next sentence classifies a present-but-invalid slot after session loss as acceptance unknown | The same invalid slot can be counted with question-mark series or omitted from the floor and curve depending on how the coordination is parsed | Write `A known-accepted pass whose slot is now absent or invalid` and then state the acceptance-unknown branch separately +MINOR | medium | §3 "Composition, and what cannot happen" / standing triviality skip | `Nothing else closes a cycle` is absolute, but the same block later says the Gate-B triviality skip ends a skipped cycle under its own rule and the standing Mechanics text records that no-pass termination | A literal downstream reader can treat the valid skip as forbidden or treat the exclusivity claim as false, weakening the promised single closure vocabulary | Scope the claim to completed-pass outcomes, for example `No other pass outcome closes a cycle`, and name the no-pass triviality skip as the separate termination path +MINOR | high | §9 "The check" | The verification text says `the five source edits, each as a pair` but its own colon-delimited enumeration contains four pairs; the recovery-source pair is a separate preceding bullet | The mechanically stated count does not match its enumeration and makes completeness of the source-edit checks ambiguous | Change this to `the other four source edits`, or move the recovery pair into the same five-item enumeration +NIT | high | Header | The artifact at committed pass 8 still reports `Gate A running (pass 6 revised)` at line 3 | The spec's own review-state metadata is stale and obscures which revision the current findings assess | Update the status to the current pass whenever the reviewed artifact is revised, or remove the pass number from durable status metadata +END OF FINDINGS (12 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-9-dispositions.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-9-dispositions.md new file mode 100644 index 0000000..d3e7c0e --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-9-dispositions.md @@ -0,0 +1,16 @@ +1 | dissolved | The snapshot/authoritative-body/malformed-stop states it attacks are deleted; its one surviving half — two records disagreeing on a five-field key — is now a question for the user, answered and recorded like any other, so it takes the loop's existing stop-and-resume transitions rather than a new terminal state. +2 | dissolved | The re-surface-before-next-pass duty it attacks is deleted. The residual paragraph is the answer: a resumed cycle continues on what it can read and the reviewer re-raises what is still true. +3 | fixed | Reverted to the frozen reading, which D4 and b6 both require: membership is answered against the set as it stood at the pass that raised the question, an explicit answer is required whatever the scope does meanwhile, and a later broadening is a new fact the next pass reads. This reverses the pass-6 choice of re-evaluation; that choice made a hold dischargeable with no user answer, which D4 forbids outright. b6 is marked kept and the reversal is stated in the (b) accounting rather than left silent. +4 | fixed | Cleanliness is defined on the logical pass with every required branch file combined; one branch's clean findings file never establishes a clean pass. +5 | dissolved | The universal `Cycle answers: none` marker that made the answer snapshot an unconditional record is deleted, so the standing "records every cycle owes" list at C:697-698 stays unchanged and the spec's claim that an answer record is conditional is true again. Verified against the standing text. +6 | dissolved | Same deletion; verified the standing list is untouched and nothing in the spec now claims otherwise. +7 | fixed | The copied `battery+check+verification` is gone from §9's heading; §9 now says the mode is read from the story header at execution and describes what each level obliges. +8 | fixed | The Severity edit reads "for every finding in the assigned fix set" and names no decline. AC 3's second half: the answer moves a finding into or out of the set and never changes what the duty demands of what is in it. +9 | fixed | e7 is now a real edit, not just a pointer: the two-tell threshold is mandatory "where the clean-completion branch did not close the pass". Marked qualified in the accounting with D2 as authority — unqualified, e7 and the ordering decide a clean two-tell pass in opposite directions. +10 | fixed | The curve edit uses Q6's own predicate, "validated a pass this session", instead of "durable proof". No pass-validity record is introduced. +11 | fixed | Ended the partition ping-pong (raised at passes 7, 8 and 9, each asking one more split). Root establishment is one condition; failure is an INCOMPLETE pass, a state §5 already defines; the report carries git's own message verbatim as the discriminator, and the spec says plainly that prose cannot partition git's failures better than git does. Shorter than what it replaced. +12 | fixed | The binding-decline overlap is named beside the demotion overlap, both in the composition paragraph and in the duties paragraph, with the same first-branch resolution: closes at or above the floor, continues below it. +13 | fixed | The reason now reads "every other pass either leaves a required repair, a finding hold or a question outstanding, or has an unmet closure precondition such as the floor". §10's item-6 quote updated to match the shipped wording. +14 | fixed | New §4 item 15: the Gate-A cadence at C:573 / W:764 becomes "revise where the severity and scope rules require a repair", so a below-floor Minor-only pass is not told to manufacture the repair the severity rule forbids. Added to the contract membership list and to §9's assert pairs. +SLIM-DOWN | Daniel's answer at the pass-9 two-tell stop | Deleted: the full-snapshot rule and its `Cycle answers` marker, the kind/artifact header fields, the authoritative-latest-body and omission-is-loss rules, the newest-first branch search and the whole recovery-passage rewrite (old item 11), the malformed-snapshot stop, the durable-validity language, and the identity-preserving re-surface duty. Kept exactly Daniel's pass-4 choice: two labels, one form, the nonce, written before the next pass, restated in the closing body, squash carry, the sameness test, the cycle binding. Replaced by one residual paragraph stating that the record is legible to a human reading history and buys no automatic recovery, which is what story criterion 7 asks it to state. +TWO-RULES | self-review, none raised by Codex | Three repaired: the hold sentence now states "one rule with two parts" explicitly (carried since pass 3); effective-severity routing is stated in one place and pointed at from the three branches that used to restate it; and the simultaneous-health-suspension sentence is framed as an instance of the per-question rule rather than an exception to it. diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-9.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-9.md new file mode 100644 index 0000000..d43117c --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-9.md @@ -0,0 +1,15 @@ +BLOCKER | high | §4 "Rules both labels share" and "Binding, and the sameness test" | The new terminal states for a missing or malformed latest snapshot and for conflicting same-key records only say to stop; they are outside the three classified suspensions and define neither a correcting input nor how a repaired body becomes authoritative | A cycle reaching either record error has no specified transition that permits another pass or closure, which violates acceptance criterion 4 and the explicit no-dead-end requirement | Define a repair-and-resume transition for each state, including who selects the valid answer, how the corrected snapshot is written, and which body becomes authoritative before the next pass +BLOCKER | high | §4 "What each is worth" and §7 "Q6" | Identity-preserving recovery must re-surface any unrecorded question or health-stop answer before another pass, but the design persists neither the suspension nor the question, while Q6 explicitly permits the originating findings slot to be absent or invalid | After both session state and that slot are lost, the cycle cannot know what it is required to ask, yet it is forbidden to run another pass; it can neither complete the required suspension nor proceed toward closure | Persist every active suspension and its exact question in the authoritative cycle snapshot, or define a conservative identity-abandonment transition that makes the lost cycle resolvable and starts a fully independent cycle +MAJOR | high | §3 "What a suspension asks" and §5(b) old-condition accounting | Membership is re-read when the answer is given and a later scope broadening automatically discharges the membership hold, contradicting settled decision 4's requirement that the surfaced finding await a user answer and current b6 at CLAUDE.md:200 and workflow-init.md:407, which the accounting marks kept and fixes the assigned set before the pass being answered | A membership-only finding can lose its hold without any explicit accept or decline, and the claimed old-condition accounting silently replaces b6 | Keep the pass-time fix set for the pending membership question and require an explicit attributable answer even if the governing scope later broadens; update the transition text while retaining b6 and decision 4 +MAJOR | high | §3 "First, clean completion" | The sentence that a clean findings file always gives a clean pass is false for a full Gate-B logical pass, where one branch may contain the clean signal while the other contains an in-set Blocker or Major | A reader can credit or close a split pass from one clean branch despite the preceding concatenation and both-branches requirements | Define cleanliness only for the validated logical pass after all required branch files are combined; state that one clean branch never establishes a clean full pass +MAJOR | high | §4 "Rules both labels share" | Only the empty snapshot gets a universal `Cycle answers: none` line; non-empty snapshots have answer records but no independent identity-and-count marker, despite the later claims that the newest body is found by that line regardless of whether it carries an answer and that a missing or malformed snapshot is detectable | A newer body that drops all answer records is not discoverable as that cycle's body, and omission of one record from a non-empty set is indistinguishable from a complete smaller snapshot, allowing stale replay or silent obligation loss | Give every cycle body a mandatory `Cycle answers: ` identity header and require exactly that many following records, including zero, before applying newest-body or malformed-snapshot rules +MAJOR | high | §4 "Rules both labels share" and the note after item 15 | Requiring every cycle body with no answers to carry `Cycle answers: none` makes the answer snapshot an unconditional record, but the spec leaves the standing "records every cycle owes" list at CLAUDE.md:697-698 and workflow-init.md:883-884 unchanged and expressly claims answer records are conditional | Clean and trivially skipped cycles can follow the surviving exhaustive-looking list, omit the required snapshot marker, and then fail recovery or appear partially adopted | Add the answer-record snapshot to the standing unconditional-record list and to its old-condition accounting, including the triviality-skip path; reserve conditionality for individual Accepted and Declined entries +MAJOR | high | Header and §9 "Verification" | The header says the story profile is the only writable copy and that no value may be copied into the spec, but §9 copies `battery+check+verification` into its heading and specifies the verification under that remembered mode | A later profile change can leave the spec prescribing stale evidence while appearing to comply with the single-source rule | Remove the copied mode value and require execution to read the current validation mode from the story header before selecting the applicable battery, check, and verification obligations +MAJOR | high | §4 "Which rules the answer modifies" and item 14 "Severity" | The spec says a decline never qualifies the resolve duty and that no rule there mentions a decline, yet the proposed Severity rule explicitly says a declined finding owes nothing while excluded | The proposed text violates acceptance criterion 3's "at no unmodified rule" half and makes the decline look like a direct waiver of Blocker or Major resolution rather than a membership decision | Limit the Severity edit to "for every finding in the assigned fix set" and leave all decline-specific effects in the membership, hold, and clean-pass rules that the answer actually modifies +MAJOR | high | §5(e) "The five tells" | The current sentence "Any two present makes stop-and-surface mandatory" remains unconditional and is marked kept, while the ordering says tells on a clean closing pass never block closure and clean completion outranks the two-tell stop | Both prompt copies would still contain one instruction requiring suspension and another requiring closure for the same clean two-tell pass, defeating settled decision 2 and acceptance criterion 1 | Qualify the tell threshold as mandatory only after the clean-completion branch has not closed the pass, and mark e7 and its consequence as qualified or replaced in the old-condition accounting +MAJOR | high | §4 item 15 and §7 "Then, per earlier pass" | The curve edit says a question-mark entry requires durable proof that the pass was valid, but Q6 defines known acceptance solely as validation remembered by the current session and explicitly creates no pass-acceptance record | Readers have no defined durable artifact that satisfies the new proof test and can either count an unavailable pass toward the floor or omit it depending on which section they follow | Replace "durable proof" with the exact current-session validation predicate used by Q6, or introduce and fully specify a durable pass-validity record and its recovery, carry, and threat semantics +MAJOR | high | §7 "First the root" | The claimed four-way diagnostic says each failure has its own fix, but "git ran and failed" still combines ownership refusal, unreadable or corrupt metadata, permissions, and other command failures with different remedies; only the common ownership case receives a fix | The shipped prompt still violates prompt-standard item 10 and can send a user with a non-ownership Git failure toward an irrelevant trust operation | Partition returned Git failures into the distinguishable cause classes and pair each with its check and repair, leaving a genuinely unclassified returned error as an explicit residual with a retry or escalation transition +MAJOR | high | §3 "Where the other two exits stand in this duty" | The text treats the only no-scope-trigger overlap between clean completion and clearly stuck as an in-set reviewer-written Blocker or Major demoted by the ceiling, but an out-of-set regenerating Blocker or Major matched by a binding decline also has no membership trigger and is clean because it is outside the fix set | A reachable decline-and-reraise conflict required by acceptance criterion 1 is omitted, so readers can wrongly attach a hold or deny clean precedence on that path | Name the binding-decline overlap alongside the demotion overlap and state that the same first-branch ordering closes it at or above the floor and continues it below the floor +MINOR | high | §3 "First, clean completion" | The reason given for "No other pass outcome closes" says every other path leaves a finding or question open, but a below-floor clean pass carrying only collected Minors or Nits has no repair or answer outstanding and continues solely because the floor is unmet | The rationale contradicts the continue branch and weakens prompt-standard item 6 by teaching the wrong reason for a non-closing state | Say that every other pass either has an unmet closure precondition such as the floor or leaves a required repair, finding hold, or question outstanding +MINOR | medium | §3 "Third, a pass that neither closes nor suspends continues" | The new instruction to rerun on the current artifact "revised or not" conflicts with the unchanged Gate-A cadence at CLAUDE.md:573 and workflow-init.md:764, which imperatively says "Each pass: validate, revise, re-run" | A below-floor clean pass with only collected Minors or Nits can prompt a manufactured revision despite the rule that those findings are never iterated | Change the standing cadence to revise only when the governing severity and scope rules require a repair, then rerun while the cycle remains open +END OF FINDINGS (14 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md new file mode 100644 index 0000000..310ce74 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -0,0 +1,387 @@ +# Gate-A (spec) working record — cycle awsf1ec771 + +Advisory, cycle-stable, per CLAUDE.md §5 optional companions. Retire at closure. +Nothing depends on it; the pass files and the repo are authoritative where this disagrees. + +- **Kind:** Gate-A spec +- **Nonce:** awsf1ec771 (drawn 2026-09-10 from /dev/urandom, 10 chars, no collision among open cycles) +- **Artifact:** `docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md` +- **Branch:** loop-rule-consolidation +- **Story:** `docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md` — profile read from its header at each pass (was high / none / battery+check+verification at pass 1) +- **Derived floor:** 3 (risk high → level 2; security none → 0; max 2 ≠ 0 → 3) +- **Hook knob:** absent (no `.context/codex-gate.floor`-style file observed) + +## Passes + +| Pass | Spec rev | Findings | Blockers | Majors | Valid | Notes | +|---|---|---|---|---|---|---| +| 1 | 510f6ce | 24 | 5 | 15 | yes | all B/M in-set; dispositions in `gate-a-spec-awsf1ec771-pass-1-dispositions.md` | +| 2 | 07c88a1 | 17 | 2 | 13 | yes | all in-set; #8's mandatory-record demand dismissed (existing no-identity rule answers it); session 01a08b0b-495e-7ee0-8641-631684a4db4e | +| 3 | 8d20b2e | 12 | 4 | 6 | yes | #5 re-raises the mandatory-record demand → dismissed a second time, residual disclosed instead; #11 collected; #12 severity corrected MINOR→MAJOR (instrument carve-out, false green); session 01a08b2f-8fce-7130-96c9-d70d94f48778 | + +| 4 | 0a5205f | 18 | 1 | 10 | yes | **SCOPE STOP surfaced to Daniel** on finding 1 (mandatory record for accepted obligations — a new contract question, raised a 3rd time and rebutting the D10 dismissal correctly); other 17 held, unrepaired, pending the answer; session 01a08b6f-09b2-7c00-99c6-9bd230c32d18 | + +| 5 | 4625679 | 17 | 0 | 15 | yes | zero Blockers; #16 collected; revision must SHRINK (3 concepts removed); session 01a08b9b-456e-7fa2-a50a-2c56ec4e08c8 | + +| 6 | 726a5de | 16 | 4 | 9 | yes | 4 findings are one structural defect (clean candidacy defined on a post-hoc predicate); session 01a08bbe-93da-7951-b694-2bd4f5bff99a | + +| 7 | de66e00 | 14 | 0 | 9 | yes | structural repair converged: Blockers 4→0, findings 16→14; session 01a08bdc-7f64-7980-87f6-768998056025 | + +| 8 | 5aef816 | 12 | 1 | 7 | yes | 5 of 12 mechanical (assert, count, header, grammar); session 01a08bfb-3cd3-7ff3-99a2-c887a75d316d | + +| 9 | 2feab0c | 14 | 2 | 10 | yes | **TWO-TELL STOP — mandatory, surfaced to Daniel.** All 14 held open; session 01a08c14-71b6-74f1-a835-8098e0f3c90a | + +| 10 | 18c742f | 20 | 2 | 15 | yes | **SECOND TWO-TELL STOP + clearly-stuck all three conditions met.** Surfaced to Daniel with a split recommendation; all 20 held open; session 01a08c43-ad6f-75d2-92ce-9f7497be3f51 | + +| 11 | c8f96b8 | 20 | 1 | 13 | yes | first pass on the narrowed 647-line spec; finding 14 names the cause — the block restates rules their own paragraphs still own; session 01a08c65-7a3e-7c80-8cd5-2565f1b61d45 | + +| 12 | 11b0e47 | 21 | 1 | 13 | yes | **two mechanisms ended by Daniel's repeat criterion**, brought in mid-loop from another project; session 01a08c7f-c573-7df1-8b31-7704ec7be064 | + +| 13 | 78e5f97 | 16 | 2 | 8 | yes | **the loop turned: B+M 14 → 10, Majors 13 → 8**; session 01a08c98-b6cd-7bc3-9556-8f1c87491031 | +| 14 | 0168f88 | — | — | — | not run | next action. Spec cut to 532 lines; design 36% of it. | + +## Pass-13 three-line report + +- **Trend:** findings …, 20, 21, **16**. Blockers …, 1, 1, **2**. Majors …, 13, 13, **8**. + Blocker+Major 20, 15, 10, 11, 15, 13, 9, 8, 12, 17, 14, 14, **10** — flat for three passes, then + the first real fall since pass 8. +- **Cluster (pass 13):** product 11 of 16; the instrument 1 (§7's table inputs); bookkeeping 4 + (two wrong counts, two partial OLD quotations). +- **require↔withdraw:** none. + +**Tells: one of five** — the Blocker count 1 → 2. Findings falling, Majors falling hard. + +**What the repeat criterion actually bought, measured rather than asserted.** Mechanism 1, the §7 +substring asserts: **3 Majors at pass 12, 0 at pass 13.** Dead, because the work moved to the plan +rather than being repaired a fifth time. Mechanism 2, the block restating what it cites: **4 → 2**, +and both survivors are a narrower shape — a *source* paragraph restating (`b3`), and a citation +that omits a source — not the block restating. Half-ended. + +**What it did not buy, and the criterion is not built to.** Pass 13's finding 2 is the pass-12 +revision's own fix regenerating: a contradiction surface was introduced with no next state. That is +the no-progress defect class, and only the transition walk catches it — which is why the pass-13 +revision brief adds one as a standing check. + +**A second signal, worth more than the count.** Pass 13's findings are mostly the spec's edges +meeting **standing** text — Mechanics "Finishing the cycle", the evidence-entry revalidation rule, +`b8`, `b3` — rather than the spec contradicting itself. The reviewer is running out of internal +problems and working the boundary, which is the shape a converging loop takes. + +## Pass-12 three-line report + +- **Trend:** findings …, 20, 20, **21**. Blockers …, 2, 1, **1**. Majors …, 15, 13, **13**. + Blocker+Major 20, 15, 10, 11, 15, 13, 9, 8, 12, 17, 14, **14**. Flat for three passes. +- **Cluster (pass 12), split by Daniel's criterion rather than the old three buckets:** + **product 10** (1 Blocker, 9 Major — the shipped rules); **harness 6** (4 Major on §7's assert + list, all of them false-green class, so they keep their severity under the existing instrument + carve-out); **bookkeeping 5** (all NIT, path citations). +- **require↔withdraw:** none this pass. + +**Tells: zero of five** by the stated definitions. The loop is not tripping its own thresholds, and +that is exactly the gap Daniel named: a flat curve with no tell is invisible to the rules. + +**The criterion Daniel brought, and what it ended.** Three lines from a sibling project: a product +Major always blocks; a harness Major blocks only where it makes a claim vacuous; **the second +finding of the same shape against the same mechanism ends that mechanism's rounds** — narrow the +claim, print the residual, move the work. The kit already carries the first two in the Mechanics +severity carve-out. The third has no counterpart, and two mechanisms here were on their fourth +round of one shape: +- **§7's assert list** — "an assertion that does not detect what it claims", at passes 8, 8-audit, + 11 and 12. Cause is structural: a spec cannot build an exact substring check for text that does + not exist yet. **Moved to the plan**, where the substrings exist; §7 now states what must be + verified and names which passages owe a pair, and prints the residual. +- **The block restating what it cites** — passes 11 and 12, four findings. **Made mechanical**: a + cited rule contributes zero predicate words to the block, which is checkable by reading. + +**The criterion is NOT a shipped rule and this record must not read as if it were.** It was applied +here as a judgement call by the agent running the cycle, on Daniel's endorsement, in the same way +any triage call is made. Nothing in `CLAUDE.md` obliged it and nothing checked it. Captured for the +kit as `docs/superpowers/stories/2026-09-10-harness-finding-termination-story.md` — intake only, +profile proposed and unconfirmed, so the criterion is a candidate rule that has been used once. + +That intake found the gap more precisely than the summary above: the existing severity ceiling +sends a non-vacuous harness Major to "Minor or below: collect, never iterate", and **"collect" +names no destination** — no ticket, no tripwire, no trace outside that pass's report. It also found +a partial counterpart nobody had named: `harden-finding` and the ledger escalate a *recurring* fix +after a finding closes, which is not a move a running loop can make. + +## Pass-11 three-line report + +- **Trend:** findings 24, 17, 12, 18, 17, 16, 14, 12, 14, 20, **20**. Blockers 5, 2, 4, 1, 0, 4, 0, 1, 2, 2, **1**. + Majors 15, 13, 6, 10, 15, 9, 9, 7, 10, 15, **13**. Blocker+Major 20, 15, 10, 11, 15, 13, 9, 8, 12, 17, **14**. +- **Cluster (pass 11):** product behaviour 14 of 20; the test instrument 6 (§7's oracle and rows, + and three wrong line-range claims in the mechanical checks). Instrument at 30% is the highest + it has been and worth watching, though product is still the cluster. +- **require↔withdraw:** one pair. Finding 4 demands the hold apply to **every** surfaced finding, + which is what the passes 6–7 revisions removed when they narrowed it to scope stops. The story's + own third duty says every surfaced finding, so the pair resolves back to the story. + +**Tells: one of five.** The finding count is flat rather than rising, Blockers fell, the cluster is +product. **The split helped**: Blocker+Major 17 → 14 and Blockers 2 → 1 across a change that also +cut 320 lines. + +**A convergent move is nameable again, and it is finding 14's.** The block restates triggers, +duties, preconditions and the severity answer that their own paragraphs still define, so each copy +holds two authorities and findings 3, 4 and 6 are that drift already happening. The repair inverts +it: the block owns the evaluation order and closure only, and cites every other rule where it +already lives. That is what the story's desired outcome asked for and what the pass-6 repair — +the one that worked — looked like. + +## Pass-10 three-line report — MANDATORY STOP, AND THE STUCK EXIT IS NOW READABLE + +- **Trend:** findings 24, 17, 12, 18, 17, 16, 14, 12, 14, **20**. Blockers 5, 2, 4, 1, 0, 4, 0, 1, 2, **2**. + Majors 15, 13, 6, 10, 15, 9, 9, 7, 10, **15**. Blocker+Major 20, 15, 10, 11, 15, 13, 9, 8, 12, **17**. +- **Cluster (pass 10):** product behaviour 15 of 20; the test instrument 2 (§9's oracle and its + row set); prose about it 3. Within product: **four on the record transport** (2, 4, 5, 6), + **three on the Q6 root check** (9, 10, 11), two on rollback and the discriminator (1, 15), + five on the ordering (3, 13, 18, 19, 20), one on sameness (7). +- **require↔withdraw:** present. Pass 10 finding 7 demands byte-exact field equality for the + sameness test, which is what the pass-2 revision removed in favour of reading it on meaning. + +**Tells: two of five, unambiguously — the threshold.** Finding count rising 14 → 20; Blocker count +failing to fall, 2 → 2. Stop-and-surface is mandatory. + +**And this time the clearly-stuck reading is satisfied on all three conditions**, which it was not +at pass 7: a plateau across ten passes (Blocker+Major never below 8, never zero); coverage +affirmable after ten readings of every section; and Blocker/Major findings regenerating from the +previous round's own repairs — findings 2, 4, 5 and 6 are the pass-9 slim-down's, and 9, 10 and 11 +are the third rework of a root check that pass 4 introduced. + +**Two-tell stop ANSWERED 2026-09-10 by Daniel: SPLIT.** The record and its transport, the +unavailable-history report, the checkout-root condition, the rollback reading and the +discriminator dissolution move to `docs/superpowers/stories/2026-09-10-record-durability-story.md` +(profile proposed, unconfirmed — Daniel decides before it is executable). The parent story's §2 +records the split and its criterion 7 moved with them, so it carries six criteria again +(commit 6cbc174). Twelve of pass 10's twenty findings are **deferred, not fixed**: 1, 2, 4, 5, 6, +7, 8, 9, 10, 11, 14, 15. Eight survive the narrowing and are fixed: 3, 12, 13, 16, 17, 18, 19, 20. + +**The cycle continues under this nonce.** The artifact was revised, as it has been at every pass; +the floor of 3 is long met; what is still owed is a clean final pass against the current artifact, +which is now the narrowed one. The pass counter continues at 11. + +**Why no convergent move is claimed this time.** At pass 6 one was nameable and it worked; at +pass 9 one was nameable and it did not — the slim-down cut ninety lines and the next pass was the +worst since pass 1 on Majors. What still regenerates is not one defect but one *subject*: the +durability of records and history across sessions, which the record transport, Q6, the root check +and the rollback paragraph all belong to. The ordering itself has converged — its five remaining +findings are small. + +## Pass-9 three-line report — MANDATORY STOP + +- **Trend:** findings 24, 17, 12, 18, 17, 16, 14, 12, **14**. Blockers 5, 2, 4, 1, 0, 4, 0, 1, **2**. + Majors 15, 13, 6, 10, 15, 9, 9, 7, **10**. Blocker+Major 20, 15, 10, 11, 15, 13, 9, 8, **12**. +- **Cluster (pass 9):** product behaviour 13 of 14 — and within it, **five on the answer-record + transport** (1, 2, 5, 6, 10) and six on the ordering; prose about it 1; the instrument 0. +- **require↔withdraw:** **one pair, and this one fits the definition.** Pass 6 asked whether a + pending membership question is frozen at surface time or re-evaluated; re-evaluation was chosen + and the frozen reading removed. Pass 9 finding 3 demands the frozen pass-time fix set back. + +**Tells: three of five present — the threshold is two, so stop-and-surface is mandatory and not +discretionary.** Finding count rising (12 → 14); Blocker count failing to fall (1 → 2); a +require↔withdraw pair. The clearly-stuck reading is **also** satisfied for the first time — a +plateau across nine passes, coverage affirmable, and Blocker/Major findings regenerating from +the previous round's own repairs — but it is not needed: the two-tell threshold stands alone. + +**Two-tell stop ANSWERED 2026-09-10 by Daniel: continue, with the answer record slimmed back to +his pass-4 choice.** Deleted: the snapshot rule and its marker, the kind/artifact header fields, +the authoritative-latest-body rule, the newest-first branch search and the recovery-passage +rewrite, the malformed-snapshot stop, the durable-validity proof, and the re-surface-before-next- +pass duty. Kept: the two labels, one form, the nonce, written before the next pass, restated in +the closing body, squash carry, the sameness test, the cycle binding. In their place one residual +paragraph: the record is legible to a human reading history and buys no automatic recovery — +which is what the story's criterion 7 asks it to state. Findings 1, 2, 5 and 6 dissolve; finding 3 +reverses the pass-6 choice back to the frozen pass-time fix set, as D4 and `b6` require. + +**Diagnosis carried to the user rather than acted on.** The closure ordering, which is what the +story asked for, has converged: its findings are narrow and its Blockers came and went. The +**answer-record transport has not.** It began as one commit-body line and has accreted an +identity line, a full-snapshot rule, an authoritative-body rule, a newest-first branch search, a +malformed-snapshot stop and a durable-validity proof — and each round's repair of it produces the +next round's findings. Five of pass 9's fourteen sit there, including both Blockers. + +## Pass-8 three-line report + +- **Trend:** findings 24, 17, 12, 18, 17, 16, 14, 12. Blockers 5, 2, 4, 1, 0, 4, 0, 1. + Majors 15, 13, 6, 10, 15, 9, 9, 7. **Blocker+Major: 20, 15, 10, 11, 15, 13, 9, 8.** +- **Cluster (pass 8):** product behaviour 9 of 12; the test instrument 2 (a false assert and a + false count in §9); prose about it 1 (stale header metadata). +- **require↔withdraw:** none. Findings 4, 6 and 7 each extend a pass-7 requirement that was + applied incompletely; none demands something an earlier pass removed. + +**Tells: one of five** — the Blocker count 0 → 1, counted conservatively as failing to fall even +though the four-pass shape is 0, 4, 0, 1. Findings falling, Majors falling, cluster on product. + +**Note for the plan phase, not acted on mid-loop:** most of the spec's growth from 462 to 940 +lines is §4's full OLD/NEW quoting of small source edits. By the story's own sizing guidance that +is plan detail. Moving it mid-loop is the restructure that produced pass 6's Blockers, so it is +recorded here for the plan to take deliberately. + +## Pass-7 three-line report + +- **Trend:** findings 24, 17, 12, 18, 17, 16, 14. Blockers 5, 2, 4, 1, 0, 4, **0**. Majors 15, 13, 6, 10, 15, 9, **9**. +- **Cluster (pass 7):** product behaviour 13 of 14; prose about it 1 (a false count in the spec's + own claim); the test instrument 0. +- **require↔withdraw:** **one pair, counted.** Pass 6 offered two fixes for the partial-adoption + guard, a central membership list or a marker on every mergeable hunk; the central list was + taken and the markers dropped. Pass 7 demands the markers. The removal was mine at pass 6's + invitation rather than pass 6's own, which is one step removed from the definition — counted + as present anyway, because the conservative reading is the honest one. + +**Tells: one of five.** Findings falling, Blockers back to zero, cluster on product behaviour. + +**The clearly-stuck exit is no longer reachable and the pass-6 note is superseded.** Its third +condition required Blocker/Major findings regenerating across repair attempts; the pass-6 repair +took Blockers to zero and the count fell, so the regeneration condition fails. Two of three are +gone, not two of three present. The pass-7 revision aims at a clean pass 8. + +- **Trend:** findings 24, 17, 12, 18, 17, 16. Blockers 5, 2, 4, 1, 0, **4**. Majors 15, 13, 6, 10, 15, 9. +- **Cluster (pass 6):** product behaviour 15 of 16; prose about it 1 (the §10 conformance claim); + the test instrument 0. +- **require↔withdraw:** none under the definition. Two near-misses, named: pass 3 required the D3 + sentence preserved in full and pass 6 says that preservation now contradicts the ordering; pass + 5 required the fourth duty over its original domain and pass 6 says the placement it produced is + circular. Both are a later pass objecting to an earlier pass's *addition*, which is the mirror of + the pair's shape rather than the shape itself. + +**Tells: one of five** — the Blocker count failing to fall (0 → 4). Findings still falling; the +cluster is product behaviour; no pair. One is not two. + +**The clearly-stuck reading, examined rather than assumed.** Two of its three conditions are now +present: a plateau is visible (12–18 findings, 10–20 Blocker+Major, for five passes, and six is +where the field saw one), and Blocker/Major findings are demonstrably regenerating — six of pass +6's sixteen trace to pass 5's own repairs, with the lineage nameable. The third, an **affirmative +judgement that coverage is sufficient**, is available. **The exit is nonetheless declined, and the +reason is the rule's own:** reporting "will not converge" is a false report where the convergent +move can be named, and it can be — four of the six regenerated findings are one defect, clean +candidacy resting on a predicate produced by a later branch, and pass 6 states its repair. One +more round is spent on that repair rather than on the exit. If the round after it regenerates +again, the third condition stops being affirmable in good faith and this exit is taken. + +## Pass-5 three-line report + +- **Trend:** findings 24, 17, 12, 18, 17. Blockers 5, 2, 4, 1, **0**. Majors 15, 13, 6, 10, 15. +- **Cluster (pass 5):** product behaviour 14 of 17; prose about it 3 (the a13 accounting, the + rollback claim, the scope-provenance sentence); the test instrument 0. +- **require↔withdraw:** none under the definition. One near-miss named rather than hidden: pass 4 + required a wrong-root **stop**, pass 5 says that stop contradicts D10 and offers removing it as + one of two fixes. That is a later pass questioning an earlier pass's addition, which is the + mirror of the pair's shape, not the shape itself. + +**Tells: zero of five.** Not stuck either — the clearly-stuck exit needs all three of a plateau, +an affirmative coverage judgement and regenerating Blocker/Major findings. The third is plainly +present (pass 4's fixes produced six of pass 5's findings), the first is approaching, and the +second cannot be affirmed: each pass still reaches material the previous ones had not read. Two +of three is not the exit. + +**Majors rose 10 → 15 while Blockers fell to 0.** Read as density, not regression: six of the +seventeen are first review of the `Accepted:` record authorised between passes, and most of the +rest say one thing — the record was carrying jobs it did not need. The pass-5 revision removes +three concepts rather than patching them. + +## Pass-4 three-line report (duty active from pass 4) + +- **Trend:** findings 24, 17, 12, 18 — **rising at pass 4**. Blockers 5, 2, 4, 1. Majors 15, 13, 6, 10. +- **Cluster (pass 4):** product behaviour 12 of 18 (the shipped ordering, decline block, Q6 text, + severity paragraph); prose about them 6 (the spec's own accounting, inventory ids, parity + rationale, rollback and dissolution claims, the §10 conformance claim); the test instrument 0. +- **require↔withdraw:** none. Findings 2 and 3 report an *incompletely applied* pass-3 decision, + not a reversal of one; finding 1 is a demand repeated after dismissal, which is not a withdrawal. + +**Tells: one of five** — the finding count rising. Blockers are at their lowest; the dominant +cluster is product behaviour, not the instrument or prose about it; no pair. One tell is not two, +so no mandatory tell-stop. **The loop was stopped by the scope stop on finding 1 instead.** + +**Scope stop ANSWERED 2026-09-10 by Daniel: accepted.** The finding enters the fix set and the +change ships a second record label, `Accepted:`, sharing the decline record's form, transport, +nonce and carry rules. Authorised scope growth beyond the story's §2, recorded here and in the +pass-4 dispositions file. This Gate-A cycle runs under the pre-change rules, so it writes no +accept record of its own — the rules it started with have none. + +Tells after pass 3 (duty starts at pass 4; tracked early): findings 24→17→12 falling; +**Blockers 5→2→4 — failing to fall, one tell present**; clusters are product-behaviour (the +shipped rules), not instrument or prose-about; no require↔withdraw pair (pass 3 narrows pass 2's +decline-suppression demand rather than withdrawing it; the mandatory-record demand is a repeat +after dismissal, not a withdrawal). One tell is not two — no mandatory stop. + +## Prompt + +**`.context/gate-a-spec-prompt.md`** — durable copy, substitute `__SHA__` and `__P__`. It carries +the pre-call checklist and the three "deliberately not here" blocks that have to stay in it, or the +reviewer re-raises deferred material. The scratchpad copy is gone with its session. + +## STATE AT HANDOFF — 2026-09-10 evening + +- **HEAD `0168f88`** on `loop-rule-consolidation`, tree clean. Spec **532 lines** (was 785). +- **Pass 13's findings are all dispositioned and applied**; dispositions in + `gate-a-spec-awsf1ec771-pass-13-dispositions.md`. Findings 15, 16, 13, 14 dissolved with the cut; + 4, 5, 12 became rows in the new §4 table; 6 became one line in §6; 10 was absorbed by §7's + narrowing; the rest fixed in §3. +- **The next action is Gate-A spec pass 14** against `0168f88`. Nothing else is pending. +- **The cut, measured:** the design is now **192 of 532 lines, 36%**, against 173 of 785, 22%. + §5 accounting 216 → 33, §7 verification 119 → 67, §4 quoting 83 → 54, §6 parity 38 → 18. The + bookkeeping did not vanish; the plan carries it, beside the edits, where it is checkable against + real files. +- **Transition walk done and clean** — every state the spec names reaches close or park, and no + state returns itself with its input consumed. That was pass 13 finding 2's defect class. + +## Next — resume procedure, in order + +1. `git -C /Users/daniel/DEVELOPMENT/APPS/dev-workflow-kit log --oneline -5` and `git status --short`. Branch is `loop-rule-consolidation`, tree must be clean. +2. Read this file's Passes table and the newest three-line report for where the loop stands. +3. `python3 .context/spec-precheck.py docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md` — exit 0, and eyeball its COUNT list against the spec's enumerations. +4. `rm -f .context/codex-reviews/gate-a-spec-awsf1ec771-pass-.md`, confirm gone. +5. Fill `__SHA__` (current HEAD, short) and `__P__` into `.context/gate-a-spec-prompt.md`, call `mcp__codex__exec` with `workingDirectory` = repo root. +6. Validate the file: last line exactly `END OF FINDINGS ( total)`, exactly `` finding lines, nothing else. Anything else is an INCOMPLETE pass — do not count it, do not act on the partial list. +7. Disposition every finding, then have a fork revise. Report the three lines to Daniel (trend, cluster, require↔withdraw) — the duty is active from pass 4 and every pass owes it, plus the floor line (floor 3, risk high, security none, read from the story header). +8. Append a row and a report section here before running the next pass. + +## Standing practices adopted mid-cycle — NONE of these is a shipped rule + +Each was applied as a judgement call on Daniel's endorsement. `CLAUDE.md` obliges none of them and +nothing checks them. They are candidates captured in +`docs/superpowers/stories/2026-09-10-harness-finding-termination-story.md`. + +1. **Repeat criterion.** The second finding of the same shape against the same mechanism ends that + mechanism's rounds: narrow the claim, print the residual, move the work. Measured: it took the + §7 substring-assert mechanism from 3 Majors to 0 in one pass. +2. **Pass ceiling.** A cycle grinding past floor + 3 without a clean pass owes a stop-and-surface. +3. **Size limit.** A design spec whose bookkeeping outweighs its design gets cut before the next + pass, not diagnosed after it. +4. **Precheck before every pass** — step 3 above. §5 already requires this and it was dropped after + pass 1; roughly a quarter of later findings were things it decides in seconds. +5. **Transition walk before every commit.** List every state the spec names with its input and next + state; confirm close or park is reachable from each. Pass 13's finding 2 was the previous + revision's own fix creating a dead end, and only this walk catches that class. + +## Daniel's decisions this cycle, in order + +| When | Decision | +|---|---| +| before pass 1 | Profile confirmed `high / none / battery+check+verification`; slot discriminator dissolved rather than shipped. | +| pass 4 scope stop | Accept the finding: ship an `Accepted:` record label beside the decline record, sharing form, transport, nonce and carry rules. | +| pass 9 two-tell stop | Continue, but slim the record back to that pass-4 choice; delete the snapshot, identity, search, authoritative-body and durable-validity machinery. | +| pass 10 two-tell stop | **Split.** Record durability moves to `docs/superpowers/stories/2026-09-10-record-durability-story.md`; this story keeps the ordering, the duty classification and the severity/health answer. | +| mid-loop, from another project | Anchor the repeat criterion in the kit; captured as the harness-finding-termination story. | +| after pass 13 | **Cut the spec to the design**; the plan carries the bookkeeping. Precheck runs before every pass. | + +## For the execution phase, not needed yet + +When the diff actually edits `CLAUDE.md` §5, the Gate-B cycle runs under rules the diff is changing. +Copy §5 to `.context/rules-.md` at cycle start and read the cycle's own rules from there. +That makes "a cycle already running finishes under the rules it started with" a fact rather than an +intention. One command; not needed while only the spec is being written. + +## Where everything lives + +| What | Path | +|---|---| +| spec under review | `docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md` | +| condition inventory (135 ids, committed) | `docs/superpowers/specs/2026-09-10-loop-rule-consolidation-condition-inventory.md` | +| this story | `docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md` | +| successor, profile unconfirmed | `docs/superpowers/stories/2026-09-10-record-durability-story.md` | +| loop-termination story, profile unconfirmed | `docs/superpowers/stories/2026-09-10-harness-finding-termination-story.md` | +| pass files and dispositions | `.context/codex-reviews/gate-a-spec-awsf1ec771-pass-*.md` | +| precheck | `.context/spec-precheck.py` | +| pass prompt | `.context/gate-a-spec-prompt.md` | + +**Two stories await a profile confirmation from Daniel** and are not executable until he gives it. diff --git a/.context/codex-reviews/gate-a-spec-pass-1-dispositions.md b/.context/codex-reviews/gate-a-spec-pass-1-dispositions.md new file mode 100644 index 0000000..c23a493 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-1-dispositions.md @@ -0,0 +1,138 @@ +# Gate A — spec — pass 1 dispositions (salvage cycle) + +15 findings (3 BLOCKER, 10 MAJOR, 2 MINOR). **14 applied; finding 12 dismissed with its +reason, human-confirmed.** That is the first dismissal in 319 findings across four cycles. + +## Where this sits against the three stopped cycles + +| Cycle | Blockers by pass | Findings pass 1 | +|---|---|---| +| Three-tier | 4 → 4 → 6, stopped | 37 | +| Two-tier + debt | 7 → 8 → 12, stopped | 30 | +| Stripped tier 3 | 4 → 14 → 15, stopped | 28 | +| **Salvage** | **3** | **15** | + +Half the finding count of any predecessor, and — the part that matters — **the three blockers +converge on one fix, and that fix was available immediately.** No finding says the mechanism +cannot exist. The mechanism is a paragraph. + +## The three blockers are one blocker + +**1, 11 (and 5 supplies the fix).** The draft wrote the record form as covering *"a pass short +of the floor, an absent bot review, a check that could not be run"* while asserting it +authorized nothing. Pass 1 rejected that as a distinction without a difference: a normative +instruction shaped *"when a human decides to proceed past X, write this"*, a required commit +form, and §6 routing the stall through it **is** a route past the gate, whatever the +surrounding sentence claims. A compliant agent reads assent + three lines as sufficient to +continue after a STOP — the waiver §1 found unbuildable, re-entering through the prose. And +11 makes it mechanical rather than a matter of reading: with below-floor cases included, +`docs/getting-started.md` ("Gate A is not skippable at any level"), +`docs/coding-workflow.md` and `docs/sparring-briefing.md` ("do not treat a satisfied human as +a substitute for a clean pass") are all falsified, and §4's claim that they stay true is +simply wrong. Both quotes verified in the working tree. + +**Finding 5 gives the fix, and it is the more embarrassing half:** the draft mis-stated its +own precedent. PR #14 was a merge past an **absent supplementary PR-bot review** on a change +that **had already passed Gate B**. Never a core-gate closure. + +**Fix, applied as new §2.0:** the form is scoped to what the precedent actually covers — +*it applies only where §5's own gates are satisfied, and never to a gate, a floor, a pass +count, or a profile-derived evidence obligation.* The shipped paragraph carries that scope +itself, including the sentence "if you are reaching for this form to get past a gate, the +answer is no". This is not a softening; it is the difference between a record and a waiver, +and it makes §4's untouched-sites claim true rather than aspirational. + +**2.** "A reader can see that a human chose" overclaims, while the same paragraph admits +nothing verifies the handle or the pause. The gate-proof Don't, applied to a record instead of +a mechanism. **Fix:** the shipped text now says the record is an **unverified assertion** and +that a reader learns only that *the commit claims* a human chose. + +## Majors and minors — applied + +**3** — §1's finding overclaimed impossibility with no defined properties. Rewritten: §1.1 +names the four properties ("safe" = grounded, guarded, bounded, attributable), §1.3 maps each +failure mode to the design class it defeats and says *why* it is structural, and §1.5 states +plainly that this is repeated structural failure, **not** exhaustion of the design space — a +class resting on authority outside the repository is untouched by this evidence. Turning the +gate-proof calibration rule on a negative claim was the right call and the artifact is more +useful for it. + +**4** — four accounting rows were not enough; the draft's wording created edges beside several +of §5's terminal actions while claiming everything was untouched. §2.3 now lists sixteen +clauses individually, including the floor, the clean-final-pass rule, the zero-finding exit, +the stuck STOP, recovery-and-STOP, INCOMPLETE discounting, the triviality skip, mode +obligations, the evidence-gap override, and work-gap-vs-setup-gap. + +**6** — rider (b) contradicted §5's "companions never validate a pass". This *was* accounted +for in the rejected design (its row 29, "narrowed in one named scope") and the accounting was +lost in the rewrite. Restored, and both prompt copies must state the narrowing in the +companions clause itself. + +**7** — a disappearing drift record discounts a pass with nothing rechecking credited passes. +Added: a pre-close audit before the final pass and again before the closing commit, with +numbering and STOP behaviour defined. + +**8** — the drift grammar was a sketch. Now complete: position, multiplicity, line-list +ordering, token escaping, empty-token handling, Gate-B `full` per branch. + +**9** — the carry chain described only Gate B's amend flow. Added §2.4: per-gate placement, +accumulation, the identical-collapse and differing-both-kept rules, both merge strategies, and +an explicit statement that **carry is instruction-backed** — nothing validates it. + +**10** — verified mechanically and correct: the tier-2 story asserted at three live sites that +tier 3 exists, closes with no passes, and unblocks work, including an open question asking +whether tier 2 is worth building "once tier 3 exists". Amended with old-condition +dispositions; both stories added to §4's site list. + +**13** — the `+check` lived only in a story comment, so a plan could implement the edits and +skip it. Promoted to a new §5 Validation: fixture, before/after, both prompt copies, and the +counterfactual, which is the observation PR #23 actually produced. + +**14** — `§4` → `§5` xref. Applied, and then the section renumber moved the backlog to §6, so +the mechanical sweep caught the same class again on the rewritten text and fixed it to §6. +Second time the sweep has paid for itself this cycle. + +**15** — the tracked-debt trigger presupposed a follow-up obligation the shipped form does not +create. Re-stated as an observable event: a human explicitly asks for follow-up and it is +later found not to have happened. + +## Finding 12 — DISMISSED, human-confirmed 2026-08-14 + +**This is the first dismissed finding across all four cycles — 319 findings in.** Labelled a +dismissal rather than "held", because in protocol terms a finding not applied is dismissed and +owes its one-line why. Daniel's reasoning, recorded verbatim as the why: + +> Finding 12 was written against the pre-§2.0 draft, where the form sat on a path past a gate +> and a forged handle would have covered unreviewed work; under the applied scope the record +> authorizes nothing and labels its handle unverified, so false attribution misstates who +> accepted a non-gate shortfall — **authorship metadata, not a trust boundary**. Axis values +> track the definition, not lens-set cheapness; **an axis that fires on any name field stops +> discriminating.** + +Security stays `none`; mode stays `battery+check`; passes 2–3 run with no lens sets. The +profile log gains no entry, because no value moved. + +### The finding as written, and the analysis it was dismissed against + + +Pass 1 argues security should return to `standard`: the form names *who decided* via a handle, +its admitted central gap is that the handle is unverified, and dropping the roles / +trust-boundaries lens on precisely the change whose risk is false attribution is +substantively inconsistent even though the syntax and derived mode are valid. + +**Against it, and this is why it is closer than pass 1 assumed:** the finding was written +against the **pre-§2.0 draft**, where the record sat on the path past a gate — there, a forged +handle would have covered an unreviewed change. Under the applied scope it authorizes nothing, +satisfies no evidence obligation, and the shipped text labels itself unverified. A false +attribution now misstates who accepted a *non-gate* shortfall. That is a real but small thing. + +**For it:** "touches no role" is not quite true of a form whose second field is a person's +identity, and lens sets are cheap. + +Surfaced rather than decided, because §5 makes a profile change proposed, human-confirmed and +logged, and this would have reversed an axis set the same day. Answered above. + +## Status + +Not clean; floor not met (1 of 3). **14 of 15 applied, 1 dismissed with its reason.** +Proceeding to pass 2. diff --git a/.context/codex-reviews/gate-a-spec-pass-1-dispositions.stopped-2tier-debt.md b/.context/codex-reviews/gate-a-spec-pass-1-dispositions.stopped-2tier-debt.md new file mode 100644 index 0000000..c6c1f56 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-1-dispositions.stopped-2tier-debt.md @@ -0,0 +1,103 @@ +# Gate A — spec — pass 1 dispositions (two-tier cycle) + +30 findings (7 BLOCKER, 21 MAJOR, 2 MINOR). None dismissed. Enum held a fourth time. +Prior three-tier cycle archived under `.stopped-3tier`. + +## Shape + +Total findings fell 40 → 30 and changed character: the stopped cycle's blockers said the +mechanism could not exist; these say a coherent mechanism has weak controls. That is a +reviewable design with holes, not an unbuildable one. + +## Blockers — my calls + +**1. Tier 3 after an adverse pass. ACCEPT, and it is the sharpest finding of the cycle.** +Nothing stops tier 3 being entered *after* a pass returned unresolved Blocker/Major, and the +marker then says "no qualifying passes" and carries none of those findings. A genuine outage +after a bad pass becomes a sanctioned way to erase review evidence and land the exact +defects the gate found. **Fix, taken:** tier 3 is prohibited once any pass in the cycle has +returned Blocker/Major that is not remediated or individually human-dispositioned in the +waiver record. Every completed pass is preserved and named in the marker. + +**19 rides with it. ACCEPT.** §5's own "clearly stuck → STOP and surface" must be stated as +*never* remediable by tier 3, or the rule that stopped the previous design becomes a waiver +path. Ironic and correct. + +**2. The STOP-as-control argument. ACCEPT — it is a rationalization and it is mine.** +§4 rests the whole no-sentinel case on a STOP that may not fire: a Gate-A closing commit is +docs-only and exempt, Gate B can read satisfied from previously counted calls, and the hook's +output goes to the **agent**, not to the authorizing human. **Fix, taken:** §4 keeps the "no +sentinel, no invariant amendment" conclusion — which stands on its own, since nothing needs +silencing — but drops the STOP from the justification entirely and states that the unchanged +hook supplies **no** waiver control. + +**3. Debt closable by logged decision. ACCEPT.** "Not in the cycle that created it" is a +one-cycle delay, not a separation. **Fix, taken:** cancellation requires a **different** +accountable handle from the one that authorized the waiver, plus an explicit stated risk +acceptance. Separation of duties, not a cooling-off. + +**4. "Checked" with no due event. ACCEPT.** **Fix, taken:** an open, unreconciled or +unreviewed debt row **blocks a subsequent tier-3 waiver in that repository**. That is a real +consequence using machinery that exists, and it stops waivers compounding silently. + +**5. Human authority is forgeable text. ACCEPT.** GOES TO DANIEL — it decides what the +feature claims, not how it is worded. + +**9. Cycle-id and debt-row schema undefined. ACCEPT.** **Fix, taken:** canonical cycle-id +format, debt-row grammar with a stable key and status field, and creation/update algorithms. +11, 10 and 30 fold in here — multi-story markers carry a list of row keys, an uncited +artifact makes a resolvable story citation a tier-3 precondition, and the repayment profile +is the one at closure, referenced by an immutable story-commit identity rather than a copied +value. + +**18. The inventory is still thematic while claiming to be paragraph-by-paragraph. ACCEPT, +and it is an overclaim of exactly the class this repo regenerates.** I asserted the method +rather than performing it. **Fix, taken:** rebuild §8 from §5 in document order, uniquely +identifying each atomic condition, with the twelve-plus omissions named (stop-and-surface, +do-not-manufacture-findings, validate-before-applying, `workingDirectory` binding, branch-file +acceptance and resume deletion, hook result-envelope limits, no-story and malformed-profile +branches, per-story aggregation, counterfactual evidence, human-confirmed profile changes, +`baseSha`/WIP mechanics). + +## Story defects — accepted, and my amendment was incomplete + +**21. ACCEPT.** The story still says tier 3 "exists today for the init-time case" while the +spec says calling that existing practice would be false. The authoritative artifact and the +design disagree about whether this is an extension or a new waiver. + +**22. ACCEPT, and it falsifies my own accounting.** The story's §1, §5 and §6 still frame the +problem as needing vocabulary for a *weaker review*, still carry five tier-2 open questions, +and still size the work around what a tier-2 pass means. My amendment claimed tier 2 moved +"in its entirety". It did not. Every one of those must be marked moved with its destination, +and §1/§6 restated around the zero-pass exception. + +## Major — accepted without further comment + +6 (role definitions: author / implementer / gate operator / waiver approver / debt accepter, +including the solo-user case), 7 (outage confirmation sources, sanitized, with a validity +window and revalidation before closure), 12 (Gate-A debt records artifact path + immutable +blob identity, since `baseSha` does not identify text passed to `exec`), 13 (original + +remediation range mapped to the tool's single-range interface), 14 (ordinary-merge identity +and base selection), 15 (corrective commit reproduces the canonical blocks, not just names +the loss), 16 (supported merge strategy becomes a tier-3 precondition, recorded), 17 (a +backward-compatible debt-servicing procedure that survives rollback), 20 (`docs/getting- +started.md` says Gate A is not skippable at any level — a missed product-claim site, plus +`coding-workflow.md`'s pipeline overview), 23 (both blocks mandatory, ordered and linked, each +validated independently), 24 (the counterfactual is tautological — rebuild it around a +load-bearing claim), 25 (adversarial negative cases: forged authorization, self-authorization, +stale outage evidence, secret-bearing reasons, unresolved findings, immediate write-off), 26 +(`AGENTS.md` omitted from prompt-conformance scope though §9 changes it — invariant 11), 27 +(I described the version bump as *already verified* by a checker that cannot run before the +commit exists: an unverified-enforcement claim), 29 (companion lifecycle for enum drift: +write order, grammar, retention, behaviour if it disappears). + +## Minor — collected + +8 (tie "sustained timeout" to the configured MCP timeout and the recovery attempt), 28 +(appending occurrence 3 *does* edit the `todos.md` item — "never edited" belongs to the +hardening ledger, not here). + +## Status + +Not clean. One question goes to Daniel (finding 5) because it decides what tier 3 claims to +be; every other fix is taken and does not need him. diff --git a/.context/codex-reviews/gate-a-spec-pass-1-dispositions.stopped-3tier.md b/.context/codex-reviews/gate-a-spec-pass-1-dispositions.stopped-3tier.md new file mode 100644 index 0000000..29c0065 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-1-dispositions.stopped-3tier.md @@ -0,0 +1,113 @@ +# Gate A — spec — pass 1 dispositions + +Cycle: reviewer-availability fallback ladder (story 2026-08-13, spec 2026-08-14). +37 findings returned; 4 BLOCKER + 28 MAJOR actioned, 5 MINOR collected. +Severity enum held on first live use of rider (b) — no out-of-enum token. + +## Blocker + +1. ACCEPT — correct and the most serious. Invariant 3's content-derived validity is + delivered by the hook's fingerprint, which is only written on `mcp__codex__review` + calls. Tier 2 makes none, so no fingerprint exists to invalidate; silencing is not + even required for the hole to open. Needs a content-identity record + pre-close + equality check. NEEDS DECISION (form of the record). +2. ACCEPT — real structural hole I introduced. The WIP body is a Gate-B artifact; both + Gate-A cycles complete before any WIP commit exists, so a Gate-A tier-2 pass has no + durable record at all. Candidate fix: the spec/plan docs commit body is the Gate-A + marker's home, carried forward into the implementation cycle's closing body. +3. ACCEPT — already-initialized projects keep their copied CLAUDE.md with the old + prohibition, so 0.9.0 ships them no ladder. Overlaps the parked "§5 version stamp" + row, which exists because this is a design of its own. NEEDS DECISION (solve here vs + scope + cross-reference). +4. ACCEPT — rider (a) introduces this: two calls, no bound snapshot, both files + structurally valid, potentially different diffs, counted as one pass. A provenance + false ✓ in the replacement for the mechanism that had one. + +## Major — accepted, mechanical or wording fixes + +18. ACCEPT (wording) — "closed set" vs accepted out-of-enum tokens is a genuine + contradiction in how I wrote it, not a reversal of the decided behaviour. Fix by + separating canonical syntax from a tolerant normalization rule. +19. ACCEPT — §9 requires the mapping be recorded; §12 says the file may not exist. The + design both requires and disclaims its own audit evidence. Make dispositions + conditionally mandatory when normalization occurs. +33. ACCEPT (MINOR, fixed anyway) — §1 says "three clean passes are required"; §5 requires + min 3 passes with only the FINAL one clean, plus a zero-finding early exit. A factual + misstatement of the procedure this design promises to preserve. +34. ACCEPT (MINOR, fixed anyway) — "Tier 1 … unchanged in every respect" is contradicted + by riders (a) and (b) in the same document. Narrow to reviewer independence. + +## Major — accepted, structural + +2, 6, 10, 22 are the four that change the shape of the design rather than its wording. + +6. ACCEPT — and it contradicts my own §1. §1 says §2.13 does not reach the mid-flight + case; §3 then lists tier 3 as "the documented human exception, unchanged". For + mid-flight, tier 3 as described does not exist. NEEDS DECISION (define vs remove). +7. ACCEPT — no entry/exit predicate for "unavailable"/"temporary". Without one the + fallback is available whenever review is inconvenient, which is the standing + arrangement the narrowing forbids. +8. ACCEPT — the safety argument rests on "human-attended" and nothing requires a human. + The reviewed actor can silence the reminders and authorize its own weaker reviewer. +9. ACCEPT — no canonical marker schema, so "carried by name" cannot be validated and two + readers produce different histories that both satisfy the prose. +10. ACCEPT — the story disclosed per-pass AND per-cycle; the design dropped the per-pass + surface without accounting. Exactly the decision-procedure Don't. NEEDS DECISION + (define a per-pass record vs amend the story with accounting). +11. ACCEPT — a Gate-B-shaped post-merge review cannot repay a missed Gate-A design review. +12. ACCEPT — "merge-base range" is not a durable identity after squash or later merges. +13. ACCEPT — the debt row must not close on a result that still carries Blocker/Major. +14. ACCEPT — "no new waiver mechanism" is an enforcement overclaim; the profile override + vocabulary is not defined for this use. AGENTS.md overclaim Don't. +15. ACCEPT — my singular max rule contradicts §5's explicit multi-story rule that each + profiled story satisfies its own obligations with no winning max. +16. ACCEPT — `full` stays supported with its demonstrated two-writer loss, so the rider + closes the backlog row without closing the defect. NEEDS DECISION (prohibit vs keep + row open). +17. ACCEPT — pass pairing and the shared single recovery attempt are undefined once one + pass is two calls. +20. ACCEPT — and this one is an evidence defect, not a doc defect. The proposed check + passes or fails identically before and after the prompt change, so it cannot + discriminate. Fails CLAUDE.md's own "confirm the wiring could have produced the + failing observation" rule. +21. ACCEPT — the verification proves hook-invisibility, not that a tier-2 pass is + identifiable from main. It can pass while disclosure is absent or stale. +22. ACCEPT — silencing every reminder is the dangerous direction under invariant 2 while + the design claims the invariant is untouched. NEEDS DECISION (scope the invariant + with rationale vs retain one visible degraded reminder). +23. ACCEPT — touching a sentinel file states no reason; no content schema is given. +24. ACCEPT — no procedure for availability returning mid-cycle, mixed-tier pass sets, or + which tier the closed cycle is finally recorded as. +25. ACCEPT — the gate prompts are Codex-facing and open with a Codex-directed skill + instruction; tier 2 executes on Claude. Prompt-standard item 1. +26. ACCEPT — "fresh" and "same-family" are unverifiable as written; no named invocation, + no context-isolation guarantee, no behaviour when the configured subagent differs. +27. ACCEPT (security lens) — a reviewer reading repo content is never told to treat that + content as data rather than instructions. Prompt injection steers the weaker reviewer. +28. ACCEPT (security lens, and the sharpest of them) — a "fresh" reviewer can read + `.context/` and inherit prior passes' conclusions through the filesystem, collapsing + the 3-pass loop into one answer repeated. +29. ACCEPT — diff-as-prompt-text has no truncation detection; `mcp__codex__review` reads + the range itself and cannot silently truncate. A clean findings file can certify a + lossy diff. +30. ACCEPT — §5's several-WIP `git reset --soft` collapse path is omitted from my + carry-forward chain, so markers in earlier WIP bodies are destroyed outside the named + hop. +31. ACCEPT — the obligation row must exist before the final pass and inside the reviewed + snapshot, or it lands unreviewed. +32. ACCEPT — a late tier-1 result can overwrite a valid tier-2 findings file, and every + shape check still passes. +5. ACCEPT — the file-first protocol requires the reviewer to WRITE; the design grants + only read. NEEDS DECISION (narrow write capability to the one path vs parent-mediated + write, which reintroduces the truncation risk file-first exists to prevent). + +## Minor — collected, not iterated + +35. Parked-row trigger has no threshold or owner. +36. No rollback procedure for an active degraded cycle or outstanding debts. +37. The "two tool names" are never actually named, and routing is version-dependent. + +## Note + +No finding was dismissed. Six carry NEEDS DECISION and are going to Daniel before pass 2, +because each leads to materially different work rather than a different sentence. diff --git a/.context/codex-reviews/gate-a-spec-pass-1-dispositions.stopped-tier3-core.md b/.context/codex-reviews/gate-a-spec-pass-1-dispositions.stopped-tier3-core.md new file mode 100644 index 0000000..464732d --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-1-dispositions.stopped-tier3-core.md @@ -0,0 +1,107 @@ +# Gate A — spec — pass 1 dispositions (stripped design) + +28 findings (4 BLOCKER, 20 MAJOR, 4 MINOR). None dismissed. Enum held a seventh time. + +## The strip worked + +| Cycle | Blockers by pass | +|---|---| +| Three-tier | 4 → 4 → 6, stopped | +| Two-tier with debt machinery | 7 → 8 → 12, stopped | +| **Stripped** | **4** | + +More important than the count: **none of the four says the mechanism cannot exist.** Every +prior cycle's blockers were "this control is unenforceable" or "this state needs the gate +that is unavailable". These four are ordinary design holes with named fixes. + +## Blockers — all accepted, fixes taken + +**3. The outage calls are not required to be canonical gate calls.** `attempt-failed` also +admits a failure that happened *after* a review actually ran — an output-write or validation +failure. So unavailability can be manufactured: send a malformed request, point at the wrong +tool, or induce a findings-write failure, and then waive a reviewer that was perfectly able +to review. This contradicts the stated "cannot run at all" boundary. **Fix:** the initial +call, the recovery and the revalidation must each be a **canonical gate call** — correct +tool, correct working directory, the current artifact or range, the cited story paths and +evidence entries — and a request-shape, validation or output-protocol failure is explicitly +**not** proof of unavailability. The marker records which canonical call class failed. + +**6. A Gate-A waiver does not require a docs-only closing tree.** Nothing stops a staged +prompt, script or code change riding inside the Gate-A closing commit under an A-spec marker, +with §4's precedence rule excusing the Gate-B STOP it would raise. That is a second gate-off +path in the dangerous direction. **Fix:** a Gate-A waiver requires a **positively determined** +docs/artifact-only tree containing the named artifact and allowed supporting prose, records +the complete prospective tree id, and refuses any mixed or product path outright — no +separately authorized Gate-B waiver riding along inside it. + +**7. Nothing atomically compares the authorized tree with what `git commit` writes.** "If the +tree changes, re-authorize" is a rule with no observation behind it, and another process can +mutate the index between check and commit. **Fix:** commit from a dedicated fixed index, +compare the committed tree id to the authorized tree id immediately afterward, and define the +mismatch path as an incident requiring re-authorization. + +**9. Gate-A authorization is consumed too early to survive the reminder it must survive.** The +close happens at the docs commit, but the hook emits the Gate-A below-floor reminder later, at +`executing-plans` — and §4 grants precedence at *one* closing transition while telling later or +resumed sessions to obey the hook. A correctly waived plan therefore cannot proceed to +implementation. **Fix:** define the continuation rule explicitly — how the next skill +invocation discovers and validates the immediately preceding A-plan marker across a resume — +rather than pretending the authorization is spent where the reminder is not. + +## Majors that change the document's claims + +**1. §2's structural argument is a false dichotomy, and it is mine.** It presents +unenforceable prose and recursively gated state as the only options, when a CI rule, an +append-only check or a protected external record could preserve the record without asking the +unavailable reviewer to review each transition. Removing the debt machinery was a sound +**cost and trust decision**; calling it structurally impossible overclaims, and it tells +future readers not to evaluate feasible alternatives. **Fix:** recast §2 as the trade-off it +is, name the alternatives considered and why each was rejected, and keep §12's trigger. + +**23. The `+check` harness does not exist as an artifact.** Withdrawing the mode override +rested on a "prompt-harness scenario", but the repository has no such harness and the design +names no driver, fixture, oracle, command or repeatability criterion. Calling an unspecified +future scenario a check is the unverified-evidence claim the withdrawal was meant to fix. +**Fix:** specify it concretely — frozen old/new prompt inputs, the exact scenario, the +observable assertion per precondition branch, and how a nondeterministic agent result is +handled — with rider (b)'s automated `IMPORTANT`-fixture check carrying the mechanical half. +If that specification cannot be made concrete, the evidence gap returns and needs a valid +**whole-mode** human-confirmed override, not the withdrawn per-portion one. + +**14. A surviving dependency on the removed machinery, in the story.** Story §5 still says the +profile "scales the repayment" while the design now defines no repayment at all. The +withdrawal is incomplete until that is amended. + +## Majors — accepted without further comment + +2 (Gate-A evidence is close to vacuous — needs a non-vacuous checklist and durable result, or +must be labelled syntax hygiene rather than compensating evidence), 4 (no transition table from +a fresh observation to an enum value), 5 (revalidation happens before an unbounded human pause; +needs a max age and a final probe **after** the answer), 8 (authorization binds to the tree but +the marker, evidence entry and decision live in the commit *message*, which is not in a tree), +10 (no fail-closed rule when a prior findings file is missing or unreadable — and this cycle +has already lost artifacts to slot collision), 11 (the decision block has no `Gate:` field but +is keyed by gate), 12 (no grammar, escaping, length bound or timezone for handles and reasons +interpolated into a line-oriented protocol), 13 (ordinary merge is permitted but only the squash +chain is specified), 15/16/17/18/19 (five more inventory dispositions wrong **by effect** — +recovery-by-source, Gate-A's two loops, the lens sets having no prompt to attach to at tier 3, +profile resolution not required at Gate A, and the skip rule marked untouched on the same +noun-based reasoning row 49 correctly rejects), 20 (`docs/pr-review-bots.md` asserts every PR +head has passed Gate B — a site I missed), 21 (`docs/coding-workflow.md`'s "What is essential" +section names two independent gates as load-bearing — also missed), 25 (rollback ignores +downstream scaffolded copies and version-keyed plugin caches), 28 (Gate B's precondition says +"the profile's mode" singular where §5 requires per-story modes, suffixes and lens unions). + +## Minors — collected + +22 (marketplace.json does **not** actually carry the two-gates claim — my site row asserts a +premise the file does not support, and the mechanical sweep should have caught it), 24 (§5.1 +cross-references §11 for validation; validation is §10), 26 ("a line on main that every reader +sees" overstates commit-body visibility), 27 ("later sessions treat the hook as authoritative" +still implies authority the hook lacks). + +## Status + +Not clean; floor not met. No new decisions needed — every fix above is takeable without +Daniel, including 23's, which is a specification task inside his standing "+check stands" +instruction. diff --git a/.context/codex-reviews/gate-a-spec-pass-1.stopped-2tier-debt.md b/.context/codex-reviews/gate-a-spec-pass-1.stopped-2tier-debt.md new file mode 100644 index 0000000..681a6b6 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-1.stopped-2tier-debt.md @@ -0,0 +1,31 @@ +BLOCKER | high | §3 lines 49-60 and §7 lines 191-198 | [risk: threats] Tier 3 can be entered after one or more qualifying passes, including a pass that exposed unresolved Blocker or Major findings, yet the effect and marker say the cycle had no qualifying passes and carry none of those findings | A genuine outage after an adverse pass becomes a sanctioned way to erase known review evidence and land the exact defects the gate found | Define the partial-cycle state explicitly; either prohibit tier 3 once any pass has returned findings, or preserve every completed pass and require all known Blocker and Major findings to be remediated or individually human-dispositioned in the waiver record +BLOCKER | high | §4 lines 73-85 and §11 line 304 | [risk + security: abuse and abuse paths] The claimed Gate-B STOP is not guaranteed: a Gate-A docs-only closing commit is exempt, and Gate B can already show a satisfied count when failed or incomplete calls were counted and the current fingerprint matches; moreover the hook exits 0 and its output is seen by the agent, not necessarily by the authorizing human | The argument that knowingly passing one STOP is the compensating control rests on a signal that may be absent, may be a false checkmark, and cannot establish human presence or authorization; this is a rationalization rather than a control | Remove the STOP from the justification and state that the unchanged hook supplies no waiver control, then add a closure-time human confirmation and a durable validation mechanism that do not depend on hook count or docs-only classification +BLOCKER | high | §7 lines 191-198 | [risk: observability] The allegedly mandatory cross-model re-review debt can be closed without any review by a logged human decision | The sole compensating review control can be erased one cycle later with the same unverified prose mechanism that created it, contradicting the settled decision that every tier-3 closure owes a re-review and making repayment optional in substance | Remove the no-review closing condition, or rename the control as optional and justify that weaker product; if exceptional debt cancellation must exist, require a separately governed approval with explicit risk acceptance that cannot occur merely by starting another cycle +BLOCKER | high | §7 lines 200-201 and §13 lines 339-341 | [risk: observability] Open debt is only “checked,” with no defined result, due event, deadline, escalation, or operation that an open row blocks, and the spec admits nothing fails if it is never reconciled or closed | A debt that can remain open forever does not compensate for a zero-review landing and provides no reliable signal that availability returned | Define the exact check, observable states, owner, due trigger, and stop condition; at minimum block a subsequent tier-3 waiver and a release or merge milestone while any due debt is unreconciled or unreviewed +BLOCKER | high | §3 lines 54-60, §5 lines 89-115, and §13 lines 336-338 | [security: trust boundaries] Human authorization is represented only by text that the same agent can write, and the spec explicitly acknowledges that nothing binds it to an actual human interaction | The gate-waiver authority can be forged at the boundary it is supposed to protect, so both disclosure and authorization are assertions by the actor being constrained | Require a direct pause for an explicit human response immediately before closure and define a verifiable approval channel or signed/protected record; if no binding mechanism is in scope, state that tier 3 provides disclosure only and does not enforce human authority +MAJOR | high | §3 lines 54-56 | [security: roles] “The actor whose work is under review” is undefined for human-authored prompts, mixed human-agent work, delegated agents, and teams, while only “the human” may enter tier 3 | A solo human who authored the change may have no eligible approver, while an agent can treat any nearby human as separate authority; inconsistent role interpretation changes whether the waiver is allowed | Define author, implementer, gate operator, waiver approver, and debt accepter, including the solo-user and multi-agent cases, and state the required separation for each +MAJOR | high | §3 lines 49-53 | [security: external systems] “Independently confirmed” has no trusted source, evidence, scope, timestamp, or expiry, and it waives even the required recovery attempt | An agent can cite a stale status page, another account’s quota, or an unauthenticated assertion to bypass the only live availability check; availability may also return before closure | Enumerate acceptable confirmation sources per cause, record source and time in sanitized form, set a short validity window, and require revalidation immediately before the closing commit +MINOR | high | §3 lines 49-53 | [risk: compatibility] “Not slow” is excluded while “sustained timeout” is allowed, but sustained has no duration beyond the same timeout and retry that a slow reviewer naturally triggers | Operators will classify the same latency differently and may use tier 3 for mere inconvenience despite the stated boundary | Tie sustained timeout to the configured MCP timeout, the one shared recovery attempt, and an explicit elapsed-time or repeated-failure rule +BLOCKER | high | §5 lines 91-119 and §7 lines 167-175 | [risk: idempotency] Neither the cycle-id format and uniqueness scope nor the todos.md debt-row schema, states, and key syntax are defined, although deduplication, conflict detection, reconciliation, and idempotency all depend on them | Implementers cannot create or update the promised durable debt deterministically, and collisions can merge unrelated waivers or overwrite the wrong debt | Specify a collision-resistant cycle identifier, canonical debt-row block with stable key and status fields, creation and update algorithms, duplicate handling, and exact validation rules +MAJOR | high | §7 lines 170-175 and 203-205 | [risk: empty path] Tier 3 is said to work at any profile, but §5 permits artifacts with no cited story while every debt identity requires a story path and the design defines one row per cited story | A no-story or unprofiled gate cycle has no row to create, so “every tier-3 closure owes one” is unsatisfiable on an existing supported path | Define a debt identity for uncited artifacts or make a resolvable story citation an explicit tier-3 precondition with migration guidance for in-flight no-story work +MAJOR | high | §5 lines 94-99 and §7 lines 203-205 | [risk: idempotency] One marker has one todos.md row key, but a multi-story closure requires one row per cited story; repeating the marker with the same gate and cycle id but different row keys is declared a conflict | The schema cannot represent the multi-story debt state the procedure requires, so some story debts will be unlinked or the close will stop permanently | Make the marker carry a canonical list of all row keys, or define one marker per story with a deduplication key that includes the story path +MAJOR | high | §7 lines 163-189 | [risk: data loss] A waived Gate-A debt records story path, cycle id, and baseSha but not the reviewed artifact path, commit or blob identity, and baseSha does not identify text passed to exec | A later repayment can review the current or wrong spec or plan and still claim to discharge the historical debt, violating the gate-matched promise | Record artifact path plus immutable blob or commit identity at closure and require the repayment worktree to materialize and verify that exact artifact before each pass +MAJOR | high | §7 lines 167-189 | [risk: compatibility] “The original range plus the named remediation range” is not mapped to the review tool’s single baseSha-to-head interface and may be non-contiguous with unrelated commits between the two ranges | A repayment may omit remediation, include unrelated work, or use separate calls whose pass count no longer represents one reviewed object | Define the exact commit graph and tool calls for contiguous and non-contiguous remediation, including head SHA, exclusions, how branch findings are combined, and what fingerprint or range each clean pass certifies +MAJOR | high | §6 lines 142-150 and §7 lines 167-175 | [risk: compatibility] Post-land reconciliation defines squash identity only; for an ordinary merge it does not say whether “landed commit” means the original closing commit or the merge commit, nor which parent is the base | Debt rows can point at different diffs for the same supported merge strategy and later re-review a merge aggregate rather than the waived artifact | Define identity and base selection separately for squash and ordinary merge, with a graph example and validation for each +MAJOR | high | §6 lines 142-146 | [risk: data loss] After a squash drops markers or evidence entries, the remedy is only a corrective commit “naming what was lost,” not a requirement to restore the complete canonical marker, human decision, debt links, and evidence entries | Naming a loss does not recreate the durable record required by story AC 2 or the profile evidence required on main | Require the corrective commit to reproduce every missing canonical block verbatim or reconstruct it from an authoritative source, and verify the restored main history against the pre-merge inventory +MAJOR | high | §6 lines 142-150 | [risk: compatibility] Rebase-merge and cherry-pick are stop conditions only after the cycle has already been waived and committed, with no pre-authorization check that the repository or merger will honor a supported strategy | A later merge choice can strand a completed waiver with unreconcilable identity, at which point “stop” cannot undo the zero-review closure | Make supported merge strategy a tier-3 precondition, record the selected strategy and responsible merger in the authorization, and define recovery if the hosting platform applies a different strategy +MAJOR | high | §13 lines 346-349 | [risk: rollback] Rolling back the prompt change removes the only procedure that can service already-created debt rows, and the proposed answer is to repay all debt before rollback even when rollback may be urgent | A security or operational rollback can either remain blocked on an unavailable reviewer or orphan obligations permanently | Preserve a backward-compatible debt servicing procedure or provide a migration and emergency rollback protocol that remains usable after the ladder text is reverted +BLOCKER | high | §8 lines 207-244 | The old-condition inventory is thematic rather than paragraph-by-paragraph and omits many operative §5 conditions, including stop-and-surface when stuck, do not manufacture findings, validate before applying, workingDirectory binding, exact branch-file acceptance and resume deletion rules, hook result-envelope limitations, profile no-story and malformed-profile branches, per-story aggregation, counterfactual evidence, human-confirmed profile changes, and exact baseSha/WIP mechanics | This directly violates the AGENTS.md decision-procedure invariant and leaves no proof that the replacement prompt will retain the omitted branches; several omissions interact with the new waiver path | Rebuild §8 from every §5 paragraph and sub-bullet in order, quote or uniquely identify each atomic condition, and mark each kept, moved, deliberately dropped, overturned, or inapplicable with the destination text +MAJOR | high | §3 and §8 | The existing “clearly stuck means STOP and surface” condition is neither inventoried nor bounded against tier 3, so the new availability exception can be read as a way to close a cycle after rising or disputed findings once the reviewer also becomes unavailable | The exact stop-and-surface rule that terminated the prior design can silently become a waiver path, defeating the old-condition accounting and the gate’s adverse-finding control | State that tier 3 is never a remedy for a stuck review or unresolved known findings and account for the stuck condition explicitly in §8 +MAJOR | high | §9 lines 246-263 | The product-claim inventory misses docs/getting-started.md, which says Gate A is not skippable at any level and prescribes unconditional three-pass Gate-A loops, and it treats only one section of docs/coding-workflow.md even though that file says every story follows the gated pipeline | After implementation, shipped guidance will still contradict the tier-3 waiver, so the deliberate overturn is not propagated to every assertion as claimed | Search by the semantic claim across all non-archive docs and prompt mirrors, add every hit including getting-started and the pipeline overview, and specify the exact sentence-level disposition for each +MAJOR | high | Story §2 lines 39-45 versus spec §3 lines 43-63 | The amended story says tier 3 already exists for the init-time case, while the spec correctly says calling that existing practice would be false and that init-time gatelessness is not a tier | The authoritative story and design disagree about whether this is an extension or a new gate waiver, undermining scope, rationale, and acceptance review | Amend the story to say tier 3 is new and distinct from §2.13, with old-condition accounting for the removed “exists today” claim +MAJOR | high | Story §1, §5, and §6 | The amended story still frames the unsolved need as vocabulary for a weaker review, leaves five tier-2 open questions active, and sizes the story around what a tier-2 pass means | Current source-of-truth text still assumes the deliberately split tier and makes the claim that it moved “in its entirety” untrue | Mark each tier-2 problem statement, open question, and sizing clause as moved with its destination, then restate the parent story solely around the zero-pass tier-3 exception +MAJOR | high | §3 lines 57-60 and §5 lines 91-115 | The closure record is ambiguous: §3 requires the logged-human-decision block, §5 defines a separate tier-3 marker, §6 carries only markers, and §11 validates only marker presence | An implementation can include either block and plausibly claim conformance, losing either authorization rationale or machine-recognizable disclosure | State that both blocks are mandatory, define their ordering and linkage, carry and deduplicate both, and validate omission of each independently +MAJOR | high | §11 lines 289-292 | The “check that fails without the change” merely asks an actor to follow old text that contains no marker and new text that does; it cannot fail for a broken authorization, carry, debt, or repayment mechanism | This is a tautological prompt-difference demonstration, not a counterfactual for the high-risk behavior the design claims to control | Choose a load-bearing claim such as preventing closure without direct human authorization or preventing debt discharge without full re-review, build a check whose prior-state observation is unsafe, and show the wiring can expose the defect +MAJOR | high | §11 lines 294-310 | [security: assets and trust boundaries] The verification matrix has no negative cases for forged authorization, self-authorization, stale or wrong-account outage evidence, secret-bearing decision reasons, unresolved prior findings, or immediate debt write-off | The validation plan exercises record plumbing but not the abuse paths that make a human gate waiver dangerous, so a cosmetically correct implementation can ship with the trust boundary open | Add named adversarial scenarios for each authority and outage boundary, assert rejection before commit, and verify commit bodies and context artifacts contain only sanitized allowed fields +MAJOR | high | §11 lines 283-287 and §9 lines 252-259 | Prompt-conformance scope lists §5, its inline mirror, and §2.13 but omits AGENTS.md even though §9 changes it and CLAUDE.md classifies AGENTS.md as prompt product requiring full Gate B | The only manual quality gate for the changed invariant prompt is silently skipped, violating invariant 11 | Include AGENTS.md and every other changed prompt artifact in the all-12-items review and name the reviewer and recorded result +MAJOR | high | §12 lines 326-328 | The unimplemented, uncommitted version bump is described as already verified by check-version-bump.sh, although AGENTS.md says that checker is useless before plugin changes and the bump are committed | This is an unverified enforcement claim and can make invariant 12 appear satisfied before the checker has any relevant commit range to inspect | Put the evidence in future tense and require running the checker against the current PR base after the WIP commit contains both plugin edits and the manifest bump +MINOR | high | §12 lines 314-318 | The compound-command occurrence is said to be appended to an existing todos.md row and simultaneously “never edited,” although appending the occurrence necessarily edits that mutable checklist item | The contradictory mutation instruction invites either a duplicate todo or failure to add the occurrence | Say explicitly that the existing todos.md item is edited in place to append occurrence 3; reserve append-only language for docs/hardening-log.md +MAJOR | high | §10 lines 267-277 and §8 line 240 | The dispositions companion becomes required for enum-drift validity, but the retained §5 lifecycle says companions are advisory, may be deleted or rebuilt, and do not participate in pass validation; the spec does not define when the mapping is written, how it is validated, or how long it must survive | A pass can be accepted under one reading and incomplete under another, and later deletion removes the only evidence that an unknown severity was conservatively handled | Amend the companion lifecycle and acceptance algorithm explicitly for enum drift, including atomic write order, exact record grammar, retention through cycle close, and behavior if the record disappears +MAJOR | high | §7 lines 156-179 | “Profile scales the repayment” does not say whether repayment uses the profile at waiver time or the mutable current story header, while the debt row intentionally records no profile value | A later profile downgrade can silently weaken repayment, while using the historical profile violates the single-writable-copy model unless an immutable reference is recorded | Choose and justify one temporal rule; if current profile governs, forbid debt-motivated downgrades or require separate approval, and if closure-time profile governs, record an immutable header or story commit identity rather than a copied value +END OF FINDINGS (30 total) diff --git a/.context/codex-reviews/gate-a-spec-pass-1.stopped-3tier.md b/.context/codex-reviews/gate-a-spec-pass-1.stopped-3tier.md new file mode 100644 index 0000000..332d887 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-1.stopped-3tier.md @@ -0,0 +1,38 @@ +BLOCKER | high | §3 "Tier 2" and §4 "What the hook does" | a tier-2 Gate-B pass has no content-derived validity after the hook is silenced: neither the findings artifact nor the tier marker binds the reviewed base, HEAD, effective index, or included worktree, so edits can land after the final pass without invalidating it | this directly opens the stale-review false ✓ that invariant 3 exists to prevent | require a content-derived snapshot record and a pre-close equality check covering the same index and worktree axes as invariant 3, or change the hook so tier-2 passes receive honest fingerprint invalidation +BLOCKER | high | §6 "Disclosure and carry-forward" | the only in-cycle tier record is a WIP commit body, but both Gate-A cycles happen before Gate B's WIP snapshot exists and the spec names no durable bridge from Gate-A spec/plan review through planning and implementation to that later commit | a Gate-A tier-2 pass can be lost on handoff and reach main indistinguishable from tier 1, violating the story's central acceptance criterion | define a durable Gate-A record available at pass time, its exact carry into the later WIP and closing commit, and its behavior when work stops before Gate B +BLOCKER | high | §8 "The narrowing, and the sites it touches" and §10 "Backlog and packaging" | the change updates the inline scaffold for future initialization but specifies no upgrade path for already-initialized projects whose copied CLAUDE.md contains the old prohibition | the exact configured projects that can lose Codex mid-flight will install 0.9.0 yet still have no sanctioned ladder, so the motivating incident remains unresolved for the target population | specify how existing projects detect and merge the template update, including rerunning workflow-init, conflict handling under invariant 9, and an observable completion check +BLOCKER | high | §9(a) "Sequential single-branch gate calls" | separate spec and quality calls are not bound to one immutable Gate-B snapshot, so HEAD, the index, or the worktree can change between them while both files remain structurally valid | the two branches can approve different diffs and be counted as one clean pass, recreating a provenance false ✓ in the replacement for full review | pin and record one base/head/content identity for the pair, forbid mutation between branches, and discard both files and rerun the pair if any identity changes +MAJOR | high | §3 "Tier 2" | repository read access is specified but the unchanged file-first protocol requires the reviewer itself to create directories and write the findings file; no narrow write capability or parent-mediated lossless write path is defined | a least-privilege subagent cannot complete a valid pass, while granting general write access expands the reviewer role beyond the settled read-only boundary | define the exact tool/capability set, allowing only the named review-artifact path, or define a truncation-safe parent handoff that preserves the file-first guarantee +MAJOR | high | §3 "Tier 3" | tier 3 is only called "the documented human exception" and "unchanged" without naming its procedure, mid-flight reach, disclosure, or whether it satisfies either gate | story AC 1 requires every tier's standing, and the existing init-time gateless choice is not a pass that closes a mid-flight cycle | cite the exact existing authority and spell out tier 3's activation, effect, record, and closing rule for both gates, or remove it from this ladder +MAJOR | high | §5 "Degraded posture" and §8 "The narrowing" | "unavailable," "temporary," and "return to tier 1" have no decision procedure: no qualifying failures, prerequisite recovery attempt, cooldown, maximum duration, or proof that tier 1 was retried is named | an implementer can choose the weaker reviewer whenever review is inconvenient and keep the allegedly temporary fallback as a standing arrangement | define entry and exit predicates, including who confirms them, which diagnostic evidence is recorded, and when continued unavailability must be surfaced again +MAJOR | high | §5 "Degraded posture" | "operator" and "human-attended" are undefined roles and no human confirmation is required before the implementing agent creates the opt-out file | the same actor whose work is being reviewed can silently disable every reminder and authorize its own weaker reviewer, an obvious abuse path across the human/agent trust boundary | require an explicit human decision to enter and extend degraded mode, name where that decision is recorded, and distinguish actions the agent may perform from decisions only the human may make +MAJOR | high | §6 "Disclosure and carry-forward" | no exact tier-marker schema or worked example is specified, including gate, artifact, pass numbers, reviewer family/model, outage reason, and whether a mixed-tier cycle is classified by its weakest pass | "tier markers and evidence entries" cannot be validated or carried "by name" consistently, and two readers can produce histories that satisfy the prose while conveying different facts | define a canonical marker block and aggregation rule with examples for Gate-A spec, Gate-A plan, Gate B, mixed tiers, and multiple cited stories +MAJOR | high | Story §2 versus design §6 | the story says tier 2 is disclosed in the pass record and again in the closing body, while the design forbids it in findings and makes the dispositions marker optional, leaving only a cycle-level WIP body | a settled disclosure surface has been silently dropped, so per-pass provenance is absent even when the cycle marker survives | identify a required pass record outside the strict findings body and carry it forward, or explicitly amend the story with old-condition accounting +MAJOR | high | §7 "Re-review obligation" | a Gate-B-shaped post-merge review is the only repayment even when the degraded cycle was Gate-A spec or plan review | reviewing the landed diff does not recover the missing cross-model design review and may miss architectural alternatives that are no longer visible after implementation | make the debt match each degraded cycle's artifact and gate, or state and justify a deliberate non-equivalence with separate obligations +MAJOR | high | §7 "What the re-review is" | "over the merge-base range" names neither refs nor durable SHAs and becomes ambiguous after squash, branch deletion, later merges, or multiple landed changes | the reviewer may inspect the wrong range or an ever-growing range, and the todo cannot establish what debt it closes | record the landed commit and exact base/head or patch identity in the obligation and define range selection for squash, merge-commit, and rebased histories +MAJOR | high | §7 "What the re-review is" | the todo closes on "the review's recorded result" even when that result contains unresolved BLOCKER or MAJOR findings | serious post-merge defects can close the only repayment control merely by being observed | require disposition and remediation or an explicit human acceptance for every Blocker/Major before the debt row closes +MAJOR | high | §7 "How the row can otherwise close" | a profile mode override changes validation evidence, but the spec treats its vocabulary as an existing waiver mechanism for review debt without naming syntax, authority, location, or compatibility rules | "No new waiver mechanism" is an enforcement overclaim and permits inconsistent or invisible debt cancellation | either define a distinct recorded re-review waiver or cite the exact existing fields and rules that already authorize this use, including who confirms it +MAJOR | high | §7 "At effective level 2" | the singular max-level rule does not account for CLAUDE.md §5's explicit multi-story rule that each profiled story satisfies its own obligations and no winning max mode is used | a cycle citing mixed profiled and unprofiled stories has no defined number, ownership, or closure rule for re-review debts | specify per-story qualification and whether one review can discharge several named obligations without collapsing their independent records +MAJOR | high | §9(a) "Sequential single-branch gate calls" | the design says concurrency is eliminated but retains full calls as a supported branch with their unsafe two-writer behavior | making single-branch calls the default does not prevent the already-demonstrated data loss whenever full is selected, so the rider closes its backlog row without closing the defect | prohibit full calls until provenance is guarded, or keep the backlog open and state that the default only reduces exposure +MAJOR | high | §9(a) old-condition accounting | the meaning of one Gate-B pass and its shared recovery budget is undefined after one full call becomes two calls: it does not say how pass numbers pair, whether a successful branch survives failure of the other, or which retry consumes the one attempt | agents can combine branches from different attempts or accidentally receive two recovery budgets, defeating clean-final-pass and retry limits | define the pair state machine, shared attempt counter, deletion set, resume rules, and terminal behavior for every spec/quality success and failure ordering +MAJOR | high | §9(b) "severity enum" | the enum is called a closed set of permitted tokens while out-of-enum tokens are explicitly accepted as complete passes | this contradicts story AC 6 and makes "valid finding line" mean two incompatible things in the acceptance rule | distinguish canonical syntax from a tolerant normalization rule and define when the normalized finding becomes valid, or make out-of-enum files incomplete while preserving their content for diagnosis +MAJOR | high | §9(b) and §12 fifth residual | §9 requires the reader to record every enum mapping in a dispositions file and claims the drift stays visible, while §12 says that companion is optional, may not exist, and can leave the mapping unrecorded | the same design both requires and disclaims its audit evidence, violating prompt-standard items 2, 7, and 11 | make the dispositions file conditionally mandatory and validated when normalization occurs, or remove the recording and visibility claims and name the loss explicitly +MAJOR | high | §11 "Check that fails without the change" | manually amending a WIP message only proves Git replaces commit text; including the marker manually passes before and after the prompt change, while omitting it fails before and after | the proposed check cannot distinguish the new instruction from the old workflow and therefore does not satisfy the profile's fails-without-change obligation | run the old and new prompts through the same closing scenario and assert only the new procedure carries the canonical marker across every hop +MAJOR | high | §11 "Named verification of the risk path" | observing that the hook counter does not move verifies only invisibility to the hook, not the stated risk path of a tier-2 pass reading as tier 1 | the validation can pass while disclosure is absent, stale, or lost at amend/squash—the exact false ✓ the design claims to close | verify an end-to-end degraded cycle and assert the tier-2 identity from main history alone, with negative cases for omitted and stale markers +MAJOR | high | §5 "Degraded posture" versus AGENTS.md invariant 2 | the prescribed response to uncertainty is to silence all hook reminders, including fingerprint invalidation and commit warnings, without reconciling that normative path with the non-negotiable "on uncertainty, fire" rule | the design sanctions the dangerous missed-reminder direction while claiming the hook invariant remains untouched | explicitly amend or scope the invariant with rationale, or retain a visible degraded reminder that fires until equivalent manual state is proven +MAJOR | high | §5 "touches codex-gate.off with a stated reason" | touching a sentinel does not state a reason, and the spec defines neither file contents nor another reason record | operators cannot follow or audit the instruction consistently, and a later reader cannot distinguish an outage fallback from a stale or init-time opt-out | give the exact command/content schema and safe handling rules for the reason, or name the separate durable record that carries it +MAJOR | high | §5 "deletes it on return to tier 1" | no transition procedure covers availability returning between passes or branches, stale gate counters, an outstanding tier-2 marker, or a final tier-1 clean pass after weaker earlier passes | re-enabling can produce contradictory hook state and lets agents erase or retain disclosure arbitrarily | define all mid-cycle transition states and state which prior passes count, which files are discarded, when the off file is removed, and how the final cycle tier is derived +MAJOR | high | §3 "running the same gate prompt" | the existing prompts are written for Codex and invoke a Codex-facing skill/output protocol, but tier 2 executes on Claude without a target-model adaptation or confirmation that the skill and instructions are available in that fresh context | literal reuse conflicts with prompt-standard item 1 and can fail before review or produce model-inappropriate behavior | define a Claude-targeted wrapper that preserves the coverage contract, names the executing model/surface, and verifies required skills are loaded +MAJOR | high | §2 decision 2 and §3 "fresh-context same-family reviewer" | the exact subagent mechanism and settings that guarantee no inherited context are absent, as are model-family identification and behavior when the configured subagent is another family or inherits parent turns | "fresh" and "same-family" are unverifiable labels, so a non-fresh self-review can be disclosed as tier 2 and accepted | name the supported Claude Code invocation, context-fork setting, model selection, verification check, and diagnostic stop states for unsupported versions or configurations +MAJOR | medium | §3 "repository read access" | the trust boundary for repository content is unstated: a reviewer with broad reads is not told to treat specs, diffs, source, prior findings, and tool-shaped text as untrusted evidence rather than instructions | a contributor can place prompt injection in the reviewed artifact and steer the weaker same-family reviewer or its tool use | add explicit data-not-instructions handling, constrain tool use, and name what content may supply authoritative instructions +MAJOR | medium | §3 "repository read access" | the reviewer can read prior findings, dispositions, and resume notes under .context, so a nominally fresh pass can inherit earlier reviewers' conclusions through the filesystem | correlated passes undermine the fresh-context benefit and can turn the three-pass loop into repetition of one answer | exclude prior review artifacts from reviewer context except explicitly required evidence, or define a sanitized view and verify it per pass +MAJOR | high | §3 Gate B input | passing git diff text through the subagent prompt has no completeness limit, truncation detection, binary-file rule, or diagnostic for a range too large for context, unlike mcp__codex__review reading the range itself | a syntactically clean findings file can certify a silently truncated or lossy diff | define lossless input construction and size checks, require the reviewer to read the pinned range directly when possible, and make incomplete input an INCOMPLETE pass with cause-specific remediation +MAJOR | high | §6 carry-forward chain versus CLAUDE.md §5 Mechanics | the chain covers one WIP amended in place but omits the existing several-WIP branch that runs git reset --soft to the first parent and creates one new closing commit | tier markers distributed across multiple WIP bodies are destroyed outside the named amend hop, violating the old-condition accounting rule | add the multi-WIP collapse path and require collecting, deduplicating, and validating every disclosure and evidence entry before the replacement commit +MAJOR | high | §7 todo creation and §3 unchanged validation | the spec does not say when the required todos.md row is written relative to the final tier-2 Gate-B pass | adding or editing the row after review changes the diff and, with the hook off, can enter the closing commit unreviewed | require the canonical obligation row before the final pass and include it in the bound reviewed snapshot; any later edit must invalidate and rerun the pass +MAJOR | medium | §3 and §5 error paths | switching from a timed-out or backgrounded tier-1 call to tier 2 does not carry the existing requirement to stop or await the original writer before deleting or reusing its slot | a late Codex result can overwrite a valid tier-2 findings file with a different run that still passes every shape check | define tier-switch cleanup as part of the single recovery budget and require proof that no prior task can still write before the tier-2 call starts +MINOR | high | §1 "Three clean passes are required per gate" | current §5 requires a minimum of three passes with only the final pass clean and permits an early exit below three only on zero findings | the problem statement misstates the decision procedure this design promises to preserve and invites an implementation that demands three clean results | restate the exact floor, final-clean rule, and zero-finding early exit +MINOR | high | §3 "Tier 1 ... Unchanged in every respect" | riders (a) and (b) change tier-1 call sequencing, validation, and severity handling in the same design | the categorical statement contradicts later requirements and obscures which old conditions need accounting | narrow "unchanged" to reviewer independence and qualification, then enumerate the protocol changes +MINOR | medium | §10 "One new parked row" | the trigger "frequent enough that the silenced reminders demonstrably cost something" has no threshold, observation, or decision owner | the row can remain parked forever or be pulled on opinion, defeating the backlog's trigger discipline | name a measurable recurrence or concrete incident threshold and who records that it fired +MINOR | medium | §12 residuals | rollback is not addressed for an active degraded cycle, an installed project reverting to pre-ladder prompts, or outstanding re-review debts after the feature is withdrawn | markers, off files, and debt rows can become orphaned with no rule for preserving or closing them | add a prompt-only rollback procedure that restores reminders, preserves disclosure, and keeps or explicitly waives outstanding obligations +MINOR | medium | §4 "the two tool names a subagent-based reviewer would arrive under" | the spec never names those two tool names and does not verify that current Claude Code routes the chosen subagent mechanism outside the hook matcher | the hook-behavior claim is version- and mechanism-dependent despite prompt-standard item 11 requiring the mechanism to be verified | name the actual tool events for the selected mechanism and add them to the scratch verification, treating an unexpected routed name as a stop condition +END OF FINDINGS (37 total) diff --git a/.context/codex-reviews/gate-a-spec-pass-1.stopped-tier3-core.md b/.context/codex-reviews/gate-a-spec-pass-1.stopped-tier3-core.md new file mode 100644 index 0000000..1eb7c55 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-1.stopped-tier3-core.md @@ -0,0 +1,29 @@ +MAJOR | high | §2 lines 34–52 | [risk: threats; risk: rollback] The structural argument presents unenforceable prose and recursively gated state as the only possible controls, but a waiver can authorize its own record in the same fixed closure and a mechanical CI rule or protected external record can preserve or constrain that record without asking the unavailable reviewer to review each transition | The false dichotomy turns a chosen simplicity and trust trade-off into an impossibility claim, so future readers are told not to evaluate feasible compensating controls even though this change removes the product's only substantive gate | Recast §2 as a deliberate cost/trust decision, document the concrete alternatives considered (same-waiver state transition, CI append-only/marker checks, protected issue or branch approval) and why they were rejected, and retain the §12 trigger without claiming all stateful controls are structurally impossible +MAJOR | high | §3.1 lines 72–82 | [risk: threats; security: assets] Gate A's required evidence is vacuous or purely syntactic: “the artifact's own self-checks” may not exist, has no defined checklist or result record, and the mechanical sweep only establishes reference and syntax consistency | A high-risk prompt change can therefore waive its only substantive quality gate after satisfying formalities that say nothing about contradictions, unsafe decisions, prompt-standard conformance, or invariant risk, while “never evidence” implies more assurance than exists | Define a non-vacuous Gate-A waiver checklist and durable result, including all 12 prompt standards for prompt artifacts, explicit invariant-by-invariant self-audit, and accountable human review of the artifact; otherwise label the sweep as syntax hygiene rather than compensating evidence +BLOCKER | high | §3.2 lines 84–105 and §7 row 72 | [risk + security: abuse and abuse paths; security: external systems] The outage calls are not required to be correctly formed executions of the exact gate against the fixed artifact/range with the required story paths, evidence entries, tool, model and working directory; `attempt-failed` also admits output-write/validation failures after a review actually ran | An author or faulty agent can manufacture availability by sending an invalid request, using the wrong account/tool, or inducing a findings-write failure, then waive a reviewer that was capable of reviewing; this contradicts the stated “cannot run at all” boundary and makes row 72's “Kept, untouched” disposition unsupported | Require the initial call, recovery and revalidation to be canonical gate calls with every existing §5 argument and current artifact/evidence, reject request/validation/output-protocol failures as proof of reviewer unavailability, and record which canonical call class failed +MAJOR | high | §3.2 lines 98–103 | [risk: compatibility; risk: observability] When a fresh observation differs, `Cause` becomes the new observation, but the source enum has no general observation value and the retry requirement changes by source; for example quota revalidation that becomes a generic failure cannot be recorded as `attempt-failed` until a recovery is also spent | Different implementations can either write an out-of-enum cause, omit the newly required recovery, or refuse a valid waiver, so the closed enum and ordered close are not deterministic | Specify a transition table from every fresh result to one enum value, the recovery budget each transition requires, and the exact stop result for an unclassifiable failure +MAJOR | high | §3.2 lines 95–103 and §3.5 lines 134–140 | [risk: concurrency; security: external systems] Revalidation occurs before an unbounded human pause, while only a “materially delayed” commit invalidates authorization and no duration or clock rule defines materiality | The reviewer can recover while the human takes minutes or hours to answer, allowing a waiver to land even though the external system is available at the actual close; two agents can also apply different delay judgments | Define a short maximum age in UTC and run the final canonical availability probe after the human answer immediately before the commit; any expiry, clock failure, or different result must restart the close +BLOCKER | high | §3.5 lines 132–142, §4 lines 153–163, and §5.1 lines 172–182 | [risk: data loss; risk + security: abuse and abuse paths] A Gate-A waiver fixes an exact prospective tree but records only the artifact blob alternative, never requires the closing commit to be docs/artifact-only, and says the waiver takes precedence over a Gate-B STOP | A staged prompt, script, or code change can ride inside the Gate-A closing commit with only an A-spec/A-plan marker, silently bypassing Gate B in the dangerous false-✓ direction that Hook invariants 2 and 3 protect | For Gate A require a positively determined docs-only tree containing only the named artifact and explicitly allowed supporting prose, record the complete prospective Git tree id, and refuse any mixed/product path until a separately authorized Gate-B waiver closes it +BLOCKER | high | §3.5 lines 132–142 and §10 line 418 | [risk: concurrency; security: assets] The design never defines how the fixed prospective tree is atomically compared with what `git commit` actually writes; “if the tree changes” is only a rule, and another process can mutate the index between the check and the commit | Authorization may disclose one tree while Git lands another, exactly the mismatch §3.5 claims to prevent, and discovering it after commit has no specified rollback or incident path | Use a dedicated fixed index or equivalent immutable snapshot for the commit, compare the committed tree id to the authorized tree id immediately afterward, and specify rollback/re-authorization behavior for a mismatch or unverifiable tree +MAJOR | high | §3.5 lines 138–142 and §5 lines 165–204 | [security: assets; security: trust boundaries] Authorization binds only to the Git tree even though the evidence entry, outage cause, adverse-finding dispositions, merge strategy, authorization identity and waiver disclosure live in the commit message, which is not part of a tree and can change after the human pause | The durable record can materially misstate what the human approved while the content binding still appears valid | Present the complete prospective marker, evidence entry and decision schema to the human, bind approval to a digest or exact normalized rendering of both tree and message fields, and restart on any message-field change +BLOCKER | high | §3.5 lines 136–141 and §4 lines 160–163 | [risk: compatibility; risk: observability] The ordered close consumes Gate-A authorization at the docs commit, but the unchanged hook emits the plan Gate-A below-floor reminder later at `executing-plans`; §4 grants precedence only at one closing transition and says later or resumed sessions treat the hook as authoritative | A plan waived correctly cannot coherently proceed to implementation after the commit, especially after a session resume: either the agent obeys the later STOP and remains stalled or silently extends an already-consumed authorization | Define the state-free Gate-A continuation rule explicitly, including how the next skill invocation discovers and validates the immediately preceding A-plan marker across a resume, or move the authorization/consumption point to the transition where the Gate-A reminder actually fires +MAJOR | high | §3.1(2) lines 64–67 and §5.1 lines 176–177 | [risk: data loss; risk: observability] The waiver relies on knowing every prior adverse finding, but it gives no fail-closed rule for a missing, overwritten, unreadable or conflicting findings/dispositions file even though the spec records that this cycle already lost predecessor artifacts through slot collision | A resumed cycle can record `Adverse findings: none` merely because the evidence was lost, erasing known Blocker/Major review evidence through the exact data-loss path the precondition is meant to stop | Require a complete pass ledger reconstruction before waiver and treat any missing or unverifiable previously valid pass/disposition as unresolved adverse evidence that blocks tier 3 +MAJOR | high | §5.2 lines 195–204 and §5.3 lines 206–210 | [risk: idempotency] §5.3 keys both records by cycle id plus gate, but the human-decision block contains no gate field and only a free-form `authorizes` value | A-spec, A-plan and B decisions sharing a cycle id cannot be deterministically matched, deduplicated or conflict-checked, so multi-WIP collapse can merge the wrong authorization or reject the right one | Add a canonical `Gate: A-spec \/ A-plan \/ B` field to the decision block and include it in the exact cross-block match +MAJOR | high | §5.1 lines 172–182 and §5.2 lines 197–204 | [risk + security: abuse and abuse paths; security: roles; security: trust boundaries] Handles, reasons and disposition text have no single-line grammar, escaping rule, length bound or canonical timestamp/timezone, yet they are interpolated into a line-oriented commit-body protocol and later “normalized” for equality | A newline or delimiter in human-supplied text can inject or spoof marker fields, ambiguous dates can match differently across agents, and oversized or sensitive rationale can corrupt disclosure despite the sanitization sentence | Define an allowlisted handle grammar, ISO-8601 UTC timestamps, one-line escaped/encoded free-text fields with size limits, and exact byte-level normalization and parsing rules +MAJOR | high | §3.1(4) line 70 and §6 lines 212–232 | [risk: compatibility; risk: observability] Ordinary merge is allowed, but every carry and pre-merge validation requirement is specified only for WIP→amend→squash and the displayed chain unconditionally ends in squash | On an ordinary merge, implementers do not know which commit/body is the durable record, what must be validated, or how multiple Gate-A and Gate-B blocks survive; one permitted strategy therefore has no complete success or error path | Add a separate ordinary-merge chain and validation matrix identifying the reachable commits on main, required merge-commit content if any, handling of branch-head changes, and the exact refusal/recovery behavior +MAJOR | high | Story §5 lines 204–207 versus spec §2 lines 49–52, §11 lines 427–437, and §13 lines 471–474 | [risk: compatibility] The story still says the profile “scales the repayment,” but the revised design defines no repayment operation or profile-scaled re-review and says the only surviving obligation is an untracked sentence | This is a surviving dependency on the removed debt/reconciliation machinery, so the story's settled-decision accounting is not honest or implementable and the withdrawn decision is incomplete | Amend the story to say only that re-review is owed and untracked with no defined scaling or repayment procedure, or restore and specify the profile-dependent repayment behavior as an explicit in-scope requirement +MAJOR | high | §7 row 33 | [risk: compatibility] The inventory says a spent recovery attempt is a tier-3 precondition, but §3.2 permits `quota-observed` and `auth-failed` after one fresh call and reserves call-plus-recovery for `attempt-failed` | The supposedly exhaustive old-condition accounting misstates the new decision procedure and can cause implementations to require or skip the recovery inconsistently | Mark the recovery rule narrowed by source and enumerate which sources consume the shared recovery attempt +MAJOR | high | §7 row 39 | [risk: compatibility] “Gate A is two runs, each its own loop” is marked kept even though tier 3 closes either run with zero passes and therefore no loop | This is another effect-level change hidden under an untouched noun, undermining the fourth inventory's claimed clause-by-clause completeness | Mark it kept for tier 1 and narrowed for tier 3: the two independent gate cycles remain, but either may close without entering/completing its loop +MAJOR | high | §7 rows 55–58 | [risk: threats; security: assets; security: trust boundaries; security: roles; security: external systems] All risk/security lens obligations are marked kept untouched, but a tier-3 cycle has no review prompt to which threats, abuse, rollback, data loss, idempotency, compatibility, observability, assets, trust boundaries, roles or external-systems questions can be appended | High-risk and security-profiled changes can close without any actor applying the questions the profile was created to require, while the inventory falsely claims no effect | Mark these conditions narrowed or inapplicable at tier 3 and decide whether the Gate-A human waiver checklist must explicitly perform the unioned lenses; if not, disclose that those lenses are waived too +MAJOR | high | §7 row 61 and §3.1 lines 61–82 | [risk: compatibility] “Stop and surface on an unresolvable profile” is marked kept untouched, but Gate A's tier-3 preconditions never require the cited story profile to resolve; only Gate B's reference to mode indirectly forces resolution | A malformed or semantically inconsistent high-risk profile can be bypassed precisely on the zero-pass path, weakening evidence and lens obligations while the inventory says the old stop survives | Make successful fresh profile resolution a common precondition for every cited story at both gates and define the no-story/unprofiled cases separately +MAJOR | high | §7 row 62 | [risk + security: abuse and abuse paths] The inventory marks the two-condition Gate-B skip rule untouched because “a waiver is not a skip,” even though row 49 correctly says the effect of tier 3 is a non-trivial change closing without review regardless of the noun | This repeats the noun-based accounting error §7 explicitly says the fourth inventory avoids and obscures a second gate-off path that bypasses triviality/profile eligibility | Mark the legacy skip rule kept within tier 1 but narrowed at the overall decision-procedure level by tier 3, cross-referencing its distinct preconditions and disclosure +MAJOR | high | §8 lines 327–356 and docs/pr-review-bots.md lines 136–143 | [risk: compatibility] The site inventory omits the assertion that every PR head has already passed Gate B, which becomes false for a tier-3-waived head | PR-bot routing will be justified by independent-review evidence that may not exist, so a supplementary review can remain non-blocking under a premise this feature invalidates | Add this site, mark the premise narrowed, and specify how PR-bot routing describes or treats a waived Gate B +MAJOR | high | §8 lines 339–349 and docs/coding-workflow.md lines 253–268 | [risk: compatibility] The coding-workflow site list stops at the stage-level passages and explicitly excludes only line 149, but misses “What is essential — keep it” naming two independent review gates as load-bearing and the minimal-adoption claims below it | The methodology will continue teaching an unconditional invariant that tier 3 deliberately overturns, violating the old-condition and semantic-drift Don'ts | Add the adapting/minimal-adoption section to the site inventory and qualify independence as the normal gate with an explicit availability waiver +MINOR | high | §8 line 344 and .claude-plugin/marketplace.json line 14 | [risk: compatibility] The spec says marketplace metadata carries the same two-independent-gates claim, but the working-tree marketplace description only lists plugin components and does not make that claim | The mechanical sweep should reject the factual premise, and an unnecessary metadata edit creates scope and release risk | Remove this site or state the distinct wording that actually needs changing and why, based on the file's current contents +MAJOR | high | §10 lines 387–403 and story header lines 6–16 | [risk: observability; security: assets] The withdrawn mode override depends on a “prompt-harness scenario,” but the repository contains no such harness and the design gives no driver, old/new prompt fixture, deterministic oracle, command, artifact path or handling for nondeterministic agent behavior | `battery+check+verification` therefore has no implementable `+check` that demonstrably fails without the change; calling an unspecified future scenario a check recreates the unverified-evidence claim the withdrawal was meant to fix | Specify the harness as a concrete artifact and command with frozen old/new inputs, observable assertions for every required branch, repeatability criteria and failure semantics, or keep the evidence gap and obtain a valid whole-mode human-confirmed override +MINOR | high | §5.1 lines 167–170 | [risk: observability] The text says §11 validates the marker and decision block independently, but validation is in §10 and §11 contains only withdrawn decisions | A reader following the required-record claim lands in the wrong section, and the advertised mechanical sweep should have caught the stale cross-reference | Change the reference to §10 +MAJOR | high | §13 lines 485–488 and invariant 8 | [risk: rollback; risk: compatibility; risk: data loss] The rollback account considers only reverting this repository's prompts and says the procedure is removed, but `/workflow-init` copies an inline §5 into downstream repositories and installed plugins live in versioned caches; reverting or downgrading this repo does not remove those deployed copies | A rollback can leave tier 3 active in initialized projects with no matching current documentation, support procedure or future fixes, while operators believe the waiver was removed | Define rollback for installed and scaffolded copies, including version/downgrade behavior, how downstream CLAUDE.md mirrors are detected and migrated, and what happens to already authorized in-flight waivers +MINOR | high | §2 lines 49–52 | [risk: observability] “A line on main that every reader sees” overstates commit-body visibility: the marker is only discoverable by inspecting history and may be buried in an ordinary merge, while §13 admits history can be rewritten | This turns durable-at-commit-time disclosure into an ambient visibility guarantee and conflicts with the gate-proof calibration invariant | Say the line is discoverable in reachable main history when commit bodies are inspected, not that every reader sees it, and name the history query or UI surface used to find it +MINOR | medium | §4 lines 160–163 | [risk: observability] Saying later sessions treat the hook as “authoritative” survives the rewrite even though the same section and invariant 1 establish that the hook emits advisory text and always exits 0 | Readers can still infer a mechanism or authority the hook does not possess, especially when the waiver is said to take precedence over it | Say the compliant-agent procedure obeys those later reminders; reserve “authoritative” for an enforced source of truth +MAJOR | high | §3.1 line 80, §6 line 221, and CLAUDE.md lines 335–372 | [risk: compatibility; security: external systems] Gate B's precondition uses singular “the profile's mode,” while current §5 requires per-story modes, suffixes, lens unions and one evidence entry per cited profiled story | A multi-story cycle can satisfy only one story's checks yet produce a superficially valid waiver marker and squash body, under-serving stricter stories | State that every cited story is freshly resolved and independently satisfies its full mode and suffix, the battery runs once, lenses are unioned, and all per-story entries are revalidated and carried +END OF FINDINGS (28 total) diff --git a/.context/codex-reviews/gate-a-spec-pass-10-dispositions.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-10-dispositions.pre-2026-08-14.md new file mode 100644 index 0000000..4327ee4 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-10-dispositions.pre-2026-08-14.md @@ -0,0 +1,33 @@ +# Gate A — spec — pass 10 dispositions + +Artifact: `docs/superpowers/specs/2026-08-05-hardening-ledger-supersession-design.md` +5 findings: 3 Major, 2 Minor. **No Blocker.** All read as correct. **None applied.** + +Reviewer: `mcp-codex-dev@1.0.1`, `codex-cli 0.144.5`, model `gpt-5.6-sol`. Same as passes 4–9. + +Passes 4 → 10: **11, 14, 16, 15, 9, 10, 5.** First pass since the cycle began with no Blocker, +and the smallest count of the run. Both human-settled changes (the row-bound floor, match +semantics) drew no objection to the decision itself — only to remnants left beside them. + +## Verified, not taken on the reviewer's word + +| # | Claim | Result | +|---|---|---| +| 1 | assertion 2's "self-contained" fence is broken | **Reproduced.** The fence defines `unwrap() { :; }` — a stub I wrote as a placeholder with the comment "as defined in assertion 1". With it, `occurrences()` reads no input and returns `0` for text that is present: checked at a shell, `occurrences('hello')` on a file containing `hello` returns **0**, not 1. Every sentinel assertion then fails and the block exits before extracting either region. **Pass 9's finding 6 was not actually applied** — the fix made the block self-contained and non-functional. | +| 2 | uniqueness remnants survive | **Confirmed**, `:588` ("an entry's locator resolves to exactly one row") and `:633` ("its locator selects one row"). | +| 5 | "nothing detects it" contradicts check 3 | **Confirmed**, `:641` against check 3's duplicate-label assertion. | + +| # | Sev | Verdict | Note | +|---|---|---|---| +| 1 | Major | **Valid, reproduced** | Mine, and the cycle's signature failure in miniature: the fix for "not self-contained" produced a block that is self-contained and does nothing. A stub is not an implementation. | +| 2 | Major | **Valid, confirmed** | Match semantics was applied to §2.2's rule and not to the two places that consume the withdrawn guarantee. Also correctly notes an open question the settlement did not answer: **is a zero-match locator a stop?** Match semantics says an entry applies to what it matches; it does not say what a locator matching *nothing* means. | +| 3 | Major | **Valid** | `D=$(mktemp -d)` is unchecked before the trap and the redirects. On failure `D` is empty and the block writes `/hardening-log.md.region`, outside the scratch dir the paragraph promises. Same class as finding 1 — the no-artifact claim rests on a step that can fail silently. | +| 4 | Minor | **Valid** | §7's stored riders are not what pass 10 actually ran: the append-only rider lists six cases, not the seven including the pre-row state, and carries no match-semantics walk. The section claims to be "verbatim in every pass prompt". | +| 5 | Minor | **Valid, confirmed** | Gate-coverage overclaim class: "nothing detects it" while check 3 rejects duplicate labels once, during this change. | + +## Read + +Nothing here reopens a settled decision. Three of the five are defects in the *mechanics* of §6's +checks (a stub helper, an unchecked `mktemp`, an overclaim about what a check proves); two are +remnants of the withdrawn uniqueness rule. One genuine open question falls out of finding 2 — +what a **zero-match** locator means — which the match-semantics settlement did not cover. diff --git a/.context/codex-reviews/gate-a-spec-pass-10.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-10.pre-2026-08-14.md new file mode 100644 index 0000000..b05b5af --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-10.pre-2026-08-14.md @@ -0,0 +1,6 @@ +MAJOR | high | §6 Check 2, assertion 2 | the supposedly self-contained fence defines `unwrap() { :; }` instead of repeating assertion 1's implementation | every `occurrences` call receives no input and returns zero, so the mandatory parity check always exits before extracting either region under both `sh` and `dash`; pass 9's finding 6 was not actually applied | repeat the full `unwrap` helper in the fence, then exercise the standalone block on scratch pre-change, one-surface, incomplete, and complete fixtures +MAJOR | high | §2.2 locator rationale and §8 first bullet | the withdrawn uniqueness rule still survives as "an unresolvable locator is a stop" and as a missing check that a locator resolves to exactly one row | match semantics deliberately permits multiple matches, while the shared convention states no stop for zero matches; these remnants leave authors unable to tell whether the stop means zero or many and make the residual coverage claim rest on a requirement the design removed | delete the exactly-one residual, define a zero-match locator as the only stop if that is intended, and state explicitly that two-or-more matches remain valid +MAJOR | high | §6 Check 2, assertion 2 temporary directory setup | `D=$(mktemp -d)` is not checked before redirects and cleanup use it | if `mktemp` fails, `D` is empty and the block attempts to create `/hardening-log.md.region` and `/workflow-init.md.region`; with sufficient permissions it leaves files outside the scratch directory, contradicting the no-artifact claim | fail immediately unless `mktemp -d` succeeds and returns a non-empty directory before installing the trap or performing redirects +MINOR | high | §7 "Gate-A riders, verbatim in every pass prompt" | the stored riders are not verbatim in pass 10 and the append-only rider still lists only six cases, omitting the newly settled half-typed pre-row state and all match-semantics consequences | the claimed persistent review mechanism can regress the row-bound floor or retain a uniqueness assumption even though this pass's external prompt happens to add those checks | synchronize §7 with the current riders, including the seventh floor case and match-semantics walk, or rename it as non-authoritative review guidance and point to the actual authoritative prompt +MINOR | high | §8 union-merge residual | "nothing detects it" is absolute even though §6 check 3 explicitly rejects duplicate `Superseded rows:` labels during this change | the spec gives incompatible accounts of what its check proves, repeating the gate-coverage overclaim class called out by AGENTS.md | qualify the residual to say that no standing check detects duplicates after the one-time check 3 has run +END OF FINDINGS (5 total) diff --git a/.context/codex-reviews/gate-a-spec-pass-11-dispositions.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-11-dispositions.pre-2026-08-14.md new file mode 100644 index 0000000..4cf231e --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-11-dispositions.pre-2026-08-14.md @@ -0,0 +1,41 @@ +# Gate A — spec — pass 11 dispositions + +Artifact: `docs/superpowers/specs/2026-08-05-hardening-ledger-supersession-design.md` +12 findings: 7 Major, 2 Minor, 3 Nit. **No Blocker.** All read as correct. **None applied.** + +Reviewer: `mcp-codex-dev@1.0.1`, `codex-cli 0.144.5`, model `gpt-5.6-sol`. Same as passes 4–10. + +Passes 4 → 11: **11, 14, 16, 15, 9, 10, 5, 12.** + +**Read the rise honestly.** Pass 10 was 5 because the artifact had stopped changing much. Pass 11 +is 12 because pass 10's fixes **added new mechanism** — an executable inert-entry assertion and a +new inert semantics — and new mechanism generates new findings. Six of the twelve land on things +that did not exist before this round (findings 1, 3, 4, 5, 10, 11). That is a different curve +from passes 4–7, where each pass found the *previous fix* unsound; here the settled decisions +drew no objection at all. + +## Reproduced, not taken on the reviewer's word + +| # | Claim | Result | +|---|---|---| +| 3 | the new inert assertion's `rc` is lost to a pipeline subshell | **Reproduced under both `sh` and `dash`.** `grep … \| while … rc=1` runs the loop in a subshell; the fence prints `INERT: 2099-01-01 missing-fp matches no row` and **exits 0**. The assertion detects and does not fail. My prototype test checked the printed output and never the exit status — a check wired to pass, written while building a check against wording-not-mechanism. | +| 11 | `**Cite, do not restate.*` has an unmatched delimiter | **Confirmed**, `:228`. My scripted edit broke it. | + +| # | Sev | Verdict | Note | +|---|---|---|---| +| 3 | Major | **Valid, reproduced** | See above. Fix: redirect a guarded temp file into `while` instead of piping, and verify an inert fixture exits nonzero under both shells. | +| 4 | Major | **Valid** | The assertion's `sed` ignores an optional `""`, so an entry whose pair matches but whose fragment matches no `finding` is inert under §2.2 and passes. | +| 5 | Major | **Valid, and genuinely new** | Matching is defined dynamically, so an **inert entry can activate later** when a row with that locator is appended — a zero-to-nonzero transition neither the settlement nor §8 contemplates. A mistyped entry left as history can silently come to govern an unrelated future row. Needs Daniel: is inertness permanent or current-only? | +| 1 | Major | **Valid** | "an unresolvable locator is a stop" survives in §2.2's rationale — a stop for a state that is now either valid (many) or inert (zero). | +| 2 | Major | **Valid** | §3.1 still explains the omitted fragment by exact-one cardinality and misquotes §2.2 as "already resolves". The worked example teaches the withdrawn procedure. | +| 6 | Major | **Valid** | §8's inventory omits two shared-region decisions — the block sitting above `Columns:`, and the label carrying one appended line per supersession. Fourth pass on this inventory; the rule-plus-list form did not stop it drifting. | +| 7 | Major | **Valid** | `docs/coding-workflow.md:204` uses `entry` for a hardening *row*, while this design reserves `entry` for supersession markers and forbids entries referencing one another. As written it reads as reopening the cut correcting-entry path. My rewrite introduced the collision. | +| 8 | Minor | **Valid** | "Nothing mechanical reads this block" is absolute while checks 1 and 3 grep, parse and position-check it during this change. Enforcement-claim class; the intended claim is about *standing* tooling. | +| 9 | Minor | **Valid** | §4's `todos.md` row does not say what the resolved source row says, which status it lands in, or where the parked item goes — the docs-drift class this repo treats as recurring. | +| 10 | Nit | **Valid** | Anchor numbering off by one: the citation phrase is anchor 20, not 19. | +| 11 | Nit | **Valid, confirmed** | Broken bold, mine. | +| 12 | Nit | **Valid** | The spec attributes the restoration to pass 8, the story to pass 7. Both are defensible (finding pass vs fix pass) and they disagree. | + +## Needs Daniel + +**Finding 5 only.** Everything else is mechanical or editorial. diff --git a/.context/codex-reviews/gate-a-spec-pass-11.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-11.pre-2026-08-14.md new file mode 100644 index 0000000..92fb17f --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-11.pre-2026-08-14.md @@ -0,0 +1,13 @@ +MAJOR | high | §2.2 "Why uniqueness is not required" | the sentence "which is why an unresolvable locator is a stop" survives the withdrawn uniqueness rule | zero-match locators are settled as inert and multiple matches are valid, so the convention now gives authors an undefined stop that contradicts both accepted cardinalities | delete the stop sentence and state only that a fragment narrows the match set while zero matches remain inert and multiple matches remain valid +MAJOR | high | §3.1 fragmentless first entry | the explanation says the fragment is omitted because the pair matches exactly one row and misstates §2.2 as saying to omit it when the pair "already resolves" | exact-one is no longer the omission criterion, and a later same-pair row explicitly joins the set, so the worked example teaches the withdrawn uniqueness procedure | say the fragment is omitted because this marker is intended to apply to every row the pair matches, with the current one-row cardinality recorded only as an observation +MAJOR | high | §6 Check 1 "No inert entry" fence | `rc=1` is assigned inside a pipeline-fed `while`, so POSIX `sh` and `dash` discard it when the subshell exits | reproduced on scratch fixtures under both shells: the fence prints `INERT: 2099-01-01 missing-fp matches no row` and exits 0, so the new assertion detects in wording but not in mechanism | feed the loop without a pipeline subshell, such as through a guarded temporary file redirected into `while`, and verify an inert fixture exits nonzero under both shells +MAJOR | high | §6 Check 1 "No inert entry" fence | the `sed` commands extract only row date and fingerprint, and the table grep ignores an optional `""` | an entry whose date and fingerprint match a row but whose fragment matches no finding is inert under §2.2 yet passes silently; reproduced under both `sh` and `dash` | parse the optional fragment and require at least one date-plus-fingerprint row whose finding contains it literally, while retaining pair-only match-all behavior when the fragment is absent +MAJOR | high | §2.2 inert-entry semantics versus "a later row simply joins the set" | an unmatched entry is said to remain inert history, but the same section defines matching dynamically so a later row with that locator joins its set | a mistyped entry left as history can silently activate later and govern an unrelated future row, a zero-to-nonzero state transition the spec neither accepts nor contains | decide whether inertness is permanent or current-only; if current-only disclose the activation risk, and if permanent define a stable eligibility rule that does not require editing history +MAJOR | high | §8 "What check 2's anchor list does not cover" | the claimed complete unanchored inventory omits shared-region decisions including that the block is above `Columns:` and that its label carries one appended line per supersession | deleting either instruction from both surfaces leaves all twenty-six anchors, both sentinels, parity, and check 3 green, so the four-state coverage account still overstates what the checks compare | add deletion-sensitive anchors for these decisions or list them as unvalidated and narrow the four-state claim accordingly +MAJOR | high | docs/coding-workflow.md:204 | the rewritten methodology uses `entry` for a hardening row and then says an entry whose text is wrong gets a supersession record, while this design reserves `entry` for supersession markers and forbids entries from referencing one another | the terminology reopens the cut correcting-entry path and leaves readers unable to tell whether a wrong supersession marker may itself be superseded | use `row` consistently: a resolution is a new row referencing the row it closes, and a row whose text is wrong gets a supersession entry +MINOR | high | §2.2 shared convention final sentence | "Nothing mechanical reads this block" is absolute, but checks 1 and 3 mechanically grep, parse, and position-check the block during this change | the shipped prompt makes a mechanism claim contradicted by its own validation, risking invariant 11's enforcement-claim failure even though the intended claim is only about standing recurrence tooling | say that no standing recurrence reader consumes the correction and explicitly distinguish the one-time landing checks that read structure and locators +MINOR | high | §4 `todos.md` change surface | "the source row resolved; the uncovered case parked with its trigger" does not specify how the still-present-tense source row is historicalized, which checkbox state lands, or the content and placement of the new parked item | an implementer can nominally satisfy the row while leaving `todos.md` saying the ledger still has no convention, recreating the docs-drift class this repository treats as recurring | specify the source row's completed wording and status, the parked hardening-removal row with its first-removal trigger, and a named read that validates both +NIT | high | §8 anchor inventory | the prose says "citing where the current answer lives" became anchor 19, but anchor 19 is the inert-entry phrase and the citation phrase is anchor 20 | the coverage audit points at the wrong mechanical evidence | change anchor 19 to anchor 20 +NIT | high | §2.2 rationale heading | `**Cite, do not restate.*` has an unmatched bold delimiter | the malformed Markdown obscures the heading in the artifact being approved | close it as `**Cite, do not restate.**` +NIT | high | §4 and §7 amendment accounting versus governing story §1 | the spec attributes restoration of the absolute rule to pass 8, while the governing story says it was amended from pass 7 | the trace for the decision-procedure reversal is internally inconsistent even though the restored rule itself agrees | choose whether attribution names the finding pass or the first pass containing the fix and use that convention consistently in all three sites +END OF FINDINGS (12 total) diff --git a/.context/codex-reviews/gate-a-spec-pass-12-dispositions.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-12-dispositions.pre-2026-08-14.md new file mode 100644 index 0000000..c2ba1e6 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-12-dispositions.pre-2026-08-14.md @@ -0,0 +1,37 @@ +# Gate A — spec — pass 12 dispositions + +Artifact: `docs/superpowers/specs/2026-08-05-hardening-ledger-supersession-design.md` +8 findings: 5 Major, 2 Minor, 1 Nit. **No Blocker.** All read as correct. **None applied.** + +Reviewer: `mcp-codex-dev@1.0.1`, `codex-cli 0.144.5`, model `gpt-5.6-sol`. + +Passes 4 → 12: **11, 14, 16, 15, 9, 10, 5, 12, 8.** + +## Reproduced + +| # | Claim | Result | +|---|---|---| +| 4 | the new layout anchor's backticks are unescaped | **Reproduced under `sh` and `dash`.** `:485` reads `"block above the \`Columns:\`"` with **no** backslash escapes, so the shell attempts `Columns:` as a command substitution: it prints `Columns:: command not found` and the anchor degrades to `block above the `. The block still exits 0. Every other backtick-bearing anchor escapes correctly. **My anchor harness could not see this** — it stripped the escapes and did a literal `grep`, never exercising shell expansion. | + +| # | Sev | Verdict | Note | +|---|---|---|---| +| 4 | Major | **Valid, reproduced** | The anchor added last round to close a coverage gap validates nothing and fails silently. Same class, mine, again — and the harness blind spot is the interesting part: fence *execution* was added to catch stubs, and it did run, but nothing asserted that a fence's failure output means failure. | +| 3 | Major | **Valid** | Check 1's first fence prints `0` then `1` and exits 0 — it can report its own falsifying observation as success. I flagged this in the pass-12 prompt as "a stated read rather than a gate"; Codex is right that describing it does not excuse it. | +| 6 | Major | **Valid, and the best finding of the pass** | All five fences still pass after the 2026-07-20 row's narration is **edited**, and after the existing first header paragraph is changed identically in both surfaces. So the validation can go green while breaking the absolute append-only rule, story AC 2's every-byte requirement, and §4's "the existing first paragraph is **untouched**". Nothing in §6 pins either. | +| 1 | Major | **Valid** | The prose says "a row appended later can never come under an entry written before it"; the test excludes only *later-dated* rows. Same-day appends (already disclosed) **and backdated appends** (not disclosed) both remain eligible, and §3.1 says a later same-locator row joining the set is *intended* — so prose and procedure disagree in both directions. | +| 2 | Major | **Valid** | §2.1's trigger sanctions an entry when *any* later change falsifies a row, while the same paragraph puts a removed hardening out of scope. An author facing a removal gets both instructions. Prompt-standards item 7. | +| 5 | Minor | **Valid — and it lands on Daniel's phrasing** | "a typo fails the battery pre-merge, and merged history cannot carry an inert entry" overclaims: this assertion is a one-time bespoke check in §6, **not** in `AGENTS.md`'s quality battery, and §8 says no standing check validates future entries. Either wire it into the battery or narrow the sentence. | +| 7 | Minor | **Valid** | Fifth consecutive pass on the unanchored inventory: deleting the inert-entry remedy or the name-the-false-claim instruction from both surfaces leaves all 29 anchors intact. | +| 8 | Nit | **Valid** | §4 quotes the landed wording without its `**` emphasis. | + +## Needs Daniel + +- **F5** — his own phrasing: wire the assertion into `AGENTS.md`'s battery, or narrow the claim. +- **F1** — backdating: accept and disclose as a second residual, or forbid and validate. +- **F6** — byte-pin the target row and the existing header paragraph, or declare them unvalidated. + +## Process note + +The reviewer wrote `.context/gate-a-pass12-harness.sh` into the repo during this pass. It is no +longer present and `.context/` is gitignored, so nothing leaked into the worktree — but a gate +call is expected to write only its findings file, and this one did not. diff --git a/.context/codex-reviews/gate-a-spec-pass-12.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-12.pre-2026-08-14.md new file mode 100644 index 0000000..dc202cb --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-12.pre-2026-08-14.md @@ -0,0 +1,9 @@ +MAJOR | high | §2.2 "An entry applies only to matching rows dated on or before the entry's own date" and §3.1's fragmentless-entry rationale | the formal date test excludes only later-dated rows, while the shared convention says a row appended later can never join and §3.1 says a later row with the same locator intentionally does join; same-day and backdated later appends both satisfy the test | an entry believed inert or bounded to existing rows can activate for a row written after it, so the convention's temporal guarantee and inert-entry consequences contradict its actual decision procedure | replace "appended later can never" with "dated later does not", state that same-day and backdated later appends remain eligible, and either accept both residuals or forbid and validate backdating +MAJOR | high | §2.1 "When a row's text no longer describes reality" | the universal later-change branch sanctions a supersession entry when any later change falsifies row text, but the same paragraph says a hardening later removed is out of scope without excluding removal from that branch | an author facing a removed hardening receives both "append an entry" and "out of scope", violating prompt-standards item 7 and leaving the parked case's boundary undefined | exclude hardening removal explicitly from the sanctioned trigger and state that this convention supplies no move for it pending the parked story +MAJOR | high | §6 Check 1 first `sh` fence | the fence prints counts but does not assert them; with the entry absent and the table row present it prints `0` then `1` and exits 0 under both `sh` and `dash` | the advertised mechanical half can report its own falsifying observation as success, repeating the report-without-gating failure class the spec calls out | capture both counts and exit nonzero unless each equals 1 +MAJOR | high | §6 Check 2 assertion 1 anchor `block above the Columns paragraph` | the backticks around `Columns:` are unescaped inside a double-quoted shell word, so both shells attempt a `Columns:` command substitution twice, print command-not-found, reduce the anchor to `block above the `, and still exit 0 on the good fixture | the newly added layout anchor does not validate the decision it names, and a shell execution failure is again only reported rather than gated | write `"block above the \`Columns:\`"` as the other backtick-bearing anchors do and retain the two-shell execution test +MINOR | high | §6 Check 1 "No inert entry — gating" | the prose says a typo fails the battery and merged history cannot carry an inert entry, but this assertion is a one-time bespoke check outside the AGENTS.md battery and §8 explicitly says no standing check validates future entries | readers can mistake a single implementation-time observation for repo enforcement and assume future inert entries are mechanically prevented | say the one-time Check 1 fails for this change and that future entries remain instruction-backed, or wire a standing checker before making the broader claim +MAJOR | high | §6 Validation and story acceptance criterion 2 | all five fences still pass after the 2026-07-20 table row's narration is edited, and they also pass after the existing first header paragraph is changed identically in both surfaces | the validation can go green while violating the absolute append-only rule, the story's every-byte requirement, and §4's authoritative instruction that the existing paragraph is untouched | add base-relative exact-byte checks for the complete target row and the pre-existing header paragraph in both surfaces, or make each an explicit named read and disclose the lack of mechanical validation +MINOR | high | §8 "What check 2's anchor list does not cover" | the claimed whole list names only fragment narrowing and its quote restriction, but deleting the explicit inert-entry remedy or the instruction to name the false claim from both surfaces leaves all 29 intended anchors and both sentinels intact | the coverage inventory overstates Check 2 and hides shared decision text that can disappear symmetrically without detection | add distinguishing anchors for those decisions or include them in the unanchored inventory; also account for the existing-header immutability decision unless the new byte check covers it separately +NIT | high | §4 `docs/coding-workflow.md` change-surface row | the quoted landed wording says `strictly append-only: a row is never edited`, while the actual normalized source says `strictly append-only: a **row** is never edited` | the rider's byte-for-byte quotation check fails even though the semantic wording matches | include the `**` emphasis markers in the quote or present it explicitly as a paraphrase +END OF FINDINGS (8 total) diff --git a/.context/codex-reviews/gate-a-spec-pass-13-dispositions.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-13-dispositions.pre-2026-08-14.md new file mode 100644 index 0000000..d2ba57a --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-13-dispositions.pre-2026-08-14.md @@ -0,0 +1,54 @@ +# Gate A — spec — pass 13 dispositions + +Artifact: `docs/superpowers/specs/2026-08-05-hardening-ledger-supersession-design.md` +6 findings: **1 Blocker**, 4 Major, 1 Minor. All read as correct. **None applied.** + +Reviewer: `mcp-codex-dev@1.0.1`, `codex-cli 0.144.5`, model `gpt-5.6-sol`. +This call left **no** stray artifact (pass 12's did); the only untracked path, `docs/research/`, +predates the session. + +Passes 4 → 13: **11, 14, 16, 15, 9, 10, 5, 12, 8, 6.** + +## The blocker — two settled instructions are in tension + +The inert assertion scans **every** entry (`grep '^- … supersedes '`). The convention says an +inert entry is **never removed and never edited** — it stands as history. So the first time the +typo case actually fires: + +- appending the corrected entry leaves the inert one in the scan → the check stays red, forever; +- editing or deleting the inert entry → violates the absolute rule the whole design exists for. + +**There is no legal state in which validation is green again.** Confirmed against the text at +`:412` (the scan) and §2.2 (inert entries stand). + +This is not a defect in either instruction on its own. "An inert entry stays as history" (the +zero-match settlement) and "the assertion stays gating over all entries" (the F5 settlement) are +individually sound and jointly unsatisfiable. **It needs Daniel.** Codex's suggestion — scope this +change's check to §3.1's mandated entry only, and turn the parked battery-wiring row into a +*design* task for enforcement compatible with immutable inert history — is one way; there are +others (e.g. gate only entries appended by the change under review). + +## Reproduced / verified + +| # | Claim | Result | +|---|---|---| +| 1 | no legal green state once an inert entry exists | **Confirmed** from the scan pattern and §2.2. | +| 6 | `[ -n "$B" ]` is ineffective | **Confirmed**: `printf '' \| cksum` → `4294967295 0`. Non-empty for empty input, so the guard never fires. | +| 3 | `BASE` silently defaults to `HEAD` | **Confirmed** at `:446` and `:458`. | + +| # | Sev | Verdict | Note | +|---|---|---|---| +| 1 | **Blocker** | **Valid** | See above. Needs Daniel. | +| 2 | Major | **Valid** | The fragment is matched against the **whole table row**, not the `finding` column it is defined to locate. Codex reproduced it: a fragment taken from the row's `ref` (`mcp-codex-dev@1.0.1`) exits 0. My prototype only ever tested fragments that happened to live in `finding`. | +| 3 | Major | **Valid, confirmed** | `${BASE:-HEAD}` means that after this change is committed, an edited row becomes **its own baseline** — the pin passes on a committed mutation. The base-relative design was right; the default undoes it. | +| 4 | Major | **Valid** | The dates assertion checks **recorded string order**, not that a date is the row's actual append day. A row appended today but labelled `2026-08-05` after a `2026-08-04` tail passes. So "backdating is forbidden **and mechanically checked**" is an overclaim — mine, written in the same edit that added the check. | +| 5 | Major | **Valid** | Anchor 22 covers `A row's date is the day it is appended` but **not** `the table is chronological: backdating a row is forbidden`. Deleting that clause from both surfaces leaves all 32 anchors and parity green — sixth consecutive pass finding a hole in this inventory. | +| 6 | Minor | **Valid, confirmed** | The `cksum` fence cannot prove it read a real base paragraph. | + +## Pattern worth naming + +Findings 2, 3, 4 and 6 are all the **same shape**: a check that is right in design and wrong in +one mechanical detail, where the detail makes it pass when it should fail. That is a better class +of defect than passes 4–7 produced (where the *design* kept being unsound), but it is the class +that survives review longest, because the check looks correct and the fixture that would expose it +is the one nobody wrote. diff --git a/.context/codex-reviews/gate-a-spec-pass-13.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-13.pre-2026-08-14.md new file mode 100644 index 0000000..92fa040 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-13.pre-2026-08-14.md @@ -0,0 +1,7 @@ +BLOCKER | high | §2.2 "inert" and §6 Check 1 "No inert entry — gating" | the all-entry assertion rejects exactly the immutable inert entry that the convention says to leave standing | when the intended typo case fires, appending the corrected entry leaves the inert one in the scan, while editing or removing it violates the absolute entry rule, so there is no legal state in which validation can become green | make this change's check validate only the mandated §3.1 entry, and replace the parked direct-wiring task with a design task for enforcement compatible with immutable inert history +MAJOR | high | §6 Check 1 inert-entry fence, fragment branch | the optional fragment is searched across the entire table row, not only the `finding` column it is defined to locate | a fragment absent from `finding` but present in `ref`, `rung`, or another column passes as non-inert; a scratch entry using `mcp-codex-dev@1.0.1` from the target row's `ref` exited 0 | parse the escaped-pipe-aware `finding` column and apply the fixed-string fragment match only to that value +MAJOR | high | §6 Check 1 base-relative row and paragraph pins | both alleged base reads silently default `BASE` to current `HEAD` | after the implementation is committed, an edited row or identically edited paragraph becomes its own baseline; a committed mutation of the protected row passed the row-pin fence with `BASE` unset | require an explicit pre-change base ref, fail when it is unset or unresolvable, and document that shallow callers must fetch that object before running the fences +MAJOR | high | §6 Check 1 "Dates never decrease down the table" | the assertion checks only non-decreasing recorded strings and cannot establish that a row's date is its actual append day | a row appended on 2026-08-11 but labelled 2026-08-05 after the existing 2026-08-04 tail passed, so the claim that the assertion mechanically enforces the backdating prohibition is false and a later backdated row can activate an older inert locator | narrow the assertion and every dependent claim to chronological string order, leaving append-day truth explicitly instruction-backed, or design standing change-relative enforcement that checks newly appended rows against their append date +MAJOR | high | §6 Check 2 anchor list and §8 unanchored inventory | anchor 22 covers only `A row's date is the day it is appended`, so deleting the separate `table is chronological: backdating a row is forbidden` decision from both surfaces leaves all 32 anchors and parity green, while §8 says only fragment narrowing is unanchored | the load-bearing prohibition can disappear identically from both prompt surfaces without the presence mechanism detecting it; blocks 5, 6, and 7 all exited 0 on that mutant | add an exact backdating/chronology anchor, update the count to 33, and correct §8's inventory +MINOR | high | §6 Check 1 first-paragraph cksum fence | `git show ... \| para ... \| cksum` does not preserve `git show` or range-extraction failure, and `[ -n "$B" ]` is ineffective because `cksum` emits a checksum even for empty input | an unavailable shallow/base object reports later as a changed ledger with raw git stderr, and the fence never actually proves that its claimed base paragraph with both delimiters was read | capture `git show` into guarded scratch input, assert exactly one start and terminator, then checksum the validated range +END OF FINDINGS (6 total) diff --git a/.context/codex-reviews/gate-a-spec-pass-14-dispositions.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-14-dispositions.pre-2026-08-14.md new file mode 100644 index 0000000..c5454f8 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-14-dispositions.pre-2026-08-14.md @@ -0,0 +1,55 @@ +# Gate A — spec — pass 14 dispositions + +Artifact: `docs/superpowers/specs/2026-08-05-hardening-ledger-supersession-design.md` +3 findings: 1 "Blocker", 1 Major, 1 Minor. **Two valid, one misdiagnosed.** None applied. + +Reviewer: `mcp-codex-dev@1.0.1`, `codex-cli 0.144.5`, model `gpt-5.6-sol`. No stray artifact. + +Passes 4 → 14: **11, 14, 16, 15, 9, 10, 5, 12, 8, 6, 3.** + +## Finding 1 — the "Blocker" is misdiagnosed, and is not a blocker + +Codex states the terminator is `/date + anchor\)\.$/` with an unescaped `+`, so the range would +never terminate and the fence would checksum to EOF. + +**The spec does not say that.** `:489` reads `/date \+ anchor\)\.$/` — the `+` **is** escaped. +Verified with `cat -A`, and both the spec form and Codex's suggested `[+]` form terminate +correctly on the awk here (`awk version 20200816`). The pass-14 harness had already run this +fence against the post-change fixture and it exited 0 with clean stderr, which it could not have +done if the range ran to EOF. + +**Dismissed as stated.** The *remedy* is still worth taking: `\+` inside an ERE is undefined +behaviour in POSIX, so `[+]` is unambiguous and costs nothing — applying it as portability +hardening, not as the bug Codex described. + +## Finding 2 — valid, reproduced, and the real defect of this pass + +The `finding`-column extraction uses `sed 's/\\|/\001/g' … sed 's/\001/|/g'`. In `sed`, `\001` in +the replacement is the literal text `001`, not a control byte, so the restore rewrites **every** +`001` in the column. + +Reproduced against a row whose `finding` is `case 001 and a pipe \| here`: + +``` +extracted: [ case | and a pipe | here ] +``` + +`case 001` became `case |`. So a fragment match can reject a valid locator or accept one that is +not in the `finding` at all. This is the pass-13 fragment fix — right in design, wrong in one +mechanical detail — with the detail one layer deeper than last time. My fixtures used a `finding` +containing neither `\|` nor `001`, so nothing exercised it. Fix: extract column 4 in one +escaped-pipe-aware `awk` pass, with no textual sentinel, and add fixtures containing both. + +## Finding 3 — valid + +The dates fence's `grep | cut | tr` producer is unchecked, so an unreadable ledger leaves an +empty temp file and the fence exits 0 while emitting a grep error. It reports chronology success +without reading its input. (The harness's stderr rule would catch the *error*, but the fence +itself must fail.) + +## Read + +Two real findings and one misdiagnosis. Both real ones are in the same fence, and both are the +now-familiar shape: correct design, one mechanical detail that makes the check pass when it +should fail. The design questions are closed — nothing in this pass touched a settled decision, +the convention prose, the story, or any claim about what the gates prove. diff --git a/.context/codex-reviews/gate-a-spec-pass-14.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-14.pre-2026-08-14.md new file mode 100644 index 0000000..5f42919 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-14.pre-2026-08-14.md @@ -0,0 +1,4 @@ +BLOCKER | high | §6 Check 1 “The existing first paragraph is untouched” | `para` uses the awk ERE `/date + anchor\)\.$/`, where `+` quantifies the preceding space instead of matching the literal plus in `date + anchor).` | the range never terminates at the paragraph, so it checksums through EOF and a correct change that adds the convention after the paragraph necessarily fails | match the plus literally, for example with `/date [+] anchor[)]\.$/`, and rerun the extracted fence under both shells against the post-change fixture +MAJOR | high | §6 Check 1 “No inert entry” finding-column extraction | the claimed `\001` placeholder is ordinary text in sed: the first substitution produces the bytes `\001`, and the restore expression replaces every `001` sequence in the finding, including unrelated text | a finding such as `case 001` is transformed to `case \|`, so fragment matching can reject a valid locator or accept a fragment that is not in the finding | extract column 4 with an escaped-pipe-aware awk routine that does not use a colliding textual sentinel, then add pass/fail fixtures containing both `\|` and `001` +MINOR | high | §6 Check 1 “Dates never decrease down the table” | the `grep \| cut \| tr` producer is unchecked, so a missing or unreadable `docs/hardening-log.md` leaves an empty temp file and the fence exits 0 while emitting a grep error on stderr | the standalone check can report chronology success without reading its input, contrary to the required guarded failure paths and clean stderr | guard the ledger read and producer status explicitly, or perform extraction and ordering in one awk invocation whose nonzero status is propagated +END OF FINDINGS (3 total) diff --git a/.context/codex-reviews/gate-a-spec-pass-15-dispositions.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-15-dispositions.pre-2026-08-14.md new file mode 100644 index 0000000..f888e06 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-15-dispositions.pre-2026-08-14.md @@ -0,0 +1,35 @@ +# Gate A — spec — pass 15 dispositions + +Artifact: `docs/superpowers/specs/2026-08-05-hardening-ledger-supersession-design.md` +4 findings: 3 Major, 1 Minor. **No Blocker.** All four valid. **None applied.** + +Reviewer: `mcp-codex-dev@1.0.1`, `codex-cli 0.144.5`, model `gpt-5.6-sol`. No stray artifact. + +Passes 4 → 15: **11, 14, 16, 15, 9, 10, 5, 12, 8, 6, 3, 4.** + +## Reproduced + +| # | Claim | Result | +|---|---|---| +| 1 | `finding()` mishandles an escaped backslash before a pipe | **Confirmed.** A row whose `finding` ends `…backslash \\` followed by the column pipe extracts as `[ends with backslash \\| next-col-value]` — the delimiter vanishes and the next column is swallowed. | +| 2 | candidate rows are never validated as complete rows | **Confirmed.** A truncated line `\| 2026-07-20 \| tt \|` matches `^\| RD \| FP \|` and counts as a match in the fragmentless branch: `n=1`, want 0. | +| 4 | the precheck story's count contradicts itself | **Confirmed.** `:64` says "Seven — six paired … and one inherited"; `:101` says "the six questions above". | + +| # | Sev | Verdict | Note | +|---|---|---|---| +| 1 | Major | **Valid** | The scanner treats *any* backslash before a pipe as escaping it, without checking whether the backslash is itself escaped. Fix: count the parity of the consecutive backslash run and split on a pipe following an even-length run. Narrow in practice — it needs a `finding` ending in a literal backslash — but it is the same shape as the last two rounds, one layer deeper again. | +| 2 | Major | **Valid, and the more consequential of the two** | Nothing requires a candidate to *be* a table row. A truncated or malformed line matching the date+fingerprint prefix satisfies the gate, so an entry can be declared non-inert by a line that is not a row. Fix: assert the full row shape — a minimum count of unescaped delimiters — before counting, in **both** branches. | +| 3 | Major | **Valid** | §8's inventory omits the §2.2 calibration that append-day truth is unverifiable and the check enforces only non-decreasing recorded dates. Delete that from both surfaces and all 33 anchors stay green while a bounded claim becomes an overclaim. **Seventh consecutive pass** on this inventory. | +| 4 | Minor | **Valid, confirmed** | My phantom-hardening edit widened that story's §5 without touching §6's count. | + +## Read + +No design finding. Nothing touched a settled decision, the convention prose, the governing story's +amendments, or any claim about what the gates prove. Findings 1 and 2 are both inside the one +`finding()`/candidate-matching routine; 3 is the inventory; 4 is a stale numeral. + +**On the inventory (finding 3).** Seven passes, seven holes, each a decision that turned out to be +anchorable. The rule-plus-list form did not stop it, and neither did "re-derive it" as an +instruction. The honest options are to stop claiming the list is complete, or to derive it +mechanically — every `**bold**` decision sentence in the shared region that is not a substring of +some anchor. That is a change to what §8 claims, so it goes to Daniel rather than being applied. diff --git a/.context/codex-reviews/gate-a-spec-pass-15.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-15.pre-2026-08-14.md new file mode 100644 index 0000000..73addb5 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-15.pre-2026-08-14.md @@ -0,0 +1,5 @@ +MAJOR | high | §6 Check 1 `finding()` | the scanner treats every backslash immediately before a pipe as escaping that pipe, without considering whether the backslash is itself escaped | a finding containing literal `\\\|` makes the delimiter disappear, merges the following column into `finding`, and can let a fragment found only outside column 3 satisfy the inert-entry gate | track the parity of each consecutive backslash run, split on a pipe after an even-length run, and add a fixture whose sought fragment appears only in the following column after `\\\|` +MAJOR | high | §6 Check 1 diff-scoped inert assertion | candidate rows are never validated as complete table rows: the fragmentless branch increments for `\| date \| fingerprint \|`, and `finding()` accepts `\| date \| fingerprint \| finding \|` despite all later columns being absent | an entry can be declared non-inert by a truncated or non-row line, so the gate can pass when no ledger row actually matches | require the expected full row shape and number of unescaped delimiters before counting a candidate, for both fragmentless and fragment-bearing entries, and add fewer-than-four-pipes and trailing-pipe-only rejection fixtures +MAJOR | high | §8 "What check 2's anchor list does not cover" | the claimed complete unanchored inventory omits the shared §2.2 decision that append-day truth is unverifiable and the check enforces only nondecreasing recorded dates | deleting that calibration from both prompt surfaces leaves all thirty-three anchors and both sentinels green while turning an explicitly bounded enforcement claim into an overclaim, contrary to invariant 11 and the gate-claims Don't | add an anchor for `Nothing can verify the append day itself` or `weaker property that dates never decrease`, update the stated count, or list this decision as unanchored +MINOR | high | `docs/superpowers/stories/2026-08-04-harden-finding-guard-scope-precheck-story.md` §6 | the suggested-size rationale still says "the six questions above" after §5 was widened to seven questions | the governing handoff story contradicts its own current inventory and can make a later designer omit the inherited supersession question | change `six` to `seven` +END OF FINDINGS (4 total) diff --git a/.context/codex-reviews/gate-a-spec-pass-16-dispositions.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-16-dispositions.pre-2026-08-14.md new file mode 100644 index 0000000..bde4898 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-16-dispositions.pre-2026-08-14.md @@ -0,0 +1,46 @@ +# Gate A — spec — pass 16 dispositions + +Artifact: `docs/superpowers/specs/2026-08-05-hardening-ledger-supersession-design.md` +7 findings: 5 Major, 2 Minor. **No Blocker.** All seven valid. **None applied.** + +Reviewer: `mcp-codex-dev@1.0.1`, `codex-cli 0.144.5`, model `gpt-5.6-sol`. No stray artifact. + +Passes 4 → 16: **11, 14, 16, 15, 9, 10, 5, 12, 8, 6, 3, 4, 7.** + +## Verified + +| # | Claim | Result | +|---|---|---| +| 1 | §7's rider still demands §8's list be "accurate and complete" | **Confirmed**, `:794`. The completeness claim was deleted from §8 and softened in §6's two cross-references — and left standing in §7, which calls itself "verbatim in every pass prompt". The deletion falsified a statement elsewhere and I did not sweep for it. | +| 3 | `pipes < 8` accepts *more* than seven columns | **Confirmed.** A 10-pipe line is accepted: I guarded the lower bound only, while the prose claims a full seven-column row. | + +| # | Sev | Verdict | Note | +|---|---|---|---| +| 1 | Major | **Valid, confirmed** | Mine. Also correct to report despite the "don't flag §8's incompleteness" instruction — this is a *wrong statement*, which that instruction explicitly still admits. | +| 3 | Major | **Valid, confirmed** | Fix: require exactly eight unescaped delimiters and a terminal one, rejecting every other shape. | +| 4 | Major | **Valid** | `[ -n "$F" ] \|\| continue` conflates "not a row" with "row whose `finding` is empty", so a fragmentless entry can be reported inert because an unrelated row has an empty `finding`. I named this risk in my own rider and shipped it anyway. Fix: return parse status separately from the value. | +| 5 | Major | **Valid** | I guarded the *base* read and left the *current-ledger* producer unguarded: a missing working-tree ledger emits stderr, yields an empty `new`, skips the loop and exits 0. | +| 6 | Major | **Valid** | Check 3 asserts `label < entry < Columns` but never that the convention's end sentinel precedes the label, nor that every entry lies inside that interval — so the block could sit above `# Hardening log`, or a second entry below the table, with all three checks green. It claims "the block is where §2.2 says, and **only** there". | +| 2 | Minor | **Valid** | §4's story row omits the pass-5 desired-outcome and pass-9 AC-2 amendments. | +| 7 | Minor | **Valid** | The residual I flagged in the prompt: check 1's row-count grep is a prefix match. Confirmed as real. | + +## The pattern this pass makes plain + +Findings by section across the last six passes have concentrated in the **validation scaffolding**, +not the design: + +``` +pass 11 12 findings §6/§7/Check-n mentions: 5 +pass 12 8 findings 9 +pass 13 6 findings 12 +pass 14 3 findings 6 +pass 15 4 findings 5 +pass 16 7 findings 13 +``` + +Pass 16: five of seven findings are in §6, and the two others are §4 and §7 bookkeeping. **§1–§3, +§5 and the convention prose have drawn no finding for four consecutive passes.** The design is +stable; the shell is not, and each hardening pass adds surface that the next pass finds defects in. + +Count trajectory since the artifact stabilised: **3, 4, 7** — oscillating, not converging, and the +rise tracks how much shell §6 has grown. diff --git a/.context/codex-reviews/gate-a-spec-pass-16.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-16.pre-2026-08-14.md new file mode 100644 index 0000000..376d22e --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-16.pre-2026-08-14.md @@ -0,0 +1,8 @@ +MAJOR | high | §7 "Gate-A riders, verbatim in every pass prompt" | rider 1 still requires §8's unanchored-decision list to be "accurate and complete", and the five stored riders are not the four riders in this pass | this directly reinstates the completeness claim the human deleted and makes the spec's stated review procedure contradict both §8 and the actual pass prompt | replace §7 with the current rider set and state only the exact anchor rule plus the deliberately partial list +MINOR | high | §4 `docs/superpowers/stories/2026-08-04-hardening-ledger-supersession-story.md` change-surface row | the authoritative row says the story gets three amendments and then names several other edits, but it omits the pass-5 desired-outcome amendment and the pass-9 AC-2 amendment that are present in the governing story | an implementer following the authoritative surface can miss two changes that remove contradictions with the settled convention, defeating the old-condition-accounting rule the row is meant to enforce | enumerate the pass-5 desired-outcome and pass-9 AC-2 amendments and their kept/changed/dropped accounting in §4 +MAJOR | high | §6 Check 1 `finding()` row-shape guard | `pipes < 8` rejects short rows but accepts rows with more than seven columns, an unescaped extra delimiter, or trailing content after an eighth delimiter while claiming to require a full seven-column row | a malformed line can still supply a nonempty finding and satisfy the no-inert-entry gate, recreating the false-green row-shape defect this pass is meant to close | require exactly eight unescaped delimiters, a terminal closing delimiter modulo allowed trailing whitespace, and reject every other shape before emitting the finding +MAJOR | high | §6 Check 1 `finding()` result contract and caller | a shape-valid row whose `finding` is empty or entirely whitespace produces the same empty command-substitution result as a malformed row, and `[ -n "$F" ] \|\| continue` discards both | a fragmentless entry locates rows by date and fingerprint, so it can be falsely reported inert solely because an unrelated finding field is empty; the parser has no way to return row validity independently of field contents | return an explicit parse-status marker separately from the finding value and let the fragmentless branch count every shape-valid row +MAJOR | high | §6 Check 1 diff-scoped inert assertion | the current-ledger `grep ... \| sort` producer and `comm -13` are unguarded, so a missing or unreadable current ledger can emit stderr, produce an empty `new` file, skip the loop, and exit 0 | the fence is required to execute standalone with guarded failures, but this failure path reports success for a check that never examined the changed ledger | validate current-ledger readability and guard extraction, sorting, and `comm` statuses without a pipeline that masks the failing producer +MAJOR | high | §6 Check 3 placement assertions | the checks assert only `label < target entry < Columns`; they never assert that the shared convention's end sentinel precedes the label or that every live entry lies inside that interval | all three checks can pass with the block above `# Hardening log`, or with a different valid entry below the table, despite Check 3 claiming the block is where §2.2 says and only there | locate the unique end sentinel, require `end < label < entry < Columns`, and reject every entry-shaped ledger line outside the unique label-to-Columns interval +MINOR | high | §6 Check 1 mechanical row count | the row-exists pattern is a prefix match, so a truncated `\| 2026-07-20 \| truncated-tool-output-read-as-complete \|` still yields the advertised row count of one | the later inert and byte-pin fences make the whole validation fail, but this fence's own stated cardinality result remains a false positive and can mislead a reader diagnosing which guarantee failed | make the count require the complete seven-column row shape or reuse the corrected row parser for the cardinality assertion +END OF FINDINGS (7 total) diff --git a/.context/codex-reviews/gate-a-spec-pass-17-dispositions.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-17-dispositions.pre-2026-08-14.md new file mode 100644 index 0000000..9fa98bc --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-17-dispositions.pre-2026-08-14.md @@ -0,0 +1,43 @@ +# Gate A — spec — pass 17 dispositions + +Artifact: `docs/superpowers/specs/2026-08-05-hardening-ledger-supersession-design.md` +6 findings: **1 Major**, 4 Minor, 1 Nit. **No Blocker.** All six valid. **None applied.** + +Reviewer: `mcp-codex-dev@1.0.1`, `codex-cli 0.144.5`, model `gpt-5.6-sol`. No stray artifact. + +Passes 4 → 17: **11, 14, 16, 15, 9, 10, 5, 12, 8, 6, 3, 4, 7, 6.** + +## The composition changed, which matters more than the count + +Pass 16: five of seven findings were in §6's shell. Pass 17, with the shell gone: **one Major**, +and it is a claim-versus-reality question about the design, not a mechanism defect. The other five +are a copied value, two miscounts, a wrong cross-reference, and markdown rendering — all mine, all +introduced by the restructure itself. + +No finding disputes a settled decision. No finding says a property is undecidable, a falsifying +observation is wired to pass, or an oracle is missing a distinction — which was rider 4's whole +purpose, and the first time that rider has come back empty. + +## Verified + +| # | Claim | Result | +|---|---|---| +| 5 | the backdating prohibition is in §2.2, not §2.1 | **Confirmed.** It sits at `:148`, inside §2.2 (which begins at `:125`). Both `:298` and `:446` cite §2.1. | +| 6 | anchors 3, 5, 8, 15, 18 render wrong | **Confirmed.** Each is wrapped in single backticks while containing backticks — e.g. `` 15. `block above the `Columns:`` `` renders with a stray trailing backtick. A reader copying the rendered list implements a weaker pattern than the prose requires. | +| 3 | "Five properties" precedes 1a–1f | **Confirmed** — six. | +| 4 | "four things" precedes five bullets | **Confirmed** — five. | + +| # | Sev | Verdict | Note | +|---|---|---|---| +| 1 | Major | **Valid — and it needs Daniel** | Two rows sharing `date` **+** `fingerprint` **+** `finding` cannot be separated: the only narrowing device is a fragment of `finding`, which they share. §2.2 says "there is no unresolvable state left for it to stop on" — that is an overclaim in exactly the class this review keeps finding, and it is now the *design* making it rather than a check. Consequence: if only one such row's narration is false, every permitted entry also marks its accurate sibling. Options: narrow the claim and record the limitation (consistent with how §8 is written), or add a discriminator, which means new locator syntax in the shared convention. | +| 2 | Minor | **Valid** | The opening says no profile value is copied into the spec; §6's "Mode" copies `battery+check` from the story header. A confirmed profile change would leave the spec stale while the header is authoritative. | +| 3 | Minor | **Valid, confirmed** | Mine, from the restructure. | +| 4 | Minor | **Valid, confirmed** | Mine — and precisely the coverage-count drift rider 1 exists to catch, committed while writing the rider. | +| 5 | Minor | **Valid, confirmed** | Two sites cite the wrong section for the precondition they rely on. | +| 6 | Nit | **Valid, confirmed** | Mine. The anchors are declared as literal data and five of them do not render as their own content. | + +## Read + +Five of six are artifacts of the restructure — the cost of moving 226 lines, and all cheap. The +one that isn't is a real limit on the locator, surfaced only once the shell stopped drowning it +out. diff --git a/.context/codex-reviews/gate-a-spec-pass-17.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-17.pre-2026-08-14.md new file mode 100644 index 0000000..622540d --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-17.pre-2026-08-14.md @@ -0,0 +1,7 @@ +MAJOR | high | §2.2 "Why uniqueness is not required" | the locator cannot express a correction to only one of two rows that share the same date, fingerprint, and finding, yet the spec claims there is no unresolvable state and says an author who wants one row can narrow the locator | if only one duplicate row's ref or other narration is false, every permitted entry also marks the accurate sibling, so the intent that a falsified row can be marked without falsifying another row is not achievable under an allowed ledger state | add a discriminator capable of separating identical findings, or explicitly narrow the supported scope and replace the no-unresolvable-state claim with this limitation +MINOR | high | opening profile authority statement and §6 "Mode" | the opening says no profile value is copied into the spec, but §6 literally copies `battery+check` from the story header | a later confirmed profile change can leave the spec carrying a stale validation mode despite the governing-story and read-fresh rules | remove the copied mode value and require the executor to resolve it from the fresh story header, or narrow the opening claim and define how parity with the authoritative header is maintained +MINOR | high | §6 Check 1 opening | "Five properties" is followed by six required items, 1a through 1f | an execution plan can reasonably implement only the advertised five and omit the named read, weakening the `battery+check` validation while still appearing to follow the spec | say six required properties, or separate 1f from the numbered property count and state explicitly that it must also run +MINOR | high | §6 Check 1d oracles | the text promises four oracle distinctions but lists five: diff scope, row shape, escaped delimiters, parse-failure versus empty finding, and read failure | this stated count obscures whether one distinction is accidental or optional and recreates the coverage-count drift the rider is meant to catch | change the count to five and keep all five bullets required +MINOR | high | §3.1 fragmentless entry and §6 Check 1e | both passages attribute the backdating prohibition to §2.1, but the prohibition and recorded-order floor are defined in §2.2 | the wrong cross-reference sends implementers and reviewers to a section that does not establish the date-bound precondition being relied on | change both references from §2.1 to §2.2 +NIT | high | §6 "The anchors" items 3, 5, 8, 15, and 18 | the inline-code delimiters are nested incorrectly, so CommonMark renders the literal backticks outside or missing from the anchors; item 15 even renders a stray trailing backtick | the anchors are declared as literal data, and a reader copying the rendered list can implement weaker or wrong presence patterns despite the prose requiring exact anchors | wrap backtick-bearing anchors in correctly delimited longer code spans, using spaces where the anchor itself begins or ends with a backtick +END OF FINDINGS (6 total) diff --git a/.context/codex-reviews/gate-a-spec-pass-18-dispositions.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-18-dispositions.pre-2026-08-14.md new file mode 100644 index 0000000..e9408d0 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-18-dispositions.pre-2026-08-14.md @@ -0,0 +1,37 @@ +# Gate A — spec — pass 18 dispositions + +Artifact: `docs/superpowers/specs/2026-08-05-hardening-ledger-supersession-design.md` +6 findings: 4 Major, 2 Minor. **No Blocker.** All six valid. **None applied.** + +Reviewer: `mcp-codex-dev@1.0.1`, `codex-cli 0.144.5`, model `gpt-5.6-sol`. No stray artifact. + +Passes 4 → 18: **11, 14, 16, 15, 9, 10, 5, 12, 8, 6, 3, 4, 7, 6, 6.** + +## Verified + +| # | Claim | Result | +|---|---|---| +| 1 | the narrowed limitation is still too narrow | **Confirmed by construction.** Enumerating every substring of `foo`: **none** fails to also match `foobar`. So two rows with *non-identical* findings can still be inseparable. My narrowing named the wrong condition. | +| 2 | §4's end sentinel is not literal text | **Confirmed.** `:335` quotes `` `…and nothing checks the difference.` `` — with a literal `…` **inside** the code span. The convention text at `:162` has no ellipsis. A check taking the documented sentinel literally can never match. | +| 6 | anchor 15 renders with a trailing space | **Confirmed.** CommonMark strips padding only when the span has **both** a leading and a trailing space. ` ``block above the `Columns:` `` ` renders as `block above the \`Columns:\` ` — trailing space included. | + +| # | Sev | Verdict | Note | +|---|---|---|---| +| 1 | Major | **Valid — and my pass-17 fix under-corrected** | The real condition is not "identical in date, fingerprint and `finding`" but "**the target row has no permitted distinguishing fragment**" — which also covers a `finding` that is a substring of its sibling's, and pairs where the only separating text contains a double quote (barred by the no-double-quote rule). Both the limitation and its reopen trigger need generalising. Daniel settled the *shape* (record a limitation, no new syntax); this is that limitation stated correctly, so I read it as inside the settlement rather than reopening it — flagging in case he disagrees. | +| 2 | Major | **Valid, confirmed** | An editorial ellipsis inside a code span that is quoted as literal sentinel text. Present since the sentinel was introduced and missed by every prior sweep, including the ones that checked the sentinel was "unchanged". | +| 3 | Major | **Valid** | Story AC 4 promises a reader can determine *which* claims no longer hold; the design supplies only *recorded* corrections, and §2.1 explicitly leaves a removed hardening unmarkable. The desired outcome was narrowed at pass 5; AC 4 was not brought with it. | +| 4 | Major | **Valid** | Check 1f's named read covers "names the claim that does not hold" and the citation resolving, but not §2.2's other two requirements: saying **which cause** (stopped holding vs never true) and **citing without restating**. An entry violating either passes every stated check. | +| 5 | Minor | **Valid** | §2.1 says the hardening "still stands" and that ejecting the row would misreport class history — but for a wrong-fingerprint or phantom row the lineage **already** misreports it. The design knowingly preserves that; the rationale claims the opposite. | +| 6 | Minor | **Valid, confirmed** | My pass-17 rendering fix was itself subtly wrong. | + +## Read + +Two of six (1 and 6) are cases where **my fix from the previous pass was subtly wrong** — the +pattern continues, at diminishing severity: pass 17's version of each was closer than pass 16's, +and this one closer still. None of the six is a mechanism defect and none touches §6's oracles; +rider 4 came back empty for the second consecutive pass. + +Findings 1, 3 and 5 are all one shape: **a claim about the design that is not exactly true** — +a limitation stated too narrowly, an acceptance criterion promising more than the design delivers, +and a rationale asserting the opposite of a case the design knowingly accepts. That is the class +worth spending passes on, and it is now the only class left. diff --git a/.context/codex-reviews/gate-a-spec-pass-18.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-18.pre-2026-08-14.md new file mode 100644 index 0000000..bded8d6 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-18.pre-2026-08-14.md @@ -0,0 +1,7 @@ +MAJOR | high | §2.2 "Why uniqueness is not required" and §8 identical-rows limitation | the claim that only rows identical in date, fingerprint, and `finding` are inseparable is false | two nonidentical findings can also be inseparable: if one finding is `foo` and the other is `foobar`, every permitted fragment of `foo` also matches `foobar`, and the no-double-quote rule creates further nonidentical cases; the stated limitation and reopen trigger therefore miss real states in which an accurate sibling must also be marked | generalize the limitation to any pair for which the target row has no permitted distinguishing fragment, and trigger reopening on the first such pair whose truth diverges +MAJOR | high | §4 "Parity, with sentinels a check can locate" | the named end sentinel is not literal text in the convention | §4 names ``…and nothing checks the difference.`` inside a code span, but the shared prose contains `and nothing checks the difference.` without the ellipsis character, so a check treating the documented sentinel literally can never pass after the change | quote an exact suffix such as `and nothing checks the difference.` or the full final line, with no editorial ellipsis inside the code span +MAJOR | high | governing story §3 AC 4 and design §4 Pass 5 amendment account | AC 4 still promises complete knowledge of which claims no longer hold, while the design supplies only recorded corrections and explicitly leaves removed hardenings unmarkable | a row can have an unrecorded false claim, and a removed hardening has no sanctioned entry at all, so the ledger cannot satisfy the authoritative story criterion even though the desired outcome was narrowed to “have been recorded” | narrow AC 4 to the claims recorded as no longer holding and account explicitly for preserving the removed-hardening exclusion +MAJOR | high | §6 Check 1f "a named read" | the human oracle omits two required distinctions in the actual entry: cause classification and cite-don't-restate | an implementation can append an entry that never says whether the claim stopped holding or was never true, or that copies the current answer into the marker, and still satisfy every stated check; both violate §2.2 and recreate the stale narration risk | require the named read to confirm the entry states the applicable cause and cites without restating the current answer +MINOR | high | §2.1 "Why supersession leaves recurrence alone" and §8 mechanical-identity limitation | the rationale says a superseded hardening still stands and that ejecting it would misreport class history even for wrong-fingerprint and phantom-hardening rows | when the fingerprint was wrong or the hardening never existed, the mechanical lineage already misreports the class or rung history; the design deliberately preserves that false lineage pending the reader redesign, but its rationale states the opposite | limit the “hardening still stands” rationale to later-stale narration, and describe wrong-fingerprint and phantom cases as knowingly preserved mechanical misclassification rather than accurate history +MINOR | high | §6 anchor 15 | the double-backtick code span renders with a trailing literal space | CommonMark strips one padding space only when the code-span content has both a leading and trailing space; ``block above the `Columns:` `` has no leading padding, so Pandoc renders `block above the `Columns:` ` rather than the intended literal ending at the backtick | write `` block above the `Columns:` `` so CommonMark removes both padding spaces and preserves the terminal backtick without adding content +END OF FINDINGS (6 total) diff --git a/.context/codex-reviews/gate-a-spec-pass-19-dispositions.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-19-dispositions.pre-2026-08-14.md new file mode 100644 index 0000000..f0061e0 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-19-dispositions.pre-2026-08-14.md @@ -0,0 +1,52 @@ +# Gate A — spec — pass 19 dispositions + +Artifact: `docs/superpowers/specs/2026-08-05-hardening-ledger-supersession-design.md` +(plus the governing story, which finding 3 lands in). +5 findings: 4 Major, 1 Minor. **No Blocker.** All five valid. **None applied** — held for Daniel +under the pinned exit ("anything new comes to Daniel"). + +Reviewer: `mcp-codex-dev@1.0.1`, `codex-cli 0.144.5`, model `gpt-5.6-sol`. Pass file valid: +6 lines, 5 finding lines, terminator exact, no stray artifact. + +Passes 4 → 19: **11, 14, 16, 15, 9, 10, 5, 12, 8, 6, 3, 4, 7, 6, 6, 5.** + +## Premises verified mechanically before judging + +| Claim | Result | +|---|---| +| the resolution floor is stated for **rows** only | **Confirmed.** §2.1:49 says "Drafting before a row exists… is below the rule's resolution". §2.2:161 says "Entries are never edited, never removed" and states **no** floor of its own. | +| the shipped prose promises narrowing **to one** | **Confirmed verbatim.** §2.2: "To narrow it to one, add `""`… pick one containing no double quote." No path is given for a row that has no permitted fragment. | +| 1d's five oracles never define entry recognition | **Confirmed.** The five are: which entries are in scope · a row from a non-row · a delimiter from an escaped pipe · parse failure from an empty field · a read failure from a clean pass. Four of the five specify **row** parsing; none defines what a well-formed *entry* is. | +| the ledger's dates are non-decreasing today | **Confirmed.** 22 rows, no decrease — so 1e passes on the current file, and finding 1's backdating half is a future-state claim, not a present failure. | + +## Verdicts + +| # | Sev | Verdict | Note | +|---|---|---|---| +| 1 | Major | **Valid — and it is a design gap, not a claim** | Two irreparable states. (a) An entry written with a mistyped locator is, by the letter of §2.2, an entry the moment it exists in the file, and entries are never removed — so appending the correction leaves the inert one inside 1d's *against-base* scope and 1d ("no entry this change appends is inert") can never pass. The convention **sanctions** that state ("append a new entry with a locator that matches, and leave the inert one standing as history"); only §6 refuses it, so §6 is what is wrong. (b) A backdated row is likewise unrepairable and 1e can never be restored — true, and out of this change's reach, since this change appends no row. | +| 2 | Major | **Valid** | The overstatement is in the **shipped** convention prose, so it propagates into every scaffolded ledger while the limitation that qualifies it stays behind in the spec's §8. An author meeting the no-permitted-fragment case is told to do something the design says is impossible, and given no fallback. Inside the pass-17 settlement (record a limitation, add no locator syntax) — it records the limitation where the instruction lives. | +| 3 | Major | **Valid — the third instance of AC 4 promising more than the design delivers** | Pass 1 narrowed it, pass 5 brought §2's desired outcome into line, pass 18 narrowed it to *recorded* corrections and named the removed-hardening exclusion. This is a fourth gap in the same criterion: where siblings are inseparable, one fragmentless entry marks an accurate row too, so a reader determines something **false** about that row — a wrong answer, not a missing one, which is worse than the case pass 18 closed. Needs a **story** amendment, and story amendments in this cycle are human-confirmed. | +| 4 | Major | **Valid — and the first oracle-coverage finding in three passes** | Rider 4 came back empty at passes 17 and 18; this is a real missing distinction. 1d quantifies over "every entry added by the change" and never says how an added entry is **recognised**, so a checker that silently drops a malformed entry-shaped line from its candidate set reports "every recognised entry is non-inert" and establishes less than the property claims — the read-failure-from-clean-pass mistake, one level up, on the half of the match the oracles never describe. | +| 5 | Minor | **Valid as a sharpening, premise partly overstated** | My wording is *"a pair separated **only** by text containing a double quote"*, and Codex's counterexample (`alpha"x` vs `alpha"y`) is separated by `x` and `y`, so the clause correctly does not apply to it. But the phrasing is pair-level and sits inside a condition I had just made per-row, which is the inconsistency worth fixing. Codex's directional example (`foo` vs `foo"`) is better: the longer row's only unique substrings all carry the quote. | + +## Read + +**The termination rule does not fire.** It says stop if pass 19's findings are *only* +claim-sharpening. Findings 3 and 5 are that class. **Findings 1, 2 and 4 are not**: one is a +validation property the convention's own text contradicts, one is an overstatement in the +**deliverable prose** rather than in the spec's narration about itself, and one is a missing +oracle distinction after two empty rider-4 passes. Two of the three would survive into the plan +and the diff. + +**Suggested fixes, if Daniel says continue:** +1. Restate 1d's bound from "no entry this change appends is inert" to "**the entry §3.1 + specifies** matches at least one row dated on or before its date" — the property this change + actually needs, decidable, and consistent with §2.2 sanctioning inert entries. Record in §8 the + two irreparable states (a landed inert entry; a backdated row) as residuals of absolutism. +2. Rewrite the shared sentence: a fragment narrows the **match set**, and singles out one row only + where a permitted distinguishing fragment exists; otherwise the entry necessarily marks every + matching sibling. Shared prose → lands in both surfaces; re-run the 33-anchor sweep after. +3. Extend AC 4 to the rows the convention can distinguish, with old-condition accounting. +4. Add a sixth oracle to 1d: bound the block, define complete entry syntax (locator, optional + fragment, date), and separate "no added entries" from "an added entry that could not be parsed". +5. Replace the double-quote example with the directional one, per-row. diff --git a/.context/codex-reviews/gate-a-spec-pass-19.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-19.pre-2026-08-14.md new file mode 100644 index 0000000..d4aaae7 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-19.pre-2026-08-14.md @@ -0,0 +1,6 @@ +MAJOR | high | §2.2 inert-entry rule and §6 checks 1d/1e | the absolute no-edit/no-remove rules make the checks' own failure states irreparable: after a mistyped inert entry is written, appending its correction leaves the inert line inside 1d's against-base scope forever, and after a backdated row is written no later append can restore non-decreasing order | a check can permanently strand the implementation in a state that cannot pass without performing the forbidden edit or removal | define and account for a pre-entry/pre-row candidate state with validation before atomic append, or explicitly permit recovery for newly written invalid records and revise the absolute-rule claims accordingly +MAJOR | high | §2.2 shared convention, sentence beginning "To narrow it to one" | the shipped prose says adding a fragment narrows a locator to one and tells the author to pick such a fragment, while the same section later establishes that some rows have no permitted distinguishing fragment and can never be narrowed to one | downstream scaffolded ledgers receive the overstrong instruction but not the limitation or its no-solution path, so an author cannot apply the convention correctly in the acknowledged edge case | change the shared sentence to say a fragment narrows the match set and singles out one row only when a permitted unique fragment exists, then state that otherwise the entry necessarily marks every matching sibling +MAJOR | medium | governing story §2 and AC 4 versus spec §8 "A row with no permitted distinguishing fragment" | the story unconditionally promises that a reader can determine which recorded claims and current-answer citations belong to any row, but §8 admits that siblings can be inseparable and one fragmentless entry then marks accurate and falsified rows alike; divergent corrections or citations cannot be represented per row under latest-wins | the design has an acknowledged reachable state in which it cannot satisfy its governing outcome, even though the current ledger has no instance | either add a discriminator that can represent divergent sibling corrections or narrow the story outcome and AC 4 explicitly to rows distinguishable by the convention, with old-condition accounting +MAJOR | high | §6 check 1d oracles | the property quantifies over every entry added by the diff, but its five oracles never define how an added entry is recognized and parsed or require an entry-shaped but malformed added line to fail; they specify row parsing in detail while leaving the other half of the match undefined | a checker can silently omit a malformed or partially written entry from its candidate set and report that every recognized entry is non-inert, establishing less than the property claims | add an oracle that bounds the supersession block, defines complete entry syntax including the optional fragment and date bound, and distinguishes no added entries from an unreadable or unparseable added-entry candidate +MINOR | high | §2.2 limitation rationale and §8 limitation bullet | "a pair separated only by text containing a double quote is barred" is too broad: findings such as `alpha"x` and `alpha"y` still have permitted distinguishing fragments `x` and `y`; the bar exists only when every distinguishing substring for the row contains a double quote | the example misstates the otherwise-correct per-row necessary-and-sufficient condition and can make readers classify separable rows as inseparable | replace it with a directional example such as `foo` versus `foo"`, where the longer row's only unique substring contains the forbidden quote, and retain the per-row asymmetry +END OF FINDINGS (5 total) diff --git a/.context/codex-reviews/gate-a-spec-pass-2-dispositions.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-2-dispositions.pre-2026-08-14.md new file mode 100644 index 0000000..cfa7109 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-2-dispositions.pre-2026-08-14.md @@ -0,0 +1,18 @@ +# Gate A — spec — pass 2 dispositions + +Artifact: `docs/superpowers/specs/2026-08-05-hardening-ledger-supersession-design.md` +11 findings: 2 Blocker, 6 Major, 3 Minor. All eleven validated as correct; all eleven applied. + +| # | Verdict | Reason | +|---|---|---| +| 1 | **Applied** | Correct, and the one pass 1 could not see because pass 1's spec said the existing paragraph was *kept*. `Never edit a row` left absolute contradicts §2.1's in-place amend, on the same surface — item 7. New §2.0 amends the phrase to `Never edit a landed row`. | +| 2 | **Applied** | Correct. §2.2's shared prose opened with `**Superseded rows.**`, which is also the block's label — so an empty scaffold shipped the label and the repo would carry two. Prose relabelled `**Correcting a landed row.**`; `**Superseded rows:**` now belongs to the block alone. | +| 3 | **Applied** | Correct, twice over. The decidability sentence said absence at the base proves landed — backwards. And presence/absence is not symmetric: a row can arrive from another cycle by merge after the base. Rewritten as a one-way test that only ever confirms landed. | +| 4 | **Applied** | Correct: "the cycle now open" was undefined. §2.1 now names cycle identity (the Gate-B cycle of §5 Mechanics), its close event (the commit replacing the `WIP:` snapshot), the no-open-cycle state (everything is landed), and concurrency (another worktree's open cycle is landed to you). | +| 5 | **Applied** | Correct and the sharpest finding of the pass: the entry restated 0.8.0's counting rules, which is the exact artifact shape that goes stale — and would have made the first entry the next thing needing supersession. Entry now names only what stopped holding and cites. Added as a standing clause in §2.2, not just fixed in place. | +| 6 | **Applied** | Correct gap introduced by pass 1's own fix. Entry-correction had no locator, shape or precedence. An entry is now located by its date plus the row it supersedes; a correcting entry retires exactly what it names and states what now holds, so nothing is restored implicitly. | +| 7 | **Applied** | Correct and subtle. Check 2 normalized whitespace, which would erase the four-space indent that is the only thing separating the format example from a live entry — the check could pass on a surface that had converted the example into entry-shaped content. Normalization now covers hard-wrap joins only; indentation compared exactly. | +| 8 | **Applied** | Correct: the greps establish cardinality, not content, so AC 4 could fail with green evidence. Check 1 split into a mechanical half and a named read, with §9 stating plainly that the content half is a human read nothing validates. | +| 9 | **Applied** | Correct — verified independently: `CLAUDE.md` has two paragraphs opening "What this does not do" (lines 168 and 390). The entry now carries a unique opening fragment. | +| 10 | **Applied** | Correct: "through the end of §2.2's prose paragraph" is not locatable in the target files, which carry no §2.2 marker. §4 now names two unique line sentinels, inclusive, and states that the indented example is inside the region. | +| 11 | **Applied** | Correct. The prohibition is *kept* — the replacement still states it — and only the *definition* of landed moves. The table said "moved", which obscured that the prohibition now sits on two surfaces. Corrected, and the duplication is named in §5 and again in §9. | diff --git a/.context/codex-reviews/gate-a-spec-pass-2-dispositions.stopped-2tier-debt.md b/.context/codex-reviews/gate-a-spec-pass-2-dispositions.stopped-2tier-debt.md new file mode 100644 index 0000000..6adeb8c --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-2-dispositions.stopped-2tier-debt.md @@ -0,0 +1,115 @@ +# Gate A — spec — pass 2 dispositions (two-tier cycle) + +34 findings (8 BLOCKER, 20 MAJOR, 6 MINOR). None dismissed. Enum held a fifth time. + +## Trajectory + +| Pass | B | M | m | Spec lines | +|---|---|---|---|---| +| 1 | 7 | 21 | 2 | 349 | +| 2 | 8 | 20 | 6 | 538 | + +Blocker count is flat-to-up, but the **kind** changed: pass 1's blockers said the +compensating controls did not exist; pass 2's say the controls I added have specific, +namable holes. Four of the eight are outright errors of mine with obvious fixes. That is +movement, not the wall the three-tier cycle hit — where each pass found the mechanism more +impossible than the last. + +## Blockers — mine, fixable, taken + +**2. My §11 matrix contradicts my §13.** The matrix says self-authorization, same-approver +acceptance, a second waiver and stale confirmation are "refused"; §13 says nothing is +enforced and an agent can write a conforming block. Both are mine. **Fix:** the matrix +describes **observable manual behaviour** — "the procedure refuses" — never a mechanism, and +drops "no decision block is producible", which is false. + +**3. My counterfactual is backwards.** Under the prior policy tier 3 does not exist, so a +compliant agent refuses *every* tier-3 closure and the advisory hook permits the commit both +before and after. Nothing makes the old state uniquely succeed. GOES TO DANIEL — §5 calls an +unobservable counterfactual a **blocking evidence gap** requiring a logged mode override, and +that is his call, not mine. + +**8. Pinned SHAs are immutable but not durable.** Story-commit and Gate-A blob SHAs taken +from WIP or pre-squash history become unreachable after amend/squash and are collectable; +and if the profile changes in the closing amend, embedding that commit's own SHA in its own +row is self-referential. **Fix:** anchor identity in `main` and make the row carry a +self-contained snapshot of what it must repay, rather than a reference that can evaporate. + +**9. The inventory is still thematic — second occurrence.** I rebuilt it to 60 numbered rows +in document order and it is *still* a compression: the named omissions include the hook's +spec/plan indistinguishability and reset timing, full companion and resume-note lifecycles, +fresh-rerun deletion scope, result-envelope residuals, exact Gate-A prompt inputs, the +Gate-B path classifier, and WIP/`baseSha` mechanics. **Fix:** derive from each operative +sentence and list item, not each theme. Noting for the record that I have now claimed +exhaustiveness twice and been wrong twice. + +## Blockers — real holes, fixes taken + +**4. Tier 3 must not waive the profile's evidence obligations.** Nothing currently requires +the battery, the check, the named verification, `+abuse-path`, or a revalidated evidence +entry before a waiver. **Fix:** tier 3 waives **reviewer passes only**; every mode-derived +obligation stays a precondition, with refusal cases for missing, stale or inadequate +evidence. + +**5. `todos.md` becomes authoritative gate state while staying prose-exempt.** Debt rows can +be deleted, edited or falsely closed in a commit Gate B never sees — and my §8 row 43 marks +prose-vs-product classification "Kept, untouched" while doing exactly that. **Fix:** either +the debt record moves somewhere gate-classified, or the classifier disposition changes +honestly. Both are real options; taking the second and marking row 43 **narrowed**. + +**6. Two handles can alternate indefinitely.** Approver waives, accepter cancels the debt, +next waiver is available again — unbounded zero-pass cycles with every rule followed. +**Fix:** a cumulative cap on consecutive waivers whose reset requires an actual completed +tier-1 repayment. The "different handle" rule alone creates this, and it was my fix for +pass-1 finding 3. + +**7 rides with it. ACCEPT** — the `reconciled` state escapes a block phrased for "open or +unreconciled", and post-land reconciliation *creates* that state. Three states: open- +unreconciled, open-reconciled, closed; refuse tier 3 for anything but closed. + +## Blocker 1 — goes to Daniel + +Every compensating control is authored by the agent whose work is being waived. Daniel +decided this knowingly (procedure + disclosure; enforcement parked with a trigger), so a +bare re-raise would not reopen it. **Pass 2 brings new information that he was not given:** +blockers 5 and 6 show the *debt* — the control his decision leaned on — is itself +bypassable in two specific ways, by editing a prose-exempt file and by alternating handles. +The question is therefore no longer "is prose enforceable" but "does the compensating +control survive contact with the abuse path". + +## Major — accepted, fixes taken without further comment + +10 (`sparring-briefing`'s rule also forbids treating a satisfied human as a substitute for a +clean pass — that clause did **not** move with tier 2 and my accounting said it did), 11 +(zero-finding early exit is overturned for tier 3, not kept), 12 ("skip ONLY trivial" is +narrowed by effect regardless of calling it a waiver), 13 (the uncited-artifact path *is* +changed and row 45 denies it), 14 (marker cannot represent completed-passes-with-dispositions +— add pass ids and per-finding dispositions), 15 (merge strategy is a precondition with no +field to record it), 16 (no source category exists for the fresh-attempt branch), 17 (solo +case makes the different-handle rule impossible or alias-satisfiable), 18 (concurrent +branches can collide on cycle id and each pass the debt check), 19 (two separate clean +reviews never review the integrated result), 20 (Gate-A repayment is undefined once findings +require revising the artifact), 21 (a corrective commit cannot make the *closing* commit +contain the blocks — AC 2 names that commit), 22 (no precedence rule for the hook's STOP +against an authorized waiver), 23 (the companion cannot be written before the hook counts — +PostToolUse fires first; reword to "before the reader credits the pass"), 24 (the row is not +self-contained as claimed), 25 (calling a waived range `reviewed` is the gate-proof Don't in +one word), 26 (more falsified sites: `getting-started` line 84, `coding-workflow` 27–29, +91–99, 108–109, 123–126, README daily-use), 28 (a status page does not prove *this* account +cannot call), 33 (the pinned story SHA is recorded and never consumed by the repayment). + +## Minor — collected + +27 (the `coding-workflow` 149 sentence is about the **quality** gate, not the review gate — +my site row is wrong and would weaken an unrelated guarantee; mark untouched), 29 +(sanitization covers `Cause` but not the free-form reason field), 30 (the first waiver's debt +blocks the next, so tier 3 advances **one** cycle per outage — the story's "work continues" +oversells it), 31 (no equality rule binds the handle/date across the two blocks), 32 (the +enum-drift line and the per-finding dispositions grammar are undefined against each other), +34 (unstated whether an open debt blocks a normal tier-1 cycle or only tier-3 eligibility). + +## Status + +Not clean. Two go to Daniel — blocker 1 (new evidence against a settled decision) and +blocker 3 (a blocking evidence gap that only a logged mode override can resolve). Every +other fix is taken. diff --git a/.context/codex-reviews/gate-a-spec-pass-2-dispositions.stopped-3tier.md b/.context/codex-reviews/gate-a-spec-pass-2-dispositions.stopped-3tier.md new file mode 100644 index 0000000..cf15b3b --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-2-dispositions.stopped-3tier.md @@ -0,0 +1,125 @@ +# Gate A — spec — pass 2 dispositions + +Cycle: reviewer-availability fallback ladder. 39 findings (4 BLOCKER, 33 MAJOR, 2 MINOR). +Severity enum held again — no out-of-enum token, second live use of rider (b). + +## The pass-2 headline + +Pass 1 found holes in the *design*. Pass 2 finds that the **mechanism cannot execute as +written**, and that one premise I gave Daniel — which he then decided on — is false +against the shipped text. Findings 1, 2, 3 and 5 are the ones that matter; the rest are +mostly downstream of them or ordinary rigour. + +## Blocker + +1. ACCEPT — and self-inflicted in the pass-1 revision. I fixed the diff-truncation finding + by having the reviewer produce the diff itself from `baseSha`/`headSha`, while in the + same section limiting it to Read/Grep/Glob. None of those can compute a git range. Gate + B tier 2 is currently unexecutable. REOPENS the capability question. +2. ACCEPT — "write to exactly one path" is stated as a capability and is not one. A + prompt-only design has no path guard; Claude Code's Write tool grants general access. + The containment claim is false as written. REOPENS Q2 (prompt-only vs shipped agent + definition). +3. ACCEPT, and it is the sharpest finding of either pass — a non-fork subagent loads the + repository's CLAUDE.md hierarchy as instructions, so "only the wrapper supplies + authoritative instructions" is false before the reviewer ever reads the artifact. The + trust boundary I wrote does not exist. +4. ACCEPT — genuinely self-referential and mine: §6 requires the debt row inside the + snapshot before the final pass; §7 requires that row to name the landed SHA, which does + not exist until the commit is written and changes again under squash. Needs a stable + pre-land identity plus a post-land reconciliation step. + +## Major — the premise reversal + +5. ACCEPT, and it overturns a decision already taken. I told Daniel that `.off` always sat + outside invariant 2, and he chose to amend the invariant on that basis. CLAUDE.md's own + §5 says of `.off`: "**the gates still apply**". So `.off` has always silenced reminders + while leaving the gates in force; prescribing it as the posture for running *weaker + gates* is a substantive change to what it means, not a scope clarification. The + amendment as drafted is the disguised-exception shape the AGENTS.md Don't names. + MUST GO BACK TO DANIEL — his decision rested on my false premise. + +## Major — accepted, mechanism and procedure + +6. ACCEPT — `.off` is workspace-global; one outage authorization silences unrelated commits + and unrelated cycles. Prescribing a global sentinel as if it were cycle-scoped. +7. ACCEPT — no handling for a pre-existing `.off` (init-time, user opt-out, another cycle); + entry overwrites its meaning and exit deletes it unconditionally. +8. ACCEPT — no concurrency rule; one cycle's exit can re-enable another's workspace. +9. ACCEPT — nothing requires an availability re-check, so tier 2 becomes the standing + arrangement by never looking. +10. ACCEPT — the two entry predicates conflict: "vendor known unavailable" permits no call, + yet the recovery attempt must be spent. +11. ACCEPT — tier 3's "unavailable **or unwarranted**" makes the broadest fail-open tier + reachable through an uncheckable adjective. "Unwarranted" should go. +12. ACCEPT — tier 3 closes with no passes, yet §7 calls it a clean cycle and §9 demands a + per-pass line for it. Incoherent; tier 3 needs its own no-pass closure shape. +13. ACCEPT — the marker schema cannot truthfully represent tier 3 (no reviewer model, no + pass count, no final-pass cleanliness). Needs per-tier schemas. +14. ACCEPT — human authority is a prose block the agent can write. §15 admits it; the design + adds no live-confirmation step. Injection or agent error fabricates the sole control. +15. ACCEPT — entry authorization lives in an ignored file that is deleted on exit, so `main` + can show degradation was claimed but never that a human authorized it. +16. ACCEPT — the inherited recovery procedure is MCP-specific (`sessionId`, + `specSessionId`/`qualitySessionId`); tier 2 runs through a different surface entirely + and has no valid one-attempt procedure. Resuming may also break fresh-context isolation. +17. ACCEPT — §6's identity rule is scoped to tier 2, but rider (a) changes tier 1 too, so + tier-1 sequential pairs can still review different content. +18. ACCEPT — `git status --porcelain` empty output without checking exit status: a git error + reads as a clean tree. Exactly the false ✓ the section exists to prevent. +19. ACCEPT — endpoint clean-checks do not exclude a concurrent writer mutating and restoring + between them. +20. ACCEPT — tier-switch cleanup covers only a prior tier-1 task, not tier-2 agents or any + other writer to the slot. +21. ACCEPT — "each pass records baseSha and headSha" names no file, actor or acceptance + check, and the findings grammar forbids metadata lines. Needs a provenance record. +22. ACCEPT — rider (a) prose and its accounting table give two different recovery graphs. +23. ACCEPT — normalization vs story AC 6's "closed set of permitted tokens". Either make + non-enum INCOMPLETE or amend the story. MUST GO BACK TO DANIEL — he decided the + behaviour; the story wording is what conflicts. +24. ACCEPT — who writes the required per-pass dispositions line, given the reviewer may write + only the findings path? Unreconciled. +25. ACCEPT — the three schemas are field lists with no literal examples or delimiters; + prompt-standards item 4. +26. ACCEPT — "the re-review ran" specifies no floor, no final-clean rule, no disposition + authority. A token pass could discharge the compensating control. +27. ACCEPT — nothing detects availability's return, so the debt trigger is observationally + empty. +28. ACCEPT — rebase-merge and cherry-pick change the landed SHA outside the squash rule. +29. ACCEPT — the squash hop depends on a merge-time actor the design never assigns. +30. ACCEPT — the old-condition inventory is partial. §5's "gates still apply", "the gate + itself is not optional", loop-closure and incomplete-pass conditions are untouched by my + tables. The AGENTS.md rule wants the complete inventory. +31. ACCEPT — more falsified sites: `plugin.json`'s description, marketplace metadata, + README's product summary. Search the claim, not the phrase — the repo's own recipe. +32. ACCEPT — the evidence covers one happy Gate-B path; Gate A, tier 3, mixed tiers, + recovery, concurrency, multi-WIP, real squash and the debt lifecycle are untested. +33. ACCEPT — restates B3 from pass 1 at product level: the installed user base is exactly the + population the incident hit. Daniel scoped this out; finding 33 argues the release claim + must then narrow too. Worth honouring in the wording. +34. ACCEPT — the slot-collision fix should land in THIS change, since the new protocol makes + per-pass records load-bearing. Stronger than my "park it" plan, and it already bit twice + this cycle. +35. ACCEPT — no rollback path; reverting the prompts while `.off` persists leaves an older + §5 saying gates apply with reminders globally silent. +36. ACCEPT (security lens) — failure shapes, causes and vendor diagnostics go into commit + history with no redaction rule; auth errors can carry account identifiers or endpoints. +37. ACCEPT — tier 3 is described as existing practice and "no new waiver mechanism", but + init-time gateless closes no active gate and a profile override changes evidence mode + rather than authorizing zero-pass closure. It IS a new exception and should say so. + +## Minor — collected + +38. "HEAD is the complete content identity" outruns the comparison — ignored files, untracked + content. The gate-proof wording rule, and I wrote it in the section that exists to + prevent exactly this. +39. Marker dedup has no identity key. + +## Verdict + +Nothing dismissed across two passes: 71 findings, 69 actioned, 0 rejected. The design is +not converging at the mechanism layer — pass 2's blockers are about whether tier 2 can run +at all, which pass 1 never reached because the design was too vague to be tested that way. +Three items go back to Daniel: finding 5 (his invariant-2 decision rested on a premise the +shipped text contradicts), findings 1-2-3 (prompt-only cannot deliver the stated +containment, reopening Q2), and finding 23 (story AC 6's wording). diff --git a/.context/codex-reviews/gate-a-spec-pass-2-dispositions.stopped-tier3-core.md b/.context/codex-reviews/gate-a-spec-pass-2-dispositions.stopped-tier3-core.md new file mode 100644 index 0000000..b50c90f --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-2-dispositions.stopped-tier3-core.md @@ -0,0 +1,141 @@ +# Gate A — spec — pass 2 dispositions (stripped design) + +29 findings (14 BLOCKER, 13 MAJOR, 2 MINOR). **None dismissed.** Enum held an eighth time. + +## The blocker count rose 4 → 14. Read it before reacting to it. + +| Cycle | Blockers by pass | +|---|---| +| Three-tier | 4 → 4 → 6, stopped | +| Two-tier with debt machinery | 7 → 8 → 12, stopped | +| **Stripped** | **4 → 14** | + +**Eight of the fourteen are defects in machinery pass 1 asked for and this pass is the first +to see** — the dedicated index (2), its parent binding (3), the authorization digest (4, 6), +the Gate-A continuation rule (7, 8), the merge-strategy scoping (9), the docs-only tree test +(1). The artifact grew 489 → 750 lines; every one of those blockers is in the new 261. This +is the "a fix introduces a subtler version of what it fixed" pattern the pass-2 prompt +explicitly hunted, and it found it. Ordinary, all with named fixes. + +**Two are the recursive shape returning from a new direction** — 12 (the "complete pass +ledger" has no authoritative enumeration source) and 13 (dropping a story citation drops the +profile). Both are the same thing the two stopped cycles died of: making the waiver safe +wants protected state, and protecting state wants the gate. **Both are taken fail-closed +instead**, which needs no new state: ambiguity refuses the waiver. That makes tier 3 refuse +more often, which is the safe direction, and it is a real answer rather than a deferral. + +**Four are the inventory and the probe** — 19, 20 (the fourth exhaustiveness claim is false +again) and 21 (the probe's positive scenario does not reach the decision point). + +So: **not stuck.** No finding says the mechanism cannot exist. + +## The watch item fired, and is taken by calibration — not by an override + +Findings 21–26 all land on §10's prompt-differential probe: the positive scenario omits the +ordered close (21), it spends the wrong call budget (22), its `OLD` assertion is not entailed +(23), it names no runnable invocation (24), it pins no runner or model against invariant 5 +(25), and — the substantive one — **three matching outputs from a probabilistic model +establish that the model emitted those tokens, not that the procedure's semantics require +them** (26). + +26 is the same unverified-evidence claim in a new costume, exactly what the resume note +watched for. **It does not force the whole-mode override**, because the fix is calibration +rather than abandonment: pin the runner and model, freeze the prompt bytes, make the parse +deterministic, add a mechanical structural oracle beside the model probe, and **state the +claim as what was observed** — a pinned model's verdicts under frozen inputs — instead of as +what a compliant reader must derive. CLAUDE.md already permits a **named verification** where +no automated test is possible, and requires the counterfactual, which the frozen `OLD` blob +supplies. Calibrated that way it is honest evidence; claimed the old way it was an overclaim. + +`battery+check+verification` therefore still stands in full and no override is proposed. +**If pass 3 falsifies the calibration too, that is the point to take a whole-mode override +to Daniel**, and not before. + +## Blockers — all accepted + +1. **Docs-only test enumerates the whole tree.** A tree is the entire repository, so the rule + as written refuses every real Gate-A tree. **Fix:** classify the **changed paths** of the + parent-tree → prospective-tree diff, naming additions, deletions, renames, modes and type + changes — and name that comparison, per the gate-proof Don't. +2. **The dedicated index is never initialized.** A fresh `GIT_INDEX_FILE` starts **empty**, so + following §3.5 literally writes a tree that deletes every tracked path not restaged. A + data-loss path introduced by pass 1's own fix. **Fix:** `git read-tree` the expected parent + first, then stage, then verify. +3. **Nothing binds the expected parent.** The index is isolated; `HEAD` is not. Another + process can advance the branch between authorization and commit, and step 7 still passes + because it compares only trees. **Fix:** expected parent inside the digest, a + compare-and-swap check that `HEAD` still equals it immediately before commit, and both + parent and tree verified after. +4. **Read-back checks the tree, never the message.** Everything authorization is said to bind + — marker, decision, checklist, evidence — lives in the message. **Fix:** read the committed + message back and compare its normalized record bytes as well as tree and parent. +6. **The authorization loop is circular.** The decision block's timestamp and reason are + *created by* the answer, yet the digest containing them must be shown *before* it. Filling + them afterwards changes the authorized bytes; pre-filling them records a predicted + decision. **Fix:** split a pre-answer **request digest** from a post-answer + **attestation** — the human returns the request digest plus their decision data, and the + committed block binds to what was returned. +7. **No A-spec continuation.** Pass 1 defined continuation only for A-plan at + `executing-plans`; the A-spec reminder at `writing-plans` has none, so a correctly waived + **spec** still cannot advance — the original stall, unsolved. **Fix:** the symmetric rule. +8. **A-plan continuation never compares the plan it is about to execute.** It validates the + blob named in the waived commit, not the current one, so a plan edited afterwards clears + the reminder on an old authorization. Second gate-off path. **Fix:** require blob equality. +9. **Branch-head immutability makes ordinary merge impossible after a Gate-A waiver**, since + plan and implementation commits necessarily follow it. **Fix:** scope immutability per + gate — Gate A binds its own commit and artifact; Gate B binds the final head. +10. **`auth-failed` is locally manufacturable.** Withhold or corrupt a credential and the + request stays canonical while the reviewer "cannot run". **Fix: the source is removed.** + A local authentication failure is a **configuration error to repair**, never an outage. +12. **The complete pass ledger has no authoritative source.** §5 records no durable pass ids + as passes occur, `.context/` slots are reusable and collide, and the hook counter is + explicitly not evidence — so a lost adverse pass is indistinguishable from no pass. + **Fix, fail-closed and state-free:** the ledger is authoritative only where it is durable + (pass ids accumulated in the WIP body at Gate B), and **any** ambiguity — a resumed + session without one, a slot whose provenance cannot be established, a gap in the sequence + — **refuses the waiver**. `Passes completed: none` may only be written when the absence + itself is established, never when it is merely observed. +13. **Dropping a story citation drops the profile.** Removing the citation immediately before + the waiver sheds the high-risk mode, its lenses and its evidence while still satisfying + the written three-case rule. **Fix:** bind to the **union** of story paths recorded across + the cycle; an unexplained disappearance is an **unresolvable profile** (stop and surface), + never `unprofiled`. +19. **The fourth exhaustiveness claim is false again.** Named omissions: Gate A's placement + before `writing-plans` / `executing-plans`; the sweep's no-side-effects and + do-not-run-quoted-commands limits; Gate B's explicit `AGENTS.md` check; the + companion-deletion rule; the zero-finding early exit as a distinct clause; and the rule + that a changed evidence entry invalidates the clean pass. **Fix:** add them, and stop + claiming exhaustiveness in the abstract — claim the derivation and name what was checked. +20. **The standing falsification lens is silently dropped at tier 3.** Rows 52 and 53 are + marked untouched, but a zero-pass Gate-B closure has no reviewer to ask them and the + checklist does not include them. **Fix:** both narrowed, both questions — including the + required repository search — added to the checklist. +21. **S1 does not reach the decision point.** It lists preconditions and an authorization but + omits the ordered close, so a compliant reader should return `REFUSE` on `NEW` and the + probe fails while the procedure is correct. **Fix:** a fully sequenced transcript, plus + negative variants for each ordered-close binding. + +## Majors and minors — accepted + +5 (digest has no algorithm, framing or field order — specify SHA-256 over a length-delimited +serialization and record it), 11 (step 5's probe count per source), 14 (mixed +profiled/unprofiled cycles need per-story status in the marker), 15 (the item-by-item +checklist evaporates into four words and is not carried — needs a bounded record inside the +digest and through both merge strategies), 16 (the free-text +"answered-with-exception" is an unbounded route around the only compensating evidence — +**removed**), 17 (`baseSha..headSha` cannot survive the amend that creates it — target by +expected parent + authorized tree + diff digest), 18 (a revert leaves the invalid marker +discoverable — needs an explicit invalidating incident marker), 22/23/24/25/26 (the probe — +see above), 27 (escaping needs backslash-parity, reverse-order unescaping, list framing and +accept/reject vectors), 28 (`status-page` is unreachable as a recorded `Cause` — **removed** +from the enum; it survives as corroboration only), 29 (the continuation clears the reminder +**whenever** derived for that unchanged plan; "once" overstates a state-free mechanism). + +**Net effect on the enum:** 10 and 28 together shrink the source enum from four values to +**two** — `quota-observed` and `attempt-failed`. Both require a canonical call; neither is +manufacturable without disabling the account itself. + +## Status + +Not clean; floor not met. Pass 3 next. No decision is owed to Daniel yet — the watch item is +taken by calibration, and the stuck test does not fire. diff --git a/.context/codex-reviews/gate-a-spec-pass-2.md b/.context/codex-reviews/gate-a-spec-pass-2.md new file mode 100644 index 0000000..0e60644 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-2.md @@ -0,0 +1,26 @@ +BLOCKER | high | parent story lines 9–16 and 62–68 | Both closure summaries say the quota stall routes through the recorded human exception, although design §2.0 and §7 say the form cannot reach a core-gate stall; both also point to §6, which is Backlog rather than the stall routing section | This restores the withdrawn operational route from outage to human record in the story that future agents will use, so the record can still be read as a zero-pass gate waiver | Remove the recorded exception from both stall-routing lists, point to §7, and state that only operational bridges or a buildable reviewing fallback can unblock gated work +MAJOR | high | parent story lines 9–12 | The story says the passes established that a sanctioned zero-pass closure cannot have the required properties in a prompt-only system, while design §1.5 explicitly says the work is not a proof of impossibility and leaves outside-authority classes untouched | The settled record is stronger than its evidence in one of the two authoritative artifacts, defeating the calibration the closure design is trying to preserve | Rephrase to the design’s bounded finding: no safe design was found in the three tried classes, for the recorded structural reasons +MAJOR | high | design §1.3 lines 50–72 versus §1.5 lines 97–103 | The headings claim failure of any design carrying compensating state and of history-derived preconditions generally, and the rationale says the gate is the only protection mechanism, but §1.5 concedes that protected external authority and history were not tried | The class-level finding outruns the examined comparison and contradicts its own disclaimer; externally protected state is still compensating state and can establish history | Qualify the failed classes as client/repository-authored state and unprotected local history, and reserve the broader labels for a claim the evidence actually supports +MAJOR | high | design §1.5 lines 101–103 and §6 lines 324–327 | A signed commit or protected-branch approval is asserted to supply grounding and attribution genuinely, solely because its authority is supposedly not the waived party’s to mint; an ordinary signer can be the author, and a branch approval proves neither reviewer unavailability nor an authorized approver without additional identity and policy comparisons | This is another gate-proof overclaim and makes the parked direction look sufficient when the named mechanisms do not establish property 1 or 4 by themselves | State the exact additional requirements: trusted signer identity and role policy for attribution, plus an independently controlled availability attestation for grounding; treat the mechanisms as candidates, not proof +MAJOR | high | design §2.1 lines 149–150 versus §2.4 lines 221–223 and §8 lines 366–370 | The shipped paragraph says the record “survives to main,” but the design later says nothing validates carry and an omitted record is undetectable | A standalone CLAUDE.md reader receives an enforcement guarantee the actual instruction-backed comparison cannot make, repeating the repository’s recurring gate-proof defect | Replace the guarantee with a normative instruction and calibrated result, such as “copy it into the squash body; nothing verifies that carry” +MAJOR | medium | design §2.0 lines 123–126, §2.1 lines 138–159, and §2.4 lines 210–218 | Applicability is phrased as “this section’s gates are satisfied,” yet §2.4 places records in individual Gate-A spec and plan commits before the later Gate-A run and Gate B can be satisfied | The load-bearing scope condition is temporally impossible under one literal reading and ambiguous under another, so agents must invent whether “gates” means the current cycle, current run, or all gates | Define applicability per completed gate run or per current lifecycle point, then make the same singular scope wording appear in the shipped paragraph and placement table +MAJOR | high | design §2.2 lines 170–176 | “Not conditioned” and “No preconditions” directly contradict §2.0’s load-bearing condition that the relevant gates already be satisfied | Calling the scope condition a non-condition weakens the very boundary that keeps the form from becoming a waiver | Say that the form has a strict applicability boundary but no conditions that authorize bypassing an obligation; do not deny that the boundary is a precondition to using it +MAJOR | high | design §2.1 lines 138–166 | “A convention consciously set aside” and “something less than the full process was acceptable” are open-ended normative examples; the exclusions protect gates and profile evidence but do not say that the form cannot authorize other mandatory CLAUDE.md or AGENTS.md obligations | A reader of only the shipped paragraph can take human assent plus the form as the sanctioned procedure for ignoring any mandatory non-gate rule, recreating the first draft’s distinction-without-a-difference outside the enumerated exclusions | Limit examples to independently optional or supplementary activities and state that the form records a decision after the fact but supplies no permission to violate any mandatory process or invariant +MAJOR | high | design §2.3 lines 179–203 and rider (b) lines 237–269 | The sixteen-row accounting collapses all companion behavior into one row and omits changed old clauses: dispositions are currently optional, may be deleted or rebuilt, contain one verdict line per finding, never participate in validation, and the findings file is the only hard requirement; the “Accept a pass only when” grammar also changes | Implementing the rider while preserving the omitted clauses yields direct contradictions and lets a supposedly complete accounting hide the decision-procedure edges most likely to be lost | Add individual kept/narrowed dispositions for every affected companion and acceptance clause, including deletion, one-line-per-finding shape, only-hard-requirement, zero-finding behavior, and the pass validator’s new dependency +MAJOR | high | design §2.3 line 200 versus §2.4 lines 221–223 | The accounting says the exception record is “validated the same way” as the evidence entry, but §2.4 says there is no validator and only a human or agent reads a prospective body | This falsely upgrades an instruction-backed carry check into validation and contradicts the mechanism description inside the same design | Name the exact manual comparison for each record type, or say both are re-read manually without calling that validation +MAJOR | high | rider (b) lines 263–269 | The audit must run “before the final pass,” but a pass is known to be final only after it returns clean; the early zero-finding exit can make even pass 2 final unexpectedly | The procedure cannot be followed as written, and a normalized earlier pass can reach the outcome-dependent final pass without the promised pre-pass audit | Require the audit before every subsequent call while any credited pass depends on normalization, and again immediately before closure +MAJOR | high | rider (b) lines 263–269 | After discounting a pass, numbering “continues from the surviving count,” which can reuse an existing pass number and overwrite a surviving findings or dispositions slot; saying passes are not renumbered does not resolve that collision | Slot identity is already known to collide, and overwriting a surviving artifact can destroy the very drift record the pre-close audit needs | Separate monotonic invocation ordinal from credited-pass count; never reuse a slot number, and compute the floor from surviving credited ordinals +BLOCKER | high | design §4 lines 275–294 versus §6 lines 338–347 | The Sites section omits todos.md, docs/hardening-log.md, the plugin CHANGELOG, and the plugin manifest even though §6 requires edits to them; it also labels the plugin manifest “checked and unchanged” while §6 requires a 0.8.2 to 0.9.0 manifest bump | The implementation surface is internally contradictory and risks violating invariant 12 or silently dropping the hardening/backlog work | List every changed file in §4 with its exact change, move the manifest out of “unchanged,” and distinguish content truth checks from files that remain byte-for-byte unchanged +MAJOR | high | design §5 lines 302–315 | The battery is said to cover parity between CLAUDE.md §5 and the workflow-init mirror, but scripts/check-invariants.sh only compares the prompt-standards checklist counts; it does not compare the two §5 copies | This is a concrete false enforcement claim, and the two shipped prompt copies can diverge while the named battery stays green | Replace the claim with the actual narrow checker behavior and specify a real extraction-and-diff parity command or add a mechanical parity check to the implementation +MAJOR | high | design §5 lines 305–317 | The discriminating check specifies a fixture and semantic Before/After verdict but no runner, prompt-harness protocol, oracle, or command; the “Before: ACCEPTED” result is inferred from prose that never defines IMPORTANT as accepted | The story’s battery+check obligation is not executable or reproducible, and its counterfactual can be asserted rather than observed | Specify the exact harness, model/input, extraction of both prompt copies, acceptance oracle, and observed prior-state run; if model interpretation cannot be deterministic, name the verification route and its limitations +MAJOR | medium | design §5 lines 305–320 | Validation covers only the missing-drift negative fixture and gives no positive or boundary cases for a valid normalized pass, malformed drift grammar, duplicate tokens or line numbers, full Gate-B per-branch placement, record deletion, or the two pre-close audits | Most of rider (b)’s live control flow can be implemented incorrectly while the sole discriminating scenario still passes | Add reject/accept pairs for the grammar and lifecycle branches, especially full-review branch asymmetry, retention discount, audit reopening, and monotonic numbering +MAJOR | high | parent story acceptance criterion 3 lines 156–177 | The latest live accounting still says docs/sparring-briefing.md’s never-exempt and human-substitution clauses are overturned and must be amended, while design §2.0 and §4 now rely on that file remaining unchanged | The story and design give opposite implementation requirements for a load-bearing no-waiver statement | Add a fifth-amendment disposition that overturns the earlier amendment and keeps those clauses unchanged because the record never substitutes for a pass +MAJOR | high | parent story §4 lines 211–236 | The affected-invariants section still says tier 3 qualifies the project’s cross-model-gate statement by making the gate waivable | This leaves a live statement that the withdrawn tier ships and directly contradicts the closure banner and design §2.0 | Amend the invariant disposition to say no qualification ships and the cross-model statement remains true unchanged +MAJOR | high | tier-2 story lines 50–64, 78–79, and 126–127 | The desired outcome refers to disclosure “wherever the ladder already discloses degradation,” AC 5 delegates its durable form to the parent’s now-narrowed human-exception AC 2, and the open question asks whether a narrowing principle the parent shipped covers tier 2; the closure says no ladder or narrowing ships | The dependent story still relies on three withdrawn decisions, so §4’s claim that no live tier-2 requirement rests on tier 3 is false | Give tier 2 its own prospective disclosure requirement and record form question, and replace both parent-ladder references with the actual one-tier baseline +MAJOR | high | tier-2 story lines 55–60 | The story infers that preventing writes to the reviewed repository also prevents exfiltration, but a reviewer that can read content and emit findings already has an output channel, and external execution may add network or environment channels | The security-high containment goal claims a protection its stated control does not provide | Separate integrity from confidentiality, enumerate output/network/environment channels, and either define an exfiltration control and check or drop the exfiltration guarantee +MAJOR | medium | tier-2 story lines 59–60 and 76–77 | “Reviews a complete range” is treated as established by an unspecified check that fails on truncation, without naming the exact source-to-reviewed-input comparison or how files outside the diff are covered | A truncation detector can cover one mutation without proving lossless input, violating the gate-proof calibration rule | Define the authoritative source set and exact byte/hash/path comparison for diff and out-of-diff inputs, and bound any residual cases instead of claiming completeness +MAJOR | high | design §2.4 lines 214–219 | Identical blocks are collapsed, while any field difference is declared a separate decision; there is no decision identifier or correction/supersession rule, so two distinct same-day decisions can collapse and a typo correction can become a second contradictory exception | The durable record can lose multiplicity or preserve stale claims, undermining the only value the form is meant to provide | Add a stable record identifier and an explicit correction/supersedes form, then deduplicate only exact copies carrying the same identifier +MINOR | medium | rider (b) lines 245–258 | The drift grammar escapes middle dot and pipe but does not escape backslash, so tokens containing backslashes can serialize ambiguously with the escape sequences themselves | A reader cannot always reconstruct the verbatim token or decide whether two records name the same drift | Define backslash escaping first and specify encode/decode order with examples for literal backslash, pipe, and middle dot +MINOR | medium | design §2.4 lines 217–223 | Squash carry says “every block” is copied but never defines the commit range or source bodies to scan, so an agent can omit Gate-A records, include records inherited from the base, or disagree about amended commits | The manual pre-merge comparison has no exact input set and therefore cannot support the carry claim consistently | Define the prospective squash range relative to the PR base and collect canonical blocks from the final reachable commits in that range before comparing them to the proposed squash body +MINOR | high | parent story lines 16, 67–68, and 258–259 | Several references were not updated after the design was renumbered: stall routing points to §6 instead of §7, and the tracked-debt backlog row points to design §5 instead of §6 | Readers land on validation or backlog text that does not support the cited claim, making the closure record harder to audit | Update the citations to design §7 for stall routing and §6 for tracked re-review debt +END OF FINDINGS (25 total) diff --git a/.context/codex-reviews/gate-a-spec-pass-2.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-2.pre-2026-08-14.md new file mode 100644 index 0000000..b0a843b --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-2.pre-2026-08-14.md @@ -0,0 +1,12 @@ +BLOCKER | high | §2 lines 26-40 | the design keeps the existing absolute `Never edit a row` instruction and then tells the same reader to amend an open-cycle row in place | the shared prompt has two incompatible actions for row D and every other in-flight correction, violating prompt-standard item 7 and leaving the central gap unresolved | change the retained sentence to prohibit editing a landed row, or state the open-cycle exception in that sentence, identically on both surfaces +BLOCKER | high | §2.2 lines 76-95 and §4 | the supposedly omit-until-first-entry block label `**Superseded rows.**` is already the first text of the always-shared §2.2 prose, while the layout and union-repair rules require a separately created bold label when entries exist | an empty scaffold already contains the label, and the repo implementation must either have no separately labelled block or duplicate the label, contradicting the settled placement and empty-state contract | make the shared convention prose unlabeled, then specify a separate dynamic block consisting of `**Superseded rows.**` plus its entries only when at least one entry exists +MAJOR | high | §2.1 lines 37-67 | self-test verdicts are (a) 2026-07-20: landed → supersession entry, as specified; (b) row D during its authoring cycle: open-cycle → in-place amend, as specified; (c) earlier completed cycle on an unmerged branch: landed → supersession entry, as specified by the definition; and (d) open-cycle row after a `WIP:` snapshot: in-place amend, as specified, but line 65 calls absence at the open-cycle base the test for being landed, which reverses case (d) and also cannot distinguish it from a row brought in from another cycle after the base | the recorded decision procedure can classify an in-flight row as landed or authorize editing a concurrent cycle's landed row, breaking append-only history | state that presence at the base proves landed, while absence does not decide authorship; define how current-cycle authorship is established for rows added after the base +MAJOR | high | §2.1 lines 37-67 | the shared convention uses `the one now open`, `appended by` a cycle, and `until that cycle closes` without defining cycle identity, its opening/closing events, the no-open-cycle state, or which cycle controls when several worktrees have cycles open | a reader of the ledger cannot reliably decide whether editing is sanctioned, and two agents can apply different boundaries to the same row | define the exact cycle and close event, say that no matching open cycle makes every existing row landed, and distinguish the current cycle from other concurrent open cycles +MAJOR | high | §3.1 lines 133-140 and story acceptance criterion 4 | the first entry restates the new 0.8.0 three-shape counting behavior even though the human-confirmed criterion moved current behavior to the cited artifact under the cite-don't-restate rule | the copied mechanism description can become the next stale ledger narration and conflicts with prompt-standard item 11's authoritative-source rule | state only that the old claim that every incomplete pass increments is now false, then point to a unique authoritative location for the current behavior +MAJOR | high | §2.2 lines 83-88 and lines 120-123 | `a wrong entry is corrected by a further entry naming it` and `retires ... by naming it` define neither a locator, an output shape, nor precedence for correcting an entry, restoring a formerly false claim, or reconciling contradictory concurrent entries | the promised append-only recovery path is not executable and can leave multiple standing corrections with no ledger-only current answer | add prose examples that uniquely identify prior entries and specify when a correcting entry retires one or several earlier entries, including restoration and concurrent-conflict cases +MAJOR | high | §7 Check 2 and §9 lines 276-282 | normalizing whitespace before the parity comparison can erase the four-space indentation that is the only distinction between the format example and a live list entry | the validation can pass when one surface has converted the guarded code example into parseable entry-shaped content, recreating the exact limitation §9 records | normalize hard-wrap joins only for prose while preserving structural newlines and leading indentation, and compare the indented example's prefix exactly +MAJOR | high | §7 Check 1 lines 221-230 | the two greps verify only an entry prefix and a unique table locator; they do not verify that the entry identifies the false claim or points to a resolvable current answer | acceptance criterion 4 can fail while the required `battery+check` evidence is green | retain the cardinality greps but add a named read verification that checks both semantic fields and resolves the citation, or define bounded checks for those required contents without pretending to parse all prose +MINOR | high | sweep — §3.1 line 134 | `CLAUDE.md §5, "What this does not do"` is not a unique citation because §5 contains two bold paragraphs with that exact title text | the entry does not unambiguously tell a ledger-only reader where the current answer lives | cite the paragraph by a unique opening fragment such as `The hook counts on PostToolUse` or by its surrounding file-protocol subsection +MINOR | high | sweep — §4 lines 160-164 | the end boundary `through the end of §2.2's prose paragraph` is not mechanically unique because §2.2 contains two prose paragraphs separated by the indented example and the target files carry no §2.2 marker | two implementations of the parity check can extract different regions, so the check is not yet specified well enough to reproduce | name exact start and end sentinels, preferably the final shared sentence inclusive, and state whether the indented example lies inside the compared region +MINOR | high | sweep and old-conditions audit — §5 lines 189-201 | the table marks the landed-row prohibition as `moved`, but the proposed skill replacement still says `never edit a landed row`; only its definition is delegated to the ledger header | the required old-conditions audit records an incorrect disposition and obscures that the prohibition remains duplicated across two prompt surfaces | mark the prohibition `kept` and say the definition of `landed` is moved to the header +END OF FINDINGS (11 total) diff --git a/.context/codex-reviews/gate-a-spec-pass-2.stopped-2tier-debt.md b/.context/codex-reviews/gate-a-spec-pass-2.stopped-2tier-debt.md new file mode 100644 index 0000000..0d075ba --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-2.stopped-2tier-debt.md @@ -0,0 +1,35 @@ +BLOCKER | high | §§3.4, 5, 7, 12–13 | [risk: threats; risk + security: abuse and abuse paths; security: trust boundaries] the same agent whose work is being waived can fabricate the authorization block, invent the accountable handle, create or omit the debt row, and later declare it closed; every compensating control is therefore controlled by the subject it constrains | §13 honestly admits this, but that admission concedes the core feature: this is an agent-operated escape hatch that can look exactly like a human exception, so the pause, disclosure, debt and next-waiver block are decorative under the abuse path they exist to contain | require a human-bound signal outside the agent-authored commit text, such as a protected-branch approval or signed approval artifact, and a mechanically checked non-closed debt guard before tier 3 ships rather than parking enforcement until after the first disputed waiver +BLOCKER | high | §11 state matrix, lines 468–474, versus §13, lines 515–527 | the matrix says self-authorization, same-approver debt acceptance, a second waiver and stale availability are refused, including the claim that no decision block is producible, while §13 says an agent can write a conforming-looking block and none of those refusals is enforced | this is a direct internal contradiction and an enforcement claim without a mechanism, violating prompt-standard item 11 and making the named verification report controls it cannot observe | either add and verify the mechanism that performs each refusal or change the matrix to the exact observable manual behavior and stop calling adversarial bypasses refused +BLOCKER | high | §11 “Check that fails without the change”, lines 445–454 | the counterfactual is backwards: under the prior policy tier 3 does not exist and the gate is mandatory, so a compliant agent refuses every tier-3 closure, while the advisory hook permits a commit both before and after the change; no interpretation makes the old state uniquely succeed because it lacks the new adverse-finding prohibition | the high-risk profile requires a check that fails without the change, and §5 says an unobservable counterfactual is a blocking evidence gap; this proposed check cannot distinguish old from new behavior | design a probe against an observable mechanism introduced by the change, or explicitly surface the evidence gap and obtain the logged validation-mode override that §5 requires +BLOCKER | high | §§3.1, 6–7 and §8 row 48 | [risk: threats] tier-3 preconditions never require the profiled battery, check, named verification, abuse-path suffix when applicable, or current evidence entry to be complete and revalidated; §6 only says to carry an evidence entry, which can carry an unsupported assertion | a high-risk Gate-B waiver can bypass both independent review and the independent evidence obligations, removing the only remaining substantive check while still producing the expected marker | state that tier 3 waives only reviewer passes, make every existing mode-derived obligation and evidence revalidation a precondition, and add refusal cases for missing, stale and inadequate evidence +BLOCKER | high | §§5.4, 7, 12 and §8 row 43 | [risk + security: abuse and abuse paths; risk: observability] `todos.md` becomes authoritative gate state even though it is mutable backlog prose and the existing Gate-B classifier exempts an ordinary root Markdown file; deleting, editing or falsely closing a debt row can therefore land without Gate B and makes the next-waiver check pass | the design silently changes `todos.md` from advisory backlog into security-relevant product state while marking prose-vs-product classification “Kept, untouched”; this opens an unreviewed debt-erasure path and makes the claimed consequence decorative | keep authoritative debt state in an immutable or gate-classified record, or explicitly change the classifier and old-condition disposition; validate every marker-to-debt transition and retain an auditable closure record +BLOCKER | high | §7 closing condition 2, lines 297–304, and story §5 settled question | [risk + security: abuse and abuse paths] a second human handle may close the “mandatory cross-model re-review debt” without any review, then the next waiver is available again | two people can alternate waiver approval and debt acceptance indefinitely with zero qualifying passes, so the supposed mandatory re-review and anti-compounding control reduce to repeated human waivers | remove no-review debt cancellation, or rename the obligation honestly and add a hard cumulative cap whose reset requires an actual tier-1 repayment +MAJOR | high | §§5.4 and 7, lines 205–209 and 306–309 | the state enum makes an unpaid row `reconciled`, but the block is phrased for rows that are `open` or unreconciled; after landed identity is added, a reconciled-but-unpaid row is neither and can escape the next-waiver guard | the normal post-land transition creates the exact state that bypasses the only stated debt consequence | define states as open-unreconciled, open-reconciled and closed, and refuse tier 3 for every status other than closed +BLOCKER | high | §§5.4, 6–7 and §13 rollback claim | [risk: data loss; risk: compatibility] story commit SHAs and Gate-A blob-or-commit SHAs taken from WIP or pre-squash history are immutable but not durable: amend and squash can leave them unreachable and garbage collection can remove them; if the profile changes in the closing amend, embedding that final commit SHA in its own row is also self-referential and impossible | the debt can lose the very profile and artifact it must repay, so rollback and delayed repayment become unserviceable despite the design claiming durable identity | use an identity computable before commit and anchored in `main`, and preserve the exact story profile and Gate-A artifact in a reachable snapshot or self-contained debt record across amend and squash +BLOCKER | high | §8, lines 315–382, against CLAUDE.md §5 lines 67–428 | the claimed one-row-per-atomic-condition inventory is still a thematic compression, not an exhaustive document-order inventory; it omits, among other operative conditions, the hook’s inability to distinguish spec from plan and reset timing, the complete companion schemas and resume-note lifecycle, fresh-rerun deletion scope and resume eligibility, result-envelope/counting residuals, the exact Gate-A prompt inputs and timing, the Gate-B path classifier, and multiple WIP/baseSha mechanics | the previous thematic inventory was a blocker, and this version still cannot demonstrate that the replaced decision procedure preserved every old condition, violating the named AGENTS.md Don’t | rebuild from each operative sentence and list item in §5, splitting every independently actionable condition into its own row before assigning a disposition +MAJOR | high | story lines 23–35 and 91–102; spec §9 lines 405–407; docs/sparring-briefing.md lines 41–44 | the amendment says the sparring rule moved entirely with same-family tier 2 because it only forbids a model fallback, but the actual rule also says not to treat a satisfied human as a substitute for a clean pass and not to design around the gates | tier 3 does exactly what that still-shipped instruction forbids, so the old-condition accounting is not honest or complete and leaves contradictory prompt text in force | keep only the genuinely same-family premise in the tier-2 story, account for the human-substitution and never-exempt clauses here, and update or deliberately overturn them in this change +MAJOR | high | §8 row 11 | “Only early exit is a zero-finding pass” is marked Kept even though tier 3 closes below the floor with zero qualifying passes and may do so after one or two completed passes | calling the operation a waiver rather than an exit does not preserve the old decision condition and hides a central deliberate overturn | mark the rule kept only for tier 1 and explicitly overturned for tier 3, then update every dependent statement +MAJOR | high | §8 row 39 and §12 lines 501–502 | “Skip ONLY trivial changes” is marked Kept solely by renaming the new zero-review path a waiver rather than a skip | operationally, a nontrivial change now proceeds without Gate B review, so the old condition is narrowed regardless of record vocabulary; the current disposition repeats the semantic loophole the accounting rule is meant to prevent | mark the condition narrowed or overturned for tier 3 and audit all skip/no-review assertions by effect rather than by the noun used +MAJOR | high | §8 row 45 versus §3.1(4) | the three profile-reading cases are marked Kept “without changing the unprofiled path”, but an uncited artifact currently runs unprofiled and tier 3 now refuses it until a citation is added | this is a real change to the uncited path and the disposition expressly denies it, so the old-condition accounting is false | mark the uncited case narrowed for tier 3 and specify whether adding a new story solely to obtain a waiver is permitted and how its profile is chosen +MAJOR | high | §§3.1(2), 3.5 and 5.2, lines 56–61, 123–125 and 170–189 | the precondition allows completed passes with remediated or human-dispositioned adverse findings and says each pass and disposition is preserved in the waiver record, but the marker can encode only `none`, or a count asserting “none adverse”, and it names no pass identifiers or finding-level dispositions | valid waiver states are unrepresentable and the durable record can falsely erase adverse review evidence | add canonical pass identifiers plus each adverse finding’s remediation or human disposition, handle and reason, and validate them against the retained pass artifacts before closure +MAJOR | high | §§3.1(5), 5.2 and 6 | the merge strategy must be “supported and recorded” before authorization, but no marker, decision block or debt-row field records it | the precondition cannot be audited, deduplicated or used reliably during reconciliation, and a later merge-method change is indistinguishable from compliant selection | add a closed-set merge-strategy field to the canonical durable record and revalidate it immediately before merge +MAJOR | high | §§3.1(1), 3.2 and 5.2 | the marker always requires a confirmed source category, but §3.2 defines sources only for the confirmation-without-attempt branch; a fresh failed call plus recovery has no canonical source category or schema for the two attempts | a principal error path cannot produce a conforming marker, encouraging invented categories or loss of the evidence that justified the waiver | define a closed source-category enum and fields for fresh attempt, recovery attempt and independently confirmed outage, including which timestamp is the pre-close revalidation +MAJOR | high | §§3.3 and 7 | [security: roles] the debt accepter must be a different accountable handle, yet the solo case says the same human is both approver and accepter and only human-from-agent separation holds | the rule is either impossible for solo users or satisfied by one person using two aliases, which manufactures the separation the text says it refuses to pretend exists | require a genuinely different human for no-review debt acceptance and disallow that closing condition in solo use, or explicitly waive separation and stop claiming separation of duties +MAJOR | high | §§5.1, 5.4 and 7 | [risk: idempotency; risk: concurrency] repository uniqueness and the open-debt check are branch-local instructions with no merge-time serialization, so two concurrent branches can choose the same cycle id or each authorize a waiver before either debt row reaches the other branch | simultaneous waivers can collide, deduplicate unrelated records or compound debt despite every participant following the procedure | recheck current `main` immediately before merge, define deterministic collision repair, and add a CI or protected merge check that serializes waiver/debt creation +MAJOR | high | §7 non-contiguous remediation, lines 290–295 | reviewing the original defective range and a later remediation range as two separately clean cycles neither reviews their integrated result nor lets the original range become clean when the remediation is what fixes it | the debt can be impossible to close or can certify two fragments while missing their interaction, violating the gate-matched repayment claim | construct and review one isolated synthetic history containing the waived change plus selected remediation on its original base, or define a linked protocol that reviews the integrated landed state without unrelated commits +MAJOR | high | §7 Gate-A repayment, lines 267–283 and 297–300 | a full Gate-A cycle normally revises the spec or plan after findings, but the debt protocol requires every pass to target the immutable waived artifact and never defines the remediated Gate-A artifact, its identity, or the carry/closure record | any real Gate-A finding makes “same artifact” and “clean final pass after remediation” incompatible, leaving the debt lifecycle undefined | specify how the original artifact is first reviewed, how a revised artifact is produced and pinned, what the final pass reviews, and what durable identities close the row +MAJOR | high | §6 lines 245–252 versus story AC 2 | a later corrective commit can restore blocks somewhere in `main` history but cannot make the already-landed squash commit, which is the cycle-closing commit, contain them | the proposed recovery does not satisfy the acceptance criterion that the cycle-closing commit body itself distinguish the degraded cycle, and debt association becomes ambiguous | validate the prospective squash body before merge and refuse the merge when blocks are absent; treat a post-merge corrective commit as incident recovery, not as AC 2 satisfaction +MAJOR | high | §4 lines 139–153 and existing hook messages in CLAUDE.md §5 | [risk: compatibility] the hook is accurately described as waiver-unaware, but the design never resolves its model-visible STOP or “only early exit” instructions against an authorized waiver; Gate B may emit STOP and Gate A may emit a below-floor STOP at execution | a compliant agent can still halt after authorization, while teaching it to ignore STOP generically weakens invariant 2’s useful loose reminders for later commits or resumed sessions | add an explicit precedence rule scoped to one recorded cycle id and one transition, including restart behavior, while leaving the hook firing and all unrelated STOP messages authoritative +MAJOR | high | §10 lines 418–428 versus CLAUDE.md §5 hook counting | the enum-drift companion must be written before the pass is counted, but PostToolUse increments the hook counter before the outer agent can inspect the returned file and write dispositions | the prescribed order is mechanically impossible under the settled no-hook-change decision and its validation case cannot be performed as written | say the record must exist before the reader accepts or credits the pass, explicitly discounting the already-incremented hook count, or change the hook to support a pending-validation state +MAJOR | high | §§5.4 and 13 lines 532–538 | [risk: rollback; risk: observability] the debt-row template is not self-contained as claimed: open and reconciled rows have no defined values for `reviewed`, `landed` or `closed-by`, the checkbox transition is unspecified, and “condition 1” or “condition 2” names labels rather than the two closing procedures | after rollback, an operator cannot reconstruct the promised lifecycle or even distinguish unpaid reconciled debt from closed debt using the row alone | provide canonical schemas for every state and embed the full repayment and exceptional-close conditions or durable references that survive rollback +MAJOR | high | §§5.4 and 7, lines 208 and 270–273 | the Gate-B target is labelled `reviewed` and repeatedly called the reviewed range even though tier 3 is defined as a zero-pass closure | this overstates exactly what the gate proved and can make later readers believe the range already received review, violating the AGENTS.md gate-proof Don’t | rename the field and prose to `waived-range` or `review-target` until a repayment actually completes, then record the separately reviewed range +MAJOR | high | §9 site inventory | the repository survey omits directly falsified assertions outside the rows it names, including `docs/getting-started.md` line 84 that Gate A’s floor is unchanged at every level and the stage-level independent-review claims in `docs/coding-workflow.md` lines 27–29, 91–99, 108–109 and 123–126; the README daily-use pipeline is also left unqualified | row 56’s overturn would ship with user-facing text still teaching the old mandatory path, recreating the docs-drift class the site survey claims to close | search by the behavior claim across all non-archive files and list each affected passage with a kept, narrowed or overturned disposition rather than relying on one row per file +MINOR | high | §9 row for docs/coding-workflow.md line 149 | the cited sentence says branch protection requiring the quality check makes “the gate” mandatory, which in context is the repo-enforced quality gate, not the cross-model review gate | changing it for tier 3 would weaken or confuse an unrelated enforcement guarantee while claiming to fix waiver documentation | mark this passage untouched or first clarify its noun to “quality check”; do not qualify it with the review waiver +MAJOR | medium | §§3.1–3.2 | [security: external systems] a vendor status page can report a broad or partial incident without proving this account and tool call cannot complete, and “revalidated” is undefined, so checking the same still-red page can authorize a waiver after the reviewer has recovered for this user | the external-system trust boundary admits a skipped review while the reviewer is actually available, contrary to the narrow definition of unavailable | require an account/tool-specific probe at pre-close whenever safe, or define which authoritative status states prove total unavailability and how partial recovery is handled +MINOR | high | §§3.2 and 5.3 | [security: assets] sanitization covers the outage confirmation but not the free-form human-decision reason or finding-level waiver dispositions destined for a public commit body | raw error payloads, customer details, private endpoints or security findings can leak through the unbounded fields even while the Cause line is sanitized | apply the same closed, sanitized disclosure policy to every public record field and provide a private reference mechanism for sensitive rationale +MINOR | high | story lines 54–58 and spec §7 lines 306–309 | [risk: observability] “work continues” reads as sustained degraded operation, but the first waiver opens debt that blocks the next Gate-A or Gate-B waiver while the same outage persists; absent no-review debt cancellation, tier 3 advances at most one cycle | operators can expect the five-day stall to be solved and encounter it again at the immediately following gate | state the one-cycle limit explicitly in the outcome and incident analysis, and add the sustained-outage state to validation +MINOR | high | §§5.2–5.3 and 5.5 | the same authorization handle and date appear in both the marker and decision block, but no equality rule says they must match and deduplication only compares whole markers | contradictory identities can survive as two individually well-formed blocks and make accountability ambiguous | define one canonical source or require exact normalized equality of handle, date, cycle id and authorization action across the two blocks +MINOR | high | §10 versus CLAUDE.md §5 companion schema | the new `Enum drift` summary adds one line per token to a dispositions file whose existing contract is one verdict-and-reason line per finding, but the design does not define how the two grammars coexist or how repeated tokens map to individual findings | an implementation can satisfy one schema only by violating the other, and later readers cannot reconstruct which findings were normalized | amend the companion schema explicitly with ordered record types and a per-finding link or line number for each normalized token +MAJOR | high | §§5.4 and 7 repayment protocol | the debt says the profile at closure governs by story SHA, but the existing gate procedure reads the story path fresh and the repayment procedure never checks out or passes the SHA-pinned version; a later profile change therefore silently steers the repayment | the immutable reference is present but not consumed, so the claimed protection against later downgrade is illusory | materialize the exact pinned story version in the repayment worktree and pass that artifact explicitly to every gate call, with a check that the path content matches the recorded identity +MINOR | medium | §7 lines 306–309 | “checked at the start of any Gate-B cycle” does not state whether an open debt blocks the normal tier-1 cycle or merely disables tier 3 for that cycle | one reading recreates a full process stall even when Codex is available; the other is only an informational check | state the transition explicitly: normal tier-1 review remains allowed, while tier-3 eligibility is false until every debt is closed +END OF FINDINGS (34 total) diff --git a/.context/codex-reviews/gate-a-spec-pass-2.stopped-3tier.md b/.context/codex-reviews/gate-a-spec-pass-2.stopped-3tier.md new file mode 100644 index 0000000..020d2e9 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-2.stopped-3tier.md @@ -0,0 +1,40 @@ +BLOCKER | high | §3 "Capabilities" and "Gate B" inputs | the tier-2 reviewer is limited to Read, Grep, Glob and one write path, but none of those tools can compute a Git diff from baseSha and headSha | Gate B's separately named tier-2 mechanism cannot execute at all, so story AC 3 is only prose | add a concrete read-only Git-range capability and its failure handling, or choose an input mechanism the stated toolset can actually consume +BLOCKER | high | §3 "Capabilities" | "write to exactly one path" is described as a real capability, but the design names no agent definition or enforcing mechanism and Claude Code's Write tool grants general path access | the core containment claim is false under the prompt-only design, exposing the repository to a same-family reviewer processing untrusted input | define the shipped reviewer agent and a verified path guard, or explicitly downgrade this to an instruction-backed rule and reassess the rejected parent-write design +BLOCKER | high | §3 "Untrusted input" | the wrapper says only it supplies authoritative instructions, but a normal non-fork Claude subagent automatically loads the repository's CLAUDE.md hierarchy as instructions | a malicious or merely conflicting reviewed repository can steer the reviewer before it interprets the artifact as data, defeating the stated trust boundary | use an invocation that demonstrably excludes repository instruction files or redesign the trust model to name which loaded instructions remain authoritative and how conflicts stop the pass +BLOCKER | high | §6 final-pass identity versus §7 "Durable identity" | the todos.md debt row must be inside the WIP snapshot before the final pass, yet it must record the landed commit SHA, which does not exist until the reviewed commit message is written and is unknowable before a later squash | the required record is self-referential and cannot be both accurate and covered by the final pass | replace the landed-SHA-in-snapshot requirement with a stable pre-land identity plus a defined post-land reconciliation record, or move durable debt tracking to a mechanism that does not change the reviewed commit +MAJOR | high | §5 "Invariant 2" | the claim that .off always sat outside invariant 2 is contrary to the shipped policy: current CLAUDE.md says .off only silences reminders and "the gates still apply," while the hook continues classification and state tracking | this is a substantive fail-open exception being disguised as scope clarification, violating the old-condition accounting rule | state explicitly that the design changes the meaning of .off and invariant 2, then account for the prior "gates still apply" condition as deliberately dropped or preserve it +MAJOR | high | §5 degraded posture | .context/codex-gate.off suppresses every hook message in the workspace, not only the authorized degraded cycle or gate | one outage authorization silently removes reminders for unrelated commits and gate cycles, broadening the false-checkmark surface far beyond the decision recorded | keep the hook audible with a degraded-specific instruction or introduce cycle-scoped silence; do not prescribe the global sentinel as though it were scoped +MAJOR | high | §5 "The reason record" and exit | the design does not handle a pre-existing .off file used for init-time inactivity, a user opt-out, or another degraded cycle; entry can replace its meaning and exit unconditionally deletes it | retries and overlapping uses can lose the prior reason or re-enable a workspace the human still intended to keep silent | define ownership, prior-state preservation, idempotent entry, and conditional exit that removes only the authorization this cycle created +MAJOR | high | §5 entry, extension and exit | the workspace-global sentinel has no concurrency rule, so one cycle can delete it while another authorized degraded cycle is still active | concurrent agents or gates can re-enable reminders prematurely or inherit an authorization that was never granted to them | serialize degraded cycles or use scoped leases with explicit owners and release rules +MAJOR | high | §5 exit and §10 "temporary" narrowing | no rule requires an availability re-check, exit proposal, or tier-1 use once Codex returns; exit and extension are left entirely to another human decision | tier 2 can become the forbidden standing arrangement simply by never observing or acting on recovery | require a check before each pass or cycle, define an expiry or review window, and stop tier-2 entry when tier 1 is available +MAJOR | high | §5 entry predicate items 1-2 | the "vendor is known unavailable" branch permits no initial call, while item 2 still requires that pass's recovery attempt to have been spent | this state cannot satisfy all entry conditions without performing pointless calls against a known outage | define separate observable predicates for attempted failure and independently verified outage, including whether recovery is waived and who verifies it +MAJOR | high | §3 tier 3 activation | "unavailable or unwarranted" is undefined and allows a human to bypass an available tier-2 review for any reason | the broadest fail-open tier is reachable through an uncheckable adjective rather than a bounded exception | remove "unwarranted" or enumerate closed reasons, evidence, and stop conditions for each +MAJOR | high | §3 tier 3, §7 and §9 | tier 3 closes without qualifying passes, yet §7 calls it a clean cycle and §9 requires every tier-3 pass to write a per-pass line | the cycle has no coherent pass count, clean signal, or final-pass semantics, so operators can produce mutually incompatible records | define tier 3 as a separate no-pass closure with explicit N/A fields and no per-pass requirement, or specify an actual human-pass protocol +MAJOR | high | §8 tier-marker schema | the canonical marker requires a reviewer model surface, same-family statement, pass count and final-pass cleanliness even for tier 3, whose authority is a human and which has no qualifying passes | the required durable disclosure cannot truthfully represent one of the three tiers | provide tier-specific canonical schemas and complete examples, including the exact tier-3 values +MAJOR | high | §3 and §8 human authority | the only protection against an agent weakening its own review is a prose block the agent can write, a limitation §15 admits without adding a live-confirmation protocol | repository prompt injection or agent error can fabricate the sole authorization control and still leave a syntactically valid record | require direct human confirmation in the active interaction before mutation, distinguish authorization from its later record, and define the stop behavior when confirmation cannot be verified +MAJOR | high | §5 reason record and §8 logged decision | entry authorization appears to live in the ignored .off file, which is deleted on exit, while the cycle marker does not include who authorized what | the durable main-history record can prove only that degradation was claimed, not that the human-authority condition was met | require the entry decision block itself in the cycle-closing commit body and carry it through the same amend and merge chain +MAJOR | high | §3 "unchanged" recovery rules | the inherited recovery procedure is MCP-specific, using sessionId, specSessionId and qualitySessionId, but tier 2 runs through Agent and has different partial-output, task-ID, stop and resume behavior | tier-2 timeouts, quota failures and partial writes have no valid one-attempt procedure, and resuming may also violate fresh-context isolation | specify tier-2 failure shapes and whether recovery creates a new agent, resumes one, or discounts the pass, with an exact shared-budget state machine +MAJOR | high | §6 and §11 cross-tier call identity | §2 says riders change tier 1 too, but the clean-worktree and pinned-head procedure is scoped to tier 2 | sequential tier-1 spec and quality calls can review different content and still be combined as one pass | define one base/head identity and before-between-after checks for the pair at every tier, using the actual semantics of mcp__codex__review +MAJOR | high | §6 "git status --porcelain produces no output" | the rule checks empty output but does not require a zero exit status | a Git error can produce no stdout and be misclassified as a clean worktree, creating the false provenance checkmark this section is meant to prevent | require successful exit plus empty output and name distinct diagnostic causes for command failure versus dirty state +MAJOR | high | §6 content stability | clean checks before and after a pass do not prevent another agent or process from changing files during the review and restoring the same clean HEAD before the re-check | the reviewer can observe a mixed transient state while every endpoint identity check passes | prohibit and verify concurrent writers for the whole pair, or review in an immutable isolated checkout keyed to the two SHAs +MAJOR | high | §6 "Tier switching" | only a prior tier-1 task is stopped; prior tier-2 agents, resumed agents, and other writers that can target the same slot are not covered | a late same-tier result can overwrite a valid file or race the next pass exactly like the tier-1 case | stop and confirm every task with access to the target or use invocation-unique targets and immutable publication +MAJOR | high | §6 pass identity record | the spec says each pass records baseSha and headSha but names no file, field, actor, or acceptance check, and the findings grammar forbids metadata | a well-formed findings file can be accepted without evidence of which range the reviewer actually used | define a separate provenance record with canonical syntax, bind it to the invocation, and validate it before counting the pass +MAJOR | high | §11 rider (a) recovery | the prose says any branch failure discards both files and reruns the pair, while the accounting table keeps single-branch resume that deletes and recreates only its own target | operators cannot tell whether a successful first branch survives or whether recovery must repeat both, so the shared budget has two conflicting transitions | choose one recovery graph and update every old condition to that same rule +MAJOR | high | §11 rider (b) versus story AC 6 | the story requires severity to be a closed set of permitted tokens, but the spec declares every other token valid after normalization | this silently redefines rather than satisfies the acceptance criterion and lets malformed output count | make non-enum tokens INCOMPLETE, or explicitly amend the story before adopting tolerant normalization +MAJOR | high | §3 capabilities and §9 per-pass disclosure | the reviewer may write exactly the findings path, but every degraded pass also requires a tier line in a dispositions file; the design never says who writes that second file or when it becomes trustworthy | the required per-pass disclosure either violates the capability boundary or becomes a parent-authored assertion misdescribed as reviewer output | name the actor, ordering, and validation rule, and reconcile the write capability with the required companion +MAJOR | high | §8 record schemas and §9 per-pass line | the tier marker, logged decision and per-pass tier line are field lists without full literal examples, delimiters, required N/A values, or duplicate handling | these are new shipped prompt outputs and fail prompt-standards item 4, while inconsistent renderings cannot be reliably carried or audited | include complete canonical examples and parsing-by-reader rules for every tier and use +MAJOR | high | §7 debt closing conditions | "the re-review ran" does not say whether the normal pass floor, final-clean rule, profile lenses, file validation, or artifact-specific prompt applies, and "dispositioned" names no decision authority | a token single pass followed by unsupported dismissals can discharge the only compensating control | define the exact tier-1 repayment protocol and who may accept each finding disposition +MAJOR | high | §7 re-review obligation | availability return is the trigger, but §15 concedes nothing detects it and no regular workflow step checks open rows | the compensating control can remain dormant forever even after the outage ends, making the high-risk debt promise observationally empty | add a mandatory availability and open-debt check at a recurring prompt entry point and define what blocks or alerts when rows are due +MAJOR | high | §7 durable identity and §9 carry chain | only squash and ordinary amend paths are discussed; rebase-merge, cherry-pick, and rewritten commits change the landed SHA without the specified squash reconciliation | debt rows and markers can point to non-landed identities or disappear under supported Git integration strategies | enumerate supported merge strategies and define identity and disclosure reconciliation for each, or prohibit the unsupported ones with a stop condition +MAJOR | high | §9 squash body | carrying markers into the squash body depends on the human or hosting platform merger, but no role, handoff, pre-merge check, or failure response is assigned | story AC 2 depends on main history, yet the author can complete every local step and still lose the only durable disclosure at merge | assign the merge-time actor, provide exact instructions and verification, and stop declaring the AC met until main is checked +MAJOR | high | §9-§11 old-condition tables | the accounting covers amend mechanics and part of rider (a), but omits §5's existing "gates still apply," "gate itself is not optional," loop closure, incomplete-pass, profile, and tier-switch conditions affected by a no-pass tier and global silence | the spec violates the explicit AGENTS.md rule for replacing a decision procedure and hides which protections were dropped | build a complete old-condition inventory from current §5 and mark every condition kept, moved, narrowed, or deliberately dropped +MAJOR | high | §10 affected sites | README's product summary, plugin.json's "two independent cross-model review gates" description, marketplace metadata, and categorical independent-review passages in docs/coding-workflow.md are outside the table | shipped product claims remain false after tier 2 and tier 3 are introduced, recreating the documentation-drift class the spec itself identifies | search for the claim rather than the prior phrase and add every falsified site to the change surface +MAJOR | high | §12 validation evidence | the proposed verification exercises one happy tier-2 Gate-B cycle and two negatives but not Gate A, tier 3, mixed tiers, recovery, global-off ownership, concurrent writers, multiple WIPs, actual squash, debt creation/closure, or rebase | a high-risk prompt change can pass its named evidence while most fail-open and data-loss branches remain untested | add a state-transition matrix covering both gates, all tiers, entry/exit, failures, concurrency, carry paths, and debt lifecycle +MAJOR | high | §13 downstream projects | the motivating case is an already-configured project losing Codex mid-flight, yet existing initialized projects retain the old prohibition and receive no ladder on plugin update | the shipped change does not solve the stated incident class for the installed user base except through an undocumented-per-project manual migration | make migration an explicit product path with version detection and instructions, or narrow the desired outcome and release claims to new scaffolds plus named manual adopters +MAJOR | high | §14 parked gate-cycle slot collision | the spec acknowledges that pass 1 already deleted a previous cycle's identically named target but defers the fix while making degraded per-pass records required | later cycles can erase audit and resume evidence on which the new disclosure protocol relies | introduce a cycle-stable unique identifier in slot names in this change and define cleanup separately +MAJOR | high | rollback path | no procedure explains what happens if the ladder prompts are reverted while .off remains, degraded cycles are open, or debt rows exist | rollback can leave an older §5 that says gates apply while reminders remain globally silent and obligations become orphaned | define rollback prerequisites that re-enable the hook, preserve disclosure and debt, and close or migrate active cycles +MAJOR | medium | §5 failure record and §8 degraded marker | the design records failure shape, cause, dates and decision-maker in ignored files and public commit history without a redaction rule | authentication errors and vendor diagnostics can contain account identifiers, endpoints, tokens, or other sensitive operational data | use a closed safe cause enum, prohibit secrets and raw error payloads, and specify redacted diagnostic storage +MAJOR | high | §3 tier 3 rationale and §8 waiver vocabulary | tier 3 is called existing practice and "no new waiver mechanism," but init-time gateless scaffolding closes no active gate and profile overrides change evidence mode rather than authorize a zero-pass gate closure | the design invents a materially broader waiver while claiming continuity, obscuring its security and governance cost | describe tier 3 as a new exception, identify its new authority and abuse surface, and account for the old no-exception conditions explicitly +MINOR | high | §6 "HEAD is then the complete content identity" | git status omits ignored files and HEAD identifies committed tracked content, not every byte present in the workspace | the categorical wording outruns the comparison and violates the gate-proof wording rule | say that HEAD identifies the reviewed committed range under the successful clean-status precondition and name ignored content as outside it +MINOR | medium | §9 marker aggregation and collapse | "collect, deduplicate and validate" gives no identity key for deciding when two similar markers are duplicates or conflicts | a manual collapse can silently discard distinct Gate-A and Gate-B records or retain contradictory copies | define the deduplication key and require conflicts to stop the close +END OF FINDINGS (39 total) diff --git a/.context/codex-reviews/gate-a-spec-pass-2.stopped-tier3-core.md b/.context/codex-reviews/gate-a-spec-pass-2.stopped-tier3-core.md new file mode 100644 index 0000000..1b2656a --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-2.stopped-tier3-core.md @@ -0,0 +1,30 @@ +BLOCKER | high | §3.5 step 2 | The Gate-A docs-only test says to enumerate every path in the prospective Git tree and require every one to be the artifact or exempt prose, but a tree contains the entire repository, including unchanged prompts, scripts, and plugin files | Taken literally, every real Gate-A tree is refused; if implemented as an unstated changed-path test, the design violates the gate-proof rule because it never names the base-to-prospective-tree comparison that actually establishes docs-only | Define the exact parent-tree versus prospective-tree diff, including additions, deletions, renames, modes, and type changes, and classify those changed paths only +BLOCKER | high | §3.5 steps 2 and 6 | The dedicated index is never initialized from the intended parent tree before selected content is staged and committed | A new GIT_INDEX_FILE starts empty, so following the procedure can create a tree that deletes every unstaged tracked path; this is a direct data-loss path | Require a fresh index path, initialize it with git read-tree at a fixed expected parent, apply the intended changes, and verify the resulting tree before authorization +BLOCKER | high | §3.5 steps 2, 6, and 7 | The procedure isolates the index but never binds or compares the expected parent or branch tip, so another process can advance HEAD between authorization and git commit | The closing commit can acquire an unauthorized parent and use the previously authorized tree to undo concurrent work while step 7 still passes because it compares only tree ids | Record the expected parent commit in the digest and use a compare-and-swap style pre-commit check that HEAD still equals it; after commit verify both parent and tree +BLOCKER | high | §3.4 and §3.5 steps 3 through 7 | Authorization is said to bind the exact marker, decision, checklist, and evidence text, but post-commit validation reads back only the tree id | A changed, omitted, duplicated, or editor-rewritten commit body can survive with the authorized tree and be declared a valid closure, defeating the durable disclosure and exact-record authorization | Compute the authorized record digest with a specified framing and hash, then read the committed message back and compare its normalized record bytes and expected parent as well as its tree +MAJOR | high | §3.4, §3.5 step 3, and §5.4 | The authorization digest has no hash algorithm, byte framing, field ordering, or stored expected value | Different agents can hash ambiguous concatenations differently, and the later instruction that the digest must still match is not implementable or auditable from the record | Specify a canonical length-delimited serialization, exact hash algorithm, and an Authorization-digest field included in the decision record +BLOCKER | high | §3.4 and §5.2 | The complete decision block, including its decision timestamp and human-supplied reason, must be inside the digest shown before the human answers, but the answer is what creates those values | Filling the real answer after the pause changes the authorized bytes and forces a restart; pre-filling them records a predicted decision time and rationale rather than the decision that occurred, so the authorization loop is circular | Separate a pre-authorization request digest from the post-answer attestation, have the human return the digest plus decision data, and bind the committed block to that returned value without requiring foreknowledge +BLOCKER | high | §4 lines 272–289 | The history-derived continuation rule exists only for an A-plan waiver at executing-plans; there is no equivalent rule for the A-spec below-floor reminder at writing-plans | A correctly waived spec still cannot advance into planning under the stated compliant-agent procedure, which leaves the original Gate-A outage stall unsolved | Define the symmetric A-spec continuation at writing-plans, bound to the current spec blob and its valid marker and decision block +BLOCKER | high | §4 lines 278–286 | A-plan continuation searches for the commit that introduced the plan and validates the blob in that commit, but never compares that blob to the plan currently about to be executed | A plan can be edited after its waived commit and the old authorization will still clear the reminder, creating a second gate-off path for unreviewed plan bytes | Resolve the latest effective plan blob at execution time and require it to equal the Waived-target blob; any later plan change must require a new Gate-A cycle +BLOCKER | high | §6 merge-strategy table | The table says any branch-head movement after authorization voids the authorization, including under ordinary merge, while the documented Gate-A flow necessarily adds plan and implementation commits after the waived spec or plan commit | Ordinary merge becomes impossible after any Gate-A waiver, or the rule must be silently ignored; restarting Gate A at the final head is also impossible because the head contains product paths forbidden by §3.5 | Scope immutability to the authorized commit and artifact for Gate A, and separately bind only the final branch head for Gate B and pre-merge validation +BLOCKER | high | §3.2 source enum, auth-failed | Authentication failure is sufficient after one canonical call, but authentication failure is locally manufacturable by removing, replacing, or withholding credentials while keeping the request itself canonical | A compliant-looking author can create the outage condition on demand and close a non-trivial change with zero review, exactly the gate-off abuse path the design says it excludes | Treat local authentication failure as a configuration error that must be repaired, or require independently verifiable evidence that valid previously working credentials and account configuration were unchanged +MAJOR | high | §3.2 and §3.5 step 5 | attempt-failed requires a canonical call plus a recovery attempt on every revalidation, but the post-answer step prescribes only one final canonical probe | One probe is insufficient under the source definition; spending another recovery appears to create a second shared recovery budget for the same pass or closure, so there is no defined conforming transport-failure path | State exactly whether step 5 repeats both calls, how those calls relate to the one-attempt-per-pass budget, and what happens for every possible first and recovery result +BLOCKER | high | §3.1(2), §5.1 Passes completed, and §12 slot collision | The claimed complete pass ledger has no authoritative enumeration source: current §5 does not durably record pass ids as passes occur, slot files omit the cycle id and collide, and the hook counter explicitly is not evidence | After a restart or collision, the procedure cannot distinguish no prior pass from a lost adverse pass, so an outage can erase findings and still be represented as Passes completed: none | Add a cycle-unique durable or protected pass ledger, or fail closed whenever any prior call or cycle ambiguity exists; do not claim completeness from reusable .context slots +BLOCKER | high | §3.1(4), §5.1 Profiles, and §7 row 60 | Tier 3 validates only stories the artifact currently cites and explicitly permits unprofiled closure, with no binding to story paths used earlier in the cycle or otherwise known to govern the work | Removing or omitting a story citation immediately before waiver drops the high-risk mode, its evidence entry, and its lenses while still following the written three-case rule | Bind the waiver to the union of story paths recorded at cycle start and every prior pass or commit, and treat removal or unexplained absence as an unresolvable profile rather than unprofiled +MAJOR | high | §5.1 Profiles and §3.1 Gate-B multi-story rule | The marker grammar allows either unprofiled or a list of story paths, but §5 supports a mixed cycle containing both profiled and unprofiled stories | Such a cycle cannot durably disclose which cited paths were profiled and which legitimately owed no evidence, so later validation cannot reconstruct the per-story obligations | Define a per-story representation that records every cited path and its profiled or unprofiled status without copying mode values +MAJOR | high | §3.1 checklist, §3.4, §5.1, and §6 | The human must answer every lens, invariant, standard, and self-check item in writing and authorize the full answered checklist, but the durable schema carries only four words such as lenses answered, and the carry chain preserves only the two blocks and evidence entries | The item-by-item compensating review disappears on amend or squash, leaving no observable basis for the human decision and making checklist completion an unverifiable assertion | Specify a bounded checklist record containing every question and answer, include it in the digest, and carry that record through amend and both merge strategies +MAJOR | high | §5.1 Waiver-checklist grammar | A result may be an answered-with-exception note even though §3.1 says every item must be answered and any unanswered item fails the precondition; no permitted exception classes are defined | A free-text exception becomes an unbounded route around the only compensating evidence required for zero-pass closure | Remove exceptions or define a closed set that cannot excuse a missing answer, failed invariant check, failed prompt standard, or missing self-check +MAJOR | high | §5.1 Waived-target, §6, and CLAUDE.md §5 Mechanics | Gate B names baseSha..headSha, but closure amends the WIP commit, changing headSha and making the named WIP object potentially unreachable; the resulting commit cannot name its own sha in its body | The durable marker on main can point to a range that no longer resolves and therefore cannot identify the bytes waived from history alone | Define the Gate-B target using non-self-referential values that survive amend, such as expected parent plus authorized tree and an exact diff digest, and validate those values after commit +MAJOR | high | §3.5 step 7 and §6 incident recovery | A tree mismatch may be handled by a revert, but reverting does not remove the invalid commit or its valid-looking Reviewer-tier marker from reachable history | The documented discovery query will continue to find a closure the design itself says is invalid, with no machine-readable invalidation relation | Before publication require history repair; if already published, add an explicit incident marker that names and invalidates the bad marker and update the discovery and validation procedure to honor it +BLOCKER | high | §7 old-condition inventory | The claimed clause-by-clause inventory omits old requirements including Gate A's exact placement before writing-plans and executing-plans, the sweep's no-side-effects and do-not-run-quoted-commands limits, Gate B's explicit AGENTS.md check, companion deletion and zero-finding rules, and the rule that a changed evidence entry invalidates the clean pass | This fails the AGENTS.md non-negotiable condition-accounting rule and makes the inventory's fourth exhaustiveness claim false again | Derive rows from every normative sentence in CLAUDE.md §5, split compressed rows where effects differ, and mark each omitted condition kept, narrowed, moved, or overturned +BLOCKER | high | §7 rows 52 and 53 versus §3.1 waiver checklist | The inventory marks the standing Gate-B falsification lens and its size, value, or position search kept untouched, but a tier-3 Gate-B closure has no successful reviewer call and the human checklist does not include either question | Tier 3 silently drops the established docs-drift defense on the path with no independent review, contradicting the inventory by effect | Mark both conditions narrowed and add both questions, including the required repository search, to the durable human waiver checklist +BLOCKER | high | §10 S1 and S1 negative scenarios | S1 expects NEW to return CLOSE after listing preconditions and human authorization, but it omits the ordered-close requirements: dedicated index fixation, exact digest authorization, post-answer canonical probe, age check, immediate commit, and read-back comparison | A compliant reader should refuse CLOSE now, so the proposed concrete check cannot satisfy the story's plus-check obligation even when the new procedure is correct | Make S1 a fully sequenced transcript reaching the precise decision point, and add negative variants for each ordered-close and binding condition +MAJOR | high | §10 S1 | S1 spends a recovery call after quota exhaustion, while §3.2 says quota-observed requires one call and only attempt-failed consumes call plus recovery | The supposedly frozen positive case violates the new call-budget procedure, so a compliant NEW verdict may be REFUSE for the wrong reason | Use one quota observation for the quota case, or change S1 to two transport failures and assert attempt-failed consistently +MAJOR | medium | §10 S2 | The expected OLD verdict ACCEPT is not entailed by old §5: acceptance requires finding lines in the format above, whose severity example is MAJOR, so a reader can reasonably treat IMPORTANT as structurally invalid even before the new closed enum | The differential may fail because the asserted prior behavior is ambiguous, not because the change is absent, so it is not a reliable counterfactual | Use an old-text behavior that is unambiguously accepted before the change, or first add a mechanically justified oracle for the old grammar +MAJOR | high | §10 Driver and Scenarios | The probe calls itself completely specified but provides no executable runner command, exact full prompt bytes and delimiters, model API settings, environment sanitization, timeout, output capture, or reproducible second-line clause matcher | Different implementers can run materially different probes and still record the expected verdicts, so the check is neither implementable as written nor repeatable evidence | Specify a pinned executable invocation and exact scenario fixtures plus deterministic parsing and assertion rules, or add a real harness artifact and test it +MAJOR | high | §10 Driver versus AGENTS.md invariant 5 | The headless agent runner is an executable dependency, yet only the model is promised to be named later; no exact runner, package, or model version is fixed in the design | The validation run can float between implementations and bypass the repository's exact-pinning invariant | Name and exactly pin every executing component used by the probe, including the CLI or API client and a non-floating model identifier +MAJOR | high | §10 What it establishes and counterfactual | Three matching outputs from the implementer's probabilistic model establish only that that model emitted those tokens under those invocations, not that a compliant reader derives the procedure's semantics or that the wiring would reveal a false semantic claim | This recreates the unverified-evidence claim the probe was introduced to fix and outruns the exact comparison actually performed, violating the gate-proof calibration rule | Calibrate the claim to observed model outputs and pair it with a deterministic structural oracle, or stop treating it as the plus-check until a test can distinguish procedure semantics without assuming the model is compliant +MAJOR | high | §5.4 field grammar and §3.5 authorization digest | Escaping and unescaping are not fully specified: unescaped delimiter detection lacks backslash-parity rules, reverse unescape order is absent, and lists and block concatenation have no canonical framing | Two readers can parse the same bytes differently or hash distinct records identically by concatenation, undermining idempotency, conflict detection, and exact authorization | Define a formal grammar with parity-aware escaping, reverse-order unescaping, list separators and path constraints, and length-delimited digest framing; add accept and reject vectors +MINOR | high | §3.2 status-page source | status-page is listed as a Cause source but is never sufficient and revalidation requires a canonical probe whose result is then recorded as quota-observed, auth-failed, attempt-failed, success, defective, or unknown | The status-page enum value is unreachable as a recorded revalidated Cause, contradicting the claim that the closed enum covers both branches | Remove status-page from the Cause enum or define a reachable combined evidence value and exact transition that preserves it +MINOR | high | §4 lines 278–288 | The design says the state-free derivation clears one reminder once, but the same marker is deliberately re-derived on every resumed or repeated executing-plans invocation and no state records prior clearance | The observability and idempotency claim is stronger than the mechanism; the authorization can clear the same reminder repeatedly for the unchanged plan | Say it clears the reminder whenever derived for exactly that unchanged plan, or add state if once-only consumption is actually required +END OF FINDINGS (29 total) diff --git a/.context/codex-reviews/gate-a-spec-pass-20-dispositions.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-20-dispositions.pre-2026-08-14.md new file mode 100644 index 0000000..d3b3918 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-20-dispositions.pre-2026-08-14.md @@ -0,0 +1,42 @@ +# Gate A — spec — pass 20 dispositions + +Artifact: `docs/superpowers/specs/2026-08-05-hardening-ledger-supersession-design.md`. +6 findings: 6 Major. **No Blocker, no Minor, no Nit.** All six valid. **All six applied.** + +Reviewer: `mcp-codex-dev@1.0.1`, `codex-cli 0.144.5`, model `gpt-5.6-sol`. Pass file valid: +7 lines, 6 finding lines, terminator exact, no stray artifact. + +Passes 4 → 20: **11, 14, 16, 15, 9, 10, 5, 12, 8, 6, 3, 4, 7, 6, 6, 5, 6.** + +## Shape of the pass + +**Five of six land on check 1d — the property pass 19 rewrote.** That is the cycle's standing +pattern (the previous pass's fix is subtly wrong) but at a severity and concentration the last +three passes did not show: pass 19 replaced an unsatisfiable property with a satisfiable one and +under-specified almost everything the replacement newly needed. None of the six is claim-sharpening, +so the termination rule does not fire. + +## Verdicts + +| # | Sev | Verdict | Applied as | +|---|---|---|---| +| 1 | Major | **Valid — the replacement property is weaker than it reads.** "No inert added entry stands uncorrected" is satisfied by any later non-inert added entry, including one marking an unrelated row. Nothing links two entries, and the only mechanism that could — an entry identity — is the wire format §2.2 refuses. So the old check caught the mistyped-locator case and the new one can pass while the target row stays unmarked. | 1d split into an explicitly **unequal** pair: a mechanical **positional floor** that says what it is, and the repair judgement moved to **1f**, which now owes five confirmations for the successor entry. §8's human-read bullet records the gap. | +| 2 | Major | **Valid, and mechanically checkable.** CommonMark needs a blank line between the entry list and the `Columns:` paragraph or `Columns:` renders inside the list item — so the prescribed layout puts a non-entry line inside the block, and an oracle demanding every added line in the block parse would reject the design's own layout. | The candidate interval is now stated (label → `Columns:`, both exclusive) and candidates are its **non-blank** lines; the structural blank is named as required rather than exempted by silence. | +| 3 | Major | **Valid — the withdrawn uniqueness guarantee can survive in the implementation.** §3.1's entry is fragmentless and matches exactly one row today, so a checker that ignored fragments or still demanded exactly one match passes this change. No oracle covered fragment-aware matching, zero/one/many, or the on-or-before boundary. | A third oracle, with named fixtures: the fragmentless locator matching **both** `2026-07-18` `docs-drift` rows, a fragment narrowing that pair to one, a fragment matching neither, and rows before / **on** / after the entry date. | +| 4 | Major | **Valid, low blast radius here.** Identical duplicate entry lines — which `merge=union` can produce — make occurrence alignment ambiguous, and in an *inert, valid, inert* sequence the verdict would depend on which occurrence a differ attributed to the base. For **this** change the base carries no block at all, so the added set is unambiguous; the property was still stated undecidably. | The alignment is pinned: the base entry sequence must remain an **exact prefix**, added entries are the remaining suffix, and the empty-prefix case for this change is named. | +| 5 | Major | **Valid — and it is the gap I named in pass 19's own dispositions.** §2.1's resolution floor is stated for rows; §2.2 said "never edited, never removed" with no floor, so when an entry becomes immutable was left to analogy — which is exactly how the amendable class came back three times at row level. 1d's whole rationale depends on the answer. | The floor is written into the **shared prose** (so it reaches every scaffolded ledger) and given a rationale paragraph in §2.1. New **anchor 35**. | +| 6 | Major | **Valid.** Entry recognition named the two prose fields without requiring content, and 1f covered only the mandated entry — so a locator-valid line with empty fields could count as the non-inert successor for an earlier typo, and no read would ever open it. | Recognition now requires **both prose fields non-empty**; 1f's new paragraph extends the four confirmations to the successor entry. | + +## Also applied, not from a finding + +Two clauses added to the shared prose across passes 19–20 (the fragment-narrowing limitation, the +entry floor) were unanchored, so check 2 could not detect their deletion from both surfaces. Both +now carry anchors — **34** and **35** — the count claims moved thirty-three → **thirty-five** at +both sites, and §8's known-unanchored list was corrected: the narrowing is anchored now, its +no-double-quote condition still is not. + +## Sweep after applying + +35 anchors, numbering contiguous, each exactly once in the convention prose, none a substring of +another, no CommonMark padding asymmetry; six fence lines, balanced; both count claims read +thirty-five. diff --git a/.context/codex-reviews/gate-a-spec-pass-20.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-20.pre-2026-08-14.md new file mode 100644 index 0000000..78bb612 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-20.pre-2026-08-14.md @@ -0,0 +1,7 @@ +MAJOR | high | §6 Check 1d "no added entry that is inert may stand uncorrected" | any later non-inert added entry satisfies the stated condition even when it marks an unrelated row, so the check cannot decide that it corrected the inert entry | the old check caught that inert typo, while the replacement can report the target correctly marked although the intended row remains unmarked | either add a repair identity that links the two entries, or move the correction relationship to a named read and state the mechanical floor only as a positional non-inert-successor check +MAJOR | high | §6 Check 1d oracle "An entry from a line that is not one" | it says every added line inside the block that does not parse as an entry must fail, but the intended Markdown layout needs an added blank line between the list entry and the following `Columns:` paragraph or CommonMark renders `Columns:` inside the list item | a literal implementation rejects the prescribed layout, while silently exempting the blank contradicts the oracle | define the candidate interval precisely and exempt the single structural blank separator while continuing to fail every nonblank unparseable content line +MAJOR | high | §6 Check 1d oracles | no oracle distinguishes fragment-aware matching, zero versus one versus many matching rows, or the on-or-before date boundary; the real §3.1 entry is fragmentless and currently unique, so a checker that ignores fragments or still requires exactly one match can pass this change | the validation can preserve the withdrawn uniqueness guarantee and misclassify a narrowed or multi-row locator while claiming match semantics | add explicit oracle fixtures for a fragmentless locator matching both `2026-07-18` `docs-drift` rows, a fragment narrowing to one, a fragment matching none, and otherwise-matching rows before, on, and after the entry date +MAJOR | medium | §6 Check 1d "entries added by this change" and "a later added entry" | when the base and current block contain identical duplicate entry lines, sequence differencing can attribute either occurrence to the base; in a sequence such as inert, valid, inert that choice changes whether the added inert entry has a later valid successor | pass or fail depends on an unspecified diff alignment, so the property is not decidable from the two file states under a convention that preserves duplicate union-merged entries | require the base entry sequence to remain an exact prefix and define the added entries as the remaining suffix, or specify another occurrence-alignment rule whose result is invariant for duplicates +MAJOR | high | §2.1 resolution floor and §2.2 "Entries are never edited" | the floor is defined only for rows; the shipped prose never says when a drafted or uncommitted line becomes an immutable entry, even though §6 relies on a mistyped locator becoming an entry immediately and forbids repairing it in place | authors can reasonably edit an uncommitted complete entry or treat a malformed near-entry as editable, recreating the amendable class and bypassing the sanctioned typo-then-append repair | state that once text exists as a complete entry-shaped line it is never edited or removed, committed or not, and that only drafting before such a line exists is below the entry rule +MAJOR | medium | §6 Check 1d entry-recognition oracle and §6 Check 1f | recognition names two remaining prose fields but never requires them to be nonempty, and the named read covers only the mandated §3.1 entry; a later locator-valid line with empty `what is false` or citation fields can therefore count as the non-inert correction for an earlier typo | the check can say an inert entry was corrected even though the replacement conveys neither the false claim nor where the current answer lives | require both prose fields to be nonempty during recognition and extend the named read to every later entry relied on to repair an inert addition +END OF FINDINGS (6 total) diff --git a/.context/codex-reviews/gate-a-spec-pass-21-dispositions.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-21-dispositions.pre-2026-08-14.md new file mode 100644 index 0000000..6401214 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-21-dispositions.pre-2026-08-14.md @@ -0,0 +1,55 @@ +# Gate A — spec — pass 21 dispositions + +Artifact: `docs/superpowers/specs/2026-08-05-hardening-ledger-supersession-design.md`. +7 findings: 6 Major, 1 Minor. **No Blocker.** All seven valid. +**Five applied. Findings 1 and 2 held for Daniel** — their fix reverses wording he specified. + +Reviewer: `mcp-codex-dev@1.0.1`, `codex-cli 0.144.5`, model `gpt-5.6-sol`. Pass file valid: +8 lines, 7 finding lines, terminator exact, no stray artifact. + +Passes 4 → 21: **11, 14, 16, 15, 9, 10, 5, 12, 8, 6, 3, 4, 7, 6, 6, 5, 6, 7.** + +## The shape that matters more than the count + +The count is rising — 5, 6, 7 — and it is rising **inside machinery this cycle added at passes 19 +and 20**. Findings 1 and 2 are both about 1d's *second half* and 1f's successor paragraph; pass +20's findings 1, 4 and 6 were too. That clause is generating roughly one Major per pass and each +fix makes it larger. The other four findings this pass are ordinary under-specification of things +the change genuinely does (parsing, comparison domain, the floor's boundary, 1c's input count) and +they closed on one edit each. + +## Verdicts + +| # | Sev | Verdict | Disposition | +|---|---|---|---| +| 1 | Major | **Valid, and it breaks the clause structurally.** §2.2 admits an inert entry aimed at *a row that never existed*. For that entry no successor naming its intended row can exist — there is no such row — while 1d demands a successor and the entry floor forbids deleting the line. The validation is then permanently unsatisfiable in a state the convention explicitly sanctions. | **HELD.** | +| 2 | Major | **Valid.** "The row the inert one meant to name" is authorial intent the ledger does not encode. Even when the row exists, 1f's fifth confirmation can be asserted but not verified from the artifact, so 1f cannot carry the judgement pass 20 moved out of 1d. | **HELD** — same fix as 1. | +| 3 | Major | **Valid.** The two prose fields are free text and ` · ` is the field separator; nothing forbade or escaped it, so a permitted entry could parse two ways. | Applied: the shared prose now forbids ` · ` inside either field, and the recognition oracle cites that as why the split is unambiguous. | +| 4 | Major | **Valid.** Fragment matching never named its comparison domain — raw cell bytes, rendered text, or an unescaped cell — so a checker and a reader can disagree exactly on the rows carrying `\|` or markup, and agree everywhere else, which is the silent direction. | Applied: shared prose pins **literal, case-sensitive, against the raw `finding` as written**; the matching oracle repeats it and gains `\|`, markup and case fixtures. | +| 5 | Major | **Valid.** §2.1 said protection starts "once the line exists", the shipped prose "once a line exists as an entry" — leaving a saved half-typed line with two defensible verdicts. It also had to be settled for 1d to be *actionable*: 1d fails on an unparseable candidate, and an author who could not repair it would be stuck. | Applied: protection begins at **complete entry shape**; partial or malformed lines are drafting and may be fixed; a complete zero-match entry is protected. **Anchor 35 reworded to match.** | +| 6 | Major | **Valid.** 1c's three-input design uses one base paragraph for both surfaces, so it proves the template untouched only by leaning on the separate fact that the surfaces are byte-identical today. | Applied: four inputs — each surface against its own base — with current parity asserted separately. | +| 7 | Minor | **Valid.** `2026-02-30` is `YYYY-MM-DD`-shaped, sorts correctly, and is not a date; it would satisfy both 1e's ordering and 1d's on-or-before bound. | Applied: calendar validity required in entry recognition and in 1e, with the failure named. | + +## Sweep after applying + +35 anchors, each exactly once in the convention prose, none a substring of another, no padding +asymmetry; six fence lines, balanced; both count claims read thirty-five; 1d still declares seven +oracles. + +## The question put to Daniel + +1d's second half — *"no inert added entry stands uncorrected by a later added entry"* — is his own +pass-19 wording. Recommendation: **drop it**, leaving 1d as *"the entry §3.1 mandates matches at +least one row dated on or before its own date"*. + +- It still passes §2.2's sanctioned typo-then-correct scenario — the mandated entry **is** the + corrected one — which was his binding constraint on the rewrite. +- It still fails the mistyped-locator case it was written for: the mandated entry does not match. +- It drops with it the successor rule, the duplicate-alignment oracle pass 20 added, and 1f's fifth + confirmation — the three sites that have produced four Majors across two passes. +- What it stops covering: an **additional** inert entry appended by the same change goes undetected. + This change appends exactly one entry, so the state is not reachable here, and §6's checks are + one-time regardless — §8 already says nothing runs on a future supersession. + +The alternative is to keep the clause and add finding 1's no-target exemption plus finding 2's +evidence requirement. That is more machinery in the region that keeps generating findings. diff --git a/.context/codex-reviews/gate-a-spec-pass-21.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-21.pre-2026-08-14.md new file mode 100644 index 0000000..22effcb --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-21.pre-2026-08-14.md @@ -0,0 +1,8 @@ +MAJOR | high | §2.2 “A locator matching nothing is inert” and §6 checks 1d–1f | the design expressly includes an inert entry for “a row that never existed,” but 1d requires every inert added entry to have a later matching successor and 1f requires that successor to name the row the inert entry meant to name | no such row or matching successor exists in this admitted state, while the entry floor forbids deleting the inert line, so the convention’s validation becomes permanently unsatisfiable | distinguish a mistyped locator aimed at an existing row from a no-target entry; require a successor only for the former and explicitly sanction the latter standing inert without one +MAJOR | high | §6 check 1f, inert-entry paragraph | “the row the inert one meant to name” is authorial intent that the ledger does not encode, and the spec names no evidence from which the reader can distinguish the real repair from an unrelated later entry | even when the intended row exists, the fifth confirmation can be asserted without being verified, so 1f cannot carry the repair judgment moved out of 1d | require named evidence that makes the intended row observable, such as unchanged claim/citation fields plus an explicit author attestation, or add a correlation mechanism and define which successor is evaluated +MAJOR | high | §2.2 entry format and §6 check 1d entry-recognition oracle | the two prose fields may contain the literal ` · `, but the convention neither forbids nor escapes it and the oracle does not say which separator occurrence divides the fields | a valid prose-only entry can be parsed multiple ways or rejected as malformed, so the property is not decidable for all entries the shared prose permits | forbid ` · ` inside the two fields or define an unambiguous split rule and add fixtures with the delimiter in each field +MAJOR | high | §2.2 fragment rule and §6 check 1d matching oracle | fragment matching never states whether it is literal and case-sensitive against raw Markdown cell bytes, rendered text, or an unescaped cell value; the row oracle’s `\|` fixture tests column splitting but not this choice | findings containing Markdown escapes or markup can match for a reader and fail for the checker, or vice versa, so the locator semantics and oracle can disagree | define the comparison domain and normalization exactly, then add fragment fixtures containing `\|`, emphasis markup, and case differences +MAJOR | medium | §2.1 “Entries take the same floor” and §2.2 shared floor prose | §2.1 says protection starts “once the line exists,” while the shipped prose says it starts once a line exists “as an entry”; neither defines the status of a saved half-typed or malformed candidate line | the requested entry-level floor still has two plausible verdicts for the half-typed case—editable draft or immutable history—and check 1d merely fails the line without saying which repair is sanctioned | state that protection begins only when a line satisfies the complete entry shape, that partial or malformed candidate lines remain drafts that may be edited or removed, and that a complete zero-match entry is protected +MAJOR | medium | §6 check 1c oracle | the “all three inputs” design necessarily uses one base paragraph as the reference for both current surfaces, so it proves both were untouched only by relying on the separately stated fact that the two base surfaces are byte-identical today | if the explicit base ref contains an independent template variation, editing that template to the ledger’s bytes passes the three-input comparison even though the template paragraph changed from its own base | read four inputs and compare each current surface with its own base version, then separately assert current parity; name the paragraph extraction boundaries +MINOR | high | §6 checks 1d and 1e date oracles | the entry recognizer requires only a `YYYY-MM-DD` shape and the ordering oracle does not require a real calendar date | values such as `2026-02-30` or `2026-99-99` can pass lexical parsing and ordering while violating the shared rule that the value is the day written, weakening the on-or-before boundary | require calendar-valid dates for rows and entries, with invalid-day and invalid-month fixtures, or explicitly narrow the mechanical property to lexical shape and record that limitation +END OF FINDINGS (7 total) diff --git a/.context/codex-reviews/gate-a-spec-pass-22-dispositions.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-22-dispositions.pre-2026-08-14.md new file mode 100644 index 0000000..b9f2e07 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-22-dispositions.pre-2026-08-14.md @@ -0,0 +1,51 @@ +# Gate A — spec — pass 22 dispositions + +Artifact: `docs/superpowers/specs/2026-08-05-hardening-ledger-supersession-design.md`. +5 findings: 3 Major, 2 Minor. **No Blocker.** All five valid (finding 1 in part). +**All five applied.** + +Reviewer: `mcp-codex-dev@1.0.1`, `codex-cli 0.144.5`, model `gpt-5.6-sol`. Pass file valid: +6 lines, 5 finding lines, terminator exact, no stray artifact. + +Passes 4 → 22: **11, 14, 16, 15, 9, 10, 5, 12, 8, 6, 3, 4, 7, 6, 6, 5, 6, 7, 5.** + +## The class changed, and that is the finding about the pass + +**No mechanism defect. No unhandled path. No invariant risk.** Narrowing 1d removed the generator: +passes 20 and 21 spent thirteen findings on machinery added at 19–20, and with that machinery gone +this pass found none in its place. What it found instead is four claims-about-the-design that are +not exactly true, plus one scope reduction. Rider 4 came back with an **over**-coverage finding +rather than a missing distinction — the first time in the cycle it has run that direction. + +## Verdicts + +| # | Sev | Verdict | Applied as | +|---|---|---|---| +| 1 | Major | **Valid in part.** The evaluation point was unstated, so check 3's "exactly one position" read as a standing property, and a duplicate arriving by a later merge looked like a permanent failure with no repair. It over-reads in one respect: §6's checks run **once**, on this change, before any merge involving it, so no state can "permanently fail" them. The duplication premise is also weaker than stated — §8 records that identical label lines **coalesced** under the configured driver in a scratch reproduction. | Check 3's cardinality oracle now names its evaluation point and points at §2.2's keep-one-label repair and §8's no-standing-check record. | +| 2 | Major | **Valid, and introduced by pass 21's own fix.** §2.2 claimed "Prose only — no machine-readable fields" while §6 requires exact field splitting, calendar parsing and a `·`-delimited grammar — and pass 21's fix *shipped* a format constraint (` · ` forbidden inside a field) into every scaffolded ledger. The decision being made is about **standing consumers**, not about whether a line can be parsed. | §2.2's Format paragraph rewritten: no standing consumer parses an entry; §6's one-time checks parse a **validation-only grammar** that binds this change and nothing after it. What is refused is designing an input format for a reader that does not exist yet. | +| 3 | Major | **Valid.** 1d now concerns one fragmentless entry and treats one match and many alike, yet its oracle still demanded a general optional-fragment matcher with escape, markup and case fixtures — driving exactly the generalized parser §2.2 declines to justify. | The matching oracle narrowed to the two load-bearing distinctions: **zero from at-least-one** (the many fixture stays — it is what stops a checker implementing the withdrawn exactly-one rule) and the **on-or-before boundary**. Fragment-narrowed matching and the comparison domain are named as *not validated by this change* and recorded in §8. | +| 4 | Minor | **Valid — the AGENTS.md decision-procedure Don't, exactly.** Removing the alignment oracle at pass 21 removed with it the exact-prefix condition, which had also asserted that base entries survive unedited and in order. The §8 record covered the inert-successor coverage and not this. | 1d gained an explicit **old-condition accounting** naming all three dropped requirements, marking the prefix rule as the non-obvious one; §8 records that no check validates pre-existing entry immutability. | +| 5 | Minor | **Valid.** "For this change the state is unreachable — it appends exactly one entry" is circular: it is unreachable only if the implementation conforms, which is what a check exists to stop assuming. | Replaced with "outside the prescribed change, but a nonconforming implementation can produce it and every stated check will accept it", and the circularity named. | + +## Sweep after applying + +35 anchors, each exactly once in the convention prose, none a substring of another, no padding +asymmetry; six fence lines, balanced; both count claims read thirty-five; 1d declares seven oracles +and has seven. + +## Termination assessment — the rule fires + +Daniel's rule: *non-clean with only claim-sharpening → termination assessment, no pass 23 by +momentum.* Findings 1, 2, 4 and 5 are claim-sharpening by his definition — an evaluation point left +unstated, a design claim that outran what the design does, an unrecorded dropped condition, a +circular word. Finding 3 is not a defect at all; it **removes** required work. That is the trigger, +read honestly. + +Trend by severity, not just count: **18** 4 Major · **19** 4 Major · **20** 6 Major · **21** 6 Major +· **22** 3 Major, and none of the three a mechanism defect. The 20–21 spike was self-inflicted and +is now removed at the root. + +**Against closing immediately:** five fixes were applied after this pass, so the revision now on +disk has not been reviewed by anything. §5 requires the final pass to be clean. A single confirming +pass 23 is what the pinned exit's own condition asks for — not momentum, since the recommendation is +to close on it whether it comes back clean **or** dispositions-only, and to stop either way. diff --git a/.context/codex-reviews/gate-a-spec-pass-22.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-22.pre-2026-08-14.md new file mode 100644 index 0000000..4e40c25 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-22.pre-2026-08-14.md @@ -0,0 +1,6 @@ +MAJOR | high | §2.2 entry floor and §6 checks 1a/3 | a complete mandated entry duplicated by a union merge is immutable under “Entries are never edited, never removed” and “keep every entry,” but check 3 requires that entry to resolve to exactly one position | the prescribed concurrent state permanently fails validation and has no sanctioned recovery | let validation accept repeated byte-identical mandated entries and define their positional effect, or explicitly permit duplicate coalescing and account for that exception to the no-remove rule +MAJOR | high | §2.2 “Format” and §6 check 1d | the claim “Prose only — no machine-readable fields” and the statement that parsing would make the entry wire format contradict 1d’s required exact grammar, calendar parsing, field splitting, row parser, and semantic fixtures | the plan can treat the syntax as non-contractual while its validation depends on parsing it exactly, obscuring the compatibility surface and overstating the boundary | say that no standing consumer reads the block but the one-time validation parses a validation-only grammar, or remove the parser-dependent properties +MAJOR | high | §6 check 1d matching oracle | the narrowed property concerns one known fragmentless mandated entry and treats one and many matches identically, yet its oracle still requires a general optional-fragment matcher to distinguish one from many and handle raw escapes, markup, and case; the text itself admits the mandated entry exercises none of these distinctions | the seven oracles over-serve the stated property and can drive a generalized parser that §2.2 says is not part of the design, while a simpler check can decide the actual property | split general convention-semantics conformance into its own property with its own falsifying observation, or narrow 1d’s oracles to the distinctions needed for the mandated fragmentless entry +MINOR | medium | §6 check 1d “One entry, and no claim about any other” | removing the duplicate-alignment oracle also silently removes its exact-prefix condition, which preserved every base entry in order; the replacement records the dropped inert-successor coverage but never records that pre-existing entry preservation is now out of the check | the old-condition accounting is incomplete under the AGENTS.md decision-procedure rule, and a future reuse of the check could allow editing, removal, or reordering despite absolute entry immutability | explicitly mark exact-prefix preservation as dropped because this base has no entries and record that no check validates pre-existing entry immutability, or retain a separate base-versus-current preservation property +MINOR | high | §8 “An extra inert entry appended alongside the mandated one” | “For this change the state is unreachable — it appends exactly one entry” is circular and contradicts the bullet’s own premise: a nonconforming diff can append the extra inert entry and every stated check will accept it | calling a mechanically possible undetected implementation error unreachable understates the recorded validation gap | replace “unreachable” with “outside the prescribed change but mechanically possible in a nonconforming implementation” +END OF FINDINGS (5 total) diff --git a/.context/codex-reviews/gate-a-spec-pass-23-dispositions.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-23-dispositions.pre-2026-08-14.md new file mode 100644 index 0000000..500bec6 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-23-dispositions.pre-2026-08-14.md @@ -0,0 +1,33 @@ +# Gate A — spec — pass 23 dispositions (CLOSING PASS) + +Artifact: `docs/superpowers/specs/2026-08-05-hardening-ledger-supersession-design.md`. +4 findings: 2 Major, 2 Minor. **No Blocker.** All four valid. **ALL FOUR HELD-NOT-FIXED**, +recorded as §8 residuals per the stop decision taken before the pass ran. + +Reviewer: `mcp-codex-dev@1.0.1`, `codex-cli 0.144.5`, model `gpt-5.6-sol`. Pass file valid: +5 lines, 4 finding lines, terminator exact, no stray artifact. + +Passes 4 → 23: **11, 14, 16, 15, 9, 10, 5, 12, 8, 6, 3, 4, 7, 6, 6, 5, 6, 7, 5, 4.** + +## Verdicts + +| # | Sev | Verdict | +|---|---|---| +| 1 | Major | **Valid, and the most consequential of the four.** 1d's narrowed matching oracle names zero-from-at-least-one and the on-or-before bound; it never says the locator's **row date** must match exactly. A checker comparing fingerprint only passes the mandated entry, the two-row fixture and all three date positions, while reporting non-inert an entry whose locator matches no row — the property 1d exists to decide. The narrowing at pass 22 removed the fixture set that had incidentally covered it. | +| 2 | Major | **Valid.** Pass 22's Format repair was right about consumers and overreached on persistence: it called the grammar validation-only and binding "nothing after" this change, while the shared prose makes the shape, the ` · ` ban, calendar dates and the completeness boundary **standing rules for every future author**. §5 and §8's discriminator bullet still carry the old broader "prose-only / nothing consumes" wording. The accurate claim is *no standing machine consumer*; the syntax is a standing human convention. | +| 3 | Minor | **Valid.** 1d's scope oracle reads plural — "The mandated entry, and only entries this change adds" — and can be quantified over every added entry, recreating the unsatisfiable property pass 21 removed. | +| 4 | Minor | **Valid.** §8's format-example bullet says shape is "the whole guard". For this change, location guards it too: 1d bounds candidates to the label-to-`Columns:` interval and check 3 rejects entry-shaped lines outside it. The claim is correct about **standing** enforcement and wrong about this change's validation. | + +## Why none was fixed + +Daniel took the stop decision **before** pass 23 ran, precisely so the closing pass could not +reopen the loop: clean or dispositions-only → normal close; anything else → §8 residuals, stated +as held-not-fixed. Repairing findings 1 and 2 would have produced another revision no pass had +seen — the exact state pass 23 existed to end, and the state that generated four of this cycle's +worst passes. All four are recorded in §8 with the concrete remedy, and findings 1 and 3 land on +the executable check, which the plan writes and **Gate B reviews against the real diff**, where a +check finally has something to be checked against. + +Two of the four (1 and 4) are the cycle's signature pattern once more: **the previous pass's fix +was subtly wrong.** Pass 22 narrowed 1d's oracle and took a needed distinction with it; pass 22 +narrowed §2.2's Format claim and left two sites carrying the old one. Diminishing, and never zero. diff --git a/.context/codex-reviews/gate-a-spec-pass-23.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-23.pre-2026-08-14.md new file mode 100644 index 0000000..97cafb8 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-23.pre-2026-08-14.md @@ -0,0 +1,5 @@ +MAJOR | high | READ — §6 check 1d “Matching, to the extent the mandated entry uses it” | the claim that only zero-versus-at-least-one and the on-or-before boundary are load-bearing omits equality on the locator’s row-date field; the typo fixture covers fingerprint equality, but a checker that ignores the superseded-row date can still pass the mandated entry, the two-row fixture, and all three eligibility-date positions | such a checker reports the entry non-inert when its date-plus-fingerprint locator matches no row, so check 1d does not decide its stated property | add a fixture where only the locator row date is wrong and state that exact row-date and fingerprint equality are validated before applying the eligibility bound +MAJOR | high | READ — §2.2 “Format”; §5; §8 discriminator-rejected bullet | the repair distinguishes a one-time parser from a standing consumer, but then overreaches by calling the grammar validation-only and binding “nothing after” this change while the shipped shared prose makes the same shape, delimiter ban, valid dates, and complete-entry boundary standing rules for every future entry; §5 still says the correction is “prose-only” and “nothing consumes” it, and §8 again says §2.2 keeps the entry deliberately prose-only | the spec still gives two incompatible answers about what persists: future authors must honor the syntax and immutability boundary, even though no standing machine consumer or validator currently reads them | say the syntax is a standing human-authored convention, while §6’s machine parser is validation-only and creates no ongoing validation or compatibility promise; replace §5/§8’s prose-only and nothing-consumes claims with the narrower no-standing-machine-consumer claim +MINOR | medium | READ — §6 check 1d first oracle | “The mandated entry, and only entries this change adds” uses a plural scope and immediately stresses that every present entry is added, contradicting the surrounding rule that only the mandated entry is evaluated and §8’s claim that an extra inert addition passes | an implementer can reasonably quantify 1d over all added entries and recreate the unsatisfiable check that the narrowing removed | define selection directly: identify the exact §3.1-mandated entry in current content, prove it was added relative to the base, and explicitly ignore every other added entry for 1d +MINOR | high | READ — §8 “The format example is not an entry” | “nothing enforces that distinction beyond its shape” and “That is the whole guard” omit the one-time location guards: 1d restricts candidates to the label-to-Columns interval, and check 3 rejects entry-shaped lines outside it | this repeats the broader no-check claim immediately after §2.2 was narrowed and misstates what this change’s validation distinguishes, even though no standing check remains afterward | qualify the claim as no standing enforcement, record that this change’s 1d/3 checks also use location, and narrow the future risk to readers or checks that scan dated lines without respecting the interval +END OF FINDINGS (4 total) diff --git a/.context/codex-reviews/gate-a-spec-pass-3-dispositions.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-3-dispositions.pre-2026-08-14.md new file mode 100644 index 0000000..858c7c8 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-3-dispositions.pre-2026-08-14.md @@ -0,0 +1,15 @@ +# Gate A — spec — pass 3 dispositions + +Artifact: `docs/superpowers/specs/2026-08-05-hardening-ledger-supersession-design.md` +8 findings: 7 Major, 1 Nit. All eight validated as correct. Seven applied; one to Daniel. + +| # | Verdict | Reason | +|---|---|---| +| 1 | **Applied, human-confirmed** | Correct, and the same class as pass 1's finding 4. Surfaced rather than settled, because an agent narrowing a condition the story marks "must preserve" is the gate-off shape `AGENTS.md` names. Confirmed decision: amend all three sites — the inherited-conditions table, the desired outcome, AC 3 — with accounting, and add the story to §4's change surface. Accounting as confirmed: **kept** — the committed record's protection, what every consumer of the rule actually relied on; **narrowed** — the textual rule, "never edit a row" → "never edit a landed row", making textual what #22's Gate B already accepted in practice (row D, amend-during-authoring, reasoning in its evidence); **dropped** — nothing. The amendment records both facts: the text changed here, and the boundary it moved to was already the practiced one. | +| 2 | **Applied** | Correct, and a reachable overlap I had not seen: an open-cycle `pending` row is governed by "resolve by appending" and by "amend in place" at once. Resolved by precedence — resolving a `pending` row always appends, in any cycle, because it is a new hardening earning its own row; amendment reaches only that row's own wording. §5's disposition for that condition now says so explicitly. | +| 3 | **Applied** | Correct. "Name what stopped being true" has no accurate value for the wrong-when-written case, which is half the settled scope. Field is now "what is false", and the entry says which of the two it is. §3.1's entry updated to state that its claim *was* true when written. | +| 4 | **Applied** | Correct. Entry locators collide exactly as row locators do — two falsification-scoped entries against one row on the same day share date-plus-row — so a correcting entry could not retire exactly one of them. Same fallback, same stop, now stated for entry locators too. | +| 5 | **Applied** | Correct and the best finding of the pass. The end sentinel is *new convention text*, so it exists in neither pre-change file; the falsifying observation claimed extraction "reports equal", and an extractor reading two absent sentinels as equal would green-light the convention missing from both surfaces. Now: a missing or non-unique sentinel is a failure, and that failure *is* the pre-change observation. Three states named and each distinguished. | +| 6 | **Applied** | Correct. Grep 1 matched on dates only and grep 2 on date-plus-fingerprint, so an entry with a wrong or missing fingerprint passed both while the spec claimed they jointly established the locator resolves. Fingerprint now inside both patterns, with the reason stated inline. | +| 7 | **Applied** | Correct: authorship can be unestablishable after interruption, handoff or compaction, and the spec gave no action there. Uncertainty now resolves to *landed* — the conservative direction, costing an unneeded entry rather than an edit to a shared record — and §9 states plainly that this is a bias, not a determination. | +| 8 | **Applied** | Correct, and precisely the kind of thing the sweep exists for: the cited fragment omitted the `**` emphasis bytes, so it was semantically unique but not byte-findable in the file it points at. Citation now quotes the source markdown exactly, and rider 1 was widened to ask for byte-for-byte including emphasis markers. | diff --git a/.context/codex-reviews/gate-a-spec-pass-3-dispositions.stopped-2tier-debt.md b/.context/codex-reviews/gate-a-spec-pass-3-dispositions.stopped-2tier-debt.md new file mode 100644 index 0000000..dbeb6bb --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-3-dispositions.stopped-2tier-debt.md @@ -0,0 +1,96 @@ +# Gate A — spec — pass 3 dispositions (two-tier cycle) — CYCLE STOPPED + +32 findings (12 BLOCKER, 19 MAJOR, 1 MINOR). None dismissed. Enum held a sixth time. + +## Trajectory + +| Pass | B | M | m | Spec lines | +|---|---|---|---|---| +| 1 | 7 | 21 | 2 | 349 | +| 2 | 8 | 20 | 6 | 538 | +| 3 | **12** | 19 | 1 | 610 | + +Blockers rising while the artifact grows. Same shape as the three-tier cycle, and §5's stuck +condition applies again. + +## The finding that settles it + +**Blocker 2 — the debt-location fix creates a recursion, and it is the fix approved last +round.** Moving the debt record to a non-`.md` path so that erasing it costs a review means +the Gate-A closing commit now stages a product-classified file. That commit is therefore no +longer docs-only, so it raises a **Gate-B** obligation — during an outage, inside the very +waiver meant to escape one. Blocker 10 completes the circle: **every** debt-state transition +edits that file, so `accepted-no-review` cannot be committed during the outage without +ignoring Gate B, and marking a Gate-B debt `repaid` after a final pass changes the +fingerprint that pass covered. Codex's conclusion is the right one: *the settled "no hook +change" constraint is not compatible with the current record location.* + +## The pattern, stated plainly + +Every compensating control added to make the waiver safe has landed in one of two states: + +1. **Unenforceable** — a prose assertion the waived party authors (blockers 15, 16; §13 + already concedes it), or +2. **Recursive** — real state that itself needs gating, and gating it needs the gate that is + unavailable (blockers 2, 10). + +That is not a sequence of fixable defects. It is what a gate waiver *is* in a prompt-only +system: the thing being waived is the only mechanism available to protect the record of the +waiver. + +## The two fixes from last round both fail + +**Blocker 13 — the scoped `mode override` is not a valid §5 profile change.** §5's grammar +permits a whole effective mode in the header and requires the header to carry the override; +there is no per-portion override semantics. I invented a fourth profile mechanism, and +Daniel confirmed it on my description. A gate reader must now either still demand `+check` +or classify the header/log combination as semantically inconsistent — which §5 says is a +STOP. + +**Blocker 14 — and the gap it was taken for is not real.** A prompt-harness scenario does +discriminate: drive an unavailable reviewer and assert *prior refusal* versus *new +conditional closure*, with negative cases per precondition. So the override was taken on a +false premise as well as in an invalid form. Both must be withdrawn. + +## Blockers — accepted, all of them + +1 (Gate-A waivers are impossible under §3.1(6): implementation-derived evidence cannot exist +before implementation, and Gate A is the gate that blocks planning — where the outage +actually bit), 2, 4 (a waiver commit can delete prior debt rows in the same index; the +current-state scan then sees no cap history), 5 (the cap has no authoritative home, +derivation, multi-story semantics or parse-failure behaviour), 6 (two `accepted-no-review` +closures leave the cap at two with nothing repayable — tier 3 permanently disabled by a +permitted sequence), 7 (§5.3's equality rule and condition 2's different-handle rule are +mutually unsatisfiable), 8 (`Waived-range` SHAs are unreachable after amend/squash — the +durability argument covered only the Gate-A blob), 9 (no debt state machine), 10, 11 +(ordinary `git revert` of the waiver commit erases the row and cap history), 13, 24 (the +inventory omits operative conditions for the **third** time — lens sets, "lenses are +questions not passes", Gate-B tool/range parameters, skip-reason rules, lower-profile +carryover, the non-enforcement residuals). + +## Majors — accepted + +3, 12, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 25, 26, 27, 29, 30, 31, 32. Notable: 25, 26 +and 30 are three more **noun-based dispositions** in the inventory of exactly the kind the +`[corrected]` rows were added to warn against — "Re-review after every fix" and "`.off`; the +gates still apply" marked kept while their terminal effect changes, and "fires full Gate B" +claiming a mechanism when the hook only emits a reminder and always exits 0. 27 is a story +defect: AC 3 still asserts `sparring-briefing.md` moved entirely, which the spec itself now +contradicts. + +## Minor + +28 — the story misstates §5's floor as "three clean passes per gate". + +## Recommendation + +Stop and simplify, rather than stop and split. **The entire blocker cluster except 1, 13 and +24 is debt machinery** — the state file, the cap, reconciliation, rollback, concurrency, +recursion. That machinery exists to make the waiver *compensated*, yet §13 already concedes +nothing forces anyone to service it. It is elaborate, recursive, unenforceable state whose +delivered value is a row someone may ignore. + +Removing it leaves a coherent, small feature: preconditions, a human pause, and a marker in +the commit body carried to `main`. No state file, no cap, no reconciliation, no rollback +problem, no recursion, and no invented profile grammar — the `+check` gap disappears with +blocker 14's harness scenario, so the override is withdrawn rather than repaired. diff --git a/.context/codex-reviews/gate-a-spec-pass-3-dispositions.stopped-3tier.md b/.context/codex-reviews/gate-a-spec-pass-3-dispositions.stopped-3tier.md new file mode 100644 index 0000000..1b97e70 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-3-dispositions.stopped-3tier.md @@ -0,0 +1,78 @@ +# Gate A — spec — pass 3 dispositions — CYCLE STOPPED + +40 findings (6 BLOCKER, 31 MAJOR, 3 MINOR). Enum held a third time. + +## Trajectory + +| Pass | Blocker | Major | Spec lines | +|---|---|---|---| +| 1 | 4 | 28 | 281 | +| 2 | 4 | 33 | 448 | +| 3 | 6 | 31 | 522 | + +116 findings across three passes, **none dismissed**. Blockers rising while the artifact +grows. This is accretion, not convergence, and §5's stop condition applies. + +## The four findings that stop the cycle + +**Pass-3 B4 — rider (a) is incompatible with "no hook change", and it breaks TIER 1.** +Splitting one Gate-B pass into sequential `spec` then `quality` calls means the unchanged +hook increments its pass counter on **each** call. Three logical passes become six counted +ones, and the hook can report the floor satisfied before a pair's second branch exists. +This is a new mechanical false ✓ in **normal cross-model operation** — nothing to do with +degraded mode. Rider (a) was bundled here as a §5 prose edit; it is not one. It is a hook +change, or it is not shippable. + +**Pass-3 B6 — debt repayment contaminates live gate state.** A later tier-1 repayment +review calls the same `exec`/`review` tools, so it increments the *current* cycle's +counters, and a Gate-B repayment stores a fingerprint of the *current* workspace rather +than the historical range it reviewed. The compensating control corrupts the control it +compensates for. + +**Pass-3 B2 — the tier-2 trust boundary is not achievable in-repo.** Restates pass-2 B3 and +closes the door on the fix Daniel and I chose: a custom agent definition does **not** stop +the repository's `CLAUDE.md` hierarchy loading as instructions, and Claude Code offers no +per-agent switch to omit it. So the prompt product under review can address its own +reviewer before that reviewer treats anything as data. An agent definition improved the tool +allowlist; it did not create the boundary. Fixing it needs a sanitized external checkout or +an independently controlled harness — neither of which is a prompt-only change. + +**Pass-3 B3 — the `.off` posture is the invariant-2 violation, not a residual.** A +workspace-global sentinel authorized for one cycle silences reminders for unrelated commits +in that workspace. That is the missed-commit direction invariant 2 names as dangerous. +Naming it in §16 does not make it compliant. + +## Two more accepted, both mine + +**B1** — tier 3 is a no-pass closure, so "the gate itself is not optional" is not KEPT as my +§11 inventory claims; it is **overturned**. Tier 3 is a gate waiver and the inventory +mislabels it. + +**M20 / M33** — my pass-2 revision **dropped the packaging and backlog section entirely**. +Occurrence 3 (story AC 8), the 0.8.2 → 0.9.0 bump, the CHANGELOG entry and the parked rows +all vanished while I was fixing other findings. A regression introduced by the fix round, +caught by the review — which is the loop working, and also the sign that the artifact is +past the size one revision pass can hold. + +## The remaining 25 MAJOR + +All accepted, none dismissed. They cluster: schema completeness (24, 25, 35, 38), `.off` +state handling and races (21, 22, 29, 30), slot naming and atomicity (23, 39), debt +lifecycle (7, 9, 27, 28, 31), provenance and authorship (13, 40), downstream contradiction +(17), prompt-standards conformance (18, 19), Gate-B range validation (15), agent identity +(16), inventory completeness (10, 11, 12), squash ownership (26), AC 9 template parity (34), +validation coverage (32), and Bash-as-capability (14). + +Not itemized further, because they are downstream of a scope decision that has to come +first. + +## Recommendation + +Stop the cycle and re-scope. The evidence is that this is an epic wearing a story's header: +at minimum the ladder, rider (a), and the tier-2 reviewer surface are three separable +specs, and two of the three need hook or harness work that the settled "prompt-only" +decision excludes. The story's §6 split rule anticipated one split; pass 3 says there are +more. + +No further pass should run until that decision is made. Passes 1-3 stand as a record; none +of them was clean, so the floor is not met and nothing here is approved. diff --git a/.context/codex-reviews/gate-a-spec-pass-3-dispositions.stopped-tier3-core.md b/.context/codex-reviews/gate-a-spec-pass-3-dispositions.stopped-tier3-core.md new file mode 100644 index 0000000..8513499 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-3-dispositions.stopped-tier3-core.md @@ -0,0 +1,102 @@ +# Gate A — spec — pass 3 dispositions (stripped design) — **CYCLE STOPPED** + +34 findings (15 BLOCKER, 18 MAJOR, 1 MINOR). **None dismissed. None applied** — the cycle is +stopped under §5's stuck condition and surfaced to Daniel. + +## The stuck test fires + +| Cycle | Blockers by pass | Outcome | +|---|---|---| +| Three-tier | 4 → 4 → 6 | stopped | +| Two-tier with debt machinery | 7 → 8 → 12 | stopped | +| **Stripped** | **4 → 14 → 15** | **stopped** | + +Artifact size across the same three passes: 489 → 750 → 1016 lines. Every round of fixes has +produced a larger crop of blockers **inside the fixes**. Pass 2's dispositions argued this was +explained by new surface and that none of its blockers said the mechanism could not exist. +Pass 3 falsifies the second half of that argument. + +## Three findings say the mechanism cannot exist — the same shape that killed two cycles + +Design §2 records what ended the debt machinery: every compensating control landed as +**unenforceable** (a prose assertion by the party being waived) or **recursive** (state that +itself needs the unavailable gate). Pass 3 lands that pair on tier 3's **core**, not its +bookkeeping. + +**7 — both surviving outage sources are still locally manufacturable.** Pass 2 removed +`auth-failed` because withholding a credential fakes an outage. The identical argument +defeats what is left: exhaust the account's quota on purpose, point at an +already-exhausted account, or block DNS to produce two transport failures. A transport +failure does **not** establish that the vendor answered, which §3.2 line 191 claims it does. +A human author may self-approve (§3.3), so inducing either source yields a **sanctioned +zero-review closure on demand**. Codex's own suggested fix ends with "explicitly acknowledge +any abuse path that cannot be closed" — the remedies it names (independently obtained +provider evidence, a bound reviewer identity) do not exist in a prompt-only system. +**Unenforceable.** + +**2 — the motivating case cannot satisfy its own central precondition.** Tier 3 exists for the +outage where *no pass could be taken*. Pass 2's fix requires the pass ledger to be +*establishable* and fails closed otherwise. But §3.1(2) itself states that session memory, +`.context/` slots and the hook counter are all non-authoritative, and at Gate A there is no +prior commit body — so in exactly the motivating scenario `Passes completed: none` can never +be established, and the waiver can never be granted. Loosening it to "absence observed" +reopens the lost-pass gate-off path the rule exists to close. **Recursive.** + +**3 — the governing-story union has the same shape.** Pass 2 bound the profile to every story +path cited anywhere in the cycle, to stop a citation being dropped to shed a `high` profile. +Prompts are not durably recorded, so the union cannot be reconstructed; fail-closing makes +every resumed waiver unusable. **Recursive.** + +## The rest are ordinary, and several are in machinery pass 2's fixes introduced + +**The git algorithm is wrong in five places** — 4 (Gate B's parent is the WIP, so §3.5 writes a +*follow-up* commit, not the amend §5 Mechanics requires, and the diff digest then covers only +post-WIP changes: a Gate-B gate-off path), 5 (no atomic ref compare-and-swap; the `HEAD == +parent` check and the commit are separate operations), 20 (squash copies a marker whose +parent, tree and message describe a commit that is then unreachable — the three-way read-back +cannot be applied to the carrier), 21 (the invalidation record reproduces `Reviewer-tier: 3` +verbatim, so §6's own discovery grep counts the invalidation as a **new apparent closure**), +23 (no rollback state defined between a failed read-back and the retry). + +**27 is the one that reframes the whole thing.** The design's highest-risk logic is now +git-level — dedicated index, parent selection, ref replacement, amend, collapse, read-back, +squash carry — and §10 deliberately exercises **no git operation at all**. A model agreeing +with prose cannot surface a data-loss, race or wrong-parent defect, so +`battery+check+verification` is unsatisfied *for the risk path this design actually +introduces*. That risk path did not exist two passes ago: it was created by the fixes. + +**Also real:** 6 (a Gate-A tree may carry unrelated specs, plans and stories that get neither +their own Gate-A decision nor Gate B — a second closure route inside the one being fixed), 10 +(`attempt-failed` can never complete the ordered close under its own recovery budget), 11 +(`authorizes` and `reason` are unbound after the answer), 12 (equality strips whitespace while +digests cover exact bytes — the mechanism proves less than the text claims), 13 (no cardinality +or ordering rules, so a crafted message can present different values to different parsers), +17 (continuation binds the artifact blob but not the governing rules, so an authorization +outlives the conditions it was granted under), 31 (the **fifth** inventory is still +incomplete), plus 8, 9, 14, 15, 16, 18, 19, 22, 24, 25, 26, 28, 29, 30, 32, 33, 34. + +**1** is a clarity defect worth noting on its own: the design's History paragraph and the +story's profile comment both cite pass numbers from the *earlier stopped cycles*, which reads +as though results from the pass now running were already known. + +## Why this is a stop and not a pass 4 + +§5's stuck condition is "clean or clearly stuck". Three signals together: + +1. **Blockers rising across every pass of every cycle of this story**, now for the third time. +2. **The unenforceable/recursive pair has reached the core.** In the two prior cycles it lived + in tier 2 and in the debt machinery — both were successfully cut out, and the cycle + restarted smaller. Findings 2, 3 and 7 are not attached to a removable sub-feature; they + are attached to *"a human may authorize a zero-pass closure"* itself. +3. **The fixes are generating the defects.** Two rounds of hardening turned a prose amendment + to §5 into a git algorithm that now needs its own throwaway-repository test harness (27) — + and §13 has said throughout that none of it constrains a non-compliant agent. + +Nothing here is dismissed and no fix is refused. The question is whether to keep hardening a +mechanism whose own §2 argument now applies to its centre — and that is Daniel's call. + +## Status + +**STOPPED and surfaced.** Floor met by count (3 passes), final pass not clean. Pass-3 fixes +are recorded and unapplied. The spec and story are **uncommitted in the working tree** at the +post-pass-2 state (`df850ab` + pass-1 and pass-2 fixes). diff --git a/.context/codex-reviews/gate-a-spec-pass-3.md b/.context/codex-reviews/gate-a-spec-pass-3.md new file mode 100644 index 0000000..7fbe6b6 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-3.md @@ -0,0 +1,24 @@ +BLOCKER | high | §4 lines 326-353 and §5.1 lines 372-377 | The path table omits `.github/workflows/ci.yml` even though the design says the new checker joins the battery and that CI mirrors that battery; the current workflow invokes and lints only the existing checker pairs | An implementation following the claimed complete site list can leave the parity check out of CI, turning the profile's discriminating check into a local convention and repeating the bookkeeping contradiction this pass is meant to catch | Add `.github/workflows/ci.yml` to §4 and require it to lint the new script and suite and run both in the invariant-check step, with the workflow's checker-count comments updated +MAJOR | high | §4 lines 332-353 | The `AGENTS.md` edit is scoped to the tree and command rows, but its Boundaries paragraph currently says there are exactly two repo-local CI checkers and their tests | Adding a third checker while leaving that architecture statement intact makes the single source of truth false in the same change | Expand the `AGENTS.md` site row to require updating the Boundaries inventory and every exact checker-count claim, not only the tree, battery, and lint rows +BLOCKER | high | §5.1 lines 372-377 | The extraction contract names no exact boundary markers and gives no required cardinality, ordering, or non-empty-span checks | Two missing, duplicated, reversed, or prematurely matched marker pairs can yield equal empty or truncated extracts and a green parity result, so the proposed mechanism can assert parity without comparing the intended prose | Specify the literal start and end sentinels for every compared region and require exactly one ordered pair per surface, a non-empty extraction, and fail-closed handling of every read or extraction error before diffing +BLOCKER | high | §5.1 lines 372-380 | A whole-§5 comparison that normalizes only leading template indentation cannot pass on the current sources: the two §5 bodies already differ materially in hook-tool mapping, incident prose, source citations, prose exemptions, and other wording | Either implementation must make broad unrelated prompt changes absent from §4, or it must compare only selected new regions while the design continues to claim that the two copies and §5 match | Define the exact shared subregions and limit the claim to their normalized equality, or deliberately synchronize the entire sections and add every resulting change to scope and review +MAJOR | high | §5.1 lines 363-368 | The statement that `scripts/check-invariants.sh` has no parity or mirror comparison anywhere is false: check 4b parses the checklist in `docs/prompt-standards.md` and the workflow-init template and fails when their item counts disagree | The correction still overstates what was verified and violates the enforcement-claim calibration rule in the section written to repair that class | Say that no textual parity check for `CLAUDE.md` §5 and its inline copy exists; separately acknowledge check 4b's narrow checklist-cardinality comparison +MAJOR | high | §5.1 lines 375-380 | Stripping unspecified leading indentation is a lossy normalization, yet the design says the resulting diff proves the copies match | Markdown indentation can change list nesting or turn prose into a code block, so two semantically different prompts can compare equal if arbitrary leading whitespace is removed | Define a fixed, mechanically verified template prefix that alone is removed on every extracted line, preserve all remaining indentation byte-for-byte, and claim only equality after that named normalization +BLOCKER | high | §5.3 lines 404-424 | The checker is defined as a source extractor, parity diff, and closed-set prose assertion, but its suite is required to accept and reject runtime `Enum drift:` records, ordering, escaping, empty severity fields, and two-branch combinations; no production parser, input interface, or invocation mode for those artifacts is specified | The listed fixtures cannot exercise the described checker, and a test-only parser would prove its own fixture grammar rather than any mechanism used by the shipped prompt or quality gate | Either specify a real artifact-validator mode with exact inputs, parsing algorithm, exit contract, and production caller, or remove parser/lifecycle fixtures and test the actual source assertions with prompt-source mutations +MAJOR | high | §5.2 lines 382-402 | The stated base counterfactual is not reproducible as written: `scripts/check-prompt-parity.sh` does not exist at `df850ab`, and a new combined checker run against base-era inputs can fail on the pre-existing §5 differences rather than on the absent closed-set rule | An exit 1 from the combined script would not establish that the discriminating assertion is load-bearing, so the profile evidence can borrow strength from an unrelated parity failure | Specify running the new checker against an isolated tree populated with the two `df850ab` prompt inputs, expose or report the closed-set assertion separately, and demonstrate that this assertion fails for the named reason while the harness itself completes successfully +MAJOR | high | reviewer-availability story lines 38-43 versus design §5.2 lines 394-402 | The story's profile note still says the salvage's check makes an out-of-enum file without a drift record ACCEPTED before and INCOMPLETE after, while the design now explicitly says no check observes after-behaviour and only checks that prose contains the rule | The story header is the authoritative profile record, so it promises behavioral counterfactual evidence that the design disclaims and can steer Gate B toward accepting the wrong check | Amend the profile note to describe the actual source-level closed-set assertion and name the PR #23 behavior only as historical verification, matching §5.2's split +BLOCKER | high | Rider (b) lines 295-315 versus CLAUDE.md §5 recovery lines 152-165 | A monotonic never-reused invocation ordinal is incompatible with single-branch recovery: after a full call where one branch file is valid and the other fails, retrying the failed branch must either reuse the original slot ordinal or create a new ordinal that has no same-number partner for the retained branch | The next pass cannot determine which two branch files form one valid full pass, so recovery can overwrite an artifact or make the both-files rule impossible to satisfy | Separate logical pass identity from attempt or invocation identity and specify filenames and pairing for partial retries, including how the reply, Todo number, hook count, credited count, and both branch dispositions map to them +MAJOR | high | Rider (b) lines 269-300 and CLAUDE.md §5 lines 112-137, 152-165 | The existing pre-call rule deletes findings targets but not optional dispositions; the rider gives no freshness binding between a findings file and its drift record | A retry or later reuse of the same slot can leave an old well-formed drift line in place and have it qualify newly written findings before anyone records the new drift | Require deleting or invalidating that slot's dispositions before every write attempt and bind a drift record to the exact findings artifact, then cover stale-record rejection in the suite +MAJOR | high | Rider (b) lines 298-309 | The audit says when to run but not what exact comparison to perform, and it requires only the drift record to be retained even though validating tokens and referenced line numbers requires the corresponding findings file | Once a findings file is deleted, replaced, or unreadable, a surviving companion cannot show that its token set, line list, or branch still describes the credited pass; an existence-only audit can preserve false credit | Require retaining and auditing both artifacts, define the exact parse-and-compare operation and failure states, and state whether any byte change to either artifact discounts the pass +MAJOR | high | Rider (b) lines 298-309 versus §2.3 lines 212 and 228 | The unconditional rule that loss of a drift record reopens the loop because the floor is unmet contradicts the kept early exit for a later zero-finding pass | If pass 1 depended on normalization and pass 2 returned zero findings, discounting pass 1 leaves a valid zero-finding early exit even though the rider says the cycle cannot close | Define the discount transition separately for cycles that already have a zero-finding pass, and update the old-condition accounting to say whether the early-exit rule is kept or deliberately narrowed +MAJOR | high | §2.3 lines 201-231 | The claimed clause-by-clause accounting omits rules directly changed by the rider: pre-call target deletion, sequential slot reuse, one-line pass replies, full Gate-B branch pairing, single-branch recovery, and WIP/amend/squash mechanics; the final `Every other...` row is the same catch-all this repository's decision-procedure rule rejects | Those omitted conditions contain the stale-companion and ordinal collisions, so the completeness claim hides exactly the dropped behavior the accounting is supposed to expose | Add individual rows for every affected file-protocol and mechanics clause, with each old condition marked kept, moved, narrowed, or dropped, and remove the catch-all as evidence of completeness +MAJOR | high | §2.4 lines 242-245 | The identity triple can collide for two genuine decisions by the same handle on the same date about the same `Not done` text | The prescribed deduplication then silently deletes one decision and treats a different reason as a copying error, so the durable record is not lossless | Give every decision a collision-resistant stable identifier or include gate, cycle or head identity and a timestamp in the canonical identity; use that identifier for deduplication +MAJOR | high | §2.4 line 245 | A correction must say `supersedes · `, but record identity also includes `Not done`, so the supersession reference does not identify its target when that handle has several records on the date | A history reader cannot deterministically know which assertion is superseded and may apply the correction to the wrong decision | Reference the full canonical record identifier in `supersedes`, preferably the stable identifier added to the record form +MAJOR | medium | §2.4 lines 243-244 | When two carried copies have the same identity triple but differ elsewhere, the design calls the difference a copying error but still says to collapse to one without defining which bytes win or whether merge must stop | Two people following the procedure can produce different squash bodies, and one can silently discard the corrected or more complete reason | Make divergent same-identity copies a hard pre-merge conflict that must be resolved by a superseding record, never an arbitrary deduplication +MAJOR | high | §1.3 lines 50-56 | The structural rationale says the gate is the only mechanism available to protect its waiver record in a prompt-only product, without retaining the repository-authored qualifier in the heading | That categorical sentence outruns §1.5's explicit admission that externally protected state was untried and can protect history | Say the gate is the only repository-controlled protection mechanism considered here, and keep external authority expressly outside the conclusion +MAJOR | high | §1.5 lines 105-113 | The sentence that a protected-branch approval `establishes neither` is ambiguous and appears false for attribution: a platform approval can establish an authenticated approver and an enforced role policy, while it still says nothing about reviewer availability | The candidate analysis blurs property 4 with property 1 and overstates what the mechanism fails to establish, violating the exact-comparison rule | State separately what authenticated branch approval can establish under a named branch policy and that it cannot establish property 1 without an independent availability attestation +MAJOR | high | §6 lines 431-436 | The backlog calls external authority `the only direction that could work`, contradicting §1.5's statements that this is not an impossibility proof, the design space is not exhausted, and external authority is only a candidate | The closure record turns an untried candidate into an exclusivity claim unsupported by the nine passes | Replace `the only direction that could work` with `one untried direction worth reconsidering` and carry forward all named identity, role-policy, and availability-attestation requirements +MAJOR | high | reviewer-availability story lines 228-234 versus design §4 lines 326-353 | The amended story says `AGENTS.md` stays true and unchanged and that design §4 lists it among files verified unchanged, but §4 actually requires editing `AGENTS.md` for the new checker | This is another bookkeeping contradiction in a settled-decision artifact and can make an implementer omit required architecture and command updates | Say only the quoted cross-model-gate sentence remains true and unchanged while the `AGENTS.md` file itself changes for checker inventory and commands +MINOR | medium | §2.4 lines 238-248 | The durable-record procedure covers squash and ordinary merge but does not state the disposition for rebase merge, fast-forward, or cherry-pick, nor a concrete post-merge correction form | The story criterion is visibility from `main` history alone, so unsupported integration paths leave the carry obligation ambiguous | Enumerate supported merge strategies and their carry rule, and define the exact corrective commit-body procedure when a record is found missing after merge +MINOR | high | Rider (b) lines 259-293 | The parser contract does not define trimming or byte identity for the severity field and distinct drift tokens, despite the finding format placing spaces around separators and allowing escaped pipes | Implementations can disagree on whether ` IMPORTANT `, `IMPORTANT`, or differently escaped spellings are one token, which changes record multiplicity and line mappings | Define field splitting on unescaped separators, whitespace handling, unescaping order before recognition, and whether distinctness is computed on raw or decoded token bytes +END OF FINDINGS (23 total) diff --git a/.context/codex-reviews/gate-a-spec-pass-3.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-3.pre-2026-08-14.md new file mode 100644 index 0000000..16db422 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-3.pre-2026-08-14.md @@ -0,0 +1,9 @@ +MAJOR | high | §1 and §4 change surface | the spec weakens `Never edit a row` to `Never edit a landed row`, but the governing story still marks the absolute rule as kept, says any solution must preserve it, and requires the append-only rule to remain after the change | the design and its acceptance source authorize opposite outcomes for row D, while the story is absent from the declared change surface | amend the story's inherited condition, desired outcome, and acceptance criterion to the landed-row boundary with explicit old-condition accounting, and add the story to §4 +MAJOR | high | §2.0–§2.1 and §5 old-conditions audit | an open-cycle `pending` row is governed both by the untouched rule to resolve it by appending a same-fingerprint row and by §2.1's rule that a row appended by the open cycle is amended in place | the two kept decision procedures produce contradictory actions in a reachable overlap state, so the disposition `Append rather than edit, to resolve a pending row — kept` is not actually preserved unambiguously | state precedence explicitly, such as making pending-row resolution always append and narrowing open-cycle amendment to corrections that are not pending resolution, then reflect that in the disposition table +MAJOR | high | §2.2 convention prose | the required wrong-when-written case is covered in §2.1 but the operational instruction says to name `what stopped being true`, which presupposes the claim was once true | an author correcting a claim that was false from its first publication has no accurate value for the required field and may omit or misdescribe that settled scope case | say to name the claim that is false, distinguishing `stopped holding later` from `was never true` where relevant, and align the format example +MAJOR | high | §2.2 entry-correction rule | an entry is said to be located by its own date plus the row it supersedes, while the same paragraph explicitly permits two independently scoped entries for one row and does not prevent both being written on the same date | those valid entries have the same locator, so a later correcting entry cannot retire exactly one of them as the precedence rule requires | add a distinguishing-fragment fallback and unresolved-locator stop for entry locators too, or define another prose locator that uniquely names the entry being retired +MAJOR | high | SWEEP §4 parity sentinels and §7 Check 2 | the terminal sentinel `and nothing checks the difference.` is new convention text and occurs in neither pre-change target, yet Check 2 says sentinel-delimited extraction on the pre-change files reports equal | a correct extractor must fail because its end sentinel is absent, while an extractor that treats two missing sentinels as equality can green-light missing convention text on both surfaces | make missing or non-unique sentinels a failure and use that absence as the pre-change falsifying observation; retain the one-surface edit as the mismatch case +MAJOR | high | SWEEP §7 Check 1 mechanical half | the entry grep matches only supersession date plus row date, while the table grep independently matches row date plus fingerprint; the first grep never verifies that the entry carries `truncated-tool-output-read-as-complete` | both greps pass if the entry has a missing or wrong fingerprint, so they do not establish the stated claim that the entry's locator resolves to exactly one row | include the expected fingerprint in the anchored entry grep, then keep the table grep to prove that exact date-plus-fingerprint pair selects one row +MAJOR | medium | §2.1 base check and authorship boundary | the spec admits that absence from the cycle base does not establish ownership but gives no action when the agent cannot establish which cycle appended the row, for example after interruption, handoff, compaction, or provenance-obscuring integration into a WIP snapshot | the core landed-versus-amend decision can remain undecidable exactly where editing the wrong row breaches the append-only contract | add an uncertainty terminal state that conservatively treats the row as landed unless open-cycle authorship can be established, and name the evidence a reader may use to establish it +NIT | high | SWEEP §3.1 citation | the cited opening fragment is not byte-for-byte present in `CLAUDE.md`: the source begins `**What this does not do.** The hook counts on `PostToolUse`` with Markdown emphasis bytes | the prose locator is semantically unique but fails the pass's explicit quoted-passage sweep and cannot be found as the claimed literal opening | quote the source Markdown exactly, including the emphasis markers, or explicitly define this locator as rendered-text rather than byte-text +END OF FINDINGS (8 total) diff --git a/.context/codex-reviews/gate-a-spec-pass-3.stopped-2tier-debt.md b/.context/codex-reviews/gate-a-spec-pass-3.stopped-2tier-debt.md new file mode 100644 index 0000000..89ce23b --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-3.stopped-2tier-debt.md @@ -0,0 +1,33 @@ +BLOCKER | high | §3.1(6), lines 69–73 | [risk: compatibility] Tier 3 is said to cover Gate A, but it requires every profile-derived battery, counterfactual check, named verification, and evidence entry before a Gate-A spec or plan closes; those are §5 obligations owed before Gate B and commonly cannot exist before implementation | The fallback cannot close the early gate that blocks planning or implementation, so it does not meet the both-gates story outcome or the motivating outage | Define gate-specific waiver prerequisites: preserve applicable Gate-A controls and defer implementation-derived evidence to Gate B without representing it as already complete +BLOCKER | high | §§4–6, especially §5.4 lines 230–235 | Adding the mandatory root `review-debt.tsv` row to a Gate-A docs commit makes that commit non-docs-only under the actual hook, so it creates a Gate-B STOP; §4 then lets the one Gate-A waiver override that STOP without a Gate-B marker or debt, while a second Gate-B waiver would be barred by the newly open debt | A supposedly single-cycle Gate-A exception either stalls anyway or silently waives an additional product-change Gate B, producing exactly the unrecorded false checkmark invariant 2 forbids | Redesign the atomic close and debt storage together, likely including hook-aware handling or an explicitly combined two-gate transaction; the settled “no hook change” constraint is not compatible with the current record location +MAJOR | high | §5.4 lines 230–235 and hook `is_docs_only` lines 788–800, commit branch lines 910–958 | The classifier reasoning is only half true: a root non-`.md` path does fall through to the full Gate-B reminder, but “erasing a debt then costs a review” overstates an advisory hook that always exits 0 and that §4 explicitly permits a waiver to overrule | This violates the gate-proof and prompt-standard enforcement-claim rules and presents a reminder as a compensating control stronger than it is | Say precisely that the staged path makes `is_docs_only` false and causes the hook to emit the full Gate-B reminder; separately state that nothing forces the review and close the resulting abuse path +BLOCKER | high | §§3.1(7), 4, 5.4, and 7 | [risk + security: abuse and abuse paths] The procedure does not forbid deleting or rewriting prior debt rows in the same index that takes a new tier-3 closure; the current-state scan then sees no open row or cap history, and the same waiver precedence excuses the Gate-B reminder caused by that erasure | A compliant-looking single commit can erase the guard and consume another waiver during the outage, so moving the state to a product-classified path does not secure it | Evaluate eligibility against the merge base and current main as well as the prospective tree, prohibit debt-record removal or mutation in a waiver commit, and fail closed on missing or malformed history +BLOCKER | high | §7 lines 305–309 and marker line 195 | The consecutive-waiver counter has no authoritative home or derivation algorithm: the marker merely asserts `k`, the TSV schema has no counter field, and the text does not define whether to scan current rows, first-parent history, all markers, cycles, or story rows | Different compliant agents can compute different values, multi-story cycles can count once or many times, and a false low value bypasses the principal control added in this revision | Define one durable source of truth, an exact repository-wide counting/reset algorithm, multi-story semantics, parse-failure behavior, and the actor/checkpoint that computes and verifies `k` +BLOCKER | high | §7 lines 305–347 | After two waivers are each closed as `accepted-no-review`, the cap is two and no open debt remains to repay; the only reset is a completed repayment, but the state machine provides no way to repay or reopen a row already closed by acceptance | A permitted sequence can permanently disable tier 3, so the cap is not actually “resettable only by repayment” under the defined states | Specify a lossless transition that permits later repayment of an accepted row and preserves the acceptance record, or make acceptance non-closing and redesign the cap around outstanding repayable debt +BLOCKER | high | §§5.3 and 7 closing condition 2 | §5.3 unconditionally requires the human-decision handle, date, and cycle id to match the marker, but `accepted-no-review` requires a different handle from the marker’s waiver approver | The only no-review debt-closing condition outside solo use is impossible to represent conformingly | Scope marker equality to the waiver-authorization decision, and define separate debt-acceptance fields that match the row and deliberately differ from `Authorized-by` +BLOCKER | high | §§5.2, 5.4, 6, and 7 Gate-B repayment | `Waived-range` points at the pre-amend WIP SHA, which becomes unreachable when the WIP is amended and is also discarded by squash; unlike the Gate-A blob, no durable object preserving the Gate-B range is specified | The later `mcp__codex__review` cannot repay a debt whose base or head has been garbage-collected, and §5.4’s durability rationale covers only the artifact-blob case | Reconcile the row to durable landed tree or patch identities before branch SHAs disappear, or preserve an explicit ref/bundle; specify and verify the mapping for amend, squash, and ordinary merge +BLOCKER | high | §5.4 lines 237–262 | The authoritative debt state lacks a complete state machine: `row-key` and `landed` have no grammar, creation timing is unstated, `open-unreconciled` to `open-reconciled` is never defined, malformed and duplicate rows have no outcome, and no transition says which fields are immutable | Open-debt eligibility, repayment targeting, idempotency, and fail-closed behavior cannot be implemented consistently from this design | Specify every state transition and invariant, immutable versus mutable columns, key generation, reconciliation for each merge strategy, duplicate/conflict handling, and that unreadable or invalid state refuses a waiver +BLOCKER | high | §§5.4 and 7 repayment/closure | Every debt-state transition edits the non-`.md` file and therefore creates a Gate-B obligation, but the protocol never sequences that obligation: `accepted-no-review` cannot be committed during the outage without ignoring Gate B, and marking a Gate-B debt `repaid` after the final pass changes the fingerprint that pass covered | Debt cannot be durably closed by the advertised procedures without either recursive review debt, a stale fingerprint, or another undocumented waiver | Design an atomic closure protocol whose final reviewed target already contains a non-premature transition, or add a purpose-built mechanism that validates and applies debt transitions without recursively creating an uncovered gate +BLOCKER | high | §13 rollback, lines 607–610 | [risk: rollback and data loss] The claim that rollback leaves rows in place is false for an ordinary `git revert` of the waiver commit, which will remove or reverse the row together with the feature unless the operator manually excludes it | A rollback can erase open debt and cap history, making the next waiver eligible; the stated self-contained recovery property disappears exactly on the path §13 discusses | Define and verify a rollback procedure that preserves or migrates debt rows and cap history, including emergency rollback, and treat missing expected rows as fail-closed +MAJOR | high | §§5.1, 7, and 13 concurrency residual | [risk: idempotency; security: external systems] Two branches can both pass the same pre-merge main check and merge concurrently, exceeding the cap or landing while another branch creates open debt; acknowledging that nothing serializes the check does not make “at most two” or the matrix’s third-waiver refusal true | The foundational compensating controls fail even when both agents follow the procedure, not only under malicious bypass | Require a serialized merge queue or protected server-side check, or weaken the guarantee and choose a concurrency-safe reservation/compare-and-swap design +BLOCKER | high | Story profile lines 4–15 and spec §11 lines 512–518 | The scoped `mode override` is not a valid §5 profile change: the header remains `battery+check+verification`, while the settled grammar permits only a whole effective mode in the header and requires that header to carry the override; there is no per-feature “tier-3 portion” override semantics | Gate readers must still require `+check`, or must classify the profile/log combination as semantically inconsistent; the exception silently invents a fourth profile mechanism | Keep the declared mode and satisfy its check, or amend the profile system first with an explicit compositional override grammar, full resulting header, human confirmation, reader rules, and old-condition accounting +MAJOR | high | §11 lines 506–518 and story profile log lines 7–15 | The claimed counterfactual gap is not real: the named outage scenario itself distinguishes the states—before the change a compliant agent refuses a tier-3 close, after it accepts one only when the authorization, marker, debt, cap, and evidence conditions hold | The design lowers evidence on a false premise and misses the most direct behavioral regression check for the product, which is prompt behavior | Use a scenario or prompt-harness check that exercises an unavailable reviewer and asserts prior refusal versus new conditional closure, including negative cases for every precondition +MAJOR | high | §§3.3–3.4, 7 condition 2, 12, and 13 | [security: assets, trust boundaries, and roles] Authorization and debt acceptance are unauthenticated caller-authored strings; a committer can invent both handles, and the only proposed enforcement is postponed until after the first disputed waiver | The gate is removed on the strength of records that provide no evidence the named humans acted, so separation of duties and approval are decorative against the abuse lens this high-risk story requires | Either add protected-branch approval or a verifiable signed human attestation now, or explicitly narrow the threat model and stop presenting handle separation as a control +MAJOR | high | §3.3 lines 110–125 | “Approver is never the implementer” provides no meaningful independence because `Implementer` is defined as the agent session, while the human author may approve their own change and solo use explicitly expects one human | The role model sounds like separation of duties without separating the change author from the waiver decision-maker | State whether self-approval by the human author is intended; if not, require an independent authorized maintainer, and if yes, describe the role as accountable self-approval rather than independence +MAJOR | high | §3.2 lines 88–106 | [security: external systems] Revalidation does not repeat the source’s defined observation: `attempt-failed` means a call plus its recovery both failed, but revalidation requires only one fresh call; it also does not say how a fresh different failure changes the source or whether the shared recovery budget applies | A transient single failure can authorize a waiver even though the normal recovery path might immediately succeed, and the recorded cause may not match what was observed | Define exact result transitions for each source, account/tool binding, whether and when the one recovery attempt is used, and update the recorded enum to the fresh observed cause +MAJOR | high | §§3.2, 3.4, and 4 closing transition | [risk: observability and idempotency] The ordering and lifetime of authorization, evidence revalidation, outage revalidation, target hashing, and commit are not defined, nor is the response to a failed or delayed closing commit | A valid answer or outage observation can be replayed after availability, evidence, target content, or merge strategy changes, especially because the hook resets on a failed non-WIP commit attempt | Specify a single ordered close protocol, require all checks again after any commit failure or material delay, and define when authorization is consumed +MAJOR | high | §5.2 `Waived-range` and `Waived-artifact` | The recorded target is not required to equal the prospective closing commit: an amend can change the tree after the WIP `headSha`, and the artifact blob, debt row, evidence entry, and merge strategy can change after authorization without a binding comparison | History can disclose a waiver for one target while landing different bytes, violating the durable-identification criterion and the gate-proof invariant | Bind authorization to the exact prospective tree or artifact blob and evidence set, forbid post-binding content changes, and recompute/re-authorize on any mismatch +MAJOR | high | §5.1 lines 170–178 and §5.3 | Collision repair by incrementing the cycle id is not local: it invalidates the matching human decision, marker, row keys, dedup key, and possibly the waived target after history rewriting, yet the design says only to increment `` | A pre-merge repair can leave contradictory records or silently carry authorization to a different identity | Define the full repair transaction and require renewed human authorization and target revalidation after every identity or history change +MAJOR | medium | §5.4 TSV schema | [risk: compatibility and data loss] The TSV has no escaping or canonical encoding even though Git paths may contain tabs or newlines and the compound `waived-target` and snapshot fields have no grammar | A valid repository path can split one debt into extra columns or rows; without a fail-closed parser rule this can hide open debt or make repayment target the wrong object | Use a format with defined escaping such as canonical JSON Lines, or define percent/base64 encoding, column validation, and fail-closed parsing +MAJOR | high | §§5.4, 9 `/workflow-init` scaffolding, and 12 packaging | Existing initialized projects will not acquire `review-debt.tsv` merely by updating the version-keyed plugin, and no migration or first-use creation rule is specified | The fallback’s authoritative state and cap are absent in exactly the installed projects this feature is meant to unblock; a missing file can be mistaken for no debt | Add an upgrade/migration path or require tier 3 to fail closed until an idempotent initialization creates and validates the file +MAJOR | high | §5.4 and §9 `/workflow-init` scaffolding | [risk: idempotency and data loss] The design says only that `/workflow-init` scaffolds the file; it does not define missing, identical, divergent, or malformed-file behavior, nor an additive merge that preserves existing debt rows | Implementing the obvious overwrite would violate invariant 9 and could erase the only durable guard | Specify inline canonical template content and idempotent merge/conflict behavior, with tests for existing rows, divergent headers, duplicate keys, and interrupted writes +BLOCKER | high | §8 old-condition inventory | The third inventory still omits operative §5 conditions, including the exact risk/security lens sets and merged abuse lens, “lenses are questions not passes,” Gate-A timing, Gate-B tool/range parameters and full-branch behavior, skip-reason/evidence rules, lower-profile-pass carryover, and the stated non-enforcement residuals | The governing decision procedure is being replaced without accounting for all old conditions again, directly violating the AGENTS.md rule whose prior two failures made this table necessary | Rebuild the inventory sentence-by-sentence from CLAUDE.md §5, include every operative clause, and disposition each by effect rather than theme +MAJOR | high | §8 row 48 versus §3.1(2) | Row 48 says “Re-review after every fix” is kept untouched, but tier 3 expressly permits a pass with an adverse finding to be remediated and then closes with no qualifying re-review of that fix | The inventory conceals a substantive weakening: the landed remediation may never be seen by an independent reviewer | Mark the condition narrowed or overturned for tier 3 and add a compensating requirement or explicit debt coverage for the remediated target +MAJOR | high | §8 rows 3 and 36 | “`.off`; the gates still apply” and “spent and still incomplete means STOP and surface” are marked kept untouched, yet tier 3 makes the gate inapplicable after the pause and permits continuation after the spent recovery; their initial diagnostic steps may remain, but their terminal effect changes | These are further noun-based dispositions of the kind the corrected rows warn against, leaving the old-condition accounting inaccurate | Mark both conditions narrowed for the authorized tier-3 path and state precisely which initial actions remain and which terminal prohibition is overturned +MAJOR | high | Story AC 3 lines 102–113 versus spec §9 lines 461–466 and `docs/sparring-briefing.md` lines 41–44 | The amended story still says `docs/sparring-briefing.md` moved to tier 2 in its entirety and its wording stays untouched, while the spec correctly observes that its human-substitution and never-exempt clauses directly forbid tier 3 and must change | The spec does not satisfy the current acceptance criterion, and the story’s old-condition accounting remains knowingly false | Amend AC 3 again: move only the same-family premise and explicitly retain and overturn the two tier-3-relevant clauses in this story +MINOR | high | Story §1 lines 19–21 | The story says §5 requires “three clean passes per gate,” but §5 requires a minimum of three passes with only the final pass clean and permits a zero-finding early exit | This misstates the baseline decision procedure the design is changing and can distort waiver and repayment requirements | Replace it with the exact floor, final-clean, and early-exit rule +MAJOR | high | §9 sites versus AGENTS.md architecture tree lines 27–59 | The new root-level authoritative `review-debt.tsv` is absent from the planned AGENTS.md architecture-tree update; §9 mentions only changing “What this project is” | Adding a meaningful product-state file without updating the single-source-of-truth tree creates the docs drift AGENTS.md explicitly warns about | Add the file and its role to the architecture tree and include that edit in the site/verification list +MAJOR | high | §11 matrix row “Debt row edited” | The row says an edit “fires full Gate B,” but the verified hook only chooses the full Gate-B reminder branch; it cannot launch a review, validate findings, or block the commit | Even after the matrix preamble’s qualification, this wording repeats the §5.4 enforcement overclaim and makes the validation matrix claim a mechanism that does not exist | Change the observable outcome to the exact emitted reminder/state behavior and add a separate row showing that an ignored reminder still permits the commit +MAJOR | high | §5.2 lines 188–207 | `` is undefined and existing slot names are reused and stored under ignored `.context/`, so `Passes completed` and `Adverse findings` cannot durably identify which invocation or finding the marker references | The disclosure can look complete on `main` while its evidence is ambiguous or gone, weakening observability and incident review | Define a cycle-unique durable pass identifier and what summary must be copied into history without exposing sensitive finding text +MAJOR | high | §5.5 deduplication and §5.4 rows | [risk: idempotency] Deduplication covers only markers; repeated or resumed closes can append duplicate debt rows and decision blocks, and no rule says whether identical rows collapse or conflicting rows fail closed | Retries can create multiple open debts for one waiver or contradictory authorities, changing eligibility and repayment behavior | Extend idempotency rules to decision blocks and debt rows, keyed by cycle and story, with atomic write and conflict semantics +END OF FINDINGS (32 total) diff --git a/.context/codex-reviews/gate-a-spec-pass-3.stopped-3tier.md b/.context/codex-reviews/gate-a-spec-pass-3.stopped-3tier.md new file mode 100644 index 0000000..48c4360 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-3.stopped-3tier.md @@ -0,0 +1,41 @@ +BLOCKER | high | [risk] §3 Tier 3 and §11 "The gate itself is not optional" | tier 3 is explicitly a no-pass closure, yet the inventory says the gate still runs and only the reviewer changes | this silently redefines a mandatory review gate into an authorization ceremony and makes the central "gates still apply" claim semantic rather than real | call tier 3 a gate waiver, mark the old mandatory-gate condition deliberately overturned, and either amend the governing invariant and all product claims or remove zero-pass closure +BLOCKER | high | [security: trust boundaries] §3 Tier 2 "Untrusted input" versus §16 | the design says only the agent definition and dispatch prompt are authoritative while also admitting that every custom subagent loads the repository's CLAUDE.md hierarchy as instructions; Claude Code provides no per-agent switch to omit those files | the prompt product being reviewed can steer its own reviewer and manufacture a clean same-family pass, so the proposed containment is not achievable with this custom-agent mechanism | use a reviewer surface that does not load repository instructions, or run from a sanitized external checkout/harness with an independently controlled system prompt; otherwise classify tier 2 as unavailable +BLOCKER | high | [risk] §6 `.off` scope and §16 residuals | a cycle-specific authorization deliberately uses a workspace-global sentinel that silences reminders for unrelated commits, and the spec leaves that false-negative path open while claiming invariant 2 is merely scoped | an unrelated commit can close with no hook warning because of another cycle's degradation, exactly the missed-commit direction invariant 2 forbids | make silence cycle-scoped in the hook or do not use `.off`; naming the invariant violation as a residual is not sufficient +BLOCKER | high | [risk] §7 plus §13 rider (a) | one logical Gate-B pass becomes two sequential `mcp__codex__review` calls, but the unchanged hook increments its pass and fresh-content counters on each call; after the second pair starts, the hook can report the three-pass floor satisfied before that pair's quality branch exists | this creates a new mechanical false checkmark and directly violates invariant 2 and §5's claim that the hook counts passes | either change the hook to bind and count completed branch pairs, or define and mechanically support separate three-pass branch loops; the settled "no hook change" decision is incompatible with this rider +BLOCKER | high | [risk] §7 content identity across Gate-B passes | the spec binds the two branches within one pass and rechecks only the final pass, but never says that a changed `headSha` invalidates every earlier pass in the current three-pass streak | fixes between passes can yield three accepted tier-2 passes over three different snapshots even though only the last snapshot received one pass, losing the current hook's fresh-fingerprint floor | require all passes that satisfy the floor to share one validated base/head identity and reset the qualifying streak whenever either identity changes +BLOCKER | high | [risk] §8 repayment protocol versus unchanged hook state | a later tier-1 debt review uses the same `exec` or `review` tools as an active gate, so its calls increment Gate-A or Gate-B counters and a historical Gate-B debt review stores a fingerprint of the current workspace rather than the historical range it reviewed | repayment passes can satisfy an unrelated live cycle and can create precisely the content-attribution false checkmark invariants 2 and 3 prohibit | run debt repayment in an isolated checkout with isolated `.context` state, or add explicit hook state separation/reset that cannot credit repayment calls to current work +MAJOR | high | [risk] §8 effective-level-2 debt rule | only effective-level-2 profiled stories receive a re-review obligation, while lower-profile and unprofiled tier-2 cycles—and even tier-3 zero-pass closures—close permanently with no future independent review | the generic shipped ladder creates a permanent no-independent-review path for much of its audience even though review is the product's only correctness control | require debt for every degraded cycle, then vary urgency or evidence by profile rather than whether independent review is ever owed +MAJOR | high | [risk + security: abuse and abuse paths] §3 Tier 3 activation | "tier 2 unavailable" has no predicate, proof, retry budget, failure taxonomy, or recorded evidence analogous to tier 1, and `not configured` appears in the cause enum despite the ladder being mid-flight-only | an agent or operator can reach the broadest fail-open tier by asserting that the local agent surface is unavailable, including after self-created configuration changes | define tier-2 availability checks and failure states, require its one recovery attempt where applicable, distinguish vendor outage from local misconfiguration, and record the result before tier 3 can be authorized +MAJOR | high | [security: roles] §8 closing condition 2 | the same logged human-decision mechanism that authorizes a degraded closure can immediately accept and erase the resulting re-review debt without any separation of role, time, or justification threshold | the debt marker is optional in substance and provides no durable pressure to restore the independent control | require a distinct later decision by a named eligible role after a cooling period or documented risk acceptance, and forbid debt acceptance in the same cycle that created it +MAJOR | high | [risk] §10 and §11 old-condition inventory | current §5 says dispositions are optional advisory notes that never participate in validation, but the design makes them load-bearing for every tier-2 provenance record and for severity normalization without listing or disposing that old condition | this is the exact silent condition loss the decision-procedure Don't requires the inventory to prevent | add the old advisory-only condition to §11, mark the precise scopes deliberately overturned, and specify the new required-file validation and recovery rules +MAJOR | high | [risk] §11 "complete inventory" | the claimed complete inventory omits material §5 conditions including TodoWrite per pass, dismissal reasons, the two separate Gate-A loops, mechanical pre-read checks, full-profile resolution and evidence-gap stops, story/evidence carriage on every Gate-B call, re-review after fixes, spec synchronization, the standing falsification lens, exact prose classification, current-profile final pass, and timeout/background handling | future replacement prose can silently drop any omitted safety condition while still citing this table as complete | derive the inventory paragraph by paragraph from §5 and mark every condition kept, moved, narrowed, or dropped rather than summarizing only selected themes +MAJOR | high | [risk] §11 and §13 rider (a) accounting | the inventory omits the old `reviewType: full` default, two-file acceptance rule, one-slot concurrency limitation, both-files deletion rule, and successful-branch single-resume recovery path even though rider (a) replaces or narrows each | AC 5 can be implemented while losing the old race and recovery protections, repeating the repository's documented decision-rewrite failure mode | add each old branch/file/recovery condition to §11 with its new disposition and make the sequential pair state machine normative +MAJOR | high | [risk] §7 and §10 Gate-A records | §10 requires every tier-2 pass to carry the §7 provenance record, but §7 defines only Gate-B `baseSha`/`headSha` identity and never binds a Gate-A pass to the exact spec or plan text later committed | a changed artifact can inherit a clean tier-2 Gate-A disclosure, and the debt row cannot identify the text that needs re-review | define Gate-A provenance with artifact path, content digest, and committed blob or commit reconciliation, and invalidate the pass when that identity changes before close +MAJOR | high | [security: assets] §3 Tier-2 tools and §16 | granting Bash to a prompt-influenced reviewer in the main checkout grants arbitrary repository writes, process execution, environment reads, and possible credential exfiltration; the instruction to use it only for a diff is not a capability boundary | the reviewer can alter the artifact it certifies or leak developer and repository assets, while a clean findings file hides the side effect | use a read-only isolated worktree/sandbox plus a narrowly scoped diff helper, or abandon the prompt-only constraint; at minimum make this an explicit tier-2 stop for sensitive repositories rather than a residual +MAJOR | high | [security: external input] §3 Gate-B pinned range | the design does not require full-hash syntax validation, commit-object resolution, `headSha == HEAD`, ancestry policy, safe shell quoting, or a `--` separator before the reviewer uses caller-supplied range values in Bash | malformed or instruction-influenced values can review the wrong range, be parsed as options, or become shell injection, and a well-formed findings file will not reveal it | specify exact validation and quoting commands for full object IDs, resolve both as commits, require the intended ancestry and current HEAD, and define binary, submodule, and LFS review behavior +MAJOR | high | [risk: compatibility] §3 shipped agent invocation | the spec names a file but not the exact plugin-scoped agent identifier, how the dispatcher proves that definition loaded, or how it rejects a project/user/managed shadow or a missing post-update registration | the wrong agent can run with different tools, model, context, and prompt while still producing the expected file shape | require explicit invocation of the plugin-scoped identifier, verify the active definition/tool/model surface before credit, and enumerate restart and precedence failure diagnostics +MAJOR | high | [risk: compatibility] §15 installed downstream projects | the out-of-scope section says old projects receive the plugin "without the ladder," but they do receive the new convention-loaded agent, which Claude Code may select from its description while their copied §5 still forbids same-family review | plugin update creates a contradictory partially deployed feature in exactly the installed base the release claim excludes | make the agent description and prompt refuse use unless the current workspace §5 contains a versioned ladder and explicit authorization, or withhold the agent until downstream policy synchronization is solved +MAJOR | high | [risk] §3 target model and skill dependency | the proposed definition's tool list omits `Skill`, no `skills:` preload is specified, and a prose "Target model" line does not configure the effective model; Claude Code model selection can be overridden by environment or invocation | the agent may be unable to run the required brainstorming method or may disclose a model surface it did not actually use | specify the exact frontmatter including model and skill preload/tool access, define how the dispatcher verifies effective model family and skill availability, and stop on mismatch +MAJOR | high | [risk] §14 validation versus invariant 11 | validation never requires the new agent, changed §5, hook messages if any, and inline scaffold template to pass all 12 prompt-standard items; the generic battery mechanically covers only two narrow spellings | a prompt-only product can ship an unusable or contradictory reviewer even with every listed check green | add a named 12-item review for every changed prompt artifact, including output example, causes, stop conditions, enforcement calibration, target-model verification, and mirror parity +MAJOR | high | [risk] §12 implementation inventory versus packaging invariant 12 | the design adds and changes files under `plugins/dev-workflow/` but lists only a description edit, not the mandatory manifest version change or the corresponding CHANGELOG entry; the current manifest is 0.8.2 while §15 assumes 0.9.0 | an unbumped plugin change remains invisible in the version-keyed installed cache, so the ladder may never reach users | add the exact manifest bump and changelog work to scope and validate the committed range with the version-bump checker against the current base +MAJOR | high | [risk: concurrency] §6 sentinel serialization | "degraded cycles are serialized" is stated as fact, but check-then-write entry and restore/delete exit have no lock or atomic ownership check, so two cycles can both enter and one can remove the other's authorization | overlapping cycles can corrupt authorization state or silently re-enable reminders mid-cycle | define an atomic owner token and compare-and-swap or lock protocol, and downgrade the claim to procedural until that mechanism exists +MAJOR | high | [risk: data loss + idempotency] §6 pre-existing `.off` preservation | embedding prior arbitrary content inside new content has no delimiter, byte-preserving encoding, mode handling, or compare-before-restore rule; edits made while degradation is active would be overwritten on exit | restoration can corrupt or destroy a user's prior opt-out state and repeated entry/exit is not demonstrably idempotent | define a separate cycle authorization file or a byte-exact sidecar backup with ownership, atomic writes, content hash comparison, and conflict-stop behavior +MAJOR | high | [risk + security: abuse and abuse paths] §7 slot naming | `` has no generation rule, allowed character set, uniqueness boundary, or path-safety validation even though it becomes a filename and a deduplication key | collisions can overwrite another cycle's evidence and path separators can escape the intended review directory | define a collision-resistant repository-scoped identifier with a strict safe alphabet, canonical validation, and exclusive slot creation +MAJOR | high | [risk: observability] §9 schemas | deduplication requires a cycle identifier and §6 says mixed-tier markers record both the weakest and final-pass tiers, but neither value appears in the tier-2 marker schema; Gate A also cannot distinguish spec from plan in the shown `Gate:` field | the carry chain cannot deduplicate reliably or reconstruct which gate/artifact and tier transition actually occurred | add cycle ID, `A-spec`/`A-plan`/`B`, weakest tier, final-pass tier, and pass identities to the canonical schema and its validation rules +MAJOR | high | [risk] §10 and story AC 7 | the squash hop explicitly carries markers, but never explicitly carries each existing profile evidence entry into the squash body; §13 merely points rider (c) back to that disclosure-only hop | squash can preserve degraded-tier disclosure while deleting the validation evidence AC 7 specifically requires on main | state and test that both every tier marker and every profiled-story evidence entry are copied into the squash body +MAJOR | high | [security: roles + external systems] §10 squash carry | the author is made responsible for merge-time body contents even when a separate maintainer or hosting automation performs the squash, and verification occurs only after the destructive rewrite with no repair path | the only durable disclosure can be lost after all local gates pass, with no defined owner able to prevent or remediate it | assign pre-merge and merge-operator responsibilities, add a blocking body check where supported, and define a post-merge corrective disclosure commit when verification fails +MAJOR | high | [risk: observability] §8 post-land reconciliation | reconciliation is a later unreviewed `todos.md` edit with no owner, deadline, idempotent row key, merge-parent rule, or recovery if it never lands | debt rows can remain unusable or point at the wrong base, making the promised recurring availability check observationally empty | define the exact row key and state machine, ordinary/squash base calculation, responsible actor, timing, retry/conflict handling, and a check that detects unreconciled rows +MAJOR | high | [risk: rollback] §8 repayment after findings | the matched Gate-B artifact is the already-landed historical diff, but remediation necessarily lands in a later commit; repeatedly reviewing only the original range will continue to see the defect and cannot establish that the fix resolves it | the normal "remediate then clean final pass" closing condition is not executable for real post-land findings | define repayment as reviewing the original change plus named remediation range, or define an explicit disposition proving the follow-up addresses each finding and review that cumulative state +MAJOR | high | [risk + security: external systems] §6 availability re-check | "availability is re-checked before every pass" names no probe, timeout, authority, freshness window, or effect on counters/files; an actual gate call is itself a pass while a status page or remembered quota can be stale or spoofed | operators cannot implement the transition consistently and can remain degraded after recovery or consume/credit unintended calls | define a side-effect-free check where possible and an exact real-call fallback with file/counter treatment, evidence source, freshness, timeout, and failure diagnostics +MAJOR | high | [risk: compatibility] §6 pre-existing `.off` | the same pre-existing sentinel may represent init-time `INACTIVE` policy or deliberate opt-out, both of which say the gates do not run, yet the design permits preserving that state while running a mid-flight weaker gate | a workspace can simultaneously assert inactive gates and a completed tier-2 cycle, producing irreconcilable policy and history | classify the prior state from the current §5 marker and sentinel contents; refuse ladder entry until init inactivity is explicitly converted to active policy, and preserve deliberate opt-out as a separate state +MAJOR | high | [risk: rollback] §15 rollback | keeping debt rows while reverting the only §5 procedure that knows when and how to service them strands obligations as inert prose | rollback can permanently remove the recovery control while appearing safe because the records remain | require open debts to be repaid or explicitly accepted before rollback, or preserve a minimal debt-processing protocol independently of the ladder feature +MAJOR | high | [risk] §14 validation evidence | the proposed counterfactual only checks whether an instructed actor writes a marker, and the matrix omits the hook's branch-pair overcount, cross-pass identity reset, debt-call state contamination, prompt-injection boundary, tier-2-to-tier-3 failures, sentinel races, rollback, and squash evidence carriage | the high-risk validation mode can pass while the principal false-checkmark and containment failures remain | add executable or scratch-workspace checks for each state transition and hook interaction, plus adversarial prompt, range, race, and merge-strategy cases +MAJOR | high | [risk] story AC 8 versus whole spec | the required occurrence-3 append for the compound `git add`/`git commit` PreToolUse incident is absent from the design's scope, site inventory, riders, and validation | a stated acceptance criterion can be silently omitted from the implementation plan | add the exact append-only `todos.md` update and a validation check that the prior row was extended rather than edited or duplicated +MAJOR | high | [risk] story AC 9 and §12 | the implementation inventory mentions `/workflow-init` only for §2.13 and never names its large inline §5 mirror as a change surface or defines a parity check | CLAUDE.md can gain the ladder and riders while newly scaffolded projects keep the old protocol, directly failing AC 9 | list the inline template explicitly for every §5 edit and validate extracted semantic parity for the ladder, branch default, severity reader, debt, and squash carry +MAJOR | high | [risk] §13 rider (b) tolerant reader | "any other token" does not define the parser boundary: empty severity, extra unescaped separators, shifted fields, multiple bad tokens, or `NO FINDINGS` mixed with findings could all be treated as normalizable rather than malformed | tolerance can turn structurally corrupt or attacker-shaped output into an accepted MAJOR finding list while violating the exact-line grammar | limit normalization to an otherwise valid six-field line with one non-empty severity field, keep all structural failures INCOMPLETE, and provide a canonical dispositions schema and examples for multiple mappings +MINOR | high | [risk] §2 narrowing rationale | `docs/prompt-standards.md` line 136 says fresh-context verifier subagents outperform self-critique, but does not establish the stronger proposition that a same-family tier-2 review "beats nothing" in this gate or quantify its independence | the design presents a supportive note as evidence for the central weakening and overstates what the cited text proves | restate this as a hypothesis to validate for this workflow or cite direct evidence comparing no review, self-review, and fresh same-family review +MINOR | high | [risk: compatibility] §12 init-time distinction | the claim that tier 2 at init would be a permanent arrangement conflicts with `/workflow-init`'s current explicitly temporary inactive state and re-enable checklist | a false permanence rationale hides the actual policy choice between unconfigured-at-init and lost-mid-flight states | state the real distinction and its tradeoff without claiming permanence, or permit the same temporary ladder under equivalent authorization +MINOR | medium | [security: assets] §9 human-decision schema | requiring `` in public commit history has no identity convention, privacy guidance, or option for an accountable non-personal handle | public repositories can disclose personal information unnecessarily while still failing to authenticate the decision | define an approved accountable identifier format, obtain consent for public recording, and avoid sensitive personal details +MAJOR | high | [risk: concurrency] §7 slot ownership and §3 Bash | stopping Claude tasks does not account for child processes a Bash-capable reviewer can background, and files are written directly to stable slots rather than invocation-unique temporaries atomically renamed | a late child can overwrite a validated pass after the task is reported stopped, defeating every shape and provenance check | forbid and detect background descendants, use invocation-unique temporary files with atomic publication and owner tokens, and validate the published file immediately before close +MAJOR | high | [risk: observability] §9-§10 record authorship | every durable marker and human-decision block is caller-authored prose with no binding to the validated findings files, dispositions, or actual human interaction, yet history is treated as proof that degradation was disclosed and authorized | an errant or compromised agent can fabricate tier, pass count, authorization, and debt status while satisfying the visible schema | bind markers to hashes of pass records and an explicit human-confirmation transcript identifier or signed/manual commit step, and calibrate all claims to "recorded assertion" unless verified +END OF FINDINGS (40 total) diff --git a/.context/codex-reviews/gate-a-spec-pass-3.stopped-tier3-core.md b/.context/codex-reviews/gate-a-spec-pass-3.stopped-tier3-core.md new file mode 100644 index 0000000..0db5e9e --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-3.stopped-tier3-core.md @@ -0,0 +1,35 @@ +MAJOR | high | design lines 9-15; story lines 6-16 | The artifact says six prior passes produced 212 findings and the story says Gate-A pass 3 already produced findings 13 and 14, while this review is pass 3 after two passes totaling 57 findings | The durable design history and profile rationale describe a future or different review cycle, so readers cannot tell which decisions and fixes actually precede this artifact | Replace both histories with the actual two-pass counts and remove all references to results from the pass now being run +BLOCKER | high | §3.1(2), lines 72-88 | The motivating zero-pass case cannot establish `Passes completed: none`: the design expressly says neither session memory, slot files, nor the hook counter authoritatively enumerate calls, and no durable cycle-start record exists before the closing marker | A fresh outage can never satisfy the central precondition, while treating local absence as proof would recreate the lost-pass gate-off path | Add a durable, non-recursive cycle-start and accepted-pass ledger or define another mechanically establishable zero-pass proof and account for how it survives resume and slot collision +BLOCKER | high | §3.1(4), lines 95-102 | The union of story paths cited in prior prompts cannot be established because prompts are not durably recorded and the design supplies no authoritative prompt inventory | An omitted earlier high-risk story silently sheds its lenses and evidence, while fail-closing on uncertainty makes resumed waivers unusable | Persist the governing-story set before calls begin in a record bound to the cycle, or narrow the rule to a source that can actually be reconstructed and state the resulting risk +BLOCKER | high | §3.5 steps 2-3 and 7; §5.1 lines 407-413; CLAUDE.md Mechanics lines 400-417 | Gate B fixes `` to the current tip, which is the WIP commit, then creates a commit on that parent; that is a follow-up commit, not the required amend that replaces the WIP, and its parent-to-tree digest covers only post-WIP changes rather than the implementation | The WIP remains in history and most of the unreviewed diff can fall outside `Waived-target`, creating a direct Gate-B gate-off path | Distinguish `expected HEAD` from `new commit parent`: at Gate B require HEAD to equal the WIP, build the full prospective tree, create the replacement commit with the WIP's parent, and atomically replace HEAD only if it still names that WIP +BLOCKER | high | §3.5 steps 6-7, lines 296-307 | The design performs a separate `HEAD == ` check and then says to write a commit, but never specifies an atomic compare-and-swap of the branch ref against the expected old HEAD | A concurrent ref move between the check and update can attach the closure to or overwrite the wrong history; the prose claims expected-parent binding without defining the exact comparison that provides it | Require a concrete atomic ref update with expected-old and new commit IDs, define failure as a restart before any history mutation, and test the race in a throwaway repository +BLOCKER | high | §3.5 step 3, lines 279-289 | A Gate-A closing tree may contain any number of unrelated `docs/**.md`, README, or MANIFEST changes even though the waiver marker names only one artifact and the Gate-A review prompt covered only that artifact | Other specs, plans, governing stories, or consequential documentation can ride inside one Gate-A waiver and receive neither their own Gate-A decision nor Gate B, which is the exact second closure route the design is meant to prevent | Restrict changed paths to the named artifact plus an explicit closed set of required support edits, or require a separately bound marker and checklist for every gate-relevant artifact in the tree +BLOCKER | high | §3.2 lines 168-193 | Both surviving outage sources remain locally manufacturable: an author can exhaust the selected account's quota, select an already exhausted account, or block network/DNS to produce two transport failures; a transport failure does not imply that the vendor answered as line 191 claims | A human author may self-approve, so deliberately inducing either source creates a sanctioned zero-review closure | Treat local transport as configuration rather than outage, bind the account/configuration to a previously working reviewer identity, require independently obtained provider evidence for quota/service failure, and explicitly acknowledge any abuse path that cannot be closed +MAJOR | high | §3.2 lines 158-216 | The design never defines how raw tool results are classified as quota, transport, service, authentication, defective, or unknown, then discards the payload and retains only the chosen enum | Two compliant agents can classify the same envelope differently and no later reader can audit whether the waiver source was legitimate | Specify exact accepted result shapes and precedence for the pinned reviewer, retain a sanitized digest or protected reference to the source result, and add accept/reject fixtures for every class +MAJOR | high | §3.2 lines 151-166 and 194-209 | Gate B normally uses one `reviewType: full` call with two branches, but `Cause` records a singular `review-spec` or `review-quality` class and the table has no rule for one branch succeeding while the other fails, nor for whether a fresh single-branch call is canonical | A partial successful review can be discarded, double-counted, or incorrectly turn availability of one branch into refusal of the other branch's waiver | Define branch-level canonical calls, artifact handling, ledger credit, recovery, and revalidation for every mixed `full` result +BLOCKER | high | §3.2 lines 194-213 and §3.5 step 6, lines 296-303 | `attempt-failed` consumes the one shared recovery during step 4, but step 6 requires another call plus recovery and simultaneously says it spends the same single attempt and must stop if that attempt is already spent | The `attempt-failed` enum can never complete the ordered close under its own budget, so a major advertised outage path still stalls | Define whether each probe is a distinct pass with its own one-recovery budget or whether the close has one total budget; then make initial observation, revalidation, and post-answer probe consistent with that scope +BLOCKER | high | §3.4 lines 246-261; §3.5 step 8; §5.2 | The request digest excludes the human decision block, and step 8 only requires the two marker lines created by the answer to match the attestation; §5.2 checks only handle, timestamp, gate, and cycle, leaving `authorizes` and `reason` unbound after the answer | The committed rationale and scope can differ from what the human actually authorized while all stated comparisons pass | Require an exact byte comparison of the entire decision block to the captured attestation and state how that attestation is retained through commit read-back +BLOCKER | high | §5.4 lines 497-512 and §3.5 step 8 | Equality strips leading and trailing spaces from unescaped fields, while all digests cover exact unescaped bytes | A committed record can differ from the bytes hashed and authorized yet pass normalized read-back, so the mechanism proves less than the text claims | Define one canonical serialization and use its exact bytes for parsing, equality, hashing, and read-back; reject non-canonical whitespace instead of normalizing it away +BLOCKER | high | §5.1-§5.4 | The grammar does not require exactly one marker, decision block, and checklist block; exactly one occurrence of every field; a fixed field order; or rejection of unknown and duplicate fields | A crafted message can present different values to different parsers or let continuation select a favorable duplicate, undermining authorization and idempotency | Specify a complete closed grammar with exact cardinality and ordering for every record and make duplicates, unknown fields, and multiple blocks hard failures +MAJOR | high | §5.4 lines 490-496 | Git paths may contain comma, spaces, middle dots, backslashes, and newlines, but path/list fields neither restrict those bytes nor define path escaping while lists split on comma-space | Some valid changed trees cannot be represented deterministically and shell parsers can disagree or split one path into several profiles or targets | Use NUL- or length-delimited path serialization internally and a reversible encoded display form, or explicitly refuse trees containing names outside a documented safe path grammar +MAJOR | high | §5.5 lines 531-533 | Oversized checklist answers may be replaced by an arbitrary path, pass id, or query, but the referenced content is not required to be committed, immutable, readable, or covered by a digest | The durable checklist can reduce to dead references while still reporting every item answered | Require referenced evidence to be in the prospective tree or another durable immutable store and include its path plus content digest in both checklist and request digests +MAJOR | high | §5.5 lines 526-538 | The fixed order does not define ordering among multiple changed prompt artifacts or multiple self-checks, and `/` itself has no grammar | Equivalent checklists can hash differently and trigger conflicts, while different implementations can omit or reorder items without a mechanical way to decide validity | Define canonical artifact ordering, self-check identifiers and ordering, exact item-key grammar, uniqueness, and the derivation of the total count +BLOCKER | high | §4 lines 355-381 | Gate-A continuation rechecks only the artifact blob and old commit records; it does not re-resolve governing profiles or bind current CLAUDE.md, AGENTS.md, prompt standards, or cited settled decisions | An unchanged artifact can clear a later reminder after its risk profile or governing rules changed, allowing a stale authorization to outlive the conditions under which it was granted | Bind the continuation to digests of every governing input or require a fresh Gate-A cycle whenever any governing input changed since the waiver commit +MAJOR | high | §6 lines 572-583 | Merge validation says a Gate-A artifact blob must remain unchanged, without defining how a later legitimate Gate-A cycle for an edited artifact supersedes the earlier waiver | Any post-waiver edit makes the branch permanently unmergeable even after a tier-1 review or a new waiver, or else an implementation will ignore the stated check | Validate the latest applicable Gate-A closure for the current blob and define explicit supersession of older waiver markers +MAJOR | high | §3.1(6), §3.5 lines 317-320, and §6 | The merge strategy is bound during the early Gate-A commit and any later change requires repeating the whole close, but Gate A cannot be repeated at the final product-bearing head because step 3 refuses product paths | Choosing or changing squash versus ordinary merge later can strand an otherwise valid branch | Remove merge strategy from the early authorization, or add a separately authorized pre-merge choice that does not require rebuilding the Gate-A closing commit +BLOCKER | high | §6 lines 548-568 | Squash copies a marker whose authorized parent, tree, and message describe the branch closing commit into a different squash commit with a different parent and message; the original closing commit is then unreachable and the copied block has no carrier/source identity | A reader of `main` cannot apply §3.5's three-way read-back to the containing commit or prove which original closure the copied assertion described, contradicting the durable-record acceptance criterion | Define a distinct carried-record envelope naming and digesting the original closing commit and its exact record, and validate that envelope before squash; do not present a verbatim marker as if it described the squash commit +BLOCKER | high | §6 lines 594-608 | An invalidation commit must reproduce missing or corrected blocks verbatim, including `Reviewer-tier: 3`, so the prescribed discovery grep finds the invalidation commit itself as a new apparent closure while subtraction removes only the named invalid SHA | Incident recovery can increase the set of apparent waivers and make the advertised discovery procedure return a false closure | Encode reproduced material so it cannot match the marker query, define exact parsing rather than grep, and add fixtures proving invalidated and invalidation commits produce the intended set +MAJOR | high | §6 lines 594-610 | The invalidation record has no grammar, digest, idempotency key, authority rule, merge carry rule, or definition of which commit SHA to name after squash versus ordinary merge | Concurrent or repeated incident recovery can conflict, disappear at squash, or invalidate the wrong record | Specify a closed invalidation schema and lifecycle for both merge strategies, including deduplication, authorization, and discovery fixtures +MAJOR | high | §3.5 lines 308-320 and §6 lines 594-604 | After step 8 detects a bad commit, HEAD may already point at that invalid commit, but the text says every step repeats without first restoring the original parent; the only repair directions appear later and differ by publication state | A retry can stack a new closure on the invalid commit, change the authorized parent, or leave bad history behind | Make failed read-back a defined rollback state: atomically restore the expected ref before retry when unpublished, otherwise stop and enter the fully specified invalidation flow +MAJOR | high | §10.1 lines 799-813 | The structural oracle has no command-line contract or stable source for OLD, NEW, and the inline mirror; a normal battery run has no PR base, while the text alternately calls these “two frozen texts” and requires old/new plus two shipped copies | The proposed shell check cannot be implemented reproducibly or run in the claimed local and CI battery contexts | Define checked-in fixtures or explicit arguments/base resolution, exact extraction boundaries, exit codes, and how CI and local runs supply every input +MAJOR | high | §10.1 rows 804-813 | Row 3 is `n/a` on OLD and row 4 passes trivially, yet the prose claims rows 1-3 fail on OLD; moreover “exactly the items §3.1 names” is interpretive unless the expected tokens and extraction are explicitly encoded | The claimed counterfactual result overstates what the oracle compares and gives implementers no deterministic expected failure set | Correct the OLD matrix and enumerate the exact byte-level assertions and expected exits for each fixture +BLOCKER | high | §10.2 S1, lines 833-845 | S1 expects NEW to return `CLOSE` but never states that the Gate-A mechanical sweep is green, even though §3.1 makes that a gate-specific precondition | The load-bearing positive probe is wired to expect authorization when its own transcript omits a required condition | Add the green sweep and every other positive precondition explicitly, then generate each negative by changing exactly one field from that complete baseline +BLOCKER | high | §10 lines 865-890 and named-verification matrix | The validation deliberately exercises no Git operation even though the new behavior's highest-risk logic is dedicated-index staging, parent selection, atomic ref replacement, amend/collapse, message read-back, and squash carry | A model agreeing with prose cannot reveal the data-loss, race, or wrong-parent defects in the algorithm, so `battery+check+verification` is not satisfied for the risk path the design actually introduces | Add a throwaway-repository shell harness that executes accept/reject cases for Gate A, Gate B amend, concurrent HEAD movement, hooks rewriting messages, multi-WIP collapse, and both merge strategies +MAJOR | high | §10.2 lines 878-883 | The counterfactual says that if OLD already allowed zero-pass closure then oracle rows 1-3 would pass, but severity closure and marker-field parity are independent changes and could fail even if OLD had some waiver route | This is precisely the gate-proof overclaim forbidden by AGENTS.md: the asserted observation does not follow from the compared condition | Limit the counterfactual to row 1 and S1, and state separate counterfactuals for severity and marker parity +MAJOR | high | §10.2 lines 857-863 | Failure semantics mention missing verdicts and run disagreement but do not explicitly fail an unexpected unanimous verdict or a verdict with the wrong quoted clause | Three consistent wrong answers can satisfy the stated repeatability rule and be recorded as evidence | Require exact token and clause identity for every OLD/NEW scenario and fail on any mismatch, extra token, malformed first line, or wrong cited clause +MAJOR | medium | §10.2 lines 824-831 | The design requires a CLI/model combination supporting exact model IDs, temperature zero, seed, tool isolation, and byte-for-byte prompt capture but names no runner or proves those controls exist; “where the runtime accepts one” also makes seed handling variable | The check may be impossible to reproduce with the selected reviewer and two implementations can satisfy materially different harnesses | Select the exact pinned runner and model in the design, document each supported control and fallback, and make unsupported isolation or determinism controls a hard failure +MAJOR | high | §7 inventory | The claimed repeatable derivation still omits normative §5 effects, including the zero-finding-no-companion rule and companion lifecycle at CLAUDE.md 121-137, the specific `.context/codex-reviews/` fingerprint exclusion at 195-198, the Gate-A `exec` versus document-`review` routing rule at 200-204 and 397-399, and the against-main `baseSha` rule at 400-401 | The decision-procedure rewrite can silently drop these conditions despite the inventory's purpose, and rows 82-87 do not close the class | Re-walk every normative sentence and clause, add separate effect-level rows for these examples, then mechanically record the source line range for every row so omissions and duplicate coverage are auditable +MINOR | high | §3.1 Gate-B table line 145 versus §3.5 | The Gate-B row says evidence is revalidated at step 3, but the ordered close performs evidence revalidation at step 4 | An implementer following the cross-reference can revalidate before the prospective tree is fixed or conclude the wrong step is missing | Change the cross-reference to step 4 and keep the ordering statement in one authoritative place +MAJOR | high | §6 ordinary-merge table lines 563-568 | The ordinary-merge check says “the branch head is the authorized one” for every waived cycle, but a Gate-A authorized commit must be followed by plan and implementation commits, so it cannot also remain the branch head | The literal rule either rejects every ordinary merge containing a Gate-A waiver or tempts implementations to skip the head check entirely | Scope the head-equality requirement to the latest Gate-B closure and separately validate reachability and supersession for Gate-A closures +MAJOR | medium | §3.5 step 6 and §5.4 timestamps | Freshness relies on a readable local UTC wall clock but does not reject backward jumps, large skew, or a decision timestamp later than the probe | A local clock change can make stale outage evidence appear younger than 15 minutes and weakens the claimed authorization-age binding | Define ordering constraints among probe, decision, and commit times, reject negative or implausible ages, and use a monotonic elapsed-time measurement within the live close where available +END OF FINDINGS (34 total) diff --git a/.context/codex-reviews/gate-a-spec-pass-4-dispositions.md b/.context/codex-reviews/gate-a-spec-pass-4-dispositions.md new file mode 100644 index 0000000..aaeadee --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-4-dispositions.md @@ -0,0 +1,97 @@ +# Gate A — spec — pass 4 dispositions (salvage cycle) + +20 findings (4 BLOCKER, 14 MAJOR, 2 MINOR). **None dismissed. 9 applied; 11 held pending one +decision from Daniel.** + +## Trajectory + +| Pass | Findings | Blockers | +|---|---|---| +| 1 | 15 | 3 | +| 2 | 25 | 2 | +| 3 | 23 | 5 | +| 4 | **20** | **4** | + +Findings declining since pass 2; blockers flat. **The composition is the signal:** eleven of +this pass's twenty — including blocker 2 — are downstream of a single element, rider (b)'s +requirement that a dispositions file record the enum drift. + +## Applied (9) + +**B1 — the scope was too narrow a "narrow".** §2.1 excluded duties from `CLAUDE.md` §5 and +`AGENTS.md`, and offered "an absent supplementary PR-bot review" as an unqualified example. But +`docs/pr-review-bots.md` makes a bot review **and** a recorded human decision mandatory whenever +that bot is under **Wait for** — so the example is a mandatory rule in some routing states, and +the form would have been the prescribed record for bypassing it. Verified at line 139. +**Fixed:** applicability is now "something no applicable rule required", the bot example is +qualified to **opportunistically routed** bots, and the shipped paragraph names the Wait-for +branch as explicitly out of reach, along with `AGENTS.md`, project docs, CI, branch policy and +platform rules. + +**B3 — `AGENTS.md` does change after all.** Invariant 11 says `scripts/check-invariants.sh` +carries "two narrow checks" and enumerates them; §5.2 makes it three. Third time this class has +bitten in this cycle (the manifest, the withdrawn-checker chain, now the invariant's own count). +**Fixed:** `AGENTS.md` invariant 11 is in §4's table, with its calibration sentence kept verbatim. + +**B4 — §4 contained mutually exclusive instructions.** The rewrite replaced the table and the +"no new file" paragraph but left the older "Two new files are added" paragraph standing, so an +implementer could legitimately resurrect the withdrawn checker. My incomplete edit. **Deleted.** + +**F5** — the closed-enum assertion is now bounded to the §5 region of each file, requires +exactly one canonical match with no fifth token, and fails closed on read or parse errors. +Unbounded, it could have been satisfied by a line in `workflow-init.md`'s surrounding command +prose, or by one naming four tokens while negating the rule. + +**F6** — "the edited regions" is now an enumerated table of five blocks with expected +occurrence counts and a stated extract-and-diff procedure. Left to judgement, two people diff +different regions and both record success. + +**F14** — post-merge restoration said "verbatim" while changing a line inside the block, which +its own identity rules classify as a hard conflict. Now: reproduce the block byte-identically, +`Ref:` included, with restoration context in prose outside it. + +**F15** — `ref:` / `Ref:` casing, unified. **F19** — the story's "two prompt edits and a +paragraph" understated the surface; withdrawn. **F20** — the tier-2 story called a Major a +Blocker. + +## Held (11) — all downstream of one decision + +**B2, F7, F8, F9, F10, F11, F12, F13, F16, F17, F18.** + +Rider (b) began as: state the severity enum as a closed set; let the reader normalize an +unrecognized token to `MAJOR`; **record the drift in that pass's dispositions file**. The last +clause is what has grown. It now needs, and pass 4 finds defects in, all of: + +- a token-identity rule with a whitespace edge case that currently contradicts its own example + (F8), and a six-field parse `CLAUDE.md` §5 never actually defines (F7); +- a **bijection** check, because the audit only verifies cited lines resolve, not that every + drifted line is cited — so one of four `IMPORTANT` findings can be recorded and the other + three silently omitted (F9); +- a freshness rule, a two-artifact audit, a discount path, and a logical-pass / attempt / + credited-count identity model to survive single-branch recovery (F11, F12, F13); +- **edits to four shipped hook reminder strings and their test assertions** (B2), which + currently instruct a retry to *delete* the slot rider (b) now says to *preserve* — the + standing falsification lens working exactly as intended; +- a `docs/hardening-log.md` append-only supersession row, because a ledger row states + categorically that dispositions never participate in pass validation (F16); +- more §2.3 rows and a rollback account covering all of it (F17, F18). + +**The alternative is to drop the record requirement.** Normalize an unrecognized token to +`MAJOR`, and stop. The drift stays permanently visible **in the findings file itself**, which +carries the original token on the finding line — that artifact is retained by §5 already, and it +is better evidence than a companion that can be deleted. Every held finding evaporates: no +identity model, no audit, no bijection, no hook-message edits, no ledger supersession, and +§5's "companions never validate a pass" is left **untouched** rather than narrowed. + +What is lost: the record made the normalization deliberate rather than silent, and let a reader +see at a glance that a pass had drifted. Against that, the findings file shows the same thing to +anyone who opens it. + +**Not taken unilaterally.** Daniel named riders (b) and (c) as the salvage; this narrows what +(b) ships. The discriminating check — closed enum + normalization, §5.2's assertion — survives +either way, so the `+check` obligation is unaffected. + +## Status + +Not clean; 4 passes taken. Held findings are unapplied and recorded. Blocked on the rider +question. diff --git a/.context/codex-reviews/gate-a-spec-pass-4-dispositions.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-4-dispositions.pre-2026-08-14.md new file mode 100644 index 0000000..0a40133 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-4-dispositions.pre-2026-08-14.md @@ -0,0 +1,28 @@ +# Gate A — spec — pass 4 dispositions + +Artifact: `docs/superpowers/specs/2026-08-05-hardening-ledger-supersession-design.md` +11 findings: 7 Major, 2 Minor, 2 Nit. All eleven validated as correct against the files. +**None applied** — the pass is not clean, and the pinned exit sends anything new to Daniel. + +Reviewer: `mcp-codex-dev@1.0.1`, `codex-cli 0.144.5`, model `gpt-5.6-sol`. Same server as +passes 1–3; the model was not recorded for those passes, so a change cannot be ruled out. + +| # | Sev | Verdict | Reason | +|---|---|---|---| +| 1 | Nit | **Valid** | Confirmed against `docs/hardening-log.md:7`. Source reads ``resolve a `pending` row``; §2.0 quotes it inside a single-backtick span, so `pending`'s own backticks are dropped. Same class as pass 3's finding 8 — a locator that is semantically right and not byte-findable. | +| 2 | Minor | **Valid** | §3.1's citation writes `` \`PostToolUse\` `` inside a single-backtick span. Backslashes do not escape code-span delimiters in markdown, and those bytes are not in `CLAUDE.md`. The pass-3 fix added the `**` markers but left the span unterminated in the other direction. | +| 3 | Nit | **Valid** | `SKILL.md:157` requires a **short** one-line escaped `finding`; §5's table carries only "one line, with `\|` escaped". One old condition unaccounted, which is exactly what rider 3 exists to catch. | +| 4 | Major | **Valid** | `CLAUDE.md` §5 Mechanics defines when a Gate-B cycle *closes* and never when it *opens*. §2.1's "the one you have open" therefore has no defined start, and a row appended during implementation before the first `WIP:` snapshot classifies both ways. Codex's six self-test verdicts otherwise all come out as the spec says. | +| 5 | Major | **Valid** | Landed is observer-relative by design ("landed to you"). Cycle B may append a supersession against a row cycle A is still entitled to amend in place, so B's entry can be stale or its locator broken on arrival. Not a contradiction in the text — an unhandled concurrent path. | +| 6 | Major | **Valid, and the pass's most important finding** | §2.2 routes entry locators through "the same fallback" — a distinguishing fragment of *the row's* `finding`. Two same-day entries against one row share that row, so they share the fragment too. The fallback cannot separate the collision it was added for, which means pass 3's finding 4 was answered in words and not in mechanism. | +| 7 | Major | **Valid** | Only one line shape is defined, and it supersedes a *row*. Nothing says how a correcting entry names the entries it retires, distinguishes itself from a row correction, or cites its answer — while §2.2 requires that path to exist. | +| 8 | Major | **Valid** | "a correcting entry retires exactly the entries it names and states what now holds" contradicts the same paragraph's cite-don't-restate rule. Two incompatible instructions on one surface — prompt-standards item 7, the same defect §2.0 exists to remove. | +| 9 | Major | **Valid** | Check 2 compares the two regions **to each other**. Two identically wrong regions that both end at the sentinel pass. The check is titled "the convention actually reached both surfaces" and proves parity, not presence — the `AGENTS.md` "never describe what a gate proves" class, on a check written to close that class. | +| 10 | Major | **Valid** | Check 1 greps the entry anywhere in the file; check 2 excludes the block, `Columns:` and the table. So §2.2's layout decisions — exactly one label, positioned above `Columns:`, absent from the template — are validated by nothing. | +| 11 | Minor | **Valid** | Story §5 still poses template reach as an open question and §4 marks invariants 11 and 12 conditional on it, while the design settles it (§4 lists `workflow-init.md` and the version bump). The authoritative source presents settled scope as unresolved. | + +## Decision needed from Daniel + +Findings 4, 5, 6, 7 and 8 are convention-text changes, not editorial fixes: they change what +the shared prose says, which lands in every scaffolded ledger. 9 and 10 change what §7 claims +to validate. 11 edits the story again. diff --git a/.context/codex-reviews/gate-a-spec-pass-4.md b/.context/codex-reviews/gate-a-spec-pass-4.md new file mode 100644 index 0000000..8404c3d --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-4.md @@ -0,0 +1,21 @@ +BLOCKER | high | §2.1 lines 158-180 and docs/pr-review-bots.md "Then the two facts separate" | The read-alone scope excludes only duties imposed by CLAUDE.md §5 or AGENTS.md, while its unqualified example is an absent supplementary PR-bot review; docs/pr-review-bots.md makes exactly that review plus a recorded human decision mandatory whenever a bot is under Wait for, and other mandatory repo or platform rules can likewise live outside those two files. | A reader can use this form as the prescribed record for bypassing a mandatory non-§5 rule, so the paragraph functions as the kind of waiver its scope claims to forbid. | Define applicability as work not required by any applicable repository, workflow, CI, branch-policy, or platform rule, and qualify the bot example as an absent review from a bot currently routed opportunistically; explicitly exclude the Wait-for branch. +BLOCKER | high | §4 lines 381-399 versus plugins/dev-workflow/hooks/codex-gate.sh lines 360-370 and codex-gate.test.sh lines 1891-1898 | The site table omits the shipped hook reminder strings and their test copies, which still instruct retries to delete the original findings slot and warn about a late writer landing in that same slot; rider (b) instead preserves attempt 1 and writes retry attempt 2 to a suffixed slot. | Users can receive two contradictory recovery procedures from shipped prompts, and following the hook message destroys or overwrites the artifacts the new identity and audit model depends on. | Add codex-gate.sh and codex-gate.test.sh to §4 and amend every failure, no-result, and backgrounding reminder plus its asserted test string to name the attempt-suffixed target and the retained prior attempt. +BLOCKER | high | §4 lines 394-399 and AGENTS.md invariant 11 lines 170-177 | The design says AGENTS.md remains unchanged, but invariant 11 currently says scripts/check-invariants.sh has exactly two narrow prompt checks and enumerates only Target model and checklist count; §5.2 adds a third closed-enum check. | Landing the design as written makes the single source of truth false and violates the standing docs-drift invariant in the artifact meant to preserve it. | Add AGENTS.md to the site table and update invariant 11's exact count and guarded-spelling inventory while retaining its no-comprehensive-check calibration. +BLOCKER | high | §4 lines 394-415 | The section first says no new file is added and the parity checker was withdrawn, then lines 412-415 direct that two new parity-checker files are added and that AGENTS.md's tree and command rows gain them. | The implementation scope has mutually exclusive instructions at the exact decision pass 4 is meant to settle; an implementer can legitimately resurrect the withdrawn checker and its invalid counterfactual. | Delete the stale two-new-files paragraph and leave one explicit statement that the only mechanical addition is an assertion inside the existing checker and suite. +MAJOR | high | §5.2 lines 451-469 | The proposed assertion is described as finding any line in each whole file that names the four tokens, but workflow-init.md contains command prose as well as the inline §5 template, and a line can name all four while negating the rule, allowing extras, or living outside the scaffolded section. | The battery can pass while the shipped §5 copy still lacks a closed permitted-token rule, so the claimed check does not establish the exact comparison §5.2 says it establishes. | Require one exact canonical closed-enum sentence inside the bounded §5 region of each file, reject zero or multiple occurrences and extra severity tokens, fail closed on read or parse errors, and make the regression fixtures cover wrong-section and open-set false positives. +MAJOR | high | §5.3 lines 477-486 | "Edited regions" has no defined boundaries, extraction command, canonical snippets, or expected occurrence count, even though the two sections legitimately differ and the verification is the sole evidence for AC 9. | Two implementers can diff different hand-selected regions and both record success while one mirrored clause is missing or placed in the wrong section. | Mark each shared insertion with unambiguous boundaries or enumerate exact canonical blocks and occurrences, then give a read-only extraction and diff procedure whose output is recorded. +MAJOR | high | §3 lines 282-316 and CLAUDE.md §5 lines 90-105, 142-150 | Normalization is limited to an "otherwise-valid six-field line", but the settled acceptance rule never defines an escape-aware six-field parser: it says finding lines and count, while the writer rule only says literal pipes are escaped and never defines how backslashes make a pipe escaped. | The structural-versus-enum boundary is undecidable for tokens containing backslashes or pipes, so the same line can be normalized by one reader and rejected as INCOMPLETE by another. | Specify a complete line grammar and decoding algorithm, including odd-versus-even preceding backslashes, exactly six fields, separator spacing, and which escapes are legal, and amend the acceptance clause rather than claiming its grammar is untouched. +MAJOR | high | §3 token identity lines 321-325 | The rule says only the one format space adjacent to each separator is stripped and no further trimming occurs, yet it declares ` IMPORTANT ` and `IMPORTANT` identical; for the first severity field the leading space is not adjacent to a preceding separator and therefore survives. | Drift identity, multiplicity, and audit outcomes depend on incompatible whitespace rules. | Choose one rule and demonstrate it byte-for-byte: either trim a precisely defined amount from both field edges and keep the equivalence example, or retain separator-only stripping and change the example so a leading space is distinct. +MAJOR | high | §3 drift grammar and audit lines 304-342 | The grammar requires one record per distinct unrecognized token, but freshness and audit only check the forward direction that every cited line resolves to the stated token; they never check that every unrecognized severity line is cited exactly once. | A dispositions file can cite one of four IMPORTANT findings, omit the other three, and still pass every stated audit while falsely claiming the pass's normalization is recorded. | Add the reverse comparison: derive all unrecognized severity occurrences from the accepted findings file and require an exact bijection with decoded token and line-number pairs in the drift records. +MAJOR | high | §3 identity lines 363-369 versus CLAUDE.md §5 lines 168-185 and the story lines 206-213 | "One hook count" is false for the motivating recovery: a successful full call with one valid branch can count once, and the successful single-branch retry is another routed PostToolUse call that can count again; the story already records that the unchanged hook counts each call. | This is an uncalibrated gate-mechanism claim and can make the visible hook floor appear satisfied by attempts rather than credited logical passes. | State that the hook may count both invocations, that its counter is not the credited count, and that the reader must discount the extra attempt regardless of the hook display; remove the one-hook-count claim. +MAJOR | high | §3 identity lines 354-370 | Pairing the latest attempt of each Gate-B branch has no invariant that both attempts reviewed the same WIP head, base SHA, prompt, profile, evidence entry, and additional context; a concurrent amend or other state change between attempts can produce a structurally valid mixed pair over different inputs. | The logical pass can be credited even though no single pass reviewed one coherent artifact and evidence state. | Capture and compare the immutable review-input identity before accepting a mixed-attempt pair, require it to match across branches, and discard the logical pass and start a new ordinal when any input changed. +MAJOR | high | §3 discount path lines 343-361 | Credited count can fall when an audit discounts an earlier logical pass, but the design never says which ordinal the reopened loop uses; reusing the missing ordinal conflicts with retained artifacts and a spent retry budget, while using the next ordinal is not stated. | Recovery after a late discount can overwrite slots, accidentally reuse an exhausted attempt, or count a replacement as the wrong logical pass. | Specify that ordinals are monotonic within a cycle and never reused after discount, that the replacement uses the next unused pass number, and how Todos and hook overcounts are reconciled. +MAJOR | medium | §3 attempt model lines 354-370 | Attempt selection is defined only for the single-branch Gate-B example; the general table says every retry gains a suffix, but it never says which file is authoritative after a Gate-A retry or a two-branch full rerun, what happens to invalid earlier artifacts, or how a dispositions companion is paired in those paths. | Readers can validate attempt 1 while the reply names attempt 2, mix full-rerun branches, or attach a drift record to the wrong attempt. | Define the selected-attempt algorithm for Gate A, full Gate B reruns, and single-branch resumes, including exact companion filenames and a rule that only the selected artifact set can earn or retain credit. +MAJOR | high | §2.4 line 263 versus lines 258-260 | The post-merge restoration row requires reproducing a missing record "verbatim" while changing its Accepted because line to begin `restores :`; that is not verbatim, and if it keeps the original ref it is a byte-different duplicate that the preceding rule classifies as a hard conflict. | The prescribed repair violates the record's own identity and deduplication rules and has no deterministic merge result. | Either copy the original block byte-identically and put restoration context outside it, or create a new uniquely referenced record with an explicit relationship to the missing ref and define that relationship alongside supersedes. +MINOR | high | §2.1 line 164 versus §2.4 lines 258-270 | The canonical form spells the identity field `ref:` while every carry, conflict, and identity rule calls it `Ref:`. | A case-sensitive collector or a person following the later spelling can fail to find or can duplicate the canonical record. | Pick one exact field label and use it consistently in the form, carry rules, supersession syntax, and any search procedure. +MAJOR | high | §4 line 390 and docs/hardening-log.md line 97 | The existing append-only ledger row states categorically that dispositions never participate in pass validation and a cycle with none behaves exactly as before; rider (b) deliberately makes a dispositions file mandatory for normalization-dependent passes, but §6 names no superseding ledger action and §4 says only "whatever §6 lands". | The change leaves a durable mechanism claim false in the repository's hardening authority, precisely the drift class the design says it audited. | Name the exact append-only supersession or follow-up row in §4 and §6, preserving the old row while explicitly narrowing its claim for normalization-dependent passes. +MAJOR | high | §8 rollback lines 551-554 | Rollback is described as two prompt reverts plus a version bump, omitting the new checker assertion and regression case; reverting the closed-enum prompt lines while leaving that assertion active makes the invariant check fail, and the design also omits the retry-related hook messages that must be changed for consistency. | The documented rollback cannot produce a green repository and can encourage a version decrease or incomplete release repair. | Define a forward-version rollback that reverts the prompt behavior, enum assertion and tests, any amended hook messages/tests, and related changelog entry while preserving append-only history through supersession rather than deletion. +MAJOR | high | §2.3 lines 203-246 and CLAUDE.md §5 lines 168-199, 418-428 | The old-condition table omits the separate Mechanics timeout/abort deletion clause and the "What this does not do" detection/provenance claims, although attempt suffixes, retained artifacts, normalization audits, and possible multiple hook counts directly change all of them. | The shipped §5 can retain old same-slot deletion instructions and an obsolete enumeration of what reader detection catches, producing internal contradictions and violating the old-condition accounting invariant. | Add explicit kept/narrowed dispositions for both clauses and include their exact resulting edits in the CLAUDE.md and workflow-init site rows. +MAJOR | medium | reviewer-availability story §6 lines 299-305 and design §4 | The amended story says the residue is "two prompt edits and a paragraph", while the design requires a shell assertion, regression suite, version and changelog work, ledger and todo updates, two story amendments, a multi-part record protocol, and substantial recovery/audit behavior. | The settled story understates the implementation and validation surface and can drive an undersized plan or reinforce the profile rationale using a scope that is no longer true. | Replace the residue summary with the actual bounded surfaces from the final §4 after its omissions are fixed, while keeping the historical story-size explanation separate. +MINOR | high | tier-2 dependent story lines 29-44 | The text says "The three blockers" but the middle item is explicitly labelled `Pass-3 M14`, i.e. a Major finding, not a Blocker. | The dependent story misstates the severity evidence used to justify its containment direction. | Say "three stopping findings" or preserve the original severity labels and counts exactly. +END OF FINDINGS (20 total) diff --git a/.context/codex-reviews/gate-a-spec-pass-4.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-4.pre-2026-08-14.md new file mode 100644 index 0000000..eb6ab44 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-4.pre-2026-08-14.md @@ -0,0 +1,12 @@ +NIT | high | SWEEP §2.0 quoted header phrase | the claimed existing phrase drops the source Markdown backticks around `pending`, so it is not byte-for-byte correct even after hard-wrap normalization | this fails Rider 1's explicit quotation check and makes the spec's own exact-text assertion false | use a double-backtick code span and quote the source fragment with `pending`'s backticks intact +MINOR | high | SWEEP §3.1 citation and inline-code syntax | the citation writes `PostToolUse` with backslash-escaped backticks inside a single-backtick code span; backslashes do not escape code-span delimiters, and those backslashes are not in `CLAUDE.md` | the locator is neither byte-for-byte source text nor one intact Markdown span, so the pass-3 citation fix does not hold and a copied ledger entry cannot be byte-found as claimed | delimit the whole citation fragment with double backticks and leave the source's single backticks around `PostToolUse` unescaped +NIT | high | SWEEP §5 old-conditions audit | the current Log format requires a `short` finding, but the disposition table keeps only one-line shape and pipe escaping | Rider 3 requires every old condition to be present, and this judgment constraint otherwise disappears from the audit | add `finding is short` to the same kept disposition or account for it separately +MAJOR | high | READ §2.1 boundary and six-case self-test | `CLAUDE.md` Mechanics defines how a WIP snapshot closes but never defines when a Gate-B cycle opens, so ownership is ambiguous before the first WIP; applying the written rule otherwise yields all six requested verdicts: case 1 landed → append §3.1's entry; case 2 own open cycle → amend row D in place; case 3 earlier completed unmerged cycle → landed and append an entry; case 4 own open cycle after WIP → amend in place; case 5 own open-cycle `pending` resolution → append a same-fingerprint resolving row; case 6 unknown authorship → treat as landed and append an entry | a row created during implementation before the first WIP can be classified as landed because no cycle is yet observably open or as open-cycle because that implementation later enters Gate B, producing opposite edit-versus-append actions | define the cycle's opening event and explicitly route a row appended before the first WIP snapshot +MAJOR | medium | READ §2.1 concurrent-cycle path | landed status is observer-relative: cycle B must treat cycle A's still-open row as landed, while cycle A remains allowed to amend that same row in place | B can append a supersession against text or a finding fragment that A subsequently changes, leaving the correction stale or its locator broken and contradicting the rationale that the boundary is exactly where the record stops being the author's own | define what freezes A's amendment right once another cycle imports or supersedes the row, or forbid consuming an open-cycle row until its authoring cycle closes +MAJOR | high | READ §2.2 entry-locator fallback | `under the same fallback` points back to a fragment of the superseded row's `finding`; two same-day entries against the same row necessarily share that row and fragment, so this fallback cannot distinguish the collision it was added to resolve | a correcting entry still cannot retire exactly one of the valid same-day pair, so the pass-3 fix does not execute its stated decision | specify that an entry locator falls back to a distinguishing fragment of the entry's own `what is false` text, with an explicit example and exact-one stop +MAJOR | high | READ §2.2 correcting-entry format | the only defined line shape supersedes a ledger row; no line shape says how a correcting entry names one or more prior entries, distinguishes entry correction from another row correction, or cites its answer | authors cannot implement the append-only correction path consistently, and readers cannot determine which earlier entries were retired as AC 4 requires | add a concrete correcting-entry example and define its locator, fields, delimiter placement, and multiple-target syntax +MAJOR | high | READ §2.2 correction semantics | `a correcting entry ... states what now holds` conflicts with the same paragraph's instruction to cite the current answer and never restate it | the shared prompt gives opposite instructions on whether the correction contains the new answer, violating prompt-standards item 7 and recreating the stale-copy risk the convention is meant to avoid | require the correcting entry to name what was false in the prior entry and cite where the current answer lives, without restating that answer +MAJOR | high | READ §7 Check 2 completeness | the parity operation compares the two target regions only to each other; two identically truncated or wrong regions containing the end sentinel pass even if §2.1 or most of §2.2 never lands | the check titled `the convention actually reached both surfaces` proves equality, not presence of the designed convention, and can satisfy `battery+check` with the feature missing from both prompts | compare each normalized region to one explicit expected region or add a named read that checks every required convention clause on both surfaces +MAJOR | high | READ §4 layout and §7 validation | Check 1 greps the entry anywhere in the ledger, while Check 2 deliberately excludes the `Superseded rows:` block, `Columns:` paragraph, and table; no validation checks one label, its required position, or that §3.1's entry is under it | an entry at the file end, a missing or duplicate label, or a block below the table passes every stated check while violating the layout and empty-template decisions | add a named structural check for exactly one label between convention prose and `Columns:`, with the entry inside that block and no label in the template +MINOR | high | READ governing story §§4–6 | the story still calls template reach an open question, marks invariants 11 and 12 conditional on that unanswered choice, and sizes the work as `possibly one inline template`, while the design definitively chooses the template and plugin changes | the authoritative change source presents settled scope as unresolved and can steer a later plan or reviewer to omit the version bump or prompt-conformance work | mark the question resolved, make invariants 11 and 12 unconditional, and update the suggested surface while preserving the recorded decision history +END OF FINDINGS (11 total) diff --git a/.context/codex-reviews/gate-a-spec-pass-5-dispositions.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-5-dispositions.pre-2026-08-14.md new file mode 100644 index 0000000..eb4c999 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-5-dispositions.pre-2026-08-14.md @@ -0,0 +1,45 @@ +# Gate A — spec — pass 5 dispositions + +Artifact: `docs/superpowers/specs/2026-08-05-hardening-ledger-supersession-design.md` +14 findings: 11 Major, 3 Minor. All fourteen read as correct. **None applied** — the pass is +not clean, and the pinned exit sends anything new to Daniel. + +Reviewer: `mcp-codex-dev@1.0.1`, `codex-cli 0.144.5`, model `gpt-5.6-sol`. Same as pass 4. +(Recorded here rather than in the findings file: CLAUDE.md §5 admits no line there that is not +a finding line or the terminator, so a model line would make the pass malformed.) + +**Seven of the eleven Majors are defects in the pass-4 fixes**, not in pre-existing text: +1, 2, 3, 4, 5, 6 and 10 all land on §7's rewritten checks or on §9's new residual bullet. The +fixes were written and reviewed in the same pass; this is what the next pass is for. + +## Verified mechanically, not taken on the reviewer's word + +| # | Claim | Result | +|---|---|---| +| 2 | a `grep -F` pattern beginning with `-` parses as an option | **Reproduced.** `grep -cF '- · …' file` → `invalid option`. With `-e` it returns `1`. | +| 3 | `grep -c` counts matching *lines*, not occurrences | **Reproduced.** `printf 'aa xx aa\n' \| grep -cF 'aa'` → `1`, not `2`. Fatal after wrap-normalization collapses a paragraph to one line. | +| 12 | an append-only claim survives outside the change surface | **Confirmed.** `docs/coding-workflow.md:204` — "strictly append-only: history is never rewritten". Not in §4. | + +| # | Sev | Verdict | Reason | +|---|---|---|---| +| 1 | Major | **Valid** | `N="$(unwrap "$F")"` captures *content*, then `grep … "$N"` uses it as a *filename*. The block is presented as commands that "must report `1`", so it has to be runnable. | +| 2 | Major | **Valid, reproduced** | The two anchors added in pass 4 to validate the format lines are exactly the two that cannot run. | +| 3 | Major | **Valid, reproduced** | Assertion 1 claims exactly-once cardinality and cannot measure it. Worse after normalization, which is the pass-4 fix that created the exposure. | +| 4 | Major | **Valid, and the pass's most important finding** | The anchor list does not support "dropping any decision §2 settles removes at least one" — the unknown-provenance default, the wrong-when-written cause, the recurrence effect, the locator fallback, the cumulative rule, cite-don't-restate and the union-merge repair are all droppable with all seven anchors intact. That claim is the AGENTS.md "never describe what a gate proves" class, committed inside the fix written to close that class. | +| 5 | Major | **Valid** | Check 3's assertions 1–3 exist only as `#` comments, and assertion 4's desired zero-match exits nonzero. Presented as mechanical; not executable. | +| 6 | Major | **Valid** | A live entry in the template without a label passes assertion 4, and §7 says check 1's named read catches it — but that read examines the repo ledger only. A second overclaim, same class as 4. | +| 7 | Minor | **Valid** | Story §1 still states the un-narrowed rule in the present tense. Defensible as pre-change framing, but it is the story's own §1 against its §1 accounting. | +| 8 | Major | **Valid** | Both fallbacks require a distinguishing fragment; neither format line shows where it goes, and the correcting line gives no per-retired-entry fragment slot. The pass-4 fix named the right field and never placed it — the same words-not-mechanism failure pass 4 found in pass 3. | +| 9 | Major | **Valid** | The six verdicts all come out as the spec says. But the cycle *close* is still defined only as "the commit that replaces the `WIP:` snapshot", so a trivial/N-A commit, an abandoned change and interleaved work have no close. Pass 4's opening fix moved the ambiguity rather than removing it. | +| 10 | Major | **Valid** | §9's new bullet asserts A's amendment "does not move" date + fingerprint. Nothing freezes either field, so B's locator can dangle, not merely lose its fallback. The residual was written as bounded without establishing the bound. | +| 11 | Major | **Valid** | Story §2 still promises a reader can tell whether *any* row describes current behaviour — the exact requirement AC 4 was amended away from. §2 was amended in pass 3 for the append-only clause and its first clause was not re-read against AC 4. | +| 12 | Minor | **Valid, confirmed** | The standing falsification lens working: `docs/coding-workflow.md:204` teaches the absolute rule and is outside §4. | +| 13 | Major | **Valid** | "Wrong when it was written" admits a row whose *hardening identity* — rung, ref, fingerprint — was never true, while the same paragraph says supersession never touches the hardening and the row keeps counting. A nonexistent hardening would stay in recurrence lineage and escalate. | +| 14 | Minor | **Valid** | The correcting line carries one row locator but the prose lets it retire "exactly the entries it names", with no requirement they concern that row. | + +## Decision needed from Daniel + +Findings 1, 2, 3 and 5 are mechanical and cheap — the check blocks need to become runnable +shell. Findings 4, 6 and 10 are overclaims and each has the two-way exit Daniel already set for +finding 10 of pass 4: validate it or state in §9 that nothing validates it. Findings 9 and 13 +are convention-text decisions. Findings 8, 11, 12 and 14 are contained edits. diff --git a/.context/codex-reviews/gate-a-spec-pass-5.md b/.context/codex-reviews/gate-a-spec-pass-5.md new file mode 100644 index 0000000..045dcc6 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-5.md @@ -0,0 +1,17 @@ +MAJOR | high | §2.0, §2.1, and docs/pr-review-bots.md lines 117–151 | PR #14 is still described as the precedent whose scope the new form follows, but #14 occurred while CodeRabbit was under Wait for and the applicable rule required an explicit recorded human decision; §2.1 now excludes every Wait-for case and says this form is not that decision | The historical derivation is false: the form covers optional omissions that #14 did not exemplify, so a reader cannot rely on the claimed precedent to understand the boundary | State that #14 is deliberately out of scope under the stronger any-applicable-rule boundary, and justify the optional-work record independently rather than calling #14 its precedent +MAJOR | medium | §2.1 “an absent review from a bot currently routed opportunistically” and docs/pr-review-bots.md lines 134–145 | The shipped example makes proceeding without an opportunistic review a human exception whose record “goes in” the commit, while the routing authority says an absent opportunistic review blocks nothing, needs no exception, and must not regain exactly that per-quiet-bot ceremony | The two instructions can produce opposite behavior for the same absent bot and §4 incorrectly claims docs/pr-review-bots.md remains true unchanged | Explicitly say ordinary opportunistic absence is not an exception and needs no record; if the form is meant only for a separately requested optional review that a human later cancels, name that narrower trigger +MAJOR | high | §3 drift-record rationale, §6 normalization backlog, and parent-story AC 6 amendment | The deletion rests on the claim that §5 retains each findings file permanently and that it “may not” be deleted, but §5 imposes no post-validation retention rule, the files live under ignored local `.context/`, and this design’s own §6 says slot collisions already destroyed predecessor findings | The original token is visible while the pass is read but is not a permanent or durable record, so the stated premise for deleting the companion record is false and conflicts with the closure record’s own account of unprotected local history | Replace the permanence argument with the narrower observation actually established—normalization is visible to the reader in the accepted file at decision time—and explicitly decide whether that ephemeral visibility is sufficient, without claiming retention §5 does not provide +MAJOR | high | §2.3 row “Recovery: one attempt per pass, shared” | This surviving row says rider (b) specifies retry filenames and pairing, but the revised rider (b) contains only severity syntax and normalization; the attempt-suffixed slot model and its pairing rules are listed in §3 as deleted | This is a direct orphan from the cut and falsely attributes operational retry behavior to the remaining rider | Delete the duplicate row or rewrite it as kept untouched with no rider reference; keep any genuinely required filename behavior attributed to the existing §5 recovery text +MINOR | high | §3 paragraph beginning “So §5’s companions…” | The paragraph says §2.3 has no companion rows, but §2.3 contains an explicit “Companions are advisory” row | The deletion accounting contradicts the table it cites and makes it unclear whether companion semantics were actually re-walked | Change the claim to say the existing companion row is retained unchanged and no new companion-validation row is added +MAJOR | high | §4 Sites, the paragraph after its table, and parent story §4 first invariant bullet | The path table correctly requires editing AGENTS.md invariant 11 from two checks to three, but the following paragraph says AGENTS.md is therefore unchanged, and the parent story likewise says the design records that AGENTS.md needs no edit at all | The implementation surface and invariant accounting disagree about a required source-of-truth edit, inviting either a stale invariant inventory or an unplanned change | Keep the invariant-11 table row and amend both “AGENTS.md unchanged” claims to distinguish the withdrawn new-checker inventory churn from the still-required narrow-check count update +MAJOR | high | §5.2 bounded-region algorithm and counterfactual | `CLAUDE.md` §5 has a next level-2 heading, but the inline `workflow-init.md` §5 template ends at its four-backtick fence and has no next level-2 heading inside that template; a line parser either fails the region or runs onward into later scaffold templates until `## Classes` | The assertion is not implementable with the stated common delimiter, and at `df850ab` it can fail because the workflow-init region is unfindable or misbounded rather than because the enum sentence is absent, defeating the claimed counterfactual | Specify per-file structural boundaries: next level-2 heading for CLAUDE.md and the closing fence of the `CLAUDE.md` template for workflow-init.md, with tests proving content after each boundary cannot satisfy or duplicate the assertion +MAJOR | medium | §3 canonical syntax and §5.2 “canonical closed-set sentence” | No exact canonical sentence or complete matching grammar is specified, yet the checker must recognize exactly one line, reject a fifth severity token, and distinguish that token from arbitrary prose; “names all four and no fifth severity token” leaves the token universe and permitted surrounding text undefined | Implementations and fixtures can choose different meanings while all claiming conformance, so the mechanical guarantee is not reproducible | Pin the exact source line byte-for-byte, or define a complete anchored grammar and the exhaustive token vocabulary used to detect an extra severity +MINOR | high | §4 scripts/check-invariants.sh row and scripts/check-invariants.sh lines 24–53, 261–280 | Adding a third prompt-conformance assertion also falsifies the checker’s own “two prompt-conformance checks below” inventory and leaves its marked mutation procedure describing only checks 4a/4b; §4 lists only insertion of the assertion | The checker would ship internally stale instructions, and future mutation evidence would not say how the new load-bearing assertion is neutered or audited | Include the checker’s count/inventory comments and a marked mutation procedure for the new assertion in the path-table change, plus the corresponding expected flipped-set record in the suite +MAJOR | high | §5.3 Parity and parent-story AC 9 | The story still requires “§5 and its inline-template mirror agree after the change,” while the design silently redefines agreement to four inserted blocks because the full sections already differ; the story criterion has no amendment recording that narrowing | The design under-implements a live acceptance criterion and violates the old-condition accounting rule it cites elsewhere | Amend AC 9 explicitly to inserted-block parity with kept/narrowed/dropped reasoning, or make the implementation satisfy the criterion’s existing whole-section reading +MINOR | medium | §5.3 parity procedure | Only the human-exception block has usable first/last anchors; “§2.4’s placement, identity and supersession rules,” the enum reader rule, and the carry sentence are concepts rather than pinned source boundaries, yet the procedure requires extracting each by its first and last line and checking exact occurrence counts | Two verifiers can select different text and both truthfully report an empty diff, so the named verification is not reproducible | Give the exact first and last source lines for every inserted block, or add explicit begin/end markers whose presence and uniqueness are themselves checked +MAJOR | high | §2.3 old-condition accounting | The table claims every §5 clause the earlier waiver-shaped text could touch was walked, but it omits at least the workspace opt-out rule that says the gates still apply, the mandatory separate Gate-A spec and plan runs, and the profile-resolution STOP; the general “any STOP” wording directly reaches the omitted STOP | This repeats the repository’s recorded failure mode: a general assurance substitutes for accounting for concrete terminal conditions, so a future edit can accidentally weaken an unlisted branch | Add individual kept/untouched rows for each omitted gate-applicability and terminal clause after re-walking all of §5, including profile parsing and the opt-out marker +MAJOR | high | §8 rollback | Removing the new assertion would reduce the narrow mechanical checks back from three to two, but the rollback list omits reverting AGENTS.md invariant 11 and the checker’s own updated inventory | The prescribed forward rollback leaves the source of truth making a false count claim and violates the same calibration requirement that brought AGENTS.md into §4 | Add a forward edit restoring invariant 11 and all checker-inventory prose to the two-check state, alongside removal of the assertion and its tests +MAJOR | high | §2.4 Ref identity and concurrent records | Two branches can independently choose the same eight-hex Ref for different decisions; the design turns their byte-different copies into a hard pre-merge conflict but provides no legal resolution, because committed records are never edited and a superseding record cannot make the two original same-Ref identities unique | A low-probability collision or deliberate duplicate can permanently deadlock the documented merge path or force an undocumented history rewrite | Define collision detection before first commit and a concurrency resolution, such as reissuing one unmerged record under a new Ref with an explicit non-supersession collision note; also define what to do if both colliding records already reached main +MAJOR | high | Rider (c), §2.4 squash merge, and parent-story AC 7 | Rider (c) says “the evidence entry” is copied on squash but never defines which entries are collected when the squash range contains several cycle-closing commits, several profiled stories, or repeated entries for one story; only exception blocks get range, identity, deduplication, and conflict rules | Squash can still erase or ambiguously duplicate the durable evidence that the rider exists to preserve | Specify the authoritative squash range and an all-entry selection, ordering, identity, duplicate, and conflict rule for evidence entries, or narrow the acceptance criterion to a single-entry case and park multi-entry squash explicitly +MAJOR | medium | §1.4 final sentence | After saying the git algorithm had five defects and no validation exercised it, the design concludes that “the machinery protected the record against accident” | No exact comparison or observed check supports that protection claim; the preceding evidence establishes only an attempted protection design with known defects | Rephrase as “the machinery was intended to protect against accident” or name the exact tested comparisons and bound the accident classes they actually covered +END OF FINDINGS (16 total) diff --git a/.context/codex-reviews/gate-a-spec-pass-5.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-5.pre-2026-08-14.md new file mode 100644 index 0000000..b8cce97 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-5.pre-2026-08-14.md @@ -0,0 +1,15 @@ +MAJOR | high | Sweep — §7 Check 2 assertion 1 | the sample assigns `N="$(unwrap "$F")"` but then passes `$N` to `grep` as a filename, while neither `unwrap` nor a temporary-copy contract is defined | as written the promised per-surface presence check cannot consume normalized content and cannot produce the required battery+check evidence | define a concrete side-effect-free normalizer that writes a temporary file, store that pathname in `N`, and clean it up, or pipe normalized content to each grep +MAJOR | high | Sweep — §7 Check 2 lines 377–378 | both `grep -cF` patterns begin with `-` and have no `-e`, so grep parses the example anchor as an option rather than a pattern | the two anchors added specifically to validate both format lines do not run portably and can make the validation fail for the wrong reason | pass each pattern with `-e`, for example `grep -cF -e '- · supersedes ' "$N"` +MAJOR | high | Sweep — §7 Check 2 assertion 1 | `grep -cF` counts matching lines, not occurrences, so two copies of an anchor on one wrap-normalized paragraph line still report `1` | the check does not establish its stated exactly-once cardinality and can miss a duplicate convention clause | count fixed-string matches rather than matching lines, using a portable occurrence-counting helper with an asserted numeric result +MAJOR | high | Sweep — §7 Check 2 lines 358–360 | the seven anchors do not support the claim that dropping any decision §2 settles removes one; removing the unknown-provenance default, wrong-when-written cause, recurrence effect, locator fallback, cumulative rule, cite-don't-restate rule, or union-merge repair leaves all seven | two identically incomplete surfaces can satisfy presence and parity while omitting a settled decision, which is the exact false-green state presence-first was added to close | either anchor every load-bearing §2 decision and assert each occurrence, or narrow the claim to the seven decisions actually covered and explicitly accept the remaining design-parity gap in §9 +MAJOR | high | Sweep — §7 Check 3 lines 417–430 | `L`, `C`, and `E` are assigned but the cardinality and `L < E < C` assertions exist only as comments, and the final grep's desired zero-match state exits nonzero | the section calls this a mechanical check that fails before and passes after, but the shown commands neither execute assertions 1–3 nor have a successful exit status for assertion 4 after the change | provide an executable shell check that validates each variable as one integer, performs both numeric comparisons, and explicitly asserts the template count equals zero +MAJOR | high | Sweep — §7 Check 3 lines 431–434 | a live supersession-shaped entry placed in the template without a label passes assertion 4, and Check 1's named read examines only the repo-ledger entry and its CLAUDE.md citation | the spec says this defect surfaces in the named read even though no validation step reads the template for stray entries, leaving an unacknowledged false green in the empty-template decision | add a template assertion/read that rejects live entry shapes outside the indented examples, or move this case to §9 as an accepted unvalidated layout defect +MINOR | high | Sweep — story §1 lines 11–13 | the story still states in the present tense that the header says never edit a row and that editing breaks that rule | after the declared story change lands this is the un-narrowed rule the amended sites audit was supposed to eliminate, so the authoritative story will contradict its own §1 accounting and AC 3 | qualify the passage as the pre-change rule or rewrite it to say a landed row cannot be edited +MAJOR | high | Read — §2.2 entry formats and locator rules | both locator fallbacks require a distinguishing text fragment, but neither example format defines where that fragment goes or, for multiple retired entries, which fragment belongs to which entry date | authors cannot execute the collision decisions unambiguously, including the live duplicate row pair and the same-day two-entry case that motivated the second format | extend the row and correcting-entry grammars with an explicit fragment position and a per-retired-entry date-plus-fragment form, then show worked collision examples +MAJOR | high | Read — §2.1 and Rider 2 self-test | the six required verdicts are internally consistent—case 1 landed → append supersession; case 2 own open cycle → amend; case 3 earlier completed cycle → landed and append; case 4 own post-WIP row → amend; case 5 pending resolution → append a resolving row; case 6 unknown authorship → landed and append—but the added pre-WIP sentence settles only a single eventual-WIP path and leaves a skipped Gate-B cycle with no snapshot-replacing close, an abandoned change with no close, and interleaved intended commits with no objective cycle identity | landed status can remain permanently or observer-dependently indeterminate on real no-snapshot and interrupted paths, moving the ambiguity from the first snapshot to intent about a future commit | define cycle close for trivial/N-A/no-snapshot commits and abandonment, and define how interleaved work is assigned to a cycle; otherwise default those states explicitly to landed +MAJOR | high | Read — §9 two-open-cycles residual | the claim that cycle A's amendment does not move date plus fingerprint is unsupported because §2.1 permits amending the open-cycle row in place without freezing either field | A can change the date or fingerprint after B writes an entry, making B's locator dangling rather than merely making a finding-fragment fallback stale, so the accepted race is not bounded and repairable as stated | either constrain open-cycle amendment once a row is shared, choose a locator A cannot change, or state the full dangling-locator residual and provide a recovery convention +MAJOR | high | Read — story §2 desired outcome versus settled AC 4 | the desired outcome still says a reader can tell whether any row describes current behaviour, while settled AC 4 and the design only expose claims that have been explicitly marked false and admit that nothing validates completeness | absence of a supersession entry cannot prove a row is current, so the story retains the stronger requirement the accepted amendment deliberately replaced | align the desired outcome with AC 4: determine which claims are recorded as no longer holding and where their current answer lives +MINOR | medium | Read — §4 change surface and `docs/coding-workflow.md:203` | the change surface omits existing prose that calls the ledger strictly append-only and says history is never rewritten | allowing an open-cycle row to be edited makes that unqualified description at least ambiguous and risks teaching the old absolute rule outside the two updated prompt surfaces | inspect every append-only claim under the standing falsification lens and qualify `docs/coding-workflow.md` with the landed boundary or record why its use of history excludes open-cycle rows +MAJOR | high | Read — §2.1 covered causes versus recurrence rule | “wrong when it was written” covers a row whose claimed hardening, rung, ref, date, or fingerprint was never true, but the same paragraph categorically says supersession never marks the hardening and the row always keeps counting | a nonexistent or misclassified hardening would remain in mechanical recurrence lineage and trigger unjustified escalation even though the convention presents it as corrected | restrict supersession to narration defects where a real hardening and its recurrence identity still stand, or specify separate handling for false hardening-identity claims and name it in the out-of-scope list +MINOR | high | Read — §2.2 correcting-entry scope | the correcting format carries one superseded-row locator but the prose allows it to retire “exactly the entries it names” without requiring all named entries to concern that same row | an author can attempt to retire entries from different rows in one line that cannot locate all of them, producing an internally ambiguous correction | state that one correcting entry may retire only entries against its single named row and require separate correcting entries for different rows +END OF FINDINGS (14 total) diff --git a/.context/codex-reviews/gate-a-spec-pass-6-dispositions.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-6-dispositions.pre-2026-08-14.md new file mode 100644 index 0000000..b91d6e2 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-6-dispositions.pre-2026-08-14.md @@ -0,0 +1,61 @@ +# Gate A — spec — pass 6 dispositions + +Artifact: `docs/superpowers/specs/2026-08-05-hardening-ledger-supersession-design.md` +16 findings: 11 Major, 4 Minor, 1 Nit. All sixteen read as correct. **None applied** — the pass +is not clean, and the pinned exit sends anything new to Daniel. + +Reviewer: `mcp-codex-dev@1.0.1`, `codex-cli 0.144.5`, model `gpt-5.6-sol`. Same as passes 4–5. +(Recorded here, not in the findings file: CLAUDE.md §5 admits no non-finding line there.) + +## Trajectory — worth reading before the table + +Passes 4, 5, 6 returned **11, 14, 16** findings. Every pass has been valid; every pass has been +applied in full; every pass has found more than the last. That is not the shape of a converging +review, and it is the fact Daniel needs most from this pass. + +The three failure modes are stable across all three passes: +1. **A fix executes in wording, not mechanism** (pass 3's entry locator → pass 4's §7 rewrite → + pass 5's shell → pass 6's finding 2, where assertion 1 became runnable and assertion 2 did + not). +2. **A fix overclaims what it validates** (findings 3, 12 here; 4, 6, 10 in pass 5). +3. **A fix falsifies a statement elsewhere that nobody re-reads** (findings 4, 5, 10 here). + +Four of pass 6's Majors (**7, 8, 9**, and part of **6**) land on the old-conditions accounting +table **written in pass 6 to satisfy the `AGENTS.md` Don't about old-conditions accounting**. +That is the sharpest signal in the cycle: the mechanism this project uses to prevent silent +condition-dropping was itself applied wrongly, in the very edit that introduced it. + +## Verified mechanically, not taken on the reviewer's word + +| # | Claim | Result | +|---|---|---| +| 1 | §2.0's quote elides source bytes behind `…` | **Confirmed.** `docs/hardening-log.md:7` reads `by appending a new row (same fingerprint,`. The `…` hides real text, not a wrap. | +| 2 | check 2's assertion 2 has no runnable command | **Confirmed.** Assertion 1 is executable shell; assertion 2 is prose from "Locate §4's two sentinels" onward. | +| 4 | §8's stored riders are stale | **Confirmed.** Rider 3 (line 544) still audits `§5` only, while pass 6 was run with a rider requiring both §5 and §2.1. Rider 1 still carries the anchor-coverage premise §7 now disclaims. | +| 5 | stale claims survive outside the change surface | **Confirmed.** `docs/superpowers/plans/2026-08-04-hardening-round-0-8-0-and-pr-21.md` lines 64, 474, 483 carry "Never edit an existing row", "Never edit a row \| **kept** — any solution must preserve it", and the pre-amendment desired outcome. | + +| # | Sev | Verdict | Reason | +|---|---|---|---| +| 1 | Nit | **Valid, confirmed** | Third consecutive pass finding a defect in this one quoted phrase. | +| 2 | Major | **Valid, confirmed** | The both-surfaces decision — the thing the whole two-surface scope rests on — has no reproducible check. Failure mode 1, again. | +| 3 | Major | **Valid** | The four-state claim still says a clause dropped from both surfaces makes an anchor report `0`. Deleting "once it merges it is landed, and stays landed", the date+fingerprint locator rule, the uniqueness stops, or the removed-hardening exclusion leaves all fourteen anchors intact — and §9's omission list names none of them. The narrowing was applied to one sentence and not to the two that depend on it. | +| 4 | Major | **Valid, confirmed** | §8 is titled "verbatim in every pass prompt" and is now behind the riders actually used. A later pass reading it would be steered back to a rejected coverage claim. | +| 5 | Minor | **Valid, confirmed** | Story §1 was pass 5 finding 7, applied only to §2; the plan file is a fourth site nobody swept. | +| 6 | Major | **Valid** | All six verdicts correct. But cases 3, a closed-cycle form of 4, and an own-row form of 6 all flip relative to the cycle rule, and §2.1's accounting names only the earlier-finished-branch flip. | +| 7 | Major | **Valid, and the most serious** | The table states the old condition as "a row you have not yet published", which the cycle rule never carried — it tested same-open-cycle ownership — then marks the materially broader not-yet-reachable rule **kept**. An old condition restated in the new rule's vocabulary cannot detect what the new rule widened. | +| 8 | Major | **Valid** | The table credits the old procedure with a cycle opening and a definite permanent exit. Pass 4 established it had no opening (pass 4 *added* one); pass 5 established the close was undefined for trivial, abandoned and interleaved work. The accounting invents guarantees, then calls the replacement a strengthening against them. | +| 9 | Major | **Valid** | §9 says the cycle rule carried the same single-writer assumption. It did not: the "landed to you" clause is exactly what stopped a second worktree amending the author's row. Reachability authorizes any holder to amend any unmerged row — a broader lost-update path, described as pre-existing. | +| 10 | Major | **Valid, confirmed** | The Scope line still says "Nothing else" while §4 now requires `docs/coding-workflow.md`, two story files, `todos.md`, the manifest and the changelog. | +| 11 | Major | **Valid** | "Contained in a commit reachable from `origin/main`" never says how a reader maps a ledger *row* to a *commit*. `git merge-base --is-ancestor` needs a commit argument the spec never supplies, and blame can return an amendment commit, several commits, or none for an uncommitted row. The decidability claim rests on a step that does not exist. | +| 12 | Major | **Valid** | "The repository always answers", "decidable in both directions" and "the one way the test misreads" are the `AGENTS.md` overclaim class. A missing or renamed `origin/main`, shallow history, rewritten remote history and command failure are all distinct from a stale ref, and several yield false *un*reachability — the breach direction. | +| 13 | Major | **Valid** | Two branches appending same-scope supersessions, or competing correcting entries retiring one predecessor with different cited answers, merge cleanly under `merge=union` into a state where neither retires the other and AC 4 cannot be satisfied. | +| 14 | Minor | **Valid** | Both fragment slots are shown as mandatory in the grammar and described as conditional in the prose; no escape rule for an embedded double quote. | +| 15 | Minor | **Valid** | `` semantics undefined — discovery, writing, or landing date — while date is half of every locator. §3.1's entry says 2026-08-05 and the change is landing later. | +| 16 | Minor | **Valid** | "Cannot share" their "what is false" text is false; two entries can be worded identically. The next paragraph already concedes free text guarantees nothing. | + +## What this needs from Daniel + +Not a fix list. The question is whether to keep applying passes on this artifact, or to cut its +scope. Findings 7, 8, 9 and 11 say the reachability replacement is not yet sound: its accounting +misdescribes what it replaced, and its central test has an undefined step. Findings 2, 3 and 4 +say the validation section is still ahead of its own mechanism. diff --git a/.context/codex-reviews/gate-a-spec-pass-6.md b/.context/codex-reviews/gate-a-spec-pass-6.md new file mode 100644 index 0000000..a707e57 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-6.md @@ -0,0 +1,13 @@ +MAJOR | high | story §3 AC rider (b), lines 224–231; design §6, lines 543–549 | Both surviving passages still say the findings file is permanent or may not be deleted, even though design §3 now states that `.context/` is ignored, retention is not required, and slot collisions have destroyed findings. | This restores the false premise that pass 5 explicitly removed and gives the parked recording mechanism a justification the settled decision rejects. | Amend the story disposition and the §6 backlog item to say only that the original token is visible in the findings file when the reader validates the pass; remove every permanence, retention, and strictly-better-evidence claim. +MAJOR | high | design §2.0, lines 131–140 | After establishing that PR #14 falls outside the form and is history rather than derivation, the section immediately says the form is scoped to what that precedent actually covers. | The two statements cannot both be true, and the latter reintroduces #14 as the normative derivation that pass 5 was meant to remove. | Replace the line with a self-grounded transition such as "So the form is scoped by the optional-work boundary below, and no further," without calling #14 a precedent. +MAJOR | high | design §2.0, lines 131–138; shipped §2.1 block, lines 185–189 | The text says that under Wait for `docs/pr-review-bots.md` requires a bot review and a recorded human decision, but that document says they are alternatives: a review permits the merge, while an absent required review may be passed only on an explicit recorded human decision; #14 had no review. | The proposed standalone CLAUDE.md paragraph misstates the repository's authoritative routing procedure and describes #14 using a conjunction its history disproves. | State the actual branch: under Wait for the review is required unless an explicit recorded human decision permits proceeding without it, and this new optional-work form is not that decision. +MAJOR | high | design §2.0, lines 140–150 | The purported load-bearing boundary is restated first as only excluding gate obligations and then as excluding only obligations "this section never imposed," whereas the shipped §2.1 block excludes anything required by any applicable source, including project docs, CI, branch policy, and the platform. | An implementer or later reviewer relying on §2.0 gets a materially weaker boundary that could admit mandatory non-gate work, despite the design calling that boundary the distinction between a record and a waiver. | Use the same broad "no applicable rule required" boundary throughout §2.0 and explicitly say that obligations originating anywhere else remain out of scope. +MAJOR | high | design §5.3, lines 503–515 | The table claims exact first and last source lines but supplies abbreviated ellipsis fragments, a conceptual "pinned line" reference, and a nonexistent fence opener containing `Human exception:`; the rider-(c) start `**On squash-merge**` appears only in the table, not in any specified shipped sentence. | The mandated extract-and-diff verification cannot resolve these anchors byte-for-byte, so AC 9 can fail mechanically or two implementers can select different blocks while claiming success. | Spell out every complete first and last line exactly as it will appear in both files, define the fenced block by its actual fence line plus adjacent content without pretending they are one line, and provide the full rider-(c) shipped sentence before pinning its anchors. +MAJOR | high | design §2.4, lines 282–299 and 301–306 | The claimed Ref-collision exit is not implementable under the record's immutability and carry rules: reissuing an already committed unmerged record under a fresh Ref leaves the old block reachable and therefore carried, the one-merged/one-unmerged state is not covered, and after both collide on main a future `supersedes ` cannot identify which record it corrects. | A normal merge, rebase, or fast-forward can preserve the duplicate identity indefinitely, after which deduplication and supersession are ambiguous despite the text claiming a defined exit. | Define collision handling for uncommitted, committed-but-unmerged, one-merged, and both-merged states; either permit an explicit pre-merge history rewrite under bounded conditions or introduce a durable disambiguated identity that carry and supersession both use. +MAJOR | medium | design §3 rider (c), lines 363–372 | Differing evidence entries for one story are resolved by "the later one in range order," but no traversal order is defined and a squash range can contain incomparable entries from merged or concurrent branch history. | Different collectors can retain different evidence, and an arbitrary traversal can silently discard the entry actually associated with the closing pass. | Define the exact ordering and ancestry rule; when entries are incomparable, require an explicit conflict that stops the squash until a human selects and records the authoritative entry. +MAJOR | high | shipped §2.1 block, lines 174–175; shipped §2.4 block, lines 263–269; design §2.4, lines 308–313 | The shipped text says nothing verifies the carry and then says nothing validates survival, while the detailed rule requires a human or agent to collect source blocks and compare them with the prospective body before merge. | The comparison does validate the prospective carry if performed; what is unverified is whether the procedure ran and whether the final merge preserved that body. Conflating those claims violates the gate-proof calibration rule and makes the pre-merge check's status ambiguous. | Say exactly that the prescribed source-to-prospective-body comparison validates the carry at that moment, but no mechanism enforces that the comparison ran or checks the post-merge result. +MAJOR | medium | design §2.1, lines 163–201; §2.4, lines 258–289 | The form is presented for optional checks, stood-down requested reviews, and courtesy steps generally, but the "full rule set" gives homes only for Gate-A and Gate-B cycles. It does not say where a legitimate decision attached to a prose-only, trivially skipped, or otherwise ungated commit lives. | A legitimate use explicitly allowed by §2.1 can have no cycle-closing spec, plan, or WIP commit, leaving the record placement undefined or making users invent a gate association. | Either narrow the form to optional decisions attached to an existing gate cycle or add an ordinary-commit row defining placement and squash/merge carry for ungated changes. +MINOR | high | design §1.2, lines 44–46; §6, lines 543–549 | The design calls prior review files under `.context/codex-reviews/` "archived" and tells a future owner to start from them, although `.gitignore` ignores `.context/*` and this same design records that slot collisions have destroyed such files. | The closure record and backlog point future readers at evidence that is not shipped with the repository and may vanish, creating orphaned references after clone or cleanup. | Treat those paths as ephemeral current-worktree aids only, or move the evidence that must remain auditable into a tracked documentation location and cite that instead. +MINOR | high | story §5 settled-question disposition, lines 296–300 | The story says design §5 keeps the tracked-debt row parked, but the tracked re-review debt is parked in design §6; §5 is validation. | The newly pinned cross-reference sends readers to the wrong decision and undermines the amendment's old-condition traceability. | Change `Design §5` to `Design §6`. +MINOR | high | design §2.3 first row, line 226 | The accounting labels the hard floor as "min 3 passes per gate," while current CLAUDE.md §5 requires three passes per run and Gate A contains two separate runs, each with its own loop. | The derivation table is supposed to map actual clauses; this shorthand can be read as three total passes across both Gate-A runs. | Say "min 3 passes per run; Gate A spec and plan are separate runs" to match the current procedure exactly. +END OF FINDINGS (12 total) diff --git a/.context/codex-reviews/gate-a-spec-pass-6.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-6.pre-2026-08-14.md new file mode 100644 index 0000000..b4d02df --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-6.pre-2026-08-14.md @@ -0,0 +1,17 @@ +NIT | high | SWEEP §2.0 "The amended phrase" | the purported existing byte quote says `by appending…`, but the ledger source says `by appending a new row`; the ellipsis replaces source bytes instead of merely hiding a hard wrap | the rider explicitly requires byte-for-byte quoted passages, and this quote still cannot be located as written after pass 5's delimiter fix | quote the complete normalized source text or label the excerpt as an elision rather than a byte quote +MAJOR | high | SWEEP §7 Check 2 assertion 2 | assertion 1 is runnable under both `sh` and `dash`, but assertion 2 provides no extraction, sentinel-cardinality, normalization, or comparison command at all | the only executable check proves anchor presence; the parity decision that both surfaces are identical modulo wrapping remains an instruction to improvise, so the spec's central both-surfaces evidence can be claimed without a reproducible check | add a runnable POSIX-shell parity block that rejects zero or multiple sentinels, extracts both inclusive regions, applies the stated normalizer, and compares them +MAJOR | high | SWEEP §7 Check 2 and §9 anchor coverage | all fourteen anchors occur exactly once in the proposed shared text, have no anchor-on-anchor substring collision, and the quoting parses under `sh` and `dash`, but the four-state claim still says a clause dropped from both surfaces makes an anchor report zero; deleting `once it merges it is landed, and stays landed`, the date-plus-fingerprint locator rule, the uniqueness stops, the mechanical-identity effects, or the removed-hardening exclusion leaves all fourteen anchors intact, and §9's supposedly explicit omission list names none of these | two identically incomplete surfaces can satisfy presence and parity, while the validation prose and residual-risk enumeration say that state is distinguished | narrow the four-state claim to deletion of an anchored clause and make §9's unanchored-decision list complete, or add anchors for every layout/behavior decision the check claims to protect +MAJOR | high | SWEEP §8 "Gate-A riders, verbatim in every pass" | the stored rider still asks whether dropping any §2 decision removes an anchor—the premise §7 now explicitly disclaims—and its old-conditions rider audits only §5, while the authoritative pass-6 rider requires both §5 and §2.1 tables | the artifact's promised reusable review protocol is already stale relative to the settled decisions and would steer a later pass back to a rejected coverage claim while omitting the new boundary audit | replace §8 with the current riders verbatim, including the accurate anchor/§9 test and both old-condition tables +MINOR | high | SWEEP governing story §1 and `docs/superpowers/plans/2026-08-04-hardening-round-0-8-0-and-pr-21.md:64,474,483` | live-sounding statements still say never edit any row, preserve that absolute rule, leave the both-surfaces question unresolved, or promise that any row can be identified as current; the story §1 site is in the declared surface and was pass 5 finding 7, while the plan contains the other exact stale claims | after implementation, the standing falsification sweep still finds the un-narrowed rule and superseded requirement, contrary to the requested no-survivors check | reframe the story problem in explicit pre-change tense and either mark the plan passages as historical snapshots or account for why archived artifacts are intentionally exempt from the sweep +MAJOR | high | SELF-TEST §2.1 cases 1–6 | verdicts are: 1 the 2026-07-20 row is landed and gets §3.1's entry; 2 row D was unreachable during PR #22 and was amendable then, but is landed now; 3 finished work on an unmerged branch is amendable; 4 a pushed but unmerged current-branch row is amendable; 5 resolving an unmerged `pending` row appends; 6 indeterminate reachability defaults to landed and appends; cases 3 and a cycle-closed form of case 4 differ from the cycle rule, and an own-current-cycle row in case 6 also flips from amend to append, but the accounting names only the earlier-finished-branch case | the new rule decides the six cases, yet its migration record does not disclose every changed verdict the old procedure decided | add the pushed/closed-cycle and known-own-row-with-unusable-remote transitions to the old-conditions accounting, with their accepted costs +MAJOR | high | OLD-CONDITIONS §2.1 row "A row you have not yet published" | the replaced rule did not test publication; it allowed amendment only for a row attributed to the currently open cycle, including after a `WIP:` snapshot could be committed or pushed, and treated an earlier completed but still unpublished cycle as landed | the table rewrites the old condition into one it never carried and then labels the materially broader not-yet-reachable rule as kept | state the old same-open-cycle ownership condition verbatim and mark the expansion to every unmerged row as deliberately changed, not kept +MAJOR | high | OLD-CONDITIONS §2.1 rows on permanent exit and cycle lifecycle | the table says the old procedure carried a cycle identity, opening, and close and had a definite permanent exit, but the immediately replaced text defined identity and one `WIP:` replacement close only; pass 4 had already established that it never defined an opening and left trivial, abandoned, and interleaved cycles without a close | the accounting invents old guarantees and calls the new merge boundary a strengthening, silently dropping the exact lifecycle defects that motivated the wholesale replacement | record identity as kept history, missing opening and incomplete close as old defects eliminated by removing cycle state, and avoid claiming those absent conditions were carried +MAJOR | high | OLD-CONDITIONS §2.1 and §9 concurrent-writer disposition | §9 says the cycle rule had the same single-writer assumption because its worktree clause only said whose rows you may not touch, but that ownership prohibition is precisely what stopped a second worktree from amending the author's row; the new reachability rule instead authorizes any holder to amend every unmerged row | the replacement introduces a broader lost-update path while describing it as an existing assumption, so the old concurrency condition is neither correctly preserved nor deliberately dropped | account for the observer-relative ownership barrier as dropped, distinguish the old stale-supersession race from the new multi-writer amendment race, and state why the new risk is accepted +MAJOR | high | READ top-level `Scope` versus §4 change surface | the scope says the deliverable is only two ledger-header surfaces and one skill phrase, “Nothing else,” while §4 also requires behavioral prose in `docs/coding-workflow.md`, two story edits, backlog state, manifest version, and changelog content | an implementer following the scope boundary can legitimately omit files that the same design later makes required, violating prompt-standard clarity and risking invariant-12/version-release drift | distinguish behavioral scope from required file/change surface, or expand the scope sentence to include every required supporting artifact +MAJOR | high | READ §2.1 reachability mechanism | “contained in a commit reachable from `origin/main`” does not define how a reader identifies the commit that contains the logical row, yet the rationale says the test is one `git merge-base --is-ancestor` away; that command needs a commit argument, and blame/search can return an unmerged amendment commit, multiple historical versions, or no commit for an uncommitted row | two readers can apply different row-to-commit mappings and take opposite edit-versus-supersede actions at the invariant-11 prompt boundary | specify a deterministic row identity and commit-selection procedure, including uncommitted rows, amended row text, duplicate date/fingerprint rows, and multiple candidate commits, then state the exact reachability command +MAJOR | high | READ §2.1 and §9 reachability claims | “the repository always answers,” “decidable in both directions,” and “the one way the test misreads” overstate the mechanism; a missing or renamed `origin/main`, shallow/partial history, unavailable candidate commit, rewritten remote history, and command errors are distinct from a merely stale ref, and some yield false non-reachability rather than a recognizable unknown unless the procedure distinguishes exit 1 from operational failure | the breach direction is not bounded to the single axis §9 enumerates, conflicting with the AGENTS.md rule against exhaustive-sounding gate claims | enumerate the checked axes and all known non-exhaustive failure classes, define which command outcomes mean unreachable versus unknown, and route every operationally indeterminate result to landed +MAJOR | high | READ §2.2 union-merge and correcting-entry concurrency | the convention repairs duplicate labels but gives no result for two branches concurrently appending same-scope supersessions or competing correcting entries that retire the same predecessor and cite different current answers; union merge keeps both, neither entry retires the other, and cumulative-not-latest-wins leaves both active | the ledger can merge cleanly into a state where a reader cannot tell which correction governs, despite AC 4 requiring the current-answer location to be determinable | define duplicate/competing-entry reconciliation and conflict precedence, or make same-target concurrent entries a stop requiring one explicit correcting entry after merge +MINOR | high | READ §2.2 entry-locator grammar | the correcting format shows an `""` after every retired-entry date, while prose makes that fragment conditional on a date-plus-row collision and never explicitly says to omit it when unique; neither row nor entry fragments define how embedded double quotes are escaped | authors can produce different line shapes or ambiguous quoted locators from the same rule, especially when the only distinguishing finding text itself contains quotes | mark both fragment slots syntactically optional, add the entry-fragment omission rule explicitly, and define a quote/escape convention or require a quote-free distinguishing substring +MINOR | medium | READ §2.2 format and §3.1 example | the supersession entry's leading `` is never defined as discovery date, writing date, landing date, or correction-effective date; the first entry uses 2026-08-05 even though the reviewed change is being finalized later | date is part of every entry locator, so inconsistent author choices create avoidable collisions and unclear chronology | define the date semantics and YYYY-MM-DD basis for supersession and correcting entries, then justify or update the first entry's date +MINOR | high | READ §2.2 entry-locator rationale | the rationale says two independently scoped entries “cannot share” their `what is false` text, but duplicate wording and same-claim concurrent entries can share it; the next paragraph correctly concedes that free-text fragments do not guarantee uniqueness | the false impossibility claim hides a live stop path and contradicts the design's own fallback limitation | replace “cannot share” with the weaker intent that this field is more likely to distinguish entries, and explicitly include identical false-text entries in the unresolved-locator stop +END OF FINDINGS (16 total) diff --git a/.context/codex-reviews/gate-a-spec-pass-7-dispositions.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-7-dispositions.pre-2026-08-14.md new file mode 100644 index 0000000..38f685b --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-7-dispositions.pre-2026-08-14.md @@ -0,0 +1,56 @@ +# Gate A — spec — pass 7 dispositions + +Artifact: `docs/superpowers/specs/2026-08-05-hardening-ledger-supersession-design.md` +15 findings: **1 Blocker**, 9 Major, 4 Minor, 1 Nit. All read as correct. **None applied.** + +Reviewer: `mcp-codex-dev@1.0.1`, `codex-cli 0.144.5`, model `gpt-5.6-sol`. Same as passes 4–6. + +## The blocker + +**Finding 9.** §2.1 says *"If that file cannot be read at that ref, treat the row as landed."* +§9 says every non-confident outcome — including **a stale ref** — "is *landed*, **by the rule in +§2.1**". §2.1 carries no such rule: a stale-but-**readable** `origin/main` yields a confident +*absent* for a row that is published on the real remote, and the convention then authorizes +editing it. That is the append-only breach the whole design exists to prevent. + +Verified: `docs/…-design.md:58` covers only *cannot be read*; `:653–658` claims stale resolves +to landed by that rule. The two disagree, and the safe reading is not the one §2.1 states. + +This is the **third instantiation of one defect**, each time in the sentence written to fix the +previous one: pass 6 finding 12 killed "the repository always answers"; the replacement narrowed +to "confident present / confident absent"; and *confident absent* is itself unsound when the ref +is stale. The undefined step keeps moving instead of closing. + +## Verified mechanically + +| # | Claim | Result | +|---|---|---| +| 9 | §2.1 and §9 disagree on a stale readable ref | **Confirmed**, lines 58 and 653–658. | +| 11 | §3.1 still teaches the cut model | **Confirmed**, line 257: "The entry is falsification-scoped, not row-scoped". The cut executed in §2.2 and not in the worked example — the exact wording-not-mechanism failure this pass was told to hunt. | +| 6 | check 2 dirties the worktree | **Confirmed**, lines 497–499 write `region.*.txt` into the repo root with no cleanup. | + +## Table + +| # | Sev | Verdict | Note | +|---|---|---|---| +| 9 | **Blocker** | **Valid, confirmed** | See above. | +| 1 | Major | **Valid** | §2.2 defines `` as the day the entry is written; §3.1 and all three checks hard-code `2026-08-05`, and implementation is later. Either the first entry violates the convention on day one, or fixing the date breaks the checks. | +| 3 | Major | **Valid** | The re-derived table omits three of the old rule's conditions — marks text never the hardening, keeps fingerprint and counting, later removal out of scope. The first is *materially changed* by the phantom-hardening case, so the omission hides a widening. Same class as pass 6's findings 7–9, on the table rewritten to fix them. | +| 4 | Major | **Valid** | §9's "complete" unanchored list is not complete: amendment of an absent row, one-way landedness, the date+fingerprint locator, the `YYYY-MM-DD` requirement, phantom-hardening, and fingerprint/counting preservation are in neither the 16 anchors nor the list. | +| 10 | Major | **Valid** | "Unknown authorship" was not widened into "ref unreadable" — the triggers are incomparable, and a *readable absent* ref flips unknown-provenance rows from landed to amendable. The changed-verdict table claiming completeness omits that case and two others. | +| 11 | Major | **Valid, confirmed** | See above. | +| 12 | Major | **Valid** | The "cover the row as it now stands" duty lives only in design rationale, outside the prose copied to both surfaces. A downstream author following the shipped convention can write a narrow second marker and make an earlier still-false claim non-governing — breaking AC 4. | +| 13 | Major | **Valid** | Supersession entries target *landed* rows, so they sit outside the unpublished-region single-writer assumption entirely; the concurrency rationale invokes an assumption that does not cover the case. And §9 still says the old rule carried the same assumption, which §2.1 now correctly says it did not — the two contradict inside one document. | +| 14 | Major | **Valid** | `merge=union` does not "overwrite"; it keeps both edited variants, producing two rows with one locator — which corrupts recurrence counts, a worse and different failure than the one §9 describes. | +| 15 | Major | **Valid** | "Damage cannot reach merged history without passing through the PR that merges it" — nothing requires publication through a PR, and `AGENTS.md` already records direct-push bypass for invariant 12. Enforcement-overclaim class, again. | +| 2 | Minor | **Valid** | Table row 1's quote inserts a `…` not present at `5e295f0` while the prose claims each row carries the old rule's own sentence verbatim. | +| 5 | Minor | **Valid** | Assertion 2 uses `grep -cF` — lines, not occurrences, and substring not line-shape — so it does not prove the sentinel cardinality it claims. The same defect pass 5 fixed in assertion 1, reintroduced in assertion 2. | +| 6 | Minor | **Valid, confirmed** | Worktree dirtied by the validation itself. | +| 7 | Minor | **Valid** | Template label check is column-zero exact; a leading-space label passes as "no label". | +| 8 | Nit | **Valid** | "Four assertions on line numbers" — assertion 4 produces no line number. | + +## Recommendation recorded with the pass + +Passes 4–7 returned **11, 14, 16, 15**. Four consecutive passes, three stable failure modes, and +the same defect class re-instantiated three times inside its own fixes. This is not a list to +grind down; the loop is not converging and the next pass should not be more of the same. diff --git a/.context/codex-reviews/gate-a-spec-pass-7.md b/.context/codex-reviews/gate-a-spec-pass-7.md new file mode 100644 index 0000000..c10f30e --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-7.md @@ -0,0 +1,17 @@ +MAJOR | high | §2.0, lines 154–157 | The example says a Gate-A spec commit may record an absent supplementary bot review, but the settled routing leaves no such eligible absence: a Wait-for absence is mandatory and must use that rule's own recorded-decision branch, while an opportunistic absence explicitly needs no record. | The example contradicts both §2.1 and docs/pr-review-bots.md and reintroduces the exact bot-review ambiguity the scope correction is meant to close. | Replace the example with a genuinely optional action that was consciously planned and then stood down, or delete it. +MAJOR | high | §2.4 shipped placement block, lines 271–279 | The shipped rule says every record lives in the commit that closes the cycle it belongs to, but the table adds records on ungated commits, which have no cycle and therefore no closing cycle commit. | A CLAUDE.md-only reader cannot place one of the design's expressly legitimate uses and may invent a gate association or omit the record. | State the ungated case in the shipped block first, then separately state placement for Gate-A and Gate-B cycles. +MAJOR | high | §2.4 shipped placement block, lines 277–279 | The claim that comparing collected blocks with the prospective body establishes that the carry is correct outruns the comparison: it proves only that the prospective body matches the blocks the operator selected, not that every in-range block was found or that the merge used the inspected body. | A missed source block still produces a successful comparison and a false assurance, violating the gate-proof calibration rule. | Say exactly that the comparison establishes equality with the collected set at that moment, and separately require and calibrate the range-completeness check. +MAJOR | high | §2.4 accumulation rules, lines 292–299 | Exception records must remain in the order recorded, but squash ranges may contain merged side branches with incomparable records and no total recorded order; unlike evidence entries, exception records get no conflict or deterministic ordering rule. | Two compliant squash preparers can produce different bodies from the same range, undermining the claimed repeatability of the identity rules. | Define a deterministic order for incomparable records, such as ancestry order with a stable Ref-based tie-break, and state that ordering carries no semantic precedence. +MAJOR | high | §2.4 correction and supersession rules, lines 295 and 312–324 | The reader is told to take the superseding record, but the model does not handle two records superseding the same target, cycles, a missing target, self-supersession, or chains; after a Ref collision it also requires a commit SHA without defining the syntax inside Accepted because. | The correction mechanism can yield several equally valid authoritative records or a reference no reader can resolve consistently. | Define target grammar including the collision-qualified SHA, require an existing older distinct target, reject cycles, and add a conflict rule for multiple superseders. +MAJOR | high | §3 rider (c), lines 377–399 | The shipped sentence requires copying every evidence entry, but the following algorithm deliberately drops earlier differing entries for the same story and carries only one. | A CLAUDE.md-only reader will preserve entries that the design tells an implementer to delete, and deleting evidence contradicts the byte-for-byte instruction the parity check anchors. | Either make the shipped sentence state the exact per-story reduction rule or preserve every entry and remove the reduction. +MAJOR | high | §3 evidence-entry selection, lines 384–397 | Later is defined as later in first-parent traversal from base to head, yet commits on merged side branches are not members of that traversal; the later text then treats first-parent reachability as a partial order without defining how mainline-to-side and side-to-side cases are classified. | The same merge graph can be read as ordered or incomparable by different operators, changing which evidence survives or whether squash stops. | Define one precise graph relation over every commit in the range, with examples for mainline versus side branch, two commits on one side branch, and two distinct side branches. +MAJOR | high | §3 canonical squash sentence, line 379 | The rationale that main's tip is the durable record and nothing else survives is categorical and false: the squash commit will cease to be the tip, and original commits may survive in PR refs, remote branches, reflogs, or clones even though they are not reachable from main. | This is another mechanism overclaim and obscures the actual property the carry rule needs: reachability from main history. | Say that the squash commit body is the durable record reachable from main and that the original range's commit bodies are not carried into main history by the squash. +MAJOR | high | §5.3 normalization rule, lines 539–545 | The design says the mirror is uniformly indented and therefore cannot share line breaks, but the working-tree template at workflow-init.md lines 192–629 is flush-left inside its four-backtick fence. The proposed whitespace collapse also erases paragraph boundaries, which are meaningful Markdown. | The parity procedure is justified by a false repository fact and can accept prompt copies that render as different paragraph or block structures. | Remove the indentation premise and compare paragraph structure and prose content with a Markdown-aware or paragraph-preserving normalization; permit wrapping differences only within the same paragraph. +MINOR | high | §5.3 placement-block anchor, line 534 | The first anchor is declared to be a sentence but stops at belongs to, before the em dash and the rest of the actual first sentence. | The extraction rule violates its own sentence-anchor contract and leaves boundary selection dependent on substring behavior. | Pin the complete first sentence through like the evidence entry. +MAJOR | high | §4 scripts/check-invariants.test.sh row and §5.2 | Adding check 4c makes every current fixture repository invalid because init_prompt_fixtures creates no CLAUDE.md and gives workflow-init.md no CLAUDE template section or canonical line, but the design never requires updating that shared initializer and every fixture builder that depends on it. | An implementation following only the named reject/accept additions will turn the entire existing suite red for an unrelated baseline failure. | Add an explicit requirement to extend init_prompt_fixtures with valid bounded §5 regions in both files before adding 4c-specific mutations. +MINOR | high | §5.2 boundary validation, lines 483–497 | The assertion rejects occurrences anywhere outside the region, but the only required boundary fixture puts content after each end boundary; nothing discriminates the start-boundary logic or an occurrence before §5. | A checker that scans from file start can pass the specified tests while violating the stated region contract. | Add before-start fixtures for both files, once with no in-region line and once with the valid in-region line, and require the expected outside-region diagnostic. +MINOR | medium | §3 rider (b) and existing CLAUDE.md Mechanics line 395 | The new canonical writer tokens are uppercase, while the unchanged Mechanics severity rule spells Blocker, Major, Minor and Nit in title case; §4 and §5.3 do not account for reconciling that nearby writer-facing vocabulary. | A literal model can emit the Mechanics spelling, which the tolerant reader then upgrades to MAJOR, turning intended Minor or Nit findings into iteration-triggering findings. | Change the Mechanics labels in both copies to the canonical uppercase spellings and add that edited sentence to the parity anchors and site accounting. +MAJOR | medium | §2.4 Ref collision table, lines 307–317 | When both colliding records are pushed but unmerged, the rule says to treat them as the next row even though that row is defined for merged records; it does not say which branch carries the follow-up, when it is written, or how both future merges retain the qualifier. | Concurrent branches can each remain locally conforming while main receives an ambiguous Ref without the promised collision record or receives two inconsistent collision records. | Add a distinct both-shared-unmerged state with a named coordinating commit or branch, required ordering before either merge, and an exact collision-record form. +MINOR | medium | §6 Enforced waiver authority backlog item, lines 567–572 | The trigger is disputed authorship of an optional-work record, but the parked work is a sanctioned zero-pass waiver requiring an availability attestation; those are different problems and the trigger does not establish renewed need for a gate waiver. | The backlog conflates attribution hardening for the shipped form with reopening the rejected availability design, making the narrow record form look like latent waiver machinery. | Split this into optional-record attribution and external-authority zero-pass research, each with its own trigger and requirements. +NIT | high | reviewer-availability story title, line 1 | Closed, with a recorded human exception reads naturally as though the story or gate was closed by an exception, while the settled result is that no exception may close a gate and only an optional-work record form survives. | The first line can recreate the central misunderstanding before the closure banner corrects it. | Rename it to state the negative closure and optional-work record, matching the design title's distinction. +END OF FINDINGS (16 total) diff --git a/.context/codex-reviews/gate-a-spec-pass-7.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-7.pre-2026-08-14.md new file mode 100644 index 0000000..d9a0eca --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-7.pre-2026-08-14.md @@ -0,0 +1,16 @@ +MAJOR | high | Sweep — §2.2 and §3.1 entry date | the convention defines `` as the day the entry is written, but the not-yet-landed entry and all three checks hard-code 2026-08-05 even though implementation is occurring after that date | the first entry would violate the convention on the day it is added, or an implementer who corrects the date would make the specified checks fail | use the actual implementation date consistently in the entry and checks, or parameterize the checks from the chosen entry date +MINOR | high | Sweep — §2.1 old-conditions table row 1 | the claimed source quote inserts a Unicode ellipsis that is not present at 5e295f0, yet the disposition calls it kept verbatim and the surrounding prose says every row carries the old rule's own sentence | the audit cannot serve as byte-accurate evidence for the replaced procedure | quote the complete source sentence exactly, with wrapping normalized only, and remove the false verbatim claim if an elision is retained +MAJOR | high | Sweep — §2.1 old-conditions table | the table omits the old rule's final conditions that supersession marks text but never the hardening, preserves fingerprint and recurrence counting, and leaves later removal out of scope | the first omitted condition is materially changed by the new phantom-hardening case, so the audit hides a semantic widening instead of accounting for it as AGENTS.md requires | add all three old conditions with kept, changed, or dropped dispositions, explicitly reconciling “never its hardening” with entries that say the hardening never existed +MAJOR | high | Sweep — §9 unanchored-decision inventory | the purported complete list omits multiple shared-region decisions with no deletion-sensitive anchor, including amendment of an absent row, one-way landedness, the date-plus-fingerprint locator, the YYYY-MM-DD requirement, the phantom-hardening case, and preservation of fingerprint and recurrence counting | both surfaces can lose any of these clauses identically while all sixteen anchors and parity still pass, contradicting the four-state coverage account | add anchors for each decision or enumerate every omitted decision accurately in §9 and narrow the validation claim accordingly +MINOR | high | Sweep — §7 check 2 assertion 2 | `grep -cF` counts matching lines and uses substring matching, so it does not prove either sentinel occurs exactly once and does not enforce the stated exact start line or end-of-line shape | two occurrences on one line or trailing text after the end sentinel can pass while the check claims exact sentinel cardinality and region boundaries | count occurrences rather than lines and use line-shape checks matching §4, such as exact start-line and end-anchored end-sentinel assertions +MINOR | high | Sweep — §7 check 2 assertion 2 | the parity command writes `region.hardening-log.md.txt` and `region.workflow-init.md.txt` in the repository root and never removes them | running the specified validation dirties the worktree and leaves artifacts outside §4's authoritative change surface | write both files under a `mktemp -d` directory with a cleanup trap, or compare streams without persistent repo-root files +MINOR | medium | Sweep — §7 check 3 assertion 4 | the template-label check rejects only the exact column-zero line `**Superseded rows:**`; a Markdown-equivalent label with leading spaces or trailing whitespace passes | the check can report that no label leaked while the scaffolded empty template visibly contains one, and §9 records only the different live-entry-without-label gap | normalize or match permitted Markdown whitespace around the label, and state any deliberately accepted variants in §9 +NIT | high | Sweep — §7 check 3 mechanical description | the text calls these “four assertions on line numbers,” but assertion 4 is a content-cardinality check on the template and produces no line number | the inaccurate description makes the check harder to audit against its actual mechanism | describe them as four layout and cardinality assertions, three of which compare line numbers +BLOCKER | high | Read — §2.1 six-case self-test | verdicts are: 1 the 2026-07-20 row is landed and gets an entry; 2 row D was amendable while absent and is now landed so a future correction gets an entry; 3 finished but unmerged work is amendable; 4 a pushed but unmerged current-branch row is amendable; 5 resolving an unpublished `pending` row still appends; 6 an unreadable ref is landed, but a known-stale readable ref has no consistent verdict because literal absence says amendable while §8 and §9 say landed; rewritten readable history likewise defeats “stays landed” | the unresolved sixth case can authorize editing a row already published on the real remote, directly breaching the narrowed append-only rule, and proves the undefined step moved into freshness and history rather than disappearing | make freshness or trust an explicit precondition for an absent verdict and define conservative behavior for stale or rewritten readable refs, then narrow or mechanize the one-way “stays landed” claim +MAJOR | high | Read — §2.1 changed-verdict and old-condition tables | “unknown authorship” is not widened into “ref unreadable”; the triggers are incomparable, and a readable absent ref flips unknown-author rows from landed to amendable; the exhaustive verdict table also omits that case, the no-cycle unpublished case, and an open-cycle row whose pair collides with a published row | the old-procedure audit overstates preservation and the table claiming every changed verdict is incomplete | mark the unknown-authorship protection as dropped or replaced, and add every omitted old-to-new verdict including no-cycle, unknown-provenance, and pair-collision states +MAJOR | high | Read — §3.1 first-entry explanation | it still says “The entry is falsification-scoped, not row-scoped,” even though §2.2 says the falsification-scoped and cumulative model is cut and entries are row-markers | the primary worked example teaches the removed model and can cause later entries to be interpreted cumulatively | describe the line as a row-marker whose content identifies the currently false claim, without saying it is not row-scoped +MAJOR | high | Read — §2.2 shared convention and §9 latest-wins residual | the requirement for a later marker to cover the row as it now stands exists only in design rationale outside the prose copied to the ledger and template, while the shipped convention merely says the last entry governs | a downstream author can follow every shipped instruction, write a narrow second marker, and make an earlier still-false claim non-governing, violating story AC 4's requirement to determine which claims no longer hold | add an explicit consolidation requirement to the shared convention prose and give it validation coverage or list it as unanchored +MAJOR | high | Read — §2.2 concurrency and §9 single-writer assumption | the spec says the unpublished-region single-writer assumption makes concurrent entries for one row rare, but supersession entries can only target landed rows and therefore sit outside that assumption; §9 also says the old rule carried the same assumption after §2.1 correctly says its worktree clause prevented the second writer | the concurrency rationale executes in wording only and the historical accounting contradicts itself | separate the unpublished-row editing assumption from a new landed-row entry-writer assumption, and state accurately that the latter is new and unenforced +MAJOR | high | Read — §9 concurrent unpublished-row failure | the stated failure is that two writers can “overwrite each other,” but `docs/hardening-log.md` uses `merge=union`, whose conflict resolution retains both edited replacements as duplicate row variants | the actual failure silently creates two rows with the same locator, corrupts recurrence counts, and forces fragment disambiguation rather than losing one writer's text | describe and handle the union-produced duplicate-row state, including how it is detected and reconciled before merge +MAJOR | high | Read — §9 single-writer consequence | the claim that the merge boundary guarantees damage cannot reach merged history without passing through the PR is false because neither Git nor this design requires publication through a PR, and AGENTS.md already records direct-push bypasses elsewhere | reviewers may rely on a review barrier that does not exist, repeating the repository's enforcement-overclaim class | qualify the statement as the normal PR workflow and say nothing enforces it, or add an actual branch-protection premise and verification +END OF FINDINGS (15 total) diff --git a/.context/codex-reviews/gate-a-spec-pass-8-dispositions.md b/.context/codex-reviews/gate-a-spec-pass-8-dispositions.md new file mode 100644 index 0000000..bfc5814 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-8-dispositions.md @@ -0,0 +1,85 @@ +# Gate A — spec — pass 8 dispositions (salvage cycle) — **TERMINATION ASSESSMENT** + +17 findings (0 BLOCKER, 10 MAJOR, 7 MINOR). **None dismissed. None applied** — Daniel's +pre-set condition fired: *"if pass 8 still oscillates on fix-of-fix findings with the Ref +surface gone, stop for a termination assessment instead of a pass 9."* + +## The data + +| Pass | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | +|---|---|---|---|---|---|---|---|---| +| Findings | 15 | 25 | 23 | 20 | 16 | 12 | 16 | 17 | +| Blockers | 3 | 2 | 5 | 4 | **0** | **0** | **0** | **0** | + +144 findings this cycle; 447 across the story's four cycles and seventeen passes. + +**Two things are true at once, and both matter.** + +**Blockers are genuinely gone — four consecutive passes.** Nothing in passes 5–8 says the +mechanism is unsafe, unbuildable, or a gate waiver in disguise. That is a real difference from +the tier-3 cycle, which died on blockers that said exactly those things. The *shipped +behaviour* has been stable and sound since pass 5. + +**The finding rate is flat at ~15/pass and has not converged for four passes.** Cutting the +drift record helped (20 → 16). Cutting `Ref:` did not: 16 → 17, and **six of the seventeen — +findings 2, 3, 4, 5, 6, 7 — are direct consequences of that cut**, including finding 2, which +correctly observes that "git identifies a record by its commit" is false under the design's own +rules: one commit may carry several records, and squash deliberately merges records from many +commits into one. So the justification for the cut was itself an overclaim. + +## The diagnosis + +**Where pass 8's findings actually live:** + +| Area | Findings | Count | +|---|---|---| +| §2.4 record identity, carry, restoration | 2, 3, 4, 5, 6 | 5 | +| §3 rider (c)'s evidence-entry ordering algorithm | 7, 8, 16 | 3 | +| §5.3 the parity extract-and-diff procedure | 9, 10, 11 | 3 | +| §5.2 checker region detail | 13, 14 | 2 | +| §2.1 examples and placement | 1, 17 | 2 | +| §2.3 accounting | 12 | 1 | +| tier-2 story | 15 | 1 | + +**Eleven of seventeen are in three sections — and those three sections specify procedures +nobody will mechanically execute.** §5.3 tells a person how to extract anchor-delimited +blocks from two files and normalize Markdown paragraphs before diffing; in practice a person +edits both copies and diffs them. §3's algorithm defines an ancestor relation over a squash +range to pick among evidence entries; §2.4 defines record identity, duplicate detection and +restoration matching. None of it is executed by any tool. All of it is prose specifying prose. + +**What actually ships is about forty lines:** one exception paragraph, one placement +paragraph, one canonical severity line, one reader rule, one carry sentence, and one shell +assertion. The design is **654 lines**. The finding rate is the specification layer being +reviewed, not the change. + +**Why both cuts behaved differently.** The drift record was a *requirement* — cutting it +removed obligations. `Ref:` was an *answer to a question the design had already asked* +("which record is this?"), so cutting it left the question standing and the answers dangling. +That is the general shape: this document keeps asking mechanical questions about an artifact +that is read by humans, and every answer generates its own findings. + +## Options + +**A — Cut the specification layer, keep the closure record and the shipped text.** §1 (which +has been stable for passes and is the story's actual deliverable), §2's shipped paragraphs +verbatim, §3's riders verbatim, §4's file list, §5's one assertion plus battery, §6's backlog. +Delete §5.3's procedure (→ "edit both copies, diff them, record that you did"), §2.4's +identity and carry algorithms (→ "copy the records across; nothing checks that you did"), and +§3's evidence-entry reduction (→ leave §5's existing carry rule alone). Roughly 654 → ~250 +lines, and the eleven algorithm findings stop existing rather than being fixed. One confirming +pass. + +**B — Keep passing.** Apply all 17, run pass 9. On four passes of evidence, expect ~15 more. + +**C — Close on dispositions.** The pinned exit permits it. Several findings are real (1, 3, +12) and would ship into planning as known gaps. + +**D — Ship only §1.** Drop the record form entirely; keep the closure record, which is the +finding the story was for. The riders (b) and (c) ship separately as the small prose changes +they always were. + +## Status + +Stopped for assessment, per instruction. Nothing applied from pass 8. Working tree holds the +post-pass-7 artifacts. diff --git a/.context/codex-reviews/gate-a-spec-pass-8-dispositions.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-8-dispositions.pre-2026-08-14.md new file mode 100644 index 0000000..8d8b243 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-8-dispositions.pre-2026-08-14.md @@ -0,0 +1,51 @@ +# Gate A — spec — pass 8 dispositions + +Artifact: `docs/superpowers/specs/2026-08-05-hardening-ledger-supersession-design.md` +9 findings: **1 Blocker**, 6 Major, 1 Minor, 1 Nit. All read as correct. **None applied.** + +Reviewer: `mcp-codex-dev@1.0.1`, `codex-cli 0.144.5`, model `gpt-5.6-sol`. Same as passes 4–7. + +## The trajectory broke + +Passes 4 → 8: **11, 14, 16, 15, 9**. First real drop, and — more telling — the *character* +changed. Pass 7 was design holes in a boundary that kept regenerating. Pass 8 is almost entirely +**sites I failed to sweep when deleting that boundary**: findings 1, 3, 5, 6 and 9 are each a +sentence written for the old design and left standing. Finding 4 is a stale count. Only +**finding 7** is a new observation about the design as it now stands, and it is narrow. + +Nothing in pass 8 says the boundary-free convention is unsound. That is the difference from +every pass since 4. + +## The blocker is a self-inflicted leftover + +**Finding 1.** `docs/coding-workflow.md:204` — *"The ledger is append-only **once an entry has +merged**: merged history is never rewritten…"*. I wrote that qualifier during the reachability +design, when a pre-merge amendable class existed. The class is gone; the qualifier reinstates it +in prose, in a file outside the spec, and it directly contradicts AC 3's literal reading. It is +a genuine blocker and a two-line revert. + +## Verified, not taken on the reviewer's word + +| # | Claim | Result | +|---|---|---| +| 1 | `coding-workflow.md:204` reinstates a pre-merge amendable class | **Confirmed**, quoted above. My edit, from the deleted design. | +| 4 | §8's unanchored list is stale | **Confirmed.** Three items it calls unanchored are anchors 13, 18 and 19 (`:382`, `:387`, `:388`). The list was written against the 16-anchor set and never re-derived for the 24. The distinguishing-fragment fallback is anchored nowhere *and* missing from the list. | +| 6 | the plan's quarantine note asserts the narrowing | **Confirmed**, `:3–10` — it says the design "narrowed it to *never edit a landed row*". Written before the restoration; the note meant to quarantine stale history now carries it. | + +| # | Sev | Verdict | Note | +|---|---|---|---| +| 1 | **Blocker** | **Valid, confirmed** | See above. Revert to absolute, keep the supersession sentence. | +| 3 | Major | **Valid** | The Scope paragraph still lists the `harden-finding` phrase while §4 and the version-bump paragraph say the skill is untouched. Two authoritative statements disagreeing about the plugin change surface — invariant-11 scope depends on it. | +| 2 | Major | **Valid** | §3.1 *still* says "falsification-scoped, not row-scoped". Pass 7 found this exact leftover; it was cited in my own applied list and the edit never landed. The worked example is the most-copied part of the document. | +| 4 | Major | **Valid, confirmed** | Gate-proof Don't: the residual list misdescribes what the check covers, in both directions. | +| 5 | Major | **Valid** | §8 still says nothing validates "that a row was landed when superseded" — a precondition from a deleted procedure, preserved as an unstated assumption. | +| 6 | Major | **Valid, confirmed** | See above. | +| 7 | Major | **Valid — the one genuinely new finding** | "No single-writer assumption is needed… nothing to overwrite" is justified only for two-branch union merges. Inserting an entry above `Columns:` is a read-modify-write on one file, so two writers in **one worktree** can lose an entry before git's union driver is ever involved. The claim is true of merges and overstated as written. | +| 8 | Minor | **Valid** | Story §5 points at "§7 check 2"; validation is now §6. Renumber fallout. | +| 9 | Nit | **Valid** | §4 claims all three story amendments use kept/narrowed/dropped; pass 1 uses kept/moved/dropped and pass 8 restored/dropped/kept. | + +## Recommendation + +Apply all nine. Eight are deletion-sweep or wording; finding 7 needs one scoped sentence +(the guarantee holds across independently committed branches; same-worktree concurrent writers +are unsupported). Then pass 9. On this trajectory the story does **not** need to park. diff --git a/.context/codex-reviews/gate-a-spec-pass-8.md b/.context/codex-reviews/gate-a-spec-pass-8.md new file mode 100644 index 0000000..97d0100 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-8.md @@ -0,0 +1,18 @@ +MAJOR | high | §2.1 lines 170–209 and §2.4 lines 266–297 | The form admits a review that was requested and later stood down, but the placement table has no path when that decision happens after the Gate-B closing commit during PR review; ordinary merge explicitly adds no merge-body copy and creating or amending a commit would reopen mandatory review work. | A legitimate use named by the shipped paragraph can have nowhere to put its required record, so the procedure is incomplete when read alone. | Add an explicit post-close placement rule for every merge strategy, or narrow the admitted examples to decisions made before the owning commit closes. +MAJOR | high | §2.4 lines 293–305 | The claim that Git identifies a record by the commit carrying it is false under the design's own states: one commit may carry several records, and squash deliberately moves records from several commits into one new commit. | Similar records become indistinguishable precisely on the accumulation and squash paths for which identity is needed, so correction, duplicate detection and audit cannot reliably name one record. | Define identity as a commit plus an occurrence discriminator that survives the supported carries, or stop claiming record identity and specify content-based/manual disambiguation with its limits. +MAJOR | high | §2.4 lines 271–312 | The settled rule that a differing duplicate is a copying error that stops the merge is absent; the shipped block says only that records accumulate, and the table handles a record found missing but never defines duplicate comparison or a pre-merge stop for differing copies. | A carry can silently mutate a record and still satisfy the written procedure, losing the one corruption guard retained after the four-state collision machinery was deleted. | Add the simple duplicate rule to the shipped block and table, define how duplicates are recognized without `Ref:`, and state the exact comparison and stop action. +MINOR | high | §2.4 line 293 | `in the order recorded` preserves a record-ordering rule even though the settled deletion says records have no meaningful order. | Implementers may invent or preserve ordering machinery that was deliberately removed, and concurrent/branched histories have no single recorded order. | Remove the ordering requirement and say that all records accumulate with order semantically irrelevant. +MINOR | high | §2.3 line 257, §2.4 lines 307–308, §3 lines 365–366, and §4 line 403 | Several orphaned references still promise the exception record's `identity rules` or `supersession rules` after those rules were cut. | The implementation surface and old-condition accounting direct readers to machinery that no longer exists, making it unclear whether deleted behavior must be rebuilt. | Replace these references with the surviving accumulation/correction/carry rules and remove `supersession` from the site table. +MAJOR | high | §2.4 line 296 | The restoration rule calls a changed block a second differing copy of `one decision`, but the new identity model says records are identified by their carrying commits; the original and follow-up copies necessarily have different carrying commits. | The procedure cannot both treat commit identity as decisive and recognize the follow-up as the same decision, so its byte-identical repair rule has no stated matching key. | State the exact content/provenance comparison that recognizes a restoration independently of commit identity, including how the original commit is named after squash. +MAJOR | high | §3 lines 365–366 | Saying exception records get the evidence entries' `same range and identity rules` conflicts with §2.4: evidence entries are keyed by story path and reduced to one winner, while exception records have no stable identifier and every record must accumulate. | Applying the evidence reduction to exceptions can drop records; refusing it leaves this sentence false and the exception algorithm unspecified. | Separate the two algorithms explicitly: share only the squash-range source, then state evidence reduction and exception accumulation/duplicate handling independently. +MAJOR | high | §3 lines 377–394 | The shipped follow-up sentence says to carry `the latest` evidence entry but does not ship the design's stop-and-human-choice branch for incomparable entries on different branches. | A reader of `CLAUDE.md` alone has no defined action for a real merge topology and can silently choose or discard an entry. | Put the incomparability stop and recorded-choice rule in both shipped copies, or define a deterministic total rule that is safe for branched ranges. +MAJOR | high | §5.3 lines 529–560 and §3 lines 388–394 | The parity inventory covers only rider (c)'s first carry sentence, not the mandatory immediately following per-story reduction sentence. | AC 9 can be recorded as passing while one shipped copy omits or changes the rule that decides which evidence is discarded. | Add the reduction/incomparability text as an anchored parity block and include it in §4's per-file change description. +MAJOR | high | §5.3 lines 529–559 | The anchor procedure is not executable as written: several listed `sentence` anchors are only sentence prefixes, their text is line-wrapped in the shipped blocks, and the procedure says to extract by anchors before applying the paragraph line-join normalization that would make those strings contiguous. | A verifier can find only the anchor-table copy, fail to find the shipped block, or silently choose a different pre-normalization scheme, defeating the uniquely-resolving parity claim. | Specify paragraph parsing and normalization before anchor matching, use complete start/end sentences, and require uniqueness after excluding the design's own anchor table. +MINOR | medium | §5.3 lines 545–551 | Joining every non-fenced paragraph line with one space erases Markdown hard line breaks made with trailing spaces or a backslash. | Two prompt copies can render with different instruction structure yet compare equal, contradicting the claim that meaningful Markdown boundaries are preserved. | Preserve explicit hard-break markers or compare parsed Markdown block/inline structure rather than normalized source lines. +MAJOR | high | §2.3 lines 232–258 and CLAUDE.md lines 72–80, 142–150, 394–396 | Old-condition accounting omits the existing Mechanics severity semantics and the Blocker/Major filtering clauses even though rider (b) directly changes how severity tokens are interpreted. | The required kept/moved/dropped audit does not state whether Title-case spellings, resolve/collect behavior, or the downstream filter survive normalization, so a rewrite can silently alter them. | Add explicit rows for the Mechanics severity line, the floor/filter behavior and writer format, marking each condition kept or narrowed and locating the new normalization step. +MINOR | high | §5.2 lines 474–484 | The line is called canonical `byte-for-byte`, but the required assertion matches it case-insensitively; the existing Title-case Mechanics sentence is a different sentence and does not require weakening the canonical-line comparison. | The check can pass a noncanonical spelling while the design claims it establishes presence of the canonical rule. | Either make the full canonical-line check case-sensitive while keeping only reader token normalization case-insensitive, or redefine the invariant precisely as byte-exact except for ASCII case and retract `byte-for-byte`. +MINOR | medium | §5.2 lines 482–496 | Region discovery fails closed only for an unfindable region; it does not require the `## 5.` opener or template closing fence to resolve uniquely. A duplicated §5 can therefore be concatenated or one can be selected while the assertion still appears exactly once. | The checker may report green when one of two shipped §5 copies lacks the closed-set rule, which outruns the stated containment claim. | Require exactly one opener and one applicable boundary per file and add duplicate-heading and duplicate-boundary reject fixtures. +MAJOR | high | dependent tier-2 story lines 85–92 | The proposed completeness proof hashes `the diff's paths` plus selected out-of-diff files, but that source set does not by itself cover deletions, file modes, renames, binary patch semantics, or prove that the path list and falsification-lens selection are complete. | The acceptance criterion claims a complete Gate-B range from a comparison that can omit review-relevant dimensions, violating the gate-proof calibration invariant. | Define the authoritative Git objects/patch representation and every compared axis, then bound any exclusions; avoid claiming completeness from current-file content hashes alone. +MINOR | high | §3 lines 377–384 | `That relation covers every pair in the range` overstates the ancestor relation and is contradicted immediately by the admitted incomparable side-branch pairs. | This is exactly the totality-style gate/mechanism overclaim AGENTS.md forbids and can mislead an implementer into assuming a winner always exists. | Say the relation is evaluated for every pair but only orders ancestor/descendant pairs; incomparable pairs take the explicit stop branch. +MINOR | medium | §2.1 lines 170–202 | The example `a review someone asked for and then stood down` can include an opportunistically routed bot review, while the same shipped block later says an absent opportunistic review needs no exception and writing one recreates forbidden ceremony. | A reader can follow either sentence and produce opposite behavior for the same supplementary-review event, undercutting the settled opportunistic routing rule. | Qualify the example as a non-opportunistic, genuinely decided optional review, or explicitly state that opportunistic bot silence remains excluded even if someone informally requested a run. +END OF FINDINGS (17 total) diff --git a/.context/codex-reviews/gate-a-spec-pass-8.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-8.pre-2026-08-14.md new file mode 100644 index 0000000..905d629 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-8.pre-2026-08-14.md @@ -0,0 +1,10 @@ +BLOCKER | high | SWEEP §4 "Change surface" and docs/coding-workflow.md:204 | the specified rewrite says the ledger is append-only only after an entry merges | this reinstates an amendable pre-merge class, so Rider 2 case 1 may be edited in place and AC 3 no longer holds literally | keep the append-only statement absolute and add the supersession move without any merged/unmerged qualifier +MAJOR | high | SWEEP §3.1 "The 2026-07-20 row" | the worked-example explanation still says the entry is "falsification-scoped, not row-scoped" | this teaches the deleted cumulative model and contradicts the settled row-marker plus latest-wins mechanism | describe the entry as a row marker whose current text identifies the falsified claim and must restate the row's complete supersession state when a later marker is added +MAJOR | high | SWEEP opening "Scope — behavioural" | scope still includes "the one phrase in harden-finding that the convention falsifies" while §4 and the version-bump paragraph say harden-finding is untouched because its absolute rule remains true | two authoritative instructions disagree about both behavior and plugin change surface, risking an unnecessary skill edit and invalid invariant-11 review scope | remove the harden-finding phrase from behavioral scope and state that only the ledger header and workflow-init mirror change behavior +MAJOR | high | SWEEP §6 Check 2 and §8 "What check 2's anchor list does not cover" | the coverage inventory calls the omit rule, stopped-versus-never-true instruction, and label-presence condition unanchored even though anchors 18, 19, and 13 cover them; deleting the closing clause is also caught by the END sentinel, while the live distinguishing-row-fragment fallback has no direct anchor and is absent from the unanchored list | the claimed four-state proof and exhaustive residual list misdescribe exactly what the check compares, violating the gate-proof Don't and allowing the duplicate-row locator decision to disappear from both surfaces unnoticed | re-derive the inventory from the literal anchors and sentinels, add a direct anchor for the distinguishing-fragment fallback or list it as unvalidated, and separate partially anchored subclauses such as the no-double-quote condition +MAJOR | high | SWEEP §8 "It adds no standing mechanical check" | the residual still says nothing validates "that a row was landed when superseded" | landed status is a deleted boundary with no role in the absolute convention, so this preserves the old decision procedure as an unstated precondition | delete the landed-status clause or replace it with the actual surviving condition, such as whether an entry locator resolves to exactly one committed row +MAJOR | high | SWEEP docs/superpowers/plans/2026-08-04-hardening-round-0-8-0-and-pr-21.md:3 | the new historical note says the supersession design narrowed the rule to "never edit a landed row" | the pass-8 restoration is already authoritative, so the very note meant to quarantine stale history reasserts the deleted boundary and fails Rider 3 | say the later design retained the absolute rule while adding a supersession mechanism, and identify only the plan's genuinely historical trust/current-behavior claims and Row D treatment +MAJOR | high | READ §2.2 "Concurrency takes the same answer" | "No single-writer assumption is needed" and "there is nothing to overwrite" are justified only with two-branch union merges | inserting an entry above Columns is a read-modify-write edit to the file, so concurrent writers in one worktree can lose an entry before Git's union driver ever participates | scope the guarantee to independently committed branches and record same-worktree concurrent writers as unsupported or specify coordination that preserves both writes +MINOR | high | SWEEP governing story §5 "Open questions" | the resolved note points to the design's §7 check 2 after validation was renumbered to §6 | §7 now contains review riders, so the governing story sends implementers to the wrong authority for the both-surfaces proof | change "§7 check 2" to "§6 check 2" +NIT | high | SWEEP §4 story change description and governing story §1 amendment notes | the spec says all three amendments have kept/narrowed/dropped accounting, but pass 1 uses kept/moved/dropped and pass 8 uses restored/dropped/kept with no explicit narrowed disposition | the mechanical compliance claim is false even though the prose substantially accounts for the detour, leaving it unclear whether "narrowed" was deliberately empty or omitted | describe the amendments as explicit old-condition accounting, or add "Narrowed: nothing" to the restoring amendment and avoid claiming identical category labels for all three +END OF FINDINGS (9 total) diff --git a/.context/codex-reviews/gate-a-spec-pass-9-dispositions.md b/.context/codex-reviews/gate-a-spec-pass-9-dispositions.md new file mode 100644 index 0000000..d578559 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-9-dispositions.md @@ -0,0 +1,72 @@ +# Gate A — spec — pass 9 dispositions (salvage cycle) — **GATE A CLOSED** + +9 findings (0 BLOCKER, 8 MAJOR, 1 MINOR). **None dismissed. All 9 applied.** + +Closed under the exit Daniel pinned before the pass ran: *one confirming pass 9; clean or +dispositions-only closes Gate A; anything else becomes named residuals and it closes anyway; +no pass 10.* Every finding turned out cheap and to touch either shipped text or an orphan, so +**all were fixed and no residuals are carried.** + +## The cut worked + +| Pass | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | +|---|---|---|---|---|---|---|---|---|---| +| Findings | 15 | 25 | 23 | 20 | 16 | 12 | 16 | 17 | **9** | +| Blockers | 3 | 2 | 5 | 4 | 0 | 0 | 0 | 0 | **0** | + +**17 → 9, and five consecutive zero-blocker passes.** Cutting the specification layer nearly +halved the finding rate in one step — the first structural change all cycle that moved it. +That is the termination assessment's diagnosis confirmed: the findings were living in prose +about prose, and deleting the prose deleted the findings rather than fixing them one at a time. + +## What pass 9 caught, and it is a good final pass + +**Two self-contradictions I introduced with the cut** — §5.2 requiring any out-of-region +occurrence to fail while its own fixture accepted a before-region copy alongside a correct one +(F2), and §2.3 still crediting the exception record with "a person re-reading the prospective +body (§2.4)" after §2.4 stopped specifying one (F8). Plus story AC 9 still promising the +enumerated anchors and extract-and-diff that the assessment deleted (F6). + +**One correction that partly reverses an earlier resolution, and should.** F1: the canonical +severity line was being matched **case-insensitively**, which meant a Title-case copy of it +would pass — so the check could go green while a shipped file omitted the uppercase rule it +exists to establish. Daniel's earlier resolution was about the **reader**, and that stands +unchanged: `CLAUDE.md` Mechanics legitimately spells severities in Title case, and a model +copying that spelling is following instructions. But the canonical line is a *different +sentence*, newly written, and it is now matched byte-for-byte. **Writer syntax exact; reader +tolerant** — which is what rider (b) always meant. + +**Three real gaps in shipped behaviour:** + +- **F7** — a decision made after its commit closed had no destination on a branch heading for + an ordinary or rebase merge: no next commit, no squash body, not yet merged. Now: **add an + empty commit for it.** It changes no content, so it raises no review obligation, and a record + with nowhere to go is a record that does not exist. +- **F9** — the reader never said whether to trim the whitespace the finding format puts around + each separator, so `MINOR ` could fail the token match and be escalated to `MAJOR` — turning a + collect-only finding into an iterate-and-fix one on an unstated parsing choice. Now trimmed, + with all-whitespace treated as empty and therefore INCOMPLETE. +- **F5** — the `+check` counterfactual said to run the assertion over a scratch tree holding + the two `df850ab` files, which would have failed other invariants and proved nothing. Now: + copy the otherwise-green worktree, confirm exit 0, swap only those two files, require exit 1 + **carrying the 4c diagnostic and no other**. The isolation is the point; without it the + evidence would have been green-for-the-wrong-reason. + +**F3, F4** — fixtures for a template-only copy and for every fail-closed path (missing start, +missing end, unreadable input, parser failure), so 4c is not the check that fails open. + +## Cycle summary + +**Nine passes, 153 findings, one dismissed** — the security-axis proposal at pass 1, dismissed +by Daniel with recorded reasoning. Across the whole story: **four cycles, seventeen passes, 456 +findings.** + +Three deletions decided by Daniel drove the convergence, and their pattern is recorded in the +design's §1.6: the drift record (a **requirement** — removing it removed obligations, 20 → 16), +the `Ref:` identifier (an **answer** — removing it left its question standing, 16 → 17), and +the specification layer (a whole **class of question** — 17 → 9). + +## Status + +**Gate A closed.** Spec, both stories and the profile change are uncommitted in the working +tree. Next: `writing-plans`, then Gate A on the plan — its own three-pass loop. diff --git a/.context/codex-reviews/gate-a-spec-pass-9-dispositions.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-9-dispositions.pre-2026-08-14.md new file mode 100644 index 0000000..3aa16b4 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-9-dispositions.pre-2026-08-14.md @@ -0,0 +1,52 @@ +# Gate A — spec — pass 9 dispositions + +Artifact: `docs/superpowers/specs/2026-08-05-hardening-ledger-supersession-design.md` +10 findings: **2 Blocker**, 6 Major, 1 Minor, 1 Nit. All read as correct. **None applied.** + +Reviewer: `mcp-codex-dev@1.0.1`, `codex-cli 0.144.5`, model `gpt-5.6-sol`. Same as passes 4–8. + +Passes 4 → 9: **11, 14, 16, 15, 9, 10.** All nine pass-8 fixes landed and none regressed — the +per-fix landing check confirmed each, and pass 9 raised no complaint about any of the nine +edits themselves. What it found is one omission beside them and a set of tightenings. + +## The two blockers + +**Finding 5 — the landing check's own blind spot.** §4's change-surface row still instructs +*"line 204's … qualified to merged history"* — the boundary-era instruction — while +`docs/coding-workflow.md:204` itself now correctly reads *"strictly append-only: a row is never +edited"*. §4 is the authoritative instruction an implementer follows, so the spec would have them +re-break the file the blocker fix just repaired. **My landing check missed this because it +grepped the target file, not the spec's instruction about the target file.** The rider was +right and was applied one level too shallow. + +**Finding 1 — a genuine conflict in the resolution-floor sentence.** "The convention governs +committed content; an uncommitted editor buffer is below its resolution" sits against "This holds +for every row without exception". A row **already appended but not yet committed** satisfies both +descriptions and has no verdict. That is the amendable class reappearing through the sentence +written to stop it reappearing — and it needs Daniel, because he specified that sentence. + +## Reproduced, not taken on the reviewer's word + +| # | Claim | Result | +|---|---|---| +| 5 | §4 still instructs the merged-history qualifier | **Confirmed**, `:271` against `coding-workflow.md:203-204`. | +| 9 | union does **not** reliably keep both labels | **Reproduced.** Scratch repo, `merge=union` on `.gitattributes`, both branches adding a `**Superseded rows:**` label: the merge **coalesces** the identical label line and keeps both distinct entries. Label count **1**, not 2. §8's categorical claim is wrong; §2.2's conditional "if a union merge leaves two labels" was right all along. | + +| # | Sev | Verdict | Note | +|---|---|---|---| +| 1 | **Blocker** | **Valid** | See above. Needs Daniel. | +| 5 | **Blocker** | **Valid, confirmed** | Pure omission; one row rewrite. | +| 3 | Major | **Valid** | A fragmentless locator unique when written can be made ambiguous by a *later* row sharing its date + fingerprint, and identical duplicate rows admit no distinguishing fragment at all. Entries are immutable, so there is no legal repair — AC 4 defeated with no move available. | +| 4 | Major | **Valid** | §3.1's entry and both date-bearing checks are pinned to `2026-08-05` while §2.2 defines `` as the day the entry is written, and the change lands later. Pass 7 raised this; it was not in the applied set. | +| 6 | Major | **Valid** | Assertion 2's fence calls `occurrences` and `unwrap` and reads `$LED`/`$TPL` without defining them. Run as the standalone block it is presented as, it fails under `sh` and `dash`. The assertions pass only when run in one shell — undocumented shared state. | +| 7 | Major | **Valid** | §8's re-derived inventory is still not complete: the opening row-claim sentence, the `YYYY-MM-DD` format and the cite-where-the-answer-lives requirement are in neither the 24 anchors nor the list, and the start sentinel anchors only `# Hardening log`. Third pass in a row on this inventory. | +| 8 | Major | **Valid** | §8 says "no lost-update case at all" while §2.2 now names the same-worktree read-modify-write race as unsupported. The pass-8 concurrency fix scoped §2.2 and left §8 asserting the unscoped claim. | +| 9 | Major | **Valid, reproduced** | See above. | +| 2 | Nit | **Valid** | "A row records a hardening" against the phantom-hardening case the design explicitly covers. "Records a hardening claim" resolves it. | +| 10 | Minor | **Valid** | Story AC 2 — "the row no longer reads as current behaviour" — reads as demanding the forbidden edit, since the design leaves every byte of the row unchanged. | + +## Read + +Two blockers, but neither says the design is unsound: one is an unswept instruction, one is a +conflict inside a single added sentence. Findings 6, 7, 8 are the same three sites tightening +across passes. Nothing here reopens the boundary. diff --git a/.context/codex-reviews/gate-a-spec-pass-9.md b/.context/codex-reviews/gate-a-spec-pass-9.md new file mode 100644 index 0000000..03fcb3c --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-9.md @@ -0,0 +1,10 @@ +MAJOR | high | §5.2 lines 452–460 and fixture row 508 | The assertion is called byte-for-byte and rider (b) makes uppercase the canonical writer syntax, yet the checker is required to match case-insensitively and the suite must accept Title-case and mixed-case copies of `Severity is one of exactly: BLOCKER \| MAJOR \| MINOR \| NIT — no other token.` | The check can pass when either shipped copy omits the canonical uppercase rule it is supposed to establish; the separate Mechanics sentence using Title-case severity names does not conflict with requiring this new normative line to be exact | Match the canonical line case-sensitively and replace the mixed-case accept fixture with reject fixtures; retain case-insensitive matching only in the reader-behaviour rule +MAJOR | high | §5.2 lines 462–476 and fixture rows 505–506 | The requirement says any occurrence outside the bounded region must fail, but the fixtures require a copy before the region plus a correct in-region copy to pass | No implementation can satisfy both rules, and the expected diagnostic for the before-only case is consequently undefined | Decide that out-of-region text is ignored, require exactly one in-region occurrence, and diagnose before-only as a missing in-region rule; alternatively reject every outside occurrence and change the accept row consistently +MAJOR | high | §5.2 fixture table lines 501–509 | The suite has `CLAUDE.md`-only as the sole one-copy case but no inverse case where only the workflow-init template contains the rule | An implementation that checks only the template passes every stated presence/absence case, so the requirement that `CLAUDE.md` independently carry the rule is not load-bearing | Add a template-only reject case with the 4c diagnostic, and record it in the 4c mutation flip set +MAJOR | high | §5.2 lines 462–476 and fixture table lines 496–509 | The assertion must fail closed on an unreadable file or an unfindable region, but the required fixtures cover neither a missing start/end boundary nor an operational read/parser failure | The error paths that prevent a traversal or parse failure from becoming a green invariant check can be omitted while the specified suite still passes, repeating the exact fail-open class checks 4a/4b already guard | Add diagnostic-isolated fixtures for missing start, missing end, unreadable input, and failure of each status-bearing tool stage introduced by 4c +MAJOR | high | §5.2 counterfactual lines 483–486 | `Check out that commit's two files into a scratch tree, run the assertion` is not a reproducible isolated procedure because 4c is embedded in the full checker, and a scratch tree containing only those files makes other invariant checks fail; therefore the claim that nothing else can supply the failure does not follow | The required `battery+check` counterfactual could be recorded green-for-the-wrong-reason, violating the gate-proof calibration and the existing test harness's diagnostic-isolation discipline | Specify an otherwise-green after-state scratch copy, replace only the two target files with their `df850ab` versions, require baseline exit 0, then require mutant exit 1 with the 4c-specific diagnostic and no other diagnostic +MAJOR | high | story AC 9 lines 238–245 versus design §5.3 lines 517–533 | The story still says the inserted blocks are enumerated with exact anchors in design §5.3 and verified by extract-and-diff, but §5.3 deliberately deletes the anchors/extraction procedure and replaces it with a human edit/diff instruction that nothing checks | This is a direct orphan from the cut machinery and leaves the acceptance criterion demanding a mechanism the final design explicitly rejects | Amend AC 9 to say both copies are edited, the edited regions are manually diffed, and the commit records that they matched, with the same explicit admission that nothing checks the claim +MAJOR | high | §2.1 lines 200–204 | The late-decision placement rule has no destination when the last branch commit is already closed, the branch has not merged, and the merge will be ordinary or rebase rather than squash: there is no next content commit, no squash body, and the follow-up-commit branch is limited to an already-merged branch | A use admitted by the shipped paragraph cannot be placed or made reachable from `main` using only that paragraph, breaking the story's durable-history criterion without creating a gate waiver | Name a destination for every non-squash late path, such as an explicit empty/follow-up branch commit or the ordinary merge commit body, and state how a rebase merge carries it +MINOR | high | §2.3 evidence-entry row line 272 versus §2.4 lines 295–313 | The old-condition table says the exception record has `a person re-reading the prospective body (§2.4)`, but §2.4 no longer contains that re-read instruction after the record procedure was cut | The accounting attributes a surviving validation step to a section that does not specify one, obscuring the deliberate decision that the record has no defined validation procedure | Remove the claimed re-read and state that the exception record has no defined revalidation, or add the re-read explicitly to the shipped paragraph if it is still required +MAJOR | medium | §3 lines 317–326 | The tolerant reader never defines whether surrounding whitespace is removed before matching the severity field, although the canonical finding format writes spaces around `\|`; a raw pipe split makes `MINOR ` and `NIT ` unrecognized and therefore `MAJOR` | A conforming Minor/Nit line can be escalated into the iterate-and-fix class depending on an implementer's unstated parsing choice | State that the reader splits on unescaped pipes and trims the format's surrounding ASCII whitespace before the case-insensitive token match, with an all-whitespace severity treated as empty and therefore INCOMPLETE +END OF FINDINGS (9 total) diff --git a/.context/codex-reviews/gate-a-spec-pass-9.pre-2026-08-14.md b/.context/codex-reviews/gate-a-spec-pass-9.pre-2026-08-14.md new file mode 100644 index 0000000..faf8356 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-pass-9.pre-2026-08-14.md @@ -0,0 +1,11 @@ +BLOCKER | high | §2.1 "The resolution floor" | committed-only scope conflicts with "every row without exception" and leaves an appended but uncommitted row outside the convention | the first rider case has no single verdict and the text recreates the amendable class the settled absolute rule forbids | remove the committed-content carve-out or state explicitly that any text already appended as a row is never edited, while pre-row drafting alone is below the rule +NIT | medium | §2.1 opening sentence | "A row records a hardening" asserts an actual hardening even though the same design supports a phantom row whose hardening never existed | the convention's premise and its covered wrong-when-written state use incompatible meanings of row | say that a row records a hardening claim or purports to record a hardening +MAJOR | high | §2.2 "Locator" | a fragmentless locator that is unique when written can become ambiguous after a concurrent or later same-date same-fingerprint row lands, and identical duplicate rows admit no distinguishing fragment at all | an immutable entry can cease to identify one row, defeating story AC 4 with no legal repair under the no-edit and no-remove rules | choose a stable immutable locator or define collision prevention and semantics for later and identical duplicates +MAJOR | high | §3.1 first entry and §6 checks 1 and 3 | the entry is fixed to 2026-08-05 although the convention defines the date as the day the entry is written and the ledger change is still unimplemented on 2026-08-10 | implementation would either violate the new convention or require ad hoc edits to the entry and both exact-date checks | parameterize all three sites with the actual implementation date or explicitly define and justify a different date semantic +BLOCKER | high | §4 change surface, docs/coding-workflow.md row | the authoritative instruction still says to qualify append-only history to merged history, while the landed pass-8 fix and current source instead make the rule absolute | an implementer is told to follow §4 and would reinstate a pre-merge amendable class, directly reversing settled decision 1 and violating prompt-standard 7 | replace the row with the actual absolute wording change and the appended supersession clause, with no merged-history qualifier +MAJOR | high | §6 check 2 assertion 2 fence | the parity fence calls occurrences and unwrap and reads LED and TPL without defining them; run as the standalone fenced command it fails under both sh and dash, although the combined assertions pass | the claimed runnable validation depends on undocumented same-shell state and can fail before testing parity | combine assertions 1 and 2 into one fence or repeat every required definition in assertion 2 +MAJOR | high | §8 "What check 2's anchor list does not cover" | the claimed two-item unanchored inventory omits settled decisions including the opening row-claim sentence, YYYY-MM-DD date format, and the requirement to cite where the current answer lives; the start sentinel anchors only "# Hardening log" | deleting those requirements identically from both surfaces leaves all 24 anchors and both sentinels green, so the bidirectional completeness claim is false | add distinguishing anchors for every omitted decision or enumerate every unanchored decision in §8 and narrow the completeness claim +MAJOR | high | §8 concurrent entries residual | "no lost-update case at all" contradicts §2.2's explicit unsupported same-worktree read-modify-write race that can lose an entry | this repeats the mechanism overclaim the pass-8 concurrency fix was meant to remove and violates the gate-claims Don't | scope the no-overwrite result to independently committed branches and retain the same-worktree lost-update exception +MAJOR | high | §8 union-merge residual | the categorical claim that two branches creating the block makes union keep both identical labels is not guaranteed; a scratch merge with the configured union driver coalesces the common label while retaining both distinct entries | the spec overstates what the merge mechanism proves and bases duplicate-label prose on an outcome that is only possible, not inevitable | say a union merge may leave duplicate labels and keep the shared convention's existing conditional wording +MINOR | medium | governing story §3 acceptance criterion 2 | "the row no longer reads as current behaviour" is literal about the row while the design deliberately leaves every byte of that row unchanged and only adds a marker above it | the acceptance criterion can be read as requiring the forbidden in-place edit rather than the intended contextual correction | require that the ledger no longer presents the row as unqualified current behaviour while the row bytes remain unchanged +END OF FINDINGS (10 total) diff --git a/.context/codex-reviews/gate-a-spec-rejected-design.stopped-tier3-core.md b/.context/codex-reviews/gate-a-spec-rejected-design.stopped-tier3-core.md new file mode 100644 index 0000000..4eccc82 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-rejected-design.stopped-tier3-core.md @@ -0,0 +1,1016 @@ +# Reviewer-availability fallback — a mid-flight human exception — Design + +**Story:** `docs/superpowers/stories/2026-08-13-reviewer-availability-fallback-story.md` +— the story header is the single writable copy of the profile. Every gate call carries that +path and reads the axes, the mode and the lens sets fresh from it; this document never +restates them as values. The mode is owed **in full**: a scoped `mode override` recorded on +2026-08-14 was **withdrawn** (§11). + +**History.** Three Gate-A passes on a three-tier design (37, 39, 40 findings) stopped under +§5's stuck condition; tier 2 and rider (a) split out to +`docs/superpowers/stories/2026-08-14-tier-2-same-family-reviewer-story.md` and +`docs/superpowers/stories/2026-08-14-sequential-branch-calls-hook-story.md`. Three further +passes on a two-tier design (30, 34, 32 findings) stopped again — blockers rising 7 → 8 → 12, +almost all of them in the re-review **debt machinery**, which this revision removes. **None +of the 212 findings across six passes has been dismissed.** + +## 1. Problem + +Both gates depend on one external reviewer. §5 requires a **minimum of three passes** per +gate with only the **final** pass clean, and permits an early exit below three only on a pass +returning zero findings. Out of quota, no pass can be taken at all, so no cycle closes and +all work stops. Observed: a five-day full process stall, 2026-08-05 to 2026-08-10. + +`/workflow-init` §2.13 does not reach this: it fires at **init time** and scaffolds a +gateless project. A configured project that loses its reviewer mid-flight falls outside it. + +**The gap: there is no authorized way to close a gate cycle when no reviewer can run.** + +## 2. Why this design carries no stateful controls + +**A cost-and-trust decision, not an impossibility claim.** Recorded because a later reader +will otherwise try to add the machinery back — and recorded *as a trade-off* because an +earlier draft of this section argued that stateful controls were structurally impossible, +which was a false dichotomy: it would have told that reader not to evaluate controls that +are in fact feasible. + +**What was tried, and what it cost.** A debt-tracking file was designed, reviewed and +removed. In `todos.md` it was Gate-B exempt, so a debt could be erased in a commit no +reviewer saw. On a non-`.md` path it became product-classified — so the **Gate-A closing +commit** staged a product file and raised a **Gate-B** obligation, during an outage, inside +the waiver meant to escape one. Every debt-state transition had that shape. A +consecutive-waiver cap built on that state inherited it, and two permitted acceptances could +disable the feature permanently. + +**Alternatives considered, and why each was rejected.** None is impossible. Each buys +tracking at a price this change declined to pay, and the price is stated so a future reader +can re-price it rather than re-derive the rejection. + +| Alternative | Why rejected | +|---|---| +| **State transition inside the same authorized closure** — the waiver authorizes its own debt row inside the tree it already fixes (§3.5 step 2), so opening the debt needs no second gate | Feasible, and the narrowest option. Rejected because it covers only *opening*: closing the row later is an unwaived transition needing its own authority, so the machinery returns at repayment time instead of being avoided. It also puts a product-classified path into a Gate-A closing tree, which §3.5 now refuses outright | +| **CI append-only / marker check** — a workflow asserting debt rows are only appended, and that every tier-3 marker has a matching open row | Feasible and mechanical. Rejected on cost and blast radius: it is new executable surface in a repo whose only executables are the hook and two checkers; like invariant 12's checker it can only run on `pull_request`, so a direct push bypasses it; and it enforces bookkeeping, not review — the marker would be guaranteed to have a row, never that anyone re-reviewed | +| **Protected external record** — a tracking issue, or a protected branch carrying the approval and the debt | Feasible, and the only option offering authority the author cannot forge. Rejected because it moves the record off `main`, where §5's entire disclosure model lives, and creates a second system that can disagree with history; it also assumes a forge, an account model and branch permissions this kit does not require of a user | + +**So the obligation survives as a sentence, not a system.** The marker states that a +cross-model re-review is owed. Nothing tracks it, nothing enforces it, and §13 says so. +What this buys instead is a record that is **discoverable in reachable `main` history** — +`git log --grep='Reviewer-tier: 3' `, and any commit-body view in a forge UI +— rather than ambient visibility every reader is exposed to. §12 parks the tracked version +with the trigger that would justify paying one of the prices above. + +## 3. Tier 3 — the mid-flight human exception + +**A new exception, stated as one.** §2.13's init-time gateless path closes no active gate, +and a profile override changes evidence mode rather than authorizing a zero-pass closure. + +### 3.1 Preconditions + +**Common to both gates:** + +1. **Tier 1 is unavailable** (§3.2). +2. **No unresolved adverse findings — and the pass ledger must be *establishable*, not + merely empty.** For every pass this cycle took, its findings file and dispositions must be + present, readable, valid under §5's acceptance rule, and every Blocker or Major in them + remediated or individually human-dispositioned. + **There is no authoritative enumeration of passes, and the design does not pretend + otherwise.** §5 records no durable pass id at the moment a pass happens; `.context/` slot + names carry no invocation-unique component and collide (this cycle has already destroyed a + predecessor's artifacts that way, §12); and the hook counter is explicitly not evidence. + So the ledger is authoritative **only where it is durable** — the ``s accumulated + in the WIP commit body at Gate B, and any recorded in a prior commit body at Gate A. + **Everything else fails closed.** The waiver is **refused** whenever the ledger cannot be + established: a resumed session with no durable record, a slot file whose provenance cannot + be tied to this cycle, a gap in the pass sequence, two artifacts disagreeing, or any + uncertainty about whether a call was made at all. `Passes completed: none` may be written + **only when the absence is established** — never when it is merely what the session can + see. Read the other way, an outage arriving after a bad pass, or after a *lost* one, is a + way to erase review evidence and land the defects the gate found. +3. **Not a remedy for a stuck review.** §5's "clearly stuck → STOP and surface" is + unaffected; disputed or rising findings are not an availability problem. +4. **Every governing story's profile resolves freshly**, by §5's three-case rule, read from + the story header at waiver time. Stated as a common precondition because a zero-pass + closure otherwise never forces resolution, and a malformed profile would be bypassed + precisely on the path with no reviewer. + **The set is the union, not the current citation.** It is every story path this cycle has + cited anywhere — the artifact now, any prior pass's prompt, any prior commit body in the + cycle — because validating only what the artifact cites *at waiver time* makes deleting a + citation a one-line way to shed a `high`-risk mode, its lens sets and its evidence + obligations while still obeying the written rule. A path that appears earlier in the cycle + and is absent now is an **unresolvable profile** — stop and surface — never `unprofiled`, + unless the removal is itself explained and human-confirmed as a profile change under §5. + Only a cycle in which **no** story was ever cited runs unprofiled, and the marker says so. +5. **The waiver checklist is answered** — see below. This is the compensating evidence, and + an unanswered item is a failed precondition. +6. **The merge strategy is squash or ordinary merge** (§6), revalidated before merge. + +**The waiver checklist.** Tier 1 buys questions being asked by someone other than the +author. Tier 3 cannot buy the independence back, so it obliges the **named human approver** +(§3.3) to ask the questions in writing instead. Five parts, each answered item by item — +answered, not asserted green: + +- the **unioned lens sets** the governing stories' profiles require (§5), applied and + answered, with risk's *abuse* and security's *abuse paths* as one question carrying both + labels; +- an **invariant-by-invariant** pass over `AGENTS.md` § Key invariants, marking each touched + and checked, or untouched; +- all **12 items of `docs/prompt-standards.md`** for every changed prompt artifact; +- **the standing falsification lens** — *which existing statements does this change falsify?* + — together with its companion: **name what this change alters the size, value or position + of**, and **grep for where each is described elsewhere**, recording the search performed + and its hits. At tier 1 these ride on the Gate-B prompt; a zero-pass closure has no prompt, + so without this they would be silently dropped on the one path with no independent reader, + and this repo's most persistent defect class is exactly the statement a change made wrong + in a file it never touched; +- the artifact's own stated self-checks, where it states any. + +**Every answer is `answered` or the precondition fails.** There is no +"answered-with-exception" and no free-text escape: a single unbounded exception field would +be a route around the only compensating evidence a zero-pass closure has. + +**The answers are a durable record (§5.5), not a summary word.** They enter the request +digest and travel the whole carry chain. An earlier draft compressed the whole checklist into +four words in the marker, which meant the item-by-item review existed only in the session +that performed it and evaporated at the first amend — leaving a claim that a review happened +and nothing a later reader could weigh. + +**Gate-specific evidence** sits on top of the checklist. An earlier draft required the full +profile-derived evidence set at both gates, which made a Gate-A waiver **impossible** — +implementation-derived evidence cannot exist before implementation, and Gate A is the gate +that blocks planning, which is where the outage actually bit. + +| Gate | What must be true | +|---|---| +| **A (spec / plan)** | The **mechanical sweep** is green — cited paths resolve, quoted passages match, stated counts agree, standalone fenced blocks parse. **This is syntax hygiene, not evidence:** it establishes that the artifact refers to real things and parses, and says nothing about contradictions, unsafe decisions, prompt-standard conformance or invariant risk. Those are the checklist's job, and the distinction is stated because an earlier draft offered the sweep as the compensating evidence itself. | +| **B (diff)** | The **battery** is green, and **every cited profiled story independently satisfies its own mode and suffix** with its own current evidence entry: the battery runs once for the cycle, each profiled story's entry is named separately, a cited **unprofiled** story owes no entry, and every entry is revalidated at §3.5 step 3 and carried into the closing body (§6). Singular "the profile's mode" is not what §5 requires of a multi-story cycle, and a marker satisfying one story's checks while citing three would look valid. | + +**Tier 3 waives reviewer passes only, never evidence.** + +### 3.2 Outage confirmation + +**Only a canonical gate call is evidence of unavailability.** A canonical call is the call +§5 already specifies for that gate: the correct tool (`mcp__codex__exec` at Gate A, +`mcp__codex__review` at Gate B), the reviewed repo root as `workingDirectory`, the **current** +artifact text or commit range, every cited story path and — at Gate B — each cited profiled +story's current evidence entry verbatim, plus the prompt's required opening and the +appended lens sets. + +**A failure of the request is not proof that the reviewer cannot run.** A malformed request, +the wrong tool, a wrong working directory, a stale artifact or range, a missing required +argument, a findings-file write failure, a count-mismatch or terminator failure, an +`INCOMPLETE` reply — each says *this call* was defective. Those are the manufacturable +failures: without this rule, unavailability can be produced on demand by sending a bad +request, and a reviewer perfectly able to review is waived. Such a result is **not** an +enum source; it is a defective call to be corrected and re-sent, and it consumes no +recovery attempt. The marker records **which canonical call class failed** — `exec`, +`review-spec`, `review-quality` — beside the source (§5.1 `Cause`). + +**Closed source enum — two values.** An earlier draft had four; two were removed because +they could be produced locally, which is the whole abuse path this section exists to close. + +| Source | Meaning | Canonical calls it requires | +|---|---|---| +| `quota-observed` | Quota exhaustion returned to a canonical call | one | +| `attempt-failed` | A canonical call **and** its recovery attempt both returned a transport- or service-level failure | two | + +**What was removed, and why it is not a gap:** + +- **`auth-failed`** — an authentication failure is **locally manufacturable**: withhold, + revoke or corrupt a credential and the request stays perfectly canonical while the + reviewer "cannot run". It is therefore **never an outage**. A local authentication failure + is a **configuration error to be repaired**, and a vendor-side credential revocation is + indistinguishable from it at the client, so both take the same answer: fix the + configuration, then call. If the account is genuinely gone, that is a setup gap under §5's + work-gap-vs-setup-gap rule, surfaced to the human — not a waiver. +- **`status-page`** — it was already never sufficient alone, and revalidation always ends in + a canonical probe whose *own* result is what gets recorded. So it could never be the + recorded `Cause`, and a value the enum cannot reach is not an enum value. A status page + remains useful as **corroboration** a human may cite in the decision `reason`; it + authorizes nothing. + +Both surviving sources require a canonical call that the vendor answered, so neither can be +produced without disabling the account itself — and disabling the account is the case above. + +**Revalidation repeats the source's own observation** immediately before the closing commit: +`quota-observed` needs one fresh canonical call; `attempt-failed` needs a fresh canonical +call **and** a fresh recovery attempt, because one call is not what that source means. + +**When the fresh observation differs, this table decides. No other reading is permitted, and +no enum value is inferred:** + +| Fresh result | Recorded `Cause` | Calls consumed | Outcome | +|---|---|---|---| +| Any canonical call **succeeds** | — | the fresh call | **Waiver refused** — the reviewer is available, and a waiver would simply be a skipped review | +| Quota exhaustion | `quota-observed` | the fresh call | Proceed | +| Transport/service failure, recovery attempt also spent and failed | `attempt-failed` | fresh call **and** recovery | Proceed | +| Transport/service failure, recovery **not yet spent** | — | the fresh call | Spend the recovery attempt, then re-enter this table with its result | +| **Authentication failure** | — | the fresh call | **Repair the configuration**, then re-enter this table. Never a waiver; if it cannot be repaired, surface it as a setup gap | +| A **defective call** (above) | — | none | Correct and re-send; not an observation of unavailability | +| Anything else — an unrecognized error shape, a result yielding no usable text, a clock or environment failure that makes the observation untimestampable | — | — | **STOP and surface**, naming the observed result verbatim. The close does not proceed and no waiver is taken | + +The recorded `Cause` is always the **revalidated** observation, never the original one. The +recovery row is the only transition that spends the shared one-attempt-per-pass budget of +§5; every other row consumes at most the one fresh call, which is not a recovery. + +Evidence is recorded **sanitized**: enum value, canonical call class and timestamp only, +under §5.4's grammar. No error payloads, endpoints or account identifiers. + +### 3.3 Roles, and what separation is available + +| Role | Who | +|---|---| +| **Author** | Whoever produced the change, human or agent | +| **Implementer** | The agent session that ran the cycle | +| **Waiver approver** | The human who authorizes tier 3 | + +The approver is a **human** and is never the implementer. **That is human-from-agent +separation, and it is the only separation this design provides.** A human author approving +the waiver on their own change is **accountable self-approval**, not independence, and the +design calls it that rather than dressing it as separation of duties — an earlier draft used +role names that implied an independence the solo case cannot deliver. + +### 3.4 Authorization — a real pause + +Before the closing commit the agent **stops, asks the human directly for an explicit +decision, and waits**. This is the **decision-question pattern this workflow already runs** +at every human checkpoint, not a new mechanism. + +**What is presented is the complete prospective record, not a summary**: the expected parent +commit, the fixed tree id, the full §5.1 marker with every field filled **except** the two +the answer itself creates, the full checklist record (§5.5), every carried evidence entry, +and the revalidated `Cause`. Binding approval to the tree alone would leave the marker, the +cause, the dispositions and the merge strategy free to change afterwards while the content +binding still looked valid — those fields live in the commit *message*, which no tree +contains. + +**Two digests, because one is circular.** The decision block's own timestamp and reason are +*created by* the answer. A single digest containing them would have to be shown before they +exist: fill them in afterwards and the authorized bytes change; pre-fill them and the record +states a predicted decision time and a predicted rationale rather than the decision that +happened. So: + +- The **request digest** (§5.4) covers the expected parent, the tree id, the marker without + its `Authorized-by` and `Authorization-digest` lines, the checklist record and every + evidence entry. It is computed at §3.5 step 3 and is what the human is shown. +- The **attestation** is the human's answer: they return **the request digest** together with + their handle, the decision timestamp, what they authorize and why. §5.2's block records + exactly that, and the marker's `Authorization-digest` carries the request digest verbatim, + which is what ties the answer to the bytes without either preceding the other. + +Nothing downstream re-derives the human's fields; everything downstream re-derives the +request digest and compares it. **Nothing verifies that the pause happened** — §13. + +### 3.5 The ordered close + +Because a valid answer or observation can otherwise be replayed after the world moves: + +1. **Preconditions checked** (§3.1). +2. **The expected parent is fixed.** `` is the current branch tip. It enters the + request digest, and steps 6 and 7 check it — an isolated index alone leaves `HEAD` free to + move, and a commit that acquires an unauthorized parent while carrying the authorized tree + **undoes concurrent work** and still passes a tree-only comparison. +3. **The prospective tree is fixed.** + - A **dedicated index** is created — `GIT_INDEX_FILE` at a path used by nothing else for + this cycle — and **initialized from `` with `git read-tree`**. A fresh index file + starts *empty*, so staging into one without this step writes a tree that **deletes every + tracked path not restaged**: a data-loss path, not a formality. + - The intended changes are applied to that index, and `git write-tree` records the + **tree id**. Nothing else writes that index before step 6. + - **At Gate A the change must be positively determined docs/artifact-only.** The test is + the **diff of ``'s tree against the prospective tree**, not an enumeration of the + tree itself — a tree contains the whole repository, so "every path qualifies" would + refuse every real commit. Every **changed** path is classified, across **additions, + deletions, renames (both sides), modifications, mode changes and type changes**, and + each must be the named artifact or prose `CLAUDE.md` exempts (`docs/**.md`, `README.md`, + `MANIFEST.md`). **Any prompt, script, code or otherwise product-classified path on + either side refuses the Gate-A waiver outright**, and no separately authorized Gate-B + waiver may ride inside a Gate-A closing commit. Otherwise a staged product change closes + under an A-spec or A-plan marker with §4's precedence excusing the Gate-B STOP it would + raise — a second gate-off path, in the direction invariants 2 and 3 exist to protect. +4. **Outage revalidated** (§3.2); **evidence revalidated** (§3.1), including every cited + profiled story's entry; then the **request digest** is computed (§5.4) over ``, the + tree id, the marker without its two answer-created lines, the checklist record and every + carried evidence entry. +5. **The human pause** (§3.4), on that request digest. The answer returns the digest plus the + decision fields. +6. **The final probe and the age check**, after the answer. The close **restarts at step 1** + if: the step-4 revalidation is more than **15 minutes** old by UTC clock; the clock is + unreadable; the probe's result differs from the recorded `Cause`; the request digest no + longer matches a recomputation; or **`HEAD` is no longer ``**. The probe repeats + the source's own observation (§3.2) — one canonical call for `quota-observed`, a call + **and** a recovery attempt for `attempt-failed` — and the recovery it spends is the same + single per-pass attempt §5 already budgets, never a second one; if that attempt is already + spent and `attempt-failed` cannot be re-established, the close **stops and surfaces**. + The pause is unbounded and the reviewer can recover inside it, which is why the probe sits + on the far side of the answer. +7. **The closing commit, immediately**, written from the dedicated index of step 3 onto + ``. +8. **Read back what actually landed** — parent, tree and message — and compare all three: + the commit's parent must be ``, its tree the authorized tree id, and the + normalized record bytes of its message (§5.4) must equal what the request digest covered, + with the two answer-created lines matching the attestation. **Any mismatch, or anything + that cannot be read back, is an incident** (§6): the commit is not a valid tier-3 closure. + Checking the tree alone was the earlier draft's error — everything authorization binds + lives in the *message*, which no tree contains, so an edited, truncated, duplicated or + hook-rewritten body would have passed. + +**Authorization binds to the request digest of step 4 plus the attestation of step 5.** If +``, the tree, the artifact, any evidence entry, the checklist record, the merge +strategy or any marker field changes after that, or the commit fails, **every step is +repeated** — authorization is consumed by one commit attempt and does not survive it. +Otherwise history could disclose a waiver for one target while landing different bytes. + +## 4. What the hook does — and what it does not + +The hook is unchanged and **waiver-unaware**. + +An earlier draft prescribed `.context/codex-gate.off`; it silences reminders for **unrelated** +commits, which is the missed-commit direction invariant 2 names as dangerous. With a +single-commit closure nothing needs silencing, so **invariant 2 is untouched and `.off` keeps +its meaning.** + +A further earlier draft claimed the hook's Gate-B STOP was itself the compensating control. +It is not: a Gate-A closing commit is docs-only and exempt; Gate B can report satisfied from +previously counted calls; the output reaches the **agent**, not the human; and the hook +**always exits 0**, so it blocks nothing in any case. **The hook supplies no waiver control in +either direction.** Where this design says a reminder "fires", it means exactly that the +classifier selects that branch and the text is emitted — never that a review is forced. + +**Precedence, narrowly.** For **one recorded cycle id and gate, at one closing transition**, +an authorized waiver takes precedence over the hook's Gate-B or below-floor STOP and the +agent proceeds. This does **not** generalize: at every other STOP, every later commit and +every resumed session, the compliant-agent procedure is to obey the reminder as before. +"Authoritative" is the wrong word for text emitted by something that always exits 0; what +the reminders have is standing in the procedure, not enforcement. + +**Gate-A continuation, across reminders that fire later than the authorization.** Both Gate-A +reminders arrive after the commit that consumed the corresponding authorization: the spec +waiver closes at the spec commit but the below-floor reminder is emitted at `writing-plans`, +and the plan waiver closes at the plan commit while its reminder is emitted at +`executing-plans` (the hook resets its Gate-A count at `writing-plans` in between). Without a +rule at **both** points, a correctly waived artifact can never advance — the agent either +stalls on a STOP it cannot clear or silently extends a spent authorization — and the original +stall this design exists to fix would survive at the spec stage, which is where it bit. + +**The rule is state-free and re-derived, never remembered.** Before proceeding past the +reminder, the agent reads history for the commit that introduced the artifact it is about to +work from, and continues **only if** all of these hold: + +| At | Artifact | Required in that commit's body | +|---|---|---| +| `writing-plans` | the spec about to be planned from | a valid §5.1 marker with `Gate: A-spec` | +| `executing-plans` | the plan about to be executed | a valid §5.1 marker with `Gate: A-plan` | + +plus, in both cases: a §5.2 decision block matching it under §5.3, and a `Waived-target` +naming that artifact — **at a blob sha equal to the artifact's content right now**. The +current blob is resolved at the moment of the check and compared; a `Waived-target` that +merely names the right *path* is not enough. Otherwise the artifact can be edited after its +waived commit and the stale authorization still clears the reminder, which is a gate-off path +for bytes no one waived. Any later edit needs a new Gate-A cycle for that artifact. + +Anything else — no marker, the wrong gate's marker, a marker for a different artifact or +cycle id, a blob that no longer matches, an unparseable block, a body that cannot be read — +is **no authorization**, and the reminder stands. A resumed session performs the same +re-derivation and reaches the same answer, because nothing is carried in session state. + +**This clears that reminder whenever it is derived for that unchanged artifact**, which is +not the same as clearing it once: the derivation is deliberately state-free, so a resumed or +repeated invocation re-derives and re-clears. Nothing records prior clearance, and claiming +one-time consumption would describe a mechanism that is not there. Changing the artifact is +what ends it. The Gate-B floor for the implementation that follows is untouched, and needs +its own tier-1 passes or its own tier-3 waiver. + +## 5. Records + +### 5.1 The tier-3 cycle marker + +Mandatory in the cycle-closing commit body. **Both** this and §5.2 are required; §10 +validates each independently. + +``` +Reviewer-tier: 3 — human exception, gate waived +Gate: A-spec | A-plan | B · cycle: +Waived-target: @ | diff +Authorized-parent: +Authorized-tree: +Profiles: none-cited | [, …] +Passes completed: none | [, …] +Adverse findings: none | # — : +Waiver-checklist: · · / answered +Merge-strategy: squash | merge +Authorization-digest: +Authorized-by: · +Cause: · · revalidated +Cross-model re-review: OWED — untracked; this line is the only record +``` + +**`Waived-target` never names a sha the commit itself will change.** At Gate A it is the +artifact path and its blob sha, which the commit lands and does not alter afterwards. At Gate +B it is `diff ` — the digest (§5.4) of the `` → `` diff. +An earlier draft wrote `..`, which cannot work at Gate B: the closure is an +**amend**, so the `headSha` named in the body is destroyed by the very commit that carries +it, leaving a durable marker pointing at an object that may be unreachable. Parent, tree and +diff digest all survive the amend and identify the same bytes. + +**`Authorized-parent` and `Authorized-tree`** are what §3.5 step 8 reads back. Both are +present because either alone can be satisfied while the other is wrong. + +**`Profiles:` lists every governing story path** (§3.1(4) — the union across the cycle, not +the current citation) with its **status only**, never its mode: §5 makes the story header the +single writable copy and keeps mode values out of commit bodies, so what the marker records +is *which stories governed and which of them owed evidence*, and a reader re-resolves the +rest. A mixed cycle carrying both kinds is ordinary and must remain reconstructible, which a +bare `unprofiled` could not express. `none-cited` is the genuinely story-less case. + +**`Waiver-checklist:`** carries the approver's handle, the digest of the §5.5 record, and the +answered count over the total — never a per-part summary word, which would be a claim about a +record instead of a handle on it. + +**`Waived-target`, never "reviewed"** — one word would otherwise overstate exactly what the +gate proved, which is the AGENTS.md gate-proof Don't. + +**`` is `//pass-

`** — cycle-unique and durable in the commit body, +because the underlying findings files live in git-ignored `.context/` under reusable slot +names and cannot be cited durably. + +**`Cross-model re-review: OWED — untracked`** is deliberate and literal. §2 explains why the +tracking system was removed; this line is what remains, and it is honest about being a +sentence rather than a mechanism. + +### 5.2 The logged human decision + +``` +Human decision: · · gate: A-spec | A-plan | B · cycle: · authorizes: · reason: +``` + +The handle, timestamp, **gate** and cycle id **must match the marker** exactly under §5.4's +normalization, or the close stops. The `gate:` field is explicit rather than inferred: §5.3 +keys both records by cycle id **and** gate, and an A-spec, an A-plan and a B decision sharing +one cycle id are otherwise indistinguishable — the collapse would merge the wrong +authorization or reject the right one. + +**Sanitization applies to `reason`** as it does to `Cause`: no raw error payloads, endpoints, +customer details or security-finding text in a public commit body; sensitive rationale is +referenced, not quoted. + +### 5.3 Idempotency + +A resumed or retried close must not append a second marker or decision block. Both are keyed +by **cycle id + gate**, present in both records: an identical repeat — byte-equal after §5.4 +normalization — collapses to one; a repeat differing in **any** field is a **conflict** and +stops the close. The same rule governs the multi-WIP collapse (§6). + +### 5.4 Field grammar + +The blocks are a line-oriented protocol carrying human-supplied text, so the grammar is +stated rather than assumed: an unescaped newline or separator in a reason would otherwise +inject or spoof a field, and two agents "normalizing" differently would disagree about +whether two records match. + +- **Encoding**: UTF-8. Each block is a contiguous run of lines with no blank line inside it. +- **``**: `[A-Za-z0-9._-]{1,64}`. An identifier that resolves to an accountable + person in the project's own terms; nothing verifies that it does (§13). +- **``**: ISO-8601 UTC to seconds — `YYYY-MM-DDThh:mm:ssZ`. No local times and no + bare dates, so no two readers can order or compare two records differently. +- **``**: `--`, slug `[a-z0-9-]{1,48}`, `` a positive + integer. +- **``**: `//pass-

`. +- **Free text** — `reason`, `authorizes`, a disposition's ``, a checklist answer: + one line, at most **200 bytes after escaping**. Text that cannot fit the bound is + **referenced, not truncated** — truncation silently changes what was approved — and text + that cannot be escaped stops the close. +- **Escaping**, applied in this order: `\` → `\\`, newline or carriage return → `\n`, + `·` (U+00B7) → `\·`, `|` → `\|`. **Unescaping applies the exact reverse order**, so a + literal `\` before an escapable character cannot be misread as an escape it never was. +- **Parsing, parity-aware.** A separator counts as a separator **only when preceded by an + even number (including zero) of consecutive backslashes** — otherwise `\\ · ` and `\ · ` + split identically and the record is ambiguous. A line splits on its first unescaped `: `; + a value splits on unescaped ` · `; a `: ` pair inside a value splits on its + first unescaped `: `. **Unescape only after every split**, never before. +- **Lists** — `Profiles:`, `Passes completed:`, `Adverse findings:` — are `, `-separated, + parity-aware in the same way, and their elements are **ordered**: story paths, pass ids and + findings each sort ascending by byte sequence before the record is written, because + otherwise two byte-different renderings of the same set would hash differently and §5.3 + would call an identical repeat a conflict. +- **Paths** in `Profiles:` and `Waived-target:` are repo-relative, `/`-separated, with no + `.` or `..` component and no leading `/`. +- **Normalization, for the §5.1/§5.2 and §5.3 equality checks**: compare the **unescaped byte + sequences** with leading and trailing ASCII spaces stripped. Nothing else — no case + folding, no Unicode normalization, no whitespace collapsing inside a value. A permissive + comparison is what lets two different handles or two different cycles match. +- **Accept/reject vectors are owed with the implementation**, not with this design: at least + one case per escape, one odd-backslash case per separator, one over-length free-text case, + one reordered-list case, and one case per marker field. A grammar with no vectors is a + grammar two implementers will read differently. + +**Digests.** All three — the **request digest**, the **checklist digest** and the +**diff digest** — are `sha256`, hex, lowercase, and computed over a **length-delimited** +serialization: each component is emitted as its decimal byte length, a `:`, then its exact +unescaped bytes, concatenated in the fixed order the defining section lists. Length framing +is not decoration — plain concatenation lets two different records hash identically by moving +a byte across a boundary, which is precisely the substitution the digest exists to detect. +The digest string appears in the record; the serialized input does not. + +### 5.5 The waiver checklist record + +The checklist (§3.1) is the only compensating evidence a zero-pass closure has, so it is a +**record**, not a claim that a record existed. It goes in the same commit body, below the +marker: + +``` +Waiver-checklist-record: + / | answered | + … +``` + +- **One line per item**, in a fixed order: the lens questions in the order §5 lists them + (risk first, then security, with the dual-labelled abuse question once), then the invariants + in `AGENTS.md` order, then `docs/prompt-standards.md` items 1–12 per changed prompt + artifact, then the falsification lens and its size/value/position companion, then each + stated self-check. +- `` is §5.4 free text — one line, 200 bytes, escaped. **The bound is the point:** an + answer that does not fit is *referenced* (a path, a pass id, a query), which keeps the + record bounded without letting an unanswerable item pass as answered. +- **`answered` is the only permitted verdict.** An item with any other verdict, or with no + line, fails the precondition (§3.1(5)) — there is no exception class, because one + free-text exception field would be a route around the whole checklist. +- The **``** is §5.4's `sha256` over the item lines in order, and it is what + the marker's `Waiver-checklist:` field and the request digest both carry. + +**It travels the whole carry chain** (§6) exactly like the two blocks: restated on amend, +collected on a multi-WIP collapse, copied into the squash body, and validated before the +merge. Without that it would exist only in the session that produced it and vanish at the +first amend — leaving `Waiver-checklist: answered` as an assertion about a document nobody +can read. + +## 6. Disclosure, and the carry chain + +- **Gate A** → the **spec or plan commit body**; carried forward into the implementation + cycle's closing body. If work stops before Gate B, that docs commit is already durable. +- **Gate B** → the **WIP commit body**. +- **The closing amend** replaces the WIP body wholesale **except** that disclosure content is + restated explicitly. Nothing is preserved automatically. +- **The multi-WIP collapse**: every block collected, deduplicated by §5.3, validated before + the replacement commit. +- **The squash body**: both blocks **and** every profiled story's evidence entry copied in. + +**Two strategies are permitted, and they have different durable records.** The chain above +is the squash chain — `Gate-A docs commit → WIP → amend → squash → main` — where the squash +body is the only thing reachable from `main`. Under **ordinary merge** every branch commit +stays reachable, so the durable record is the **branch commit that already carries the +blocks**, and the merge commit carries nothing: + +| | Squash | Ordinary merge | +|---|---|---| +| Durable record | the squash commit body | the Gate-A docs commit, or the amended Gate-B closing commit, reachable from `main` via the merge's second parent | +| Validated before the merge | the prospective squash body carries both blocks and every profiled story's evidence entry | every waived cycle's block pair and evidence entries are present in a reachable branch commit, and the branch head is the authorized one | +| Several waived cycles on one branch | collapsed per §5.3 into one body | left as separate commits, one block pair each — **no collapse**, because each already has its own durable commit | +| Merge-commit body | n/a | **no requirement, and no restatement.** A restated block that drifts creates two records on `main` that can disagree, with nothing to arbitrate | +| A block missing, found before the merge | merge refused | merge refused | +| A block missing, found after the merge | incident recovery (below) | incident recovery (below) | + +**What immutability binds, and what it does not.** §3.5's "authorization is consumed by one +commit attempt" binds **that commit** — its parent, its tree, its message. It does **not** +freeze the branch. It cannot: a Gate-A waiver is followed by a plan commit and then the whole +implementation, so a rule voiding the authorization whenever the branch head moves would make +Gate A's waiver unusable under either strategy, and re-running the close at the final head is +impossible anyway, since that head contains the product paths §3.5 step 3 refuses. So: + +- **Gate A** binds its own commit and its artifact blob. Later commits on the branch are + expected and change nothing. What must still hold at merge time is that the waived commit is + reachable and the artifact blob is unchanged (§4's continuation check is the same test). +- **Gate B** binds the final branch head, because the Gate-B closure *is* the head. A head + that moves after authorization voids it, and the merge is refused until the close restarts. + +Under either strategy the record is found the same way — +`git log --grep='Reviewer-tier: 3' ` — which is why §2 names that query +rather than claiming ambient visibility. + +**Validated before the merge**, per the rows above, and **the merge is refused if any block, +checklist record or evidence entry is missing**. Story AC 2 names the *cycle-closing commit*; +a later corrective commit cannot put blocks into a commit that already landed, so it satisfies +the criterion for no cycle. + +**Incident recovery, and why a revert is not enough.** A commit that fails §3.5 step 8, or a +closure discovered after the fact to be missing its records, leaves a **valid-looking marker +in reachable history**: reverting the content does not remove the commit, and +`git log --grep='Reviewer-tier: 3'` keeps finding a closure this design calls invalid. So: + +- **Before the branch is published**, repair history — amend or reset — so the invalid marker + never becomes reachable from `main`. This is the preferred path and the usual one. +- **Once published**, history is not rewritten. A **`Reviewer-tier-3-invalidated:`** record is + committed instead, naming the invalid commit sha, the check that failed, and the handle and + timestamp of whoever recorded it. It reproduces the missing or corrected blocks verbatim so + history is not silent. +- **The discovery query is therefore two-step, and the design says so rather than pretending + the grep is sufficient**: collect the tier-3 markers, then collect the invalidation records + and subtract the shas they name. A marker whose commit appears in an invalidation record is + not a closure. This is a convention read by a human, enforced by nothing. + +Either way the incident is recorded as a **disclosure failure, not as compliance**. + +**Rebase-merge and cherry-pick are refused before the waiver**, not discovered after it. + +**Never in the findings file.** + +## 7. Old-condition inventory — CLAUDE.md §5 + +Fifth attempt. **The previous four each claimed exhaustiveness and each was wrong** — three +were thematic compressions, and the fourth, which claimed a clause-by-clause derivation, still +omitted six conditions (rows 82–87 below). + +**So this version claims a derivation, not a property.** It was produced by walking every +normative sentence of `CLAUDE.md` §5 in order — including the Profiles and Mechanics +subsections — and emitting a row per sentence that constrains behaviour, splitting a sentence +whose clauses have different effects. That is a procedure a reader can repeat and check. It is +**not** a guarantee of completeness: nothing mechanical derives these rows, §5 changes, and +four predecessors were confident and incomplete. If you are relying on this table, re-walk §5. + +Every disposition is judged **by effect, not by the noun used**: calling an operation a waiver +rather than a skip does not preserve a condition about skipping. Rows a previous version got +wrong are **[corrected]**; rows a previous version **omitted entirely** are **[added]**. + +| # | §5 condition | Disposition | +|---|---|---| +| 1 | Two independent cross-model gates | **Narrowed** — tier 1 unchanged; tier 3 waives | +| 2 | Gates are advisory but mandatory | **Overturned for tier 3** | +| 3 | `.off` opt-out; "the gates still apply" | **[corrected] Narrowed** — the initial diagnostic meaning is kept and the file is untouched, but after an authorized waiver the gate does not apply to that one closure. A previous version marked this untouched | +| 4 | Hard floor: min 3 passes per gate | **Kept** at tier 1; **N/A** at tier 3 | +| 5 | Blocker/Major only drive the loop | **Kept** | +| 6 | Hook counts passes; cannot read findings | **Kept**; §4 adds it bears on tier 3 not at all | +| 7 | Hook cannot tell the spec run from the plan run; resets at `writing-plans` | **Kept, untouched** | +| 8 | Gate A is instruction-backed; a satisfied count is not a clean review | **Kept** | +| 9 | TodoWrite a task per pass | **Kept, untouched** | +| 10 | Fix Blocker/Major after each pass | **Kept** | +| 11 | Final pass must be clean | **Kept**; **N/A** at tier 3 | +| 12 | Stuck → STOP and surface | **Kept**, explicitly not waivable (§3.1(3)) | +| 13 | Only early exit is a zero-finding pass | **[corrected] Overturned for tier 3** | +| 14 | Don't manufacture findings to pad | **Kept, untouched** | +| 15 | Codex is advisory — validate before applying | **Kept, untouched** | +| 16 | Dismissed finding → one-line why | **Kept**; §5.1 requires handle and reason for a waiver-time disposition | +| 17 | Findings go to a FILE, not the response | **Kept, untouched** | +| 18 | Pass the repo root as `workingDirectory` | **Kept, untouched** | +| 19 | Slot naming; no invocation-unique component | **Kept**; the collision row stays open (§12) | +| 20 | One finding per line; escape literal pipes | **Kept**; rider (b) pins the severity token | +| 21 | No blank lines, headings, prose or continuations | **Kept**; rider (b) keeps structural failures INCOMPLETE | +| 22 | Exact terminator | **Kept, untouched** | +| 23 | `NO FINDINGS` for a clean pass | **Kept, untouched** | +| 24 | Reply is one line per branch, or `INCOMPLETE` | **Kept, untouched** | +| 25 | Gate B takes one file per branch under `full` | **Kept, untouched** — rider (a) moved out | +| 26 | Delete every target before each call; confirm gone | **Kept, untouched** | +| 27 | A surviving target → stop and name the cause | **Kept, untouched** | +| 28 | One pass at a time | **Kept** as a stated limitation | +| 29 | Companions are advisory; never validate a pass | **Narrowed in one named scope** — rider (b) (§10) | +| 30 | Resume note is cycle-stable, not pass-named | **Kept, untouched** | +| 31 | Acceptance: readable, exact terminator, exactly `` lines, nothing else | **Kept, untouched** | +| 32 | Anything else is INCOMPLETE and discounted | **Kept, untouched** | +| 33 | Recovery: one attempt per pass, shared | **[corrected] Narrowed** — the shared one-attempt budget is kept and is never doubled, but the precondition it feeds is **per source** (§3.2): `attempt-failed` consumes call **plus** recovery, `quota-observed` needs one canonical call, and a defective call consumes nothing. §3.5 step 6's post-answer probe spends the *same* attempt, not a second one. A previous version stated a spent recovery as a blanket tier-3 precondition | +| 34 | Recovery is a fresh re-run deleting exactly what it rewrites | **Kept, untouched** | +| 35 | Resume preferred only when the write failed; pass `sessionId` + `reviewType` | **Kept, untouched** | +| 36 | Spent and still incomplete → STOP and surface | **[corrected] Narrowed** — the diagnostic steps are kept, but tier 3 permits continuation after a spent attempt instead of terminal escalation. A previous version marked this untouched | +| 37 | Hook counting residuals are not evidence | **Kept, untouched**; §4 relies on it | +| 38 | Discount every incomplete pass whatever the counter says | **Kept, untouched** | +| 39 | Gate A is two runs, each its own loop | **[corrected] Kept at tier 1; narrowed at tier 3** — the two independent cycles remain and tier 3 is available to each, but either may close with **zero passes** and therefore without entering its loop at all. A previous version marked this kept outright, which is the effect-level change hidden under an untouched noun | +| 40 | Prompt opens with the brainstorming directive | **Kept, untouched** | +| 41 | One broad prompt, re-run each pass over the revised artifact | **Kept, untouched** | +| 42 | Append intent, artifact text, and which invariants it touches | **Kept, untouched** | +| 43 | Every finding with severity and confidence; you filter, Codex never does | **Kept, untouched** | +| 44 | Mechanical sweep before each read pass | **Kept**, and promoted to a Gate-A waiver precondition (§3.1) — explicitly as **syntax hygiene**, never as the compensating evidence | +| 45 | Optional focused per-dimension passes on large artifacts | **Kept, untouched** | +| 46 | Gate B: tests green, before commit | **Kept**; the battery is a Gate-B waiver precondition | +| 47 | Gate B tool and args: `instruction`, `whatWasImplemented`, `baseSha`; `full` runs both | **Kept, untouched** | +| 48 | `baseSha` from a WIP commit; `WIP:` naming; amend to close; `reset --soft` for several | **Kept**, and §6 adds what the collapse must carry | +| 49 | Skip ONLY trivial changes | **[corrected] Narrowed** — by effect a non-trivial change can now close without review; the noun differs, the effect does not | +| 50 | Re-review after every fix | **[corrected] Narrowed for tier 3** — §3.1(2) permits closing after a remediation that no independent reviewer ever sees. The marker records the remediation and the `OWED` line covers it; a previous version marked this untouched | +| 51 | A fix changing specified behaviour updates the spec in the same commit | **Kept, untouched** | +| 52 | The standing falsification lens — which existing statements does this diff falsify? | **[corrected] Narrowed** — kept at tier 1, where it rides on the Gate-B prompt. A zero-pass closure has no prompt, so it is asked by the human in §3.1's checklist instead; a previous version marked it untouched while tier 3 dropped it entirely on the one path with no independent reader | +| 53 | Name what the diff changes the size, value or position of, and grep for where each is described elsewhere | **[corrected] Narrowed** — same as row 52, and the checklist record (§5.5) must carry the search performed and its hits, not just the answer | +| 54 | Prose-vs-product classification | **Kept, untouched** — no file changes classification in this design | +| 55 | Risk lens set: threats, abuse, rollback, data loss, idempotency, compatibility, observability | **[corrected] Narrowed** — kept at tier 1; at tier 3 there is no reviewer prompt to append them to, so they are asked by the named human in §3.1's waiver checklist instead. Answered by a different actor, in a different artifact — not untouched | +| 56 | Security lens set: assets, trust boundaries, roles, external systems, abuse paths | **[corrected] Narrowed** — same as row 55, and §13 records that the answers are self-audit rather than independent review | +| 57 | Both axes → union appended once, abuse carrying both labels | **[corrected] Narrowed** — the union-once rule and the single dual-labelled abuse question are kept, but at tier 3 they govern the checklist, since there is no prompt to append to | +| 58 | Lenses are different questions, not more passes | **Kept** — the checklist asks the same questions and adds no passes; at tier 3 there are none to add | +| 59 | Profile: story header is the single writable copy | **Kept** — no value is copied anywhere | +| 60 | Three profile-reading cases | **[corrected] Narrowed for tier 3** — the three cases are kept, but §3.1(4) changes what they are read against: the **union of story paths across the cycle**, not the artifact's citation at waiver time. Only a cycle that never cited a story is `none-cited`; a path that appears earlier and is gone at the waiver is **unresolvable**, not unprofiled. A previous version said tier 3 "needs no citation", which made dropping one a way to shed a `high`-risk profile | +| 61 | Stop and surface on an unresolvable profile | **Kept**, and **strengthened**: §3.1(4) makes fresh resolution of every cited story a tier-3 precondition at **both** gates. A previous version left it kept-untouched while the zero-pass path never forced resolution at all, so a malformed high-risk profile was bypassable precisely where no reviewer looked | +| 62 | Gate-B triviality skip needs two independent conditions | **[corrected] Narrowed** — the two-condition rule is kept **within tier 1**, but at the level of the decision procedure tier 3 opens a second route by which a non-trivial change closes without review, with its own preconditions (§3.1) and its own disclosure (§5.1). Marking it untouched because "a waiver is not a skip" is exactly the noun-based accounting row 49 rejects | +| 63 | The skip reason is recorded in the commit body | **Kept, untouched**; the waiver marker is a separate record | +| 64 | A skip removes the review, never the evidence | **Kept**, and tier 3 follows the same principle (§3.1) | +| 65 | A cycle citing several stories aggregates per story | **Kept**, and applied at the waiver: §3.1's Gate-B row requires per-story satisfaction, one battery, unioned lenses, one entry per profiled story | +| 66 | What the author owes by mode | **Kept**, gate-appropriately (§3.1) | +| 67 | A named verification may substitute for an automated check | **Kept, untouched** | +| 68 | The counterfactual, and the wiring that could produce it | **Kept**, and applied to this change (§10) | +| 69 | An unobservable counterfactual is a blocking evidence gap | **Kept, untouched** — and **not** invoked here: §10 specifies the probe that makes the counterfactual observable, and §11 records the withdrawn override that assumed it was not | +| 70 | A fabricated test satisfies nothing | **Kept, untouched** | +| 71 | Evidence entry in the commit body, revalidated before close | **Kept**; the marker joins it in every hop | +| 72 | Every Gate-B call carries each cited story path + evidence entry verbatim | **Kept, untouched** | +| 73 | Work gap vs setup gap | **Kept, untouched** | +| 74 | Profile changes are proposed, human-confirmed, logged | **Kept, untouched** | +| 75 | An axis change voids prior overrides | **Kept, untouched** | +| 76 | Passes under a lower profile still count; only the final clean pass must be current | **Kept, untouched** | +| 77 | Fold a mid-cycle profile edit into the WIP by amend | **Kept, untouched** | +| 78 | The closing message carries the validated evidence entry | **Kept**, and now the marker too | +| 79 | Timeout/abort handling; one retry is the shared attempt | **Kept**; feeds §3.1 | +| 80 | §5's stated non-enforcement residuals — nothing checks which file was read, whether the header moved mid-call, whether lens sets were appended | **Kept, untouched**, and §13 adds this design's own | +| 81 | **The gate itself is not optional** | **OVERTURNED for tier 3**, deliberately and by name | +| 82 | **[added]** Gate A runs on the spec right after brainstorming (before `writing-plans`) and on the plan before `executing-plans`/`subagent-driven-development` | **Kept** — the placement is unchanged; §4 adds the continuation rule at both of those transitions because that is where the reminders arrive | +| 83 | **[added]** The pre-pass mechanical sweep settles only what a machine can decide **without side effects**, and does **not** run commands quoted in the artifact — a quoted command may be destructive or an intentional failure | **Kept, untouched**, and load-bearing at tier 3: §3.1 promotes the sweep to a precondition, so its no-side-effects limit now guards a path with no reviewer. Omitted by the fourth inventory | +| 84 | **[added]** Gate B checks the diff against `AGENTS.md` explicitly | **[corrected] Narrowed** — at tier 3 no reviewer performs it, so the invariant-by-invariant part of §3.1's checklist carries it. Omitted by the fourth inventory, which marked the invariants only as lens content | +| 85 | **[added]** Companions may be deleted or rebuilt; the findings file plus terminator is the only hard requirement | **Kept** — and rider (b) (§9) is the one named scope where a companion's **absence** discounts a pass, already recorded at row 29 | +| 86 | **[added]** The only early exit below the floor is a pass with **zero** findings; don't manufacture findings to pad | **Kept** as a distinct clause at tier 1, and **N/A** at tier 3, which is not an early exit but a closure with no pass at all. Rows 13 and 14 covered the padding half and the exit half separately; neither stated the clause itself | +| 87 | **[added]** If revalidation changes the evidence entry, the clean pass no longer covers what is being committed — fix, re-review, close on the entry that pass validated | **[corrected] Narrowed for tier 3** — there is no clean pass to invalidate, so the analogue is §3.5: a changed entry after step 4 voids the authorization and the close restarts. Omitted by the fourth inventory | + +## 8. Sites + +Found by searching the **behaviour claim**, which is how the passages below were found after +earlier drafts listed one section of one file. + +| Site | Change | +|---|---| +| `CLAUDE.md` §5 | Tier 3, the records, the carry chain, riders (b) and (c) | +| `/workflow-init` inline §5 mirror | The same edits — story AC 9, verified by extracted parity | +| `/workflow-init` §2.13 | Init-time scope stated; **gateless answer unchanged** | +| `docs/getting-started.md` ~106 | **"Gate A is not skippable at any level."** Directly falsified — highest priority | +| `docs/getting-started.md` ~84, ~34, ~53 | Floor and skippability stated unconditionally | +| `docs/coding-workflow.md` § *The two gates…* | "advisory but mandatory… not optional"; **heading not renamed** | +| `docs/coding-workflow.md` ~19, ~27–29, ~91–99, ~108–109, ~123–126 | Pipeline and stage-level independent-review claims | +| `docs/coding-workflow.md` ~253–268 § *What is essential — keep it* / *Minimal viable adoption* | Names "the two independent review gates" as load-bearing and "one independent code-review gate" as a day-one minimum. Tier 3 makes both waivable, so the passage gains the exception rather than continuing to teach an unconditional invariant. Missed by earlier drafts, which stopped at the stage-level passages | +| `docs/pr-review-bots.md` ~141 | "**Every head reaching a PR has already passed Gate B**" — the premise under which an absent opportunistic-bot review blocks nothing. A tier-3-waived head falsifies it, so the premise is **narrowed** and the routing states how a waived Gate B is treated: the PR bots become the *only* automated review of that head, which is a reason to read them, not a reason to promote them to **Wait for** | +| `docs/sparring-briefing.md` ~41–44 | **"Advisory, never exempt… do not treat a satisfied human as a substitute for a clean pass."** Tier 3 is exactly that, so this is **overturned here**, not moved to the tier-2 story as the story's first amendment wrongly said | +| `README.md` product summary and daily-use pipeline | The mandatory-gate claim gains the exception | +| `plugins/dev-workflow/.claude-plugin/plugin.json` | Its description carries the same claim | +| `.claude-plugin/marketplace.json` — the **top-level `description`**, not the plugin entry | "Cross-model review workflow: **independent Gate A/B review**, …". The plugin entry at line 14 lists components and makes no such claim; an earlier draft cited the wrong field, and this pass's mechanical sweep caught it | +| `AGENTS.md` § *What this project is* | Gains a clause admitting the waiver | + +**`docs/coding-workflow.md` ~149 is NOT changed** — its "makes the gate mandatory" is the +repo-enforced **quality** gate, not the review gate. An earlier draft listed it, which would +have weakened an unrelated guarantee. + +**Two new files are added, and both need tree entries.** §10.1's structural oracle is an +executable check — `scripts/check-waiver-prose.sh` — with a regression suite beside it, +`scripts/check-waiver-prose.test.sh`, matching the convention the two existing repo-local +checkers already follow. Both go in `AGENTS.md`'s architecture tree and in its § Commands +battery row, because that row claims to be what CI runs. This is stated explicitly because an +earlier draft of this design added a file and omitted the tree, and a later one asserted "no +new file is added" — which §10.1 falsified the moment the oracle became mechanical. The +alternative was to keep the check judgement-based, which is what pass 2 rejected. + +**Not changed:** the four same-model prohibition sites. They forbid a same-model *reviewer*; +a human exception is not a model reviewing its own work. + +## 9. Riders + +**(b) Canonical syntax, and a tolerant reader.** **Canonical:** severity is +`BLOCKER | MAJOR | MINOR | NIT`, uppercase. **Reader:** normalization applies **only** to an +otherwise-valid six-field line whose severity field is non-empty and unrecognized; it maps to +`MAJOR`. Every **structural** failure stays INCOMPLETE. + +The pass is valid only if that pass's dispositions file records the drift: + +- **Grammar**, a distinct ordered record type alongside the file's existing + one-verdict-per-finding contract: `Enum drift: → MAJOR · lines [, …]`, citing + the finding lines normalized. +- **Timing:** the record must exist **before the reader credits the pass** — not before the + hook counts it, since `PostToolUse` fires before the agent can read the returned file. +- **Retention:** until the cycle closes. **If it disappears**, the pass it qualified is + **discounted** — §5's existing answer for an unverifiable pass. + +Motivating incident: PR #23's Gate-B pass 3 returned all four findings at `IMPORTANT`. + +**(c) Squash-merge carry** — §6, covering both markers and evidence entries. + +## 10. Validation evidence + +**Battery** — the `AGENTS.md` quality command, green. + +**Prompt conformance** — all 12 items of `docs/prompt-standards.md` for every changed prompt +artifact: `CLAUDE.md` §5, the `/workflow-init` mirror, §2.13, and **`AGENTS.md`**, which §8 +changes and which `CLAUDE.md` classifies as product. The reviewer and result are named. + +**Check that fails without the change — two halves, one mechanical and one observational.** +An earlier draft claimed the counterfactual was unobservable and took a scoped mode override; +**both the claim and the override are withdrawn**. A later draft replaced them with "a +prompt-harness scenario", which named no runner, no fixture and no oracle — the same +unverified-evidence claim in a new place. A third specified scenarios but still rested the +whole result on a probabilistic model's agreement with itself. + +**So the check is split, and each half is claimed at exactly its own strength.** + +### 10.1 The structural oracle — mechanical, deterministic, the load-bearing half + +A shell check over the two frozen texts, run by the same battery that runs everything else. +It decides only questions a parser decides, and it is the half that can **fail closed**: + +| Assertion | On `OLD` | On `NEW` | +|---|---|---| +| §5 contains a closure path whose preconditions are an enumerated list | absent | present, and the enumeration has exactly the items §3.1 names | +| The severity enum is stated as a closed set of tokens, not shown by example | absent | present, all four tokens | +| Every marker field §5.1 defines appears in the template §5 ships | n/a | all present, none extra | +| The `/workflow-init` inline mirror and §5 agree on all of the above | trivially | byte-identical after the documented extraction | + +**Failing without the change is structural, not interpretive**: run against `OLD`, rows 1–3 +fail. This half needs no model, reruns identically forever, and is what the evidence entry +leads with. + +### 10.2 The prompt-differential probe — observational, and claimed as such + +The oracle establishes that the prose *says* the thing. It cannot establish that a reader +*acts* on it, which is the risk this change actually carries. The probe addresses that, and +its claim is bounded accordingly. + +**Frozen inputs**, by blob sha, never paraphrased: `OLD` = `git show :CLAUDE.md`, +`NEW` = `git show :CLAUDE.md`. + +**Driver, pinned** — invariant 5 applies to anything that executes in a run, and a probe is a +run. The evidence entry names the exact CLI package at an exact version, an exact model id +(never an alias), temperature 0, a fixed seed where the runtime accepts one, a fixed timeout, +and a sanitized environment: an empty scratch directory, no project files, no `CLAUDE.md` on +any parent path, no MCP servers, no tools. The prompt is assembled from the frozen blob plus a +scenario file byte-for-byte, with a documented delimiter, and both are stored with the run +output. Anything unpinned makes two implementers' probes incomparable while both record the +expected verdicts. + +**Scenarios.** Each supplies the §5 text, a **fully sequenced transcript** reaching one exact +decision point, and one question demanding a first-line verdict token. + +| # | Transcript reaches | Question | `OLD` | `NEW` | +|---|---|---|---|---| +| **S1** | The end of §3.5 step 6: mid Gate-A-spec cycle; pass 1 complete, 3 findings, all remediated, ledger establishable; **one** canonical `mcp__codex__exec` call returned quota exhaustion (so `quota-observed`, one call — not a recovery); profile resolves; checklist answered, all items; parent fixed; dedicated index read-tree'd and written; changed paths docs-only; request digest computed; human returned the digest and authorized; post-answer probe repeated quota exhaustion 4 minutes later; squash | *"Under the procedure above and nothing else, may the closing commit be made now? Answer `CLOSE` or `REFUSE` on the first line, then quote the clause you relied on."* | `REFUSE` | `CLOSE` | +| **S1-neg-1…13** | S1 with **exactly one** thing wrong: unremediated Major · stuck review · unresolvable profile · story path dropped since cycle start · unestablishable ledger · an unanswered checklist item · rebase-merge · a product path among the changed paths · post-answer probe **succeeds** · revalidation 40 minutes stale · `HEAD` moved off the fixed parent · index not initialized from the parent · digest recomputes differently | same | `REFUSE` | `REFUSE`, **quoting that clause** | +| **S2** | A findings file with four findings at severity `IMPORTANT`, correct count and terminator, **and** a dispositions file present but carrying **no** enum-drift record | *"Is this pass acceptable? Answer `ACCEPT` or `INCOMPLETE` on the first line, then quote the clause."* | `ACCEPT` | `INCOMPLETE` | + +**S1's call budget matches §3.2**: one canonical call, `quota-observed`. An earlier version +spent a recovery attempt there, which the new procedure only permits for `attempt-failed` — +so the frozen positive case violated the procedure it was meant to confirm, and `NEW` could +have returned `REFUSE` for the right reason and the wrong finding. + +**S2's `OLD` verdict is the weak one, and is stated as weak.** Old §5 says findings come "one +per line in the format above", whose example severity is `MAJOR`; a careful reader could call +`IMPORTANT` structurally invalid *before* the change. So S2 is **corroborating, not +load-bearing** — 10.1's row 2 carries the rider (b) claim mechanically, and S2 is reported +with its ambiguity named rather than counted as a clean differential. + +**The thirteen negative variants are what make the probe non-vacuous.** S1 alone passes on any +prose that merely permits closure; each negative fails if the corresponding clause is stated +loosely, so the probe can fail for a reason other than "the text changed". + +**Repeatability.** Every scenario runs **three times per input** in independent sessions, and +**all three must agree**. A split verdict is a **failure**, not a retry: it is evidence the +prose does not determine the answer. The remedy is to fix the prose. + +**Failure semantics.** Any oracle assertion failing, or any S1 / S1-neg scenario missing its +verdict → the check fails → `battery+check+verification` is unmet → **Gate B is not called**. +No majority rule, no rerun budget, no partial credit. S2 is reported, not gating. + +**What each half establishes — calibrated to the comparison actually performed.** + +- **10.1 establishes** that the shipped text contains the named structures and that the two + copies agree. That is a string comparison, and the claim is exactly that large. +- **10.2 establishes** that **one pinned model, under frozen inputs, emitted these verdicts in + three of three runs**. It does **not** establish that the procedure's semantics compel the + verdict, that a different model would agree, or that a *compliant* reader must conclude the + same — a model is not a proof of meaning, and three agreeing samples from one distribution + are not three witnesses. Read as evidence it is real; read as proof of semantics it would be + the overclaim §13 and the gate-proof Don't both forbid. +- **Neither half establishes enforcement.** A non-compliant agent is prevented from nothing + (§13). No git operation, hook or script is exercised, because this change contains none. + +**The counterfactual, and the wiring that could produce it.** If the claim were false — if the +old procedure already authorized a zero-pass closure — 10.1's rows 1–3 would pass on `OLD` and +S1 would return `CLOSE` on `OLD`. The wiring can produce those observations: `OLD` is the real +prior blob supplied whole and unedited, the oracle's assertions are evaluated against it by +the same code path, and the verdict token is read from the model's own first line rather than +supplied by the probe. + +**The result is recorded in the evidence entry**: both blob shas, the pinned CLI version and +model id, the run count, the oracle result, and the per-scenario verdicts. + +**Named verification.** The matrix describes **observable behaviour of the procedure followed +correctly**. It does **not** claim enforcement: §13 states an agent departing from the +procedure can produce a conforming-looking commit, and no row should be read as a mechanism. + +| Case | Observed outcome when the procedure is followed | +|---|---| +| Gate-A spec waiver | marker **and** decision block in the spec commit body; identifiable from history alone | +| Gate-B waiver | both blocks survive WIP → amend → squash; readable from `main` alone | +| Decision block absent | non-conforming; each block validated independently | +| Handle / timestamp / gate / cycle mismatch across blocks | close stops (§5.2) | +| Repeated or resumed close | identical blocks collapse; conflicting blocks stop the close | +| Multi-WIP collapse | all blocks collected and validated | +| Prospective squash body missing a block, the checklist record or an evidence entry | **merge refused** before it lands (§6) | +| Waiver after an unremediated Major | procedure refuses (§3.1(2)) | +| A checklist item unanswered | procedure refuses; there is no exception verdict (§3.1(5), §5.5) | +| Waiver during a stuck review | procedure refuses (§3.1(3)) | +| Gate-A waiver, sweep red | procedure refuses (§3.1) | +| Gate-B waiver, battery red | procedure refuses (§3.1) | +| Tree, entry, checklist record or marker field changes after authorization | authorization void; all steps repeated (§3.5) | +| Committed parent, tree **or message bytes** differ from what was authorized | **incident** — not a valid closure (§3.5 step 8); repaired before publication, or invalidated by record after (§6) | +| A **changed path** at Gate A is a prompt, script or code path — added, deleted, renamed, mode- or type-changed | Gate-A waiver refused outright (§3.5 step 3) | +| Reviewer recovers **during** the human pause | the post-answer probe catches it; waiver refused (§3.5 step 6) | +| Revalidation older than 15 minutes, or `HEAD` no longer the fixed parent | close restarts at step 1 (§3.5 step 6) | +| Dedicated index not initialized from the parent | the written tree deletes unstaged tracked paths — the data-loss path §3.5 step 3 exists to prevent | +| A prior pass's findings file missing or truncated | counted as unresolved adverse evidence; waiver refused (§3.1(2)) | +| A governing story's profile unresolvable, **or a story path dropped since cycle start** | stop and surface — not a waiver, and never `unprofiled` (§3.1(4)) | +| Pass ledger not establishable — resumed session, colliding slot, sequence gap | waiver refused; `Passes completed: none` may not be written (§3.1(2)) | +| A defective call — wrong tool, stale range, failed findings write | not an outage observation; corrected and re-sent, no recovery consumed (§3.2) | +| Fresh failure matching no enum row | STOP and surface; no waiver (§3.2) | +| A-plan marker present at `executing-plans`, plan blob unchanged | that reminder is cleared, whenever derived, for that plan (§4) | +| A-spec marker present at `writing-plans`, spec blob unchanged | that reminder is cleared, whenever derived, for that spec (§4) | +| Waived artifact edited after its waived commit | blob mismatch; no authorization, the reminder stands (§4) | +| A-spec marker found where an A-plan marker is required | no authorization; the reminder stands (§4) | +| `attempt-failed` revalidated with one call only | insufficient; the source needs call **and** recovery (§3.2) | +| Authentication failure at any point | configuration error to repair — never a waiver (§3.2) | +| Revalidation succeeds | waiver refused — the reviewer is available | +| Rebase-merge selected | refused **before** the waiver | +| Hook at the closing commit | emits its reminder and **exits 0**; it forces nothing | +| Enum drift with / without the companion record | valid / INCOMPLETE | +| Companion deleted before close | the pass it qualified is discounted | +| Structurally broken finding line | INCOMPLETE, never normalized | + +## 11. Withdrawn + +Recorded because both were confirmed decisions, and a reader of the history will otherwise +find them and assume they hold. + +- **The scoped `mode override`** on the story's profile log. §5's grammar permits a whole + effective mode in the header and requires the header to carry an override; there is no + per-portion override, so recording one invented a mechanism the profile system does not + have. `battery+check+verification` is owed **in full**, and §10 supplies the `+check`. +- **The debt record's move to a Gate-B-classified path.** Moot — the debt machinery is gone + (§2) — and it was also recursive. + +## 12. Packaging and backlog + +- `todos.md`: **occurrence 3** added to the compound-commands row (story AC 8) — `git add` + and `git commit` in one Bash call, empty staged set at `PreToolUse`, loose STOP; observed on + PR #23's close. Same shape as occurrence 2 and, like it, a **false positive**. The existing + item is **edited in place**; append-only-never-edit is `docs/hardening-log.md`'s rule. +- `prompt-vague-criteria` closes. `unverified-enforcement-claim` **stays open**, re-pointed at + the hook story. +- **New parked row: tracked re-review debt.** The obligation currently lives only as the + marker's `OWED` line. *Trigger: a tier-3 waiver whose re-review is found never to have + happened, or the third waiver in one repository — whichever comes first.* §2's table prices + the three alternatives that were considered and rejected, so whoever takes this row starts + by re-pricing one of them rather than rediscovering the option space. +- **New parked row: enforced waiver authority** — signed commit or protected-branch approval. + *Trigger: the first tier-3 record whose authorization is disputed or unattributable.* +- **New parked row:** the hook's `is_docs_only` exempts **any** `.md` path outside a prompt + directory, broader than §5's prose list (`docs/**.md`, `README.md`, `MANIFEST.md`). Found + while siting the removed debt store. *Trigger: a root `.md` file acquiring gate-relevant + state.* +- **Gate-cycle slot collision — trigger fired, row stays open.** This cycle's first pass + deleted its predecessor's findings file and dispositions before the surviving 44 artifacts + were archived by hand. +- **New parked row:** tier-2 counting and containment, pointing at the tier-2 story. +- **`AGENTS.md` § Commands and § Architecture both change** — `scripts/check-waiver-prose.sh` + and its suite join the battery row, the lint row (`shellcheck --shell=sh`) and the tree + (§8). The battery row is the one CI mirrors, so a check added to one and not the other is + the drift invariant 11's prose-conformance checks cannot see. +- Version **0.8.2 → 0.9.0** with a `plugins/dev-workflow/CHANGELOG.md` entry — invariant 12. **This will be + verified** by `scripts/check-version-bump.sh` against the PR's base *after* the WIP commit + contains both the plugin edits and the manifest bump; run before then it reports clean, + uselessly. + +## 13. What this does not do + +Residuals known at design time. The list is **not** exhaustive. + +- **Tier 3 waives the gate.** A change closed this way has had **no** independent review. +- **The re-review obligation is a sentence.** Nothing tracks it, nothing schedules it, nothing + fails if it never happens. §2 explains why that is the honest form rather than a defect, and + §12 parks the tracked version with its trigger. +- **The authorization pause is procedure, not enforcement.** Nothing verifies that it happened + or that the handle belongs to whoever answered. Every block in §5 is a **recorded + assertion**; an agent that skips the pause and writes the blocks produces a + conforming-looking commit. +- **The only separation is human-from-agent.** A human author approving their own change is + accountable self-approval, and nothing checks even that. +- **The waiver checklist is self-audit.** §3.1 makes the profile's lens sets, the invariants + and the 12 prompt standards get *asked* at tier 3, by a named human, in writing. It does + not make them get asked by someone independent of the work — that is precisely what tier 1 + buys and tier 3 cannot. A checklist answered honestly is worth having; a checklist answered + to clear a gate is worth nothing, and nothing distinguishes them from history. +- **Every "procedure refuses" row in §10 is behaviour of a compliant agent**, never a + mechanism preventing a non-compliant one. +- **The hook forces nothing.** It emits a reminder and exits 0. +- **Nothing detects availability's return** outside the pre-close revalidation. +- **`main`'s history can be rewritten.** A force-push can remove a marker. +- **Rollback is not one operation, because the prompts ship to three places.** Reverting this + repository removes the procedure *here* and leaves existing markers standing as records of + what happened — that part is simple, and no state file exists to migrate. The other two are + not: + - **Installed plugin copies** live under a version-keyed cache path, so a revert on `main` + reaches no machine until each user updates. Until then a machine keeps running the tier-3 + procedure with no signal that it was withdrawn — the same staleness invariant 12 exists + for. Rollback therefore means a **version bump that removes the feature**, announced in + `plugins/dev-workflow/CHANGELOG.md`, not a revert commit. + - **Scaffolded downstream copies.** `/workflow-init` writes an inline §5 into each + initialized repository, and those copies are the user's file, edited thereafter. Nothing + detects them and nothing migrates them: a revert here leaves tier 3 documented and live + in every project already initialized, with no matching upstream documentation. The honest + answer is that downstream removal is **manual and unprompted** — the CHANGELOG entry is + the only notification, and this is a cost of invariant 8's inline-template design, not a + defect introduced here. + - **In-flight authorizations** are unaffected either way: authorization is consumed by one + commit attempt (§3.5) and nothing survives a session, so there is no authorized-but-unspent + state for a rollback to strand. diff --git a/.context/codex-reviews/gate-a-spec-resume.2026-08-03-hardening-round.md b/.context/codex-reviews/gate-a-spec-resume.2026-08-03-hardening-round.md new file mode 100644 index 0000000..71720e1 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-resume.2026-08-03-hardening-round.md @@ -0,0 +1,75 @@ +# Gate A — spec cycle — resume note + +**Cycle:** hardening round over the 0.8.0 cycle and PR #21. +**Spec:** `docs/superpowers/specs/2026-08-03-hardening-round-0-8-0-and-pr-21-design.md` +**Story:** `docs/superpowers/stories/2026-08-03-hardening-round-0-8-0-and-pr-21-story.md` + +## State + +Three valid passes run and accepted (file-first protocol satisfied each time: terminator +exact, counts matched, no non-finding body lines). The 3-pass floor is **met**; the +final-pass-clean requirement is **not**. + +| Pass | Findings | Major | File | +|---|---|---|---| +| 1 | 19 | 12 | `gate-a-spec-pass-1.md` + dispositions | +| 2 | 15 | 11 | `gate-a-spec-pass-2.md` + dispositions | +| 3 | 19 | 14 | `gate-a-spec-pass-3.md` + dispositions | +| 4 | 12 | 5 | `gate-a-spec-pass-4.md` + dispositions | +| 5 | 9 | 4 | `gate-a-spec-pass-5.md` + dispositions | +| 6 | 10 | 4 | `gate-a-spec-pass-6.md` + dispositions | + +**Paused at pass 3**, not through exhaustion: two named story exits tripped (machinery in the +skill edit; §5 insertions that were paragraphs rather than clauses), both pre-declared by +Daniel as exits rather than judgement calls. He chose to re-scope and resume: the +`harden-finding` change was split into its own story, the §5 additions were cut back to one +sentence each, and the loop continued from pass 4 on the revised artifact. + +Every finding of passes 1–6 was accepted; none was dismissed. Seven reopened human-confirmed +decisions and were taken back to Daniel rather than applied unilaterally. + +## Prior cycle's artifacts + +The 0.8.0 classifier cycle's spec-slot files were **archived, not destroyed**, before this +cycle's first call: 17 files renamed to `*.2026-07-30-classifier-cycle.md`. The parked row +"Each Gate cycle destroys the previous cycle's review record" documents that loss; it did not +recur here. + +## Resuming + +Awaiting Daniel's decision on the split. Whatever the shape: + +- Pass 3's findings **4, 7, 8, 12, 14, 17, 19** and minors **2, 3, 15, 16** stand regardless + of how the work is divided — they are defects in the current spec and story text, not + artifacts of scope. +- Finding **12** is the one to carry first: the Row B precheck verdict is wrong (C5 is inside + the 2026-07-19 Don't), which is the single-row error pass 2 identified, committed inside the + spec that fixes it. +- Any resumed cycle restarts pass numbering at 1 against the revised artifact. The three + passes recorded here do not carry over — they reviewed a spec that no longer exists in that + form. + +## Environment observation, for the `CLAUDE_CODE_MCP_AUTO_BACKGROUND_MS` todos row + +`CLAUDE_CODE_MCP_AUTO_BACKGROUND_MS=0` for the whole session. Wall clock taken from `date` +stamps immediately before and after each call: + +| Pass | Start | End | Duration | Returned | +|---|---|---|---|---| +| 1 | not stamped | not stamped | **not timed** | foreground, `success: true` | +| 2 | 20:07:53 | 20:15:31 | **458 s** | foreground, `success: true` | +| 3 | 20:20:27 | 20:29:37 | **550 s** | foreground, `success: true` | +| 4 | 08:12:55 | 08:25:32 | **757 s** | foreground, `success: true` | +| 5 | 08:32:40 | 08:43:43 | **663 s** | foreground, `success: true` | +| 6 | 08:46:56 | 08:59:56 | **780 s** | foreground, `success: true` | + +Five timed calls, all far past 120 s, none auto-backgrounded, each returning an ordinary +envelope the hook could read. + +**No control run was made** with the variable unset, so this is a correlation observed under +one setting, not a demonstration that the variable is what held the calls in the foreground. + +**Correction, recorded because it is the round's own subject matter:** the spec cited this note +for four durations including "671 s". This note recorded only passes 2 and 3 at the time, and +663 s is the correct figure for pass 5 — 671 was arithmetic error. Caught at Gate-A spec +pass 6, in evidence attached to a rule about unsupported coverage claims. diff --git a/.context/codex-reviews/gate-a-spec-resume.md b/.context/codex-reviews/gate-a-spec-resume.md new file mode 100644 index 0000000..4c69b21 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-resume.md @@ -0,0 +1,90 @@ +# Gate-A (spec) resume note — review-loop-economics + +Cycle-stable, per CLAUDE.md §5's optional-companions rule. Written 2026-08-29 because the +session is about to be compacted. Advisory: nothing depends on it existing, and the repo state +plus the pass files are authoritative wherever this disagrees. + +## Where things stand + +- **Branch:** `review-loop-economics`. Working tree clean at `c513094`. +- **Spec:** `docs/superpowers/specs/2026-08-28-review-loop-economics-design.md`, **revision 11** + (`c513094`), 648 lines. +- **Story:** `docs/superpowers/stories/2026-08-28-review-loop-economics-pass-floor-story.md`, + **eleven** acceptance criteria. Profile `risk high · security none · + battery+check+verification` — **read it from the header, never from here**. +- **Slot discriminator: `rle`.** Never write a bare `gate-a-spec-pass-N` slot: doing so once + destroyed a previous cycle's findings file in this very cycle. + +## Gate-A spec loop — where the counter actually is + +**Pass 10 has RUN and its findings are UNPROCESSED.** A hold arrived to pause after pass 9, but +the pass-10 call was already in flight and returned; the file is valid and on disk. + +| Pass | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | +|---|---|---|---|---|---|---|---|---|---|---| +| findings | 27 | 30 | 54 | 40 | 33 | 34 | 32 | 33 | 28 | **27** | +| Blockers | 5 | 2 | 0 | 4 | 0 | 3 | 9 | 4 | 2 | **2** | +| Blocker/Major | 24 | 22 | 43 | 31 | 30 | 26 | 29 | 28 | 22 | **23** | + +Files: `.context/codex-reviews/gate-a-spec-rle-pass-{1..10}.md`, all validated (terminator exact, +counts matching, no non-finding lines). **These files are the durable record of this loop** and +are gitignored — do not clear `.context/` without extracting them. + +**Floor is long met. NOT clean** — 23 Blocker/Major at pass 10. Zero tells at pass 9; pass 10's +tells are **not yet computed**. + +## Next step, concretely + +1. **Read `gate-a-spec-rle-pass-10.md`** — 2 Blocker, 21 Major, 2 Minor, 2 Nit — and work the + Blocker/Major. Validate each finding against the tree before applying; Codex is advisory. +2. Compute and report the **three lines** (trend, cluster, require↔withdraw) — the duty is active + from pass 4 onward, and **two tells make stop-and-surface mandatory, not discretionary**. +3. Revise, commit, then **pass 11**. + +**Order of operations, learned the hard way — four occurrences:** *issue the gate call first, then +write the report.* Never write "running pass N" in a message unless the call is already in flight. +No turn ends on an announcement. + +## The pass call + +`mcp__codex__exec`, `workingDirectory` = repo root. Delete the target and confirm it gone first. +The instruction must **open with the `superpowers:brainstorming` directive**, carry the intent, the +settled decisions as INPUTS, the **risk-high lens set once** (threats, abuse, rollback, data loss, +idempotency, compatibility, observability), a mechanical-verification demand (line citations, +quotes against **both** §5 copies, stated counts against their own enumerations, re-run §6.1's +grep), coverage-first with `NO FINDINGS` as the clean signal, and the file protocol writing to +`.context/codex-reviews/gate-a-spec-rle-pass-

.md`. Reply is one line. + +Earlier pass instructions are recoverable from this session's transcript; the shape above is what +matters. + +## Standing rules for this cycle + +- **Contract questions route to the sparring session `dev-workflow-kit-56`**, never to a dialogue. + Daniel is reached only through it. Spec and plan approval are delegated to that session + (Daniel, 2026-08-28); **§5's named human confirmations are not** — profile changes, + STOP-and-surface closures, human exceptions, and any contract question the sparring session + cannot ground in Daniel's recorded decisions. +- Record decisions as **"sparring session, under Daniel's 2026-08-28 delegation"** — never imply + Daniel reviewed something he did not. +- **Owed at the final clean pass:** an explicit **coverage statement**, including the spec's + claims about the `/workflow-init` mirror. Those were verified mechanically on 2026-08-29 — + template fenced 192–778 with §5 at 257–777, the prose-exemption rationale confirmed **absent** + from the template (that is §7's prerequisite, not a defect), all 21 §6.2 passages present + exactly once in the template, rows 20 and 21's citations confirmed in both copies. **Re-verify + before asserting**, since the spec has changed since. +- After Gate A closes: `superpowers:writing-plans`, then Gate A on the plan, then implementation. + +## Open, and not to be lost + +- **The conditions artifact** — `…-review-loop-economics-conditions.md` — is **not yet written**. + It carries the row-by-row kept/moved/dropped dispositions for §6.2's twenty-one passages, + produced **once against frozen final text** and gated by its own Gate-A review before any + replacement text is written. +- **Prediction ledger and the four announce-then-idle occurrences** live in + `docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md`. The §6.2-split prediction scored + partly right; the "well below 20 Blocker/Major" prediction was **not met** at 22 and is recorded + as not met. +- **The lesson worth keeping**, if any of this reaches the field record: a restructuring guard + asking "did a decision move?" misses the case where **a rule survives in outline and loses its + force** — pass 9 found eight of those. And a grep for the phrasing you expect is not a check. diff --git a/.context/codex-reviews/gate-a-spec-resume.stopped-2tier-debt.md b/.context/codex-reviews/gate-a-spec-resume.stopped-2tier-debt.md new file mode 100644 index 0000000..65d84f3 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-resume.stopped-2tier-debt.md @@ -0,0 +1,68 @@ +# Gate A — spec cycle — RESUME NOTE (cycle-stable) + +Cycle: reviewer-availability fallback, **two-tier** design (tier 1 + tier 3 human exception). +Story: `docs/superpowers/stories/2026-08-13-reviewer-availability-fallback-story.md` +Spec: `docs/superpowers/specs/2026-08-14-reviewer-availability-fallback-design.md` +Both committed at `6c3e175`. Working tree clean as of this note. + +## State + +**Pass 1 complete and validated** — 30 findings (7 BLOCKER, 21 MAJOR, 2 MINOR), none +dismissed. File: `gate-a-spec-pass-1.md`. Dispositions with every call recorded: +`gate-a-spec-pass-1-dispositions.md`. + +**Floor NOT met.** Minimum three passes with a clean final pass. Two more passes minimum, +and the pass-1 fixes are not yet applied to the spec. + +The previous three-tier cycle is archived under `.stopped-3tier` (passes 1–3 + dispositions ++ its resume note). Older cycles are under `.pre-2026-08-14`. Do not reuse those slots +without archiving first — this cycle already destroyed one predecessor's artifacts before +that was noticed. + +## Next action: apply pass-1 fixes to the spec, then run pass 2 + +All 30 are accepted. Decisions already taken (do not re-litigate): + +1. **F1 + F19** — tier 3 is prohibited once any pass returned unremediated Blocker/Major; + completed passes are preserved and named in the marker; §5's "clearly stuck → STOP and + surface" is never remediable by tier 3. +2. **F2** — §4 keeps its "no sentinel, no invariant amendment" conclusion but **drops the + Gate-B STOP from the justification**, and states the unchanged hook supplies no waiver + control. (The STOP may not fire: docs-only exemption, prior counts, and it reaches the + agent not the human.) +3. **F3** — debt cancellation requires a **different** accountable handle from the waiver + authorizer, plus stated risk acceptance. +4. **F4** — an open/unreconciled debt row **blocks a subsequent tier-3 waiver** in that + repository. +5. **F5 (human-confirmed)** — §5 requires an interactive pause for an explicit human + response immediately before closure, named as **the decision-question pattern this + workflow already runs**, not a new invention. §13 states plainly that nothing verifies + the pause happened or that the recorded handle answered. **Escalation path recorded:** + signed commit / protected-branch approval, *trigger — the first tier-3 record whose + authorization is disputed or unattributable.* +6. **F9 + F10 + F11 + F30** — canonical cycle-id format; debt-row grammar with stable key + and status; multi-story markers carry a list of row keys; a resolvable story citation + is a tier-3 precondition; repayment uses the **closure-time** profile, referenced by an + immutable story-commit identity, never a copied value. +7. **F18** — rebuild §8 from §5 **in document order**, uniquely identifying each atomic + condition. Named omissions to restore: stop-and-surface, do-not-manufacture-findings, + validate-before-applying, `workingDirectory` binding, branch-file acceptance and resume + deletion, hook result-envelope limits, no-story and malformed-profile branches, + per-story aggregation, counterfactual evidence, human-confirmed profile changes, + `baseSha`/WIP mechanics. + +Majors 6, 7, 12–17, 20, 23–27, 29 and minors 8, 28 are each accepted with a stated fix in +the dispositions file — apply as written. + +## Story is already fixed + +Findings 21 and 22 are **done** and committed: tier 3 is stated as new (not an extension of +§2.13), §1 is reframed around closure rather than weaker review, and all five open questions +are individually marked settled or moved. Do not re-amend for those. + +## Standing riders for every pass + +Mechanical sweep before each read pass; unioned risk+security lens sets appended once with +abuse carrying both labels; severity enum `BLOCKER|MAJOR|MINOR|NIT`; findings to file with +the exact terminator; delete the target and confirm gone before each call; one recovery +attempt per pass; do not let the reviewer read `.context/codex-reviews/`. diff --git a/.context/codex-reviews/gate-a-spec-resume.stopped-3tier.md b/.context/codex-reviews/gate-a-spec-resume.stopped-3tier.md new file mode 100644 index 0000000..9baf0da --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-resume.stopped-3tier.md @@ -0,0 +1,45 @@ +# Gate A — spec cycle — RESUME NOTE (cycle-stable) + +Cycle: reviewer-availability fallback ladder. +Story: `docs/superpowers/stories/2026-08-13-reviewer-availability-fallback-story.md` +Spec: `docs/superpowers/specs/2026-08-14-reviewer-availability-fallback-design.md` + +## State at interruption (2026-08-14) + +**STOPPED at pass 3 under CLAUDE.md §5's stuck condition. The floor is NOT met — no pass +was clean, so nothing here is approved and the spec must not proceed to `writing-plans`.** + +Passes 1, 2, 3 all ran, all validated structurally (terminator, count, grammar, enum), all +returned Blocker/Major. 116 findings, **zero dismissed**. Findings and dispositions: +`gate-a-spec-pass-{1,2,3}.md` and `gate-a-spec-pass-{1,2,3}-dispositions.md`. + +Blockers rose 4 → 4 → 6 while the spec grew 281 → 448 → 522 lines. + +## Why it stopped, in one line each + +- Rider (a) double-counts the hook's pass counter at **tier 1** — a hook change, not a §5 + prose edit. +- Debt repayment reviews contaminate the live cycle's counters and fingerprint. +- The tier-2 trust boundary is unachievable in-repo: repository `CLAUDE.md` loads into any + custom agent and there is no per-agent switch. +- The workspace-global `.off` posture is an invariant-2 violation rather than a residual. + +## Awaiting + +A scope decision from Daniel. The recommendation on the table is to split: the ladder, +rider (a), and the tier-2 reviewer surface are at least three specs, and two need hook or +harness work the "prompt-only" decision excludes. + +## Not yet done, if the cycle resumes as-is + +The pass-2 revision **dropped** the packaging/backlog section: story AC 8's occurrence-3 +append, the 0.8.2 → 0.9.0 bump, the CHANGELOG entry, and the parked rows are all missing +from the current spec and must come back regardless of how the scope decision lands. + +## Repo state + +Story committed (`e420420`), amended in the working tree for AC 6 — **uncommitted**. +Spec committed at its pass-0 text (`3c5712c`); the working tree holds the pass-2 revision — +**uncommitted**. Prior cycles' 44 `gate-a-spec-*` artifacts were preserved under a +`.pre-2026-08-14` suffix; pass 1 of this cycle destroyed the previous cycle's +`gate-a-spec-pass-1.md` and its dispositions before that preservation was put in place. diff --git a/.context/codex-reviews/gate-a-spec-resume.stopped-tier3-core.md b/.context/codex-reviews/gate-a-spec-resume.stopped-tier3-core.md new file mode 100644 index 0000000..79fdd4d --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-resume.stopped-tier3-core.md @@ -0,0 +1,103 @@ +# Gate A — spec cycle — RESUME NOTE (cycle-stable) + +Cycle: reviewer-availability fallback, **stripped** design — tier 1 + tier 3 human exception, +no debt machinery. +Story: `docs/superpowers/stories/2026-08-13-reviewer-availability-fallback-story.md` +Spec: `docs/superpowers/specs/2026-08-14-reviewer-availability-fallback-design.md` +Base commit `df850ab`; **pass-1 fixes are now applied in the working tree, uncommitted.** + +## State + +**Pass 1 complete and validated** — 28 findings (4 BLOCKER, 20 MAJOR, 4 MINOR), none +dismissed. `gate-a-spec-pass-1.md` + `-dispositions.md`. + +**All 28 pass-1 fixes applied** (2026-08-14). Spec 489 → 750 lines. What changed: + +- §2 recast as a **cost/trust trade-off** with a three-row alternatives table (F1); "every + reader sees" → discoverable via `git log --grep` (F26). +- §3.1 rebuilt: complete-pass-ledger fail-closed rule (F10), fresh profile resolution as a + common precondition (F18), a **waiver checklist** carrying the lens sets, invariants and + the 12 prompt standards (F2, F17), Gate-B per-story modes/entries (F28); the mechanical + sweep relabelled **syntax hygiene, not evidence**. +- §3.2 rebuilt: **canonical gate call** definition + defective-call exclusion (F3), and a + full fresh-observation → enum transition table with a STOP row (F4). +- §3.4/§3.5: complete prospective record presented; **authorization digest** over tree *and* + message (F8); dedicated index + **committed-tree comparison** as step 7 (F7); positively + determined docs-only Gate-A tree (F6); **15-minute max age + post-answer probe** (F5). +- §4: **Gate-A continuation rule** at `executing-plans`, state-free and re-derived (F9); + "authoritative" removed (F27). +- §5: `Authorized-tree`, `Profiles` (paths only), `Waiver-checklist` fields; `gate:` added to + the decision block (F11); **new §5.4 field grammar** — handles, ISO-8601 UTC, escaping, + 200-byte bound, byte-level normalization (F12); §11 → §10 xref (F24). +- §6: **ordinary-merge chain and validation matrix** beside the squash chain (F13). +- §7: rows 33, 39, 44, 55–58, 61, 62, 65 corrected (F15–F19). +- §8: `docs/pr-review-bots.md` ~141 and `coding-workflow.md` ~253–268 added (F20, F21); + marketplace.json row **re-pointed at the top-level `description`** (F22 — see below). +- §10: the **prompt-differential probe** specified concretely (F23) — driver, frozen + `OLD`/`NEW` blob shas, S1 + 8 negative variants + S2, 3 runs all-must-agree, hard failure + semantics, no new repo file. **The watch item is resolved; no whole-mode override needed.** +- §13: rollback split three ways — this repo, version-keyed caches, scaffolded copies (F25); + new residual naming the checklist as self-audit. +- Story: §2 "re-review debt" → "untracked re-review obligation"; §5 "profile scales the + repayment" overturned (F14), both with fourth-amendment accounting. + +**One deviation from the recorded dispositions, deliberate:** F22 said drop the +`marketplace.json` row because the file does not carry the two-gates claim. The mechanical +sweep found that inspected the *plugin entry* (line 14); the **top-level `description`** +(line 4) does read "independent Gate A/B review". Took F22's stated alternative — re-point +the row — rather than dropping it. + +**Pass 2 complete and validated** — 29 findings (14 BLOCKER, 13 MAJOR, 2 MINOR), none +dismissed, **all applied**. Spec 750 → 1016 lines. Enum shrank to two sources; §3.5 became an +eight-step close with expected-parent binding, `read-tree` init, changed-path classification +and a three-way read-back; §4 gained a second continuation point; §5.4 gained parity-aware +escaping and sha256 digests; §5.5 the checklist record; §6 an incident-invalidation +convention; §7 rows 82–87; §10 split into a mechanical oracle and a pinned-model probe. + +**Pass 3 complete and validated** — 34 findings (15 BLOCKER, 18 MAJOR, 1 MINOR), none +dismissed, **none applied**. + +## ⛔ CYCLE STOPPED 2026-08-14 — §5 stuck condition, surfaced to Daniel + +Blockers by pass: **4 → 14 → 15**, artifact 489 → 750 → 1016 lines. Three pass-3 blockers are +the **unenforceable/recursive** pair that ended both prior cycles, now landing on tier 3's +core rather than on a removable sub-feature: + +- **7** — both surviving outage sources are still locally manufacturable (exhaust the quota, + point at a spent account, block DNS); a self-approving human author therefore has a + sanctioned zero-review closure on demand. Codex's own fix says to acknowledge the abuse path + that cannot be closed. +- **2** — in the motivating outage, `Passes completed: none` can never be *established*, so + the waiver can never be granted in the case it exists for; loosening it reopens the + lost-pass gate-off path. +- **3** — the governing-story union cannot be reconstructed, because prompts are not durable. + +Plus **27**: two rounds of fixes turned a prose amendment into a git algorithm whose +highest-risk logic §10 exercises not at all, so the mode's evidence is unsatisfied for a risk +path the fixes themselves created. + +**Do not run pass 4 without a decision from Daniel.** Full triage in +`gate-a-spec-pass-3-dispositions.md`. + +Two earlier cycles are archived and must not be overwritten: `.stopped-3tier` (three-tier +design, passes 1–3) and `.stopped-2tier-debt` (two-tier with debt machinery, passes 1–3). +Older cycles are under `.pre-2026-08-14`. **Archive before reusing any slot** — this story has +already destroyed one predecessor's artifacts through slot collision. + +## Why this cycle looked different, and then did not + +After pass 1 this read: blockers three-tier 4 → 4 → 6 (stopped); two-tier+debt 7 → 8 → 12 +(stopped); stripped **4** — *"and none of the four says the mechanism cannot exist."* That +held through pass 2 and **failed at pass 3**, where findings 2, 3 and 7 put the +unenforceable/recursive pair on tier 3's core. The earlier optimism is kept here rather than +deleted, because it was the reasoning that justified spending passes 2 and 3, and a future +reader deciding whether to restart should see what it was based on. + +## Standing riders for every pass + +Mechanical sweep before each read pass (it has now caught a wrong finding total, a stale +cross-reference, a wrong `marketplace.json` field and a bare `CHANGELOG.md` path); unioned +risk+security lens sets appended once, abuse carrying both labels; severity enum +`BLOCKER|MAJOR|MINOR|NIT`; findings to file with the exact terminator; delete the target and +confirm gone before each call; one recovery attempt per pass; the reviewer must not read +`.context/codex-reviews/`. diff --git a/.context/codex-reviews/gate-a-spec-rle-pass-1.md b/.context/codex-reviews/gate-a-spec-rle-pass-1.md new file mode 100644 index 0000000..620d49e --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-rle-pass-1.md @@ -0,0 +1,28 @@ +BLOCKER | high | §4 and story §§2–3 | the dropped docs-only arm remains normative in the story: desired outcome 1 promises docs-only floor 1 and acceptance criterion 8 requires a docs-only or trivial floor-1 cycle on this branch, while the fresh profile is risk high/security none and therefore requires floor 3 | the story can be accepted only by violating the settled mapping or by citing another story, whose aggregation is itself undefined | replace both docs-only references with the exact profile-zero predicate and make the end-to-end demonstration use a cycle whose complete cited-story set actually licenses floor 1 +BLOCKER | high | §6 "Old-conditions accounting" | the accounting covers only twelve floor-wording sites, not the scope-stop, clearly-stuck, two-tell, three universal-duty, severity, human-exception, squash-carry, profile, and commit-closing passages that parts 2 and 3 must rewrite | implementation can silently drop old conditions despite directly claiming compliance with the AGENTS.md decision-procedure invariant | enumerate every rewritten passage and mark each pre-existing condition kept, moved, or deliberately dropped before implementation +BLOCKER | high | §4 "Mechanism and its residual" and §4.1 | no lifecycle rule writes or removes `.context/codex-gate.floor` when a cycle starts, the profile changes, the next cycle has floor 3, a branch is switched, or the policy is rolled back | a floor 1 left by one trivial cycle persists into a later high or unprofiled cycle and makes the hook miss two required passes, violating the loose-in-the-firing-direction invariant | require deriving and synchronizing the knob to 1 or 3 before every pass and closure check, verify the stored value, and define cleanup on cycle end, branch change, and rollback +BLOCKER | high | §2.1 versus CLAUDE.md Mechanics "Scope, and it is narrow" | explicit decline is designed to let a surfaced finding stop blocking continuation or closure, but existing Mechanics says neither human assent nor its record may answer a STOP, an unclean final pass, or let an agent close or continue a cycle | the two shipped instructions would issue opposite actions for the same declined scope finding | include both Mechanics copies in old-condition accounting and distinguish the new specific-decline procedure from the optional human-exception form without weakening other mandatory stops +BLOCKER | high | §4 and story §5 "Which floor governs a cycle citing several stories" | the spec leaves the story's explicit multi-story floor question unanswered even though current §5 supports multi-story cycles | an agent may choose the lowest cited floor and under-review a cycle that also cites a high, standard, or unprofiled story | define the cycle predicate deterministically so floor 1 is available only when the complete cited-story set satisfies the settled zero rule and every other set yields 3 +MAJOR | high | §4 and §10 "gate-off lever" | disclosure and an agent-written closing line are the only checks on the same agent that chooses the floor; neither the gate prompt nor a pass report exposes the parsed axes, derived floor, or actual knob value before closure | an agent seeking to skip review can write 1, omit or forge the later provenance line, and leave no independent observation in the diff or review artifact | within the prompt-only scope, require each pass report and gate context to state the story path, freshly read axes, derived floor, and observed knob value, with mismatch stopping before the pass counts +MAJOR | medium | §4 mechanism | `.context/codex-gate.floor` is one checkout-global mutable value and the spec states no serialization rule for concurrent Gate-A, Gate-B, or subagent cycles with different profiles | the last writer can change another cycle's closure threshold and both cycles can report internally plausible counts | forbid concurrent cycles in one checkout or define isolated state and an ownership check that makes cross-cycle writes detectable +MAJOR | high | §3 "The rule" versus "The decision procedure" | the first rule categorically demotes narration, mechanism prose, and instrument internals by subject, while the next rule says severity is determined by an in-system reader's changed decision rather than artifact kind | reviewers receive two incompatible classifiers and can use the categorical sentence to demote consequence-changing findings | make the named-reader-and-changed-decision test the sole classifier and present artifact kinds only as non-normative examples that are Minor when that consequence test fails +MAJOR | high | §3 lines 112–121 | false green is the sole instrument exception, but a false red, a test that blocks a valid release, or instrument logic that drives an unnecessary product rewrite also changes an in-system decision | consequential instrument defects can be forced to Minor even though the settled rule is consequence-based | retain severity for every instrument defect that changes an in-system decision, using false green as one example rather than the only exception +MAJOR | high | §3 "The boundary case" | the blanket claim that rationale prose in rule files never flips a decision contradicts prompt-standards item 6, which requires rationale precisely because models follow and scope rules better when the why is present | a wrong why-clause can alter whether an agent applies a rule, yet the spec mandates demotion without running its own reader test | apply the same named-reader-and-changed-decision test to rationale prose and remove the categorical Minor result +MAJOR | high | §3 "Expected effect" | the claim that PR #23's scratch-harness findings would demote almost entirely is unsupported by `7bbdb14`, whose body describes numerous harness defects that could make checks pass for wiring reasons—the proposed false-green carve-out | the spec promises an economic effect that its own exception preserves and has not classified finding by finding | classify the cited 27 findings with the proposed consequence test or delete the projected demotion claim +MINOR | high | §3 field-record rationale | the spec says two of three fic2 pass-5 story-criterion findings were parked and one was acted on, but the field report says all four pass-5 findings were parked and identifies the acted-on activation-boundary condition as a pass-4 item | the factual example used to justify consequence-based severity misstates its cited evidence | say that the three pass-5 story findings were parked while a separate pass-4 activation-boundary finding on the same artifact was acted on +MAJOR | high | §2.1 "Explicitly declined" | the spec calls the decline durable but names the optional, advisory, gitignored `-dispositions.md` companion as its natural home; existing §5 allows that file to be deleted or rebuilt and it disappears on a fresh checkout | the load-bearing permission can be lost, rewritten, or unavailable, causing either repeated stops or an unverifiable close | define a required durable schema and destination for the user's decision, plus squash carry and recovery behavior, and keep the optional companion only as a cache +MAJOR | high | §2.1 "specific finding" | findings have no stable identifier and the spec gives no rule for recognizing the same declined finding after rewording, relocation, splitting, merging, or recurrence | an agent can treat a materially new Blocker as the declined item, or re-ask the same item forever, so the binding is neither safe nor idempotent | define a cycle-stable finding key and state that ambiguity, splits, and consequence changes create undecided findings rather than inheriting a decline +MAJOR | high | §2.1 and existing scope/action rule | only the decline branch is defined; if the user accepts an out-of-set Minor or Nit, it enters the assigned set while the action rule says collect and never iterate and the surfaced-finding duty keeps the pass unclean | the cycle has no legal next action and cannot close or repair without violating one of the rules | specify the accepted-expansion transition for each severity and identify which user decision, new cycle, or repair obligation resolves the surfaced finding +MAJOR | high | §5.1 "Q6's durable half" | a curve written only in the closing commit body cannot restore history to pass 4 of the same still-open cycle after a fresh checkout, cleared `.context/`, or machine move | the claimed durable half does not mitigate Q6's in-flight availability gap and overstates what the commit record proves | describe the closing curve only as post-close observability, or define a pre-close durable record that the resumed pass actually reads +MAJOR | high | §5 Q6 unavailable-history rule | the proposed diagnostic names a few total-loss examples but no checks or distinct remedies, and it omits partial, malformed, stale, unreadable, and concurrently deleted pass history | agents can silently compute tells from an incomplete curve or report generic degradation without resolving the cause, contrary to prompt-standards item 10 | define availability per tell, enumerate distinguishable causes with checks and fixes, and require partial history to be reported rather than treated as complete or wholly absent +MAJOR | high | §10 final risk bullet | "the new ones bind afterwards" gives no activation boundary for cycles already in flight when the prompt version changes | banked passes, earlier severities, open dispositions, and closing duties can be evaluated under different rule sets with no reproducible answer | state whether a cycle pins the rules at start or migrates on the next pass, and define how banked passes, declines, reports, and provenance transition +MAJOR | high | §7 and §9 rollout scope | updating the inline template changes new or manually reinitialized projects only; existing initialized downstream `CLAUDE.md` copies do not update when the plugin cache version changes | most installed users can keep the old three-pass and severity rules indefinitely while believing the plugin update applied them | specify the upgrade path, CHANGELOG instruction, and idempotent `/workflow-init` merge behavior, including the warning for an active cycle +MAJOR | medium | §5.1 versus story acceptance criterion 2 | the spec says closing commit bodies generically carry curves, while the story limits the requirement to the closing commit of a Gate-B cycle | implementers cannot tell whether Gate-A spec and plan cycles must record counts, and the economics record can omit the very Gate-A loops cited in the problem | state the settled scope explicitly in the spec and both prompt copies, including where each covered cycle's closing body lives +MAJOR | high | §5.1 pinned count form | the form does not define whether a full Gate-B pass sums its two branch files, records branches separately, deduplicates repeated findings, excludes incomplete attempts, or writes zero for a zero-finding pass | curves from different authors are not comparable and the two-tell trend can change with an arbitrary counting convention | define the source files and aggregation rule for every accepted, incomplete, recovered, branch-specific, and zero-finding pass +MAJOR | high | §5.1 and existing Mechanics squash rule | the existing squash carry names only evidence entries and human-exception records; the new count curve and non-default-floor provenance are not added to it | squash merging discards the records the spec calls durable, so main-history economics become a self-selected subset again | extend and account for the squash-carry rule in both copies to include every new closing-body record +MAJOR | high | §9 scope and invariant 12 | the design changes `plugins/dev-workflow/commands/workflow-init.md` but names no manifest version bump or CHANGELOG update in its implementation surface | the planned change fails the plugin-version invariant or ships a version whose inventory omits the behavior change | add the exact manifest and CHANGELOG obligations to scope and verification while preserving the no-hook-code boundary +MAJOR | high | §8 evidence plan and story acceptance criterion 7 | no verification compares every changed rule across the live §5 and inline template; the named differential check covers only two floor sentences and the invariant checker covers only a narrow severity spelling | one copy can retain a stale exception, precedence, provenance, or migration rule while the planned evidence stays green | require a scoped parity verification over every changed condition and list each deliberate variance explicitly +MAJOR | high | §1 "four ways" and "four standing duties" | the stated counts do not match §5 as a whole, which also has mandatory stops for failed target deletion, exhausted incomplete-pass recovery, unresolvable profiles, and blocking evidence/setup gaps, plus other standing obligations | the advertised consolidated precedence is partial while presenting itself as the complete cycle decision procedure, so interactions with omitted stops remain undefined | either scope the count explicitly to the four local loop outcomes or include every §5 stop and standing duty in the precedence and old-condition accounting +MINOR | high | §8 evidence plan | the plan invokes the battery and two named verifications but never requires a manual pass over all 12 prompt-standards items, even though invariant 11 says the mechanical checks are only a floor | prompt changes can satisfy every named check while failing the repository's actual prompt gate | add an explicit 12-item review of both edited prompt regions and record any deliberate exceptions +MINOR | high | §4 and §10 terminology | a floor of 1 does not turn either gate off: it removes two mandatory passes but still requires a valid clean pass, while `.context/codex-gate.off` suppresses reminders without changing policy | "gate-off lever" overstates the mechanism and obscures which risk is actually being accepted | call it an unverified floor-reduction or under-review lever and state exactly which obligations remain +END OF FINDINGS (27 total) diff --git a/.context/codex-reviews/gate-a-spec-rle-pass-10.md b/.context/codex-reviews/gate-a-spec-rle-pass-10.md new file mode 100644 index 0000000..6934615 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-rle-pass-10.md @@ -0,0 +1,28 @@ +BLOCKER | high | §1 lines 52-60 versus §5 lines 307-340 | §1 forbids closure whenever unavailable prior-pass files prevent proving every earlier in-set Blocker or Major resolved, while §5 makes an absent prior file unrecoverable and expressly says unavailable history is not a stop condition | a resumed cycle with lost history can neither close nor take a defined terminal or recovery path, so it loops forever | define a fail-safe terminal procedure, such as reconstructing resolution from a durable source or abandoning and restarting under rules that preserve the unresolved duties +BLOCKER | high | §2.1 "It appears in every record"; §4.1 provenance grammar; §5.1 curve grammar; §6.2 passage list | the nonce is mandatory in every cycle record, but neither pinned machine-readable form nor the skipped-loop form has a nonce field, and the passage list omits other cycle records such as evidence entries and optional resume or disposition records | records carried through squash cannot be bound to one cycle as promised, and the conditions artifact is structurally unable to account for every passage the nonce rule rewrites | add the nonce to every record grammar and skipped form, enumerate every existing record passage affected, and make the conditions artifact verify those bindings +MAJOR | high | opening lines 9-14 versus §§4.1, 5.1 and 11 | the document still categorically says concrete formats live in the plan, although revision 11 returns two concrete durable formats to the spec | a plan author can treat the grammars as downstream detail despite the later text making them contract | qualify the opening statement with the two durable-interface exceptions or rewrite it to distinguish replacement wording from machine-consumed formats +MAJOR | high | §4.1 provenance grammar lines 256-267 | the grammar has no production for N or a quoted path, and the prose escape rule covers quotes but not backslashes, newlines, or other Git-valid path bytes, so comma, brace, semicolon, quote, and backslash combinations are not uniquely parseable | P8 cannot implement one parser that accepts every repo-relative path the spec claims to cover | define N as the allowed floor tokens and add a complete length-delimited or JSON-style path production with exhaustive escaping +MAJOR | high | story criterion 11 lines 233-236 versus §4.1 lines 256-264 | the required demonstration string `floor 3 per ` is not an instance of the pinned grammar, which requires a braced entry with level plus the hook-threshold clause | satisfying the acceptance criterion literally violates the spec interface, while satisfying the interface violates the criterion's stated example | replace the criterion text with one complete valid provenance line for this story and its actual knob state +MAJOR | high | §4 and §4.1 versus story criterion 1 lines 119-120 | the story requires every pass report to state the floor and the risk and security axes read, but the spec requires only the floor and per-story numeric level and never obliges the report to expose both axis values | an implementation can satisfy the spec while failing the acceptance criterion and hiding an incorrect max-level derivation | require every pass report to name each cited story's risk and security values as read, separately from the durable closing provenance grammar +MAJOR | high | §5.1 curve form lines 351-358 | MODELS gives no delimiter grammar for multiple pass-model entries, no representation for several contributing models in one logical pass, and no escaping or token grammar for model names; SPEC and the repeated count lists are descriptive rather than productions | examples required by the surrounding properties cannot be parsed deterministically, especially split Gate-B passes assembled from different models | pin complete productions for pass sets, count vectors, model vectors, multi-model logical passes, separators, escaping, and vector-length agreement +MAJOR | high | §5.1 "Each of the three loops" versus §6 conditions-artifact gate | §6 adds a separately reviewed Gate-A artifact with its own floor and clean-pass acceptance condition, creating another paid review loop, but the curve taxonomy and story criterion 2 account only for the main spec, plan, and Gate-B loops | the new gate's cost disappears from the P8 measurement whose purpose is to measure Gate-A economics | either define the conditions review as a named fourth curve-bearing loop with an unambiguous label or explicitly justify and accept its exclusion in the story and P8 contract +MAJOR | high | §2.1 nonce constraint lines 130-141 | collision-resistant is not made operational: the allowed length includes four base-36 characters, generation entropy is unspecified, and no atomic reservation or collision check exists | concurrent cycles can select the same infix, misidentify each other's files as their own, and overwrite or combine records despite the refuse-on-foreign-target rule | require a fixed high-entropy generation method and length plus atomic collision detection and regeneration before any record is written +MAJOR | high | §2.1 nonce recovery lines 137-141 | recovery names working state and commit history but gives no rule for disagreement, multiple candidate nonces, a malformed source, or one source pointing at a different reviewed revision | a restart can inherit the wrong cycle and make an old decline suppress a new finding, or needlessly discard valid passes | require agreement where both sources exist and define fail-safe source validation, precedence, ambiguity handling, and new-cycle behavior +MAJOR | high | §2.1 working record and §5/§6.2 concurrent-cycle rule | only findings-slot names gain the nonce infix and collision refusal; the advisory working record that carries the nonce and live declines, plus fixed-name resume and disposition companions, remain able to collide between concurrent cycles | one cycle can overwrite the very state used to recover another cycle's identity or decline decision | give every cycle-owned working artifact the nonce namespace and the same refuse-rather-than-overwrite semantics, or prohibit and detect same-kind concurrent cycles before any write +MAJOR | high | §6 lines 397-403 | the conditions artifact must earn a clean pass at the derived floor "like any other," but the universal loop rule retains a zero-finding early exit below the floor | reviewers cannot tell whether a zero-finding first pass approves the accounting artifact or whether this special gate requires reaching the numeric floor | state explicitly whether the universal zero-finding exit applies and make the wording agree with the ordinary loop contract +MAJOR | high | §6 conditions-artifact gate lines 397-416 | the gate has no feedback rule for a review finding that changes the passage list, a disposition, the final replacement text, or this governing spec | replacement text can be written after the conditions artifact becomes clean while the spec or planned final text it was checked against has changed without renewed Gate A | require affected artifacts to be updated and their relevant Gate-A loops rerun before the no-replacement gate can open +MAJOR | high | §10 unknown-start fallback lines 572-579 | floor 3 is called the stricter reading, but an old-rule loop whose user knob is a valid value above 3 can owe more than 3 | an unknown-start fallback can reduce a pre-existing obligation and close early while claiming to fail safe | compare all plausible old and new obligations and use the strictest, or discard the ambiguous loop and start a new cycle under one known rule set +MAJOR | high | §10 unknown-start fallback versus §§2.1-5 | the fallback claims to cover every touched part but maps only floor, severity, suspensions, declines, and curves; it leaves nonce recovery, slot collision handling, provenance and pass-report duties, Q6 disclosure, hook precedence, cited-set changes, and closure-resolution evidence undecided | partially adopted or origin-unknown loops receive a hybrid procedure with unsafe and contradictory defaults | enumerate a fallback for every new normative obligation and reconcile it with §2.1's separate rule that no recoverable nonce starts a new cycle +MAJOR | high | §10 rollback lines 594-600 | the in-repo rollback says the activation rule governs loops in flight after a revert, but reverting the shipping commit removes the very prompt text that tells an agent to preserve the old rules | an active loop can silently switch procedures or become origin-unknown during rollback, invalidating banked passes and records | require the revert to carry a temporary migration instruction or require all active cycles to stop and restart under an explicitly selected rule set +MAJOR | high | §4.1 named hook residual lines 237-240 | the residual names only the redundant-warning direction at hook line 933, but hook lines 947 and 967 also call the knob-derived ratio a hard minimum, and lines 956 and 973 can emit a satisfied checkmark below the derived floor when the knob is lower | the design discloses the safe false-red case while omitting the dangerous false-green message that can encourage early closure | disclose both directions and every affected message class in both shipped copies, emphasizing that even a hook checkmark does not satisfy the derived floor +MAJOR | high | §4.2 cited-set removal lines 292-303 | the spec lets a cited story be removed and the floor lower but never states that accepted in-set Blocker or Major findings discovered while that story was cited remain repair obligations | deleting a citation can be read as deleting the unresolved duty and become a gate-off path | state that cited-set changes only recompute profile-derived obligations and never discharge already accepted in-set findings +MAJOR | high | §1 lines 52-58 | the closing report is said to establish each finding's resolution "from the per-pass findings files," but those files contain the finding and suggested fix, not the applied resolution or evidence that it worked | the spec overstates what the gate artifact proves and leaves the closure check's actual evidence source undefined | name the artifact, diff, disposition, or verification that establishes resolution and describe the check as author-performed and unverified where that is the truth +MAJOR | high | §10 gate-off routes lines 587-593 versus §2.1 lines 124-128 | the list is labeled routes known today but omits the explicitly known fabricated-decline route, which can release a real hold and is mechanically unchecked | the abuse inventory understates a newly introduced gate-off surface even while claiming to enumerate known routes | add fabricated or altered decline records and nonce misbinding to the known routes while retaining the non-exhaustive disclaimer +MAJOR | high | story §4 affected invariants | the story omits prompt invariants 9 and 10 even though the spec explicitly relies on invariant 9 for downstream adoption and invariant 10 for the open reader-kind list | Gate A and implementation can satisfy the story's declared invariant surface without checking two governing constraints the change touches | add both invariants with the concrete spec obligations they govern +MAJOR | high | story criterion 5(b) versus §3 lines 155-158 | the criterion asks whether something takes a different decision but does not separately require naming the consuming act, so it is weaker than the settled two-part test | a shipped copy can pass the acceptance criterion after naming only a changed decision, recreating the exact demotion loophole revision 11 repaired | require both an in-system consuming act and the decision that act changes, and say failure to name either imposes the Minor-or-below ceiling +MAJOR | high | story criteria 5(f) and 6 versus §2.1 lines 92-123 | no acceptance criterion checks that a decline is a distinct record type from a human exception, is available only to a scope-stop finding, and never qualifies the in-set Blocker or Major resolution duty | the plan can collapse the two records or turn decline into a general waiver while still satisfying durability and Q1-Q6 at a high level | add explicit bidirectional acceptance checks for record-type distinctness, scope-stop-only availability, and the unqualified in-set resolution duty +MINOR | high | §11 lines 614-625 | revision 11 says all eight weakened details are repaired but identifies only the both-required severity defect and the two returned formats, with no eight-item enumeration | the restructuring audit cannot mechanically verify its own count or show that each repair was restored rather than merely mentioned elsewhere | enumerate the eight pass-9 defects with their repaired section and disposition +MINOR | high | §6.2 row 13 | the `Gate B — Code` citations are `:335/:519`, but the headings are at `CLAUDE.md:331` and template `:515`; the cited lines instead describe why fixes invalidate review | the required mechanical citation check fails and a later conditions author can inspect the wrong passage | correct the pair to `:331/:515` +NIT | high | opening line 25 and §§8, 9 references to 192-line divergence | a conventional zero-context diff of the cited §5 regions currently has 96 removed plus 90 added lines, 186 changed lines, not 192 | repeated quantitative context is stale and undermines later parity claims | recompute the number from a stated command or remove the unstable count +NIT | high | story criterion 5 line 175 | the criterion introduces "Four properties" but enumerates five severity properties, (a) through (e), plus a separate record property (f) | the criterion's own stated count is false and complicates mechanical acceptance accounting | say five severity properties plus one record property, or renumber and regroup the list +END OF FINDINGS (27 total) diff --git a/.context/codex-reviews/gate-a-spec-rle-pass-11.md b/.context/codex-reviews/gate-a-spec-rle-pass-11.md new file mode 100644 index 0000000..4eb68c5 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-rle-pass-11.md @@ -0,0 +1,39 @@ +BLOCKER | high | §1 lines 54-63 | the resolution list says a recorded decline can establish resolution for an in-set Blocker or Major even though §2.1 and the settled decision forbid declining any in-set Blocker or Major | an agent can treat the scope-only decline mechanism as a waiver of the mandatory resolve duty and close with a serious finding unfixed | remove decline from the resolution sources for in-set Blocker and Major findings and state that it only classifies a scope-stop finding out of set +MAJOR | high | §1 lines 57-60 | findings files are said to establish which findings are in-set, but their six-field protocol contains severity and finding text, not scope membership or the user's accept-or-decline decision | a closing inventory reconstructed from findings files alone can silently include declined findings or omit accepted ones | say the files establish the raw finding inventory and require the advisory disposition or decision record plus the assigned fix set to establish membership +BLOCKER | high | §1 lines 66-73 | the new unavailable-history stop has no state-changing human answer: every answer merely resumes the same cycle while the missing inventory still prevents closure, and an in-set finding cannot be declined | a lost findings file leaves the cycle permanently unable either to close or to take a defined conservative restart or abandonment path | define the allowed human decisions and their effects, including how inventory is reconstructed or how a new cycle starts without silently discharging unresolved findings +MAJOR | high | story criterion coverage versus §1 lines 54-75 | no acceptance criterion requires the new closing resolution inventory, the inability-to-establish-resolution stop, or its report of unreadable passes and unaccounted findings | implementation can satisfy all eleven criteria while omitting revision 12's central terminal-path repair | add a criterion that exercises closure with complete history and with missing history and checks the exact stop, disclosure, and resumption behavior +MAJOR | high | §4.2 lines 338-345 | the set-removal rule preserves only an accepted finding said to have entered through a user's answer, excluding ordinary in-set Blocker and Major findings that never needed a scope decision | removing a cited story can be read as discharging a serious finding discovered inside the original assigned set | state that every already in-set Blocker and Major survives cited-story removal regardless of how it entered the set +MAJOR | high | §4 lines 225-230 and §4.1 line 307 | the iff rule says floor 1 when every cited story is profiled at level 0, which is vacuously true for an empty cited set, while the worked no-story instance assigns floor 3 | two conforming readers can derive different floors for a no-story cycle | require a non-empty cited set for floor 1 and classify an empty set explicitly as unprofiled floor 3 +MAJOR | high | story criterion 1 lines 119-121 versus §4.1 lines 242-260 | the criterion requires every pass report to state the floor, both raw axes, and the stories supplying them, but the spec requires only the floor in each pass report and defines story levels only for the closing provenance line | the spec can be implemented while pass-time derivation remains unauditable and criterion 1 fails | add a pass-report contract carrying each cited path with its risk and security values plus the resulting floor +MAJOR | high | §4.1 provenance grammar lines 283-309 | the repeated-entry production contains a literal comma with no following whitespace production, but worked instances two and three contain comma-space | two of the five instances do not parse under the pinned grammar they claim to demonstrate | change the production to include the literal separator comma-space or remove the spaces from every instance and parser +MAJOR | high | §4.1 path grammar lines 292-299 | any non-bare repo path is promised a quoted representation, but quoted paths only escape backslash and quote and therefore cannot encode a legal path containing a newline or other line-breaking control character in this one-line record | the claimed single form cannot represent every valid cited story path and a generated record can become multi-line and unparsable | define a complete escaped lexical grammar for control characters or explicitly reject unrepresentable cited paths with a stop condition +MINOR | high | §4.1 KNOB production and required properties lines 293 and 327-328 | positive integer is broader than what the hook accepts because the hook's shell numeric comparison rejects values outside its integer range and silently keeps the default | provenance can record a huge digit string as the hook threshold even though the hook treats the file as unusable | define KNOB usability by the hook's observed parse result and record unusable whenever its numeric test does not accept the value +MAJOR | high | §2.1 lines 145-166 and §5.1 lines 400-415 | the spec says the nonce leads both pinned forms and appears in every cycle record, but the curve form starts with Loop and the skipped-loop form contains no nonce at all | curves and skip records cannot be attributed to one cycle and the every-record rule is unsatisfiable | prefix both curve variants with cycle NONCE and update the grammar, examples, squash carry, and P8 parser contract +MAJOR | high | §5.1 MODELS production lines 404-413 | model identifiers permit only letters, digits, dot, underscore and hyphen, but this repo already records actual provider-qualified models such as moonshotai/kimi-k3 in commit baa75c1 | the required never-recalled actual model cannot be represented by the pinned grammar | support slash and every character the tool can report through a quoted and escaped model token rather than narrowing real identifiers +MAJOR | high | §5.1 curve grammar lines 400-407 | the machine-parsed form never defines n or p, uses a prose ellipsis instead of a repetition production, and leaves comma-and-range syntax, ordering, overlap, and count-to-pass arity informal | independent writers and the P8 parser can accept different records or map counts to the wrong pass numbers | provide complete productions for nonnegative counts and canonical nonoverlapping pass lists and require Findings, Blockers, model entries, and expanded SPEC to have equal arity +MAJOR | medium | §5.1 line 415 | a generic rule says any Loop variant can be recorded as skipped without naming the existing authorization that makes a skip legitimate | an agent can mistake a syntactically valid Gate-A skipped record for permission to bypass a Gate-A loop even though the story says Gate A is never skippable | restrict the skipped form to the Gate-B triviality rule or explicitly state the only inactive-gate case that permits each Gate-A skip +MAJOR | high | §2.1 nonce generation lines 145-161 | the nonce has a randomness requirement but no behavior when a uniform random source is missing, unreadable, or returns an invalid or short value | an environment failure invites an ad hoc timestamp or commit-derived fallback that the same section forbids, or leaves cycle startup undefined | require validation of the generated nonce and stop and surface when a compliant nonce cannot be produced +MAJOR | high | §2.1 nonce recovery lines 155-161 | candidate is not scoped to a particular active artifact or history position, so normal history containing prior cycle nonces can yield multiple candidates or disagree with the current advisory record | recovery can restart a valid cycle repeatedly or attach a decline to the wrong historical cycle | define exactly which commit bodies and working-record path are candidate sources, their precedence, and how closed prior cycles are excluded +MINOR | high | §6.2 row 18 and lines 529-537 | the slot grammar is called optionally infixed and says bare slots remain valid for a single-cycle case, while every new cycle must generate a nonce and every record it writes must carry it | a literal reader can write a new bare slot and violate cycle attribution despite following the local naming rule | reserve bare names explicitly for legacy pre-activation recovery and require the nonce infix for every cycle governed by this design +MAJOR | high | §2.1 lines 168-172 versus codex-gate.sh lines 758-764 | the required property is phrased in terms of the resulting commit message beginning WIP, but the hook never reads the resulting message and recognizes only a command string containing a -m argument followed by wip | an amend using --no-edit or -F preserves a WIP subject yet the hook treats it as closure and resets the pass counters | specify the hook-recognizable command-input shape as a contract requirement or change scope to make the hook inspect the actual commit message +MAJOR | high | §5 lines 367-382 | absent, partial, malformed, unreadable, and stale are not mutually exclusive and no precedence is defined, such as a partial file from another cycle or a stale malformed file | one prior-pass artifact can select incompatible computations and remedies, defeating the diagnostic-state requirement | define a deterministic classification order and the single report and remedy for every overlapping combination +MINOR | high | §5 lines 373-375 | full disk is named as a cause of an unreadable existing file even though it normally prevents or truncates a write and therefore belongs to absent, partial, or malformed state | the diagnostic sends the user to the wrong branch and leaves actual read failures such as I/O errors unspecified | place ENOSPC under write failure or malformed recovery and enumerate real unreadable causes with distinct checks +MAJOR | high | §6 lines 454-479 | frozen final text is neither identified by commit or tree nor rechecked immediately before replacement, and feedback is triggered only by review findings rather than intervening source changes | the sole old-condition guard can pass against one pair of §5 copies and then authorize rewriting different text after a rebase, concurrent edit, or later spec repair | bind the conditions artifact to hashes of both source regions and fail before replacement unless both still match, then revalidate the accounting against the completed replacement +MAJOR | medium | §6 conditions-artifact gate lines 454-466 | the new conditions artifact has its own Gate-A loop but the spec does not define whether it is a Gate-A spec cycle, which nonce and slot namespace it uses, or how its pass counter is separated from the already-run spec and plan loops | its mandated clean floor and own curve can inherit prior passes or collide with another Gate-A loop while still appearing compliant | name its cycle type and lifecycle explicitly, including fresh nonce, pass numbering, findings slots, provenance, and curve label +MAJOR | high | §10 unknown-start fallback lines 637-649 | the list claimed to cover every touched part omits provenance, nonce and identity recovery, per-pass derivation reports, Q6 degraded-history reporting, and the new closure-resolution inventory | an unknown-start loop has no conservative behavior for several revision 12 obligations despite the completeness claim | enumerate every new obligation and give the strict fallback for each, or remove the completeness claim and define a single conservative restart procedure that covers them +MAJOR | high | §8 risk verification lines 591-596 | recomputing a floor from the stories listed in the provenance line never checks that the listed set equals all stories the artifact actually cites | the verification passes the named gate-off route where a high-risk story is omitted and a falsely low floor is internally consistent with the incomplete line | derive the expected citation set independently from the artifact and compare membership before recomputing levels and floor +MAJOR | high | §8 prompt-standards verification lines 604-606 | eleven checklist items are scoped to changed prompt regions and only item 7 is applied to the whole result, while invariant 11 requires each changed prompt to pass all twelve items | global failures in target model, output structure, stop conditions, duplicated rules, or calibrated emphasis can be declared passed because they sit just outside the edited spans | review both complete resulting prompt artifacts against all twelve items and use changed-region notes only as supplementary evidence +MAJOR | high | §8 prompt-standards obligation and current two §5 copies | the spec never requires a target-model statement in CLAUDE.md or in the scaffolded CLAUDE template, and neither resulting prompt currently names its executing target as checklist item 1 requires; the command's own line 9 does not travel into the generated file | the planned prompt-standard pass cannot truthfully pass and invariant 11 remains violated in the shipped template | add a target-model statement and model-page revalidation obligation to both generated CLAUDE surfaces +MAJOR | high | §3 lines 206-208 versus §4 lines 235-238 and current prose exemption | the spec calls the consequence test a consistency check for the path-level prose exemption, yet its own hardening-log example proves a docs Markdown file can be operational input while the current exemption categorically gives docs Markdown no Gate B | the claimed common principle yields opposite treatment for the same operational text and can skip review of behavior-driving documentation | either narrow the path exemption by operational consumption or stop claiming kinship and account explicitly for the mismatch +MAJOR | high | §10 gate-off routes lines 657-666 | the known-today list omits the new author-written closure inventory and resolution assertions, which nothing verifies and which can falsely claim an unresolved in-set Blocker or Major was fixed | a newly created silent-close route is absent from the disclosure intended to stop readers mistaking the design for guarded | add fabricated or incomplete resolution accounting and membership accounting to the known routes and state that no mechanism checks them +MAJOR | high | §10 rollback lines 667-671 | the rollback paragraph says any loop crossing a revert lands in unknown-start, contradicting the preceding activation rule that a loop with establishable starting rules finishes under them | a recoverable in-flight cycle can be needlessly restarted or inconsistently switched merely because a revert occurred | make unknown-start conditional on failure to establish the original rules and state how nonce and recorded revision prove the known-start case across a revert +MAJOR | high | §4.1 lines 256-260, story criterion 11, and passive-metrics story lines 34-52 | the spec says the first eligible floor-1 cycle is P8's first checkpoint and that P8 reads it, but the named P8 story contains only the after-roughly-three-cycles economics question and no floor-1 provenance checkpoint | the deferred demonstration has no durable owner and can be omitted while both this story and P8 appear complete | add the floor-1 checkpoint and its trigger to the P8 story and include that file in this change's surface +MAJOR | high | story criterion 8 lines 224-229 | the criterion says CI enforces both the plugin version and CHANGELOG update, but check-version-bump.sh checks only the manifest version and no invariant script references CHANGELOG | a missing changelog entry can pass CI while the acceptance criterion reports mechanical enforcement | either add a real changelog consistency check with tests or state that only the version bump is CI-enforced and make changelog review-backed +MINOR | high | story criterion 5 line 176 | the criterion announces four properties but enumerates six lettered requirements from a through f | count-based review and later references can omit two requirements while appearing to satisfy the stated total | change four to six or separate the record requirement and state the resulting count accurately +MINOR | high | opening Surfaces lines 24-27 | the stated 192-line divergence is not reproduced by the cited ranges; a line-based no-index diff reports 95 deletions and 99 additions, 194 differing lines | parity evidence begins from a false mechanical baseline and later reductions can be misreported | define the counting metric and update the number to the reproducible result +NIT | medium | §6.2 row 19 versus §5.1 lines 438-441 | row 19 says three record types are added to squash carry, while the target paragraph separately names decline records, provenance, curves, and skipped-loop records | the accounting table understates its own disposition count and invites one carry obligation to be dropped | call it four record types or explicitly define the skipped-loop record as the curve variant being counted +MINOR | high | §11 lines 690-723 | the section says narration of earlier-revision errors was dropped, but §4.1, §6, and §11 itself retain extensive narration of revision 10, pass 2, pass 8, pass 9, and earlier errors | the restructuring summary makes a falsifiable scope claim that the current document contradicts | narrow the claim to the specific historical narration actually removed or delete the remaining revision-history prose +MAJOR | high | story-to-spec agreement for criterion 5f and §2.1 | criterion 5f covers durable cycle-bounded declines but does not require uniform nonce generation, every-record propagation, source disagreement handling, multiple-candidate handling, or conservative identity loss | the implementation can satisfy the criterion while omitting the operational properties that are supposed to prevent cross-cycle decline reuse | extend the criterion with observable cases for generation failure, recovery from each source, disagreement, multiple candidates, and nonce presence on every record form +MAJOR | high | story-to-spec agreement for criterion 9 and §6 | criterion 9 requires only that accounting exists and covers rewritten rules; it does not require the new accounting artifact's clean-floor gate, own curve, frozen-input binding, or feedback and spec-reopen rules | the conditions artifact can meet the story while bypassing the safeguards that justify deferring dispositions out of the spec | add acceptance checks for the artifact gate, its curve and provenance, input revision binding, regeneration, and spec Gate-A reopening +MINOR | high | §4.1 STORY-SET production lines 288-325 | the grammar calls the value a set but permits duplicate paths and even conflicting levels for the same path, with no canonical ordering | a parseable provenance line can represent no coherent cited-story state and different P8 implementations can deduplicate differently | require each normalized path exactly once with one level and define deterministic ordering +END OF FINDINGS (38 total) diff --git a/.context/codex-reviews/gate-a-spec-rle-pass-12.md b/.context/codex-reviews/gate-a-spec-rle-pass-12.md new file mode 100644 index 0000000..05fa840 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-rle-pass-12.md @@ -0,0 +1,25 @@ +BLOCKER | high | §1.1 and §4 pinned curve | the spec says the nonce leads every pinned form and that every curve carries it, but both the ordinary curve and skipped-loop forms begin with the loop label and contain no nonce | P8 cannot attribute a curve to a cycle, and the curve cannot satisfy the retained nonce contract | prefix both curve variants with the same `cycle ;` production and add matching examples +MAJOR | high | §4 curve grammar | the purported machine grammar leaves `` and `

` undefined, defines `` only in prose, gives no ordering, overlap, or count-cardinality grammar, and supplies no complete valid curve instance | independent P8 implementations can accept different records or misassociate counts with passes | define every terminal and list constraint formally and include full single-model, per-pass-model, gapped-pass, multi-model, and skipped examples +MAJOR | high | §4 `` production | `[A-Za-z0-9._-]+` rejects provider-qualified model identifiers such as `moonshotai/kimi-k3`, which this repository already records in commit history, despite requiring the value exactly as reported | a valid gate pass can have no valid curve representation, forcing omission or falsification of the reviewer model | make model values quoted and escaped, or define a delimiter-safe encoding that preserves every reported identifier including slash and plus +BLOCKER | high | §1.1, §5.2, and §10 split accounting | the retained global rule that the nonce appears in every cycle record still rewrites the human-exception record, squash carry, `baseSha` and WIP closing-body block, evidence-entry record, and optional companions, yet all five passages were removed from the current twelve-passage accounting and assigned only to the successor; §1.1 simultaneously says the optional-companion rules are unchanged | parts 1 and 2 would ship unaccounted changes to old decision procedures and depend on successor work despite claiming independence | either narrow nonce scope to the provenance and curve records or restore every still-touched passage to this story as a shared passage and account for all old conditions now +BLOCKER | high | §5.1, §5.2, and §10 clearly-stuck split | §5.1 correctly identifies the digit-free pass-1 sentence in the clearly-stuck paragraph as false under floor 1 and changing, but §5.2 omits that paragraph and §10 assigns it solely to the successor, whose pass floor is expressly out of scope | the floor implementation rewrites a passage whose clean-precedence and surfacing conditions are accounted only in a later story, exactly the dropped-condition failure the governing Don't forbids | put the clearly-stuck paragraph in both stories' shared passage lists and perform the floor-related old-conditions accounting in this story +MAJOR | high | §7 parity verification | the promised parity scope covers §5.1, §5.2, and §§2–4 but omits the nonce generation, recovery, working-record, and collision rules in §1.1 | the two shipped copies can diverge on cycle identity while the named parity verification still passes | explicitly include every §1.1 rule and every passage it reaches in the parity verification +MAJOR | high | §4 squash-carry seam and successor story §§3–4 | the parent says the successor must add the decline record without dropping provenance, curves, or skip records, but the successor has no settled input or acceptance criterion requiring decline records to survive squash or preserving the parent's additions | the successor can replace the shared carry sentence and erase the only durable inputs P8 reads | add a settled successor input and acceptance criterion that the carry list is extended, never replaced, and includes both stories' record types +MAJOR | high | §9 unknown-start seam and successor story | the parent says the successor extends the unknown-start fallback, but the successor neither names that shared passage nor requires preservation of the floor, severity, and curve entries | the later consolidation can replace the fallback and silently remove this story's stricter behavior | add the activation passage to the successor's shared-passage inventory and require additive extension with all parent entries preserved verbatim in meaning +MAJOR | high | §10 and successor story §5 | the spec says retaining the nonce settles the successor's question, while the successor still lists whether the nonce belongs as an open question and says it may introduce its own | a paid decision is explicitly reopened and two incompatible nonce scopes or formats can be designed | move nonce ownership to the successor's settled inputs and state that it consumes the parent's nonce without redefining it +MAJOR | high | successor story §4 decisions 7–9 | the successor claims every paid-for decision is carried, but its decline-record inputs omit the five-field identity and sameness test, uncertainty behavior, decision date, squash survival, advisory working transport, and unverified-assertion boundary that revision 12 had already settled | the successor design must rediscover or can silently drop the very constraints the split promises not to reopen | copy each settled record property into the successor's settled-input table, leaving only wording and layout to its design +BLOCKER | high | §9 unknown-start fallback | the complete-list sentence is syntactically broken at `every and`, and its supposedly exhaustive list omits the retained nonce, provenance, recovery, working-record, and collision duties | a loop with unknown starting rules has no determinate handling for cycle identity and records, so the activation acceptance criterion is not implementable | repair the sentence and enumerate the strict fallback for every retained duty, including generation or abandonment of cycle identity and whether each record is owed +MAJOR | high | §9 downstream partial adoption | the surviving text says taking the severity rule without the ordering violates §1's coupling argument, but ordering is successor-owned and §1 couples severity to the profile-derived floor | this is an orphan from part 3 that contradicts the reduced scope and falsely makes parts 1 and 2 depend on the successor | replace `ordering` with the actual coupled floor rule and describe the concrete partial-adoption consequence for floor plus severity only +MAJOR | high | §9 rollback | the rollback paragraph says every loop spanning a revert lands in the unknown-start case, while the activation rule immediately above says a running loop uses its starting rules whenever those rules can be established | known-start loops receive two conflicting procedures and may unnecessarily restart or apply the wrong rule set | make unknown-start conditional on inability to establish the start version; otherwise preserve the stated started-under behavior across a revert +MAJOR | high | §3.2 cited-set changes | the spec categorically calls adding a story a raise, removing one a lowering, and says lowering drops floor, lenses, and evidence mode together, which is false for multi-story sets where another story can keep each dimension unchanged and evidence modes are per-story rather than one winning mode | agents can spend or release passes and evidence obligations on a set change the aggregation rules do not license | require one new pass under the changed set, then independently rederive the floor, unioned lenses, and each remaining story's evidence obligations; state that any dimension may stay unchanged +MAJOR | high | parent story criterion 1 versus spec §3 and §3.1 | the story requires every pass report to state the derived floor, both axes read, and the source story for each, but the spec requires only that the floor be stated and reserves story levels for the closing provenance line | implementation can satisfy the spec while failing the parent acceptance criterion and leaving in-progress derivation unauditable | add the axes and cited-story source fields as required per-pass-report properties in both copies +BLOCKER | high | P8 handoff and §§3.1, 4, 6 | the named consumer `2026-08-04-passive-metrics-over-the-ledger-story.md` has no acceptance criterion for review-loop metrics or the floor-1 checkpoint and still expressly says required curves are undecided; it never names either pinned grammar | the consumer used to justify pinning does not actually promise to parse the records, so the economics measurement and first floor-1 observation can be dropped while every listed criterion passes | add P8 to the affected surfaces and update its desired outcome and acceptance criteria with both grammars, nonce grouping, the first eligible floor-1 checkpoint, and the self-reported-data limit +MAJOR | high | §3.1 `unusable` knob state and prompt standard 10 | the new diagnostic value merges unreadable, empty-after-whitespace, zero, non-digit, numeric-overflow, and other parse failures without enumerating checks or user-owned remedies | agents and P8 cannot distinguish an intentional invalid value from an unreadable artifact, and the shipped prompt fails the binding diagnostic-state checklist item | define the hook-equivalent classification, name how each cause is distinguished, keep the agent's action as leave untouched and disclose, and direct any repair to the user +MINOR | high | parent story §2 named out of scope | calling part 1 a prompt-only policy `over` the floor knob conflicts with the settled rule that derivation never reads or acts on that knob | a plan reader can infer that the knob participates in the obligation despite the later corrective criterion | say the prompt-only floor policy is independent of the existing user knob, which remains only the hook reminder threshold +MINOR | medium | §1 coupling argument | the hardening-log example proves that a path-derived docs arm is wrong, but it does not prove floor and severity must ship atomically because the floor rule can carry that rationale without the severity procedure being present | the stated reason does not support retaining both parts in one already costly change and makes the partial-adoption risk hard to state accurately | either name a concrete inconsistent state produced by shipping only one part or describe the relationship as shared rationale rather than atomic coupling +MINOR | high | reduction cross-references throughout spec and parent story | renumbering left stale targets: intro §7 should be §6; curve §5.1 should be §4; residual §10 should be §9; story criterion 6 should be 8; §4.2 should be §3.2; kinship §3 should be §2; story criterion 7 should be 9; the §7 seam should be §6; judgement §3 should be §2; parent `Section A` should point to spec §2; and parent criterion 8 plus §8 verification should be criterion 10 plus spec §7 | readers and the implementation plan are sent to unrelated rules, including the wrong acceptance criteria for accounting and parity | update every target after the reduction and rerun the cross-reference scan +MINOR | high | intro surfaces count | the two cited §5 ranges produce 95 deleted-side and 99 added-side diff lines, 194 total, not the stated 192 | the spec's quantified baseline and later parity claims are mechanically false | replace 192 with the verified metric and define the counting method, or remove the volatile count +NIT | high | §5.2 paragraph after the table | `Row 18` survived renumbering even though the table now has rows 1–12 and the naming paragraph is row 11 | the explanation points at a nonexistent row and obscures which old conditions it qualifies | change it to row 11 +NIT | high | §5.2 Gate-B passage citation | the row quotes `Gate B — Code` but cites lines 335 and 519, while that heading is at 331 and 515; the cited lines are only the later fixed-three rationale | the mechanical citation check does not land on the quoted passage | cite 331 and 515, or relabel the row as the `where the 3 come from` sentence +MINOR | medium | §5.2 row 5 and twelve-passage claim | the table is explicitly a list of passages rewritten by this change, but the Gate-B triviality-skip row says it is merely adjacent and checked for consistency without specifying any changed rule | the claimed twelve rewritten passages and the conditions-artifact scope are not falsifiable from the design | either state the exact floor-related edit to the skip passage or remove it from the rewrite count while retaining it as an explicit consistency check +END OF FINDINGS (24 total) diff --git a/.context/codex-reviews/gate-a-spec-rle-pass-13.md b/.context/codex-reviews/gate-a-spec-rle-pass-13.md new file mode 100644 index 0000000..1cf409d --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-rle-pass-13.md @@ -0,0 +1,33 @@ +BLOCKER | high | §9 unknown-start fallback, lines 531-540 | the normative sentence is syntactically incomplete at "every and the curve duty treated as owed", so one strict fallback duty has no object | an agent that cannot establish the starting rules cannot determine the safe fallback and may omit an obligation on the gate-off path | restore the missing object and enumerate each strict fallback duty in a grammatical list +BLOCKER | high | §3 lines 122-127 and current CLAUDE.md lines 406-422 | "everything else" yields floor 3 without preserving the existing rule that a present but unresolvable profile stops and surfaces rather than falling back as unprofiled | a malformed or contradictory high-risk profile can be treated as an ordinary floor-3 cycle and proceed without its required lenses and evidence | state that profile resolution runs first, that malformed profiles still stop, and that only no-story or no-profile cases are unprofiled +BLOCKER | high | §3.2 lines 243-263 | a cited-set change that leaves the unanimity verdict unchanged is not required to receive a further pass under the current set | adding a security-high or otherwise differently-profiled story while the floor remains 3 can close on a clean pass that never saw that story, its lenses, or its evidence obligations | require the final clean pass to run against the current cited set after every set change, independently of whether the derived floor changes +BLOCKER | high | §4 lines 269-325 and P8 criterion lines 77-86 | the pinned curve records total findings and Blockers only, while the handed-off question asks whether consequence-keyed severity changes finding distributions | Major-to-Minor demotions, the principal expected effect, leave both recorded series unchanged, so P8 cannot answer the accepted handoff | record counts for all four severities or at least Blocker and Major separately, and define the exact distribution P8 will compare +BLOCKER | high | §3.1 lines 142-146, §5.2, and current pass-report paragraphs at CLAUDE.md lines 133-146 and template lines 330-343 | this change adds floor, axes, and source-story fields to pass reports but the eighteen-passage accounting omits the existing pass-report passage and §10 assigns it only to the successor | the parent can rewrite or contradict a live decision procedure without either accounting covering its existing three-line carrier, activation, and tell duties | add the pass-report passage to both accountings and make the successor extend the parent's fields rather than replace them +BLOCKER | high | §5 lines 335-369 | the conditions artifact is not required to account for the two existing copies separately even though the spec establishes that they differ on 192 diff lines | one combined disposition can preserve the root rule while silently dropping a template-only condition, violating the named AGENTS.md decision-procedure invariant | require a per-copy old-condition inventory, or an explicit union after mechanically diffing each listed passage, before parity rewriting +BLOCKER | high | §5.2 row 18, §10 lines 596-600, and successor story lines 105-120 | the parent says the baseSha and WIP closing-amend block is rewritten by both changes, §10 calls it successor-only, and the successor's exhaustive three shared seams omit it | the successor can replace the parent's provenance and curve carry while still satisfying its own split instructions | decide ownership consistently and add this block as a shared extend-not-replace seam in the successor if both changes really touch it +BLOCKER | high | §9 rollback rule, lines 561-568 | activation wins when the start is determinable, but no mandated first-pass artifact records the governing rule revision or a trustworthy start point; the nonce records identity, not rule version | a loop spanning a revert cannot mechanically choose activation versus strict fallback and may finish under the wrong floor or severity rule | record the governing policy commit or equivalent immutable rule-version identifier at cycle start and define the comparison used on recovery +BLOCKER | high | §§1.1, 3, and 4 versus current CLAUDE.md lines 215-222 | the spec treats Gate-A spec, Gate-A plan, and Gate B as loops of one cycle without saying whether they share one nonce, while current shipped text explicitly calls the two Gate-A loops separate cycles and Gate B a third | implementations can generate one nonce or three, breaking slot isolation, history recovery, and P8 grouping in mutually incompatible ways | define the cycle boundary and nonce lifetime explicitly and account for the current three-cycle terminology in both copies +MAJOR | high | §1.1 lines 62-67 | history is a nonce recovery source, but the spec gives no rule for selecting candidates from a repository containing many prior cycle records | normal history will contain multiple nonces, so recovery either always starts a new cycle or can attach to an unrelated one depending on an invented selector | define the exact commit range, story identity, loop label, and uniqueness test used to select the historical candidate +MAJOR | high | §5.2 lines 422-430 | the slot infix is called optional and bare slots remain valid, yet every activated cycle must have a nonce and the same paragraph says every cycle with a nonce uses the infix | agents can choose the bare form under the claimed single-cycle exception, violating the every-record nonce rule and reopening concurrent overwrite races | reserve bare slots explicitly for pre-activation legacy cycles and require nonce-infixed slots for every cycle started under the new rules +MAJOR | high | §10 lines 596-600 | the text claims eleven successor-only passages but its enumeration names only eight passage units and includes clearly-stuck and baseSha passages that §5.2 says the parent also rewrites | the mechanically asserted split count cannot be reproduced and masks missing or misassigned accounting rows | enumerate all eleven individually and label shared passages separately from successor-only passages +MAJOR | high | §5.2 row 18 and lines 422-430 | the paragraph introduced as "Row 18's change" describes optional findings-slot names and collision refusal, which are rows 11 and 12, not row 18's baseSha and closing-body block | the baseSha block's old conditions and actual replacement obligations remain unspecified despite appearing covered | relabel the paragraph for rows 11 and 12 and add a separate row-18 accounting description +MAJOR | high | §6 lines 434-439 and §5.2 | the spec explicitly changes the template's "What counts as prose" passage by adding the confused-reader rationale, but that existing passage is absent from the eighteen-passage accounting | template-specific conditions in the edited exemption can be lost while the conditions artifact still claims completeness | add the prose-exemption passage as a template-side accounting row and state the deliberate one-copy-to-parity change +MAJOR | high | cross-references at lines 26, 217, 222, 337, 412, 436, 493, 520, and 576 | nine references point to the wrong section or acceptance criterion: §7 for the §6 docs inventory, §5.1 for the §4 curve, §10 for §9 routes, criterion 6 for criterion 8, nonexistent §4.2 for §3.2, §3 for §2 kinship twice, criterion 7 for criterion 9, and §7 for the §6 seam | plan authors and the conditions gate are directed to unrelated requirements, making claimed coverage non-reproducible | correct every reference and mechanically recheck the criterion numbering after edits +MAJOR | high | §3.1 lines 196-197 and 225-236 | unreadable, empty, non-numeric, and out-of-range are named but no mutually exclusive production procedure, precedence, or hook-equivalent normalization is defined; the hook removes all whitespace before testing | the same file, such as internal whitespace, a signed value, or an overflowing decimal, can produce different provenance causes and a record that does not describe what the hook does | specify byte-level checks in order, including whitespace removal, numeric domain, overflow behavior, and the distinct fix for each cause +MAJOR | high | §3.1 grammar lines 186-198 | the machine grammar leaves "positive integer" undefined and permits any positive derived floor even though the only valid floors are 1 and 3 | parsers cannot decide lexical validity consistently and a floor 2 record is grammatically accepted despite being impossible under the contract | define the integer token exactly and constrain the derived floor production to 1 or 3 +MAJOR | high | §3.1 examples, lines 205-213 | "one instance per variant" is false after revision 14 because only unusable(non-numeric) is shown; unreadable, empty, and out-of-range have no instances | three newly split diagnostic states can be encoded incorrectly without any example exposing it | add a valid instance for every unusable cause and narrow the claim to the variants actually enumerated +MAJOR | high | §4 lines 276-304 | §4 supplies no concrete main-form curve instance whose SPEC, counts, and model coverage can be parsed, and the skipped form is stated outside the grammar with no alternate production | the claimed fully defined machine interface cannot be tested character by character or checked against prompt-standards item 4 | add main-form examples covering singleton, ranges, gaps, per-pass models, and mixed-model calls, add the skipped production, and show matching count cardinalities +MAJOR | high | §4 model grammar, lines 286-300 | the model token allows plus even though plus separates contributing models, and it has no quoting form for a verbatim reported identifier containing another excluded delimiter | pass 1 foo+bar has two parses, while identifiers containing spaces, colons, commas, semicolons, or parentheses cannot be recorded verbatim as required | exclude plus from the bare token and define a quoted escaped model form with an unambiguous delimiter grammar +MAJOR | high | §4 lines 296-300 | per-pass model entries must cover every SPEC pass, but the contract does not require exactly one entry per pass, forbid extra or duplicate pass numbers, or fix their order | two parsers can assign different models to the same pass while accepting the same record | require the per-pass keys to equal the SPEC expansion exactly once each in ascending pass order +MAJOR | medium | §4 skipped-loop rule, lines 302-304 | "legitimately skipped" does not say that this record form grants no permission and the grammar permits every loop label, including both Gate-A loops that current policy does not allow to skip | an agent can treat existence of the skip form as a new Gate-A bypass | state that only an independently applicable skip rule authorizes the record and identify which loop types currently have one +MAJOR | high | §3.2 lines 258-263 versus current CLAUDE.md lines 435-448 | a lowering is said to drop "the evidence mode" even though multi-story cycles deliberately have one mode and evidence entry per profiled story, and "high ... lowered to trivial" omits the security axis required for level 0 | lowering or removing one story can be read as discharging unrelated stories' evidence or granting floor 1 while security remains nonzero | describe per-story evidence removal and use the exact level-0 condition of risk trivial plus security none +MAJOR | high | §§3 and 3.1 | "cited story" is never defined even though the artifact also links historical, successor, and measurement stories that do not all govern the cycle | agents can derive different floors and provenance sets from the same artifact, or omit a higher-risk governing story as merely contextual | define the authoritative governing-story set and how it is supplied to every gate call, distinct from incidental document links +MAJOR | medium | §1 lines 37-41 and §2 lines 103-105 | the spec uses the system-consumed hardening ledger to prove docs paths can change behavior, then calls consequence severity a consistency check for the unchanged path-level prose exemption that categorically exempts that ledger | the two policies classify the same operational text oppositely, so the rationale cannot guide a reviewer and hides an existing gate gap | state that the path exemption is a coarser accepted exception rather than an analog, or stop and scope a separate alignment decision +MAJOR | high | §7 evidence plan | no verification parses the two pinned interfaces or exercises their semantic cardinality, quoting, cause, and model variants even though P8 depends on them mechanically | malformed records can pass the story's evidence gate and make the later read-only measurement fail after deployment | add a differential parser verification with valid and invalid fixtures for every provenance and curve production +MAJOR | high | §3.1 lines 142-146 and story criterion 1 lines 122-124 | the contract constrains what a pass report contains but never requires a report after each pass; current shipped text starts mandatory reporting only at pass 4 | the derived floor can remain invisible during the floor-1 pass where the spec says checking is cheapest, while the acceptance wording is still satisfied vacuously | require one outer-agent pass report after every valid pass and define how unprofiled and no-story axes are represented +MAJOR | high | §7 prompt-standards verification and docs/prompt-standards.md item 1 | the changed root CLAUDE prompt and its scaffolded CLAUDE template still contain no target-model statement, a known existing condition, while the spec requires a fresh twelve-item pass and provides no exception or repair | invariant 11 cannot truthfully pass after this prompt change | add the target model to both shipped prompt copies or record and approve a checklist exception in the governing standard +MINOR | high | §3.1 quoted-path rule, lines 200-203 | a quoted path permits arbitrary whitespace but recognizes escapes only for backslash and quote, so a legal Git path containing a newline cannot fit the promised single provenance line | an exotic story path makes the durable record unrepresentable or splits it into multiple lines that P8 cannot parse | define an escape for line breaks and controls or explicitly reject such story paths with a stop condition +MINOR | medium | §4 curve semantics, lines 281-291 | the grammar does not require each Blocker count to be less than or equal to its corresponding finding count | internally impossible curves are syntactically valid and can contaminate P8 without a parser error | add the per-pass cross-field invariant and an invalid example +MINOR | high | §3.1 knob recording | the contract does not state when the mutable user knob is sampled for the closing provenance line | a user edit between body generation and commit can leave history claiming a reminder threshold the hook did not use at closure | pin the sample point and require regeneration or disclosure if the knob changes before the closing commit +MINOR | high | §3.1 residual hook citation, lines 165-168 | the sentence says all three cited messages render a ratio, but codex-gate.sh line 933 renders no ratio and only interpolates the minimum | the mechanical disclosure overstates the exact hook output it claims to have checked | distinguish the no-fingerprint minimum statement from the ratio-bearing floor messages +END OF FINDINGS (32 total) diff --git a/.context/codex-reviews/gate-a-spec-rle-pass-14.md b/.context/codex-reviews/gate-a-spec-rle-pass-14.md new file mode 100644 index 0000000..5777582 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-rle-pass-14.md @@ -0,0 +1,19 @@ +MAJOR | high | §2 “One predicate” and “Unanimity across a cited set” | the unanimity rule is not scoped to a non-empty, fully resolvable cited set: an empty set vacuously satisfies “every cited story”, while “any other set yields 3” can also swallow a set containing a present-but-unresolvable profile | no-story work can be read as floor 1 and a malformed cited profile can silently default to 3, contradicting both settled outcomes | state that floor 1 requires a non-empty cited set whose every member resolves at level 0; no story or any cited unprofiled story yields 3, and any present-but-unresolvable member stops and surfaces +MAJOR | high | §2.4 “A profile or cited set that moves mid-cycle” | a cited-set change that leaves the unanimity verdict unchanged is not required to receive a clean pass under the changed set; for example, adding a high-risk story to a set already at floor 3 leaves the floor unchanged while adding lenses, evidence and review scope | a cycle can close on a pass that never reviewed the newly cited high-risk story | require the cited set to be re-read each pass and require the final clean pass to run against the current set whenever its members change, even when the numeric floor does not +MAJOR | high | §2.4 “downward lowers the floor and releases nothing else” | “nothing else” no longer names what remains bound and conflicts with §5’s current per-story aggregation: removing a story legitimately removes that story’s lenses and evidence duty, while an already accepted in-set Blocker or Major must remain open | the plan must choose between retaining obsolete profile obligations and discharging accepted findings, so a prior normative decision has become ambiguous | say explicitly that obligations are recomputed from the current cited set, but removing a story never discharges an accepted in-set Blocker or Major because user acceptance, not the story citation, put it in the fix set +MAJOR | high | §2.3 and §4; parent-story criteria 2 and 3; P8 criterion 5 | the spec moves both concrete grammars to the plan, but the parent still requires one pinned form and says the spec states which, while P8 says both pinned forms are specified in this design | the parent story is unsatisfiable as written and P8 has two competing candidates for the authoritative parser contract | either restore both grammars here or amend the parent and P8 criteria to make the reviewed implementation plan the explicit authoritative owner of the forms +MAJOR | high | §4 “Both branches of one logical pass” | “same artifact revision” has no identity rule, and the spec no longer says what happens to each branch when the artifact changes between them | the plan must decide whether equality means commit identity or content identity and may credit, discard or combine mismatched branches inconsistently | retain the decision-level rule that revision identity is the tracked reviewed commit and that a mismatch makes the completed branch an incomplete pass while the later branch starts a new pass; leave only the concrete encoding to the plan +MINOR | high | §4 “The model each pass ran under is recorded” | one logical pass may be assembled from several branch calls, but the surviving property does not require every contributing model to be recorded when those calls used different models | P8 can attribute a mixed-model pass to one convenient model and lose the independence evidence the field exists to preserve | require every contributing branch model to be recorded, using undetermined for each contributor that cannot be established +MINOR | medium | §4 “A legitimately skipped cycle records the skip” | the rule does not identify which cycle types can legitimately be skipped, while current §5 permits the behavioral-triviality skip only for Gate B and says Gate A is not skippable | the plan can emit a skip form for a mandatory Gate-A cycle or treat an absent artifact as a skipped cycle rather than no cycle | bind the property to Gate B when §5’s existing two-condition skip applies and state that a nonexistent spec or plan creates no cycle and no skip record +MINOR | high | §2.3 knob field and §8 conditional knob verification | the provenance record has no sampling time for a knob that the user may create, remove or edit mid-cycle, and the byte-identical verification treats a legitimate concurrent user edit as evidence that the agent violated the no-write rule | history can misstate the threshold that produced reminders, or validation can block despite the agent leaving the user-owned file alone | define the provenance value as an explicit close-time snapshot and make the verification report concurrent change as attribution-undetermined rather than claiming the agent wrote it +MAJOR | high | §5 “The advisory working record is a cycle record too” | the text invokes “§5’s per-cycle infix and refuse-on-collision rules”, but neither current §5 copy contains those rules and this slim no longer states when the infix is mandatory, what owns a target, or what must be refused | concurrent or resumed cycles can overwrite one another’s findings and working records while the plan invents collision semantics that were previously normative | keep the exact infix spelling in the plan, but restore the properties that a cycle with a nonce uses it in every shared slot, the bare slot is reserved for the legacy single-cycle case, and a target owned by another nonce is refused rather than overwritten +MAJOR | medium | §5 “Recovery has two sources” | “candidate” is undefined and no lifecycle or eligibility rule separates the active cycle from the many prior nonces normally present in history or from stale optional working records | normal history can always present multiple candidates, causing every resume to mint a new identity and making recovery progressively less effective | define the cycle type, artifact and open-state keys that make a nonce a candidate, plus when the working record is removed or retired at closure +MAJOR | medium | §5 nonce generation | no terminal behavior is defined when randomness is unavailable, produces an invalid nonce, or collides with an existing cycle; the later refuse-on-collision reference only prevents overwrite | an agent can fall back to a timestamp or commit despite the prohibition, or remain unable to start a required cycle with no mandated surface | require that no cycle starts without a valid unique random nonce, retry generation only within a bounded policy, then stop and surface without a deterministic fallback +MAJOR | high | §5 and parent-story acceptance criteria | nonce generation, recovery, advisory-record ownership and collision behavior are normative spec obligations, but none of the parent’s ten criteria covers them; the curve and provenance criteria only require fields in the two closing records | the story can be accepted while the nonce system is missing or unable to preserve identity, violating the requested agreement from spec back to criteria | add a criterion covering nonce creation, one nonce per each of the three cycles, recovery, collision refusal and the bounded record set that must carry it +MAJOR | high | §6 “reviewed before any replacement text is written” versus §§4–5 | the conditions artifact is required to undergo a profiled Gate-A review with a clean-pass acceptance condition, which creates another review cycle even though the spec repeatedly fixes the topology at exactly Gate-A spec, Gate-A plan and Gate B with three nonces | the extra cycle has no defined nonce, provenance, curve or closing body and contradicts the settled three-cycle decision | make the conditions review part of one named existing Gate-A cycle, or classify it as a non-cycle verification whose acceptance semantics are stated without invoking a fourth Gate-A cycle +MAJOR | high | parent story criterion 1 “one derived value governs every loop of the cycle” | the parent criterion still encodes Gate-A spec, Gate-A plan and Gate B as loops of one cycle, although current §5 and revision 15 define them as three separate cycles | the binding criterion can reintroduce one shared nonce or one closing record and makes the topology correction incomplete across in-scope artifacts | change the criterion to say the same derived value governs each of the three separate cycles because each derives from the same cited-story set +NIT | high | parent story criterion 5 opening sentence | it says “Four properties” and then defines five independently mandatory properties, labeled a through e | a checklist reader can treat the fifth property as accidental or make a false count claim | change “Four” to “Five” +NIT | high | parent story criterion 1 “Section A’s reachability test” | the referenced Section A no longer exists in the parent or current slimmed spec | the rationale points to a stale structure and cannot be mechanically followed | replace it with the current spec’s §3 severity-semantics reference +MAJOR | high | parent story criteria 2 and 10; current spec evidence plan; commit 8bc04d6 | the story requires this branch’s actual Gate-A spec and plan commits to demonstrate the pinned curve and floor-3 provenance, but the evidence plan does not require retrofitting those bodies and the current spec commit records only aggregate Gate-A counts with no nonce, floor, story levels, per-pass finding and Blocker curve, or models | the implementation can satisfy every shipped-text rule and still fail the parent story’s end-to-end acceptance on its existing history | add an explicit plan and verification obligation to amend or reconstruct the spec and plan closing commits into the final forms, then verify all three cycle bodies before closure +NIT | high | lines 15–17 hook citation claim | codex-gate.sh line 94 does contain the stated regex, but gate_citation supplies the adoption predicate and the policy citation used by core gate-status messages, not “every reminder’s citation”; result-classification reminders either hard-code section 5 or carry no policy citation | the heading constraint is valid but its mechanism is overstated, contrary to the gate-claim Don’t | say the heading must remain matched because the function detects adoption and builds the policy citation for the core gate-status reminders +END OF FINDINGS (18 total) diff --git a/.context/codex-reviews/gate-a-spec-rle-pass-15.md b/.context/codex-reviews/gate-a-spec-rle-pass-15.md new file mode 100644 index 0000000..a06735c --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-rle-pass-15.md @@ -0,0 +1,14 @@ +MAJOR | high | §2.4 "A profile or cited set that moves mid-cycle", lines 126-130 | "A lowering drops the floor" is false for a profile lowering from level 2 to level 1, where the one predicate still yields floor 3; the paragraph recognizes that raises can leave the floor unchanged but omits the symmetric lowering case | an agent can state floor 1 and close after too few passes even though a current cited profile remains above level 0 | distinguish any profile lowering from a derived-floor change: every profile change requires the final clean pass under the current profile, while only a change making every cited member level 0 lowers the floor to 1 +MAJOR | high | §4 `` production, lines 203-205 | `[^ ;:,()]+` includes `+`, although `+` is also the contributing-model separator, and there is no quoted form for a verbatim reported identifier containing any excluded delimiter | one valid curve can have multiple parses, while another reported model value can have no representation, so P8 cannot consume the pinned interface deterministically | exclude `+` from the bare token and add one quoted, escaped model production that can preserve every reported identifier +MAJOR | high | §4 `` and `` productions, lines 203-205 | the per-pass form never requires its pass keys to equal the expanded `` exactly once in ascending order, and the properties do not require every model contributing a split logical pass to be recorded | a curve can omit, duplicate or add pass-model mappings or attribute a multi-branch pass to one convenient model while remaining grammatically valid | require the per-pass keys to be an exact ordered bijection with `` and require each key to carry every contributing branch model, using `undetermined` per unknown contributor +MAJOR | high | §8 Evidence plan, lines 335-360 | no verification parses the two pinned machine interfaces or exercises their semantic constraints; confirming this branch's three eventual instances cannot cover quoted paths, unusable knob causes, gapped ranges, mixed models, skipped cycles or rejected cardinalities | malformed forms can pass this story's evidence gate and fail only when P8 tries to consume deployed history | add a named mechanical grammar verification with accepted and rejected fixtures for every provenance and curve production, including count-to-pass and model-to-pass cardinality +MAJOR | high | §8 prompt-standards verification, lines 350-354 | eleven checklist items are scoped to changed prompt regions and only item 7 is applied to the whole result, while invariant 11 requires each changed prompt artifact to pass all twelve items as a complete prompt | global failures in target, structure, stop conditions, duplicated rules or emphasis can be declared passed because they sit outside the edited spans | run all twelve items against each complete resulting prompt and use changed-region notes only as supplementary evidence +MAJOR | high | §7 implementation surface and §8 prompt-standards obligation; current `CLAUDE.md` and the generated `CLAUDE.md` template | neither resulting prompt states its executing target model as checklist item 1 requires; `workflow-init.md:9` names the target only for the outer command and those bytes are not scaffolded | the mandated twelve-item pass cannot truthfully succeed and the shipped template remains nonconformant with invariant 11 | require a target-model statement in both resulting CLAUDE surfaces and revalidate the applicable current prompting guidance +MAJOR | high | §8 branch-closing-body reconstruction, lines 356-360 | the text says this branch's spec and plan cycles closed before the pinned forms existed, but the current spec cycle is still in Gate A and no plan artifact or Gate-A plan cycle exists in branch history | the plan is told to reconstruct nonexistent closed bodies and can perform unnecessary history surgery or falsely claim the three-cycle demonstration already exists | require the current spec body to take the form when this cycle actually closes, then require the future plan and Gate-B bodies to be written in the form at their normal closures; reserve reconstruction for a body that demonstrably already closed +MAJOR | high | parent story criterion 1 versus spec §2.4 | criterion 1 does not observe the new rule that any cited-set membership change requires a final clean pass against the current set even when the numeric floor is unchanged, nor that removing a citation cannot discharge an already accepted in-set Blocker or Major | both shipped copies can omit the revision-16 membership repair and the story can still be accepted, allowing closure without review of a newly cited high-risk story | extend criterion 1 with both membership-change consequences and make the unchanged-floor case explicitly checkable +MAJOR | high | parent story criterion 2 versus spec §4 lines 215-223 | criterion 2 does not observe the rules that split branches form one summed logical pass only on the same tracked reviewed commit and that a revision mismatch ends the first pass as incomplete before the later branch starts a new one | both shipped copies can omit the artifact-identity rule while the story still passes, letting one curve entry describe two revisions | extend criterion 2 to require branch summation, tracked-commit equality and the mismatch-to-incomplete transition +MINOR | high | §7 lines 306-308 versus `workflow-init.md:546-551` | the claim that the template lacks the prose-exemption rationale is false: the template already says explanatory docs describe the product rather than being it and contrasts them with prompts that are the product; only the more specific confused-human-reader versus broken-behaviour clause is absent | the plan can treat an existing condition as missing, duplicate rationale and corrupt the old-conditions accounting | name the exact missing human-cost clause and preserve the rationale already present +NIT | high | opening lines 15-17 versus `codex-gate.sh:93-109` and result-classification messages | `gate_citation` does not build every reminder's citation; it controls heading-based adoption and supplies the policy citation for policy-bearing gate-status messages, while several classification reminders hard-code section 5 or carry no policy citation | the heading constraint is correct but its stated mechanism overclaims what the hook reads and does | describe the pattern as controlling heading-based adoption and the policy citation used by the core gate-status reminders +NIT | high | parent story criterion 5 opening at line 180 | the criterion says "Four properties" and then defines five mandatory properties labelled a through e | the count claim is false and makes the fifth condition look accidental | change "Four" to "Five" +NIT | high | parent story criterion 1 line 137 | "Section A's reachability test" points to a section that no longer exists in the parent story or revision-16 spec | the binding rationale cannot be followed mechanically | replace the stale name with the current spec's §3 severity-semantics reference +END OF FINDINGS (13 total) diff --git a/.context/codex-reviews/gate-a-spec-rle-pass-16.md b/.context/codex-reviews/gate-a-spec-rle-pass-16.md new file mode 100644 index 0000000..24cec15 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-rle-pass-16.md @@ -0,0 +1,5 @@ +BLOCKER | high | §8 "fresh twelve-item docs/prompt-standards.md pass" | the spec correctly says invariant 11 binds each resulting prompt as a complete artifact, then knowingly leaves checklist item 1 false by routing the missing Target model declarations instead of fixing them | the implementation can follow the spec exactly only by shipping two changed prompt artifacts that fail a non-negotiable invariant, so the prompt-conformance gate cannot legitimately pass | add and verify the target-model declaration in both the repository CLAUDE.md and the scaffolded CLAUDE.md template in this change, or change the invariant explicitly; a follow-up route does not satisfy invariant 11 +MAJOR | high | §8 "No reconstruction is needed" and §5 cycle nonce | the open Gate-A spec cycle predates the nonce rule and has no valid nonce created at cycle start; its existing rle slot discriminator is deterministic and only three characters, so an eight-to-sixteen-character random nonce added now would be late-created provenance rather than a native cycle record | the branch cannot truthfully demonstrate the required cycle-start identity in its spec closing body, and pretending otherwise fabricates the provenance that criteria 6 and 11 rely on | either restart the spec cycle under a freshly generated nonce and count only the restarted cycle's passes, or state a legacy exception and move the first fully compliant spec-cycle demonstration to a later cycle instead of claiming native coverage here +MAJOR | high | §2.4 lines 126-140 | the first paragraph says a lowering drops the floor, lenses and evidence mode together and closes after one further pass, while the repair below says profile changes can leave the floor unchanged and that the consequence attaches to any profile or set change; it also omits the requirement that the further pass be clean and all other closure duties be satisfied | the two shipped copies can encode different transition procedures, including treating a level-2-to-level-1 change as a floor drop or closing after a pass that still has Blocker or Major findings | replace the raise and lowering paragraph with one rule: every profile or cited-set change recomputes current obligations and requires one further clean pass under them, prior valid passes keep counting, and closure still requires every ordinary duty +MAJOR | high | parent story criterion 1 lines 113-120 versus spec §2 lines 38-50 | the parent criterion still says floor 3 otherwise and names only higher-profile or unprofiled members, so it does not preserve the settled present-but-unresolvable stop and can be read to accept the silent fallback the spec explicitly forbids | implementation can satisfy the binding story while violating the spec and continue under floor 3 when a malformed or unresolved profile should stop, potentially omitting the profile's real lenses and evidence duties | amend criterion 1 to limit floor 3 to no-story, unprofiled, and resolvable-above-level-0 cases and explicitly require any present-but-unresolvable member to stop and surface +END OF FINDINGS (4 total) diff --git a/.context/codex-reviews/gate-a-spec-rle-pass-17.md b/.context/codex-reviews/gate-a-spec-rle-pass-17.md new file mode 100644 index 0000000..08b4528 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-rle-pass-17.md @@ -0,0 +1,7 @@ +BLOCKER | high | §2.3 `` grammar, §8 `pre-rule`, and parent criteria 6 and 11 | `pre-rule` contains a hyphen and cannot match `[a-z0-9]{8,16}`, yet §2.3 says one machine-extractable form covers every case and the parent requires this branch's records to use that pinned form | the mandated current-cycle records fail their own parse check, cannot satisfy the branch acceptance criteria, and cannot be consumed by P8 as specified | add an explicit legacy cycle-id production and align both grammars, their properties, the parent criteria, and P8, or define a valid migration that does not claim a late-created value was generated at cycle start +MAJOR | high | §8 "This branch's own three closing bodies" versus §10 "From the commit that ships them" | §8 treats only the already-running spec cycle as pre-rule and says the later plan and Gate-B bodies are native, but §10 does not make the rules bind when the spec closes: the Gate-A plan cycle starts before the implementation commit ships them, and Gate B also starts before the closing amend if that amend is the shipping commit | the branch can claim native nonce provenance and demonstrations for cycles that began under the old rules, so activation and acceptance have no single answer | define an explicit early-adoption point for this branch's later cycles, or apply the legacy treatment to every cycle begun before the shipping commit and align the demonstrations and criteria +MAJOR | high | §2.4 "A lowering additionally drops the lens sets and the evidence mode" | lowering a profile does not categorically drop both: a mode-only override changes evidence without changing axis-derived lenses, and security high to standard retains the security lens set while changing evidence obligations | shipped text derived from this rule can tell an agent to omit lenses that the current axes still require, under-reviewing the changed artifact | say that floor, lenses, and evidence are recomputed independently from the current axes and override, and only obligations no longer required are dropped +MAJOR | high | parent story §5 "Does a profile change mid-cycle move the floor" and acceptance criterion 1 | the answered question still states only that a raise costs a further pass, and criterion 1 observes cited-set changes but not profile changes, while spec §2.4 requires every profile change in either direction to cost a further clean pass even when the floor is unchanged | an implementation can satisfy the parent while allowing closure immediately after a lowering or same-floor profile change, violating the settled rule | update the answered question and criterion 1 to require one further clean pass with all closure duties after any profile change in either direction +NIT | high | parent story §2 opening | it says "Three outcomes" but now enumerates two and immediately explains that the former third moved to the successor | readers are sent looking for a missing current outcome before the scope-narrowing note repairs the count | change "Three" to "Two" +NIT | high | parent story criterion 5 opening | it says "Four properties" but enumerates five, including the load-bearing reviewing-pass exclusion as item (e) | the completeness count contradicts the criterion and can make item (e) look appended rather than required | change "Four" to "Five" +END OF FINDINGS (6 total) diff --git a/.context/codex-reviews/gate-a-spec-rle-pass-18.md b/.context/codex-reviews/gate-a-spec-rle-pass-18.md new file mode 100644 index 0000000..adc4e27 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-rle-pass-18.md @@ -0,0 +1,3 @@ +MAJOR | high | §2.3 lines 110-120; §4 lines 224-230; §5 lines 261-269; §8 lines 390-408; parent criteria 6 and 11 | the pre-rule alternative is grammatical, but the semantic contract is not exception-aware: the required properties still say every provenance record and curve carries its cycle nonce, the nonce section says every record carries it, and the story requires this branch's three closing bodies to carry distinct identifiers and name their cycles, while §8 and §10 require all three branch cycles to use the non-identifying `cycle none (pre-rule)` field and defer the first real nonce to a future cycle | this branch cannot satisfy its acceptance criteria or demonstrate the attribution property, and the promised first post-ship nonce verification has no in-scope cycle or named deferred vehicle | qualify the nonce and attribution properties for post-rule cycles, revise both branch-demonstration criteria to state exactly what the pre-rule records prove, and assign the first post-ship nonce verification to a named future checkpoint +NIT | high | parent criterion 1 line 147 and criterion 11 line 264 | the story cites `Section A` and an unqualified `§8`, but the current story has neither; the intended targets are the design's severity section and evidence plan | the mechanically requested cross-reference check fails and readers cannot follow either reference within the artifact that contains it | replace `Section A` with `part 2` or `design §3`, and qualify the latter as `design §8` +END OF FINDINGS (2 total) diff --git a/.context/codex-reviews/gate-a-spec-rle-pass-19.md b/.context/codex-reviews/gate-a-spec-rle-pass-19.md new file mode 100644 index 0000000..1ee360a --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-rle-pass-19.md @@ -0,0 +1,8 @@ +MAJOR | high | §5 "The cycle nonce" and parent-story criteria 6 and 11 | categorical claims still say both record types carry a nonce, every cycle holds an identifier, a run of all three produces three nonces, and this branch's provenance names its cycle, despite the new rule that every branch cycle is pre-rule, has no nonce, and is not cycle-attributable | the plan can satisfy the categorical acceptance text only by violating the bounded pre-rule exception, so revision 20 still encodes both sides of pass 18's contradiction | qualify every nonce and attribution claim as post-rule and make criterion 11 say the branch carries a non-attributable pre-rule cycle field +MAJOR | high | §8 "The nonce is verified at a named later checkpoint" and parent-story criterion 6 | "the first cycle started after the implementation commit" does not identify a cycle type or artifact and has no total ordering when sibling cycles can start concurrently | the only promised real-nonce checkpoint can be claimed against different cycles or have no uniquely auditable target, leaving real attribution unverified | name one concrete cycle by type and artifact or define a durable serial selection rule and the record that proves it was selected +MAJOR | high | §3 "Expected effect", §10 "The expected demotion", §4 curve grammar, and P8 criterion 5 | the spec says P8 measures how much consequence-keyed severity demotes findings, but the only durable per-pass form records total Findings and Blockers, not Majors or subject categories | a Major demoted to Minor changes neither stored number, so P8 cannot observe the dominant demotion class or distinguish severity improvement from unrelated pass-count movement | add durable Major or combined Blocker/Major counts and any subject split needed by the claim, or narrow the P8 claim to the pass-shape economics the form can actually observe +MAJOR | high | §2.3 and §4 pinned grammars | the supposedly machine-pinned grammars leave `` and `` as prose-only placeholders and define `` with an ellipsis; neither gives a closed character and escape language, and a quoted path containing a line break is not encoded despite the record being one line | independent writers and the P8 parser can accept different strings or split one record into several lines, defeating the one-form and machine-extractable properties | replace the placeholders with exact productions for allowed characters and escapes, including an explicit encoding or rejection rule for line breaks and other control characters +MINOR | high | parent-story criterion 8 versus spec §7 | the story says the user-facing-site inventory is in the spec and each cited line is checkable there, while the spec explicitly delegates the site list to the plan and cites no individual user-facing lines | Gate A cannot verify this acceptance claim from the rules artifact, and the ownership split makes a falsified documentation site easier to omit | put the complete cited site inventory in the spec or amend the criterion to make the plan the authoritative, checkable owner +NIT | high | parent story lines 269 and 304 | the historical note still calls the provenance demonstration "criterion 8", which is now the packaging/docs criterion, and the invariant note calls the accounting "part 3" after part 3 was split out | readers are directed to an unrelated criterion and a nonexistent current part | refer to the criterion by title or current number and point the accounting note to design §6 +NIT | high | successor story §4 "Two passages this story shares with the parent" | the section announces two shared passages, enumerates two, and then adds "A third" | the stated count disagrees with its own inventory and can make the clearly-stuck passage look outside the required dual accounting | rename the subsection to three passages and enumerate all three uniformly +END OF FINDINGS (7 total) diff --git a/.context/codex-reviews/gate-a-spec-rle-pass-2.md b/.context/codex-reviews/gate-a-spec-rle-pass-2.md new file mode 100644 index 0000000..bbc1028 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-rle-pass-2.md @@ -0,0 +1,31 @@ +BLOCKER | high | story criterion 5 at lines 141-148 versus spec §3 | the story still requires artifact-kind demotion and preserves severity only for a false green, while the settled decision and spec make consequence the classifier and also preserve false-red consequences | no implementation can satisfy both the current acceptance criterion and the settled contract | revise criterion 5 and desired outcome 2 to use the same named-reader-and-changed-decision test and bidirectional instrument consequence rule as §3 +BLOCKER | high | §6 "Old-conditions accounting" | the section enumerates passages but never lists the conditions each passage currently requires or marks any condition kept, moved, or deliberately dropped; it delegates the actual accounting to a future plan | the spec claims compliance with the non-negotiable AGENTS.md decision-procedure invariant without performing the accounting that prevents dropped conditions | add a per-passage condition inventory and disposition to this spec before approving the replacement procedure +MAJOR | high | §6 floor-wording count | the claimed twelve-site inventory omits three changing pairs: "below 3" at CLAUDE.md:79 and template:279, "each its own 3-pass loop" at CLAUDE.md:300 and template:485, and "where the 3 come from" at CLAUDE.md:335 and template:519 | implementation can leave fixed-three instructions in both shipped copies while the spec's own count and verification still pass | add all six occurrences, correct the stated site/change totals, and prescribe each replacement +MAJOR | high | §6 first table | the "every passage rewritten" enumeration omits profile-reading and profile-change passages that §4.1-§4.2 must change, including "Reading the profile", "Changing a profile", and the Gate-A/Gate-B loop paragraphs containing the omitted numeric sites | old profile-resolution, banked-pass, final-pass, and gate-specific conditions can be lost or left contradictory | enumerate every additionally rewritten passage and account for all of its current conditions +MAJOR | high | §8 parity verification | the verification walks only §6's first table, excluding the second floor-site table and newly inserted lifecycle, exposure, history-shape, curve-format, provenance, and activation rules | both copies can diverge on a new load-bearing rule while the planned parity check reports success | make the verification inventory every changed or inserted rule, including every site in both §6 tables and each new standalone rule +MAJOR | high | §4.1 "The knob's lifecycle" versus README.md:128-132 | the design takes ownership of an existing documented per-workspace user override, overwrites its prior positive value with 1 or 3, then deletes it without preserving or migrating that value | an update silently destroys a user's configured floor, including a stricter floor above 3, creating data loss and backward incompatibility | define ownership and migration for pre-existing overrides, preserving or explicitly retiring the prior value instead of silently overwriting and deleting it +MAJOR | high | §4.1 remove-on-close | cleanup is defined only as part of a Gate-B closing amend, but the same file governs the separate Gate-A spec and plan cycles, whose closure events are skill invocations rather than closing amends | a floor 1 from Gate A can leak into the next Gate-A or Gate-B cycle, or cleanup before the closure check can make the hook evaluate the wrong floor | define the write, closure-check, and cleanup ordering for each of Gate-A spec, Gate-A plan, and Gate B +MAJOR | high | §4.1 write/read verification | the spec says to verify the stored value but gives no terminal action for a failed write, failed readback, parsed mismatch, or failed removal, and it does not require a final readback immediately before closure | an old 1 can remain effective after an attempted write of 3, or an agent can change the gitignored value after the last truthful pass report and close under the lower floor | require mismatch or unreadability to stop before counting or closing, and require a closure-time comparison of parsed axes, derived floor, and actual knob value +MAJOR | medium | §4.1 concurrency limitation | "cycles are sequential by construction" is unsupported: current §5 serializes calls on one slot but neither prevents overlapping Gate-A and Gate-B cycles nor cycles run by separate agents in one checkout | one cycle can overwrite or remove another cycle's floor and both can present internally plausible reports | replace the unsupported construction claim with an explicit no-overlap rule and a stop when another cycle appears to own the checkout-global knob +MAJOR | high | §10 gate-off residual | the disclosed attack considers only writing the knob incorrectly, not omitting a higher-risk cited story, minting or editing a level-0 profile, or presenting an incomplete cited-story set; all are controlled by the same agent and can make the derived value and provenance internally consistent | an agent seeking to skip review can evade the exposure check without any reported mismatch | disclose the whole caller-authored trust chain and require each pass to show the complete cited set plus the existing human-confirmation and profile-log basis, while retaining the stated lack of independent enforcement +MAJOR | high | §10 activation boundary | "the rules it started with" has no durable policy-version or start marker, and optional resume notes cannot recover it after a fresh checkout, machine move, or context loss | an in-flight cycle cannot determine whether to use old severity, floor, exit, curve, and decline duties, so the non-retroactivity rule is not executable or auditable | make the governing prompt revision determinable from an existing durable record and define the safe action when that record is absent +MAJOR | high | §4.1 and §10 rollback and abandonment | cleanup occurs only on successful closure; no behavior is defined for an abandoned cycle, a blocking profile/setup stop, or rollback to prompt text that does not know the new lifecycle while a 1 remains | a reduced floor can survive precisely when the rule responsible for removing it is no longer running | define cleanup or restoration on abnormal termination and rollback, with absence or 3 as the fail-safe state +MAJOR | high | §5 stale-history shape | the findings-file format and slot names carry no cycle identity, and cycle start does not clear all prior-pass slots, so a pass-4 reader cannot distinguish a stale valid file from a valid file in the current cycle | a foreign curve can be accepted as current history despite the rule saying stale files are absent | add a mechanically observable cycle boundary or start-of-cycle slot cleanup and state the conservative action when provenance cannot be established +MAJOR | high | §5 history-shape diagnostics versus prompt-standards item 10 | the four categories name symptoms but provide no checks and per-cause remedies; "malformed or unreadable" merges distinct causes, and "stale" has no available discriminator | the shipped prompt would report a diagnostic state that its reader cannot resolve, violating the binding checklist | for each cause, name the distinguishing check, the affected tells, and the remedy or stop action +MAJOR | high | §5.1 pinned curve form | only `reviewType: full` is defined; the existing protocol also permits separate `spec` and `quality` calls and single-branch recovery, while the hook counts calls rather than logical branch pairs | authors can count two branch calls as two passes or one, producing incomparable curves and potentially satisfying the floor with only one reviewer dimension | define the logical pass and curve aggregation for full calls, paired single-branch calls, single-branch recovery, retries, and incomplete attempts, matching the unit the floor credits +MAJOR | medium | §5.1 "what P8 and any future resumption need" | the mandated durable curve covers Gate-B cycles only, while the floor and severity change also govern Gate-A spec and plan loops and the economics story describes review loops generally | the deferred measurement remains unable to observe Gate-A distributions, so the spec overstates how completely the new record answers the economic question | narrow the P8 and durability claims explicitly to Gate B and state the unmeasured Gate-A axis without changing the settled Gate-B-only record requirement +MAJOR | high | §8 risk-path verification versus story criterion 8 | the verification requires a level-0 cycle to write 1 and close, but this branch cites a risk-high story and criterion 8 explicitly forbids a floor-1 demonstration here, deferring the first such checkpoint to P8 | the evidence plan either violates the profile-derived floor or manufactures the level-0 fixture the story rejects | make this branch's verification exercise only the floor 3 path and leave the first real floor-1 lifecycle observation at the settled P8 checkpoint +MAJOR | high | §7 and §9 rollout scope | the plan omits user-facing text already falsified by the new rule: docs/getting-started.md:31-40 and :51-58 require three passes, :84 says Gate A is unchanged at every level, and README.md:128-132 presents the floor file as a persistent user knob | users can follow shipped documentation into the old floor or have their configured knob silently repurposed despite installing the new version | include the falsified guidance and knob documentation in the implementation surface and old-condition review +MINOR | medium | §2.1 commit-body transport during Gate A | a decline is called a recorded decision that later passes recognize, but a Gate-A spec or plan commit does not exist until the cycle closes and the advisory dispositions file is expressly not the record | an interrupted Gate-A cycle has no durable in-flight identity record and can re-ask or lose the disposition before closure | state that absence of the closing record on resume restores the hold and requires re-confirmation, or define an existing durable in-flight carrier without making the advisory companion authoritative +MAJOR | high | §2.1 finding identity | sameness is decided only by location and defect, yet the next sentence says a materially reworded finding is new; severity, consequence, and suggested-fix changes can materially change a finding while location and defect still match | a newly consequential Blocker can inherit an earlier decline silently | include consequence-changing fields in the identity test and make any disagreement or material change a new held finding +MINOR | high | §2.1 decline record form | `` has no pinned serialization for the required pass, slot, line, location, and defect quotation, nor an escaping rule for delimiter or newline characters | two agents cannot reliably write, grep, compare, or squash-carry the same identity | specify one exact single-line identifier grammar and escaping rule with a complete example +MAJOR | high | §2.1 "explicitly declined" | the design calls the decision attributable while reusing an unverified caller-authored commit-body transport; unlike the existing human-exception section, it never says what the decline record does not verify | history can be read as proof that the named human actually declined the finding even though no mechanism checks that claim | state that the record is an unverified assertion of a user decision and that its loop effect is instruction-backed, not authenticated +MAJOR | high | §2 precedence table | "every surfaced finding still open has been explicitly declined" is self-contradictory because §2.1 defines a decline as releasing the finding so it is no longer open | an agent can read the central clean-versus-scope cell as impossible to satisfy or as permission to close with some genuinely open findings | say clean wins only when every scope-stop finding has been explicitly declined and no other surfaced finding remains open +MINOR | high | §3 "Why not a list of demotable artifact kinds" | the spec says one of three `fic2` pass-5 findings was acted on, but the field report says the three pass-5 story-criterion findings were parked and identifies the acted-on activation-boundary item as a pass-4 finding | the evidence offered for consequence-based severity is factually misclassified | distinguish the three parked pass-5 findings from the separate acted-on pass-4 activation-boundary finding +MAJOR | high | story criteria 3 and 8 versus §4.1 | criterion 3 requires provenance only for a non-default floor, criterion 8 requires this default-floor-3 cycle to carry provenance, and the spec requires provenance for every cited story in every cycle | the story has no single checkable provenance requirement and implementation can be rejected whichever interpretation it follows | align both criteria with the spec's one chosen provenance rule and state whether default 3 records are mandatory +MINOR | high | §4.2 lowering | "a high cycle lowered to trivial can close on one pass" conflicts with the immediately preceding rule that the final clean pass must run under the current profile after earlier passes were banked | a reader may close on an old-profile pass instead of running the required post-lowering clean pass | say it can close after one further clean pass under the lowered profile, regardless of the total banked count +MINOR | high | §4.1 exposure requirement versus prompt-standards item 4 | the new mandatory pass-report structure names three values but provides no concrete output example or pinned ordering | agents can produce incomparable or incomplete exposure lines while believing the format is satisfied | add a literal example showing story paths, parsed axes, derived floor, read-back value, and mismatch state +MINOR | high | §5.1 historical-curve claim | `baa75c1` records per-pass finding totals but no per-pass Blocker totals, while the surrounding requirement defines the curve as both finding and Blocker counts and says `3cdd075` and `baa75c1` "did" it | the precedent set is overstated and can be read as evidence for a format that commit does not carry | qualify `baa75c1` as a findings-only partial precedent and keep `3cdd075` as the complete-form precedent +NIT | high | §1.1 first block quote | the quoted sentence ending "once the question is answered." exists verbatim in the template, but CLAUDE.md:105 continues with an em dash instead of that period, and the spec does not state this one-copy punctuation variance | the mandated exact quote check fails against one of the two claimed source copies | quote only the byte-identical shared text or state the punctuation variance explicitly +MINOR | high | §8 evidence plan | the plan relies on the battery's narrow invariant checker but never requires the manual all-12-item prompt-standards review that AGENTS.md invariant 11 makes the actual gate | the implementation can pass every named verification while failing a binding prompt criterion | add an explicit recorded twelve-item review of every changed prompt region in both copies +END OF FINDINGS (30 total) diff --git a/.context/codex-reviews/gate-a-spec-rle-pass-20.md b/.context/codex-reviews/gate-a-spec-rle-pass-20.md new file mode 100644 index 0000000..9592554 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-rle-pass-20.md @@ -0,0 +1,7 @@ +MAJOR | high | parent story criterion 6, lines 226-230 | the story still assigns the real-nonce checkpoint to "the first cycle started after the implementation commit", while spec §8 assigns it to the Gate-A spec cycle of the next story whose spec is written after that commit lands | the two authoritative artifacts select different later cycles, so the checkpoint can be claimed against a cycle the spec did not designate | mirror the spec's exact Gate-A-spec-cycle checkpoint and its same-nonce and distinct-from-other-cycles checks in criterion 6 +MAJOR | high | parent story criterion 11, lines 253-257 | the criterion still says this branch's pre-rule provenance line names the cycle, but the pinned line contains only `cycle none (pre-rule)` and spec §§4-5 explicitly say that value is not cycle-attributable | satisfying the criterion requires either inventing attribution the grammar cannot carry or accepting a closing record that fails its stated demonstration | say the branch records the non-attributable pre-rule cycle field and demonstrates every other provenance field, reserving cycle attribution for the named post-rule checkpoint +MAJOR | high | spec §4 "Records Majors...", spec §10 "P8 measures it", and P8 §1 and criterion 5 | adding Major counts does not make the durable data sufficient to measure how much consequence-keyed severity demotes findings: the `fic2` baseline commit `3cdd075` says no Blocker or Major occurred after pass 2, the committed field report records Majors 6, 2 and 5 on passes 3-5 and supplies no Major counts for passes 6-7, and the subject material the spec relies on remains only in gitignored findings files | P8 cannot establish a trustworthy pre-rule Blocker/Major baseline or distinguish rule-driven demotion from a different mix of findings, so the promised economics conclusion is not answerable from its read-only git inputs | either commit and reconcile a verified full baseline plus durable consequence or subject coding sufficient for the counterfactual, or narrow P8 and every claim about it to comparing self-reported aggregate severity shares without attributing the difference to this rule +MINOR | high | spec §3 "The amount is a prediction, not a measurement", §4 "must not present it as measurement", and §10 "P8 measures it" | the spec both calls P8 the measurement that measures the demotion and forbids P8 from presenting its unchecked self-reported curves as measurement | the plan and P8 receive incompatible reporting standards, so the same aggregate comparison can be represented as established evidence or as an admitted self-report | use one term consistently, preferably a self-reported comparison or analysis unless a validating mechanism is added +MAJOR | high | spec §8 prompt-standards verification, lines 400-411 | the required whole-artifact twelve-item pass covers the resulting root `CLAUDE.md` and the resulting scaffolded template but omits `plugins/dev-workflow/commands/workflow-init.md` as the outer command prompt, even though this change edits that prompt artifact and the text itself acknowledges its separate target-model declaration | the plan can close while never checking the changed command as a complete prompt, violating invariant 11 and leaving command-level contradictions or stop/output defects outside the promised review | require a fresh twelve-item whole-artifact pass for the outer `workflow-init.md` command in addition to the root prompt and the embedded scaffolded prompt +NIT | high | spec §4 line 233 and §5 line 280 | both pre-rule explanations point to §8 for the rule governing cycles that began before shipping, but the activation rule is in §10; §8 is the evidence plan and only refers onward to §10 | readers are sent to the demonstration procedure instead of the normative activation rule | change both cross-references from §8 to §10 +END OF FINDINGS (6 total) diff --git a/.context/codex-reviews/gate-a-spec-rle-pass-21.md b/.context/codex-reviews/gate-a-spec-rle-pass-21.md new file mode 100644 index 0000000..ce43b53 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-rle-pass-21.md @@ -0,0 +1,10 @@ +MAJOR | high | parent story criterion 11, lines 257-261 | the criterion says this branch's provenance line names the cycle, while criterion 2, criterion 6, and spec §§4-5 say `cycle none (pre-rule)` neither names nor attributes a cycle | implementation cannot satisfy both statements without inventing the prohibited late provenance or accepting a closing record that fails the criterion | replace `naming the cycle` with `carrying the reserved pre-rule cycle field` and reserve attribution for the named post-rule checkpoint +MAJOR | high | spec §10 lines 468-474 and successor story §4 lines 113-116 | the unknown-start fallback claims to cover every part this change touches but omits the provenance-line duty, and the successor repeats the same four-item list of floor, severity, curve, and nonce | an agent can close an unknown-start cycle without provenance, breaking the parent criterion's history-only reconstruction and leaving P8 no floor record to parse | add the provenance-line duty at its strictest reading to the fallback and mirror that addition in the successor's inherited list +MAJOR | high | spec §4 lines 237-249 and §10 lines 495-498 | aggregate Findings, Blockers, and Majors from different post-rule cycles make recorded severity mixes comparable, not `demotion`; there is no paired old-rule classification or durable subject material, and §10's comparison-only limit contradicts §4's stronger claim | P8 can attribute a change caused by different artifacts or finding mixes to the severity rule and report the causal conclusion the spec otherwise forbids | describe the output consistently as a comparison of recorded severity shares, or add durable paired counterfactual evidence that can identify actual demotions +MINOR | high | spec §4 lines 242-249 | the claim that the `fic2` baseline supports total-volume comparison and nothing finer omits its complete per-pass Blocker series, present in `3cdd075` and repeated in the committed field report | P8 is instructed to discard a legitimate baseline dimension and the stated source limit is factually wrong | say the baseline supports per-pass Findings and Blockers, but not complete Majors, a complete combined Blocker/Major distribution, subject coding, or a demotion inference +MINOR | high | spec §4 grammar lines 217-224 | ` := [^ ;:,()+"]+` accepts tabs and other control characters even though the next rule says any identifier containing a control character is written `undetermined` | a parser following the pinned grammar can accept a non-single-line or non-greppable identifier that the prose declares unrepresentable | exclude every whitespace and control character from `` and add tab and control-byte rejection cases to the parse check +MINOR | high | spec §2.3 lines 99-126 | `` permits the same path more than once, including duplicate entries carrying conflicting levels, and no semantic uniqueness rule rejects that form | a grammar-valid provenance line can fail to determine which level the cited story had, defeating history-only reconstruction and floor verification | require each normalized path exactly once in the set and make conflicting or repeated entries invalid parse cases +MINOR | medium | spec §7 lines 354-358 versus workflow-init template lines 546-551 | the spec says the template lacks the describe-versus-product rationale, but the current template already says explanatory docs describe the product rather than being it and prompts are product | the plan starts from a false pre-change claim and can add duplicated prompt text, contrary to the token-lean standard | name the exact missing clause if the intended addition is the confused-reader consequence, or drop this claimed divergence and planned seam +MINOR | high | successor story §4 lines 80-98 versus the fic2 field report lines 40-50 and 216-221 | the successor says every settled input is recorded in the field report, but that report explicitly leaves Q3-Q5 unanswered and presents Q6 only as candidate answers | a future design or gate agent is sent to a source that records the opposite provenance state and may reopen or accept decisions without their actual confirmation record | cite the commit bodies that contain the later confirmations and state precisely which earlier inputs, if any, the field report records as settled +NIT | high | spec §4 lines 231-234 | the pre-rule exception points to §8, but the normative activation rule that decides whether a cycle is pre-rule is §10; §8 is the evidence plan | a reader following the cross-reference lands on a verification procedure instead of the governing rule | change the cross-reference from §8 to §10 +END OF FINDINGS (9 total) diff --git a/.context/codex-reviews/gate-a-spec-rle-pass-22.md b/.context/codex-reviews/gate-a-spec-rle-pass-22.md new file mode 100644 index 0000000..b1a3da9 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-rle-pass-22.md @@ -0,0 +1,5 @@ +MAJOR | high | parent story §3 criterion 11, lines 257-263 | the criterion still requires the pre-rule provenance line to name the cycle and immediately says the same line does not name its cycle because `cycle none (pre-rule)` attributes nothing | the acceptance criterion is internally unsatisfiable unless implementation invents the prohibited late identity or ignores one of its requirements | replace `naming the cycle` with `carrying the reserved pre-rule cycle field` and leave real attribution to the named post-rule checkpoint +MAJOR | high | spec §3 "Expected effect", §4 "What the curve makes answerable", §10 final risk, parent story §2 and criterion 11, and P8 story §1 and criterion 5 | §4 correctly says cross-cycle severity mixes cannot measure demotion, but §3 still routes the amount of demotion to P8, §10 says P8 checks the expected demotion and that recorded values can show less demotion, and the parent and P8 still call this an economic measurement of the rule's effect without requiring the cross-artifact confound; P8 also names only fic2's Findings curve although the verified baseline has complete Findings and Blockers plus partial Majors | P8 can satisfy its story by reporting an unsupported causal demotion or economics result, or can discard baseline dimensions the spec now permits | align every site on a comparison of recorded Findings, Blockers and available Majors only, require the self-report and different-artifact confounds, forbid a demotion or rule-effect figure without paired classification, and state the fic2 partial-Major limit +MAJOR | high | successor story §4 "Three passages this story shares with the parent", lines 117-120, versus spec §10 lines 481-488 | the successor's inherited unknown-start list still names floor, severity, curve and nonce but omits the provenance-line duty, even though revision 23 says the parent touches five parts and the successor must extend that same list | the successor can replace the fallback while believing it preserved the parent and silently drop provenance, leaving an unknown-start cycle without the history record P8 and criterion 3 require | add the provenance-line duty to the inherited list and require the successor's old-conditions accounting to preserve all five parent duties before extending it +MINOR | high | spec §7 lines 365-371 versus §9 line 475 | §7 withdraws the template-rationale prerequisite and says no seam is opened, while §9 still excludes reconciliation only beyond `§7's one seam` | the implementation scope remains contradictory and can reintroduce the duplicate template edit revision 23 explicitly removed | replace the stale one-seam reference with an exclusion of reconciliation beyond the rules this change actually edits +END OF FINDINGS (4 total) diff --git a/.context/codex-reviews/gate-a-spec-rle-pass-23.md b/.context/codex-reviews/gate-a-spec-rle-pass-23.md new file mode 100644 index 0000000..423b205 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-rle-pass-23.md @@ -0,0 +1,3 @@ +MAJOR | high | parent story criterion 11, lines 258-266 | the criterion still says the pre-rule provenance line names the cycle, then says its reserved `cycle none (pre-rule)` field attributes nothing and demonstrates no identifier | the acceptance criterion remains internally unsatisfiable and can make implementation invent the prohibited late identity or reject the correct pre-rule record | replace `naming the cycle` with `carrying the reserved pre-rule cycle field`, leaving attribution to the named post-rule checkpoint +MAJOR | high | spec §4 lines 243-246; parent story §2 lines 98-106 and criteria 2 and 11 lines 169-172 and 274-278; P8 story §1 lines 34-41 | revision 24's comparison-only correction is still contradicted by text saying the curve is kept to answer whether severity demotes findings, the rule's effect appears across cycles, P8 can measure the problem and the economics later, and P8 receives an economic measurement | a P8 implementer can still report an unsupported causal demotion or rule-effect result from self-reported curves over different artifacts, violating the settled decision and the gate-overclaim Don't | replace every remaining effect or measurement claim with comparison of recorded pass counts and severity mixes, and point each handoff to the paired-classification limit plus both confounds +END OF FINDINGS (2 total) diff --git a/.context/codex-reviews/gate-a-spec-rle-pass-24.md b/.context/codex-reviews/gate-a-spec-rle-pass-24.md new file mode 100644 index 0000000..dca1655 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-rle-pass-24.md @@ -0,0 +1,8 @@ +MAJOR | high | parent story criterion 11, lines 260-281 | the criterion calls this branch's provenance proof end to end and says the derivation, knob and provenance path work, but it simultaneously defers attribution, calls the demonstrated reserved cycle field the field it cannot demonstrate, and the spec makes knob preservation not applicable when the knob is absent | acceptance can be marked satisfied for nonce attribution and a present or unusable knob path this branch never exercises | narrow the branch proof to the floor-3 derivation, grammar, reserved pre-rule cycle-field production and absent-knob case; leave nonce attribution and non-absent knob behavior to explicit later checkpoints +MINOR | high | spec §4 opening, lines 207-210, and parent story criterion 2, lines 152-169 | both summarize the required curve as Findings and Blockers even though the pinned grammar also mandates Majors and the later comparison relies on that series | a plan can preserve the explicit two-series wording and still appear story-compliant while producing no forward Major series | name Findings, Blockers and Majors everywhere the curve requirement is summarized +MINOR | high | spec §5, lines 294-302 | the rationale says a record that cannot be attributed to a cycle is unusable by the analysis, then defines every pre-rule record as unattributable while the design and P8 rely on the pre-rule fic2 baseline | readers receive incompatible rules about whether the required pre-rule records can participate in the comparison | scope the unusability claim to distinguishing concurrent post-rule cycles and state that a pre-rule record remains usable as an unattributed baseline with that limitation +MINOR | high | spec §3 lines 195-201, §4 lines 207-210 and 249-286, and parent story criterion 3 lines 181-187 | summary passages still call P8 a measurement, call the missing Gate-A record unmeasured, and require only the singular confound, while the settled contract is a comparison of recorded mixes with both the self-reported-curve and different-artifact confounds named | a plan or P8 report can follow the weaker summary, omit one limit, and imply measurement of the rule's effect | use comparison or analysis, change unmeasured to unrecorded, and name both confounds at every summary site +MINOR | high | spec §4, lines 256-260 | the sentence says the closing commit and committed field report both carry Majors for only some passes, but 3cdd075 carries no Major series at all; only the field report records Majors, for passes 1-5 | a source audit can look in the closing commit for data that is not there and misstate the baseline's provenance | say that both sources carry all totals and Blockers, that the closing commit carries no Majors, and that the field report carries Majors only for passes 1-5 +MAJOR | high | spec §4 quoted-model rule, lines 223-230 | a model identifier containing a control character is replaced by undetermined in the pinned line but its raw value is required in prose; literal control data, especially NUL, is not safely representable in a Git commit message | the error path intended to preserve observability can instead prevent cycle closure or create an unsafe malformed record | require a printable escaped encoding such as JSON or hexadecimal for the original value, or stop and surface when it cannot be encoded +MINOR | high | spec §8 lines 458-463 versus §5 lines 304-326 and parent criterion 6 lines 222-237 | the later checkpoint requires its nonce to differ from any other cycle's, while the normative collision and uniqueness rule is only among open cycles | a nonce valid under the shipped rule can fail its checkpoint merely by matching a closed historical cycle | require difference from every other open cycle at the checkpoint, or deliberately strengthen the normative rule and recovery procedure to global uniqueness +END OF FINDINGS (7 total) diff --git a/.context/codex-reviews/gate-a-spec-rle-pass-25.md b/.context/codex-reviews/gate-a-spec-rle-pass-25.md new file mode 100644 index 0000000..b3fd0e6 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-rle-pass-25.md @@ -0,0 +1,6 @@ +MAJOR | high | §8 lines 435-448; scripts/check-invariants.sh:331-337 | adding the required scaffolded `Target model:` line leaves `workflow-init.md` with its existing outer declaration plus a second column-1 declaration inside the inline template, while check 4a explicitly rejects every claiming file whose declaration count is not exactly one | the natural implementation fails the required quality battery and invariant 11, forcing an unplanned checker change or wording workaround after the design is approved | require a nested declaration form that remains a valid target statement when scaffolded without becoming a second counted declaration in the source, or add the checker and its tests to scope and specify region-aware validation +MAJOR | high | parent story §1 lines 14-17 | the story infers that "the substance converged" from the cited finding and Blocker curves even though CLAUDE.md says those curves do not measure coverage and hardening-log supersession entries 78, 84, 85 and 87 explicitly identify that same inference as unsupported or false | the brainstorming agent consumes an overstated empirical premise when deciding that the fixed floor buys no additional substantive review | remove the convergence claim or support it with independent evidence that measures coverage and substantive convergence rather than the count curves +MAJOR | high | §8 lines 450-469 and parent story criterion 11 lines 260-271 | the criterion first names two out-of-reach cases, the cycle identifier and the non-absent user-knob clause, but later says the cycle field is the one thing it cannot demonstrate and that every other field is shown; the spec likewise claims every field except the nonce is demonstrated | an acceptance reader can treat the emitted `absent` token as proof of present, unusable and preservation behavior that the branch explicitly does not exercise | distinguish the demonstrated absent-knob production from the unexercised non-absent knob behavior everywhere, and keep both limitations named consistently in the criterion and verification claim +MINOR | high | §4 lines 263-269 | the absolute claim that neither historical source carries per-finding subject material is false of the committed field report, which classifies the four pass-5 findings by subject, reproduces one verbatim, summarizes the other three and also carries aggregate pass-2 subject clusters | P8 can be told that no finer evidence exists and omit the partial qualitative baseline that is actually durable, even though it still cannot derive a complete subject series or demotion figure | say that neither source carries a complete per-finding subject classification, then name the partial subject evidence and the comparisons it cannot support +MINOR | medium | §2.3 lines 103-108 and §4 lines 225-236 | the quoted path and quoted-model productions constrain escapes and control characters but never require non-empty content, so the pinned grammar admits `""` where a story path or model identifier is required | a P8 parser can accept a provenance record with no source story or a curve with no reviewer identity, defeating the reconstructibility and observability properties the formats exist to provide | require at least one decoded non-control character; stop on an empty path and encode an empty or unavailable model as `undetermined` +END OF FINDINGS (5 total) diff --git a/.context/codex-reviews/gate-a-spec-rle-pass-26.md b/.context/codex-reviews/gate-a-spec-rle-pass-26.md new file mode 100644 index 0000000..2db0315 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-rle-pass-26.md @@ -0,0 +1,5 @@ +MAJOR | high | spec §8 lines 443-455 and parent story criterion 8 lines 249-258 | the settled item-1 n/a is not scoped to the scaffolded artifact: the spec also endorses the root `CLAUDE.md` having no declaration, while the story says any checklist item may be recorded n/a with a reason | an implementation or acceptance review can treat this as a general escape from prompt-standards instead of the single model-agnostic-template exception, weakening invariant 11 for other items and artifacts | state only that item 1 is n/a for the scaffolded `CLAUDE.md`, remove the root-file endorsement and the any-item rule, and keep every other item and artifact subject to the ordinary checklist +MAJOR | high | spec §8 lines 461-479 | line 465 says the branch demonstrates every field of both pinned forms except the nonce, but lines 474-479 correctly name two undemonstrable-here fields: the cycle identifier and the knob clause's non-absent form | the closing verification can be read as satisfying present or unusable knob behavior that this knob-absent branch does not exercise | revise the earlier every-field claim to name both exceptions and preserve the undemonstrable-here reasons consistently +MAJOR | high | spec §2.3 lines 96-98 and §5 lines 316-318 | the pinned grammar caps a nonce at 16 characters, but the generation requirement says only at least 8 and therefore permits longer values | a conforming generator can create a cycle identifier that every required provenance line and curve must reject as malformed, preventing valid closure or losing cycle attribution | require generated nonces to be 8-16 characters everywhere, or widen the grammar and its parse fixtures to the intended upper bound +MINOR | medium | spec §4 lines 223-229 | the reserved missing-model token `undetermined` is also accepted by ``, so the grammar does not distinguish an unknown reviewer from an actual model identifier whose literal name is `undetermined` | a valid provider-qualified model can be recorded as missing, making the model-attribution record ambiguous for P8 and human readers | reserve bare `undetermined` for the sentinel and require an actual identifier with that spelling to use the quoted-model production +END OF FINDINGS (4 total) diff --git a/.context/codex-reviews/gate-a-spec-rle-pass-27.md b/.context/codex-reviews/gate-a-spec-rle-pass-27.md new file mode 100644 index 0000000..0260d6c --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-rle-pass-27.md @@ -0,0 +1,7 @@ +MAJOR | high | spec §2.1 "The derived floor controls whether a cycle may close" and §2.4 "closing requires the floor" | the absolute closure rule omits the existing zero-findings below-floor exit in both §5 copies and the legitimate Gate-B skip that §4 itself requires to be recorded | either the implementation silently removes valid exits or ships contradictory closure instructions, so an agent can run unowed passes or refuse a valid Gate-B skip | state the precedence explicitly, preserving the zero-findings exit and a legitimately skipped Gate-B cycle as exceptions to the ordinary derived-floor closure requirement +MAJOR | high | spec §8 item-1 n/a passage and parent story criterion 8 "Prompt conformance is judged item by item" | the n/a repair is still not scoped consistently: the spec says every item binds the other two changed artifacts and then says item 1 never bound the root `CLAUDE.md`, while the story still permits any checklist item to be recorded n/a whenever a reason establishes inapplicability | the implementation and acceptance review have no single conformance matrix and can extend the one settled exception to another item or artifact, weakening invariant 11 | say exactly that only item 1 for the scaffolded `CLAUDE.md` is the reasoned n/a, that items 2-12 bind it, that all 12 bind the outer command, and separately state the root file's intended item-by-item scope without the contradictory every-item claim +MAJOR | high | parent story §5 answered question "Which floor governs a cycle citing several stories" | "any other set yields 3" includes a set containing a present-but-unresolvable profile, contradicting criterion 1 and the settled stop-and-surface rule | a reader following the story's recorded answer can silently default an invalid profile to floor 3 instead of stopping, the exact gate-off interpretation the criterion forbids | replace "any other set" with the precise higher-profile or unprofiled cases and state alongside them that any present-but-unresolvable member stops and surfaces +MAJOR | high | spec §7 "Packaging" and parent story criterion 8 "which CI enforces" | invariant 12 and `scripts/check-version-bump.sh` enforce only a changed manifest version; they do not check for a `CHANGELOG.md` entry, so both passages overstate the gate by attributing the changelog requirement to invariant 12 or CI | a plugin change can pass CI without the required changelog entry while the implementation and reviewer are told the gate will catch that omission | state that CI enforces the version difference and that the changelog entry is a separate review-backed acceptance requirement, or add and test a real changelog check +MAJOR | medium | spec §5 "It appears in every record the cycle writes" and parent story criterion 6 | the universal record requirement never defines its record set: §5 introduces cycle fields only for the provenance line and curve, while a cycle also writes findings slots, the advisory working record, evidence entries, skip records and sometimes human-exception records | implementations can disagree about where the nonce must appear, leaving some records unattributable or changing existing record formats and old-condition accountings that another implementation leaves untouched | enumerate the record types covered and how each carries the identifier, including whether a nonce in a slot name satisfies the rule, or narrow the requirement explicitly to the two new commit-body records plus the working record and nonce-keyed slots +NIT | high | spec §5 nonce generation bullet | the sentence contains the duplicated punctuation "usable infix — — never derived" | the copy-editing error interrupts an already dense normative sentence but changes no decision | remove one em dash +END OF FINDINGS (6 total) diff --git a/.context/codex-reviews/gate-a-spec-rle-pass-28.md b/.context/codex-reviews/gate-a-spec-rle-pass-28.md new file mode 100644 index 0000000..099a181 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-rle-pass-28.md @@ -0,0 +1,7 @@ +MAJOR | high | spec §2.4 "closing requires the floor as currently derived" and parent story §5 profile-change answer | the categorical closure statement remains unchanged even though §2.1 now says the floor replaces only the number and preserves the zero-finding below-floor exit and legitimate Gate-B skip | an implementation following the profile-change rule can still delete a preserved exit or ship two contradictory closure procedures, causing unowed passes or refusal of a valid skip | qualify this as the ordinary floor-based closure path in both artifacts and restate that zero findings may close below it while a legitimate Gate-B skip removes the review +MINOR | high | parent story §2 "prompt-only policy over the already-shipped .context/codex-gate.floor knob" | this still describes the profile-derived floor as policy over the knob, contradicting the spec and criterion 4 that derivation never reads or acts on the knob and the knob remains only the hook reminder threshold | a plan reader can make the knob an input to the obligation and repurpose user-owned state despite the settled prompt-only design | say the profile-derived policy is independent of the existing knob, which remains solely the hook's reminder threshold +MAJOR | high | spec §5 named nonce record set versus parent story criterion 6 "present in every record that cycle writes" | the spec now limits nonce-bearing records to provenance, curve or skip record, findings slots, and the advisory working record while expressly excluding evidence entries and human-exception records, but the parent criterion still requires every record without those exclusions | implementation can satisfy one artifact only by violating the other, either changing out-of-scope record formats or failing the story's acceptance test | mirror the spec's exact record list and exclusions in criterion 6, including that a nonce in the slot name satisfies the findings and working-record requirement +MAJOR | high | spec §8 prompt-standards passage and parent story criterion 8 | §8 first says invariant 11 binds every changed prompt artifact and requires a complete twelve-item pass over root CLAUDE.md, then says whether invariant 11 reaches that file is not ruled on; with the sole n/a reserved for item 1 of the scaffolded CLAUDE.md, root item 1 has no defined disposition | the conformance evidence cannot be completed without silently deciding the explicitly unsettled root scope, adding an unrequired declaration, or creating a second n/a forbidden by the story | honor the settled non-ruling by removing the claim that invariant 11 binds root CLAUDE.md and defining any root audit as outside this decision, while keeping the complete pass for the scaffolded artifact and outer command +MAJOR | high | spec §8 later nonce checkpoint and parent story criteria 6 and 11 | "the Gate-A spec cycle of the next story whose spec is written" names neither a concrete artifact nor a durable selector, so concurrent sibling specs have no total ordering despite the text claiming this fixes that ambiguity; it also still selects a cycle after a revert even though §10 correctly says that cycle would run under the restored old rules and owe no nonce | the only promised real-nonce demonstration can be claimed by different cycles, missed by all of them, or falsely fail after rollback, leaving post-rule attribution unverified | designate a concrete story path in a durable record when the rules land or define an auditable tie-breaker, and condition the checkpoint on the new rules still governing that cycle +NIT | high | spec §8 prompt-standards n/a rationale | the quoted AGENTS.md example "typecheck: n/a — no typed sources" is not the table's text, which is "n/a — no typed sources (shell + markdown)" under a separate typecheck column | the analogy remains understandable, but the requested source verification cannot reproduce the quotation exactly | quote the actual cell text or remove quotation marks and paraphrase it +END OF FINDINGS (6 total) diff --git a/.context/codex-reviews/gate-a-spec-rle-pass-29.md b/.context/codex-reviews/gate-a-spec-rle-pass-29.md new file mode 100644 index 0000000..b807f3d --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-rle-pass-29.md @@ -0,0 +1,4 @@ +MAJOR | high | spec §2.4 lines 138-165 and parent story criterion 1/open-question answer | the opening qualification preserves a legitimate Gate-B skip, but the following rules still say any profile or cited-set change costs a further pass and requires a final clean pass regardless; lowering to a unanimously level-0 set during an active cycle can make an already-behaviourally-trivial diff legitimately skippable, so these categorical duties contradict the untouched skip | an implementation can force a review and pass after the settled rule has removed that review, leaving two incompatible closure procedures in both shipped copies | qualify every one-further-pass/final-clean-pass statement with the floor-governed review path and state that a newly legitimate Gate-B skip closes with its skip record and no further pass +MAJOR | high | parent story criterion 6 lines 227-243 versus spec §5 lines 333-340 | the story now requires a named record set but never names its members, while the spec fixes the set as provenance line, per-pass curve or skip record, findings slots, and advisory working record and excludes evidence entries and human-exception records | the parent can accept shipped text that names a smaller or different set, allowing nonce-free slots or working records and defeating the collision/data-loss protection while claiming story/spec agreement | mirror the spec's exact record list and exclusions in criterion 6, including whether the nonce is carried in the slot name or record body +MINOR | medium | spec §8 lines 509-519 and parent story criterion 6 lines 240-243 | the self-identifying obligation still assigns discharge to "the first" qualifying cycle even though concurrent sibling histories have no total order and no atomic shared discharge state; two closers can each see no prior discharge, or each can defer to the other, and the text neither permits duplicate discharges nor defines the history-visible test | the one promised real-nonce demonstration can be duplicated or missed, so its audit record is not deterministic under the concurrency the repair explicitly targets | define discharge idempotently: every qualifying closer with no discharge in its reachable history records one, concurrent duplicate discharges are valid, and any valid discharge satisfies the obligation +END OF FINDINGS (3 total) diff --git a/.context/codex-reviews/gate-a-spec-rle-pass-3.md b/.context/codex-reviews/gate-a-spec-rle-pass-3.md new file mode 100644 index 0000000..4c79d8a --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-rle-pass-3.md @@ -0,0 +1,55 @@ +MAJOR | high | §4.1 "The agent writes only when no user value exists" | no write order is defined for the marker and floor; a crash after writing the floor but before writing the marker leaves agent state that the next cycle must classify as user-set | an agent-derived floor, including 1, can become permanent and exempt from overwrite or cleanup | require marker-first creation, verify the marker, then write and verify the floor so every interrupted prefix fails toward recognizable agent state +MAJOR | high | §4.1 "Both are removed at cycle close" | no removal order is defined; removing the marker before the floor can leave an unmarked agent-written floor when the second removal fails | partial cleanup converts agent state into a supposedly user-owned knob and violates the never-remove or overwrite distinction | require floor-first removal followed by marker removal, verify both absent, and define the terminal state for each partial failure +MAJOR | high | §4.1 "A marker-bearing file found at cycle start is stale agent state" and lines 280-284 | the marker may belong to another live cycle, which §4.1 later admits can overlap, yet the start rule unconditionally calls it stale and overwrites it | one session can lower another higher-risk cycle's effective floor or remove its active floor at close | make a foreign or concurrently active marker a stop-and-surface state and define how the marker identifies the owning cycle before any overwrite +MAJOR | high | §4.1 "A floor file with no marker beside it is user-set" | the rule declares every unmarked file effective even though the cited hook ignores zero, empty, negative and non-numeric values and silently uses 3 | the agent's reported effective floor can disagree with the hook and pass accounting can use the wrong target | define the same positive-integer parse as the hook, report raw and effective values separately, and give invalid user content a terminal action +MAJOR | high | §4.1 "Failure has a terminal action" | naming only failed write, read-back, mismatch or removal does not enumerate the distinct causes, checks and remedies required by prompt-standards item 10; it also does not say whether the failed operation concerned the floor or marker | operators cannot distinguish permissions, directory-at-path, disk-full, concurrent writer and malformed-content cases, so the new prompt fails the binding diagnostic standard | enumerate checks and remedies for each file and each failure class, including the safe state to leave behind +MAJOR | high | §4.1 "cleanup attaches to each cycle type's own closing event" | cleanup is not ordered relative to the spec commit, plan commit or Gate-B amend | cleanup before the commit makes the hook re-read its default floor, while cleanup after the commit means a failed removal is discovered only after the cycle has already closed and cannot obey the stated stop | specify a pre-close verification, the exact commit or amend step, post-close cleanup, and the recovery record when post-close cleanup fails +MAJOR | high | §4.2 lines 293-305 | "closing requires meeting the floor as currently derived" and the one-pass lowering example ignore the settled rule that a user-set workspace floor is the effective floor | a user floor above the derivation can be silently bypassed after a profile lowering | state closure against the current effective floor, use the derived floor only when no user knob exists, and qualify the lowering example accordingly +MAJOR | high | §4.2 "current profile at each pass" | the spec never instructs an active agent-owned floor and marker to be rewritten and verified when a profile or cited-story set changes mid-cycle | the hook can continue enforcing the old value while reports and closure reason from the new derivation | add the mid-cycle transition for both files, including write order, read-back, provenance update and failure handling; leave a user knob untouched +MAJOR | high | §4.1 "Provenance" and story criterion 3 | every cycle must leave a floor trace, but the only normal form is `floor N per ` and §5 explicitly supports artifacts citing no story | an unprofiled no-story cycle has no legal provenance line, so "every cycle" is unsatisfiable and omission is indistinguishable from forgetting | define a pinned no-story or unprofiled provenance form and include it in both prompt copies and verification +MINOR | high | §4.1 "one entry per cited story" | a unanimous multi-story floor is derived from the cited set, but separate `floor N per ` lines make a level-0 member of a mixed set look as though it independently derived floor 3 | history misstates the derivation and makes omissions harder to audit | record the complete cited set and unanimity result in one pinned form, or make each line explicitly say it is a member of the named set +MAJOR | high | §10 "Abandonment and rollback" | an abandoned or blocked level-0 cycle deliberately leaves both marker and floor behind until another cycle starts | unrelated commits and resumed work in the interval are evaluated with the abandoned floor of 1, creating the accidental-persistence gate-off path §10 claims the lifecycle removes | define an abandonment and blocking-stop cleanup or safe-floor transition instead of relying on a future cycle to self-heal +MAJOR | high | §10 "Rollback to prompt text that predates this change" | rollback leaves the agent-written floor in place and the old prompt treats that file exactly like a user override while ignoring the marker | a derived floor of 1 becomes persistent under rules that no longer know how to clean or distinguish it | require rollback to remove verified agent-owned state or restore the safe default before the prompt rollback lands +MAJOR | high | §10 "recoverable from git plus file mtime" | a git commit timestamp and a gitignored file's mtime do not reliably establish creation order; checkout, copy, clock skew and `touch` all change that relation | an old loop can be classified as new or a new loop as old, changing which floor and duties apply | use an explicit cycle start record carrying the adopting rules revision, and treat missing or conflicting identity as a stop or conservative restart without claiming mtime proves order +MAJOR | high | §10 "treat the loop as new and re-derive, which costs passes and never skips them" | under the new rule a restarted level-0 loop can require one pass where the old in-flight loop required three | the fallback can skip old-rule passes, directly contradicting its stated safety property | require the conservative maximum of the old and new floor when start provenance is unavailable, or restart the old-rule floor explicitly +MAJOR | high | §10 "from the commit that ships them" | that commit exists only in this repository; the inline template is adopted later in unrelated repositories whose history contains no plugin shipping commit | scaffolded agents cannot evaluate the activation rule, and the two copies cannot be in parity without an unstated variance | define activation by the commit that adopts this §5 text in the current repository and state any root-versus-template wording variance explicitly +MAJOR | high | §10 "The first is bounded by §4.1's read-back" | read-back proves only that the bytes written equal the bytes requested; it does not compare the value with the cited profiles | an agent intentionally writing an unlicensed floor of 1 passes the claimed bound, so the mechanism overstates what it controls | say read-back detects write integrity only and leave semantic licensing explicitly unguarded or add an independent profile comparison +MAJOR | high | §10 gate-off lever enumeration | the list omits the cheapest marker-specific bypass: an agent can write 1 without a marker or remove the marker after writing, causing its own floor to be classified as user-set and exempt from cleanup | the two-state ownership claim and provenance can be forged without even falsifying a story profile | add marker omission or removal and false workspace-knob attribution to the disclosed residuals and require pass reports to expose marker state +MAJOR | high | §4.1 exposure and closure | reports read the value on each pass, but there is no final pre-close re-read of the effective floor, marker ownership, current profiles and cited set before the closing action | a floor or profile can change after the last pass and before close without another pass, while `.context/` changes do not invalidate Gate B | require an immediate pre-close semantic revalidation and another pass whenever it changes the effective obligations +MAJOR | high | §2.1 "Identity" | the spec first defines sameness as matching location and defect, then says a change to severity, consequence or suggested fix makes a new finding | the same later finding is simultaneously released and held, so an agent can choose the cheaper branch | define one non-contradictory identity predicate containing every field that matters +MAJOR | high | §2.1 decline record form and "Identity" | the durable record stores only location and defect, but later sameness depends on severity, consequence and suggested fix, whose original values disappear with gitignored findings files | a re-raised Blocker or materially stronger consequence can be mistaken for the declined finding after local artifacts are cleared | store the complete original finding line or all identity fields verbatim in the commit-body record +MAJOR | high | §2.1 "for the remainder of it" and squash carry | decline records persist into commit and squash history but carry no cycle identifier or explicit prohibition on reuse by later cycles | the same location and defect in a future cycle can inherit an old decline, turning a one-cycle exception into a permanent waiver | bind the record to an explicit cycle identity and state that historical decline records never discharge a later cycle +MAJOR | high | §2.1 accepted branch | acceptance ends the hold and Minor or Nit is collected, but the text does not state that the triggering pass still cannot be credited clean even though the settled exception to that duty is only an explicit decline | an accepted Minor can retroactively make the surfaced pass clean, broadening the exception beyond the settled decision | state that acceptance resumes the loop but the surfaced pass remains unclean and a later clean pass is required; only recorded decline can requalify that pass +MINOR | high | §3 decision procedure | "if you cannot name an in-system reader and a changed decision, the finding is Minor" assigns every non-gating issue to Minor and leaves no path to the existing Nit class | stylistic Nits are systematically inflated to Minor and the closed four-level vocabulary loses one decision boundary | make the reachability test decide gating versus non-gating, then retain the existing Minor-versus-Nit distinction within the non-gating branch +MAJOR | high | §1.2 "Blocker/Major must resolve is definitional" | a clean pass having no new Blocker or Major is not proof that a previously surfaced Blocker or Major was resolved; reviewer variance can simply fail to repeat it | the consolidation can drop the independent resolve duty and close over an unfixed earlier finding | keep resolution as a separate standing obligation and define clean completion as requiring both no new gating findings and all prior accepted gating findings resolved +MINOR | medium | §2 precedence table clean completion versus two-tell stop | clean is said to win, but the spec does not preserve the existing pass-4 reporting duty or say that triggering tells still appear in the closing report | implementers may read precedence as suppressing the tells rather than only suppressing the stop | state that clean completion cancels the suspension only; the three-line report and tell disclosure still run +MINOR | high | §5 line 319 and its table | the text says "Four shapes" but enumerates absent, partial, malformed, unreadable and stale | the stated count fails its own enumeration and the parity verification will inherit the wrong inventory | change the count to five +MAJOR | high | §5 stale row and lines 330-336 | stale is defined as a slot carrying no cycle identity, then the spec admits current slots carry none; the proposed per-cycle infix is not added to the existing exact slot grammar or to §6.2's rewritten-passage inventory | every current history file is stale by the stated check, while implementations may invent incompatible slot names and miss files | specify the new slot grammar and cycle identity in both findings-protocol copies, define migration, and account for every condition of that rewritten passage +MAJOR | high | §5 malformed-history treatment | the suggested remedy says a malformed historical pass may be rerun from its session ID, but at pass 4 the artifact may already differ from the revision reviewed in that earlier pass | a rerun can write current findings into an old pass slot and fabricate the historical trend | keep the historical slot unavailable unless the exact reviewed revision and branch can be proven and reproduced; otherwise rerun only as a new current pass +MAJOR | high | §5.1 separate and recovery calls | the curve counts logical passes while the kept hard-floor rule says passes are counted by the hook, but the hook counts each call; two branch calls or a recovery call can satisfy the hook floor before three logical passes exist | the newly documented valid call shapes create a false floor satisfaction path | make closure use validated logical-pass count whenever calls and passes differ, and explicitly forbid treating the hook count as the floor in those shapes +MAJOR | high | §5.1 separate-branch rule | no rule binds the spec and quality branches of one logical pass to the same artifact revision or forbids edits between their calls | two reviews of different states can be summed into one valid pass and one curve entry | record and compare a shared revision or fingerprint for both branches and invalidate the pair if the artifact changes between them +MAJOR | high | §5.1 curve form and squash carry | all three loop types use the same unlabeled `Findings ... Blockers ...` line, and squash carry can place several such lines in one commit body | P8 and later readers cannot tell Gate-A spec, Gate-A plan and Gate-B curves apart, defeating the all-three-loops measurement | prefix the pinned form with gate, artifact type and cycle identity, and define ordering when several records are carried +MAJOR | high | §5.1 "durable half across cycles" | finding and Blocker counts are author-written and no check compares them with validated pass files, yet the economics follow-up will consume them as measurements | fabricated or mistaken curves are durable but observationally indistinguishable from real ones | state that curves are unverified assertions and add a closing verification that derives each entry from the validated files before recording it +MINOR | high | §5.1 call/pass divergence annotation | `(N calls, M passes)` has no pinned placement, label or association with a particular curve despite the claim that the overall form is parseable | different authors can emit mutually incompatible records that P8 cannot parse reliably | include call and logical-pass counts in the single pinned curve grammar and provide examples for full, separate-branch and recovery shapes +MINOR | medium | §5 partial-history treatment | "compute historical tells over the passes present" does not say when too few comparable points remain to decide rising or failing-to-fall | a single surviving pass can be silently treated as a trend or as no tell depending on the agent | define the minimum observations for each historical tell and otherwise report that tell uncomputable +MAJOR | high | §6.2 row "Both gates are a LOOP with a HARD FLOOR" | the row omits the hook's inability to judge cleanliness or distinguish Gate-A runs, satisfied-count warning, fix-after-each-pass duty, keep-going-or-stuck branch, anti-padding rule, advisory validation and dismissal reason | "every other requirement kept verbatim" can pass while multiple old conditions disappear, violating the governing Don't | enumerate every condition in the current paragraph and mark each kept, moved, changed or dropped +MAJOR | high | §6.2 row "What a loop absorbs, and what stops it" | the row omits the both-correction-and-new-question precedence, the requirement that ancestry never supplies set membership, and the explicit resume-on-answer branch | a rewrite can lose exactly the contract-decision and resumption conditions this consolidation is meant to preserve | add those conditions individually and state their dispositions +MAJOR | high | §6.2 row "Recognizing clearly stuck" | the row compresses the three-part test and omits missing-one-means-continue, the six-pass observation's non-threshold status, stated coverage judgement, the known-unreviewed prohibition and genuine repair attempts | a replacement can weaken the exit while still being marked "all kept" | enumerate each predicate, qualifier and consequence separately +MAJOR | high | §6.2 row "Surfacing does not close the cycle" | the disposition says decline is an exception only to the third condition, while §2.1 and the settled decision require the qualification at all three universal rules | the accounting licenses an implementation that leaves the finding open or the resolve duty unqualified after a recorded decline | mark the explicit-decline qualification on open status, resolution duty and clean-pass credit individually +MAJOR | high | §6.2 row "From pass 4 onward" | the inventory omits activation at pass 4, the contents of each of the three lines, all five tell definitions, the report-the-triggering-tells duty, handoff to the user and clearly-stuck independence | the old decision procedure can be replaced while silently dropping the same conditions the field report says were previously lost | enumerate every line, tell, threshold input and terminal consequence individually +MAJOR | high | §6.2 row "The Gate-B triviality skip" | the row omits the profiled versus unprofiled evidence branches, battery obligation, evidence-entry placement and unchanged pre-existing triviality judgement | a rewritten skip can remove evidence or broaden unprofiled behavior while still being called fully kept | list and disposition each skip precondition, evidence duty and record location +MAJOR | high | §6.2 row "Severity" | the table says the old Major design-flaw rule and Blocker/Major resolve rule are kept, but the consequence-keyed classifier deliberately narrows which design flaws may remain Major and explicit decline qualifies resolution | the accounting hides two real changes as preservation, defeating criterion 6 | mark the classifier and decline qualification as changes to the old conditions rather than additions beside untouched rules +MAJOR | high | §6.1 and §6.2 completeness claim | §6.1 says all fourteen floor sites change, but §6.2 has no rows for the pass-acceptance paragraph, Gate-A heading, Gate-B re-review sentence or lens-set paragraph; it also omits the findings-slot passage changed by the discriminator and the close mechanics changed by cleanup | the inventory is not every rewritten passage despite claiming completeness | add a condition row for every changed surrounding passage, not only the twelve selected lead-ins +MAJOR | high | §7 line 469 | the spec calls README and docs paths "Gate-B N/A" even though the implementation also changes CLAUDE.md and a plugin prompt, and current §5 says any mixed commit fires full Gate B | the spec misdescribes the exact gate classifier it is forbidden to overstate | say those paths are prose individually but the mixed implementation commit receives full Gate B +MAJOR | high | §7 falsified-statement table | `docs/coding-workflow.md:79-80` also says Gate A's floor is the same at every profile level, but it is absent from the fix set | shipped user-facing methodology remains directly false after the change | add that site to the table, implementation surface and differential verification +MAJOR | high | §8 risk-path verification | checking only that the floor and marker are absent after close passes when the lifecycle never wrote either file | the named verification can report green with the entire derived-floor mechanism missing | require evidence that both files existed with verified contents during a cycle and were absent afterward +MAJOR | high | §8 risk-path verification and settled user-knob rule | no verification exercises an existing unmarked user floor through start, passes and close | the most destructive regression, overwriting or deleting the user's knob, can pass every planned check | add a differential state case proving the exact user bytes and absence of a marker survive unchanged +MAJOR | high | §8 risk-path verification versus §10 activation | §10 says all loops already in flight, including this spec's Gate A, finish under old rules and new rules begin only at the shipping commit, but §8 relies on this branch's own closes to exercise the new lifecycle | no pre-shipping cycle on the branch is licensed to run the behavior the evidence claims to verify | define a post-adoption verification cycle or change the activation boundary so the named evidence can actually execute under the new rules +MINOR | high | §8 "No automated test is possible for prose" | the absolute claim is contradicted by the same section's mechanical grep, parity checks and proposed lifecycle observations; only semantic adjudication lacks a complete automated oracle | the evidence rationale overstates a limitation and invites weaker verification than the testable mechanics permit | say no comprehensive automated semantic test is available and automate the mechanical portions +MINOR | high | opening "Surfaces edited" | a direct zero-context diff of the cited §5 ranges is 186 added-plus-deleted lines across 20 hunks, not 192 lines across 20 hunks | the mechanically checkable baseline count is stale in a spec that makes count accuracy a review requirement | update 192 to the observed count and state the counting method +MAJOR | high | story §2 outcome 2 versus story criterion 5 and spec §3 | the desired outcome still says an instrument keeps gating only for false green, while the settled criterion and spec require the carve-out in both false-green and false-red directions | two authoritative inputs prescribe different severity behavior | update the story outcome to the bidirectional consequence rule before implementation +MAJOR | high | story §2 "without judgement calls" versus spec §10 line 567 | the story requires a judgement-free severity decision, while the spec explicitly says the reachability test needs judgement | the spec cannot satisfy the stated desired outcome even if implemented exactly | reconcile the story's success claim with the settled consequence test by naming the remaining judgement boundary +MINOR | high | story §5 Open questions | multi-story floor aggregation and mid-cycle profile movement remain labeled open after revision 3 encodes settled answers for both | plan authors and reviewers can reopen or choose alternatives to decisions the user declared closed | mark both questions answered with the exact settled rules +MINOR | medium | invariant 11 and prompt-standards item 1 | the edited root CLAUDE.md prompt still has no target-model statement, and the spec does not account for one; the outer workflow-init command's target line does not appear in the scaffolded CLAUDE.md bytes | the changed prompt cannot be shown to pass all twelve binding checklist items | add a concise target-model statement to each CLAUDE surface or explicitly justify and verify how the enclosing command satisfies the template prompt +MINOR | high | opening "Surfaces edited" versus §7 Packaging | the opening inventory omits the manifest version bump and changelog entry that §7 later makes mandatory, and also omits the newly falsified coding-workflow page | scope readers can build an incomplete plan from the first declared surface list | make the opening inventory exhaustive and consistent with the packaging and falsified-doc sections +END OF FINDINGS (54 total) diff --git a/.context/codex-reviews/gate-a-spec-rle-pass-30.md b/.context/codex-reviews/gate-a-spec-rle-pass-30.md new file mode 100644 index 0000000..210e2e8 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-rle-pass-30.md @@ -0,0 +1,4 @@ +BLOCKER | high | §2.4 lines 143-148 and 169-172; parent story criterion 1 | the new "every rule in this section" precondition makes the later non-discharge rule inapplicable exactly when citation removal makes Gate B skippable, even though an accepted in-set Blocker or Major remains in the fixed repair set independently of the citation | removing the citation can now close by skip with a serious accepted finding unresolved, reopening the gate-off path the story expressly forbids | scope the exception only to the further-pass and final-clean-pass duties; state that skip removes the review but never an existing accepted Blocker/Major repair obligation +BLOCKER | high | §2.4 lines 143-148 versus §2.3, §4, CLAUDE.md lines 424-433 and the template lines 603-612 | the statement that the skip reason and skip record "are what govern instead" omits the existing battery and per-profile evidence duties and the new every-cycle provenance line | an implementation can treat a newly skipped cycle as owing only those two records, closing without required validation or evidence and leaving the floor unreconstructible from history | enumerate every duty that survives a skip: the battery, one evidence entry per cited profiled story, the provenance line, the skip reason and the skip record in place of the curve +MAJOR | high | §2.4 lines 143-148 versus CLAUDE.md lines 424-439 and the template lines 603-618 | the restatement defines a legitimate skip as a behaviourally trivial diff plus a unanimously level-0 cited set, omitting the existing judgment-based eligibility of an unprofiled story and therefore mixed cited sets whose members are each skip-eligible | a cited-set or profile transition to an unprofiled case is forced through another Gate-B pass even though the settled skip rule removes that review, producing a false red that blocks a valid trivial change | reference the existing §5 eligibility rule without narrowing it, or state both member cases and retain the requirement that every cited story be individually skip-eligible +END OF FINDINGS (3 total) diff --git a/.context/codex-reviews/gate-a-spec-rle-pass-31.md b/.context/codex-reviews/gate-a-spec-rle-pass-31.md new file mode 100644 index 0000000..efe05e4 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-rle-pass-31.md @@ -0,0 +1,2 @@ +BLOCKER | high | §2.1 lines 75-79 and §2.4 lines 138-147 | the spec still restates §5's exits, and §2.4 says a legitimate skip is "unaffected by a moving profile" even though §5 makes profiled Gate-B skip eligibility depend on the current profile and cited set; this also contradicts §2.4's next paragraph delegating the skip to §5 and the settled rule that only the pass-count number is replaced | an agent consuming this rule can preserve a formerly valid skip after a profile or cited-set change revokes eligibility, omitting Gate B on a now nontrivial or higher-risk cycle | remove the exit summaries from §2.1 and §2.4; state only that the derived value replaces the pass-count number and that the further-pass rules apply while §5 says a gate is running, leaving §4's new skip-record duty as the sole skip-specific addition +END OF FINDINGS (1 total) diff --git a/.context/codex-reviews/gate-a-spec-rle-pass-32.md b/.context/codex-reviews/gate-a-spec-rle-pass-32.md new file mode 100644 index 0000000..e7bb3eb --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-rle-pass-32.md @@ -0,0 +1,3 @@ +BLOCKER | high | parent story criterion 1 at lines 142-147 | the criterion still says any cited-set or profile change requires a final clean pass, without the spec's "while §5 says a gate is running" scope | an implementation satisfying the criterion would reintroduce a review after a legitimate Gate-B skip, while an implementation preserving §5 would fail its parent criterion | qualify the duty to cycles that §5 says are still running a gate and refer to §5 for whether the gate runs; keep the accepted in-set Blocker/Major clause separate +MAJOR | high | §4 lines 228-307, §5 lines 349-358, and successor settled input 10 | the curve requires numeric counts for every valid pass, but a recovered cycle can retain one nonce candidate after its prior pass files and count history are unavailable; only the model field has an undetermined production, and recovery says a single candidate keeps identity | the closing agent must fabricate counts, violate the grammar, or leave the cycle unable to close, and P8 can ingest a false severity mix | add an explicit unavailable-history production and P8 treatment, or define a compatible recovery transition that never claims missing valid-pass counts +END OF FINDINGS (2 total) diff --git a/.context/codex-reviews/gate-a-spec-rle-pass-33.md b/.context/codex-reviews/gate-a-spec-rle-pass-33.md new file mode 100644 index 0000000..26215e9 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-rle-pass-33.md @@ -0,0 +1,2 @@ +MAJOR | high | §4 "A count that cannot be recovered..." and P8 acceptance criterion 5 | the spec turns a count-level unknown into a pass-level exclusion, while the P8 story never requires any `?` handling; when one series is `?` and the others are known, excluding the pass from every comparison discards usable data and leaves the deferred consumer without the settled rule | P8 can shrink or bias different comparisons and cannot reproduce the intended count-level treatment of unknowns | exclude only the affected count from comparisons that require it, retain known series where valid, report exclusions per series, and mirror that rule in the parent curve criterion and P8 criterion 5 +END OF FINDINGS (1 total) diff --git a/.context/codex-reviews/gate-a-spec-rle-pass-34.md b/.context/codex-reviews/gate-a-spec-rle-pass-34.md new file mode 100644 index 0000000..276a97f --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-rle-pass-34.md @@ -0,0 +1,2 @@ +MINOR | high | §4 "A count that cannot be recovered..." | the rationale says the series most often lost is not the same across cycles, but neither this artifact nor any cited record establishes that distribution and future cycles may lose the same series | the per-series rule is supported with a categorical empirical claim the evidence cannot bear, weakening the rules-only artifact even though preserving known values is sufficient rationale | say the lost series can differ across cycles and the discard can therefore be uneven, or delete that clause and rest the rule on retaining known values +END OF FINDINGS (1 total) diff --git a/.context/codex-reviews/gate-a-spec-rle-pass-4.md b/.context/codex-reviews/gate-a-spec-rle-pass-4.md new file mode 100644 index 0000000..5e372c0 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-rle-pass-4.md @@ -0,0 +1,41 @@ +BLOCKER | high | story criterion 3 versus §4.1 | the current story still requires one floor entry per cited story and reverses the settled meanings by recording `floor N per workspace knob; profile derivation M`, while revision 4 requires a profile-derived floor plus a separately labelled hook reminder threshold | no implementation can satisfy both the authoritative story and the settled no-floor-file design | update criterion 3 to the revision-4 provenance contract, including unprofiled and multi-story forms, before planning +BLOCKER | high | story criterion 4 versus §§4.1 and 10 | the current story still requires both shipped copies to say the floor is agent-written in gitignored state and that writing 1 is the cheapest gate-off lever, the exact mechanism revision 4 deletes | implementing the spec would fail an unchecked acceptance criterion, while implementing the criterion would recreate the rejected mechanism | replace criterion 4 with the residual-risk disclosure for the prompt-only derived floor and state explicitly that this design writes no floor file +BLOCKER | high | §2 precedence table, clean completion × scope stop | the hold is stated as lasting while a finding is “unresolved and undeclined”; after the user accepts a Minor or Nit it remains undeclined and has no repair resolution, so the table still blocks closure despite §2.1 saying either answer ends the hold and accepted Minor or Nit is merely collected | the accepted direction can deadlock or silently acquire a different rule from the settled both-directions decision | make the scope hold depend on the decision being unanswered; after either answer release that hold, then apply severity so only accepted unresolved Blocker or Major blocks closure +BLOCKER | high | §2.1 and §6.2 row 6 | §2.1 says a recorded decline releases the finding, but row 6 says decline changes only clean-pass credit, leaves the finding open, and does not waive the resolve duty; §1.2 also requires every surfaced Blocker or Major to be resolved with no decline qualification | a declined Blocker or Major is simultaneously released and permanently closure-blocking, and the old-conditions accounting licenses the wrong implementation | state the explicit-decline effect consistently at each universal rule named by the story, while preserving the resolve duty for accepted Blocker or Major only +MAJOR | high | story §5 “Open questions” | multi-story floor aggregation and mid-cycle profile movement remain presented as open questions even though revision 4 and the settled inputs choose unanimity, current-profile derivation, banked-pass counting, and current-floor closure | plan authors can reopen or choose alternatives to decisions the user declared closed | mark both questions answered and copy the exact settled rules into the story +MAJOR | high | §4.1 provenance form | `floor N per ` has no legal representation for an artifact citing no story, and one line per story misstates a mixed set because a level-0 member does not independently derive the set’s floor of 3 | the “every cycle” provenance duty is unsatisfiable for one supported profile case and ambiguous for multi-story cycles, making omissions hard to audit | pin separate no-story, cited-unprofiled, single-story, and complete-cited-set forms, with the whole set named wherever unanimity determines the floor +MINOR | high | §4.1 claim about `$floor` at hook lines 933–973 | `$floor` does not appear only inside advisory `note` messages: hook lines 946 and 966 use it in branch conditions that select below-floor versus satisfied output | the mechanical explanation is false even though the broader conclusion that the hook only emits advisory output is correct | say the value controls the hook’s advisory threshold and resulting branch/message, and cite both the comparisons and interpolations +MAJOR | high | §4.2 “On lowering” | the claim that a high cycle lowered to trivial can close on one pass ignores the kept rule that the final clean pass must run under the current profile; a clean pass banked before the lowering cannot become the final current-profile pass | a mid-cycle lowering can retroactively close on evidence gathered under superseded lenses and mode, contradicting the spec’s own raise analysis and current §5 | state that banked passes still count toward the lowered floor but at least one clean pass must run after the confirmed profile change before closure +MAJOR | high | §§4.1 and 10 marker-deletion rationale | deleting the marker does not remove the camouflage path: an agent can still write `1` directly to the gitignored user knob, omit or forge the provenance line, and obtain a satisfied hook message with no diff showing who wrote the knob | the spec claims the dangerous reminder-suppression route was removed when it merely stopped being part of the sanctioned mechanism | disclose unauthorized knob mutation as an unverified attack, require pass reports to state raw/effective knob state, and avoid claiming marker deletion removed camouflage +MAJOR | high | §10 “The gate-off surface, honestly enumerated” | the purportedly complete four-item list omits the easiest existing review-avoidance paths: falsely claiming Gate-B triviality, claiming a zero-finding early exit, writing `codex-gate.off`, lowering the user knob, removing adoption signals, or fabricating pass validation and decline records | readers can treat a partial attack list as complete and miss the actual ways an agent that wants to skip review would act | either scope the list explicitly to non-exhaustive derived-floor provenance attacks or enumerate the remaining policy, reminder, adoption, early-exit, and record-forgery paths with their observability limits +MINOR | high | §10 statement-attack paragraph | the text says all four residuals live in the commit body, but story-profile minting lives in a story file and omitted or incomplete cited sets are detected by comparing artifact and prompt inputs, not by reading the body alone | the observability claim directs reviewers to an incomplete source and overstates what history exposes | name the actual surface for each attack and state that the commit body is only an author-written cross-reference +MAJOR | high | §3 reachability test | “a gate” is allowed as the in-system reader without excluding the gate currently discovering the defect, so any Gate-A finding can name Gate A’s own decision to raise that finding and thereby escape demotion | the classifier is self-satisfying for spec and plan narration and can preserve the mis-calibration this change is meant to reduce | require a downstream operational reader and changed decision independent of the review that surfaced the finding +MINOR | high | §3 decision procedure | the rule assigns every non-reachable issue to Minor and provides no path to the existing Nit class | stylistic Nits are systematically inflated and the closed four-level vocabulary loses one boundary | make reachability decide gating versus non-gating, then retain an explicit Minor-versus-Nit distinction inside the non-gating branch +MAJOR | high | §2.1 one-cycle binding | `` has no defined construction, uniqueness rule, or comparison with the active Gate-A spec, Gate-A plan, or Gate-B cycle | a persisted decline can be reused across later cycles when human-readable ids collide or an agent cannot determine which cycle a record belongs to | define a canonical cycle identity and the exact boundary check that makes historical records inert, failing toward a new finding when identity is absent or ambiguous +MINOR | high | §5 “Four shapes” table | the passage claims four shapes but enumerates absent, partial, malformed, unreadable, and stale | the diagnostic inventory fails its own mechanically checkable count | change four to five +MAJOR | medium | §5 current-pass cluster tells | the spec says the two current-pass cluster tells can jointly trigger the two-tell stop but never defines what count or proportion constitutes clustering, whether one finding can form a cluster, or how a pass split between instrument and prose sets both tells | the same pass can mandatorily stop or continue depending on agent taste, including the no-history path this section is meant to settle | define the predicate for each cluster tell and the zero-, one-, and mixed-category cases +MAJOR | high | §5 missing-history diagnostic table versus prompt-standards item 10 | the “Treatment” cells do not enumerate the causes, distinguishing checks, and cause-specific fixes for absent, partial, or stale history; fresh checkout, wrong root, deletion, incomplete write, and foreign-cycle reuse collapse into the same labels | the new shipped prompt would violate the binding diagnostic-state standard and can send a resumed loop down the wrong recovery path | add cause, check, and remedy branches for each state, explicitly preserving uncomputable when the cause cannot be distinguished +MAJOR | high | §5 stale-history mitigation and §6.2 | the proposed per-cycle slot infix changes the exact findings-slot grammar at both §5 copies, but no new grammar, migration rule, or condition-inventory row covers that passage | implementations can invent incompatible names, current bare slots become ambiguously stale, and the claimed old-conditions accounting is incomplete | specify the slot grammar and cycle discriminator in both copies, define handling of existing bare slots, and add the rewritten findings-protocol passage to §6.2 +MAJOR | high | §5 malformed-history treatment | a malformed earlier pass may be rerun from its session id without proving the artifact and branch revision still equal the earlier pass’s input | a resumed call can write current findings into an old pass slot and fabricate the historical curve | permit recovery into the old slot only when the exact original revision and branch are proven; otherwise keep it unavailable and run a new current pass +MAJOR | high | §5.1 valid-pass ordering | incomplete passes are excluded while the next sentence says a zero entry preserves “position equals pass number” | after any incomplete pass the list position is only a valid-pass ordinal, so consumers mislabel every later value | either include an explicit incomplete placeholder in the pinned grammar or say positions are valid-pass ordinals and record original pass numbers separately +MAJOR | high | §5.1 branch/artifact-revision binding | “same artifact revision” has no defined identifier or check, and when a revision changes between branches the spec says the second starts a new pass without saying how its missing peer branch is obtained or how the first orphan is classified | branches from different revisions can still be summed, or a lone second branch can be credited as a pass | pin the revision evidence each branch records, invalidate the old pair on mismatch, and require both branches against the new revision before one logical pass is valid +MAJOR | high | §5.1 call count versus logical-pass floor | the spec distinguishes hook calls from logical passes but does not state unambiguously that closure uses validated logical-pass count when separate branch calls or recovery calls make the hook counter larger | two calls for one Gate-B pass can satisfy a floor of 3 before three valid logical passes exist | make the derived floor compare only against validated logical passes and state that the hook count never licenses closure when the numbers differ +MAJOR | high | §5.1 pinned curve versus docs/coding-workflow.md:261–280 | the new pinned curve records findings and Blockers but omits the existing requirement to record the actual reviewer model beside each pass count | the durable pass record loses the family-independence evidence current methodology requires and cannot be interpreted after provider switches | include per-pass model identity or an explicitly adjacent model vector in the pinned grammar and preserve the undetermined case +MINOR | high | §5.1 call/pass annotation | `(N calls, M passes)` has no pinned placement or association with a particular labelled curve despite the claim that the form is parseable | squash bodies with multiple loops can attach the annotation to the wrong curve | include call and pass totals inside each curve’s exact grammar and examples +MAJOR | medium | §5.1 skipped Gate-B path | “all three loops” and “each loop records” never define the curve record for a legitimately skipped Gate B, where there are zero valid review passes but a closing commit still exists | authors can omit Gate B silently or emit incompatible empty records, weakening P8’s measurement and skip observability | pin a labelled `Gate B: skipped` or zero-pass form tied to the existing skip reason and evidence entry +MAJOR | high | §6.2 row 7 versus §2 clean/two-tell precedence | row 7 says the existing “any two makes stop-and-surface mandatory” consequence is all kept, while §2 changes that consequence so clean completion wins | the accounting hides a real qualification and can produce shipped prose that simultaneously mandates closure and suspension | mark the terminal consequence changed for the clean-pass conflict while keeping the three-line report and tell disclosure mandatory +MINOR | high | §6.2 row 5 | the disposition says the clean-precedence sentence is preserved verbatim in §2’s table, but the table replaces “a clean completion takes precedence over this exit” with “Clean wins” plus meta-commentary | the old-condition record makes a mechanically false preservation claim | mark it moved and semantically preserved, or include the original sentence verbatim +MAJOR | high | §6.2 row 13 | the disposition says re-review-after-fix reasoning generates the floor’s number, but that rule generates however many passes repairs require and does not derive either policy value 1 or 3 | the replacement rationale overstates the mechanism and can teach an implementer to conflate repair passes with the profile-derived minimum | derive 1 or 3 only from the profile predicate and retain re-review-after-fix as an independent invalidation duty +MAJOR | high | §6.2 completeness claim | the inventory omits the “Changing a profile” passage that §4.2 must amend and the findings-file slot passage that §5 changes with a discriminator; it also does not account for adding the prose-exemption rationale to the template | conditions such as human confirmation, log append, override invalidation, current-profile final pass, WIP amend, slot validation, and delete-before-call can be dropped while the spec claims full accounting | add a row for every surrounding passage the implementation rewrites and disposition each existing condition individually +MINOR | high | opening “Surfaces edited” | a direct zero-context diff of the cited §5 ranges is 186 added-plus-deleted lines across 20 hunks, not 192 lines across 20 hunks | the mechanically checkable parity baseline is stale | change 192 to 186 and state the counting method +MAJOR | high | §7 falsified-statement table | the table misses docs/getting-started.md:44–45, which says the hook reports when the Gate-A floor is unmet, and :58–59, which makes the hook’s satisfied message the trigger to close Gate B; after this change the knob threshold may disagree with the derived floor in either direction | user guidance can require extra passes at level 0 or authorize closure below a derived floor of 3 | add both sites and rewrite them so normative closure comes from §5’s derived floor while hook output is only the knob-based reminder +MAJOR | high | §8 risk-path verification timing | the verification requires this branch’s closing commits and their final provenance lines, but mode-derived evidence must exist before Gate B and the Gate-B closing amend does not exist until after the final pass; committing first also resets the hook cycle | the named risk verification cannot be completed at the point §5 requires it, so its evidence can be circular or unreviewed | split the check into pre-Gate-B verification of existing Gate-A bodies plus the proposed Gate-B message, then a post-close audit with an explicit reopen path if the actual body differs +MAJOR | high | §8 user-knob verification | the test requires an existing user-set floor but gives no executable path when the workspace has none; an agent creating or changing one would violate the settled rule that the knob is purely the user’s | a mandatory named verification can be impossible on a normal workspace or can induce the very unauthorized write it is meant to detect | make it conditional on a pre-existing user knob or require an explicit user-created fixture, and define the no-knob observation separately +MINOR | high | §8 “No automated test is possible for prose” | the absolute claim is contradicted by the same section’s grep, exact-form, parity, revision-diff, and commit-body checks; only semantic correctness lacks a comprehensive automated oracle | the rationale invites weaker verification of mechanics that are automatable | say no comprehensive automated semantic test is available and automate or mechanically verify the structural parts +MAJOR | high | §§8 and 10 versus story criteria 2 and 9 | §10 says this design’s in-flight Gate-A spec loop finishes under old rules, while the story requires this branch’s own commits to carry new labelled curves and floor provenance for every loop that ran | the activation rule and branch-level acceptance evidence point in opposite directions, so the demonstration may be omitted or performed under rules that supposedly do not bind | state which new records this branch voluntarily owes despite old loop semantics, or provide a post-activation cycle that satisfies the criteria without rewriting the in-flight floor +MAJOR | high | §10 unknown in-flight fallback | treating an unidentifiable loop as new does not necessarily “cost passes and never skip them”: a level-0 loop begun under the old fixed floor of 3 can be reclassified under the new floor of 1 | two banked obligations can disappear on the exact ambiguous migration path that is claimed to fail conservatively | require the conservative maximum of the possible old and new floors when start rules are unknown, while still applying the current profile’s lenses and final-clean-pass duty +MAJOR | medium | §10 downstream adoption | “rules bind from the `/workflow-init` run that writes them” does not cover skip, overwrite, or partial merge outcomes from invariant 9’s required diff-and-ask flow | a downstream project can be classified as newly governed after only part of the consolidated decision procedure was accepted | bind activation to the destination §5 containing the complete new rule set and treat partial or unverified adoption as old or conservatively new with an explicit diagnostic +MAJOR | high | §8 prompt-conformance evidence | the evidence plan acknowledges that the mechanical checker is not coverage but never requires a fresh twelve-item `docs/prompt-standards.md` review of each changed prompt surface | the principal quality invariant for this prompt product can remain unverified even when every listed evidence bullet passes | add an explicit item-by-item prompt-standards review for root §5 and the scaffolded copy, recording deliberate differences and failures +MINOR | high | invariant 11 and prompt-standards item 1 | the edited root CLAUDE.md and the scaffolded CLAUDE template still contain no target-model statement, and the command’s outer target line is not among the bytes written downstream | both changed prompt artifacts still fail a literal binding checklist item | add a concise target-model statement to each CLAUDE surface or explicitly establish why the generated prompt satisfies item 1 +MAJOR | high | opening “Surfaces edited” and §7 | the declared implementation surface omits the story edits required to remove stale criteria and open questions, and omits the manifest version bump and changelog entry that §7 later makes mandatory | a plan built from the opening scope can leave authoritative contradictions and fail packaging | make the opening surface inventory exhaustive, including story reconciliation, plugin manifest, changelog, both prompt copies, and every falsified user-facing statement +END OF FINDINGS (40 total) diff --git a/.context/codex-reviews/gate-a-spec-rle-pass-5.md b/.context/codex-reviews/gate-a-spec-rle-pass-5.md new file mode 100644 index 0000000..c58a743 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-rle-pass-5.md @@ -0,0 +1,34 @@ +MAJOR | high | opening claim at lines 8-9 | the spec says no file the hook reads is written by the design, but CLAUDE.md is an edited surface and gate_citation reads CLAUDE.md on every hook invocation | the scope and compatibility analysis excludes a real hook input and can miss adoption or citation effects of the prompt edit | say that no hook state or control file is written and explicitly acknowledge that the hook reads the edited CLAUDE.md for adoption and policy citation +MAJOR | high | §4 "Why the text must say the floor governs both gates" and story criterion 1 | the cited hook locations operate on the workspace reminder threshold, not the independently derived policy floor, so the hook's single floor variable does not establish that the derived value governs both gates | the rationale conflates the two values immediately before §4.1 insists they are distinct, violating the exact-comparison rule | make both-gate scope a prompt-policy decision and describe the hook citation only as evidence that reminders use one separate threshold +MAJOR | high | §4.1 lines 263-266 | the claim that the floor variable appears only inside advisory note messages is mechanically false because lines 946 and 966 compare pass counts against it to select the below-floor or satisfied branch | the spec overstates the hook mechanism and obscures the actual way the knob changes reminder behavior | state that the variable is used only to select and render advisory outcomes and cite both comparisons as well as the messages +MINOR | high | opening surfaces paragraph lines 15-17 | the two stated §5 ranges currently differ by 186 added-plus-deleted lines across 20 hunks, not 192 lines across 20 hunks | the mechanical baseline and the parity verification start from a stale count | replace 192 with 186 or define and reproduce a different counting method +NIT | high | §1.1 first blockquote | the quoted sentence ending "question is answered." is exact in the template but not in CLAUDE.md, where the text continues with an em dash instead of that period, and the spec does not disclose the quote difference | a mechanically asserted quotation is not present in both copies as required by this review protocol | quote the common substring without altered punctuation or identify the one-copy punctuation difference +MAJOR | high | story criterion 6 lines 174-184 versus spec §1.2 | the criterion requires the decline qualification at the Blocker/Major-resolve duty, while the settled decision and §1.2 say the decline never qualifies that duty and applies only because a declined finding is out of set | no implementation can satisfy both the story and the spec, so acceptance can force the exact resolve-duty exception the design forbids | rewrite the criterion to require the qualification only at the hold and clean-credit rules and state that in-set Blocker/Major resolution remains unconditional +MAJOR | high | §2.1 lines 97-99 and §6.2 row 6 | the spec promises qualification at three universal rules without naming a coherent set: §1.2 says the resolve duty is not qualified, while row 6 says the decline touches only the third of four listed conditions | the plan cannot know which passages to edit, and parity can be achieved while one required rule is omitted or the resolve duty is weakened | enumerate each qualified rule by exact current passage and give its replacement disposition, keeping the in-set resolve duty explicitly outside that list +MAJOR | high | §2.1 "The duty is qualified for... a finding the user has explicitly declined" | the decline record is not restricted to a finding surfaced for out-of-set membership, yet the argument later assumes every declined finding is out of set | an already in-set Blocker or Major could be relabelled "declined" and released, bypassing the settled resolve duty | permit this record only when the surfaced decision is whether the assigned set expands and the user's answer is that the finding remains out of set; forbid it for known in-set findings +MAJOR | high | §2.1 commit-body transport | a Gate-A spec or plan commit does not exist while its loop is running, but the commit body is declared the record and the dispositions file is explicitly only advisory | a Gate-A decline cannot become an authoritative same-cycle record before later passes need to test it, so the hold has no defined discharge path | define where the authoritative pending record lives before the Gate-A closing commit and how it is copied atomically into that body, or change the transport to one that exists during the loop +MAJOR | high | §2.1 cycle identifier definition | Gate A reviews changing raw text before its closing commit exists, so the current HEAD does not identify the artifact revision; Gate B can have distinct WIP snapshots or concurrent branches with the same parent, so a WIP parent is not unique either | declines can collide across concurrent cycles or bind to a different revision, contradicting the claimed one-cycle isolation and undermining §5.1's same-revision test | define a stable collision-resistant cycle id created at cycle start and carried unchanged through resumes, records, WIP amendments and squash, with an explicit artifact-revision field separate from it +MAJOR | high | §1.2 closing-report resolution duty | the closing report must name every earlier in-set Blocker and Major and its resolution, but no durable inventory, identity rule, or output form is defined for those findings, while prior slot history may be absent by §5 | an interrupted or resumed cycle can produce a quiet final pass without enough information to discharge the new precondition | define the source, identity and exact report form for the full in-set Blocker/Major resolution inventory, including unavailable-history handling +MAJOR | high | §3 self-exclusion clause | the proposed operative wording excludes "the gate currently finding the defect", not the reviewing pass, even though the explanation says gates remain legitimate operational readers and only the current pass is excluded | a reader can exclude the entire Gate A or Gate B mechanism and demote a genuine rule defect, or count the gate and defeat demotion, depending on interpretation | say that the specific reviewer invocation raising the finding cannot count, while a gate consuming the rule during normal operation can +MAJOR | high | §3 named-reader test | the purported stack-neutral test enumerates only a gate, skill, rule, escalation procedure or scaffolded template as readers | in downstream application repos an API handler, runtime, compiler, job, database consumer or external system can change product behavior yet fail the enumerated test and be forced to Minor | ask for any in-system component or consumer whose decision or behavior changes, and keep the prompt-specific subjects only as examples +MAJOR | medium | §3 decision procedure and §6.2 row 14 | the reachability test determines only Minor versus "keeps its severity", while the retained Mechanics definitions separately choose Blocker versus Major, but the spec never pins their ordering as one decision tree | two readers can apply the old subject definitions before or after the new test and reach different severities despite story criterion 5 requiring exactly one procedure | state one ordered algorithm: test operational consequence first, then classify reachable consequences with the retained Blocker/Major definitions, otherwise Minor +MAJOR | high | §4.1 common provenance form | one `floor N per path` entry per cited story falsely attributes a set-derived floor to each story individually; for example a level-0 story in a mixed set would be recorded as independently producing floor 3 | history cannot reconstruct the unanimity calculation and P8 can misclassify which profiles caused the floor | pin one set-level record listing every cited story and each profile state, followed by the single floor derived from the whole set +MAJOR | high | §4.1 knob clause and §8 conditional verification | §4.1 includes the knob clause only when a knob is "set", but §8 requires disclosure whenever the file exists; an empty, unreadable, zero, negative or nonnumeric file has no valid threshold M to disclose | the verification is unsatisfiable or will misstate the hook's effective default on invalid-file paths | define valid, invalid and unreadable knob states, record the hook's effective threshold plus the file state, and make the conditional verification use the same predicate as the hook +MINOR | high | §5 lines 338-348 | the text announces four shapes but the table enumerates five: absent, partial, malformed, unreadable and stale | the stated count fails its own enumeration and weakens the promised mechanical inventory | change four to five +MAJOR | high | §5 malformed-history row | the row claims to use the pass-acceptance checks but lists only terminator, count and non-finding-line checks, omitting field count, empty fields, severity structure, wrong-path and full-Gate-B two-branch validity | an invalid prior pass can enter the curve and alter trend or two-tell decisions, producing a false stop or false continuation | reference the complete existing acceptance procedure normatively and list only additional history-specific checks +MAJOR | high | §5 stale-history convention and §6.2 row 18 | an infixed name such as gate-a-spec-rle-pass-5 is not one of the three exact slot grammars currently specified, and no rule maps the arbitrary infix to the cycle identifier that stale detection is supposed to compare | implementers can preserve the old grammar or add an incompatible convention, while foreign files remain indistinguishable and concurrent writers still collide | formally replace the slot grammar with a defined cycle-id component, specify its encoding and expected-value source, and mark the old names changed rather than kept +MAJOR | high | §5.1 pinned curve form | the only displayed curve omits the required pass-number mapping and per-pass model, and no complete grammar shows how calls-versus-passes or noncontiguous pass numbers are encoded | the supposedly machine-greppable record has multiple valid interpretations and story criterion 2 can pass with an unparseable form | provide one exact complete example and grammar covering loop label, pass numbers, finding counts, Blocker counts, model per pass and optional call count +MAJOR | high | §5.1 example Gate-A plan loop | the example shows two nonzero Gate-A plan passes, which cannot close under floor 3 and does not declare a level-0 story or zero-finding early exit | the prompt's output example teaches an invalid close in the same section defining durable evidence | use a floor-valid three-pass example or include an explicit floor-1 provenance context that makes the two-pass example valid +MAJOR | medium | §5.1 model recording | one logical Gate-B pass can be assembled from separate spec and quality calls, but the spec defines one model "each pass ran under" and does not handle branches run under different models | model independence and later economics analysis become ambiguous exactly on the recovery and split-call paths the design supports | require both branch models when they differ, or declare such a split incomplete until both branches run under one recorded model +MAJOR | high | §6.2 rows 4 and 6 | the accounting says the old conditions are kept, but a decline changes both "the finding stays open" and "the loop resumes on the revised artifact"; row 6 instead says only its third listed condition changes | the replacement can silently drop the revised-artifact condition and fail to qualify the open-finding condition, violating the old-condition Don't | mark each affected condition changed, state that a decline may resume on an unchanged artifact, and show the exact qualification at the open and clean-credit sentences +MAJOR | high | §7, §8 parity scope and §10 gate-off residual | story criterion 4 requires the residual routes and no-guard disclosure in both shipped §5 copies, but the spec leaves that material in design-risk §10 and parity explicitly covers only rules newly inserted by §§2-5 | implementation can omit the required disclosure from both prompts and still satisfy the stated parity verification | name the exact §5 insertion in both copies and add it to the parity and prompt-standards verification inventory +MAJOR | high | §8 revalidation timing lines 582-588 | the spec requires the final amended commit to be re-read before the cycle closes, but the amend is itself the hook-observed closing event and that commit does not exist before the amend | the prescribed ordering is impossible; post-amend drift is discovered only after the reviewed cycle has already been reset | prepare and verify the exact final message before the closing amend, then define a reopen-and-review path if the resulting commit differs +MAJOR | high | §8 no-floor-file verification | comparing pre-cycle and closing state can show persistence or byte identity but cannot prove that no transient floor-file write and removal occurred on any path | the evidence plan claims stronger write-history coverage than its observation can provide | narrow the claim to no instructed or persistent write and inspect the implementation text, or add an actual write-observation mechanism without touching the user's knob +MAJOR | high | §8 prompt-standards verification | the pass is scoped to changed prompt regions, although invariant 11 binds the resulting prompt and item 7 requires whole-prompt contradiction checking; the curve and degraded-history outputs also lack compliant complete examples | a locally clean paragraph can ship inside a globally contradictory or format-incomplete CLAUDE.md and still be recorded as a twelve-item pass | review each complete resulting prompt as well as every edited region and require exact output examples for each new record and report form +MAJOR | high | §10 downstream partial adoption | the three changes are justified as coupled, yet downstream projects are allowed to bind whichever fragments their CLAUDE.md happens to contain, with no rule for combinations such as the new decline release plus the old resolve or clean duties | a user-approved partial merge can create a deadlock or an unintended gate-off path indefinitely | define an atomic compatibility boundary or require the old rule set with floor 3 until all coupled clauses are present, while preserving workflow-init's ask-before-merge behavior +MAJOR | high | §10 unknown-start fallback | floor 3 is defined when a loop's starting rules cannot be established, but severity, precedence, decline records, curve duties and closure semantics are left undecided | an in-flight or resumed loop can mix old and new rules even while conservatively using three passes | specify the complete fallback rule set, not only its floor, and state how the agent identifies or reports an indeterminate start +MAJOR | high | story acceptance coverage | no acceptance criterion observes the settled prohibition on writing .context/codex-gate.floor or the conditional byte-preservation of a user's knob | an implementation can reintroduce the removed marker/write mechanism and still satisfy every criterion | add an observable criterion that no agent path writes, removes or rewrites the knob and that an existing valid knob is preserved and disclosed +MAJOR | high | story criterion 5 coverage | the criterion does not require exclusion of the reviewing pass, although that exclusion is necessary to prevent every Gate-A finding from naming itself as its reader | the spec's core anti-self-reference obligation can be dropped while all four stated severity properties still pass | add the current reviewing invocation exclusion and the continued legitimacy of gates as operational readers to criterion 5 +MAJOR | medium | story acceptance coverage for §10 | no criterion covers activation, the floor-3 unknown-start fallback, rollback compatibility or downstream partial adoption, all of which are normative spec obligations | the plan can omit the compatibility behavior without any acceptance failure | add a criterion covering in-flight local loops, unknown starts and downstream adoption outcomes in both full and partial-update cases +MAJOR | high | story criterion 2 versus spec §5.1 | the criterion observes only per-pass finding and Blocker counts in a pinned form; it does not cover the spec's required pass numbers, per-pass models, same-artifact binding, incomplete-pass exclusion, calls-versus-passes disclosure or skipped-Gate-B line | those normative curve semantics can be dropped while the story still passes | extend criterion 2 to enumerate each required field and the skip case, including split-branch and noncontiguous-pass examples +END OF FINDINGS (33 total) diff --git a/.context/codex-reviews/gate-a-spec-rle-pass-6.md b/.context/codex-reviews/gate-a-spec-rle-pass-6.md new file mode 100644 index 0000000..5b12817 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-rle-pass-6.md @@ -0,0 +1,35 @@ +BLOCKER | high | §2.1 "The duty is qualified for, and only for, a finding the user has explicitly declined" | an accepted Minor or Nit ends the settled answer-wait hold, but the no-clean rule remains unqualified because the text grants that qualification only to declines; this contradicts the later claim that accepted Minors never block a clean pass | an accepted Minor either forces an extra iteration or leaves closure governed by contradictory rules | key the hold rules to "awaiting the user's answer"; after any answer apply ordinary severity, so accepted Minor/Nit can close, accepted Blocker/Major must be repaired and re-reviewed, and declined findings leave the set +BLOCKER | high | story criterion 6 versus §§1.2 and 2.1 | the criterion still names the Blocker/Major-resolve duty as one of three universal rules that must be qualified, while the settled design and spec explicitly say a decline never qualifies that duty and instead name the resumption rule | no implementation can satisfy the current story, the settled decision, and the spec simultaneously | rewrite criterion 6 to use the same answer-wait hold rules as the settled design and state that the resolve duty applies only after an accepted Blocker/Major enters the set +MAJOR | high | §2.1 accepted branch | "the pass that surfaced it is still not clean until it does" makes the finding-bearing pass become clean retroactively when an accepted Blocker/Major is repaired | closure can reuse a pass that predates the repair even though §5 requires re-review after every fix | state that the surfaced pass never counts as clean and that a later pass over the repaired artifact must be clean +MAJOR | high | §6.2 row 6 | the accounting says the decline "touches only the third condition", but §2.1 says it changes the surfaced-finding-open rule, the no-clean rule, and the resumption rule | the implementation plan can follow the accounting and omit two required edits, violating the old-conditions Don't | mark each old condition separately: open and no-clean are bounded by the answer-wait state, resumption is made explicit for either answer, and the resolve duty is kept unchanged for accepted in-set Blocker/Major findings +MAJOR | high | §2.1 "Where the decision is recorded" | the paragraph says the commit carrying the record does not exist while the loop runs, then says a Gate-B record lands during the loop by amending the already-existing WIP commit | implementers get opposite timing instructions and may defer a Gate-B record until close or overwrite it during the closing amend | distinguish Gate A, where the eventual commit does not yet exist, from Gate B, where the WIP commit is the mid-cycle durable carrier and the final amend must preserve the record +MAJOR | high | §2.1 "Identity" | Gate A has no reviewed-artifact commit until closure, and separate worktrees or branches can run the same loop name from the same HEAD; Gate-B cycles can likewise share a WIP parent, so the claimed identifier is neither always defined nor unique and the statement that no concurrent cycles share it is false | a decline from another cycle can be mistaken for the current cycle's decision and silently release a finding | define an explicit stable cycle identifier that includes the artifact path and a collision-resistant per-cycle token, persist it in the working record, and refuse reuse +MAJOR | high | §5.1 "Both branches ... same artifact revision" | the spec says the cycle identifier identifies the reviewed Gate-B revision, but §2.1 defines that identifier with the WIP commit's parent; the parent is the base, not the amended WIP commit or tree being reviewed | spec and quality branches from different WIP revisions can be summed into one apparently valid pass | record and compare the exact reviewed WIP commit or tree for each branch independently of the cycle identifier +MAJOR | high | §3 decision-procedure quote versus the following boundary paragraph | the proposed rule says a failed reader test makes the finding exactly Minor, while the next paragraph says the test decides only "Minor-or-below" versus keeping severity | Nit findings are promoted to Minor and two incompatible boundary procedures ship | say the failed test caps the finding at Minor, with the existing definitions choosing Minor versus Nit; passing the test preserves the independently assigned severity +MAJOR | high | §3 illustrative reader list | a "rule" and a "scaffolded template" do not read text or take decisions, despite being offered as examples of in-system readers | reviewers can satisfy the consequence test by naming an artifact rather than an executing consumer, preserving inflated severity | list only actors or procedures that consume text in operation, such as the outer gate procedure, a skill runner, the escalation procedure, or the scaffolder consuming a template +MAJOR | high | §4 "Why the text must say the floor governs both gates" | the rationale assumes the Gate-A spec loop, Gate-A plan loop, and Gate-B cycle cite the same story set, while §6.2 row 12 says the two Gate-A runs derive independently and real plans or implementation cycles may cite additional stories | a cycle can inherit another loop's lower floor and under-review a higher-risk cited set | derive the floor independently from the complete cited-story set of each loop; use the shared predicate, not an assumed shared set, as the reason both gate types follow the same rule +BLOCKER | high | story criterion 1 | the criterion still says the prompt-derived value is the hook's single floor consumed at hook lines 946 and 966, but revision 6 correctly establishes that those comparisons use the separate workspace reminder threshold | the authoritative story still encodes the rejected hook-policy coupling and cannot be met by the no-write design | remove the hook citation as the basis for the obliged floor and state that each policy loop derives its own value from its cited-story set while the hook threshold remains independent +MAJOR | high | §4.1 workspace-knob provenance | the required numeric `M` is undefined when `codex-gate.floor` exists but is empty, malformed, non-positive, unreadable, or too large for the shell comparison; the hook ignores or mishandles those states rather than treating the raw contents as a valid threshold | the commit can record a non-threshold as a threshold or omit the fact that the hook fell back to 3 | pin forms that distinguish a valid effective threshold from invalid or unreadable knob contents and record the hook's effective fallback without treating raw invalid data as `M` +MAJOR | high | §4.1 provenance forms versus story criteria 3 and 9 | the first pinned-looking form attributes the floor to one story path, the next requires a set with levels, and criterion 9 still expects `floor 3 per `; no single exact machine-extractable grammar covers multi-story, no-story, unprofiled, knob-absent, and knob-present cases | P8 cannot parse the history reliably and a compliant set-derived record can fail the story's literal criterion | define one labelled grammar with exact variants for every cited-set and knob state, then use that same grammar in both story criteria and all examples +MINOR | high | §4.1 "under revision 4 a floor file is always a user artifact" | the preceding sentence admits any agent can write the file, so its authorship is not knowable and it is not always a user artifact | readers may attribute an agent-written reminder change to the user | say the policy treats the knob as user-controlled and this design never writes it, while authorship of an existing file is unverified +NIT | high | §5 "Four shapes" | the table enumerates five shapes: absent, partial, malformed, unreadable, and stale | the diagnostic inventory's own count is false | change the count to five or merge two genuinely equivalent states +MAJOR | high | §5 malformed row | "the reader tolerates the severity token's casing and nothing else" drops the existing rule that every non-empty unrecognized severity token is normalized to MAJOR, and the stated structural checklist does not preserve the special `NO FINDINGS` body for zero | valid current files can become INCOMPLETE or clean files can be classified as malformed, silently replacing old conditions | state the complete existing acceptance procedure, including unescaped-pipe parsing, the zero-body exception, case-insensitive recognized tokens, and non-empty unknown-token normalization to MAJOR +MAJOR | high | §5 diagnostic-shape table | the text claims each shape has its own check and remedy, but absent and partial have no recovery action, unreadable combines permissions, decoding, and alleged full-disk causes without distinct checks or fixes, and malformed only conditionally suggests a rerun | the shipped prompt fails prompt-standards item 10 and can send an operator down the wrong recovery path | enumerate the actual causes within each state, give a discriminating check and bounded remedy for each, and use an admitted unknown for causes the prompt cannot distinguish +MINOR | high | §5 malformed-row recovery | "the pass may be re-runnable" when a sessionId remains omits the existing single shared recovery budget | an agent can resume or rerun repeatedly after malformed output | qualify every resume or rerun by an unspent one-attempt recovery budget and say the pass remains incomplete once that budget is spent +MAJOR | high | §5 stale row and discriminator discussion | the table treats an absent discriminator as stale, yet the widened grammar keeps bare slots valid and delete-and-confirm is presented as sufficient evidence that a freshly written bare slot belongs to the current cycle | valid current-cycle bare files are either discarded as stale or accepted under a rule the table says is invalid | treat a missing discriminator as unknown only when current-cycle provenance cannot be established; explicitly accept a bare file freshly produced after this cycle's confirmed deletion +MAJOR | high | §5 slot grammar versus §6.2 row 18 | §5 says the optional infix widens and replaces the grammar, while row 18 says it is merely a convention on top and that the original three names remain unchanged | the old-conditions accounting directs the plan to implement the opposite grammar decision | make row 18 say the grammar is widened and list the exact new forms while preserving the sequential-per-slot limitation +MAJOR | high | §5 discriminator convention | an infix is only recommended and no uniqueness rule or occupied-target refusal exists; a repeated `rle`-style infix can still collide, after which the mandatory pre-call deletion destroys the earlier file | review evidence can be lost and concurrent or resumed cycles can consume one another's slots | derive the infix from the explicit unique cycle identifier, validate its filename alphabet, and stop rather than delete when a target already belongs to another cycle +MINOR | medium | §5.1 pinned curve form | the form requires every contributing model for a multi-model logical pass but supplies no exact syntax or example for mapping two branch models to one pass | two compliant authors can produce incompatible curves and the claimed pinned greppable form is not actually pinned | add a literal example and grammar for per-pass multi-model mappings and for the optional calls-versus-passes suffix +MAJOR | high | §6.2 row 1 and §5.1 call/pass distinction | row 1 keeps "counted by the hook" verbatim even though the hook counts calls and §5.1 correctly says a logical Gate-B pass can contain two calls | three branch or recovery calls can satisfy the apparent policy floor before three validated logical passes exist | mark hook counting as moved to advisory reminder state and state that closure compares the derived floor only with validated logical passes +MAJOR | high | §6.2 row 13 | the disposition still says re-review-after-fix reasoning "generates the number", but that duty generates however many repair passes are needed and derives neither policy value 1 nor 3 | the replacement repeats the gate-mechanism overclaim the Don't forbids and can conflate an unbounded repair duty with a profile minimum | derive 1 or 3 only from the profile predicate and retain re-review after fixes as an independent duty that may exceed the floor +MAJOR | high | §7 falsified user-facing statements | the inventory omits `docs/getting-started.md:58-59`, where a hook `satisfied (3/3)` message is the trigger for the real amend, and leaves line 35's "below floor" terminology unexplained | with a user knob of 1 and an obliged floor of 3, the guide can authorize closure after the hook reports satisfied; with a level-0 cycle it can teach the redundant reminder as policy debt | add these sites and make closure depend on §5's derived logical-pass floor and clean-final-pass rule, with hook messages labelled only as reminder-threshold status +MAJOR | high | §8 closing revalidation | the staged tree plus the message "about to be written" is claimed to be exactly what amend will commit, but the index can change concurrently and prepare-commit-msg, commit-msg, or other commit hooks can change or reject the result | the actual closing commit can differ from the state re-read before the amend while the spec reports it as verified | verify the written commit after amend and reopen the cycle on any mismatch, or freeze and compare the exact index tree and final message on both sides without claiming the pre-amend check is a guarantee +MINOR | medium | §8 knob-survival and no-write checks | start/end comparison cannot distinguish a design write from a concurrent user edit to the user's own knob, so a legitimate user change makes the evidence fail even though the implementation complied | the cycle can block on an unattributable concurrent mutation | require the knob to remain stable for this verification, detect and report concurrent user mutation as inconclusive, and never attribute the writer +MAJOR | high | §10 "gate-off surface" | the text says routes beyond the four listed are pre-existing and not created here, but this design creates an author-written, unchecked decline record with real release effect; forging the claimed handle or decision is a new bypass route | an agent can release a scope-stop finding without the user answering while the risk section falsely implies the new routes are exhausted | add forged or misattributed decline records to the non-exhaustive new residuals and state that narrow scope limits blast radius but does not verify authorization +MAJOR | high | §10 activation versus story criteria 2 and 9 and §8 | §10 says this design's in-flight Gate-A loop runs wholly under the old rules, under which curves and floor provenance are not owed, while the story requires this branch's own spec and plan commits to carry those new records and §8 assumes they exist | a plan can correctly follow activation and still fail branch acceptance or leave the verification with no input | explicitly require this branch to produce the records as acceptance evidence despite their not being closure obligations for the in-flight loop, or move the demonstration to a post-activation cycle +NIT | high | story criterion 5 | it says "Four properties" and enumerates (a) through (e), five properties | the criterion's stated count is mechanically false | change "Four" to "Five" +MAJOR | high | story acceptance-criteria coverage of multi-story floors | no acceptance criterion requires the settled unanimity rule that floor 1 applies only when every cited story is profiled at level 0 | an implementation can satisfy all ten criteria while using first-story, minimum, or maximum-story selection incorrectly | add an observable criterion for complete-set derivation, including mixed and unprofiled cited sets +MAJOR | high | story acceptance-criteria coverage of decline durability | no acceptance criterion requires a decline to use the distinct commit-body record, carry all five sameness fields, bind one cycle only, or remain separate from a human exception | the implementation can satisfy Q1-Q6 while making declines session-only, reusable forever, or indistinguishable from authorizations | add an observable criterion covering record type, transport, identity fields, attribution calibration, and one-cycle effect +MINOR | high | opening "Surfaces edited" and story acceptance criteria | the opening surface list omits the manifest version bump and CHANGELOG entry later required by invariant 12, and no criterion covers those packaging edits or the three user-facing documentation corrections | a plan derived from the declared surface or a criteria-only acceptance check can omit required shipped files | include every required file class in the implementation surface and add acceptance coverage for user docs plus plugin version and changelog consistency +MINOR | high | §5 Q6 disclosure | the new mandatory degraded-sensitivity report has no literal three-line example showing computable tells, unavailable tells, shape, cause, and affected pass numbers | agents can produce incompatible reports and the changed prompt region does not meet prompt-standards item 4 | add an exact pass-report example for at least absent, partial, and malformed history while keeping it outside the findings-file protocol +END OF FINDINGS (34 total) diff --git a/.context/codex-reviews/gate-a-spec-rle-pass-7.md b/.context/codex-reviews/gate-a-spec-rle-pass-7.md new file mode 100644 index 0000000..c466d57 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-rle-pass-7.md @@ -0,0 +1,33 @@ +BLOCKER | high | §2.1 lines 123-150 | the qualification list says an answer makes the surfaced finding no longer open and makes a pass carrying it clean, while the accepted branch says pass cleanliness never changes and an accepted Blocker or Major remains unresolved until repaired | the two shipped copies would give incompatible closure and resolution instructions and cannot satisfy story criterion 6 | distinguish the decision hold from finding resolution; keep the surfacing pass permanently unclean, keep an accepted Blocker or Major open until fixed, and qualify only closure while the membership decision is pending +MAJOR | high | §6.2 row 6 | the accounting says all four old conditions are kept, says the fourth condition carries the change, then says the decline touches only the third, while §2.1 says three rules are modified | the required kept-moved-dropped audit cannot guide an implementation and can certify a rewrite that drops or alters the wrong conditions | rewrite row 6 after settling §2.1 so every old condition has one consistent disposition and points to the exact shipped sentence that changes it +BLOCKER | high | §2 clean-completion precedence and §6.2 row 7 | clean completion is declared to outrank the two-tell stop, but row 7 keeps verbatim the current rule that any two tells make stop-and-surface mandatory and identifies no qualification at that rule | a clean pass with two tells is simultaneously required to close and required to stop, so a core settled decision remains unimplementable | mark the two-tell threshold as changed and add the clean-completion exception at the mandatory-threshold sentence in both copies +MAJOR | high | §6.1 and §6.2 row 1 | the claimed fourteen-site floor inventory omits the live floor-specific sentence at CLAUDE.md:77 and workflow-init.md:277, “if pass 3 still finds Blocker/Major”, even though that trigger also changes under a floor of 1 | the implementation can leave a floor-3 decision branch in both shipped prompts while the generated inventory and evidence still report complete coverage | add both sites, raise the total to sixteen, map them to row 1, and replace the pass-3 trigger with a derived-floor formulation +BLOCKER | high | §6.2 row 1 “counted by the hook” | row 1 keeps the headline claim that the mandatory floor is counted by the hook, but the design itself says the hook counts calls against an independent reminder threshold and cannot identify valid logical passes or the derived profile floor | agents can treat a 3-call hook ratio, including incomplete calls, as proof of a 1-or-3 logical-pass obligation, recreating the false clearance the design warns about | replace the claim with an agent-maintained count of validated logical passes from findings files and state that the hook call counter is advisory and may use a different threshold +BLOCKER | high | §6.2 row 13 | the row preserves “the hook invalidates the prior pass” and replaces “where the 3 come from” with a claimed lower bound of one pass per fix, but the hook performs no edit-event invalidation and the rule requires one re-review per repair round, not per individual fix | this contradicts invariant 3’s content-derived mechanism and can overcount or over-oblige passes when several findings are fixed together | describe the exact comparison: changed content will fail the next fingerprint comparison until a re-review records it, and each batch of repairs requires a re-review independently of the floor +BLOCKER | high | §5 lines 447-460 versus §6.2 row 18 | §5 correctly says the infix replaces the slot grammar, while row 18 says it is merely a convention beside the grammar and that an infixed slot already satisfies the old names | the plan has two mutually exclusive instructions for the same protocol and can preserve the very invalid grammar the change is meant to fix | make row 18 say the three grammars are replaced by the optional-infix forms and account separately for bare-slot compatibility, collision refusal, and uniqueness +BLOCKER | high | §4.1 lines 371-375 and codex-gate.sh:933-973 | a level-0 cycle is meant to close after one pass, but the model-visible hook message still says it MUST reach the hook’s threshold and that the only early exit is zero findings; calling this merely a redundant warning supplies no action rule for the contradictory instruction | at the new floor’s main success case the agent can be commanded to run two unowed passes, so the economics outcome and prompt-standards item 7 both fail | require both §5 copies to state explicit precedence and action: the derived obligation controls closure, the hook ratio is only a reminder threshold, and a below-threshold reminder is disregarded once the derived floor and ordinary closure rules are met +MAJOR | high | §4.1 lines 332-338 | the assertion that the knob has never made an agent owe a pass is inferred solely from the hook exiting 0, even though the hook injects model-visible additionalContext containing MUST language selected by that knob | non-blocking execution is confused with lack of behavioural influence, repeating the mechanism-overclaim class governed by AGENTS.md | narrow the claim to “the hook never blocks the tool call” and separately acknowledge that its prompt can influence the agent unless §5 gives the derived rule explicit precedence +BLOCKER | high | §4.1 lines 342-369 | the provenance block is called the only machine-extractable form, then no-story, unprofiled, and unusable-knob cases receive different forms; the examples also alternate between “level L” and a bare “high”, and no mixed profiled-plus-unprofiled set form is pinned | story criterion 3 cannot be parsed reliably and valid multi-story cycles have no canonical historical record | define one grammar with explicit tokens for story-set entries, no-story, unprofiled, knob absent, knob numeric, and knob unusable, use it for every case, and provide one parser-oriented example per variant within that grammar +MAJOR | high | story criterion 1 versus §4.2 | the story requires one derived value to govern the Gate-A spec loop, Gate-A plan loop, and Gate-B cycle, while §4.2 re-derives from the current profile on every pass and expressly lets a profile change move the floor | a valid profile change makes the criterion impossible to satisfy even if the spec is implemented exactly | change the criterion and matching prose to require one derivation rule over the current cited set, while allowing the resulting value to change under the specified mid-cycle transition +MAJOR | high | §4 and §4.2 | profile-value changes are handled, but adding or removing a cited story mid-cycle is not, even though unanimity can move the floor in either direction without any profile changing | a higher-risk story can enter without forcing the current-profile clean pass, or be removed to lower the floor without a defined approval and counting rule | define cited-set changes as floor-affecting transitions with scope approval, fresh derivation, already-valid-pass treatment, and a final clean pass under the current complete set +BLOCKER | high | §2.1 lines 188-195 and §5.1 lines 515-519 | the claimed unique cycle identifier is Gate A’s loop name plus a repository commit and Gate B’s WIP parent; sibling or restarted cycles can share those values, and the WIP parent identifies the base rather than the reviewed WIP artifact | a decline can leak into a later cycle and separate Gate-B branches can be merged as one logical pass even when they reviewed different trees | define an immutable collision-resistant cycle nonce recorded at cycle start and in history, encode it safely for slots, and track the actual reviewed WIP commit or tree separately for same-revision checks +MAJOR | high | §5 lines 452-467 | collision refusal depends on knowing whether an occupied target was written by this cycle, but neither the filename nor file contents provide that fact, and delete-and-confirm is still described as a guarantee even though the spec admits it destroyed a closed cycle’s file | an agent can either delete evidence it does not own or refuse its own recoverable attempt, causing data loss and non-idempotent recovery | define a mechanically checkable ownership marker tied to the immutable cycle id and delete only a target proven to belong to the current attempt; otherwise choose a unique infix or stop without deleting +MAJOR | high | §5 lines 438-465 | the stale row treats every slot lacking a discriminator as absent, while the replacement grammar explicitly keeps bare slots valid | a cycle that legitimately used bare slots loses all historical tells after interruption even when those validated files are its own | either require a discriminator for every cycle whose history may be read later or define evidence by which the outer agent can attribute a bare slot to the current cycle +MINOR | high | §5 lines 429-438 | the text promises four diagnostic shapes but enumerates absent, partial, malformed, unreadable, and stale | the count claim is mechanically false and weakens confidence in the item-10 audit | say five shapes or merge two shapes and update the table accordingly +MAJOR | high | §5 absent recovery row | “No recovery” contradicts the standing single-branch resume path, which can recover an absent file when the call ran and its sessionId is still available | a recoverable write failure is misclassified as permanently lost, causing an unnecessary discarded pass and extra review | split historical absence with no session from a fresh missing-write result with a resumable session, preserving the one-attempt shared budget +MAJOR | high | §5 partial and stale recovery rows | partial and stale have treatments but no per-cause recovery actions despite the section promising each shape its own check and remedy | the shipped diagnostic leaves the agent unable to choose between preserving valid present passes, resuming a missing branch, selecting a new slot, or stopping on collision | add explicit recovery for each shape and state which actions consume the shared recovery attempt +MAJOR | high | §5.1 lines 507-539 | the curve requires one numeric entry for every valid pass, but Q6 permits prior valid pass files to be absent, unreadable, malformed, or unattributable and defines no representation or closure rule for that case | a cycle can be allowed to close while being unable to produce the mandatory pinned curve, so story criterion 2 is unsatisfiable on the designed recovery path | define a pinned unknown-with-cause entry for unavailable valid-pass counts or make missing curve inputs a named closure stop; align that choice with Q6’s continue bias +BLOCKER | high | §2.1 lines 155-166 | a Gate-B decline is to be amended into the WIP immediately, but the design does not require the amend command itself to carry a WIP-prefixed message; the hook recognizes WIP only from the command’s -m argument and otherwise treats amend as cycle-closing | recording the settled decline can reset Gate-B counters and close the cycle before ordinary rules are satisfied | pin the mid-cycle amend form so its command visibly retains the WIP prefix, preserve all existing body records, and reserve the non-WIP amend for final closure +MAJOR | medium | §2.1 decline record form | consequence and fix share one line with unlengthed textual delimiters while both are copied verbatim, so a field containing the delimiter text can make the five-field identity record parse more than one way | sameness can be judged against the wrong bytes and silently release or re-hold a finding | put each identity field on its own labelled line and define escaping for label-like content, then make the sameness test consume that exact canonical form +MAJOR | high | §2.1 lines 155-217 | during Gate A the only pre-commit copy of a decline is the dispositions companion, yet the spec explicitly preserves that file’s advisory, optional, deletable status | an interruption or cleanup can erase the user’s decision and the resumed cycle cannot know whether the hold was released | make the decline entry mandatory working state until the Gate-A commit is written, include it in resume recovery, and allow safe fallback to re-ask only when that state is genuinely unavailable +MAJOR | high | §5.1 squash carry | the carry rule adds curves, provenance, and decline records but omits the “Gate B: skipped” record and its required skip reason | after squash, main history can show neither a curve nor the designed explanation for a legitimate skipped cycle, defeating story criterion 2 and P8 | add the labelled skip record and its reason to the explicit squash-carry set +MAJOR | high | story acceptance criteria | no criterion observes the required user-facing documentation corrections, plugin version plus changelog work, or the settled per-pass floor report; all ten criteria can pass while those spec obligations are omitted | spec-to-story traceability fails in the spec-to-story direction and invariant-12-adjacent release work can be incomplete without failing the story | add checkable criteria for the falsified docs, manifest-version and changelog parity, and a floor statement in every pass report +MAJOR | high | §4.1 and §8 | every pass report must state the floor, but no pass-report form is shown and the evidence plan validates only commit-body provenance | the settled observable can drift across passes while all planned checks remain green | pin the pass-report form, including no-story, unprofiled, and unusable-knob cases, and include each actual pass report in the provenance verification +MINOR | high | §1 surfaces and §8 parity baseline | the spec says the two §5 regions differ on 192 lines across 20 hunks, but the cited current ranges produce 90 additions plus 96 deletions across 20 hunks, or 186 changed lines | the parity baseline and out-of-scope claim are stale, so later verification can misreport an unexplained six-line delta | correct the count to 186 and state the counting method +MINOR | high | story-criterion citations at lines 25, 473, 549, 681, and 690 | the current ten-criterion story places parity at 9, Q1-Q6 at 6, old-condition accounting at 8, and the branch provenance demonstration at 10, but the spec cites 7, 5, 6, 9, and 7 respectively | reviewers and the plan are routed to the wrong acceptance obligations despite the request to read the current story fresh | update those references to 9, 6, 8, 10, and 9 and recheck every remaining criterion citation +MAJOR | high | §8 lines 673-689 | the endpoint comparison explicitly cannot observe a floor file created and removed during a cycle, yet the next bullet says the same risk-path verification checks “no floor file is written on any path” unconditionally | the evidence overclaims the behaviour it proves, violating prompt-standards item 11 and the gate-proof Don’t | describe the verification only as an endpoint persistence check and leave the stronger no-write statement as an instruction-backed design rule with no comprehensive detector +MAJOR | high | §8 knob-preservation verification | byte identity requires a captured start value, but the plan defines no baseline capture or durable recovery, and it also includes unreadable files among unusable knob values even though their bytes cannot be compared | the verification can be asserted from memory or becomes impossible on interruption and unreadable-knob paths | record a start checksum and readability state in non-knob working evidence, compare it at close, and report an explicit unverified result when bytes were never readable +MAJOR | high | §8 lines 708-720 | the text says a changed close invalidates the clean pass, but Gate B reviews a git range and does not review the final commit message; the curve and closing provenance are necessarily written after the final pass | re-running Gate B cannot establish the commit-body claims and can create a circular curve whose final entry changes on every rerun | separate diff review from commit-body verification: rerun Gate B only for reviewed content changes, rerun the named verifier for message changes, and state that neither proves the resulting commit after the admitted race +MAJOR | high | §10 lines 777-789 | the threat analysis says all four residual routes are statement attacks living in the commit body, but profile minting edits a story file and cited-set omission can occur in the artifact or gate prompt before any commit body exists | the proposed revalidation does not cover every surface the threat claim assigns to it, so a gate-off route can bypass the named check | classify each route by its actual writable surfaces and make the risk verification inspect story headers, artifact citations, gate inputs, and commit provenance as applicable +MAJOR | high | §6.2 row 14 versus §3 | row 14 calls the reachability test “the classifier”, while §3 says it is only a Minor-or-below ceiling and that the existing definitions still choose Blocker versus Major | an implementer can replace the severity procedure with the reachability test and lose the settled ceiling-not-floor semantics | call it a preliminary ceiling test in row 14, then preserve the existing consequence definitions as the second stage of one composite procedure +END OF FINDINGS (32 total) diff --git a/.context/codex-reviews/gate-a-spec-rle-pass-8.md b/.context/codex-reviews/gate-a-spec-rle-pass-8.md new file mode 100644 index 0000000..d0d95df --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-rle-pass-8.md @@ -0,0 +1,34 @@ +BLOCKER | high | §1.2 lines 57-68 and §5 lines 481-541 | closure may proceed after prior findings files become unavailable by merely naming the unreadable passes, even though the spec says every in-set Blocker and Major must be inventoried from those files and separately proven resolved | a cycle can close with an earlier mandatory finding unresolved, violating the settled resolve duty and the loose firing direction | make inability to reconstruct the Blocker/Major inventory a closure hold, or add another durable complete inventory whose recovery is defined; keep Q6's continue bias limited to the two-tell statistics +BLOCKER | high | §6.2 passage list | the nineteen-passage list omits existing decision procedures this design explicitly rewrites, including “Before each call, delete every target file” at CLAUDE.md:199 and template:386 and the WIP creation, amend, closing-body and evidence rules at CLAUDE.md:500-517 and template:679-696 | the deferred dispositions artifact can cover every listed row while silently dropping old deletion, collision, amend and body-preservation conditions, so the split does not satisfy the old-conditions Don't | add every actually rewritten source passage to the frozen inventory before producing dispositions, including the target-deletion and Mechanics commit-body paragraphs +MAJOR | high | §6.2 lines 649-666 | the future conditions artifact is only said to be “reviewed”; the spec does not name who reviews it, the acceptance protocol, how findings are resolved, what prevents the plan from crossing the gate, or what happens when review changes the supposedly frozen spec or the old §5 text | the artifact can be present yet unapproved, stale, or self-contradictory while implementation proceeds, making the deferral gate non-operational | pin the reviewer and file protocol, require a clean review before replacement prose, record the exact source revisions of both §5 copies, and define re-freeze, regeneration and re-review whenever either input changes +MAJOR | high | §6.1 lines 616-638 | the claimed fourteen-site inventory misses CLAUDE.md:77 and template:277, where “if pass 3 still finds Blocker/Major” hard-codes the old floor; the permitted grep misses these because its patterns do not match “pass 3” | both shipped copies can retain a pass-3 trigger under floor 1 while the inventory, row mapping and parity verification report complete coverage | add the two sites, raise the total to sixteen, map them to the floor paragraph, and replace the trigger with a derived-floor formulation +BLOCKER | high | §4.1 lines 364-412 | the promised single provenance grammar is contradicted by required forms outside it: lines 405-412 prescribe a colon-based unusable form, `floor 3 (no story cited)`, and an unbraced unprofiled form, while line 399 gives `B (high)`, which the ENTRY production rejects | story criterion 3's machine-extractable history has several incompatible encodings and some prescribed examples cannot parse | delete every alternate form and express unusable, no-story, unprofiled and mixed sets only through the grammar and valid examples +MAJOR | high | §4.1 STORY-SET grammar | `` has no lexical boundary or escaping rule even though commas, braces, leading spaces and strings resembling ` (level 0)` are valid path content, and set ordering is not defined despite the claim of a canonical form | two different cited sets can serialize ambiguously and the P8 reader cannot parse all valid repositories or compare equivalent sets reliably | define a length-prefixed or escaped path token and a deterministic ordering for entries, then include adversarial-path examples +MAJOR | medium | §4.1 KNOB grammar versus codex-gate.sh:120-127 | `` treats every positive numeral as usable, but the hook accepts only values its shell arithmetic comparison can represent; an overflow-sized digit string falls back to 3 | provenance can report a numeric reminder threshold the hook did not adopt | define usability by the hook's actual parse and arithmetic result, or bound KNOB to the portable numeric range and record out-of-range values as unusable +MAJOR | high | §4.1 precedence text at lines 414-424 | the proposed shipped sentence says the hook ratio “controls nothing”, while the same section correctly establishes that it controls which model-visible advisory message fires | the two copies would contain an internal mechanism contradiction and repeat the gate-overclaim class | say the hook threshold controls reminder selection but never controls the cycle's closure obligation +MINOR | high | §4.1 residual at lines 426-437 | the residual is attributed to codex-gate.sh:933 as though one message carried it, but the conflicting hard-minimum and early-exit instructions also occur at :947 and :967; the categorical claim that every level-0 one-pass close draws a 1/3 warning is also false when the user knob is 1 or reminders are off | the shipped disclosure and verification can inspect the wrong message set and overstate when the residual appears | name all three conflicting message sites and qualify the example to an enabled hook with threshold 3 +MAJOR | high | §4 and §4.2 versus story criterion 1 | the story says one derived value governs all three loops, while §4.2 permits that value to change per pass, and neither artifact defines what happens when a story citation is added or removed mid-cycle even though unanimity can change the floor without a profile edit | Gate-A spec, Gate-A plan and Gate B can derive different sets or let a newly added high-risk story avoid a clean pass under the current obligations | define one authoritative current cited-story set for the whole cycle, make citation-set changes human-approved floor transitions with a new clean pass, and rewrite the story criterion to require one derivation rule rather than an immutable value +MAJOR | high | §2.1 lines 146-188 versus story criterion 5(f) | every answer is said to end the hold and all branches are called symmetric, but the only durable form is `Declined finding`; no accepted-decision record exists | an accepted Minor or Nit produces no repair and no durable fact showing the hold ended, so a resumed cycle can either re-hold it or close on an answer it cannot establish | define and carry a single decision record with accepted or declined verdict, or separately pin the accepted form and its recovery semantics +BLOCKER | high | §2.1 lines 160-180 | the pinned mid-cycle command `git commit --amend -m "WIP: "` replaces the whole commit message with that one subject while the adjacent rule requires every nonce, decline and other body record to be carried forward | recording one decline can erase earlier decisions and cycle identity, defeating durability and making recovery choose the wrong cycle | pin a command form that visibly retains WIP and supplies every preserved body paragraph, such as repeated message paragraphs generated from the existing body, and verify the resulting message before the amend +MAJOR | high | §2.1 cycle nonce at lines 201-217 | the nonce has no standalone commit-body production, active-nonce marker, generation procedure, or Gate-A pre-commit recovery source; “recover from recorded history” cannot work for an interrupted Gate-A loop whose first commit is created only at closure | resumed cycles can mint a new identity unnecessarily or select among multiple historical nonces without a deterministic answer, losing valid passes and hold decisions | pin a `Cycle nonce:` record, define collision-resistant generation, identify the active nonce, and define mandatory working-state recovery for Gate A before falling back to a new cycle +MAJOR | high | §2.1 unrecoverable-nonce transition | “starts a new cycle” does not say how old pass slots, the Gate-B WIP body, advisory hook counters, decline records and the previous cycle's incomplete curve are segregated or retired | the new cycle can inherit reminders and records from the abandoned one or overwrite its evidence, producing false provenance and non-idempotent resume behavior | define an atomic logical restart procedure that preserves old artifacts, records abandonment, selects a fresh nonce and excludes all old counts and decisions from the new curve +MAJOR | high | §2.1 lines 209-211 versus §5.1 lines 577-581 | §2.1 separates the cycle nonce from the reviewed commit or tree, but §5.1 again says the reviewed artifact is “the commit the cycle identifier names”; a body-only WIP amend also changes the commit id without changing the reviewed tree | two branches can be split or combined using the wrong identity, and nonce semantics are conflated with artifact equality immediately after being separated | give every logical pass an explicit reviewed-artifact field, choose commit id or tree id with message-only-amend semantics stated, and never derive it from the cycle nonce +MAJOR | high | §5 stale and slot-grammar rules at lines 500-527 | the table defines an absent discriminator as stale, the prose later says absent does not mean stale, bare slots remain valid, and the optional-infix grammar does not state whether new cycles must use their nonce | a valid bare-slot cycle can lose all history after interruption, while two readers classify the same file differently | separate foreign, unattributable and legacy-bare states, require nonce-infixed slots for every new cycle, and reserve bare slots only for explicitly defined legacy handling +MAJOR | high | §5 lines 514-522 | collision refusal depends on knowing whether an occupied file “the cycle did not write”, but neither the filename nor file contents identify the writer or invocation; a late concurrent writer using the same cycle nonce is indistinguishable from the current attempt | recovery can delete foreign evidence or trust a correctly shaped file from the wrong call | add an invocation identity or ownership record that is checked before deletion or reuse, and stop without mutation when ownership is not provable +MAJOR | high | §5 Q6 absent row versus existing §5 recovery | the absent row says there is no recovery, while the standing shared recovery path can resume or rerun a call that ran but failed to write its slot when its session id remains available; the malformed row acknowledges exactly that possibility | a recoverable review is discarded and an extra expensive pass is forced solely because the missing-write symptom was put in the wrong row | split historical loss with no session from a fresh missing-write result and preserve the one-attempt resume or rerun path for the latter +MAJOR | high | §5 Q6 table | partial and stale entries give reporting treatments but no cause-specific fix despite promising each shape its own check and remedy, and unreadable recovery does not define what happens if the environment cannot be repaired within the cycle | agents cannot choose among preserving valid branches, resuming a missing branch, selecting a fresh slot, abandoning history or stopping, violating prompt-standards item 10 | add a bounded remedy and terminal action for every diagnostic shape and tie any retry to the existing shared recovery budget +MINOR | high | §5 lines 491-500 | the section promises four diagnostic shapes but enumerates absent, partial, malformed, unreadable and stale | the mechanically false count undermines the diagnostic audit | say five shapes, or merge two and update every later reference +MAJOR | high | §5.1 lines 543-592 | the curve is called pinned and greppable but only examples are supplied; there is no grammar for noncontiguous passes, per-pass multiple models, mixed call and logical-pass counts, incomplete-pass annotations, or the skip form | two compliant authors can emit incompatible lines and P8 cannot reliably consume the record story criterion 2 requires | define one formal curve grammar covering every listed production and give a parsing example for each variant +MAJOR | high | §5 and §5.1 unavailable-history path | Q6 allows previously accepted pass files to become absent, unreadable or unattributable, while the closing curve requires exact counts for every valid pass and defines neither an unknown entry nor a closure stop | a cycle can be allowed to continue but cannot satisfy its mandatory closing record, or it can invent counts from memory | persist accepted counts in cycle working state, define an explicit unknown-with-cause production, or hold closure until the mandatory curve can be reconstructed +MAJOR | high | §5.1 squash carry | the carry set adds declines, floor provenance and curves but omits the required `Gate B: skipped` record and its skip reason | a squash can erase the only explanation for a legitimately absent Gate-B curve, leaving main history indistinguishable from an omitted gate | add the labelled skip record and its reason to the explicit squash-carry rule in both copies +MAJOR | high | §8 risk-path verification at lines 757-778 | the spec correctly admits an endpoint absence check cannot detect a transient floor-file write, then claims the same verification checks the stronger “no floor file is written on any path” rule unconditionally | the evidence can be reported green for behavior it never observed, violating prompt-standards item 11 and the gate-proof Don't | limit the evidence claim to endpoint persistence and knob byte preservation, and label the universal no-write rule instruction-backed with no comprehensive detector +MAJOR | high | §8 lines 797-809 | the final provenance and curve live in a commit message written after the clean Gate-B pass, but the remedy for a changed proposed message says to re-review; Gate B reviews a git range, not that message, and rerunning it cannot validate commit-body claims | the process can enter a circular re-review loop or falsely imply Codex covered data outside its review object | route message changes through the named commit-body verifier, reserve Gate-B reruns for reviewed content changes, and state that the final commit still has the admitted read-to-write race +MAJOR | high | §10 gate-off analysis | the spec says all four residual routes are statement attacks living in the commit body, but profile minting edits a story file and cited-set omission can occur in the spec, plan or gate input before any commit body exists | the risk verification can inspect only provenance and miss the actual writable surface of a floor-lowering attack | map each route to its real surfaces and verify story headers, artifact citations, gate inputs and closing provenance as applicable +MAJOR | high | story criterion 4 versus §8 | the story requires both shipped copies to say nothing checks the stated floor against profiles, while §8 explicitly mandates a verification that recomputes every provenance floor from those profiles | the design simultaneously requires a check and requires the product to say no check exists, creating another enforcement overclaim | distinguish “no mechanical guard” from the author-run named verification and use that precise statement in the story, spec and both prompt copies +MAJOR | high | story acceptance criteria versus §§4.1, 7 and 8 | no story criterion observes the required per-pass floor statement, the enumerated user-facing documentation corrections, or the plugin version and changelog work; all current criteria can pass while those spec obligations are omitted | spec-to-story traceability fails in the spec-to-story direction and release-critical work can disappear from the plan without failing the story | add checkable criteria for the pass-report form, falsified docs, manifest version change and matching changelog entry +MINOR | high | story §4 “Affected AGENTS.md invariants” | the story omits invariants 9 and 10 even though its rollout depends on non-overwriting adoption and its severity rationale explicitly depends on stack-neutrality; the spec and review brief both name them | reviewers following the story's invariant list can skip two governing constraints | add invariants 9 and 10 with the concrete surfaces they govern +MINOR | high | spec story-criterion references | the current ten-criterion story places Q6 at 6, old-condition accounting at 8, parity at 9 and the branch provenance demonstration at 10, but the spec cites criteria 5, 6, 7 and 9 for those obligations | plan traceability points reviewers to the wrong acceptance requirements despite the instruction to read the current story fresh | update §5 line 535 to 6, §6 line 611 to 8, both parity references to 9, and §8 line 770 to 10 +MINOR | high | opening surface statement at lines 20-26 | the current cited §5 ranges differ by 90 additions plus 96 deletions across 20 hunks, or 186 changed lines, not 192 | the parity baseline and out-of-scope delta are mechanically stale | change 192 to 186 and state that the count is additions plus deletions from a zero-context diff of the cited ranges +MAJOR | medium | §7 and §10 rollback lens | activation and partial adoption are discussed, but no rollback procedure exists for a bad rule after it has been accepted into a downstream CLAUDE.md, which invariant 9 prevents the plugin from silently restoring | a prompt regression can leave downstream loops stuck or mis-calibrated even after the plugin is reverted, with no defined recovery or preservation of in-flight records | define rollback as an offered reverse diff through workflow-init or a documented manual procedure, state how in-flight cycles finish, and preserve user content and cycle records +MAJOR | high | §4.1 and §8 pass-report observability | the spec requires the derived floor in every pass report, but only the commit-body grammar is tied to verification and no evidence step reads actual pass reports or handles an unavailable pass report | the settled “derived and stated” duty can fail throughout the loop while every planned check and closing record passes | state that the same grammar is mandatory in pass reports, preserve those reports in cycle working evidence, and include them in the provenance verification or explicitly report the verification gap +END OF FINDINGS (33 total) diff --git a/.context/codex-reviews/gate-a-spec-rle-pass-9.md b/.context/codex-reviews/gate-a-spec-rle-pass-9.md new file mode 100644 index 0000000..8a7c8df --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-rle-pass-9.md @@ -0,0 +1,29 @@ +BLOCKER | high | §3 lines 143-145 | revision 9 required both an in-system reader and a changed decision, but revision 10 says “If you can name neither”, which demotes only when both are absent and lets a finding retain severity when just one exists | the central consequence test no longer implements the settled severity boundary and can keep non-behavioural findings iterating | replace it with “If you cannot name both an in-system reader and the changed decision” and retain the two required outputs explicitly +BLOCKER | high | §4.1 lines 231-241 versus story criterion “Every cycle's floor is reconstructible” | revision 10 moves the provenance grammar to the plan even though the story expressly requires one pinned form and says “the spec states which”; machine-extractable, one form and listed fields admit incompatible serializations and do not define escaping or parsing | the reviewed spec cannot satisfy its story and the plan is left to decide the durable interface P8 consumes | restore the complete provenance grammar to the spec or amend the story through the decision process before deferring it +MAJOR | high | §5.1 lines 294-307 and §11 lines 528-533 | the curve's concrete form moved to the plan while the remaining properties never require one pinned greppable or machine-parseable serialization | a plan can satisfy every listed property with incompatible free-form lines, leaving P8 unable to consume the all-three-loop record required by the story | retain a normative property requiring one pinned greppable grammar with defined field boundaries and escaping, or keep the grammar in the spec +MAJOR | high | §2.1 lines 124-129 | revision 9's collision-resistant nonce and exact `[a-z0-9]{4,16}` constraint became only “restricted so it is safe as a slot infix”, which cannot be checked without the removed lexical and collision rules | unsafe characters or a colliding identifier can alias slots and let a decline escape its intended cycle | restore a checkable alphabet, length, collision requirement and generation property in the spec +MAJOR | high | §2.1 lines 108-122 | the required properties no longer say the decline is a record type distinct from the human-exception form, and revision 9's mandatory reason and pass/slot/line provenance disappeared without being listed as moved or deliberately dropped | the plan can encode a decline as a human exception or omit why and where it arose, contradicting the settled distinction and weakening later re-identification | require a separately labelled decline record and disposition reason, state whether source-pass provenance is kept or deliberately dropped, then let the plan choose only syntax +MAJOR | high | §2.1 lines 108-129 | “recorded in history” cannot recover the nonce of an interrupted Gate-A loop because its spec or plan commit does not exist until closure, while the unnamed “advisory working record” has no required path, durability or resume lookup | an interruption forces a new cycle despite recoverable state, losing valid passes and decisions and making restart behaviour non-idempotent | name the pre-commit nonce record for Gate A, require it to survive resume, and define deterministic recovery before starting a new cycle +MAJOR | high | §5 lines 262-264 | revision 9 named instrument clustering and prose clustering as the two tells computable from the current pass, but revision 10 retains only the count “two” | a plan may choose either history-dependent tell and still claim the property, changing whether the mandatory two-tell stop remains reachable | name the two current-pass tells in the contract +MAJOR | high | §10 lines 492-497 | “the stricter reading of every part this change touches” replaced revision 9's explicit fallback mapping for severity, suspensions, decline records and curve duty | stricter has no unique meaning for these different rules, so a loop with unknown starting rules can under-review, release a hold or omit a record while claiming the fallback | restore the per-part fallback table: floor 3, no severity demotion, suspensions binding, no decline release without a record, and curve owed +MAJOR | high | §6 lines 324-345 | the deferred conditions artifact is merely “reviewed” before replacement, with no reviewer, clean acceptance condition, frozen-source identifiers, or refreeze rule if review changes the spec | method plus a passage list does not operationally satisfy the old-conditions Don't because the plan still decides what review is sufficient and can proceed on stale or unresolved dispositions | define the conditions gate's reviewer, clean result, exact source revisions and regeneration/re-review trigger while keeping the dispositions in the separate frozen artifact +MAJOR | high | §4.2 lines 243-253 | only profile-value changes are defined; adding, removing or correcting a cited story mid-cycle is unhandled even though unanimity derives the floor from the whole cited set | a newly cited high-risk or unprofiled story can arrive after passes were banked without a required clean pass under the new set | treat cited-set changes as floor transitions, derive from the current complete set and require a subsequent clean pass under that set +MAJOR | high | §5 lines 266-272 | “absent” is declared unrecoverable, but existing §5 permits the single shared recovery attempt when a review ran and failed to write its file and its sessionId remains available | a fresh missing-write result is misclassified as historical loss, discarding a recoverable pass and contradicting the standing recovery procedure | split historical absence from a current missing write and preserve session resume or rerun for the latter within the existing one-attempt budget +MAJOR | high | §5 lines 274-277 and §6.2 lines 393-397 | the per-cycle infix is required only after a bare slot is already occupied, so two concurrent new cycles can both observe an empty bare slot and choose it; ownership is also not provable for a late writer using the same cycle identity | one cycle can overwrite or accept another call's well-formed findings, losing evidence with no structural signal | require nonce-infixed slots for every new cycle, reserve bare slots for explicit legacy handling and add invocation ownership before deletion or reuse +MAJOR | high | §5.1 lines 294-316 | squash carry names decline records, provenance and curves but does not require carrying a legitimately skipped Gate-B record and its separately required skip reason | a squash can erase the only explanation for a missing Gate-B curve and make a valid skip indistinguishable from an omitted gate | add both the labelled skip record and its reason to the explicit squash-carry set in both copies +MAJOR | high | story criterion “The gate-off residual” versus spec §8 lines 451-456 | the story requires both shipped copies to say “nothing checks” the stated floor against profiles, while the evidence plan requires a named verification that recomputes every provenance floor from those profiles | the change simultaneously claims no check exists and relies on one, repeating the enforcement-overclaim class | distinguish no mechanical guard from the author-run named verification in the story, spec and shipped wording +MAJOR | high | story final provenance criterion versus spec §§4, 8 and P8 story | the story makes the first eligible floor-1 cycle a P8 checkpoint, but this spec never requires that checkpoint and the named P8 story contains no floor-1 provenance obligation | the criterion is delegated to a vehicle that cannot satisfy it, so the only promised floor-1 demonstration can disappear | add the checkpoint and its observation to this spec and the P8 story, or keep the demonstration in a vehicle already bound by this change +MAJOR | high | §7 lines 408-422 versus story acceptance criteria | the six user-facing documentation corrections are spec obligations but no acceptance criterion observes them | every story criterion can pass while README, getting-started and coding-workflow continue teaching the old floor and knob semantics | add a criterion enumerating the falsified user-facing statements and requiring their corrected post-change meaning +MAJOR | high | §7 lines 433-434 versus story acceptance criteria | the plugin version bump and changelog entry are implementation obligations but no acceptance criterion covers either; the affected-invariants list mentions only the bump | the story can be marked complete with an unreleased template change or without release notes even though the spec requires both | add a criterion for a changed manifest version and matching changelog entry +MAJOR | high | §4.1 lines 204-205 and §8 | every pass report must state the derived floor, but no story criterion or evidence step observes pass reports; only closing commit provenance is checked | the settled derived-and-stated duty can fail throughout all three loops while every acceptance criterion and closing verification passes | add a pass-report criterion and preserve or inspect the reports, or explicitly narrow the obligation to the durable closing record +MAJOR | high | §10 versus story activation criterion | the story requires activation, unknown-start fallback and partial-adoption consequences in both shipped copies, but the parity evidence scopes inserted rules to §§2-5 plus only §10's residual disclosure | a plan can omit the rest of §10 from both copies and still pass the stated parity verification | require every §10 activation and partial-adoption rule in both copies and include them in the itemized parity check +MAJOR | medium | §10 “Risks and activation” | activation and partial adoption are described, but no rollback path exists after a bad rule has been accepted into a downstream CLAUDE.md that the scaffolder may not overwrite | reverting the plugin leaves downstream projects running the bad prompt indefinitely or forces unsafe manual replacement during an active cycle | define an offered reverse diff or documented manual rollback, preserve user content and state how in-flight cycles finish +MAJOR | high | §2.1 lines 120-122 | the spec says narrowness “makes” an author-written, unauthenticated decline record safe, but narrowness only limits blast radius and does not establish that the user decided anything | one forged record can release the exact finding that holds closure while readers are told the mechanism is safe | say narrowness bounds consequence rather than supplies safety, and state the accepted authenticity risk or require confirmation on recovery +MAJOR | high | §§2.1, 4.1 and 5.1 | revision 9 required the nonce in every record carried thereafter, but revision 10 only says it is recorded somewhere in history; provenance and curve properties have no shared cycle key | after squash, multiple Gate-A or Gate-B cycles cannot be reliably joined to their floors, cited sets and curves, so P8 cannot measure pass economics per cycle | require the same cycle nonce and loop label in provenance, curve, decline and skip records +MINOR | high | story §4 “Affected AGENTS.md invariants” | invariants 9 and 10 are absent even though rollout relies on non-overwriting scaffolding and the severity reader list relies on stack neutrality; the review brief explicitly names both as touched | reviewers following the story's own invariant list can skip two governing constraints | add invariants 9 and 10 with the surfaces they govern +MINOR | high | spec lines 322 and 460 | the ten-criterion story places old-condition accounting at criterion 8 and parity at criterion 9, but the spec cites criteria 6 and 7 | plan traceability sends reviewers to the Q1-Q6 and activation criteria instead of the obligations being claimed | update the references to criteria 8 and 9 or number the story criteria explicitly +NIT | high | story severity criterion line 173 | it announces “Four properties” and then enumerates six labelled items, five severity properties plus one record property | the mechanically false count weakens criterion auditing and invites a reader to stop before the later requirements | say six properties or separate the five severity properties from the record requirement +MINOR | high | §4.1 lines 218-220 | the named one-pass residual cites only codex-gate.sh:933, but the enabled threshold-3 success path actually emits the conflicting hard-minimum instructions at lines 947 and 967; line 933 is the no-fingerprint branch | verification can inspect the wrong hook message and miss the contradictions a legitimate level-0 close encounters | cite lines 947 and 967 as well, and describe line 933 separately +MINOR | high | opening lines 16-20 | gate_citation is also an adoption predicate, so losing the heading can silence the hook entirely when codex-gate.on is absent rather than merely degrade a citation; several diagnostic reminders do not print the policy citation at all | the spec understates the compatibility risk and overstates how widely the citation is used | say the pattern controls heading-based adoption and citations on policy-bearing gate reminders, with generic fallback only when another adoption marker keeps the hook active +MINOR | high | §1 lines 36-38 | revision 9's separate setup gap disappeared from the purported list of §5's other mandatory procedure stops, leaving only the blocking evidence gap | the scope boundary can be read as exhaustive and a setup gap may be mistakenly treated as one of the reordered loop exits | restore “blocking evidence or setup gap” or state that the examples are non-exhaustive +END OF FINDINGS (28 total) diff --git a/.context/codex-reviews/gate-a-spec-vision-pass-1-dispositions.md b/.context/codex-reviews/gate-a-spec-vision-pass-1-dispositions.md new file mode 100644 index 0000000..e647318 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-vision-pass-1-dispositions.md @@ -0,0 +1,29 @@ +# Gate A (spec) — dark-factory vision — pass 1 dispositions +52 findings: 2 BLOCKER, 39 MAJOR, 10 MINOR, 1 NIT. Unprofiled run (the artifact cites no story). + +## Fixed in this pass (commit "docs(vision): Gate A pass 1") +- F4 MAJOR pipeline ended at "Merge/Deploy" vs non-goal 5 → terminus is Merge; deploy named as outside. +- F5 MAJOR §9 "no human involved" vs decision 5 → sequenced: automatic landing from step 6, human merge until then. +- F24 MINOR Sample-Gate→Audit drawn unconditional → drawn / not-drawn branches added. +- F28 MAJOR §5 dim 2 claimed the profile drives the pass floor → shipped CLAUDE.md has a fixed 3-pass floor; corrected to lenses+evidence today, floor after step 1. +- F29 MINOR "one question back" misstated the intake skill → "one question round". +- F30 MINOR "5–7 are new" vs 8 listed dimensions → "5–8". +- F31 MINOR gap 2 omitted the Wave dimension → added. +- F36 MAJOR stage labels incoherent (step 5 "Stage 3" but carried event loops) → stage 4 starts at step 4 with the Intake-Loop per decision 6; step 5 is clock loops only; step 6 owns the freigegeben-wake. +- F37 MINOR E2E clock loop absent from the canonical pipeline → added as an asynchronous loop with its edge back to the pool. +- F38 MAJOR "main is never red" — gate overclaim, the class AGENTS.md forbids by name → replaced with the bounded claim (no candidate lands whose smoke fails; E2E can still find main-level failures). +- F46 MAJOR cited gap-sweep files sit under a git-ignored docs/research/ and are in no commit → citation now says they are local notes and the durable record is the list itself. +- F47 MINOR "non-goal 4" does not resolve (§8 bullets are unnumbered) → cited by title. +- F48 MINOR "non-goal 1" ditto → cited by title. +- F49 MAJOR "a human override" unconstrained vs CLAUDE.md §5 ("supplies no permission") → constrained to human-owned knobs. +- F52 NIT "Every node is a loop with its own fresh context" false for stores, humans and the queue → narrowed to model-operated nodes. + +## Collected, not iterated (Minor) +- F23 MINOR undefined `N` in the decision-queue throttle. Real, but naming a knob and its range is step-4/5 mechanism, not vision text. +- F50 MINOR Phase 0 has no empty-pool outcome. Belongs to the Phase 0 / workflow-init story. +- F51 MINOR "prior wave is fully merged" undefined for rejected/split/deferred items. Belongs to the wave-control story. + +## Routed to the human (loop paused, findings open) +Decision-touching: F1, F2, F3, F9, F10, F11, F19, F21, F22, F27. +Unowned mechanism, proposed for §11/§7 rather than specification here: +F6, F7, F8, F12, F13, F14, F15, F16, F17, F18, F20, F25, F26, F32, F33, F34, F35, F39, F40, F41, F42, F43, F44, F45. diff --git a/.context/codex-reviews/gate-a-spec-vision-pass-1.md b/.context/codex-reviews/gate-a-spec-vision-pass-1.md new file mode 100644 index 0000000..10d5497 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-vision-pass-1.md @@ -0,0 +1,53 @@ +BLOCKER | high | §10 "Second sweep", lines 342-347 | the proposed first blocking hook directly contradicts Key Invariant 1, which requires the plugin hook always to exit 0; depending only on local state does not create an exception to that invariant | a step-4 or step-5 story following this vision would ship a mechanism the repository explicitly forbids and could make the workflow unusable when its classifier is wrong | keep plugin hooks advisory and put enforcement in a repo-owned command or required CI check, or make an explicit prerequisite architecture meta-story resolve the invariant before this item can enter a build step +BLOCKER | high | §2 decision 8, lines 108-112, and §10 "Deliberately not adopted", lines 325-328 | the vision says a same-model-family fresh agent is the shipped tier-two fallback, but the shipped workflow says "When no independent reviewer is available, be gateless — not self-reviewed", and §10 itself says same-model review is not adopted and the adversarial gate stays cross-model | later stories would treat a forbidden same-family pass as satisfying a mandatory gate, defeating the product's cross-model independence guarantee | replace the shipped-fallback claim with the actual rule: use another model family through an operational bridge or stop until an independent reviewer is available; if a future same-family tier remains desired, name it as an unshipped rule change with its own prerequisite story +MAJOR | high | §2 decision 5, lines 77-79 | "the adversarial gate checks every merge" and "no merge skips both gates" do not account for the shipped Gate-B triviality skip, so a trivial change not drawn by the Sample-Gate can skip both Gate B and human audit | the end-state safety claim is false on an already-supported path and automated merging could land a change with neither review named by the decision | state whether the end state removes the Gate-B skip or make every Gate-B-skipped change mandatory for Sample-Gate audit, and assign that rule to step 6 +MAJOR | high | §1 pipeline, line 39, versus §8 non-goal 5 and §10 line 380 | the pipeline ends at "Merge/Deploy" although two later settled statements say deployment is out of scope and the factory ends at merge | the canonical overview gives the future build path a station the scope explicitly excludes | change the pipeline terminus to Merge and show release/deploy, if mentioned at all, outside the factory boundary +MAJOR | high | §9 "Three test layers", lines 283-287, versus §2 decision 5 line 79 and §7 step 6 lines 231-232 | step 5's merge queue says a green candidate lands with no human involved, while decision 5 says merge remains human until step 6 ships and holds | the build sequence silently enables auto-merge one story earlier than the settled rollout allows | specify that step 5 prepares and smoke-checks the candidate but requires human merge, and that step 6 alone switches the green outcome to automatic landing +MAJOR | high | §2 decisions 2, 4, and 6 plus §9 "Phase 0 exists once" | classification is declared the mandatory first station before anything else, but Phase 0 must build tree v1 and goals/out-of-scope before the architecture verdict and before production; decision 2 also says the tree is initially built from the pool | the bootstrap order is cyclic and a later story must invent whether initial pool items are classified before or after the truth needed to classify them | define Phase 0 as an explicit bootstrap precondition outside the per-item station order, then state that classification is the first station after Phase 0 +MAJOR | high | §2 decision 6, lines 80-93, versus decision 3 lines 69-73 | the Intake-Loop is said to trigger exactly one thing and only attach cards, yet classification must also set status to klassifiziert and may create an architecture meta-story plus a wait relationship | write authority and atomicity are undefined on the architecture-breaking path, undermining the claimed guard against shared-field mutation | define one classification transaction with its complete allowed outputs, including card, status, meta-story, and dependency link, while preserving the no-production property +MAJOR | high | §5 "Verdicts", lines 196-200, versus §2 decisions 4 and 7 | `freigeben` is listed as a classification verdict even though Freigabe is a separate human-owned production decision that sets `freigegeben`; the two near-identical terms are not distinguished | an implementation can convert a classifier recommendation directly into the production-triggering status and bypass the only human trigger | rename the classification outcome to a non-authorizing state such as `freigabe-faehig` and define the separate human or standing-rule transition to `freigegeben` +MAJOR | medium | §2 decision 2 versus §9 "Phase 0 exists once", lines 248-254 | decision 2 says architecture is built from the pool initially, while the settled existing-project path says tree v1 is read from the codebase | the source-of-truth rule differs by paragraph and a Phase-0 story cannot know whether pool intent or current code wins when they disagree | scope "built from the pool" to greenfield projects and define reconciliation and precedence for existing projects using both codebase reality and pool intent +MAJOR | medium | §2 decision 1, lines 60-63, versus §3 stage 4, line 143 | decision 1 limits platform mechanics to a scheduled or loop clock wake-up, but stage 4 adds an event-triggered wake when a story becomes freigegeben without waiting for a tick | the hosting boundary for the proactive wake is undefined and may require a second platform mechanism forbidden by the decision | define event delivery as another supported clock adapter under the same platform boundary, or state that the event only shortens the next scheduled tick rather than starting a separate session +MAJOR | high | §2 decisions 7 and 8 | decision 8 says human eyes are never O(stories), but decision 7 explicitly starts Freigabe at per-story granularity and decision 5 keeps every merge human until step 6 | the maturity ladder's safe initial state contradicts its scaling invariant, so later stories cannot tell whether per-story human review is permitted during rollout | qualify the O(waves + exceptions) rule as the end-state target and explicitly permit the per-story bootstrap rungs until evidence authorizes promotion +MAJOR | high | §9 "parallelism budget", lines 271-275, §2 decision 7, and §11 pool-storage question | `lanes` and other human-owned knobs are required to live in a committed pool header and be protected by repository mechanics, while the pool may still be an external board | the external-storage option cannot satisfy the stated location, commit history, hook protection, or atomic-write assumptions | either settle on an in-repo pool or restate the requirements in storage-neutral terms with versioned audit history, compare-and-set operations, and an enforcement adapter for each allowed backend +MAJOR | high | §2 decision 7 and §4 "execution plan", lines 159-166 | Freigabe approves a rendered wave plan, but that plan is never stored and is recomputed every tick from mutable priorities, dependencies, waves, tree, lanes, and throttles | the clock can execute a materially different plan from the one the human approved, especially after concurrent input changes | bind every approval or notice to an input-version digest and require re-render plus re-approval when approval-relevant inputs change, while keeping the plan itself a generated view +MAJOR | high | §2 decision 8, lines 116-121, and decision 4 | an architecture merge mechanically reclassifies cards, but the document does not revoke or suspend an already-freigegeben card or stop an already-running lane whose verdict became stale | the clock can pull stale approved work, and an in-flight lane can continue against an architecture it was never classified against | define architecture-version binding on cards and lanes, atomically move affected approved items to a non-pullable reclassification state, and specify the stop or revalidation rule for running lanes +MAJOR | high | §2 decision 6 and §5 "Verdicts" | the write-authority rule names only `klassifiziert` and `freigegeben`, while the card introduces teilen, rückfragen, warten-auf, ablehnen, unclassified, AC-change-needed, and split-child paths without transition owners or terminal semantics | undefined transitions invite exactly the free-form status edits the vision says break automated resolution | add a storage-neutral state machine naming every state, authorized operation, precondition, side effect, and terminal or resumable outcome +MAJOR | high | §5 dependency dimension and `warten-auf` verdict, lines 192 and 196 | dependencies have no behavior for cycles, self-dependencies, missing or deleted targets, a target that is rejected or split, or a dependency satisfied while the waiter is concurrently edited | waves can deadlock forever or release work against the wrong predecessor | require cycle detection and define dependency identity, retargeting after split, failure propagation, and the atomic transition when all dependencies become satisfied +MAJOR | high | §2 decision 6 and §10 run-lock adoption | debounced story-appearance events can be delivered twice or overlap, but the only proposed lock is for orchestrator ticks and no per-item claim, idempotency key, or compare-and-set rule protects Intake-Loop outputs | duplicate cards, duplicate meta-stories, or conflicting architecture verdicts can be created before production even starts | define stable item IDs, idempotent classification writes, and atomic claiming or version checks for overlapping intake batches +MAJOR | high | §10 "run lock", lines 348-351 | a lock with stale-lock cleanup is not sufficient concurrency control: a paused but live owner, especially one in the specified usage-limit countdown, can be declared stale and run concurrently with its replacement | double orchestrators can mutate the same pool and open duplicate lanes despite the lock's stated purpose | use an atomic project-scoped lease with heartbeat and fencing token, keep it valid during usage holds, and require every mutating operation to reject an obsolete fence +MAJOR | high | §2 decision 8 and §9 "Architecture churn", lines 256-263, versus §5 split rules | batching every architecture-relevant item into one meta-story per wave can combine independent subsystems, mixed profiles, and different architecture verdicts, all of which §5 says must split | the O(waves) scaling guard creates stories the classifier itself would reject as oversized or mixed-risk | define batching by compatible branch and profile groups and revise the bound to the number of architecture batches per wave, or state the explicit exception and how its risk is contained +MAJOR | medium | §9 "The tree is the parallelism map", lines 264-270 | "disjoint branches" and "same branch" do not define ancestor-descendant overlap, shared-root files, cross-cutting changes, multi-branch stories, or a lane whose touched branch set expands during implementation | apparently disjoint lanes can conflict or violate dependency boundaries after they start | define branch-claim overlap semantics, shared nodes, multi-branch locking, and a mandatory stop plus reclassification when observed scope exceeds the claim +MAJOR | high | §4 "Waves structure a new project", lines 167-174, versus §2 decision 7 | the next wave may open automatically when the prior wave is fully merged, but wave-level Freigabe is defined as a human "Go" unless the human has explicitly raised the knob to standing auto-Freigabe | automatic wave advance can become an unapproved production trigger | make automatic opening conditional on a committed standing rule for that wave or keep it as a rendered proposal awaiting the required wave-level Go +MAJOR | medium | §2 decisions 7 and 8 | maturity knobs only specify how human presence is reduced after good P8 evidence; no adverse-evidence rule lowers the automation rung, raises sample rates, or revokes standing auto-Freigabe | a system that degrades after promotion stays autonomous despite the text saying presence is a dial rather than a one-way ratchet | define downgrade triggers, authority, hysteresis, and the immediate safe state after audit rejection, repeated smoke failure, or metric deterioration +MINOR | high | §9 "parallelism budget", lines 276-279 | the human-decision throttle uses an undefined `N`, distinct from the defined `lanes` value and also reused informally for worktree count | the orchestrator cannot deterministically decide when the decision queue blocks new lanes | name a separate committed knob, its default and valid range, and the exact comparison including equality +MINOR | high | §1 pipeline, lines 35-38 | Sample-Gate and Audit are drawn as an unconditional linear pair, with no explicit not-selected edge from Sample-Gate to Merge-Queue | the ordinary unsampled path is absent from the canonical flow and later stories may incorrectly require every PR to enter Audit | draw selected and not-selected branches, with both paths converging on Merge-Queue +MAJOR | high | §10 audit rejection rule, lines 358-361, and §9 merge queue | audit selection and approval are not bound to a content fingerprint, and no rule invalidates them after a lane fix, Gate-B fix, rebase, or smoke failure | a changed PR can inherit an audit result for different bytes, while a previously rejected PR's mandatory re-audit can race later revisions | bind draw and audit records to the reviewed candidate fingerprint and define which mutations require Gate B, resampling, and mandatory re-audit before queue entry +MAJOR | high | §2 decision 9, lines 122-134 | the AC fingerprint mechanism never defines its baseline, canonical AC-block serialization, storage, or the durable proof that an authorized pool round-trip occurred | Gate B cannot distinguish an authorized AC correction from reward hacking, particularly when the pool may be external | define canonicalization, the lane-opening baseline, versioned mutation record, and exact comparison and authorization evidence supplied to Gate B +MAJOR | medium | §10 model-strength knob, lines 352-354 | "reviewer is never weaker than the builder" depends on an ordering across model families that the document never defines, including unknown, fallback, or newly released models | the guard can silently accept a weaker reviewer or stop every run depending on an implementer's ad hoc ranking | require a project-owned role-to-model policy with an explicit strength order and a fail-closed unknown-model path, while retaining cross-family independence +MAJOR | high | §5 classification dimension 2 and §7 step 1 | §5 says the current profile drives the pass floor, but shipped CLAUDE.md still has a hard three-pass floor and explicitly says lenses do not add passes; profile-derived floors are only the in-flight step-1 work | the vision misstates current kit behavior and can cause later steps to assume economics controls have already shipped | say that the profile currently drives lenses and evidence and will drive the floor only after step 1 lands and is activated +MINOR | high | §5 classification dimension 3, lines 184-185 | "too thin → one question back" misstates the shipped intake skill, which allows one question round containing multiple targeted questions | a later classification story may narrow intake and leave several independent gaps unresolved while claiming parity | change this to "one question round" and preserve the shipped stop-without-writing behavior when grounding remains insufficient +MINOR | high | §5 introductory count, lines 179-194 | the text says dimensions 5-7 are new but the numbered list adds four dimensions, 5-8 | the mechanically stated count is false and obscures that Wave is also new | change the claim to "5-8 are new" +MINOR | high | §6 gap 2, lines 204-206 | the classification gap lists scope, architecture, and dependency but omits the new Wave dimension from §5 | the gap inventory and build decomposition can declare classification complete without implementing wave assignment | add Wave to gap 2 and its owning step +MAJOR | high | §7 step 4 | one promised story combines pool storage and state, classification, architecture truth and projection, Phase 0, waves, Freigabe-adjacent mechanics, and part of a blocking enforcement mechanism | this violates the document's own split rule for multiple independent subsystems and cannot pass through one coherent spec-plan-PR cycle without silent scope invention | decompose step 4 into ordered stories with named interfaces, at minimum pool/state, classification, architecture bootstrap/projection, and wave/Freigabe integration +MAJOR | high | §7 step 5 | one promised story combines scheduler adapters, polling and events, drift audit, PR processing, E2E execution, merge coordination, doctor, notifications, run locking, usage holds, budgets, and possibly enforcement | these are independent subsystems and failure domains, contradicting the one-story-per-step promise and §5's epic split rule | split step 5 into separate stories behind a shared clock-loop contract, with merge queue and E2E as distinct consumers +MAJOR | high | §7 build path | no numbered step explicitly owns decision 7's rendered wave plan, Freigabe state machine, maturity knob, standing rule, or approval invalidation | a central pipeline station can remain unbuilt even when all six roadmap steps are marked done | add explicit ownership, preferably a separate Freigabe and wave-control story after pool state and before autonomous scheduling +MAJOR | high | §7 build path and §2 decision 9 | no numbered step explicitly owns AC read-only enforcement, the pool round-trip, the AC fingerprint check, or spec-delta presentation; §10 assigns only AC IDs and test reports to steps 1 and 4 | the reward-hacking guard can be omitted while the build path appears complete | assign all decision-9 mechanics to a named story and order it before any autonomous lane execution +MAJOR | high | §3 maturity ladder versus §7 steps 5-6 | step 5 is labeled Stage 3 but includes event loops, which §3 defines as Stage 4, while step 6 is separately labeled Stage 4 and decision 6 calls the intake event loop the first Stage-4 instance | maturity, rollout, and ownership use incompatible stage boundaries | keep polling and scheduled loops in step 5, move event-triggered wakes to step 6 or explicitly redefine the stage table and all dependent labels +MINOR | high | §1 pipeline versus §7 step 5 and §9 "Three test layers" | the full E2E clock loop is a named station-like loop elsewhere but is absent from the fenced canonical pipeline | the required station-consistency check fails and the overview hides the only post-merge integration loop | add an asynchronous E2E-Loop branch after merges or wave close, including its failure edge back to Story-Pool +MAJOR | high | §9 "Three test layers", lines 280-290 | "main is never red" overstates what the merge-queue smoke comparison proves, because the full E2E suite deliberately runs only nightly or at wave close and may later fail on already-merged main | operators and later stories may treat smoke success as a guarantee of full-system health, the exact gate-overclaim class AGENTS.md forbids | state the bounded claim: the queue does not land a candidate that fails the configured smoke command; E2E can still discover main-level failures later +MAJOR | high | §9 "Three test layers", lines 290-291 | a delayed E2E failure is required to carry the trace ID of "the suspect merge", but a nightly or wave-close run can cover many merges and the document defines no attribution method | the automatically filed story can accuse the wrong change and create misleading dependencies or rollback work | carry the tested merge range and all candidate trace IDs, mark attribution unknown unless deterministically isolated, and assign bisect or triage behavior to step 5 +MAJOR | high | §9 merge queue, lines 283-290 | the rebase stage has no terminal path for conflicts, an unclean worktree, a branch deleted while queued, or a retry invalidated by a newer main | queue progress and lane ownership are undefined before smoke can even run | define a rebase-failure artifact and return transition, retry cap, branch-version check, and whether the queue continues or pauses for each class +MAJOR | high | §10 "documented scars", lines 332-337 | run-artifact retention is declared a constraint "from day one" but has no owning build step, even though step 2 starts persistent SQLite and trace artifacts | the first analytics story can reproduce the cited unbounded-disk P0 defect before a later story notices it | assign retention, compaction, deletion authority, and failure behavior to step 2 as part of the initial artifact store +MAJOR | high | §10 "documented scars", lines 332-335 | planning artifacts are required to have a defined home that cannot contaminate implementation branches, but no build step owns or even chooses that home | separate-context handoffs can leak scratch plans into lane diffs or lose the only copy of an approved plan | assign the artifact-home contract to the orchestrator story and define tracked versus ephemeral locations and branch visibility +MAJOR | high | §10 analytics adoption and documented scars, lines 304-337 | the adopted metrics are written post-run and the dashboard is parked, yet real-time cost visibility is retained as the control that prevents a token furnace | post-run analysis detects overspend only after it happened and therefore does not satisfy the stated preventive constraint | separate passive P8 analytics from live budget visibility and assign live counters, alerts, and stop thresholds to the budget or dashboard story before autonomy rises +MAJOR | high | §10 analytics lines 304-308 and minimum-review signal lines 355-357 | analytics writes are non-fatal, but review cost and duration later decide whether a pass counts; the missing, partial, duplicate, or unattributable-metric path is not defined | a telemetry failure can either launder an under-floor pass or halt gates indefinitely, and the non-fatal promise gives no safe answer | let run telemetry fail non-fatally but treat missing or unattributable gate measurements as INCOMPLETE, deduplicate by pass ID, and exclude incomplete windows from maturity evidence +MAJOR | medium | §10 architecture projection, lines 312-317 | the generated machine-readable projection has no freshness binding to the AGENTS.md architecture version and no required behavior on parse failure, ambiguity, or cycle detection failure | the scheduler can open lanes from stale or invalid topology while still claiming AGENTS.md is the only source | regenerate against a recorded architecture digest before classification and each tick, and stop affected scheduling with a diagnostic artifact when generation or validation fails +MAJOR | high | §10 "Second sweep", lines 339-340 | the cited `docs/research/2026-08-30-*-gap-sweep.md` files exist only locally under an ignored `docs/research/` directory and `git ls-files` returns none of them | a future reader cloning commit 23d8889 cannot inspect the evidence the vision says its second sweep used | move the cited gap sweeps to a tracked evidence location or replace the citation with tracked artifacts or stable public sources +MINOR | high | §2 decision 7, line 107 | `non-goal 4` does not resolve to a numbered item because §8 uses unnumbered bullets | the mechanical cross-reference is invalid even though the intended autonomy-evidence bullet can be inferred by position | cite `§8 "No autonomy expansion ahead of measured evidence"` or number the non-goals +MINOR | high | §10 "Deliberately not adopted", line 329 | `non-goal 1` does not resolve to a numbered item because §8 uses unnumbered bullets | the mechanical cross-reference is invalid and can drift if bullets are reordered | cite `§8 "No daemon or server-side runner"` or number the non-goals +MAJOR | medium | §1 vision prose, lines 49-52 | "a human override" is left unconstrained even though shipped CLAUDE.md §5 says a human exception supplies no permission to bypass gates, STOP states, profile evidence, or AGENTS.md invariants | later stories may turn an audit or operator action into a generic escape hatch from mandatory rules | define overrides as changes only to explicitly human-owned product knobs and state that mandatory workflow and invariant obligations retain their existing terminal actions +MINOR | high | §9 Phase 0, lines 248-255 | the greenfield bootstrap has no empty-pool outcome even though tree v1 is derived from the pool and filling the pool is allowed to be empty | Phase 0 cannot establish architecture or decide whether to wait, ask for goals, or produce a minimal root tree | define the empty-pool terminal state and the minimum human inputs required before Phase 0 can complete +MINOR | high | §4 wave progression, lines 167-174 | "prior wave is fully merged" has no definition when a wave contains rejected, split, deferred, waiting, or permanently failed items | wave close can block forever or advance while unresolved work is silently abandoned | define which terminal statuses count as resolved for wave closure and how carried-over items are reassigned +NIT | medium | §1 vision prose, lines 42-49 | "Every node is a loop with its own fresh context" is literally false for static storage, Freigabe and Audit humans, and the merge operation shown as nodes in the same diagram | the universal wording blurs which stations require fresh model sessions and which are state stores or human actions | narrow the claim to model-operated stations and label stores, human gates, and mechanical queue operations separately +END OF FINDINGS (52 total) diff --git a/.context/codex-reviews/gate-a-spec-vision-pass-2-dispositions.md b/.context/codex-reviews/gate-a-spec-vision-pass-2-dispositions.md new file mode 100644 index 0000000..d61517d --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-vision-pass-2-dispositions.md @@ -0,0 +1,41 @@ +# Gate A (spec) — dark-factory vision — pass 2 dispositions +34 findings: 1 BLOCKER, 25 MAJOR, 7 MINOR, 1 NIT. Unprofiled run. +Trend: pass 1 = 52 findings / 41 Blocker+Major; pass 2 = 34 / 26. + +## Fixed (all Blocker/Major, plus the cheap Minors) +BLOCKER 1 — gateless plus automatic merge could land an author-only candidate during a +reviewer outage. Decision 5 now makes reviewer availability an input to merge +authorization: a gateless candidate never lands automatically. + +Corrections of pass-1 corrections (absorbed, in scope): +- P2-4 my decision-2 edit left AGENTS.md/code/pool as a three-way ambiguity → the three are now named and kept distinct. +- P2-5 my decision-1 edit promised the event "shortens the next tick", which nothing can observe → named as stage-3 polling. +- P2-20 my 5b sat under a clock-loop contract that does not fit a serial station → contract narrowed. +- P2-15 my epic split made the preamble's "each step is one story" false → the unit is the leaf. +- P2-33 step 4's "(decisions 2 and 4)" went stale when 4b/4d/4e were added. +- P2-30 my decision-8 rewrite overstated the evidence: three fallback designs failed, only the first was the same-family tier-2 shape. Corrected — this is the overclaim class AGENTS.md names. + +Pre-existing contradictions: +- P2-2 §1 called Freigabe human-only while decision 7 lets a standing rule write it → the pipeline names the authorization state, and decision 6 names the standing-rule operation. +- P2-3 decision 8's knob-everywhere vs §8's four protected paths → knobs are routine touchpoints only. +- P2-6 a meta-story could spawn a meta-story forever → a meta-story is the amendment and produces no further one. +- P2-7 "churn blocks branches, never the factory" vs a global mandatory-stop throttle → the global brake is stated and the guarantee narrowed. +- P2-23 "every clock loop starts report-only" vs the E2E loop filing a story → report-only defined as no code and no status change; filing is the one write. +- P2-29 "filling the pool is always consequence-free" — totality overclaim → "never triggers production", with what it does cost stated. + +Missing stations and owners: +- P2-9/10/11/12/13 pipeline gained PR, Bewertungs-Loop, Drift-Audit, the vet preflight and the judge sidecar. +- P2-16/17 step 2 split into 2a/2b/2c; the tracked P8 story stays read-only over the ledger and git, and the vision's analytics, tracing, spec-delta and live cost move to a new 2c. Five stale "[step 2]" references repointed. +- P2-19 step 3 owns the plan interface; the working dry-run lands after 4d and 5a. +- P2-21/22 step 6 split into 6a-6d, giving the promotion path past level-0 an owner. +- P2-28 the blocking hook also meets invariants 2 and 4; the meta-story must say which it amends and which it preserves. +- P2-31/32 the gap-sweep paths are no longer cited, since a clone cannot reach them. +- P2-34 §11's heading no longer claims every question has an owner; the three parked ones are marked. + +Recorded in §11 rather than specified (the approved shape): P2-8 Gate B vs the rebase [5b], +P2-14 the hardening ledger has no station [4b/5d], P2-18 promotion evidence [6d], +P2-24 wave-close ordering [4d], P2-25 AC canonicalization [4e], P2-26 the stopped lane [4e], +P2-27 given/when/then and the per-AC test report [4e], P2-9 PR readiness before the queue [5d]. + +## Collected, not iterated +None outstanding — every pass-2 Minor and the Nit were cheap enough to fix in place. diff --git a/.context/codex-reviews/gate-a-spec-vision-pass-2.md b/.context/codex-reviews/gate-a-spec-vision-pass-2.md new file mode 100644 index 0000000..cd6079f --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-vision-pass-2.md @@ -0,0 +1,35 @@ +BLOCKER | high | §2 decisions 5 and 8, lines 93-140 | the reviewer-unavailable state is declared gateless, but the end state also auto-merges undrawn candidates and promises that no merge skips both Gate B and human audit; no rule changes Sample-Gate or merge behavior when the cross-family reviewer disappears | an outage can send an author-only candidate down the undrawn path and land it automatically, contradicting both the four-eyes principle and decision 5 | make reviewer availability a merge-authorization input: when Gate B is unavailable, halt automatic merge or force every candidate through a human audit until an independent family returns +MAJOR | high | §1 pipeline lines 26-30 and §2 decisions 6-7 lines 96-121 | Freigabe is labeled a human-only production trigger and decision 6 authorizes only human freigeben or loop klassifiziert status operations, while decision 7 says a standing rule automatically writes freigegeben | the automatic maturity rung has neither the actor nor the defined operation required by the settled write-authority rule, so implementations can either bypass the seam guard or never auto-approve | describe Freigabe as the sole authorization state rather than an always-human actor, and name the standing-rule transition as a distinct authorized operation with its preconditions and audit record +MAJOR | high | §2 decision 8 lines 138-140 versus §8 lines 295-297 | every mandatory human touchpoint is said to have a knob that lowers human presence, while the non-goals prohibit removing the human from mandatory stops, profile confirmations, scope changes, and architecture meta-stories | no knob can both lower presence at every mandatory touchpoint and keep those listed touchpoints permanently human, so later stories must silently choose which settled sentence wins | limit maturity knobs to routine approval or sampling touchpoints and state that the four non-goal paths remain fixed human escalations without a presence-lowering knob +MAJOR | high | §2 decision 2 lines 77-84 | AGENTS.md is called project truth, the architecture tree is subordinate to the pool, and code is then called the fact when sources disagree, without defining the precedence or state of the AGENTS.md tree in that three-way disagreement | classification reads AGENTS.md and can issue an architecture verdict that conflicts with the code reality the same decision says wins | distinguish normative approved architecture, observed code state, and desired pool intent, then define which mismatches block classification and what Phase 0 must reconcile before AGENTS.md becomes authoritative +MAJOR | high | §2 decision 1 lines 70-76 and §3 stage 4 line 168 | when a platform has no event trigger, the text says the event still shortens the next scheduled tick, but no mechanism can observe or reschedule from that event; decision 6 and the stage table still classify this fallback as proactive | on platforms lacking event delivery the promised stage-4 wake silently degrades to ordinary polling with an undefined latency | state that this fallback is stage-3 polling with the configured tick latency, or name the actual adapter that receives the event and reschedules the tick +MAJOR | high | §2 decision 3 lines 85-89 and §5 architecture verdict lines 221-222 | an architecture-breaking story creates an architecture meta-story through the same classifier, but nothing prevents that meta-story, whose purpose is itself to break or extend the tree, from generating another meta-story | the factory can recurse forever before any architecture change reaches Gate A | define a recognized architecture-meta-story type whose high-risk classification authorizes tree amendment without spawning another meta-story, while retaining the branch lock and human escalation +MAJOR | high | §9 architecture churn lines 313-320 versus parallelism throttle lines 328-336 | architecture churn is promised to block only touched branches while other branches keep flowing, but any mandatory stop in any lane globally forbids every new lane from opening | an architecture stop can drain all running lanes and freeze the whole factory, contradicting the branch-scoped settled flow rule | scope architecture-related mandatory-stop throttling to the claimed branches, or explicitly state that the global stop rule overrides the branch-only guarantee and narrow that guarantee +MAJOR | high | §9 merge queue lines 337-349 and §11 audit-binding gap lines 553-555 | the queue rebases onto newer main after Gate B, then runs only smoke before landing; the document records invalidation for human audit but no equivalent binding or re-review rule for the adversarial Gate-B result | the rebased candidate and its composed diff may differ from what Gate B reviewed, so the claim that the adversarial gate checks the merge outruns the actual comparison in violation of the AGENTS.md gate-claim rule | bind Gate B to the exact post-rebase candidate and re-run it whenever the candidate or base changes, then run smoke and audit against that same immutable candidate +MAJOR | high | §1 pipeline lines 31-40 and §7 step 5d line 285 | Sample-Gate draws PRs even though no station creates a PR, and PR-bot processing exists only as an unsequenced clock sub-story with no readiness edge into Sample-Gate or Merge-Queue | automatic merge can run before PR creation, required CI readiness, or any configured wait-for bot processing, and step 5d has no defined consumer | add explicit Open-PR and PR-readiness or bot-processing stations before Sample-Gate and queue entry, with the shipped routing rules and final-head checks as inputs +MAJOR | high | §1 fenced pipeline versus §2 decision 2, §5 dimension 6, §6 gap 3, and §9 lines 313-320 | the Bewertungs-Loop is a named continuous station elsewhere but is absent from the canonical pipeline, and neither 4b nor 4c unambiguously owns its ongoing re-evaluation, batching, and post-merge reclassification behavior | a build can finish the bootstrap projection and classifier while omitting the continuous loop that creates architecture meta-stories | add the Bewertungs-Loop and its stale-verdict edges to the pipeline and assign its runtime contract explicitly to 4b or 4c +MINOR | high | §1 fenced pipeline versus §3 line 167 and §7 step 5d | the drift-audit loop is a named stage-3 station but is absent from the fenced pipeline | the requested station-parity check fails and the reverse-traceability consumer has no visible place in the factory | add an asynchronous Drift-Audit loop with its report or pool-story edge and its 5d owner +MINOR | high | §1 fenced pipeline versus §1 lines 54-56 and §10 lines 380-382 | the judge or watchdog is described as the deliberate process-observing exception and an adopted supervisor but is not represented in the canonical pipeline | liveness escalation and repeated-smoke handling depend on an invisible actor whose inputs and affected loops cannot be traced from the overview | show the watchdog as a sidecar over the model loops and queue, with liveness-only inputs and human-escalation outputs +MINOR | high | §1 fenced pipeline versus §10 lines 374-379 | the adopted mechanical vet preflight must run before every model pass but is absent from the canonical pipeline | a future story can implement the visible model stations without the prerequisite check and still appear pipeline-complete | add a cross-cutting Vet preflight before each model-operated node, or annotate those nodes with the common preflight contract +MAJOR | medium | §1 pipeline and the shipped workflow architecture in AGENTS.md and docs/coding-workflow.md | the pipeline carries Gate findings straight toward merge and never invokes harden-finding or the append-only hardening ledger, although self-hardening is one of the kit's three defining mechanisms | autonomous repairs can fix instances without recording fingerprints or escalating recurring classes, silently dropping a shipped load-bearing workflow property | add a findings-disposition and hardening edge for accepted actionable gate and PR findings before merge readiness, or explicitly assign an equivalent ledger consumer to an existing station +MINOR | high | document preamble lines 3-6 and §7 heading line 250 | the document says each build step is one story, but steps 4 and 5 are expressly epics made of nine sub-stories | the decomposition's stated unit and its actual units disagree, which makes references such as step 4 and step 5 ambiguous as story owners | say each leaf build item becomes a story and that numbered epic containers are ordering groups, then use leaf IDs for ownership +MAJOR | high | §7 step 2 lines 255-256 | one numbered item combines the already-separate loop-rule-consolidation successor story with the already-separate P8 passive-metrics story despite the one-story-per-leaf rule | sequencing, profile, Gate A, and completion cannot be attributed to one story, and a marked-done step could hide one unfinished half | split step 2 into two ordered leaf steps with their existing story paths and separate completion conditions +MAJOR | high | §7 step 2 and §10 lines 367-370, 424-426, 454-456 | local SQLite writes, per-stage tracing, spec-delta capture, gate cost instrumentation, and harness-token measurement are assigned to P8, but the tracked P8 story explicitly keeps analysis-only, no new state, no instrumentation, and nothing written back | the vision silently replaces settled scope in an existing story and makes its build order impossible as written | preserve P8 as its tracked read-only analysis and create a separate telemetry or observability story before it for the new state, trace, spec-delta, and live measurement contracts +MAJOR | high | §2 decisions 7-8, §8 lines 298-299, and the tracked P8 story | autonomy promotion is conditioned on P8 evidence, but P8 only analyzes ledger recurrence and compares review-severity mixes across different artifacts; it defines no measure of Freigabe accuracy, audit escapes, merge safety, or per-knob promotion threshold | the central safety condition can be declared satisfied by evidence that does not measure the behavior being automated | assign each autonomy knob an outcome metric, adverse-event measure, sample window, and promotion threshold in a new evaluation story rather than treating review-economics P8 as universal evidence +MAJOR | high | §4 execution-plan view lines 184-190 and §7 step 3 | step 3 owns printing a tick execution plan before steps 4a-4d define the pool, dependency graph, tree projection, waves, lanes, approval binding, and before 5a defines the scheduler tick | step 3 cannot produce the promised plan from interfaces that do not yet exist and will have to invent or later rewrite its contract | make step 3 define only the orchestrator role and plan interface, then implement the real dry-run after 4d and 5a, or reorder the dependent stories +MAJOR | high | §7 step 5 lines 278-285 and §1 lines 47-49 | all 5a-5d stories are said to sit behind one clock-loop contract, but 5b is explicitly a serial mechanical merge-queue operation rather than a clock loop; tick, usage-limit hold, and report-only semantics do not naturally apply to it | the new sub-story split still gives the queue the wrong lifecycle contract and leaves its actual trigger and ownership undefined | move 5b outside the clock-loop epic or define a narrower shared run contract and a separate event-driven queue contract +MAJOR | high | §7 step 6 lines 286-289 and §5 split rule lines 227-231 | step 6 remains one story combining the freigegeben event adapter, Sample-Gate drawing and audit state, and graduated automatic merge, which are independent failure domains and interfaces | the new decomposition applies the epic split rule to steps 4 and 5 but violates it at the final and highest-risk automation step | split step 6 into ordered leaves for event wake, sample and audit, and automatic-merge rollout, with explicit interfaces and profiles +MAJOR | medium | §2 decision 5, §7 step 6, and §1 end-state claim | the build path ends after automation is enabled first only for level-0 stories, but no later leaf or explicit maturity operation owns expansion to standard or high stories required by the sampled-audit end state | completing every listed story can still leave per-merge human approval for most work while the vision appears built | define the profile-threshold knob and evidence-driven promotion path inside a named step-6 leaf, or add subsequent rollout leaves through the stated end state +MAJOR | high | §7 steps 5a and 5c plus §9 lines 352-355 and §10 lines 447-448 | every clock loop starts report-only, yet the E2E clock loop automatically creates a pool story on failure; filing a story is a durable pool mutation, not report-only behavior | step 5 cannot satisfy both its shared starting contract and its E2E outcome, and the first run may mutate production-control state without the required maturity authorization | distinguish report generation from authorized operational writes and specify whether auto-filing is initially a proposed artifact, a human-confirmed write, or a separately raised mutation knob +MAJOR | medium | §4 lines 197-201 and §9 lines 337-355 | automatic next-wave opening is tied to prior-wave completion while the wave-close E2E run is merely nightly or at close, with no ordering or failure rule between E2E, the as-built view, closure, and the standing opening rule | the next wave can start before the previous wave's integration result exists, and a later E2E failure has no defined wave assignment or reopening effect | define the wave-close transaction order and state whether E2E gates closure, delays automatic opening, or files carry-over work into a named wave without reopening +MAJOR | high | §2 decision 9 lines 147-159 and §11 recorded gaps | the AC comparison still lacks the lane-opening baseline, canonical block serialization, storage, and durable proof that a pool round-trip authorized a changed fingerprint, and none of the 21 new gaps owns those prerequisites | Gate B cannot distinguish an authorized human correction from reward hacking, especially with an external pool | add a 4e-owned gap naming canonicalization, baseline version, mutation record, and the exact evidence Gate B compares +MAJOR | medium | §2 decision 9 lines 147-153 | when a lane stops for an AC change, the document does not say whether its partial implementation is discarded, quarantined, or resumed after human mutation and reclassification | preserving it lets builder-influenced work survive under changed goals, while discarding it can lose valid work; either choice also affects the later Gate-B baseline | name this stopped-lane state and assign 4e the disposition and resume rule for its worktree, plan, evidence, and review counters +MAJOR | high | §10 lines 443-446 versus §7 steps 1 and 4e and the tracked review-economics story | given-when-then ACs and the per-AC test report are owned by step 1 or 4, but step 1's existing story does not contain them and 4e names IDs, immutability, fingerprinting, and round-trip only | the machine-checkable goal report can be omitted while both cited owners truthfully finish their stated scopes | assign given-when-then normalization and the per-AC test-report artifact explicitly to 4e or to a new predecessor leaf; remove step 1 as an owner unless its settled story is separately amended +MAJOR | high | §10 blocking-hook prerequisite lines 407-416 versus AGENTS.md invariants 1, 2, and 4 | the prerequisite accounts only for amending invariant 1, but a blocking hook also inherits invariant 2's fire-on-uncertainty direction and invariant 4's POSIX-sh and optional-jq constraints; the document does not decide whether parse, storage, or environment uncertainty blocks writes | amending exit status alone can turn the advisory false-positive policy into denial of every protected write or produce a second invariant violation | require the meta-story to account for all existing hook conditions, explicitly defining fail-open or fail-closed behavior for each local failure and preserving or amending invariants 2 and 4 deliberately +MINOR | high | §1 line 21 and §2 decision 6 lines 96-109 | filling the pool is called always consequence-free even though it spends model tokens, writes cards and statuses, may create a meta-story, and can enqueue human questions; later prose only establishes that it triggers no production | the totality claim overstates the mechanism in the same way AGENTS.md warns against for gates and checks | narrow the phrase to production-consequence-free or state exactly that pool insertion cannot open a lane or merge code +MINOR | high | §2 decision 8 lines 128-131 and the cited reviewer-availability fallback design | the sentence says the same-family tier-2 reviewer was designed three times, but the cited record says three zero-pass fallback designs were tried and only the first design contained a same-family tier-2 reviewer | the historical rationale attached to a settled gate rule is factually wider than its evidence | say three fallback designs across nine passes failed, one of which was the same-family tier-2 design +MAJOR | high | §10 second-sweep citation line 403: docs/research/2026-08-30-godarkfactory-gap-sweep.md | the cited file exists locally but is ignored by .gitignore and absent from git ls-files, so a clone of the reviewed commit cannot reach it | the required cited-path reachability check fails for this source and later readers cannot inspect it | track the file in a durable evidence location or remove the path citation and cite only reachable source material +MAJOR | high | §10 second-sweep citation line 403: docs/research/2026-08-30-transcript-gap-sweep.md | the cited file exists locally but is ignored by .gitignore and absent from git ls-files, so a clone of the reviewed commit cannot reach it | the required cited-path reachability check fails independently for the second source and later readers cannot inspect it | track the file in a durable evidence location or remove the path citation and cite only reachable source material +NIT | high | §7 step 4 heading lines 258-277 | the parenthetical says the epic covers decisions 2 and 4, but its new leaves explicitly implement decisions 6, 7, and 9 as well | the decision cross-reference became stale when 4b, 4d, and 4e were added | list decisions 2, 4, 6, 7, and 9 or omit the non-exhaustive parenthetical +MINOR | high | §11 heading and lines 472-485 | the section claims its open questions are owned by stories, but hidden verification is unscheduled and the dashboard and shortcut questions are merely parked without any numbered build leaf or story owner | open load-bearing work can remain outside the roadmap while the ownership heading claims complete decomposition | give each parked question a named future story or backlog owner, or narrow the heading to questions with owners and separate the unowned parking lot +END OF FINDINGS (34 total) diff --git a/.context/codex-reviews/gate-a-spec-vision-pass-3-dispositions.md b/.context/codex-reviews/gate-a-spec-vision-pass-3-dispositions.md new file mode 100644 index 0000000..2038f56 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-vision-pass-3-dispositions.md @@ -0,0 +1,51 @@ +# Gate A (spec) — dark-factory vision — pass 3 dispositions +32 findings: 2 BLOCKER, 22 MAJOR, 8 MINOR. Unprofiled run. + +## STATUS: loop stopped and surfaced. Nothing fixed in this pass. +Two of CLAUDE.md §5's five tells are present, which makes stop-and-surface +mandatory rather than discretionary. The findings stay OPEN, no pass is +credited clean, the floor and the clean-final-pass rule still stand, and the +loop resumes on whatever the human decides. This is not an exit from the gate. + +## The three lines (CLAUDE.md §5, pass-4-onward reporting, applied early) +1. TREND — pass 1: 52 findings, 2 Blocker, 41 Blocker+Major. + pass 2: 34 findings, 1 Blocker, 26 Blocker+Major. + pass 3: 32 findings, 2 Blocker, 24 Blocker+Major. +2. CLUSTER — three groups. (a) Ownership of the decomposition: 11 of the 24 + Blocker/Major say no §7 leaf owns something the document names elsewhere + (Bewertungs-Loop runtime, the judge, the vet preflight, the dry-run + implementation, reviewer-availability detection, the as-built view, + given/when/then). Seven of those eleven are consequences of the five + stations pass 2 added. (b) Claims checked against shipped kit behaviour: + process-pr-review, pr-review-bots.md, the P8 story, the hardening ledger — + genuinely new coverage, not regeneration. (c) Merge-queue edge states. +3. REQUIRE-WITHDRAW — one pair. Pass 2's blocker fix said "halt automatic + merge or force every candidate through a human audit". That was + implemented. Pass 3's blocker says the human-audit half is a + human-authorized zero-pass closure, which the cited fallback record rejects + and §1's own override rule forbids. + +## The tells +PRESENT: the Blocker count failed to fall (2 -> 1 -> 2). +PRESENT: one require-withdraw pair. +NOT PRESENT: findings rising (52 -> 34 -> 32, falling but flattening); +clustering on a test instrument (no instrument in this artifact). +UNCLEAR: clustering on prose about the decomposition rather than the +decomposition itself — arguable, since ownership assignment IS this +document's product. + +## The mechanism worth naming +The document grew 416 -> 442 -> 564 -> 680 lines across the three passes. Each +round's fix adds prose; added prose names mechanisms; named mechanisms need +owners; missing owners are the next round's findings. That is CLAUDE.md §5's +sizing guidance made visible — "prefer smaller specs with named interfaces and +let the plan carry the detail". Roughly half of pass 3 is that self-generated +surface; the other half is Codex reaching shipped files it had not checked +before, which is real coverage and argues against calling this a plateau. + +## Open, nothing acted on +BLOCKER: p3-1 (reviewer outage vs a sanctioned zero-pass closure), +p3-11 (a clock loop graduating to fix permission bypasses classification and +Freigabe — decisions 4, 6, 7). +MAJOR: p3-2..p3-9, p3-12, p3-13, p3-15, p3-18, p3-20..p3-25, p3-28..p3-31. +MINOR: eight, uncollected. diff --git a/.context/codex-reviews/gate-a-spec-vision-pass-3.md b/.context/codex-reviews/gate-a-spec-vision-pass-3.md new file mode 100644 index 0000000..9c55189 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-vision-pass-3.md @@ -0,0 +1,33 @@ +BLOCKER | high | §2 decisions 5 and 8, §1 override rule, and cited reviewer-availability fallback | temporary reviewer unavailability makes the cycle gateless and offers human audit as the alternative to waiting, which reads as a human-authorized zero-pass closure even though the cited fallback record rejects every sanctioned zero-pass closure and §1 says human overrides never waive a gate obligation | an unavailable Gate A can be bypassed to create the candidate and an unavailable Gate B can be replaced by manual audit, reopening the exact fallback question decision 8 says is closed | distinguish project-level explicitly inactive degraded mode from a runtime outage; make audit non-authorizing for an active gate and require the independent gate to resume before the cycle can close, unless a separate settled story deliberately changes the shipped rule +MAJOR | high | §2 decision 5 and §7 leaves 6b–6d | none of the step-6 leaves owns reviewer-availability detection, the gateless marker, the forced-audit or wait transition, or the condition that reauthorizes landing when a reviewer returns | all four automation leaves can finish while merge authorization still ignores the new input that was added to repair pass 2's blocker | assign the complete availability state machine and authorization check to a specific leaf, including outage during a pass, restoration, and the authority choosing wait versus audit +MAJOR | high | §1 Intake-Loop and Bewertungs-Loop, §2 decision 6, and §5 dimension 6 | Intake is said to attach the initial architecture verdict, §5 says that verdict comes from the Bewertungs-Loop, and the fenced order puts Bewertungs after Intake with no return edge | two fresh-context nodes can each believe the other produces the load-bearing verdict, or Intake can emit a complete card before its declared source has run | define one producer for the initial verdict and show the request and result artifact edges; reserve the other loop for later staleness re-evaluation if that is the intended split +MAJOR | high | §7 leaves 4b and 4c versus §1 Bewertungs-Loop | adding Bewertungs to the pipeline did not assign its continuous pool-versus-tree evaluation, meta-story batching, or post-architecture-merge reclassification runtime to either leaf; 4b owns proactive intake and 4c owns only bootstrap plus projection | the pool, classifier, and projection can all ship while the newly visible continuous loop remains unbuilt | add the Bewertungs-Loop contract to 4b or 4c, or split it into its own leaf with triggers, inputs, outputs, and staleness behavior +MAJOR | medium | §1 fenced pipeline and §7 step 3 | Takt says it wakes the model-based orchestrator, but the orchestrator is not a node in the fenced pipeline even though the Vet applies only to model-operated nodes above and the judge applies to the shown model loops | the controller itself can fall outside the fresh-context, preflight, artifact-handoff, and watchdog contracts while every visible station appears covered | show an Orchestrator node between Takt and lane execution, or state explicitly that Takt includes it and that every cross-cutting contract applies to that embedded model node +MAJOR | high | §7 step 3 and §10 orchestrator dry-run owner | step 3 is limited to the execution-plan interface and says the working dry-run lands after 4d and 5a rather than here, but no later leaf owns that implementation and §10 still assigns it to step 3 | the numbered build path can complete with only an interface, or step 3 must be reopened after it was already completed | either reorder step 3 after 4d and 5a or split a later dry-run implementation leaf and repoint §10 to it +MAJOR | high | §7 step 5 preamble and leaf 5d | the epic says its stories are separate failure domains, yet 5d combines drift auditing and PR-bot processing, two independent loops with different inputs, writes, retry behavior, and shipped predecessors | the leaf violates both its own failure-domain rule and §5's split rule, making one story and one profile cover unrelated subsystems | split 5d into one drift-audit leaf and one PR-polling or processing leaf behind the shared 5a contract +MAJOR | high | §1 Judge/watchdog, §7 step 5, and §10 judge adoption | no 5a–5d leaf owns the judge implementation, its process-signal inputs, notification outputs, or its supervision of non-clock queue work; step 5 is only an ordering group | the liveness sidecar and repeated-failure escalation can be omitted while every leaf story is truthfully complete | create a judge leaf or explicitly extend a named leaf with its full sidecar contract and queue coverage +MAJOR | high | §1 Vet preflight, §7 leaves 4a–4e, and §10 Vet bracket | §10 assigns Vet, numeric split thresholds, projection checks, and a pool-storage data point collectively to step 4, but step 4 is not a story owner and no leaf owns the cross-cutting preflight before every existing and future model node | completing 4a, 4b, and 4c can still leave the universal Vet station unimplemented, while later nodes assume it exists | split the bracket across 4a, 4b, and 4c and give the cross-node Vet runner its own explicit leaf or shared contract owner +MINOR | high | §10 graduated auto-merge bracket and §7 leaves 6b–6c | the coarse step-6 owner hides two different artifacts: 6c clearly owns mechanical merge thresholds, while no leaf explicitly owns punchlist generation even though the punchlist is the artifact human audit reads | sampled audit can ship with draw and verdict state but without its promised review input | assign thresholding to 6c and punchlist generation plus its binding to 6b, or add a separate artifact leaf +BLOCKER | high | §10 clock-loop fix-permission knob versus decisions 4, 6, and 7 | a clock loop is allowed to graduate from report-only to fix permission, which implies it may change product code directly, but the settled decisions make classified Freigabe the only production trigger and say merely filing a pool item cannot open a lane | a matured drift, E2E, or PR loop can bypass classification, flags, the rendered wave plan, and human or standing-rule authorization | define fix permission only as permission to create or advance a normally classified and freigegeben story through the factory, or remove direct fixing as a loop maturity state +MAJOR | high | §3 stage 3, §7 leaf 5d, and §10 report-only default | PR-bot processing is classified as a clock loop that starts report-only, but the shipped process-pr-review command fixes actionable regressions in the PR, replies to threads, hardens findings, commits, and rechecks CI | either the clock node cannot perform the station named processing, or it violates the shared report-only contract on its first run | separate a report-only PR poller that wakes the existing in-lane processing station from the processing station itself, and name the artifact and trigger between them +MAJOR | high | §11 hardening-ledger gap versus plugins/dev-workflow/commands/process-pr-review.md item 5 | the recorded gap says no station disposes of accepted bot findings into the ledger, but the shipped PR command already runs harden-finding for accepted actionable bot findings; the same gap omits findings from sampled human audits entirely | the vision duplicates an existing bot path while leaving its new highest-value audit-escape signal without recurrence hardening | narrow the missing path to accepted Gate A and Gate B findings, preserve the shipped bot route, and include accepted sampled-audit findings with an explicit owner +MINOR | high | §10 model-strength owner and §11 model-strength-order gap | the bracket cites step 1 even though the tracked review-economics story contains pass-floor and severity work, not a model-strength knob; step 3's leaf description also never names the knob | a cited existing story can finish without the adopted feature and later readers cannot tell whether step 3 alone was intended to own it | remove step 1 as an owner and add the knob explicitly to step 3, or create a new leaf whose story can define and validate the cross-family order +MAJOR | high | §11 Given/when/then and per-AC test-report gap | the text correctly says neither tracked step 1 nor leaf 4e owns the machine-checkable goal condition, then ends the same gap with [4e] and the section preamble says recorded gaps have owning steps | the gap is marked handled by an owner the prose itself proves cannot hold it, so the build can finish without the requirement | either extend 4e's leaf definition to own normalization and the per-AC report, or mark this gap explicitly unowned and add a new leaf +MINOR | high | §11 opening count of unowned questions | the section says exactly three topics lack numbered owners, but the later Given/when/then gap explicitly says it is unowned, and the autonomy-downgrade gap also calls material details unowned | the required stated-count check fails and the decomposition understates work outside its roadmap | recount all unowned items after resolving the contradictory brackets, then state the actual number or avoid an exact count +MINOR | high | §11 autonomy-downgrade gap versus decision 8 and leaf 6d | the gap says nothing raises human presence on bad evidence and that downgrade triggers are unowned, while decision 8 now requires the dial to rise on adverse evidence and 6d explicitly owns downgrade triggers | the newest repair leaves stale prose that makes 6d both owner and non-owner and obscures which details are genuinely still open | say the desired upward transition and trigger owner are settled in 6d, then limit the open gap to authority, hysteresis, thresholds, and immediate safe state +MAJOR | high | §10 documented scars and harness redirect versus §7 leaves 2b–2c and the tracked P8 story | the text still calls real-time cost visibility our P8 and waits for P8 to measure harness tokens, although P8 explicitly permits no instrumentation or new state and reads only ledger and git; live counters, traces, and per-session token measurement belong to 2c | the cost control and harness-router prerequisite can wait forever for data their named story is forbidden to collect | replace both P8 attributions with 2c telemetry and state the exact counter or trace field each later decision consumes +MINOR | high | §11 dashboard open question | the dashboard is said to build on P8 trace and analytics, but 2b P8 has neither traces nor run analytics; those were deliberately split into 2c | the dependency label points the parked design session at the wrong artifact and can make it widen P8 again | name 2c trace and analytics plus 2b's read-only ledger metrics as separate inputs +MAJOR | high | §11 planning-artifact-home gap versus CLAUDE.md §5 and the tracked docs/superpowers workflow | the gap requires plans and specs never to contaminate an implementation branch, while the shipped workflow tracks those artifacts in the repository and requires a behavior-changing Gate-B fix to update its spec in the same commit | taken literally, the proposed home makes the shipped same-commit traceability rule impossible; taken loosely, contaminate has no checkable meaning | define contamination as transient run output or uncommitted scratch state, while keeping durable specs and required deltas tracked on the implementation branch, or name the deliberate shipped-rule amendment +MAJOR | high | §10 rollup-branch rationale versus §9 bounded merge and E2E claims | the text says per-candidate smoke plus wave-close E2E gives the same integration checkpoint as a rollup branch, but smoke is incremental before each merge and wave-close E2E runs after changes are already on main | the rationale overclaims equivalence and hides the loss of a pre-main, whole-wave, atomic integration point that §9 itself admits | state the exact weaker replacement: incremental composed-candidate smoke before each landing plus post-merge wave detection, with no atomic whole-wave checkpoint +MAJOR | high | §10 main-whole-waves release-tag claim | release tags can mark completed waves but cannot make main contain only whole waves because the merge queue has already landed individual candidates throughout the wave | a project following the stated shortcut still exposes partial waves on main, so the claimed solution does not satisfy the named requirement | say tags provide release boundaries while main remains incremental, and leave projects requiring whole-wave main outside this design rather than claiming tags enforce it +MAJOR | high | §9 smoke gate and §11 merge-queue terminal paths | the queue defines green and red but not missing or unverified smoke command, timeout, cancellation, signal termination, malformed output, infrastructure failure, or an inconclusive result | an automatic merge implementation can treat not-red as green and violate the bounded claim that a failing configured smoke never lands | require an explicit successful exit from a verified smoke command as the only green state and route every absent, failed, timed-out, or indeterminate state to a named non-landing artifact and retry policy in 5b +MAJOR | medium | §7 leaf 5b and §11 merge-queue terminal paths | the recorded queue gaps omit duplicate enqueue, worker crash, and restart windows before smoke, between smoke and landing, and after landing but before queue state is persisted | redelivery can re-run side effects, lose a green candidate, or return an already-merged branch to a lane while the serial queue still appears healthy | assign 5b durable candidate identity, idempotent enqueue, fenced queue ownership, and recovery states for every crash boundary +MAJOR | high | §7 leaves 5a and 6a plus §10 run lock | the run lock is defined per orchestrator tick and its gap discusses double-firing clocks, but 6a adds at-least-once event wakes that can overlap a scheduled tick or fan out when a batch or wave marks many stories freigegeben | stage 4 can start concurrent orchestrators despite the stage-3 exclusion rule, or queue many redundant sessions that race on pool and lane state | make scheduled and event wakes share one lease and fencing domain, and give 6a event deduplication, coalescing, and redelivery semantics +MINOR | high | §7 heading | the heading still says each step equals one story even though steps 2, 4, 5, and 6 are ordering groups whose lettered leaves are the stories | the stale count unit contradicts the corrected preamble and keeps step references ambiguous | change the heading to say each leaf equals one story and numbered entries are ordering groups +MINOR | high | §9 Phase 0 mini-wave claim versus the shipped Gate-B triviality skip and §11 triviality gap | a single story is said to skip no stage, but a behaviorally trivial eligible change may skip Gate B today and the end-state rule for that skip is explicitly still open | the totality claim denies an existing path and masks the exact unresolved branch §11 records | say the views collapse to one line without removing any otherwise-applicable stage; preserve explicit shipped skips and future decisions +MAJOR | high | §1 PR station versus docs/pr-review-bots.md and process-pr-review | the pipeline says to wait for the routed review bots, but the authoritative Wait for list is empty and both enabled findings bots are opportunistic and must never block; the shipped command waits only for bots in Wait for | an implementation following the overview can hang forever waiting for Greptile or CodeRabbit and contradict the station it claims already exists | say apply the routing table: wait only for Wait for bots, currently none, and process opportunistic posts present at the time plus later follow-ups +MAJOR | high | §10 blocking-hook interim enforcement | before the invariant amendment, a repo-owned command or required CI check is called enforcement for protected pool status, lanes, and architecture writes, but neither prevents an agent from editing or consuming unauthorized local state during the running factory | the system can schedule or build from a forbidden mutation long before CI rejects a later commit, so mechanical write protection is being claimed where only instruction or delayed detection exists | call the interim mechanisms authorized-write tooling and merge-time detection, state their bypass window, and reserve write protection for the blocking mechanism after its invariant decision +MAJOR | high | §10 minimum review cost or duration and §7 leaf 2c | a below-floor cost or duration is made a new INCOMPLETE cause, but 2c is defined as telemetry and instrumentation and owns neither the threshold policy nor the required changes to both shipped gate-prompt copies; current INCOMPLETE semantics do not inspect cost or duration | instrumentation can ship without the authorization rule, or a telemetry story can silently change gate closure and trigger invariants 11 and 12 without that scope being visible | split measurement from policy: let 2c emit attributable measurements and give a prompt-rule leaf the thresholds, trust model, missing-data behavior, mirrored edits, and version bump +MAJOR | medium | §10 as-built view and §11 wave-close ordering versus §7 leaves 4d and 5a–5d | the wave-close as-built view is assigned only to coarse step 4/5 brackets and is absent from every leaf definition and the fenced pipeline | wave closing, E2E, and automatic next-wave opening can all be implemented while the promised code-grounded view is never generated | assign generation and consumption to concrete leaves, such as 5d for producing the view and 4d for requiring it in the wave-close transition, and show its edge at wave close +MINOR | medium | §7 final sentence of step 6 | merge stays human until 6c ships and holds, but holds has no named evidence, window, success threshold, or owner distinct from 6d's later promotion evidence | the level-0 automatic-merge go-live condition is not checkable, so a future story can enable it immediately or wait indefinitely | define holds as an explicit 6c validation condition or replace it with the precise prerequisite already intended +END OF FINDINGS (32 total) diff --git a/.context/codex-reviews/gate-a-spec-vision-pass-4-dispositions.md b/.context/codex-reviews/gate-a-spec-vision-pass-4-dispositions.md new file mode 100644 index 0000000..8bab75f --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-vision-pass-4-dispositions.md @@ -0,0 +1,85 @@ +# Gate A (spec) — dark-factory vision — pass 4 dispositions +26 findings: 2 BLOCKER, 14 MAJOR, 9 MINOR, 1 NIT. Unprofiled run. + +## STATUS: gate NOT closed. +The author's closing condition was "if the pass returns ONLY leaf-ownership +findings, record them in §11 and close; a NEW Blocker/Major that is not +leaf-ownership stops the loop". Pass 4 returned two non-leaf-ownership +Blockers and seven non-leaf-ownership Majors, so the condition is not met and +the close decision goes back to the author. Everything fixable was fixed; no +pass is credited clean. + +## The three lines +1. TREND — p1: 52 findings, 2 Blocker, 41 Blocker+Major. + p2: 34 / 1 / 26. p3: 32 / 2 / 24. p4: 26 / 2 / 16. +2. CLUSTER — three groups. (a) Leaf ownership, 7 of 16 Blocker+Major: a station + §1 draws that no §7 leaf owns. (b) The overclaim class AGENTS.md forbids by + name, 5 findings: rollup-branch equivalence, release tags solving whole-wave + main, "enforcement" for the interim write path, "no stage is skipped", and + "an audit finding is by definition something both gates let through". + (c) Two Blockers that are implementation errors of the pass-3 ruling, not + new questions. +3. REQUIRE-WITHDRAW — none this pass. The pass-3 pair did not recur: pass 4 + corrects WHERE the pass-3 ruling was applied, never asks to withdraw it. + +## Tells +Findings rising: NO (52 -> 34 -> 32 -> 26). +Blocker count failing to fall: YES (2 -> 1 -> 2 -> 2) — still present. +Instrument cluster: N/A, no test instrument in this artifact. +Prose-about cluster: arguable; the overclaim findings are about the product's +own claims, which is this document's substance rather than commentary on it. +Require-withdraw pair: NO — cleared this pass. +One tell, down from two, so §5's mandatory two-tell stop no longer fires. The +stop here is the author's own closing condition, not that rule. + +## Fixed — the two Blockers, both my own errors implementing the pass-3 ruling +- p4-1 The ruling ("work waits") was written as "the candidate waits in the + merge queue", but Gate B runs at Verify, before PR, Sample-Gate and the + queue, and a Gate-A outage precedes any candidate. The wait now happens + where the outage happens. The ruling is unchanged; only its placement was + wrong. +- p4-2 I called `.context/codex-gate.off` a declared project-level gateless + state. Shipped CLAUDE.md says it suppresses the hook reminder per workspace + and "the gates still apply", and .gitignore excludes it so a clone never + sees it. The declaration is the tracked INACTIVE notice /workflow-init + writes into the project's CLAUDE.md. Corrected, and neither file is said to + authorize anything. + +## Fixed — seven non-leaf-ownership Majors +p4-3 one producer of the architecture verdict (Intake asks, Bewertungs-Loop +answers, edge drawn) · p4-4 the Spec-Loop no longer intakes a second time after +classification already did · p4-6 the orchestrator named as a model node the +cross-cutting contracts cover · p4-7 story stages are serial; only work inside +a stage fans out · p4-21 the rollup replacement named as weaker instead of +equivalent · p4-22 release tags give boundaries, not a whole-wave main · +p4-23 the interim path is authorized-write tooling and merge-time detection, +with its bypass window stated, not "enforcement". + +## Fixed — nine Minors and the Nit +p4-8 §7 heading names the leaf as the unit [SEE CORRECTION BELOW] · p4-16 6d owns the downgrade +triggers the entry called unowned · p4-17 the §11 opening no longer states a +false count · p4-19 the dashboard builds on 2c, not read-only P8 · p4-20 a +mini-wave bypasses no station, but the triviality skip still exists · +p4-24 an audit finding is not by definition past both gates · p4-25 decision 8 +no longer says "every mandatory touchpoint" one sentence before listing four +that carry no knob · p4-26 step 5's title admits 5b is not a clock loop. + +## Recorded in §11, per the author's scope decision (7 leaf-ownership items) +Leaf definitions extended: 4c gained the Bewertungs-Loop runtime (p4-5); +4e gained given/when/then normalization and the per-AC test report (p4-15); +the ledger entry routes gate findings to 2a rather than 4b, which is +pre-production and never sees a review loop (p4-18). +New "Stations drawn in §1 that no leaf yet owns" block, six entries with a +candidate owner each: the vet preflight, the judge/watchdog, the working +orchestrator dry-run, the as-built view (p4-9, p4-10, p4-11, p4-14), plus +punchlist generation and model strength per role (p4-12, p4-13). + +## CORRECTION, written at close (pass 5) +This file's "Fixed — nine Minors and the Nit" section OVERSTATED what landed. +The second pass-4 edit script aborted on a non-matching string BEFORE its +write, so seven edits it had reported as applied were never written to disk: +p4-8 (§7 heading), p4-16 (downgrade owner), p4-17 (§11 opening count), +p4-19 (dashboard P8 attribution), p4-20 (mini-wave "no stage is skipped"), +p4-25 (decision 8's "every mandatory touchpoint"), p4-26 (step 5's title). +Pass 5 re-reported all seven, which is how the miss surfaced. They were applied +at close. The pass-4 REPORT was wrong; the pass-4 findings file was not. diff --git a/.context/codex-reviews/gate-a-spec-vision-pass-4.md b/.context/codex-reviews/gate-a-spec-vision-pass-4.md new file mode 100644 index 0000000..fa1987a --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-vision-pass-4.md @@ -0,0 +1,27 @@ +BLOCKER | high | §2 decision 5 versus §1 pipeline | a reviewer outage leaves the candidate in Merge-Queue even though Gate B runs at Verify before PR, Sample-Gate, and queue entry, and a Gate-A outage happens before any candidate exists | the decided wait path has no reachable state in the canonical flow, so an implementation must either bypass an unfinished gate to enqueue or wait somewhere the decision does not permit | keep the wait decision but place work at the gate or lane where the outage occurs, or move the independent-review check into the queue and align every pipeline reference +BLOCKER | high | §2 decisions 5 and 8, `.context/codex-gate.off` | the file is described as a declared project-level gateless state, but shipped CLAUDE.md calls it a per-workspace hook opt-out whose presence does not remove the gate obligation, `.gitignore` excludes it, and workflow-init declares degraded mode only by pairing it with a tracked INACTIVE line | a clone loses the marker and starts emitting reminders, while code that treats the marker itself as merge authorization would infer semantics the shipped mechanism explicitly does not provide | distinguish the tracked project declaration from the per-workspace reminder-suppression file and define clone behavior without attributing gate authorization to `codex-gate.off` +MAJOR | high | §1 Intake-Loop and Bewertungs-Loop, §5 dimension 6 | Intake is said to produce and attach the initial architecture verdict while §5 says that verdict comes from the Bewertungs-Loop, which is drawn after Intake with no request or result edge | two fresh-context stations can each expect the other to produce the load-bearing verdict, or Intake can attach a card before its declared source runs | name one producer of the initial verdict and show the artifact edge, reserving the other loop for later re-evaluation if that is the intended split +MAJOR | high | §1 fenced pipeline, Spec-Loop | the new Intake-Loop already classifies and attaches the card before Freigabe, but the later Spec-Loop still includes shipped `intake`, whose job is to create and profile the story before design | an approved pool item can be intaken twice, producing duplicate or conflicting story and profile artifacts and making the canonical station order disagree with decision 4 | remove intake from the post-Freigabe Spec-Loop or explain that this is a different operation with a different artifact name +MAJOR | high | §7 leaves 4b and 4c versus §1 Bewertungs-Loop | no leaf owns the continuous pool-versus-tree evaluation, meta-story batching, or post-architecture-merge reclassification runtime; 4b owns proactive intake and 4c owns only bootstrap plus projection | every defined leaf can finish while the pipeline's Bewertungs-Loop and §6 gap 3 remain unbuilt | add the Bewertungs-Loop contract to a compatible leaf or define a separate leaf for the station +MAJOR | medium | §1 fenced pipeline and §7 step 3 | Takt wakes the model-operated orchestrator, but the orchestrator is not a distinct pipeline node even though Vet applies to model nodes above and the judge supervises the shown model loops | the controller can fall outside the fresh-context, preflight, artifact-handoff, and watchdog contracts while the overview appears complete | show the Orchestrator as a station or state explicitly that it is embedded in Takt and covered by every cross-cutting contract +MAJOR | high | §9 “The tree is the parallelism map” | “Within a story, stages fan out to subagents” contradicts the arrow-ordered pipeline and artifact-only handoffs in which plan depends on spec, build on plan, and Verify on the diff | an orchestrator can parallelize causally dependent stages and review or build incomplete predecessor artifacts | say that story stages remain serial while independent work within a stage may fan out, preserving Gate B's already-parallel review branches as the concrete example +MINOR | high | §7 heading | “each step = one story” contradicts the document preamble and §7 itself, where steps 2, 4, 5, and 6 are ordering groups and only their lettered leaves are stories | step references remain ambiguous and coarse step brackets can be mistaken for valid story ownership | change the heading to “each leaf = one story; numbered entries are ordering groups” +MAJOR | high | §7 step 3 and §10 dry-run adoption | step 3 owns only the execution-plan interface and says the working dry-run lands after 4d and 5a rather than in that story, but no later leaf owns the implementation while §10 still assigns it to step 3 | the ordered roadmap can complete with an interface and no runnable dry-run, or require reopening an already-landed story | reorder the step-3 story after its dependencies or add a later leaf for the working dry-run and repoint §10 +MAJOR | high | §1 Vet preflight, §7 leaves, and §10 Vet adoption | the universal Vet runner before every model-operated node is assigned only to coarse step 4, while 4a–4c divide pool storage, classification, and projection and none owns the cross-node preflight station | all leaf stories can complete while later nodes assume a mechanical preflight that nobody builds | assign the universal runner to a specific leaf or shared-contract leaf and separately map its pool, split-threshold, and projection checks to their component owners +MAJOR | high | §1 Judge/watchdog, §7 leaves 5a–5e, and §10 judge adoption | the liveness sidecar over model loops and the non-clock merge queue is assigned only to coarse step 5 and appears in no leaf definition | every stage-3 leaf can finish without the watchdog, push notifications, or repeated-failure escalation that other sections rely on | create a judge leaf or explicitly extend one leaf with ownership of all supervised nodes and notification outputs +MINOR | high | §10 punchlist adoption versus §7 leaves 6b and 6c | punchlist generation is promised as the artifact sampled audit reads but is assigned only to coarse step 6; 6b owns draw and audit state and 6c owns merge thresholds, with neither owning the punchlist | sampled audit can ship without its promised review input while every leaf remains complete | assign punchlist generation and candidate binding explicitly to 6b, leaving thresholds with 6c +MINOR | high | §10 model-strength adoption and §11 model-strength gap | the feature is assigned to step 1 even though the tracked review-economics story contains no model-strength knob, and step 3's definition also omits it | both cited stories can finish without the adopted project knob or the order needed to evaluate its reviewer-not-weaker rule | remove step 1 as an owner and add the feature to step 3, or create a dedicated leaf +MAJOR | high | §10 as-built view and §7 leaves 4a–5e | the wave-close as-built view is assigned only to coarse steps 4 and 5 and is absent from every leaf definition and from the fenced pipeline | wave closing, E2E, and standing auto-open can all be implemented without generating or consuming the promised code-grounded view | assign generation to a concrete audit or clock leaf and make 4d consume it in the wave-close transition +MAJOR | high | §11 “Given/when/then and the per-AC test report” | the entry proves that step 1 does not contain the requirement and that 4e covers only IDs, immutability, and round-trip, then marks the machine-checkable goal condition unowned while still attaching owner `[4e]` | the section's owner marker falsely classifies an acknowledged roadmap hole as handled | extend 4e's leaf definition to own the normalization and per-AC report, or mark the item genuinely unscheduled and add a leaf +MINOR | high | §11 “Autonomy downgrade” versus §7 leaf 6d | the gap says downgrade triggers are unowned even though 6d explicitly owns downgrade triggers and the entry itself ends with step 6 | later stories cannot tell whether the trigger decision is settled or still missing, and the unowned-question count depends on the answer | say that 6d owns triggers and narrow the open details to authority, hysteresis, thresholds, and immediate safe state +MINOR | high | §11 opening count | the section says exactly three topics lack numbered owners, but later entries explicitly call the machine-checkable goal condition and downgrade details unowned | the mechanically stated count is false under the section's own prose and understates undecomposed work | resolve the contradictory owner markers, then recount or remove the exact count +MAJOR | high | §11 hardening-ledger entry and §7 leaf 4b | accepted Gate A and Gate B findings are assigned to 4b, whose complete scope is pre-production classification and whose station does not participate in either review loop | the classifier story can truthfully close without adding any gate-to-hardening route, leaving one of the named finding sources unrecorded | assign gate findings to the loop-rule or gate prompt owner, or create a dedicated hardening-route leaf; keep sampled-audit findings with 6b +MINOR | high | §11 dashboard open question | the dashboard is said to build on “P8 trace/analytics”, but the tracked P8 story is read-only over the ledger and git and explicitly adds no traces, instrumentation, or run-analytics state | the parked design points at a source forbidden to provide one of its named inputs and can accidentally widen P8 again | name 2c trace and run analytics separately from 2b's passive ledger metrics +MINOR | high | §9 Phase 0 mini-wave claim | “no stage is skipped” conflicts with the shipped Gate-B triviality skip that §11's triviality-gap entry explicitly records | the totality wording denies an existing path and obscures the exact end-state branch still assigned to step 6 | say no otherwise-applicable stage is removed merely because the wave has one story +MAJOR | high | §10 rollup-branch rationale | per-candidate smoke before each incremental landing plus E2E after a wave is already on main is claimed to give the same integration checkpoint as a pre-main rollup branch | the replacement has no atomic whole-wave candidate and detects some integration failures only after partial-wave main has shipped, so the equivalence overstates what the mechanisms prove | state the exact weaker replacement: incremental composed-candidate smoke plus post-merge wave E2E, with no atomic whole-wave checkpoint +MAJOR | high | §10 release-tag shortcut | release tags are said to solve a project requirement that “main = whole waves only” even though the queue lands individual candidates to main throughout a wave | tags can mark release boundaries but cannot retroactively prevent partial waves from existing on main | say tags provide release boundaries while main remains incremental and leave whole-wave-only main outside this design +MAJOR | high | §10 blocking-hook interim path | before the invariant amendment, a repo-owned command or required CI check is called the enforcement for mechanical write protection | CI can reject a later merge and a voluntarily invoked command can detect or mediate a write, but neither generally prevents an agent from editing or consuming unauthorized local state during a running factory | call the interim path authorized-write tooling or merge-time detection, state its bypass window, and reserve write protection for the blocking mechanism after the invariant decision +MINOR | high | §11 hardening-ledger rationale | a sampled-audit finding is said to be “by definition something both gates let through”, while the same section records that Gate B can be legitimately skipped and Gate A reviews spec and plan rather than the final candidate bytes | the rationale overstates what the finding proves and can misclassify an audit escape as evidence against reviews that never inspected that defect | say an audit finding escaped the workflow checks that actually ran, and only call it a Gate-B escape when Gate B reviewed that candidate +MINOR | high | §2 decision 8 | “Every mandatory human touchpoint carries a maturity knob” is immediately contradicted by the four mandatory paths that the next sentence says carry no knob and stay human at every rung | the universal wording makes the settled exception set internally inconsistent | change the subject to every routine human touchpoint and leave the four permanent paths unchanged +NIT | high | §7 step 5 heading and leaf 5b | the ordering group is titled “Stage 3: clock loops” even though 5b is explicitly a non-clock merge-queue station | the roadmap taxonomy disagrees with its own leaf and makes stage membership harder to audit | rename the group to stage-3 loops and supporting stations, or move 5b under a correctly named ordering group +END OF FINDINGS (26 total) diff --git a/.context/codex-reviews/gate-a-spec-vision-pass-5-dispositions.md b/.context/codex-reviews/gate-a-spec-vision-pass-5-dispositions.md new file mode 100644 index 0000000..07854b6 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-vision-pass-5-dispositions.md @@ -0,0 +1,55 @@ +# Gate A (spec) — dark-factory vision — pass 5 dispositions — CYCLE CLOSED +15 findings: 0 BLOCKER, 7 MAJOR, 7 MINOR, 1 NIT. Unprofiled run. + +## Exit taken +The author's second exit: non-leaf-ownership Majors were present, so the gate +closes as a vision on his scope disposition rather than on a clean pass. + +One correction to that exit's stated rationale, because the record should not +carry a false reason. It anticipated "a document generating non-leaf findings +faster than it sheds them". The data says the opposite: findings 52 -> 34 -> +32 -> 26 -> 15, Blockers 2 -> 1 -> 2 -> 2 -> 0. The loop converged. The gate +closes because leaf mechanics are out of scope, not because the loop failed. + +## The three lines +1. TREND — p1 52/2/41 · p2 34/1/26 · p3 32/2/24 · p4 26/2/16 · p5 15/0/7. +2. CLUSTER — (a) §11's own entries lagging §7 after pass 4 extended two leaves; + (b) ownership granularity — coarse step numbers where the document makes the + leaf the unit; (c) two factual errors about tracked repo artifacts, both + mine. +3. REQUIRE-WITHDRAW — none. Cleared in pass 4 and did not return. + +## Tells +Findings rising: NO. Blocker failing to fall: NO — cleared this pass, 2 -> 0. +Instrument cluster: N/A. Prose-about cluster: partly, and by now expected — the +remaining findings are about the record rather than the design. +Require-withdraw: NO. Zero tells at close, down from two at pass 3. + +## Fixed — one deliberate exception to "fix nothing further" +Two findings were factually false statements about tracked repository files +that I introduced in earlier passes. Closing while asserting them would have +published something untrue about the repo, which the exit rule did not +contemplate. Both corrections are narrow and neither changes what the author +decided. +- p5-2 (MAJOR) Decision 8 said the same-family tier-2 question was "closed with + a negative answer". It is not. What closed with a negative answer was a safe + *sanctioned zero-pass closure*. A same-family tier-2 reviewer is a different + question and is still open as a tracked, unshipped story + (docs/superpowers/stories/2026-08-14-tier-2-same-family-reviewer-story.md); + the fallback design's own §7 routes to it — "a weaker review is still a + review" — if its containment proves buildable. My pass-1 fix corrected a + false claim that the kit ships tier-2 and overcorrected into declaring it + dead. Decision 8's substance is untouched: another family, or none. +- p5-1 (MAJOR) "it is git-ignored so a clone never sees it" is true of this kit + repo only. /workflow-init has target projects ignore /.context/codex-reviews/ + specifically, so in a scaffolded project the codex-gate.off marker is visible + and committable. Narrowed to say which repo is which. + +## Recorded, unresolved at close (13) +Appended to §11 as a dated "Unresolved at close" block, each with its leaf: +labeled examples have no owner [unowned]; the orchestrator product is wider +than step 3 [3]; "this closes the parallelism blind spot" overstates smoke +[5b/5c]; two §11 entries lag §7 after 4e and 6d were extended [4e, 6d]; +"promotion evidence" names one owner for knobs belonging to three [4d, 6b, 6d]; +four top-level questions still cite coarse steps [6b, 4a, 5a, 5c]; three known +wordings kept as they are [editorial]. diff --git a/.context/codex-reviews/gate-a-spec-vision-pass-5.md b/.context/codex-reviews/gate-a-spec-vision-pass-5.md new file mode 100644 index 0000000..143b740 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-vision-pass-5.md @@ -0,0 +1,16 @@ +MAJOR | high | §2 decision 5, `.context/codex-gate.off` | the claim that the marker is git-ignored so a clone never sees it is true only in this kit repository; `/workflow-init` writes the marker in target projects but tells them to ignore only `/.context/codex-reviews/`, not `codex-gate.off` | a target project can accidentally commit and clone the supposedly per-workspace switch, so the shipped-behaviour claim and clone semantics are false | make `/workflow-init` merge an explicit ignore rule for `/.context/codex-gate.off` while preserving the tracked `codex-gate.on`, or narrow the vision and record the missing ignore rule as owned work +MAJOR | high | §2 decision 8, lines 183–190 | the cited fallback design closed only safe sanctioned zero-pass closure, not the same-family tier-2 question; its §7 explicitly routes the stall to the still-tracked `2026-08-14-tier-2-same-family-reviewer-story.md` if containment proves buildable | the vision calls a live tracked design question closed and says reopening it would need a story even though that story already exists, leaving contradictory project truth about whether tier 2 is allowed to proceed | either close or supersede the tier-2 story with old-condition accounting, or narrow decision 8 to the zero-pass result and explicitly disposition the separate contained-tier-2 question +MINOR | high | §2 decision 8, “Every mandatory human touchpoint” | the sentence says every mandatory human touchpoint has a maturity knob, then immediately says four mandatory paths have no knob and stay human at every rung | the universal rule and its exception set cannot both guide later leaf specs, risking knobs on permanent stops or inconsistent interpretations | change the subject to every routine human touchpoint and keep the four permanent paths unchanged +MAJOR | medium | §1 human overrides and §2 decision 7 Freigabe feedback | human overrides and repeatedly difficult Freigabe decisions are required to become labeled examples that tighten rules or intake, but no §7 leaf owns capturing, storing, routing, or consuming those examples and §11 does not record the gap | the roadmap can complete while the promised learning feedback path does not exist, so adverse human evidence never changes the factory | assign the labeled-example feedback contract to explicit leaves such as 2a and 4b, or state that it is manual practice rather than a build-path capability +MAJOR | high | §6 gap 4 versus §7 step 3 | gap 4 defines the orchestrator product as question routing, artifact handoff, a prediction ledger, and batched human decisions, while step 3 owns only the role, artifact handoff, and execution-plan interface | step 3 can close without three capabilities that the current-state gap says are part of the orchestrator, leaving the decomposition incomplete | add question routing, the prediction ledger, and batched human-decision handling to step 3 or remove them from the vision after an explicit decision +MINOR | high | §7 heading | “each step = one story” contradicts the preamble and §7 itself, where steps 2, 4, 5, and 6 are ordering groups and their lettered leaves are the story owners | step-level references can be mistaken for valid leaf ownership, undermining the document’s rule for what later becomes a story | change the heading to “each leaf = one story; numbered entries may be ordering groups” +NIT | high | §7 step 5 heading | the group is titled “Stage 3: clock loops” although 5b is explicitly a non-clock merge-queue station | the roadmap taxonomy contradicts its own leaf and obscures whether supporting stations belong to the group | rename the group to “Stage 3 loops and supporting stations” or equivalent +MINOR | high | §9 Phase 0 mini-wave | “no stage is skipped” still contradicts the shipped Gate-B triviality skip and the Sample-Gate path where an undrawn PR bypasses Audit | the totality claim hides real branches that the same document assigns to step 6 and reintroduces the overclaim pass 4 said was removed | say only that one-item roadmap and execution-plan views collapse to one line; do not claim that every stage executes +MAJOR | high | §9 “This closes the parallelism blind spot” | per-candidate smoke proves only that the configured smoke command passed on the composed candidate; it does not close the broader case where independently green lanes compose incorrectly outside smoke coverage, which the later E2E loop may discover only after merge | downstream stories can treat cross-lane integration as solved and under-specify E2E or smoke coverage, overstating exactly what the queue proves | replace “closes” with the bounded claim that the queue adds a pre-merge check over the configured smoke surface and leaves broader integration detection to post-merge E2E +MINOR | high | §11 process dashboard open question | the dashboard is said to build on “P8 trace/analytics”, but §7 and the tracked P8 story explicitly keep P8 read-only over the ledger and git with no traces or run-analytics instrumentation | the parked dashboard depends on a source that is forbidden to produce its named inputs and can accidentally widen 2b | name 2c as the trace and run-analytics source, keeping P8 as a separate passive-metrics input if needed +MINOR | high | §11 “Autonomy downgrade” | the entry says downgrade triggers are unowned even though §7 leaf 6d explicitly owns “its downgrade triggers” | later stories cannot tell whether triggers are already assigned or still missing, and the entry falsely classifies handled scope as open | state that 6d owns triggers and limit the open part to any genuinely unowned authority, hysteresis, and immediate-safe-state decisions +MAJOR | high | §10 acceptance-criteria adoption and §11 “Given/when/then and the per-AC test report” | §7 leaf 4e now explicitly owns given/when/then normalization and the per-AC test report, but §11 still says 4e covers only IDs, immutability, and round-trip and declares the goal condition unowned; §10 also retains the stale owner `[step 1/4]` | the same requirement is simultaneously owned and unowned, so the leaf can be reopened unnecessarily or another story can duplicate it | remove the stale §11 entry and change §10’s owner to `[4e]` +MINOR | high | §11 “Stations drawn in §1 that no leaf yet owns” | the block is not limited to stations drawn in §1: the as-built entry explicitly says it is not drawn there, and punchlists, model strength, AC canonicalization, stopped-lane policy, promotion evidence, and wave-close ordering are concerns rather than drawn stations | the heading makes the station-to-leaf coverage audit mechanically false and hides which entries are pipeline stations versus cross-cutting gaps | split the actual drawn stations from other unowned concerns or rename the block to cover both without claiming they are all drawn in §1 +MAJOR | high | §11 “Promotion evidence” owner `[6d]` | the entry requires outcome and adverse-event evidence for every autonomy knob, but 6d is defined only as promotion of automatic merge from level 0 to standard and high stories; Freigabe and wave-opening knobs belong to 4d and the sample-rate knob belongs to 6b | those knobs can ship or widen without the evidence contract even though the document says evidence is a precondition on every one | map evidence ownership to 4d, 6b, and 6d respectively, or explicitly widen 6d into a cross-cutting promotion-governance leaf +MINOR | high | §11 top-level open-question owners | the Sample-Gate rule, pool storage, clock platform configuration, and E2E cadence cite coarse ordering groups `step 6`, `step 4`, and `step 5` even though this document makes leaves the unit of ownership | recording mechanics under a non-story group does not assign them to a later Gate-A artifact, so the gaps can be lost when the epics split | repoint them to 6b, 4a, 5a, and 5c respectively, leaving only the explicitly parked dashboard portion unowned +END OF FINDINGS (15 total) diff --git a/.context/codex-reviews/gate-b-0.8.0-dispositions.md b/.context/codex-reviews/gate-b-0.8.0-dispositions.md new file mode 100644 index 0000000..bb6e7b7 --- /dev/null +++ b/.context/codex-reviews/gate-b-0.8.0-dispositions.md @@ -0,0 +1,214 @@ +# Gate-B 0.8.0 — triage of the ten open Blocker/Major, into exactly two bins + +Cycle-level record. Supersedes the pass-1-only companion; it is cycle-qualified because +the conventional `-dispositions.md` name is occupied by the 26 Jul cycle. + +Rule applied: **fix now everything cheap and in-diff; disposition formally only what has a +named home outside this PR** — a `todos.md` row with a trigger, or a follow-up story. One +line each. + +--- + +## BIN 1 — FIXED IN THIS DIFF (7) + +| # | Finding | What was done | +|---|---|---| +| 1 | **P9-12 · unroutable flush + unescaped event** (SPEC-2, QUALITY-4) | **A correctness bug in this diff, not an inheritance.** Reproduced: with a pending disclosure, an event named `Bogus\` emitted `"hookEventName":"Bogus\","` — the backslash escapes the closing quote and Claude Code receives **invalid JSON**. Two new defects combined: `flush_notes` ran unconditionally so an unroutable event reached `emit`, and `$event` was interpolated raw beside two escaped fields. Both fixed; 5-row oracle in both runners, written first, failed first. | +| 2 | **QUALITY-7 · composed order contradicts spec §6** | Spec §6 is explicit — *"disclosure first, then the per-occurrence message"* — and the code appended the disclosure instead, with the golden freezing the inversion. Added `note_front`; goldens updated. The implementation drifted, so the implementation moved. | +| 3 | **QUALITY-1 (residual) · anchor-prefixed near miss** | The pass-1 fix guarded `${b%%\n*}` on the anchor prefix, which left it running in full for a block that *does* start with the anchor — 5.9 s at 150 KB, 23.1 s at 300 KB. Head now bounded to 4096 chars. Two oracle rows (both `\n` placements), written first, failed first. One stated limit: a ~4070+ character tool name counts instead of being discarded. | +| 4 | **P9-2 / P9-4 · plan-contract drift** (SPEC-7, QUALITY-8/9) | Docs-drift class, cost sentences. A2's comma rule is now scoped to the walked prefix, naming the trailing-comma-after-selected-element as never seen rather than refused; A3 now says **four** whitespace sites, naming `_token_ends` as the fourth. | +| 5 | **QUALITY-11 / SPEC-9 · unknown-tool role split** | A prompt fix inside a PR that already reviews prompts. `additionalContext` now carries only model-owned next steps (report it, treat passes as uncounted); install / `.mcp.json` / mapping / namespace all moved to `systemMessage`, per spec §6's field split. Goldens updated. | +| 6 | **SPEC-9 / QUALITY-3 · rollback hard stop** | Recorded as an explicit **amendment to plan Step 5**, not left as commit prose. Names what was established (cache is version-keyed, `installed_plugins.json` selects, 0.7.1 bytes match the resolved commit), what was not (no switch executed after activating 0.8.0), why executing it is the wrong trade here (mutates a shared plugin environment; 0.7.1 is already active), and reinstates the hard stop for any later release shipping with 0.8.0 active. | +| 7 | **P9-6 · escaped `type` value** (SPEC-1, QUALITY-3) | Resolved the way Gate A's own disposition allowed — *state it, add a fixture*. Boundary now stated in spec §3.1 as a known gap in that section's own contract, pinned by a characterization row asserting `no-result`, and carried in `todos.md` with a trigger. Fail-closed, so a real result is discarded rather than miscounted; still a wrong verdict on a legal payload, and the spec says so. | + +## BIN 2 — DISPOSITIONED, each with a named home (4) + +| # | Finding | Home | +|---|---|---| +| 8 | **QUALITY-2 / SPEC-11 · `skipval` walks containers character-by-character** — 3.2 s at 200 KB, 11.5 s at 400 KB; only the 1 Mi-unit ceiling stops it | `todos.md` → *"Locator: `skipval` walks containers one character at a time"*. Trigger: a slow-hook report, or any change raising the ceiling. Three candidate fixes, all design calls — a patch would be picking one silently. | +| 9 | **SPEC-3 / QUALITY-5 · A5 marker matrix breadth** (single emitter pair, missing cross-family failure combinations, P9-9's named skip) | `todos.md` → *"A5 marker matrix and A6 composition coverage are narrower than the approved plan"*. Trigger: a disclosure/advice bug the current rows miss. Needs a selective `rm` shim. | +| 10 | **SPEC-4 / QUALITY-6 · A6 composition breadth** (not exact-tested against every emitting branch) | Same `todos.md` row as #9 — one row, because the two share a fixture-and-shim build-out. | +| 11 | **QUALITY-12 · `hardening-log.md:26` teaches the old mechanism** | `todos.md` → *"The hardening ledger has no supersession convention"*. The ledger's header forbids editing a row **and** limits rows to one per hardening, so there is no sanctioned in-diff move — that missing convention is the actual defect. Trigger: the next row falsified by a later change; this is the second. | + +## Minors — collected, never iterated + +SPEC-5, SPEC-6, QUALITY-5, QUALITY-6, QUALITY-7, QUALITY-13 … QUALITY-19. Two are worth a +line because they are honesty items rather than coverage items: QUALITY-18/19 concern the +fixture README's provenance wording (isolation, and "edited by hand … exactly seven field +values and nothing else" sitting beside a description of mechanical replacement). Neither +changes behaviour; both are candidates for the next prose pass. + +## Closure rule agreed for this cycle + +If pass 3's only Blocker/Major are re-raises of the four dispositions above, the cycle +closes **clean-with-dispositions**, and the closure record names each dismissed finding +with its home. Any *new* Blocker/Major: fix and continue, floor unchanged. + +--- + +# Pass 3 — VALID, and it was not clean: seven NEW Blocker/Major + +Both branch files valid (spec 15, quality 15; terminators present, counts matched, no extra +lines, distinct branch-appropriate content). **The reply carried a spurious extra line** — +the quality branch reported `gate-b-spec | pass 3 | 16 findings` alongside its own — which +is the write-race shape §5 warns about. Checked rather than assumed: the two files are not +identical, each holds findings of its own branch's flavour, and the spec file's mtime is the +later of the two. Treated as a mis-report, not a race; the pass stands. + +Per the closure rule, the four dispositions were re-raised as expected and cost nothing. But +pass 3 also raised **seven new Blocker/Major**, so the rule's second clause fired: fix and +continue, floor unchanged. + +## New in pass 3 — all fixed + +| Finding | Verified how | Fix | +|---|---|---| +| **Reserved-name mapping hijack** (QUALITY-1) | Reproduced: `reviewTool=Bash` made a `git commit` **count** a Gate-B pass instead of resetting the cycle; `execTool=Skill` counted a skill invocation as Gate A. Mapped cases precede the native cases, and those are the only two out-of-namespace names the matcher delivers. | Parser now requires `mcp__codex__*`; three rows, including one proving a legitimate in-namespace mapping still counts. **This diff introduced the contract it contradicted**, so it is in-scope, not inherited. | +| **Second quadratic: the record accumulator** (SPEC-1, QUALITY-3) | Reproduced: pretty-printed payloads cost 0.35 s at 4k lines, 0.75 s at 8k, 2.69 s at 16k — independent of `skipval`. | The *finding* was that plan/CHANGELOG/evidence named only `skipval`. Both paths are now named in all four places and the `todos.md` row covers both. Chunked accumulation reduces but does not remove it, and the two paths share a fix only if the scan stops indexing with `substr` — so the deferral is extended, not invented. | +| **4096 bound narrows a normative class** (SPEC-2, QUALITY-2) | The bound was documented in the plan and CHANGELOG but not in **spec §4**, which still said the anchor covers any mapped name at any length. | Stated in spec §4 with its reason and its reinstatement condition, and pinned by two rows: a genuine notice immediately inside the cutoff (→ `backgrounded`) and immediately outside (→ `unrecognized`). | +| **The plan's battery omits `HOOK_SH`** (SPEC-3, QUALITY-4) | True: `AGENTS.md` was updated at pass 2, the plan's own copy was not — so the plan still prescribed the exact run that produced the false dash claim. | Both the fenced battery and the dash bullet now set `HOOK_SH`, with the reason. | +| **A1 still describes `readstr` as an `esc`-flag loop** (SPEC-4, QUALITY-15) | Plan-internal contradiction created by the pass-1 rewrite. | A1 now states the parity rule as the contract and the anchored `match()` as the mechanism, noting the language and returned bytes are unchanged. | +| **The plan calls its ten message strings final literal text** (SPEC-5, QUALITY-5) | Six differ from the shipped strings after P9-28…37. | The plan now says it is **not** their source of truth and names which findings moved them; `codex-gate.sh` and the literal goldens are authoritative. | +| **Evidence reports 437/437** (SPEC-6, QUALITY-6) | True and stale — the suite is now **451/451, 0 failures, 1 named skip**, under both shells. The implementer claim of 449 was also wrong. | Corrected, and the base-hook counterfactual totals are now labelled as a record of that run rather than of the current row set. | + +## Dismissed with reason + +- **"The suite leaks an untracked repo-root file `0` containing `400000`."** Not reproducible + and not ours: a clean-status check before and after a full run under **both** shells shows + no change, and no such file exists. It was created by the reviewer's own ad-hoc probing + during the pass. Worth recording because a suite that dirties the checkout would move the + Gate-B fingerprint (invariant 3) — so this was checked rather than waved off. + +--- + +# The bounded sweep (2026-08-03) — three claims, full inventory, one home each + +Authorized after pass 6 widened instead of narrowing. The diagnosis was that three *claims* +were restated across many artifacts and each pass reached a copy the last had not. The fix +is inventory-first, all sites together, one normative home with pointers — not another pass. + +## Claim 1 — the namespace boundary + +**Normative home:** the design's **decision 1** (`…failed-codex-call-counts-as-a-pass-design.md`), +now carrying the sentence *"This paragraph is the single normative statement of the namespace +boundary; every other mention points here rather than restating it."* + +| Site | State before the sweep | +|---|---| +| design, decision 1 | correct, but ended in a broken fragment `…), and always could not.` left by the pass-4 edit — **repaired, and made the home** | +| plan, Task 6 Step 4 insertion (×2) | corrected at pass 5 | +| `workflow-init.md` cause-2 remedy | *"the hook never fires on them at all"* — **fixed** | +| `workflow-init.md` preflight sample output | *"where the hook never fires"* — **fixed** | +| `workflow-init.md` inline `CLAUDE.md` template | correct; **gained a pointer** to decision 1 | +| `codex-gate.sh` parser comments (×2) | *"a tool that never fires"* — **fixed** | +| `codex-gate.sh` namespace guard comment | correct (pass 3) | +| unknown-tool `systemMessage` + its golden | correct (pass 4) | +| `README.md`, `CHANGELOG.md` | correct (pass 4) | +| `hooks.json` matcher, suite comments | statements of fact, not the claim | + +The claim is precise only when it says **both** halves: outside `mcp__codex__*` a name is +either never delivered, **or** — for the reserved `Bash`/`Skill` the matcher does deliver — +hijacks that lifecycle event. The parser refuses both. + +## Claim 2 — the trust boundary for third-party tools + +**Normative home:** the design's §3.1 *"Why fail-closed is right here"* paragraph. + +| Site | State before the sweep | +|---|---| +| design §3.1 rationale | *"a **mapped** third-party tool may legitimately return…"* — **fixed** | +| design, decision text | same wording — **fixed** | +| design §6 `no-result` diagnosis | mapped-only — **fixed**, now names both routes | +| design §6 `unrecognized` cause list | mapped-only — **fixed** | +| `UNVERIFIED_MSG` in hook **and** its golden | *".context/codex-gate.tools names it, and unmapping it removes the gate"* — **fixed**: the file names it *only if mapped*, and a server registered under the default `codex` name reaches the gates with no mapping | +| `NORESULT_CTX`/`MSG` | corrected at P9-33 | +| story `no-result` criterion | corrected at pass 5 | + +The precise form: a third-party tool reaches the gates **either** through a mapping **or** +as a server registered under the default `codex` name, so an absent mapping does not rule +it out — which is exactly what decision 1's own remedy creates. + +## Claim 3 — when a pending disclosure is flushed + +**Normative home:** the design's §5.2 state-transition table. + +| Site | State before the sweep | +|---|---| +| design §5.2 transition table | *"pending \| any unsuppressed hook event"* — **fixed** to *routed* hook event | +| design §5.2 rationale | *"an unrelated later event has nothing to flush"* — **fixed** | +| design §6 composition | *"would collide again on the next event"* — **fixed** | +| plan A5 marker table | *"present \| present \| any event"* — **fixed** | +| `codex-gate.sh` routing gate comment | already correct (the P9-12 fix) | + +The P9-12 fix was right; only its restatements lagged. Since that fix, `flush_notes` runs +**only for a routed event** — an unroutable payload preserves the debt for a later routed one. + +## Also in this sweep + +- **Malformed-routing narrowed** in the design. It said such a payload "cannot route"; that + holds **with** `jq`, and **without** it the `grep` fallback can read a tool name out of a + document malformed elsewhere, so the same payload routes and classifies — normally landing + in `unrecognized`, counted and disclosed. The divergence is `field()`'s and predates this + change; the plan already said so and the design now does too. +- **The seeded-state TIMEOUT rows are restored**, not re-dispositioned. `timeout` (the + captured `shape2-executor-timeout` envelope) now runs the same seeded matrix as the other + discarded classes — 16 rows across both gates, both name sources, both runners. The plan + required a seeded timeout preservation row; a classification-only row had been substituted + without a kept/moved/dropped disposition, which is an AGENTS.md Don't. + +Suite: **451 → 467** assertions, 0 failures, 1 named skip, under both shells. + +--- + +# CLOSURE — clean-with-dispositions (Gate-B pass 9, 2026-08-03) + +**Nine passes. Pass 9 returned ZERO Bucket C — nothing new — on both branches.** + +- **Spec branch: 4 findings, all Bucket A.** Dispositions-only, which is the pinned exit. +- **Quality branch: 22 findings — 4 Bucket A, 1 Bucket B, 0 Bucket C.** The Bucket B was a + one-clause over-generalization in the namespace home ("the plan points here rather than + restating it") — Task 6 necessarily quotes the literal text it authors for a shipped + surface. Narrowed after the pass; **docs-only, no behaviour, prompt, or test touched.** + Recorded rather than glossed: it was not re-reviewed. + +Both branches independently reproduced the evidence rather than reading it: **467 passes, +0 failures, 1 named skip under `sh` AND under `dash` separately**, invariants 123/123, +version-bump 36/36 against BASE, strict plugin validation — all from their own checkout. + +## The four dismissed findings, each with its home + +| # | Finding | Home | Trigger | +|---|---|---|---| +| 1 | **Two quadratic locator paths** — `skipval` walks containers one character at a time (`substr(s,i,1)` is O(len) per call in BWK awk): 3.2 s at 200 KB, 11.5 s at 400 KB. Independently, the record accumulator `s = s $0 "\n"` rebuilds the whole input once per line: 0.35 s at 4k lines, 2.69 s at 16k. Only the 1 Mi-unit ceiling stops either. | `todos.md` → *"Locator: TWO quadratic paths — `skipval`'s container walk and the record accumulator"* | a slow-hook report, or any change that raises the ceiling | +| 2 | **A5 marker-matrix breadth** — runs through one emitter pair, omits mixed disclosure/background-advice write-failure combinations, and carries P9-9's pending-delete row as a *named* skip because one permission governs both operations on `.context/` and a directory at the pending path is not seen as pending. | `todos.md` → *"A5 marker matrix and A6 composition coverage are narrower than the approved plan"* | a disclosure or advice bug the current rows do not catch | +| 3 | **A6 composition breadth** — exact-tested for a carried disclosure with failure, a silent Bash event and one jq-free pair, not against every emitting branch. | same row as #2 — they share a fixture-and-shim build-out | as above | +| 4 | **`docs/hardening-log.md:26` teaches the pre-0.8.0 counting rule.** No sanctioned in-diff move exists: the ledger header forbids editing a row *and* permits one row per hardening. The missing supersession convention is the actual defect. | `todos.md` → *"The hardening ledger has no supersession convention"* | the next row falsified by a later change — this is the second | + +## Why each is a disposition rather than a dodge + +#1's three candidate fixes (jump-based walk, length-scaled work budget, lower ceiling) are +each design calls with different trade-offs; patching one silently would pick for the reader. +#2 and #3 need a selective `rm` fault shim and per-branch composition goldens — a coverage +build-out, not a correction. #4 cannot be done inside this diff without breaking one of the +ledger's own two rules. + +## What the cycle actually produced + +Four defects that would have shipped, each found by review and each fixed with an oracle +written first: a **denial of service** (150 KB → 10.9 s, now 0.46 s, three quadratics), +**invalid emitted JSON** (`"hookEventName":"Bogus\","` from an unroutable event plus one +unescaped field), a **false ✓ configuration hijack** (`reviewTool=Bash` made a `git commit` +*count* a Gate-B pass instead of resetting the cycle), and a **spec-order inversion** frozen +by its own golden. Plus five overclaims in the release evidence, each caught and narrowed. + +## awk portability — now established, on PR #21 + +The one gap the commit body names as outstanding is closed by the PR it points at. +CI run 30794079534 (`quality`) **passed in 1m43s** on `ubuntu`, which supplies **mawk** +where the local battery runs BWK awk, and which runs the hook suite twice — `HOOK_SH=sh` +and `HOOK_SH=dash`. So the locator's `match(/^([^"\\]|\\(.|\n))*"/)`, the bounded +`printf '%.4096s'` head and the whole 467-row matrix are now exercised against a second +awk implementation and a second shell. That is shell **and** awk-implementation coverage, +which 467/467 locally was not. + +PR: https://github.com/dsnger/dev-workflow-kit/pull/21 diff --git a/.context/codex-reviews/gate-b-0.8.0-pass-1-dispositions.md b/.context/codex-reviews/gate-b-0.8.0-pass-1-dispositions.md new file mode 100644 index 0000000..040c1d2 --- /dev/null +++ b/.context/codex-reviews/gate-b-0.8.0-pass-1-dispositions.md @@ -0,0 +1,86 @@ +# Gate-B pass 1 (0.8.0 cycle) — dispositions, spec + quality branches + +Advisory companion. Not a findings file, not part of pass validation. Both branch files +validated: terminator present, counts matched (9 and 12), no extra lines, nothing else. + +**Cycle-qualified filename on purpose.** The conventional slot name +`gate-b-spec-pass-1-dispositions.md` is already occupied by the 26 Jul cycle's record. +Slot names carry no cycle component, so writing there would have destroyed it — the same +collision §5 documents for findings files, reached through the companion instead. + +**21 findings, 16 Blocker/Major. Cycle STOPPED and surfaced to Daniel rather than fixed**, +per his standing instruction for this task: stop rather than squeeze if Gate B opens a new +front. It did. Two findings are design-level and were verified independently here; two more +are defects in prompt text written during this same session. + +## Verified independently before reporting + +- **QUALITY-1 — CONFIRMED, and worse than "a performance concern".** The locator's + character-at-a-time `RAWSTR` concatenation and record-at-a-time `s` accumulation are + super-linear. Measured here, one hook invocation, valid under-ceiling payloads: + 10 KB → 0.2 s · 50 KB → 1.3 s · 150 KB → **10.9 s**. Size ×3 costs time ×8.4. The + advertised backstop is 1 Mi *units*, roughly 7× the 150 KB case, so a perfectly legal + result stalls the synchronous hook for minutes. Codex reported ~900 KB still running at + 30 s from its own probe; the curve measured here agrees. This is not exotic input — a + full Gate-B review result is routinely tens to hundreds of KB. Neither bound bounds + *time*, and the code comments only ever disclaimed bounding *memory*. +- **QUALITY-2 — CONFIRMED, and it falsifies evidence already committed.** Every general + runner invokes the hook as `sh "$HOOK"` / `/bin/sh "$HOOK"` (lines 62, 77-79, 388, 625, + 825-827). On this machine `/bin/sh` is **bash 3.2.57 in sh-mode**, not dash. So + `dash codex-gate.test.sh` runs the *harness* under dash while every hook invocation runs + under bash. Only the ~7 explicit invariant-1 rows (1182-1186, 1720-1725) actually execute + the hook with `dash`. The claim "435 green under sh and dash" — in the CHANGELOG, in the + commit body evidence, and in the Task 5/6 execution-log notes — overstates what was + tested for the other 428 rows. The dash-specific rows are real and did catch the + invariant-1 defect; the generalization from them is what is false. + +## Defects in text written this session + +- **SPEC-7 — VALID.** The P9-30 reason clause added to four shipped prompts says a findings + file left by a failed attempt "is well-formed and correctly terminated, so no later check + can tell it apart". A *partial* write can be malformed and IS caught by the + terminator/count checks. Only a stale, complete, valid-looking file evades provenance. A + fix for an item-11 overclaim introduced a new item-11 overclaim — the four-round pattern + AGENTS.md documents, reproduced inside the change that cites it. +- **QUALITY-12 — VALID.** The rewritten unknown-tool `additionalContext` tells Claude to + install the pinned server and edit `.context/codex-gate.tools`. Spec §6 assigns + operator-only remedies to `systemMessage`, and the hook's own header comment states that + rule. The role split is documented and not implemented in the prompt written here. + +## Standing-lens findings — the diff falsifies live statements elsewhere + +- **QUALITY-8 — VALID.** §5's Mechanics timeout bullet still says an abort "may already + have moved the hook's counter (the pinned server returns its own timeouts as ordinary + results)". A recognized immediate-first timeout no longer moves it. The B1 rewrite fixed + the "What this does not do" paragraph and never reached this one — same file, same + section, different paragraph. +- **QUALITY-9 / QUALITY-10 — VALID.** `todos.md` still carries this defect as open work and + describes the hook as inspecting only the tool name; its preflight item states a failure + mode 0.8.0 falsifies and misstates the C1 residual. +- **QUALITY-11 — VALID.** `docs/superpowers/specs/2026-07-20-codex-file-first-output.md` + still states the hook keys only on tool name and increments after failed reviews. Two + approved specs now describe incompatible behaviour. + +## Known Gate-A deferrals, now confirmed against the implementation + +- **SPEC-1** = P9-6 — escaped `type` value compared as raw bytes, so `text` reads as + `no-result`. +- **SPEC-2** = P9-12 — `flush_notes` runs on unroutable payloads. Codex adds a detail the + Gate-A finding did not: on the jq-free path the unrouted event string is interpolated + into JSON **unescaped**. +- **SPEC-8** = P9-2 + P9-4 — the plan's own contract text disagrees with the shipped + whitespace sites and comma handling. +- **SPEC-3 / SPEC-4 / QUALITY-4** — the A5 marker matrix and A6 composition coverage are + narrower than the plan requires; P9-9's skip is real and still printed by name. + +## Needs Daniel, not a code fix + +- **SPEC-9 / QUALITY-3 — the rollback hard stop.** Daniel chose "document, don't execute" + today. Codex is right that the plan states a release-blocking condition, and right that + the finding's own remedy allows "obtain and record an explicit amendment". That amendment + currently exists only as evidence prose; to satisfy the finding it must be recorded as an + amendment to the plan's Step 5, the way story criterion 10 was amended. + +## Minor — collected, never iterated + +SPEC-5, SPEC-6, QUALITY-5, QUALITY-6, QUALITY-7. diff --git a/.context/codex-reviews/gate-b-fic2-parked-review-economics.md b/.context/codex-reviews/gate-b-fic2-parked-review-economics.md new file mode 100644 index 0000000..e415485 --- /dev/null +++ b/.context/codex-reviews/gate-b-fic2-parked-review-economics.md @@ -0,0 +1,228 @@ +# Parked: opening-evidence package for the review-economics story + +Written 2026-08-26 during Gate-B cycle `fic2` on branch `field-intake-canvas-a1-a5`. +Not a findings file — it participates in no pass validation. It exists because the +material below was produced by a live loop and would otherwise die with the session. + +**Disposition (Daniel, 2026-08-26): revert and park.** The cycle does not absorb the §5 +surgery these findings imply. The assigned fix set contracts back to the eleven +dispositions plus the two loop rules that survived the earlier 12-pass cycle, plus this +cycle's verified error fixes. Everything below moves to a story, unanswered. + +## What was reverted, and why it is not a retreat + +Two clauses were written into CLAUDE.md §5 and the `/workflow-init` template mirror +during this cycle and then taken back out: + +- **Q1 — clean completion outranks the two-tell stop.** A Blocker/Major-free pass at or + above the floor would close even with two or more tells present, the tells going into + the closing status report rather than blocking the close. +- **Q2 — a declined expansion has a defined exit.** A finding the user declines to bring + in-set leaves the cycle as a recorded out-of-scope item, with the Blocker/Major-resolve + duty scoped to in-set findings. + +Both were correct answers to real defects. Pass 1 found the defects; Daniel confirmed both +answers. Pass 2 then found that shipping them requires qualifying **three §5 rules nobody +proposed changing** — the universal Blocker/Major-resolve duty, the rule that a surfaced +finding stays open with resolution unwaived, and the rule that no pass carrying it counts +as clean. That is the scope expansion the cycle declined. + +## The three questions that go to the story unanswered + +- **Q3.** Does a *scope stop* outrank clean completion, or the reverse? "Clean completion + outranks both exits" was written, but §5 has three exits — the scope stop, the + clearly-stuck exit, and the two-tell stop. If scope is included, the sentence + contradicts the rule that even a Minor opening a contract question must stop. +- **Q4.** Does a decline bind *later passes in the same cycle*? As drafted it governed one + stop only, so the same out-of-set Blocker could stop every subsequent pass and re-ask the + same question indefinitely. +- **Q5.** Is qualifying the three universal rules with an in-set boundary acceptable at + all? Yes makes Q2 shippable and is §5 surgery; no means Q2 cannot ship in that form. + +## The instrument is broken, and that is the most reusable finding here + +The `battery+check` evidence for a prose-only change was a decision matrix: N review +states, complete inputs, one expected output each, scored against the old text and the new +text and against both copies independently. Two defects were found in it by Gate B, and +both are properties of the *technique*, not of this instance: + +1. **A state's inputs must include every input the rule reads.** Rows 3 and 10 carried + identical recorded inputs and different expected outputs, because the user's expansion + answer was never an input column. A matrix that omits an input cannot distinguish the + states that input separates, and it will still look complete. +2. **A counterfactual must distinguish ABSENT from CONTRADICTORY.** The entry claimed the + parent commit was "contradictory" on one state. `git show 17d5ad3:CLAUDE.md` has no + two-tell rule at all — only the undefined phrase "clearly stuck" at line 78. The + contradiction existed solely in an intermediate draft produced *inside this cycle*, so + the check reported a failure mode the prior state could not produce. + +Both survived a full Gate-B pass before being caught on the next one. + +## Loop economics, measured on this cycle + +| Pass | Findings | Blocker | Major | +|---|---|---|---| +| 1 | 14 | 3 | 5 | +| 2 | 24 | 4 | 13 | + +Four of the five tells present at pass 2: finding count rising; Blocker count failing to +fall; findings clustering on the **instrument** (6 of 24, on the matrix rather than on the +rules it scores); findings clustering on **prose about** the rules (9 of 24 — CHANGELOG, +closure record, ledger entry, a parked story). No require↔withdraw pair: pass 2 narrowed +what pass 1 required, which is a qualification and not a withdrawal. + +The reporting duty formally begins at pass 4. These were readable at pass 2, which is the +argument for the duty starting earlier — or for the trend being computed rather than +narrated. + +## Cross-references + +- Findings: `.context/codex-reviews/gate-b-{spec,quality}-fic2-pass-{1,2}.md` +- The matrix as it now stands (nine states, both clauses removed): + `docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md`, section "The named + verification behind the `battery+check` entry" +- Slot-name variance for this cycle: `fic2` discriminates it from the 2026-08-17 cycle + whose records occupy the bare `gate-b-{spec,quality}-pass-N` slots. + +--- + +## Parked at pass 5 (Daniel, 2026-08-26: bounded fix, then close) + +Four findings from Gate-B pass 5 were parked rather than repaired. All four are true. All four +are about the *instrument* or about the entry conditions of unstarted work, and none is about +what §5 tells an agent to do. Pass 5 returned **zero product-behaviour findings**. + +### The one that is this story's subject matter, verbatim + +**A prose rule's check demands a fixture per predicate.** Pass 4 found that the nine-state matrix +could not see the reporting duty at all — no input it read changed when that duty was deleted. A +five-state table was added. Pass 5 then found that the new table takes the *tell count* as a +precomputed input rather than deriving it from raw observations, so a draft could delete or invert +one of the five tell definitions and every row would keep its expected result. The repair asked for +is raw input columns exercising each of the five tell predicates, against both prompt copies at both +revisions: roughly 5 × 2 × 2 hand-scored fixtures, for a rule whose entire product surface is one +paragraph. + +The finding is correct and the repair is disproportionate. That gap — **what differential evidence +a prose-only rule can actually carry, and where the cost of the fixture exceeds the value of the +coverage** — is the question. `battery+check` says "a check that fails without the change" and says +nothing about how much instrument a one-paragraph rule is worth. + +Sequence worth keeping: pass 4 asked for the table, pass 5 asked for it to be rebuilt, and the thing +being measured did not change between them. + +### Three on a parked story's acceptance criterion + +`docs/superpowers/stories/2026-08-17-arms-race-remedy-as-procedure-story.md`. Its kept/moved/dropped +criterion should enumerate individually, per pass 5: the three line contents; all five tell +definitions; and the threshold's consequence (report the triggering tells, hand the decision to the +user). Its problem statement also says §5 carries "one recognition heuristic and a terminal action" +while the same story later inventories two stop paths — a contradiction in the story's own opening. + +Each is true. Each refines the entry conditions of work nobody has begun. One condition from this +group *was* applied rather than parked — the pass-4 activation boundary — because a future rewrite +could otherwise move the duty to pass 1 while checking off every other listed condition. + +### Cycle shape at the point of closing + +| Pass | Findings | Blocker | Major | +|---|---|---|---| +| 1 | 14 | 3 | 5 | +| 2 | 24 | 4 | 13 | +| 3 | 12 | 0 | 6 | +| 4 | 3 | 0 | 2 | +| 5 | 6 | 0 | 5 | + +Two stop-and-surfaces, at pass 2 (four tells) and pass 5 (three tells). Only the pass-5 stop was +required by the shipped duty, which begins at pass 4; the pass-2 stop was an early surface chosen +because the tells were already readable, which is the argument the duty's activation boundary +invites rather than a rule it enforces. The pass-2 stop produced a revert; the pass-5 stop produced +this bounded close. Both were decided by the maintainer, neither by the loop. + +--- + +## Observation recorded 2026-08-26: pass counter disagreed with the pass record + +Raw facts only. **The cause is UNDIAGNOSED**, and no attribution to any existing ledger row +is made here — attribution without diagnosis is the class the ledger polices. + +- Cycle: `fic2`, branch `field-intake-canvas-a1-a5`, closed 2026-08-26 at commit `3cdd075`. +- Passes actually run and validated: **7**. Each wrote both branch files, each file carried a + well-formed terminator and a count matching its finding lines: + `.context/codex-reviews/gate-b-{spec,quality}-fic2-pass-{1..7}.md` — 14 files on disk. +- What the gate hook reported at the closing commit: **"only 1/3 mcp__codex__review pass(es) + since the last commit"**. +- What it reported at each intermediate amend: **"1 recorded pass(es) this cycle"**, from the + amend following pass 1 onward. The value did not advance across passes 2 through 7. +- One intermediate amend instead reported **"no fingerprint is recorded for this cycle"**. +- `.context/codex-gate.passCount` read `7` at the start of the session, before this cycle began; + that value belongs to the previous cycle and was not re-read afterwards. +- Every one of the 7 calls returned `success: true` with a normal result envelope. None timed + out, none was aborted, none returned an `INCOMPLETE` reply. +- The cycle used a slot-name discriminator (`fic2`) rather than the bare + `gate-b-{spec,quality}-pass-N` names. Whether that is related is **not established** — the + hook is documented as never reading the findings file at all. +- Each pass was followed by `git commit --amend` on a single `WIP:`-prefixed commit. + +Not diagnosed, and deliberately not guessed at: whether the counter was reset, never +incremented, incremented and overwritten, or read from a different key than it was written to. +Nobody inspected the hook's state files during the cycle, so there is no evidence either way. + +Consequence for this cycle: none. `CLAUDE.md` §5 says the counter is not evidence and that every +incomplete pass is discounted regardless of what it says; the close rested on the 14 validated +findings files, not on the counter. The observation matters for the instrument, not for this +artifact. + +--- + +## Parked from PR #25 review (Daniel, 2026-08-26): the unavailable-history gap + +Greptile raised one P1 on PR #25 against `CLAUDE.md:143`, the pass-4-onward reporting duty. +Thread: https://github.com/dsnger/dev-workflow-kit/pull/25#discussion_r3864986788 + +**Split verdict.** + +*The stated mechanism is FALSE.* The claim was that "the mandatory findings format cannot +retain all of that history". It can. §5 mandates one findings file per pass per branch at a +**pass-numbered** slot, one finding per line, severity as a leading closed-set field, and a +terminator carrying the count. Both historical inputs the three-line report needs are +therefore derivable from the mandated artifacts alone — trend by counting finding lines and +`^BLOCKER` lines per pass, require↔withdraw by comparing across those same files. Run over +this cycle it reproduced the reported figures exactly (findings 14, 24, 12, 3, 6, 6, 2; +Blockers 3, 4, 0, 0, 0, 0, 0) with no optional artifact consulted. The deletion rule does not +erase history either: §5 deletes only the slot about to be written, and slots are numbered. + +*The gap is REAL, and it is availability rather than format.* Where prior-pass files are +genuinely absent — a fresh checkout, a cleared `.context/`, another machine, a cycle resumed +elsewhere — §5 defines **no behaviour** for the report from pass 4 onward. The cycle-stable +resume note is not the fallback: it is explicitly optional, and `CLAUDE.md:224` says "Nothing +depends on it existing." An agent in that position must invent the trend, omit the line, or +decide for itself, and the two-tell threshold is mandatory on top of whatever it decides. + +**Why parked rather than fixed:** closing it needs new normative §5 content, which is the +surgery this cycle declined twice. It joins Q3-Q5 above as a fourth open question of the same +shape — a gap in the shipped rule whose repair is a contract decision. + +**Q6.** What does the pass-4-onward report do when the prior-pass record is unavailable? The +candidate answers are not obviously equal: report the lines that *are* computable and say +which are not; treat unavailable history as a stop condition of its own; make the resume note +mandatory for cycles that cross a session boundary (which changes an artifact §5 currently +calls advisory); or start the duty's clock at the first pass of the *current* record rather +than of the cycle. + +## Raw observation, undiagnosed: CodeRabbit plan metadata disagrees with the routing file + +Recorded 2026-08-26, not attributed and not acted on. + +- CodeRabbit's run configuration on PR #25 reports **`Plan: Pro Plus`** (Run ID + `cdbb25fa-8c3d-45fb-a259-6b973b2ea965`, review profile CHILL). +- `docs/pr-review-bots.md`'s CodeRabbit row records **Plan: Free** "(per Daniel)", and states + that the earlier Pro Plus reading "was observed on PR #1 only and no longer describes the + account". +- These disagree. **The cause is UNDIAGNOSED.** Two candidates, not distinguished: the plan + actually changed since that row was written, or the run-configuration metadata is + unreliable. +- Why it matters: that row's review-limit reasoning — and part of the argument for routing + CodeRabbit opportunistically rather than blocking on it — rests on the Free reading. +- Not this PR's business; `docs/pr-review-bots.md` is untouched by PR #25. A docs-only + follow-up can correct it **after** diagnosis, not before. diff --git a/.context/codex-reviews/gate-b-pass-1-dispositions.md b/.context/codex-reviews/gate-b-pass-1-dispositions.md new file mode 100644 index 0000000..94df7c9 --- /dev/null +++ b/.context/codex-reviews/gate-b-pass-1-dispositions.md @@ -0,0 +1,15 @@ +# Gate B pass 1 — dispositions (reviewer: moonshotai/kimi-k3 via OpenRouter flip, effort max) + +Spec branch (3): +1. MINOR README.md:160 two→three count — FIX (same defect as quality MAJOR #1). +2. MINOR spec §4 "checked for truth and unchanged" falsified by this commit's edits to docs/coding-workflow.md, docs/sparring-briefing.md (and now README.md) — FIX: temporal scoping note; false sentence shipped by this commit, repo overclaim rule applies. +3. NIT plan lacks a task for the two docs files — COLLECT: plan is a historical artifact; the docs work is task scope recorded in the closing message instead. + +Quality branch (7): +1. MAJOR README.md:160 stale count — FIX to "three". +2. MAJOR ci.yml:84 stale count in step comment — FIX to "three". +3. MINOR coding-workflow.md:247 schema enumeration factually wrong vs mcp-codex-dev@1.0.1 — FIX: name the absent key (no provider key) instead of enumerating present ones; false sentence introduced by this diff. +4. MINOR checker terminator pins "2.2", renumbering breaks 4c repo-wide — COLLECT: reviewer itself says deliberate fail-loud, documented at :549-553, no change required. +5. MINOR sev_case chmod 644 restores CLAUDE.md only; @LOCK@ currently CLAUDE.md-only — COLLECT: latent, dead code today; backlog candidate. +6. NIT checker comment "three space-separated integers" (four fields, one string) — COLLECT. +7. NIT CHANGELOG exception-form bullet asymmetric with severity bullet — COLLECT. diff --git a/.context/codex-reviews/gate-b-passes-2-4-dispositions.md b/.context/codex-reviews/gate-b-passes-2-4-dispositions.md new file mode 100644 index 0000000..af56b31 --- /dev/null +++ b/.context/codex-reviews/gate-b-passes-2-4-dispositions.md @@ -0,0 +1,34 @@ +# Gate B passes 2–4 — dispositions (reviewer: moonshotai/kimi-k3 via OpenRouter flip, effort max, all passes) + +Pass 2 (spec 3 / quality 4, no Blocker/Major): +- MINOR switch-surface table listed 3 of 4 resolution layers — FIXED (coding-workflow.md table → four rows; sparring-briefing.md list extended), amended in c5c48da. +- MINOR spec §4 parenthetical decay warning — COLLECT (accurate today; re-verify on next edit of any listed file). +- NIT plan Step-7 4c=13 stale vs re-measured 19 — COLLECT (plan is frozen; its wrap-up records "13, then 16, then 19"). +- NIT AC-9 parity confirmation — informational, no fix owed. +- Re-collected pass-1 NITs (checker comment field count, sev_case chmod, CHANGELOG bullet asymmetry) — remain COLLECT. + +Pass 3 (spec 4 / quality 6, one MAJOR): +- MAJOR CHANGELOG 0.9.0 "three switch surfaces" falsified by the pass-2 four-surface fix — FIXED (count-free enumeration naming all four surfaces), amended in 2353409. +- MINOR plan fence range 192–629 (now 192–699), :325 (now :341), Step-7 4c=13, "three switch surfaces" phrasing — COLLECT (frozen planning artifact; wrap-up self-corrects the flip count; the plan records its writing-time state). +- Re-collected NITs as above. + +Pass 4 (spec 3 / quality 4, no Blocker/Major — CLEAN FINAL PASS; first attempt timed out at the 2400s MCP limit, orphaned reviewers killed, this is the single Mechanics retry): +- MINOR no inbound link from README/getting-started to the reviewer-model-selection section — COLLECT (docs reachability; backlog candidate). +- MINOR resume path (`codex exec resume`) also carries --model, so a mid-outage model switch silently changes the reviewer on a resumed session — COLLECT (one-sentence docs addition; backlog candidate; verified by reviewer against dist/services/codex-executor.js:88). +- MINOR spec §4 parenthetical decay — COLLECT (repeat of pass 2). +- Re-collected: checker comment field count, sev_case chmod asymmetry, CHANGELOG bullet thinness — COLLECT. + +## Bridge revert — 2026-08-16 ~14:00, on Daniel's decision + +Reverted to the codex native subscription default (Daniel purchased codex credits). +Cost was the driver: OpenRouter frontier pricing at our pass sizes (~$76 for passes +2-4 including the timed-out pass-4 first attempt). Reverted: flip line +`model_provider = "openrouter"` removed from ~/.codex/config.toml; +`model_reasoning_effort` restored to "medium"; per-repo +`.mcp/mcp-codex-dev.config.json` (kimi-k3 pin) deleted. Kept: the fenced +`[model_providers.openrouter]` block — proven inert without the flip; it is the +standing outage bridge. auth.json sha256 953551b4… verified unchanged across the +whole bridge lifecycle. Effective at the next Claude Code session per the +startup-cache caveat (mcp-codex-dev resolves its model chain per project root at +server startup); backups ~/.codex/config.toml.bak-2026-08-16{,-pre-openrouter} +remain. diff --git a/.context/codex-reviews/gate-b-quality-fic2-pass-1.md b/.context/codex-reviews/gate-b-quality-fic2-pass-1.md new file mode 100644 index 0000000..5a1edc7 --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-fic2-pass-1.md @@ -0,0 +1,10 @@ +BLOCKER | high | commit 6c70839 body, paragraphs opening `Check — named verification` and `Counterfactual`; docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:133-172 | The mandatory `battery+check` evidence names current-copy mirror parity as the check even though the closure record itself says parity would pass if both copies omitted the rule, and its counterfactual falsely says the parent copies defined the assigned fix set using the current pass: `git show 17d5ad3:CLAUDE.md` and the matching template show that the entire absorb/stuck block was absent | The stated check does not fail on the actual prior state, so this Gate-B call lacks the mode-required differential evidence and its positive result can mask a false premise | Replace the commit-body check with the nine-state old/new matrix, explicitly record expected and observed results for both copies, and make the counterfactual distinguish an absent definition from an inclusive definition before revalidating the evidence +MAJOR | high | CLAUDE.md:121-145; plugins/dev-workflow/commands/workflow-init.md:317-342; plugins/dev-workflow/CHANGELOG.md:47-57 | A Blocker/Major-free pass at or above the floor is required to close, while any two reporting tells are required to stop and surface, and the tells can coexist with a clean pass, for example a rising Minor/Nit count clustered on prose | The same pass can require two incompatible terminal actions despite the text claiming the rules do not compete | State explicit precedence between clean completion and the two-tell stop in both prompt copies and align the changelog description +MAJOR | high | CLAUDE.md:87-103; plugins/dev-workflow/commands/workflow-init.md:287-304 | The scope procedure says an out-of-set correction resumes once the user says whether the set includes it, but it defines no disposition when the user says no while the finding remains Blocker/Major and the severity rule still requires resolution | Agents can loop on the same stop, silently expand scope, or continue with an unresolved mandatory finding | Define the declined-expansion path, including how the finding is recorded or handed back, whether it is dismissed, and when the gate cycle and pass credit resume +MAJOR | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:242-253; docs/hardening-log.md:76,80-81,88 | The closure says four malformed supersession entries were appended during this cycle, but only three are added in the reviewed range; line 76 and its `the entry immediately above` phrase came from pre-existing commit `baa75c1` | The round's count and its claimed governing-entry audit are false, undermining the durable closeout record | Change the round count to three and distinguish the pre-existing fourth entry, or restate the scope precisely and re-audit only the entries actually appended by this round +MAJOR | high | docs/hardening-log.md:43-45,58-63,90; docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:248-253 | The latest purportedly self-contained governing entry says the reporting duty belongs to the `prompt-vague-criteria` guard, which both creates a cross-row dependency and restates the current answer instead of only naming what is false and citing its source | The append-only repair is itself malformed under the ledger header, so the claim that every affected row has a compliant governing entry is untrue | Append a new governing entry for the target row that completely names its three stale claims and cites the two current prompt blocks without assigning the reporting duty to another ledger guard; leave line 90 as history +MAJOR | high | docs/superpowers/stories/2026-08-17-field-intake-canvas-a1-a5-report-story.md:53-67; docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:105-122 | The amended criterion requires the quoted guard of every nearest prior row examined, but the closure quotes only `mechanical-check-skipped-before-review`; it summarizes `prompt-diagnostic-cause-unnamed` and gives no guard quotation for `rewrite-drops-prior-condition` | The moved closure record exists but does not contain the evidence the amended acceptance criterion says was kept, so the criterion and kept/moved accounting remain unmet | Add the verbatim operative guard and the recorded in-scope or out-of-scope result for every examined prior row, or narrow the criterion with explicit accounting if a row was not part of the precheck +MINOR | high | CLAUDE.md:132-138; plugins/dev-workflow/commands/workflow-init.md:329-335; docs/prompt-standards.md:41-42 | The new mandatory three-line status-report format is described but never shown as a literal example, contrary to prompt-standards checklist item 4 | Different agents can format or combine the three lines inconsistently, weakening the tell comparison the duty exists to support | Add the same compact literal three-line example to both prompt copies +MINOR | medium | AGENTS.md:19-51; docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:1 | The change establishes `docs/field-reports/` as a durable input and closure surface, but the meaningful architecture tree still omits that directory even though the adding-files rule says the tree is a drift surface | Future reviewers and maintainers do not see a now-required artifact class in the repository's single source of truth | Add `docs/field-reports/` to the architecture tree with its committed-input and closure-record role, or explicitly document why it is excluded from the meaningful surface +MINOR | high | docs/superpowers/plans/2026-08-15-reviewer-availability-salvage.md:165,173,359; docs/superpowers/plans/2026-08-01-gate-pass-result-classification.md:57,507 | Adding 84 lines to CLAUDE.md and 73 to the inline template invalidates existing numeric citations that previously landed on the intended §5 text, such as workflow-init line 295 and CLAUDE.md line 150, but the standing-lens pass left those references unchanged | Readers following the cited positions now land in unrelated prompt rules, recreating the docs-drift class this repository explicitly guards | Replace affected numeric references with stable paragraph-opening anchors, or append a historical-location note where closed artifacts must remain immutable, after auditing every CLAUDE.md and workflow-init.md numeric citation +END OF FINDINGS (9 total) diff --git a/.context/codex-reviews/gate-b-quality-fic2-pass-2.md b/.context/codex-reviews/gate-b-quality-fic2-pass-2.md new file mode 100644 index 0000000..8539ebd --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-fic2-pass-2.md @@ -0,0 +1,12 @@ +MAJOR | high | CLAUDE.md:95-100,133-137; plugins/dev-workflow/commands/workflow-init.md:295-300,329-333 | The declined-expansion clause says an out-of-scope Blocker leaves the cycle recorded and unresolved, but the later unqualified surfacing rule says the resolve rule is not waived and every Blocker/Major resolves; the pass-1 repair was not propagated through the existing condition. | A user declining an out-of-scope Blocker produces two mandatory actions, so the loop can either violate the user's scope decision or remain unable to resume. | Scope the later surfacing paragraph to in-set findings or explicitly distinguish stuck/tell surfacing from scope-expansion surfacing in both mirrors. +MAJOR | high | CLAUDE.md:92-100; plugins/dev-workflow/commands/workflow-init.md:292-300 | Recording a declined finding is defined only for the current stop; the rule does not say what happens when the same unresolved finding is reported again on a later full-coverage pass. | The same out-of-scope Blocker can stop every subsequent pass and force the same scope question repeatedly, so the promised path by which the loop resumes without it is not stable. | State that the decline and its record govern later re-reports in this cycle unless scope or evidence changes, and define how the agent cites/disposes the recurrence without asking again. +MAJOR | high | CLAUDE.md:164-168; plugins/dev-workflow/commands/workflow-init.md:355-359 | The replacement precedence paragraph says clean completion outranks “both exits” after discussing a scope stop, a clearly-stuck exit, and the independent two-tell stop, without naming which two; if scope is included it contradicts the rule that even a Minor opening a new contract question must stop. | A Blocker/Major-free pass can either close or surface an out-of-scope structural/contract question, leaving the exact pairing this paragraph claims to settle ambiguous. | Name the clearly-stuck and two-tell exits explicitly and state whether a scope stop outranks clean completion. +MAJOR | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:158-170 | Matrix rows 3 and 10 have identical inputs—floor met, out of set, correction ancestry, MAJOR, and all stop-condition fields n/a—but prescribe different actions (“resolve once the set is settled” versus “record, not resolve”) because the user's expansion decision is not an input. | The named check violates its own complete-input/one-output contract and cannot validate the newly defined declined-expansion path. | Add an expansion-decision input and make row 3 explicitly pending or accepted while row 10 is explicitly declined, then rescore both copies and revisions. +MAJOR | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:170-175; commit 37ee4c8 evidence body | State 11 and the commit evidence classify the parent as contradictory because clean completion and the two-tell threshold allegedly both existed, but `git show 17d5ad3:CLAUDE.md` and the parent inline template contain no two-tell rule at all. | The central counterfactual reports a failure mode the parent could not produce, so the claimed ABSENT-versus-CONTRADICTORY distinction and battery+check evidence are false. | Mark the parent result absent/clean-close, explain that the contradiction existed only in an intermediate draft, and rewrite the score and commit evidence accordingly. +MAJOR | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:146-170 | The eleven-state check never exercises the explicit below-floor plus two-tells branch: row 7 is below the floor with no tell input and row 11 has two tells only at the floor; the table also has no tell-count field despite the commit body's complete-input claim. | The check can pass if “below the floor the threshold binds” is omitted or reversed in either mirror, so battery+check does not validate one of the new precedence clause's two branches. | Add a tell-count input and a below-floor state with at least two tells, with stop-and-surface as the expected result in both new copies. +MAJOR | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:122-128 | The closure record says it quotes the `rewrite-drops-prior-condition` row's guard, but the quoted standing lens (“which existing statements does this diff falsify?”) is from the adjacent `docs-drift` row; the actual rewrite row cites AGENTS.md's kept/moved/dropped accounting rule. | One of the three required verbatim guard checks was performed against the wrong source, so the amended story criterion and the claimed old-condition accounting evidence are not satisfied. | Quote the actual `rewrite-drops-prior-condition` ref verbatim and record the in/out result against that guard; keep the docs-drift lens as a separate check if desired. +MAJOR | high | docs/hardening-log.md:90-91 | The new last-entry-governs record says four claims “do not hold and none held when written,” but its fourth item is that the ref predates the reporting duty and declined-expansion path: a later-added path cannot make an omission false when the row was written, and the preceding governing entry correctly assigns the reporting duty to the other fingerprint. | The supposedly self-contained governing entry violates the header's requirement to distinguish never-true claims from claims that later stopped holding and now makes the ledger's authoritative narration stale. | Separate never-true faults from later obsolescence, include only additions this fingerprint's guard owns, and preserve the boundary that the reporting duty belongs to `prompt-vague-criteria`. +MAJOR | high | docs/superpowers/stories/2026-08-17-arms-race-remedy-as-procedure-story.md:24-43,72-74,101-105 | The parked follow-up describes §5 as having one recognition path and requires future accounting only for the “clearly stuck” clause, while the same diff adds an independent two-tell mandatory stop expressly not conditioned on that clause; mentioning the reporting duty later does not add its threshold and precedence to the kept/moved/dropped duty. | A future implementation following this story can replace the instrument/non-convergence procedure while silently dropping the two-tell stop or clean-completion precedence, recreating the exact decision-procedure defect AGENTS.md forbids. | Enumerate both stop paths, the floor and clean-completion precedence, and require old-condition accounting for both before the story changes their shared decision path. +MINOR | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:258-269 | The record first calls append “the only” repair the ledger convention supplies for these malformed entries, then correctly admits the header prescribes append only for mistyped locators and says nothing about other malformed entries. | The closure record presents an ungoverned judgment as a governed procedure and contradicts itself about why the repair is valid. | Describe append as the deliberately chosen analogy pending the recorded policy gap, or extend the header before claiming the convention supplies it. +NIT | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:152-178 | The method says action agreement is checked everywhere and every one of the eleven rows has an Action cell, but the score says the texts agree “in all nine rows.” | The stated check count is internally inconsistent and obscures whether the two control rows' action outputs were checked. | Change the action count to eleven or explicitly say it refers only to the nine scope/exit-scored rows and separately account for the two controls. +END OF FINDINGS (11 total) diff --git a/.context/codex-reviews/gate-b-quality-fic2-pass-3.md b/.context/codex-reviews/gate-b-quality-fic2-pass-3.md new file mode 100644 index 0000000..0e9fe71 --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-fic2-pass-3.md @@ -0,0 +1,11 @@ +MAJOR | high | plugins/dev-workflow/CHANGELOG.md:39-40 | The 0.10.0 entry says every finding that opens a structural or contract question is outside the assigned scope, but the shipped rule and matrix state 5 allow that finding to be in-set and stop because novelty overrides ancestry | The release note gives users a different scope classifier from the prompt this version ships | Say that a new structural or contract question stops the loop regardless of assigned-set membership, without classifying it as out of set +MAJOR | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:122-128 | The closure still claims to quote the `rewrite-drops-prior-condition` row's ref but quotes the adjacent `docs-drift` standing lens; the actual row cites AGENTS.md's kept, moved, or dropped accounting rule | One of the three required guard checks remains sourced from the wrong row, so the amended story criterion and its claimed in-scope result are unsupported | Quote the actual `rewrite-drops-prior-condition` ref verbatim and record the in or out result against that guard, keeping the standing lens as a separate check if useful +MAJOR | high | docs/hardening-log.md:91 | The last-entry-governs record is not an accurate self-contained description of its target: the unconditional absorb rule is in the target row's `ref`, not its `finding`; the reporting duty belongs to the `prompt-vague-criteria` guard rather than this guard; and a later-added duty cannot be a claim that was false when the row was written | The ledger's authoritative correction violates its own field, ownership, and timing rules, so later recurrence decisions will read stale narration as governing | Preserve history and append a new governing entry that assigns the set-membership fault to `ref`, removes the reporting-duty claim from this fingerprint, and distinguishes never-true faults from later obsolescence +MAJOR | high | docs/superpowers/stories/2026-08-17-arms-race-remedy-as-procedure-story.md:40-43,72-74,101-105 | The story explicitly considers whether the surviving reporting tells are the answer for instrument non-convergence, yet its kept, moved, or dropped acceptance criterion inventories only the recurrence step and the clearly-stuck clause, omitting the pass-4 reporting duty and its two-tell stop | A future implementation can replace the shared non-convergence decision path while silently dropping the reporting threshold, recreating the decision-procedure defect AGENTS.md forbids | Add the existing reporting duty, its three-line carrier, five tells, and two-tell stop to the old-condition inventory the future story must account for +MINOR | high | CLAUDE.md:133-146; plugins/dev-workflow/commands/workflow-init.md:330-343 | The new mandatory three-line status-report format is described but never shown as a literal example, contrary to prompt-standards checklist item 4 | Agents can combine or label the lines inconsistently, weakening cross-pass tell comparison | Add the same compact literal three-line example to both prompt copies +MINOR | high | CLAUDE.md:1-3; plugins/dev-workflow/commands/workflow-init.md:192-195 | The changed root CLAUDE prompt and scaffolded CLAUDE template do not name their executing target model; workflow-init's outer target-model line is not part of the file it writes | Both changed prompt artifacts fail prompt-standards checklist item 1 and therefore AGENTS.md invariant 11 | Add `Target model: Claude via Claude Code` to the root prompt and inside the inline CLAUDE template +MINOR | medium | AGENTS.md:50-58; docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:1 | The diff adds a durable closure and named-verification artifact under `docs/field-reports/` but leaves that directory out of AGENTS.md's meaningful architecture tree despite the adding-files drift rule | Maintainers and reviewers cannot discover this now-required artifact surface from the repository's single source of truth | Add `docs/field-reports/` to the architecture tree with its committed-input and closure-record role +MINOR | high | docs/superpowers/plans/2026-08-15-reviewer-availability-salvage.md:90,155-173 | Inserting the new §5 blocks moved numeric CLAUDE.md and workflow-init.md citations that previously landed on the intended text; for example workflow-init line 295 now lands in the scope rule instead of the finding format and CLAUDE.md line 150 lands in the reporting rationale instead of pass acceptance | Readers following the locators reach unrelated rules, reproducing the size, value, or position docs-drift class the standing lens requires this diff to check | Replace affected numeric locators with paragraph-opening anchors or add commit-qualified historical-location notes after auditing every numeric citation to both changed files +MINOR | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:1 | The title says the round closed on 2026-08-17 even though the same file records repairs and governing ledger work dated 2026-08-18 and 2026-08-26 and the closing commit is still WIP | The durable chronology makes later in-round repairs look external to an already closed round | Remove the close date until the cycle actually closes or update it to the eventual closure date +MINOR | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:254-265 | The record first calls append the only repair the convention supplies for these malformed entries, then correctly admits that the convention supplies append only for mistyped locators and names no repair for this shape | The closure presents an ungoverned analogy as a governed procedure and contradicts its own parked-gap explanation | Describe append as the deliberately chosen analogy pending the recorded policy gap instead of as a repair the convention supplies for this defect +END OF FINDINGS (10 total) diff --git a/.context/codex-reviews/gate-b-quality-fic2-pass-4.md b/.context/codex-reviews/gate-b-quality-fic2-pass-4.md new file mode 100644 index 0000000..c14ddfd --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-fic2-pass-4.md @@ -0,0 +1,4 @@ +MAJOR | high | docs/superpowers/stories/2026-08-17-arms-race-remedy-as-procedure-story.md:43-47,76-78,105-109 | The pass-3 repair names the reporting duty in narrative, but the formal kept/moved/dropped acceptance criterion still requires accounting only for the recurrence step and the clearly-stuck clause; the story's own open question confirms the reporting duty sits beside that clause rather than inside it | A future implementation can satisfy the acceptance checklist while dropping the three-line carrier, five tells, or two-tell mandatory stop, so the same decision-procedure loss remains possible | Amend the acceptance criterion itself to require accounting for both stop paths and explicitly include the reporting duty's three-line carrier, five tells, and any-two stop +MAJOR | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:138-204 | The only named battery+check verification scores scope and action, and its matrix has no pass-number, three-line carrier, or tell inputs, so deleting or reversing the pass-4 reporting duty and its any-two mandatory stop in both prompt copies would leave every one of the nine results unchanged | One of the two assigned loop rules has no differential coverage even though prompt text has no stronger executable test, allowing the evidence entry to remain green across a regression in that rule | Add old/new cases for pre-pass-4 and pass-4-onward reports with zero, one, and two tells; assert the three required outer-status lines and mandatory stop at two tells against CLAUDE.md and the template independently, then revalidate the commit evidence +MINOR | high | CLAUDE.md:153-155; plugins/dev-workflow/commands/workflow-init.md:344-346; .context/codex-reviews/gate-b-fic2-parked-review-economics.md:17-28 | The root prompt says all "rules above" do not compete while the template narrows this to two and explicitly discusses only absorb-vs-stop versus clearly-stuck; because the reporting duty is immediately above, the root wording also asserts away its unresolved competition with clean completion that the parked record says was deliberately not shipped | The two product copies are not semantically aligned and the root copy overstates a precedence decision this cycle explicitly parked | Without deciding the parked precedence question, make both copies name only the absorb-vs-stop and clearly-stuck rules as the non-competing pair +END OF FINDINGS (3 total) diff --git a/.context/codex-reviews/gate-b-quality-fic2-pass-5.md b/.context/codex-reviews/gate-b-quality-fic2-pass-5.md new file mode 100644 index 0000000..9e0fc45 --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-fic2-pass-5.md @@ -0,0 +1,5 @@ +MAJOR | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:193-199 | The R1-R5 table takes `Tells present` as a precomputed verdict instead of supplying the raw observations the rule classifies: finding trend, Blocker trend, cluster category, and require↔withdraw state and identity; it also names the three-line contents only as an expected phrase rather than exercising them | Deleting or reversing a tell definition, omitting Blocker counts, or weakening one required line can leave every R result unchanged, so the named check does not cover the full reporting duty even though it is the only differential evidence for that rule | Replace the tell-count oracle with raw-input columns and cases that derive and exercise each of the five tells and each exact line, then rescore both revisions and both prompt copies +MAJOR | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:223-225 | After the verification grew from one nine-state table to a nine-state plus five-state method, its limitation still says only that `these nine states` are not exhaustive while calling the enumeration `the whole method` | The five reporting states' non-exhaustiveness is left unstated, so the added table can be overread as complete and the text violates AGENTS.md's rule that mechanism limitations name the checked axes and say whether the list is exhaustive | State separately that neither the nine scope/action states nor the five reporting-duty states is exhaustive, and name the raw axes the second table covers and omits +MAJOR | high | docs/superpowers/stories/2026-08-17-arms-race-remedy-as-procedure-story.md:76-83 | The criterion names the `three-line report` and `five tells` only as bundles and omits the threshold consequence that §5 explicitly requires: report the triggering tells and hand the decision to the user | A future procedure can delete a line's required contents, alter a tell definition, or stop generically without reporting why and still satisfy the four labels, recreating the silent-condition-drop defect this criterion exists to prevent | Enumerate the three line contents, all five tell definitions, and the threshold's report-tells and user-handoff consequence individually for kept, moved, or dropped accounting +MINOR | medium | docs/superpowers/stories/2026-08-17-arms-race-remedy-as-procedure-story.md:24-30 | The problem statement says §5 now carries `one recognition heuristic and a terminal action`, but the same story later correctly inventories two separate stop paths: the clearly-stuck test and the any-two-tells reporting path | The parked story starts from a contradictory inventory of the decision procedure it must later preserve, increasing the chance that the reporting path is re-derived or dropped | Narrow `one recognition heuristic` to the clearly-stuck path specifically, or describe both current stop paths here +END OF FINDINGS (4 total) diff --git a/.context/codex-reviews/gate-b-quality-fic2-pass-6.md b/.context/codex-reviews/gate-b-quality-fic2-pass-6.md new file mode 100644 index 0000000..38ca047 --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-fic2-pass-6.md @@ -0,0 +1,5 @@ +MAJOR | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:193-199,223-230 | The repaired limitation says the reporting-duty table covers only pass number, tell count, and carrier, but R5 also varies whether the clearly-stuck reading is available and uses that input to establish that the two-tell stop has no such precondition. | The named battery+check description still understates what its own mechanism checks, so the accepted repair remains incomplete under AGENTS.md's requirement to name the actual checked axes and say whether the list is exhaustive. | Add clearly-stuck-reading availability to the covered axes and state explicitly whether that input-axis list is exhaustive, while retaining that the five row enumeration is not exhaustive. +MINOR | high | .context/codex-reviews/gate-b-fic2-parked-review-economics.md:63-66,128-134 | Both cycle tables report two Blockers for pass 1, but the preserved branch files contain two BLOCKER lines in gate-b-spec-fic2-pass-1.md and one in gate-b-quality-fic2-pass-1.md, for three total. | The companion note's claimed full cycle shape is factually wrong, even though correcting the count does not change the pass-2 rising-Blocker tell. | Change pass 1's Blocker count to 3 in both tables. +MINOR | high | .context/codex-reviews/gate-b-fic2-parked-review-economics.md:74-76,136-138 | The note calls pass 2 one of two mandatory stop-and-surfaces while also stating that the reporting duty formally begins at pass 4; the four tells observed at pass 2 therefore did not yet trigger the shipped duty's mandatory threshold. | The cycle narrative contradicts the activation boundary that this pass correctly added to the parked story criterion and overstates a maintainer-chosen early surface as rule-required. | Describe pass 2 as an early maintainer-chosen surface, or cite a different rule that made it mandatory; reserve the shipped reporting-duty mandate for pass 4 onward. +MINOR | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:47-52; .context/codex-reviews/gate-b-fic2-parked-review-economics.md:128-136 | The closure record says pass 5 had exactly two tells, but the cycle table shows Blockers stayed at 0 from pass 4 to pass 5, which satisfies §5's unqualified tell "the Blocker count failing to fall"; together with the rising finding count and prose clustering this is three tells, as the parked note itself says. | Two durable records give different tell counts for the same stop, making the claimed exercise of the reporting rule internally inconsistent even though either count exceeds the threshold. | Change the closure record to three tells and name the flat Blocker count alongside the rising finding count and prose cluster. +END OF FINDINGS (4 total) diff --git a/.context/codex-reviews/gate-b-quality-fic2-pass-7.md b/.context/codex-reviews/gate-b-quality-fic2-pass-7.md new file mode 100644 index 0000000..8d781f4 --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-fic2-pass-7.md @@ -0,0 +1,2 @@ +NO FINDINGS +END OF FINDINGS (0 total) diff --git a/.context/codex-reviews/gate-b-quality-harden-pass-1.md b/.context/codex-reviews/gate-b-quality-harden-pass-1.md new file mode 100644 index 0000000..d21c3cb --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-harden-pass-1.md @@ -0,0 +1,5 @@ +MAJOR | high | plugins/dev-workflow/commands/workflow-init.md:716 | The rung-P change sharpens item 11 only in the repository copy of docs/prompt-standards.md; the inline checklist that /workflow-init scaffolds still ends item 11 after the old mechanism-naming rule and omits both new hardenings | Newly initialized projects do not receive the hardening this plugin release claims to ship, leaving the recurring class unaddressed in the actual downstream product | Mirror the downstream-neutral source-citation and fourth-correction guidance into the inline item 11 template, preserving the 12-item count +MAJOR | high | CLAUDE.md:237; plugins/dev-workflow/commands/workflow-init.md:414; docs/hardening-log.md:34 | The new standing lens says no parity check, resync, or self-grep can find an invalidated untouched statement and that asking for such statements is what finds them; this categorically rules out repository-wide semantic/parity checks and reads as a detection guarantee, although the addition is only prompt text and nothing enforces either the ask or its success | The hardening reproduces the unverified-enforcement-claim class it is meant to avoid and can make users treat the lens as complete rather than judgment-backed | State that ordinary edit-local checks did not find these incidents, that the targeted ask found these four, and explicitly say the lens is prompt guidance with no deterministic enforcement or completeness guarantee; mirror that wording in both §5 copies and the ledger +MINOR | high | docs/hardening-log.md:32 | The new row labels this the FIFTH occurrence, but the anchored lineage before this change contains the 2026-07-18 first occurrence, the 2026-07-19 recurrence, and the 2026-07-25 third occurrence; the 2026-07-26 row explicitly resolves that same pending third fingerprint rather than recording a fourth defect | The recurrence signal used to choose escalation rungs is off by one and contradicts the ledger's own resolution semantics | Relabel this as the FOURTH occurrence unless a distinct fourth pre-2026-07-27 incident is identified and cited in the row +MINOR | high | docs/prompt-standards.md:102; docs/hardening-log.md:32 | The claim that no deterministic rung reaches the class is justified only by saying any grep for matcher terms has false positives; ruling out one naive grep does not establish that no deterministic comparison, scoped allowlist, or source-derived check can cover any relevant spelling, and the ledger further calls recurrence count “the signal” without saying it is non-blocking | Item 11 overclaims the limits of enforcement while establishing a rule against exactly that behavior, and future hardening decisions may prematurely reject a fitting mechanical rung | Narrow this to “no current deterministic rung covers this paraphrase shape,” state that the ledger count is a human review signal rather than an enforcing check, and record only the concrete limitations of candidate checks actually evaluated +END OF FINDINGS (4 total) diff --git a/.context/codex-reviews/gate-b-quality-harden-pass-2.md b/.context/codex-reviews/gate-b-quality-harden-pass-2.md new file mode 100644 index 0000000..3c7c3b7 --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-harden-pass-2.md @@ -0,0 +1,2 @@ +MAJOR | high | plugins/dev-workflow/CHANGELOG.md:28-29 | The 0.7.1 release note still says categorically that “no parity check or self-grep reaches” untouched stale statements, even though this pass deliberately narrowed the two §5 copies and ledger row to the edited-path checks that actually ran and admits that a narrow check could pin a spelling. | The shipped release note repeats the unverified-enforcement/impossibility claim this hardening is meant to stop, misrepresenting the lens as necessary because deterministic checks cannot reach the class. | Match the scoped wording used elsewhere: say the parity diffs/resyncs/self-greps used in this cycle inspected edited paths and missed these files, and that no comprehensive check currently covers arbitrary semantic drift. +END OF FINDINGS (1 total) diff --git a/.context/codex-reviews/gate-b-quality-harden-pass-3.md b/.context/codex-reviews/gate-b-quality-harden-pass-3.md new file mode 100644 index 0000000..68ca63c --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-harden-pass-3.md @@ -0,0 +1,2 @@ +MAJOR | high | plugins/dev-workflow/commands/workflow-init.md:731 | The downstream checklist says to delete a mechanism claim after a “third or fourth correction,” while the authoritative docs/prompt-standards.md:98 rule sets the threshold at the fourth correction; the two required copies therefore do not agree in substance. | Initialized repositories receive a different decision rule from this repository, so a claim may be deleted one recurrence earlier depending on which copy an agent reads, defeating the settled parity requirement for these inline self-contained templates. | Change the downstream wording to the incident-neutral but equivalent “has needed a fourth correction, delete the claim rather than refine it again.” +END OF FINDINGS (1 total) diff --git a/.context/codex-reviews/gate-b-quality-harden-pass-4.md b/.context/codex-reviews/gate-b-quality-harden-pass-4.md new file mode 100644 index 0000000..8d781f4 --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-harden-pass-4.md @@ -0,0 +1,2 @@ +NO FINDINGS +END OF FINDINGS (0 total) diff --git a/.context/codex-reviews/gate-b-quality-pass-1-dispositions.md b/.context/codex-reviews/gate-b-quality-pass-1-dispositions.md new file mode 100644 index 0000000..256b8d7 --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-pass-1-dispositions.md @@ -0,0 +1,6 @@ +# Gate B — quality branch — pass 1 dispositions + +2 findings (1 Major, 1 Minor). Both accepted. Neither requires a spec change. + +1 MAJOR intake lets a human correction set any value with no recompute or bounds — ACCEPT, sharpest finding of the pass. "A correction may set any value, up or down" was written for the AXES and silently licensed an arbitrary Validation value too, so intake could write security `high` with `battery` — a header §5 then classifies as unresolvable and stops on. The spec already bounds this; the prompt did not carry it. Fixed in the intake skill: an axis correction RECOMPUTES the mode, an explicit mode override may raise or lower but never drops `+abuse-path` while security is `high`, and an intake-time override writes the documented `mode override` log entry. +2 MINOR docs/coding-workflow.md "lighter questions" — ACCEPT. Same defect the spec branch reported as its finding 3; one fix closes both. Two independent reviewers landing on the same sentence is the reason it gets rewritten rather than argued with. diff --git a/.context/codex-reviews/gate-b-quality-pass-1.pre-2026-08-16.md b/.context/codex-reviews/gate-b-quality-pass-1.pre-2026-08-16.md new file mode 100644 index 0000000..f5938b1 --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-pass-1.pre-2026-08-16.md @@ -0,0 +1,10 @@ +MAJOR | high | docs/superpowers/specs/2026-08-05-hardening-ledger-supersession-design.md:287,753 | the amended spec still says backdating is mechanically checked and later says a check enforces non-decreasing dates, contradicting line 176 and §8 lines 832-837, which say no chronology validation exists | the deliberate correction leaves the same false enforcement claim alive elsewhere in the changed specification, so readers receive mutually exclusive guarantees | replace both claims with the instruction-backed reality and re-sweep the spec and plan by meaning for equivalent chronology-enforcement claims +MAJOR | high | scratch c1core.awk, C1d scope selection (`if (!(cand[i] in ATBASE)) { m = i; break }`) | C1d identifies the mandated entry as the first interval line absent anywhere in BASE instead of identifying the entry by its mandated date, row date, and fingerprint | a sanctioned mistyped-entry-then-corrected-entry sequence checks the earlier typo and fails even though the design says the corrected mandated entry must pass; an unrelated newly added entry before the mandated one can likewise determine the verdict | parse all added candidates and select the unique candidate matching D, ROWDATE, and FP, with an explicit ambiguity failure; add a fixture with a mistyped entry followed by its corrected mandated entry +MINOR | high | scratch build-fixtures.sh and matrix.sh, `rows-escapes` C1a/C1d rows | the escape-heavy row is unrelated to the mandated locator, so both checks can pass even if a naive row parser rejects or mis-splits that added row | the advertised PASS does not kill the escaped-pipe or trailing-backslash parser defects it claims to cover, weakening the row-parser evidence | put the escape-heavy finding on the locator-matching row or add a property that requires the unrelated complete row to parse successfully +MAJOR | high | scratch matrix.sh, `if [ -s "$D/.matrix.err" ]; then _st=$((_st + 1)); fi` | stderr is folded into the same nonzero status used for an expected semantic FAIL | when a broken check exits 0 but emits a shell/runtime error, every matrix row expecting FAIL is accepted, so infrastructure failure can mask the exact genuine disagreement the matrix is meant to expose | treat non-empty stderr as an unconditional matrix disagreement before comparing the check's original exit status with PASS/FAIL +MINOR | high | scratch check-c2.sh, C2b parity comparison after `join_paragraphs` | C2b normalizes ordinary line breaks before `files_equal`, so it cannot establish the requested byte-identical shared region | a wrap-only divergence between the live ledger and inline template passes despite violating the two-surface parity contract | compare the raw extracted regions with `cmp -s`; keep paragraph joining only for the anchor-presence check +MAJOR | high | scratch build-fixtures.sh sub_line() and after_line() | fixture builders never assert that their source anchor matched exactly once or that the generated file differs from the baseline | a PASS fixture can silently remain identical to baseline and appear to prove a boundary or quantification behavior its wiring never exercised | require exactly one source-anchor match and compare every generated fixture with baseline before running the matrix +MAJOR | high | scratch c1core.awk parse_entry() | the parser splits the entire entry on ` · ` and requires exactly four parts before recognizing the optional quoted fragment | the shipped convention forbids that separator only in the two prose fields, so a valid fragment containing it is rejected and the checker implements a narrower grammar than the product | parse the locator and optional fragment before splitting the two prose fields, then add a passing fragment fixture containing the separator +MINOR | high | scratch check-c2.sh one_line_contains() and region() | sentinel cardinality is counted by matching lines instead of literal occurrences | two copies of a sentinel on one line satisfy the advertised exactly-once check | count literal occurrences across each file and require exactly one before extracting the region +MINOR | medium | scratch c1core.awk entry parsing loops | whole-array deletion via `delete e` is an awk extension rather than POSIX awk syntax | portability depends on the host awk implementation even though the harness is presented as POSIX and shell-switching alone does not validate the awk dialect | clear arrays portably with `for (k in e) delete e[k]`, or explicitly document and test the required awk implementations +END OF FINDINGS (9 total) diff --git a/.context/codex-reviews/gate-b-quality-pass-10-dispositions.md b/.context/codex-reviews/gate-b-quality-pass-10-dispositions.md new file mode 100644 index 0000000..ad5553e --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-pass-10-dispositions.md @@ -0,0 +1,10 @@ +# Gate B — quality branch — pass 10 dispositions + +4 findings, all Major. All accepted. + +1 MAJOR Task 5's profiled-only branch contradicts itself across Steps 5-7 — ACCEPT, same defect the spec branch reported as its finding 4. Steps 5, 6 and 7 are now consistently branched: paths for every cited story, `Evidence:` entries created/quoted/preserved only for profiled ones, and an optional `Verification:` record for this unprofiled story that is never presented or handled as mode-derived evidence. +2 MAJOR `plugins/dev-workflow/commands/process-pr-review.md` still skips Gate B on fix size alone — ACCEPT, and it is the most consequential finding of the pass: a SHIPPED command, outside every file this change edited, that would let a one-line fix on a security-relevant profiled story bypass the narrowed skip contract entirely. The sweep for pre-existing invalidated sentences is what surfaced it; no parity or grep of my own edits would have. Rewritten to resolve the cited story profile first, permit a skip only at effective level 0 for a profiled story, keep the judgement call for an unprofiled one, and record the battery, the skip reason and any evidence entry. +3 MAJOR the getting-started Opting-out item — ACCEPT, same as spec-1; one fix. +4 MAJOR docs/coding-workflow.md still describes the Gate-B skip as judgement-only — ACCEPT. Third pre-existing site. Now carries the profiled effective-level-0 condition, the unprofiled branch, and the fact that a skip removes review but never evidence. + +Class note: three of this pass's eight findings are sentences that were TRUE before this change and false after it, in files the change never touched. That is a distinct defect class from "fix one site, miss its sibling" — no parity check, resync or self-grep can catch it, only a reviewer asked to look for it. diff --git a/.context/codex-reviews/gate-b-quality-pass-10.md b/.context/codex-reviews/gate-b-quality-pass-10.md new file mode 100644 index 0000000..da27814 --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-pass-10.md @@ -0,0 +1,2 @@ +MINOR | high | docs/hardening-log.md:75 | The appended supersession entry cites the current-answer locations and then restates the corrected per-tool fallback rule, contrary to this ledger's explicit instruction that the final field cite where the answer lives without restating it. | The correction can itself become stale and recreate the narration-drift class this supersession is meant to contain. | Leave this complete entry unchanged; append a later entry with the same row locator that fully states the original row's top-level-only probe was wrong when written and cites only docs/coding-workflow.md's model-record paragraph and the matching docs/sparring-briefing.md rule as the current-answer locations. +END OF FINDINGS (1 total) diff --git a/.context/codex-reviews/gate-b-quality-pass-11-dispositions.md b/.context/codex-reviews/gate-b-quality-pass-11-dispositions.md new file mode 100644 index 0000000..4c7391a --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-pass-11-dispositions.md @@ -0,0 +1,6 @@ +# Gate B — quality branch — pass 11 dispositions + +1 finding, Major. Accepted. + +1 MAJOR the rewritten process-pr-review decision has no no-story branch and omits the all-cited-stories condition — ACCEPT. My pass-10 rewrite fixed the size-only skip and introduced two singular branches, leaving the commonest shape of all — a review-fix PR citing no story — unclassifiable, and permitting a multi-story fix to be decided story-by-story. Both are now explicit, with evidence entries still scoped to cited profiled stories. + Worth noting where this came from: pass 10's fix to this file was itself prompted by the pre-existing-sentence sweep, and it introduced this gap. Two consecutive rounds of finding-then-breeding on the same file argue for reading a rewritten decision procedure against every input shape it can receive, not only the one that motivated the rewrite. diff --git a/.context/codex-reviews/gate-b-quality-pass-11.md b/.context/codex-reviews/gate-b-quality-pass-11.md new file mode 100644 index 0000000..b3d980d --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-pass-11.md @@ -0,0 +1,2 @@ +MINOR | high | docs/hardening-log.md:76 | The governing supersession says "the entry immediately above", directly violating the ledger grammar that entries never reference one another. | The durable correction is itself invalid and requires another corrective append, obscuring the row's governing state. | Append a new locator-identical supersession that describes only the 2026-08-16 row's false probe and ends with citations to the two current-answer locations, without mentioning either earlier entry. +END OF FINDINGS (1 total) diff --git a/.context/codex-reviews/gate-b-quality-pass-12-dispositions.md b/.context/codex-reviews/gate-b-quality-pass-12-dispositions.md new file mode 100644 index 0000000..64986a7 --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-pass-12-dispositions.md @@ -0,0 +1,7 @@ +# Gate B — quality branch — pass 12 dispositions + +2 findings (1 Major, 1 Minor). Both accepted. (The spec branch was CLEAN this pass — NO FINDINGS.) + +1 MAJOR docs/coding-workflow.md claims documentation-only changes carry no gate "since there is no code diff" — ACCEPT, and the sentence is mine from pass 10: I corrected the skip half of it and carried its false rationale forward unexamined. Wrong twice — prompt Markdown (the policy file, skills, commands, agent definitions, hook messages, inline templates) fires FULL Gate B precisely because it is the product, and "no diff exists" is not why explanatory prose is exempt. Rewritten to name the explanatory-only exemption, state the real reason (a wrong sentence costs a confused reader, not broken behaviour), and say plainly that prompt artifacts are not prose. +2 MINOR the `$policy` fallback "this project's review policy" is non-navigable in a marker-only setup — ACCEPT the string half, DECLINE the test half. The reminder now ends "if you cannot locate and check that rule, run the remaining passes", which resolves the unsafe-direction guess without teaching the hook anything: it is invariant 2 (loose in the firing direction) applied to the author's fallback, and it still restates no rule. Still one string, still inside the human's waiver, still profile-agnostic. + Declined: adding a marker-only test assertion for the clause. §5 says Minor findings are collected and not iterated on, the waiver covers a string rather than the hook's test surface, and a new assertion pinning reminder WORDING would make future wording fixes fail a test that checks phrasing rather than behaviour — the two existing assertions match only `below floor|floor NOT met` for exactly that reason. Recorded here as a deliberate non-fix, not an oversight. diff --git a/.context/codex-reviews/gate-b-quality-pass-12.md b/.context/codex-reviews/gate-b-quality-pass-12.md new file mode 100644 index 0000000..8d781f4 --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-pass-12.md @@ -0,0 +1,2 @@ +NO FINDINGS +END OF FINDINGS (0 total) diff --git a/.context/codex-reviews/gate-b-quality-pass-13-dispositions.md b/.context/codex-reviews/gate-b-quality-pass-13-dispositions.md new file mode 100644 index 0000000..94a23ab --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-pass-13-dispositions.md @@ -0,0 +1,5 @@ +# Gate B — quality branch — pass 13 dispositions + +1 finding, Minor. Accepted — same defect the spec branch reported as a Major, so it is treated at the higher severity and fixed rather than collected. + +1 MINOR stage-8 says agent definitions categorically fire Gate B — ACCEPT. Both branches converging on one sentence, at different severities, is a useful signal on its own: the claim was wrong in the same way to two independent readers. One fix closes both, verified against the hook's literal `is_prompt_path` regex rather than against prose describing it. diff --git a/.context/codex-reviews/gate-b-quality-pass-13.md b/.context/codex-reviews/gate-b-quality-pass-13.md new file mode 100644 index 0000000..f0201ff --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-pass-13.md @@ -0,0 +1,2 @@ +MINOR | high | docs/coding-workflow.md:132-134 | The rewritten stage-8 summary says agent definitions categorically fire full Gate B, but the hook matches only `CLAUDE.md`/`AGENTS.md` and paths under `.claude/`, `plugins/`, `skills/`, or `commands/` at any depth; a bare `agents/` directory is not matched. | Readers can rely on Gate B for an agent-definition layout that the hook actually classifies as docs-only, recreating the exact path-classification overclaim the detailed §5 text avoids. | Name the hook's literal matched filenames/directories and any-depth rule here, and qualify agent definitions as covered only when they live under `.claude/` or `plugins/` (or explicitly note that bare `agents/` is not detected). +END OF FINDINGS (1 total) diff --git a/.context/codex-reviews/gate-b-quality-pass-14.md b/.context/codex-reviews/gate-b-quality-pass-14.md new file mode 100644 index 0000000..8d781f4 --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-pass-14.md @@ -0,0 +1,2 @@ +NO FINDINGS +END OF FINDINGS (0 total) diff --git a/.context/codex-reviews/gate-b-quality-pass-15-dispositions.md b/.context/codex-reviews/gate-b-quality-pass-15-dispositions.md new file mode 100644 index 0000000..8b67988 --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-pass-15-dispositions.md @@ -0,0 +1,7 @@ +# Gate B — quality branch — pass 15 dispositions + +1 finding, Major. Accepted. (The spec branch was CLEAN this pass — NO FINDINGS.) + +1 MAJOR the one-story profiled branch makes effective level 0 SUFFICIENT for a skip — ACCEPT, and it is the most serious defect of the whole Gate-B cycle, because it widens the gate-off path the change promises never to widen. My pass-11 rewrite listed eligibility per citation shape and left the fix's own triviality stated only in the no-story branch, with the catch-all reading as a fall-through. Read literally, a substantial logic fix could skip Gate B because its cited story happened to be profiled `trivial`/`none` — a NEW way to skip a gate, invented by a command while the policy it cites was being narrowed. + Restructured so the shape cannot be misread: the substantial-fix rule leads ("every substantial fix runs Gate B — no profile makes a substantial fix skippable"), and a skip requires BOTH a trivial fix AND every cited story eligible, stated as a conjunction rather than as branches a reader completes by fall-through. + Class note: this is the third defect introduced into this same file by a fix to it (pass 10 size-only skip -> pass 11 missing branches -> pass 15 sufficiency error). The pattern is that each rewrite preserved the NEW rule and lost an OLD condition that lived in prose the rewrite replaced. A decision procedure being rewritten should be diffed for conditions DROPPED, not only checked for the condition added. diff --git a/.context/codex-reviews/gate-b-quality-pass-15.md b/.context/codex-reviews/gate-b-quality-pass-15.md new file mode 100644 index 0000000..9d4e57d --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-pass-15.md @@ -0,0 +1,2 @@ +MAJOR | high | plugins/dev-workflow/commands/process-pr-review.md:141-155 | The one-story profiled branch makes effective level 0 sufficient for a Gate-B skip without also requiring the review fix itself to be trivial; the catch-all then sends substantial fixes to Gate B only when they fall outside the preceding eligibility branches. | A substantial logic or path-changing bot fix can now bypass Gate B merely because its cited story was originally profiled `trivial`/`none`, widening the pre-existing trivial-change exemption and violating the settled no-new-gate-off-path decision. | Keep the existing diff-level triviality condition in every branch: require both a trivial fix and story eligibility for a skip, and state unambiguously that every substantial fix runs Gate B regardless of profile. +END OF FINDINGS (1 total) diff --git a/.context/codex-reviews/gate-b-quality-pass-16.md b/.context/codex-reviews/gate-b-quality-pass-16.md new file mode 100644 index 0000000..8d781f4 --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-pass-16.md @@ -0,0 +1,2 @@ +NO FINDINGS +END OF FINDINGS (0 total) diff --git a/.context/codex-reviews/gate-b-quality-pass-2-dispositions.md b/.context/codex-reviews/gate-b-quality-pass-2-dispositions.md new file mode 100644 index 0000000..7ff3dd7 --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-pass-2-dispositions.md @@ -0,0 +1,6 @@ +# Gate B — quality branch — pass 2 dispositions + +2 findings (1 Major, 1 Minor). Both accepted. Finding 1 changes specified behaviour, so the SPEC is updated in the same commit per §5. + +1 MAJOR §5 sends the skip reason to the profile log, which has no event kind for it — ACCEPT. A real contradiction between two shipped prompts: the log defines exactly three event kinds (axis change, mode override, adoption) and a skip is none of them. Of the two offered fixes, a fourth event kind is the wrong one — the log is a record of PROFILE CHANGES, and a skip changes no profile value; it is a per-cycle decision. The skip reason therefore goes where the other per-cycle durable record already lives: the commit body, beside the evidence entry. That contradicts the approved spec's sentence putting it in the profile log, so the spec is corrected in this same commit — the §5 rule about a Gate-B fix that alters specified behaviour, applied to its own repo. +2 MINOR CHANGELOG omits `abuse` from the risk lens enumeration — ACCEPT. The changelog is read as the release contract; an incomplete lens list understates what ships. diff --git a/.context/codex-reviews/gate-b-quality-pass-2.STALE-other-cycle-preserved.md b/.context/codex-reviews/gate-b-quality-pass-2.STALE-other-cycle-preserved.md new file mode 100644 index 0000000..d808f40 --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-pass-2.STALE-other-cycle-preserved.md @@ -0,0 +1,4 @@ +MAJOR | high | CLAUDE.md:462 and plugins/dev-workflow/commands/workflow-init.md:652 | The new `stale-fingerprint STOP` label is narrower than the hook branch it claims to describe: codex-gate.sh:941-943 enters that STOP when the current fingerprint is `unavailable`, the recorded fingerprint is `unavailable`, or the two usable values differ; computation or storage failure can therefore reach it even when content is unchanged | The shipped prompts again turn a broader gate mechanism into a false single-cause label, violating AGENTS.md's exact-comparison rule and steering readers away from the machinery-failure cases the hook message deliberately preserves | Call it the cannot-confirm STOP and, if retaining detail, state the exhaustive comparison as current or recorded fingerprint unavailable or unequal +MINOR | high | CLAUDE.md:465, plugins/dev-workflow/commands/workflow-init.md:655, docs/superpowers/specs/2026-08-14-reviewer-availability-fallback-design.md:197, and plugins/dev-workflow/CHANGELOG.md:33 | `It lands whichever it draws` and `the hook is advisory and always exits 0, so the commit lands` infer successful commit creation from this hook's exit status | Exit 0 establishes only that this hook does not block the commit attempt; Git, another hook, or repository state can still reject it, so the correction introduces a new guarantee its cited mechanism does not provide | Say that this hook never blocks the commit attempt or that its reminder is advisory, without claiming the commit necessarily lands +NIT | high | CLAUDE.md:458 and plugins/dev-workflow/commands/workflow-init.md:648 | The mirror differs beyond the necessary repo-specific invariant citations: `is_docs_only` becomes "the hook's docs-only test," and the outcome articles also change | Although the two versions currently retain the same meaning, these extra differences make future parity checks noisier and fail the requested confirmation that divergence is limited to invariant-number wording | Use identical mechanism and outcome wording in both copies, changing only the invariant-number references that cannot be scaffolded safely +END OF FINDINGS (3 total) diff --git a/.context/codex-reviews/gate-b-quality-pass-3-dispositions.md b/.context/codex-reviews/gate-b-quality-pass-3-dispositions.md new file mode 100644 index 0000000..8d24410 --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-pass-3-dispositions.md @@ -0,0 +1,6 @@ +# Gate B — quality branch — pass 3 dispositions + +2 findings (1 Major, 1 Nit). Both accepted; both duplicate the spec branch's findings 3 and 4, which is corroboration rather than noise. + +1 MAJOR the plan's Task 2 §5 insertion still routes the skip reason to the profile log — ACCEPT. Same defect the spec branch reported; one fix closes both. The plan's embedded block is synced to the corrected text, and the correction is stated rather than silently swapped so a reader sees why. +2 NIT the MANIFEST.md verdict swallows the following process paragraph — ACCEPT. Same as the spec branch's Nit; formatting restored. diff --git a/.context/codex-reviews/gate-b-quality-pass-3.md b/.context/codex-reviews/gate-b-quality-pass-3.md new file mode 100644 index 0000000..6ef1801 --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-pass-3.md @@ -0,0 +1,2 @@ +MAJOR | high | Gate-B evidence / 42227feb commit body | No current evidence entry carrying docs/superpowers/stories/2026-08-13-reviewer-availability-fallback-story.md is quoted to this pass or present in the WIP commit body; the task supplies detached evidence claims, while CLAUDE.md §5 requires each profiled story's evidence entry itself to carry the story path and be revalidated and quoted verbatim | The review cannot bind the claimed battery and check to this profiled story, and adding the mandatory entry later would produce an unreviewed closing record | Add the story path and named evidence to the WIP commit body, revalidate it, and supply that exact entry verbatim on the next Gate-B pass +END OF FINDINGS (1 total) diff --git a/.context/codex-reviews/gate-b-quality-pass-4-dispositions.md b/.context/codex-reviews/gate-b-quality-pass-4-dispositions.md new file mode 100644 index 0000000..8efd95f --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-pass-4-dispositions.md @@ -0,0 +1,6 @@ +# Gate B — quality branch — pass 4 dispositions + +2 findings (1 Major, 1 Minor). Both accepted; both were reported independently by the spec branch, which is corroboration on the two defects that mattered most this round. + +1 MAJOR intake's correction path demands a second confirmation while promising no second pause — ACCEPT, same defect as spec-1. The branch's suggested fix (permit a bounded settlement pause and synchronize skill, spec, plan and CHANGELOG) was NOT taken: relaxing the single-round rule would change an approved decision and ripple through four artifacts to buy something the derivation already gives for free. Taking the cheaper route instead — the proposal states the derivation, so one answer settles axes, override and mode together — leaves the approved single-round contract intact and needs no spec change. +2 MINOR the plan's "verbatim" Profiles block is missing the two paragraphs added in pass 1 — ACCEPT, same as spec-3. Synced from CLAUDE.md and verified equal, together with the intake block. diff --git a/.context/codex-reviews/gate-b-quality-pass-4.md b/.context/codex-reviews/gate-b-quality-pass-4.md new file mode 100644 index 0000000..bb2a53a --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-pass-4.md @@ -0,0 +1,4 @@ +MAJOR | high | CLAUDE.md:458-466; plugins/dev-workflow/commands/workflow-init.md:648-656; plugins/dev-workflow/CHANGELOG.md:29-35; docs/superpowers/specs/2026-08-14-reviewer-availability-fallback-design.md:192-198,644-646; docs/superpowers/plans/2026-08-15-reviewer-availability-salvage.md:292-294 | This fourth review round still finds a hook-mechanism overclaim: “fires,” “routes,” and “draws” collapse internal branch selection into observable output despite the pre-routing adoption exit at codex-gate.sh:104 and the opt-out suppression at line 318; docs/prompt-standards.md:98-99 expressly requires deleting a mechanism claim when it needs a fourth correction rather than refining it again | Keeping and multiplying the state-machine paraphrase violates the repo's prompt standard and recreates the exact recurring defect that standard was added to stop, making another synonym-level regression likely | Delete the detailed hook-state narration from all five restatements; keep the policy result that an empty diff creates no review obligation and, where operational context is useful, say only that the advisory hook may still emit a Gate-B reminder and point maintainers to codex-gate.sh +MINOR | high | docs/superpowers/specs/2026-08-14-reviewer-availability-fallback-design.md:646 | “What no hook does is read the record” is not the exact comparison performed: input_field reads the whole Bash command at codex-gate.sh:905, so an inline exception body is consumed before is_commit and is_wip_commit scan the command; the missing behavior is semantic recognition or validation of the record | The new limitations wording fails AGENTS.md's exact-mechanism rule and can mislead future work about what data the hook receives | Replace “read” with “recognize or validate as a human-exception record” +NIT | high | CLAUDE.md:459; plugins/dev-workflow/commands/workflow-init.md:649; plugins/dev-workflow/CHANGELOG.md:30; docs/superpowers/specs/2026-08-14-reviewer-availability-fallback-design.md:193 | “PreToolUse list” is imprecise because no path list comes from the PreToolUse payload; codex-gate.sh:914-916 derives a staged-path list from Git at invocation time | The terminology makes an already over-detailed explanation harder to verify against the hook | If the mechanism sentence remains, call this the staged-path list computed during PreToolUse +END OF FINDINGS (3 total) diff --git a/.context/codex-reviews/gate-b-quality-pass-5-dispositions.md b/.context/codex-reviews/gate-b-quality-pass-5-dispositions.md new file mode 100644 index 0000000..d42dbc6 --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-pass-5-dispositions.md @@ -0,0 +1,40 @@ +# Gate-B pass 5 — dispositions (quality branch) + +Range f9ed886..556cbfe. Quality branch only: pass 4's spec MINORs were fixed alongside +its MAJOR, so the spec branch is re-run in the closing round rather than against a tree +already superseded. 1 finding, MAJOR, validated file. + +## Applied — scratch harness + +1. **MAJOR — the unreadable-ledger C2 row records only C2a.** ACCEPTED and **reproduced**: + `check-c2.sh` on the directory fixture emitted `FAIL C2a: the ledger is a readable file` + and then returned, with no C2b line anywhere in the output. The plan (Task 3 Step 6) + requires *both* labels to fail on that input. + + **This was self-inflicted, by pass 3's own fix.** The early bail-out I added at pass 3 — + correct in itself, since handing awk a directory is host-dependent — cut C2b's assertions + out of that run. Worse, the `matrix.sh` comment I wrote in the same pass still asserted + "both emit their own FAIL lines", which had been true *before* the bail-out and was false + after it. A fix introduced a stale claim about itself, in the check whose whole purpose is + catching stale claims. + + Fixed: the guard now records C2b's own verdict before stopping — on readability, so no + parser is reached and the host-dependence stays closed. Both labels now emit a FAIL line + (`--- C2: 2 failure(s)`), stderr is still empty, and C3's single label still fails as the + plan requires. The `matrix.sh` comment is corrected and now states plainly what the row's + exit status can and cannot establish, and records the earlier version as the cautionary + case. + +## The pattern, now four for four + +Every pass of this cycle found the same defect class in the evidence harness: **a fixture or +check reporting an expected failure that would report it just as loudly with the tested +property removed.** Passes 2 (unasserted appends), 3 (calendar dates unreachable), 4 (locator +equality unreachable), 5 (C2b unreachable). The fix is always the same shape — make the check +reach the property — and the recurrence is itself the finding: this is what the hardening +ledger exists to record, and it is a candidate row of its own. + +## Committed diff + +Unchanged. This fix touched only the never-committed harness, so `556cbfe` still stands and +nothing was amended. The closing round reviews that same commit with the corrected harness. diff --git a/.context/codex-reviews/gate-b-quality-pass-5.md b/.context/codex-reviews/gate-b-quality-pass-5.md new file mode 100644 index 0000000..81d6787 --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-pass-5.md @@ -0,0 +1,5 @@ +MAJOR | high | commit 949db3fa:3-4 | The commit body copies the story's Risk, Security, and Validation values even though CLAUDE.md:302-304 makes the story header the single writable profile copy and CLAUDE.md:375-377 says an evidence entry carries the story path and named evidence but not the mode value | A later profile change can leave the durable commit record and the review context carrying a plausible stale profile, steering the wrong lenses or evidence obligations | Delete the parenthesized profile summary and retain only the story path plus the named evidence, reading the profile fresh from the cited story +MINOR | high | plugins/dev-workflow/commands/workflow-init.md:649-650 | The scaffolded prompt tells a downstream reader to read codex-gate.sh, but workflow-init does not scaffold that file and the installed plugin's version-keyed cache path is explicitly not an API; this conflicts with docs/prompt-standards.md:92-98's self-contained inline-template exception | An initialized project receives an unactionable pointer and cannot inspect the claimed authority from its repository, so the deletion leaves a small information dead end | Remove the codex-gate.sh pointer from both mirrored copies and keep the self-contained action rule that reminder output is not evidence of reopening and only an empty diff is exempt +MINOR | high | docs/sparring-briefing.md:45-50 | The standing-lens copy still directs the advisor to read the configured reviewer value at pass time, while the new CHANGELOG entry at lines 37-42 correctly says mcp-codex-dev caches its model chain at server startup and current configuration can name a different model | An advisor following this prompt can record the wrong reviewer identity and misjudge the family-independence rule that gives the gate its value | Mirror docs/coding-workflow.md:252-259 here: record the model the pass ran under and probe the reviewer wherever its cached model can differ from current configuration +MINOR | high | commit 949db3fa:22-23 | The revalidated evidence entry says three ordinary outcomes are named in the corrected prose, but this pass deliberately deleted that outcome enumeration | The durable validation record claims coverage of text that no longer exists, contrary to CLAUDE.md:375-379's revalidation rule, and misstates what the battery supports | Delete the final sentence; the preceding battery result and discriminating allow-empty check remain sufficient for the battery+check profile +END OF FINDINGS (4 total) diff --git a/.context/codex-reviews/gate-b-quality-pass-6-dispositions.md b/.context/codex-reviews/gate-b-quality-pass-6-dispositions.md new file mode 100644 index 0000000..800c0be --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-pass-6-dispositions.md @@ -0,0 +1,6 @@ +# Gate B — quality branch — pass 6 dispositions + +1 finding, Major. Accepted. (The spec branch was CLEAN this pass — `NO FINDINGS`.) + +1 MAJOR the ordered-override resolver has no case for a log with no axis change — ACCEPT, and it is a hole in a rule I added at plan pass 2 to close a different hole. "Only the latest `mode override` recorded AFTER the latest `axis change`" is unsatisfiable when no axis change exists — which is precisely the ordinary intake-time override the same design explicitly supports. A literal gate reader would classify every such profile as unresolvable and stop, so the supported path would be broken by the rule meant to protect it. + Reworded in the spec, both §5 copies and the plan block: the latest `mode override` explains a mismatch when its direction is compatible; if the log ALSO holds an `axis change`, the override must postdate the latest one; a log with no axis change is the intake-time case and resolves on its own. Parity re-verified — TEMPLATE PARITY: OK, PLAN PARITY: OK. diff --git a/.context/codex-reviews/gate-b-quality-pass-6.md b/.context/codex-reviews/gate-b-quality-pass-6.md new file mode 100644 index 0000000..71f7567 --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-pass-6.md @@ -0,0 +1,2 @@ +MINOR | high | plugins/dev-workflow/CHANGELOG.md:31 | The 0.9.1 entry says both corrected prompt copies point at `codex-gate.sh`, but CLAUDE.md:458-461 and workflow-init.md:648-651 contain no such pointer, and the same entry later says the mechanism claim was deleted | The release note falsely describes the shipped prompt and leaves the hook-mechanism accuracy branch not fully closed | Delete “pointing at `codex-gate.sh` instead,” while retaining the accurate statement that the copies do not restate the hook decision logic +END OF FINDINGS (1 total) diff --git a/.context/codex-reviews/gate-b-quality-pass-7-dispositions.md b/.context/codex-reviews/gate-b-quality-pass-7-dispositions.md new file mode 100644 index 0000000..2021640 --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-pass-7-dispositions.md @@ -0,0 +1,6 @@ +# Gate B — quality branch — pass 7 dispositions + +1 finding, Major. Accepted. + +1 MAJOR the `axis change` example contradicts the rule it teaches — ACCEPT. The line recorded `risk ↑, security ↓` but described the effect as "swaps the risk lens set for none", which is neither movement's consequence: raising risk ADDS the risk lens set, lowering security may REMOVE the security one. Its reason ("no auth surface after all") also explained only the security drop, leaving the risk increase unmotivated. Second example-versus-rule contradiction this cycle — the first was the `mode override` example at pass 2 — which says the examples deserve the same scrutiny as the prose they illustrate, since an executing agent copies the example. + Rewritten so the reason explains both movements (an irreversible truncation found, the apparent auth surface unreachable) and the downstream clause names both lens effects. diff --git a/.context/codex-reviews/gate-b-quality-pass-7.md b/.context/codex-reviews/gate-b-quality-pass-7.md new file mode 100644 index 0000000..42e5d4a --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-pass-7.md @@ -0,0 +1,5 @@ +MAJOR | high | plugins/dev-workflow/CHANGELOG.md:27-43; docs/hardening-log.md | The release entry identifies accepted, actionable Greptile and CodeRabbit findings, but this range contains no `harden-finding` result or appended ledger row even though `plugins/dev-workflow/commands/process-pr-review.md:167-170` requires every fixed bot finding matching an existing class or warranting a new one to take that path; these findings match the recurring `docs-drift` and `unverified-enforcement-claim` classes | The core self-hardening mechanism loses recurrence and rung evidence, so the same class can return without the escalation the workflow promises | Run `dev-workflow:harden-finding` for each accepted actionable bot finding, perform its anchored guard-scope check, and record the resulting hardening or named pending prerequisite +MINOR | high | docs/sparring-briefing.md:45-53 | The retained heading and lead say the reviewer is whatever is configured and direct the advisor to read its value from configuration, while the new sentences say the running server may use a different cached model and must be probed; this prompt artifact therefore gives contradictory attribution instructions, contrary to prompt-standards item 7 | An advisor can still record the configured model instead of the model that actually ran and can misjudge the family-independence rule that gives the gate its value | Rewrite the heading and lead around the model the pass actually ran under, treating configuration only as the expected input and probing as the authority wherever cache state can differ +MINOR | high | plugins/dev-workflow/CHANGELOG.md:34 | “The two copies are now byte-identical” contradicts the approved story and design, which explicitly say the two §5 copies are not byte-identical and require parity only for the edited regions; the current whole-section hashes also differ while the changed human-exception regions match | The release note overstates parity and can make a future maintainer assume unrelated legitimate differences were synchronized or mechanically guarded | Say that the edited human-exception regions are byte-identical, not the two copies as a whole +MINOR | high | plugins/dev-workflow/CHANGELOG.md:35-37; todos.md:175-180 | The changelog records a distinct PR #24 chained `git add && git commit` observation of the compound-command timing gap, but the untouched backlog still ends at “Occurrence 3” and says there have been three occurrences | The standing recurrence record and its count are stale, so later prioritization or hardening decisions start from incomplete field evidence | Add the PR #24 event as occurrence 4 and update the total, or explicitly state why this deliberate reproduction is excluded from the occurrence count +END OF FINDINGS (4 total) diff --git a/.context/codex-reviews/gate-b-quality-pass-8-dispositions.md b/.context/codex-reviews/gate-b-quality-pass-8-dispositions.md new file mode 100644 index 0000000..0f41484 --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-pass-8-dispositions.md @@ -0,0 +1,5 @@ +# Gate B — quality branch — pass 8 dispositions + +1 finding, Major. Accepted; identical to the spec branch's finding, found independently and with the wider file list (it named the spec and plan as well as both §5 copies). + +1 MAJOR the multi-story aggregation rule contradicts the corrected call contract and the unprofiled-story compatibility promise — ACCEPT. Same defect, same fix: each cited PROFILED story satisfies its mode and suffix with its own named evidence entry; each unprofiled cited story contributes only its path and keeps today's behaviour. Applied to the spec, the plan block, and both §5 copies, with parity re-verified (TEMPLATE PARITY OK, PLAN §5 PARITY OK) and a grep confirming no unqualified phrasing survives anywhere. diff --git a/.context/codex-reviews/gate-b-quality-pass-8.md b/.context/codex-reviews/gate-b-quality-pass-8.md new file mode 100644 index 0000000..9c59bda --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-pass-8.md @@ -0,0 +1,5 @@ +MAJOR | high | docs/coding-workflow.md:241-243 | The timing-facts paragraph still says the model chain is resolved once per project root "at server startup", contradicting lines 255-259 and pinned mcp-codex-dev@1.0.1: startup preloads only the launch root, while any other root is loaded and cached on its first call. | A reader can still infer that a post-start edit never affects any root, so the secondary caching correction remains internally inconsistent and operationally wrong. | Say that the launch root is loaded at startup and each other project root on its first call, then cached until restart. +MAJOR | high | commit 852520ef:Evidence block | The durable evidence entry verifies only the empty-record-commit correction; it contains no named source-level verification or counterfactual for the reviewer-model caching correction also carried by this range. | The profiled Gate-B record does not cover the full change the pass is asked to validate, so the model-selection claim can merge without its stated pinned-server evidence. | Add the mcp-codex-dev@1.0.1 source inspection: config.js checks cachedByProjectRoot before reading configuration and inserts the resolved result afterward, with the counterfactual that an already-loaded root would otherwise reread an edited model on its next call. +MAJOR | high | docs/sparring-briefing.md:45-55; docs/coding-workflow.md:252-261 | Both rules say to "probe the reviewer" when configuration may differ but define neither a deterministic probe nor where the promised pass record lives; an ad-hoc model self-report is not authoritative. | The reader cannot reliably carry out the new rule, so reviewer-family independence can still be recorded from the wrong model despite the wording correction. | Name an actionable probe, such as mcp__codex__health with the same workingDirectory and inspection of the cached effective review-model override with top-level fallback, and name the durable pass-record destination and fields. +MINOR | high | plugins/dev-workflow/CHANGELOG.md:49-50 | The release note says docs/coding-workflow.md was corrected in this branch's parent commit, but parent d5b47bd still contains the old text and commit 852520ef changes both that file and docs/sparring-briefing.md. | The changelog misstates the release history and sends future audits to a commit that does not contain the claimed correction. | State that both documents are corrected in the 0.9.1 commit/range, or name the actual commit after the history is finalized. +END OF FINDINGS (4 total) diff --git a/.context/codex-reviews/gate-b-quality-pass-9-dispositions.md b/.context/codex-reviews/gate-b-quality-pass-9-dispositions.md new file mode 100644 index 0000000..10dc892 --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-pass-9-dispositions.md @@ -0,0 +1,6 @@ +# Gate B — quality branch — pass 9 dispositions + +2 findings, both Major. Both accepted. + +1 MAJOR the plan's Task 5 obliges this cycle's own UNPROFILED story to produce and preserve an evidence entry — ACCEPT, and the irony is the point: the plan demonstrated the exact tightening the compatibility guarantee forbids, on itself. Steps 5-7 are now explicitly scoped to cited PROFILED stories; an unprofiled story contributes its path and nothing more unless adopted through §6's confirmed procedure. The closing template's line is relabelled `Verification:` rather than `Evidence:`, because recording what was run is honest practice while calling it evidence would claim an obligation this story does not carry. +2 MAJOR the cycle-closing Mechanics clause is singular while the call and aggregation contracts are per-story — ACCEPT. In a mixed or multi-profiled cycle an agent would preserve one entry at amend time and silently discard the others, destroying durable evidence the final clean pass had consumed. Now: one validated entry per cited profiled story, none for unprofiled ones, mirrored in both §5 copies with parity re-verified for the Mechanics block as well as the Profiles block. diff --git a/.context/codex-reviews/gate-b-quality-pass-9.md b/.context/codex-reviews/gate-b-quality-pass-9.md new file mode 100644 index 0000000..f6a8e34 --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-pass-9.md @@ -0,0 +1,2 @@ +BLOCKER | high | docs/coding-workflow.md:261; docs/sparring-briefing.md:54; docs/hardening-log.md:106 | The health call does return mcp-codex-dev@1.0.1's cached per-root config, but `checks.config.effective.model` is not necessarily the model a gate uses: `getToolConfig` selects `tools.exec.model` or `tools.review.model` before the top-level model, and `CODEX_DEV_REVIEW_MODEL` is stored specifically at `tools.review.model` | A Gate-B pass using the documented review-only override can still be attributed to the wrong model, reproducing the reviewer-independence evidence defect this change is meant to close; the ledger row and commit evidence also overclaim the probe | For Gate A read `checks.config.effective.tools.exec.model` with fallback to `checks.config.effective.model`; for Gate B read `checks.config.effective.tools.review.model` with that fallback, define the no-explicit-model case, and amend every live copy plus the ledger and commit evidence +END OF FINDINGS (1 total) diff --git a/.context/codex-reviews/gate-b-quality-pr15-pass-1.md b/.context/codex-reviews/gate-b-quality-pr15-pass-1.md new file mode 100644 index 0000000..0b6fb8c --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-pr15-pass-1.md @@ -0,0 +1,3 @@ +BLOCKER | high | plugins/dev-workflow/.claude-plugin/plugin.json:4 | The range changes a shipped plugin command but leaves the manifest at the base's 0.7.0; `sh scripts/check-version-bump.sh 03901f3` fails, contrary to invariant 12 and the claimed green verification. | CI's PR-only version-bump check rejects this range, and an installed version-keyed cache would not receive the corrected command. | Bump the plugin version (normally to 0.7.1 for these fixes) and add the matching newest CHANGELOG entry. +MAJOR | high | docs/superpowers/plans/2026-07-26-risk-security-validation-profiles.md:602 | The accepted `git add -A` finding was replaced with the literal placeholder `git add `, not the explicit path list the implementer claims; this is not an executable command and leaves the agent to rediscover staging scope. | The instruction-bearing plan remains unfollowable at the snapshot step and can still lead to omitted intended files or unrelated staged files, so the original safety defect is not actually closed. | Replace the placeholder with the complete explicit list of intended paths for the snapshot, retaining the preceding `git status` check. +END OF FINDINGS (2 total) diff --git a/.context/codex-reviews/gate-b-quality-pr15-pass-2.md b/.context/codex-reviews/gate-b-quality-pr15-pass-2.md new file mode 100644 index 0000000..9dcf3ef --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-pr15-pass-2.md @@ -0,0 +1,2 @@ +MAJOR | high | docs/superpowers/specs/2026-07-26-risk-security-validation-profiles-design.md:394-395 | The new surface summary says `process-pr-review` keys the skip on the story profile “not on fix size,” but the settled rule and the shipped command require both an eligible story profile and a behaviourally trivial fix. | The approved spec now contradicts the shipped prompt on the exact two-condition rule this pass is meant to preserve, so a later implementation could legitimately drop the behavioural-triviality condition. | Rewrite the parenthetical to say the skip requires both behavioural triviality and every cited story's eligibility. +END OF FINDINGS (1 total) diff --git a/.context/codex-reviews/gate-b-quality-pr15-pass-3.md b/.context/codex-reviews/gate-b-quality-pr15-pass-3.md new file mode 100644 index 0000000..2a7fc06 --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-pr15-pass-3.md @@ -0,0 +1,6 @@ +MAJOR | high | docs/superpowers/plans/2026-07-26-risk-security-validation-profiles.md:15 | The Architecture summary says process-pr-review keys the skip on the story profile, omitting the independent requirement that the fix itself be behaviourally trivial | A future implementer following this plan can treat an eligible profile as sufficient and create the new gate-off path the settled design forbids | State that the skip requires both a behaviourally trivial fix and eligibility of every cited story +MAJOR | high | docs/superpowers/stories/2026-07-26-risk-security-validation-profiles-story.md:50 | AC 6 still specifies only effective-level eligibility and does not require the change itself to be behaviourally trivial | The implementation can satisfy the story acceptance criteria while allowing a behaviour-changing fix to skip Gate B | Add behavioural triviality as the first condition and retain effective level 0 for every cited profiled story as the second +MAJOR | high | plugins/dev-workflow/CHANGELOG.md:48 | The process-pr-review release note says the skip is decided from the story profile and “not from the fix's size alone,” but never says behavioural triviality remains mandatory | User-facing release documentation describes profile eligibility as the deciding condition and silently drops half of the shipped two-condition rule | Say the fix must be behaviourally trivial by effect and every cited story must be eligible +MAJOR | high | docs/superpowers/plans/2026-07-26-risk-security-validation-profiles.md:599 | The replacement for `git add -A` explicitly stages the implementation snapshot but omits `plugins/dev-workflow/hooks/codex-gate.sh`, even though the same plan's waiver requires that hook reminder change in this cycle | Following the plan leaves an intended plugin change outside the WIP commit and therefore outside the committed Gate-B range, so the closing commit can omit it or carry it unreviewed | Add `plugins/dev-workflow/hooks/codex-gate.sh` to the explicit staging list +MAJOR | high | docs/superpowers/plans/2026-07-26-risk-security-validation-profiles.md:675 | The amended fix-loop instruction claims `git status --short` lists exactly what the current fix touched, but it lists every dirty path in the worktree and cannot distinguish pre-existing user changes | An agent following the comment can fold unrelated dirty-worktree changes into the WIP amend, expanding the reviewed and eventually committed scope | Record the pre-fix status and compare it after the fix, then stage only paths attributable to the fix after inspecting their diffs +END OF FINDINGS (5 total) diff --git a/.context/codex-reviews/gate-b-quality-pr15-pass-4.md b/.context/codex-reviews/gate-b-quality-pr15-pass-4.md new file mode 100644 index 0000000..94b3627 --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-pr15-pass-4.md @@ -0,0 +1,3 @@ +MAJOR | high | docs/superpowers/plans/2026-07-26-risk-security-validation-profiles.md:676-683 | The fix loop tells the author, only after editing, to compare `git status --short` with “the status you recorded before the fix,” but no preceding step records that baseline | The attribution check cannot be followed as written, so an author can still stage an unrelated dirty path or omit a fix path while believing the reviewed range is complete | Add an explicit pre-fix `git status --short` capture before each accepted fix, then compare the post-fix output to that captured baseline before naming paths to `git add` +MINOR | high | plugins/dev-workflow/CHANGELOG.md:40-41; docs/superpowers/plans/2026-07-26-risk-security-validation-profiles.md:448-449,588-591,711-713 | Four summaries still describe the Gate-B skip only as narrowing to effective level 0 and leave the independent behavioural-triviality condition implicit in “triviality skip” | These mirrors preserve the exact dropped-condition wording class this pass is meant to eliminate and can again be copied as though profile eligibility alone authorizes a skip | State both conditions in each summary: the change is behaviourally trivial and every cited profiled story is at effective level 0 +END OF FINDINGS (2 total) diff --git a/.context/codex-reviews/gate-b-quality-pr15-pass-5.md b/.context/codex-reviews/gate-b-quality-pr15-pass-5.md new file mode 100644 index 0000000..a16246e --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-pr15-pass-5.md @@ -0,0 +1,2 @@ +MAJOR | high | docs/superpowers/plans/2026-07-26-risk-security-validation-profiles.md:682-684 | The new baseline capture is commented out, and even if the leading `#` were removed the comparison runs immediately with no fix step between capture and comparison; `/tmp/pre-fix-status` can therefore be missing or stale, while an existing dirty worktree is never meaningfully compared after the edit | The attribution step either reports an error hidden by `|| true`, compares against another run's file, or reports no delta, so the author cannot reliably distinguish this fix from pre-existing dirt and may stage the wrong paths | Make the capture an executable command using a freshly created temporary file, place an explicit `` step between capture and comparison, diff the post-edit status against that file, and clean the temporary file afterwards +END OF FINDINGS (1 total) diff --git a/.context/codex-reviews/gate-b-quality-pr15-pass-6.md b/.context/codex-reviews/gate-b-quality-pr15-pass-6.md new file mode 100644 index 0000000..bfb6950 --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-pr15-pass-6.md @@ -0,0 +1,3 @@ +MAJOR | high | docs/superpowers/plans/2026-07-26-risk-security-validation-profiles.md:683-690 | The baseline records only `git status --short`, so changing a path that was already dirty with the same status code leaves the before and after listings identical; the recipe's claim that the delta is the accepted fix is false for exactly the pre-dirty-worktree case it says it supports | The agent can omit an accepted fix from the explicit `git add`, amend the old tree, and re-review a range that does not contain the fix | Capture and compare content-derived snapshots (for example trees written through throwaway indexes, including untracked content) and derive changed paths from those trees, or require and verify a clean worktree before each fix +MAJOR | high | docs/superpowers/plans/2026-07-26-risk-security-validation-profiles.md:683-690 | The executable sequence does not stop when `mktemp`, the baseline `git status`, or `diff` fails; `diff` exit 1 is expected when a fix exists, but exit 2 is an error and is currently followed by cleanup and staging just like success | A failed capture or comparison can be silently treated as an authoritative delta, allowing the wrong paths to be staged and reviewed | Wrap baseline creation and capture in explicit success checks, preserve cleanup with a trap, and handle `diff` statuses explicitly so 0/1 are interpreted and any status greater than 1 aborts before staging +END OF FINDINGS (2 total) diff --git a/.context/codex-reviews/gate-b-quality-pr15-pass-7.md b/.context/codex-reviews/gate-b-quality-pr15-pass-7.md new file mode 100644 index 0000000..4456d41 --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-pr15-pass-7.md @@ -0,0 +1,4 @@ +MAJOR | high | plugins/dev-workflow/commands/process-pr-review.md:159-161 | The skip branch says to run the battery and record the skip reason, but only explicitly requires recording an evidence entry for profiled stories; it never says to record the battery result for an unprofiled story, unlike CLAUDE.md §5's explicit “records the reason and the battery result” rule. | A literal execution of this shipped command can close an unprofiled skip with no durable battery result, so the operational prompt diverges from the settled profiled/unprofiled evidence contract. | Rewrite the sentence to require recording the skip reason and battery result for every skipped cycle, then additionally require one mode-derived evidence entry per cited profiled story and none for unprofiled stories. +MINOR | high | docs/superpowers/stories/2026-07-26-risk-security-validation-profiles-story.md:50-57 | The amended acceptance criterion now states both skip-eligibility conditions but still records only the skip reason; it omits the settled evidence split that each cited profiled story owes an evidence entry while an unprofiled story owes no mode-derived entry. | The story can read as satisfied by an implementation that gets eligibility right but loses the author-evidence requirement, leaving its source requirement weaker than the spec, plan, and §5. | Extend this criterion with the profiled/unprofiled evidence split and the always-run battery requirement. +MINOR | medium | docs/getting-started.md:80-83 | “what a trivial profile unlocks is the Gate-B skip” makes behavioural triviality only implicit in the preceding sentence and frames effective level 0 as the thing that unlocks the skip, even though the settled rule says the profile supplies only one of two independent conditions and creates no gate-off path. | A reader skimming the operative sentence can treat an eligible profile as sufficient and skip Gate B for a behaviour-changing diff. | State in the same sentence that Gate B may be skipped only when the change is behaviourally trivial and the profiled story is at effective level 0; describe the profile as supplying eligibility, not unlocking the skip. +END OF FINDINGS (3 total) diff --git a/.context/codex-reviews/gate-b-quality-pr15-pass-8.md b/.context/codex-reviews/gate-b-quality-pr15-pass-8.md new file mode 100644 index 0000000..81eb5de --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-pr15-pass-8.md @@ -0,0 +1,2 @@ +MAJOR | high | docs/superpowers/specs/2026-07-26-risk-security-validation-profiles-design.md:375-380 | §7 still says an unprofiled story owes “no commit-body entry,” contradicting §5 lines 275-279, process-pr-review lines 159-162, and story AC 6, which require every skipped unprofiled cycle to record its skip reason and battery result in the commit body | An implementer following the compatibility section can close an unprofiled skip without the durable battery result this pass was meant to require, so the four stated sources do not yet agree | Replace “no commit-body entry” with “no mode-derived evidence entry,” and explicitly preserve the skip reason plus battery-result record for an unprofiled skip +END OF FINDINGS (1 total) diff --git a/.context/codex-reviews/gate-b-quality-pr15-pass-9.md b/.context/codex-reviews/gate-b-quality-pr15-pass-9.md new file mode 100644 index 0000000..8d781f4 --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-pr15-pass-9.md @@ -0,0 +1,2 @@ +NO FINDINGS +END OF FINDINGS (0 total) diff --git a/.context/codex-reviews/gate-b-quality-pr15b-pass-1.md b/.context/codex-reviews/gate-b-quality-pr15b-pass-1.md new file mode 100644 index 0000000..8d781f4 --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-pr15b-pass-1.md @@ -0,0 +1,2 @@ +NO FINDINGS +END OF FINDINGS (0 total) diff --git a/.context/codex-reviews/gate-b-quality-rle-pass-1.md b/.context/codex-reviews/gate-b-quality-rle-pass-1.md new file mode 100644 index 0000000..fde52e7 --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-rle-pass-1.md @@ -0,0 +1,3 @@ +BLOCKER | high | CLAUDE.md:920-924; plugins/dev-workflow/commands/workflow-init.md:1104-1108 | the rule says to capture the full reviewed head commit from each branch result, but mcp-codex-dev@1.0.1 returns no head SHA and the required findings-file and reply formats provide no field for one; the existing call guidance also permits symbolic headSha = HEAD | separate spec and quality calls can resolve HEAD to different commits while both leave the agent with no result-bound object name to compare, so the branches can be summed into one logical pass despite reviewing different revisions | resolve HEAD to a full 40-character object name before every call, pass that explicit value as headSha, retain the invocation value with its result, and compare those captured values before summing branches +BLOCKER | high | plugins/dev-workflow/CHANGELOG.md:53-61 | the 0.11.0 entry claims nonce-less cycles gain a deterministic slot discriminator when bare slots already exist, but the shipped rules reserve bare names for every no-nonce cycle and contain no discriminator procedure; the implementer explicitly deferred that rule | users are told this release prevents the exact bare-slot overwrite that remains possible, so the release notes overstate the shipped isolation contract and contradict the production prompt | remove the discriminator and own-path claims from 0.11.0, or ship the deferred discriminator rule in both prompt copies with its required semantics and parity +END OF FINDINGS (2 total) diff --git a/.context/codex-reviews/gate-b-quality-rle-pass-2.md b/.context/codex-reviews/gate-b-quality-rle-pass-2.md new file mode 100644 index 0000000..c5d4414 --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-rle-pass-2.md @@ -0,0 +1,14 @@ +BLOCKER | high | CLAUDE.md:654-665; plugins/dev-workflow/commands/workflow-init.md:840-851 | The advertised four-case profile procedure is neither exhaustive nor mutually exclusive: a cited path that resolves to nothing satisfies neither case 3, which requires the path to exist, nor case 4, which requires a profile to be present, even though case 4 lists that failure | A missing citation has no defined answer, so agents can silently default, invent a fifth case, or disagree on a stop condition that controls the review floor | Define one path-resolution case covering missing paths, broken symlinks, non-regular files, and unreadable files; reserve the present-profile case for syntactic and semantic profile failures, and mirror the exact procedure in both copies +NIT | high | plugins/dev-workflow/commands/process-pr-review.md:159-161 | The cross-reference calls a present-but-unresolvable profile section 5's third case, but the changed procedure makes it the fourth case | The reader is sent to the unreadable-path case instead of the malformed-profile case, weakening an operational instruction whose distinction now matters | Update the reference to the fourth case, or replace the ordinal with the stable name of the rule +BLOCKER | high | CLAUDE.md:856-863; plugins/dev-workflow/commands/workflow-init.md:1040-1047; plugins/dev-workflow/hooks/codex-gate.sh:119-127 | The non-numeric diagnostic tests content after trimming only trailing newlines, while the hook deletes all whitespace before parsing; for example, bytes containing `1 2` are documented as unusable but make the hook use threshold 12 | The provenance record can state the wrong knob value and wrong cause, so its promised decidable check and cause-specific fix do not describe the mechanism it records | Define classification from the hook's exact all-whitespace normalization, or change the hook and its tests to implement the documented raw-content rule +BLOCKER | high | CLAUDE.md:856-863; plugins/dev-workflow/commands/workflow-init.md:1040-1047; plugins/dev-workflow/hooks/codex-gate.sh:119-127 | The out-of-range cause says a value exceeds the hook's accepted maximum, but the hook defines no maximum and delegates numeric acceptance to the shell's integer comparison | The cause is not portably decidable and `lower it` is not an actionable fix; the same oversized decimal can be accepted or rejected by different shell implementations | Define and enforce an explicit portable maximum in the hook and both prompt copies with an exact check and fix, or replace this cause with an accurate comparison-failure state that names the environment-dependent check +BLOCKER | high | CLAUDE.md:863-866,903-910; plugins/dev-workflow/commands/workflow-init.md:1047-1050,1087-1094 | The rejected-model rule requires prose saying which bytes were rejected while also saying the raw value must not be reproduced, but it defines no safe canonical representation, distinguishing check, or cause-specific fix | An agent can embed the unsafe control characters while trying to identify them, omit the required information, or produce prose that a later reader cannot compare, violating the diagnostic-state standard | Require a deterministic safe representation such as lowercase hexadecimal octets, explicitly prohibit raw control characters in that representation, and give the check, fix, and a filled example +BLOCKER | high | CLAUDE.md:804-809,936-944; plugins/dev-workflow/commands/workflow-init.md:988-993,1120-1128 | The retained old routing bullet still defines `headSha` as symbolic `HEAD`, contradicting the new rule to resolve and pass a full 40-character object name immediately before every call | An agent can follow the earlier instruction and retain equal strings reading `HEAD` even when the two calls resolve to different commits, defeating the revision-separation rule | Rewrite the earlier bullet to require the full object name resolved immediately before each call and reserve symbolic `HEAD` only for explaining how `baseSha` is computed +BLOCKER | high | CLAUDE.md:932-946; plugins/dev-workflow/commands/workflow-init.md:1116-1130; docs/superpowers/specs/2026-08-28-review-loop-economics-design.md:309-317 | The text says both branches reviewed the same commit or artifact revision, but the stated mechanism compares only retained `headSha` request inputs and explicitly receives no reviewed revision in the result | Equal request targets prove only that the calls were asked to review the same object name, not that the reviewer read it; summing branches under the stronger claim violates the repository rule against describing more than a gate actually compares | State that both branches must have been requested against the same full object name and that equality proves only equal request targets; if actual reviewed-revision identity is required, add an attestation mechanism that reports and validates it +BLOCKER | high | docs/getting-started.md:51-58 | `Verification is by content` implies that Gate B verifies the included bytes were reviewed, while the hook only compares the current content fingerprint with the fingerprint recorded on a counted call | Users can treat fingerprint equality as evidence that Codex read those bytes, repeating the exact gate-proof overclaim prohibited by AGENTS.md | Say that the hook verifies only fingerprint equality and a fresh counted-call streak, and explicitly state that this does not prove Codex read the fingerprinted bytes +BLOCKER | high | CLAUDE.md:826-849,883-916; plugins/dev-workflow/commands/workflow-init.md:1010-1033,1067-1100; docs/prompt-standards.md:44-45 | The new provenance, curve, and skip output contracts provide grammars but no filled valid examples | Prompt-standard item 4 requires output structure with an example; without instances, the shipped prompt leaves quoting, singleton-model, split-pass, unavailable-count, and skip realizations to inference and risks incompatible durable records | Add concrete valid examples for the provenance line, uniform-model and per-pass or split curve forms, an unavailable count, and a skipped cycle, including at least one quoted value +MAJOR | high | CLAUDE.md:96-113; plugins/dev-workflow/commands/workflow-init.md:303-320 | The rule mandates one path per entry when a `Story:` header cites several stories but never defines whether entries are repeated headers, a delimiter-separated field, Markdown list items, or how paths are quoted | The cited set controls every gate floor, so two conforming agents can parse the same multi-story header differently, omit a higher-risk story, or stop a valid artifact | Pin one multi-entry header grammar, including delimiter, repetition, quoting, escaping, duplicate handling, and whitespace, and show valid one-story and multi-story examples plus a malformed example +MINOR | high | CLAUDE.md:886-900; plugins/dev-workflow/commands/workflow-init.md:1070-1084; docs/superpowers/specs/2026-08-28-review-loop-economics-design.md:232-246 | The model grammar permits a lone model as an alternative to per-pass mappings but never states that the lone value applies to every pass enumerated by the pass specification | A record intended to identify the model for each pass has ambiguous singleton semantics, and a future parser cannot know whether the shorthand applies to all passes or only one unlabeled pass | State explicitly that the lone-model branch applies the same model to every enumerated pass and include a multi-pass example using that shorthand +MINOR | high | docs/superpowers/specs/2026-08-28-review-loop-economics-design.md:131-135 | The spec says the deferred P8 measurement `parses` the provenance form in the present tense, while the release notes and prompt correctly say no parser exists today | The implementation status is internally contradictory and can make reviewers mistake a planned consumer for validation already supplied by this release | Change the sentence to say the deferred measurement is intended to parse or will parse the form +MINOR | high | plugins/dev-workflow/CHANGELOG.md:53-60 | The release-note headline says findings slots take a per-cycle infix, then immediately documents that nonce-less cycles keep bare names and can still collide | The headline overstates the collision isolation shipped in 0.11.0 and obscures the exact residual the same entry says is deferred | Qualify the claim to nonce-holding cycles and state up front that legacy or nonce-less cycles retain bare slots and lack the per-cycle deletion boundary +END OF FINDINGS (13 total) diff --git a/.context/codex-reviews/gate-b-quality-rle-pass-3.md b/.context/codex-reviews/gate-b-quality-rle-pass-3.md new file mode 100644 index 0000000..b7187bc --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-rle-pass-3.md @@ -0,0 +1,10 @@ +BLOCKER | high | CLAUDE.md:109-116; plugins/dev-workflow/commands/workflow-init.md:316-323; docs/superpowers/specs/2026-08-28-review-loop-economics-design.md:4; docs/superpowers/plans/2026-08-29-review-loop-economics-plan-a-rules.md:14; docs/superpowers/plans/2026-08-30-review-loop-economics-plan-b-records.md:17; docs/superpowers/plans/2026-08-30-review-loop-economics-plan-c-rollout.md:14 | The new Story-header grammar accepts only bare or double-quoted paths, but the governed spec and all three contributing plans wrap their paths in Markdown backticks, which are neither form | Once the new rules ship, the kit's own normal artifacts are malformed and every Gate-A or Gate-B derivation over them must stop instead of deriving a floor | Either admit the established backtick code-span spelling and define any trailing-text policy, or atomically migrate every existing header and every producer to the newly pinned spelling +MAJOR | high | CLAUDE.md:661-666; plugins/dev-workflow/commands/workflow-init.md:847-852; docs/superpowers/plans/2026-08-29-review-loop-economics-plan-a-rules.md:110 | The four unreadable-story sub-causes are not distinguishable in the stated existence, type, then link-resolution order: ordinary existence and regular-file tests follow symlinks, so a dangling symlink is classified as nonexistent or non-regular before the link check | The promised mutually exclusive diagnostic procedure is not executable as written, so agents can report the wrong cause and apply the wrong fix; the accounting row's claim of a decidable order is also false | Define an explicit lstat or symlink-first procedure, then test target existence and readability, and update the accounting row to match the actual partition +MAJOR | high | CLAUDE.md:870-884; plugins/dev-workflow/commands/workflow-init.md:1054-1068; plugins/dev-workflow/hooks/codex-gate.sh:119-127 | The knob paragraph says its four causes are classified by what the hook does, but the hook first gates on -f, so a broken symlink never reaches the read, while a failed cat in the cat-to-tr pipeline yields the same empty value as an empty file because only tr's status survives | Provenance can label states as unreadable that the claimed mechanism treats as absent or empty, making the pinned cause field an inaccurate account of hook behavior and leaving no check that reliably separates the advertised causes | Either specify a separate provenance classifier with explicit symlink, readability, read-status, stripping, and integer-comparison checks, or align the cause tokens and description with the hook's observable branches +MAJOR | high | plugins/dev-workflow/commands/process-pr-review.md:148-161 | After correctly saying absence in PR and commit records cannot substitute for the governing Story-header set, the next sentence again declares that absence from those two records to be section 5's no-story-cited case | The command contains opposite rules for the same skip decision and can let non-authoritative PR metadata relax a profile-governed Gate-B obligation | Delete the surviving section-5 equivalence and state only that record absence selects this command's local unprofiled judgment without establishing or changing the governing set +MAJOR | high | CLAUDE.md:912-937; plugins/dev-workflow/commands/workflow-init.md:1096-1121 | The pinned curve production can emit only the token undetermined, while the accompanying rule requires the body to name one of three causes and, for UNREPRESENTABLE, record byte length and rejected offsets; no production or keyed companion form says where or how those values are encoded | Different writers cannot produce one parseable form, and the deferred metrics consumer cannot associate the required diagnostics with a pass, so the record is not actually pinned for these states | Extend the grammar with a per-pass diagnostic form carrying the cause and required metadata, define its association and ordering, and add a filled undetermined example that satisfies it +MAJOR | high | CLAUDE.md:899-944; plugins/dev-workflow/commands/workflow-init.md:1083-1128 | The skip record is only cycle-field plus skipped-see-skip-reason, but no grammar, identifier, or placement rule links it to one particular free-form reason even though the squash rule says to carry the reason it points at | A commit containing several skipped cycles can contain several reasons with no decidable mapping, so the mandated record and squash carry are ambiguous | Put the reason directly in a pinned escaped field on the skip record, or define a cycle-field-keyed reason record and require exactly one matching reason per skipped cycle +MAJOR | high | CLAUDE.md:968-981; plugins/dev-workflow/commands/workflow-init.md:1152-1165; docs/superpowers/stories/2026-08-28-review-loop-economics-pass-floor-story.md:170-173; docs/superpowers/specs/2026-08-28-review-loop-economics-design.md:309-317; docs/superpowers/plans/2026-08-30-review-loop-economics-plan-b-records.md:487-498; docs/superpowers/plans/2026-08-29-review-loop-economics-plan-a-rules.md:111 | The revised mechanism establishes only that two requests carried equal headSha values, yet its closing sentence still describes merging two revisions and the story, spec, and Plan B still require or claim that both branches reviewed the same revision; Plan A's accounting says the condition was narrowed without accounting for those surviving sources | The shipped prompt still overstates what the request comparison proves and the implementation no longer satisfies the cited story criterion or design requirement | Replace every reviewed-same-revision claim with issued-against or aimed-at the same requested commit and explicitly preserve the no-review-evidence limit, or add an independently verified reviewed-revision attestation that meets the existing requirement +MAJOR | high | docs/superpowers/plans/2026-08-30-review-loop-economics-plan-c-rollout.md:970-980; docs/superpowers/plans/2026-08-30-review-loop-economics-plan-c-rollout.md:1047; docs/superpowers/plans/2026-08-30-review-loop-economics-plan-c-rollout.md:1292-1308 | The new note drops Tasks 19 and 20 and says the rle form is only a plan-local exception, but Task 20's heading remains executable and Task 23 still says those tasks admit rle into the shipped rule | A later executor entering at Task 20 or following Task 23 can ship the explicitly rejected discriminator production, reopening the slot-safety and source-of-truth contradiction the drop was meant to close | Mark Task 20 and its body as dropped, update Task 23 to cite the plan-local exception directly, and revise the task accounting and earlier discriminator prose so no live instruction says the shipped grammar admits rle +MINOR | high | plugins/dev-workflow/commands/process-pr-review.md:164-166 | The command calls a present-but-unresolvable profile section 5's third case, but the revised four-case taxonomy makes it case 4; case 3 is an unreadable story path | The prescribed stop remains correct, but the wrong cross-reference sends readers to the wrong diagnostic procedure | Change third case to fourth case +END OF FINDINGS (9 total) diff --git a/.context/codex-reviews/gate-b-quality-rle-pass-4.md b/.context/codex-reviews/gate-b-quality-rle-pass-4.md new file mode 100644 index 0000000..1921e0e --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-rle-pass-4.md @@ -0,0 +1,10 @@ +BLOCKER | high | CLAUDE.md:871-875; plugins/dev-workflow/commands/workflow-init.md:1054-1059; plugins/dev-workflow/hooks/codex-gate.sh:119-127 | The surviving load-bearing claim that every unusable knob value leaves the hook default standing is false: the hook captures the file through shell command substitution, which drops NUL bytes, so a file containing the bytes `1` followed by NUL is accepted as floor 1 under both `sh` and `dash` instead of leaving floor 3 | The provenance observer can correctly classify the raw file as unusable while the hook silently uses a non-default threshold, so the new record and explanatory claim disagree with actual behavior and violate the mechanism-claim invariant | Either remove or narrow the universal claim and disclose the normalization gap, or expand scope to make the hook read and validate bytes without losing NUL and add the corresponding `sh` and `dash` regression case +MAJOR | high | CLAUDE.md:858-875; plugins/dev-workflow/commands/workflow-init.md:1042-1059; docs/prompt-standards.md:53-61 | Deleting the knob walkthrough left the provenance grammar still requiring one of four `` values and saying they need different fixes, but the shipped prompts now provide neither checks that distinguish those values nor the fixes; pointing abstractly to the hook does not tell the record-writing observer how to classify the raw file | Downstream agents cannot deterministically produce the pinned required field, and the prompt now fails checklist item 10 for a diagnostic state with multiple causes | Restore an explicit observer-side classifier with mutually exclusive checks and a fix per cause, or redesign the record to carry only a state the observer can determine without the deleted procedure +MAJOR | high | CLAUDE.md:661-666; plugins/dev-workflow/commands/workflow-init.md:847-852 | The unreadable-story diagnostic still says to distinguish nonexistence, non-regular type, and broken symlink by testing existence, type, then link resolution; ordinary existence and regular-file tests follow symlinks, so a dangling symlink is consumed by an earlier branch and its dedicated cause is unreachable | Agents report the wrong cause and remedy for a broken citation, defeating the reason the four sub-causes were split and leaving prompt-standards item 10 unsatisfied | Test symlink identity or perform an lstat-equivalent check before target existence and regular-file readability, then update the stated order and the Plan A accounting row +MINOR | high | CLAUDE.md:657-682; plugins/dev-workflow/commands/workflow-init.md:843-868 | The heading promises a four-case profile partition, but the four entries omit the normal state where the story is readable and its present profile resolves successfully | The surrounding rules imply what happens on that path, but this purported decision table is not exhaustive and invites readers to treat a valid profile as an unspecified fifth state | Add the readable-and-resolvable case with its proceed-and-derive action, or label the list explicitly as exceptional states rather than a complete partition +MAJOR | high | CLAUDE.md:905-924; plugins/dev-workflow/commands/workflow-init.md:1089-1108; docs/superpowers/specs/2026-08-28-review-loop-economics-design.md:242-256; docs/superpowers/plans/2026-08-30-review-loop-economics-plan-b-records.md:451-465 | The model-cause deletion did not leave one pinned contract: the shipped copies make `undetermined` cover both an unknown identifier and an unrepresentable one and make its explanation optional free prose, then still say the body records the source and rejected bytes; the spec and Plan B retain the older control-character-only rule and mandatory source-and-byte account | Writers cannot know whether diagnostics are required, a genuinely unknown identifier has no rejected bytes to report, and the authoritative spec, implementation plan, and shipped record grammar now disagree | Choose sentinel-only or keyed diagnostics, remove the contradictory leftover sentence, and update the spec, Plan B, and both prompt copies to exactly the same semantics +MAJOR | high | CLAUDE.md:890-931; plugins/dev-workflow/commands/workflow-init.md:1074-1115 | A skipped-cycle record is still only `; : skipped (see skip reason)`, while the skip reason is free-form and carries no cycle field, identifier, adjacency rule, or other mapping | A commit containing multiple skipped cycles can contain multiple reasons with no decidable way to tell which reason the squash-carry rule says each skip record points at | Put the reason in an escaped field on the skip record, or define a cycle-field-keyed skip-reason record with exactly one match per skipped cycle +MAJOR | high | CLAUDE.md:815-818,951-968; plugins/dev-workflow/commands/workflow-init.md:999-1002,1135-1152 | Logical-pass agreement compares only the captured `headSha`, although Gate B reviews the range selected by both `baseSha` and `headSha`; separate branch calls or a single-branch recovery can therefore use equal heads with different bases and still be summed | The curve can merge findings from different diffs into one logical pass even though the new rule claims equality identifies one reviewed artifact | Capture and compare the full `(baseSha, headSha)` pair for every branch, keep both values with each result, and treat a mismatch in either value as an incomplete pass boundary +MAJOR | high | docs/superpowers/plans/2026-08-30-review-loop-economics-plan-c-rollout.md:970-980,1047-1053,1297-1301,1427-1443 | Tasks 19 and 20 are marked dropped, but live Task 23 and Self-Review still state that those tasks add the deterministic `rle` production to the shipped grammar; the new notes merely tell readers to reinterpret those later instructions as superseded | The plan still gives mutually contradictory instructions at the exact task being executed and falsely claims its shipped slot grammar admits the current cycle's filenames | Rewrite Task 23 and Self-Review to name the approved plan-local exception directly and remove every assertion that the dropped tasks change the shipped rule +NIT | high | docs/superpowers/plans/2026-08-30-review-loop-economics-plan-c-rollout.md:979 | The Task 19 drop note ends with an orphaned `§5` token | The fragment looks like a truncated cross-reference and adds avoidable ambiguity to an already dense supersession note | Delete the stray token or complete the intended citation +END OF FINDINGS (9 total) diff --git a/.context/codex-reviews/gate-b-quality-rle-pass-5.md b/.context/codex-reviews/gate-b-quality-rle-pass-5.md new file mode 100644 index 0000000..9a5f304 --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-rle-pass-5.md @@ -0,0 +1,11 @@ +BLOCKER | high | CLAUDE.md:657-686; plugins/dev-workflow/commands/workflow-init.md:843-872 | The advertised five-case profile partition is not mutually exclusive: case 0 is vacuously true for an empty cited set and overlaps case 1, while a multi-story set can simultaneously contain an unprofiled member, an unreadable member, and a readable-but-unresolvable member and therefore match several cases with conflicting run and stop answers | The profile reader can choose the permissive branch for a set that also contains a mandatory-stop condition, changing whether the gate runs and violating the claimed partition | Define predicates over the whole set with precedence: no citations; any unreadable member; otherwise any present-but-unresolvable profile; otherwise any unprofiled member; otherwise a non-empty set whose profiles all resolve +BLOCKER | high | CLAUDE.md:862-877; plugins/dev-workflow/commands/workflow-init.md:1046-1061; docs/superpowers/specs/2026-08-28-review-loop-economics-design.md:130-139,476-481; docs/superpowers/plans/2026-08-30-review-loop-economics-plan-b-records.md:381-392; docs/superpowers/plans/2026-08-30-review-loop-economics-plan-c-rollout.md:1115-1126,1214-1226 | Removing the cause token did not leave a producible knob field: the shipped prompts still require a writer to choose a positive integer or `unusable` but give no rule for a readable file containing nonnumeric, nonpositive, empty, or binary content, while Plan B still says to name the cause and the spec and Plan C still require coverage or classification of the removed unusable causes | Every cycle owes this provenance field, so an observer can be forced to guess, emit a token outside the pinned grammar, or follow the stale cause contract; this is the stranded classification the cut was intended to eliminate | Either define and test one observer-owned coarse predicate that does not claim to reproduce hook behavior, or narrow/remove the knob field; then delete every remaining cause-level duty from the spec and all three plans +MAJOR | high | docs/superpowers/specs/2026-08-28-review-loop-economics-design.md:130-137 | The withdrawal note still states that a file containing `1`, NUL, `2` is accepted by the hook as twelve under three shells | This is an explicit surviving description of what the hook accepts, despite the stated decision to delete this repeatedly corrected mechanism surface; it creates the fifth copy that can drift and defeats prompt-standards item 11's withdrawal rule | Replace the behavioral example with a mechanism-free statement that the previous classification was falsified by an observed counterexample and was therefore withdrawn +BLOCKER | high | CLAUDE.md:907-922; plugins/dev-workflow/commands/workflow-init.md:1091-1106; docs/superpowers/specs/2026-08-28-review-loop-economics-design.md:246-260; docs/superpowers/plans/2026-08-30-review-loop-economics-plan-b-records.md:453-467 | The model-failure cut is incomplete: both shipped copies still say a control character makes an identifier unrepresentable and that an undetermined or unrepresentable identifier becomes `undetermined`, immediately before claiming nothing describes how failure happens; the spec and Plan B retain the older control-character rule and additionally require source and rejected-byte prose for which the grammar has no field | The pinned curve contract disagrees across its authority, plan, and shipped copies, so writers cannot know when `undetermined` applies or whether diagnostic prose is mandatory | Keep only the sentinel's cause-free meaning and raw-value safety rule in every copy, deleting the failure classification and source/rejected-byte duty from the prompt, spec, and Plan B +BLOCKER | high | CLAUDE.md:929-932; plugins/dev-workflow/commands/workflow-init.md:1113-1116; docs/superpowers/specs/2026-08-28-review-loop-economics-design.md:268-271; docs/superpowers/plans/2026-08-30-review-loop-economics-plan-b-records.md:469 | The shared cycle field is not a sufficient or pinned link from a skip record to its reason: every pre-rule cycle uses the same `cycle none (pre-rule)` value, nonce collisions are explicitly admitted, no skip-reason production or example says where its cycle field goes, and the spec and Plan B do not carry the new link rule at all | A commit with multiple pre-rule skips or colliding fields cannot resolve records to reasons, so the claim that one record and one reason per cycle makes pairing unambiguous is false and the squash-carry instruction can preserve the wrong reason | Pin a skip-reason record and example in the spec, plan, and prompts with an additional deterministic pairing rule such as immediate adjacency or an artifact/occurrence key, and state the residual for collisions instead of claiming unconditional uniqueness +MAJOR | high | CLAUDE.md:952-969; plugins/dev-workflow/commands/workflow-init.md:1136-1153; docs/superpowers/specs/2026-08-28-review-loop-economics-design.md:313-323; docs/superpowers/stories/2026-08-28-review-loop-economics-pass-floor-story.md:170-175; docs/superpowers/plans/2026-08-30-review-loop-economics-plan-b-records.md:489-502 | Logical-pass agreement still retains and compares only `headSha`, although Gate B reviews the range selected by both `baseSha` and `headSha`; two branches can use the same head with different bases and still be summed | The curve can merge findings from different diffs into one logical pass, so equality of the retained value does not establish that the calls were aimed at the same reviewed artifact | Capture, retain, and compare the full `baseSha` plus `headSha` pair for every branch and recovery, and end the logical pass when either component differs +MINOR | high | docs/superpowers/plans/2026-08-29-review-loop-economics-plan-a-rules.md:102-111 | The replacement-accounting table says nothing outside its eleven passages was edited, then adds rows 12 and 13 for two additional edited procedures; row 12 also says case 3 gained a decidable test order even though the shipped rule now expressly prescribes no order | The plan's required kept/moved/dropped audit is internally false, so it cannot serve as evidence that the rewrites preserved every old condition | Recast the eleven-passage statement as historical, count all thirteen current rows, and update row 12 to record the deliberate no-order rule and its rationale +MINOR | high | plugins/dev-workflow/CHANGELOG.md:42-48 | The release note says every cycle gets one provenance line and one per-pass curve, but a skipped cycle gets a skip record with no counts instead of a curve | Users reading the release contract are told a record exists that the shipped prompt explicitly tells skipped cycles not to write | Say that every cycle writes one provenance line and either a per-pass curve or a skip record +NIT | high | docs/superpowers/plans/2026-08-30-review-loop-economics-plan-c-rollout.md:972-980,1049-1052 | The dropped-task notes still say Task 23 and Self-Review claim Tasks 19 and 20 ship the `rle` production even though both later sites were corrected, and the Task 19 note ends with an orphaned `§5` token | These stale supersession instructions and fragment make the plan appear internally inconsistent after the claimed cleanup | Remove the obsolete warnings about the now-corrected later sites and delete or complete the stray `§5` +NIT | high | plugins/dev-workflow/commands/process-pr-review.md:163-167 | The command calls the present-but-unresolvable branch "§5's fourth case", but the zero-based list labels it case 4 and places it fifth | The inline condition prevents a behavioral error, but the ordinal cross-reference is ambiguous precisely because the cases were renumbered from zero | Refer to it as `§5 case 4` rather than the fourth case +END OF FINDINGS (10 total) diff --git a/.context/codex-reviews/gate-b-rle-pass-5-dispositions.md b/.context/codex-reviews/gate-b-rle-pass-5-dispositions.md new file mode 100644 index 0000000..8f07d81 --- /dev/null +++ b/.context/codex-reviews/gate-b-rle-pass-5-dispositions.md @@ -0,0 +1,57 @@ +# Gate-B cycle `rle` — dispositions at close + +The cycle closed on §5's **clearly-stuck** exit (Daniel, 2026-09-02), not on a clean pass. +Every pass-5 finding gets a line. Four were repaired in the closing round because they were +unambiguously defects of mine; the rest are recorded open, each with why. + +## Repaired in the closing round +- spec-2 / quality-5 — the shared `` does not link a skip record to its reason, + since every pre-rule cycle writes `cycle none (pre-rule)`. **Fixed**: adjacency is the link — + the reason is the text immediately following the record — and the cycle-field claim is + withdrawn in as many words. +- spec-5 / quality-4 — the model-failure cut was incomplete. **Fixed**: the source-and- + rejected-bytes obligation is withdrawn in the spec and in Plan B, dated. +- spec-10 — Task 23 requires the `rle` exception in the closing body while Task 24's "complete + list" omitted it. **Fixed**: the list now carries the exception and the standing decision + record, items 5 and 6. +- spec-9 — the dropped Tasks 19/20 were half-propagated. **Fixed**: Plan C's accounting row 19 + and its Task-12 changelog text no longer say the release ships the discriminator. + +## Declined by human decision 2026-09-02 — a chosen cost, not a missed defect +- spec-3 — "`unusable` and `undetermined` intentionally collapse distinct causes and provide + neither a discriminating check nor a per-cause fix." Correct, and deliberate. Four successive + attempts to state the classification behind those tokens were each wrong in a different way, + the last demonstrably so. The record now says THAT a knob was unusable and no longer WHY. +- spec-4 — "Profile case 3 says its four causes need different fixes but supplies no per-cause + fixes and expressly declines to prescribe a discriminating check." Same decision. The order + previously prescribed could not reach the broken-symlink cause, because ordinary existence and + regular-file tests follow symlinks. +Both are the requirement Daniel withdrew. A gate cannot clear a finding whose resolution the +human has declined, which is the require-withdraw pair that made this stop mandatory. + +## Open at close, with reasons +- spec-1 / quality-1 — the five-case partition is not mutually exclusive for **mixed** cited + sets: case 0 is vacuously true of an empty set, and a multi-story set can satisfy several + cases at once. Real, and the third distinct defect found in this one list across three + rounds. Owner: the loop-rule consolidation story. +- quality-2 — removing the cause token left no rule for a readable file whose content is + neither a positive integer nor absent, so the writer has no producible value. Real; a + consequence of the withdrawal above and inseparable from it. +- quality-3 — the withdrawal note explaining why the hook description was deleted is itself a + description of the hook. Correct, and self-consuming: no version of that paragraph survives + its own rule. Recorded rather than attempted a fifth time. +- spec-7 — the shipped nonce-recovery rule narrows the spec's single-candidate rule. +- spec-8 — the story requires every cycle that ran to carry its curve; Plan C's closing body + excludes the four pre-rule Gate-A cycles, whose record is the field report instead. +- quality-6 — logical-pass agreement compares only `headSha`, though a range is selected by + `baseSha` too. +- spec-11, spec-12, spec-13, quality-7, quality-8, quality-9, quality-10 — Minor and Nit: + stale accounting sentences, a release note that omits the skip-record case, and a "fourth + case" label left over from the pre-zero-based numbering. Collected, not iterated. + +## Why the cycle closed here +Five passes: 16/9 · 29/17 (discounted) · 25/15 · 25/15 · 23/16 findings/Blocker+Major. +Blocker+Major never returned to its pass-1 level and rose on the last pass. Each round's +repair produced the next round's findings on the same mechanism. Coverage is affirmatively +sufficient: across five passes the reviewers covered both prompt copies, the spec, the story, +all three plans, the hook source and the user docs; no materially unreviewed area is known. diff --git a/.context/codex-reviews/gate-b-rle-resume.md b/.context/codex-reviews/gate-b-rle-resume.md new file mode 100644 index 0000000..2f87eb7 --- /dev/null +++ b/.context/codex-reviews/gate-b-rle-resume.md @@ -0,0 +1,66 @@ +# Gate-B cycle `rle` — resume note (cycle-stable, per CLAUDE.md §5) + +Advisory human note. Not a findings file; participates in no pass validation. + +**Artifact:** the combined A+B+C diff. **Base:** ab8ac98ef7eccf06d9d24a74f201b06016228dd5 +(fixed at cycle open, never moved). **WIP tip at write time:** 203cf6b. +**Slots:** `gate-b--rle-pass-

.md` — a RECORDED plan-local naming exception +under the old rules, approved by Daniel 2026-09-02. Reason: this cycle is pre-rule and cannot +mint a nonce, and the bare family already holds 30 files that delete-before-call would destroy +(re-inventoried; the plan's "61" was stale, and the plan said to re-count rather than trust it). + +## Pass ledger +| pass | findings | Blocker | Major | counted? | +|---|---|---|---|---| +| 1 | 16 (spec 14 + quality 2) | 4 | 5 | yes | +| 2 | 29 (spec 16 + quality 13) | 15 | 2 | **NO** — both files structurally valid, but the reply contradicted itself: each reviewer reported the other branch INCOMPLETE, mistaking its counterpart's legitimate file for a foreign write. Findings acted on; pass not credited | +| 3 | 25 (spec 16 + quality 9) | 6 | 9 | yes | + +Counted passes: 2 of a floor of 3. + +## STATUS AT PASS 3: stopped and surfaced. + +### The three lines +1. TREND — findings 16 -> 29 -> 25; Blocker+Major 9 -> 17 -> 15; Blockers 4 -> 15 -> 6. + Nothing has returned to its pass-1 level. +2. CLUSTER — two groups, and both are about corrections rather than about the original work. + (a) **My own fixes**: the knob description, the profile-case partition, the model-cause + enumeration, the Story-header syntax, the process-pr-review clarification. Each was + introduced or rewritten by a previous pass and is wrong again in a subtler way. + (b) **Upstream artifacts the shipped prompts now contradict**: the story, the spec at + revision 36, Plan B and Plan C still assert what the prompts were corrected to stop + asserting. +3. REQUIRE-WITHDRAW — no clean pair. Pass 3 refines what pass 2 demanded rather than + withdrawing it. + +### Tells +- Finding count failing to fall: **present** (16 -> 29 -> 25). +- Blocker count failing to fall: **present** (4 -> 15 -> 6). +- Instrument cluster: N/A, no test instrument. +- Prose-about cluster: **present** — a large share of pass 3 is about the plan, spec and story + artifacts rather than the shipped prompts. +- Require-withdraw pair: absent. +Three tells. Two make stop-and-surface mandatory. + +### A second, independent trigger +`docs/prompt-standards.md` item 11: **"when a claim about a mechanism needs a fourth +correction, delete the claim rather than refine it a fifth time."** The knob paragraph has now +been written three times and is wrong a third time; the model-cause enumeration twice. This is +the named remedy for exactly this shape, and it points at deletion, not a fourth refinement. + +### The sharpest single defect +The `Story:` syntax I pinned in pass 2 accepts a bare or double-quoted path. **Every** +`Story:` header in this repository uses Markdown backticks, and several carry trailing prose +after the path. Under the rule now shipped, all of them are malformed — including the three +plans governing this very cycle. A rule that invalidates its own governing artifacts is not a +refinement problem. + +## Open, nothing acted on at the stop +Spec: 5 Blocker, 2 Major. Quality: 1 Blocker, 7 Major. Files: +`.context/codex-reviews/gate-b-{spec,quality}-rle-pass-3.md`. + +## Evidence state (valid, revalidated at pass 3) +Battery green after every round. Check-that-fails-without-the-change: the four pass-1-Minor +assertions. Named risk-path verification: DISCHARGED — union of the three plans' Story headers +is one story, Risk high / Security none, level 2, floor 3; the provenance line matches the +pinned grammar and its floor is licensed by that level. diff --git a/.context/codex-reviews/gate-b-spec-fic2-pass-1.md b/.context/codex-reviews/gate-b-spec-fic2-pass-1.md new file mode 100644 index 0000000..eda1353 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-fic2-pass-1.md @@ -0,0 +1,6 @@ +BLOCKER | high | CLAUDE.md:98; CLAUDE.md:154; plugins/dev-workflow/commands/workflow-init.md:298; plugins/dev-workflow/commands/workflow-init.md:345 | The new scope rule requires a correction to stay inside the assigned fix set, but both copies later say a long correction aimed at the last correction does not stop and that a small correction-of-a-correction is absorbed, without carrying the in-set predicate. | An out-of-scope repair can be silently absorbed solely because of its ancestry, defeating the requested absorb-vs-stop boundary and contradicting the same block's stop rule. | Qualify both later examples in both mirrors with the requirement that the repair remains inside the assigned fix set. +BLOCKER | high | commit 6c70839 body:12 | The battery+check counterfactual is false: parent 17d5ad3 contains no assigned-fix-set definition in either CLAUDE.md or the workflow-init template, so `git show HEAD~1:CLAUDE.md` cannot demonstrate the quoted prior wording; the named mirror-parity read would also pass before the change. | The required check does not fail without the change, leaving the profiled story's mandatory evidence inadequate and this Gate-B call without its stated precondition. | Amend the evidence entry to cite and perform a real old-versus-new decision check, such as the nine-state matrix recorded at docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:140, and state the actual parent observation. +MINOR | high | plugins/dev-workflow/CHANGELOG.md:48 | The 0.10.0 entry first says any Blocker/Major-free pass has satisfied the clean-final-pass rule and should close, then only at line 50 limits clean completion to at-or-above the floor. | The shipped entry contains two incompatible answers for a below-floor pass carrying only Minor/Nit findings, recreating the gate-off reading the next sentence claims to close. | Add the at-or-above-floor qualification to the first clean-completion sentence so both statements give one answer. +MINOR | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:244; todos.md:71 | The claim that this round appended four cross-referencing supersession entries conflates three entries added in the reviewed range with the `docs-drift` entry at docs/hardening-log.md:76, which was already present in base commit 17d5ad3 from commit baa75c1; four such phrases exist in the current file, but only three belong to this round. | The closure record and parked gap misstate the incident count and credit this round with a malformed entry and repair inherited from an earlier cycle. | State the range-accurate count of three, or explicitly distinguish four currently present as one pre-existing plus three appended by this round. +MINOR | high | docs/hardening-log.md:90 | The final governing entry restates the current answer ("This row's guard covers ...; the reporting duty belongs ...") even though the header at line 44 requires only naming what is false and citing where the current answer lives. | The purported append-only repair is itself malformed under the ledger convention, so the closure claim at docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:250 is not fully true and the restatement can drift again. | Preserve append-only history by appending a new self-contained governing entry that names every fault in the target row but cites the current rule instead of restating it. +END OF FINDINGS (5 total) diff --git a/.context/codex-reviews/gate-b-spec-fic2-pass-2.md b/.context/codex-reviews/gate-b-spec-fic2-pass-2.md new file mode 100644 index 0000000..19a9e90 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-fic2-pass-2.md @@ -0,0 +1,14 @@ +BLOCKER | high | commit 37ee4c8 body; docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:169-180 | State 11 falsely calls the parent contradictory: `git show 17d5ad3:CLAUDE.md` and the parent workflow-init template contain no two-tell threshold, so the old clean-final-pass rule simply closes this at-floor Blocker/Major-free pass and the claimed 0-of-9 score is false | the mandatory `battery+check` evidence is not a truthful counterfactual against the committed prior state | make state 11 an old/new agreement control, change the open set to the actual eight states, and rescore both parent blobs +BLOCKER | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:146-170 | The matrix claims complete inputs but has no tell-count column, and states 3 and 10 record identical inputs while producing different outputs because the user's expansion answer is also absent | several new-text exits change when two tells are present, and the declined-expansion result cannot be distinguished from the pre-answer stop, so the 9-of-9 score is not determined | add tell-count and expansion-answer inputs to every state, split pre-answer from post-decline states, and rerun all four text variants +BLOCKER | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:152-178 | The action check is false for state 10: parent §5 requires every Major to resolve, while the new declined-expansion rule changes the action to record rather than resolve; one shared Action cell hides that change while line 178 says old and new agree | the named verification does not test the central behavioral change introduced by the declined-expansion repair | split old and new action columns and score the transition honestly +BLOCKER | high | CLAUDE.md:76-80,95-100,133-137,507-508; plugins/dev-workflow/commands/workflow-init.md:276-280,295-300,329-333,686-687 | The declined out-of-set Blocker/Major path says recording discharges the finding, but unchanged rules still say to fix every Blocker/Major, require surfaced findings to remain open with resolution unwaived, and prevent a pass carrying the repeated finding from being clean; no rule makes a decline binding when the same finding is reported later | a declined Blocker can remain mandatory, stop every later pass again, or create duplicate records, so the requested defined path can still spin | qualify all universal resolve and surfacing rules with the in-set boundary and define how later re-reports of an already declined finding are acknowledged without reopening or rerecording it, in both mirrors +MAJOR | high | CLAUDE.md:92-109,152-168; plugins/dev-workflow/commands/workflow-init.md:292-311,349-359 | The replacement precedence paragraph says clean completion outranks “both exits” without separating the scope stop from the clearly-stuck and two-tell exits; an at-floor pass with only an out-of-set Minor is simultaneously required to stop for scope and allowed to close as Blocker/Major-free | the paragraph overclaims complete precedence and leaves a real scope-versus-close pairing ambiguous or lets clean completion bypass a contract question | limit clean precedence to the clearly-stuck and two-tell exits, state scope-stop precedence separately, and add this pairing to the matrix +MAJOR | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:122-128 | The old-condition accounting calls the Blocker/Major filter kept, but the new declined path narrows the parent's universal “fix Blocker/Major after each” duty to in-set findings and never marks that condition kept, moved, or deliberately dropped | the durable record violates AGENTS.md's decision-procedure accounting rule and conceals the behavior change behind the declined path | list every prior loop condition and mark the universal resolve duty explicitly narrowed, including its final-pass and surfacing consequences +MAJOR | high | docs/hardening-log.md:91 | The new last-entry-governs record does not describe its target row accurately: the unconditional absorb rule is in the target row's `ref`, not its `finding`; the reporting duty belongs to the `prompt-vague-criteria` guard rather than this guard; and later reporting/decline additions were not false when the row was written, while the citation omits the reporting paragraph | the required self-contained governing repair is itself stale narration and fails the ledger header's field, timing, and current-answer requirements | append a new governing entry with faults assigned to the actual fields, distinguish never-true claims from later staleness, and cite each current prompt paragraph without restating it +MAJOR | high | plugins/dev-workflow/CHANGELOG.md:33-40 | The 0.10.0 entry says every new structural or contract question is outside scope, while shipped §5 and matrix state 5 allow such a finding to be in the assigned set and stop because novelty overrides ancestry | the CHANGELOG misdocuments the rule this manifest version ships | describe novelty as an independent stop condition that does not determine assigned-set membership +MINOR | high | CLAUDE.md:138-144; plugins/dev-workflow/commands/workflow-init.md:335-341 | The mandatory pass-4-onward three-line status format is described but never shown with a literal example, contrary to prompt-standards checklist item 4 | agents can combine or label the lines inconsistently, weakening the comparison the reporting duty exists to support | add the same compact three-line example to both prompt copies +MINOR | high | CLAUDE.md:1-3; plugins/dev-workflow/commands/workflow-init.md:192-195 | The changed root CLAUDE prompt and the scaffolded CLAUDE template do not name their executing target model; workflow-init's outer `Target model` line is not part of the file it writes | the changed prompt artifacts fail prompt-standards checklist item 1 and therefore AGENTS.md invariant 11 | add `Target model: Claude via Claude Code` to the root prompt and inline template +MINOR | high | docs/superpowers/plans/2026-08-15-reviewer-availability-salvage.md:155-173,354-359; docs/superpowers/plans/2026-08-01-gate-pass-result-classification.md:57-59,507-509 | The inserted §5 blocks moved numeric CLAUDE.md and workflow-init.md citations that these files still present as locators; for example workflow-init line 295 now lands on the declined-expansion clause instead of the finding format and CLAUDE.md line 150 lands on the tell threshold instead of pass acceptance | readers following the cited positions reach unrelated rules, reproducing the size/value/position docs-drift class the standing lens requires this diff to check | replace affected numeric locators with stable paragraph-opening anchors or add commit-qualified historical-location notes, and audit every numeric citation to both changed files +MINOR | medium | AGENTS.md:50-59; docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:1 | The diff adds a durable closure and named-verification artifact under `docs/field-reports/` but leaves that directory out of AGENTS.md's declared meaningful architecture tree | maintainers and reviewers cannot discover the new required artifact surface from the repository's single source of truth | add `docs/field-reports/` to the architecture tree with its committed-input and closure-record role +MINOR | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:1 | The title says the round closed on 2026-08-17 even though the same round contains governing ledger work dated 2026-08-18 and 2026-08-26 and is still undergoing Gate B | the closure record's date is stale and makes later repairs look external to the round | update the close date when the cycle actually closes or remove the date until then +END OF FINDINGS (13 total) diff --git a/.context/codex-reviews/gate-b-spec-fic2-pass-3.md b/.context/codex-reviews/gate-b-spec-fic2-pass-3.md new file mode 100644 index 0000000..b178122 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-fic2-pass-3.md @@ -0,0 +1,3 @@ +MAJOR | high | docs/hardening-log.md:91 | The last and therefore governing supersession entry assigns the unbounded absorb rule to the target row's `finding`, but that rule is in its `ref`; it also assigns the reporting duty to this `prompt-missing-stop-condition` row even though that duty belongs to the `prompt-vague-criteria` guard, contrary to the ledger's lines 58-62 last-entry rule and the story's grounded-evidence outcome at lines 24-25 | The governing record describes the target row inaccurately, so a recurrence reader can evaluate the wrong field and guard scope | Append a new self-contained governing entry that assigns the observation-base fault to `finding`, the missing assigned-fix-set predicate and floor contradiction to `ref`, and leaves the reporting duty with `prompt-vague-criteria` +MAJOR | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:122, docs/hardening-log.md:115-116 | The third guard-scope bullet names `rewrite-drops-prior-condition` but quotes the standing-falsification lens from the following `docs-drift` row; the named row's actual guard is the AGENTS.md decision-procedure accounting rule, so story acceptance criterion lines 53-67 does not have the required quoted guard for every examined row | The closure record does not establish which guard was actually examined | Quote the actual `rewrite-drops-prior-condition` guard and re-confirm the explicit in/out result, or name `docs-drift` if that was the row actually examined +END OF FINDINGS (2 total) diff --git a/.context/codex-reviews/gate-b-spec-fic2-pass-4.md b/.context/codex-reviews/gate-b-spec-fic2-pass-4.md new file mode 100644 index 0000000..8d781f4 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-fic2-pass-4.md @@ -0,0 +1,2 @@ +NO FINDINGS +END OF FINDINGS (0 total) diff --git a/.context/codex-reviews/gate-b-spec-fic2-pass-5.md b/.context/codex-reviews/gate-b-spec-fic2-pass-5.md new file mode 100644 index 0000000..d14d8a6 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-fic2-pass-5.md @@ -0,0 +1,3 @@ +MAJOR | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:185-204 | The five-state verification supplies an already-computed tell count but never identifies which of §5's five tell predicates are true or supplies the findings trend, Blocker trend, cluster category, and require↔withdraw input from which that count must be derived. A draft can delete or invert one tell definition while the fixture still declares 0, 1, or 2 tells and every R row keeps its expected result. | The named battery+check evidence can stay green while part of the shipped reporting duty regresses, so it does not yet cover the whole rule it claims to score. | Add raw input columns or focused rows that exercise each of the five tell predicates and derive the tell count, including a decisive case for each tell against both prompt copies; then rescore and update the commit-body counterfactual. +MAJOR | high | docs/superpowers/stories/2026-08-17-arms-race-remedy-as-procedure-story.md:76-83 | The acceptance criterion's exhaustive list of reporting-duty conditions names the three-line report, status-report carrier, five tells, and any-two mandatory stop, but does not require the duty's pass-4 activation boundary to be accounted for individually. Calling the path the “pass-4-onward reporting duty” merely names that condition in prose even though the next sentence says naming the duty does not satisfy the criterion. | A future implementation could move the duty to pass 1 or pass 5 while still checking off every individually enumerated condition, silently changing when reporting and the two-tell stop apply. | Add “the duty starts at pass 4 and applies on every later pass” as its own condition that must be marked kept, moved, or deliberately dropped. +END OF FINDINGS (2 total) diff --git a/.context/codex-reviews/gate-b-spec-fic2-pass-6.md b/.context/codex-reviews/gate-b-spec-fic2-pass-6.md new file mode 100644 index 0000000..98ec5a3 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-fic2-pass-6.md @@ -0,0 +1,3 @@ +MINOR | high | .context/codex-reviews/gate-b-fic2-parked-review-economics.md:65,130 | Both cycle tables record pass 1 with 2 Blockers, but the referenced pass-1 spec and quality findings files contain 3 Blockers in total. | The companion note's full-cycle Blocker curve is factually false, so its claimed cycle shape is not internally consistent even though the pass-2 rising-Blocker tell and stop decision remain unchanged. | Change the pass-1 Blocker value from 2 to 3 in both tables. +MINOR | high | commit 3ad6a74 body:10-14 | The evidence says there is one table per shipped rule and describes the nine-state table as scoring only absorb-vs-stop scope and action, but that table's states 8 and 9 also score the separately added clearly-stuck criterion; the second table alone scores the reporting duty. | The required evidence entry understates the first table's actual coverage and gives a false one-table-per-rule account of a verification that spans absorb scope, the stuck exit, and the reporting duty. | Describe the tables by their real coverage: the first scores scope/action plus the stuck-exit conditions, while the second scores the reporting-duty conditions; remove or qualify the one-per-rule claim. +END OF FINDINGS (2 total) diff --git a/.context/codex-reviews/gate-b-spec-fic2-pass-7.md b/.context/codex-reviews/gate-b-spec-fic2-pass-7.md new file mode 100644 index 0000000..7d8e3ce --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-fic2-pass-7.md @@ -0,0 +1,3 @@ +MINOR | high | commit 5786a9b body:10-16; docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:159-169 | The repaired evidence correctly says matrix states 8 and 9 score all three clearly-stuck conditions, then incorrectly groups stuck-reading availability among inputs the first table does not read; those states derive that availability from Plateau, Coverage sufficient, and B/M regenerating | The required commit-body evidence is internally contradictory and misstates why deleting the reporting duty leaves the nine-state matrix unchanged | Say only pass number, tell count, and carrier are absent from the first table; clarify that the first table uses stuck availability for the clearly-stuck exit but cannot test whether it is a precondition for the independent two-tell stop +MINOR | high | todos.md:241-243 | Item 9's handback disposition names only the clearly-stuck and scope-stop paths as the gate-loop stop-and-surface cases, but this diff adds a third independent path at CLAUDE.md:133-146 and plugins/dev-workflow/commands/workflow-init.md:330-343: the pass-4-onward two-tell stop, for which a clearly-stuck reading is explicitly not a precondition | A future implementation following this durable disposition can omit context-percentage handling from handbacks produced by the new two-tell stop | Make the examples non-exhaustive or include the pass-4-onward two-tell stop +END OF FINDINGS (2 total) diff --git a/.context/codex-reviews/gate-b-spec-harden-pass-1.md b/.context/codex-reviews/gate-b-spec-harden-pass-1.md new file mode 100644 index 0000000..e3f928f --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-harden-pass-1.md @@ -0,0 +1,4 @@ +MAJOR | high | docs/prompt-standards.md:92-96 | The new unconditional “cite the source; do not paraphrase it” rule conflicts with invariant 8 and this change's settled requirement that `/workflow-init` carry a downstream-usable inline copy of §5: an installed template cannot point at this repo's `CLAUDE.md`, and replacing required inline rules with such pointers would break scaffolding | Future prompt edits can comply with item 11 only by violating the inline-template architecture or can treat the standard as knowingly inapplicable without an exception | Scope the rule to cases where the authoritative source is available to the executing reader and no inline/self-contained copy is required; explicitly preserve invariant 8's inline-template exception +MINOR | high | CLAUDE.md:238-241; plugins/dev-workflow/commands/workflow-init.md:414-417; docs/hardening-log.md:34 | The new lens and ledger categorically say no parity check, resync, or self-grep “can” find untouched statements, but a targeted deterministic consistency check can catch a particular stale spelling just as check 4b catches one docs-drift spelling; what does not exist is a comprehensive deterministic check for arbitrary semantic drift | The hardening repeats the unverified-enforcement/coverage overclaim it is meant to prevent and obscures the same scope-sensitive rung reasoning used to refuse escalation | Replace “can find/reach” with “no existing check or grep limited to edited paths found these” and state that no comprehensive deterministic rung covers arbitrary semantic drift +MINOR | high | docs/hardening-log.md:34 | The `docs-drift` row never states the limitation of its actual `P std` rung: the standing lens is prompt text only, nothing enforces that a reviewer asks or answers it, and it cannot guarantee discovery of falsified statements; unlike the other two rows, it also does not plainly say that no comprehensive deterministic rung fits this mechanism | The ledger overstates “asking” as the hardening without recording what the rung does not cover, contrary to the requested blind-spot accounting and the house standard established by adjacent rows | Add that the lens is instruction-backed, no mechanism enforces or validates its use/result, and no comprehensive deterministic rung reaches arbitrary semantic drift +END OF FINDINGS (3 total) diff --git a/.context/codex-reviews/gate-b-spec-harden-pass-2.md b/.context/codex-reviews/gate-b-spec-harden-pass-2.md new file mode 100644 index 0000000..b52bfc5 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-harden-pass-2.md @@ -0,0 +1,2 @@ +MINOR | high | plugins/dev-workflow/CHANGELOG.md:28 | The new 0.7.1 entry still categorically says “no parity check or self-grep reaches” untouched stale statements, even though the hardened §5 copies and ledger correctly limit that claim to the particular edited-path checks that ran and admit a narrow check could pin a spelling. | The release note commits the same unverified-enforcement overclaim this change is meant to harden and contradicts the ledger’s stated limitation. | Scope the sentence to the checks that ran, or say no comprehensive check covers arbitrary semantic drift. +END OF FINDINGS (1 total) diff --git a/.context/codex-reviews/gate-b-spec-harden-pass-3.md b/.context/codex-reviews/gate-b-spec-harden-pass-3.md new file mode 100644 index 0000000..26774fb --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-harden-pass-3.md @@ -0,0 +1,2 @@ +MAJOR | high | plugins/dev-workflow/commands/workflow-init.md:731 | The downstream prompt-standard says to delete a mechanism claim after a "third or fourth correction", while docs/prompt-standards.md:99 requires deletion only when the claim needs a fourth correction; the two copies therefore do not agree in substance. | Initialized projects can apply the hardening one recurrence earlier than this repo, so the supposedly synchronized prompt standard has two different decision thresholds. | Keep the downstream copy free of repo-specific incident details, but give it the same fourth-correction threshold as the authoritative copy. +END OF FINDINGS (1 total) diff --git a/.context/codex-reviews/gate-b-spec-harden-pass-4.md b/.context/codex-reviews/gate-b-spec-harden-pass-4.md new file mode 100644 index 0000000..8d781f4 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-harden-pass-4.md @@ -0,0 +1,2 @@ +NO FINDINGS +END OF FINDINGS (0 total) diff --git a/.context/codex-reviews/gate-b-spec-pass-1-dispositions.md b/.context/codex-reviews/gate-b-spec-pass-1-dispositions.md new file mode 100644 index 0000000..6725b8d --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-pass-1-dispositions.md @@ -0,0 +1,97 @@ +# Gate-B pass 1 — dispositions + +Range f9ed886..f964bc5, `reviewType: full`. Both branch files validated after a +single-branch spec resume (see the pass-1 note at the end). + +Union of the two branch files, deduped. Blocker/Major resolved; Minor/Nit collected. + +## Applied — committed diff + +1. **MAJOR (both branches) — spec still claims backdating is mechanically checked.** + ACCEPTED, verified independently. `grep` over the spec found two survivors of the + four-line correction: line 287 "Backdating is forbidden and mechanically checked", + and line 753 "while the check enforces only non-decreasing recorded dates". Both + contradicted the corrected line 176, §8's "Check 1e is therefore not implemented … + No chronology validation exists", the plan's line 87 and `todos.md`. Fixed both. + Re-swept by claim rather than by phrase, per the AGENTS.md Don't: every remaining + hit is a correct *negative* claim or a conditional about the unimplemented 1e. + +2. **MAJOR — §8 says an extra inert entry is "detected by nothing".** ACCEPTED after + independent verification. This finding came from the *clobbered* first spec file + and appears in neither validated branch file; it is recorded here as the reader's + own observation, not as a counted pass-1 finding. Spec Check 3's cardinality oracle + requires the entry to "resolve to exactly one position", and `check-c3.sh` + implements it as `nent != 1` — so a second entry-shaped line *is* caught, once, on + this change. Narrowed the headline to check 1d and named check 3 as what sees it; + also corrected the companion sentence "every stated check will accept it". + +## Applied — scratch harness (never committed; it is the evidence, not the product) + +3. **MAJOR (both branches) — C1d identifies the mandated entry positionally.** + ACCEPTED. It took "the first candidate not present at BASE", so an unrelated or + mistyped earlier entry would decide the verdict. Spec §8 states the opposite order: + *identify the §3.1-mandated entry in current content, then prove it was added*. + Rewritten to select by date + row date + fingerprint, require exactly one match, + and only then assert it is absent at BASE. `entry-good` now passes because the + entry is the right one rather than because it happens to come first. + Consequence: two fixtures mutate the entry's own date, so `matrix.sh` now passes + the entry date for those rows — which makes each test its advertised dimension + (the on-or-before bound) instead of testing findability. + +4. **MAJOR (both branches) — `matrix.sh` folds stderr into the exit status.** + ACCEPTED as a real masking path, though verified LATENT: no fixture row emits + anything on stderr today, so no current verdict rested on the fold. Stderr is now + an unconditional disagreement, decided before the PASS/FAIL comparison. + +5. **MAJOR (quality) — fixture builders assert nothing about their anchor.** + ACCEPTED as real, verified LATENT: every fixture was checked and all sixteen differ + from `baseline.md`. `sub_line`/`after_line` now assert the anchor matched exactly + once and that the result differs from its input. + +6. **MAJOR (both branches) — `parse_entry` splits the whole line on ` · `.** + ACCEPTED. §2.2 forbids the separator only in the two prose fields, so a fragment + quoting a `finding` that contains ` · ` is legal and was being rejected — a grammar + narrower than the product's. Rewritten to parse the locator left-to-right and bound + the fragment by its quotes, splitting only the remainder. No effect on this change's + evidence: the mandated entry is fragmentless. + +7. **MINOR (both branches) — sentinel cardinality counted by line, not occurrence.** + ACCEPTED; cheap and correct. Now uses `occurs_exactly 1`. + +8. **MINOR (quality) / MAJOR (spec) — C2b compares regions after `join_paragraphs`.** + ACCEPTED. Byte-identity is what the two-surface parity requirement claims, and + normalizing line breaks away could pass a wrap-only divergence. Now compares the + raw regions. Verified latent first: the two regions are byte-identical as they + stand (70 lines each). + +9. **MAJOR|medium (spec) / MINOR|medium (quality) — `delete e` is not POSIX awk.** + ACCEPTED. Correct: whole-array delete is a common extension. Replaced with a + `clear()` helper deleting key by key. + +10. **MAJOR (spec) / MINOR (quality) — `rows-escapes` exercises no parser.** + ACCEPTED. The escape-heavy row carries no mandated locator, so C1a and C1d never + had to parse it. `build-fixtures.sh` now asserts the row round-trips through + `rowline` byte-identically — which it can only do if the escaped pipe was not read + as a delimiter, the doubled backslash was, and the digits were not coerced. + +## Checked and NOT a finding + +- **Spec line 832's "twenty-two current rows".** A first count said 19 and was wrong: + it split on `|` naively and mis-parsed rows carrying escaped pipes. Re-counted with + the escape-aware parser: 22. The spec is correct; nothing changed. + +## Pass-1 protocol note + +The `full` call returned four protocol lines, not two: both reviewers wrote both +slots. The spec branch reported `spec | 7 / quality | 5`; the quality branch reported +`9 / 9`; the disk held `9 / 9`, and both files read as quality-axis reviews. The spec +branch's seven spec-compliance findings had been overwritten — §5's stated race, with +every mechanical check still passing. Treated as an INCOMPLETE pass, not counted; the +single recovery attempt was spent on a write-only resume of the spec session +(`reviewType: spec`, its own `sessionId`), deleting only the failed branch file. That +returned one clean line and a valid 7-finding file. + +**Passes 2 onward run as two sequential single-branch calls** — `reviewType: spec`, +then `reviewType: quality` — each naming only its own slot. `full` shares one +`additionalContext` between two parallel reviewers, and naming both slots in it is +what let each write both. diff --git a/.context/codex-reviews/gate-b-spec-pass-1.md b/.context/codex-reviews/gate-b-spec-pass-1.md new file mode 100644 index 0000000..3ef4a3a --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-pass-1.md @@ -0,0 +1,19 @@ +BLOCKER | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:29; todos.md:185-209 | Item 6 has two dispositions: an independently actionable todos row with a named trigger and a parked story, although the closure table declares only the story | The intake story requires exactly one of the five permitted disposition forms for every item, so two backlog owners violate its core invariant | Retain either the parked story or the todos row as the disposition and remove or reduce the other to a non-actionable backlink +BLOCKER | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:51-57 | Item 9 is rejected on the false claims that no consumer artifact carries context percentages and chat evidence is unverifiable by construction; consumer docs/handoff-cowork.md:10,72, its A4 plan at lines 231 and 697, and machine-readable session JSONLs all contain the evidence | The round adopted a rejection without exhausting its cited evidence and thereby discarded a supported field claim | Withdraw the rejection, verify the named artifacts and session records, and give item 9 one grounded disposition +BLOCKER | high | CLAUDE.md:96-106; plugins/dev-workflow/commands/workflow-init.md:292-301; plugins/dev-workflow/CHANGELOG.md:31-44 | The new stuck rule treats 0-1 Blockers plus few-line findings as proof of substantive convergence and gives exact 2848 or approximate 2800 size and an 8-Blocker start, while the cited consumer taxonomy says exact size was not recorded, the measured series starts at 11, low Blockers do not measure coverage, and late Blockers were semantic | The shipped stop criterion re-adopts claims its own evidence refutes and can stop review while a subsystem remains unreviewed | Keep the Blocker curve only as a non-decisive heuristic, require an explicit coverage judgment, and remove the unsupported size, starting count, convergence proof, and causal claims +BLOCKER | high | todos.md:487-499 | The /capture-finding row was closed and folded into Finding A without the required choice first being put to Daniel; the originating session contains no intervening human decision | This autonomously resolves a human-owned acceptance criterion and open question that could materially change the round's scope | Revert or re-park the closure, ask Daniel to choose build now versus re-park or fold, and record that decision only afterward +BLOCKER | high | docs/superpowers/stories/2026-08-17-field-intake-canvas-a1-a5-report-story.md:4; CLAUDE.md:384-399 | The proposed battery+check counterfactual searches the prior file for vocabulary selected from the new prose, so an equivalent prior rule using synonyms and a substantively divergent mirror would still pass; the assertion that either prior rule necessarily would have produced a grep hit is false | Although the check reads an independent prior commit, its self-supplied lexical oracle cannot adjudicate the semantic prose defect required by battery+check | Replace it with a named semantic comparison that enumerates every old and new decision condition, checks substantive root-template parity, and names an observation that would actually occur if the claim were false +MAJOR | high | todos.md:137-160; docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:116-120 | Item 1 infers that a session-start unavailable fingerprint makes the cycle permanently unhealable and that count 21 proves a fingerprint was computed and stored, but the hook recomputes and writes on each accepted review and increments the count independently of fingerprint availability | Observations were converted into a hook mechanism defect that was not verified, and the stated limitations omit the two decisive uncertainties | Preserve only the observed files and messages, add later-pass computability and reviewed-state contents to the unknowns, and characterize the hook write and reset paths before naming a cause +MAJOR | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:61-69 | Part 3 is a multi-line causal narrative rather than the required one-line result, and unchanged committed baselines are presented as proof that the pass floor, terminators, WIP mechanics, quote-before-edit, and severity discipline all work as designed | The same baseline history is compatible with any of those mechanisms being broken and proves only the committed boundary bytes, not their causes or transient per-run state | Replace the paragraph with one calibrated line reporting the unchanged baselines and cited matches while disclaiming causal attribution to the five rules +MAJOR | high | todos.md:441-460 | The item 3 note conflates three different envelope shapes into a one-field NO_CREDITS remedy and says task_complete distinguishes success-with-progress but no answer, even though the cited evidence contains a no-file degeneration with task_complete and another silent death without it | The proposed upstream scope does not fit all three required instances, and a credit code cannot validate a success-true response's completeness | Narrow the candidate to missing credit-error data and separately record the output-completeness shape with the actual artifact, answer, log-growth, and completion diagnostics it needs +MAJOR | high | todos.md:399-430; docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:107-114 | Item 4 calls the operator step discoverable only by measurement and a documentation gap even though fixed-point README.md:79-104 already documents the variable, value, startup timing, failure mechanics, and workaround; it also says the variable buys only C1 while the same row says foregrounding lets successful long calls reach and be counted by the hook | The adopted row misstates both the shipped documentation and the operational effect its own evidence is meant to motivate | Recast the row as preflight automation and discoverability beyond existing setup docs, and separate false-count protection from keeping long successful calls in the foreground +MAJOR | high | todos.md:223-231 | Item 11 repeats that every large field tranche split and incurred a costly round-trip but cites no tranche records, sizes, or measurements and caveats only the proposed threshold | This adopts an unverified factual premise with a caveat, contrary to the round's verify-or-reject requirement | Cite and bound the actual tranche split records or park only the proposal and explicitly reject the frequency and cost premise as unverified +MAJOR | high | docs/hardening-log.md:111 | The immutable prompt-vague-criteria row repeats the unsupported 2848, 34-pass, and 8-Blocker claims, equates the curve with substantive convergence, declares the condition to be stuck, and omits the cited evidence's coverage and late-semantic-Blocker limits | The ledger overstates what its prompt guard proves and will count the unsupported hardening as a held rung in later recurrence decisions | Use the ledger's supersession mechanism to retract the unsupported claims and append an accurate current record that treats the curve as insufficient without coverage judgment +MAJOR | high | docs/superpowers/stories/2026-08-17-arms-race-remedy-as-procedure-story.md:15-26,80-87 | The new parked story says the kit has neither a recognition signal nor remedy and describes section 5 as containing only the old bare clearly-stuck clause, while this diff simultaneously adds a Blocker-curve non-convergence heuristic | The story's starting state is already false and can cause future work to duplicate or replace the just-shipped decision branch without preserving it | State that section 5 now has a field-specific heuristic and terminal action, narrow the missing capability to the general signal, moves, rung integration, and durable state, and make preservation or replacement explicit +MAJOR | medium | CLAUDE.md:82-87; plugins/dev-workflow/commands/workflow-init.md:282-287 | The absorb and stop cases overlap: a finding can both correct the immediately preceding correction and open a new structural or contract question, but the rule gives no precedence between absorb and stop | Readers can legitimately take opposite actions on the exact boundary the new decision procedure was meant to settle | Make the cases mutually exclusive and state explicitly that a new structural or contract question overrides correction ancestry +MINOR | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:28-31,73-75 | The disposition table labels prompt codification as the disposition for items 5 and 8 even though prompt text is not one of the five allowed forms and the same record later attributes a hardening-log row to each item | The closure record obscures whether each item received exactly one permitted disposition, independently of the separately required prompt codification | Point each table row to its hardening-log disposition and describe the section 5 and template edits separately as required codification +MINOR | high | docs/hardening-log.md:110 | The prompt-missing-stop-condition row defines its exact guard as including the sentence that preserves the three-pass floor, then says the guard says nothing about pass counts | The contradictory scope description can make a later recurrence incorrectly treat loss of the preserved floor as outside the guard | Supersede the immutable row with wording that says the rule introduces no new pass threshold while explicitly guarding preservation of the existing floor +MINOR | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:41-49 | Item 7 is rejected because no kit artifact is read when a session is framed, overlooking the scaffolded CLAUDE.md that Claude Code reads as session instructions and other workflow prose that can prescribe handoff framing | The rejection rests on a false product-boundary premise rather than a verified reason not to adopt the field proposal | Re-evaluate the proposal against the scaffolded CLAUDE.md surface or reject it on a defensible scope or value reason +NIT | high | docs/superpowers/stories/2026-08-04-harden-finding-guard-scope-precheck-story.md:64,110-111 | The story now says and lists eight open questions but its sizing rationale still says the seven questions above | This is a stale count in a durable design record | Change seven to eight +NIT | high | CLAUDE.md:96-100 | The stuck criterion says the paragraph above sends that state to the user, but the immediately preceding paragraph is the new absorb-versus-structural-stop rule; the clearly-stuck terminal action is two paragraphs above | The internal pointer is misleading and unnecessarily differs from the clearer inline-template wording | Replace the reference with that is the stuck state to surface to the user +END OF FINDINGS (18 total) diff --git a/.context/codex-reviews/gate-b-spec-pass-1.pre-2026-08-16.md b/.context/codex-reviews/gate-b-spec-pass-1.pre-2026-08-16.md new file mode 100644 index 0000000..1fea951 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-pass-1.pre-2026-08-16.md @@ -0,0 +1,8 @@ +MAJOR | high | docs/superpowers/specs/2026-08-05-hardening-ledger-supersession-design.md:287 | The residual analysis still says backdating is "mechanically checked", although the corrected shared convention, §8, the plan, and todos.md all say check 1e was not implemented and no chronology validation exists | This is the same stale enforcement claim the deliberate four-line spec correction was meant to remove, so the authoritative design still overstates the protection behind the date-bound argument | Replace "mechanically checked" with the instruction-backed reality and audit the surrounding residual claim so it does not imply that non-decreasing dates are validated +MAJOR | high | scratch c1core.awk:C1d scope | C1d identifies the mandated entry as the first interval candidate whose whole line appears nowhere in BASE, rather than identifying the required §3.1 entry | An unrelated or mistyped entry inserted earlier is checked instead, so the named check can fail on a valid mandated entry or pass without proving that entry is non-inert; this also conflicts with the sanctioned typo-then-correct sequence | Select the mandated entry by its required fields or exact expected content, then independently prove that selected entry was added against BASE +MAJOR | high | scratch matrix.sh:row() | Non-empty stderr is folded into the check status with `_st=$((_st + 1))` and then any nonzero status satisfies an expected-FAIL row | A check that incorrectly exits zero can still be recorded as the advertised rejection merely because it emitted a shell or read error, so the fixture matrix can mask a genuine disagreement instead of proving the property failed for its stated reason | Treat any non-empty stderr as a matrix disagreement independently; only compare the original exit status with PASS or FAIL when stderr is empty, and print the captured stderr on failure +MAJOR | high | scratch build-fixtures.sh:rows-escapes and matrix.sh:rows-escapes | The escapes fixture appends an unrelated `fixture-escapes` row, while C1a and C1d require only at least one row matching the unchanged 2026-07-20 locator and silently ignore any unrelated line that `rowfields` rejects | Both assertions pass even if escaped-pipe or trailing-backslash parsing of the fixture row is broken, so the advertised oracle against a naive splitter or numeric coercion is not exercised and the matrix overstates its coverage | Put the escaped pipe, digits, and trailing backslashes in the locator-matching row used by C1a and C1d, or add a predicate that explicitly proves the fixture row parsed with the expected fields before counting the PASS +MAJOR | high | scratch c1core.awk:parse_entry() | The parser splits the entire entry on ` · ` and requires exactly four parts before it recognizes the optional quoted fragment | The shipped convention forbids that separator only in the two prose fields, so a valid fragment containing ` · ` is rejected and the one-time checker implements a narrower grammar than the product | Parse the locator and optional quoted fragment before splitting the two prose fields, and add a passing fixture whose fragment contains the separator +MINOR | high | scratch check-c2.sh:one_line_contains() and region() | Sentinel cardinality is counted by matching lines rather than by literal occurrences | Two copies of a sentinel on one line satisfy the advertised exactly-once check, leaving a small gap in the parity oracle | Count literal occurrences across each file and require one before extracting the region +MAJOR | medium | scratch c1core.awk:entry parsing loops | `delete e` uses whole-array deletion, which is an awk extension rather than POSIX awk syntax | A strict POSIX awk can reject the scratch check even though the harness is required to be portable under the target shells, making the evidence environment-dependent | Clear arrays portably with `for (k in e) delete e[k]` +END OF FINDINGS (7 total) diff --git a/.context/codex-reviews/gate-b-spec-pass-10-dispositions.md b/.context/codex-reviews/gate-b-spec-pass-10-dispositions.md new file mode 100644 index 0000000..0b3e112 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-pass-10-dispositions.md @@ -0,0 +1,8 @@ +# Gate B — spec branch — pass 10 dispositions + +4 findings, all Major. All accepted. + +1 MAJOR the getting-started "One trivial change: just commit" opt-out still keys the skip on triviality alone — ACCEPT. Second pre-existing sentence invalidated by this change, twenty lines below the one fixed at pass 9; the summary was corrected while the opt-out list underneath it kept teaching the old rule. Qualified the same way: profiled stories skip Gate B only at effective level 0 and still owe the battery and the recorded reason; unprofiled stories keep the judgement call; Gate A is skippable at no level. +2 MAJOR the plan's quoted Mechanics snippet is still singular — ACCEPT. Synced from the shipped clause, along with the CHANGELOG sentence that repeated it. +3 MAJOR the plan's parity check searches for an anchor the fix deleted — ACCEPT, and the finding proved itself: running the documented command produced `EMPTY CAPTURE: M_REPO` rather than the claimed agreement. That is the guard from pass-2 doing its job on the plan's own instructions — without it the check would have diffed two empty captures and reported success. Anchors updated to the current wording, and the check now returns BOTH HALVES AGREE. +4 MAJOR Gate-B Steps 6 and 7 still demand the evidence entry unconditionally — ACCEPT; the quality branch reported the same. Both steps now scope entries to cited PROFILED stories: every cited story contributes its path, only profiled ones an entry, and this cycle's unprofiled story contributes neither. diff --git a/.context/codex-reviews/gate-b-spec-pass-10.md b/.context/codex-reviews/gate-b-spec-pass-10.md new file mode 100644 index 0000000..d76c43f --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-pass-10.md @@ -0,0 +1,2 @@ +MINOR | high | docs/hardening-log.md:75 | The new supersession entry does not follow the ledger's own “where the current answer is” rule: after citing the two current sources it restates their answer, and it says the superseded row cites `CODEX_DEV_REVIEW_MODEL` even though that row never names that surface. | The correction is itself stale-prone and contains a false provenance claim, undermining the append-only record this task required. | Append a later entry for the same locator that fully describes the row's superseded status but ends with citations only; do not edit either existing line. +END OF FINDINGS (1 total) diff --git a/.context/codex-reviews/gate-b-spec-pass-11-dispositions.md b/.context/codex-reviews/gate-b-spec-pass-11-dispositions.md new file mode 100644 index 0000000..ed550a0 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-pass-11-dispositions.md @@ -0,0 +1,7 @@ +# Gate B — spec branch — pass 11 dispositions + +2 findings (1 Major, 1 Minor). Both accepted. + +1 MAJOR the hook's below-floor reminder still says "proceed only if this change is trivial" — ACCEPT, and it tripped the story's scope guard (no hook changes), so it was surfaced to the human rather than decided here. Waiver granted, with the reasoning recorded in the plan's decision record: the edit REMOVES a rule statement from the hook and defers to the policy file, so the hook reads no profile and gains no machinery — the inverse of what the guard protects. Shipping a reminder that asserts a now-false rule at the decision moment would be the unverified-enforcement-claim class in its most harmful position, and §5's skip rule changing means its echo in the hook changes in the same commit. + Constraint honoured: the replacement points at the policy without paraphrasing the rule, so no second sync surface is created. Codex's suggestion to "adjust its existing test expectation" was checked and dismissed — no test asserts that clause; the two hook tests match only `below floor|floor NOT met`. +2 MINOR process-pr-review leaves the no-story and all-stories cases implicit — ACCEPT; the quality branch raised the same as a Major. Three explicit branches now: no story cited (§5's unprofiled judgement rule), one cited story, and several cited stories where a skip needs every one individually eligible. diff --git a/.context/codex-reviews/gate-b-spec-pass-11.md b/.context/codex-reviews/gate-b-spec-pass-11.md new file mode 100644 index 0000000..776d257 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-pass-11.md @@ -0,0 +1,2 @@ +MINOR | high | docs/hardening-log.md:76 | The new governing supersession entry explicitly refers to "the entry immediately above", contrary to the ledger rule at lines 62-63 that entries never reference one another; it also spends part of the current problem field diagnosing that prior entry instead of describing only the located row as it now stands. | The pass-11 correction is itself nonconforming, and the current status depends on cross-entry narration the format expressly forbids. | Append another entry with the same locator (do not edit either existing entry) that says only that the target row's hardening names the wrong top-level probe and was wrong when written, then end with the two current-document citations without restating the rule or mentioning earlier entries. +END OF FINDINGS (1 total) diff --git a/.context/codex-reviews/gate-b-spec-pass-12.md b/.context/codex-reviews/gate-b-spec-pass-12.md new file mode 100644 index 0000000..8d781f4 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-pass-12.md @@ -0,0 +1,2 @@ +NO FINDINGS +END OF FINDINGS (0 total) diff --git a/.context/codex-reviews/gate-b-spec-pass-13-dispositions.md b/.context/codex-reviews/gate-b-spec-pass-13-dispositions.md new file mode 100644 index 0000000..0ce3515 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-pass-13-dispositions.md @@ -0,0 +1,7 @@ +# Gate B — spec branch — pass 13 dispositions + +1 finding, Major. Accepted. The quality branch reported the same defect independently (as a Minor). + +1 MAJOR the rewritten stage-8 rule overstates what the hook detects — ACCEPT, and it is the cycle's sharpest irony: the pass-12 fix for a gate-overclaim introduced a gate-overclaim. I wrote that agent definitions and hook messages categorically fire full Gate B; the hook's classifier is `(^|/)(CLAUDE|AGENTS)\.md$|(^|/)(\.claude|plugins|skills|commands)/`, so a bare `agents/foo.md` or `hooks/reminder.md` is classified docs-only and gets no gate. + This is AGENTS.md's named most-persistent class — "never describe what a gate proves without checking what it actually compares" — and it regenerated inside a correction of itself, exactly as the invariant-file warns ("each correction introduced a subtler version of the same claim"). + Fixed by naming the literal matched set: `CLAUDE.md` and `AGENTS.md` themselves, plus `.claude/`, `plugins/`, `skills/`, `commands/` at any depth; agent definitions and hook messages are covered because they sit under a matched parent; a bare top-level `agents/` or `hooks/` matches nothing. Claim verified against the matcher in the hook rather than against §5's prose about it. diff --git a/.context/codex-reviews/gate-b-spec-pass-13.md b/.context/codex-reviews/gate-b-spec-pass-13.md new file mode 100644 index 0000000..9e526b8 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-pass-13.md @@ -0,0 +1,2 @@ +MAJOR | high | docs/coding-workflow.md:130 | The rewritten stage-8 rule says agent definitions and hook messages categorically fire full Gate B, but the hook only recognizes Markdown prompts at `(CLAUDE\|AGENTS).md` or beneath a `.claude`, `plugins`, `skills`, or `commands` directory at any depth; a bare `agents/foo.md` or `hooks/reminder.md` is classified docs-only. | The user-facing workflow overstates what the gate detects and can give authors false confidence that prompt artifacts in unmatched layouts were reviewed. | State the exact matched directory names and any-depth rule here, and explicitly note that agent definitions are covered only through matched parents (for example `.claude/agents` or `plugins/*/agents`) while bare `agents/` and `hooks/` are not. +END OF FINDINGS (1 total) diff --git a/.context/codex-reviews/gate-b-spec-pass-14-dispositions.md b/.context/codex-reviews/gate-b-spec-pass-14-dispositions.md new file mode 100644 index 0000000..1e2b1f5 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-pass-14-dispositions.md @@ -0,0 +1,8 @@ +# Gate B — spec branch — pass 14 dispositions + +1 finding, Major. Accepted. (The quality branch was CLEAN this pass — NO FINDINGS.) + +1 MAJOR the corrected classifier description is STILL an overclaim — ACCEPT. My pass-13 fix named `is_prompt_path`'s regex as if it were the whole match set and said a bare `agents/` or `hooks/` "matches nothing"; `is_docs_only` also routes non-Markdown paths, mixed commits and an empty/undeterminable list to the reminder, so `hooks/codex-gate.sh` IS seen. "Fire full Gate B" also blurred an advisory hook that only prints and always exits 0 with the policy requirement itself. + FOURTH round of this class on ONE paragraph (pass 12: false rationale; pass 13: categorical prompt claim; pass 14: partial-matcher claim + enforcement blur). AGENTS.md predicts the shape exactly — "each correction introduced a subtler version of the same claim" — and this cycle reproduced it in a paragraph being rewritten to fix that very class. + So the fix is NOT a more precise enumeration. A user-facing methodology overview has no business restating a classifier: every restatement is a new copy that can drift from the source and a new chance to overclaim. The paragraph now says the policy requires Gate B for prompt artifacts, that the policy file spells out which paths count, that a reminder hook classifies independently, that both err toward firing, and that the hook only reminds and is not what makes the review happen. No path list, no enforcement claim, nothing to drift. + Generalisable lesson for the ledger: when a claim about a mechanism needs a fourth correction, delete the claim rather than refine it — the doc that needs it can cite the source instead. diff --git a/.context/codex-reviews/gate-b-spec-pass-14.md b/.context/codex-reviews/gate-b-spec-pass-14.md new file mode 100644 index 0000000..e466e26 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-pass-14.md @@ -0,0 +1,2 @@ +MAJOR | high | docs/coding-workflow.md:133-137 | The paragraph presents the `is_prompt_path` regex as the hook's complete match set and says a bare top-level `agents/` or `hooks/` “matches nothing,” but `is_docs_only` also sends every non-Markdown path, every mixed commit, and an empty/undeterminable path list down the Gate-B reminder branch; for example `hooks/codex-gate.sh` is seen despite its parent directory not matching `is_prompt_path`. “fire full Gate B” also blurs the policy requirement with an advisory hook that only emits a reminder and always exits 0. | Readers can wrongly conclude that any artifact under bare `agents/` or `hooks/` is invisible and that the hook itself invokes or enforces Gate B, so the claimed literal description still outruns the source. | Name this as the prompt-path submatcher only, qualify the bare-directory statement to Markdown prompt files, add the other `is_docs_only` fall-through cases (non-`.md`, mixed, empty/unknown), and say they trigger an advisory Gate-B reminder while §5 requires the review. +END OF FINDINGS (1 total) diff --git a/.context/codex-reviews/gate-b-spec-pass-15.md b/.context/codex-reviews/gate-b-spec-pass-15.md new file mode 100644 index 0000000..8d781f4 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-pass-15.md @@ -0,0 +1,2 @@ +NO FINDINGS +END OF FINDINGS (0 total) diff --git a/.context/codex-reviews/gate-b-spec-pass-16.md b/.context/codex-reviews/gate-b-spec-pass-16.md new file mode 100644 index 0000000..8d781f4 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-pass-16.md @@ -0,0 +1,2 @@ +NO FINDINGS +END OF FINDINGS (0 total) diff --git a/.context/codex-reviews/gate-b-spec-pass-2-dispositions.md b/.context/codex-reviews/gate-b-spec-pass-2-dispositions.md new file mode 100644 index 0000000..cea4096 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-pass-2-dispositions.md @@ -0,0 +1,60 @@ +# Gate-B pass 2 — dispositions (both branches) + +Range f9ed886..8647b8d. Run as **two sequential single-branch calls** — `reviewType: +spec`, then `reviewType: quality` — each naming only its own slot, after pass 1's +both-write race. Both branches reviewed the same commit; the quality call ran before any +pass-2 fix was amended, so the pass covers one tree. Both files validated. + +Spec branch: 3 findings (1 MAJOR, 2 MINOR). Quality branch: 1 finding (1 MAJOR). + +## Applied — committed diff + +1. **MAJOR (spec) — §6 still mandates check 1e while §8 says it was never implemented.** + ACCEPTED, verified. §6 read "**Six** properties, 1a through 1f. All must hold", and + 1e was written as live ("this checks recorded order", "this is a floor under it") + while §8 says "Check 1e is therefore not implemented … No chronology validation + exists". A surviving synonym of pass 1's finding 1, in the section that defines what + the checks are. Recast: §6 now says five must hold for this change (1a–1d, 1f) and + names 1e as specified-but-deliberately-unimplemented; 1e's own oracle moved to the + conditional, closing with "a floor that does not exist today". + +## Applied — scratch harness and the evidence entry + +2. **MINOR (spec) — the counterfactual mis-describes C1d's diagnostics.** ACCEPTED. The + entry enumerated "a parse, locator or eligibility diagnostic" while its own preceding + bullet quotes `undecidable interval`, and an empty block yields "the interval holds no + candidate line". Rewritten to enumerate the diagnostic belonging to each falsified + dimension, including both absence cases. + +3. **MINOR (spec) — `rows-two-matching` is built by `cp` + append and asserted nowhere**, + while the entry claimed every fixture is asserted to differ from its input. ACCEPTED + as a real overclaim I introduced in pass 1. Added an assertion that the locator + matches exactly two complete rows. Mutation-tested: suppressing the append gives + "rows-two-matching carries 1 complete row(s) matching the locator, want 2". + +4. **MAJOR (quality) — `entry-bad-date-locator`'s appended row is not asserted.** + ACCEPTED. Its advertised dimension is a calendar-invalid locator on an entry that + would *otherwise* match a complete row; lose the append and it still fails, for the + weaker "locator matches nothing" reason `entry-wrong-rowdate` already covers. Added an + assertion that exactly one complete row carries 2026-02-30. + **Extended beyond the reported site:** `entry-outside-block` has the identical + unasserted-append shape — lose its append and the fixture becomes "no entry anywhere", + which is not what it advertises. Asserted that the entry occurs exactly once and below + `Columns:`. Both mutation-tested; each fails with its own diagnostic. + +## Found by the author while validating, not reported by either branch + +5. **The evidence entry claimed "All 28 matrix rows".** `grep -c '^row '` gives **27**. + The number was introduced in the pass-1 rewrite of the entry and never verified — + the same stated-count defect this cycle exists to address. Corrected to 27, and the + arithmetic checked: 18 C1d rows (12 FAIL + 6 PASS) + 5 C1a + 2 C1b + 2 C1c = 27. + +6. **"Twelve constructed ledgers" was loose.** One of the twelve is + `ledger-unreadable`, a *directory* standing in for an unreadable ledger, not a ledger. + Reworded to "twelve constructed inputs … eleven of them ledgers, the twelfth a + directory standing in for an unreadable one". + +## Collected, not iterated on + +Nothing outstanding: both MINORs above were fixed because they concerned the accuracy of +the evidence entry, which goes into the commit body as a durable claim. diff --git a/.context/codex-reviews/gate-b-spec-pass-2.RACED-discarded.md b/.context/codex-reviews/gate-b-spec-pass-2.RACED-discarded.md new file mode 100644 index 0000000..2766a83 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-pass-2.RACED-discarded.md @@ -0,0 +1,16 @@ +BLOCKER | high | closing-commit Evidence entry quoted in the Gate-B request; CLAUDE.md:392-407 | The supplied named verification does not meet battery+check because its two observations can both pass without this change: the pre-existing paragraph is still unchanged at 17d5ad3, and a comparison of the two new regions finds no asymmetry when both regions are absent; the normative oracle is derived from the two edited copies rather than an independent requirement | The story requires a check that fails without the change and CLAUDE.md makes an unobservable or self-supplied counterfactual a blocking evidence gap, so a green battery cannot close this cycle | Derive the expected new clauses from the story and field evidence, assert each clause semantically in both prompt copies, and demonstrate that the same verification fails at 17d5ad3 and passes at HEAD while retaining the separate old-paragraph preservation check +MAJOR | high | CLAUDE.md:99-105; plugins/dev-workflow/commands/workflow-init.md:295-301; plugins/dev-workflow/CHANGELOG.md:37-40 | The stuck criterion requires coverage only to be judged and stated, not judged sufficient, then declares the stuck state from a post-six-pass 0-1 Blocker snapshot with individually small findings; it requires neither adequate artifact coverage nor observed non-convergence, despite the changelog claiming the caveat prevents stopping over an unreviewed subsystem | A reviewer can state that a subsystem is unreviewed and still satisfy the literal criterion, or stop a still-converging review, so the new rule licenses the exact unsafe clearance its caveat says it prevents | Require a stated basis for sufficient coverage with known unreviewed areas disqualifying the stuck state, and require an observed plateau or regeneration across passes before the numeric and sizing signals may trigger it; mirror the rule and narrow the changelog claim +MAJOR | high | CLAUDE.md:76-105; plugins/dev-workflow/commands/workflow-init.md:276-301 | The new precedence clause resolves only a finding that is simultaneously a correction-of-a-correction and a new structural question; it does not resolve a correction-of-a-correction that also meets the stuck criterion, for which one rule says absorb, fix, and keep looping while the other says surface, nor does the stuck rule exclude a remaining Major even though the preserved loop says every Major is fixed after each pass | The same review state has two incompatible mandatory actions, so the absorb/stop procedure remains ambiguous despite the requested precedence hardening and can bypass a preserved condition | State the precedence among correction ancestry, the stuck criterion, and the Blocker/Major fix rule, including whether any unresolved Major or an absorbable correction disqualifies the stuck state, and apply the same wording in both copies +MAJOR | high | todos.md:137-144,160-162; docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:127-129; plugins/dev-workflow/hooks/codex-gate.sh:823-840 | Item 1 says every counted pass overwrites the state file unconditionally and therefore proves the hash was uncomputable on every pass, but the hook makes the state write best-effort with failure ignored; a computed hash followed by a failed replacement write can leave the earlier unavailable value, and the row itself admits the intermediate values were not observed | This adopts a cycle-wide mechanism conclusion that neither the consumer state nor the hook supports, contrary to the verify-before-adopting requirement, and can misdirect the parked hook investigation | Describe the observation as an unknown computation-or-persistence failure, explicitly allow stale recorded state after a failed write, and leave unobserved passes' computability unknown +MAJOR | high | docs/field-reports/2026-08-16-canvas-a1-a5-field-report.md:43-47,75-81; todos.md:473-500 | Item 3 requested an upstream quota-reason field for success-false error envelopes, but its disposition adds a second upstream output-completeness contract for a Kimi success-true no-artifact failure that the row itself calls a different defect | The report expressly forbids additions beyond its eleven items, so packaging a new defect and upstream ask inside item 3 expands the intake scope and makes the item's one disposition carry two distinct asks | Keep the item-3 candidate limited to the quota and error-reason envelope request; retain the third cited case only as an explicit non-instance or record it through a separately authorized intake rather than adopting it here +MAJOR | medium | todos.md:486-496 | The item-3 note infers that only artifact plus answer plus log growth together distinguish completion from abandonment, although the cited cases establish only that task_complete by itself is insufficient and do not establish that three-part combination as necessary, sufficient, or exhaustive | This is another unverifiable mechanism claim inside the proposed upstream contract and can prescribe the wrong interface to the upstream project | Remove the sufficiency claim and state only the observed failure shapes and the insufficiency of task_complete, leaving the completeness signal's contract for upstream design +MAJOR | high | docs/superpowers/stories/2026-08-17-arms-race-remedy-as-procedure-story.md:15-24 | The parked story says the count oscillates while the substance is already settled, even though this change's own corrected evidence says no instrument measured coverage and a low count can coexist with an unreviewed subsystem | The story seeds its future procedure with the same unsupported convergence inference the stuck-rule correction was meant to withdraw, so later implementation can treat an unreviewed instrument as settled | State only the observed oscillation and same-shape correction lineage, and carry the one-consumer and no-coverage limits into the problem statement +MINOR | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:46-57; todos.md:10-14 | Item 7's rejection says the reactive-only policy requires recurrence before a shipped prompt may change, but that policy permits a change whenever a finding surfaces through real use and prohibits proactive sweeps; it states no recurrence threshold | The rejection's portability and evidence concerns may be defensible, but attributing a nonexistent threshold to repo policy leaves its sole disposition resting partly on a false premise | Remove the recurrence-as-policy claim and ground the rejection directly in the unmeasured transferability and downstream-cost evidence, or park it behind the proposed recurrence trigger +MINOR | high | todos.md:217-230 | Item 9's cited handoff document proves that the consumer asks for context percentage, but not that it is the one number deciding whether to continue a session; the row later admits that its reliability and ability to steer the decision are unsettled | This adopts a causal and determinative claim beyond the cited evidence and makes the parked question internally contradictory | Describe context percentage as a requested and potentially relevant input, and leave its reliability and decision weight as the explicit parked question +MINOR | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:147-154 | The closure record still calls item 4 a documentation gap even though its earlier pass-1 correction, the current todos row, and README step 2b all say the variable is already documented and the remaining gap is preflight or operator reach | The durable closure gives future readers two incompatible diagnoses and can trigger redundant documentation work | Replace documentation gap with preflight or operator-reach gap and cite the already-complete README guidance +MINOR | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:156-160 | The closure says the round did not read the hook at either item-1 site, while its preceding correction and the todos row explicitly derive recomputation and overwrite behavior from the counted-pass branch | This makes the round's evidence trail internally contradictory and obscures which mechanism conclusions were source-verified | Say that the counted-pass branch was read but the tree_hash failure seams and runtime persistence outcome were not established +MINOR | medium | todos.md:192-205 | Item 6 says the remedy was verified at three sites, but the three citations are one taxonomy maxim and two records from the same consumer cycle, not three independent observations that the procedure works | Counting restatements of one consumer incident as verifications overstates the evidence behind the parked story | Say the remedy is recorded in three artifacts from one consumer, distinguish the taxonomy statement from the observed cycle, and retain the one-consumer evidence limit +MINOR | high | docs/superpowers/stories/2026-08-17-field-intake-canvas-a1-a5-report-story.md:53-54; docs/hardening-log.md:112-113 | Neither newly appended ledger row states the guard-scope precheck result and quoted candidate guard required by the story; each says only that it is the first base-class row and records the new guard, while the actual candidate precheck and quotations live only in the separate dispositions document | The acceptance criterion makes the immutable ledger row itself the durable recurrence record, so the required scope evidence is absent from the artifact future hardening runs consult | Append superseding ledger entries whose refs name the no-match result and the inspected candidate guards with their quoted scope, or revise the story criterion through the normal decision process if the external closure record is intentionally authoritative +MINOR | high | docs/hardening-log.md:43-45,79 | The prompt-missing-stop-condition Superseded entry restates the supposed current rule as what was meant instead of limiting itself to the false target claim and citing the current answer | The ledger's own format expressly forbids restating the current answer because that creates another stale narration, which is the exact defect this entry is correcting | Remove the what-was-meant restatement and leave only the contradiction that made the target ref false plus the citations to the two current prompt copies +NIT | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:36-38 | The exact-disposition explanation says items 5 and 8 are two rules landing in one paragraph, but the diff places them in two separate adjacent paragraphs in both prompt copies | The closure record inaccurately describes the position and shape of the codification it is meant to audit | Say the two rules land as adjacent paragraphs in one section edit and one release cycle +END OF FINDINGS (15 total) diff --git a/.context/codex-reviews/gate-b-spec-pass-2.md b/.context/codex-reviews/gate-b-spec-pass-2.md new file mode 100644 index 0000000..121abc1 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-pass-2.md @@ -0,0 +1,15 @@ +BLOCKER | high | docs/superpowers/stories/2026-08-17-field-intake-canvas-a1-a5-report-story.md:4; CLAUDE.md:392-407 | The supplied battery+check entry does not test the behavior introduced by the prompt rules: its first observation can pass when both rules are absent, and its second compares two copies authored by this change, so it also passes when the same clause is omitted from both. The counterfactual is therefore still self-supplied and is not wired where the defect could appear. | The story requires battery+check, and CLAUDE.md makes inadequate evidence a blocking work gap; copy preservation and mirror parity establish authorship consistency, not that the decision procedure gives the required absorb, stop, and stuck outcomes. | For this pure prompt-text change, use an independently derived semantic decision matrix or controlled target-model evaluation that fails against 17d5ad3 and passes against HEAD, covering correction-only, novel-structural-only, both-at-once, plateau plus a small correction-of-a-correction, and insufficient-coverage cases; retain the battery results alongside it. +BLOCKER | high | CLAUDE.md:99-105; plugins/dev-workflow/commands/workflow-init.md:295-301; plugins/dev-workflow/CHANGELOG.md:31-40 | The stuck criterion says only that coverage is "judged" and stated before permitting a stop; it never requires the judgment to be that coverage is sufficient, and it does not make a known unreviewed subsystem disqualifying. It also permits stopping from a current 0-1 Blocker count plus small remaining fixes after about six passes without requiring an observed plateau or regeneration. | A reviewer can literally disclose inadequate coverage and then stop, so the new caveat does not prevent the exact false-clearance case it names; this also overstates the changelog claim that the rule recognizes nonconvergence without allowing an unreviewed area to masquerade as clean. | Require an explicit, supported judgment that coverage is sufficient, state that any known materially unreviewed area forbids the stuck exit, and require an observed plateau or regenerated semantic lineage across passes rather than a single low-count snapshot. +MAJOR | high | CLAUDE.md:76-105; plugins/dev-workflow/commands/workflow-init.md:276-301 | The precedence clause resolves only a finding that is simultaneously a correction-of-a-correction and a new structural question. It does not resolve a small correction-of-a-correction that also satisfies the later plateau rule: the absorb rule says fix it and keep looping, while the stuck rule says surface and stop; the latter can also be read to permit an individually small unresolved Major despite the pre-existing instruction to fix every Blocker and Major after each pass. | The two added decision procedures can prescribe opposite actions for the same review state, so implementations of the prompt can diverge on precisely the plateau edge case the change is meant to settle. | Add explicit cross-rule precedence stating whether a genuine stuck plateau overrides correction ancestry, and state whether any unresolved Blocker or Major disqualifies the stuck exit; mirror the same clause in workflow-init. +MAJOR | high | todos.md:137-144,160-162; docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:127-129; plugins/dev-workflow/hooks/codex-gate.sh:823-840 | Item 1 concludes that a session-start unavailable value proves tree_hash was uncomputable on every counted pass. The hook recomputes on each pass, but its state-file replacement is best-effort and ignored on failure, so a later computed hash can fail to replace an earlier unavailable value; the row itself admits that later-pass values are unknown. | This adopts a causal claim that neither the consumer artifact nor the hook establishes, contrary to the intake requirement to adopt nothing unverifiable and the repository rule against overstating what a mechanism proves. | Record the supported result only: the persisted value was unavailable at the observed endpoints, while later computation and persistence are unknown; include stale state or write failure among the possible causes. +MAJOR | medium | docs/field-reports/2026-08-16-canvas-a1-a5-field-report.md:43-47,75-81; todos.md:473-500 | Item 3's disposition adds a second upstream contract request for an output-completeness field on success:true responses. The report requested a machine-readable failure or quota reason for success:false envelopes; the success:true no-artifact observation is a different defect, even though the story required the third evidence instance to be considered. | The intake brief explicitly says not to add beyond the listed requests, so carrying an evidence instance does not authorize expanding it into another upstream proposal. | Keep the success:true observation as a limitation or non-instance of the requested failure-code proposal, or obtain a separate intake disposition before proposing an output-completeness contract. +MAJOR | high | todos.md:486-496 | Item 3 says only an artifact, an answer, and log growth together distinguish completion from failure or a still-running call. The cited consumer evidence establishes that task_complete alone was insufficient in the observed cases; it does not establish that this three-part combination is necessary, sufficient, or exhaustive. | The row turns a few observations into a universal mechanism claim, repeating the repository's known overclaim defect and prescribing an upstream shape the evidence did not verify. | Describe the observed envelope and artifact combinations without "only" or sufficiency language, and leave the completeness contract as an upstream design question. +MAJOR | high | docs/superpowers/stories/2026-08-17-cross-model-review-arms-race-story.md:15-24 | The problem statement says finding counts oscillate while the substance is already settled. The cited consumer taxonomy explicitly says coverage was not measured and a low Blocker count can coexist with an unreviewed subsystem, so the evidence cannot establish that the substance was settled. | This recreates the clearance inference that the corrected stuck-rule caveat is intended to prohibit and gives the parked story an unsupported factual premise. | Replace "substance is already settled" with the verified observation that counts and same-shape lineages oscillated, and state that neither convergence nor coverage was established. +MINOR | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:46-57; todos.md:10-14 | Item 7's rejection says lack of recurrence is not what the reactive-only policy requires before changing a shipped prompt. The policy permits hardening any finding surfaced by real use and forbids proactive sweeps; it does not impose a recurrence threshold, as the same intake's first-occurrence prompt changes for items 5 and 8 also demonstrate. | The transferability objection may support rejection, but attributing a recurrence requirement to the policy misstates an existing project rule. | Remove the policy claim and base the rejection solely on the verified lack of stack-neutral transferability, or park the item behind an explicitly chosen recurrence trigger rather than presenting that trigger as existing policy. +MINOR | high | todos.md:217-230 | Item 9 calls context percentage "the one number that decides" whether work continues in-session or moves to a fresh session. The cited handoff document proves that operators requested the percentage, but not that it alone decides the workflow; the row later concedes that reliability and decision weight are unsettled. | The parked trigger is founded on an unsupported causal importance claim and contradicts the row's own evidence boundary. | Describe context percentage as a potentially relevant decision input whose reliability and weight remain to be established. +MINOR | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:130-133,147-154; todos.md:454-459 | The closure record correctly says item 4 is not a documentation gap because README step 2b documents CODEX_GATE_COMMAND, but later says item 4 remains valid "as a documentation gap." | The two statements give the same report item incompatible dispositions and falsify the corrected characterization of the current README. | Replace the stale phrase with the supported operator-reach or preflight-discoverability gap. +MINOR | high | docs/hardening-log.md:58-61,78,113 | The Superseded entry for the stuck-rule row identifies the exact-size, Blocker-start, and convergence claims as false, but omits another false claim in its target row: the target says post-pass-six totals oscillated between 2 and 7, while the verified record and current prompt say 2 to 19. The target row's reference also lacks the new coverage precondition. | Under this ledger's last-entry-governs rule, a supersession must describe the prior row as it now stands; leaving a known false quantitative claim and incomplete current pointer unaccounted for makes the retraction itself inaccurate. | Append another Superseded entry, without editing prior rows, that identifies every false claim including the 2-to-7 range and points to the current rule containing the coverage condition. +MINOR | high | docs/hardening-log.md:43-45,79 | The Superseded entry for the absorb-rule row restates a current interpretation—"the rule introduces no new pass threshold"—instead of limiting itself to what was false and where the current answer lives; that wording is also risky beside the new "after about six passes" guidance. | The ledger format explicitly forbids restating the current answer because such summaries drift, and a row plus its retraction in the same commit is allowed only if the append-only correction follows that format. | Append a correcting Superseded entry that says only which contradiction in the original row is false and cites the two prompt locations for the current answer. +MINOR | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:18-20,44; todos.md:527-534 | The disposition report and closure row still count two reasoned rejections, while the final table has one rejection disposition: item 7. Item 9 was changed to a parked row, and item 11 is a parked row whose frequency and cost premise is rejected, not a second rejection disposition. | The reported disposition totals no longer match the required exactly-one-disposition accounting after item 9 changed categories. | Change both totals to one reasoned rejection and describe item 11 as a parked row with a rejected premise. +NIT | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:36-38; CLAUDE.md:82,99 | The closure record says items 5 and 8 land as two rules in one paragraph, but the implementation places them in two separate adjacent paragraphs. | This is a small but concrete position claim falsified by the diff and weakens the closure record's precision. | Say the rules land in two adjacent paragraphs in CLAUDE.md section 5 and its template mirror. +END OF FINDINGS (14 total) diff --git a/.context/codex-reviews/gate-b-spec-pass-3-dispositions.md b/.context/codex-reviews/gate-b-spec-pass-3-dispositions.md new file mode 100644 index 0000000..2436905 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-pass-3-dispositions.md @@ -0,0 +1,62 @@ +# Gate-B pass 3 — dispositions (both branches) + +Range f9ed886..c22019b, two sequential single-branch calls. Both files validated. +Spec branch: 2 findings. Quality branch: 4 findings. + +**Format note:** the quality branch used the severity token `IMPORTANT`, which is not one +of §5's four (Blocker/Major/Minor/Nit). The file is otherwise well-formed — four finding +lines, correct terminator, matching count — so the pass was accepted and all four were +treated as Major-equivalent on their content. The pass-4 prompts name the enum explicitly. + +## Applied — committed diff + +1. **MAJOR (spec) — the format-example residual understates the guard.** ACCEPTED, + verified. The residual said "nothing enforces that distinction beyond its shape", while + §8's own held-item list already recorded that claim as wrong: for this change C1d + confines candidates to the label-to-`Columns:` interval and C3 rejects entry-shaped + lines outside it. This is the *understatement* direction of the repo's signature defect, + which the pass-3 prompt explicitly asked to be checked. Rewrote the residual to say both + — two checks guard it once, shape alone guards it thereafter — and marked the §8 held + item as corrected rather than held. Also corrected that bullet's own stated count + ("Four things … deliberately held, not fixed"), which had become false: three of the + four have since been acted on, two by the checks and one by this correction. Marked the + two check-actioned items `*Done:*` with what the check now does. + +2. **IMPORTANT (quality) — the spec's `todos.md` change-surface row names only check 1d**, + while the landed `todos.md` row covers both entry validation *and* the unimplemented + chronology check 1e, matching §8. ACCEPTED, verified by reading both. The spec was stale + against what shipped. Updated the row to name both properties. + +## Applied — scratch harness + +3. **IMPORTANT (quality) — `FNR == NR` misdispatches when the ledger is empty.** + ACCEPTED and **reproduced**: with an empty ledger and the real BASE as second input, the + check reported `undecidable interval: 0 label(s), 1 Columns: paragraph(s)` — BASE's + shape. Every BASE line had landed in the ledger's array, and the `nl == 0` + "empty or unreadable" branch was unreachable for any non-empty BASE, which is always. + Replaced with `FILENAME == ARGV[1]` / `FILENAME == ARGV[2]`. Re-tested: the same input + now reports "the ledger under test is empty or unreadable". + +4. **IMPORTANT (quality) — both calendar-validation fixtures fail without calendar + validation.** ACCEPTED. `entry-bad-date-entry` was sought with the unmutated + `ENTRY_DATE`, and `entry-bad-date-locator` with a hard-coded `ROWDATE`, so identification + failed first and `isdate()` was never reached. Fixed three ways: `ROWDATE` is now + overridable like `ENTRY_DATE`; the matrix passes the mutated date(s) per row; and the + invalid entry date moved from `2026-02-30` to **`2026-08-32`**, which must be both + calendar-invalid *and* lexically on or after the row date, or the on-or-before bound + rejects it and the row again proves nothing. + **Mutation-tested:** replacing `isdate` with a shape-only check flips exactly these two + rows to passing and leaves the other twenty-eight unchanged. + +5. **IMPORTANT (quality) — awk on the directory fixture is host-dependent.** ACCEPTED. + Confirmed this host runs BSD awk (version 20200816), which reads a directory as empty and + says nothing; GNU awk writes a "is a directory: skipped" warning to stderr, which the + runner and matrix both treat as an infrastructure disagreement rather than the expected + failure. So the sh/dash evidence was awk-dependent in a way neither shell run could + reveal. All three checks now stop before any parsing once an input is unreadable, after + the readable-file assertion has recorded the failure. + +## Why pass 3 is not the closing pass + +Both branches returned Major-equivalent findings, so the cycle continues. Every fix above +is amended into the WIP commit, and pass 4 reviews the result. diff --git a/.context/codex-reviews/gate-b-spec-pass-3.md b/.context/codex-reviews/gate-b-spec-pass-3.md new file mode 100644 index 0000000..5ab71a5 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-pass-3.md @@ -0,0 +1,9 @@ +BLOCKER | high | CLAUDE.md:82-83,90-92,451-453; plugins/dev-workflow/commands/workflow-init.md:282-283,290-293,635-637; docs/field-reports/2026-08-16-canvas-a1-a5-field-report.md:130-136 | The new absorb rule applies to any correction-of-a-correction and orders "absorb it, fix it, keep looping", while the preserved Mechanics rule says Minor and Nit findings are collected and never iterated; saying that the Blocker/Major filter still stands does not choose which instruction wins | This changes Part 3's explicitly untouched severity discipline and gives the shipped prompt contradictory actions for a Minor or Nit repair-of-a-repair | Make ancestry decide scope only and leave fix-versus-collect to the existing severity rule, or explicitly restrict absorb-and-fix to Blocker/Major findings; mirror the correction in the inline template +BLOCKER | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:122-147; CLAUDE.md:400-415 | The five-case decision matrix is not a valid 0-of-5 versus 5-of-5 check: cases 1 and 4 omit severity even though the old and new text prescribe different actions by severity, and case 4 supplies only a plateau plus a correction while omitting the sufficient-coverage judgment and standing-Blocker/Major condition required to decide whether the loop may stop | The observed first-run failure of case 5 proves that row could reject a draft, but it does not make the other under-specified rows decisive; the stated limitation omits this gap, so the required battery+check counterfactual remains unestablished and CLAUDE.md calls that a blocking evidence gap | Rewrite the matrix with complete inputs and one expected output per case, including the severity dimension and all three stuck conjuncts, then evaluate those cases against the base and corrected prompt and update the evidence entry to the observed result +MAJOR | high | todos.md:157-166; docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:167-169; plugins/dev-workflow/hooks/codex-gate.sh:823-841 | Item 1 still makes both unsupported inferences the corrected row was meant to remove: it says a pass count of 21 proves the fingerprint was computable and stored, and later says the hash was uncomputable at every pass of the other cycle; the hook increments the pass count independently and writes the fingerprint best-effort, so the observations distinguish neither computation from persistence nor a later successful computation from a failed write | These statements adopt an unverifiable mechanism claim despite the intake's verify-or-reject rule and contradict the row's own endpoint-only uncertainty at todos.md:143-148 | Keep only the observed pass count, messages, and persisted endpoint states, and state that per-pass computation and persistence remain unknown in both the todo row and closure record +MAJOR | high | docs/hardening-log.md:79-80; docs/hardening-log.md:43-45,58-63 | The prompt-missing-stop-condition supersession chain is not in the prescribed form: line 79 restates what the corrected rule means, and line 80 tries to repair it by explicitly referring to "the entry immediately above" even though entries must never reference one another and each governing entry must stand on its own | The last entry is the durable correction later readers rely on, so an order-dependent explanation of another malformed entry defeats the ledger's append-only current-answer procedure | Append a new self-contained supersession for the target row that names only the target row's internal contradiction and cites the two current prompt locations, without restating the answer or mentioning another entry +MAJOR | high | docs/hardening-log.md:78,81; docs/hardening-log.md:58-61 | The latest governing prompt-vague-criteria supersession calls itself a fourth fault not named by the earlier entry and records only that new range and stale-ref defect; under the ledger's last-entry-governs rule this retires the earlier correction of the exact-size, Blocker-start, and false-convergence claims instead of describing the row as it now stands | A later recurrence reader is left with a governing correction that no longer retracts three known-false claims, so the append-only ledger presents the bad hardening as partly current | Append one standalone final supersession that names all four false claims and the obsolete ref in the target row and cites the current prompt locations without restating their answer or referring to prior entries +MINOR | high | todos.md:209-211; docs/superpowers/stories/2026-08-17-arms-race-remedy-as-procedure-story.md:21-34; docs/superpowers/stories/2026-08-04-harden-finding-guard-scope-precheck-story.md:97-105 | The new todo row says the kit states the arms-race idea in exactly one place, but the resulting tree also states it in the new arms-race story and in the amended guard-scope story | This is the standing-lens count drift the review was required to catch and makes a newly adopted evidence claim false on arrival | Qualify the count as the state before this intake or remove the exact-count claim and describe the invariant-checker row as the prior local statement +MINOR | high | plugins/dev-workflow/CHANGELOG.md:31-40,46-50 | The 0.10.0 entry summarizes the stuck rule as reading the Blocker curve after about six passes and merely stating a coverage judgment, omitting that the judgment must affirm sufficient coverage, that a known materially unreviewed area forbids exit, and that no Blocker/Major may remain by choice; it then says no threshold is claimed despite the earlier "after about six passes" phrasing | This shipped description reproduces the insufficient-coverage wording rejected in pass 2 and can be read as a weaker, time-triggered exit than the prompt actually implements | Summarize the three conjunctive conditions and describe six passes only as the single field observation, not as when the reader should stop +MINOR | medium | todos.md:227-235 | Item 9 says the consumer "enforces it by protocol rather than by hope", but the cited mechanism is two imperative lines in a handoff document and nothing verifies that an agent supplies or requests the self-reported context percentage | Calling an instruction enforcement overstates what the cited mechanism proves, the repository's recurring defect class, while the same row acknowledges that the value's reliability is unsettled | Call it a documented instruction-backed protocol and state that no mechanism verifies compliance or the reported percentage +END OF FINDINGS (8 total) diff --git a/.context/codex-reviews/gate-b-spec-pass-4-dispositions.md b/.context/codex-reviews/gate-b-spec-pass-4-dispositions.md new file mode 100644 index 0000000..25efa4d --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-pass-4-dispositions.md @@ -0,0 +1,56 @@ +# Gate-B pass 4 — dispositions (both branches) + +Range f9ed886..d83f5e5, two sequential single-branch calls. Both files validated, both +using the correct severity enum after pass 3's `IMPORTANT` drift. +Spec branch: 1 MINOR. Quality branch: 1 MAJOR + 1 MINOR. + +## Applied — scratch harness + +1. **MAJOR (quality) — `entry-wrong-fingerprint` and `entry-wrong-rowdate` never reach the + locator-equality loop they exist to exercise.** ACCEPTED and **reproduced**: both failed + with `C1d: 0 candidate(s) are the mandated entry`, i.e. at *identification*, because the + matrix left `FP` and `ROWDATE` at their defaults while the fixture mutated them. Deleting + both locator equality comparisons outright left both rows still failing — so they proved + nothing about the property they name. + + This is the third instance of one defect class this cycle (after the append assertions in + pass 2 and the calendar dates in pass 3): **a fixture that mutates part of the locator must + pass the mutated value in, or the check rejects it before the advertised dimension.** Fixed: + `FP` is now overridable, `row()` takes a sixth positional, and both rows pass their own + mutated locator. They now fail with the locator diagnostic + (`matches 0 row(s) by locator, 0 eligible`), and deleting the equality comparison flips + five rows including these two. + + **It also falsified a claim I had committed.** At pass 3 I marked spec §8 held item 1 + `*Done:*` — "the check compares both fields before the bound, and `entry-wrong-rowdate` is + that fixture". The check did compare them; the fixture did not exercise it. That `*Done:*` + note is now rewritten to say what is actually true and to record the mutation test. + +## Applied — committed diff + +2. **MINOR (spec) — §1 describes the pre-change gap in the present tense** ("neither move is + sanctioned", "The gap is live") in the same document whose §2 sanctions the move. ACCEPTED, + verified. Normally a MINOR is collected rather than acted on, but this cycle was already + going to pass 5 for finding 1, and it is the same "a change falsifies its own spec" family + the whole cycle has been about. Added an explicit pre-change framing sentence — which the + story already carries — and moved the section to the past tense. + +3. **MINOR (quality) — two sites still call the correction "prose-only".** ACCEPTED. §8's own + held item 2 names both sites and states the accurate wording ("no standing machine + consumer": the syntax *is* a standing convention every future author must honour, and §6 + adds only a validation-only parser making no ongoing compatibility promise). Corrected both, + and marked held item 2 done. With that, all four of §8's held items are closed, so the + bullet's own summary sentence was corrected too — it had said the second still stands. + +## Note on the pattern + +Three of this cycle's four passes found the same shape of defect in the evidence harness: a +fixture reporting an expected failure that would have reported it just as loudly with the +tested property removed. The fix each time was to make the check reach the property. The +general lesson is recorded in the harness comments and in the evidence entry, not just in +the individual fixes. + +## Why pass 4 is not the closing pass + +The quality branch returned a MAJOR. All fixes above are amended into the WIP commit +(556cbfe), and pass 5 reviews the result. diff --git a/.context/codex-reviews/gate-b-spec-pass-4.md b/.context/codex-reviews/gate-b-spec-pass-4.md new file mode 100644 index 0000000..c3cffc0 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-pass-4.md @@ -0,0 +1,8 @@ +MAJOR | high | CLAUDE.md:82-94; plugins/dev-workflow/commands/workflow-init.md:282-294 | Item 5's "inside the assigned fix set" predicate was dropped: both prompt copies make correction ancestry sufficient for absorption unless the finding also opens a structural or contract question | The field report permits absorbing corrections-of-corrections only inside the assigned fix set; the implemented rule can silently expand assigned work for an out-of-set correction that is neither structural nor contractual | Restore the assigned-fix-set condition in both prompt copies and cover inside-versus-outside-set cases in the decision matrix +MAJOR | high | CLAUDE.md:103-113; plugins/dev-workflow/commands/workflow-init.md:299-309; docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:138 | The new stuck test admits an ordinary clean completion: matrix case 5 has only a MINOR finding and no Blocker or Major, yet declares the stuck exit available | The unchanged severity rule says Minor and Nit are collected and never iterated, so this state has already converged for gate purposes and must finish clean rather than surface a false "will not converge" report | Give clean completion precedence and require a stuck case to contain unresolved iterating work after genuine repair attempts; replace case 5 with such a state +MAJOR | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:122-144 | The seven-case matrix does not satisfy its own complete-input contract: cases 6 and 7 omit ancestry, case 6 supplies neither a specific current severity nor an unambiguous third stuck-condition value and has no action, and case 5 cites a field Blocker while testing a synthetic Minor | The claimed new-text score of 7 of 7 and identical action decisions for all seven are unsupported, so the named check does not establish the prompt's decision coverage or satisfy battery+check as claimed | Rebuild every row with explicit assigned-set status, ancestry, severity, all three stuck inputs where applicable, and one action, then rescore only determinate outputs +MINOR | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:146-155 | The validation narrative was not updated when the matrix grew and reordered: the insufficient-coverage counterfactual is now case 6 but lines 146-151 call it case 5, and the limitation calls seven states "these five states" | These stale identifiers contradict the proposed closing evidence entry and make the check's wiring and stated enumeration boundary unreliable | Change the counterfactual reference to case 6 and the limitation count to seven after the matrix is corrected +MAJOR | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:177-179; todos.md:137-170; plugins/dev-workflow/hooks/codex-gate.sh:833-841 | The closure record concludes that the fingerprint was uncomputable on every counted pass, although the adopted todo and hook show that endpoint unavailable state cannot distinguish computation failure from a computable hash whose best-effort write failed | This adopts an unverifiable mechanism claim and directly violates the round's verify-before-adopt rule and AGENTS.md's prohibition on overstating what a gate proves | State only the observed endpoint values, counts and messages; leave per-pass computation and persistence unknown as the todo already does +MAJOR | high | docs/hardening-log.md:84; docs/hardening-log.md:117 | The final governing prompt-vague-criteria supersession is neither format-compliant nor complete: it restates corrected values and the coverage answer instead of only citing them, while omitting the row's unsupported 34-pass claim and the obsolete "after about six passes" and "every remaining finding individually fixable" criteria | Under the ledger's last-entry-governs rule, this entry retires earlier corrections and leaves a later recurrence reader with an incomplete account of what is false, recreating the defect the appended entry was meant to close | Append another self-contained supersession that names every false claim in the target row without supplying corrected answers, and cite CLAUDE.md plus the inline mirror for the current rule +MINOR | high | docs/superpowers/stories/2026-08-17-arms-race-remedy-as-procedure-story.md:24-27; plugins/dev-workflow/CHANGELOG.md:41-45 | The parked story describes the current §5 heuristic as "after about six passes read the Blocker curve," turning the six-pass field observation into an operational timer | The changelog explicitly says six passes is one observation and not when stopping becomes authorized; this false description can make the later procedure replace the wrong current condition | Describe the trigger as a plateau visible across passes and keep six passes only as the non-normative field observation +END OF FINDINGS (7 total) diff --git a/.context/codex-reviews/gate-b-spec-pass-5-dispositions.md b/.context/codex-reviews/gate-b-spec-pass-5-dispositions.md new file mode 100644 index 0000000..db79602 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-pass-5-dispositions.md @@ -0,0 +1,7 @@ +# Gate B — spec branch — pass 5 dispositions + +1 finding, Major. Accepted. + +1 MAJOR the missing-reason / impossible-answer path relabels a second pause instead of removing it — ACCEPT, and it is the sharpest kind of finding this repo gets: I fixed a contradiction by renaming it. "The story stays unwritten until the human supplies it" IS awaiting a second response, whatever the sentence after it claims, and mislabelling a mechanism is the defect class AGENTS.md names as most persistent here. + Fixed by making the wording true rather than the mechanism different: an answer that cannot be recorded **ends this intake attempt** — say which rule it collides with, stop without writing, hold no pause open. The human's corrected answer arrives at a NEW intake run that opens its own single round with that answer in hand. What the contract forbids is a second round inside one run, not the human coming back — and that is exactly what the grounding floor has always done. + Of the reviewer's two routes, the "revise the contract to permit one validation-repair pause" option was not taken: it would change an approved decision and ripple through skill, spec, plan and CHANGELOG to legitimise something the existing stop-without-writing shape already handles. diff --git a/.context/codex-reviews/gate-b-spec-pass-5.md b/.context/codex-reviews/gate-b-spec-pass-5.md new file mode 100644 index 0000000..2426fff --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-pass-5.md @@ -0,0 +1,14 @@ +BLOCKER | high | CLAUDE.md:72-80,116-120; plugins/dev-workflow/commands/workflow-init.md:272-280,312-316 | The new clean-completion sentence says any Blocker/Major-free pass has satisfied the clean-final-pass rule and should close, but the preserved floor allows a below-three-pass exit only when the pass has zero findings; pass 1 with a Minor now has opposite required outcomes | This silently weakens the mandatory three-pass floor and makes the shipped decision procedure internally contradictory | Qualify the Blocker/Major-free close as applying only after the floor is met, while retaining zero findings as the sole below-floor exit, and mirror that qualification +BLOCKER | high | CLAUDE.md:82-100; plugins/dev-workflow/commands/workflow-init.md:282-300; docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:133-139 | The mandatory assigned-fix-set predicate is never defined or bound to an authoritative artifact, and an out-of-set correction that opens no question is told to stop even though the loop is said to resume only once the question is answered | A reader cannot consistently classify the key scope input or resume the expressly covered no-question branch, so the item-5 procedure leaves real states undecided | Define how the assigned fix set is established and how ambiguity is handled, then give the out-of-set/no-question branch an explicit handback and resume condition in both prompt copies and test boundary cases +BLOCKER | high | CLAUDE.md:76-78,85-87,106-123; plugins/dev-workflow/commands/workflow-init.md:276-278,285-287,302-319; docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:141 | The loop still requires every Blocker/Major to be resolved and re-reviewed, while the stuck exit requires a current regenerating Blocker/Major and matrix case 7 simultaneously says to resolve it and report non-convergence | Stopping before repair violates the mandatory resolve rule, but repairing and rerunning may produce the clean pass that takes precedence, leaving no unambiguous reachable stuck transition | State whether the terminal pass may hand back with its current finding unresolved or specify what post-repair evidence authorizes stopping, then give case 7 one action consistent with that transition +BLOCKER | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:122-146 | Matrix case 6 omits pass number or floor status, so its Minor-only state can mean continue below the three-pass floor or close after it; with the missing input supplied, old section 5 already decides the action, contradicting old 0 of 8, and all eight Action cells are populated despite the seven-row denominator | The named check does not meet its complete-input and one-output contract, and the quoted evidence entry repeats scores that the matrix cannot support | Add floor status to every relevant case and rescore scope, exit, and action per question, correcting the old score and action denominator in both the matrix narrative and closing evidence +MAJOR | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:154-166 | The narrative says the matrix rejected three drafts, but it records an actual matrix failure only for case 8; cases 2, 3, and 6 were added after later Gate-B passes found defects, with no cited run against preserved earlier drafts | Retrospective regression cases are useful but do not establish that this verification independently rejected those drafts, so the counterfactual claim overstates its evidence | Say which draft the matrix actually rejected and describe the other cases as retrospective additions, or cite reproducible runs of those cases against the preserved drafts +MAJOR | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:113-120,144-171 | The final named verification replaces the earlier paragraph-diff and mirror-parity check but applies the eight cases only to the root CLAUDE.md text at the base and HEAD, never to the workflow-init inline template | The counterfactual is not wired at the second location where the required rule could be omitted or diverge, so it cannot substantiate that both prompt copies implement items 5 and 8 | Apply and score the same cases independently against the inline template, including a counterfactual with only one copy corrected, and record the mirror results and limitation +MAJOR | high | docs/hardening-log.md:83,117 | The latest supersession for the immutable prompt-missing-stop-condition row says its pass-count contradiction is the whole falsehood, but that row also states every correction-of-a-correction is absorbed and omits the now-essential inside-assigned-fix-set predicate | The governing ledger chain still preserves the exact overbroad absorption rule that the shipped prompt and matrix case 3 reject, so later recurrence analysis can apply a false guard | Append a new self-contained supersession naming both the internal pass-count contradiction and the missing assigned-set predicate, citing the two shipped prompt blocks +MAJOR | high | todos.md:129-170 | The item-1 title categorically says consumer Gate-B cycles cannot record a usable fingerprint, while its evidence establishes only two closing endpoint shapes and expressly leaves per-pass computation, persistence, stored values, and cause unknown | The actionable summary outruns the verified evidence and can misdirect later hook work toward a generalized computation or storage defect | Retitle the row around the two observed cycle closes with missing or unavailable persisted fingerprints and retain the stated mechanism unknowns +MAJOR | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:184-190; docs/hardening-log.md:85 | The closure says the erroneous stuck-rule figures came from the consumer memory note, but that note records 19 passes rather than the ledger's erroneous 34; the final supersession itself says 34 came from counting pass artifacts | The closure's provenance account remains factually false even after the figures were superseded, undermining the claim that every adopted statement was traced to its evidence | Separate the provenance: identify which stale figures came from the memory note, state that 34 came from artifact counting, and keep the current taxonomy as the stricter record +MAJOR | high | docs/superpowers/stories/2026-08-17-arms-race-remedy-as-procedure-story.md:24-30,101-104 | The story first says this change already gives section 5 a three-condition stuck heuristic, then leaves an open question asking whether the procedure supplies section 5's missing criterion and whether that is the same edit already made | The parked disposition is internally stale and can reopen or duplicate a decision the same story records as completed | Rewrite the open question to concern only the remaining future-procedure uncertainty, without calling the already-added section-5 criterion missing +MINOR | high | plugins/dev-workflow/CHANGELOG.md:27-46 | The release note says both shipped rules were minted in field use, although the consumer record minted the absorb-versus-stop ruling while the exact three-condition stuck heuristic was constructed during intake from limited field measurements after the original interpretation was retracted | The provenance wording makes the new heuristic sound field-validated more broadly than the cited evidence supports | Attribute the absorb rule to field use and describe the stuck heuristic as derived during intake from the stated field observations and limitations +MINOR | medium | docs/hardening-log.md:117 | The row says the absorb-versus-stop rule was field-minted across approximately 150 gate passes, but approximately 150 is the consumer project's total gate-pass count, not evidence that this ruling was observed or applied across all of them | The sentence inflates the observation base for the rule and overstates what the consumer record proves | Say the rule was field-minted in a consumer that had run approximately 150 passes, or cite the actual passes on which the rule was observed +MINOR | high | docs/hardening-log.md:85 | The final supersession says the exact line-count and total-pass-count claims did not hold, but the cited correction establishes that those figures were unsupported or sourced from a different counting method rather than demonstrating the numeric alternatives needed to prove them false | The ledger converts an evidence-quality failure into a stronger factual refutation than its cited sources support, contrary to the intake rule to adopt nothing unverifiable | Characterize those two claims as unsupported or misattributed while reserving false for the three rule interpretations the prompt text actually contradicts +END OF FINDINGS (13 total) diff --git a/.context/codex-reviews/gate-b-spec-pass-6-dispositions.md b/.context/codex-reviews/gate-b-spec-pass-6-dispositions.md new file mode 100644 index 0000000..907beaf --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-pass-6-dispositions.md @@ -0,0 +1,76 @@ +# Gate-B pass 6 — dispositions (spec branch) + the systematic check-claim sweep + +Range f9ed886..556cbfe. Spec branch: 4 findings, all verified accurate before acting. +Quality branch was not run at this commit — the sweep below changes the spec, so both +branches run together at pass 7. + +**Human decision, taken at the stop:** replace incremental site-by-site correction with one +systematic sweep of every check-claim sentence in §4, §6 and §8, fixed as a single edit set, +with the inventory recorded here. Reason: findings 1 and 2 below were *introduced by this +cycle's own earlier fixes* — one site corrected, its synonym left standing — which is the +defect class `AGENTS.md` names as the repo's most persistent, recurring inside the review +meant to catch it. Patching instances was reproducing it. + +## The sweep — method + +`grep -nE 'check ?1|check ?2|check ?3|C1[a-f]|catch|detect|validat|enforce|prove|guard|confine| +reject|nothing (checks|consumes)'` over §4 (370–403), §6 (419–706) and §8 (743–931), then every +hit read against what the harness actually does. A second pass grepped the scope-overclaim +vocabulary — `anywhere`, `whole file`, `entire file`, `unconstrained` — because that is the +wording the earlier misses hid behind. + +## Inventory — claim → what the harness does → verdict + +| # | Site | Claim | Harness | Verdict | +|---|---|---|---|---| +| 1 | §6 preamble | "each check gives the observation that fails without the change" | `C1b`/`C1c` pass on the untouched base by construction; `1e` unimplemented | **WRONG — fixed** | +| 2 | §6 `1d` oracle | "the mandated entry, and only entries this change **adds**" | selects by date+rowdate+fingerprint, *then* proves absent at base | **WRONG — fixed** | +| 3 | §6 Check 3 intro | "check 1 matches the entry anywhere in the file" | `C1a`/`C1d` confine candidates to the label-to-`Columns:` interval; an absent/duplicated/reversed interval exits 2 | **WRONG — fixed** | +| 4 | §8 held item 3 | marked `*Done:*` at pass 3 | the *check* was fixed then; the oracle wording only now | **INCOMPLETE — fixed** | +| 5 | §4:393 | "check 2 treats a missing end sentinel as a failure, anchors fail first" | `C2a` runs before `C2b`; both fail pre-change | accurate | +| 6 | §6:512 | "no check validates pre-existing entry immutability" | nothing does | accurate | +| 7 | §6:593 | "check 1 would pass with the template untouched" | `C1a`/`C1b`/`C1d` read only the ledger; `C1c`'s paragraph is identical either way | accurate | +| 8 | §6 Check 3 "what it does not do" | "confirms only that the template carries no *label*" | `tpl_label_count_is_zero` counts the label alone | accurate | +| 9 | §8:744-754 | "three checks run once, nothing standing" | true; correctly says "beyond check 1's one-time assertion" | accurate | +| 10 | §8:790-792 | "a live entry in the template without a label is caught by nothing" | `C3` tests the label only; `C2`'s region ends at the sentinel, above the block | accurate | +| 11 | §6:486 | 1c's four-input guard rationale | delimiters asserted present and unique in all four inputs before any comparison | accurate | +| 12 | §6:548, 587 | "nothing validates fragment-narrowed matching / entry prose" | no fragment in this change; `1f` is a human read | accurate | + +Four wrong, eight accurate. The four were fixed in one edit set; nothing else in §4, §6 or §8 +makes a check-claim that the harness contradicts. + +## The fixes + +1. **§6 preamble** now splits falsifying observations into two kinds and names the labels: + `1a`, `1d`, `2a`, `2b`, `3` fail on the untouched base; `1b` and `1c` are protective and + green there by construction; `1e` is unimplemented; `1f` is a named read, not an executable + label. Matches the plan's own per-label table exactly, which was checked before writing. +2. **§6 `1d`'s scope oracle** now states the two-step order — identify the §3.1 entry by its + fields, *then* prove it absent at base — and says why the order must be stated rather than + inferred from this change, where the base carries no block and the steps coincide. +3. **§6 Check 3's introduction** now says what check 1 already establishes (interval + confinement, undecidable-interval handling) and limits check 3 to the actual remainder: + the sentinel lower bound, entry-shaped lines outside the interval, blank-line structure, + and the template label. +4. **§8 held item 3** now records that the check was corrected at pass 3 and the oracle wording + only at pass 6 — with the gap named, since the item stood marked done in between. + +## Finding 3 — the plan. Human decision: leave it. + +Plan line 243 specifies `entry-bad-date-entry` as `2026-02-30`; the harness uses `2026-08-32`. +**The plan is not amended.** It landed in `f9ed886`, outside the reviewed range, and is the +record of what was intended at execution time — this repo does not rewrite executed plans to +match later rules (§4's change-surface row says exactly that of another plan). No `todos.md` +row either: a historical record needs no reconciliation. + +**Why the harness diverged, stated in full and carried into the commit body:** `2026-02-30` is +lexically *earlier* than the locator's `2026-07-20` row, so with calendar validation removed it +is still rejected — by the on-or-before bound — and the fixture proves nothing about the +property it names. The fixture value must be calendar-invalid **and** lexically on or after the +locator row date. `2026-08-32` is both. Mutation-tested: replacing `isdate` with a shape-only +check flips exactly that row and `entry-bad-date-locator`, leaving the other 28 unchanged. + +## State after this pass + +Battery green. Harness green: 30 matrix rows, 0 disagreements under `sh` and `dash`; C1/C2/C3 +0 failures on the real tree. Fixes amended into the WIP; pass 7 runs both branches. diff --git a/.context/codex-reviews/gate-b-spec-pass-6.md b/.context/codex-reviews/gate-b-spec-pass-6.md new file mode 100644 index 0000000..d2f4d1b --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-pass-6.md @@ -0,0 +1,10 @@ +BLOCKER | high | CLAUDE.md:87-93; plugins/dev-workflow/commands/workflow-init.md:287-293; docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:154 | The assigned-fix-set definition automatically includes every Blocker/Major finding in the pass being answered, so a new out-of-scope Major is in the set by definition and matrix case 3's Major with `In set: no` is unreachable | The core item-5 boundary silently expands the cycle to every new serious finding, while the claimed 0-of-7 to 7-of-7 check relies on an impossible state and cannot validate that boundary | Define the set before classifying the new pass, for example as approved story/plan scope plus previously accepted repair obligations, then test a newly reported out-of-set Blocker/Major against that fixed set in both copies and the matrix +BLOCKER | high | CLAUDE.md:119-141; plugins/dev-workflow/commands/workflow-init.md:315-338; docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:138-166 | At pass 4 or later a Blocker/Major-free pass at the floor must close, but the same pass can have rising Minor totals plus an instrument/prose cluster and therefore meet the mandatory two-tell stop; no precedence resolves close versus surface, and the matrix calls its inputs complete while omitting every tell | A converged gate has two opposite mandatory outcomes, and case 6 plus the quoted battery+check evidence do not actually establish that the new prompt decides all relevant states | State whether clean completion overrides the two-tell rule or exclude clean passes from that trigger, mirror the decision, and add clean/non-clean two-tell cases to the matrix and counterfactual +MAJOR | high | CLAUDE.md:130-141,169-187; plugins/dev-workflow/commands/workflow-init.md:327-338,356-374 | The new three-line duty never identifies who emits the `pass report`, where those lines live, or an example format, while the only explicit pass reply nearby is required to contain ONLY one line and the findings file forbids extra prose | A compliant reader can either invalidate the findings protocol by adding the three lines or omit the new duty by treating the one-line reply as controlling, so the maintainer extension is not operationally deterministic | Name the outer Claude status report as the carrier, explicitly exclude the Codex reply and findings file, and show the exact three-line format in both prompt copies +MAJOR | high | docs/superpowers/stories/2026-08-17-field-intake-canvas-a1-a5-report-story.md:53-54; docs/hardening-log.md:121-122 | The story requires every ledger row appended this round to state the guard-scope precheck outcome and quote the examined prior guard, but both new rows merely say they are the first base-class row and describe their new guards; the no-match result and nearest-guard readings exist only in the separate closure record | The two harden-finding dispositions do not meet an explicit acceptance criterion at the artifact where future recurrence logic reads them | Record the no-prior-match outcome and the quoted guard-scope comparison in a ledger-compatible append-only correction for each row, or obtain an explicit story amendment if the closure record is intended to own this evidence +MAJOR | high | docs/hardening-log.md:89; docs/hardening-log.md:121 | The final governing `prompt-missing-stop-condition` supersession says the target row's `finding` states unconditional absorption, but the finding only says absorption and handback were both defensible; the unconditional rule is in the target row's `ref`, and the same entry also associates item 8's reporting duty with the item 5 guard | The last entry governs the immutable row, so the durable correction is not self-containedly accurate and blurs the report's one-disposition ownership of the maintainer extension | Append a new governing entry that attributes the missing assigned-set condition to `ref`, keeps all three actual faults self-contained, and leaves the reporting duty attached only to the item-8 guard unless shared ownership is explicitly intended +MAJOR | high | docs/hardening-log.md:62,76,80-81,88 | Four entries added in this diff explicitly refer to another supersession entry as `the entry immediately above`, `the earlier entry above`, or `the two above`, despite the ledger rule that entries never reference one another | Those permanent cross-links make the append-only correction chain order-dependent and directly violate the governing format; later entries making them historical does not make the malformed lines disappear | Do not present these four as governing; because complete entries are protected even before commit, obtain maintainer direction for repairing them or define and apply a policy-compatible correction mechanism that does not rely on another entry +MINOR | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:24; todos.md:129 | Item 1's disposition pointer still quotes the old categorical row title, while the corrected row now begins `Two consumer Gate-B cycles closed with no usable fingerprint persisted` | The eleven-item audit no longer resolves its item-1 owner by exact search and misses the title change the standing lens was meant to propagate | Replace the stale quoted title with the current row title +MINOR | high | plugins/dev-workflow/CHANGELOG.md:47-50,60-64 | The first release-note paragraph says any Blocker/Major-free pass has satisfied the clean-final-pass rule and closes, while a later bullet correctly limits that closure to at or above the floor | The changelog preserves the exact below-floor contradiction fixed in the shipped prompts and gives readers two incompatible descriptions of 0.10.0 | Add the at-or-above-floor qualifier to the first statement so both descriptions agree +MINOR | high | docs/field-reports/2026-08-16-canvas-a1-a5-dispositions.md:47-50; CLAUDE.md:133-139 | The closure record counts pass 4 requiring clean-completion precedence and pass 5 bounding that rule by the floor as a require-withdraw pair, although the prompt defines such a pair as demanding what an earlier pass removed and pass 5 only qualified the earlier requirement | The record overstates one of the tells used to claim the new reporting duty was exercised immediately, contrary to its verified-claims standard | Remove that pair from the tell count or cite an actual earlier removal that pass 5 required restoring +END OF FINDINGS (9 total) diff --git a/.context/codex-reviews/gate-b-spec-pass-7-dispositions.md b/.context/codex-reviews/gate-b-spec-pass-7-dispositions.md new file mode 100644 index 0000000..9cbafbc --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-pass-7-dispositions.md @@ -0,0 +1,49 @@ +# Gate-B pass 7 — dispositions (spec branch) + +Range f9ed886..68c99c5, after the systematic sweep. 2 findings, both verified, both fixed. +Quality branch not run at that commit — these fixes change the spec, so both branches run +together at pass 8. + +## Applied + +1. **MAJOR — the parity narration understates what 2b proves.** §4 said the two files are + "byte-identical **modulo hard-wrap position**", and §6's 2a oracle said "Both surfaces are + hard-wrapped **at different columns**". Verified against the tree: the two delimited regions + are byte-identical including line breaks — `cmp` clean, 70 lines each — because the text is + generated once and inserted into both. And since pass 3, 2b compares them **raw**, so a + rewrap of one surface alone fails it. + + **Another stale claim created by one of this cycle's own fixes**: "modulo wrap position" was + accurate while 2b compared paragraph-joined views, and became an understatement the moment + pass 3 changed 2b to a raw comparison. Third instance of that shape this cycle. + + Fixed both sites. §6's oracle now separates the two directions explicitly — wrap-insensitive + **presence** for 2a, because anchors straddle line breaks *within* a surface; byte-exact + **parity** for 2b, with a note that the wrap-insensitivity must not leak from one into the + other. + +2. **MINOR — the C1d oracle names a fixture that does not exist.** It cited "the fragmentless + locator matching **both** `2026-07-18` `docs-drift` rows" as the many-match fixture. The + harness has no such fixture; `rows-two-matching` synthesises a second row carrying the + *mandated* `2026-07-20` locator. Verified both halves: the docs-drift pair really is repeated + in the ledger (checked with an escape-aware parser — it is the only repeated date+fingerprint + pair, so §2.2's non-uniqueness claim stands), but the mandated entry's locator does not match + it, so it cannot exercise C1d's many-match branch at all. The spec had conflated "the ledger + contains a non-unique pair" with "that pair is the fixture". Fixed to name the constructed + fixture and say why the real pair cannot serve. + +## The sweep's vocabulary was too narrow — widened + +Pass 6's sweep grepped for check names and enforcement verbs. **Neither pass-7 finding names a +check**: one is a claim about the artifacts ("byte-identical modulo wrap"), the other about a +fixture. The sweep was re-run with a widened vocabulary — parity/identity/wrap terms, and every +fixture name in the harness cross-checked against every fixture the spec names. That found +exactly these two and nothing further. + +Recorded because the lesson generalises: a check-claim sweep that greps only for *check names* +misses claims about what the checked artifacts are, which is the same defect one level down. + +## State + +Battery green. Harness green: 30 matrix rows, 0 disagreements under `sh` and `dash`. Fixes +amended into the WIP (`5f24776`). Pass 8 runs both branches. diff --git a/.context/codex-reviews/gate-b-spec-pass-7.md b/.context/codex-reviews/gate-b-spec-pass-7.md new file mode 100644 index 0000000..f842675 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-pass-7.md @@ -0,0 +1,5 @@ +MINOR | high | docs/sparring-briefing.md:50-53; docs/coding-workflow.md:241-258; plugins/dev-workflow/CHANGELOG.md:39-43 | The new “at server startup” timing is broader than pinned mcp-codex-dev@1.0.1: startup preloads only the server cwd's resolved project root, while any other workingDirectory is loaded and cached on its first tool call | In an alternate repo or worktree not yet seen by the server, a post-startup config edit does take effect, so the corrective guidance and release note are themselves false on that supported path | Describe the boundary as “on first load per resolved project root (the launch root is preloaded at server startup), then cached until restart,” and synchronize all three copies +MINOR | high | plugins/dev-workflow/CHANGELOG.md:35-48 | The entry says docs/prompt-standards.md item 11's fourth-correction rule required deleting the mechanism claim and that “this entry describes no state machine,” yet the preceding bullet narrates the chained-command timing, pre-add index read, and empty-path-list branch | The release note contradicts its own stated requirement and recreates the drift-prone hook narration this correction cycle says it deleted | Delete the chained-command failure-mode bullet; retain the observable empty-commit behavior here and leave the existing detailed event-timing account in todos.md +MINOR | high | docs/sparring-briefing.md:50-53; plugins/dev-workflow/CHANGELOG.md:38-43 | The requested task is limited to correcting the empty record-only commit claim in CLAUDE.md and its workflow-init mirror plus the required version bump, but this range also folds in a separate CodeRabbit model-selection correction and changes an additional prompt artifact | An unrelated prompt rule is shipping under an evidence entry whose named verification covers only the empty-commit claim, expanding the reviewed product surface beyond the stated task | Remove the model-selection edits from this range and handle them in a separately authorized change, or explicitly add that correction to the task/spec and its evidence obligations +MINOR | high | plugins/dev-workflow/CHANGELOG.md:38-43 | The 0.9.1 correction note names only docs/coding-workflow.md, although the 0.9.0 entry says both docs/coding-workflow.md and docs/sparring-briefing.md shipped the bad configured-value rule and this commit actually corrects the sparring copy | The release history does not accurately identify both affected copies or which one this release changes | If this secondary correction remains in 0.9.1, name both documents and state that the current range updates docs/sparring-briefing.md while the parent commit updated docs/coding-workflow.md +END OF FINDINGS (4 total) diff --git a/.context/codex-reviews/gate-b-spec-pass-8-dispositions.md b/.context/codex-reviews/gate-b-spec-pass-8-dispositions.md new file mode 100644 index 0000000..5f4b1a6 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-pass-8-dispositions.md @@ -0,0 +1,7 @@ +# Gate B — spec branch — pass 8 dispositions + +1 finding, Major. Accepted. The quality branch reported the identical defect independently. + +1 MAJOR the multi-story aggregation rule still says EVERY cited story satisfies a mode and supplies an evidence entry — ACCEPT. My pass-7 fix corrected the call-contract paragraph and missed its sibling two paragraphs above it, so the subsection contradicted itself: one paragraph said an unprofiled cited story contributes only a path, the other said each cited story owes an entry. An author following the aggregation paragraph in a mixed cycle would try to manufacture evidence for an unprofiled story and block Gate B — tightening exactly the in-flight behaviour the rollout promises to leave alone. + Qualified in all four places (spec, repo §5, inline §5 copy, plan block): each cited PROFILED story satisfies its mode and suffix with its own entry; a cited unprofiled story has no mode and owes no entry, contributing only its path. + Class note, third instance this cycle: fix one site, miss its sibling. A grep for the changed claim now follows every wording fix, not just a resync of the copies. diff --git a/.context/codex-reviews/gate-b-spec-pass-8.md b/.context/codex-reviews/gate-b-spec-pass-8.md new file mode 100644 index 0000000..7d8c0e9 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-pass-8.md @@ -0,0 +1,5 @@ +MAJOR | high | docs/hardening-log.md:105; plugins/dev-workflow/CHANGELOG.md:43-50 | The only appended hardening row processes Greptile's empty-commit overclaim under `unverified-enforcement-claim`; the accepted CodeRabbit reviewer-model correction has no `harden-finding` result, although `plugins/dev-workflow/commands/process-pr-review.md:167-170` requires every accepted actionable finding matching an existing class to take that path and this config-versus-cache mismatch matches `docs-drift` under `docs/hardening-taxonomy.md:31-38` | The pass-7 MAJOR is only half addressed, so the ledger loses the CodeRabbit recurrence and any rung decision for it | Run `dev-workflow:harden-finding` for the reviewer-model finding and append its `docs-drift` result, or record the skill's grounded reason if it determines no row is warranted +MINOR | high | plugins/dev-workflow/CHANGELOG.md:43-50; commit 852520ef Evidence | This range changes the reviewer-model attribution rule, but the durable evidence entry names only the battery and the empty-commit hook check; it omits the stated verification of caching behavior against pinned `mcp-codex-dev@1.0.1` | The evidence record does not cover one of the two behavior corrections actually shipped in the reviewed range, so a later reader cannot reconstruct why the model guidance is trusted | Add the pinned-server source verification as a named evidence item in the commit body and revalidate that exact entry before the next pass +MINOR | high | docs/coding-workflow.md:241-243 | The first timing paragraph still says the model chain is resolved once per project root "at server startup"; pinned `mcp-codex-dev@1.0.1` preloads only the launch root at startup and loads any other resolved project root on its first call, as the later paragraph at lines 255-259 correctly states | The document remains internally contradictory and still tells users too broadly that post-startup edits cannot affect a not-yet-seen root | Say it is resolved on first load per resolved project root, with the launch root preloaded at startup, then cached until restart +MINOR | high | plugins/dev-workflow/CHANGELOG.md:49-50 | The entry says `docs/coding-workflow.md` was corrected in this branch's parent commit and only `docs/sparring-briefing.md` was corrected here, but base `d5b47bd2` contains the old coding-workflow wording and commit `852520ef` changes both files | The release history misattributes which commit introduced the correction and contradicts the actual fixed range | Attribute both document corrections to `852520ef` or remove the unsupported parent-versus-current split +END OF FINDINGS (4 total) diff --git a/.context/codex-reviews/gate-b-spec-pass-9-dispositions.md b/.context/codex-reviews/gate-b-spec-pass-9-dispositions.md new file mode 100644 index 0000000..0998770 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-pass-9-dispositions.md @@ -0,0 +1,6 @@ +# Gate B — spec branch — pass 9 dispositions + +1 finding, Major. Accepted. + +1 MAJOR docs/getting-started.md still says "Trivial changes skip the ceremony" — ACCEPT, and it is the most interesting finding of the pass because the sentence is PRE-EXISTING and untouched by this change. It was true before profiles and is false after them: it keys skipping on triviality alone, implies the whole ceremony can go, and would let a risk-trivial but security-relevant story skip Gate A. My change invalidated a sentence I never edited, which no parity check or resync could ever have caught — only a reader comparing old prose against new rules. + Qualified: a `trivial` profile unlocks the GATE-B skip only, only at effective level 0, with the battery still owed and Gate A's floor unchanged; an unprofiled story keeps the prior judgement-based skip. diff --git a/.context/codex-reviews/gate-b-spec-pass-9.md b/.context/codex-reviews/gate-b-spec-pass-9.md new file mode 100644 index 0000000..88dd178 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-pass-9.md @@ -0,0 +1,3 @@ +BLOCKER | high | docs/coding-workflow.md:261; docs/sparring-briefing.md:55; docs/hardening-log.md:106; plugins/dev-workflow/CHANGELOG.md:49 | The named health probe reads only `checks.config.effective.model`, but pinned `mcp-codex-dev@1.0.1` resolves gate models through `getToolConfig`: `tools.exec.model` or `tools.review.model` overrides the top-level model, including the documented `CODEX_DEV_REVIEW_MODEL` Gate-B surface; if neither level names a model, Codex may select its own default and health establishes no model at all. | A pass can still record the wrong reviewer family, so the 0.9.1 correction and ledger claim remain false and check 2's evidence does not validate the shipped rule. | Resolve the cached per-tool value (`tools.exec.model` for Gate A or `tools.review.model` for Gate B, falling back to `model`), define a verifiable route for the no-explicit-model case, and update both documents plus the changelog, ledger claim, and evidence before re-review. +NIT | high | todos.md:179 | The occurrence-4 update records a separate compound-command/docs-only false-positive incident; it is neither one of PR #24's two requested bot fixes nor the required version bump or hardening row per fix. | This adds unrelated backlog churn to a deliberately surgical correction and makes the commit harder to attribute to the stated task. | Remove this hunk from the range or split it into separately scoped work. +END OF FINDINGS (2 total) diff --git a/.context/codex-reviews/gate-b-spec-pr15-pass-1.md b/.context/codex-reviews/gate-b-spec-pr15-pass-1.md new file mode 100644 index 0000000..9f843dd --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-pr15-pass-1.md @@ -0,0 +1,3 @@ +MAJOR | high | docs/superpowers/stories/2026-07-26-risk-security-validation-profiles-story.md:24 | The desired outcome still says `trivial` legitimizes the Gate-B skip, silently retaining the old risk-only condition even though amended AC 6 and the shipped command require effective level 0 (risk `trivial` and security `none`) in addition to a behaviorally trivial fix | The approved story gives two conflicting definitions of skip eligibility, so a later implementer can treat a security-relevant trivial-risk story as eligible and reintroduce the gate-off path this fix is meant to close | Rewrite the desired-outcome sentence to say effective level 0 legitimizes the documented skip, while preserving the separate behaviorally-trivial-fix condition +MINOR | high | plugins/dev-workflow/commands/process-pr-review.md:159 | The mixed-story decision procedure says “an unprofiled story on the old judgement call,” omitting the verb and leaving it unclear whether such a story is eligible, ineligible, or merely evaluated by that judgement | Two readers can resolve mixed profiled/unprofiled citations differently, making the all-stories condition unfollowable in exactly the branch this rewrite added | Change it to “an unprofiled story is eligible on the old judgement call” +END OF FINDINGS (2 total) diff --git a/.context/codex-reviews/gate-b-spec-pr15-pass-2.md b/.context/codex-reviews/gate-b-spec-pr15-pass-2.md new file mode 100644 index 0000000..cdf1332 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-pr15-pass-2.md @@ -0,0 +1,2 @@ +MAJOR | high | docs/superpowers/specs/2026-07-26-risk-security-validation-profiles-design.md:394 | The new surfaces-touched sentence says the Gate-B skip decision keys on the story profile “not on fix size,” which drops the independent requirement that the fix itself be behaviourally trivial and contradicts the story §2 and process-pr-review’s explicit two-condition rule. | A future implementation following the design can allow an effective-level-0 profile to skip Gate B for a behaviour-changing fix, creating the forbidden new gate-off reach. | State that profile eligibility is one independent condition and that the fix must also be behaviourally trivial (judged by effect, not line count). +END OF FINDINGS (1 total) diff --git a/.context/codex-reviews/gate-b-spec-pr15-pass-3.md b/.context/codex-reviews/gate-b-spec-pr15-pass-3.md new file mode 100644 index 0000000..1f5781f --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-pr15-pass-3.md @@ -0,0 +1,4 @@ +MAJOR | high | docs/superpowers/plans/2026-07-26-risk-security-validation-profiles.md:15 | The architecture summary says `process-pr-review` keys the Gate-B skip on the story profile, omitting the independent requirement that the fix itself be behaviourally trivial. | This restates the two-condition rule as profile-only and contradicts both the settled no-new-gate-off decision and the shipped command. | State that a skip requires both a behaviourally trivial fix and eligibility of every cited story. +MAJOR | high | docs/superpowers/stories/2026-07-26-risk-security-validation-profiles-story.md:50 | AC 6 says the Gate-B skip is available only at effective level 0, but unlike §2 it omits that the change itself must also be behaviourally trivial. | The acceptance criterion permits a behavioural change to skip Gate B solely because its story profile is eligible, weakening the documented rule and making the story internally inconsistent. | Amend AC 6 to require both a behaviourally trivial change and effective level 0 for every cited profiled story, while preserving the unprofiled-story behavior. +MAJOR | high | plugins/dev-workflow/CHANGELOG.md:48 | The `process-pr-review` release note says the skip is decided from the story profile rather than fix size alone, replacing the old condition instead of stating that profile eligibility and behavioural triviality are both required. | The shipped behavior is documented as profile-only, so readers can infer that an eligible profile makes a non-trivial fix skippable. | Rewrite the bullet to say the fix must be behaviourally trivial and every cited story must be eligible; clarify that effect, not line count, determines triviality. +END OF FINDINGS (3 total) diff --git a/.context/codex-reviews/gate-b-spec-pr15-pass-4.md b/.context/codex-reviews/gate-b-spec-pr15-pass-4.md new file mode 100644 index 0000000..c73bd30 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-pr15-pass-4.md @@ -0,0 +1,2 @@ +MAJOR | high | docs/superpowers/specs/2026-07-26-risk-security-validation-profiles-design.md:259-276 | The rewritten skip section still says risk `trivial` “supplies the recorded reason” and that “a skipped story is still mode `battery`” with an evidence entry; those statements collapse the behavioural-triviality condition into the profile and incorrectly include unprofiled skips, even though the next paragraph and §7 allow an unprofiled story to skip while owing no mode-derived evidence or entry | The spec remains internally inconsistent with the settled two-condition rule and compatibility guarantee, so an implementer can treat the profile value as the skip justification or manufacture evidence for an unprofiled story | State that risk `trivial` contributes only one part of profiled-story eligibility, and scope the evidence paragraph explicitly to a skipped profiled story; say separately that an unprofiled skip records only its reason and battery result +END OF FINDINGS (1 total) diff --git a/.context/codex-reviews/gate-b-spec-pr15-pass-5.md b/.context/codex-reviews/gate-b-spec-pr15-pass-5.md new file mode 100644 index 0000000..6db0223 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-pr15-pass-5.md @@ -0,0 +1,3 @@ +MAJOR | high | CLAUDE.md:300 and plugins/dev-workflow/commands/workflow-init.md:475 | The canonical and scaffolded §5 copies still say the skip reason sits “beside the evidence entry” and that “its entry still lands” before only afterward saying an unprofiled story owes no mode-derived evidence; unlike the rewritten spec, these sentences do not scope the entry to profiled stories or distinguish an unprofiled story's battery result. | An agent following either product prompt can manufacture an evidence entry for an unprofiled skip, contradicting the compatibility guarantee and the settled profiled/unprofiled split. | Mirror the spec's explicit split in both copies: a skipped profiled story lands its evidence entry; a skipped unprofiled story records only the skip reason and battery result, with no mode-derived entry. +MAJOR | high | docs/superpowers/plans/2026-07-26-risk-security-validation-profiles.md:682 | The claimed baseline capture is inside a shell comment, so the executable sequence starts by diffing `/tmp/pre-fix-status` without creating it; `\|\| true` then masks the missing-file failure, and a stale pre-existing file can silently attribute the wrong paths when the worktree was already dirty. | The fix loop cannot reliably distinguish pre-existing dirt from the current fix and may stage unrelated changes or omit intended ones from the reviewed WIP snapshot. | Before editing, execute a real capture to a newly created path (for example `baseline_file=$(mktemp)` followed by `git status --short > "$baseline_file"`), compare against that exact path after editing, and remove it after the diff. +END OF FINDINGS (2 total) diff --git a/.context/codex-reviews/gate-b-spec-pr15-pass-6.md b/.context/codex-reviews/gate-b-spec-pr15-pass-6.md new file mode 100644 index 0000000..2b2ee72 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-pr15-pass-6.md @@ -0,0 +1,3 @@ +MAJOR | high | docs/superpowers/plans/2026-07-26-risk-security-validation-profiles.md:683-690 | The fix-loop compares only `git status --short` summaries, so if an accepted fix further edits a path that was already dirty with the same status code, the before and after lines are identical and the claimed delta omits that fix | On an already-dirty worktree the agent can fail to attribute and stage the accepted fix, so the next Gate-B pass may review the pre-fix range despite the recipe claiming pre-existing dirt cancels safely | Capture content-level state before editing (for example with a temporary index/tree or per-path hashes/diffs), compare it after editing, and retain the resulting path set through explicit staging +MINOR | high | docs/superpowers/plans/2026-07-26-risk-security-validation-profiles.md:688 | `git status --short \| diff "$baseline" -` is a pipeline whose status is only `diff`'s status in POSIX shell, so a failing `git status` can be masked; additionally, the expected changed case makes `diff` exit 1, which aborts the advertised sequence under `set -e` before cleanup and staging | The supposedly executable failure-transparent loop can either continue with invalid comparison input or stop on the normal case, leaving the baseline behind and the fix unstaged | Write the post-edit status/content snapshot with a separately checked command, compare the two files while explicitly treating diff status 1 as “different” and statuses greater than 1 as errors, then clean up with a trap +END OF FINDINGS (2 total) diff --git a/.context/codex-reviews/gate-b-spec-pr15-pass-7.md b/.context/codex-reviews/gate-b-spec-pr15-pass-7.md new file mode 100644 index 0000000..e7356ba --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-pr15-pass-7.md @@ -0,0 +1,2 @@ +MAJOR | high | docs/superpowers/plans/2026-07-26-risk-security-validation-profiles.md:683 | The simplified staging recipe names the files edited but does not protect pre-existing changes within those same files: `git add ` stages each file wholesale, while `git diff --cached --name-only` verifies only path names and cannot reveal unrelated hunks. The following explanation explicitly invokes an already-dirty target path, which is precisely the case this recipe cannot safely handle. | Following the plan can silently fold a user's pre-existing same-file work into the WIP snapshot and final commit, so the staged content is broader than the accepted review fixes even though the advertised verification passes. | Keep the path-based simplification, but add a pre-edit/pre-staging guard: if any target path was already dirty before the fix, stop and ask or stage only the intended patch and inspect `git diff --cached`; use whole-path `git add` only when the target path was clean before the fix. +END OF FINDINGS (1 total) diff --git a/.context/codex-reviews/gate-b-spec-pr15-pass-8.md b/.context/codex-reviews/gate-b-spec-pr15-pass-8.md new file mode 100644 index 0000000..8d781f4 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-pr15-pass-8.md @@ -0,0 +1,2 @@ +NO FINDINGS +END OF FINDINGS (0 total) diff --git a/.context/codex-reviews/gate-b-spec-pr15-pass-9.md b/.context/codex-reviews/gate-b-spec-pr15-pass-9.md new file mode 100644 index 0000000..8d781f4 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-pr15-pass-9.md @@ -0,0 +1,2 @@ +NO FINDINGS +END OF FINDINGS (0 total) diff --git a/.context/codex-reviews/gate-b-spec-pr15b-pass-1.md b/.context/codex-reviews/gate-b-spec-pr15b-pass-1.md new file mode 100644 index 0000000..8d781f4 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-pr15b-pass-1.md @@ -0,0 +1,2 @@ +NO FINDINGS +END OF FINDINGS (0 total) diff --git a/.context/codex-reviews/gate-b-spec-rle-pass-1.md b/.context/codex-reviews/gate-b-spec-rle-pass-1.md new file mode 100644 index 0000000..525d98a --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-rle-pass-1.md @@ -0,0 +1,15 @@ +BLOCKER | high | plugins/dev-workflow/commands/process-pr-review.md:148 | the PR-review command still derives cited stories from PR and commit bodies although the new governing-set rule says only contributing artifacts' Story headers are citations | an unrelated body mention can grant or deny skip eligibility while the provenance line records a different header-derived set, including a gate-off path where the actual contributing story is missed | derive the set from the contributing plans' Story headers, compare it with the governing spec header as section 5 requires, and use that same set for skip and record duties +BLOCKER | high | CLAUDE.md:842-851 and CLAUDE.md:881-895 | the pinned grammars introduce diagnostic outputs without the checks and fixes prompt-standard item 10 requires: the numeric range and test for an out-of-range knob are undefined, and a rejected model identifier becomes undetermined while the promised source and rejected-byte description has no pinned field or form | agents cannot classify these states consistently, and the supposedly one-form records can omit required diagnostics or emit mutually incompatible free-form additions, breaking invariant 11 in both shipped prompt copies | define the knob's accepted numeric range and distinguishing check with a fix per cause, and add grammar fields or a separate pinned diagnostic record for each undetermined-model cause and its safe byte description in both copies +MAJOR | high | CLAUDE.md:653-660 and plugins/dev-workflow/commands/workflow-init.md:837-844 | the three-case profile procedure has no case for a cited story path that exists but cannot be read, because the agent can establish neither that a profile is absent nor that one is present but unresolvable | this state derives no floor and fires no stated stop, so an unreadable high-risk story can silently fall into the unprofiled floor-3 path | add an explicit unreadable-cited-story stop with its distinguishing check and fix in both copies +MAJOR | high | CLAUDE.md:82-85 and CLAUDE.md:363-371 | the text first says there are three cycles and later correctly says there are three cycle kinds with one cycle per run, including a separate Gate-A plan cycle for every contributing plan | a multi-plan change can follow the three-cycle wording and omit nonces or commit-body record pairs for additional plan cycles; revision 36 repeats the stale cardinality while the changelog says five cycles produce five records | consistently say three kinds of cycle, retain one record pair per cycle run, and correct the corresponding revision-36 three-cycle and three-line claims +MAJOR | high | plugins/dev-workflow/CHANGELOG.md:53-61 | the 0.11.0 entry says a no-nonce cycle with existing bare slots takes a deterministic discriminator, but the author explicitly deferred that rule and neither shipped prompt copy admits such a slot | the release notes claim functionality the release does not ship and direct readers toward a filename that the authoritative grammar rejects | remove the discriminator exception and its this-cycle claim from 0.11.0, and describe only the shipped nonce infix plus bare pre-rule form until the successor story lands +MAJOR | high | CLAUDE.md:821-825 and docs/superpowers/specs/2026-08-28-review-loop-economics-design.md:93-96,121-133 | the prompt and two nearby spec properties say the deferred metrics work or P8 parses the records in the present tense even though the corrected spec header and changelog say that consumer does not exist yet and this repo ships no parser | the same change both denies and asserts an enforcement mechanism, so readers can incorrectly treat the grammar as machine-consumed today | change all remaining present-tense parse claims to intended future consumption and state that today's evidence is only the named grep plus comparisons +MAJOR | high | CLAUDE.md:917-920 and plugins/dev-workflow/commands/workflow-init.md:1101-1104 | the curve rule requires the commit body to say when logical-pass counts differ from hook-call counts, but neither pinned record grammar contains a hook-call count or a field that can express that difference | the required economics datum has no specified output form, so producers cannot comply consistently and a future consumer cannot recover the distinction the rule promises | add a pinned hook-call or branch-count field and its cardinality semantics to the curve, or remove the unsupported body-recording claim and explicitly leave the difference unrecorded +MINOR | high | CLAUDE.md:380-395 and plugins/dev-workflow/commands/workflow-init.md:574-589 | the nonce rule says it appears in every record the cycle writes, then excludes evidence entries and human-exception records that the same cycle can write | the universal wording contradicts its own named boundary and can cause agents to invent nonce variants for unrelated records | say every cycle-keyed record introduced by this contract carries the nonce, then retain the explicit four-record list and exclusions +MINOR | high | CLAUDE.md:868-897 and plugins/dev-workflow/CHANGELOG.md:42-48 | the prompt and changelog say every cycle records a per-pass curve and one of each record, but the pinned special case says a skipped cycle writes a skip record and no counts | the two instructions disagree about the required closing body for a legitimate skipped cycle | say every cycle writes a provenance line plus either a curve or a skip record, and reserve a curve for non-skipped cycles +MINOR | high | docs/superpowers/specs/2026-08-28-review-loop-economics-design.md:468-480 | the revised evidence rule says invalid strings must fail the grep and that every constraint grep cannot decide is checked, then says grep decides lexical productions only and the four comparison constraints are not proven complete | semantic-invalid but lexically valid strings cannot fail the grep, and an explicitly incomplete comparison list cannot support the every-constraint claim | require lexically invalid fixtures to fail grep, require semantic-invalid fixtures to fail named field comparisons, and replace every constraint with every identified constraint +MINOR | high | docs/getting-started.md:58 | the new counter explanation presents three discarded result shapes as the reason counted calls differ from calls made and says the hook keeps the fingerprint streak without qualifying best-effort state, but unmapped or unattributable calls are also uncounted and failed counter or state writes can lose or restart the tally | users can treat advisory numbers as exhaustive and durable even though the hook deliberately exits zero through routing and persistence failures | identify the three shapes as recognized routed-result exclusions, mention unattributable calls, and describe the tally and streak as best-effort persisted state +MINOR | medium | CLAUDE.md:85-87 and plugins/dev-workflow/commands/workflow-init.md:292-294 | the floor rationale says exactly one value exists because there is one source, but the procedure immediately requires reading and comparing the spec header and every contributing plan header | uniqueness follows only after multiple governing headers agree, so the singular-source explanation conflicts with the actual stop-on-disagreement mechanism | say one value exists only when all governing headers resolve to the same cited set and that disagreement stops +MINOR | high | CLAUDE.md:119-124 and plugins/dev-workflow/CHANGELOG.md:32-34 | the replacement says the hook ratio controls nothing and that the hook counts passes, although the ratio selects which advisory message fires and the hook counts qualifying routed calls rather than validated logical passes | this overstates the mechanism and contradicts the later rule that explicitly distinguishes logical passes from hook-counted calls | say the ratio controls no gate obligation or closure decision and that the hook counts qualifying routed gate calls +MINOR | high | docs/superpowers/specs/2026-08-28-review-loop-economics-design.md:3 | the change alters revision 36 after its header says Gate-A passes 1-33 reviewed that revision, without advancing the revision or marking the post-review correction | later readers cannot tell the approved revision-36 bytes from the finding-driven grammar-check edit made during rollout | bump the revision and update references, or annotate the header with the exact post-Gate-A correction and its review provenance +END OF FINDINGS (14 total) diff --git a/.context/codex-reviews/gate-b-spec-rle-pass-2.md b/.context/codex-reviews/gate-b-spec-rle-pass-2.md new file mode 100644 index 0000000..e6a7043 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-rle-pass-2.md @@ -0,0 +1,17 @@ +BLOCKER | high | CLAUDE.md:654-675; plugins/dev-workflow/commands/workflow-init.md:840-861 | The four profile-reading cases are not exhaustive: a cited path that does not exist satisfies neither case 3, which requires an existing path, nor case 4, which requires a present Profile line; case 3 also names several causes without giving their individual fixes | A missing cited story can derive neither a floor nor the required stop, creating an undefined gate path and violating prompt-standard diagnostic completeness | Make case 3 cover every failure to resolve to a readable regular file, explicitly including absent paths, broken symlinks, directories, and permissions, and pin a decidable check and fix for each; reserve case 4 for a readable story whose present Profile line is invalid +BLOCKER | high | CLAUDE.md:849-863; plugins/dev-workflow/commands/workflow-init.md:1033-1047; plugins/dev-workflow/hooks/codex-gate.sh:118-128 | The floor-knob diagnostic defines out-of-range by a hook-defined accepted maximum, but the hook defines no maximum and README.md:130 and docs/getting-started.md:85 instead promise that any positive integer is accepted | No reviewer can decide this cause or apply its fix, and sufficiently large values depend on unspecified shell integer behavior rather than a checkable contract | Define an explicit numeric range shared by the prompt, user docs, and hook, or remove the out-of-range cause and specify the actual platform-bound failure behavior with a deterministic check and fix +BLOCKER | high | CLAUDE.md:849-863; plugins/dev-workflow/commands/workflow-init.md:1033-1047; plugins/dev-workflow/hooks/codex-gate.sh:118-128 | The knob procedure says to trim trailing newlines before testing digits, while the hook deletes all whitespace; for example embedded whitespace can be reported non-numeric while the hook concatenates the digits and uses the resulting floor, and whitespace-only content is not classified by the stated empty check | The pinned provenance can report a different effective knob and floor from the mechanism it is supposed to explain, falsifying the record and the floor decision | Specify and use one normalization algorithm byte-for-byte in both diagnostic copies and the hook, then give fixtures for embedded, leading, trailing, and whitespace-only content +BLOCKER | high | CLAUDE.md:896-910; plugins/dev-workflow/commands/workflow-init.md:1080-1094 | The undetermined model diagnostic still collapses every failure into one token and merely asks for free-form source and rejected bytes; it does not enumerate distinct causes or pin a check and fix for missing configuration, failed resolution, or an unrepresentable identifier | The evidence cannot support a decidable correction or prove which routing failure occurred, so the new machine-oriented record fails prompt-standard item 10 | Enumerate mutually exclusive model-resolution causes with a concrete observation and fix for each, and pin the safe serialized diagnostic form rather than leaving it as unconstrained prose +BLOCKER | high | plugins/dev-workflow/commands/process-pr-review.md:148-160; CLAUDE.md:96-118 | process-pr-review still derives its story set from PR and commit bodies and then equates absence in that record with the section 5 no-story case, even though section 5 makes contributing artifact Story headers the sole governing set and says body paths contribute nothing to it | An unrelated or missing body reference can grant or deny the Gate-B triviality skip independently of the artifacts that govern the cycle, leaving a gate-off decision on a conflicting citation source | Derive skip eligibility and the recorded floor from the contributing plan Story headers, or pin a separate non-relaxing cross-reference rule and precedence that cannot override or substitute for the governing set +BLOCKER | high | Current Gate-B evidence entry; docs/superpowers/plans/2026-08-30-review-loop-economics-plan-c-rollout.md:1316-1324 | The stated high-risk verification is explicitly deferred until task 24 after this Gate-B pass, and its planned check recomputes from the singular Story header of the artifact reviewed rather than the union of contributing plan headers required for Gate B | The old workflow requires battery, check, and named risk verification evidence before every Gate-B call; this pass therefore lacks a mandatory risk observation and the deferred procedure can verify the wrong governing profile | Construct the proposed provenance line now, recompute its floor from the complete Gate-B governing header set, state the observable failure condition and result in the evidence entry, and only then restart the pass +BLOCKER | high | docs/superpowers/plans/2026-08-29-review-loop-economics-plan-a-rules.md:80-103; docs/superpowers/plans/2026-08-30-review-loop-economics-plan-b-records.md:97-112,490-498 | The old-conditions accounting was not updated for either pass-1 decision rewrite: the profile partition changed from three cases to four, and the headSha rule changed from taking a value reported by each call to resolving and passing a value before the call; Plan B still states the now-known-false returned-head procedure | There is no auditable kept, moved, or dropped disposition for predecessor conditions, and an authoritative implementation plan still tells an executor to assume a return value the tool does not provide | Add accounting rows for both rewrites, preserve every predecessor condition explicitly, and mark the returned-head text superseded by the resolve-before-call and preserve-passed-value procedure +MAJOR | high | docs/superpowers/plans/2026-08-30-review-loop-economics-plan-c-rollout.md:125-126,1185-1257,1282-1286 | Tasks 19 and 20 are intentionally dropped, but the plan still directs executors to ship the deterministic slot discriminator, claims the accounting row is added, and says task 23 relies on those tasks to admit the rle name | The authoritative rollout plan contradicts both the deliberate boundary and the shipped prompt and changelog, so a later executor can accidentally ship deferred behavior or reject the intended plan-local exception | Mark tasks 19 and 20 dropped, replace their claimed shipped effects with the explicit human-approved plan-local naming exception, and update the accounting, task 23 dependency, and self-review claims consistently +MINOR | high | plugins/dev-workflow/commands/process-pr-review.md:160 | The revised fourth profile case is still called section 5's third case | The command points reviewers at the wrong branch of the decision procedure and obscures the distinct unreadable-file and invalid-profile diagnostics | Change the ordinal to fourth case and name the condition instead of relying only on an ordinal +MINOR | high | CLAUDE.md:381-396; plugins/dev-workflow/commands/workflow-init.md:567-582 | The nonce rule first says it appears in every record the cycle writes, then excludes evidence entries and human-approved exception records | The universal statement makes compliant exception records simultaneously noncompliant and leaves recovery audits ambiguous | Scope the universal to every cycle-keyed working, provenance, curve, and skip record named by the rule, then retain the explicit exclusions +MINOR | high | CLAUDE.md:883-915; plugins/dev-workflow/commands/workflow-init.md:1067-1099; plugins/dev-workflow/CHANGELOG.md:42-48 | The prompt says every cycle records a per-pass curve and the changelog says one provenance and one curve line per cycle, but a skipped cycle writes a skip record and has no pass counts | User-facing release behavior promises a curve that the specified skip path cannot produce | State that each cycle writes one provenance line and then either one curve line for a reviewed cycle or one skip line for a skipped cycle +MINOR | high | docs/superpowers/specs/2026-08-28-review-loop-economics-design.md:469-481 | The grammar verification says invalid constructed strings must not match grep and that every constraint grep cannot decide is checked, but semantic-invalid strings are lexically expected to match and the following check list expressly says it is not proven complete | The acceptance test is internally unsatisfiable or overclaims its coverage, so passing evidence cannot establish the stated requirement | Require lexically invalid fixtures to fail grep, require semantic-invalid fixtures to fail the named comparison checks, and replace every constraint with every identified semantic constraint unless the list is made exhaustive +MINOR | medium | docs/getting-started.md:58 | The counter explanation attributes the difference from calls made to three withheld result shapes and says the hook keeps the streak, omitting unattributable or unmapped calls and the best-effort persistence failure paths | A user can infer that the displayed number exhaustively and durably counts every other review call, which the hook does not guarantee | Present the three shapes as examples among exclusions and describe the streak as best-effort persisted qualifying routed calls +MINOR | medium | CLAUDE.md:86-118; plugins/dev-workflow/commands/workflow-init.md:272-304 | The text says a cycle has exactly one floor because there is one source, then permits several contributing artifact headers and only establishes uniqueness after comparing them for equality | The rationale overstates the input model and can cause an implementation to select one header instead of performing the required set comparison | Say there is one floor only after every governing header has been read and all derived floors agree +MINOR | high | CLAUDE.md:120-122; plugins/dev-workflow/commands/workflow-init.md:306-308; plugins/dev-workflow/CHANGELOG.md:30-33 | The docs say the combined-review ratio controls nothing and the hook counts passes, although the ratio controls which advisory message fires and the hook counts qualifying routed tool calls rather than validated logical passes | These absolutes misdescribe the unchanged mechanism while claiming to correct user-facing semantics | Limit the claim to controlling no gate obligation or closure decision and call the observed values qualifying routed calls rather than passes +NIT | high | docs/superpowers/specs/2026-08-28-review-loop-economics-design.md:3 | The fixed range changes the approved spec text but leaves its revision header at Revision 36 | The revision marker no longer uniquely identifies the bytes reviewers are approving | Increment the revision and describe the post-Gate-A correction in the revision note +END OF FINDINGS (16 total) diff --git a/.context/codex-reviews/gate-b-spec-rle-pass-3.md b/.context/codex-reviews/gate-b-spec-rle-pass-3.md new file mode 100644 index 0000000..4d8caf7 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-rle-pass-3.md @@ -0,0 +1,17 @@ +BLOCKER | high | CLAUDE.md:109-116; plugins/dev-workflow/commands/workflow-init.md:316-323; docs/superpowers/plans/2026-08-29-review-loop-economics-plan-a-rules.md:14 | The new pinned `Story:` syntax accepts only bare or double-quoted paths, but this change's spec and all three contributing plans use the repository's established Markdown-code-span form with backticks | Those governing headers are malformed under the rule being shipped, so a post-rule cycle stops before deriving any floor; the claimed union and proposed provenance line were not derived from headers that satisfy the new syntax | Either admit and normalize the established backtick form, including any supported trailing annotation, or migrate every producer and existing artifact before making the new syntax binding +BLOCKER | high | plugins/dev-workflow/commands/process-pr-review.md:148-165 | The skip procedure still defines its cited set from PR and commit-body records and never actually requires the section 5 governing set; after saying record absence cannot substitute for that set, lines 160-161 immediately equate the same absence with section 5's no-story case | A missing or incomplete PR record can still grant the unprofiled Gate-B skip without checking the contributing artifacts, preserving the gate-off path the correction was meant to close | Derive and require eligibility from the governing `Story:` headers first, use PR and commit-body citations only as an additional non-relaxing record check, and delete the contradictory no-story equivalence +BLOCKER | high | docs/superpowers/stories/2026-08-28-review-loop-economics-pass-floor-story.md:158-178; docs/superpowers/specs/2026-08-28-review-loop-economics-design.md:512-531; docs/superpowers/plans/2026-08-30-review-loop-economics-plan-c-rollout.md:1345-1364 | The story and revision-36 spec require this branch's Gate-A spec, Gate-A plan, and Gate-B closing bodies to carry native final-form records, while Plan C explicitly excludes every Gate-A record and plans to write only the Gate-B pair | An explicit acceptance criterion remains unsatisfied, and the spec's claimed three-body demonstration cannot be true after the Gate-A commits have already closed without those forms | Reconcile the acceptance source before closing: either supply the required records in their specified commits, or revise the story and spec to approve the field-report destination and bound the branch demonstration to the Gate-B cycle +BLOCKER | high | CLAUDE.md:661-669; plugins/dev-workflow/commands/workflow-init.md:847-855 | Profile case 3 says missing path, non-regular path, broken symlink, and unreadable regular file are distinguished by testing existence, type, and link resolution in that order, but that order cannot reach the broken-symlink cause: target-following existence reports false, while lstat-style existence followed by type reports the symlink as non-regular | The diagnostic is not decidable as promised and violates prompt-standard item 10; a broken link receives the wrong cause and fix even though this round specifically introduced a separate cause for it | Test for a symlink before target-following existence and regular-file checks, define the exact precedence for every state, and distinguish read failure after the single read attempt from the pre-read path tests +BLOCKER | high | CLAUDE.md:921-934; plugins/dev-workflow/commands/workflow-init.md:1105-1118; CLAUDE.md:478-486 | The three `undetermined` causes are not mutually exclusive: a call that errors before model resolution also carries no model identifier, satisfying both NONE-REPORTED and RESOLUTION-FAILED; moreover, an errored call is an incomplete pass and therefore cannot legitimately appear in the valid-pass curve at all | The same observation can produce two cause tokens, and the new curve can be instructed to record a pass that the surrounding protocol requires excluding, breaking both diagnostic completeness and curve integrity | Define disjoint tests with explicit precedence, and reserve `undetermined` for a validated pass whose running model cannot be recovered; omit errored or otherwise incomplete calls from the curve as the existing pass-validity rule requires +MAJOR | high | docs/superpowers/stories/2026-08-28-review-loop-economics-pass-floor-story.md:170-173; docs/superpowers/specs/2026-08-28-review-loop-economics-design.md:309-317; docs/superpowers/plans/2026-08-30-review-loop-economics-plan-b-records.md:487-498 | The authoritative story, spec, and Plan B still require or claim that both branches reviewed the same revision and that each call reports a reviewed head, while the shipped prompts now correctly establish only that equal request values aimed both calls at one commit | The source artifacts still demand an impossible observation and support the stronger reviewed-revision claim that this round says it removed, so later maintenance can restore the overclaim or reject compliant records | Rewrite every source occurrence to the retained-request-input rule, explicitly state that it proves neither branch reviewed the commit, and retain the mismatch terminal behavior against the passed values +MAJOR | high | docs/superpowers/plans/2026-08-30-review-loop-economics-plan-c-rollout.md:970-980,1047-1076,1292-1296,1422-1438 | Only Task 19 is marked dropped; Task 20 remains an executable replacement task, Task 23 still says both tasks admit the discriminator into the shipped rule, and the self-review still claims the rule gains that production | A later executor can still ship the deliberately deferred discriminator or conclude the current plan-local `rle` exception violates the prompt, despite the author decision to drop both tasks | Mark Task 20 dropped, remove or clearly fence both replacement bodies, rewrite Task 23 around the approved plan-local exception, update the self-review, and remove the orphan `§5` fragment at line 980 +MINOR | high | plugins/dev-workflow/commands/process-pr-review.md:164-166 | The command still calls a present-but-unresolvable profile section 5's third case, but the new four-case procedure makes it case 4 | The cross-reference directs an executor to the unreadable-file branch instead of the invalid-profile branch and obscures the distinct diagnostics | Change the ordinal to case 4 and name the condition directly so future renumbering cannot stale it again +MINOR | high | CLAUDE.md:384-399; plugins/dev-workflow/commands/workflow-init.md:578-593 | The nonce rule says the nonce appears in every record the cycle writes, then excludes evidence entries and human-exception records that the same cycle writes | A compliant evidence or exception record is simultaneously required to carry and forbidden from carrying the nonce, leaving the named record set internally contradictory | Scope the universal to the four cycle-keyed record families this change governs, then keep the explicit exclusions for other records +MINOR | high | CLAUDE.md:901-944; plugins/dev-workflow/commands/workflow-init.md:1085-1128; plugins/dev-workflow/CHANGELOG.md:42-48 | The prompt says every cycle records a per-pass curve and the changelog promises one curve per cycle, but the specified skipped-cycle production writes a skip record with no counts instead | The release description and main rule promise a record a legitimate skip cannot produce, making a compliant skipped cycle look incomplete | State consistently that every cycle writes one provenance line plus either a curve for a reviewed cycle or a skip record for a skipped cycle +MINOR | high | README.md:128-130; docs/getting-started.md:87-88; plugins/dev-workflow/hooks/codex-gate.sh:118-128 | The user docs still describe the floor knob as accepting a positive integer, including the categorical phrase `any positive integer`, while the hook accepts digits only when the invoking shell's integer comparison succeeds and falls back to 3 for larger positive values it cannot compare | Users can write a mathematically valid positive value that the hook silently ignores, contradicting the corrected out-of-range explanation | Describe the actual contract in both docs: whitespace is deleted, digits must remain, and the shell's positive-integer comparison must succeed; avoid `any` because no fixed maximum exists +MINOR | high | docs/superpowers/plans/2026-08-29-review-loop-economics-plan-a-rules.md:105-111; docs/superpowers/plans/2026-08-30-review-loop-economics-plan-b-records.md:489-498 | New accounting row 13 inventories only five predecessor conditions and omits the mismatch consequences and conservative rationale: mismatched branches are not one pass, the completed branch is incomplete and excluded, the later branch starts a new pass, and merging revisions is the unsafe direction | Although the shipped prompt retains those clauses, the required old-conditions accounting no longer guards them against a later rewrite, violating the repository's decision-procedure accounting rule | Add the omitted predecessor conditions to row 13 and mark each kept against the request-value comparison +MINOR | medium | CLAUDE.md:657-682; plugins/dev-workflow/commands/workflow-init.md:843-868 | The heading claims an exhaustive four-case profile partition, but a cited path yielding a readable story with a valid resolvable profile satisfies none of the four numbered cases | The ordinary profiled path is handled elsewhere, but a reader treating this block as the complete decision procedure has no numbered outcome for the normal state | Add the valid-profile case or rename the block to make clear that it lists only exceptional and unprofiled states +MINOR | high | CLAUDE.md:123-125; plugins/dev-workflow/commands/workflow-init.md:330-332; plugins/dev-workflow/CHANGELOG.md:32-33; docs/getting-started.md:60 | The correction still says the hook ratio controls nothing and that the hook counts passes; mechanically the threshold selects which advisory message fires, while the counters track qualifying routed tool calls subject to mappings, recognized envelopes, fail-open classification, and best-effort state persistence | These absolutes preserve the same mechanism overclaim the change is meant to remove and can make a user treat the displayed ratio as validated pass accounting | Say the threshold controls no gate obligation or closure decision but does control reminder selection, and call the numerator qualifying routed calls with the relevant limits rather than passes +NIT | high | docs/superpowers/specs/2026-08-28-review-loop-economics-design.md:3 | The fixed range changes the approved spec text but leaves its identity at Revision 36 and Gate-A passes 1-33 | The revision marker no longer uniquely identifies the bytes being reviewed and the task's cited revision can refer to two different texts | Increment the revision and record that the post-Gate-A corrections changed the approved text +NIT | high | docs/getting-started.md:53-59 | The new bounded explanation still says any included content change flips the hook back to unsatisfied, but the hook compares a checksum that may fall back to 32-bit `cksum`; equality establishes only equal fingerprints, not that no changed input collided | This reintroduces a categorical content guarantee in the paragraph added to remove one, especially on the weakest supported checksum fallback | Describe invalidation as a fingerprint comparison and say a differing fingerprint is unsatisfied; do not claim every content change must produce a different digest +END OF FINDINGS (16 total) diff --git a/.context/codex-reviews/gate-b-spec-rle-pass-4.md b/.context/codex-reviews/gate-b-spec-rle-pass-4.md new file mode 100644 index 0000000..a525889 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-rle-pass-4.md @@ -0,0 +1,17 @@ +BLOCKER | high | docs/superpowers/stories/2026-08-28-review-loop-economics-pass-floor-story.md:158-178,281-293; docs/superpowers/specs/2026-08-28-review-loop-economics-design.md:514-539; docs/superpowers/plans/2026-08-30-review-loop-economics-plan-c-rollout.md:1350-1372 | The story and revision-36 spec require this branch's Gate-A spec, Gate-A plan, and Gate-B closing bodies to carry native final-form records, but Plan C explicitly excludes all four Gate-A cycles' records and closes with only the Gate-B pair | The branch cannot satisfy the stated provenance and curve acceptance criteria, and the promised three-cycle demonstration is replaced by a field report the authoritative requirements do not admit | Reconcile the acceptance sources before closing: either put the required records in the specified cycle commits, or revise the story and spec to approve the field-report destination and bound the branch demonstration to Gate B +BLOCKER | high | CLAUDE.md:843-875; plugins/dev-workflow/commands/workflow-init.md:1027-1059; plugins/dev-workflow/hooks/codex-gate.sh:119-127; docs/prompt-standards.md:66-74 | Deleting the knob-classification walkthrough leaves the four required `CAUSE` tokens and the duty to name one, but supplies no mutually exclusive checks or cause-specific fixes; the hook only collapses rejected inputs to its default and emits no cause, and `out-of-range` has no defined bound | Writers cannot deterministically produce the pinned provenance field for broken links, read failures, whitespace-normalized input, or oversized digit strings, so the supposedly machine-extractable record can vary while invariant 11's diagnostic-state rule is violated | Define a complete observer-side classification with precedence, exact checks, and per-cause fixes that matches the hook, or replace the four-way field with the single distinction the mechanism actually exposes +BLOCKER | high | CLAUDE.md:905-924; plugins/dev-workflow/commands/workflow-init.md:1089-1108; docs/prompt-standards.md:66-74 | The model-cause deletion leaves `undetermined` covering both an identifier that cannot be determined and one that cannot be represented, while making the explanation optional free prose and providing no checks or cause-specific fixes; the surviving rejected-bytes sentence is inapplicable when no identifier was obtained | The scaffolded prompt reports a diagnostic state that its reader cannot resolve, violating invariant 11, and different writers can make materially different recovery decisions while emitting the same curve token | Keep `undetermined` in the pinned curve if desired, but define outside the grammar a mutually exclusive diagnostic procedure with a check and fix for every cause and make the companion explanation mandatory when the token is used +BLOCKER | high | CLAUDE.md:657-682; plugins/dev-workflow/commands/workflow-init.md:843-868; docs/superpowers/plans/2026-08-29-review-loop-economics-plan-a-rules.md:110; docs/prompt-standards.md:66-74 | Profile case 3 promises to distinguish missing paths, non-regular paths, broken symlinks, and unreadable regular files by testing existence, type, then link resolution, but target-following existence or type tests consume a broken symlink before the link check; the four causes also are not paired with their own fixes | The profile reader cannot execute the promised diagnostic procedure and may report the wrong stop cause, breaking invariant 11 on a state that controls the gate floor and evidence duties | Use an explicit symlink-first or lstat-based partition, classify the single read attempt's failures, and pair every resulting cause with its own check and fix in both copies and the accounting row +MAJOR | high | docs/superpowers/plans/2026-08-30-review-loop-economics-plan-c-rollout.md:1336-1339; CLAUDE.md:108-121; plugins/dev-workflow/commands/workflow-init.md:315-328 | Task 24 still tells the closer to recompute Gate B's provenance from the singular Story header of the artifact reviewed, although Gate B reviews a diff and the shipped rule derives from all contributing plan headers after comparing every expected artifact | The required risk-path verification can bless a floor derived from one convenient header while missing an absent, malformed, or disagreeing contributing header | Re-read the spec and all three contributing plan headers at closure, compare the expected sets, and recompute the provenance floor from the complete Gate-B union +MAJOR | high | CLAUDE.md:688-694,890-931; plugins/dev-workflow/commands/workflow-init.md:874-880,1074-1115; plugins/dev-workflow/commands/process-pr-review.md:169-174 | The pinned skip record says only `skipped (see skip reason)`, while the skip reason remains free-form and has no cycle field, key, adjacency, or cardinality rule linking it to that record | When squash carry gathers several skipped cycles into one body, no reader or future parser can decide which reason a given nonce-bearing skip record points to, so the mandated record is not actually attributable | Put the escaped reason in the skip production or define an exactly-one cycle-field-keyed reason record and make the squash carry preserve that association +MAJOR | high | docs/superpowers/stories/2026-08-28-review-loop-economics-pass-floor-story.md:170-173; docs/superpowers/specs/2026-08-28-review-loop-economics-design.md:309-319; CLAUDE.md:951-960; plugins/dev-workflow/commands/workflow-init.md:1135-1152 | The authoritative story still requires separate calls to have reviewed the same tracked revision, while the implementation and corrected spec establish only that equal request values aimed the calls at one commit and explicitly cannot prove what either reviewer read | The implementation does not meet the cited acceptance criterion, and future maintenance can reject a compliant record or restore the impossible reviewed-revision claim from the unchanged story | Amend the story criterion to require both calls to be issued against the same full commit value, retain the mismatch terminal behavior, and state there that request equality is not reviewed-content evidence +MAJOR | high | docs/superpowers/plans/2026-08-30-review-loop-economics-plan-c-rollout.md:970-980,1047-1052,1297-1301,1350-1365,1397-1421,1427-1441 | The drop note says `rle` is a plan-local naming exception that must be recorded in the closing body and field report, but Task 23 and Self-Review still claim the dropped tasks add it to the shipped grammar, Task 24's complete body omits the exception, and Task 25 schedules only the curve update | The author-approved replacement for dropped Tasks 19 and 20 is not actually implemented or durably recorded, leaving the current cycle's non-grammar slot names justified only by contradictory plan prose | Rewrite Task 23 and Self-Review around the plan-local exception, add the promised exception record to the closing body and field report, and remove the stale shipped-production claims +MINOR | high | CLAUDE.md:384-399; plugins/dev-workflow/commands/workflow-init.md:578-593 | The nonce rule says it appears in every record the cycle writes and then says it is not required in evidence entries and human-exception records that the same cycle writes | A cycle following the literal universal and the explicit exclusions receives opposite instructions, undermining the named-record-set boundary | Scope the first sentence to the four cycle-keyed record families this change governs, then retain the exclusions for other records +MINOR | high | CLAUDE.md:892-931; plugins/dev-workflow/commands/workflow-init.md:1076-1115; plugins/dev-workflow/CHANGELOG.md:42-48 | The prompts say every cycle records its own per-pass curve and the changelog promises one curve per cycle, but a legitimately skipped cycle writes a skip record with no curve or counts | A compliant skipped cycle is described as missing a universally required record, and the release note overstates what the pinned contract produces | State consistently that every cycle writes a provenance line plus either a per-pass curve for a reviewed cycle or the pinned skip record for a skipped cycle +MINOR | high | README.md:128-131; docs/getting-started.md:87-89; plugins/dev-workflow/hooks/codex-gate.sh:119-127 | The user docs describe the knob as accepting a positive integer, including `any positive integer`, but the hook accepts a digit string only when the invoking shell's integer comparison succeeds; sufficiently large positive strings fail that comparison and leave the default standing | Users can supply a mathematically valid value the corrected docs promise is accepted while the hook silently ignores it | Describe the actual normalization and shell-comparison contract, or define and enforce a portable numeric maximum shared by the hook, grammar, and user docs +MINOR | medium | CLAUDE.md:657-682; plugins/dev-workflow/commands/workflow-init.md:843-868 | The heading claims four profile cases and four answers, but a cited path yielding a readable story with a valid resolvable profile satisfies none of the numbered cases | The ordinary profiled path is handled elsewhere, but this block presents itself as the complete partition and gives no numbered outcome for the normal state | Add the valid-profile outcome or rename the block to say it enumerates only unprofiled and failure states +MINOR | high | CLAUDE.md:123-132; plugins/dev-workflow/commands/workflow-init.md:330-339; plugins/dev-workflow/CHANGELOG.md:32-35; plugins/dev-workflow/hooks/codex-gate.sh:119-127,947-973 | The corrected text still says the reminder threshold controls nothing and that the hook counts passes, although the threshold selects which advisory message fires and the numerator counts qualifying routed calls subject to tool mappings, recognized result envelopes, and best-effort state | These absolutes contradict the same prompt's later limitations and can make a reader treat the displayed ratio as validated pass accounting | Say the threshold controls no gate obligation or closure decision but does control reminder selection, and call the numerator qualifying routed calls rather than passes +MINOR | high | docs/superpowers/plans/2026-08-29-review-loop-economics-plan-a-rules.md:105-111; docs/superpowers/plans/2026-08-30-review-loop-economics-plan-b-records.md:487-500 | The added old-conditions row for branch agreement inventories only five predecessor conditions and omits the mismatch consequences and rationale that the implementation retains: the branches are not one pass, the completed branch is incomplete, the later branch starts a new pass, and merging revisions is the unsafe direction | The required accounting no longer guards those live conditions against a later rewrite, contrary to the story's complete-accounting criterion and AGENTS.md's decision-procedure rule | Add every omitted predecessor condition to row 13 and mark each kept against the passed-value comparison +NIT | high | docs/superpowers/specs/2026-08-28-review-loop-economics-design.md:3,11-16,92-96,310-319,470-485 | The implementation changes the approved spec's mechanism and evidence wording while leaving its identity at Revision 36 and Gate-A passes 1-33 | Revision 36 no longer names one stable byte set, so the task's cited revision can refer to both the pre-fix and post-fix documents | Increment the revision and record that the post-Gate-A corrections changed the approved text +NIT | high | docs/getting-started.md:53-60; plugins/dev-workflow/hooks/codex-gate.sh:220-221,287-298 | The rewritten user explanation still says any included content change flips Gate B back to unsatisfied, but the hook may fall back to 32-bit `cksum`, so equality establishes only equal fingerprints and a content change can collide | The paragraph added to bound the mechanism still states a categorical content guarantee that its weakest supported checksum cannot provide | Describe invalidation as fingerprint comparison: a differing fingerprint is unsatisfied, while an equal one is only an equal digest under the selected checksum +END OF FINDINGS (16 total) diff --git a/.context/codex-reviews/gate-b-spec-rle-pass-5.md b/.context/codex-reviews/gate-b-spec-rle-pass-5.md new file mode 100644 index 0000000..5ca1df1 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-rle-pass-5.md @@ -0,0 +1,14 @@ +BLOCKER | high | CLAUDE.md:657-675; plugins/dev-workflow/commands/workflow-init.md:843-861 | The advertised five-case partition is not mutually exclusive and does not define one outcome for mixed cited sets: case 0 also vacuously matches an empty set, while a set containing an unprofiled story plus an unreadable or unresolvable story matches both a run case and a stop case. | The same input can license proceeding unprofiled or require stopping, so a multi-story change can be under-reviewed. | Make case 0 explicitly require a non-empty cited set and define aggregate precedence: any unreadable or unresolvable member stops; otherwise any unprofiled member takes the unprofiled path; otherwise derive the profiled floor. +BLOCKER | high | CLAUDE.md:892,929-932; plugins/dev-workflow/commands/workflow-init.md:1076,1113-1116; docs/superpowers/specs/2026-08-28-review-loop-economics-design.md:295-298 | The shared `` does not unambiguously link a skip record to its reason: every pre-rule cycle uses `cycle none (pre-rule)`, multiple skipped cycles can be carried into one squash body, and several cycles of the same kind can therefore share both the field and record prefix. | A carried skip record can resolve to the wrong reason or to several reasons, contradicting the claimed one-record/one-reason pairing. | Put a unique reason key in both records or embed the reason in the skip record; otherwise explicitly constrain bodies so only one matching pre-rule cycle can occur and define collision handling. +BLOCKER | high | docs/prompt-standards.md:66-73; CLAUDE.md:862,872-877,907-922; plugins/dev-workflow/commands/workflow-init.md:1046,1056-1061,1091-1106 | The shipped `unusable` and `undetermined` diagnostic states intentionally collapse distinct causes and provide neither a discriminating check nor a per-cause fix. | This violates prompt-standard item 10 and therefore invariant 11; the story permits only item 1 to be n/a for the scaffolded prompt. | Restore cause-distinguishing diagnostics with checks and fixes, or revise the governing standard and story through an explicit approved exception before shipping the collapsed states. +BLOCKER | high | docs/prompt-standards.md:66-73; CLAUDE.md:664-670; plugins/dev-workflow/commands/workflow-init.md:850-856 | Profile case 3 says its four unreadable-path causes need different fixes but supplies no per-cause fixes and expressly declines to prescribe a discriminating check. | An agent cannot reliably distinguish and repair the reported state, so the changed prompt independently fails prompt-standard item 10 and invariant 11. | Specify a non-following-link observation procedure and pair each resulting cause with its own corrective action. +MAJOR | high | docs/superpowers/specs/2026-08-28-review-loop-economics-design.md:253-260; docs/superpowers/plans/2026-08-30-review-loop-economics-plan-b-records.md:460-467; CLAUDE.md:914-922; plugins/dev-workflow/commands/workflow-init.md:1098-1106 | The authoritative spec and Plan B still require an unrepresentable model identifier to record its source and rejected bytes, while both shipped prompts now say `undetermined` records no cause. | The supposedly verbatim pinned contract has two incompatible forms, and an implementer following the spec emits data the shipped rule says not to emit. | Apply the same approved withdrawal to the spec and Plan B pinned grammar, or restore the source-and-byte description duty in both prompt copies. +MAJOR | high | docs/superpowers/plans/2026-08-30-review-loop-economics-plan-b-records.md:381-392; docs/superpowers/specs/2026-08-28-review-loop-economics-design.md:476-478; docs/superpowers/plans/2026-08-30-review-loop-economics-plan-c-rollout.md:1125-1127,1220-1226 | Cause-specific knob obligations remain after `` was removed: Plan B still says to name the cause, and the spec and Plan C still require coverage or classification of each pinned unusable-knob cause. | The reduced grammar cannot supply the distinctions these rules and checks demand, leaving the exact stranded obligation the cut was intended to eliminate. | Remove or rewrite every cause-specific obligation so it tests only `unusable`, or restore a defined cause production consistently across all sources and prompts. +MAJOR | high | docs/superpowers/specs/2026-08-28-review-loop-economics-design.md:369-378; CLAUDE.md:401-420; plugins/dev-workflow/commands/workflow-init.md:595-614; docs/superpowers/plans/2026-08-30-review-loop-economics-plan-b-records.md:243-263 | The shipped nonce-recovery rule changes the spec's single-candidate-from-either-source rule by making kind, artifact, and open status insufficient and requiring a candidate to be positively linked to the run by a working record the run itself wrote, but no independent run identity or check defines that link. | Recovery after interruption is circular: the process must already know which record is its own before the record can recover its identity, so valid cycles can restart and split their records. | Keep the spec's defined single-candidate rule, or define a non-circular run-to-record credential and update the spec in the same commit. +BLOCKER | high | docs/superpowers/stories/2026-08-28-review-loop-economics-pass-floor-story.md:158-181; docs/superpowers/specs/2026-08-28-review-loop-economics-design.md:518-537; docs/superpowers/plans/2026-08-30-review-loop-economics-plan-b-records.md:56-62; docs/superpowers/plans/2026-08-30-review-loop-economics-plan-c-rollout.md:1353-1368 | The story requires every cycle that ran to carry its curve and says Gate B alone fails; Plan B identifies five cycles, but Plan C's complete closing-body procedure explicitly excludes records for all four Gate-A cycles. | The branch cannot satisfy its own acceptance criterion or the spec's promised verification, so the requested per-cycle durable record is missing for most contributing cycles. | Put each Gate-A record in its cycle's closing body as required, or formally revise the story and spec to authorize the field report as the durable substitute before executing the close. +MAJOR | high | docs/superpowers/plans/2026-08-30-review-loop-economics-plan-c-rollout.md:125-126,649-657,970-979,1047-1052,1430-1445 | Dropping Tasks 19 and 20 was not propagated through Plan C: current accounting and the Task 12 changelog still say the release ships the deterministic discriminator, the dropped-task notes incorrectly say corrected Task 23 text remains stale, and Self-Review still assigns the shipped slot rule to the dropped tasks. | The executable plan gives mutually contradictory answers about whether `rle` is a shipped production or a plan-local exception, defeating the required old-condition accounting and release audit. | Rewrite every live accounting, changelog, supersession note, and Self-Review statement to one consistent decision: Tasks 19 and 20 are dropped and `rle` is plan-local only. +BLOCKER | high | docs/superpowers/plans/2026-08-30-review-loop-economics-plan-c-rollout.md:976-979,1298-1302,1353-1361 | Task 23 requires the plan-local `rle` naming exception to be recorded in the closing commit body, but Task 24's expressly complete closing-body list omits that record and says nothing else is carried. | No executable task schedules the sole authority for the nonconforming slot names, so closure either violates Task 23 or violates Task 24's complete-list assertion. | Add the approved `rle` exception record to Task 24's closing-body list and its validation, or remove the claim that the closing body records it. +MINOR | high | docs/superpowers/plans/2026-08-29-review-loop-economics-plan-a-rules.md:105-111; CLAUDE.md:664-670; plugins/dev-workflow/commands/workflow-init.md:850-856 | Plan A's mandatory old-condition accounting says profile case 3 has a decidable order of tests, while the shipped rule explicitly says no test order is prescribed. | The accounting is false and cannot demonstrate that the rewritten decision procedure preserved, moved, or deliberately dropped every condition. | Change the accounting to state that ordered tests were deliberately dropped and explain the replacement observation rule. +MINOR | high | docs/superpowers/plans/2026-08-30-review-loop-economics-plan-c-rollout.md:127,1102-1105,1247-1254 | Plan C's old-condition table still says cardinality is the one grammar condition decided by counting, although Task 21 rejects that claim and Task 22 names four comparison-only constraints. | The plan retains the exact enforcement overclaim its later task says it corrected, weakening the audit trail for what the grammar check proves. | Update row 21 to name field comparisons generally and point to the non-exhaustive four-item list instead of claiming cardinality is the sole case. +NIT | high | docs/superpowers/plans/2026-08-30-review-loop-economics-plan-c-rollout.md:57-59,1214-1228 | Plan C says Task 13 reads the floor-knob file, but Task 13 changes the model-recording destination; the knob observation is Task 22 item 4a. | A reader following the cross-reference is sent to the wrong procedure and cannot verify the no-write evidence duty efficiently. | Replace `Task 13` with `Task 22 item 4a`. +END OF FINDINGS (13 total) diff --git a/.gitignore b/.gitignore index f4fea79..3a7454e 100644 --- a/.gitignore +++ b/.gitignore @@ -10,8 +10,23 @@ _unrelated-seo-work/ # Hook state. The adoption marker `.context/codex-gate.on` is meant to be committed # and shared; the mutable counters are per-clone scratch and must not be — a committed # passCount hands every fresh clone a pre-counted Gate-B pass. +# +# `codex-reviews/` is tracked from 2026-09-10, for the same reason `docs/field-reports/` +# is: three field reports exist only because someone hand-carried per-pass counts out of +# this directory before a `.context/` clear destroyed them. Tracked, those counts survive +# on their own and `grep -cE '^(BLOCKER|MAJOR|MINOR|NIT) \|'` re-derives any curve. This +# does NOT affect the Gate-B fingerprint: `codex-gate.sh` excludes `.context/` from all +# three of its components by pathspec and `git rm --cached`, independent of this file, and +# its own comment records that "`.context/` is committed in some projects". +# +# KNOWN DIVERGENCE, deliberate and local: `CLAUDE.md` §5 ("Why `.context/codex-reviews/`") +# still says to ignore that path, and the `/workflow-init` template still writes that +# instruction into target projects. Both are unchanged on purpose — editing §5 is a +# two-copy prompt change with a full Gate-B cycle and a plugin version bump, which is a +# separate decision. This repo deviates; nothing shipped does. .context/* !.context/codex-gate.on +!.context/codex-reviews/ # Local drafts and reading notes — working material, not repo content. `docs/field-reports/` # is deliberately NOT here: field reports are tracked, because a field-intake round cites diff --git a/todos.md b/todos.md index 1133bb9..cd92822 100644 --- a/todos.md +++ b/todos.md @@ -631,8 +631,11 @@ backlog. surviving prior file is indistinguishable from a fresh one. The consequence is that a second cycle in the same repo silently erases the first cycle's findings artifacts. **Observed, not theorised:** the result-classification cycle's pass-1 call deleted the - 2026-07-26 profiles cycle's 11 KB `gate-a-spec-pass-1.md`. `.context/` is git-ignored, - so it is unrecoverable. §5 anticipates *concurrent* calls racing on one slot and says + 2026-07-26 profiles cycle's 11 KB `gate-a-spec-pass-1.md`. `.context/` was git-ignored, + so that file is unrecoverable. **Partly mitigated 2026-09-10:** `.context/codex-reviews/` + is tracked in this repo, so a slot overwritten after a commit is now recoverable from git + — the destroy-before-delete window between two commits is not, and target projects still + ignore the path, so the row stands. §5 anticipates *concurrent* calls racing on one slot and says so; it does not cover *sequential cycles* reusing them. Note the dispositions and resume-note companions have the same property. Any fix has to keep the pre-call delete — that check is load-bearing — so it is about naming (a cycle component in the slot) or @@ -672,8 +675,10 @@ backlog. - [ ] **A recording mechanism for severity normalization.** Rider (b) normalizes an unrecognized severity token to `MAJOR` and records nothing. The drift is visible to the reader at the moment the pass is validated — the findings file carries the original token - on the finding line — but nothing is durable: `.context/` is git-ignored and slot - collisions have destroyed findings here (row above). A companion record was designed and + on the finding line — but nothing is durable *in a target project*: `.context/` is + git-ignored there and slot collisions have destroyed findings here (row above). In this + repo `.context/codex-reviews/` is tracked from 2026-09-10, which makes the original token + recoverable here and nowhere else. A companion record was designed and **cut**, at a measured cost: it needed a token-identity rule, a bijection audit, a logical-pass/attempt/credited-count identity model, edits to four shipped hook reminder strings, and a `docs/hardening-log.md` supersession row. From 17cf2e9de77388ed0ada3d3c2531d089f2bef87c Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Fri, 11 Sep 2026 10:36:30 +0200 Subject: [PATCH 023/181] docs(specs): apply Gate-A pass 14; extend the coherence instruction (item 22) Pass 14: 10 findings (1 Blocker, 5 Major, 3 Minor, 1 Nit), all applied. Blocker+Major 10 -> 6, the lowest of the cycle. Zero of five tells. SCOPE STOP on finding 6, answered by Daniel: B, bounded. The partial-adoption question was surfaced rather than absorbed, because the repair is a source edit to a paragraph the story never set out to change and the spec admits the finding is true of the artifact, so no dismissal was available. Daniel's answer: extend the existing coherence rule to the new closure logic and its dependent edits, describe the effect correctly as an instruction to the agent, and build neither a checking system nor any record-durability solution. Shipped as section 4 item 22, with section 9 rewritten to claim only what that sentence does: it is an instruction, not a guard; nothing detects a partial adoption; a project that does not adopt the sentence is not reached by it. Four corrections Daniel made to the recommendation that preceded the answer, kept because each was right and each had been argued the other way: - a documented gap is not an accepted risk -- the user project runs the copied CLAUDE.md, not this spec; - calling the existing rule a guard was the overclaim AGENTS.md names as this repo's most persistent defect; - "remove restatement" never meant "refuse every expansion", and the compatibility of the rules this change introduces belongs to this change; - B costs no extra pass, since a Blocker and five Majors oblige one regardless. The other nine: 1 BLOCKER the block widened c8 from a paragraph that does not own it; the sentence becomes a citation and item 21 makes the widening at the source, recording c8 as replaced rather than kept 2 MAJOR the block invented a withdrawal path for a decline, which D7 does not admit; removed, reconsideration is a later cycle's 3 MAJOR a parked cycle had no named next state; continue now returns it to suspended-awaiting-answer while any answer is outstanding 4 MAJOR section 7's oracle failed a legitimate re-raised health stop; it now fails only a return with its reading unconsumed 5 MAJOR the verification residual read as exhaustive; predicate derivation and reachable-combination completeness are named as also unestablished 7 MINOR the (g) text pointed at the wrong authority for fix-set membership 8 MINOR the curve-reproduction claim omitted the Majors series 9 MINOR a health-only surface had no stated hold discharge 10 NIT section 4 item 7 quoted a sentence occurring in neither copy; the source wording is used and its three-line wrap named (C 129-131, W 336-338) Counts updated: 20 -> 22 edits, 19 replacements + 3 additions, six outside the inventoried passages and sixteen inside. Verified mechanically: precheck exit 0, 22 table rows, the kind tally matches, the withdrawal path is gone, and all three newly quoted sentences plus item 22's target paragraph occur in both copies. Spec 532 -> 578 lines. No gate: every staged path is docs/**.md or the cycle's own working record under .context/, which section 5 exempts as prose. The spec is still in Gate A and not approved. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-14.md | 11 ++ .../gate-a-spec-awsf1ec771-resume.md | 59 +++++++- ...26-09-10-loop-rule-consolidation-design.md | 136 ++++++++++++------ 3 files changed, 160 insertions(+), 46 deletions(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-14.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-14.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-14.md new file mode 100644 index 0000000..4deb998 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-14.md @@ -0,0 +1,11 @@ +BLOCKER | high | §3 "The four standing duties, classified" | The block makes a repeatedly re-raised dismissal count as regeneration and says a recurring false positive reaches the clearly-stuck suspension, but the retained authoritative c8 reading still requires genuine repair attempts with each round's fix producing the next, while a dismissed false positive receives no repair | A product-behaviour Blocker repeatedly dismissed and re-raised can remain unclean with only one tell, so following c8 leaves the cycle unable to close or suspend; following the block silently broadens c8 from a different authority | Either edit c8 at its source to admit validated dismissals and name that semantic change in §4 and the old-condition accounting, or add a separately defined suspension path; do not redefine regeneration only in the ordering block +MAJOR | high | §3 "What a suspension asks, and what ends it" | The proposed withdrawal path semantically reverses a decline during the same cycle even while saying D7 admits no exception, and b7's planned "minus this cycle's declines" still excludes the finding after the alleged withdrawal | This contradicts settled D7 and gives the next pass incompatible instructions about whether the finding takes the ordinary route or remains outside the assigned fix set | Remove withdrawal from this cycle; keep the original decline binding through closure and require reconsideration in a later cycle, as D7 specifies +MAJOR | high | §3 "Composition, and what cannot happen" | When a health answer parks a suspension that still has unanswered membership or structural questions, "restarted only by an explicit later continue" does not state whether that continue returns to suspended-awaiting-answer or starts another pass, while the composition rule forbids a pass until every answer resumes it | A multi-exit cycle can either run over an unanswered question or remain parked with no explicit next-state transition, so the two-exits-at-once path is not executable without inference | State that later continue moves a parked cycle to suspended-awaiting-answer while any other answer remains outstanding, and starts the next pass only after all required answers resume it +MAJOR | high | §7 "The oracle" | The oracle rejects the same stop returning "whether or not an input was consumed," but §3 explicitly permits the same health suspension after continue when a new validated pass recomputes the reading from new data | A legitimate re-raised two-tell or clearly-stuck result is classified as a failed transition, so the named verification disagrees with the behavior it is meant to verify | Fail only an immediate return of the same consumed reading without an intervening validated pass; allow the same stop after the required post-answer recomputation +MAJOR | high | §7 "The named verification of the risk path" | After narrowing the next-state table to transitions whose predicates are assumed, the spec says the omitted coverage consists of logical-pass validation and final-acceptance preconditions, but it also leaves predicate derivation and completeness of the reachable predicate combinations unestablished | The residual reads as exhaustive and can make the evidence entry overclaim coverage of clean/scope/health interaction paths that the table assumes rather than verifies | State that the residual list is non-exhaustive and explicitly include predicate derivation and reachable-combination completeness among what the table does not establish, without adding the parked fixture-per-predicate mechanism +MAJOR | high | §9 "One residual this change owns" | The ordering block and its twenty coupled source edits can be partially adopted downstream with no coherence guard, and the spec itself confirms a project can take the clean predicate without the scoped resolve duty or only one meaning of clean and still run | A downstream cycle can close over an open duty or loop under contradictory predicates after an ordinary partial merge, failing the compatibility and partial-adoption risk path for this high-risk prompt change | Extend the existing semantic partial-adoption stop to treat this closure block and its coupled source edits as one closure contract, independently of the deferred record-specific guard +MINOR | high | §4 passage (g) replacement | The shipped severity text says fix-set membership is "a separate predicate the closure ordering defines," while §3's authority table says b7 and b8 in the absorb paragraph are its one definition and the block only cites them | The prompt points readers to the wrong authority and makes the spec's one-definition claim internally false | Replace that attribution with "the absorb paragraph defines" or an equivalent citation to b7 and b8 +MINOR | high | §4 passage (g) replacement | The rationale claims that counting finding lines and leading BLOCKER fields reproduces the self-reported curve, but the live curve grammar in both copies also records a Majors series | The proposed check cannot reproduce the complete curve it claims to make checkable, so an incorrect Major count survives that procedure | Include leading MAJOR fields in the reproduction rule and keep subject-cluster classification separate from mechanical severity counts +MINOR | medium | §3 "What a suspension asks, and what ends it" | Holds attach to every surfaced finding regardless of suspension, but the only enumerated hold-discharge answers are the one or two scope-stop answers and the health continue-or-stop answer is explicitly not counted among them | For a clearly-stuck-only finding the text does not visibly say what discharges its attached hold, contrary to AC 2 and D4, even though the separate health suspension still blocks immediate closure | State whether the shared health answer discharges holds created by a health-only surface, while remaining additional to any scope-stop answers +NIT | high | §4 item 7 | The quoted phrase "no rule about how a cycle closes is restated" occurs in neither prompt copy; the source instead says "Every other rule stated here about how a cycle closes stands as written, and none of them is restated" across CLAUDE.md:128-131 and workflow-init.md:335-338 | The required quotation check fails mechanically even though the intended source sentence is identifiable and wraps across four lines in both copies | Use the source wording or label the current text as a paraphrase, and note its four-line wrap in each copy +END OF FINDINGS (10 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index 310ce74..1e979a5 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -38,7 +38,63 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 12 | 11b0e47 | 21 | 1 | 13 | yes | **two mechanisms ended by Daniel's repeat criterion**, brought in mid-loop from another project; session 01a08c7f-c573-7df1-8b31-7704ec7be064 | | 13 | 78e5f97 | 16 | 2 | 8 | yes | **the loop turned: B+M 14 → 10, Majors 13 → 8**; session 01a08c98-b6cd-7bc3-9556-8f1c87491031 | -| 14 | 0168f88 | — | — | — | not run | next action. Spec cut to 532 lines; design 36% of it. | +| 14 | c06dd69 | 16→**10** | 2→**1** | 8→**5** | yes | spec byte-identical to 0168f88 (532 lines); **lowest B+M of the cycle, 6**; **SCOPE STOP surfaced on finding 6** (partial-adoption guard = a 21st edit, outside the twenty); session 01a08f70-6d1e-7cf1-baa3-6df3cc59bc32 | +| 15 | (pending) | — | — | — | not run | next action. All 10 pass-14 findings applied; spec 532 → 578 lines, §4 now 22 edits | + +## Pass-14 three-line report — SCOPE STOP, and the loop is converging + +**Floor line (owed every pass):** derived floor **3**; risk **high**, security **none**; read fresh +at this pass from `docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md`. One cited +story, profiled and resolvable at level 2. Floor long met; what is owed is a clean final pass. + +- **Trend:** findings …, 20, 21, 16, **10**. Blockers …, 1, 1, 2, **1**. Majors …, 13, 13, 8, **5**. + Blocker+Major 20, 15, 10, 11, 15, 13, 9, 8, 12, 17, 14, 14, 10, **6** — **the lowest of the + cycle**, and the second consecutive fall. Findings are also the lowest since pass 3. +- **Cluster (pass 14):** product 6 of 10 (findings 1, 2, 3, 6, 7, 9); the instrument 3 (§7's + oracle and its residual, plus the (g) rationale's counting claim); bookkeeping 1 (the NIT's + failed quotation check). +- **require↔withdraw:** none. Finding 2 demands *removing* the withdrawal path, which is a later + pass objecting to an earlier pass's addition — the mirror of the pair's shape, not the shape. + +**Tells: zero of five.** Findings falling hard, Blockers falling, cluster on product, no pair. The +clearly-stuck reading is not reachable either: its third condition needs Blocker/Major regenerating +across repair attempts, and this round's fixes reduced them. + +**The pass-13 note held.** Pass 13 predicted the reviewer was "running out of internal problems and +working the boundary". Pass 14 confirms it: findings 1, 4, 7 and 10 are the spec's edges meeting +standing text (`c8`, §7's own oracle, `b7`/`b8`, a §5 sentence that wraps across four lines), not +the spec contradicting itself. + +**SCOPE STOP on finding 6, surfaced rather than absorbed.** The finding asks that the existing +semantic partial-adoption stop be extended to cover the closure block and its twenty coupled +edits. It is **true of the artifact** — §9 says so itself: "That is an admitted unsafe state, not a +guarded one, and it is this change's own." So a validated dismissal is unavailable; a dismissal +says a finding is false, and this one is not. And the repair is a **twenty-first edit** to a +paragraph the story never set out to change, which leaves the assigned fix set. Under the absorb +rule that stops the loop and goes to the user. The question is a genuine either/or: **ship the +admitted residual, or spend one more source edit on a guard.** + +**SCOPE STOP ANSWERED 2026-09-10 by Daniel: B — accept, bounded.** In his words, to be carried +into the text: *"eine begrenzte Erweiterung der vorhandenen Konsistenzregel auf die neue +Abschlusslogik und ihre abhängigen Änderungen … beschreibe die Wirkung korrekt als Anweisung an +den Agenten. Daraus soll weder ein neues Prüfsystem noch eine zusätzliche Record-Durability-Lösung +entstehen."* Shipped as **§4 item 22**. Four corrections he made to the recommendation that +preceded it, recorded because each was right and each had been argued the other way here: +documented is not accepted, since the user project runs the copied `CLAUDE.md` and not this spec; +the existing rule is an **instruction**, not a guard, and calling it one was the overclaim +`AGENTS.md` names; "remove restatement" never meant "refuse every expansion", and the +compatibility of the rules this change introduces belongs to this change; and B carries **no extra +pass**, because one Blocker and five Majors oblige a further pass regardless. + +**The other nine are in-set and repairable**, and one is worth naming because it is a settled +decision being contradicted: finding 2 shows the block invented a **withdrawal** path for a +decline, which **D7** does not admit. Removing it restores the settled input rather than deciding +anything new. + +**A precheck gap, recorded not fixed.** The precheck reported "OLD blocks checked: 0" and finding +10 is exactly what an OLD-block check would have caught — a §4 quotation occurring in neither +copy. It is a mechanical check the reviewer spent judgement on. Candidate for the precheck; not a +new rule, and not this cycle's to build. ## Pass-13 three-line report @@ -363,6 +419,7 @@ nothing checks them. They are candidates captured in | pass 10 two-tell stop | **Split.** Record durability moves to `docs/superpowers/stories/2026-09-10-record-durability-story.md`; this story keeps the ordering, the duty classification and the severity/health answer. | | mid-loop, from another project | Anchor the repeat criterion in the kit; captured as the harness-finding-termination story. | | after pass 13 | **Cut the spec to the design**; the plan carries the bookkeeping. Precheck runs before every pass. | +| pass 14 scope stop | **B — accept, bounded.** Extend the existing coherence *instruction* to the closure block and its coupled edits (§4 item 22). No checker, no new mechanism, no record-durability work. | ## For the execution phase, not needed yet diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index 1b16d7f..c351ab1 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -28,8 +28,8 @@ cycle and which merely *suspend* it, in what order a pass is read so the ranking rather than asserted, what any set of suspensions at once does, and which of the four standing duties participate in that ordering versus gate it as preconditions. With it: the answer to what a severity demotion does to the loop-health counts, and the standing sentences the ordering -falsifies or leaves ambiguous if they are not edited at their source — **five outside the -inventoried passages and fifteen inside them, eighteen replacing a sentence and two adding +falsifies or leaves ambiguous if they are not edited at their source — **six outside the +inventoried passages and sixteen inside them, nineteen replacing a sentence and three adding one**. **What does not:** the pass floor and severity semantics, which the parent shipped and this spec @@ -60,11 +60,11 @@ answer produces; and the raw-severity rule for the health measures. **It defines duty or severity rule of its own** — each keeps its one definition where that definition already lives, and where one had to change to agree with the ordering it changed **at its source**. -**Twenty source edits**, counting **one edit per contiguous replacement or addition at one +**Twenty-two source edits**, counting **one edit per contiguous replacement or addition at one site**, which is the unit because a single sentence can be replaced once and a list added to once -without either being two. §4 lists all twenty. The **closure-ordering block is an addition -beside them**, not one of the twenty, so a reader counting changes to the two copies counts -twenty-one. +without either being two. §4 lists all twenty-two. The **closure-ordering block is an addition +beside them**, not one of the twenty-two, so a reader counting changes to the two copies counts +twenty-three. --- @@ -148,10 +148,11 @@ author's judgement, carrying the one-line why this section already requires of a finding, that the finding is not true of the artifact. A dismissal does not rewrite the pass that found it, and the later clean pass is still owed. **A dismissal is not a decline**: a dismissal says the finding is false, a decline is the user's decision that a **true** finding stays outside -the set, and only the second is an answer at a membership stop. **A dismissal the reviewer keeps -re-raising across passes is regeneration**, counting toward the clearly-stuck reading's third -condition — so a false positive that returns every pass reaches that suspension instead of -continuing forever. The **hold** a surfaced finding places on closure **participates in the +the set, and only the second is an answer at a membership stop. A dismissal the reviewer keeps +re-raising across passes is covered by the clearly-stuck reading's **third condition as that +paragraph states it**, which is where that condition is defined and where it was widened to admit +it — so a false positive that returns every pass reaches that suspension instead of continuing +forever, and this sentence cites the reading rather than extending it. The **hold** a surfaced finding places on closure **participates in the ordering**: it gates closing while it stands, and is discharged by the answers that surface requires. **It attaches to every surfaced finding, whichever suspension surfaced it** — clean completion creates none, because it wins before anything is surfaced. **No-clean-credit** — no @@ -166,15 +167,17 @@ single-trigger finding, both for one carrying both. That is **one rule with two answers and which way each may go, and neither is a test the other has to pass. Where a health suspension applies to the same pass, its shared continue-or-stop answer is **additional** to those and not counted among them, the health readings asking about the loop rather than about -this finding. At a **membership stop** the answer is **accept**, the finding joining the fix set +this finding — and where the only surface was a health reading, that same answer is what +discharges the hold it created, there being no scope-stop answer for it to be additional to. +At a **membership stop** the answer is **accept**, the finding joining the fix set where Mechanics Severity governs it, or **decline**, the finding staying outside and binding so for the rest of this cycle. A later answer that contradicts a decline **does not reverse it**: the decline **remains binding**, the contradiction is **surfaced to the user as information**, and the **cycle continues** — nothing here turns one answer into another, since that would let a -finding be moved out of the set and back into it to escape what it owes inside it, and **D7** -admits no exception. What the user may do is **withdraw** the decline, which is not a reversal by -these rules but a fresh explicit decision by the same authority that declined; the finding then -takes the ordinary route at the next pass that raises it. Either answer is an **explicit, +finding be moved out of the set and back into it to escape what it owes inside it. **There is no +withdrawal inside the cycle that declined**, because **D7** binds a decline for the remainder of +its cycle and admits no exception; a reconsideration is a later cycle's, where D7 gives the +decline no effect at all and the finding takes the ordinary route. Either answer is an **explicit, attributable decision on that specific finding** — never silence, never a general remark about scope, never inferred, because a fix set changed by inference is a fix set nobody chose. **Membership is answered against the set as the absorb paragraph fixes it for the pass that @@ -192,8 +195,14 @@ scope requires one, that repair comes before the post-answer pass, since a pass unrepaired in-set Blocker or Major spends a look on text the rules already say must change. **Stop parks the cycle**: open, not running, spending no passes, restarted only by an explicit later continue — a distinct state from the suspended-awaiting-answer one it was in -before the answer. Nothing a parked cycle wrote is a closing commit, and a parked cycle nobody -restarts is a human's to resolve, exactly as the nonce rules already say of open cycles. +before the answer. **That continue restarts the cycle and never skips an answer**: where any +question the suspension raised is still outstanding, it returns the cycle to +suspended-awaiting-answer, and only once every answer the composition rule requires has been +given does the next pass run. So a cycle parked with an unanswered membership or question stop +cannot be continued into a pass, and cannot sit parked with no transition either — the continue +is always available and always moves it. Nothing a parked cycle wrote is a closing commit, and a +parked cycle nobody restarts is a human's to resolve, exactly as the nonce rules already say of +open cycles. **Composition, and what cannot happen.** Every **question** is answered on its own and the loop resumes only when every answer resumes it — accept or decline at a membership stop, a decision at @@ -262,7 +271,7 @@ above — the ordering claims no in-session consequence it does not state there. ## 4. The standing sentences edited at their source -Twenty edits. Each row names the sentence, where it lives, what changes and why, and whether it +Twenty-two edits. Each row names the sentence, where it lives, what changes and why, and whether it is a replacement or an addition. **No OLD or NEW text and no line numbers appear here**: the plan quotes each sentence from the real file, writes the replacement, and re-greps the site, for the reason §7 gives. Working around any of these would ship two instructions that disagree. @@ -275,7 +284,7 @@ reason §7 gives. Working around any of these would ship two instructions that d | 4 | the lens paragraph's unchanged-list | the profiles section, C and W | it asserts the Blocker/Major filter and the clean-final-pass rule are unchanged, which the ordering falsifies; scoped to **the lens sets**, which is what that paragraph is about | replacement | | 5 | the Gate-A cadence | the Gate-A section, C and W | "validate, revise, re-run" makes a revision unconditional, telling a Minor-only pass to manufacture the repair the severity rule forbids; the revision becomes conditional on a repair being required | replacement | | 6 | `a17`–`a19`, the floor paragraph's closure sentences | passage (a), C and W | they state closure and the early exit in the paragraph that owns the floor; trimmed to point at the ordering, which states them once | replacement | -| 7 | `a13` | passage (a) | "no rule about how a cycle closes is restated" becomes categorically false once the ordering exists; scoped to that paragraph, which is what it was written to police | replacement | +| 7 | `a13` | passage (a) | the source reads "Every other rule stated here about how a cycle closes stands as written, and none of them is restated — a summary is where their conditions would get dropped", **wrapping across three lines in each copy** (C 129–131, W 336–338), so the plan greps it in parts or unwrapped rather than as one line. It becomes categorically false once the ordering exists; scoped to that paragraph, which is what it was written to police | replacement | | 8 | `a16` | passage (a) | "fix Blocker/Major after each" stands unscoped beside a Severity bullet item 3 now scopes; it points at that rule instead of restating an unscoped version | replacement | | 9 | `b3` | passage (b), C and W | it points at Mechanics and then **restates the four severity actions**, giving them two definitions; it becomes a pure pointer. W's pointer also changes target, matching C (§6) | replacement | | 10 | `b7`, the fix-set definition | passage (b) | singular "the approved story or plan" leaves a cycle governed by several with no set at all; it becomes their **union**, plus obligations already accepted, **minus this cycle's declines** | replacement | @@ -289,6 +298,8 @@ reason §7 gives. Working around any of these would ship two instructions that d | 18 | the five-tells pointer | passage (e) | the passage says nothing about what its answer does; one sentence names the stop a suspension and defers to the ordering | addition | | 19 | the handed-over severity question | Mechanics · Severity, passage (g) | the paragraph says the demotion/loop-health question is unsettled and mandates a stop; replaced by the answer, which is the (g) text below | replacement | | 20 | the unknown-start strict-reading list | passage (i) | the list of what a cycle takes at its strictest omits the suspensions this change ships; it gains **every suspension binding, since unknown starting rules cannot waive an open hold** — the reason carried inline, as prompt-standards item 6 requires of the clause itself | addition | +| 21 | `c8`, the third condition itself | passage (c) | a **semantic widening, made at the source rather than from the block**: `c8` requires Blocker/Major findings regenerating "across genuine repair attempts, each round's fix producing the next", which a **validated dismissal** never satisfies — it gets no repair and produces no fix, so a false positive the reviewer re-raises every pass leaves the cycle unable to close and unable to suspend. It gains that case: a finding the author has validly dismissed and the reviewer re-raises across passes counts toward the third condition, the re-raise standing in for the regenerating fix. The old-condition accounting records `c8` as **replaced**, not kept | replacement | +| 22 | the one-contract paragraph's membership list | Mechanics, the "These records are one contract" paragraph, C and W — **outside the ten inventoried passages** | it names the nonce, the slot naming, the provenance line, the curve, the carry rule and the unknown-start activation semantics, and reaches none of this change; it gains **the closure ordering block together with the §4 edits it depends on**, so a project holding some of them, or versions of them that disagree, meets the stop that paragraph already states. **What this is: the same instruction to the agent, over a wider membership.** It is not a checker and builds none; nothing mechanical detects a partial adoption, and a project that does not adopt this sentence is not reached by it. Scope is exactly the coupled set §9 names and nothing else | addition | **Passage (g)'s replacement is design rather than bookkeeping, so it is stated here in full:** @@ -299,11 +310,12 @@ reason §7 gives. Working around any of these would ship two instructions that d finding total and in its cluster, and a Blocker demoted to Minor is still a Blocker to the curve. Cleanliness and the resolve duty read the effective severity, after the ceiling (the closure ordering above). Two reasons for the split. The curve must stay derivable from the - findings files alone — counting finding lines and leading `BLOCKER` fields per pass - reproduces it, which is the only thing that makes a self-reported curve checkable. And the - demotion is the author's judgement about **the finding's repair severity**, never about - which findings the fix set contains — a separate predicate the closure ordering defines, and - one this must not be read as touching; a loop spending passes on findings the author keeps + findings files alone — counting finding lines and leading `BLOCKER` and `MAJOR` fields per + pass reproduces its three series, which is the only thing that makes a self-reported curve + checkable; the subject clusters are a judgement per finding and no count reproduces them. + And the demotion is the author's judgement about **the finding's repair severity**, never + about which findings the fix set contains — a separate predicate the absorb paragraph + defines, and one this must not be read as touching; a loop spending passes on findings the author keeps demoting is exactly what the prose-cluster tell exists to surface, and lowering the counts by that same judgement would hide it. ``` @@ -312,6 +324,13 @@ Two sentences are **not** edited and are named so nobody looks for them: the "Co into the squash body" sentence inside the human-exception block, and the "records every cycle owes" list. This change ships no record. +**Items 21 and 22 were added at Gate-A pass 14.** Item 21 answers a Blocker: the block was +widening `c8` from a paragraph that does not own it. Item 22 answers a Major on Daniel's decision +of 2026-09-10, as a **bounded extension of the existing coherence instruction** — no checker, no +new mechanism, and nothing about record durability, which stays with the successor story. Item 22 +is the only edit in this table that touches a passage no earlier revision named, and §9 states +what it is worth. + --- ## 5. The passage map, and who carries the accounting @@ -326,7 +345,7 @@ here. This table says what happens to each inventoried passage, so the map stays |---|---|---| | (a) the floor paragraphs | edited — closure sentences trimmed to a pointer, the no-restating prohibition scoped, the per-pass fix command pointed at Mechanics | 6, 7, 8 | | (b) what a loop absorbs | edited — owns both triggers and the fix set, so every qualification the ordering needs is made **here**, which is what keeps one definition per rule | 9, 10, 11, 12, 13, 14, 15 | -| (c) recognizing clearly stuck | edited — keeps its three-condition reading, stops carrying evaluation order; its precedence sentence moves into the block unchanged | 16 | +| (c) recognizing clearly stuck | edited — keeps its three-condition reading, stops carrying evaluation order; its precedence sentence moves into the block unchanged, and the third condition itself is widened at this source to admit a re-raised validated dismissal | 16, 21 | | (d) from pass 4 onward | **no longer edited.** The unavailable-history block moved to the successor with **D10** | — | | (e) the five tells | edited — the threshold is read after clean completion, and a pointer says what its answer does | 17, 18 | | (f) the two rules above do not compete | **unchanged.** "The two rules above" still names the absorb rule and the stuck reading; the block sits before both and adds no third rule between them | — | @@ -411,12 +430,27 @@ across every required branch file, and that each final-acceptance precondition t held — the floor, the cited set and profile, and the evidence entry's revalidation. **No fixture per predicate is built**; that question is parked in the story's §2 and is not reopened. +**That list is not exhaustive, and reading it as exhaustive is how the evidence entry would +overclaim.** Two further things the table does not establish, named because they are the ones a +reader would otherwise assume it covers: **how each predicate was derived** — the table takes a +row's predicates as given and checks the transition out of them, so a wrongly derived predicate +produces a row that passes; and **whether the enumerated rows cover every reachable combination** +of the clean, scope and health predicates — nothing enumerates that space, so a missing +combination is invisible rather than failing. Both stay unestablished here: closing either is the +parked fixture-per-predicate question, which this change does not reopen. **The evidence entry +states the table's claim at this width**, not wider. + **The oracle.** A row **fails** when its required answer does not produce a **distinct** resumable -or closed state — the same stop returning, whether or not an input was consumed — or when it -closes with **any** closure precondition unmet: an in-set Blocker or Major, a standing hold, an -unanswered question, the derived floor, a changed evidence entry, or a cited-set or profile change -during the pass. Naming only the first would pass the exact no-progress defect AC 4 cites from -the parent cycle. +or closed state — **the same stop returning with its reading unconsumed, that is, without an +intervening validated pass run after the answer** — or when it closes with **any** closure +precondition unmet: an in-set Blocker or Major, a standing hold, an unanswered question, the +derived floor, a changed evidence entry, or a cited-set or profile change during the pass. Naming +only the first would pass the exact no-progress defect AC 4 cites from the parent cycle. **The +consumption clause is what keeps the oracle and §3 in agreement**: §3 says continue consumes the +reading that raised the suspension and a further health suspension needs it recomputed over a +pass run after the answer, so the *same* two-tell or clearly-stuck result **after** such a pass is +new data and a legitimate row, not a failed transition. An earlier wording failed it "whether or +not an input was consumed", which would have classified that legitimate case as a defect. **Evidence entry**, in the closing commit body, names: the battery run; every pair the plan built with its counts in each copy and each tree, and every presence check beside them; the §6 parity @@ -467,7 +501,7 @@ repo's most persistent defect. The transport that could carry it left with the r `plugins/dev-workflow/commands/workflow-init.md` is under `plugins/`, so `plugins/dev-workflow/.claude-plugin/plugin.json` goes `0.11.0 → 0.12.0` with a `plugins/dev-workflow/CHANGELOG.md` entry: a minor bump, the template gaining a closure - ordering and twenty edited or extended sentences. + ordering and twenty-two edited or extended sentences. - **Invariant 4 / the hook.** Untouched: `plugins/dev-workflow/hooks/codex-gate.sh` is not edited, and the §5 heading it greps (`Cross-Model Review`) does not move. @@ -503,24 +537,36 @@ decision of 2026-09-10, with the evidence in §1: record it was named for. **Moved to the plan** — the implementation artifact, not a story: the per-condition disposition -list for all 135 inventoried conditions (§5); the OLD and NEW text of all twenty edits with their +list for all 135 inventoried conditions (§5); the OLD and NEW text of all twenty-two edits with their line ranges (§4); the parity divergence list and the extraction-and-diff (§6); and every verification fragment with its counts (§7). Each is work this change still owes; none of it is work a spec can do correctly, because all four are checked against files the plan edits. -**One residual this change owns rather than moves: partial adoption of the narrowed set.** The -set is mutually dependent — the ordering's clean predicate needs §4 item 3's boundary, its two -senses of *clean* need items 1 and 2, and its triggers need items 10 through 14 — and **nothing -catches a downstream merge that takes some of it**. Said exactly: the live one-contract paragraph -names the nonce, the slots, the provenance line, the curve, the carry rule and the unknown-start -semantics, and **it does not name this block or any of its coupled edits**, so its coherence rule -does not reach them. An earlier revision of this spec said that rule caught them; it does not, and -a project can take the clean predicate without the scoped resolve duty, or either sense of *clean* -without the other, and run. **That is an admitted unsafe state, not a guarded one, and it is this -change's own.** An earlier revision assigned the guard to the successor; that was wrong on the -face of the successor's own scope, which is record durability and **excludes the closure ordering -by name**, so the assignment named an owner that had not taken it — the same defect as claiming a -mechanism that does not exist. Nothing here builds one. +**Partial adoption of the narrowed set — answered by §4 item 22, and what that answer is worth.** +The set is mutually dependent: the ordering's clean predicate needs §4 item 3's boundary, its two +senses of *clean* need items 1 and 2, and its triggers need items 10 through 14. **The coupled set +is exactly items 1, 2, 3, 10, 11, 12, 13, 14 and the block**, and item 22 names that set in the +live one-contract paragraph, which until now reached only the nonce, the slots, the provenance +line, the curve, the carry rule and the unknown-start semantics. + +**Stated as what it is.** That paragraph is **an instruction to the agent**: a project whose text +carries some members and not others, or versions that disagree, **stops and has a human complete, +revert or reconcile the adoption before running a gate under it**. Item 22 widens whom that +sentence is about. It is not a guard and not a mechanical check, and this change builds neither — +**nothing detects a partial adoption**, and the stop happens only where an agent reads the +sentence and acts on it. + +**What it therefore does not buy, said rather than implied.** A project that adopts the block +without adopting item 22 is not reached at all, which is the partial-adoption case applied to the +rule against partial adoption; the existing paragraph has the same property and item 22 neither +worsens nor repairs it. Nor does it detect a *silent* half-merge in a project that did adopt it — +it obliges a stop once someone notices, which is a different thing from noticing. **What it +removes is the narrower state this spec previously admitted**: that the coherence rule did not +name this material at all, so an agent reading the sentence and willing to act on it had nothing +to act on. Two earlier revisions got this wrong in opposite directions — one claimed the existing +rule already caught the block, and one assigned the guard to the successor story, whose scope is +record durability and **excludes the closure ordering by name**. Both named a mechanism that did +not exist; item 22 names a sentence that does, and claims only what that sentence does. **Out of scope and parked**, unchanged: From af05d72ac579610c3f68dd27c80b09a2deeb72b1 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Fri, 11 Sep 2026 12:04:47 +0200 Subject: [PATCH 024/181] docs(specs): bounded rollback after Gate-A pass 15; decide, don't enumerate Pass 15 hit the two-tell threshold (findings 10 -> 12, Blockers 1 -> 1) and six of its twelve findings regenerated from pass 14's own repairs. Surfaced; the reviewer's reading and Daniel's decision: a bounded rollback rather than another round of twelve single repairs. THREE CORRECTIONS TO THE OPTION I HAD RECOMMENDED, all taken. - My "B" would have moved open rule semantics into the plan. It cannot: a plan checks text against files, it does not decide behaviour. The behaviour questions stay decided in the design; only the exact wording and its verification move. - My "A" was misframed as "repair all twelve", which treats Minor and Nit as mandatory repairs. Section 5 says they are collected and never iterated. - Daniel's consent to widen the coherence instruction was NOT an instruction to introduce item-number lists. Those lists produced findings 5, 6 and 10. They are removed rather than corrected. DECIDED IN THE DESIGN (the three the ordering needs to be unambiguous): - Gate-A vs Gate-B closure (finding 1, Blocker). Eligibility is not closure; the cycle closes at the commit carrying its reviewed artifact, which is the Gate-B amend for a Gate-B cycle and the spec or plan commit for a Gate-A cycle, that gate having no WIP snapshot. The evidence-entry precondition is scoped to Gate B, the entry being about the diff. A commit the hook reads as cycle-closing is named as a separate matter: an observation about the counter, not a closure under these rules, so the standing non-WIP warnings stay true (finding 2). - Which severity the clearly-stuck reading takes (finding 3). Every loop-health reading takes the reviewer-written severity, the clearly-stuck regeneration condition included. Stated as the rule -- what a cycle owes versus what it observes about itself -- so no list of readings has to be kept complete. - What discharges the resolve duty (finding 4). Repair or validated dismissal, and what a dismissal is, move to Mechanics Severity via item 3. The block cites and defines none of it. - The health-only hold clause pass 14 added is removed (finding 11): a health reading surfaces no finding and creates no hold; what it leaves outstanding is its own continue-or-stop question. ROLLED BACK (repeated definitions and completeness claims, the generators): - Every stated edit total, in sections 1, 2, 4, 8 and 9. A count over spans that merge and split disagrees with its own table at the next revision, which is what finding 10 caught: items 16 and 21 replaced overlapping spans. They are merged into one contiguous replacement, and no total is claimed anywhere. - Section 9's "exactly items 1, 2, 3, 10-14" coupled set. It was wrong -- it omitted several edits the block plainly depends on. Replaced by the rule (an edit is coupled when the block cites or depends on it) with the plan deriving membership against the real files. - Item 21's shipped text carries that description, not item numbers, which do not exist outside this spec. Its paragraph's opening noun broadens from "records" to the rules and records a cycle runs under, since what is added is not a record (finding 6). - Section 7's presence-only exception, which listed item numbers, becomes a rule read off the Kind column (finding 7). - Section 7's oracle stops re-enumerating the closure preconditions and reads them from the block -- a re-enumeration is a second definition that drifts, and it was already missing two (findings 8 and 9). - Section 2's claim that each of D1-D8 is cited individually (finding 12, Nit; removed as part of the same cut, not as its own repair round). No known behavioural contradiction was renamed as a residual. Minor and Nit triggered no repair rounds of their own. Verified mechanically: precheck exit 0; section 4 numbering is 1-21 with no gap; zero occurrences of any edit total or item-number coupled set; the passage map no longer cites the merged row twice; all four behaviour decisions present in the shipped block. Spec 578 -> 598 lines. No gate: every staged path is docs/**.md or the cycle's own working record under .context/, which section 5 exempts as prose. The spec is still in Gate A. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-15.md | 13 ++ .../gate-a-spec-awsf1ec771-resume.md | 49 +++++- ...26-09-10-loop-rule-consolidation-design.md | 150 ++++++++++-------- 3 files changed, 146 insertions(+), 66 deletions(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-15.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-15.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-15.md new file mode 100644 index 0000000..c9c083e --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-15.md @@ -0,0 +1,13 @@ +BLOCKER | high | §3 "What a pass is read from" / "First, clean completion" | The ordering applies to Gate-A spec, Gate-A plan and Gate B, but it unconditionally reads the Gate-B-only evidence-entry revalidation rule and declares Gate B's closing WIP amend to be "the closure itself"; current §5 instead gives Gate-A cycles their spec/plan commit, while only Gate B has that evidence-and-amend path | A clean Gate-A pass has no executable closure under the stated ordering and no suspension to take, leaving the cycle unable to close and unable to suspend | Scope the evidence precondition and closing operation by cycle kind, citing the existing Gate-A spec/plan commit and Gate-B amend authorities +MAJOR | high | §3 "First, clean completion" | "No other pass outcome" closes and "nothing is closed" before the closing amend contradict the two standing non-WIP warnings at C 751–753 and 825–826 and W 937–939 and 1009–1010, which say a non-WIP commit is read as cycle-closing and discards the accumulated passes; §4 names neither passage | A misnamed snapshot can close the operational cycle over open duties while the new block claims that path cannot exist, so AC 1 and the old-passage list are incomplete | Distinguish valid clean closure from the hook-observed premature close, and name the coupled edits to both standing passages or revise the block so both sources remain true +MAJOR | high | §3 "What a pass is read from" / "Composition, and what cannot happen" | The block says Mechanics Severity settles the field read by every finding-derived predicate and later relies on the clearly-stuck reading using reviewer-written severity, but §4's (g) replacement assigns reviewer-written severity only to counts, clusters and tell thresholds; c8's regenerating Blocker-or-Major condition is not one of the five tells and receives no field | A below-floor pass with an in-set finding demoted below Major can either suspend as clearly stuck or continue depending on an unstated severity choice, and the claimed source citation is false | Make the Mechanics Severity source explicitly assign reviewer-written severity to c8's clearly-stuck severity test, while leaving the block as a citation +MAJOR | high | §3 "The four standing duties, classified" | The block says Mechanics Severity states what discharges the Blocker/Major-resolve duty, then the block itself introduces and defines repair versus validated dismissal; §4 item 3 only scopes the existing duty to the assigned fix set and does not put that discharge rule at its claimed source | The one-authority claim is false and the plan can leave Mechanics saying only "both must resolve" while closure depends on an interpretation defined elsewhere | Extend item 3 so Mechanics owns the repair-or-validated-dismissal discharge rule, and have the block cite that source rather than define the duty's discharge semantics +MAJOR | high | §9 "Partial adoption of the narrowed set" | The set claimed to be "exactly" items 1, 2, 3, 10–14 plus the block omits edits the block directly needs: item 15 prevents the old one-answer resume, item 16 removes the old clearly-stuck no-clean rule and duplicate precedence, item 17 qualifies the mandatory two-tell stop, item 19 supplies the cited severity split, and item 21 supplies the cited dismissal case, with further standing contradictions handled by items 4–9 and 20 | Item 22 can declare a materially partial adoption coherent while the adopted prompt still contains rules that disagree, defeating the authorised compatibility instruction | Recompute the coupled membership from every source rule whose old text conflicts with or is cited by the block, then make item 22 and §9 name that complete bounded set without adding a checker +MINOR | high | §4 item 22 | The live paragraph begins "These records are one contract", but item 22 is only an addition to its membership list and adds the closure ordering and source-rule edits, none of which are records | The shipped paragraph's opening classification becomes false and its later "these records" and "them" references ambiguously govern a mixed set, so §9 overstates the accuracy of the widened instruction | Broaden the paragraph's opening noun as part of item 22, or add a separately labelled closure-coherence sentence, and update the edit kind and dependent counts accordingly +MAJOR | high | §7 "The check" | The spec classifies item 22 as an addition, but the exhaustive presence-only exception names only items 18 and 20 and the block; item 22 is therefore left under the replacement-pair rule even though it has no old wording to prove gone | The plan has no valid prescribed counterfactual for the authorised one-contract edit, so the evidence can omit it or fabricate an impossible old-wording count | Add item 22 to the presence-only checks if it remains an addition, or require a discriminating replacement pair if item 22 is corrected to rewrite the paragraph's opening +MAJOR | high | §7 "The oracle" | The oracle's exhaustive closure-precondition list omits the clean-final-pass/no-clean-credit condition, even though §3 says a pass carrying a scope-stop trigger is never credited clean and only a later pass is judged anew | A transition row can answer a membership stop with decline and close immediately on that same trigger-bearing pass while satisfying every condition the oracle enumerates | Require every closing row to enter from the clean-completion or zero-finding branch, and explicitly reject closure after a scope answer until a later validated pass establishes eligibility +MAJOR | high | §7 "The oracle" | The oracle rejects closure for a cited-set or profile change only "during the pass", while §3 and the standing floor rules also gate a change after the pass and before the closing commit or amend | A row can pass verification while closing on a pass run under a stale profile or cited set, including a same-floor profile change that still owes a further pass | Extend the oracle's change window through final acceptance and the actual closing operation, matching the block's "in between" rule +MINOR | high | §2 "Twenty-two source edits" / §4 items 16 and 21 | The stated counting unit is one contiguous replacement at one site, but item 16 replaces passage (c) "from the third condition to the end" while item 21 separately replaces that same third condition c8 | The twenty-two-edit total double-counts an overlapping source span under its own unit, making the totals in §§1, 2, 4, 8 and 9 and the plan's OLD/NEW obligation inconsistent | Narrow item 16 to the text after c8, or merge items 16 and 21 and update every affected count +MINOR | high | §3 "What a suspension asks, and what ends it" | The duties paragraph says a hold attaches to every surfaced finding, and the composition paragraph says the two-tell stop surfaces tells rather than a finding, yet the repaired health-only clause says that surface created a hold | The four-duty classification has two incompatible hold domains, so a next-state table can model a two-tell-only suspension either as an unanswered health question or as an additional finding hold | Keep the health question as the suspension's own outstanding answer, or explicitly broaden the hold definition everywhere; do not call a no-finding health surface a surfaced-finding hold in only this clause +NIT | high | §2 "Settled inputs" | The spec says each settled decision D1 through D8 is cited below, but D1, D6 and D8 never appear individually after that claim; only the range notation at the claim itself contains them | The promised traceability to the story's settled decisions is mechanically false, making it harder to verify that only-clean closure, membership-only decline and explicit attribution came from fixed inputs | Add D1, D6 and D8 citations at the corresponding shipped rules, or weaken the claim that every decision is cited below +END OF FINDINGS (12 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index 1e979a5..9b8ecb6 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -39,7 +39,54 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 13 | 78e5f97 | 16 | 2 | 8 | yes | **the loop turned: B+M 14 → 10, Majors 13 → 8**; session 01a08c98-b6cd-7bc3-9556-8f1c87491031 | | 14 | c06dd69 | 16→**10** | 2→**1** | 8→**5** | yes | spec byte-identical to 0168f88 (532 lines); **lowest B+M of the cycle, 6**; **SCOPE STOP surfaced on finding 6** (partial-adoption guard = a 21st edit, outside the twenty); session 01a08f70-6d1e-7cf1-baa3-6df3cc59bc32 | -| 15 | (pending) | — | — | — | not run | next action. All 10 pass-14 findings applied; spec 532 → 578 lines, §4 now 22 edits | +| 15 | 17cf2e9 | 10→**12** | 1→**1** | 5→**7** | yes | **TWO-TELL STOP — mandatory, surfaced to Daniel.** B+M 6→8. **Six of twelve regenerate from pass 14's own repairs** (3, 4, 6, 7, 10, 11). All 12 held open; session 01a08fb2-6e8e-73d0-9d2b-8fb92d1571dd | +| 16 | — | — | — | — | not run | blocked on the two-tell answer | + +## Pass-15 three-line report — MANDATORY TWO-TELL STOP + +**Floor line:** derived floor **3**; risk **high**, security **none**; read fresh at this pass from +`docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md`. One cited story, profiled, +level 2. + +- **Trend:** findings …, 21, 16, 10, **12**. Blockers …, 1, 2, 1, **1**. Majors …, 13, 8, 5, **7**. + Blocker+Major 20, 15, 10, 11, 15, 13, 9, 8, 12, 17, 14, 14, 10, 6, **8** — the fall reversed. +- **Cluster (pass 15):** product 7 of 12 (1, 2, 3, 4, 5, 6, 11); the instrument 3 (§7's presence-only + list and two oracle gaps); bookkeeping 2 (the 16/21 double-count, the D1/D6/D8 citation claim). +- **require↔withdraw:** none under the definition. Finding 11 objects to the health-only hold clause + that pass 14 finding 9 **asked for**, which is a later pass questioning an earlier pass's + addition — the mirror of the pair's shape, not the shape. Named rather than hidden. + +**Tells: two of five — the threshold. Stop-and-surface is mandatory, not discretionary.** The +finding count rose 10 → 12; the Blocker count failed to fall, 1 → 1. The cluster is product and +there is no pair, so it is exactly two. + +**The clearly-stuck reading is also satisfied on all three conditions**, though it is not needed: +a plateau across fifteen passes (B+M never zero, never below 6); coverage affirmable after fifteen +readings; and Blocker/Major findings regenerating from the previous round's own repairs — **six of +twelve**, nameably: 3 and 4 are the citations pass 14's block edits lean on, 6 and 7 are item 22's +own mechanics, 10 is the item 16/21 overlap pass 14 created, 11 is the health-hold clause pass 14 +finding 9 required. + +**What the pass says about pass 14's repairs, measured.** Three of the four changes made to item 21 +and 22 carry a defect: the coupled set §9 calls "exactly" items 1, 2, 3, 10–14 **omits** items 15, +16, 17, 19 and 21, which the block also depends on (finding 5); item 22 adds non-records to a +paragraph opening "These records are one contract" (finding 6); and items 16 and 21 replace +overlapping spans, double-counting under the spec's own counting unit (finding 10). The +enumeration in finding 5 was copied from the spec's earlier sentence rather than recomputed — the +exact defect `AGENTS.md` names as "a rewrite reliably preserves the condition that motivated it +and silently loses the others". + +**One mechanism is on its fourth round, and it is the same one the repeat criterion was brought in +for.** Findings 1, 3 and 4 are one shape: **the block cites a rule whose source does not say what +the block claims it says.** Gate-A closure (1), which severity field `c8` reads (3), what +discharges the resolve duty (4). The same shape produced finding 14 at pass 11, four findings at +pass 12 and finding 1 at pass 14. Each round repaired the instance and left the mechanism. + +**Finding 1 is a real design defect regardless of what is decided**, and it is the one that cannot +be deferred: the block declares the Gate-B closing amend to be "the closure itself", but a Gate-A +cycle closes with its spec or plan commit and has no amend path — so a clean Gate-A pass under this +ordering can neither close nor suspend, which is precisely the state AC 4 exists to forbid. This +cycle is itself a Gate-A cycle. ## Pass-14 three-line report — SCOPE STOP, and the loop is converging diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index c351ab1..6ff28b3 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -28,9 +28,9 @@ cycle and which merely *suspend* it, in what order a pass is read so the ranking rather than asserted, what any set of suspensions at once does, and which of the four standing duties participate in that ordering versus gate it as preconditions. With it: the answer to what a severity demotion does to the loop-health counts, and the standing sentences the ordering -falsifies or leaves ambiguous if they are not edited at their source — **six outside the -inventoried passages and sixteen inside them, nineteen replacing a sentence and three adding -one**. +falsifies or leaves ambiguous if they are not edited at their source, which §4 lists row by row +**without claiming a total** — a count over spans that merge and split is bookkeeping the plan +re-derives against the files. **What does not:** the pass floor and severity semantics, which the parent shipped and this spec reads as given, and everything §9 lists as moved or parked. @@ -46,8 +46,8 @@ says what went and where. ## 2. Settled inputs -The story's §4 table is the design's starting point and is not restated here; each decision is -cited below as **D1**…**D8**. **D9**, **D9b**, **D9c** and **D10** stay settled and move to the +The story's §4 table is the design's starting point and is not restated here; decisions are cited +as **D1**…**D8** at the rules they settle, and no claim is made that each appears individually. **D9**, **D9b**, **D9c** and **D10** stay settled and move to the successor with the material they govern (§9). Two implementation facts the parent cycle established are read as given: a pass's cleanliness is a fact about what that pass found and is **never rewritten** — an answer changes whether the *cycle* may close; and the findings files @@ -60,11 +60,10 @@ answer produces; and the raw-severity rule for the health measures. **It defines duty or severity rule of its own** — each keeps its one definition where that definition already lives, and where one had to change to agree with the ordering it changed **at its source**. -**Twenty-two source edits**, counting **one edit per contiguous replacement or addition at one -site**, which is the unit because a single sentence can be replaced once and a list added to once -without either being two. §4 lists all twenty-two. The **closure-ordering block is an addition -beside them**, not one of the twenty-two, so a reader counting changes to the two copies counts -twenty-three. +**The closure-ordering block is an addition beside the source edits**, not one of them. §4 lists +the edits; **no total is stated here or there**, because the unit — one contiguous replacement at +one site — is not stable across revisions that merge or split a span, and a stated total then +disagrees with its own table. The plan counts what it writes. --- @@ -87,8 +86,10 @@ who finds a rule defined here rather than cited has found a defect. one branch alone is already an incomplete pass. **Which severity field each of them reads is settled in Mechanics, Severity**, which is where that split lives and is not repeated here. Beyond the findings, closure reads the **derived floor** and the final-acceptance preconditions -the floor section states, the **evidence entry's revalidation rule** — a changed entry meaning -the clean pass no longer covers what is being committed — and any **hold still standing**; the +the floor section states, and any **hold still standing**; **in a Gate-B cycle it also reads the +evidence entry's revalidation rule** — a changed entry meaning the clean pass no longer covers +what is being committed — which is a Gate-B precondition because the entry is about the diff +being committed and a Gate-A cycle produces none. The scope triggers read the **current assigned fix set** as the absorb paragraph defines it, and the answers already given; the clearly-stuck reading adds its own coverage judgement. **A line in one branch file and a line in the other are distinct findings for holds and answers**, so a @@ -108,9 +109,16 @@ the set as that paragraph reads them, settled before any branch below runs, whic this order executable rather than asserted. A pass with **zero** findings is clean whatever the floor, because a floor buys further looks at an artifact that keeps yielding findings, and one yielding none has already given what those looks were for. **A clean pass at or above the derived floor, or a zero-finding pass, makes -the cycle eligible to close**, under those final-acceptance preconditions; **the closing amend -Mechanics · Finishing the cycle describes is the closure itself**, so nothing is closed before -it, and a profile, cited set or evidence entry that changes in between still gates it. A plateau +the cycle eligible to close**, under those final-acceptance preconditions. **Eligibility is not +closure: the cycle closes when the commit carrying its reviewed artifact is made, and which +commit that is depends on the cycle kind** — for a **Gate-B** cycle the closing amend +Mechanics · Finishing the cycle describes, for a **Gate-A** cycle the commit of the reviewed spec +or plan, that gate having no WIP snapshot and no amend. Nothing is closed before that commit, and +a profile, cited set or — in a Gate-B cycle — evidence entry that changes in between still gates +it. **A commit the hook reads as cycle-closing is a separate matter**: Mechanics warns that a +non-`WIP` commit mid-cycle resets the hook's counters, which is an observation about the counter +and not a closure under these rules — a cycle with an unmet precondition is not closed by being +committed over, and the warning stands as written. A plateau or tells on that pass go into the closing report and never block it, because reporting "will not converge" on a converged loop is a false report. **No other pass outcome makes a cycle eligible to close**, because every other pass leaves a required repair, a hold or a question outstanding, @@ -141,18 +149,10 @@ to run again. it gates closing, discharged by the count of valid logical passes reaching it with the last of them clean, or by the zero-finding exit. The **Blocker/Major-resolve duty** is a **precondition on closure and on any pass being clean**: Mechanics Severity, scoped to the assigned fix set, -states what it demands and what discharges it, and a validated pass finding **no in-set Blocker -or Major at effective severity** is what shows it discharged — the clean predicate's own wording, -so the two cannot drift. It is discharged by a **repair** or by a **validated dismissal**: the -author's judgement, carrying the one-line why this section already requires of a dismissed -finding, that the finding is not true of the artifact. A dismissal does not rewrite the pass that -found it, and the later clean pass is still owed. **A dismissal is not a decline**: a dismissal -says the finding is false, a decline is the user's decision that a **true** finding stays outside -the set, and only the second is an answer at a membership stop. A dismissal the reviewer keeps -re-raising across passes is covered by the clearly-stuck reading's **third condition as that -paragraph states it**, which is where that condition is defined and where it was widened to admit -it — so a false positive that returns every pass reaches that suspension instead of continuing -forever, and this sentence cites the reading rather than extending it. The **hold** a surfaced finding places on closure **participates in the +states what it demands and **what discharges it — a repair or a validated dismissal — and what a +dismissal is**, all at that source; a validated pass finding **no in-set Blocker or Major at +effective severity** is what shows it discharged, the clean predicate's own wording, so the two +cannot drift. The **hold** a surfaced finding places on closure **participates in the ordering**: it gates closing while it stands, and is discharged by the answers that surface requires. **It attaches to every surfaced finding, whichever suspension surfaced it** — clean completion creates none, because it wins before anything is surfaced. **No-clean-credit** — no @@ -167,8 +167,10 @@ single-trigger finding, both for one carrying both. That is **one rule with two answers and which way each may go, and neither is a test the other has to pass. Where a health suspension applies to the same pass, its shared continue-or-stop answer is **additional** to those and not counted among them, the health readings asking about the loop rather than about -this finding — and where the only surface was a health reading, that same answer is what -discharges the hold it created, there being no scope-stop answer for it to be additional to. +this finding. **A health reading surfaces no finding and therefore creates no hold**; what it +leaves outstanding is its own continue-or-stop question, which the composition rule below holds +the cycle on until it is answered. That is why the duties paragraph attaches a hold to every +surfaced *finding* and this one does not give a two-tell surface a hold of its own. At a **membership stop** the answer is **accept**, the finding joining the fix set where Mechanics Severity governs it, or **decline**, the finding staying outside and binding so for the rest of this cycle. A later answer that contradicts a decline **does not reverse it**: @@ -271,8 +273,11 @@ above — the ordering claims no in-session consequence it does not state there. ## 4. The standing sentences edited at their source -Twenty-two edits. Each row names the sentence, where it lives, what changes and why, and whether it -is a replacement or an addition. **No OLD or NEW text and no line numbers appear here**: the plan +The standing sentences this change edits, each row naming the sentence, where it lives, what +changes and why. **No total is claimed and the rows are not numbered as a closed set**: a count +over spans that can be merged or split is bookkeeping that has to be re-derived at every revision, +and three of pass 15's findings were that bookkeeping disagreeing with itself. The plan establishes +the edit set against the real files, which is where a span is decidable. **No OLD or NEW text and no line numbers appear here**: the plan quotes each sentence from the real file, writes the replacement, and re-greps the site, for the reason §7 gives. Working around any of these would ship two instructions that disagree. @@ -280,7 +285,7 @@ reason §7 gives. Working around any of these would ship two instructions that d |---|---|---|---|---| | 1 | the findings-file protocol's clean sentence | the gate-prompt template both gates paste, C and W | it calls a `NO FINDINGS` file a *clean pass*; renamed a **clean findings file**, which is what it always described, so §3's predicate is not read as redefining it | replacement | | 2 | the Gate-A clean-signal sentence | the Gate-A section, C and W | it makes the `NO FINDINGS` signal the only route to clean; it becomes what lets a pass be read as clean without inspecting it, since a pass carrying Minors alone is clean and could never produce that file | replacement | -| 3 | the Severity bullet's resolve duty | Mechanics · Severity, C and W | the only statement of the duty and the only one without a scope; scoped to **the assigned fix set**, the boundary **D5** implied and no sentence carried. It names the set and never a decline — AC 3's second half, since naming one here would read as a waiver of resolution rather than the membership decision it is | replacement | +| 3 | the Severity bullet's resolve duty | Mechanics · Severity, C and W | the only statement of the duty and the only one without a scope; scoped to **the assigned fix set**, the boundary **D5** implied and no sentence carried. It also gains **what discharges the duty — a repair or a validated dismissal — and what a dismissal is** (the author's judgement, carrying the one-line why already required, that the finding is not true of the artifact; it does not rewrite the pass that found it, the later clean pass is still owed, and it is **not** a decline, which is the user's decision that a *true* finding stays outside the set). The block cites this and defines none of it. It names the set and never a decline — AC 3's second half, since naming one here would read as a waiver of resolution rather than the membership decision it is | replacement | | 4 | the lens paragraph's unchanged-list | the profiles section, C and W | it asserts the Blocker/Major filter and the clean-final-pass rule are unchanged, which the ordering falsifies; scoped to **the lens sets**, which is what that paragraph is about | replacement | | 5 | the Gate-A cadence | the Gate-A section, C and W | "validate, revise, re-run" makes a revision unconditional, telling a Minor-only pass to manufacture the repair the severity rule forbids; the revision becomes conditional on a repair being required | replacement | | 6 | `a17`–`a19`, the floor paragraph's closure sentences | passage (a), C and W | they state closure and the early exit in the paragraph that owns the floor; trimmed to point at the ordering, which states them once | replacement | @@ -293,23 +298,26 @@ reason §7 gives. Working around any of these would ship two instructions that d | 13 | `b12` | passage (b) | it resumes on the membership answer alone; it says the answer ends the hold and **defers to the ordering** for what the pass does next | replacement | | 14 | `b13`, the question trigger | passage (b) | *new* is undefined, so an answered question re-raised stops the loop again on every pass; *new* excludes a question already answered in this cycle | replacement | | 15 | `b17`–`b18` | passage (b) | they state their own version of what ends a suspension; they name it a suspension and defer to the ordering | replacement | -| 16 | the clearly-stuck closure sentences | passage (c), from the third condition to the end | they carry evaluation order in the paragraph that owns the reading; the paragraph keeps its reading, and the precedence sentence moves into the block **word for word**, which is what satisfies **D3** | replacement | +| 16 | passage (c), from the third condition to the end | passage (c) | **one contiguous replacement covering both changes to that span**, so nothing is counted twice. The closure sentences carry evaluation order in the paragraph that owns the reading: the paragraph keeps its reading, and the precedence sentence moves into the block **word for word**, which satisfies **D3**. Within the same span the third condition `c8` is **widened at this source rather than from the block**: it requires Blocker/Major findings regenerating "across genuine repair attempts, each round's fix producing the next", which a **validated dismissal** never satisfies — no repair, no fix — so a false positive the reviewer re-raises every pass would leave the cycle unable to close and unable to suspend. It gains that case, the re-raise standing in for the regenerating fix. The accounting records `c8` as **replaced**, not kept | replacement | | 17 | `e7`, the two-tell threshold | passage (e) | unqualified, it and the ordering decide a clean two-tell pass at or above the floor in opposite directions; it is read **after** the clean-completion branch. Authority **D2** | replacement | | 18 | the five-tells pointer | passage (e) | the passage says nothing about what its answer does; one sentence names the stop a suspension and defers to the ordering | addition | | 19 | the handed-over severity question | Mechanics · Severity, passage (g) | the paragraph says the demotion/loop-health question is unsettled and mandates a stop; replaced by the answer, which is the (g) text below | replacement | | 20 | the unknown-start strict-reading list | passage (i) | the list of what a cycle takes at its strictest omits the suspensions this change ships; it gains **every suspension binding, since unknown starting rules cannot waive an open hold** — the reason carried inline, as prompt-standards item 6 requires of the clause itself | addition | -| 21 | `c8`, the third condition itself | passage (c) | a **semantic widening, made at the source rather than from the block**: `c8` requires Blocker/Major findings regenerating "across genuine repair attempts, each round's fix producing the next", which a **validated dismissal** never satisfies — it gets no repair and produces no fix, so a false positive the reviewer re-raises every pass leaves the cycle unable to close and unable to suspend. It gains that case: a finding the author has validly dismissed and the reviewer re-raises across passes counts toward the third condition, the re-raise standing in for the regenerating fix. The old-condition accounting records `c8` as **replaced**, not kept | replacement | -| 22 | the one-contract paragraph's membership list | Mechanics, the "These records are one contract" paragraph, C and W — **outside the ten inventoried passages** | it names the nonce, the slot naming, the provenance line, the curve, the carry rule and the unknown-start activation semantics, and reaches none of this change; it gains **the closure ordering block together with the §4 edits it depends on**, so a project holding some of them, or versions of them that disagree, meets the stop that paragraph already states. **What this is: the same instruction to the agent, over a wider membership.** It is not a checker and builds none; nothing mechanical detects a partial adoption, and a project that does not adopt this sentence is not reached by it. Scope is exactly the coupled set §9 names and nothing else | addition | +| 21 | the one-contract paragraph | Mechanics, the "These records are one contract" paragraph, C and W — **outside the ten inventoried passages** | its coherence stop reaches the nonce, the slots, the provenance line, the curve, the carry rule and the unknown-start semantics, and reaches none of this change. Two changes, one contiguous rewrite of the paragraph's opening and its membership: the opening noun broadens from *records* to **the rules and records a cycle runs under**, since what is added is not a record; and the membership gains **the closure ordering together with every source edit it cites or depends on**, stated as that description and **not as a list of item numbers** — the numbering is this spec's working aid, it is not in the shipped text, and a shipped list would have to be re-derived whenever a span merges. **What this is: the same instruction to the agent, over a wider membership.** Not a checker, and none is built. **The plan establishes the membership against the real files** — every edit the block cites or depends on — which is where dependence is decidable | replacement | **Passage (g)'s replacement is design rather than bookkeeping, so it is stated here in full:** ``` - **The demotion changes what a cycle must resolve, never what it counts.** The per-pass - counts, the finding clusters and the tell thresholds read the severity the reviewer wrote in - the findings file, before the ceiling is applied: a demoted finding still counts in the - finding total and in its cluster, and a Blocker demoted to Minor is still a Blocker to the - curve. Cleanliness and the resolve duty read the effective severity, after the ceiling (the - closure ordering above). Two reasons for the split. The curve must stay derivable from the + **The demotion changes what a cycle must resolve, never what it observes.** **Every + loop-health reading** — the per-pass counts, the finding clusters, the tell thresholds and + the clearly-stuck reading's regenerating-Blocker-or-Major condition — reads the severity the + reviewer wrote in the findings file, before the ceiling is applied: a demoted finding still + counts in the finding total and in its cluster, and a Blocker demoted to Minor is still a + Blocker to the curve and still regeneration to that condition. Cleanliness and the resolve + duty read the effective severity, after the ceiling (the + closure ordering above). The line is **what the cycle owes versus what it observes about + itself**, which is why no list of readings has to be kept complete here. Two reasons for the + split. The curve must stay derivable from the findings files alone — counting finding lines and leading `BLOCKER` and `MAJOR` fields per pass reproduces its three series, which is the only thing that makes a self-reported curve checkable; the subject clusters are a judgement per finding and no count reproduces them. @@ -324,12 +332,12 @@ Two sentences are **not** edited and are named so nobody looks for them: the "Co into the squash body" sentence inside the human-exception block, and the "records every cycle owes" list. This change ships no record. -**Items 21 and 22 were added at Gate-A pass 14.** Item 21 answers a Blocker: the block was -widening `c8` from a paragraph that does not own it. Item 22 answers a Major on Daniel's decision -of 2026-09-10, as a **bounded extension of the existing coherence instruction** — no checker, no -new mechanism, and nothing about record durability, which stays with the successor story. Item 22 -is the only edit in this table that touches a passage no earlier revision named, and §9 states -what it is worth. +**Item 21 was added at Gate-A pass 14 on Daniel's decision**, as a **bounded extension of the +existing coherence instruction** — no checker, no new mechanism, and nothing about record +durability, which stays with the successor story. It is the only edit in this table touching a +passage no earlier revision named, and §9 states what it is worth. The `c8` widening that pass 14 +first carried as a separate row was **merged into item 16 at pass 15**, the two replacing one +contiguous span. --- @@ -345,7 +353,7 @@ here. This table says what happens to each inventoried passage, so the map stays |---|---|---| | (a) the floor paragraphs | edited — closure sentences trimmed to a pointer, the no-restating prohibition scoped, the per-pass fix command pointed at Mechanics | 6, 7, 8 | | (b) what a loop absorbs | edited — owns both triggers and the fix set, so every qualification the ordering needs is made **here**, which is what keeps one definition per rule | 9, 10, 11, 12, 13, 14, 15 | -| (c) recognizing clearly stuck | edited — keeps its three-condition reading, stops carrying evaluation order; its precedence sentence moves into the block unchanged, and the third condition itself is widened at this source to admit a re-raised validated dismissal | 16, 21 | +| (c) recognizing clearly stuck | edited — keeps its three-condition reading, stops carrying evaluation order; its precedence sentence moves into the block unchanged, and the third condition itself is widened at this source to admit a re-raised validated dismissal | 16 | | (d) from pass 4 onward | **no longer edited.** The unavailable-history block moved to the successor with **D10** | — | | (e) the five tells | edited — the threshold is read after clean completion, and a pointer says what its answer does | 17, 18 | | (f) the two rules above do not compete | **unchanged.** "The two rules above" still names the absorb rule and the stuck reading; the block sits before both and adds no third rule between them | — | @@ -398,8 +406,10 @@ showing the old wording gone. Each half runs in **both copies** and against **bo tree and the parent tree, so every assertion is observed passing where the change exists and failing where it does not. A one-sided presence check is not enough: a copy carrying the new wording **and** the old one satisfies it, which is the two-instructions-that-disagree failure §4 -exists to prevent. An edit that only adds to a list — §4 items 18 and 20, and the block itself — -is checked by **presence alone**, because nothing is being replaced. **The plan builds each pair +exists to prevent. **An edit that only adds, replacing no wording, is checked by presence alone**, because there is +no old text whose absence could be counted; which edits those are is read off §4's `Kind` column +rather than listed here, a list of item numbers being the bookkeeping that has to be re-derived +whenever a span merges. **The plan builds each pair against the real files and runs both directions there**, under two constraints: a counted fragment must be **single-line in the file it is grepped from**, since one spanning a line break makes `grep -F` count zero and read as a failure; and the new wording must therefore be @@ -442,10 +452,15 @@ states the table's claim at this width**, not wider. **The oracle.** A row **fails** when its required answer does not produce a **distinct** resumable or closed state — **the same stop returning with its reading unconsumed, that is, without an -intervening validated pass run after the answer** — or when it closes with **any** closure -precondition unmet: an in-set Blocker or Major, a standing hold, an unanswered question, the -derived floor, a changed evidence entry, or a cited-set or profile change during the pass. Naming -only the first would pass the exact no-progress defect AC 4 cites from the parent cycle. **The +intervening validated pass run after the answer** — or when it closes on anything other than the +route the block states. **The closure conditions are read from the block and not re-enumerated +here**: a re-enumeration is a second definition that drifts, and pass 15 found this list already +missing two of them. Concretely the row must **enter closure from the clean-completion or +zero-finding branch** — so a pass carrying a scope-stop trigger cannot close on the answer to that +trigger, no-clean-credit being the clean predicate's own second half — and every precondition the +block names must hold **through the closing commit**, not merely during the pass, which is the +window the block's own "in between" wording fixes. Naming only the distinct-state half would pass +the exact no-progress defect AC 4 cites from the parent cycle. **The consumption clause is what keeps the oracle and §3 in agreement**: §3 says continue consumes the reading that raised the suspension and a further health suspension needs it recomputed over a pass run after the answer, so the *same* two-tell or clearly-stuck result **after** such a pass is @@ -501,7 +516,7 @@ repo's most persistent defect. The transport that could carry it left with the r `plugins/dev-workflow/commands/workflow-init.md` is under `plugins/`, so `plugins/dev-workflow/.claude-plugin/plugin.json` goes `0.11.0 → 0.12.0` with a `plugins/dev-workflow/CHANGELOG.md` entry: a minor bump, the template gaining a closure - ordering and twenty-two edited or extended sentences. + ordering and the edited or extended sentences §4 lists. - **Invariant 4 / the hook.** Untouched: `plugins/dev-workflow/hooks/codex-gate.sh` is not edited, and the §5 heading it greps (`Cross-Model Review`) does not move. @@ -537,17 +552,22 @@ decision of 2026-09-10, with the evidence in §1: record it was named for. **Moved to the plan** — the implementation artifact, not a story: the per-condition disposition -list for all 135 inventoried conditions (§5); the OLD and NEW text of all twenty-two edits with their +list for all 135 inventoried conditions (§5); the OLD and NEW text of every §4 edit with its line ranges (§4); the parity divergence list and the extraction-and-diff (§6); and every verification fragment with its counts (§7). Each is work this change still owes; none of it is work a spec can do correctly, because all four are checked against files the plan edits. -**Partial adoption of the narrowed set — answered by §4 item 22, and what that answer is worth.** -The set is mutually dependent: the ordering's clean predicate needs §4 item 3's boundary, its two -senses of *clean* need items 1 and 2, and its triggers need items 10 through 14. **The coupled set -is exactly items 1, 2, 3, 10, 11, 12, 13, 14 and the block**, and item 22 names that set in the -live one-contract paragraph, which until now reached only the nonce, the slots, the provenance -line, the curve, the carry rule and the unknown-start semantics. +**Partial adoption of the narrowed set — answered by §4 item 21, and what that answer is worth.** +The set is mutually dependent, and **the rule that says which edits belong is stated rather than +enumerated**: an edit is coupled when the block **cites it or depends on it** — the boundary its +clean predicate reads, the sentences that give *clean* its two senses, the triggers it reads, the +source rules whose old text the ordering falsifies, and the severity and dismissal rules it cites. +Item 22 carries that description into the live one-contract paragraph, which until now reached +only the nonce, the slots, the provenance line, the curve, the carry rule and the unknown-start +semantics. **An earlier revision of this section named the members as a list of item numbers and +the list was wrong** — it omitted several edits the block plainly depends on — which is why the +rule is stated and **the plan derives the membership against the real files**, where dependence is +decidable and the numbering does not exist. **Stated as what it is.** That paragraph is **an instruction to the agent**: a project whose text carries some members and not others, or versions that disagree, **stops and has a human complete, @@ -557,8 +577,8 @@ sentence is about. It is not a guard and not a mechanical check, and this change sentence and acts on it. **What it therefore does not buy, said rather than implied.** A project that adopts the block -without adopting item 22 is not reached at all, which is the partial-adoption case applied to the -rule against partial adoption; the existing paragraph has the same property and item 22 neither +without adopting item 21 is not reached at all, which is the partial-adoption case applied to the +rule against partial adoption; the existing paragraph has the same property and item 21 neither worsens nor repairs it. Nor does it detect a *silent* half-merge in a project that did adopt it — it obliges a stop once someone notices, which is a different thing from noticing. **What it removes is the narrower state this spec previously admitted**: that the coherence rule did not @@ -566,7 +586,7 @@ name this material at all, so an agent reading the sentence and willing to act o to act on. Two earlier revisions got this wrong in opposite directions — one claimed the existing rule already caught the block, and one assigned the guard to the successor story, whose scope is record durability and **excludes the closure ordering by name**. Both named a mechanism that did -not exist; item 22 names a sentence that does, and claims only what that sentence does. +not exist; item 21 names a sentence that does, and claims only what that sentence does. **Out of scope and parked**, unchanged: From 5d12884e008c59fda818accb70ccbc075b32deb9 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Fri, 11 Sep 2026 12:24:33 +0200 Subject: [PATCH 025/181] docs(specs): the two health readings hold differently; fix stale item refs The pass-15 rollback over-generalised. It wrote "a health reading surfaces no finding and therefore creates no hold", which is true of the two-tell stop and false of the clearly-stuck reading -- the block itself says at the composition paragraph that the clearly-stuck reading surfaces findings, and the duties paragraph attaches a hold to every surfaced finding. Three sentences in one block disagreed, and it decides which answer lifts a suspension rather than being a wording question. Corrected within the existing rules, no new mechanism: - the two-tell stop surfaces tells and no finding, so it creates no hold; what it leaves outstanding is its own continue-or-stop question, and the composition rule holds the cycle on that; - the clearly-stuck reading does surface findings, and each takes a hold like any other surfaced finding, discharged by every answer its own surface requires -- the scope-stop answers where the finding also carries a trigger, with continue-or-stop additional there, and where it carries neither trigger that continue-or-stop answer is the only one its surface asks for and is what discharges the hold. So no surfaced finding lacks a discharging answer and no finding-less surface gets a hold nothing could discharge. Also: two stale "Item 22" references in section 9, left when that row became item 21. Corrected here rather than in a round of their own. Worth recording why they survived the last commit's check: that grep was case-sensitive and the occurrences are capitalised, so the "0 occurrences" claim was produced by a check that could not see them. The claim was wrong, not the cleanup. No Minor or Nit triggered a revision round. Verified: precheck exit 0; case-insensitive grep for the item-number and total claims returns 0; the three hold sentences now read consistently. Spec 598 -> 604 lines. No gate: docs/**.md plus the cycle's working record, prose-exempt under section 5. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- ...026-09-10-loop-rule-consolidation-design.md | 18 ++++++++++++------ 1 file changed, 12 insertions(+), 6 deletions(-) diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index 6ff28b3..259e436 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -167,10 +167,16 @@ single-trigger finding, both for one carrying both. That is **one rule with two answers and which way each may go, and neither is a test the other has to pass. Where a health suspension applies to the same pass, its shared continue-or-stop answer is **additional** to those and not counted among them, the health readings asking about the loop rather than about -this finding. **A health reading surfaces no finding and therefore creates no hold**; what it -leaves outstanding is its own continue-or-stop question, which the composition rule below holds -the cycle on until it is answered. That is why the duties paragraph attaches a hold to every -surfaced *finding* and this one does not give a two-tell surface a hold of its own. +this finding. **The two health readings differ in what they surface, and therefore in what they +hold.** The **two-tell stop surfaces tells and no finding**, so it creates no hold; what it leaves +outstanding is its own continue-or-stop question, which the composition rule below holds the cycle +on until it is answered. The **clearly-stuck reading does surface findings** — those its third +condition is about — and each takes a hold like any other surfaced finding, discharged by **every +answer its own surface requires**: the scope-stop answers where that finding also carries a +trigger, the continue-or-stop answer being additional there; and where it carries neither trigger, +that continue-or-stop answer is the only answer its surface asks for and is what discharges the +hold. So no surfaced finding is left without a discharging answer, and no surface without a +finding is given a hold nothing could discharge. At a **membership stop** the answer is **accept**, the finding joining the fix set where Mechanics Severity governs it, or **decline**, the finding staying outside and binding so for the rest of this cycle. A later answer that contradicts a decline **does not reverse it**: @@ -562,7 +568,7 @@ The set is mutually dependent, and **the rule that says which edits belong is st enumerated**: an edit is coupled when the block **cites it or depends on it** — the boundary its clean predicate reads, the sentences that give *clean* its two senses, the triggers it reads, the source rules whose old text the ordering falsifies, and the severity and dismissal rules it cites. -Item 22 carries that description into the live one-contract paragraph, which until now reached +Item 21 carries that description into the live one-contract paragraph, which until now reached only the nonce, the slots, the provenance line, the curve, the carry rule and the unknown-start semantics. **An earlier revision of this section named the members as a list of item numbers and the list was wrong** — it omitted several edits the block plainly depends on — which is why the @@ -571,7 +577,7 @@ decidable and the numbering does not exist. **Stated as what it is.** That paragraph is **an instruction to the agent**: a project whose text carries some members and not others, or versions that disagree, **stops and has a human complete, -revert or reconcile the adoption before running a gate under it**. Item 22 widens whom that +revert or reconcile the adoption before running a gate under it**. Item 21 widens whom that sentence is about. It is not a guard and not a mechanical check, and this change builds neither — **nothing detects a partial adoption**, and the stop happens only where an agent reads the sentence and acts on it. From c792383074e8ee376e8eefabbe0a50bfaf5feffb Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Fri, 11 Sep 2026 12:41:59 +0200 Subject: [PATCH 026/181] docs(specs): apply Gate-A pass 16; two disagreements with standing section 5 text Pass 16: 8 findings (1 Blocker, 4 Major, 2 Minor, 1 Nit). Blocker+Major 8 -> 5, the lowest of the cycle. One tell of five (Blockers flat at 1), so no mandatory stop; the clearly-stuck exit is unreachable with the curve at a cycle low. Two findings are a new shape for this cycle: the spec disagreeing with STANDING section 5 text rather than with itself. Both verified against the source before acting. - The Mechanics WIP-naming warning says a snapshot named anything else "closes the cycle". That is an event-derived closure path beside the ordering's content-and-precondition-derived one, and section 4 did not name it. New row 21 edits it: the hook reads such a commit as closing and discards the counted passes, while the cycle stays open until the ordering's conditions hold. The warning keeps its force; the cost of the mistake is the lost pass credit. - The (g) replacement said health readings take "the severity the reviewer wrote", but the standing Reader paragraph case-folds and reads any non-empty unrecognized token as MAJOR. A finding written IMPORTANT would have counted for the pass and vanished from the curve. Health readings now take the reader-normalized pre-ceiling severity, and the same rewrite fixes the Minor beside it: totals count finding lines and clusters use no severity at all. The Blocker: cleanliness of a later pass was being read as proof that an earlier pass's resolve duty was discharged, which contradicts the settled fact that findings files establish inventory and not resolutions -- an omitted finding would have closed a cycle. The duty is now discharged per finding and tracked across the cycle, and closure reads two things: the final pass is clean, and no in-set Blocker or Major raised anywhere in this cycle is still undischarged. Also decided: a finding surfaced by both a membership stop and the clearly-stuck reading carries two hold components, each ended by its own answer, so decline still releases the membership hold exactly as D5 requires while the health answer ends the other; resumption remains the composition rule's. And a contradictory answer to a decline changes neither membership, cycle state nor any outstanding question, and is never itself a resumption. Minor 7 (whether the two-branch concatenation strips terminators and NO FINDINGS lines) is collected and stands open. Nit 8 was fixed inside a clause already being rewritten, not as a round of its own: the 58 percent is correct for the four bookkeeping sections (456 of 785) and the word "remaining" was wrong. Verified: precheck exit 0; section 4 numbering contiguous 1-22; case-insensitive grep for stale item references and the old percent wording returns 0. The new row cost no count update, which is what removing the totals bought. Spec 604 -> 625. No gate: docs/**.md plus the cycle's working record, prose-exempt under section 5. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-16.md | 9 ++ .../gate-a-spec-awsf1ec771-resume.md | 37 +++++++- ...26-09-10-loop-rule-consolidation-design.md | 87 ++++++++++++------- 3 files changed, 99 insertions(+), 34 deletions(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-16.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-16.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-16.md new file mode 100644 index 0000000..dd9d88b --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-16.md @@ -0,0 +1,9 @@ +BLOCKER | high | §3 "The four standing duties, classified" | The block first defines cleanliness solely from the current logical pass, then says a current pass with no in-set Blocker or Major "shows" every earlier resolve duty discharged; that contradicts §2's settled fact that findings files establish inventory rather than resolutions, because a reviewer can simply omit a still-unrepaired, undismissed prior finding | A later pass can be classified clean and close the cycle while a required Blocker or Major repair remains open, defeating the duty that is supposed to gate both cleanliness and closure | Keep pass cleanliness pass-local, track the outstanding repair-or-validated-dismissal duty as a separate closure precondition, and remove the claim that absence from a later findings file proves discharge +MAJOR | high | §3 "First, clean completion" and §4 edit list | The block says a non-WIP commit merely resets the hook's counters and is not closure under the rules, but the standing Mechanics sentence at CLAUDE.md:825–826 and workflow-init.md:1009–1010 says a misnamed pre-review snapshot "closes the cycle"; §4 does not name that sentence for replacement even though the block says the warning stands as written | The shipped copies retain a second event-derived closure path that directly contradicts the new content-and-precondition-derived closure rule, so an agent can treat an accidental commit as closing over open duties | Add this Mechanics sentence in both copies to §4 and change it to say the hook reads the commit as closing and discards its counters while the rule-governed cycle remains open until the block's closure conditions hold +MAJOR | high | §3 "What a suspension asks, and what ends it" and §4 item 13 | For a clearly-stuck finding that also raises a membership stop, the new text makes one finding hold persist until both the membership answer and the shared health answer are given, while settled D4 and D5 say any answer ends a surfaced finding's hold and specifically that decline releases it; item 13 independently promises that the membership answer "ends the hold" | The two copies can be implemented either with a decline-released hold plus a separate health suspension or with one combined hold that decline does not release, so the duty classification and next-state oracle disagree on the same dual-suspension path | Preserve the decided clearly-stuck hold by modelling the membership and clearly-stuck hold components separately: decline ends the membership hold immediately, the health answer ends the clearly-stuck hold, and composition still prevents resumption until every required answer is present +MAJOR | high | §4 item 19 "Passage (g)'s replacement" | The proposed rule says health readings use "the severity the reviewer wrote" and that counting leading BLOCKER and MAJOR fields reproduces the curve, but the standing Reader rule at CLAUDE.md:488–497 and workflow-init.md:680–689 accepts case variants and normalizes every other non-empty severity token to MAJOR | A valid finding written as `IMPORTANT` or a case variant can count as Major under the pass reader but disappear from the promised reproducible health curve, changing tell thresholds and clearly-stuck decisions between readers | State that severity-sensitive health readings use the existing reader-normalized severity before the ceiling, and make the derivability sentence count normalized fields rather than only literal reviewer-written tokens +MAJOR | high | §3 "What a suspension asks, and what ends it" | The re-raised-decline rule says an attempted reversal leaves the decline binding but "the cycle continues", while the later composition and parked-state rules require an explicit health `continue` and every other outstanding answer before another pass; on a below-floor clearly-stuck re-raise, an attempted `accept` is neither `continue` nor an answer to any other pending question | The attempted reversal can be read as resuming the loop and bypassing the clearly-stuck hold or another concurrent suspension, contrary to the composition rule | Say the contradictory answer is surfaced but does not change membership, the current cycle state, or any outstanding question; then defer all resumption to the composition rule +MINOR | high | §4 item 19 "Passage (g)'s replacement" | The universal claim that every loop-health reading "reads the severity" is false for the finding total and subject clusters, which the same replacement later derives respectively from all finding lines and from a judgement per finding | The source paragraph gives two incompatible computation models, inviting the plan or a later reader to severity-filter totals or clusters and hide the demoted findings the rule intends to retain | Say every health reading observes the reviewer-produced findings before the ceiling, and only readings that use severity consume the reader-normalized pre-ceiling severity +MINOR | high | §3 "What a pass is read from" | Reading two validated Gate-B findings files "as their concatenation" does not say whether protocol-only lines are removed; literal concatenation includes two terminators and can include `NO FINDINGS` from one branch beside actual findings from the other, even though those lines are not findings | A full pass with one empty branch and one non-empty branch can be treated as structurally invalid or have inconsistent counts despite both source files validating | Define the logical concatenation as the concatenation of extracted finding lines after per-file validation, treating a branch's `NO FINDINGS` body as an empty finding sequence and excluding each terminator +NIT | high | §1 narrowing rationale | The spec says the former 785-line version contained 173 lines of design and calls the remainder 58 percent, but 785 minus 173 is 612 lines, or 78 percent | The mechanical accounting used to justify moving bookkeeping to the plan is numerically false and fails the requested stated-count check | Correct 58 percent to 78 percent or state which narrower 58-percent subset is being measured +END OF FINDINGS (8 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index 9b8ecb6..b399967 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -40,7 +40,42 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 13 | 78e5f97 | 16 | 2 | 8 | yes | **the loop turned: B+M 14 → 10, Majors 13 → 8**; session 01a08c98-b6cd-7bc3-9556-8f1c87491031 | | 14 | c06dd69 | 16→**10** | 2→**1** | 8→**5** | yes | spec byte-identical to 0168f88 (532 lines); **lowest B+M of the cycle, 6**; **SCOPE STOP surfaced on finding 6** (partial-adoption guard = a 21st edit, outside the twenty); session 01a08f70-6d1e-7cf1-baa3-6df3cc59bc32 | | 15 | 17cf2e9 | 10→**12** | 1→**1** | 5→**7** | yes | **TWO-TELL STOP — mandatory, surfaced to Daniel.** B+M 6→8. **Six of twelve regenerate from pass 14's own repairs** (3, 4, 6, 7, 10, 11). All 12 held open; session 01a08fb2-6e8e-73d0-9d2b-8fb92d1571dd | -| 16 | — | — | — | — | not run | blocked on the two-tell answer | +| 16 | 5d12884 | 12→**8** | 1→**1** | 7→**4** | yes | **B+M 8→5, lowest of the cycle.** One tell (Blockers flat). Five B/M applied; Minors 7 and Nit 8 collected. Findings 2 and 4 are the spec disagreeing with *standing* §5 text, not with itself; session 01a08fff-fadf-72f2-994d-618048bea473 | +| 17 | — | — | — | — | not run | next action | + +## Pass-16 three-line report + +**Floor line:** derived floor **3**; risk **high**, security **none**; read fresh from +`docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md`. One cited story, level 2. + +- **Trend:** findings …, 16, 10, 12, **8**. Blockers …, 2, 1, 1, **1**. Majors …, 8, 5, 7, **4**. + Blocker+Major 20, 15, 10, 11, 15, 13, 9, 8, 12, 17, 14, 14, 10, 6, 8, **5** — **the lowest of the + cycle**, past the pass-14 low of 6. +- **Cluster (pass 16):** product 7 of 8; bookkeeping 1 (the 58% claim); **the instrument 0**, the + first zero-instrument pass since pass 9. +- **require↔withdraw:** none. Finding 3 objects to the hold split made *for* pass 15's finding 11, + which is a later pass questioning a repair — the mirror of the pair's shape, not the shape. + +**Tells: one of five** — the Blocker count flat at 1. Findings fell hard, Majors fell, the cluster +is product. One is not two, so no mandatory stop. The clearly-stuck exit is not reachable either: +its first condition needs a plateau and the curve is at a cycle low. + +**What changed in the loop's shape, and it is worth naming.** Findings 2 and 4 are the first in +this cycle where the spec disagrees with **standing §5 text rather than with itself** — the +Mechanics `WIP:` warning claiming a non-WIP commit "closes the cycle", and the Reader paragraph +normalizing severity tokens the (g) replacement assumed were raw. Both were invisible while the +spec was still contradicting itself. That is the boundary work pass 13 predicted, arriving two +passes after the rollback made room for it. + +**The rollback is holding.** No finding asked for a restored enumeration, no finding hit a stated +total, and the one new source edit (item 21) cost no count update — which is what removing the +totals bought. + +**Minor 7 and Nit 8 are collected, not repaired as their own round.** Nit 8's number was correct +and its wording was not: 58% measures the four bookkeeping sections (456 of 785), not "the +remaining" after the design's 173. Fixed in the same clause while the surrounding sentence was +already open; no revision round was spent on it. Minor 7 — whether the two-branch concatenation +strips terminators and `NO FINDINGS` lines — stands open and is recorded here. ## Pass-15 three-line report — MANDATORY TWO-TELL STOP diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index 259e436..fb91e31 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -14,8 +14,9 @@ take it. **Narrowed twice on 2026-09-10, and §9 lists what moved.** First the record-durability subject went to a successor story. Then the bookkeeping went to the plan: this spec was 785 lines of -which the design was 173, and the remaining 58% — quoted OLD and NEW text, per-condition -dispositions, the parity divergence list, the verification substrings — took roughly half the +which the design was 173, and **58% of it sat in four bookkeeping sections** — quoted OLD and NEW +text, per-condition dispositions, the parity divergence list, the verification substrings, 456 +lines between them — which took roughly half the findings of every Gate-A pass while describing work the plan performs against real files. **It is moved, not dropped.** The plan is the carrier and each section below names what it owes. @@ -115,10 +116,12 @@ commit that is depends on the cycle kind** — for a **Gate-B** cycle the closin Mechanics · Finishing the cycle describes, for a **Gate-A** cycle the commit of the reviewed spec or plan, that gate having no WIP snapshot and no amend. Nothing is closed before that commit, and a profile, cited set or — in a Gate-B cycle — evidence entry that changes in between still gates -it. **A commit the hook reads as cycle-closing is a separate matter**: Mechanics warns that a -non-`WIP` commit mid-cycle resets the hook's counters, which is an observation about the counter -and not a closure under these rules — a cycle with an unmet precondition is not closed by being -committed over, and the warning stands as written. A plateau +it. **A commit the hook reads as cycle-closing is a separate matter**: a non-`WIP` commit mid-cycle +makes the hook read the cycle as closed and discards the passes counted so far, which is an +observation about the counter — **the cycle itself stays open until the conditions above hold**, +so an accidental commit destroys the pass credit and closes nothing. The Mechanics sentence that +says such a snapshot "closes the cycle" is edited to say that, since left as written it is a +second, event-derived closure path beside this one. A plateau or tells on that pass go into the closing report and never block it, because reporting "will not converge" on a converged loop is a false report. **No other pass outcome makes a cycle eligible to close**, because every other pass leaves a required repair, a hold or a question outstanding, @@ -150,9 +153,14 @@ it gates closing, discharged by the count of valid logical passes reaching it wi them clean, or by the zero-finding exit. The **Blocker/Major-resolve duty** is a **precondition on closure and on any pass being clean**: Mechanics Severity, scoped to the assigned fix set, states what it demands and **what discharges it — a repair or a validated dismissal — and what a -dismissal is**, all at that source; a validated pass finding **no in-set Blocker or Major at -effective severity** is what shows it discharged, the clean predicate's own wording, so the two -cannot drift. The **hold** a surfaced finding places on closure **participates in the +dismissal is**, all at that source. **It is discharged per finding and tracked across the cycle, +never inferred from a later pass.** A findings file establishes the **inventory** of what that +pass found and not the resolution of anything, so a later pass that does not mention an earlier +in-set Blocker or Major says nothing about whether it was repaired or dismissed; reading its +absence as discharge would let an omission close a cycle. **Closure therefore reads two things, +not one**: the final pass is clean, *and* no in-set Blocker or Major raised anywhere in this cycle +is still undischarged. A pass's cleanliness stays a fact about that pass's own findings, which is +what §5 already says it is. The **hold** a surfaced finding places on closure **participates in the ordering**: it gates closing while it stands, and is discharged by the answers that surface requires. **It attaches to every surfaced finding, whichever suspension surfaced it** — clean completion creates none, because it wins before anything is surfaced. **No-clean-credit** — no @@ -171,17 +179,25 @@ this finding. **The two health readings differ in what they surface, and therefo hold.** The **two-tell stop surfaces tells and no finding**, so it creates no hold; what it leaves outstanding is its own continue-or-stop question, which the composition rule below holds the cycle on until it is answered. The **clearly-stuck reading does surface findings** — those its third -condition is about — and each takes a hold like any other surfaced finding, discharged by **every -answer its own surface requires**: the scope-stop answers where that finding also carries a -trigger, the continue-or-stop answer being additional there; and where it carries neither trigger, -that continue-or-stop answer is the only answer its surface asks for and is what discharges the -hold. So no surfaced finding is left without a discharging answer, and no surface without a -finding is given a hold nothing could discharge. +condition is about — and each takes a hold like any other surfaced finding. **Where one finding is +surfaced by both, it carries two hold components and each is discharged by its own answer, which +is what keeps D4 and D5 exact**: the **membership** component ends on the membership answer in +either direction, a decline releasing it as **D5** requires and as the absorb paragraph's own +sentence says; the **clearly-stuck** component ends on the reading's continue-or-stop answer. +Neither answer discharges the other's component, and where the finding carries no scope-stop +trigger the continue-or-stop answer is the only one its surface asks for and discharges the only +component there is. **Resumption is still the composition rule's**, which waits for every +outstanding answer — so separating the components changes what each answer discharges and never +lets one answer resume the loop alone. So no surfaced finding is left without a discharging +answer, and no surface without a finding is given a hold nothing could discharge. At a **membership stop** the answer is **accept**, the finding joining the fix set where Mechanics Severity governs it, or **decline**, the finding staying outside and binding so for the rest of this cycle. A later answer that contradicts a decline **does not reverse it**: -the decline **remains binding**, the contradiction is **surfaced to the user as information**, -and the **cycle continues** — nothing here turns one answer into another, since that would let a +the decline **remains binding** and the contradiction is **surfaced to the user as information**, +changing neither membership, nor the cycle's state, nor any outstanding question — **whether the +loop resumes is decided by the composition rule below and by nothing here**, so a contradictory +answer is never itself a resumption and cannot step past a hold or a health question still +awaiting its own answer. Nothing here turns one answer into another, since that would let a finding be moved out of the set and back into it to escape what it owes inside it. **There is no withdrawal inside the cycle that declined**, because **D7** binds a decline for the remainder of its cycle and admits no exception; a reconsideration is a later cycle's, where D7 gives the @@ -309,24 +325,29 @@ reason §7 gives. Working around any of these would ship two instructions that d | 18 | the five-tells pointer | passage (e) | the passage says nothing about what its answer does; one sentence names the stop a suspension and defers to the ordering | addition | | 19 | the handed-over severity question | Mechanics · Severity, passage (g) | the paragraph says the demotion/loop-health question is unsettled and mandates a stop; replaced by the answer, which is the (g) text below | replacement | | 20 | the unknown-start strict-reading list | passage (i) | the list of what a cycle takes at its strictest omits the suspensions this change ships; it gains **every suspension binding, since unknown starting rules cannot waive an open hold** — the reason carried inline, as prompt-standards item 6 requires of the clause itself | addition | -| 21 | the one-contract paragraph | Mechanics, the "These records are one contract" paragraph, C and W — **outside the ten inventoried passages** | its coherence stop reaches the nonce, the slots, the provenance line, the curve, the carry rule and the unknown-start semantics, and reaches none of this change. Two changes, one contiguous rewrite of the paragraph's opening and its membership: the opening noun broadens from *records* to **the rules and records a cycle runs under**, since what is added is not a record; and the membership gains **the closure ordering together with every source edit it cites or depends on**, stated as that description and **not as a list of item numbers** — the numbering is this spec's working aid, it is not in the shipped text, and a shipped list would have to be re-derived whenever a span merges. **What this is: the same instruction to the agent, over a wider membership.** Not a checker, and none is built. **The plan establishes the membership against the real files** — every edit the block cites or depends on — which is where dependence is decidable | replacement | +| 21 | the `WIP:`-naming warning's closure claim | Mechanics · `baseSha`, C and W — **outside the ten inventoried passages** | "A pre-review snapshot named anything else reads as a real commit and **closes the cycle**, discarding the passes you just accumulated" is an **event-derived closure path** beside the ordering's content-and-precondition-derived one, and the two decide an accidental commit in opposite directions. It becomes: the hook reads such a commit as closing and discards the counted passes, **while the cycle itself stays open until the ordering's closure conditions hold**. The warning keeps its force — the cost of the mistake is the lost pass credit — and stops claiming a closure the rules do not grant | replacement | +| 22 | the one-contract paragraph | Mechanics, the "These records are one contract" paragraph, C and W — **outside the ten inventoried passages** | its coherence stop reaches the nonce, the slots, the provenance line, the curve, the carry rule and the unknown-start semantics, and reaches none of this change. Two changes, one contiguous rewrite of the paragraph's opening and its membership: the opening noun broadens from *records* to **the rules and records a cycle runs under**, since what is added is not a record; and the membership gains **the closure ordering together with every source edit it cites or depends on**, stated as that description and **not as a list of item numbers** — the numbering is this spec's working aid, it is not in the shipped text, and a shipped list would have to be re-derived whenever a span merges. **What this is: the same instruction to the agent, over a wider membership.** Not a checker, and none is built. **The plan establishes the membership against the real files** — every edit the block cites or depends on — which is where dependence is decidable | replacement | **Passage (g)'s replacement is design rather than bookkeeping, so it is stated here in full:** ``` **The demotion changes what a cycle must resolve, never what it observes.** **Every - loop-health reading** — the per-pass counts, the finding clusters, the tell thresholds and - the clearly-stuck reading's regenerating-Blocker-or-Major condition — reads the severity the - reviewer wrote in the findings file, before the ceiling is applied: a demoted finding still - counts in the finding total and in its cluster, and a Blocker demoted to Minor is still a - Blocker to the curve and still regeneration to that condition. Cleanliness and the resolve + loop-health reading observes the findings as the reviewer produced them, before the ceiling + is applied** — so a demoted finding still counts in the finding total and in its cluster, and + a Blocker demoted to Minor is still a Blocker to the curve and still regeneration to the + clearly-stuck reading's third condition. Where such a reading uses severity at all it takes + the **reader-normalized pre-ceiling severity**, which is the severity the Reader paragraph + above already produces from the findings file — case-folded, and a non-empty unrecognized + token read as `MAJOR` — never the raw token, so a finding written `IMPORTANT` counts for the + curve exactly as it counts for the pass. Cleanliness and the resolve duty read the effective severity, after the ceiling (the closure ordering above). The line is **what the cycle owes versus what it observes about itself**, which is why no list of readings has to be kept complete here. Two reasons for the split. The curve must stay derivable from the - findings files alone — counting finding lines and leading `BLOCKER` and `MAJOR` fields per - pass reproduces its three series, which is the only thing that makes a self-reported curve - checkable; the subject clusters are a judgement per finding and no count reproduces them. + findings files alone — the finding total counts finding lines and the Blocker and Major + series count the lines whose normalized severity is each, which is the only thing that makes + a self-reported curve checkable; the subject clusters use no severity at all, being a + judgement per finding that no count reproduces. And the demotion is the author's judgement about **the finding's repair severity**, never about which findings the fix set contains — a separate predicate the absorb paragraph defines, and one this must not be read as touching; a loop spending passes on findings the author keeps @@ -338,7 +359,7 @@ Two sentences are **not** edited and are named so nobody looks for them: the "Co into the squash body" sentence inside the human-exception block, and the "records every cycle owes" list. This change ships no record. -**Item 21 was added at Gate-A pass 14 on Daniel's decision**, as a **bounded extension of the +**Item 22 was added at Gate-A pass 14 on Daniel's decision**, as a **bounded extension of the existing coherence instruction** — no checker, no new mechanism, and nothing about record durability, which stays with the successor story. It is the only edit in this table touching a passage no earlier revision named, and §9 states what it is worth. The `c8` widening that pass 14 @@ -563,12 +584,12 @@ line ranges (§4); the parity divergence list and the extraction-and-diff (§6); verification fragment with its counts (§7). Each is work this change still owes; none of it is work a spec can do correctly, because all four are checked against files the plan edits. -**Partial adoption of the narrowed set — answered by §4 item 21, and what that answer is worth.** +**Partial adoption of the narrowed set — answered by §4 item 22, and what that answer is worth.** The set is mutually dependent, and **the rule that says which edits belong is stated rather than enumerated**: an edit is coupled when the block **cites it or depends on it** — the boundary its clean predicate reads, the sentences that give *clean* its two senses, the triggers it reads, the source rules whose old text the ordering falsifies, and the severity and dismissal rules it cites. -Item 21 carries that description into the live one-contract paragraph, which until now reached +Item 22 carries that description into the live one-contract paragraph, which until now reached only the nonce, the slots, the provenance line, the curve, the carry rule and the unknown-start semantics. **An earlier revision of this section named the members as a list of item numbers and the list was wrong** — it omitted several edits the block plainly depends on — which is why the @@ -577,14 +598,14 @@ decidable and the numbering does not exist. **Stated as what it is.** That paragraph is **an instruction to the agent**: a project whose text carries some members and not others, or versions that disagree, **stops and has a human complete, -revert or reconcile the adoption before running a gate under it**. Item 21 widens whom that +revert or reconcile the adoption before running a gate under it**. Item 22 widens whom that sentence is about. It is not a guard and not a mechanical check, and this change builds neither — **nothing detects a partial adoption**, and the stop happens only where an agent reads the sentence and acts on it. **What it therefore does not buy, said rather than implied.** A project that adopts the block -without adopting item 21 is not reached at all, which is the partial-adoption case applied to the -rule against partial adoption; the existing paragraph has the same property and item 21 neither +without adopting item 22 is not reached at all, which is the partial-adoption case applied to the +rule against partial adoption; the existing paragraph has the same property and item 22 neither worsens nor repairs it. Nor does it detect a *silent* half-merge in a project that did adopt it — it obliges a stop once someone notices, which is a different thing from noticing. **What it removes is the narrower state this spec previously admitted**: that the coherence rule did not @@ -592,7 +613,7 @@ name this material at all, so an agent reading the sentence and willing to act o to act on. Two earlier revisions got this wrong in opposite directions — one claimed the existing rule already caught the block, and one assigned the guard to the successor story, whose scope is record durability and **excludes the closure ordering by name**. Both named a mechanism that did -not exist; item 21 names a sentence that does, and claims only what that sentence does. +not exist; item 22 names a sentence that does, and claims only what that sentence does. **Out of scope and parked**, unchanged: From c2d5fe1c0a70340c46f1b78a6746f831ea68b5b9 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Fri, 11 Sep 2026 12:55:07 +0200 Subject: [PATCH 027/181] docs(specs): separate the cycle-wide resolve duty from per-pass cleanliness The pass-16 Blocker repair was incomplete and shipped a contradiction. The duties paragraph still called the resolve duty "a precondition on closure and on any pass being clean", while the sentences added right after it made cleanliness a fact about the current pass alone. The two decide one concrete case in opposite directions: pass 1 raises an in-set Major that is neither repaired nor dismissed, pass 2 finds nothing. The old wording makes pass 2 unclean; the new decision makes it clean but forbids closure. Corrected in the decided direction, with no new mechanism and no record work. The duty is a precondition on CLOSURE only. It is not a second test on whether a pass is clean, and the shipped text now walks that exact case: pass 2 is clean, the cycle is eligible if pass 2 is at or above the floor, and the cycle still cannot close while the Major is undischarged, because closure reads both things. The zero-finding route changes nothing there -- it is a route to eligibility, and eligibility reads the duty like every other final-acceptance precondition. What the open Major does not do is make pass 2 unclean. Row 21 also records that the sibling non-WIP sentence in the profile-change paragraph needs no edit: it already says such a commit "reads to the hook as the cycle closing", which claims no closure. Recorded so the next pass does not re-raise it. TWO CLAIMS FROM THE PASS-16 REPORT WITHDRAWN, both in the cycle record. - That the WIP warning was newly visible. It was raised at pass 15, finding 2, which named both sites. The pass-15 revision changed the block's prose and did not add the section 4 row, so pass 16's finding 2 is a regeneration from an incomplete repair of mine, not boundary work arriving. Finding 4 is genuinely first-seen; finding 2 is not. - That the spec no longer contradicts itself. The pass-16 repair itself shipped the contradiction this commit fixes. The record also now states what the falling curve does not show: every count describes the revision the pass reviewed, never the repair made after it. Pass 16's five Blocker/Majors do not assess their own fix, and the contradiction that fix introduced was found by reading rather than by the curve. Verified: precheck exit 0; zero occurrences of the old coupling; the worked case is present in the shipped block. Spec 625 -> 634 lines. No gate: docs/**.md plus the cycle's working record, prose-exempt under section 5. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-resume.md | 32 +++++++++++++++---- ...26-09-10-loop-rule-consolidation-design.md | 17 +++++++--- 2 files changed, 39 insertions(+), 10 deletions(-) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index b399967..20fe287 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -60,17 +60,37 @@ Nothing depends on it; the pass files and the repo are authoritative where this is product. One is not two, so no mandatory stop. The clearly-stuck exit is not reachable either: its first condition needs a plateau and the curve is at a cycle low. -**What changed in the loop's shape, and it is worth naming.** Findings 2 and 4 are the first in -this cycle where the spec disagrees with **standing §5 text rather than with itself** — the -Mechanics `WIP:` warning claiming a non-WIP commit "closes the cycle", and the Reader paragraph -normalizing severity tokens the (g) replacement assumed were raw. Both were invisible while the -spec was still contradicting itself. That is the boundary work pass 13 predicted, arriving two -passes after the rollback made room for it. +**What changed in the loop's shape — and one claim here was withdrawn the same day.** Findings 2 +and 4 disagree with **standing §5 text rather than with the spec itself**: the Mechanics `WIP:` +warning claiming a non-WIP commit "closes the cycle", and the Reader paragraph normalizing +severity tokens the (g) replacement assumed were raw. + +**Withdrawn: that these were newly visible.** The `WIP:` warning was raised at **pass 15, finding +2**, which named both sites (C 751–753 and 825–826). The pass-15 revision changed the block's +prose and **did not add the §4 row**, so pass 16's finding 2 is a **regeneration from an +incomplete repair of mine**, not boundary work arriving. Pass 16's row 21 closes it, and records +that the sibling passage needs no edit because it already says "reads **to the hook** as the cycle +closing". Finding 4 is genuinely first-seen; finding 2 is not, and saying so in the pass report +was wrong. + +**Also withdrawn: "the spec no longer contradicts itself".** Pass 16's own Blocker repair shipped +a contradiction — the duties paragraph kept the resolve duty as a precondition "on any pass being +clean" while the sentences after it made cleanliness pass-local. The two decide the case *pass 1 +carries an open Major, pass 2 finds nothing* in opposite directions. Corrected before pass 17: the +duty is a precondition on **closure** only, and that case is now walked through the ordering in the +shipped text — pass 2 is clean, the cycle is eligible, and it still cannot close. **The rollback is holding.** No finding asked for a restored enumeration, no finding hit a stated total, and the one new source edit (item 21) cost no count update — which is what removing the totals bought. +**What the falling curve does and does not show.** Every count here describes the revision the +pass **reviewed**, never the repair made after it. Pass 16's five Blocker/Majors do not assess +their own fix in `c792383`, and the contradiction that fix introduced is the proof: it was found +by reading, not by the curve. The number to watch next is not the level but **whether pass 17's +first Blocker comes out of this repair again or reaches something unreviewed** — three passes of a +flat Blocker count say less than one pass's answer to that. + **Minor 7 and Nit 8 are collected, not repaired as their own round.** Nit 8's number was correct and its wording was not: 58% measures the four bookkeeping sections (456 of 785), not "the remaining" after the design's 173. Fixed in the same clause while the surrounding sentence was diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index fb91e31..46ed461 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -151,7 +151,7 @@ to run again. **The four standing duties, classified.** The **derived floor** is a **precondition on closure**: it gates closing, discharged by the count of valid logical passes reaching it with the last of them clean, or by the zero-finding exit. The **Blocker/Major-resolve duty** is a **precondition -on closure and on any pass being clean**: Mechanics Severity, scoped to the assigned fix set, +on closure**: Mechanics Severity, scoped to the assigned fix set, states what it demands and **what discharges it — a repair or a validated dismissal — and what a dismissal is**, all at that source. **It is discharged per finding and tracked across the cycle, never inferred from a later pass.** A findings file establishes the **inventory** of what that @@ -159,8 +159,17 @@ pass found and not the resolution of anything, so a later pass that does not men in-set Blocker or Major says nothing about whether it was repaired or dismissed; reading its absence as discharge would let an omission close a cycle. **Closure therefore reads two things, not one**: the final pass is clean, *and* no in-set Blocker or Major raised anywhere in this cycle -is still undischarged. A pass's cleanliness stays a fact about that pass's own findings, which is -what §5 already says it is. The **hold** a surfaced finding places on closure **participates in the +is still undischarged. + +**The duty is not a second test on whether a pass is clean, and the two are not run together.** A +pass is clean on its own findings, which is what §5 already says cleanliness is. Stated as the +case that separates them, because a reader who conflates them decides it wrongly: **pass 1 raises +an in-set Major; it is neither repaired nor dismissed; pass 2 finds nothing.** Pass 2 **is** clean, +and if it is at or above the floor the cycle is **eligible** — and the cycle **still cannot +close**, because the second thing closure reads is unmet while that Major is undischarged. **The +zero-finding route changes nothing here**: it is a route to eligibility, and eligibility reads the +duty like every other final-acceptance precondition. What the open Major does *not* do is make +pass 2 unclean, retroactively or otherwise. The **hold** a surfaced finding places on closure **participates in the ordering**: it gates closing while it stands, and is discharged by the answers that surface requires. **It attaches to every surfaced finding, whichever suspension surfaced it** — clean completion creates none, because it wins before anything is surfaced. **No-clean-credit** — no @@ -325,7 +334,7 @@ reason §7 gives. Working around any of these would ship two instructions that d | 18 | the five-tells pointer | passage (e) | the passage says nothing about what its answer does; one sentence names the stop a suspension and defers to the ordering | addition | | 19 | the handed-over severity question | Mechanics · Severity, passage (g) | the paragraph says the demotion/loop-health question is unsettled and mandates a stop; replaced by the answer, which is the (g) text below | replacement | | 20 | the unknown-start strict-reading list | passage (i) | the list of what a cycle takes at its strictest omits the suspensions this change ships; it gains **every suspension binding, since unknown starting rules cannot waive an open hold** — the reason carried inline, as prompt-standards item 6 requires of the clause itself | addition | -| 21 | the `WIP:`-naming warning's closure claim | Mechanics · `baseSha`, C and W — **outside the ten inventoried passages** | "A pre-review snapshot named anything else reads as a real commit and **closes the cycle**, discarding the passes you just accumulated" is an **event-derived closure path** beside the ordering's content-and-precondition-derived one, and the two decide an accidental commit in opposite directions. It becomes: the hook reads such a commit as closing and discards the counted passes, **while the cycle itself stays open until the ordering's closure conditions hold**. The warning keeps its force — the cost of the mistake is the lost pass credit — and stops claiming a closure the rules do not grant | replacement | +| 21 | the `WIP:`-naming warning's closure claim | Mechanics · `baseSha`, C and W — **outside the ten inventoried passages** | "A pre-review snapshot named anything else reads as a real commit and **closes the cycle**, discarding the passes you just accumulated" is an **event-derived closure path** beside the ordering's content-and-precondition-derived one, and the two decide an accidental commit in opposite directions. It becomes: the hook reads such a commit as closing and discards the counted passes, **while the cycle itself stays open until the ordering's closure conditions hold**. The warning keeps its force — the cost of the mistake is the lost pass credit — and stops claiming a closure the rules do not grant. **The sibling sentence in the profile-change paragraph was checked and needs no edit**: it already says such a commit "reads **to the hook** as the cycle closing", which is the counter observation and claims no closure | replacement | | 22 | the one-contract paragraph | Mechanics, the "These records are one contract" paragraph, C and W — **outside the ten inventoried passages** | its coherence stop reaches the nonce, the slots, the provenance line, the curve, the carry rule and the unknown-start semantics, and reaches none of this change. Two changes, one contiguous rewrite of the paragraph's opening and its membership: the opening noun broadens from *records* to **the rules and records a cycle runs under**, since what is added is not a record; and the membership gains **the closure ordering together with every source edit it cites or depends on**, stated as that description and **not as a list of item numbers** — the numbering is this spec's working aid, it is not in the shipped text, and a shipped list would have to be re-derived whenever a span merges. **What this is: the same instruction to the agent, over a wider membership.** Not a checker, and none is built. **The plan establishes the membership against the real files** — every edit the block cites or depends on — which is where dependence is decidable | replacement | **Passage (g)'s replacement is design rather than bookkeeping, so it is stated here in full:** From 12fe2e62d3c3bd0bbf3660932bbf22cd26f63284 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Fri, 11 Sep 2026 13:13:38 +0200 Subject: [PATCH 028/181] =?UTF-8?q?docs(context):=20Gate-A=20pass=2017=20?= =?UTF-8?q?=E2=80=94=20mandatory=20two-tell=20stop,=20findings=20held=20op?= =?UTF-8?q?en?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 9 findings (1 Blocker, 7 Major, 1 Minor), Blocker+Major 5 -> 8. Two tells: the finding count rose and the Blocker count failed to fall for a fourth pass. Surfaced with every finding open per section 5; nothing repaired, nothing closed. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-17.md | 10 ++++ .../gate-a-spec-awsf1ec771-resume.md | 51 ++++++++++++++++++- 2 files changed, 60 insertions(+), 1 deletion(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-17.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-17.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-17.md new file mode 100644 index 0000000..a40b480 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-17.md @@ -0,0 +1,10 @@ +BLOCKER | high | §3 "First, clean completion" | Gate-A closure is defined as the moment the reviewed spec or plan is committed, but the superpowers brainstorming flow commits the spec before Gate A is invoked and this pass itself reviews the already-committed c2d5fe1; §4 names no sequencing edit or no-change closing operation for a clean pass over an existing commit | Such a clean Gate-A pass has no future closing transition and, because clean completion outranks suspension, can remain open without a compliant close or suspend path | Either make Gate-A final acceptance close the already-identified artifact commit after the post-pass rechecks, or change the Gate-A handoff so the reviewed artifact is committed only after the clean pass +MAJOR | high | §2 "Settled inputs" and §3 opening | The spec says it defines no duty, says the block is authoritative only for order and closure, and calls any duty defined there a defect, yet the block itself defines the surfaced-finding hold, its components and discharge, composition, and cannot-co-occur pairs, which the governing one-authority decision assigns to the block | The plan cannot tell whether to preserve those rules in the block or move or cite them, so either the only definition can be removed or duplicate authorities can ship | Amend the meta-contract to name every rule the block owns—order, closure, hold, composition, cannot-co-occur pairs, and the clearly-stuck precedence exception—and reserve the citation-only claim for external triggers, severity rules, and closure preconditions +MAJOR | high | §3 "First, clean completion" and the worked case | The first branch says eligibility exists only under final-acceptance preconditions and later says an unmet closure precondition prevents eligibility, while the pass-1 Major and pass-2 zero-finding case says pass 2 is eligible even though the resolve precondition is unmet | On a clean at-floor pass with an open earlier duty plus a health tell, one reading lets clean completion win and suppress the suspension while the other reaches the suspension branch, so the fixed order does not decide the path | Define eligibility independently as clean plus floor-or-zero, define closure as eligibility plus every precondition plus the kind-specific closing transition, and state whether an eligible pass with a failed precondition is tested for suspensions before continuing +MAJOR | high | §4 edit set, passage (b) condition b16 | The live source says that when a correction also opens a structural or contract question "the new question wins" and "novelty overrides correction ancestry", but the new block says an out-of-set correction that opens that question raises both scope triggers and requires both answers; §4 leaves b16 unchanged | An agent can follow the surviving source and ask only the structural question, omitting the membership decision and silently absorbing work outside the assigned fix set | Add b16 to the source edits and make clear that novelty overrides correction ancestry only; membership and question stops still compose when both predicates hold +MAJOR | high | §4 edit set, Gate-B coverage instruction | Item 2 repairs the Gate-A clean-signal sentence, but the Gate-B instruction at CLAUDE.md:597–599 and workflow-init.md:788–789 still says to report every finding and "say NO FINDINGS if clean"; under §3 a clean pass may contain Minors and Nits | The retained sentence can make a Gate-B reviewer suppress real Minor and Nit findings or emit a false zero-finding file, violating coverage-first and changing the early-exit decision | Name and edit the Gate-B sentence so NO FINDINGS is emitted only when the branch has zero findings, while every actual finding is still written +MAJOR | high | §4 item 19 "Passage (g)'s replacement" | The replacement says a nonstandard severity such as IMPORTANT counts for the curve "exactly as it counts for the pass", then says the curve reads pre-ceiling normalized severity while pass cleanliness and the resolve duty read post-ceiling effective severity | When the ceiling demotes that normalized Major, the two counts intentionally differ, so the comparison reinstates the superseded reading and can lead an agent to apply the ceiling to the health curve | Say that IMPORTANT normalizes to Major for the curve under the Reader rule and remove the claim that it counts the same for the pass outcome +MAJOR | high | §4 omitted source edit, Mechanics curve rationale | The live curve paragraph at CLAUDE.md:938–941 and workflow-init.md:1122–1125 says Majors are recorded because "the severity rule moves the Blocker/Major line rather than the total", while item 19 now makes demotion leave the pre-ceiling Blocker and Major series unchanged; §4 names no edit to that paragraph | Both prompt copies retain a source-level explanation that tells readers the ceiling changes the curve, contradicting the central severity and loop-health decision | Add this passage to §4 and replace the rationale with one based on preserving the reviewer-normalized pre-ceiling severity mix +MAJOR | high | §4 item 22 "one-contract paragraph" | The shipped membership rule is only "the closure ordering together with every source edit it cites or depends on"; a downstream prompt contains neither the spec's edit set nor the prior text, and "depends on" supplies no live semantic test for deciding whether a rule is a member | Even an agent that reads and obeys the authorised coherence instruction cannot reliably detect that a source rule such as b16 or the curve rationale was only partially adopted, so the stop instruction is not executable on its stated membership | State a semantic membership test in the shipped paragraph—for example, every live rule whose value can change ordering inputs, branch selection, holds or answers, or closure conditions—without enumerating item numbers or adding a checker +MINOR | high | §3 "What a suspension asks, and what ends it" | The clearly-stuck reading surfaces "those [findings] its third condition is about", but that condition spans a regeneration chain across several passes and the text does not select the current live recurrence rather than already repaired or dismissed predecessors | Different readers can create one hold for the current finding or reopen holds for every historical member of the chain, changing the required answer count and delaying resumption | Define the surfaced set as the current pass's live findings that satisfy the regeneration reading, with discharged predecessors used only as history for the reading +END OF FINDINGS (9 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index 20fe287..60fb662 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -41,7 +41,56 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 14 | c06dd69 | 16→**10** | 2→**1** | 8→**5** | yes | spec byte-identical to 0168f88 (532 lines); **lowest B+M of the cycle, 6**; **SCOPE STOP surfaced on finding 6** (partial-adoption guard = a 21st edit, outside the twenty); session 01a08f70-6d1e-7cf1-baa3-6df3cc59bc32 | | 15 | 17cf2e9 | 10→**12** | 1→**1** | 5→**7** | yes | **TWO-TELL STOP — mandatory, surfaced to Daniel.** B+M 6→8. **Six of twelve regenerate from pass 14's own repairs** (3, 4, 6, 7, 10, 11). All 12 held open; session 01a08fb2-6e8e-73d0-9d2b-8fb92d1571dd | | 16 | 5d12884 | 12→**8** | 1→**1** | 7→**4** | yes | **B+M 8→5, lowest of the cycle.** One tell (Blockers flat). Five B/M applied; Minors 7 and Nit 8 collected. Findings 2 and 4 are the spec disagreeing with *standing* §5 text, not with itself; session 01a08fff-fadf-72f2-994d-618048bea473 | -| 17 | — | — | — | — | not run | next action | +| 17 | c2d5fe1 | 8→**9** | 1→**1** | 4→**7** | yes | **TWO-TELL STOP — mandatory, surfaced to Daniel.** B+M 5→8. **Three findings (4, 5, 7) are one mechanism on its third pass**: a standing §5 sentence the block falsifies that §4 does not name. All 9 held open; session 01a0901b-ecf3-73b2-886c-5b3c7dc63a74 | +| 18 | — | — | — | — | not run | blocked on the two-tell answer | + +## Pass-17 three-line report — MANDATORY TWO-TELL STOP + +**Floor line:** derived floor **3**; risk **high**, security **none**; read fresh from +`docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md`. One cited story, level 2. + +- **Trend:** findings …, 10, 12, 8, **9**. Blockers …, 1, 1, 1, **1**. Majors …, 5, 7, 4, **7**. + Blocker+Major …, 12, 17, 14, 14, 10, 6, 8, 5, **8**. Four passes at 6, 8, 5, 8. +- **Cluster (pass 17):** product **9 of 9**. The instrument 0, prose-about 0 — the cleanest + cluster of the cycle, and it is not good news: the findings are all about shipped behaviour. +- **require↔withdraw:** none. Finding 6 objects to a sentence pass 16 added; finding 8 objects to + the rule that replaced the enumeration **without asking for the enumeration back** — it names + "without enumerating item numbers or adding a checker". Both are the mirror shape, not the shape. + +**Tells: two of five — the threshold. Stop-and-surface is mandatory.** The finding count rose +8 → 9 and the Blocker count failed to fall, 1 → 1 for the fourth pass running. + +**The question asked before this pass has an answer, and it is the unwelcome one.** The test set +was: does pass 17's Blocker come out of the last repair, or reach something unreviewed? **It comes +out of the repair.** Finding 1 is Gate-A closure — introduced at pass 15, refined at 16, refined +again at 17 — on its third round. Findings 2, 3 and 6 are also pass 15/16 repairs. + +**One mechanism is on its third pass and produced a third of this pass alone.** Findings 4, 5 and 7 +are one shape: **a standing §5 sentence the block falsifies, which §4 does not name.** `b16`'s +"novelty overrides correction ancestry"; the Gate-B coverage instruction's "say `NO FINDINGS` if +clean"; the curve rationale's "the severity rule moves the Blocker/Major line". The same shape was +pass 16 finding 2 (the `WIP:` warning) and pass 15 finding 2 (the same warning, both sites). So: +passes 15, 16 and 17, and three instances in this pass. + +**Its cause is structural, like the §7 assert list at pass 12.** §4 claims to name every passage +the change touches. That claim cannot be established by a spec against text the same change is +rewriting — each pass greps differently and finds another. The assert list had the same shape and +the repeat criterion ended it in one pass, by moving the enumeration to where the text exists and +printing the residual. + +**But three of them are KNOWN contradictions and cannot be deferred**, which is the line the +reviewer drew and it holds: `b16`, the Gate-B coverage sentence and the curve rationale are live +sentences that disagree with the block *today*. Renaming them a residual would be exactly the +"no known behavioural contradiction becomes a residual" failure. The criterion applies to the +**search** for further ones, never to these three. + +**The clearly-stuck reading is now satisfied on all three conditions**, stated because it is +reachable and not because it is being taken: a plateau — B+M 6, 8, 5, 8 over four passes and never +zero in seventeen; coverage affirmable after seventeen readings; and Blocker/Major regenerating +from the previous round's own repairs, nameably, in four of this pass's nine. + +**Minor 9 collected**, not repaired: which findings the clearly-stuck reading surfaces when the +regeneration chain spans several passes. ## Pass-16 three-line report From bc082363614a52b1fb30833f26e056f20e6cbb85 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Fri, 11 Sep 2026 14:06:53 +0200 Subject: [PATCH 029/181] docs(specs): produce the non-active target text; stop describing the change four times Answer to the pass-17 mandatory two-tell stop: interrupt the repair mode. Neither option I offered was taken, and the reviewer was right to reject both. THREE CORRECTIONS, each verified against the source before acting. - My "B" was not the clearly-stuck exit. CLAUDE.md:242 says surfacing does not close the cycle, credits no pass as clean and leaves findings open. "End Gate A and move to Gate B" would have been a change to the review procedure wearing an exit's name. - The C1 precedent does not show what I said. The field report records that the Gate-B cycle receiving the relocated work closed as NOT converged, on the clearly-stuck exit, with no clean pass and none claimed. It is a decision to relocate, not a demonstrated convergence. Citing it as a success was the overclaim AGENTS.md names as this repo's most persistent defect, in my own report. - My diagnosis was too broad. A spec CAN determine the affected sites; the standing text is on disk and greppable. What it cannot do is keep four descriptions of a future text in agreement while all four are being revised -- the ordering block, the edit table, the source sentences and the rationales. That is the generator, and it is what seventeen passes were paying for. WHAT SHIPS INSTEAD. One non-active target-text file carrying the section 5 passages as they will read, 476 lines, each section marked NEW, REPLACED or CARRIED. Nothing is installed; CLAUDE.md and the workflow-init template are untouched and no rule in it governs any running cycle. The design spec stops re-narrating that text: 635 -> 385 lines, keeping the settled inputs, the passage map, parity, verification, the invariants touched and what moved or is parked. All nine pass-17 findings are resolved in the target text. None deferred, none renamed a residual, and the Minor got no round of its own. - Blocker: a Gate-A cycle closes on the author's recorded acceptance of the revision the clean pass reviewed, a commit that ordinarily already exists, and no new revision is made to close one -- a new revision being one no pass read. - Eligibility is now its own test (clean at or above floor, or zero-finding) and closure is eligibility plus every precondition plus the kind-specific closing act; an eligible pass with an unmet precondition lands on the continue branch. - The block's opening names the six things it owns and cites everything else. - The three standing sentences the block falsifies are written out: b16's both-triggers case, the Gate-B coverage instruction, the curve's Majors rationale. - The claim that a nonstandard token counts for the curve "exactly as it counts for the pass" is gone; the two are meant to differ where the ceiling demotes. - The one-contract paragraph carries a semantic membership test a downstream reader can apply, replacing a rule that named an edit set existing only here. - The clearly-stuck reading surfaces the current pass's live findings only. Gate A stays OPEN and is re-aimed at the target text with these findings as its basis. The rules are not activated. Gate B reviews the implementation diff later, and neither file is that diff. No transition to Gate B is claimed on a clearly-stuck exit. Verified: precheck exit 0; section 6 confirmed present after the restructure (it was dropped by the first splice and restored); zero stale item-number references; the passage map points at target-text sections. Design 635 -> 385, target text 476 new. No gate: docs/**.md plus the cycle's working record, prose-exempt under section 5. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-resume.md | 52 +- ...26-09-10-loop-rule-consolidation-design.md | 397 +++------------ ...-10-loop-rule-consolidation-target-text.md | 476 ++++++++++++++++++ 3 files changed, 600 insertions(+), 325 deletions(-) create mode 100644 docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index 60fb662..9000a84 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -5,7 +5,9 @@ Nothing depends on it; the pass files and the repo are authoritative where this - **Kind:** Gate-A spec - **Nonce:** awsf1ec771 (drawn 2026-09-10 from /dev/urandom, 10 chars, no collision among open cycles) -- **Artifact:** `docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md` +- **Artifact:** `docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md` from pass 18 on + (passes 1–17 reviewed `docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md`; the cycle + continues under this nonce, the artifact having been restructured rather than replaced) - **Branch:** loop-rule-consolidation - **Story:** `docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md` — profile read from its header at each pass (was high / none / battery+check+verification at pass 1) - **Derived floor:** 3 (risk high → level 2; security none → 0; max 2 ≠ 0 → 3) @@ -42,7 +44,53 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 15 | 17cf2e9 | 10→**12** | 1→**1** | 5→**7** | yes | **TWO-TELL STOP — mandatory, surfaced to Daniel.** B+M 6→8. **Six of twelve regenerate from pass 14's own repairs** (3, 4, 6, 7, 10, 11). All 12 held open; session 01a08fb2-6e8e-73d0-9d2b-8fb92d1571dd | | 16 | 5d12884 | 12→**8** | 1→**1** | 7→**4** | yes | **B+M 8→5, lowest of the cycle.** One tell (Blockers flat). Five B/M applied; Minors 7 and Nit 8 collected. Findings 2 and 4 are the spec disagreeing with *standing* §5 text, not with itself; session 01a08fff-fadf-72f2-994d-618048bea473 | | 17 | c2d5fe1 | 8→**9** | 1→**1** | 4→**7** | yes | **TWO-TELL STOP — mandatory, surfaced to Daniel.** B+M 5→8. **Three findings (4, 5, 7) are one mechanism on its third pass**: a standing §5 sentence the block falsifies that §4 does not name. All 9 held open; session 01a0901b-ecf3-73b2-886c-5b3c7dc63a74 | -| 18 | — | — | — | — | not run | blocked on the two-tell answer | +| 18 | — | — | — | — | not run | **two-tell stop ANSWERED 2026-09-11: interrupt the repair mode.** Target text produced (`…-target-text.md`, 476L); design cut 635→385L. Pass 18 aims at the target text | + +## TWO-TELL STOP ANSWERED 2026-09-11 — interrupt the repair mode, Gate A stays open + +**Neither A nor B as I put them.** The reviewer rejected both and was right on three verified +points, each checked against the source before acting: + +1. **My B was not the clearly-stuck exit.** `CLAUDE.md:242` — "Surfacing does not close the cycle + … no pass is credited as clean". "End Gate A and move to Gate B" would have been a change to + the review procedure wearing an exit's name. +2. **The C1 precedent does not show what I said.** `docs/field-reports/2026-08-30-gate-a-rle-plan-cycles.md:127` + — the Gate-B cycle that received the relocated work "closed as not converged … on the + clearly-stuck exit, with no clean pass and none claimed". It records a decision to relocate, + never a convergence. I cited it as a success; that is the overclaim `AGENTS.md` names, in my + own report. +3. **My diagnosis was too broad.** "A spec cannot determine the affected sites" is false — the + standing text is on disk and greppable. What a spec cannot do is keep **four descriptions of a + future text** in agreement while all four are being revised: the block, the edit table, the + source sentences and the rationales. That is the actual generator. + +**What was done instead.** One **non-active target-text** file, `…-target-text.md`, carrying the +§5 passages **as they will read** — 476 lines, marked NEW / REPLACED / CARRIED per section. The +design spec stops re-narrating them: 635 → 385 lines, keeping the settled inputs, the passage map, +parity, verification, invariants and what moved. Four descriptions become one text plus its +reasons. + +**All nine pass-17 findings resolved in the target text**, none deferred and none renamed a +residual: +- **1 (Blocker)** Gate-A closure: a Gate-A cycle closes on the author's recorded acceptance of the + revision the clean pass reviewed — a commit that ordinarily already exists — and **no new + revision is made to close one**, a new revision being one no pass has reviewed. +- **2** the block's opening now names the six things it owns and reserves citation for the rest. +- **3** eligibility is defined on its own (clean at or above floor, or zero-finding) and closure is + eligibility plus preconditions plus the closing act; an eligible pass with an unmet precondition + lands on the third branch, which the branch now says. +- **4, 5, 7** the three standing sentences the block falsifies — `b16`, the Gate-B coverage + instruction, the curve's Majors rationale — are written out in §F. +- **6** the "counts exactly as it counts for the pass" claim is gone; the two counts are meant to + differ wherever the ceiling demotes. +- **8** §G carries a **semantic** membership test a downstream reader can apply, replacing "every + source edit it cites or depends on", which named an edit set existing only in this repository. +- **9** the clearly-stuck reading surfaces the current pass's live findings only; discharged + predecessors are history it consults. + +**Gate A is not closed and is not claimed closed.** It re-aims at the target text with these +findings as its basis, the rules stay unactivated, and Gate B reviews the implementation diff +later. Pass 18 is the next action. ## Pass-17 three-line report — MANDATORY TWO-TELL STOP diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index 46ed461..b5bda49 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -67,313 +67,67 @@ one site — is not stable across revisions that merge or split a span, and a st disagrees with its own table. The plan counts what it writes. --- - -## 3. The closure ordering — the block that ships - -It sits in §5 **immediately before** the paragraph "**What a loop absorbs, and what stops it**", -in both copies, byte-identical. It states the ordering once; the paragraphs after it keep their -triggers and point at it. Verbatim as it will ship: - -``` -**How a cycle ends — one ordering, stated here and referenced everywhere else.** A pass is read -in a fixed order, because every rule bearing on one decision — may this cycle close — otherwise -qualifies the others and the ranking survives only in a reader's head. **This paragraph is -authoritative for that order and for closure, and for nothing else.** Every trigger, duty and -severity rule it names keeps its one definition where that definition already lives; a reader -who finds a rule defined here rather than cited has found a defect. - -**What a pass is read from.** Every finding-derived predicate reads the validated findings file -**or files** of the logical pass as their **concatenation** — a `full` Gate-B pass has two, and -one branch alone is already an incomplete pass. **Which severity field each of them reads is -settled in Mechanics, Severity**, which is where that split lives and is not repeated here. -Beyond the findings, closure reads the **derived floor** and the final-acceptance preconditions -the floor section states, and any **hold still standing**; **in a Gate-B cycle it also reads the -evidence entry's revalidation rule** — a changed entry meaning the clean pass no longer covers -what is being committed — which is a Gate-B precondition because the entry is about the diff -being committed and a Gate-A cycle produces none. The -scope triggers read the **current assigned fix set** as the absorb paragraph defines it, and the -answers already given; the clearly-stuck reading adds its own coverage judgement. **A line in -one branch file and a line in the other are distinct findings for holds and answers**, so a -`full` pass asks twice rather than risk resuming over one it never asked about. **Within one -running cycle an answer binds to the finding or question as the pass that raised it recorded -them** — which is what an agent running the cycle can do with nothing written down. Recognising -the same finding or question across a lost session needs a record these rules do not ship. - -**First, clean completion.** §5 uses *clean* in two senses and now says which is which: a **clean -findings file** is the `NO FINDINGS` signal the protocol defines, and a **clean pass** is the -predicate here, read on the logical pass with every required branch file combined, so one -branch's clean file never establishes a clean pass. **A pass is clean** when its findings carry -no in-set Blocker or Major at effective severity and **no scope-stop trigger** — the two the -absorb paragraph defines, read there and not redefined here, each already carrying the -qualification an answer given in this cycle puts on it. Both are properties of the findings and -the set as that paragraph reads them, settled before any branch below runs, which is what makes -this order executable rather than asserted. A pass with **zero** findings is clean whatever the floor, because a floor buys further -looks at an artifact that keeps yielding findings, and one yielding none has already given what -those looks were for. **A clean pass at or above the derived floor, or a zero-finding pass, makes -the cycle eligible to close**, under those final-acceptance preconditions. **Eligibility is not -closure: the cycle closes when the commit carrying its reviewed artifact is made, and which -commit that is depends on the cycle kind** — for a **Gate-B** cycle the closing amend -Mechanics · Finishing the cycle describes, for a **Gate-A** cycle the commit of the reviewed spec -or plan, that gate having no WIP snapshot and no amend. Nothing is closed before that commit, and -a profile, cited set or — in a Gate-B cycle — evidence entry that changes in between still gates -it. **A commit the hook reads as cycle-closing is a separate matter**: a non-`WIP` commit mid-cycle -makes the hook read the cycle as closed and discards the passes counted so far, which is an -observation about the counter — **the cycle itself stays open until the conditions above hold**, -so an accidental commit destroys the pass credit and closes nothing. The Mechanics sentence that -says such a snapshot "closes the cycle" is edited to say that, since left as written it is a -second, event-derived closure path beside this one. A plateau -or tells on that pass go into the closing report and never block it, because reporting "will not -converge" on a converged loop is a false report. **No other pass outcome makes a cycle eligible -to close**, because every other pass leaves a required repair, a hold or a question outstanding, -or has an unmet closure precondition — and closing over any of those is the failure this ordering -exists to prevent. The one termination that is not a pass outcome is the Gate-B triviality skip, -which runs no passes and is outside this ordering. - -**Second, only a pass that is not a clean completion can suspend** — that order is what makes -"clean completion outranks the two-tell stop" executable rather than asserted. Three suspensions, -by the names their paragraphs use and read by those paragraphs: the **scope stop**, raised by -either trigger above — a **membership stop** by the first, a **question stop** by the second; the -**clearly-stuck exit**; and the **two-tell stop**. A suspension waives nothing. Any non-empty set -of them can apply to one pass: **one surface, every reason reported, every question asked**, -because a reason left out is a decision made by omission. A finding the clearly-stuck reading -surfaces that also carries either trigger takes the scope stop's answers at that same surface, so -it is not asked twice; the two-tell stop surfaces tells and not a finding. - -**Third, a pass that neither closes nor suspends continues** — the loop runs another pass on -the **current** artifact, revised where the severity and scope rules require a repair and -unrevised where they do not. A below-floor clean pass lands here **only where no suspension -applies to it**; where one does, the second branch has already taken it, because clean completion -did not close the pass and only closing outranks a suspension. So does a pass whose only findings -are Minors and Nits, which are collected and never iterated and may leave nothing to revise. It -is a branch and not an inference, because "does not close" read alone says nothing about whether -to run again. - -**The four standing duties, classified.** The **derived floor** is a **precondition on closure**: -it gates closing, discharged by the count of valid logical passes reaching it with the last of -them clean, or by the zero-finding exit. The **Blocker/Major-resolve duty** is a **precondition -on closure**: Mechanics Severity, scoped to the assigned fix set, -states what it demands and **what discharges it — a repair or a validated dismissal — and what a -dismissal is**, all at that source. **It is discharged per finding and tracked across the cycle, -never inferred from a later pass.** A findings file establishes the **inventory** of what that -pass found and not the resolution of anything, so a later pass that does not mention an earlier -in-set Blocker or Major says nothing about whether it was repaired or dismissed; reading its -absence as discharge would let an omission close a cycle. **Closure therefore reads two things, -not one**: the final pass is clean, *and* no in-set Blocker or Major raised anywhere in this cycle -is still undischarged. - -**The duty is not a second test on whether a pass is clean, and the two are not run together.** A -pass is clean on its own findings, which is what §5 already says cleanliness is. Stated as the -case that separates them, because a reader who conflates them decides it wrongly: **pass 1 raises -an in-set Major; it is neither repaired nor dismissed; pass 2 finds nothing.** Pass 2 **is** clean, -and if it is at or above the floor the cycle is **eligible** — and the cycle **still cannot -close**, because the second thing closure reads is unmet while that Major is undischarged. **The -zero-finding route changes nothing here**: it is a route to eligibility, and eligibility reads the -duty like every other final-acceptance precondition. What the open Major does *not* do is make -pass 2 unclean, retroactively or otherwise. The **hold** a surfaced finding places on closure **participates in the -ordering**: it gates closing while it stands, and is discharged by the answers that surface -requires. **It attaches to every surfaced finding, whichever suspension surfaced it** — clean -completion creates none, because it wins before anything is surfaced. **No-clean-credit** — no -pass carrying a scope-stop trigger is credited as clean — also participates, and is a fact about -that pass that nothing discharges, a later pass being judged on its own findings. It is not a -second test beside the clean predicate but that predicate's second half, which is why it is -stated in its words. - -**What a suspension asks, and what ends it.** **A hold ends when every answer its finding -requires has been given, in whichever direction each is given** — one **scope-stop** answer for a -single-trigger finding, both for one carrying both. That is **one rule with two parts**, how many -answers and which way each may go, and neither is a test the other has to pass. Where a health -suspension applies to the same pass, its shared continue-or-stop answer is **additional** to -those and not counted among them, the health readings asking about the loop rather than about -this finding. **The two health readings differ in what they surface, and therefore in what they -hold.** The **two-tell stop surfaces tells and no finding**, so it creates no hold; what it leaves -outstanding is its own continue-or-stop question, which the composition rule below holds the cycle -on until it is answered. The **clearly-stuck reading does surface findings** — those its third -condition is about — and each takes a hold like any other surfaced finding. **Where one finding is -surfaced by both, it carries two hold components and each is discharged by its own answer, which -is what keeps D4 and D5 exact**: the **membership** component ends on the membership answer in -either direction, a decline releasing it as **D5** requires and as the absorb paragraph's own -sentence says; the **clearly-stuck** component ends on the reading's continue-or-stop answer. -Neither answer discharges the other's component, and where the finding carries no scope-stop -trigger the continue-or-stop answer is the only one its surface asks for and discharges the only -component there is. **Resumption is still the composition rule's**, which waits for every -outstanding answer — so separating the components changes what each answer discharges and never -lets one answer resume the loop alone. So no surfaced finding is left without a discharging -answer, and no surface without a finding is given a hold nothing could discharge. -At a **membership stop** the answer is **accept**, the finding joining the fix set -where Mechanics Severity governs it, or **decline**, the finding staying outside and binding so -for the rest of this cycle. A later answer that contradicts a decline **does not reverse it**: -the decline **remains binding** and the contradiction is **surfaced to the user as information**, -changing neither membership, nor the cycle's state, nor any outstanding question — **whether the -loop resumes is decided by the composition rule below and by nothing here**, so a contradictory -answer is never itself a resumption and cannot step past a hold or a health question still -awaiting its own answer. Nothing here turns one answer into another, since that would let a -finding be moved out of the set and back into it to escape what it owes inside it. **There is no -withdrawal inside the cycle that declined**, because **D7** binds a decline for the remainder of -its cycle and admits no exception; a reconsideration is a later cycle's, where D7 gives the -decline no effect at all and the finding takes the ordinary route. Either answer is an **explicit, -attributable decision on that specific finding** — never silence, never a general remark about -scope, never inferred, because a fix set changed by inference is a fix set nobody chose. -**Membership is answered against the set as the absorb paragraph fixes it for the pass that -raised the question**: a later broadening is a new fact the **next** pass reads and never -discharges a standing hold, a hold discharged by a scope change being a hold nobody answered. At -a **question stop** the answer is the user's decision on the question and membership does not -change; an out-of-set finding that opened one is a membership stop as well. **Decline is -available only at a membership stop**, that being the only stop whose question is whether a -finding belongs to the set. The **clearly-stuck and two-tell readings** ask **continue or stop**. -**Continue consumes the reading that raised the suspension**: a further health suspension needs -that reading recomputed over a pass run after the answer, which is new data — so continue -produces a distinct next state, and the same reading cannot return the same stop unanswered. It -permits an **unrevised** artifact **only where no repair is owed**; where effective severity or -scope requires one, that repair comes before the post-answer pass, since a pass run over an -unrepaired in-set Blocker or Major spends a look on text the rules already say must change. -**Stop parks the cycle**: open, not running, spending no passes, restarted only by an -explicit later continue — a distinct state from the suspended-awaiting-answer one it was in -before the answer. **That continue restarts the cycle and never skips an answer**: where any -question the suspension raised is still outstanding, it returns the cycle to -suspended-awaiting-answer, and only once every answer the composition rule requires has been -given does the next pass run. So a cycle parked with an unanswered membership or question stop -cannot be continued into a pass, and cannot sit parked with no transition either — the continue -is always available and always moves it. Nothing a parked cycle wrote is a closing commit, and a -parked cycle nobody restarts is a human's to resolve, exactly as the nonce rules already say of -open cycles. - -**Composition, and what cannot happen.** Every **question** is answered on its own and the loop -resumes only when every answer resumes it — accept or decline at a membership stop, a decision at -a question stop, continue at the health readings; one stop answer parks the whole suspension, -because a loop resumed over an unanswered question decides it by running. **The clearly-stuck and -two-tell readings raise one question between them, not two**, both asking continue or stop, so -one answer carrying every reason ends both — an instance of the sentence before it, not an -exception. **Two pairings cannot occur**, and no rule ranks them: clean completion and a **scope -stop**, since that stop's triggers are the clean predicate's own second half, so a pass raising -one is not clean; and a zero-finding pass and any suspension, since it has nothing to surface, -nothing regenerating, no cluster and no require↔withdraw pair. **Clean completion and the -clearly-stuck exit can**, and the overlap is admitted rather than argued away: the two read -different severity fields, as Mechanics · Severity sets out, so an **in-set** Blocker or Major the -ceiling demotes below Major, or an **out-of-set** one this cycle has declined, can regenerate -across passes on a pass that is clean. **The order decides it and no new rule is needed.** The -clearly-stuck paragraph's own precedence sentence is stated here rather than there, because -precedence is evaluation order and this paragraph is where evaluation order is stated once; its -opening words point back to that paragraph, which is where the reading itself lives. That third -condition is what makes a plateau rather than a finish, and it is why **a clean completion takes -precedence over this exit**: a Blocker/Major-free pass **at or above the floor** has satisfied the -clean-final-pass rule — collect the Minors and Nits and close — and reporting "will not converge" -on a converged loop is a false report. Below the floor the pass **suspends**, clean completion -having not closed it. -``` - -Where the block maps onto the criteria: the three branches are AC 4, the duties paragraph -AC 2, the composition sentences AC 1. - -**One authority per rule.** The block cites nine rules and defines none of them. Each has exactly -one definition in the shipped text, and where that definition had to change to agree with the -ordering, it changed **at its source** (§4) rather than being restated here: - -| Rule the block cites | Its one definition | -|---|---| -| membership trigger | `b11`, the absorb paragraph, qualified by §4 | -| question trigger | `b13`, the absorb paragraph, qualified by §4 | -| the assigned fix set, and what is in it | `b7` and `b8`, the absorb paragraph, both edited by §4 | -| effective vs reviewer-written severity | the (g) replacement, in Mechanics · Severity | -| what a Blocker, Major, Minor or Nit demands | Mechanics · Severity, scoped by §4 | -| the derived floor and final-acceptance preconditions | the floor section | -| the evidence entry's revalidation rule | the profiles section, unedited | -| the clearly-stuck reading | the clearly-stuck paragraph | -| the two-tell threshold | the five-tells paragraph, qualified by §4 | - -The scope stop's two triggers are `b11` and `b13` read separately, because an in-set finding -that opens a question can neither join nor stay outside the set and needs its own answer. - -**One rule runs the other way, and it is the one exception to the table.** The clearly-stuck -exit's **precedence** against a clean completion is not part of that exit's reading; it is -evaluation order, which is the block's own subject. So `c9`'s sentence — the one **D3** requires -preserved verbatim — **moves into the block** rather than being cited from where it stood, and -the clearly-stuck paragraph keeps only its reading and points forward. Stated in both places, it -would be exactly the drift this table exists to prevent; stated only in the clearly-stuck -paragraph, it would put evaluation order somewhere the ordering does not govern. The sentence's -words are untouched by the move, and the block says whose sentence it is where it quotes it. - -**Two things the ordering names and does not define, both the successor's** (§9): what makes a -later finding *the same one* this cycle declined, and what form carries an answer into the commit -body. The ordering says what an answer does; **D9b** and **D9** say how it is recognised across -sessions and written down. **Within one cycle the block supplies what it needs and says so**: a -line in each branch file is a distinct finding for holds and answers, and an answer binds to the -finding or question as the pass that raised it recorded them. Both sentences are in the block -above — the ordering claims no in-session consequence it does not state there. +## 3. Where the text is + +**The §5 text this change ships is written out in full, as it will read, in +`docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md`.** That file is the +artifact; this one is the decisions behind it. + +**Why the split, and it is the pass-17 answer.** Until pass 17 this spec carried the ordering +block, a twenty-two row edit table, the source sentences and the rationales as **four parallel +descriptions of one change**, which had to agree with one another simultaneously — so a repair to +any one of them fell out of step with the other three, and a pass found the gap. Seventeen Gate-A +passes, a Blocker count flat at 1 for four of them, and three findings in pass 17 alone from one +mechanism: a standing sentence the block falsifies that the edit table did not name. + +**What was wrong with the diagnosis that produced this, said because it was carried with force.** +The claim was that a spec *cannot* determine which sentences a change touches. That is too broad: +the standing text is on disk and greppable. What a spec cannot do is keep four descriptions of a +*future* text in agreement while all four are being revised. Writing the future text once removes +three of them, and the remaining enumeration question — whether every affected site was found — +is a sweep against real files, which §7 assigns to the plan and §I of the target text states as +an open residual rather than a closed claim. + +**What this spec still owns**, and none of it is in the target text: the settled inputs it reads +(§2), the passage map (§5), parity (§6), verification (§7), the invariants touched (§8), and what +moved or is parked (§9). + +**This is not a Gate-A close.** §5's clearly-stuck exit says surfacing closes nothing, credits no +pass as clean and leaves the findings open. Gate A stays open and re-aims at the target text with +pass 17's findings as its basis; the rules are not activated; Gate B reviews the implementation +diff later, and neither file is that diff. --- -## 4. The standing sentences edited at their source - -The standing sentences this change edits, each row naming the sentence, where it lives, what -changes and why. **No total is claimed and the rows are not numbered as a closed set**: a count -over spans that can be merged or split is bookkeeping that has to be re-derived at every revision, -and three of pass 15's findings were that bookkeeping disagreeing with itself. The plan establishes -the edit set against the real files, which is where a span is decidable. **No OLD or NEW text and no line numbers appear here**: the plan -quotes each sentence from the real file, writes the replacement, and re-greps the site, for the -reason §7 gives. Working around any of these would ship two instructions that disagree. - -| # | Sentence | Where | Change, and why | Kind | -|---|---|---|---|---| -| 1 | the findings-file protocol's clean sentence | the gate-prompt template both gates paste, C and W | it calls a `NO FINDINGS` file a *clean pass*; renamed a **clean findings file**, which is what it always described, so §3's predicate is not read as redefining it | replacement | -| 2 | the Gate-A clean-signal sentence | the Gate-A section, C and W | it makes the `NO FINDINGS` signal the only route to clean; it becomes what lets a pass be read as clean without inspecting it, since a pass carrying Minors alone is clean and could never produce that file | replacement | -| 3 | the Severity bullet's resolve duty | Mechanics · Severity, C and W | the only statement of the duty and the only one without a scope; scoped to **the assigned fix set**, the boundary **D5** implied and no sentence carried. It also gains **what discharges the duty — a repair or a validated dismissal — and what a dismissal is** (the author's judgement, carrying the one-line why already required, that the finding is not true of the artifact; it does not rewrite the pass that found it, the later clean pass is still owed, and it is **not** a decline, which is the user's decision that a *true* finding stays outside the set). The block cites this and defines none of it. It names the set and never a decline — AC 3's second half, since naming one here would read as a waiver of resolution rather than the membership decision it is | replacement | -| 4 | the lens paragraph's unchanged-list | the profiles section, C and W | it asserts the Blocker/Major filter and the clean-final-pass rule are unchanged, which the ordering falsifies; scoped to **the lens sets**, which is what that paragraph is about | replacement | -| 5 | the Gate-A cadence | the Gate-A section, C and W | "validate, revise, re-run" makes a revision unconditional, telling a Minor-only pass to manufacture the repair the severity rule forbids; the revision becomes conditional on a repair being required | replacement | -| 6 | `a17`–`a19`, the floor paragraph's closure sentences | passage (a), C and W | they state closure and the early exit in the paragraph that owns the floor; trimmed to point at the ordering, which states them once | replacement | -| 7 | `a13` | passage (a) | the source reads "Every other rule stated here about how a cycle closes stands as written, and none of them is restated — a summary is where their conditions would get dropped", **wrapping across three lines in each copy** (C 129–131, W 336–338), so the plan greps it in parts or unwrapped rather than as one line. It becomes categorically false once the ordering exists; scoped to that paragraph, which is what it was written to police | replacement | -| 8 | `a16` | passage (a) | "fix Blocker/Major after each" stands unscoped beside a Severity bullet item 3 now scopes; it points at that rule instead of restating an unscoped version | replacement | -| 9 | `b3` | passage (b), C and W | it points at Mechanics and then **restates the four severity actions**, giving them two definitions; it becomes a pure pointer. W's pointer also changes target, matching C (§6) | replacement | -| 10 | `b7`, the fix-set definition | passage (b) | singular "the approved story or plan" leaves a cycle governed by several with no set at all; it becomes their **union**, plus obligations already accepted, **minus this cycle's declines** | replacement | -| 11 | `b8`, the membership test | passage (b) | it tests a finding against "that scope" independently, so a broadened scope puts back a finding `b7` keeps out; it tests against **the assigned fix set as `b7` computes it** | replacement | -| 12 | `b11`, the membership trigger | passage (b) | unqualified, it and the clean predicate decide a re-raised declined finding in opposite directions; it excludes a finding this cycle has already declined | replacement | -| 13 | `b12` | passage (b) | it resumes on the membership answer alone; it says the answer ends the hold and **defers to the ordering** for what the pass does next | replacement | -| 14 | `b13`, the question trigger | passage (b) | *new* is undefined, so an answered question re-raised stops the loop again on every pass; *new* excludes a question already answered in this cycle | replacement | -| 15 | `b17`–`b18` | passage (b) | they state their own version of what ends a suspension; they name it a suspension and defer to the ordering | replacement | -| 16 | passage (c), from the third condition to the end | passage (c) | **one contiguous replacement covering both changes to that span**, so nothing is counted twice. The closure sentences carry evaluation order in the paragraph that owns the reading: the paragraph keeps its reading, and the precedence sentence moves into the block **word for word**, which satisfies **D3**. Within the same span the third condition `c8` is **widened at this source rather than from the block**: it requires Blocker/Major findings regenerating "across genuine repair attempts, each round's fix producing the next", which a **validated dismissal** never satisfies — no repair, no fix — so a false positive the reviewer re-raises every pass would leave the cycle unable to close and unable to suspend. It gains that case, the re-raise standing in for the regenerating fix. The accounting records `c8` as **replaced**, not kept | replacement | -| 17 | `e7`, the two-tell threshold | passage (e) | unqualified, it and the ordering decide a clean two-tell pass at or above the floor in opposite directions; it is read **after** the clean-completion branch. Authority **D2** | replacement | -| 18 | the five-tells pointer | passage (e) | the passage says nothing about what its answer does; one sentence names the stop a suspension and defers to the ordering | addition | -| 19 | the handed-over severity question | Mechanics · Severity, passage (g) | the paragraph says the demotion/loop-health question is unsettled and mandates a stop; replaced by the answer, which is the (g) text below | replacement | -| 20 | the unknown-start strict-reading list | passage (i) | the list of what a cycle takes at its strictest omits the suspensions this change ships; it gains **every suspension binding, since unknown starting rules cannot waive an open hold** — the reason carried inline, as prompt-standards item 6 requires of the clause itself | addition | -| 21 | the `WIP:`-naming warning's closure claim | Mechanics · `baseSha`, C and W — **outside the ten inventoried passages** | "A pre-review snapshot named anything else reads as a real commit and **closes the cycle**, discarding the passes you just accumulated" is an **event-derived closure path** beside the ordering's content-and-precondition-derived one, and the two decide an accidental commit in opposite directions. It becomes: the hook reads such a commit as closing and discards the counted passes, **while the cycle itself stays open until the ordering's closure conditions hold**. The warning keeps its force — the cost of the mistake is the lost pass credit — and stops claiming a closure the rules do not grant. **The sibling sentence in the profile-change paragraph was checked and needs no edit**: it already says such a commit "reads **to the hook** as the cycle closing", which is the counter observation and claims no closure | replacement | -| 22 | the one-contract paragraph | Mechanics, the "These records are one contract" paragraph, C and W — **outside the ten inventoried passages** | its coherence stop reaches the nonce, the slots, the provenance line, the curve, the carry rule and the unknown-start semantics, and reaches none of this change. Two changes, one contiguous rewrite of the paragraph's opening and its membership: the opening noun broadens from *records* to **the rules and records a cycle runs under**, since what is added is not a record; and the membership gains **the closure ordering together with every source edit it cites or depends on**, stated as that description and **not as a list of item numbers** — the numbering is this spec's working aid, it is not in the shipped text, and a shipped list would have to be re-derived whenever a span merges. **What this is: the same instruction to the agent, over a wider membership.** Not a checker, and none is built. **The plan establishes the membership against the real files** — every edit the block cites or depends on — which is where dependence is decidable | replacement | - -**Passage (g)'s replacement is design rather than bookkeeping, so it is stated here in full:** - -``` - **The demotion changes what a cycle must resolve, never what it observes.** **Every - loop-health reading observes the findings as the reviewer produced them, before the ceiling - is applied** — so a demoted finding still counts in the finding total and in its cluster, and - a Blocker demoted to Minor is still a Blocker to the curve and still regeneration to the - clearly-stuck reading's third condition. Where such a reading uses severity at all it takes - the **reader-normalized pre-ceiling severity**, which is the severity the Reader paragraph - above already produces from the findings file — case-folded, and a non-empty unrecognized - token read as `MAJOR` — never the raw token, so a finding written `IMPORTANT` counts for the - curve exactly as it counts for the pass. Cleanliness and the resolve - duty read the effective severity, after the ceiling (the - closure ordering above). The line is **what the cycle owes versus what it observes about - itself**, which is why no list of readings has to be kept complete here. Two reasons for the - split. The curve must stay derivable from the - findings files alone — the finding total counts finding lines and the Blocker and Major - series count the lines whose normalized severity is each, which is the only thing that makes - a self-reported curve checkable; the subject clusters use no severity at all, being a - judgement per finding that no count reproduces. - And the demotion is the author's judgement about **the finding's repair severity**, never - about which findings the fix set contains — a separate predicate the absorb paragraph - defines, and one this must not be read as touching; a loop spending passes on findings the author keeps - demoting is exactly what the prose-cluster tell exists to surface, and lowering the counts by - that same judgement would hide it. -``` - -Two sentences are **not** edited and are named so nobody looks for them: the "Copy every record -into the squash body" sentence inside the human-exception block, and the "records every cycle -owes" list. This change ships no record. - -**Item 22 was added at Gate-A pass 14 on Daniel's decision**, as a **bounded extension of the -existing coherence instruction** — no checker, no new mechanism, and nothing about record -durability, which stays with the successor story. It is the only edit in this table touching a -passage no earlier revision named, and §9 states what it is worth. The `c8` widening that pass 14 -first carried as a separate row was **merged into item 16 at pass 15**, the two replacing one -contiguous span. +## 4. The edits, by site + +**No sentence is quoted here and no replacement is written here** — the target text carries every +one, in final form, grouped by passage. This table exists so the passage map in §5 has something +to point at and so a reader can see the shape of the change without reading the whole target text. + +| Passage or site | What changes | +|---|---| +| the closure ordering | **new**, and it is section A of the target text | +| (a) the floor paragraphs | §H | +| (b) what a loop absorbs | §B | +| (c) recognizing clearly stuck | §C | +| (e) the five tells | §D | +| Mechanics · Severity | the resolve duty is scoped and gains its discharge rule; the handed-over question is replaced by its answer | +| Mechanics · `baseSha` | the `WIP:` warning stops claiming a closure the rules do not grant | +| Gate B, the coverage instruction | `NO FINDINGS` only when the branch found none | +| Mechanics, the curve's Majors rationale | rewritten on the pre-ceiling reading | +| Mechanics, the one-contract paragraph | membership widened, with a semantic test a downstream reader can apply | +| (i) when these rules bind | §H | +| the Gate-A section and the gate-prompt template | the two senses of *clean* are separated; the cadence makes revision conditional | +| the profiles section, the lens paragraph | its unchanged-list is scoped to the lens sets | + +**Two sentences are deliberately not edited**, named so nobody looks for them: the "Copy every +record into the squash body" sentence inside the human-exception block, and the "records every +cycle owes" list. This change ships no record. + +**No total is claimed, here or in the target text.** A count over spans that merge and split +disagrees with its own table at the next revision, which is what pass 15 found. The plan counts +what it writes. --- @@ -385,17 +139,17 @@ is the only place it can be checked against the file being changed. **Story acce 5 is satisfied by that list, not by this section**, and nothing here is dropped by being absent here. This table says what happens to each inventoried passage, so the map stays complete at ten. -| Passage | This change | Edits (§4) | +| Passage | This change | Target text | |---|---|---| -| (a) the floor paragraphs | edited — closure sentences trimmed to a pointer, the no-restating prohibition scoped, the per-pass fix command pointed at Mechanics | 6, 7, 8 | -| (b) what a loop absorbs | edited — owns both triggers and the fix set, so every qualification the ordering needs is made **here**, which is what keeps one definition per rule | 9, 10, 11, 12, 13, 14, 15 | -| (c) recognizing clearly stuck | edited — keeps its three-condition reading, stops carrying evaluation order; its precedence sentence moves into the block unchanged, and the third condition itself is widened at this source to admit a re-raised validated dismissal | 16 | +| (a) the floor paragraphs | edited — closure sentences trimmed to a pointer, the no-restating prohibition scoped, the per-pass fix command pointed at Mechanics | §H | +| (b) what a loop absorbs | edited — owns both triggers and the fix set, so every qualification the ordering needs is made **here**, which is what keeps one definition per rule | §B | +| (c) recognizing clearly stuck | edited — keeps its three-condition reading, stops carrying evaluation order; its precedence sentence moves into the block unchanged, and the third condition itself is widened at this source to admit a re-raised validated dismissal | §C | | (d) from pass 4 onward | **no longer edited.** The unavailable-history block moved to the successor with **D10** | — | -| (e) the five tells | edited — the threshold is read after clean completion, and a pointer says what its answer does | 17, 18 | +| (e) the five tells | edited — the threshold is read after clean completion, and a pointer says what its answer does | §D | | (f) the two rules above do not compete | **unchanged.** "The two rules above" still names the absorb rule and the stuck reading; the block sits before both and adds no third rule between them | — | -| (g) Mechanics · Severity, the handed-over question | edited — the unsettled statement and its interim report-and-stop duty are replaced by the answer, in both copies, removing the one deliberate story-path divergence | 19 | +| (g) Mechanics · Severity, the handed-over question | edited — the unsettled statement and its interim report-and-stop duty are replaced by the answer, in both copies, removing the one deliberate story-path divergence | §E | | (h) recording a human exception | **no longer edited.** The answer-record block that was to follow it moved to the successor with **D9** | — | -| (i) when these rules bind | edited — the strict-reading list is added to, not rewritten | 20 | +| (i) when these rules bind | edited — the strict-reading list is added to, not rewritten | §H | | (j) the squash carry | **no longer edited.** It was to name the answer record, which moved; this change ships no record for it to carry | — | **Three reversals are recorded here rather than left as silent narrowings**, because each was @@ -418,7 +172,7 @@ pre-existing wording differences are deliberate and stay, which are not and are **performs the extraction and diff**, passage by passage, against the real files. One divergence is decided here because it is a correctness call rather than a wording one: W's `b3` pointer names "the severity rule" on the inventory's reasoning that W has no Mechanics section, which is -false, so W takes C's wording (§4 item 9). +false, so W takes C's wording (`b3`, target text §H). The same extraction runs a second check within each copy: that `b11` and `b13` as edited say what the block cites them as saying, **comparing the complete predicates and not a shared phrase** — @@ -427,7 +181,6 @@ directions. A condition in the block and not in the source ships two triggers th in the source and not in the block means the block cites a rule it has not read. --- - ## 7. Verification **The mode is read from the story's header at execution**, never from here — the same rule this @@ -520,7 +273,6 @@ repo's most persistent defect. The transport that could carry it left with the r **admitted gap of this change**: unowned, unguarded, open to whoever picks it up. --- - ## 8. AGENTS.md invariants touched - **Invariant 11 — `docs/prompt-standards.md`, all 12 items.** Most at risk: item 6, every @@ -529,7 +281,7 @@ repo's most persistent defect. The transport that could carry it left with the r makes a cycle eligible to close, that a zero-finding pass is clean whatever the floor, and that decline is available only at a membership stop. The exemption was wrong twice over: item 6 admits no "settled elsewhere" clause, and a scaffolded copy cannot reach the story the reasons - were said to live in. §4 item 20 carries its reason in the shipped clause for the same reason. + were said to live in. The unknown-start clause carries its reason inline for the same reason. Then **item 8 (token-lean), which an earlier revision claimed on the wrong ground** — it said the block replaces closure sentences rather than adding beside them, while the block was in fact restating triggers, duties, preconditions and the severity answer that their own @@ -561,7 +313,6 @@ and no bare basename. An abbreviated citation fails the path-existence check a r mechanically, and five of them did. --- - ## 9. Moved, out of scope, and parked **Moved to `docs/superpowers/stories/2026-09-10-record-durability-story.md`** on Daniel's @@ -575,7 +326,7 @@ decision of 2026-09-10, with the evidence in §1: - the rollback reading — **what an open cycle owes when a revert removes the rules it started under**. This change does not answer it and no longer claims to. An earlier revision said the activation paragraph already supplies the transition; that is wrong in a specific way. A revert - of this change removes the ordering, the source edits **and §4 item 20's suspension-binding + of this change removes the ordering, the source edits **and the unknown-start list's suspension-binding extension together**, so the stricter reading such a cycle would fall back to is itself part of what the revert takes away, and it no longer mentions the suspensions that cycle is holding. What remains is the pre-change §5 — the loose ordering this story exists to replace — read by a @@ -593,12 +344,12 @@ line ranges (§4); the parity divergence list and the extraction-and-diff (§6); verification fragment with its counts (§7). Each is work this change still owes; none of it is work a spec can do correctly, because all four are checked against files the plan edits. -**Partial adoption of the narrowed set — answered by §4 item 22, and what that answer is worth.** +**Partial adoption — answered by the one-contract paragraph (target text §G), and what that answer is worth.** The set is mutually dependent, and **the rule that says which edits belong is stated rather than enumerated**: an edit is coupled when the block **cites it or depends on it** — the boundary its clean predicate reads, the sentences that give *clean* its two senses, the triggers it reads, the source rules whose old text the ordering falsifies, and the severity and dismissal rules it cites. -Item 22 carries that description into the live one-contract paragraph, which until now reached +The target text's §G carries a semantic membership test into that paragraph, which until now reached only the nonce, the slots, the provenance line, the curve, the carry rule and the unknown-start semantics. **An earlier revision of this section named the members as a list of item numbers and the list was wrong** — it omitted several edits the block plainly depends on — which is why the @@ -607,14 +358,14 @@ decidable and the numbering does not exist. **Stated as what it is.** That paragraph is **an instruction to the agent**: a project whose text carries some members and not others, or versions that disagree, **stops and has a human complete, -revert or reconcile the adoption before running a gate under it**. Item 22 widens whom that +revert or reconcile the adoption before running a gate under it**. §G widens whom that sentence is about. It is not a guard and not a mechanical check, and this change builds neither — **nothing detects a partial adoption**, and the stop happens only where an agent reads the sentence and acts on it. **What it therefore does not buy, said rather than implied.** A project that adopts the block -without adopting item 22 is not reached at all, which is the partial-adoption case applied to the -rule against partial adoption; the existing paragraph has the same property and item 22 neither +without adopting that paragraph is not reached at all, which is the partial-adoption case applied to the +rule against partial adoption; the existing paragraph has the same property and the widening neither worsens nor repairs it. Nor does it detect a *silent* half-merge in a project that did adopt it — it obliges a stop once someone notices, which is a different thing from noticing. **What it removes is the narrower state this spec previously admitted**: that the coherence rule did not @@ -622,7 +373,7 @@ name this material at all, so an agent reading the sentence and willing to act o to act on. Two earlier revisions got this wrong in opposite directions — one claimed the existing rule already caught the block, and one assigned the guard to the successor story, whose scope is record durability and **excludes the closure ordering by name**. Both named a mechanism that did -not exist; item 22 names a sentence that does, and claims only what that sentence does. +not exist; §G names a sentence that does, and claims only what that sentence does. **Out of scope and parked**, unchanged: diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md new file mode 100644 index 0000000..403f4bb --- /dev/null +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -0,0 +1,476 @@ +# §5 loop-rule consolidation — target text, NOT ACTIVE + +**Date:** 2026-09-11 · **Status:** target text under Gate A, not approved and **not installed** +**Story:** `docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md` +**Design:** `docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md` +**Profile:** read from the story header at every pass, never from here. + +## What this file is, and why it exists + +**This is the §5 text as it will read after this change** — not a description of edits to it. +`CLAUDE.md` and the `/workflow-init` template are **untouched**; nothing here is active, and no +rule below governs any running cycle. + +It was produced at the pass-17 mandatory two-tell stop, on the reading that the cycle's cost was +**four parallel descriptions of one change** — the ordering block, an edit table, the source +sentences and the rationales — which had to agree with each other simultaneously, so a repair to +any one of them fell out of step with the other three. Seventeen Gate-A passes and a Blocker count +flat at 1 for four of them. Collapsing the four into **one concrete text** removes three of them. +The design spec keeps the decisions and their reasons; it no longer re-narrates this text as +future work. + +**This is not a Gate-A close and claims to be none.** §5's clearly-stuck exit says surfacing does +not close a cycle, no pass is credited clean and the findings stay open; that is exactly the state +this file was written in. Gate A stays open and is re-aimed at this text, with pass 17's nine +findings as the basis. The rules are not activated by this file. Gate B reviews the real +implementation diff later, and this file is not that diff. + +**How to read a section.** Each is marked **NEW**, **REPLACED** or **CARRIED**. A CARRIED passage +is reproduced from the current state unchanged, because the block cites it and a reader has to see +what it says; it is here to be read, not to be edited. Both prompt copies take every NEW and +REPLACED section **byte-identical**. + +**What is deliberately not here:** the record-durability material (successor story), and the +per-condition disposition, parity diff and verification fragments (the plan, against real files). + +--- + +## A. The closure ordering — NEW + +Sits immediately before "**What a loop absorbs, and what stops it**" in both copies. + +``` +**How a cycle ends — one ordering, stated here and referenced everywhere else.** A pass is read +in a fixed order, because every rule bearing on one decision — may this cycle close — otherwise +qualifies the others and the ranking survives only in a reader's head. **This paragraph owns +exactly six things**: the evaluation order, closure and its eligibility, the hold a surfaced +finding places and what discharges it, the composition of several suspensions, the pairs that +cannot co-occur, and — as a stated exception, because precedence is evaluation order — the +clearly-stuck precedence sentence quoted into it below. **Everything else it names it cites**: +the scope triggers, the assigned fix set, every severity rule, and each closure precondition +keep their one definition in the paragraph that owns them, and a reader who finds one of *those* +defined here has found a defect. + +**What a pass is read from.** Every finding-derived predicate reads the validated findings file +**or files** of the logical pass as **the concatenation of their finding lines after each file +has been validated separately** — a `full` Gate-B pass has two, one branch alone is already an +incomplete pass, each file's terminator is not a finding line, and a branch whose body is +`NO FINDINGS` contributes an empty sequence rather than a line. **Which severity field each +predicate reads is settled in Mechanics · Severity**, which is where that split lives and is not +repeated here. Beyond the findings, closure reads the **derived floor**, the final-acceptance +preconditions the floor section states, the **resolve duty's standing over this cycle**, and any +**hold still standing**; **in a Gate-B cycle it also reads the evidence entry's revalidation +rule** — a changed entry meaning the clean pass no longer covers what is being committed — which +is a Gate-B precondition because the entry is about the diff being committed and a Gate-A cycle +produces none. The scope triggers read the **current assigned fix set** as the absorb paragraph +defines it, and the answers already given; the clearly-stuck reading adds its own coverage +judgement. **A line in one branch file and a line in the other are distinct findings for holds +and answers**, so a `full` pass asks twice rather than risk resuming over one it never asked +about. **Within one running cycle an answer binds to the finding or question as the pass that +raised it recorded them** — which is what an agent running the cycle can do with nothing written +down. Recognising the same finding or question across a lost session needs a record these rules +do not ship. + +**First, clean completion, and eligibility is its own test.** §5 uses *clean* in two senses and +now says which is which: a **clean findings file** is the `NO FINDINGS` signal the protocol +defines, and a **clean pass** is the predicate here, read on the logical pass with every required +branch file combined, so one branch's clean file never establishes a clean pass. **A pass is +clean** when its findings carry no in-set Blocker or Major at effective severity and **no +scope-stop trigger** — the two the absorb paragraph defines, read there and not redefined here, +each already carrying the qualification an answer given in this cycle puts on it. A pass with +**zero** findings is clean whatever the floor, because a floor buys further looks at an artifact +that keeps yielding findings, and one yielding none has already given what those looks were for. + +**Eligibility is exactly this and nothing more: a clean pass at or above the derived floor, or a +zero-finding pass.** It is a property of the pass. **Closure is eligibility plus every closure +precondition holding plus the cycle's closing act**, and the preconditions are properties of the +*cycle*, not of the pass — so an eligible pass whose preconditions are unmet closes nothing, and +is not thereby made unclean. Keeping the two apart is what lets the order decide the case where +they disagree, which the branches below do. + +**The closing act, by cycle kind.** A **Gate-B** cycle closes with the closing amend +Mechanics · Finishing the cycle describes. A **Gate-A** cycle has no WIP snapshot and no amend: +it closes on **the author's recorded acceptance of the revision the clean pass reviewed**, which +is a commit that in the ordinary case already exists — the pass reviewed a committed revision — +so the closing record goes in that commit where it is still the branch tip and in the next commit +on the branch where it is not. **No new revision of the artifact is made to close a Gate-A +cycle**, because a new revision is one no pass has reviewed. Nothing is closed before that act, +and a profile, cited set or — in a Gate-B cycle — evidence entry that changes in between still +gates it. **A commit the hook reads as cycle-closing is a separate matter**: a non-`WIP` commit +mid-cycle makes the hook read the cycle as closed and discards the passes counted so far, which +is an observation about the counter — the cycle itself stays open until the conditions above +hold, so an accidental commit destroys the pass credit and closes nothing. A plateau or tells on +an eligible pass go into the closing report and never block it, because reporting "will not +converge" on a converged loop is a false report. **No other pass outcome makes a cycle eligible +to close**, because every other pass leaves a required repair, a hold or a question outstanding — +and closing over any of those is the failure this ordering exists to prevent. The one termination +that is not a pass outcome is the Gate-B triviality skip, which runs no passes and is outside +this ordering. + +**Second, only a pass that is not a clean completion can suspend** — that order is what makes +"clean completion outranks the two-tell stop" executable rather than asserted. Three suspensions, +by the names their paragraphs use and read by those paragraphs: the **scope stop**, raised by +either trigger above — a **membership stop** by the first, a **question stop** by the second; the +**clearly-stuck exit**; and the **two-tell stop**. A suspension waives nothing. Any non-empty set +of them can apply to one pass: **one surface, every reason reported, every question asked**, +because a reason left out is a decision made by omission. A finding the clearly-stuck reading +surfaces that also carries either trigger takes the scope stop's answers at that same surface, so +it is not asked twice; the two-tell stop surfaces tells and not a finding. + +**Third, a pass that neither closes nor suspends continues** — the loop runs another pass on the +**current** artifact, revised where the severity and scope rules require a repair and unrevised +where they do not. **An eligible pass with an unmet closure precondition lands here**: clean +completion did not close it, and being clean it cannot suspend, so the loop continues on whatever +the unmet precondition requires — most often a repair still owed from an earlier pass. A +below-floor clean pass lands here too, **only where no suspension applies to it**; where one does, +the second branch has already taken it, because clean completion did not close the pass and only +closing outranks a suspension. So does a pass whose only findings are Minors and Nits, which are +collected and never iterated and may leave nothing to revise. It is a branch and not an inference, +because "does not close" read alone says nothing about whether to run again. + +**The four standing duties, classified.** The **derived floor** is a **precondition on closure**: +it gates closing, discharged by the count of valid logical passes reaching it with the last of +them clean, or by the zero-finding exit. The **Blocker/Major-resolve duty** is a **precondition on +closure**: Mechanics · Severity, scoped to the assigned fix set, states what it demands and what +discharges it — a repair or a validated dismissal — and what a dismissal is, all at that source. +**It is discharged per finding and tracked across the cycle, never inferred from a later pass.** A +findings file establishes the **inventory** of what that pass found and not the resolution of +anything, so a later pass that does not mention an earlier in-set Blocker or Major says nothing +about whether it was repaired or dismissed; reading its absence as discharge would let an omission +close a cycle. **The duty is not a second test on whether a pass is clean, and the two are not run +together.** A pass is clean on its own findings. Stated as the case that separates them, because a +reader who conflates them decides it wrongly: **pass 1 raises an in-set Major; it is neither +repaired nor dismissed; pass 2 finds nothing.** Pass 2 **is** clean, and at or above the floor it +is **eligible** — and the cycle **still cannot close**, the duty being unmet; it continues on the +third branch until that Major is discharged. What the open Major does *not* do is make pass 2 +unclean. The **hold** a surfaced finding places on closure **participates in the ordering**: it +gates closing while it stands, and is discharged by the answers that surface requires. **It +attaches to every surfaced finding, whichever suspension surfaced it** — clean completion creates +none, because it wins before anything is surfaced. **No-clean-credit** — no pass carrying a +scope-stop trigger is credited as clean — also participates, and is a fact about that pass that +nothing discharges, a later pass being judged on its own findings. It is not a second test beside +the clean predicate but that predicate's second half, which is why it is stated in its words. + +**What a suspension asks, and what ends it.** **A hold ends when every answer its finding requires +has been given, in whichever direction each is given** — one **scope-stop** answer for a +single-trigger finding, both for one carrying both. That is **one rule with two parts**, how many +answers and which way each may go, and neither is a test the other has to pass. Where a health +suspension applies to the same pass, its shared continue-or-stop answer is **additional** to those +and not counted among them, the health readings asking about the loop rather than about this +finding. **The two health readings differ in what they surface, and therefore in what they hold.** +The **two-tell stop surfaces tells and no finding**, so it creates no hold; what it leaves +outstanding is its own continue-or-stop question, which the composition rule below holds the cycle +on until it is answered. The **clearly-stuck reading surfaces findings** — **the findings of the +pass being read that satisfy its regeneration condition, and only those**; earlier members of a +regeneration chain that were repaired or dismissed are history the reading consults and never +findings it re-surfaces, so no discharged finding takes a second hold. Each surfaced finding takes +a hold like any other. **Where one finding is surfaced by both, it carries two hold components and +each is discharged by its own answer, which is what keeps D4 and D5 exact**: the **membership** +component ends on the membership answer in either direction, a decline releasing it as **D5** +requires; the **clearly-stuck** component ends on the reading's continue-or-stop answer. Neither +answer discharges the other's component, and where the finding carries no scope-stop trigger the +continue-or-stop answer is the only one its surface asks for and discharges the only component +there is. **Resumption is still the composition rule's**, which waits for every outstanding +answer. At a **membership stop** the answer is **accept**, the finding joining the fix set where +Mechanics · Severity governs it, or **decline**, the finding staying outside and binding so for +the rest of this cycle. A later answer that contradicts a decline **does not reverse it**: the +decline **remains binding** and the contradiction is **surfaced to the user as information**, +changing neither membership, nor the cycle's state, nor any outstanding question — **whether the +loop resumes is decided by the composition rule below and by nothing here**, so a contradictory +answer is never itself a resumption and cannot step past a hold or a health question still +awaiting its own answer. Nothing here turns one answer into another, since that would let a +finding be moved out of the set and back into it to escape what it owes inside it. **There is no +withdrawal inside the cycle that declined**, because **D7** binds a decline for the remainder of +its cycle and admits no exception; a reconsideration is a later cycle's, where D7 gives the +decline no effect at all and the finding takes the ordinary route. Either answer is an +**explicit, attributable decision on that specific finding** — never silence, never a general +remark about scope, never inferred, because a fix set changed by inference is a fix set nobody +chose. **Membership is answered against the set as the absorb paragraph fixes it for the pass that +raised the question**: a later broadening is a new fact the **next** pass reads and never +discharges a standing hold, a hold discharged by a scope change being a hold nobody answered. At a +**question stop** the answer is the user's decision on the question and membership does not +change; an out-of-set finding that opened one is a membership stop as well. **Decline is available +only at a membership stop**, that being the only stop whose question is whether a finding belongs +to the set. The **clearly-stuck and two-tell readings** ask **continue or stop**. **Continue +consumes the reading that raised the suspension**: a further health suspension needs that reading +recomputed over a pass run after the answer, which is new data — so continue produces a distinct +next state, and the same reading cannot return the same stop unanswered. It permits an +**unrevised** artifact **only where no repair is owed**; where effective severity or scope requires +one, that repair comes before the post-answer pass, since a pass run over an unrepaired in-set +Blocker or Major spends a look on text the rules already say must change. **Stop parks the +cycle**: open, not running, spending no passes, restarted only by an explicit later continue — a +distinct state from the suspended-awaiting-answer one it was in before the answer. **That continue +restarts the cycle and never skips an answer**: where any question the suspension raised is still +outstanding, it returns the cycle to suspended-awaiting-answer, and only once every answer the +composition rule requires has been given does the next pass run. So a cycle parked with an +unanswered membership or question stop cannot be continued into a pass, and cannot sit parked with +no transition either — the continue is always available and always moves it. Nothing a parked +cycle wrote is a closing commit, and a parked cycle nobody restarts is a human's to resolve, +exactly as the nonce rules already say of open cycles. + +**Composition, and what cannot happen.** Every **question** is answered on its own and the loop +resumes only when every answer resumes it — accept or decline at a membership stop, a decision at +a question stop, continue at the health readings; one stop answer parks the whole suspension, +because a loop resumed over an unanswered question decides it by running. **The clearly-stuck and +two-tell readings raise one question between them, not two**, both asking continue or stop, so one +answer carrying every reason ends both — an instance of the sentence before it, not an exception. +**Two pairings cannot occur**, and no rule ranks them: clean completion and a **scope stop**, since +that stop's triggers are the clean predicate's own second half, so a pass raising one is not clean; +and a zero-finding pass and any suspension, since it has nothing to surface, nothing regenerating, +no cluster and no require↔withdraw pair. **Clean completion and the clearly-stuck exit can**, and +the overlap is admitted rather than argued away: the two read different severity fields, as +Mechanics · Severity sets out, so an **in-set** Blocker or Major the ceiling demotes below Major, +or an **out-of-set** one this cycle has declined, can regenerate across passes on a pass that is +clean. **The order decides it and no new rule is needed.** The clearly-stuck paragraph's own +precedence sentence is stated here rather than there, because precedence is evaluation order and +this paragraph is where evaluation order is stated once; its opening words point back to that +paragraph, which is where the reading itself lives. That third condition is what makes a plateau +rather than a finish, and it is why **a clean completion takes precedence over this exit**: a +Blocker/Major-free pass **at or above the floor** has satisfied the clean-final-pass rule — +collect the Minors and Nits and close — and reporting "will not converge" on a converged loop is a +false report. Below the floor the pass **suspends**, clean completion having not closed it. +``` + +--- + +## B. Passage (b) — what a loop absorbs — REPLACED, in part + +Three sentences change; the rest of the paragraph is carried. The changed sentences, final: + +**`b7`, the assigned fix set.** +``` +**The assigned fix set is fixed before the pass you are answering: it is the union of the scope +every approved story or plan governing this change assigns to this cycle, plus repair obligations +you already accepted in earlier passes, minus every finding this cycle has declined.** +``` +*Why:* the singular "the approved story or plan" leaves a cycle governed by several with no set +at all, and a declined finding has to leave the set or the resolve duty reaches it. + +**`b8`, the membership test.** +``` +A finding is in-set when repairing it stays inside **the assigned fix set as `b7` computes it** — +never merely because it arrived in the current pass, which would put every new finding in the set +by definition and leave the boundary deciding nothing. +``` +*Why:* tested against "that scope" independently, a broadened scope puts back a finding `b7` +keeps out. + +**`b11`, the membership trigger.** +``` +**A correction that leaves that set stops the loop like any other out-of-scope finding** — except +one this cycle has already declined, which is outside the set by that decision and raises no +trigger — even when it opens no new question at all, and it resumes the moment the user says +whether the set now includes it. +``` +*Why:* unqualified, this and the clean predicate decide a re-raised declined finding in opposite +directions. + +**`b12`.** Gains, at its end: +``` +Any answer ends the hold; **what the pass does next is the closure ordering's**, not this +sentence's. +``` + +**`b13`, the question trigger.** *New* is qualified: +``` +…opens a **new structural or contract question** — new meaning not already answered in this cycle, +so an answered question raised again stops nothing… +``` + +**`b16`, novelty over ancestry.** The live sentence says "the new question wins and the loop +stops"; it gains one clause: +``` +**When a finding is both** — it corrects the last correction *and* opens a new structural or +contract question — **the new question wins and the loop stops**: novelty overrides correction +ancestry, because absorbing on ancestry is exactly how a contract decision gets made without +anyone choosing it. **Novelty overrides ancestry and nothing else: where the finding is also out +of set, both triggers hold and both answers are owed**, since a question answered about a finding +nobody placed in or out of the set leaves its membership decided by default. +``` +*Why (pass 17 finding 4):* without it an agent follows the surviving source, asks only the +structural question, and absorbs out-of-set work silently. + +**Carried unchanged:** the paragraph's opening, `b3` becomes a pure pointer (below), and every +sentence not named here. + +--- + +## C. Passage (c) — recognizing clearly stuck — REPLACED, from the third condition + +``` +…and **Blocker or Major findings that keep regenerating across genuine repair attempts**, each +round's fix producing the next — **or a finding the author has validly dismissed that the reviewer +re-raises across passes**, the re-raise standing in for the regenerating fix, since a dismissal +gets no repair and produces none and a false positive that returns every pass would otherwise +leave the cycle unable to close and unable to suspend. That third condition is what makes a +plateau rather than a finish. **Where this reading and a clean completion both apply, the closure +ordering decides it** — the precedence sentence lives there, because precedence is evaluation +order. +``` + +*Everything after that in the live paragraph — the precedence sentence and the below-the-floor +sentence — moves into the block above, word for word, which is what satisfies **D3**.* + +--- + +## D. Passage (e) — the five tells — REPLACED, in part + +**`e7`, the threshold.** Gains one clause: +``` +**Any two present makes stop-and-surface mandatory, not discretionary** — read **after** the +clean-completion branch of the closure ordering, which outranks it (**D2**) — and you report the +tells and hand the decision to the user… +``` + +**A pointer is added** at the end of the passage: +``` +**What the answer does** is the closure ordering's: this stop is a **suspension**, it closes +nothing, and continue or stop is answered there. +``` + +--- + +## E. Mechanics · Severity — REPLACED, in part + +**The resolve duty, which is the only statement of it.** +``` +- **Severity:** Blocker (wrong/unsafe/breaks invariant) · Major (design flaw → rework) → both + must resolve, **for every finding in the assigned fix set as the absorb paragraph computes it**. + **A finding is resolved by a repair or by a validated dismissal** — the author's judgement, + carrying the one-line why this section already requires, that the finding is not true of the + artifact. A dismissal does not rewrite the pass that found it and the later clean pass is still + owed; **a dismissal is not a decline**, a dismissal saying the finding is false and a decline + being the user's decision that a **true** finding stays outside the set. Minor · Nit → collect, + never iterate. +``` + +**The handed-over question, replaced by its answer.** +``` + **The demotion changes what a cycle must resolve, never what it observes.** **Every loop-health + reading observes the findings as the reviewer produced them, before the ceiling is applied** — + so a demoted finding still counts in the finding total and in its cluster, and a Blocker demoted + to Minor is still a Blocker to the curve and still regeneration to the clearly-stuck reading's + third condition. Where such a reading uses severity at all it takes the **reader-normalized + pre-ceiling severity**, which is what the Reader paragraph above already produces from the + findings file — case-folded, and a non-empty unrecognized token read as `MAJOR` — never the raw + token, so a finding written `IMPORTANT` enters the curve as a Major. **The ceiling is applied + after that and only to what the cycle owes**: cleanliness and the resolve duty read the + effective severity, so the same finding can be a Major to the curve and a Minor to the fix set, + and that difference is the point rather than a discrepancy. The line is **what the cycle owes + versus what it observes about itself**, which is why no list of readings has to be kept complete + here. Two reasons for the split. The curve must stay derivable from the findings files alone — + the finding total counts finding lines and the Blocker and Major series count the lines whose + normalized severity is each, which is the only thing that makes a self-reported curve checkable; + the subject clusters use no severity at all, being a judgement per finding that no count + reproduces. And the demotion is the author's judgement about **the finding's repair severity**, + never about which findings the fix set contains — a separate predicate the absorb paragraph + defines, and one this must not be read as touching; a loop spending passes on findings the + author keeps demoting is exactly what the prose-cluster tell exists to surface, and lowering the + counts by that same judgement would hide it. +``` +*Why the last sentence of the previous draft went (pass 17 finding 6):* it claimed `IMPORTANT` +counts for the curve "exactly as it counts for the pass", which is false wherever the ceiling +demotes it — the two counts are meant to differ. + +--- + +## F. The three standing sentences this change falsifies — REPLACED + +Each is a live sentence that the block makes wrong. All three are **known contradictions** and +none is deferred. + +**1. The `WIP:` naming warning** (Mechanics · `baseSha`). +``` +A pre-review snapshot named anything else reads to the hook as a real commit: the hook treats the +cycle as closed and **discards the passes you just accumulated**, while the cycle itself stays +open until the closure ordering's conditions hold. The cost of the mistake is the lost pass +credit, not a close nobody intended. +``` +*The sibling sentence in the profile-change paragraph needs no edit* — it already says such a +commit "reads **to the hook** as the cycle closing", which claims no closure. + +**2. The Gate-B coverage instruction** (Gate B section). +``` +Same coverage rule as Gate A: put "report every finding with severity and confidence; write +`NO FINDINGS` only when the branch found none" in `additionalContext`, with the same one-line +format. You filter to Blocker/Major, Codex never does. +``` +*Why (pass 17 finding 5):* "say `NO FINDINGS` if clean" plus the ordering's clean-pass-with-Minors +tells a reviewer to emit an empty file over real Minors. + +**3. The curve's Majors rationale** (Mechanics, the per-pass curve). +``` +**Majors are recorded as well as Findings and Blockers**, because the three series are read +before the ceiling and the mix among them is what a later reader compares; the ceiling moves what +a cycle owes and leaves these counts alone, so totals and Blockers alone could not show even a +change in the mix. +``` +*Why (pass 17 finding 7):* the live rationale says the severity rule "moves the Blocker/Major +line rather than the total", which contradicts the pre-ceiling decision above. + +--- + +## G. The one-contract paragraph — REPLACED + +``` +**These rules and records are one contract, and a partial adoption breaks it.** The nonce, the +slot naming, the provenance line, the curve, the carry rule, the unknown-start activation +semantics **and the closure ordering together with every rule it reads** depend on one another, +and the requirement is that the adopted definitions **agree**, not merely that all of them are +present. **Membership is decided by a test a reader can apply to the text in front of them, with +no list to consult: a live rule belongs to this contract when changing it would change an input +the closure ordering reads, which branch a pass takes, what a hold is or what discharges it, or +whether a cycle may close.** A curve without a cycle field cannot be attributed, a slot rule +without a nonce has nothing to key on, a carry rule naming records a project does not produce is +inert, and a clean predicate without the fix-set boundary it reads decides membership by accident. +**A project whose text carries some of them and not others, or carries all of them in versions +that disagree, stops and has a human complete, revert or reconcile the adoption before running a +gate under it.** +``` + +**What this is, stated so nothing reads it as more:** an **instruction to the agent**, over a +wider membership and now with a test that a downstream reader can actually apply — the earlier +draft's "every source edit it cites or depends on" named an edit set that exists only in this +repository's spec, so a downstream agent willing to obey it had nothing to check against (pass 17 +finding 8). It is **not a checker**; nothing mechanical detects a partial adoption, a project that +does not adopt this paragraph is not reached by it, and it obliges a stop once someone notices +rather than doing the noticing. + +--- + +## H. The remaining source edits — REPLACED, and each is one sentence + +- **the gate-prompt template's clean sentence** — a `NO FINDINGS` file is renamed a **clean + findings file**, which is what it always described. +- **the Gate-A clean-signal sentence** — the `NO FINDINGS` signal becomes what lets a pass be read + as clean without inspecting it, rather than the only route to clean; a pass carrying Minors + alone is clean and could never produce that file. +- **the lens paragraph's unchanged-list** — scoped to **the lens sets**, since the ordering + falsifies its claim that the Blocker/Major filter and the clean-final-pass rule are unchanged. +- **the Gate-A cadence** — "validate, revise, re-run" becomes conditional on a repair being + required, since unconditional it tells a Minor-only pass to manufacture the repair the severity + rule forbids. +- **`a13`** — "Every other rule stated here about how a cycle closes stands as written, and none + of them is restated" is scoped to that paragraph, which is what it was written to police; + categorically false once the ordering exists. *(It wraps across three lines in each copy.)* +- **`a16`** — "fix Blocker/Major after each" points at Mechanics · Severity instead of restating + an unscoped version. +- **`a17`–`a19`** — the floor paragraph's closure sentences are trimmed to point at the ordering. +- **`b3`** — becomes a pure pointer to Mechanics · Severity, stopping its restatement of the four + severity actions; W's pointer also changes target to match C. +- **`b17`–`b18`** — they name their stop a suspension and defer to the ordering. +- **the unknown-start strict-reading list** — gains **every suspension binding, since unknown + starting rules cannot waive an open hold**, with the reason carried inline. + +--- + +## I. What this text does not settle + +- **Minor, collected and open (pass 17 finding 9 is resolved above; this is the remainder):** + nothing in this file establishes that the edit set is complete. It is the sites known at pass + 17. **The plan sweeps both copies against the rule** — a live sentence the ordering falsifies + gets edited — and prints what it found; a spec cannot establish that claim against text the same + change rewrites, which is the mechanism that produced a finding at passes 15, 16 and 17. +- **Partial adoption is instructed against, never detected.** §G says so in its own words. +- **Which exit a cycle took is not observable from history.** The transport left with the record + (successor story) and no story has taken it. An admitted gap, unowned. From 96c3611d6cf3f1c087c0df4272e8c4b8549c36fc Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Fri, 11 Sep 2026 14:19:06 +0200 Subject: [PATCH 030/181] =?UTF-8?q?docs(specs):=20target=20text=20?= =?UTF-8?q?=E2=80=94=20six=20changed=20sentences=20in=20passage=20(b),=20n?= =?UTF-8?q?ot=20three?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The precheck's COUNT note caught a stated count disagreeing with its own enumeration before a read pass was spent on it: passage (b)'s header said three sentences change and the section lists six (b7, b8, b11, b12, b13, b16). Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../specs/2026-09-10-loop-rule-consolidation-target-text.md | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index 403f4bb..c757276 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -235,7 +235,8 @@ false report. Below the floor the pass **suspends**, clean completion having not ## B. Passage (b) — what a loop absorbs — REPLACED, in part -Three sentences change; the rest of the paragraph is carried. The changed sentences, final: +Six sentences change — `b7`, `b8`, `b11`, `b12`, `b13`, `b16` — and the rest of the paragraph is +carried. The changed sentences, final: **`b7`, the assigned fix set.** ``` From ef9a50480edde6035a86e0cd7aef0ec3be812159 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Fri, 11 Sep 2026 14:36:27 +0200 Subject: [PATCH 031/181] =?UTF-8?q?docs(specs):=20apply=20Gate-A=20pass=20?= =?UTF-8?q?18;=20finish=20the=20target=20text,=20=C2=A7H=20becomes=20real?= =?UTF-8?q?=20text?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Pass 18 is the first pass on the target text. 15 findings, ZERO Blockers -- the first zero-Blocker pass since pass 7. One tell of five (the finding count rose 9 -> 15); the Blocker count fell, the cluster is product, no pair. No mandatory stop. WHY THE COUNT ROSE, and it is not the usual reason. Findings 6, 7 and 8 each name a specific standing sentence by file and line -- b12's immediate-resumption clause, c18's no-clean-credit, the surfacing sentence's "the loop resumes on whatever the user decides". That is the same mechanism that produced one finding at pass 15, one at 16 and three at 17, but those said "the edit table does not name a passage" and these say "this sentence and that sentence disagree, here are both". The restructure made the mechanism decidable instead of recurring, which is what it was for. AND FINDING 1 SAYS THE RESTRUCTURE WAS HALF-DONE, correctly. Sections A-G were concrete text and drew concrete findings; §H was still paraphrase -- "is renamed", "becomes conditional", "are trimmed" -- so no reviewer could read the future prompt there. Finding 2 caught the consequence: §B said six sentences change and listed b3 as carried, while §H replaced b3 and b17-b18. Two sections of one artifact giving incompatible instructions. §H is now twelve literal replacements, each the final wording the plan installs, including the three standing sentences findings 6-8 named. §B and §H are split cleanly: b7, b8, b11, b13, b16 in §B; b3, b12, b17-b18, c18, a13, a16, a17-a19 and the five non-inventoried sites in §H. Verified: no id appears in both. The rest, applied: - the suspension gate reads INELIGIBILITY, not cleanliness. A clean pass below the floor is not a clean completion and can suspend; an eligible pass cannot. Cleanliness alone never decides it (finding 4). - the exhaustive reason for ineligibility gains the unmet floor, since a below-floor Minor-only pass leaves nothing outstanding and is still ineligible (finding 5). - a declined finding is no longer a route into the clearly-stuck reading: that condition admits regeneration across repairs and a re-raised validated dismissal, and a decline is neither (finding 9). - "a Minor to the fix set" becomes "a Minor at effective severity"; the ceiling moves what a cycle owes, never whether a finding is in the set (finding 10). - §G's membership test gains the production, identity and transport of records, so it stops excluding the squash-carry rule the same paragraph names as a member (finding 11). - the design spec's ownership sentence is pulled to the same six-part boundary the target text's §A states, and says the two must not drift (finding 12). - a re-raised valid dismissal stays discharged for the resolve duty; the recurrence creates the clearly-stuck hold alone (finding 14). - health observations on an eligible-but-blocked pass go to that pass's ordinary report, there being no closing report (Minor 13). - the marker legend is scoped to A-H and §I is labelled as this file's metadata (Nit 15). Minor 3 (a17-a19 and b17-b18 are multi-sentence spans, so "each is one sentence" was false) is answered by §H no longer claiming it. No Minor or Nit got a round of its own. Verified: both prechecks exit 0; fences balanced at 54; every COUNT note checked against its enumeration; §B and §H share no condition id. Target text 476 -> 562, design 385 -> 389. No gate: docs/**.md plus the cycle's working record, prose-exempt under section 5. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-18.md | 16 ++ .../gate-a-spec-awsf1ec771-resume.md | 48 ++++- ...26-09-10-loop-rule-consolidation-design.md | 10 +- ...-10-loop-rule-consolidation-target-text.md | 181 +++++++++++++----- 4 files changed, 203 insertions(+), 52 deletions(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-18.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-18.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-18.md new file mode 100644 index 0000000..399a237 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-18.md @@ -0,0 +1,16 @@ +MAJOR | high | opening "§5 text as it will read" | The artifact says it contains the final §5 text and that CARRIED passages are reproduced, but §H supplies only paraphrases such as "is renamed", "becomes conditional", and "are trimmed", while the carried source around them is absent | Gate A cannot review the actual future prompt, exact replacement wording, byte parity, or prompt-standard conformance, and the plan is left to invent behavior that this text claims is already decided | Replace each §H paraphrase with the literal final replacement passage in its standing context and reproduce any carried passage needed to compose the final text, without restoring an edit table or total +MAJOR | high | §B "Six sentences change" | §B says only b7, b8, b11, b12, b13, and b16 change and "the rest" is carried, yet §H also replaces b3 and b17–b18; §B even puts "b3 becomes a pure pointer" under "Carried unchanged" | The same artifact gives incompatible instructions about which source sentences survive, so implementation and parity can omit or duplicate three behavior-bearing edits | Reconcile §B with §H and classify b3 and b17–b18 as replacements in the single concrete passage +MINOR | high | §H "each is one sentence" | The heading claims every remaining source edit is one sentence, but its own enumeration includes a17–a19 and b17–b18, which are three- and two-sentence spans | The stated edit granularity is mechanically false and can make the plan count or replace the wrong unit | Remove the one-sentence claim or split the grouped spans into literal sentence-level replacements +MAJOR | high | §A "Second, only a pass that is not a clean completion can suspend" | The block says an eligible pass with an unmet precondition cannot suspend "being clean", but then says a below-floor clean pass can suspend; cleanliness therefore both forbids and permits suspension, with eligibility silently doing the real work | A below-floor clean pass carrying two tells can be sent to either suspension or continuation, so the requested ordering does not decide that branch | State the suspension gate in terms of the intended predicate, apparently ineligibility, and remove the contradictory claim that cleanliness itself bars suspension +MAJOR | high | §A "No other pass outcome makes a cycle eligible" | Its reason says every other pass leaves a repair, hold, or question outstanding, but the third branch expressly includes a below-floor pass whose only findings are Minors or Nits and "may leave nothing to revise" | The text gives a false reason for ineligibility and can make an agent invent a repair or hold instead of recognizing that only the floor remains owed | Add an unmet floor to the exhaustive reason or narrow the claim so the Minor/Nit-only below-floor case is truthful +MAJOR | high | §B b12 "it resumes the moment" | The standing b12 sentence, which wraps across CLAUDE.md:205–208 and workflow-init.md:412–415, says the loop resumes the moment membership is answered; appending "what the pass does next is the closure ordering's" does not remove that command, while §A says a concurrent health or question answer can still hold or park the cycle | A membership answer can either resume immediately or remain suspended, violating D1 and D4 composition on the exact membership-plus-clearly-stuck case | Replace the immediate-resumption clause itself with a pointer saying the membership answer ends only its scope-stop component and resumption follows the composition rule +MAJOR | high | §C carried "no pass is credited as clean" | The standing surfacing sentence, wrapping across CLAUDE.md:243–244 and workflow-init.md:446–447, remains un-replaced even though §A permits a below-floor clean pass to take the clearly-stuck suspension and the design explicitly says c18 is replaced | The same surfaced pass is both clean and forbidden clean credit, so later eligibility and floor handling depend on which passage the agent follows | Replace c18 at its source with wording consistent with per-pass cleanliness and the ordering, preserving only the resolve and hold duties that still apply +MAJOR | high | §C carried "the loop resumes on whatever the user decides" | The same standing sentence remains live, but §A says only continue resumes and stop parks the cycle | A clearly-stuck stop answer both resumes and parks, defeating the distinct next state required by acceptance criterion 4 | Replace the carried resumption clause with a pointer to the ordering's continue-versus-stop transitions +MAJOR | high | §A "an out-of-set one this cycle has declined, can regenerate" | The overlap example says a declined out-of-set finding can satisfy the clearly-stuck reading on a clean pass, but §C's third condition admits only regeneration across genuine repair attempts or a validly dismissed finding, and §E expressly says a decline is not a dismissal | A re-raised decline can either trigger clearly-stuck suspension or remain outside the set and follow the clean branch, so D7's idempotent later-pass behavior is not decidable | Remove the declined-finding example from the overlap or explicitly add and justify a third-condition route consistent with D7 +MAJOR | high | §E "a Minor to the fix set" | The severity paragraph says the same finding can be "a Major to the curve and a Minor to the fix set", then says the ceiling never changes which findings the fix set contains | The wording conflates severity with membership and can make an agent remove a demoted but still in-set finding, changing scope triggers and the resolve duty | Say "Minor at effective severity" or "Minor to what the cycle owes" while keeping fix-set membership unchanged +MAJOR | high | §G "Membership is decided by a test" | The semantic test includes only rules whose change affects an ordering input, branch, hold, discharge, or closure, yet the paragraph explicitly names the squash carry rule as a contract member and that post-close transport changes none of those things | A downstream reader applying only the promised test excludes one of the paragraph's own named members, so partial adoption of the carry contract is accepted instead of stopped | Extend the semantic test to changes in the production, identity, or transport of contract records, or otherwise state a single test that includes every explicitly named member +MAJOR | high | design §2 and §8 "defines no ... duty" | The design still says the block is authoritative only for evaluation order and closure and cites every other rule, while target §A now expressly owns the surfaced-finding hold and discharge, suspension composition, impossible pairings, and clearly-stuck precedence | The decisions document and target text disagree on authority, so the plan can move or duplicate rules that the target says must live in the block | Update the design's ownership statements to the same six-part boundary stated in target §A +MINOR | high | §A "plateau or tells on an eligible pass go into the closing report" | Eligibility is separately defined as insufficient for closure, and the block later sends an eligible pass with an unmet precondition to continuation, where no closing report exists | Health observations on an eligible-but-blocked pass have no truthful report destination even though the text says they never block | Scope this sentence to a pass that actually closes and direct observations from an eligible-but-blocked pass to its ordinary pass report +MAJOR | medium | §C "validly dismissed ... re-raises across passes" | The text does not say whether a re-raised valid dismissal remains discharged for the resolve duty: it calls the current recurrence a surfaced finding that takes a hold, says discharged predecessors never take a second hold, and defines resolution per finding | A repeated false positive can either require a new dismissal on every pass or remain resolved with only a health hold, so the requested idempotency trace and closure precondition differ | State explicitly whether the prior valid dismissal continues to discharge the re-raised finding and exactly which new hold, if any, the current recurrence creates +NIT | high | opening "Each is marked NEW, REPLACED or CARRIED" | Section I has no such marker, and no top-level section is marked CARRIED despite the legend's unqualified claim | The artifact's reading instructions are literally false and leave the status of §I unclear | Scope the claim to proposed-text sections A–H and label §I as non-shipping metadata, or mark every section consistently +END OF FINDINGS (15 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index 9000a84..1911799 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -44,7 +44,53 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 15 | 17cf2e9 | 10→**12** | 1→**1** | 5→**7** | yes | **TWO-TELL STOP — mandatory, surfaced to Daniel.** B+M 6→8. **Six of twelve regenerate from pass 14's own repairs** (3, 4, 6, 7, 10, 11). All 12 held open; session 01a08fb2-6e8e-73d0-9d2b-8fb92d1571dd | | 16 | 5d12884 | 12→**8** | 1→**1** | 7→**4** | yes | **B+M 8→5, lowest of the cycle.** One tell (Blockers flat). Five B/M applied; Minors 7 and Nit 8 collected. Findings 2 and 4 are the spec disagreeing with *standing* §5 text, not with itself; session 01a08fff-fadf-72f2-994d-618048bea473 | | 17 | c2d5fe1 | 8→**9** | 1→**1** | 4→**7** | yes | **TWO-TELL STOP — mandatory, surfaced to Daniel.** B+M 5→8. **Three findings (4, 5, 7) are one mechanism on its third pass**: a standing §5 sentence the block falsifies that §4 does not name. All 9 held open; session 01a0901b-ecf3-73b2-886c-5b3c7dc63a74 | -| 18 | — | — | — | — | not run | **two-tell stop ANSWERED 2026-09-11: interrupt the repair mode.** Target text produced (`…-target-text.md`, 476L); design cut 635→385L. Pass 18 aims at the target text | +| 18 | 96c3611 | 9→**15** | 1→**0** | 7→**12** | yes | **first pass on the target text. ZERO BLOCKERS, first since pass 7.** Finding 1 names the restructure as half-done: §H is still paraphrase, not text. Findings 6,7,8 name three standing sentences by line number — the 15/16/17 mechanism, now findable; session 01a09068-de4a-7071-97a7-82cc4774544f | +| 19 | — | — | — | — | not run | next action, after §H is written out as literal text | + +## Pass-18 three-line report — first pass on the target text + +**Floor line:** derived floor **3**; risk **high**, security **none**; read fresh from +`docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md`. One cited story, level 2. + +- **Trend:** findings …, 12, 8, 9, **15**. Blockers …, 1, 1, 1, **0**. Majors …, 7, 4, 7, **12**. + Blocker+Major …, 14, 10, 6, 8, 5, 8, **12**. +- **Cluster (pass 18):** product 11 of 15; **prose about the artifact 4** — the opening's two + claims, §H's "one sentence", and the design spec disagreeing with the target text. Instrument 0. +- **require↔withdraw:** none. Finding 1 asks for *more literal text*, which is the opposite of + asking the removed enumeration back. + +**Tells: one of five** — the finding count rose 9 → 15. **The Blocker count fell to zero**, the +cluster is product, no pair. One is not two: no mandatory stop. + +**Zero Blockers, first since pass 7.** Read narrowly: no path the reviewer traced leaves a cycle +unable to close and unable to suspend. That is the one thing the ordering exists to prevent and +it is the first pass in eleven where nothing hit it. + +**Why the count rose, and it is not the usual reason.** The artifact changed shape. Findings 6, 7 +and 8 each name **a specific standing sentence by file and line** — `b12`'s immediate-resumption +clause (C 205–208), `c18`'s no-clean-credit and the surfacing sentence's "the loop resumes on +whatever the user decides" (C 243–244). That is the same mechanism that produced one finding at +pass 15, one at 16 and three at 17 — but those were "§4 does not name a passage", and these are +"this concrete sentence and that concrete sentence disagree, here are both". **The restructure +made the mechanism decidable rather than recurring**, which is what it was for. + +**And finding 1 says the restructure is half-done, correctly.** §§A–G are concrete text and drew +concrete findings. **§H is still paraphrase** — "is renamed", "becomes conditional", "are +trimmed" — so the reviewer cannot read the future prompt there, and finding 2 catches the +consequence: §B says six sentences change and lists `b3` as carried, while §H replaces `b3` and +`b17`–`b18`. Two sections of one artifact giving incompatible instructions. **The remedy is to +finish the job in the direction already chosen**, not to reconsider it. + +**The rest cluster into wording that the eligibility split left behind** (4, 5, 13 — the +suspension gate's predicate, the exhaustive reason for ineligibility, where a blocked pass's +health observations go), **three semantic questions** (9 the declined-finding overlap example, 10 +"a Minor to the fix set" conflating severity with membership, 14 whether a re-raised valid +dismissal stays discharged), **one on §G's test** (11 — it excludes the squash-carry rule, which +the same paragraph names as a member), and **one on the design spec** (12 — its ownership +sentence still says the block is authoritative only for order and closure). + +**Minors 3 and 13 and Nit 15 are collected**, and 3 and 15 sit inside sections being rewritten +anyway, so they cost no round of their own. ## TWO-TELL STOP ANSWERED 2026-09-11 — interrupt the repair mode, Gate A stays open diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index b5bda49..d574120 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -57,9 +57,13 @@ establish the **inventory** of findings, not their resolutions. What the table does not settle, this spec decides in the section that uses it: the evaluation order and the file set each predicate reads; the duties' classification; which stop each of the scope stop's two triggers raises and what each answer does; what a clearly-stuck or two-tell -answer produces; and the raw-severity rule for the health measures. **It defines no trigger, -duty or severity rule of its own** — each keeps its one definition where that definition already -lives, and where one had to change to agree with the ordering it changed **at its source**. +answer produces; and the raw-severity rule for the health measures. **The block owns exactly six +things** — the evaluation order, closure and its eligibility, the hold and what discharges it, +the composition of several suspensions, the pairs that cannot co-occur, and the clearly-stuck +precedence sentence as a stated exception — **and defines no trigger, no severity rule and no +closure precondition of its own**; each of those keeps its one definition where it already lives, +and where one had to change to agree with the ordering it changed **at its source**. The target +text's §A states the same six-part boundary in its own opening, and the two must not drift. **The closure-ordering block is an addition beside the source edits**, not one of them. §4 lists the edits; **no total is stated here or there**, because the unit — one contiguous replacement at diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index c757276..2b3f06e 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -25,7 +25,9 @@ this file was written in. Gate A stays open and is re-aimed at this text, with p findings as the basis. The rules are not activated by this file. Gate B reviews the real implementation diff later, and this file is not that diff. -**How to read a section.** Each is marked **NEW**, **REPLACED** or **CARRIED**. A CARRIED passage +**How to read a section.** Sections **A**–**H** carry proposed text and each is marked **NEW** or +**REPLACED**; a passage reproduced unchanged is labelled CARRIED where one appears. **§I is not +proposed text** — it is this file's own metadata and ships nowhere. A CARRIED passage is reproduced from the current state unchanged, because the block cites it and a reader has to see what it says; it is here to be read, not to be edited. Both prompt copies take every NEW and REPLACED section **byte-identical**. @@ -100,15 +102,21 @@ gates it. **A commit the hook reads as cycle-closing is a separate matter**: a n mid-cycle makes the hook read the cycle as closed and discards the passes counted so far, which is an observation about the counter — the cycle itself stays open until the conditions above hold, so an accidental commit destroys the pass credit and closes nothing. A plateau or tells on -an eligible pass go into the closing report and never block it, because reporting "will not -converge" on a converged loop is a false report. **No other pass outcome makes a cycle eligible -to close**, because every other pass leaves a required repair, a hold or a question outstanding — -and closing over any of those is the failure this ordering exists to prevent. The one termination +**the pass that closes** go into the closing report and never block it, because reporting "will +not converge" on a converged loop is a false report; on an eligible pass that does **not** close +they go into that pass's ordinary report, there being no closing report to carry them. **No other pass outcome makes a cycle eligible +to close**, because every other pass either leaves a required repair, a hold or a question +outstanding **or has not reached the floor** — and closing over any of those is the failure this +ordering exists to prevent. The floor is named separately because a below-floor pass whose only +findings are Minors leaves nothing outstanding and is still not eligible. The one termination that is not a pass outcome is the Gate-B triviality skip, which runs no passes and is outside this ordering. -**Second, only a pass that is not a clean completion can suspend** — that order is what makes -"clean completion outranks the two-tell stop" executable rather than asserted. Three suspensions, +**Second, only a pass that is not a clean completion can suspend — and *clean completion* here +means the whole first branch, eligibility included.** So a clean pass **below** the floor is not +a clean completion and can suspend, while an **eligible** pass cannot, whatever its preconditions +do. That is what makes "clean completion outranks the two-tell stop" executable rather than +asserted, and cleanliness alone never decides it. Three suspensions, by the names their paragraphs use and read by those paragraphs: the **scope stop**, raised by either trigger above — a **membership stop** by the first, a **question stop** by the second; the **clearly-stuck exit**; and the **two-tell stop**. A suspension waives nothing. Any non-empty set @@ -164,7 +172,10 @@ on until it is answered. The **clearly-stuck reading surfaces findings** — **t pass being read that satisfy its regeneration condition, and only those**; earlier members of a regeneration chain that were repaired or dismissed are history the reading consults and never findings it re-surfaces, so no discharged finding takes a second hold. Each surfaced finding takes -a hold like any other. **Where one finding is surfaced by both, it carries two hold components and +a hold like any other. **A re-raised valid dismissal stays discharged for the resolve duty** — the +dismissal was the resolution and a reviewer repeating the finding does not undo it, so no second +dismissal is owed — and what the recurrence creates is the **clearly-stuck hold alone**, ended by +that reading's continue-or-stop answer. **Where one finding is surfaced by both, it carries two hold components and each is discharged by its own answer, which is what keeps D4 and D5 exact**: the **membership** component ends on the membership answer in either direction, a decline releasing it as **D5** requires; the **clearly-stuck** component ends on the reading's continue-or-stop answer. Neither @@ -219,9 +230,11 @@ that stop's triggers are the clean predicate's own second half, so a pass raisin and a zero-finding pass and any suspension, since it has nothing to surface, nothing regenerating, no cluster and no require↔withdraw pair. **Clean completion and the clearly-stuck exit can**, and the overlap is admitted rather than argued away: the two read different severity fields, as -Mechanics · Severity sets out, so an **in-set** Blocker or Major the ceiling demotes below Major, -or an **out-of-set** one this cycle has declined, can regenerate across passes on a pass that is -clean. **The order decides it and no new rule is needed.** The clearly-stuck paragraph's own +Mechanics · Severity sets out, so an **in-set** Blocker or Major the ceiling demotes below Major can regenerate across passes on +a pass that is clean. **A declined finding is not a route into that reading**: the third condition +admits regeneration across repair attempts and a re-raised validated dismissal, and a decline is +neither — it is the user's decision that a *true* finding stays outside the set, which **D7** binds +for the cycle. **The order decides it and no new rule is needed.** The clearly-stuck paragraph's own precedence sentence is stated here rather than there, because precedence is evaluation order and this paragraph is where evaluation order is stated once; its opening words point back to that paragraph, which is where the reading itself lives. That third condition is what makes a plateau @@ -235,8 +248,10 @@ false report. Below the floor the pass **suspends**, clean completion having not ## B. Passage (b) — what a loop absorbs — REPLACED, in part -Six sentences change — `b7`, `b8`, `b11`, `b12`, `b13`, `b16` — and the rest of the paragraph is -carried. The changed sentences, final: +This passage's replacements are split across two sections, so neither is the whole list: **`b7`, +`b8`, `b11`, `b13` and `b16` are here**, and **`b3`, `b12` and `b17`–`b18` are in §H** with the +other site-by-site replacements. Every other sentence in the paragraph is carried. The five here, +final: **`b7`, the assigned fix set.** ``` @@ -266,12 +281,6 @@ whether the set now includes it. *Why:* unqualified, this and the clean predicate decide a re-raised declined finding in opposite directions. -**`b12`.** Gains, at its end: -``` -Any answer ends the hold; **what the pass does next is the closure ordering's**, not this -sentence's. -``` - **`b13`, the question trigger.** *New* is qualified: ``` …opens a **new structural or contract question** — new meaning not already answered in this cycle, @@ -291,8 +300,8 @@ nobody placed in or out of the set leaves its membership decided by default. *Why (pass 17 finding 4):* without it an agent follows the surviving source, asks only the structural question, and absorbs out-of-set work silently. -**Carried unchanged:** the paragraph's opening, `b3` becomes a pure pointer (below), and every -sentence not named here. +**Carried unchanged:** the paragraph's opening and every sentence named in neither this section +nor §H. --- @@ -356,8 +365,9 @@ nothing, and continue or stop is answered there. findings file — case-folded, and a non-empty unrecognized token read as `MAJOR` — never the raw token, so a finding written `IMPORTANT` enters the curve as a Major. **The ceiling is applied after that and only to what the cycle owes**: cleanliness and the resolve duty read the - effective severity, so the same finding can be a Major to the curve and a Minor to the fix set, - and that difference is the point rather than a discrepancy. The line is **what the cycle owes + effective severity, so the same finding can be a Major to the curve and a Minor **at effective + severity**, and that difference is the point rather than a discrepancy. It stays in the fix set + either way; the ceiling moves what the cycle owes for it and never whether it is in. The line is **what the cycle owes versus what it observes about itself**, which is why no list of readings has to be kept complete here. Two reasons for the split. The curve must stay derivable from the findings files alone — the finding total counts finding lines and the Blocker and Major series count the lines whose @@ -420,8 +430,11 @@ semantics **and the closure ordering together with every rule it reads** depend and the requirement is that the adopted definitions **agree**, not merely that all of them are present. **Membership is decided by a test a reader can apply to the text in front of them, with no list to consult: a live rule belongs to this contract when changing it would change an input -the closure ordering reads, which branch a pass takes, what a hold is or what discharges it, or -whether a cycle may close.** A curve without a cycle field cannot be attributed, a slot rule +the closure ordering reads, which branch a pass takes, what a hold is or what discharges it, +whether a cycle may close, or the production, identity or transport of any record this section +obliges a cycle to write.** The last clause is why the squash carry belongs: it moves no pass and +decides no branch, and a record that does not survive the merge is a record the cycle did not +produce. A curve without a cycle field cannot be attributed, a slot rule without a nonce has nothing to key on, a carry rule naming records a project does not produce is inert, and a clean predicate without the fix-set boundary it reads decides membership by accident. **A project whose text carries some of them and not others, or carries all of them in versions @@ -439,29 +452,101 @@ rather than doing the noticing. --- -## H. The remaining source edits — REPLACED, and each is one sentence - -- **the gate-prompt template's clean sentence** — a `NO FINDINGS` file is renamed a **clean - findings file**, which is what it always described. -- **the Gate-A clean-signal sentence** — the `NO FINDINGS` signal becomes what lets a pass be read - as clean without inspecting it, rather than the only route to clean; a pass carrying Minors - alone is clean and could never produce that file. -- **the lens paragraph's unchanged-list** — scoped to **the lens sets**, since the ordering - falsifies its claim that the Blocker/Major filter and the clean-final-pass rule are unchanged. -- **the Gate-A cadence** — "validate, revise, re-run" becomes conditional on a repair being - required, since unconditional it tells a Minor-only pass to manufacture the repair the severity - rule forbids. -- **`a13`** — "Every other rule stated here about how a cycle closes stands as written, and none - of them is restated" is scoped to that paragraph, which is what it was written to police; - categorically false once the ordering exists. *(It wraps across three lines in each copy.)* -- **`a16`** — "fix Blocker/Major after each" points at Mechanics · Severity instead of restating - an unscoped version. -- **`a17`–`a19`** — the floor paragraph's closure sentences are trimmed to point at the ordering. -- **`b3`** — becomes a pure pointer to Mechanics · Severity, stopping its restatement of the four - severity actions; W's pointer also changes target to match C. -- **`b17`–`b18`** — they name their stop a suspension and defer to the ordering. -- **the unknown-start strict-reading list** — gains **every suspension binding, since unknown - starting rules cannot waive an open hold**, with the reason carried inline. +## H. The remaining replacements, written out — REPLACED + +Each is the final wording. Nothing here is a paraphrase; the plan installs these strings. + +**`b3`** — Mechanics is pointed at, and the four severity actions are no longer restated beside it. +Both copies take C's wording; W's "the severity rule" pointer was written on the inventory's +reasoning that W has no Mechanics section, which is false. +``` +keep it here rather than handing it back, then act on it by its severity exactly as +Mechanics · Severity says. +``` + +**`b12`** — the immediate-resumption command is **replaced**, not appended to. It is the only +sentence in the standing text that resumes a loop on one answer, and the ordering resumes on all +of them. +``` +and the membership answer ends that finding's membership hold; **what the pass does next is the +closure ordering's**, which resumes only when every answer outstanding on that surface has been +given. +``` + +**`b17`–`b18`** — the stop is named a suspension and defers to the ordering rather than restating +what ends it. +``` +Stopping this way is **not an exit from the gate**: it is a **suspension** in the closure +ordering's sense, the floor, the Blocker/Major filter and the clean-final-pass rule all stand, +and what the answer does is stated there — what the stop prevents is a loop committing you to a +design you never chose, which is a different failure from an unfinished review. +``` + +**`c18` and the surfacing sentence** — the two clauses findings 7 and 8 name. The resolve and hold +duties are kept; the blanket no-clean-credit and the one-answer resumption go, because the +ordering decides both and decided them differently. +``` +**Surfacing does not close the cycle, and that is what makes this reachable.** You surface *with +the finding still open* — the resolve rule is not waived, the hold stands until its answers are +given, and **what the answer does is the closure ordering's**: continue resumes, stop parks. +**A pass is credited clean or not on its own findings**, as that ordering defines cleanliness; +surfacing a finding that carries a scope-stop trigger is what withholds the credit, and no +credit is withheld for surfacing alone. Reading this as "stop instead of fixing" would put the +exit in competition with the rule that every Blocker and Major resolves, and then nothing could +satisfy both. +``` + +**`a13`** — scoped to its own paragraph, which is what it was written to police. *(It wraps +across three lines in each copy: C 129–131, W 336–338.)* +``` +Every other rule stated **in this paragraph** about how a cycle closes stands as written, and +none of them is restated — a summary is where their conditions would get dropped. +``` + +**`a16`** — points at Mechanics instead of carrying an unscoped copy. +``` +Open a TodoWrite "Codex pass N" per pass; resolve Blocker/Major after each as +Mechanics · Severity requires. +``` + +**`a17`–`a19`** — the floor paragraph stops stating closure and the early exit, and points at the +ordering that states them once. Three sentences become one. +``` +What a clean final pass and the zero-finding early exit mean for closing is stated once in the +closure ordering; this paragraph states the floor and nothing else. Codex is advisory — validate +before applying; dismissed finding → one-line why. +``` + +**The Gate-A clean-signal sentence** — the signal stops being the only route to clean. +``` +Ask for one line per finding and a literal `NO FINDINGS` when a pass found none — that explicit +signal is what lets a pass be read as clean without inspecting it, and a pass carrying only +Minors is clean too and could never produce that file: +``` + +**The gate-prompt template's clean sentence**, in the block both gates paste. +``` +A **clean findings file** is the single body line `NO FINDINGS` with `END OF FINDINGS (0 total)`. +``` + +**The Gate-A cadence** — revision becomes conditional, since unconditional it tells a Minor-only +pass to manufacture the repair the severity rule forbids. +``` +Each pass: validate, revise **where a repair is required**, re-run. +``` + +**The lens paragraph's unchanged-list** — scoped to the lens sets, which is what that paragraph is +about. +``` +Lenses change what a pass asks, never how many a cycle owes: **the lens sets** leave every other +rule in this section alone. +``` + +**The unknown-start strict-reading list** — gains the suspensions, reason inline. +``` +…the curve duty owed, the nonce duties at their strictest, and **every suspension binding, since +starting rules that cannot be established cannot be read as having waived an open hold**. +``` --- From 0350b82a46fa14c377bafba17e4ded0efcdba4d4 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Fri, 11 Sep 2026 14:42:12 +0200 Subject: [PATCH 032/181] docs(context): prepare the cycle record for handoff The durable prompt file still aimed at the pre-pass-18 artifact. A fresh session following the resume procedure would have run pass 19 against the design spec instead of the target text, and nothing would have caught it -- the design spec is a real file and the call would have succeeded. Rewritten for the target text, with the artifact change stated at the top. Also in the cycle record: - STATE AT HANDOFF rewritten for 2026-09-11: HEAD, the artifact change and why the nonce continues, the full Blocker and Blocker+Major series, and the note that pass 19 is the first pass over text that is complete -- section H was still paraphrase when pass 18 read it. - The resume procedure now names the target text in step 3 and says to run the precheck on both files; two false counts have been caught that way before a pass was spent on them. - A new section, "How this cycle is being run", so a fresh session does not have to rediscover it: findings are routed through an external reviewer who has corrected the agent at every stop since pass 14; verify every citation before acting on it, with the three that were wrong named; Minor and Nit get no round; no known contradiction becomes a residual; no new mechanism. - Where-everything-lives now distinguishes the artifact from the design input. - The decisions table gains passes 14, 15 and 17. - Two errors in the record itself: a German word left in an English file, and a reference to an item number that no longer exists. The prompt file is gitignored, so it survives a context clear on this machine and not a fresh clone; the record says so rather than leaving it to be discovered. This resume note and every pass file are tracked. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-resume.md | 75 ++++++++++++------- 1 file changed, 49 insertions(+), 26 deletions(-) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index 1911799..e8e5dc8 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -606,35 +606,55 @@ after dismissal, not a withdrawal). One tell is not two — no mandatory stop. ## Prompt -**`.context/gate-a-spec-prompt.md`** — durable copy, substitute `__SHA__` and `__P__`. It carries -the pre-call checklist and the three "deliberately not here" blocks that have to stay in it, or the -reviewer re-raises deferred material. The scratchpad copy is gone with its session. - -## STATE AT HANDOFF — 2026-09-10 evening - -- **HEAD `0168f88`** on `loop-rule-consolidation`, tree clean. Spec **532 lines** (was 785). -- **Pass 13's findings are all dispositioned and applied**; dispositions in - `gate-a-spec-awsf1ec771-pass-13-dispositions.md`. Findings 15, 16, 13, 14 dissolved with the cut; - 4, 5, 12 became rows in the new §4 table; 6 became one line in §6; 10 was absorbed by §7's - narrowing; the rest fixed in §3. -- **The next action is Gate-A spec pass 14** against `0168f88`. Nothing else is pending. -- **The cut, measured:** the design is now **192 of 532 lines, 36%**, against 173 of 785, 22%. - §5 accounting 216 → 33, §7 verification 119 → 67, §4 quoting 83 → 54, §6 parity 38 → 18. The - bookkeeping did not vanish; the plan carries it, beside the edits, where it is checkable against - real files. -- **Transition walk done and clean** — every state the spec names reaches close or park, and no - state returns itself with its input consumed. That was pass 13 finding 2's defect class. +**`.context/gate-a-spec-prompt.md`** — durable copy, substitute `__SHA__` and `__P__`. **Rewritten +at pass 18 for the new artifact.** It carries the pre-call checklist and the "deliberately not +here" blocks, which must stay in or the reviewer re-raises deferred material. +**It is gitignored** (`.context/*` admits only `codex-gate.on` and `codex-reviews/`), so it +survives a context clear on this machine and not a fresh clone. This resume note and every pass +file are tracked from 2026-09-10 and do survive. + +## STATE AT HANDOFF — 2026-09-11 midday + +- **HEAD `ef9a504`** on `loop-rule-consolidation`, tree clean, 18 passes run, none clean. +- **THE ARTIFACT CHANGED AT PASS 18.** Passes 1–17 reviewed the design spec; from pass 18 the + artifact is `docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md` (562 lines) + and the design spec (389 lines) is an input. Same nonce: restructured, not replaced. +- **Pass 18's fifteen findings are all dispositioned and applied**, dispositions implicit in commit + `ef9a504`'s body, which names each by number. No Minor or Nit got a round of its own. +- **Next action: Gate-A spec pass 19** against `ef9a504`. Nothing else is pending. +- **Where the loop stands.** B+M by pass: 20, 15, 10, 11, 15, 13, 9, 8, 12, 17, 14, 14, 10, 6, 8, + 5, 8, 12. Blockers: 5, 2, 4, 1, 0, 4, 0, 1, 2, 2, 1, 1, 2, 1, 1, 1, 1, **0**. Pass 18 is the + first zero-Blocker pass since pass 7 and the first pass on a fully concrete artifact — §H was + still paraphrase when pass 18 read it, so pass 19 is the first pass over text that is complete. +- **Three structural interventions moved this cycle, and nothing else did:** the split at pass 10, + the cut at pass 13, the target-text restructure at pass 17/18. Ordinary repair rounds never did. ## Next — resume procedure, in order -1. `git -C /Users/daniel/DEVELOPMENT/APPS/dev-workflow-kit log --oneline -5` and `git status --short`. Branch is `loop-rule-consolidation`, tree must be clean. +1. `git -C /Users/daniel/DEVELOPMENT/APPS/dev-workflow-kit log --oneline -5` and `git status --short`. Branch `loop-rule-consolidation`, tree must be clean. 2. Read this file's Passes table and the newest three-line report for where the loop stands. -3. `python3 .context/spec-precheck.py docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md` — exit 0, and eyeball its COUNT list against the spec's enumerations. -4. `rm -f .context/codex-reviews/gate-a-spec-awsf1ec771-pass-.md`, confirm gone. +3. `python3 .context/spec-precheck.py docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md` — exit 0 — and again on the design spec. Eyeball each COUNT note against its enumeration; two false counts have been caught this way before a pass was spent on them. +4. `rm -f .context/codex-reviews/gate-a-spec-awsf1ec771-pass-.md`, confirm gone. **Only paths carrying this nonce.** 5. Fill `__SHA__` (current HEAD, short) and `__P__` into `.context/gate-a-spec-prompt.md`, call `mcp__codex__exec` with `workingDirectory` = repo root. -6. Validate the file: last line exactly `END OF FINDINGS ( total)`, exactly `` finding lines, nothing else. Anything else is an INCOMPLETE pass — do not count it, do not act on the partial list. -7. Disposition every finding, then have a fork revise. Report the three lines to Daniel (trend, cluster, require↔withdraw) — the duty is active from pass 4 and every pass owes it, plus the floor line (floor 3, risk high, security none, read from the story header). -8. Append a row and a report section here before running the next pass. +6. Validate: last line exactly `END OF FINDINGS ( total)`, exactly `` finding lines, nothing else. Anything else is INCOMPLETE — do not count it, do not act on the partial list. +7. Disposition every finding. Report to Daniel: the floor line (floor 3, risk high, security none, read from the story header) **and** the three lines (trend, cluster, require↔withdraw) — owed by every pass from 4 on. +8. Append a row and a report section here **before** running the next pass. + +## How this cycle is being run — read before deciding anything + +- **Daniel routes findings through an external reviewer**, who has corrected the agent's + recommendation at every stop since pass 14 and been right each time. Expect a recommendation to + be revised rather than executed, and **state options with their downsides rather than arguing + for one**. +- **Verify every citation before acting on it.** Three of the agent's own claims were wrong and + caught this way: that the clearly-stuck exit permits moving to Gate B (`CLAUDE.md:242` says + surfacing closes nothing), that the C1 precedent shows convergence (the field report says that + cycle closed *not converged*), and that a case-sensitive grep had found every stale reference. +- **Minor and Nit never get a repair round of their own.** They are collected; fixing one inside a + sentence already being rewritten is not a round. +- **No known behavioural contradiction may be renamed a residual.** +- **No new mechanism, no checker, no record-durability work** — all three excluded by name. + ## Standing practices adopted mid-cycle — NONE of these is a shipped rule @@ -664,7 +684,9 @@ nothing checks them. They are candidates captured in | pass 10 two-tell stop | **Split.** Record durability moves to `docs/superpowers/stories/2026-09-10-record-durability-story.md`; this story keeps the ordering, the duty classification and the severity/health answer. | | mid-loop, from another project | Anchor the repeat criterion in the kit; captured as the harness-finding-termination story. | | after pass 13 | **Cut the spec to the design**; the plan carries the bookkeeping. Precheck runs before every pass. | -| pass 14 scope stop | **B — accept, bounded.** Extend the existing coherence *instruction* to the closure block and its coupled edits (§4 item 22). No checker, no new mechanism, no record-durability work. | +| pass 14 scope stop | **B — accept, bounded.** Extend the existing coherence *instruction* to the closure block and its coupled edits (now the one-contract paragraph, target text §G). No checker, no new mechanism, no record-durability work. | +| pass 15 | **Bounded rollback**, not another twelve-repair round: totals, item-number lists and enumerated precondition lists removed; behaviour stays decided in the design. | +| pass 17 two-tell stop | **Interrupt the repair mode.** Produce a non-active **target-text** version; Gate A stays open, no transition to Gate B, no clearly-stuck close claimed. | ## For the execution phase, not needed yet @@ -677,7 +699,8 @@ intention. One command; not needed while only the spec is being written. | What | Path | |---|---| -| spec under review | `docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md` | +| **artifact under review (from pass 18)** | `docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md` | +| design — decisions and reasons, an input | `docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md` | | condition inventory (135 ids, committed) | `docs/superpowers/specs/2026-09-10-loop-rule-consolidation-condition-inventory.md` | | this story | `docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md` | | successor, profile unconfirmed | `docs/superpowers/stories/2026-09-10-record-durability-story.md` | From 6d52fccc287ee6e89b557cb317706dd0a94120c0 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Fri, 11 Sep 2026 17:50:22 +0200 Subject: [PATCH 033/181] docs(specs): precheck repairs before pass 19 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Three repairs to the target text, none of them a new decision. a20 restored. "don't manufacture findings to pad" lives in both active copies (CLAUDE.md:135, workflow-init.md:342) and had fallen out of the target text silently: the §H block labelled `a17`-`a19` reproduces a21 and a22, so the string it installs spans a17-a22 and a20 sat inside it unaccounted. Moved to §A unchanged, beside the zero-finding rule it qualifies. The §H label now reads `a17`-`a22` and records where a20 went and that a21/a22 are carried; no sentence count is claimed, condition ids not being sentence numbers, and the per-condition accounting stays the plan's. §B no longer ships the immediate resumption §H replaces. The inventory puts `it resumes the moment the user says whether the set now includes it` at b12, which §H replaces; §B's b11 block quoted it as final wording, so the two sections instructed the plan differently. The b11 quote now ends at b11 and points at §H for the continuation. "being clean it cannot suspend" corrected to "being eligible". The second branch states two paragraphs earlier that a clean pass below the floor can suspend; cleanliness alone never blocks a suspension and eligibility does. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- ...9-10-loop-rule-consolidation-target-text.md | 18 ++++++++++++------ 1 file changed, 12 insertions(+), 6 deletions(-) diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index 2b3f06e..48ddca9 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -81,7 +81,8 @@ clean** when its findings carry no in-set Blocker or Major at effective severity scope-stop trigger** — the two the absorb paragraph defines, read there and not redefined here, each already carrying the qualification an answer given in this cycle puts on it. A pass with **zero** findings is clean whatever the floor, because a floor buys further looks at an artifact -that keeps yielding findings, and one yielding none has already given what those looks were for. +that keeps yielding findings, and one yielding none has already given what those looks were for; +don't manufacture findings to pad. **Eligibility is exactly this and nothing more: a clean pass at or above the derived floor, or a zero-finding pass.** It is a property of the pass. **Closure is eligibility plus every closure @@ -128,7 +129,7 @@ it is not asked twice; the two-tell stop surfaces tells and not a finding. **Third, a pass that neither closes nor suspends continues** — the loop runs another pass on the **current** artifact, revised where the severity and scope rules require a repair and unrevised where they do not. **An eligible pass with an unmet closure precondition lands here**: clean -completion did not close it, and being clean it cannot suspend, so the loop continues on whatever +completion did not close it, and being eligible it cannot suspend, so the loop continues on whatever the unmet precondition requires — most often a repair still owed from an earlier pass. A below-floor clean pass lands here too, **only where no suspension applies to it**; where one does, the second branch has already taken it, because clean completion did not close the pass and only @@ -275,9 +276,12 @@ keeps out. ``` **A correction that leaves that set stops the loop like any other out-of-scope finding** — except one this cycle has already declined, which is outside the set by that decision and raises no -trigger — even when it opens no new question at all, and it resumes the moment the user says -whether the set now includes it. +trigger — even when it opens no new question at all, ``` +*The sentence continues with `b12`'s replacement, which is in §H* — the trailing resumption clause +is `b12`, not `b11`, and quoting its live wording here as final would ship the immediate +resumption the ordering replaces. + *Why:* unqualified, this and the clean predicate decide a re-raised declined finding in opposite directions. @@ -509,8 +513,10 @@ Open a TodoWrite "Codex pass N" per pass; resolve Blocker/Major after each as Mechanics · Severity requires. ``` -**`a17`–`a19`** — the floor paragraph stops stating closure and the early exit, and points at the -ordering that states them once. Three sentences become one. +**`a17`–`a22`** — the floor paragraph stops stating closure and the early exit, and points at the +ordering that states them once. `a20` moves to §A unchanged, beside the zero-finding rule it +qualifies; `a21` and `a22` are carried unchanged, reproduced here because the plan installs one +contiguous string. The per-condition accounting is the plan's. ``` What a clean final pass and the zero-finding early exit mean for closing is stated once in the closure ordering; this paragraph states the floor and nothing else. Codex is advisory — validate From 8714e4334db7f0a30023481fd1623b488a29fb1a Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Fri, 11 Sep 2026 18:10:27 +0200 Subject: [PATCH 034/181] =?UTF-8?q?docs(context):=20Gate-A=20pass=2019=20?= =?UTF-8?q?=E2=80=94=20one=20tell,=20findings=20open?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 13 findings, 1 Blocker, 10 Majors, 2 Minors. Valid pass: terminator and count check out, no extra lines. One tell of five (Blockers 0 -> 1); no mandatory stop. Only 6 of 13 are against the target text. Five are against the design spec, four of those stale internal references the pass-17/18 cut left behind, and two are against standing §5 text. Findings not yet dispositioned; routed to the reviewer first. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-19.md | 14 +++++++ .../gate-a-spec-awsf1ec771-resume.md | 37 ++++++++++++++++++- 2 files changed, 50 insertions(+), 1 deletion(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-19.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-19.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-19.md new file mode 100644 index 0000000..921060d --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-19.md @@ -0,0 +1,14 @@ +BLOCKER | high | §A "Nothing is closed before that act" | The closing-time recheck names profile, cited set and Gate-B evidence, but not the reviewed artifact revision or the assigned fix set, even though the same block says a later scope broadening is "a new fact the next pass reads" and expressly allows Gate-A acceptance to be written in a later commit | An artifact edit or scope expansion after an eligible pass can reach the closing act without that promised next pass, closing the cycle over content or duties no clean pass reviewed | Snapshot the reviewed artifact revision and assigned fix set with the pass and require both still match at the closing act; otherwise make the pass non-final and run the next pass +MAJOR | high | §A "no scope-stop trigger" | Cleanliness reads triggers with the qualification "an answer given in this cycle" puts on them, while the duties paragraph later says no-clean-credit is a fact about the original pass that "nothing discharges"; the first wording does not limit the answer to one given before that pass | After a user answers a membership or structural question, one reading retroactively removes the original pass's trigger and makes that already-surfaced pass clean, while the other permanently withholds its clean credit | State that a pass snapshots triggers using only answers already given before that pass; later answers discharge holds but never rewrite that pass's cleanliness +MAJOR | medium | §B b11 "raises no trigger" | The declined-finding exception says a finding already declined "raises no trigger" without limiting that claim to the membership trigger, while b13 still requires any finding that opens a new unanswered structural or contract question to raise the question stop | A declined finding that later exposes a genuinely new contract question can either close as trigger-free or suspend for the unanswered question | Say it raises no membership trigger on account of being out of set, while leaving the independent b13 question trigger intact +MAJOR | medium | standing Mechanics "neither a human's assent nor this record" | Both standing copies retain the sentence, wrapping across C 1014–1015 and W 1198–1199, that "on a STOP you still stop, and neither a human's assent nor this record lets an agent close or continue a cycle", while §A makes the prescribed human `continue` answer resume a clearly-stuck or two-tell suspension | A downstream agent can follow the retained sentence and refuse the exact continuation transition the new ordering requires | Qualify the standing sentence to generic assent or the human-exception record, and point prescribed suspension answers to the closure ordering +MAJOR | high | §A "author's recorded acceptance" | The Gate-A closing act depends on "recorded acceptance" and a "closing record", but neither term is tied to an existing mandatory commit-body form or given an exact operation; the standing provenance line and curve formats record floor and counts, not acceptance of a named revision | One reader can treat the already-existing commit as acceptance, another can add arbitrary prose, and a third can keep the cycle open because no specified act occurred | Define which already-required Gate-A commit-body write constitutes recorded acceptance of the reviewed revision, without introducing a separate durability mechanism +MAJOR | high | §G "when changing it would change" | The one-contract membership test quantifies over an unspecified hypothetical change: a cosmetic change leaves any rule out, while changing any unrelated live rule into "close immediately" makes it a member, so the same shipped prompt does not determine the counterfactual a downstream reader must apply | Partial adopters can include or exclude rules such as lens, validation or reporting rules while each claims to have applied the test, defeating the authorised coherence stop | Define membership from the live rule's current semantic output or dependency on the named consumers, including record production, identity and transport, rather than from an unconstrained imagined edit +MAJOR | high | target-text promise "the §5 text as it will read" | Four fenced replacements still contain editorial ellipses that are absent from both standing copies: b13 begins `…opens`, §C begins `…and`, §D e7 ends `user…`, and §H's unknown-start fragment begins `…the curve`; yet the artifact says the plan installs these strings byte-identically as final text | Gate A still cannot review the exact bytes or even tell whether each ellipsis is installed punctuation or an omitted carried splice, so the plan must invent part of the supposedly concrete prompt | Write each affected final sentence without editorial ellipses, or explicitly delimit splice markers from the bytes the plan installs +MAJOR | high | §C "moves into the block above, word for word" | The standing below-floor sentence, which wraps across C 238–241 and W 441–444, says a Blocker/Major-free pass carrying a Minor "keeps looping", but §A says such a below-floor clean pass suspends whenever a health suspension applies and continues only otherwise | Following §C's word-for-word instruction would restore an unconditional continuation beside the new conditional suspension branch and recreate two answers for the same pass | Scope the word-for-word claim to the preserved precedence sentence and mark the below-floor condition as replaced and split across §A's suspension and continuation branches +MAJOR | high | design §8 "authoritative for the evaluation order and for closure" | Although design §2 and target §A say the block owns six subjects, design §8 still says it owns only evaluation order and closure and "cites every other rule"; the block actually defines hold discharge, suspension composition, impossible pairings and the clearly-stuck precedence exception | The plan can move or duplicate those four authorities while still satisfying one of the design's conflicting ownership statements | Update the §8 invariant analysis to the same six-part ownership boundary used in design §2 and target §A +MAJOR | high | design §7 "§4's `Kind` column" | The verification procedure says additions are identified from §4's `Kind` column, but §4 has only `Passage or site` and `What changes` columns | The plan has no specified source for deciding which edits receive presence-only checks versus old-and-new discriminating pairs, so required counterfactual coverage can be silently omitted | Point to an existing classification or require the plan to classify each concrete edit as add or replace before generating its checks +MAJOR | high | design §7 "passage (g)'s old sentence is the one site" | The counterfactual paragraph says passage (g) is the only site where the parent contains wording the change removes, while the target explicitly replaces standing text in §§B, C, D, E, F and H as well | The evidence design understates its replacement counterfactuals and can produce a green check that never proves those old instructions disappeared | State the addition counterfactual separately and require every replacement identified from the concrete target to show old wording present in the parent and absent after the change +MINOR | high | design §7 "keeps the oracle and §3 in agreement" | Design §7 says §3 defines how `continue` consumes a health reading, but §3 is "Where the text is" and contains no transition rule; that rule is in target text §A | The named verification points its reader at a section that cannot supply the oracle it claims to match | Replace the stale §3 citation with the target text's §A transition paragraph +MINOR | high | design §8 "carry theirs inline in §3" | Design §8 twice cites §3 as the location of the three inline reasons and of a table checking one-definition authority, but §3 contains neither those reasons nor a table; the reasons are in target §A and the site table is in design §4 | The prompt-standard audit is mechanically untraceable at the cited locations and can be accepted without checking the actual text | Cite target §A for the reasons and the actual design section that owns the relevant site or authority check +END OF FINDINGS (13 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index e8e5dc8..c176142 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -45,7 +45,42 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 16 | 5d12884 | 12→**8** | 1→**1** | 7→**4** | yes | **B+M 8→5, lowest of the cycle.** One tell (Blockers flat). Five B/M applied; Minors 7 and Nit 8 collected. Findings 2 and 4 are the spec disagreeing with *standing* §5 text, not with itself; session 01a08fff-fadf-72f2-994d-618048bea473 | | 17 | c2d5fe1 | 8→**9** | 1→**1** | 4→**7** | yes | **TWO-TELL STOP — mandatory, surfaced to Daniel.** B+M 5→8. **Three findings (4, 5, 7) are one mechanism on its third pass**: a standing §5 sentence the block falsifies that §4 does not name. All 9 held open; session 01a0901b-ecf3-73b2-886c-5b3c7dc63a74 | | 18 | 96c3611 | 9→**15** | 1→**0** | 7→**12** | yes | **first pass on the target text. ZERO BLOCKERS, first since pass 7.** Finding 1 names the restructure as half-done: §H is still paraphrase, not text. Findings 6,7,8 name three standing sentences by line number — the 15/16/17 mechanism, now findable; session 01a09068-de4a-7071-97a7-82cc4774544f | -| 19 | — | — | — | — | not run | next action, after §H is written out as literal text | +| 19 | 6d52fcc | 15→**13** | 0→**1** | 12→**10** | yes | precheck caught a silent loss before the pass: `a20` ("don't manufacture findings to pad") had fallen out of the target text; restored to §A. Two known contradictions repaired in the same commit. **Only 6 of 13 findings are against the target text**; 5 are against the design spec and 2 against standing §5; session 01a0912a-7329-7450-bdb3-da043e88e7bb | +| 20 | — | — | — | — | not run | next action, after pass 19's findings are dispositioned | + +## Pass-19 three-line report — first pass on a target text with no paraphrase left + +**Floor line:** derived floor **3**; risk **high**, security **none**; read fresh from +`docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md`. One cited story, level 2. + +- **Trend:** findings …, 8, 9, 15, **13**. Blockers …, 1, 1, 0, **1**. Majors …, 4, 7, 12, **10**. + Blocker+Major …, 10, 6, 8, 5, 8, 12, **11**. +- **Cluster (pass 19):** product 7 of 13 (1, 2, 3, 4, 5, 6, 8); **prose about the artifact or the + design 4** (7 the ellipses against "nothing here is a paraphrase", 9 the design's ownership + sentence, 12 and 13 two stale `§3` citations); the instrument 2 (10, 11 — §7's verification). +- **require↔withdraw:** none. Named near-miss: finding 10 reports that design §7 cites a `Kind` + column the pass-17/18 cut removed, but its fix offers "require the plan to classify" instead of + restoring the column — a dangling reference reported, not the removed table demanded back. The + mirror shape, not the shape. + +**Tells: one of five** — the Blocker count failed to fall, 0 → 1. Findings fell 15 → 13, Majors +fell 12 → 10, the cluster is product, no pair. One is not two: no mandatory stop. + +**Where the findings landed, and it is the new thing this pass shows.** Only **6 of 13** are +against the target text (1, 2, 3, 5, 6, 7). **Five are against the design spec** (9, 10, 11, 12, +13) and **two against standing §5 text** (4, 8). Four of the five design findings are stale +internal references the pass-17/18 cut created — `§4`'s removed `Kind` column, `§3` cited twice for +material that now lives in target §A, and an ownership sentence §2 already replaced. That is the +restructure's own debris, mechanical to fix and not a design question. + +**The precheck earned its place this pass.** It caught `a20` — a live behavioural prohibition in +both copies — vanishing inside a replacement span whose label said `a17`–`a19`. Nothing in the +pass-19 findings names it, so had it not been caught here it would have shipped. + +**The Blocker is about the closing-time recheck's list.** §A gates the closing act on a changed +profile, cited set or evidence entry, and not on the reviewed artifact revision or the assigned fix +set — while the same block says a later broadening is "a new fact the **next** pass reads". Verified +against the text: the list is those three and no more. ## Pass-18 three-line report — first pass on the target text From 5174d7abf6c5de09c70e9ea46e3336b51827bb8d Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Fri, 11 Sep 2026 18:46:06 +0200 Subject: [PATCH 035/181] docs(specs): apply Gate-A pass 19 in one repair round MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit All 13 findings repaired or collected together, design spec and target text in one round rather than two cycles. §B becomes the whole absorb paragraph, contiguous. Splitting its replacements across §B and §H left one sentence half in each, which is how the two sections came to instruct the plan differently about the same clause. The merge also removed two defects it exposed: the shipped text would have carried the inventory id b7 into CLAUDE.md, where it means nothing, and §H's claim that b12 is the only standing sentence resuming a loop on one answer was false of b18. 1, 2, 5 share one nexus in §A. The closing-act recheck now gates on the assigned fix set and the reviewed artifact revision as well as the profile, cited set and evidence entry; sameness is read on the artifact and the duties rather than on the branch tip, so writing the closing body is not itself a change. A pass's cleanliness is settled on the answers standing when it ran. The Gate-A closing act is the commit body carrying the provenance line and the curve, records Mechanics already obliges, so no new mechanism is introduced. 3, 4, 7, 8 are solved as complete passages. b11's declined-finding exception is scoped to the membership trigger, leaving the question trigger intact. The human-exception sentence separates a prescribed continuation answer from blanket assent; §F now carries four standing sentences. Every editorial ellipsis is replaced by real context. §C's word-for-word claim covers the precedence sentence only, the below-floor sentence being replaced and split across two branches. 6 reads §G's membership test on what a rule states rather than on an imagined edit to it. 9, 10, 11, 12, 13 become brief pointers to the authoritative site. The counterfactual splits into the block's absence and every replacement's old-wording-gone half, which is not one site. The transition walk caught one self-inflicted defect: §A said a Gate-A cycle has no amend, then put the closing body into an existing commit. Now it is an amend of the message alone, leaving the tree untouched. Gate A stays open; pass 19 was not clean. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-resume.md | 35 +++ ...26-09-10-loop-rule-consolidation-design.md | 38 ++- ...-10-loop-rule-consolidation-target-text.md | 277 ++++++++++-------- 3 files changed, 217 insertions(+), 133 deletions(-) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index c176142..3b2ade8 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -48,6 +48,41 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 19 | 6d52fcc | 15→**13** | 0→**1** | 12→**10** | yes | precheck caught a silent loss before the pass: `a20` ("don't manufacture findings to pad") had fallen out of the target text; restored to §A. Two known contradictions repaired in the same commit. **Only 6 of 13 findings are against the target text**; 5 are against the design spec and 2 against standing §5; session 01a0912a-7329-7450-bdb3-da043e88e7bb | | 20 | — | — | — | — | not run | next action, after pass 19's findings are dispositioned | +## Pass-19 dispositions — one connected repair round, on the reviewer's order + +All 13 repaired or collected in one round; **no separate cycles for design and target text**, which +was the reviewer's call against splitting them. Three of my report's claims were corrected by him +and are recorded because each was right: "without the precheck `a20` would have shipped" is +unproven — pass 19 read the already-repaired text, and what is proven is only that the precheck +caught a real loss in time; "the five design findings are mechanical" understates findings 9 and +11, which are responsibilities and evidence duties, and 9 repeats an incompletely repaired pass-18 +finding; and "only six hit the target text" draws an artificial line, since finding 8 objects to +§C's own "word for word" instruction. + +| Findings | Disposition | +|---|---| +| 1, 2, 5 | **§A, one nexus.** The closing-act recheck gains the assigned fix set and the reviewed artifact revision; sameness reads the artifact and the duties, never the branch tip, so writing the closing body is not a change. Cleanliness is settled on the answers standing when the pass ran. The Gate-A closing act is the commit body carrying the provenance line and curve — records Mechanics already obliges, no new mechanism. | +| 3, 4, 7, 8 | **Complete target passages.** `b11`'s exception scoped to the membership trigger; the human-exception sentence distinguishes prescribed continuation answers from blanket assent (§F, now four sentences); every ellipsis replaced by real context; §C's word-for-word claim scoped to the precedence sentence, the below-floor sentence marked replaced and split. | +| 6 | **§G reads the rule's present content**, not an imagined edit to it. No membership list. | +| 9, 10, 11, 12, 13 | **Design spec, brief pointers to the authoritative site.** §8 points at §2's boundary rather than restating it and at §4's table; §7's `Kind` column becomes the plan classifying against real files; the counterfactual splits into the block's absence and every replacement's old-wording-gone half; two stale `§3` citations point at target §A. | + +**§B is now the whole absorb paragraph, contiguous.** The split across §B and §H is what let the +two sections instruct the plan differently about one sentence, so the split is gone rather than +patched. Two things fell out of the merge: the shipped text would have carried the inventory id +`` `b7` `` into `CLAUDE.md`, where it means nothing — now "as just defined" — and §H's "the only +sentence in the standing text that resumes a loop on one answer" was false of `b18` and is gone, +as the reviewer noted it could be. + +**One defect the transition walk caught, self-inflicted:** the repaired §A said a Gate-A cycle has +"no amend" and two sentences later put the closing body into an existing commit, which is an amend. +Now: no WIP snapshot **to replace**, and the closing body is an amend of the message alone, leaving +the tree and so the reviewed revision untouched. + +**Observation, not repaired and not raised by the pass:** a cycle whose reviewer re-raises a validly +dismissed in-set Major every pass can suspend (§C's third condition admits it) but can never be +clean, so it can only ever **park**. That is a reachable stated outcome rather than a dead end, and +§C's own rationale anticipates it. + ## Pass-19 three-line report — first pass on a target text with no paraphrase left **Floor line:** derived floor **3**; risk **high**, security **none**; read fresh from diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index d574120..c755fd0 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -200,9 +200,10 @@ tree and the parent tree, so every assertion is observed passing where the chang failing where it does not. A one-sided presence check is not enough: a copy carrying the new wording **and** the old one satisfies it, which is the two-instructions-that-disagree failure §4 exists to prevent. **An edit that only adds, replacing no wording, is checked by presence alone**, because there is -no old text whose absence could be counted; which edits those are is read off §4's `Kind` column -rather than listed here, a list of item numbers being the bookkeeping that has to be re-derived -whenever a span merges. **The plan builds each pair +no old text whose absence could be counted; **which edits those are the plan classifies against the +real files**, an add-only edit being one whose site carries no wording the change removes. Neither +this section nor §4 classifies them: a list of item numbers here is the bookkeeping that has to be +re-derived whenever a span merges. **The plan builds each pair against the real files and runs both directions there**, under two constraints: a counted fragment must be **single-line in the file it is grepped from**, since one spanning a line break makes `grep -F` count zero and read as a failure; and the new wording must therefore be @@ -218,10 +219,14 @@ guess reads as a failed check rather than a wrong one. **Nothing verifies that complete or that the plan's fragments discriminate.** The enumeration moved to where the text exists; the completeness claim did not, because nothing supports it. -The counterfactual is **ABSENT and is claimed as absent**: the parent carries no ordering block, -and passage (g)'s old sentence is the one site where the parent is present and the change removes -it. Nothing is claimed as "contradictory" — the second of the two defects Gate B found in the -`fic2` instrument. +**The counterfactual splits, and stating it as one understates what is owed.** For the **ordering +block** it is **ABSENT and is claimed as absent**: the parent carries no such block, so no old +wording of it can be shown to disappear and presence alone is the check. For **every replacement** +the parent carries the old wording and the change removes it, so each owes the +old-wording-gone half of its pair — and the replacements are not one site: the target text +replaces standing wording in §§B, C, D, E, F and H, passage (g) among them. **Which sites those +are the plan derives from the target text**, where the concrete replacements live. Nothing is +claimed as "contradictory" — the second of the two defects Gate B found in the `fic2` instrument. **The named verification of the risk path** (story AC 4) is a **next-state table**, written in the plan and quoted by the closing commit body. **Its claim is narrow and stated as such: it covers @@ -254,8 +259,9 @@ trigger, no-clean-credit being the clean predicate's own second half — and eve block names must hold **through the closing commit**, not merely during the pass, which is the window the block's own "in between" wording fixes. Naming only the distinct-state half would pass the exact no-progress defect AC 4 cites from the parent cycle. **The -consumption clause is what keeps the oracle and §3 in agreement**: §3 says continue consumes the -reading that raised the suspension and a further health suspension needs it recomputed over a +consumption clause is what keeps the oracle and the shipped text in agreement**: the target text's +§A says continue consumes the reading that raised the suspension and a further health suspension +needs it recomputed over a pass run after the answer, so the *same* two-tell or clearly-stuck result **after** such a pass is new data and a legitimate row, not a failed transition. An earlier wording failed it "whether or not an input was consumed", which would have classified that legitimate case as a defect. @@ -281,19 +287,19 @@ repo's most persistent defect. The transport that could carry it left with the r - **Invariant 11 — `docs/prompt-standards.md`, all 12 items.** Most at risk: item 6, every constraint in the shipped block carrying its reason in the same sentence — **including the - three exempted until pass 6**, which now carry theirs inline in §3: that no other pass outcome - makes a cycle eligible to close, that a zero-finding pass is clean whatever the floor, and that - decline is available only at a membership stop. The exemption was wrong twice over: item 6 + three exempted until pass 6**, which now carry theirs inline in the target text's §A: that no + other pass outcome makes a cycle eligible to close, that a zero-finding pass is clean whatever + the floor, and that decline is available only at a membership stop. The exemption was wrong twice over: item 6 admits no "settled elsewhere" clause, and a scaffolded copy cannot reach the story the reasons were said to live in. The unknown-start clause carries its reason inline for the same reason. Then **item 8 (token-lean), which an earlier revision claimed on the wrong ground** — it said the block replaces closure sentences rather than adding beside them, while the block was in fact restating triggers, duties, preconditions and the severity answer that their own paragraphs still defined, which is two authorities per copy. The claim now rests on what the - block does: it is authoritative for the evaluation order and for closure and **cites** every - other rule where that rule is defined, so each has one definition in the shipped text and §3's - table is the check. Then item 3 (the stop answer produces a named state, **parked**, with its - own restart transition). + block does: **it owns the six things §2 names** — not restated here, one statement of that + boundary being the point — and **cites** every other rule where that rule is defined, so each + has one definition in the shipped text; **§4's site table is the check**. Then item 3 (the stop + answer produces a named state, **parked**, with its own restart transition). - **Don't: "Never replace a decision procedure without accounting for its old conditions."** Satisfied by the plan's per-condition disposition list against the committed inventory (§5), with this spec's passage map keeping the ten-passage set complete. diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index 48ddca9..e3aae66 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -79,7 +79,12 @@ defines, and a **clean pass** is the predicate here, read on the logical pass wi branch file combined, so one branch's clean file never establishes a clean pass. **A pass is clean** when its findings carry no in-set Blocker or Major at effective severity and **no scope-stop trigger** — the two the absorb paragraph defines, read there and not redefined here, -each already carrying the qualification an answer given in this cycle puts on it. A pass with +each already carrying the qualification **an answer given before that pass ran** puts on it. +**A pass's cleanliness is settled on what it found and on the answers standing when it ran**, and +a later answer never rewrites it: an answer discharges the holds it was asked for and leaves the +pass that raised them exactly as clean or unclean as it was, which is the same fact the duties +paragraph states of no-clean-credit. Reading a later answer back onto an earlier pass would let a +cycle close on a pass that was surfaced, answered and never re-run. A pass with **zero** findings is clean whatever the floor, because a floor buys further looks at an artifact that keeps yielding findings, and one yielding none has already given what those looks were for; don't manufacture findings to pad. @@ -92,14 +97,23 @@ is not thereby made unclean. Keeping the two apart is what lets the order decide they disagree, which the branches below do. **The closing act, by cycle kind.** A **Gate-B** cycle closes with the closing amend -Mechanics · Finishing the cycle describes. A **Gate-A** cycle has no WIP snapshot and no amend: -it closes on **the author's recorded acceptance of the revision the clean pass reviewed**, which -is a commit that in the ordinary case already exists — the pass reviewed a committed revision — -so the closing record goes in that commit where it is still the branch tip and in the next commit -on the branch where it is not. **No new revision of the artifact is made to close a Gate-A -cycle**, because a new revision is one no pass has reviewed. Nothing is closed before that act, -and a profile, cited set or — in a Gate-B cycle — evidence entry that changes in between still -gates it. **A commit the hook reads as cycle-closing is a separate matter**: a non-`WIP` commit +Mechanics · Finishing the cycle describes. A **Gate-A** cycle has no WIP snapshot to replace, so +its closing act is **the commit body carrying this cycle's provenance line and its per-pass +curve, naming there the revision the clean pass reviewed** — the two records Mechanics already +obliges every cycle to write, and writing them is the author's recorded acceptance of that +revision. **No separate act, form or record is introduced for Gate-A closure**, because a cycle +that owes two records already has somewhere to say what it accepted. That body goes in the commit +the clean pass reviewed where it is still the branch tip — **an amend of the message alone, which +leaves the tree and so the reviewed revision untouched** — and in the next commit on the branch +where it is not. **No new revision of the artifact is made to close a Gate-A cycle**, because a +new revision is one no pass has reviewed. Nothing is closed before that act, and what changes in +between still gates it: the **profile**, the **cited set**, the **assigned fix set**, the +**revision of the artifact the clean pass reviewed**, and — in a Gate-B cycle — the **evidence +entry**. Any of them differing at the closing act from what that pass read makes that pass +non-final and owes another, which is the same answer a mid-pass change already gets. **Sameness is +read on the artifact and the duties, never on the branch tip**: writing the closing body is itself +a commit, so a commit that only records the closure changes neither, while an edit to the reviewed +text or a broadening of the set changes one and costs the pass. **A commit the hook reads as cycle-closing is a separate matter**: a non-`WIP` commit mid-cycle makes the hook read the cycle as closed and discards the passes counted so far, which is an observation about the counter — the cycle itself stays open until the conditions above hold, so an accidental commit destroys the pass credit and closes nothing. A plateau or tells on @@ -247,93 +261,125 @@ false report. Below the floor the pass **suspends**, clean completion having not --- -## B. Passage (b) — what a loop absorbs — REPLACED, in part - -This passage's replacements are split across two sections, so neither is the whole list: **`b7`, -`b8`, `b11`, `b13` and `b16` are here**, and **`b3`, `b12` and `b17`–`b18` are in §H** with the -other site-by-site replacements. Every other sentence in the paragraph is carried. The five here, -final: - -**`b7`, the assigned fix set.** -``` -**The assigned fix set is fixed before the pass you are answering: it is the union of the scope -every approved story or plan governing this change assigns to this cycle, plus repair obligations -you already accepted in earlier passes, minus every finding this cycle has declined.** -``` -*Why:* the singular "the approved story or plan" leaves a cycle governed by several with no set -at all, and a declined finding has to leave the set or the resolve duty reaches it. - -**`b8`, the membership test.** -``` -A finding is in-set when repairing it stays inside **the assigned fix set as `b7` computes it** — -never merely because it arrived in the current pass, which would put every new finding in the set -by definition and leave the boundary deciding nothing. -``` -*Why:* tested against "that scope" independently, a broadened scope puts back a finding `b7` -keeps out. - -**`b11`, the membership trigger.** -``` -**A correction that leaves that set stops the loop like any other out-of-scope finding** — except -one this cycle has already declined, which is outside the set by that decision and raises no -trigger — even when it opens no new question at all, -``` -*The sentence continues with `b12`'s replacement, which is in §H* — the trailing resumption clause -is `b12`, not `b11`, and quoting its live wording here as final would ship the immediate -resumption the ordering replaces. - -*Why:* unqualified, this and the clean predicate decide a re-raised declined finding in opposite -directions. - -**`b13`, the question trigger.** *New* is qualified: -``` -…opens a **new structural or contract question** — new meaning not already answered in this cycle, -so an answered question raised again stops nothing… -``` - -**`b16`, novelty over ancestry.** The live sentence says "the new question wins and the loop -stops"; it gains one clause: -``` -**When a finding is both** — it corrects the last correction *and* opens a new structural or -contract question — **the new question wins and the loop stops**: novelty overrides correction -ancestry, because absorbing on ancestry is exactly how a contract decision gets made without -anyone choosing it. **Novelty overrides ancestry and nothing else: where the finding is also out -of set, both triggers hold and both answers are owed**, since a question answered about a finding -nobody placed in or out of the set leaves its membership decided by default. -``` -*Why (pass 17 finding 4):* without it an agent follows the surviving source, asks only the -structural question, and absorbs out-of-set work silently. - -**Carried unchanged:** the paragraph's opening and every sentence named in neither this section -nor §H. +## B. Passage (b) — what a loop absorbs — REPLACED, whole passage + +**The whole paragraph is written out here, as it will read.** Earlier revisions split its +replacements across two sections and left one sentence half in each, which is how §B and §H came +to instruct the plan differently about the same clause. Nothing about this passage is stated +anywhere else in this file. + +**Changed:** `b3`, `b7`, `b8`, `b11`, `b12`, `b13`, `b16`, `b17`–`b18`. **Carried:** every other +sentence. The rationales follow the text. + +``` +**What a loop absorbs, and what stops it — a question of scope, not of action.** A finding +that corrects the correction you just made **and stays inside the assigned fix set** is +**inside this loop's scope**: keep it here rather than handing it back, then act on it by its +severity exactly as Mechanics · Severity says. Ancestry decides where a finding belongs; it +never decides what you do with it, and it grants no Minor or Nit a repair round it would not +otherwise get. **The assigned fix set is fixed before the pass you are answering: it is the +union of the scope every approved story or plan governing this change assigns to this cycle, +plus repair obligations you already accepted in earlier passes, minus every finding this cycle +has declined.** A finding is in-set when repairing it stays inside **the assigned fix set as +just defined** — never merely because it arrived in the current pass, which would put every new +finding in the set by definition and leave the boundary deciding nothing. Where membership is +genuinely unclear treat the finding as **outside**, which costs a question and never a silent +expansion. **A correction that leaves that set stops the loop like any other out-of-scope +finding** — except one this cycle has already declined, which is outside the set by that +decision and **raises no membership trigger on that account**, its membership being the one +question already answered — even when it opens no new question at all, and the membership +answer ends that finding's membership hold; **what the pass does next is the closure +ordering's**, which resumes only when every answer outstanding on that surface has been given. +A finding that opens a **new structural or contract question** — new meaning not already +answered in this cycle, so an answered question raised again stops nothing — stops the loop and +goes to the user — **size is not the test, novelty of the question is**, so a structural finding +that is genuinely small still stops it, while a long correction still aimed at the last +correction does not — provided that correction, too, stays inside the set, which its ancestry +never supplies on its own. **When a finding is both** — it corrects the last correction *and* +opens a new structural or contract question — **the new question wins and the loop stops**: +novelty overrides correction ancestry, because absorbing on ancestry is exactly how a contract +decision gets made without anyone choosing it. **Novelty overrides ancestry and nothing else: +where the finding is also out of set, both triggers hold and both answers are owed**, since a +question answered about a finding nobody placed in or out of the set leaves its membership +decided by default. Stopping this way is **not an exit from the gate**: it is a **suspension** +in the closure ordering's sense, the floor, the Blocker/Major filter and the clean-final-pass +rule all stand, and what the answer does is stated there — what the stop prevents is a loop +committing you to a design you never chose, which is a different failure from an unfinished +review. +``` + +**The field-mint parenthetical that closes this paragraph in `CLAUDE.md` is carried unchanged and +is absent from the template**, which is a pre-existing parity divergence this change neither +creates nor removes. The plan's divergence list carries it. + +*Why `b3`:* the four severity actions are restated beside a pointer to the section that defines +them, which is two authorities per copy. The template took C's wording on the inventory's +reasoning that it has no Mechanics section, which is false. + +*Why `b7`:* the singular "the approved story or plan" leaves a cycle governed by several with no +set at all, and a declined finding has to leave the set or the resolve duty reaches it. + +*Why `b8`:* tested against "that scope" independently, a broadened scope puts back a finding the +set-defining sentence keeps out. It reads "as just defined" and not by condition id, an inventory +id being this file's vocabulary and meaningless in the installed prompt. + +*Why `b11` (pass 19 finding 3):* unqualified, the exception and the clean predicate decide a +re-raised declined finding in opposite directions. Scoped to the **membership** trigger, since +that is the question the decline answered; the question trigger in the next sentence is +independent and reaches a declined finding like any other. + +*Why `b12`:* it is the immediate-resumption command, **replaced** rather than appended to, because +the ordering resumes on every outstanding answer and not on one. + +*Why `b16` (pass 17 finding 4):* without the added clause an agent follows the surviving source, +asks only the structural question, and absorbs out-of-set work silently. + +*Why `b17`–`b18`:* the stop is named a suspension and defers to the ordering rather than restating +what ends it. --- ## C. Passage (c) — recognizing clearly stuck — REPLACED, from the third condition -``` -…and **Blocker or Major findings that keep regenerating across genuine repair attempts**, each -round's fix producing the next — **or a finding the author has validly dismissed that the reviewer -re-raises across passes**, the re-raise standing in for the regenerating fix, since a dismissal -gets no repair and produces none and a false positive that returns every pass would otherwise -leave the cycle unable to close and unable to suspend. That third condition is what makes a -plateau rather than a finish. **Where this reading and a clean completion both apply, the closure -ordering decides it** — the precedence sentence lives there, because precedence is evaluation -order. -``` - -*Everything after that in the live paragraph — the precedence sentence and the below-the-floor -sentence — moves into the block above, word for word, which is what satisfies **D3**.* +**The three-condition sentence entire**, so nothing here begins mid-clause. **Changed within it:** +the third condition gains the re-raised-dismissal clause. The plateau and coverage conditions are +carried word for word. Of the two sentences after it, the first is the live sentence's opening +clause left standing alone — its second half, the precedence sentence, moves into the block — and +the second is new. + +``` +So this exit needs three things **together**, and a missing one means keep going: a plateau +visible across passes (six or more is where the field saw one); an **affirmative judgement that +coverage is sufficient**, stated — a known materially unreviewed area forbids this exit outright, +and disclosing it does not license it; and **Blocker or Major findings that keep regenerating +across genuine repair attempts**, each round's fix producing the next — **or a finding the author +has validly dismissed that the reviewer re-raises across passes**, the re-raise standing in for +the regenerating fix, since a dismissal gets no repair and produces none and a false positive that +returns every pass would otherwise leave the cycle unable to close and unable to suspend. That +third condition is what makes a plateau rather than a finish. **Where this reading and a clean +completion both apply, the closure ordering decides it** — the precedence sentence lives there, +because precedence is evaluation order. +``` + +*What follows in the live paragraph, and the two are not treated alike (pass 19 finding 8).* The +**precedence sentence** moves into the block above **word for word**, which is what satisfies +**D3**. The **below-the-floor sentence does not move at all — it is replaced**: it says a +Blocker/Major-free pass below the floor carrying a Minor "keeps looping", while the ordering +splits that case, such a pass **suspending** where any suspension applies to it and **continuing** +where none does. Its two halves live in the ordering's second and third branches, and no copy of +the live wording survives beside them — carrying it word for word would install an unconditional +continuation next to the conditional one and give the same pass two answers. --- ## D. Passage (e) — the five tells — REPLACED, in part -**`e7`, the threshold.** Gains one clause: +**`e7`, the threshold.** Gains one clause; the sentence is given entire. ``` **Any two present makes stop-and-surface mandatory, not discretionary** — read **after** the clean-completion branch of the closure ordering, which outranks it (**D2**) — and you report the -tells and hand the decision to the user… +tells and hand the decision to the user, and the "clearly stuck" reading above is not a +precondition for it. ``` **A pointer is added** at the end of the passage: @@ -389,9 +435,9 @@ demotes it — the two counts are meant to differ. --- -## F. The three standing sentences this change falsifies — REPLACED +## F. The four standing sentences this change falsifies — REPLACED -Each is a live sentence that the block makes wrong. All three are **known contradictions** and +Each is a live sentence that the block makes wrong. All four are **known contradictions** and none is deferred. **1. The `WIP:` naming warning** (Mechanics · `baseSha`). @@ -423,6 +469,21 @@ change in the mix. *Why (pass 17 finding 7):* the live rationale says the severity rule "moves the Blocker/Major line rather than the total", which contradicts the pre-ceiling decision above. +**4. The human-exception scope sentence** (Mechanics, recording a human exception). It wraps +across C 1014–1015 and W 1198–1199. +``` +Those have their own terminal actions and this paragraph changes none of them: on a STOP you +still stop, and **neither a human's general assent nor this record** lets an agent close or +continue a cycle. **The answer a suspension asks for is not assent of that kind**: continue and +stop are the answers the closure ordering prescribes, given on the question that suspension +raised, and what each produces is stated there. +``` +*Why (pass 19 finding 4):* the live sentence says no human answer lets an agent continue a cycle, +while the ordering makes **continue** the prescribed answer that restarts a parked one. Left as +it stands, an agent following it refuses the exact transition the ordering requires. The +distinction the repair draws is between a human waving a rule through — which this paragraph +still forbids — and answering the question a suspension actually asked. + --- ## G. The one-contract paragraph — REPLACED @@ -433,10 +494,14 @@ slot naming, the provenance line, the curve, the carry rule, the unknown-start a semantics **and the closure ordering together with every rule it reads** depend on one another, and the requirement is that the adopted definitions **agree**, not merely that all of them are present. **Membership is decided by a test a reader can apply to the text in front of them, with -no list to consult: a live rule belongs to this contract when changing it would change an input -the closure ordering reads, which branch a pass takes, what a hold is or what discharges it, +no list to consult, and the test reads what a rule states rather than what changing it would do: +a live rule belongs to this contract when what it says determines or supplies an input the +closure ordering reads, which branch a pass takes, what a hold is or what discharges it, whether a cycle may close, or the production, identity or transport of any record this section -obliges a cycle to write.** The last clause is why the squash carry belongs: it moves no pass and +obliges a cycle to write.** Asking instead what an imagined edit would do decides nothing, because +any rule can be edited into deciding a branch and none decides one when edited cosmetically, so +membership would follow the edit a reader pictured rather than the text in front of them. The +last clause is why the squash carry belongs: it moves no pass and decides no branch, and a record that does not survive the merge is a record the cycle did not produce. A curve without a cycle field cannot be attributed, a slot rule without a nonce has nothing to key on, a carry rule naming records a project does not produce is @@ -459,32 +524,8 @@ rather than doing the noticing. ## H. The remaining replacements, written out — REPLACED Each is the final wording. Nothing here is a paraphrase; the plan installs these strings. - -**`b3`** — Mechanics is pointed at, and the four severity actions are no longer restated beside it. -Both copies take C's wording; W's "the severity rule" pointer was written on the inventory's -reasoning that W has no Mechanics section, which is false. -``` -keep it here rather than handing it back, then act on it by its severity exactly as -Mechanics · Severity says. -``` - -**`b12`** — the immediate-resumption command is **replaced**, not appended to. It is the only -sentence in the standing text that resumes a loop on one answer, and the ordering resumes on all -of them. -``` -and the membership answer ends that finding's membership hold; **what the pass does next is the -closure ordering's**, which resumes only when every answer outstanding on that surface has been -given. -``` - -**`b17`–`b18`** — the stop is named a suspension and defers to the ordering rather than restating -what ends it. -``` -Stopping this way is **not an exit from the gate**: it is a **suspension** in the closure -ordering's sense, the floor, the Blocker/Major filter and the clean-final-pass rule all stand, -and what the answer does is stated there — what the stop prevents is a loop committing you to a -design you never chose, which is a different failure from an unfinished review. -``` +**Passage (b) is not among them** — it is written out whole in §B, which is the only place this +file states anything about it. **`c18` and the surfacing sentence** — the two clauses findings 7 and 8 name. The resolve and hold duties are kept; the blanket no-clean-credit and the one-answer resumption go, because the @@ -548,10 +589,12 @@ Lenses change what a pass asks, never how many a cycle owes: **the lens sets** l rule in this section alone. ``` -**The unknown-start strict-reading list** — gains the suspensions, reason inline. +**The unknown-start strict-reading list** — gains the suspensions, reason inline. The whole +dash-delimited list is given, the last item being the addition. ``` -…the curve duty owed, the nonce duties at their strictest, and **every suspension binding, since -starting rules that cannot be established cannot be read as having waived an open hold**. +— at minimum floor 3, severity classified without the demotion, the provenance-line duty owed, +the curve duty owed, the nonce duties at their strictest, and **every suspension binding, since +starting rules that cannot be established cannot be read as having waived an open hold** — ``` --- From 5f6e81ce7eeed3b7dd5dd29b03d7211f5d07a116 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Fri, 11 Sep 2026 18:58:28 +0200 Subject: [PATCH 036/181] =?UTF-8?q?docs(context):=20Gate-A=20pass=2020=20?= =?UTF-8?q?=E2=80=94=20B+M=2011=20to=206,=20findings=20open?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 9 findings, 1 Blocker, 5 Majors, 3 Minors. Valid pass. One tell of five (Blockers flat at 1); no mandatory stop. Majors halved and Blocker+Major fell to 6, the second-best of the cycle. Three of the nine regenerate from pass 19's own repairs, one of them structurally: the closing-time preconditions added to the block collide with the block's own six-thing ownership boundary, which design §2 also still states as citation-only. The Blocker is pass 19's Gate-A closing act. Verified against CLAUDE.md: neither the provenance-line grammar nor the curve grammar has a field for a revision, and the provenance line admits no informal variant, so naming the reviewed revision in them is not expressible. Findings not yet dispositioned; routed to the reviewer first. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-20.md | 10 ++++ .../gate-a-spec-awsf1ec771-resume.md | 47 ++++++++++++++++++- 2 files changed, 56 insertions(+), 1 deletion(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-20.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-20.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-20.md new file mode 100644 index 0000000..7cbf7df --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-20.md @@ -0,0 +1,10 @@ +BLOCKER | high | §A "The closing act, by cycle kind" | Gate-A closure requires the commit body carrying the fixed provenance line and curve to name the revision the clean pass reviewed, but neither standing record grammar has a revision field, each forbids an informal variant, and the non-tip path puts the body in a different commit; "revision" is also left undefined between artifact content and commit identity | A Gate-A cycle whose clean pass reviewed an already-committed non-tip revision has no specified closing act it can perform and, because the pass is eligible, no suspension path, so the cycle remains unable to close or suspend | Define a checkable Gate-A closing act using the existing fixed forms and say exactly how the reviewed artifact revision is identified on both the tip-amend and later-commit paths, or change the sequencing so commit placement itself identifies it +MAJOR | high | §H "Surfacing does not close the cycle" | The replacement still says every surfaced finding is "still open" and that the resolve rule is not waived, while §A newly says a re-raised valid dismissal stays discharged and creates only a clearly-stuck hold | A repeatedly re-raised false positive can require a second dismissal or repair under §H but only a continue-or-stop answer under §A, so the valid-dismissal path is not idempotently decided | Qualify the surfacing sentence so an already validly dismissed recurrence remains resolved for the resolve duty while its newly surfaced health hold stays open +MAJOR | high | §A "Everything else it names it cites" | The opening says every closure precondition keeps its definition outside the block and that finding one defined here is a defect, but the block itself newly makes the reviewed artifact revision and assigned fix set closing-time preconditions and defines their mismatch effect; the design §2 repeats the citation-only claim, while no standing source defines either closing-time test | The authority boundary and one-contract membership are false, so the plan can duplicate, omit, or relocate two conditions that now decide closure | Either include these closing-time sameness tests in the block-owned side of the six-part boundary or define them at an existing source and leave only an exact citation in §A; make design §2 agree +MAJOR | high | design §7 "each final-acceptance precondition" | The required separate verification checks enumerate only the floor, cited set and profile, and Gate-B evidence revalidation, but target §A now also gates the closing act on an unchanged assigned fix set and reviewed artifact revision | The named verification can pass while the exact pass-19 blocker recurs: scope or artifact content changes after the eligible pass and before closure without the promised next pass | Add the assigned-fix-set and reviewed-artifact-revision closing-time checks to the plan obligation, while leaving fragments and counts to the plan as intended +MAJOR | high | standing Mechanics "Finishing the cycle" | Both live copies still say "after the final clean pass, close it with git commit --amend" and do not reference the closure ordering, although §A makes a clean eligible pass insufficient when any cycle precondition is unmet and claims the ordering is referenced everywhere else; §F does not replace this sentence | A reader entering through Mechanics can amend and close Gate B over an unresolved earlier Major, standing hold, changed evidence entry, or other unmet precondition | Replace the lead-in with a pointer that the amend is performed only when the closure ordering reaches Gate B's closing act, leaving Mechanics authoritative only for the amend operation +MAJOR | high | §H "this paragraph states the floor and nothing else" | The proposed replacement makes a mechanically false scope claim: the same standing paragraph still states hook-threshold handling, the floor-knob rule and TodoWrite duty, and the replacement itself immediately retains validation and dismissal-recording duties | The prompt simultaneously instructs the agent to preserve non-floor duties and declares that the paragraph contains none, undermining the claimed one-authority structure and making those duties easier to ignore during partial adoption | Narrow the claim to the removed closure sentences, for example that this span no longer defines completion or the zero-finding exit, without describing the whole paragraph as floor-only +MINOR | high | proposed text references "D2", "D4", "D5" and "D7" | Editorial decision labels from the story remain inside the text that will be installed byte-identically, but the scaffolded downstream prompt ships neither the story nor a D1-D8 table | A downstream reader sees unresolvable authorities in the operative ordering, contrary to the target's shipped-text framing and the self-contained-template exception in prompt standard 11 | Remove the D-label citations from proposed text and retain the already-inline substantive rules and reasons +MINOR | high | §A "closing report" | The ordering introduces a "closing report" distinct from an "ordinary report" without defining either carrier, while the standing text defines only each pass's status report to the user and fixed commit-body records | An agent can put plateau or tell material into the commit body, omit it from the required pass report, or invent a new report form, so the pass-that-closes reporting path is not checkable | Call it the closing pass's existing status report to the user and use the same carrier for eligible-but-blocked pass observations +MINOR | high | design §6 "b3, target text §H" | The parity section points the b3 wording decision to target text §H, but §H expressly says passage (b) is not there and b3 now appears only in §B | The parity handoff sends the plan to a section that disclaims ownership of the sentence, weakening the exact extraction check | Change the stale target citation from §H to §B +END OF FINDINGS (9 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index 3b2ade8..1d48860 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -46,7 +46,52 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 17 | c2d5fe1 | 8→**9** | 1→**1** | 4→**7** | yes | **TWO-TELL STOP — mandatory, surfaced to Daniel.** B+M 5→8. **Three findings (4, 5, 7) are one mechanism on its third pass**: a standing §5 sentence the block falsifies that §4 does not name. All 9 held open; session 01a0901b-ecf3-73b2-886c-5b3c7dc63a74 | | 18 | 96c3611 | 9→**15** | 1→**0** | 7→**12** | yes | **first pass on the target text. ZERO BLOCKERS, first since pass 7.** Finding 1 names the restructure as half-done: §H is still paraphrase, not text. Findings 6,7,8 name three standing sentences by line number — the 15/16/17 mechanism, now findable; session 01a09068-de4a-7071-97a7-82cc4774544f | | 19 | 6d52fcc | 15→**13** | 0→**1** | 12→**10** | yes | precheck caught a silent loss before the pass: `a20` ("don't manufacture findings to pad") had fallen out of the target text; restored to §A. Two known contradictions repaired in the same commit. **Only 6 of 13 findings are against the target text**; 5 are against the design spec and 2 against standing §5; session 01a0912a-7329-7450-bdb3-da043e88e7bb | -| 20 | — | — | — | — | not run | next action, after pass 19's findings are dispositioned | +| 20 | 5174d7a | 13→**9** | 1→**1** | 10→**5** | yes | **B+M 11→6, Majors halved.** Three of nine (3, 6, 9) regenerate from pass 19's own repairs; the Blocker is pass 19's Gate-A closing act, which names a revision in two record forms that have no field for one; session 01a0915d-6a5b-7633-a60a-881f0022816d | +| 21 | — | — | — | — | not run | next action, after pass 20's findings are dispositioned | + +## Pass-20 three-line report + +**Floor line:** derived floor **3**; risk **high**, security **none**; read fresh from +`docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md`. One cited story, level 2. + +- **Trend:** findings …, 9, 15, 13, **9**. Blockers …, 1, 0, 1, **1**. Majors …, 7, 12, 10, **5**. + Blocker+Major …, 6, 8, 5, 8, 12, 11, **6** — level with the pass-14 low, above pass 16's 5. +- **Cluster (pass 20):** product 7 of 9 (1, 2, 3, 5, 6, 7, 8); the instrument 1 (4, §7's named + verification); prose about it 1 (9, a stale §H pointer in design §6). +- **require↔withdraw:** none. **Three near-misses, all the same shape and all named:** findings 3, + 6 and 9 object to additions pass 19 made — the mirror of the pair's shape, a later pass + questioning an earlier pass's addition, not a demand for something removed. + +**Tells: one of five** — the Blocker count flat at 1. Findings fell 13 → 9, Majors halved 10 → 5, +the cluster is product. One is not two: no mandatory stop. The clearly-stuck exit is not reachable +either, its first condition needing a plateau where the curve just fell. + +**The repair round worked and left a third of the next pass behind.** B+M 11 → 6 is the second-best +of the cycle. But **findings 3, 6 and 9 are pass 19's own repairs regenerating**, and one of them +is structural rather than wording: repairing pass-19 finding 1 put two closing-time preconditions +**inside** the block, which collides with the block's own "owns exactly six things, cites +everything else" boundary — and design §2 repeats the citation-only claim. That is the same +mechanism as the §7 assert list at pass 12 and the four-descriptions problem at pass 17: a repair +made at one site falsifies a claim another site still makes. + +**The Blocker is pass 19's repair failing on the records it cites.** Verified against `CLAUDE.md`: +the provenance-line grammar is `; floor per ; hook reminder threshold +` and the curve grammar carries cycle, kind, passes, models and counts — **neither has a +field for a revision**, and the provenance line states outright that there is no informal variant. +So "the commit body carrying the provenance line and curve, **naming there the revision the clean +pass reviewed**" asks two fixed forms to say something they cannot express. On the non-tip path the +body also lands in a different commit from the reviewed one, so placement does not identify it +either. + +**Finding 5 is the standing-sentence mechanism again, on its fifth pass.** Mechanics · Finishing +the cycle still reads "after the final clean pass, close it with `git commit --amend`" at C 827 and +W 1011, which the ordering falsifies — a clean eligible pass with an unmet precondition must not +close — and **§F does not replace it**. Same shape as passes 15, 16, 17 and 19. + +**Findings 7 and 9 are the same defect class the §B merge already caught once.** Editorial +vocabulary leaking into shipped text: four `**D2**`/`**D5**`/`**D7**` labels sit inside fenced +blocks that install byte-identically, and the scaffolded template ships neither the story nor a +D-table. The merge caught `` `b7` `` doing this at pass 19; nobody swept for the rest. ## Pass-19 dispositions — one connected repair round, on the reviewer's order From 8dc22fb8a74ed6f0493e1d937451fd7f7ce4f341 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Fri, 11 Sep 2026 19:47:34 +0200 Subject: [PATCH 037/181] docs(specs): apply Gate-A pass 20; correct two claims of my own MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Two claims in my pass-20 report were wrong and are corrected here. Finding 6 is not a pass-19 repair regenerating. git log -S puts the sentence in ef9a504, the pass-18 repair, unchanged since; pass 19 relabelled the span around it without touching it. The honest split is 1, 3, 4 and 9 from pass 19's repairs, and 2, 5, 6, 7, 8 pre-existing and missed earlier. The Blocker diagnosis was too strong. §A requires the revision named in the commit body, not inside the provenance line or the curve, and their fixed grammars forbid no accompanying prose in the same body. Missing fields alone prove no impossible closure. The real defect was narrower: the naming was undetermined. The closing act now states it as prose beside the two untouched lines, carrying the artifact path and the full 40-character object name, and says outright that it records and does not prove. Both placements are written out, and walking them found two gaps now closed: the tip-path amend supersedes the name the line records, so the text says a reader takes the tree of the commit the body sits in; and "the next commit on the branch" is now "the next commit made", which may carry only the record on the allowance Mechanics already grants. 3 and 4: the finality checks belong to the closure decision the block owns, so the citation claim is corrected rather than the tests moved. Design §2 points at §A instead of restating the boundary, two copies of it being what drifted, and §7's evidence duty covers every closure condition the block states. 2 and 5 are one mechanism and both are replaced rather than qualified. Surfacing no longer reopens a discharged resolve duty; Finishing the cycle becomes §F's fifth sentence, stating the operation while the ordering states the permission. 6 drops "and nothing else". 7, 8 and 9 collected: four D-labels removed from shipped blocks, the closing report bound to the existing status report, design §6's b3 pointer moved to §B. Gate A stays open; pass 20 was not clean. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-resume.md | 47 ++++++++ ...26-09-10-loop-rule-consolidation-design.md | 26 +++-- ...-10-loop-rule-consolidation-target-text.md | 101 +++++++++++++----- 3 files changed, 134 insertions(+), 40 deletions(-) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index 1d48860..3302d8a 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -49,6 +49,53 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 20 | 5174d7a | 13→**9** | 1→**1** | 10→**5** | yes | **B+M 11→6, Majors halved.** Three of nine (3, 6, 9) regenerate from pass 19's own repairs; the Blocker is pass 19's Gate-A closing act, which names a revision in two record forms that have no field for one; session 01a0915d-6a5b-7633-a60a-881f0022816d | | 21 | — | — | — | — | not run | next action, after pass 20's findings are dispositioned | +## Pass-20 dispositions — and two of my claims corrected, both verified wrong + +**Correction 1 — the provenance claim was wrong.** I reported findings 3, 6 and 9 as pass 19's own +repairs regenerating. **Finding 6 is not**: `git log -S` puts "this paragraph states the floor and +nothing else" in **`ef9a504`**, the pass-18 repair, unchanged since. My pass-19 edit relabelled the +span around it and never touched the sentence. The honest split, which keeps new repair damage +apart from old misses: + +| Origin | Findings | +|---|---| +| **pass 19's own repairs** | 1 (the Blocker), 3, 4, 9 — four, not three | +| **pre-existing, missed by earlier passes** | 2, 5, 6, 7, 8 | + +**Correction 2 — the Blocker diagnosis was too strong.** I wrote that the provenance and curve +grammars have no revision field and concluded the closing act was impossible. **§A requires the +naming in the commit *body*, not inside those two lines**, and their fixed grammars forbid no +accompanying prose in the same body. The missing fields alone therefore prove nothing. What is +real is narrower and still a defect: **how that naming looks and what it identifies was +undetermined.** The repair states it as prose beside the two untouched lines, carrying the +artifact path and the full 40-character object name, and says outright that it records and does +not prove. + +| Findings | Disposition | +|---|---| +| 1 | **§A closing act rewritten.** Two fixed lines untouched; the revision naming is prose in the same body; both placements written out; explicitly not evidence and not durable provability. | +| 3, 4 | **The finality checks belong to the closure decision §A owns.** The citation claim is corrected rather than the tests moved. Design §2 now **points at** §A instead of restating the boundary — two copies of it are what drifted — and §7's evidence duty covers **every** closure condition the block states, the plan reading the set off the block rather than a list repeated in §7. | +| 2, 5 | **One mechanism, both replaced.** Surfacing no longer reopens a discharged resolve duty (resolution is repair **or** valid dismissal). Mechanics · Finishing the cycle becomes §F's fifth sentence: it states the **operation**, the ordering states the **permission**. | +| 6 | **"and nothing else" deleted.** The pointer to the ordering carries it. | +| 7, 8, 9 | **Collected, no round of their own.** Four `D2`/`D5`/`D7` labels removed from shipped blocks with the rule kept inline; "closing report" bound to the pass's existing status report; design §6's `b3` pointer moved §H → §B. | + +**§F is now five standing sentences, and three of them are one mechanism** — an entry point other +than the ordering carrying an unqualified instruction. Naming the mechanism is why each is +replaced rather than given an exception to point at. + +**The two Gate-A closing cases were walked concretely, as the reviewer required, and both needed a +fix before they worked:** +- **Reviewed commit still at the tip.** The amend rewrites that commit's object name, so the + naming line records a name that no longer resolves. Now stated, with what a reader does instead: + read the tree of the commit the body sits in, which the amend preserves. +- **Reviewed commit behind the tip.** "The next commit on the branch" read two ways — the commit + that already follows the reviewed one, or the next one made. Now "the **next commit made**", and + it may carry only the record, on the empty-commit allowance Mechanics already grants a + human-exception record (`CLAUDE.md:995`, `workflow-init.md:1179`, verified). + +**The shipped text is now free of this file's own vocabulary.** No inventory ids, no D-labels, no +editorial ellipses inside any fenced block — checked mechanically, not by eye. + ## Pass-20 three-line report **Floor line:** derived floor **3**; risk **high**, security **none**; read fresh from diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index c755fd0..477685b 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -57,13 +57,14 @@ establish the **inventory** of findings, not their resolutions. What the table does not settle, this spec decides in the section that uses it: the evaluation order and the file set each predicate reads; the duties' classification; which stop each of the scope stop's two triggers raises and what each answer does; what a clearly-stuck or two-tell -answer produces; and the raw-severity rule for the health measures. **The block owns exactly six -things** — the evaluation order, closure and its eligibility, the hold and what discharges it, -the composition of several suspensions, the pairs that cannot co-occur, and the clearly-stuck -precedence sentence as a stated exception — **and defines no trigger, no severity rule and no -closure precondition of its own**; each of those keeps its one definition where it already lives, -and where one had to change to agree with the ordering it changed **at its source**. The target -text's §A states the same six-part boundary in its own opening, and the two must not drift. +answer produces; and the raw-severity rule for the health measures. **The block's ownership +boundary is stated once, in the target text's §A opening, and is deliberately not restated here** — +two copies of it are what let them drift, which is pass 20 finding 3. What this spec records is the +decision behind it: the block defines no trigger and no severity rule of its own, and every closure +precondition **that has a source of its own** keeps its one definition there, changed **at that +source** where it had to change to agree with the ordering. **The closing-time sameness tests are +the exception in substance and not in principle**: they borrow no rule and have no other source, +being part of the closure decision the block owns. **The closure-ordering block is an addition beside the source edits**, not one of them. §4 lists the edits; **no total is stated here or there**, because the unit — one contiguous replacement at @@ -176,7 +177,7 @@ pre-existing wording differences are deliberate and stay, which are not and are **performs the extraction and diff**, passage by passage, against the real files. One divergence is decided here because it is a correctness call rather than a wording one: W's `b3` pointer names "the severity rule" on the inventory's reasoning that W has no Mechanics section, which is -false, so W takes C's wording (`b3`, target text §H). +false, so W takes C's wording (`b3`, target text §B). The same extraction runs a second check within each copy: that `b11` and `b13` as edited say what the block cites them as saying, **comparing the complete predicates and not a shared phrase** — @@ -234,9 +235,12 @@ answer-state transitions once the predicates producing them are established**, w `fic2` defect — a state's inputs must include every input the rule reads — answered by restricting the claim rather than by widening the table. So the plan writes **separate named checks** for what the table therefore does not establish: that a logical pass was validated -across every required branch file, and that each final-acceptance precondition the block cites -held — the floor, the cited set and profile, and the evidence entry's revalidation. **No fixture -per predicate is built**; that question is parked in the story's §2 and is not reopened. +across every required branch file, and that **every closure condition the block states held at the +closing act**. That set is **not enumerated here** — an enumeration is how this section came to +name three of them while the block states more, which is pass 20 finding 4. **The plan reads the +set off the block and writes one check per condition**, and fails where the block states a +condition the plan has no check for. **No fixture per predicate is built**; that question is parked +in the story's §2 and is not reopened. **That list is not exhaustive, and reading it as exhaustive is how the evidence entry would overclaim.** Two further things the table does not establish, named because they are the ones a diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index e3aae66..6dc411f 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -49,9 +49,12 @@ exactly six things**: the evaluation order, closure and its eligibility, the hol finding places and what discharges it, the composition of several suspensions, the pairs that cannot co-occur, and — as a stated exception, because precedence is evaluation order — the clearly-stuck precedence sentence quoted into it below. **Everything else it names it cites**: -the scope triggers, the assigned fix set, every severity rule, and each closure precondition -keep their one definition in the paragraph that owns them, and a reader who finds one of *those* -defined here has found a defect. +the scope triggers, the assigned fix set, every severity rule, and every closure precondition +**that has a source of its own** keep their one definition in the paragraph that owns them, and a +reader who finds one of *those* defined here has found a defect. **The closing-time sameness +tests below are defined here and are not an exception to that**: they borrow no rule and have no +other source, being part of the closure decision itself — the second of the six things this +paragraph owns — rather than a precondition stated elsewhere and read from here. **What a pass is read from.** Every finding-derived predicate reads the validated findings file **or files** of the logical pass as **the concatenation of their finding lines after each file @@ -98,14 +101,32 @@ they disagree, which the branches below do. **The closing act, by cycle kind.** A **Gate-B** cycle closes with the closing amend Mechanics · Finishing the cycle describes. A **Gate-A** cycle has no WIP snapshot to replace, so -its closing act is **the commit body carrying this cycle's provenance line and its per-pass -curve, naming there the revision the clean pass reviewed** — the two records Mechanics already -obliges every cycle to write, and writing them is the author's recorded acceptance of that -revision. **No separate act, form or record is introduced for Gate-A closure**, because a cycle -that owes two records already has somewhere to say what it accepted. That body goes in the commit -the clean pass reviewed where it is still the branch tip — **an amend of the message alone, which -leaves the tree and so the reviewed revision untouched** — and in the next commit on the branch -where it is not. **No new revision of the artifact is made to close a Gate-A cycle**, because a +its closing act is **the commit body that carries this cycle's provenance line, its per-pass curve +and, beside them, one line naming the revision the clean pass reviewed**. The first two are records +Mechanics already obliges every cycle to write and **neither is altered here** — their grammars are +fixed and neither has a field for a revision. The third is **ordinary prose in the same body**, +which those grammars neither supply nor forbid. Writing that body is the author's recorded +acceptance of that revision. + +**What the naming line says:** the artifact's repository-relative path, and the **full +40-character object name** of the commit the clean pass reviewed, an abbreviation being ambiguous +across repositories and across time. **What it is worth, said rather than implied:** it records +what the author accepted and **establishes nothing about what the pass actually read**, which no +reader can check from a commit body. It is a record, not evidence, and no durable provability is +claimed for it. + +**Both placements are written out, because they identify the revision differently.** Where the +reviewed commit is **still the branch tip**, the body goes into it as **an amend of the message +alone**: the amend gives that commit a new object name and leaves its tree untouched, so what the +pass reviewed changes as a commit and not as content. The naming line records the object name the +commit carried **when the pass read it**, which the amend then supersedes — **a reader wanting the +reviewed content reads the tree of the commit the body sits in**, that tree being the one the pass +read. Where the reviewed commit is **no longer the tip**, the body goes into the **next commit made +on the branch** — not the commit that happens to follow the reviewed one, which is already in the +past — and that may be **a commit carrying only the record**, which changes no content and so +raises no review obligation, exactly as Mechanics already allows for a human-exception record with +nowhere else to go. Placement says nothing on this path, so the naming line is the only thing +tying the closure to the revision, and it is written the same way. **No new revision of the artifact is made to close a Gate-A cycle**, because a new revision is one no pass has reviewed. Nothing is closed before that act, and what changes in between still gates it: the **profile**, the **cited set**, the **assigned fix set**, the **revision of the artifact the clean pass reviewed**, and — in a Gate-B cycle — the **evidence @@ -117,9 +138,11 @@ text or a broadening of the set changes one and costs the pass. **A commit the h mid-cycle makes the hook read the cycle as closed and discards the passes counted so far, which is an observation about the counter — the cycle itself stays open until the conditions above hold, so an accidental commit destroys the pass credit and closes nothing. A plateau or tells on -**the pass that closes** go into the closing report and never block it, because reporting "will -not converge" on a converged loop is a false report; on an eligible pass that does **not** close -they go into that pass's ordinary report, there being no closing report to carry them. **No other pass outcome makes a cycle eligible +**the pass that closes** go into **that pass's status report to the user** — the carrier the +three-line duty already names, and no second report form is introduced — and never block it, +because reporting "will not converge" on a converged loop is a false report; on an eligible pass +that does **not** close they go into the same place, a pass reporting what it read of the loop +whether or not it closes. **No other pass outcome makes a cycle eligible to close**, because every other pass either leaves a required repair, a hold or a question outstanding **or has not reached the floor** — and closing over any of those is the failure this ordering exists to prevent. The floor is named separately because a below-floor pass whose only @@ -191,9 +214,9 @@ a hold like any other. **A re-raised valid dismissal stays discharged for the re dismissal was the resolution and a reviewer repeating the finding does not undo it, so no second dismissal is owed — and what the recurrence creates is the **clearly-stuck hold alone**, ended by that reading's continue-or-stop answer. **Where one finding is surfaced by both, it carries two hold components and -each is discharged by its own answer, which is what keeps D4 and D5 exact**: the **membership** -component ends on the membership answer in either direction, a decline releasing it as **D5** -requires; the **clearly-stuck** component ends on the reading's continue-or-stop answer. Neither +each is discharged by its own answer**: the **membership** +component ends on the membership answer **in either direction**, a decline releasing it exactly as +an accept does; the **clearly-stuck** component ends on the reading's continue-or-stop answer. Neither answer discharges the other's component, and where the finding carries no scope-stop trigger the continue-or-stop answer is the only one its surface asks for and discharges the only component there is. **Resumption is still the composition rule's**, which waits for every outstanding @@ -206,9 +229,9 @@ loop resumes is decided by the composition rule below and by nothing here**, so answer is never itself a resumption and cannot step past a hold or a health question still awaiting its own answer. Nothing here turns one answer into another, since that would let a finding be moved out of the set and back into it to escape what it owes inside it. **There is no -withdrawal inside the cycle that declined**, because **D7** binds a decline for the remainder of -its cycle and admits no exception; a reconsideration is a later cycle's, where D7 gives the -decline no effect at all and the finding takes the ordinary route. Either answer is an +withdrawal inside the cycle that declined**, a decline binding for the remainder of +its cycle and admitting no exception; a reconsideration is a later cycle's, where that decline has +no effect at all and the finding takes the ordinary route. Either answer is an **explicit, attributable decision on that specific finding** — never silence, never a general remark about scope, never inferred, because a fix set changed by inference is a fix set nobody chose. **Membership is answered against the set as the absorb paragraph fixes it for the pass that @@ -248,7 +271,7 @@ the overlap is admitted rather than argued away: the two read different severity Mechanics · Severity sets out, so an **in-set** Blocker or Major the ceiling demotes below Major can regenerate across passes on a pass that is clean. **A declined finding is not a route into that reading**: the third condition admits regeneration across repair attempts and a re-raised validated dismissal, and a decline is -neither — it is the user's decision that a *true* finding stays outside the set, which **D7** binds +neither — it is the user's decision that a *true* finding stays outside the set, and it binds for the cycle. **The order decides it and no new rule is needed.** The clearly-stuck paragraph's own precedence sentence is stated here rather than there, because precedence is evaluation order and this paragraph is where evaluation order is stated once; its opening words point back to that @@ -377,7 +400,7 @@ continuation next to the conditional one and give the same pass two answers. **`e7`, the threshold.** Gains one clause; the sentence is given entire. ``` **Any two present makes stop-and-surface mandatory, not discretionary** — read **after** the -clean-completion branch of the closure ordering, which outranks it (**D2**) — and you report the +clean-completion branch of the closure ordering, which outranks it — and you report the tells and hand the decision to the user, and the "clearly stuck" reading above is not a precondition for it. ``` @@ -435,10 +458,12 @@ demotes it — the two counts are meant to differ. --- -## F. The four standing sentences this change falsifies — REPLACED +## F. The five standing sentences this change falsifies — REPLACED -Each is a live sentence that the block makes wrong. All four are **known contradictions** and -none is deferred. +Each is a live sentence that the block makes wrong. All five are **known contradictions** and +none is deferred. **Three of them share one mechanism** — an entry point other than the ordering +carrying an unqualified instruction — which is why each is **replaced** rather than given an +exception to point at. **1. The `WIP:` naming warning** (Mechanics · `baseSha`). ``` @@ -484,6 +509,22 @@ it stands, an agent following it refuses the exact transition the ordering requi distinction the repair draws is between a human waving a rule through — which this paragraph still forbids — and answering the question a suspension actually asked. +**5. The `Finishing the cycle` lead-in** (Mechanics · `baseSha`). It wraps across C 827–828 and +W 1011–1012. +``` +**Finishing the cycle:** once the closure ordering reaches a Gate-B cycle's closing act — an +eligible pass with every closure precondition holding, never a clean pass on its own — close it +with `git commit --amend -m ""`; that replaces the WIP commit, and the hook reads +the amend as the real cycle-closing commit. This section states the operation and never whether +the cycle may close. +``` +*Why (pass 20 finding 5):* the live sentence says "after the final clean pass, close it with +`git commit --amend`", which is a complete instruction to whoever enters through Mechanics — and +under the ordering a clean eligible pass is not enough, an unmet precondition leaving the cycle +open. A reader entering there could amend over a standing hold, an undischarged Major or a changed +evidence entry. The repair makes this section the **operation** and the ordering the **permission**, +which is the same split the other two sentences of this mechanism take. + --- ## G. The one-contract paragraph — REPLACED @@ -532,8 +573,11 @@ duties are kept; the blanket no-clean-credit and the one-answer resumption go, b ordering decides both and decided them differently. ``` **Surfacing does not close the cycle, and that is what makes this reachable.** You surface *with -the finding still open* — the resolve rule is not waived, the hold stands until its answers are -given, and **what the answer does is the closure ordering's**: continue resumes, stop parks. +the finding still open* — the resolve rule stands over it exactly as Mechanics · Severity states +it, which means **repaired or validly dismissed**, so a recurrence of one already validly +dismissed is already resolved and owes no second dismissal and no repair; the hold stands until +its answers are given, and **what the answer does is the closure ordering's**: continue resumes, +stop parks. **A pass is credited clean or not on its own findings**, as that ordering defines cleanliness; surfacing a finding that carries a scope-stop trigger is what withholds the credit, and no credit is withheld for surfacing alone. Reading this as "stop instead of fixing" would put the @@ -560,8 +604,7 @@ qualifies; `a21` and `a22` are carried unchanged, reproduced here because the pl contiguous string. The per-condition accounting is the plan's. ``` What a clean final pass and the zero-finding early exit mean for closing is stated once in the -closure ordering; this paragraph states the floor and nothing else. Codex is advisory — validate -before applying; dismissed finding → one-line why. +closure ordering. Codex is advisory — validate before applying; dismissed finding → one-line why. ``` **The Gate-A clean-signal sentence** — the signal stops being the only route to clean. From b9fc78662ca38ded25c46fe2ce0da996141ac706 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Fri, 11 Sep 2026 20:01:31 +0200 Subject: [PATCH 038/181] =?UTF-8?q?docs(context):=20Gate-A=20pass=2021=20?= =?UTF-8?q?=E2=80=94=20cycle-low=20findings,=20zero=20Blockers,=20open?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 6 findings, 0 Blockers, 5 Majors, 1 Minor. Valid pass. Findings are the lowest of the cycle and Blocker+Major ties the pass-16 low at 5. Zero tells on the plain reading, one on the conservative one; either way no mandatory stop. Finding 2 is my own overclaim from the pass-20 round: the closing act says a reader takes the tree of the commit the body sits in, "that tree being the one the pass read". CLAUDE.md:553 says mcp__codex__exec reviews the text passed to it and not the git tree, and the sentence contradicts my own disclaimer two paragraphs above. Same pattern AGENTS.md records, where each correction carries a subtler version of the claim it removed. Finding 1 is a sixth falsified standing sentence, surfaced because the prompt was asked to look for one: both copies still say to re-run each Gate-A pass over the revised artifact because the artifact changes between passes, while the ordering continues on an unrevised artifact where no repair is owed. Findings 4 and 5 are the design spec falling behind the target text, both from earlier repair rounds of mine. Findings not yet dispositioned; routed to the reviewer first. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-21.md | 7 +++ .../gate-a-spec-awsf1ec771-resume.md | 43 ++++++++++++++++++- 2 files changed, 49 insertions(+), 1 deletion(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-21.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-21.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-21.md new file mode 100644 index 0000000..7485498 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-21.md @@ -0,0 +1,7 @@ +MAJOR | high | standing Gate A "each pass over the revised artifact" | Both standing copies still require every Gate-A pass to re-run over "the revised artifact" because "the artifact changes between passes", while target §A expressly continues or resumes on an unrevised artifact when no repair is owed; §H changes only the later cadence sentence | A below-floor Minor/Nit-only pass or a post-health-answer pass with no repair due receives incompatible instructions and can manufacture a change or refuse the required next pass | Replace this sixth falsified standing sentence in both copies so each pass reads the current artifact, revised only where the severity and scope rules require repair +MAJOR | high | §A "that tree being the one the pass read" | The tip-placement branch says the commit tree is "the one the pass read", contradicting both §A's immediately preceding statement that the naming line establishes nothing about what the pass actually read and the standing Gate-A rule that `mcp__codex__exec` reviews the TEXT passed to it, not the git tree | A reader can treat the amended commit tree as reviewed evidence even though the new closing act is deliberately only an author-written record, recreating the gate-proof overclaim AGENTS.md forbids | Say only that a message-only amend preserves the commit tree and the artifact content selected by the closing-time sameness check; do not claim the pass read that tree +MAJOR | high | §A "owns exactly six things" | The ownership boundary omits the pass-input model that the next paragraph defines: logical-pass branch-file composition, per-file validation, `NO FINDINGS` as an empty sequence, and branch-line identity; design §2 likewise names "the file set each predicate reads" separately from evaluation order | The block defines a seventh authority while declaring it outside its own exhaustive boundary, so the plan and §G membership reading can duplicate, drop, or fail to reconcile the rules that decide what findings every predicate consumes | Name the pass-input/read model in the block's ownership boundary, or explicitly and unambiguously make it part of one already named owner, and keep the design's ownership account consistent +MAJOR | high | design §4–§5 passage map | The design says passage (h) is "no longer edited" and gives no human-exception site row, but target §F item 4 replaces h19; its Mechanics `baseSha` row also describes only the WIP-warning change although target §F item 5 separately replaces the Finishing-the-cycle lead-in | The design and target disagree on two live source replacements, so the plan's per-condition accounting can preserve the old human-assent prohibition or the clean-pass-only amend instruction and reintroduce the exact standing contradictions §F repairs | Update the existing site and passage maps to include these two scoped source replacements while keeping the deferred answer-record material and every other part of passage (h) unchanged +MAJOR | high | design §9 "an edit is coupled when the block cites it or depends on it" | The design still defines contract membership through source edits the block cites or depends on, exactly the repository-only test target §G says was rejected; §G instead tests the present meaning of each live rule and separately includes production, identity, and transport of records | The decisions document and shipped text yield different contract memberships, so a downstream reader and the implementation plan can accept different partial adoptions even while each follows its stated authority | Replace the design's edit/dependence formulation with the same live-rule semantic membership test stated in §G, without adding an enumeration or checker +MINOR | high | §B "The whole paragraph is written out here, as it will read" | The fenced final passage ends before the C-only field-mint parenthetical that still closes the standing CLAUDE.md paragraph; §B mentions that parenthetical only afterward as editorial prose, despite claiming every other sentence is carried and the whole paragraph is written out | The literal proposed block is not the complete future CLAUDE.md paragraph, so a plan treating the fence as the promised whole replacement can delete the deliberate divergence or a reviewer can approve bytes that are not actually the final passage | Describe the fence as the common replacement span and explicitly leave the adjacent C-only parenthetical untouched, rather than claiming the whole paragraph is reproduced +END OF FINDINGS (6 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index 3302d8a..e542b36 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -47,7 +47,48 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 18 | 96c3611 | 9→**15** | 1→**0** | 7→**12** | yes | **first pass on the target text. ZERO BLOCKERS, first since pass 7.** Finding 1 names the restructure as half-done: §H is still paraphrase, not text. Findings 6,7,8 name three standing sentences by line number — the 15/16/17 mechanism, now findable; session 01a09068-de4a-7071-97a7-82cc4774544f | | 19 | 6d52fcc | 15→**13** | 0→**1** | 12→**10** | yes | precheck caught a silent loss before the pass: `a20` ("don't manufacture findings to pad") had fallen out of the target text; restored to §A. Two known contradictions repaired in the same commit. **Only 6 of 13 findings are against the target text**; 5 are against the design spec and 2 against standing §5; session 01a0912a-7329-7450-bdb3-da043e88e7bb | | 20 | 5174d7a | 13→**9** | 1→**1** | 10→**5** | yes | **B+M 11→6, Majors halved.** Three of nine (3, 6, 9) regenerate from pass 19's own repairs; the Blocker is pass 19's Gate-A closing act, which names a revision in two record forms that have no field for one; session 01a0915d-6a5b-7633-a60a-881f0022816d | -| 21 | — | — | — | — | not run | next action, after pass 20's findings are dispositioned | +| 21 | 8dc22fb | 9→**6** | 1→**0** | 5→**5** | yes | **findings the lowest of the cycle; B+M 5 ties the pass-16 low; zero Blockers.** Finding 2 is my own overclaim from the pass-20 round; finding 1 is the sixth falsified standing sentence, found because the prompt asked for one; session 01a09195-8b00-7693-b42c-e85888188d3e | +| 22 | — | — | — | — | not run | next action, after pass 21's findings are dispositioned | + +## Pass-21 three-line report + +**Floor line:** derived floor **3**; risk **high**, security **none**; read fresh from +`docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md`. One cited story, level 2. + +- **Trend:** findings …, 15, 13, 9, **6**. Blockers …, 0, 1, 1, **0**. Majors …, 12, 10, 5, **5**. + Blocker+Major …, 8, 5, 8, 12, 11, 6, **5** — level with the pass-16 low, and **6 findings is the + lowest of the cycle**, past pass 16's 8. +- **Cluster (pass 21):** product 3 (1, 2, 3); prose about the design or this artifact 3 (4, 5, 6); + the instrument 0. **A tie rather than a cluster**, named because it is the first pass where + product does not dominate outright. +- **require↔withdraw:** none. Near-miss: finding 2 objects to a sentence the pass-20 round itself + added — a later pass questioning an earlier pass's addition, the mirror of the pair's shape. + +**Tells: zero clearly, one on the conservative reading.** Findings fell 9 → 6, Blockers fell 1 → 0, +no pair, no instrument cluster. The only candidate is the prose-about tie at 3 of 6; counted or +not, one is not two, so **no mandatory stop**. The clearly-stuck exit is not reachable either — its +first condition needs a plateau and the curve is at a cycle low. + +**Finding 2 is mine, and it is the defect `AGENTS.md` calls this repo's most persistent.** The +pass-20 round removed an overclaim from the closing act and introduced a smaller one in the same +sentence: "a reader wanting the reviewed content reads the tree of the commit the body sits in, +**that tree being the one the pass read**". Verified false at `CLAUDE.md:553` — `mcp__codex__exec` +"reviews the TEXT you pass, not the git tree" — and it contradicts my own sentence two paragraphs +above saying the record establishes nothing about what the pass read. Exactly the pattern +`AGENTS.md` records: *each correction introduced a subtler version of the same claim*. + +**Finding 1 is the sixth falsified standing sentence, and it was found because the prompt asked.** +`CLAUDE.md:554–555` and `workflow-init.md:746–747` still say to re-run the Gate-A prompt "each pass +over the revised artifact … because the artifact changes between passes", while §A continues on an +**unrevised** artifact where no repair is owed. §H edits the cadence sentence and not this one. The +mechanism is alive at passes 15, 16, 17, 19, 20 and 21; adding "look for a sixth" to the prompt is +what surfaced it. + +**Findings 4 and 5 are the design falling behind the target text**, both from my own repair rounds: +the passage map still says (h) is "no longer edited" while §F item 4 replaces a sentence in it, the +`baseSha` row names only the WIP warning while §F item 5 replaces a second sentence there, and §9 +still defines contract membership by "an edit is coupled when the block cites it or depends on it" +— the repository-only test §G was rewritten to replace. ## Pass-20 dispositions — and two of my claims corrected, both verified wrong From 2ff9f247db08ce028717852808578e725893ddf1 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Fri, 11 Sep 2026 20:29:11 +0200 Subject: [PATCH 039/181] docs(specs): apply Gate-A pass 21 across whole passages MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The test this round is whether the objected claims are gone from their passages and whether the design pointers resolve to this version, not whether six spots were touched. Finding 1 accepted with its reason repaired. The Gate-A broad-prompt instruction becomes §F's sixth standing sentence: it keeps the breadth demand and the ban on narrowing per-dimension, allows an unrevised artifact where no repair is owed, and replaces "the artifact changes between passes" with reasons that still hold when it did not. No edit is to be manufactured to justify a pass. Finding 2 accepted and the whole closing description rewritten. It said in several places that the pass had read a commit or its tree; deleting only the flagged clause would have left the claim standing next to it. The paragraph now describes the source revision the author attributes and the content a message-only amend leaves alone, and says outright that Gate A hands the reviewer text rather than a git range, so no commit name could record what was reviewed. The sameness list is reworded the same way. The empty-record commit is now stated as a precedent whose reasoning transfers, not as an allowance that already covers Gate-A closure. Finding 3 treated as an unclear boundary rather than a proven seventh item. The count and the ordinal back-reference are gone, the read model is named inside the evaluation the block owns, and externally defined rules stay at their sources. No replacement completeness list. Findings 4 and 5 corrected together in the design. The site and passage maps name the human-exception sentence, the Finishing-the-cycle lead-in and the broad-prompt instruction; §9 points at §G's membership test instead of retelling it, and records that the edits-cited-or-depended-on formulation was decidable only against this repository. Finding 6: §B is the passage's common replacement span, which drops the whole-paragraph claim while keeping the C-only parenthetical marked. Gate A stays open; pass 21 was not clean. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- ...26-09-10-loop-rule-consolidation-design.md | 30 ++-- ...-10-loop-rule-consolidation-target-text.md | 129 +++++++++++------- 2 files changed, 92 insertions(+), 67 deletions(-) diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index 477685b..e5fd15f 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -118,12 +118,13 @@ to point at and so a reader can see the shape of the change without reading the | (c) recognizing clearly stuck | §C | | (e) the five tells | §D | | Mechanics · Severity | the resolve duty is scoped and gains its discharge rule; the handed-over question is replaced by its answer | -| Mechanics · `baseSha` | the `WIP:` warning stops claiming a closure the rules do not grant | +| Mechanics · `baseSha` | two sentences: the `WIP:` warning stops claiming a closure the rules do not grant, and the `Finishing the cycle` lead-in performs the amend only where the ordering permits closing | +| Mechanics, recording a human exception | one sentence: a prescribed continuation answer is distinguished from blanket assent. The answer-record material stays moved to the successor | | Gate B, the coverage instruction | `NO FINDINGS` only when the branch found none | | Mechanics, the curve's Majors rationale | rewritten on the pre-ceiling reading | | Mechanics, the one-contract paragraph | membership widened, with a semantic test a downstream reader can apply | | (i) when these rules bind | §H | -| the Gate-A section and the gate-prompt template | the two senses of *clean* are separated; the cadence makes revision conditional | +| the Gate-A section and the gate-prompt template | the two senses of *clean* are separated; the cadence makes revision conditional; the broad-prompt instruction stops assuming the artifact is revised between passes, keeping its breadth demand | | the profiles section, the lens paragraph | its unchanged-list is scoped to the lens sets | **Two sentences are deliberately not edited**, named so nobody looks for them: the "Copy every @@ -153,7 +154,7 @@ here. This table says what happens to each inventoried passage, so the map stays | (e) the five tells | edited — the threshold is read after clean completion, and a pointer says what its answer does | §D | | (f) the two rules above do not compete | **unchanged.** "The two rules above" still names the absorb rule and the stuck reading; the block sits before both and adds no third rule between them | — | | (g) Mechanics · Severity, the handed-over question | edited — the unsettled statement and its interim report-and-stop duty are replaced by the answer, in both copies, removing the one deliberate story-path divergence | §E | -| (h) recording a human exception | **no longer edited.** The answer-record block that was to follow it moved to the successor with **D9** | — | +| (h) recording a human exception | **edited in one sentence only** — a prescribed continuation answer is distinguished from blanket assent, which the ordering makes the restart of a parked cycle. The answer-record block that was to follow it stays moved to the successor with **D9**, and nothing else in the passage changes | §F | | (i) when these rules bind | edited — the strict-reading list is added to, not rewritten | §H | | (j) the squash carry | **no longer edited.** It was to name the answer record, which moved; this change ships no record for it to carry | — | @@ -300,9 +301,10 @@ repo's most persistent defect. The transport that could carry it left with the r the block replaces closure sentences rather than adding beside them, while the block was in fact restating triggers, duties, preconditions and the severity answer that their own paragraphs still defined, which is two authorities per copy. The claim now rests on what the - block does: **it owns the six things §2 names** — not restated here, one statement of that - boundary being the point — and **cites** every other rule where that rule is defined, so each - has one definition in the shipped text; **§4's site table is the check**. Then item 3 (the stop + block does: **it owns the evaluation of a pass and what follows from it, as the target text's §A + opening states that boundary** — not restated here, one statement of it being the point — and + **cites** every other rule where that rule is defined, so each has one definition in the shipped + text; **§4's site table is the check**. Then item 3 (the stop answer produces a named state, **parked**, with its own restart transition). - **Don't: "Never replace a decision procedure without accounting for its old conditions."** Satisfied by the plan's per-condition disposition list against the committed inventory (§5), @@ -359,14 +361,14 @@ verification fragment with its counts (§7). Each is work this change still owes work a spec can do correctly, because all four are checked against files the plan edits. **Partial adoption — answered by the one-contract paragraph (target text §G), and what that answer is worth.** -The set is mutually dependent, and **the rule that says which edits belong is stated rather than -enumerated**: an edit is coupled when the block **cites it or depends on it** — the boundary its -clean predicate reads, the sentences that give *clean* its two senses, the triggers it reads, the -source rules whose old text the ordering falsifies, and the severity and dismissal rules it cites. -The target text's §G carries a semantic membership test into that paragraph, which until now reached -only the nonce, the slots, the provenance line, the curve, the carry rule and the unknown-start -semantics. **An earlier revision of this section named the members as a list of item numbers and -the list was wrong** — it omitted several edits the block plainly depends on — which is why the +The set is mutually dependent, and **the membership rule is stated once, in the target text's §G, +and is deliberately not restated here** — a second telling of it is a second definition, which is +pass 21 finding 5. §G reads what a live rule **states** rather than what changing it would do, and +it reaches a wider set than the paragraph did before: previously only the nonce, the slots, the +provenance line, the curve, the carry rule and the unknown-start semantics. **An earlier revision +of this section named the members as a list of item numbers and the list was wrong**, and a later +one defined membership here by the edits the block cites or depends on — a test decidable only +against this repository's own spec. Both are why the rule is stated and **the plan derives the membership against the real files**, where dependence is decidable and the numbering does not exist. diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index 6dc411f..9fce2a7 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -44,17 +44,20 @@ Sits immediately before "**What a loop absorbs, and what stops it**" in both cop ``` **How a cycle ends — one ordering, stated here and referenced everywhere else.** A pass is read in a fixed order, because every rule bearing on one decision — may this cycle close — otherwise -qualifies the others and the ranking survives only in a reader's head. **This paragraph owns -exactly six things**: the evaluation order, closure and its eligibility, the hold a surfaced -finding places and what discharges it, the composition of several suspensions, the pairs that -cannot co-occur, and — as a stated exception, because precedence is evaluation order — the -clearly-stuck precedence sentence quoted into it below. **Everything else it names it cites**: +qualifies the others and the ranking survives only in a reader's head. **This paragraph owns how a +pass is evaluated and what follows from that**: the evaluation order, **what each predicate is +read from**, closure and its eligibility, the hold a surfaced finding places and what discharges +it, the composition of several suspensions, the pairs that cannot co-occur, and — as a stated +exception, because precedence is evaluation order — the clearly-stuck precedence sentence quoted +into it below. **No number is put on that list.** Whether the read model counts as part of the +evaluation order or as a thing beside it is a question of wording, not of authority, and a count +turns that wording into an argument about whether the paragraph has overrun its own boundary. **Everything else it names it cites**: the scope triggers, the assigned fix set, every severity rule, and every closure precondition **that has a source of its own** keep their one definition in the paragraph that owns them, and a reader who finds one of *those* defined here has found a defect. **The closing-time sameness tests below are defined here and are not an exception to that**: they borrow no rule and have no -other source, being part of the closure decision itself — the second of the six things this -paragraph owns — rather than a precondition stated elsewhere and read from here. +other source, being part of the closure decision this paragraph owns rather than a precondition +stated elsewhere and read from here. **What a pass is read from.** Every finding-derived predicate reads the validated findings file **or files** of the logical pass as **the concatenation of their finding lines after each file @@ -102,39 +105,40 @@ they disagree, which the branches below do. **The closing act, by cycle kind.** A **Gate-B** cycle closes with the closing amend Mechanics · Finishing the cycle describes. A **Gate-A** cycle has no WIP snapshot to replace, so its closing act is **the commit body that carries this cycle's provenance line, its per-pass curve -and, beside them, one line naming the revision the clean pass reviewed**. The first two are records -Mechanics already obliges every cycle to write and **neither is altered here** — their grammars are -fixed and neither has a field for a revision. The third is **ordinary prose in the same body**, -which those grammars neither supply nor forbid. Writing that body is the author's recorded -acceptance of that revision. - -**What the naming line says:** the artifact's repository-relative path, and the **full -40-character object name** of the commit the clean pass reviewed, an abbreviation being ambiguous -across repositories and across time. **What it is worth, said rather than implied:** it records -what the author accepted and **establishes nothing about what the pass actually read**, which no -reader can check from a commit body. It is a record, not evidence, and no durable provability is -claimed for it. - -**Both placements are written out, because they identify the revision differently.** Where the -reviewed commit is **still the branch tip**, the body goes into it as **an amend of the message -alone**: the amend gives that commit a new object name and leaves its tree untouched, so what the -pass reviewed changes as a commit and not as content. The naming line records the object name the -commit carried **when the pass read it**, which the amend then supersedes — **a reader wanting the -reviewed content reads the tree of the commit the body sits in**, that tree being the one the pass -read. Where the reviewed commit is **no longer the tip**, the body goes into the **next commit made -on the branch** — not the commit that happens to follow the reviewed one, which is already in the -past — and that may be **a commit carrying only the record**, which changes no content and so -raises no review obligation, exactly as Mechanics already allows for a human-exception record with -nowhere else to go. Placement says nothing on this path, so the naming line is the only thing -tying the closure to the revision, and it is written the same way. **No new revision of the artifact is made to close a Gate-A cycle**, because a -new revision is one no pass has reviewed. Nothing is closed before that act, and what changes in -between still gates it: the **profile**, the **cited set**, the **assigned fix set**, the -**revision of the artifact the clean pass reviewed**, and — in a Gate-B cycle — the **evidence -entry**. Any of them differing at the closing act from what that pass read makes that pass -non-final and owes another, which is the same answer a mid-pass change already gets. **Sameness is -read on the artifact and the duties, never on the branch tip**: writing the closing body is itself -a commit, so a commit that only records the closure changes neither, while an edit to the reviewed -text or a broadening of the set changes one and costs the pass. **A commit the hook reads as cycle-closing is a separate matter**: a non-`WIP` commit +and, beside them, one line naming the source revision the author accepts for this cycle**. The +first two are records Mechanics already obliges every cycle to write and **neither is altered +here** — their grammars are fixed and neither has a field for a revision. The third is **ordinary +prose in the same body**, which those grammars neither supply nor forbid. + +**The naming line is an attribution by the author, and is written as one.** It carries the +artifact's repository-relative path and the **full 40-character object name** of the commit being +accepted, an abbreviation being ambiguous across repositories and across time. **It says which +revision the author accepts; it says nothing about what the reviewer was given.** Gate A hands the +reviewer text rather than a git range, so no commit name could record what was reviewed and none +is offered as doing so. A reader of this line learns what was accepted and, about the review +itself, nothing. + +**Both placements are written out, because they carry the attribution differently.** Where the +named commit is **still the branch tip**, the body goes into it as **an amend of the message +alone**: the amend gives that commit a new object name and **leaves its tree untouched**, so the +content the line attributes is unchanged by the act of recording it. The line names the commit as +it stood before the amend and that name stops resolving, which costs nothing on this path — the +content sits in the tree of the commit the body ends up in. Where the named commit is **no longer +the tip**, the body goes into the **next commit made on the branch**, not the commit that happens +to follow the named one, which is already in the past. That may be **a commit carrying only the +record**: it changes no content, so it raises no review obligation. Mechanics grants such a commit +today for a human-exception record with nowhere else to go — **that is the precedent for this and +not a permission that already covers it**, the reasoning transferring where the allowance does +not. Placement attributes nothing on this path, so the naming line carries the attribution alone, +and it is written the same way. **No new revision of the artifact is made to close a Gate-A +cycle**, a new revision being one no pass has run against. Nothing is closed before that act, and +what changes in between still gates it: the **profile**, the **cited set**, the **assigned fix +set**, the **artifact as it stood when the pass was run against it**, and — in a Gate-B cycle — +the **evidence entry**. Any of them differing at the closing act makes that pass non-final and +owes another, which is the same answer a mid-pass change already gets. **Sameness is read on the +artifact and the duties, never on the branch tip**: writing the closing body is itself a commit, +so a commit that only records the closure changes neither, while an edit to the artifact or a +broadening of the set changes one and costs the pass. **A commit the hook reads as cycle-closing is a separate matter**: a non-`WIP` commit mid-cycle makes the hook read the cycle as closed and discards the passes counted so far, which is an observation about the counter — the cycle itself stays open until the conditions above hold, so an accidental commit destroys the pass credit and closes nothing. A plateau or tells on @@ -284,12 +288,17 @@ false report. Below the floor the pass **suspends**, clean completion having not --- -## B. Passage (b) — what a loop absorbs — REPLACED, whole passage +## B. Passage (b) — what a loop absorbs — REPLACED, common span -**The whole paragraph is written out here, as it will read.** Earlier revisions split its -replacements across two sections and left one sentence half in each, which is how §B and §H came -to instruct the plan differently about the same clause. Nothing about this passage is stated -anywhere else in this file. +**This is the passage's common replacement span, written out as it will read** — the run both +copies take byte-identical, from the paragraph's opening through its closing rationale. **It stops +short of the paragraph's last sentence**, the field-mint parenthetical, which closes this +paragraph in `CLAUDE.md`, is absent from the template, and is left exactly as each copy has it. +That divergence is pre-existing; this change neither creates nor removes it, and the plan's +divergence list carries it. Earlier revisions split this passage's replacements across two +sections and left one sentence half in each, which is how §B and §H came to instruct the plan +differently about the same clause. Nothing about this passage is stated anywhere else in this +file. **Changed:** `b3`, `b7`, `b8`, `b11`, `b12`, `b13`, `b16`, `b17`–`b18`. **Carried:** every other sentence. The rationales follow the text. @@ -331,10 +340,6 @@ committing you to a design you never chose, which is a different failure from an review. ``` -**The field-mint parenthetical that closes this paragraph in `CLAUDE.md` is carried unchanged and -is absent from the template**, which is a pre-existing parity divergence this change neither -creates nor removes. The plan's divergence list carries it. - *Why `b3`:* the four severity actions are restated beside a pointer to the section that defines them, which is two authorities per copy. The template took C's wording on the inventory's reasoning that it has no Mechanics section, which is false. @@ -458,10 +463,10 @@ demotes it — the two counts are meant to differ. --- -## F. The five standing sentences this change falsifies — REPLACED +## F. The six standing sentences this change falsifies — REPLACED -Each is a live sentence that the block makes wrong. All five are **known contradictions** and -none is deferred. **Three of them share one mechanism** — an entry point other than the ordering +Each is a live sentence that the block makes wrong. All six are **known contradictions** and +none is deferred. **Four of them share one mechanism** — an entry point other than the ordering carrying an unqualified instruction — which is why each is **replaced** rather than given an exception to point at. @@ -523,7 +528,25 @@ the cycle may close. under the ordering a clean eligible pass is not enough, an unmet precondition leaving the cycle open. A reader entering there could amend over a standing hold, an undischarged Major or a changed evidence entry. The repair makes this section the **operation** and the ordering the **permission**, -which is the same split the other two sentences of this mechanism take. +which is the split every sentence of this mechanism takes. + +**6. The Gate-A broad-prompt instruction** (Gate A section). It wraps across C 553–555 and +W 745–747. +``` +Use ONE broad prompt, re-run it unchanged each pass over the **current** artifact — revised where +a repair is owed and unrevised where none is, an edit made to justify a pass being no reason to +run one. Don't narrow per-dimension: new findings surface because the artifact changed, because +an answer given since the last pass changed what the rules require of it, or because a broad +prompt reaches what the last reading did not. +``` +*Why (pass 21 finding 1):* the live wording says to re-run "over the revised artifact … because +the artifact changes between passes", which the ordering's third branch contradicts — that branch +continues on an **unrevised** artifact wherever no repair is owed. A pass whose only findings were +Minors, or one run after a health answer with nothing left to repair, receives two instructions +and can satisfy the live one only by manufacturing a change. **The replacement keeps the breadth +demand**, which is what that sentence exists for and the reason it must not be narrowed +per-dimension; it drops only the claim that a revision always precedes a pass, and replaces the +reason with one that still holds when none did. --- From 17fcf7a11b4e83cbba0afa1060caa3e349768547 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Fri, 11 Sep 2026 20:47:44 +0200 Subject: [PATCH 040/181] =?UTF-8?q?docs(context):=20Gate-A=20pass=2022=20?= =?UTF-8?q?=E2=80=94=20mandatory=20two-tell=20stop,=20findings=20held=20op?= =?UTF-8?q?en?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 13 findings, 1 Blocker, 5 Majors, 6 Minors, 1 Nit. Valid pass. Two tells of five, which is the threshold: the finding count rose 6 to 13 and the Blocker count failed to fall, 0 to 1. Stop-and-surface is mandatory. All 13 are held open and unrepaired pending the answer. Blocker+Major moved 5 to 6. The doubling is six Minors and a Nit, which collect and never iterate, so the stop is surfaced on the flat B+M stretch of 6, 5, 6 rather than on the headline count. One mechanism is on its third consecutive pass: the Gate-A closing act. Pass 20's Blocker, pass 21's finding 2 and pass 22's Blocker plus findings 2, 8 and 12 all sit in that paragraph, each out of the previous round's repair of it. Three of the thirteen are defects in sentences the last round wrote. Verified before reporting: the closed cycle-record set and nonce duty at CLAUDE.md:384-390, the legacy bare-slot reservation at CLAUDE.md:390, the Gate-A human-exception destination at CLAUDE.md:988, and suspension classifications restated outside §A in the target text. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-22.md | 14 ++++++ .../gate-a-spec-awsf1ec771-resume.md | 49 ++++++++++++++++++- 2 files changed, 62 insertions(+), 1 deletion(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-22.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-22.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-22.md new file mode 100644 index 0000000..60ade13 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-22.md @@ -0,0 +1,14 @@ +BLOCKER | high | §A "source revision the author accepts" | Neither Gate-A placement nor the closing-time sameness test requires the named commit's blob at the artifact path to equal the current artifact that the pass read; the standing Gate-A cadence permits revisions and re-runs without an intervening commit, so a message-only tip amend or record-only later commit can name and preserve a stale artifact revision | The cycle can close while its required attribution names content other than the accepted current artifact; if equality is intended implicitly, an eligible pass over uncommitted final repairs has neither a valid closing act nor a suspension path | Require the named commit's artifact content to equal the closing-time artifact and state how unchanged final bytes reach such a commit before closure, while continuing to say that this author attribution is not evidence of what the reviewer received +MAJOR | high | §A "ordinary prose in the same body" | The new source-revision line is deliberately a cycle record but carries only an artifact path and object name, while standing Mechanics says the nonce "appears in every record the cycle writes" and names a closed set of provenance, curve, findings-slot and working records | Several Gate-A cycles recorded in one body cannot unambiguously associate each attribution with its cycle, and the proposed full adoption already violates the record-identity contract that §G tells partial adopters to reconcile | Reconcile the new attribution with the existing cycle-field rule, using the existing cycle identity rather than inventing a separate durability mechanism, or explicitly narrow the standing every-record claim everywhere it is relied on +MAJOR | high | standing Mechanics "a Gate-A cycle in the spec or plan commit" | The unchanged human-exception placement requires a Gate-A exception in the spec or plan commit, but §A's new non-tip placement requires the cycle-closing body in the next commit on the branch and forbids changing the already-past artifact revision | A non-tip Gate-A cycle that also has a human-exception record receives incompatible commit destinations, so the target's second closing placement does not agree with the standing text it will sit beside | Replace the standing Gate-A exception destination with a reference to the same tip versus non-tip closing-body placement §A defines, without changing the record's form or force +MAJOR | high | §A "This paragraph owns how a pass is evaluated" | The claimed single authority is restated outside §A: §B says the membership answer ends the hold and calls the scope stop a suspension, §D says the two-tell stop is a suspension that closes nothing and takes continue or stop, and §H says the hold lasts until its answers and that continue resumes while stop parks | The shipped prompt again has parallel definitions of hold discharge, closure classification and suspension transitions, contradicting the story's once-stated ordering and allowing a later source edit to change one copy while the ordering remains unchanged | Keep trigger definitions at their source paragraphs but replace the duplicated classifications and transitions with pure references to the ordering +MAJOR | high | §F "re-run it unchanged each pass" | A Gate-A prompt cannot remain unchanged while also containing the current artifact text that this same sentence requires it to review; "unchanged" does not distinguish the stable broad review instructions from the per-pass artifact payload | A literal reader can reuse the previous pass's artifact text, producing a valid findings file for stale content that the closing-time check may then treat as the current pass | Say to keep the broad review question and review dimensions unchanged while replacing the artifact payload with the current artifact on every pass +MAJOR | high | §G "a slot rule without a nonce has nothing to key on" | This rationale contradicts the standing findings protocol, which explicitly reserves bare gate-and-pass slot names for legacy cycles with no nonce | An agent applying the one-contract paragraph can reject a valid pre-rule cycle or a correctly retained legacy slot rule as an incoherent partial adoption | Say that a slot rule without a nonce cannot distinguish sibling cycles, while preserving the standing bare-slot key for the legacy single-cycle case +MINOR | high | §G "a record that does not survive the merge is a record the cycle did not produce" | Squash omission makes a record unreachable from the squash commit and main history; it does not undo the cycle's production of the record in the branch commit | The rationale overstates what the carry rule establishes and conflates production with later transport, the exact kind of mechanism overclaim AGENTS.md forbids | Say that an uncarried record is unavailable from the squash commit and main history, which is the narrower claim the standing carry paragraph actually supports +MINOR | high | §A "that name stops resolving" | Amending a commit moves the branch reference to a new commit object but does not make the old object name immediately stop resolving; the old commit remains addressable while its object is retained and may only become unavailable after it is unreachable and pruned | Readers are given a mechanically false account of Git and cannot tell whether the old object name is expected to work when inspecting the attribution | Say that the old name stops being referenced by the branch and may later become unresolvable after pruning +MINOR | high | design §4 "This change ships no record" | The design uses absence of a new record to justify leaving two record-related standing sentences untouched, while target §A and the review brief explicitly define the Gate-A source-revision closing act as a record | The design and target disagree about the change's product surface, so the plan can omit this record from contract reconciliation and old-condition accounting | Narrow the sentence to the deferred Accepted and Declined record and acknowledge the Gate-A attribution record without adding durability or transport work +MINOR | high | §A and §C "That third condition is what makes a plateau rather than a finish" | §C leaves the opening clause standing as its own sentence while §A reproduces the entire original precedence sentence including that same clause, even though §C says only the second half moved and the design says the sentence moved unchanged | One rationale is installed twice under two section owners, contradicting the target's passage-assignment requirement and weakening the claimed single source for precedence | Preserve the verbatim precedence sentence in §A and give the third condition a nonduplicating local reason or pointer in §C +MINOR | high | §G "A curve without a cycle field cannot be attributed" | The categorical claim conflicts with standing Mechanics, which says missing cycle attribution is a limitation rather than a disqualification because a human reading one cycle's records can know their cycle, while the nonce only usually removes that contextual work | The one-contract rationale presents probabilistic and context-dependent attribution as impossible, overstating what the cycle field supplies | Say that a curve without a cycle field cannot be reliably distinguished when records from several cycles are read together +MINOR | high | §A "no commit name could record what was reviewed" | Passing text instead of a git range means the gate result supplies no reviewed commit identity and a commit name cannot prove the payload, but it does not make an author-written source attribution impossible; the paragraph itself demonstrates that commit names can record an author's claim | The sentence conflates recording with establishing evidence immediately after insisting on that distinction, leaving the new closing act's semantics internally unstable | Say that the gate result supplies no commit identity and that this naming line records only the accepted source revision, never evidence of the review payload +NIT | high | §B "Nothing about this passage is stated anywhere else in this file" | §A repeatedly states facts about the absorb passage's assigned fix set, two scope triggers and membership hold, and §E also describes the set that the absorb paragraph computes | The section's own exclusivity claim is mechanically false even if it intended only to say that the replacement bytes appear once | Scope the claim to the common replacement span rather than to anything about the passage +END OF FINDINGS (13 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index e542b36..8534215 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -48,7 +48,54 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 19 | 6d52fcc | 15→**13** | 0→**1** | 12→**10** | yes | precheck caught a silent loss before the pass: `a20` ("don't manufacture findings to pad") had fallen out of the target text; restored to §A. Two known contradictions repaired in the same commit. **Only 6 of 13 findings are against the target text**; 5 are against the design spec and 2 against standing §5; session 01a0912a-7329-7450-bdb3-da043e88e7bb | | 20 | 5174d7a | 13→**9** | 1→**1** | 10→**5** | yes | **B+M 11→6, Majors halved.** Three of nine (3, 6, 9) regenerate from pass 19's own repairs; the Blocker is pass 19's Gate-A closing act, which names a revision in two record forms that have no field for one; session 01a0915d-6a5b-7633-a60a-881f0022816d | | 21 | 8dc22fb | 9→**6** | 1→**0** | 5→**5** | yes | **findings the lowest of the cycle; B+M 5 ties the pass-16 low; zero Blockers.** Finding 2 is my own overclaim from the pass-20 round; finding 1 is the sixth falsified standing sentence, found because the prompt asked for one; session 01a09195-8b00-7693-b42c-e85888188d3e | -| 22 | — | — | — | — | not run | next action, after pass 21's findings are dispositioned | +| 22 | 2ff9f24 | 6→**13** | 0→**1** | 5→**5** | yes | **TWO-TELL STOP — mandatory, surfaced to Daniel.** Findings more than doubled and the Blocker returned. Three findings (5, 8, 12) are defects in sentences the pass-21 round itself wrote; the Gate-A closing act has now taken a Blocker or Major at passes 20, 21 and 22, each out of the previous repair. All 13 held open; session 01a091bb-b1cd-7562-98d1-a50647a30370 | +| 23 | — | — | — | — | not run | blocked on the two-tell answer | + +## Pass-22 three-line report — MANDATORY TWO-TELL STOP + +**Floor line:** derived floor **3**; risk **high**, security **none**; read fresh from +`docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md`. One cited story, level 2. + +- **Trend:** findings …, 13, 9, 6, **13**. Blockers …, 1, 1, 0, **1**. Majors …, 10, 5, 5, **5**. + Blocker+Major …, 12, 11, 6, 5, **6**. The three-pass fall 13 → 9 → 6 reversed in one pass. +- **Cluster (pass 22):** product 9 of 13 (1, 2, 3, 4, 5, 6, 7, 8, 11, 12 — ten on the strict + count); prose about the design or this artifact 3 (9, 10, 13); the instrument 0. +- **require↔withdraw:** none. Near-miss named: finding 12 objects to the sentence the pass-21 round + wrote to answer pass 21's finding 2 — a second pass on one sentence, but it asks for a third + narrower wording rather than the removed one back. + +**Tells: two of five — the threshold. Stop-and-surface is mandatory, not discretionary.** The +finding count rose 6 → 13 and the Blocker count failed to fall, 0 → 1. + +**What the count hides: B+M barely moved.** 5 → 6. The doubling is six new Minors and a Nit, which +collect and never iterate. Read on Blocker+Major alone this is the flattest stretch of the cycle — +6, 5, 6 — and that is the reading the stop is being surfaced on, not the headline number. + +**One mechanism is on its third consecutive pass and it is the Gate-A closing act.** Pass 20's +Blocker, pass 21's finding 2 and pass 22's Blocker and findings 2, 8 and 12 all sit in that one +paragraph, and **each came out of the previous round's repair of it**: +- pass 20 said the closing act named a revision in forms with no field for one; +- pass 21 said the repair claimed the pass had read the named commit's tree; +- pass 22 says the rewritten paragraph still does not require the named commit's content to equal + the artifact at closing time (Blocker), carries no cycle nonce though standing Mechanics names a + **closed set** of cycle records and demands the nonce in every one (`CLAUDE.md:384–390`, + verified), is mechanically wrong about git in "that name stops resolving" (an amended commit's + old object stays addressable until pruned), and overshoots in the other direction with "no commit + name could record what was reviewed". + +**Three of the thirteen are defects in sentences this last round wrote** — 5, 8 and 12. That is the +regeneration condition, nameable and not inferred. + +**Findings verified mechanically before this report:** the closed record set and nonce duty at +`CLAUDE.md:384–390`; the legacy bare-slot reservation at `CLAUDE.md:390`; the human-exception +Gate-A destination at `CLAUDE.md:988`; suspension classifications restated outside §A at target +text lines 337 and 415. + +**The clearly-stuck reading is also satisfiable and is not being taken.** Its three conditions: +a plateau on Blocker+Major (6, 5, 6, and never zero in twenty-two passes); coverage affirmable +after twenty-two readings; and Blocker/Major regenerating from the previous round's repairs, which +findings 1 and 5 name. It is not needed — the two-tell threshold stands alone — and taking it would +close nothing, since surfacing credits no pass clean. ## Pass-21 three-line report From d6052feecbd620eadda542666bcc426dba1ea453 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sat, 12 Sep 2026 11:17:54 +0200 Subject: [PATCH 041/181] docs(specs): apply Gate-A pass 22 by rolling back the attribution line MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The two-tell stop is answered with a deliberate reversal rather than a fourth round of hardening the same paragraph. Rolled back: the source-revision line, its artifact-path-plus-object-name form, and every sentence about old commit names and their later readability. It was meant to sharpen the closing decision and instead added an attribution duty with consequences for record identity, commit placement and content equality, which the story's acceptance criteria never asked for. What replaces it answers the Blocker rather than renaming it. Gate A closes by writing the closing commit, or the closing message of one that exists, carrying only the records the cycle already owes. The closing commit must carry at the artifact path the content the final pass was run against, unchanged. That is a condition on content and never on a commit name, and committing already-reviewed but uncommitted content is not a new revision. Three cases are decided and the safe git sequence is left to the plan: matching content at the tip closes by amending that commit's message; matching content committed but behind the tip closes with a new commit that changes nothing at the artifact path; final content still uncommitted is committed unchanged and closed in that commit. The third had no answer before, so an eligible pass over uncommitted repairs could neither close nor suspend. Walking case two caught a reading that invited restoring an older copy over newer content, now forbidden. Findings 2, 8, 9 and 12 dissolve with the rollback. 3 becomes §F entry 7, the Gate-A human-exception destination following the closing act. 4 replaces the transitions restated in §D and §H with references. 5 scopes "unchanged" to the review question and dimensions, the artifact payload being replaced every pass. 6, 7 and 11 are one edit in §G, preserving the legacy bare-slot reservation and narrowing two rationales to what they establish. 10 and 13 collected. Gate A stays open; pass 22 was not clean. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-resume.md | 55 +++++++++ ...26-09-10-loop-rule-consolidation-design.md | 4 +- ...-10-loop-rule-consolidation-target-text.md | 116 ++++++++++-------- 3 files changed, 125 insertions(+), 50 deletions(-) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index 8534215..5486a47 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -51,6 +51,61 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 22 | 2ff9f24 | 6→**13** | 0→**1** | 5→**5** | yes | **TWO-TELL STOP — mandatory, surfaced to Daniel.** Findings more than doubled and the Blocker returned. Three findings (5, 8, 12) are defects in sentences the pass-21 round itself wrote; the Gate-A closing act has now taken a Blocker or Major at passes 20, 21 and 22, each out of the previous repair. All 13 held open; session 01a091bb-b1cd-7562-98d1-a50647a30370 | | 23 | — | — | — | — | not run | blocked on the two-tell answer | +## TWO-TELL STOP ANSWERED 2026-09-12 — deliberate rollback, not a fourth hardening round + +**Neither A, B nor C as I put them.** The reviewer proposed a fourth option and corrected his own +earlier recommendation to build the source-revision line: it was meant to *sharpen* the closing +decision and instead introduced a new attribution duty with consequences for record identity, +commit placement and content equality — an expansion the story's acceptance criteria never asked +for. Hardening it sentence by sentence was judged worse than taking it back. + +**Two of my claims were corrected, and both corrections hold:** +1. *"Stable apart from the closing act"* was too strong. Findings 4, 5 and 6 are elsewhere — + double-defined transitions, a possibly stale review payload, and the handling of valid legacy + cycles. +2. *"Three passes on one paragraph = three times one defect shape"* does not follow. Pass 20's + Blocker rationale was itself partly disproven and pass 21 had a Major there, not a Blocker. + **What is proven is repeated overclaiming** — and that licenses no outsourcing of the closing + transition, which is why C was refused: "an act is needed, the plan picks which" leaves the rule + text non-executable, and the repeat criterion can end the added attribution mechanics without + ending the closure itself. + +**What was rolled back.** The source-revision line, the artifact-path-plus-object-name form, and +every sentence explaining old commit names and their later readability. Recorded here as a +deliberate reversal of a design decision, not as a defect quietly dropped. + +**What replaced it, and the Blocker is answered rather than renamed.** Gate A closes by writing +the closing commit — or the closing message of one that exists — carrying **only the records the +cycle already owes**. **The closing commit must carry, at the artifact path, the content the final +pass was run against, unchanged.** That content condition is the link the Blocker said was +missing, and it is a condition on content, never on a commit name. Committing already-reviewed but +uncommitted content is explicitly **not** a new revision. + +**Three cases decided, the safe git sequence left to the plan:** + +| Case | Closing act | +|---|---| +| matching content already at the branch tip | amend that commit's message | +| matching content committed, its commit no longer the tip | a new commit changing nothing at the artifact path, records in its body — **nothing restored or rewritten** | +| final content still uncommitted | commit it unchanged and close in that commit | + +The third had **no answer before**: an eligible pass over repairs nobody had committed could +neither close nor suspend. The walk also caught a reading of case 2 that invited restoring an older +copy over newer content, which is now forbidden in the text. + +| Findings | Disposition | +|---|---| +| 1 (Blocker) | **Answered by the content condition**, not deferred. | +| 2, 8, 9, 12 | **Dissolve with the rollback** — no new record, so no nonce duty and no "ships no record" conflict; no commit-name claims left to be wrong about git or to overshoot. | +| 3 | **§F entry 7.** The Gate-A human-exception destination follows the closing act instead of naming the spec or plan commit, which is that commit on only one of three paths. | +| 4 | **Transitions outside §A replaced by references.** §D no longer says the stop closes nothing or names its answers; §H drops "continue resumes, stop parks". | +| 5 | **§F entry 6 rewritten.** "Unchanged" now scopes to the review question and dimensions; the artifact payload is replaced every pass, so no pass runs against last pass's text. | +| 6, 7, 11 | **§G, one paragraph, one edit.** The legacy bare-slot reservation is preserved; the squash-carry and curve-attribution rationales are narrowed to what they actually establish. | +| 10, 13 | **Collected.** 13's exclusivity claim was scoped in a clause already open; 10 — the precedence clause installed in both §A and §C — stands open and gets no round. | + +**§F is now seven standing sentences, five of them one mechanism.** That count is the honest +measure of this change's reach into standing text, and it has grown at every pass that looked. + ## Pass-22 three-line report — MANDATORY TWO-TELL STOP **Floor line:** derived floor **3**; risk **high**, security **none**; read fresh from diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index e5fd15f..2b1491d 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -119,7 +119,7 @@ to point at and so a reader can see the shape of the change without reading the | (e) the five tells | §D | | Mechanics · Severity | the resolve duty is scoped and gains its discharge rule; the handed-over question is replaced by its answer | | Mechanics · `baseSha` | two sentences: the `WIP:` warning stops claiming a closure the rules do not grant, and the `Finishing the cycle` lead-in performs the amend only where the ordering permits closing | -| Mechanics, recording a human exception | one sentence: a prescribed continuation answer is distinguished from blanket assent. The answer-record material stays moved to the successor | +| Mechanics, recording a human exception | two sentences: a prescribed continuation answer is distinguished from blanket assent, and the Gate-A destination follows the closing act rather than naming the spec or plan commit. The answer-record material stays moved to the successor | | Gate B, the coverage instruction | `NO FINDINGS` only when the branch found none | | Mechanics, the curve's Majors rationale | rewritten on the pre-ceiling reading | | Mechanics, the one-contract paragraph | membership widened, with a semantic test a downstream reader can apply | @@ -154,7 +154,7 @@ here. This table says what happens to each inventoried passage, so the map stays | (e) the five tells | edited — the threshold is read after clean completion, and a pointer says what its answer does | §D | | (f) the two rules above do not compete | **unchanged.** "The two rules above" still names the absorb rule and the stuck reading; the block sits before both and adds no third rule between them | — | | (g) Mechanics · Severity, the handed-over question | edited — the unsettled statement and its interim report-and-stop duty are replaced by the answer, in both copies, removing the one deliberate story-path divergence | §E | -| (h) recording a human exception | **edited in one sentence only** — a prescribed continuation answer is distinguished from blanket assent, which the ordering makes the restart of a parked cycle. The answer-record block that was to follow it stays moved to the successor with **D9**, and nothing else in the passage changes | §F | +| (h) recording a human exception | **edited in two sentences** — a prescribed continuation answer is distinguished from blanket assent, which the ordering makes the restart of a parked cycle; and the Gate-A destination follows the closing act, the spec-or-plan commit being that commit on only one of three closing paths. The answer-record block that was to follow it stays moved to the successor with **D9**, and nothing else in the passage changes | §F | | (i) when these rules bind | edited — the strict-reading list is added to, not rewritten | §H | | (j) the squash carry | **no longer edited.** It was to name the answer record, which moved; this change ships no record for it to carry | — | diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index 9fce2a7..fd47e4f 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -104,34 +104,37 @@ they disagree, which the branches below do. **The closing act, by cycle kind.** A **Gate-B** cycle closes with the closing amend Mechanics · Finishing the cycle describes. A **Gate-A** cycle has no WIP snapshot to replace, so -its closing act is **the commit body that carries this cycle's provenance line, its per-pass curve -and, beside them, one line naming the source revision the author accepts for this cycle**. The -first two are records Mechanics already obliges every cycle to write and **neither is altered -here** — their grammars are fixed and neither has a field for a revision. The third is **ordinary -prose in the same body**, which those grammars neither supply nor forbid. - -**The naming line is an attribution by the author, and is written as one.** It carries the -artifact's repository-relative path and the **full 40-character object name** of the commit being -accepted, an abbreviation being ambiguous across repositories and across time. **It says which -revision the author accepts; it says nothing about what the reviewer was given.** Gate A hands the -reviewer text rather than a git range, so no commit name could record what was reviewed and none -is offered as doing so. A reader of this line learns what was accepted and, about the review -itself, nothing. - -**Both placements are written out, because they carry the attribution differently.** Where the -named commit is **still the branch tip**, the body goes into it as **an amend of the message -alone**: the amend gives that commit a new object name and **leaves its tree untouched**, so the -content the line attributes is unchanged by the act of recording it. The line names the commit as -it stood before the amend and that name stops resolving, which costs nothing on this path — the -content sits in the tree of the commit the body ends up in. Where the named commit is **no longer -the tip**, the body goes into the **next commit made on the branch**, not the commit that happens -to follow the named one, which is already in the past. That may be **a commit carrying only the -record**: it changes no content, so it raises no review obligation. Mechanics grants such a commit -today for a human-exception record with nowhere else to go — **that is the precedent for this and -not a permission that already covers it**, the reasoning transferring where the allowance does -not. Placement attributes nothing on this path, so the naming line carries the attribution alone, -and it is written the same way. **No new revision of the artifact is made to close a Gate-A -cycle**, a new revision being one no pass has run against. Nothing is closed before that act, and +it closes by **writing the closing commit — or the closing message of one that already exists — +carrying the records this cycle already owes**: its provenance line and its per-pass curve, in the +forms Mechanics fixes, **neither altered and neither joined by a further record, line or form**. +Writing that body once every closure condition holds is the closing act, and nothing before it +closes anything. + +**The closing commit carries, at the artifact path, the content the final pass was run against, +unchanged.** That is the condition tying the closure to what was reviewed, and it is a condition +on **content**, never on a commit name. **Committing already-reviewed content that was not yet +committed is not a new revision** — a new revision is one no pass has run against, and these are +the bytes the clean pass read. **This says what the closing commit contains and nothing about what +the reviewer was given**: Gate A hands the reviewer text rather than a git range, so no part of +this act is offered as evidence of the review payload. + +**Three cases, all decided here; the safe git sequence for each belongs to the plan.** +- **The matching content is already at the branch tip.** Close by **amending that commit's + message**. The content is untouched, so the condition holds by construction. +- **The matching content is committed but its commit is no longer the tip.** The content at the + artifact path is unchanged since the final pass — otherwise the sameness precondition has + already failed and nothing closes — so close with a **new commit that changes nothing at that + path** and carries the records in its body. **Nothing is restored or rewritten**: the content is + already there, and a commit that put an older copy back would be a revision no pass has run + against. +- **The final content is still uncommitted.** **Commit it unchanged** and close in that commit. + This case had no answer before: an eligible pass over repairs nobody had committed could + neither close nor suspend. + +**A human-exception record this cycle owes goes in the commit its closing act uses**, so the two +never land in different places on the second and third paths. **No new revision of the artifact is +made to close a Gate-A cycle**, a new revision being one no pass has run against. Nothing is +closed before that act, and what changes in between still gates it: the **profile**, the **cited set**, the **assigned fix set**, the **artifact as it stood when the pass was run against it**, and — in a Gate-B cycle — the **evidence entry**. Any of them differing at the closing act makes that pass non-final and @@ -297,8 +300,9 @@ paragraph in `CLAUDE.md`, is absent from the template, and is left exactly as ea That divergence is pre-existing; this change neither creates nor removes it, and the plan's divergence list carries it. Earlier revisions split this passage's replacements across two sections and left one sentence half in each, which is how §B and §H came to instruct the plan -differently about the same clause. Nothing about this passage is stated anywhere else in this -file. +differently about the same clause. **The replacement bytes for this passage appear here and +nowhere else in this file** — other sections cite what this paragraph defines, as §A cites the fix +set and the two triggers, and citing is not a second copy. **Changed:** `b3`, `b7`, `b8`, `b11`, `b12`, `b13`, `b16`, `b17`–`b18`. **Carried:** every other sentence. The rationales follow the text. @@ -412,8 +416,8 @@ precondition for it. **A pointer is added** at the end of the passage: ``` -**What the answer does** is the closure ordering's: this stop is a **suspension**, it closes -nothing, and continue or stop is answered there. +**What the answer does** is the closure ordering's, which is where this stop's place among the +suspensions and what its answer produces are both stated. ``` --- @@ -463,10 +467,10 @@ demotes it — the two counts are meant to differ. --- -## F. The six standing sentences this change falsifies — REPLACED +## F. The seven standing sentences this change falsifies — REPLACED -Each is a live sentence that the block makes wrong. All six are **known contradictions** and -none is deferred. **Four of them share one mechanism** — an entry point other than the ordering +Each is a live sentence that the block makes wrong. All seven are **known contradictions** and +none is deferred. **Five of them share one mechanism** — an entry point other than the ordering carrying an unqualified instruction — which is why each is **replaced** rather than given an exception to point at. @@ -533,11 +537,12 @@ which is the split every sentence of this mechanism takes. **6. The Gate-A broad-prompt instruction** (Gate A section). It wraps across C 553–555 and W 745–747. ``` -Use ONE broad prompt, re-run it unchanged each pass over the **current** artifact — revised where -a repair is owed and unrevised where none is, an edit made to justify a pass being no reason to -run one. Don't narrow per-dimension: new findings surface because the artifact changed, because -an answer given since the last pass changed what the rules require of it, or because a broad -prompt reaches what the last reading did not. +Use ONE broad prompt: **its review question and dimensions stay the same every pass, while the +artifact text it carries is always the current one**. Re-running it over an **unrevised** artifact +is legitimate wherever no repair is owed, an edit made to justify a pass being no reason to run +one. Don't narrow per-dimension: new findings surface because the artifact changed, because an +answer given since the last pass changed what the rules require of it, or because a broad prompt +reaches what the last reading did not. ``` *Why (pass 21 finding 1):* the live wording says to re-run "over the revised artifact … because the artifact changes between passes", which the ordering's third branch contradicts — that branch @@ -545,8 +550,22 @@ continues on an **unrevised** artifact wherever no repair is owed. A pass whose Minors, or one run after a health answer with nothing left to repair, receives two instructions and can satisfy the live one only by manufacturing a change. **The replacement keeps the breadth demand**, which is what that sentence exists for and the reason it must not be narrowed -per-dimension; it drops only the claim that a revision always precedes a pass, and replaces the -reason with one that still holds when none did. +per-dimension; it drops only the claim that a revision always precedes a pass. +*And (pass 22 finding 5):* an earlier wording said to re-run the prompt "unchanged", which a Gate-A +prompt cannot be while also carrying the artifact text it reviews. **Unchanged** now scopes to the +question and the dimensions; the payload is replaced every pass, so no pass can be run against +last pass's text. + +**7. The human-exception destination** (Mechanics, recording a human exception). It wraps across +C 988–989 and W 1172–1173; only the Gate-A clause changes. +``` +**Which commit:** an ungated change records it in that commit; a Gate-A cycle in the commit its +closing act uses; a Gate-B cycle in the WIP commit, restated by the closing amend. +``` +*Why (pass 22 finding 3):* the live clause sends a Gate-A cycle's exception record to "the spec or +plan commit", which is the closing commit only on the first of the three closing paths. On the +other two the record and the closure would land in different commits. Naming the closing act +instead keeps them together on all three without changing the record's form or force. --- @@ -566,10 +585,12 @@ obliges a cycle to write.** Asking instead what an imagined edit would do decide any rule can be edited into deciding a branch and none decides one when edited cosmetically, so membership would follow the edit a reader pictured rather than the text in front of them. The last clause is why the squash carry belongs: it moves no pass and -decides no branch, and a record that does not survive the merge is a record the cycle did not -produce. A curve without a cycle field cannot be attributed, a slot rule -without a nonce has nothing to key on, a carry rule naming records a project does not produce is -inert, and a clean predicate without the fix-set boundary it reads decides membership by accident. +decides no branch, and a record that does not survive the merge is unreachable from the squash +commit and from `main`'s history. A curve without a cycle field cannot be told from another +cycle's where several are read together, a slot rule without a nonce cannot keep sibling cycles +apart — the bare names staying reserved for the legacy single-cycle case they already serve — a +carry rule naming records a project does not produce is inert, and a clean predicate without the +fix-set boundary it reads decides membership by accident. **A project whose text carries some of them and not others, or carries all of them in versions that disagree, stops and has a human complete, revert or reconcile the adoption before running a gate under it.** @@ -599,8 +620,7 @@ ordering decides both and decided them differently. the finding still open* — the resolve rule stands over it exactly as Mechanics · Severity states it, which means **repaired or validly dismissed**, so a recurrence of one already validly dismissed is already resolved and owes no second dismissal and no repair; the hold stands until -its answers are given, and **what the answer does is the closure ordering's**: continue resumes, -stop parks. +its answers are given, and **what the answer does is the closure ordering's**. **A pass is credited clean or not on its own findings**, as that ordering defines cleanliness; surfacing a finding that carries a scope-stop trigger is what withholds the credit, and no credit is withheld for surfacing alone. Reading this as "stop instead of fixing" would put the From bb65f17087ad803e8bfddec007143faf48161b9a Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sat, 12 Sep 2026 11:31:05 +0200 Subject: [PATCH 042/181] =?UTF-8?q?docs(context):=20Gate-A=20pass=2023=20?= =?UTF-8?q?=E2=80=94=20zero=20tells,=20findings=20open?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 9 findings, 0 Blockers, 7 Majors, 2 Minors. Valid pass. Zero tells of five, so no mandatory stop and the loop continues by the rules. The count is not the finding. Four of the nine are defects in text written the day before: 1, 2 and 3 sit in the closing-act paragraph the rollback produced, and 5 is in the §H sentence the pass-22 round rewrote. The closing-act paragraph has now produced a Blocker or Major at passes 20, 21, 22 and 23, with an overclaim at each of the last three. The rollback made the paragraph smaller and did not end its rounds. Two of this round's three are self-contradictions inside one paragraph. Line 109 forbids the closing body being joined by a further record while line 134 requires an owed human-exception record in that commit. Cases 1 and 2 both fire where an earlier commit introduced the content and later commits left that path alone. Finding 4 is new ground rather than regeneration: standing preconditions exist whose own action is stop and surface, while §A's third branch tells an eligible pass with any unmet precondition to run another pass. That is also why the clearly-stuck exit is not affirmable, its coverage condition needing a judgement this pass disproves. Findings not yet dispositioned; routed to the reviewer first. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-23.md | 10 ++++ .../gate-a-spec-awsf1ec771-resume.md | 48 ++++++++++++++++++- 2 files changed, 57 insertions(+), 1 deletion(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-23.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-23.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-23.md new file mode 100644 index 0000000..0f69d32 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-23.md @@ -0,0 +1,10 @@ +MAJOR | high | §A "Three cases, all decided here" | The first two Gate-A cases are not mutually exclusive: when an earlier commit introduced the reviewed content and later commits left that path unchanged, the branch tip carries the matching content while "its commit" is also no longer the tip; under the content-only design there is no identified commit that makes the second predicate exclusive | The same closing state permits both amending the tip and creating a new commit, so the promised exhaustive decision procedure does not produce exactly one closing act | Define the cases using mutually exclusive observable repository states, such as whether the reviewed bytes already equal the artifact at `HEAD`, and do not rely on an unidentified content-owning commit +MAJOR | high | §A "neither joined by a further record, line or form" | The Gate-A closing-act paragraph forbids any record beyond the provenance line and curve, but §A later requires an owed human-exception record in that same closing commit and standing Mechanics requires that record when its trigger holds | A Gate-A cycle with a human exception cannot satisfy both instructions, so it cannot perform the stated closing act without dropping an owed record or violating the new prohibition | Say that closure introduces no new closure-specific record while preserving every other record the cycle already owes, including a triggered human-exception record +MAJOR | high | §A "condition tying the closure to what was reviewed" | The text calls the compared content "what was reviewed" and "the bytes the clean pass read", although the standing Gate-A mechanism establishes only the text placed in the `mcp__codex__exec` request and this paragraph itself says the act establishes nothing about the review payload | The closing condition is presented as stronger review evidence than the mechanism supplies, reproducing the gate-proof overclaim AGENTS.md expressly forbids | Define sameness against the artifact text included in the final-pass request and keep the narrower statement that this comparison is not evidence of what the reviewer actually consumed +MAJOR | high | §A "an eligible pass with an unmet closure precondition lands here" | The third branch says every such pass runs another pass, but standing final-acceptance rules include conditions whose own required action is to stop and surface, including governing headers that disagree and a readable profile that is present but unresolvable | An agent can be told both to stop and to run the next pass even though no floor can be derived, so the ordering makes surviving standing stop rules ambiguous and §F misses this additional contradiction | Distinguish a blocking precondition's own stop-and-surface action from the three named suspensions and say the third branch runs another pass only when the unmet precondition's source permits a pass +MAJOR | high | §H "with the finding still open" | The replacement still categorically says every surfaced finding is open, then says in the same sentence that a re-raised valid dismissal is already resolved and owes neither repair nor another dismissal | A repeatedly re-raised false positive has incompatible resolve states and can be needlessly repaired or dismissed again instead of requiring only its new suspension answer | Say the cycle and any newly created hold remain open while an already validly dismissed finding remains resolved for the resolve duty +MAJOR | high | §A "the clearly-stuck hold alone" | A re-raised valid dismissal can independently open a new structural or contract question, which §B makes a question stop and §A elsewhere says adds its own answer to the health answer; therefore the recurrence does not always create only the clearly-stuck hold | The word "alone" can suppress the question-stop component and let the loop resume without the user's decision on the new contract question | Limit "alone" to the resolve-duty consequence and state that independently triggered scope-stop components still compose normally +MINOR | high | §A, §B and §E "one definition" | The declared ownership boundary is mechanically false: §A restates the resolve duty's exact repair-or-dismissal discharge while §E calls itself the only statement, and §B restates that the membership answer ends its hold and that the stop is a suspension even though §A claims those classifications and discharges | The target again installs parallel descriptions of the same decisions, defeating its passage-assignment claim and recreating the drift mechanism this consolidation is meant to remove | Choose the owner promised by the design for each rule and turn every other occurrence into a pure reference that does not repeat the classification, discharge, or transition +MAJOR | medium | §G "Membership is decided by a test" | "Determines or supplies an input the closure ordering reads" has no stated directness boundary: the file-validation rules directly supply finding lines, while tool routing, lens selection and broad-prompt wording also shape those lines and can therefore be included or excluded depending on whether a reader follows causal inputs transitively | Two downstream readers with only the shipped prompt can derive different contract memberships and therefore disagree on whether the same partial adoption must stop | State whether rules that merely influence production of findings are members and bound the semantic relation by an observable direct criterion, without restoring a closed list or a checker +MINOR | high | §A and §C "That third condition is what makes a plateau rather than a finish" | §C leaves this rationale as a sentence beside the clearly-stuck predicate while §A reproduces it again as the opening of the moved precedence sentence | One rationale is assigned to two proposed sections despite the target's single-passage requirement, adding a duplicate authority and avoidable prompt weight | Keep the full precedence sentence in §A and replace §C's duplicate clause with a non-repeating pointer, or keep the rationale only at its predicate source and move only the precedence clause +END OF FINDINGS (9 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index 5486a47..0b389eb 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -49,7 +49,53 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 20 | 5174d7a | 13→**9** | 1→**1** | 10→**5** | yes | **B+M 11→6, Majors halved.** Three of nine (3, 6, 9) regenerate from pass 19's own repairs; the Blocker is pass 19's Gate-A closing act, which names a revision in two record forms that have no field for one; session 01a0915d-6a5b-7633-a60a-881f0022816d | | 21 | 8dc22fb | 9→**6** | 1→**0** | 5→**5** | yes | **findings the lowest of the cycle; B+M 5 ties the pass-16 low; zero Blockers.** Finding 2 is my own overclaim from the pass-20 round; finding 1 is the sixth falsified standing sentence, found because the prompt asked for one; session 01a09195-8b00-7693-b42c-e85888188d3e | | 22 | 2ff9f24 | 6→**13** | 0→**1** | 5→**5** | yes | **TWO-TELL STOP — mandatory, surfaced to Daniel.** Findings more than doubled and the Blocker returned. Three findings (5, 8, 12) are defects in sentences the pass-21 round itself wrote; the Gate-A closing act has now taken a Blocker or Major at passes 20, 21 and 22, each out of the previous repair. All 13 held open; session 01a091bb-b1cd-7562-98d1-a50647a30370 | -| 23 | — | — | — | — | not run | blocked on the two-tell answer | +| 23 | d6052fe | 13→**9** | 1→**0** | 5→**7** | yes | **zero tells, no mandatory stop.** But findings 1, 2 and 3 are all in the closing-act paragraph the rollback rewrote the day before, and the paragraph has now produced a Blocker or Major at passes 20, 21, 22 and 23; session 01a094e9-76bd-7521-9c93-c11491c14202 | +| 24 | — | — | — | — | not run | next action, after pass 23's findings are dispositioned | + +## Pass-23 three-line report + +**Floor line:** derived floor **3**; risk **high**, security **none**; read fresh from +`docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md`. One cited story, level 2. + +- **Trend:** findings …, 9, 6, 13, **9**. Blockers …, 1, 0, 1, **0**. Majors …, 5, 5, 5, **7**. + Blocker+Major …, 12, 11, 6, 5, 6, **7**. Five passes in the 5–7 band. +- **Cluster (pass 23):** product 8 of 9; prose about the artifact 1 (finding 9, the duplicated + rationale); the instrument 0. +- **require↔withdraw:** none. Finding 9 re-raises pass 22's finding 10, which was **collected as a + Minor**, not removed — a reviewer repeating a collected finding is not a withdrawal. + +**Tells: zero of five.** Findings fell 13 → 9, the Blocker went to 0, the cluster is product, no +pair. **No mandatory stop**, and by the rules the loop continues on the third branch. + +**What the numbers do not say, stated because nothing obliges me to and it is the real finding.** +**Four of the nine are defects in text written the day before** — findings 1, 2 and 3 are in the +closing-act paragraph the rollback produced, finding 5 is in the §H sentence the pass-22 round +rewrote. And the closing-act paragraph has now produced a Blocker or Major at **passes 20, 21, 22 +and 23**, four consecutive rounds, with the **overclaim specifically at 21, 22 and 23**: +- pass 21: the repair claimed the pass had read the named commit's tree; +- pass 22: it claimed no commit name could record what was reviewed, overshooting the other way; +- pass 23: the content condition calls the compared bytes "what was reviewed" and "the bytes the + clean pass read", while Gate A only ever establishes the text placed in the request. + +**The rollback made the paragraph smaller and did not end its rounds.** That is the honest reading +of a changed approach that was tried once. + +**Two of this round's three are self-contradictions inside one paragraph**, which is new and worse +than an overclaim: line 109 forbids the closing body "joined by a further record, line or form" +while line 134 requires an owed human-exception record in that same commit (finding 2); and cases +1 and 2 both fire where an earlier commit introduced the content and later commits left that path +alone — the tip carries matching content *and* its commit is not the tip (finding 1). + +**Finding 4 is genuinely new ground and not regeneration.** Standing preconditions exist whose own +required action is **stop and surface** — disagreeing governing headers, an unresolvable cited +profile, an unreadable `Story:` header (`CLAUDE.md:90`) — and §A's third branch tells an eligible +pass with any unmet precondition to run another pass. An agent is told both to stop and to run. +**That is why the clearly-stuck exit is not affirmable**: its second condition needs a stated +judgement that coverage is sufficient, and pass 23 still reached material no earlier pass had. + +**The clearly-stuck reading stands at two of three** — a plateau (B+M 5–7 across five passes, +never zero in twenty-three) and nameable regeneration (findings 1, 2, 3, 5) — with coverage +unaffirmable. Two of three is not the exit. ## TWO-TELL STOP ANSWERED 2026-09-12 — deliberate rollback, not a fourth hardening round From 0d7e4b43e11a5b59780170cd32bf9edb94b675e3 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sat, 12 Sep 2026 11:50:09 +0200 Subject: [PATCH 043/181] docs(context): correct the pass-22 stop record; hold before any round 24 The reviewer challenged the claim that the pass-22 mandatory stop was answered, and the challenge is correct on both counts. The record showed only that Daniel forwarded the reviewer's counter-proposal and that the agent executed it. No separate decision by Daniel is recorded, and the decisions table carried no pass-22 row while every earlier stop has one. The section heading claimed otherwise and is corrected; the table now carries an explicit OPEN row. A second error of the agent's, same family: the pass-23 report justified continuing with "by the rules the loop continues on the third branch". That branch lives in the target text, which is not installed. The rule in force is CLAUDE.md:217-218, under which the loop resumes once the question is answered. Continuing was justified by the rule this change proposes rather than the rule that governs. Pass 23 was therefore run while a mandatory stop may have stood unanswered. Its findings are real and its file validates, so nothing is discarded; what is not established is that the loop was entitled to resume. No pass 24 runs and no artifact edit is made until the answer is recorded. Also added: a disposition proposal for the seven pass-23 Majors, unexecuted, so the reviewer's parking criterion can be answered on evidence. None of the seven needs attribution mechanics or an outsourced closing decision, and the first removes mechanics by deciding the closing case on observable state at HEAD rather than on commit history. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-resume.md | 52 ++++++++++++++++++- 1 file changed, 51 insertions(+), 1 deletion(-) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index 0b389eb..59fb56a 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -52,6 +52,33 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 23 | d6052fe | 13→**9** | 1→**0** | 5→**7** | yes | **zero tells, no mandatory stop.** But findings 1, 2 and 3 are all in the closing-act paragraph the rollback rewrote the day before, and the paragraph has now produced a Blocker or Major at passes 20, 21, 22 and 23; session 01a094e9-76bd-7521-9c93-c11491c14202 | | 24 | — | — | — | — | not run | next action, after pass 23's findings are dispositioned | +## Pass-23 disposition PROPOSAL — not executed, awaiting the pass-22 stop answer + +The reviewer's precondition for any round 24: the seven Majors dispositioned concretely, the +behavioural decisions made **in the target text** rather than left to the plan, and **open parking +recommended if any of them needs attribution mechanics or an outsourced closing decision again**. +This is that disposition, written so the parking question can be answered on evidence. **No edit +has been made.** + +| # | Proposed disposition | Needs mechanics or deferral? | +|---|---|---| +| 1 | **Two cases on observable current state, not three on commit history.** Does the artifact content at `HEAD` equal the reviewed text? **Yes → amend `HEAD`'s message. No → commit the reviewed text unchanged and close in that commit.** The old case 2 collapses into the first, the old case 3 into the second, and no commit-identity question is asked. | **No.** It removes mechanics rather than adding any. | +| 2 | **"Closure introduces no new record *type*; every record the cycle already owes is preserved"**, which is what the prohibition was meant to say. One sentence. | **No.** | +| 3 | **Sameness is against the artifact text included in the final pass's review request**, not "what was reviewed" or "the bytes the clean pass read". The disclaimer stays and stops having to undo a claim made two sentences earlier. | **No.** | +| 4 | **A precondition whose own source prescribes stop-and-surface is not one of the three suspensions and is not a continue.** The third branch runs another pass **only where the unmet precondition's source permits one**; otherwise the cycle takes that source's stop. A decision, stated in §A. | **No** — but it is a real behavioural decision and belongs in the text, as the reviewer says. | +| 5 | **The cycle and any newly created hold stay open; a finding already validly dismissed stays resolved for the resolve duty.** Replaces the categorical "with the finding still open". | **No.** | +| 6 | **"Alone" scopes to the resolve-duty consequence.** Independently triggered scope-stop components compose normally, so a re-raised dismissal that also opens a new contract question still stops for it. | **No.** | +| 8 | **Bound §G's relation to what a rule's own statement determines**, excluding rules that merely influence how findings get produced. Tool routing, lens selection and prompt wording shape the output and are not members; the file-validation rules say what counts as a finding line and are. No closed list, no checker. | **No** — this is the one where a list would be the easy wrong answer, and the proposal refuses it. | + +**Minors 7 and 9 get no round.** 7 (the "one definition" claim being mechanically false) sits in +sentences items 2, 3, 4 and 6 already open, so it costs nothing there; 9 (the rationale installed +in both §A and §C) stands open and is recorded. + +**Assessment, offered as one and not as a conclusion:** none of the seven needs attribution +mechanics or an outsourced closing decision, and item 1 **removes** mechanics. On the reviewer's +own criterion that argues for a bounded round rather than parking — but this is the agent's +reading, and the reviewer has corrected this agent's reading at every stop since pass 14. + ## Pass-23 three-line report **Floor line:** derived floor **3**; risk **high**, security **none**; read fresh from @@ -97,7 +124,29 @@ judgement that coverage is sufficient, and pass 23 still reached material no ear never zero in twenty-three) and nameable regeneration (findings 1, 2, 3, 5) — with coverage unaffirmable. Two of three is not the exit. -## TWO-TELL STOP ANSWERED 2026-09-12 — deliberate rollback, not a fourth hardening round +## TWO-TELL STOP — ANSWER NOT INDEPENDENTLY RECORDED. Read this before trusting the heading below. + +**Corrected 2026-09-12 after the reviewer challenged it.** This section was first written as "TWO-TELL +STOP ANSWERED". **What is on record is narrower**: Daniel forwarded the reviewer's counter-proposal, +and the agent executed it. **No separate decision by Daniel is recorded**, and the decisions table +below carries no pass-22 row while every earlier stop — 4, 9, 10, 14, 17 — has one. The standing +rule in force is `CLAUDE.md:217–218`: the loop **resumes once the question is answered**, and that +answer is the user's. + +**A second error of the agent's, in the same family.** The pass-23 report justified continuing with +"by the rules the loop continues on the third branch". **That branch is in the target text, which is +not installed.** The rule in force is the standing §5 sentence above. Continuing was justified by +the rule this change proposes rather than the rule that governs — which is the shape of mistake the +whole cycle exists to prevent. + +**Consequence, stated and not minimised:** pass 23 was run while a mandatory stop may have stood +unanswered. Its findings are real and its file validates, so nothing is discarded; what is not +established is that the loop was entitled to resume. **Until Daniel answers, no pass 24 runs and no +artifact edit is made.** + +**What was executed under that unconfirmed answer** — the rollback — **stands as work and not as an +authorised decision**, and the reviewer's own assessment of it is that the direction was right and +the concrete draft is still contradictory. **Neither A, B nor C as I put them.** The reviewer proposed a fourth option and corrected his own earlier recommendation to build the source-revision line: it was meant to *sharpen* the closing @@ -1038,6 +1087,7 @@ nothing checks them. They are candidates captured in | pass 14 scope stop | **B — accept, bounded.** Extend the existing coherence *instruction* to the closure block and its coupled edits (now the one-contract paragraph, target text §G). No checker, no new mechanism, no record-durability work. | | pass 15 | **Bounded rollback**, not another twelve-repair round: totals, item-number lists and enumerated precondition lists removed; behaviour stays decided in the design. | | pass 17 two-tell stop | **Interrupt the repair mode.** Produce a non-active **target-text** version; Gate A stays open, no transition to Gate B, no clearly-stuck close claimed. | +| pass 22 two-tell stop | **OPEN — no decision recorded.** Daniel forwarded the reviewer's counter-proposal (deliberate rollback of the source-revision line) and the agent executed it; no separate answer from Daniel is on record. Pass 23 was run under that unconfirmed answer. Nothing further runs until this row is filled. | ## For the execution phase, not needed yet From 6913c35ed1d9c2614b5ffce11a3bd6eb3a86ce51 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sat, 12 Sep 2026 11:56:45 +0200 Subject: [PATCH 044/181] docs(specs): apply Gate-A pass 23 under bounded authorisation MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The pass-22 stop is answered continue, the rollback is confirmed, and the record is corrected: a missing decision entry showed only that no unambiguous confirmation existed in the material, not that none was given. The governing rule for a two-tell stop is the tells paragraph at CLAUDE.md:263-268, not CLAUDE.md:217-218, which is the absorb paragraph and governs the scope stop. Seven Majors repaired with three precisions. The closing act now fixes the order in one paragraph: every closure condition is established first, among them that the artifact as it stands carries the final pass's request text unchanged, and only then is the act performed. The repository reading decides how a cycle closes and never whether it may, and restoring the request's old text over an intervening edit is explicitly not a way through. Three history-based cases become two told apart by current state, exhaustive and exclusive once the precondition holds. Closure now introduces no new kind of record and excuses none, which is what the earlier prohibition was meant to say and was not. Sameness compares against the text included in the final pass's review request, never against what the reviewer consumed. An unmet precondition whose own source prescribes stop-and-surface keeps its blocking effect, stated positively: the cycle stays open, that source decides the repair or answer, and no further pass runs while its block stands. No new state, no new procedure. The surfacing sentence no longer calls every surfaced finding open while saying a re-raised valid dismissal is resolved. "Alone" is deleted from the clearly-stuck hold sentence. §G's membership test is read on the sentence and never on the section: no paragraph is exempt as a paragraph, and the examples follow the criterion rather than replacing it. Minors 7 and 9 collected, no round. Gate A stays open; pass 23 was not clean. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-resume.md | 79 +++++++++++++---- ...-10-loop-rule-consolidation-target-text.md | 86 +++++++++++-------- 2 files changed, 112 insertions(+), 53 deletions(-) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index 59fb56a..7e172a6 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -52,7 +52,42 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 23 | d6052fe | 13→**9** | 1→**0** | 5→**7** | yes | **zero tells, no mandatory stop.** But findings 1, 2 and 3 are all in the closing-act paragraph the rollback rewrote the day before, and the paragraph has now produced a Blocker or Major at passes 20, 21, 22 and 23; session 01a094e9-76bd-7521-9c93-c11491c14202 | | 24 | — | — | — | — | not run | next action, after pass 23's findings are dispositioned | -## Pass-23 disposition PROPOSAL — not executed, awaiting the pass-22 stop answer +## Pass-23 dispositions — EXECUTED 2026-09-12 under Daniel's bounded authorisation + +**Three precisions the reviewer added to the proposal below, all applied:** + +1. **The closing cases decide only *how*, never *whether*.** The order is now written into the + contiguous paragraph: **first** every closure condition is established — among them that the + artifact as it now stands carries the final pass's request text unchanged — **then** the act is + performed. Where the artifact has moved, sameness has already failed and another pass is owed, + and **putting the request's old text back is explicitly not a way through**. That was the hole: + a `HEAD` comparison on its own would have let an intervening edit be papered over by restoring + the payload. +2. **Item 4 states the stop's effect positively.** "Neither suspension nor continue" only said what + it is not. The text now says: **the cycle stays open, the precondition's own source decides what + must be repaired or answered, and no further pass runs while its block stands** — with no new + named state and no procedure of its own, because the source rule already carries both. +3. **Item 8 admits no exemption by section.** The direct/indirect split stands, but the test is read + **on the sentence, never on the paragraph it sits in**: a sentence inside a routing or prompt + paragraph that fixes a valid input or an owed file **is** a member; a sentence anywhere that only + influences the findings is not. The examples now follow the criterion instead of replacing it. + +**Item 6 was simplified further on the reviewer's reading:** "alone" is deleted outright rather than +scoped, with one clause saying independently carried triggers raise their own stop as usual. + +**Minor 7 stands open and collected**, not repaired: §A spells out the resolve duty's discharge +while citing §E as its only statement. It is in no sentence this round opened, so it gets no round. +**Minor 9** (the rationale in both §A and §C) likewise. + +**The walk, before the commit.** With the precondition established, the two cases are exhaustive and +exclusive by construction: either `HEAD` carries that text at the artifact path or it does not. A +working-tree edit elsewhere does not reach the condition; a working-tree edit *to the artifact* +fails the precondition before any case is chosen. The stop-and-surface branch has no dead end, +since the source rule carries its own exit. + +### The proposal as it stood before those precisions + + The reviewer's precondition for any round 24: the seven Majors dispositioned concretely, the behavioural decisions made **in the target text** rather than left to the plan, and **open parking @@ -124,29 +159,35 @@ judgement that coverage is sufficient, and pass 23 still reached material no ear never zero in twenty-three) and nameable regeneration (findings 1, 2, 3, 5) — with coverage unaffirmable. Two of three is not the exit. -## TWO-TELL STOP — ANSWER NOT INDEPENDENTLY RECORDED. Read this before trusting the heading below. +## TWO-TELL STOP — ANSWERED 2026-09-12 by Daniel, after the record was corrected + +**What the record established, and what it did not.** This section was first written as "TWO-TELL +STOP ANSWERED" on the strength of Daniel forwarding the reviewer's counter-proposal, which the agent +then executed. **A missing decision entry does not show that Daniel never agreed** — what was +established is only that **no unambiguous confirmation existed in the material**, the decisions table +carrying no pass-22 row while every earlier stop has one. The fix was a clear answer, not a +reconstruction of fault. -**Corrected 2026-09-12 after the reviewer challenged it.** This section was first written as "TWO-TELL -STOP ANSWERED". **What is on record is narrower**: Daniel forwarded the reviewer's counter-proposal, -and the agent executed it. **No separate decision by Daniel is recorded**, and the decisions table -below carries no pass-22 row while every earlier stop — 4, 9, 10, 14, 17 — has one. The standing -rule in force is `CLAUDE.md:217–218`: the loop **resumes once the question is answered**, and that -answer is the user's. +**The governing rule, cited correctly on the second try.** The agent first cited `CLAUDE.md:217–218`, +which is the **absorb paragraph** and governs the **scope** stop. The two-tell stop is governed by the +tells paragraph at `CLAUDE.md:263–268`: *"Any two present makes stop-and-surface mandatory, not +discretionary — you report the tells and hand the decision to the user."* The conclusion — do not +continue unanswered — was right under either; the reason was wrong. **A second error of the agent's, in the same family.** The pass-23 report justified continuing with "by the rules the loop continues on the third branch". **That branch is in the target text, which is -not installed.** The rule in force is the standing §5 sentence above. Continuing was justified by -the rule this change proposes rather than the rule that governs — which is the shape of mistake the -whole cycle exists to prevent. +not installed.** Continuing was justified by the rule this change proposes rather than the rule that +governs — the shape of mistake the whole cycle exists to prevent. -**Consequence, stated and not minimised:** pass 23 was run while a mandatory stop may have stood -unanswered. Its findings are real and its file validates, so nothing is discarded; what is not -established is that the loop was entitled to resume. **Until Daniel answers, no pass 24 runs and no -artifact edit is made.** +**Consequence, stated and not minimised:** pass 23 was run before the stop's answer was recorded. Its +findings are real and its file validates, so nothing is discarded; what was not established at the +time was that the loop was entitled to resume. -**What was executed under that unconfirmed answer** — the rollback — **stands as work and not as an -authorised decision**, and the reviewer's own assessment of it is that the direction was right and -the concrete draft is still contradictory. +**Daniel's answer, 2026-09-12:** the rollback of the source-revision line is **confirmed**, and the +stop is answered **continue**. Authorised: **one bounded repair round** for the seven Majors under +three stated precisions, then **pass 24**, then **report and a fresh decision — no automatic further +round**. Minor and Nit get no round of their own. Gate A stays open; no activation and no transition +to Gate B. **Neither A, B nor C as I put them.** The reviewer proposed a fourth option and corrected his own earlier recommendation to build the source-revision line: it was meant to *sharpen* the closing @@ -1087,7 +1128,7 @@ nothing checks them. They are candidates captured in | pass 14 scope stop | **B — accept, bounded.** Extend the existing coherence *instruction* to the closure block and its coupled edits (now the one-contract paragraph, target text §G). No checker, no new mechanism, no record-durability work. | | pass 15 | **Bounded rollback**, not another twelve-repair round: totals, item-number lists and enumerated precondition lists removed; behaviour stays decided in the design. | | pass 17 two-tell stop | **Interrupt the repair mode.** Produce a non-active **target-text** version; Gate A stays open, no transition to Gate B, no clearly-stuck close claimed. | -| pass 22 two-tell stop | **OPEN — no decision recorded.** Daniel forwarded the reviewer's counter-proposal (deliberate rollback of the source-revision line) and the agent executed it; no separate answer from Daniel is on record. Pass 23 was run under that unconfirmed answer. Nothing further runs until this row is filled. | +| pass 22 two-tell stop | **Continue, bounded.** Rollback of the source-revision line confirmed. One repair round for the seven pass-23 Majors under three precisions — the closing cases decide only *how* to close and never override an intervening artifact change; the stop-and-surface precondition keeps its blocking effect stated positively; §G's membership criterion admits no blanket exemption by section. Then pass 24, then a fresh decision. Recorded late: the answer was given 2026-09-12 after the agent found no confirmation on record. | ## For the execution phase, not needed yet diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index fd47e4f..2d03597 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -106,37 +106,39 @@ they disagree, which the branches below do. Mechanics · Finishing the cycle describes. A **Gate-A** cycle has no WIP snapshot to replace, so it closes by **writing the closing commit — or the closing message of one that already exists — carrying the records this cycle already owes**: its provenance line and its per-pass curve, in the -forms Mechanics fixes, **neither altered and neither joined by a further record, line or form**. -Writing that body once every closure condition holds is the closing act, and nothing before it -closes anything. - -**The closing commit carries, at the artifact path, the content the final pass was run against, -unchanged.** That is the condition tying the closure to what was reviewed, and it is a condition -on **content**, never on a commit name. **Committing already-reviewed content that was not yet -committed is not a new revision** — a new revision is one no pass has run against, and these are -the bytes the clean pass read. **This says what the closing commit contains and nothing about what -the reviewer was given**: Gate A hands the reviewer text rather than a git range, so no part of -this act is offered as evidence of the review payload. - -**Three cases, all decided here; the safe git sequence for each belongs to the plan.** -- **The matching content is already at the branch tip.** Close by **amending that commit's - message**. The content is untouched, so the condition holds by construction. -- **The matching content is committed but its commit is no longer the tip.** The content at the - artifact path is unchanged since the final pass — otherwise the sameness precondition has - already failed and nothing closes — so close with a **new commit that changes nothing at that - path** and carries the records in its body. **Nothing is restored or rewritten**: the content is - already there, and a commit that put an older copy back would be a revision no pass has run - against. -- **The final content is still uncommitted.** **Commit it unchanged** and close in that commit. - This case had no answer before: an eligible pass over repairs nobody had committed could - neither close nor suspend. +forms Mechanics fixes, **neither of them altered**. **Closure introduces no new kind of record, and +it excuses none**: every other record this cycle owes, a human-exception record among them, is +owed and written exactly as before. Writing that body once every closure condition holds is the +closing act, and nothing before it closes anything. + +**The order is fixed, and the second step never repairs the first.** **First** every closure +condition is established, among them that **the artifact as it now stands carries the text that +went into the final pass's review request, unchanged**. **Only then** is the closing act performed, +and the shape it takes is read off the repository as it stands. **That reading decides how a cycle +closes, never whether it may**: where the artifact has moved since that request, the sameness +condition has already failed, the pass is not final and another is owed — and **putting the request's +old text back is not a way through**, being a revision no pass has run against. + +**Two cases, told apart by the repository's current state rather than by which commit introduced +what; the safe git sequence for each belongs to the plan.** +- **`HEAD` already carries that text at the artifact path.** Close by **amending `HEAD`'s message** + to add the records. Nothing at the artifact path moves. +- **`HEAD` does not, the reviewed text being still uncommitted.** **Commit it unchanged** and close + in that commit. This case had no answer before: an eligible pass over repairs nobody had + committed could neither close nor suspend. + +**What the sameness condition is, and all it is.** It compares the artifact against **the text +included in the final pass's review request** — not against a commit name, and **not against what +the reviewer consumed**, which nothing here establishes: Gate A hands the reviewer text rather than +a git range, so no part of this act is offered as evidence of the review payload. **A human-exception record this cycle owes goes in the commit its closing act uses**, so the two -never land in different places on the second and third paths. **No new revision of the artifact is -made to close a Gate-A cycle**, a new revision being one no pass has run against. Nothing is +never land in different places. **No new revision of the artifact is made to close a Gate-A +cycle**, a new revision being one no pass has run against — and committing already-reviewed text +that was never committed is not one. Nothing is closed before that act, and what changes in between still gates it: the **profile**, the **cited set**, the **assigned fix -set**, the **artifact as it stood when the pass was run against it**, and — in a Gate-B cycle — +set**, the **artifact measured against the final pass's request text**, and — in a Gate-B cycle — the **evidence entry**. Any of them differing at the closing act makes that pass non-final and owes another, which is the same answer a mid-pass change already gets. **Sameness is read on the artifact and the duties, never on the branch tip**: writing the closing body is itself a commit, @@ -174,7 +176,13 @@ it is not asked twice; the two-tell stop surfaces tells and not a finding. **current** artifact, revised where the severity and scope rules require a repair and unrevised where they do not. **An eligible pass with an unmet closure precondition lands here**: clean completion did not close it, and being eligible it cannot suspend, so the loop continues on whatever -the unmet precondition requires — most often a repair still owed from an earlier pass. A +the unmet precondition requires — most often a repair still owed from an earlier pass. +**Where that precondition's own source prescribes stop-and-surface instead** — a profile present +but unresolvable, governing headers that disagree, a `Story:` header that cannot be read — **the +cycle stays open, that source decides what must be repaired or answered, and no further pass runs +while its block stands.** It is neither a suspension nor a continue, and it needs no name and no +procedure of its own: the source rule already carries both, and the ordering's part is to send the +reader there rather than to run a pass over a cycle another rule has stopped. A below-floor clean pass lands here too, **only where no suspension applies to it**; where one does, the second branch has already taken it, because clean completion did not close the pass and only closing outranks a suspension. So does a pass whose only findings are Minors and Nits, which are @@ -219,8 +227,9 @@ regeneration chain that were repaired or dismissed are history the reading consu findings it re-surfaces, so no discharged finding takes a second hold. Each surfaced finding takes a hold like any other. **A re-raised valid dismissal stays discharged for the resolve duty** — the dismissal was the resolution and a reviewer repeating the finding does not undo it, so no second -dismissal is owed — and what the recurrence creates is the **clearly-stuck hold alone**, ended by -that reading's continue-or-stop answer. **Where one finding is surfaced by both, it carries two hold components and +dismissal is owed — and what the recurrence creates is the **clearly-stuck hold**, ended by +that reading's continue-or-stop answer. Any trigger the recurrence independently carries raises its +own stop as usual. **Where one finding is surfaced by both, it carries two hold components and each is discharged by its own answer**: the **membership** component ends on the membership answer **in either direction**, a decline releasing it exactly as an accept does; the **clearly-stuck** component ends on the reading's continue-or-stop answer. Neither @@ -581,7 +590,15 @@ no list to consult, and the test reads what a rule states rather than what chang a live rule belongs to this contract when what it says determines or supplies an input the closure ordering reads, which branch a pass takes, what a hold is or what discharges it, whether a cycle may close, or the production, identity or transport of any record this section -obliges a cycle to write.** Asking instead what an imagined edit would do decides nothing, because +obliges a cycle to write.** **Read it on the sentence, never on the section the sentence sits in.** +A sentence is a member when **it itself** fixes one of those things — what counts as a valid +finding line, which files or records are owed, what ends a hold. It is not a member when it only +shapes what a review produces, as the choice of reviewer, the lens set and the wording of a prompt +do: those change the findings without deciding what a finding *is* or what the ordering may do with +one. **No paragraph is exempt as a paragraph** — a sentence inside a routing or prompt paragraph +that fixes a valid input or an owed file is a member, and a sentence anywhere that only influences +the findings is not. The examples follow the test; they do not stand in for it. +Asking instead what an imagined edit would do decides nothing, because any rule can be edited into deciding a branch and none decides one when edited cosmetically, so membership would follow the edit a reader pictured rather than the text in front of them. The last clause is why the squash carry belongs: it moves no pass and @@ -617,9 +634,10 @@ duties are kept; the blanket no-clean-credit and the one-answer resumption go, b ordering decides both and decided them differently. ``` **Surfacing does not close the cycle, and that is what makes this reachable.** You surface *with -the finding still open* — the resolve rule stands over it exactly as Mechanics · Severity states -it, which means **repaired or validly dismissed**, so a recurrence of one already validly -dismissed is already resolved and owes no second dismissal and no repair; the hold stands until +the cycle and the new hold still open* — the resolve rule stands over the finding exactly as +Mechanics · Severity states it, which means **repaired or validly dismissed**, so a recurrence of +one already validly dismissed **stays resolved** and owes neither a second dismissal nor a repair; +the hold stands until its answers are given, and **what the answer does is the closure ordering's**. **A pass is credited clean or not on its own findings**, as that ordering defines cleanliness; surfacing a finding that carries a scope-stop trigger is what withholds the credit, and no From 79f2eb23a7ea35a7a16c6f9e7344e98a39e3fe35 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sat, 12 Sep 2026 12:11:27 +0200 Subject: [PATCH 045/181] =?UTF-8?q?docs(context):=20Gate-A=20pass=2024=20?= =?UTF-8?q?=E2=80=94=20cycle-low=20B+M,=20findings=20open,=20decision=20ow?= =?UTF-8?q?ed?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 8 findings, 1 Blocker, 3 Majors, 3 Minors, 1 Nit. Valid pass. One tell of five, so no mandatory stop. Blocker+Major at 4 and Majors at 3 are both the lowest of the cycle. The Blocker is new ground and is this repo's own documented failure class. The closing act chooses between its two cases by reading HEAD and the working artifact and never the index, while git commit --amend commits the index. A different blob or mode staged at the artifact path would be committed as the closing act while the text states that nothing at the path moved. AGENTS.md:93 records the same omission being fixed in the hook's fingerprint. Two of the eight come out of this round's repairs. Finding 2 says the intervening-edit-and-restore prohibition required by precision 1 cannot be observed by a content-only test and contradicts the later sentence about committing already-reviewed text. Finding 3 says precision 2's exception arrives after the third branch has already said "continues". Neither is a reason to undo a precision; both are about placement. Finding 4 is a pass-21 enumeration going stale: design §7 lists the replacement sections without §G, which is marked REPLACED. Findings 5 and 6 are collected Minors re-raised, 6 for the third time. They get no round under the standing decision. No repair round was authorised beyond pass 24. Findings are open and the decision is Daniel's. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-24.md | 9 ++++ .../gate-a-spec-awsf1ec771-resume.md | 50 ++++++++++++++++++- 2 files changed, 58 insertions(+), 1 deletion(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-24.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-24.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-24.md new file mode 100644 index 0000000..da66377 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-24.md @@ -0,0 +1,9 @@ +BLOCKER | high | §A "HEAD already carries that text" | The Gate-A cases inspect the working artifact and HEAD but never the effective index, while the standing closing operation `git commit --amend -m` commits the current index; HEAD and the working artifact can both carry the final request text while the index carries a different staged blob or mode at that path, and the second case can likewise commit an index entry different from the reviewed working text | The closing act can replace or commit the artifact with unreviewed content while declaring that nothing at the path moved, closing the gate in the dangerous direction AGENTS.md invariant 3 calls out explicitly | Keep the content-sameness measure unchanged, but make preservation of the reviewed artifact through the resulting commit an operational condition of both acts, including the effective-index state; leave the safe command sequence to the plan +MAJOR | high | §A "putting the request's old text back is not a way through" | The sameness test compares only the artifact as it stands with the text in the final-pass request, so it cannot distinguish continuously unchanged text from text changed and then restored byte-for-byte; the prohibition also calls the restored text a revision no pass ran against, while the later definition says committing already-reviewed text is not a new revision | The same observable state both satisfies and fails closure depending on which sentence the agent follows, and the text claims historical detection that its content-only mechanism does not provide | Define sameness solely as current equality to the request text and remove the unobservable intervening-edit and restoration prohibition, since no attribution or persistence mechanism is in scope +MAJOR | high | §A "Third, a pass that neither closes nor suspends continues" | The branch opens with an unqualified rule that every non-closing, non-suspending pass continues, then says an eligible pass with a source-prescribed stop-and-surface precondition is neither a suspension nor a continue and runs no further pass | A reader can either run another review under an unreadable or contradictory governing state or obey the source stop, so the pass-23 repair still leaves two incompatible next states | Qualify the third-branch rule and its eligible-pass sentence with the source-prescribed stop exception before stating that all remaining cases continue +MAJOR | high | design §7 "For every replacement" | The design says every replacement owes an old-wording-gone counterfactual and then identifies replacements in target §§B, C, D, E, F and H, omitting §G even though §G is explicitly REPLACED and design §4 maps the standing one-contract paragraph to it | The plan can follow the verification passage and omit the discriminating removal check for the old one-contract rule, allowing old and new membership rules to survive together while the promised evidence still passes | Add target §G to the replacement counterfactual obligation or remove the partial section enumeration and point only to the target's REPLACED markers +MINOR | high | §A "Everything else it names it cites" | The ownership claim is still false: §A repeats repair-or-validated-dismissal while §E calls itself the only resolve-duty statement, and §B and §H restate suspension and hold-discharge effects that §A says it owns | The target preserves parallel authorities for the same decisions, recreating the drift mechanism the consolidation is meant to remove and violating its own passage-assignment claim | Keep each definition at its promised owner and turn the other occurrences into references that do not repeat the classification, discharge, or transition +MINOR | high | §A and §C "That third condition is what makes a plateau rather than a finish" | §C leaves this rationale as a standalone sentence while §A reproduces it again as the opening of the verbatim moved precedence sentence | One passage is assigned to two proposed sections despite the target's one-section rule, adding duplicate prompt weight and a second sentence that can drift | Keep the verbatim precedence sentence in §A and replace §C's duplicate clause with a non-repeating pointer, or keep the rationale only at the predicate source and move only the precedence clause +MINOR | high | §A "No other pass outcome makes a cycle eligible to close" | The stated exhaustive reason is false: an at-floor pass that re-raises an in-set effective-Major finding already validly dismissed can be unclean while owing no repair, hold, or question, and an answered scope-stop pass can likewise remain immutably unclean after every hold ends | The operative eligibility rule still rejects those passes, but its rationale says every rejection has an outstanding duty or a floor shortfall and can mislead an agent into closing or skipping the required rerun | Add the third case the predicate actually reads: a pass can remain ineligible because its own fixed cleanliness result is false even after all cycle duties and answers are settled +NIT | high | §A "Below the floor the pass suspends, clean completion having not closed it" | The same section expressly defines clean completion to include eligibility, so a nonzero clean pass below the floor is not a clean completion at all | The terminology gives clean completion two meanings inside the ordering and makes the precedence explanation harder to apply literally | Say "the clean pass having failed eligibility" or otherwise use the section's defined term consistently +END OF FINDINGS (8 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index 7e172a6..93f8b3d 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -50,7 +50,55 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 21 | 8dc22fb | 9→**6** | 1→**0** | 5→**5** | yes | **findings the lowest of the cycle; B+M 5 ties the pass-16 low; zero Blockers.** Finding 2 is my own overclaim from the pass-20 round; finding 1 is the sixth falsified standing sentence, found because the prompt asked for one; session 01a09195-8b00-7693-b42c-e85888188d3e | | 22 | 2ff9f24 | 6→**13** | 0→**1** | 5→**5** | yes | **TWO-TELL STOP — mandatory, surfaced to Daniel.** Findings more than doubled and the Blocker returned. Three findings (5, 8, 12) are defects in sentences the pass-21 round itself wrote; the Gate-A closing act has now taken a Blocker or Major at passes 20, 21 and 22, each out of the previous repair. All 13 held open; session 01a091bb-b1cd-7562-98d1-a50647a30370 | | 23 | d6052fe | 13→**9** | 1→**0** | 5→**7** | yes | **zero tells, no mandatory stop.** But findings 1, 2 and 3 are all in the closing-act paragraph the rollback rewrote the day before, and the paragraph has now produced a Blocker or Major at passes 20, 21, 22 and 23; session 01a094e9-76bd-7521-9c93-c11491c14202 | -| 24 | — | — | — | — | not run | next action, after pass 23's findings are dispositioned | +| 24 | 6913c35 | 9→**8** | 0→**1** | 7→**3** | yes | **B+M 4 and Majors 3, both the lowest of the cycle.** One tell. The Blocker is new ground: the closing act reads `HEAD` and the working artifact and never the **index**, which is what `git commit` commits — `AGENTS.md:93`'s own documented failure class; session 01a0950c-f003-79d3-985f-73e2885c9621 | +| 25 | — | — | — | — | not run | **awaiting Daniel's decision**; no automatic follow-up round was authorised | + +## Pass-24 three-line report + +**Floor line:** derived floor **3**; risk **high**, security **none**; read fresh from +`docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md`. One cited story, level 2. + +- **Trend:** findings …, 13, 9, **8**. Blockers …, 1, 0, **1**. Majors …, 5, 7, **3**. + Blocker+Major …, 11, 6, 5, 6, 7, **4** — **the lowest of the cycle**, past pass 16's 5; Majors at + 3 are also a cycle low. +- **Cluster (pass 24):** product 6 of 8 (1, 2, 3, 5, 7, 8); the instrument 1 (4, design §7's + counterfactual list); prose about the artifact 1 (6, the duplicated rationale). +- **require↔withdraw:** none. Named: finding 2 asks for the **removal** of the intervening-edit + prohibition this round added on the reviewer's explicit instruction — a later pass objecting to an + addition, the mirror of the pair's shape. + +**Tells: one of five** — the Blocker count rose 0 → 1. Findings fell, Majors fell hard, the cluster +is product, no pair. **No mandatory stop.** + +**The Blocker is new ground and it is this repo's own documented failure class.** The closing act +decides between its two cases by reading **`HEAD` and the working artifact**, and never the +**index** — while `git commit --amend` commits the index. `HEAD` and the working tree can both +carry the final request text while a different blob or mode sits staged at that path, and the +closing act would then commit unreviewed content while stating that nothing at the path moved. +`AGENTS.md:93` says exactly this about the hook's own fingerprint: *"The index component exists +because `git commit` commits the index: without it, staging a change and reverting the file on disk +read as unchanged and reported satisfied."* The target text reinvented the omission the hook was +fixed for. + +**Two of the eight come out of this round's repairs, and one of them from the precision itself.** +Finding 2 says the intervening-edit-and-restore prohibition — required by precision 1 — **cannot be +observed by a content-only test** (unchanged text and text restored byte-for-byte are the same +state) **and contradicts the later sentence** saying committing already-reviewed text is not a new +revision. Finding 3 says precision 2's stop-and-surface exception arrives after the third branch has +already stated an unqualified "continues", so the qualifier is in the wrong place. Neither is a +reason to undo the precisions; both are about where and how they are stated. + +**Finding 4 is a pass-21 enumeration going stale**, the same mechanism as the `Kind` column: design +§7 lists the replacement sections as §§B, C, D, E, F and H, and **§G is marked REPLACED** at target +line 581. Confirmed by reading both. + +**Findings 5 and 6 are collected Minors being re-raised** — 5 for the second time (pass 23's 7), 6 +for the **third** (pass 22's 10, pass 23's 9). They get no round under the standing decision, and +the repetition is recorded rather than treated as new. + +**The §A closing act has now produced a Blocker or Major at passes 20, 21, 22, 23 and 24** — five +consecutive. Against that: Blocker+Major and Majors are both at cycle lows, and this pass's Blocker +is the first in that paragraph that is **not** a repair of a repair. ## Pass-23 dispositions — EXECUTED 2026-09-12 under Daniel's bounded authorisation From 31bbe8f9671fc45ae3fdd7bb77cde070a4911ab6 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sat, 12 Sep 2026 12:20:09 +0200 Subject: [PATCH 046/181] docs(specs): apply Gate-A pass 24; withdraw the restoration requirement MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Precision 1 from the previous round is withdrawn on the reviewer's own recommendation and is not defended here. It folded two different requirements into one: that the artifact as it now stands equals the final pass's request text, and that it was never changed since. The second does not follow from the first and no content comparison can observe it. Current equality is kept as the condition; a byte-identical restoration satisfies it, stated in the text as a decision with its reason, and every other closure condition stands independently. My pass-24 report called finding 2 a question of placement. That was wrong: the two requirements are substantively different. The Blocker is answered by making the duty reach into the commit the act produces. The equality is read on the artifact while the commit is written from the effective index, so an act that does not carry the same content through has checked the condition without performing it. Both cases owe the same thing: the closing commit contains at the artifact path exactly the text the condition was read against. Safe commands and evidence go to the plan; no new fingerprint and no new record. The third branch is qualified before it states anything. The source-prescribed stop is settled first, then the branch covers what remains, rather than an exception hung behind a contradicting instruction. Design §7 drops the stale section list. The old-wording-gone obligation now reaches every passage the target text marks REPLACED, with no second enumeration beside the markers, and the sentence records that the removed list had omitted §G. Findings 5, 6, 7 and 8 collected, no round; 6 is on its third re-raise. Gate A stays open; pass 24 was not clean. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-resume.md | 30 +++++++++++ ...26-09-10-loop-rule-consolidation-design.md | 8 +-- ...-10-loop-rule-consolidation-target-text.md | 54 +++++++++++-------- 3 files changed, 68 insertions(+), 24 deletions(-) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index 93f8b3d..0486479 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -53,6 +53,36 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 24 | 6913c35 | 9→**8** | 0→**1** | 7→**3** | yes | **B+M 4 and Majors 3, both the lowest of the cycle.** One tell. The Blocker is new ground: the closing act reads `HEAD` and the working artifact and never the **index**, which is what `git commit` commits — `AGENTS.md:93`'s own documented failure class; session 01a0950c-f003-79d3-985f-73e2885c9621 | | 25 | — | — | — | — | not run | **awaiting Daniel's decision**; no automatic follow-up round was authorised | +## Pass-24 dispositions — and a recommendation of the reviewer's own, withdrawn + +**Precision 1 from the pass-23 round is withdrawn by the reviewer and is not defended here.** It +had required the closing act to forbid an intervening artifact change being papered over by +restoring the request's old text. **That folded two different requirements into one**: + +| Requirement | Status | +|---|---| +| the artifact **as it now stands** equals the final pass's request text | **kept** — this is the condition | +| the artifact was **never changed since**, restoration included | **withdrawn** — it does not follow from the first, and no content comparison can observe it | + +**My pass-24 report called finding 2 a question of placement and wording. That was wrong**, and the +reviewer's correction is the reason: the two requirements are substantively different, and the +second was introduced by the precision rather than by the design. **A byte-identical restoration now +satisfies the content condition**, stated in the text as a decision with its reason, and **every +other closure condition stands independently of it.** + +| Finding | Disposition | +|---|---| +| 1 (Blocker) | **The duty now reaches into the commit the act produces.** The equality is read on the artifact, the commit is written from the **effective index**, so an act that does not carry the same content through has checked the condition without performing it. Both cases owe: the closing commit contains, at the artifact path, exactly the text the condition was read against. Safe commands and evidence to the plan; **no new fingerprint, no new record**. | +| 2 | **Historical prohibition removed**, per the withdrawal above. Current equality only, with no claim about reviewer consumption and none about gapless unchangedness. | +| 3 | **The stop case is settled before the third branch states anything** — "Third, and the branch is qualified before it is stated", then the stop, then "Otherwise, a pass that neither closes nor suspends continues". Not another exception hung behind a contradicting instruction. | +| 4 | **The stale section list is gone from design §7.** The obligation now reaches **every passage the target text marks REPLACED**, with no second enumeration beside the markers — and the sentence records that the old list omitted §G, so the reason survives the fix. | +| 5, 6, 7, 8 | **Collected, no round.** 6 is on its third re-raise and 5 on its second; both are recorded as repeats rather than treated as new. | + +**The walk before the commit.** Case 1's amend writes from the index, so a different staged blob at +the artifact path would be committed — the new duty is what catches it, and the plan builds the +sequence. Case 2 is unchanged. The stop case is unreachable from the third branch by construction. +Every replaced passage does carry a REPLACED marker, so item 4's pointer resolves. + ## Pass-24 three-line report **Floor line:** derived floor **3**; risk **high**, security **none**; read fresh from diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index 2b1491d..59b8a2b 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -225,9 +225,11 @@ exists; the completeness claim did not, because nothing supports it. block** it is **ABSENT and is claimed as absent**: the parent carries no such block, so no old wording of it can be shown to disappear and presence alone is the check. For **every replacement** the parent carries the old wording and the change removes it, so each owes the -old-wording-gone half of its pair — and the replacements are not one site: the target text -replaces standing wording in §§B, C, D, E, F and H, passage (g) among them. **Which sites those -are the plan derives from the target text**, where the concrete replacements live. Nothing is +old-wording-gone half of its pair. **The obligation reaches every passage the target text marks +REPLACED, and no list of them is kept here** — a second enumeration beside the markers is the +bookkeeping that goes stale, which it did: the list this sentence used to carry omitted §G while +§G was marked REPLACED. **The plan reads the markers off the target text**, where the concrete +replacements live. Nothing is claimed as "contradictory" — the second of the two defects Gate B found in the `fic2` instrument. **The named verification of the risk path** (story AC 4) is a **next-state table**, written in the diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index 2d03597..fad3b2f 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -111,27 +111,37 @@ it excuses none**: every other record this cycle owes, a human-exception record owed and written exactly as before. Writing that body once every closure condition holds is the closing act, and nothing before it closes anything. -**The order is fixed, and the second step never repairs the first.** **First** every closure -condition is established, among them that **the artifact as it now stands carries the text that -went into the final pass's review request, unchanged**. **Only then** is the closing act performed, -and the shape it takes is read off the repository as it stands. **That reading decides how a cycle -closes, never whether it may**: where the artifact has moved since that request, the sameness -condition has already failed, the pass is not final and another is owed — and **putting the request's -old text back is not a way through**, being a revision no pass has run against. +**The order is fixed.** **First** every closure condition is established, among them that **the +artifact as it now stands is identical to the text that went into the final pass's review +request**. **Only then** is the closing act performed, and the shape it takes is read off the +repository as it stands. **That reading decides how a cycle closes, never whether it may**: where +the artifact differs from that request text the condition has failed, the pass is not final, and +another is owed. + +**The condition is current equality, and deliberately nothing more.** It does **not** say the +artifact went untouched in between: text edited and then restored byte for byte satisfies it. That +is a decision and not an oversight — a content comparison cannot tell those two states apart, and a +condition nobody can check is a condition nobody applies. **Every other closure condition stands +independently and is not relaxed by this one.** It likewise says nothing about **what the reviewer +consumed**: Gate A hands the reviewer text rather than a git range, so no part of this act is +offered as evidence of the review payload. + +**The content must survive into the commit the act produces, and carrying it there is part of the +act.** The equality above is read on the artifact, while **the commit is written from the effective +index** — so an act that does not carry that same content through has checked the condition without +performing it. Both cases below owe the same thing: **the commit that closes the cycle contains, at +the artifact path, exactly the text the condition was read against.** The safe command sequence and +whatever demonstrates it belong to the plan; **the duty belongs here**, and it needs no new +fingerprint and no new record. **Two cases, told apart by the repository's current state rather than by which commit introduced what; the safe git sequence for each belongs to the plan.** - **`HEAD` already carries that text at the artifact path.** Close by **amending `HEAD`'s message** - to add the records. Nothing at the artifact path moves. + to add the records, leaving that path as it stands. - **`HEAD` does not, the reviewed text being still uncommitted.** **Commit it unchanged** and close in that commit. This case had no answer before: an eligible pass over repairs nobody had committed could neither close nor suspend. -**What the sameness condition is, and all it is.** It compares the artifact against **the text -included in the final pass's review request** — not against a commit name, and **not against what -the reviewer consumed**, which nothing here establishes: Gate A hands the reviewer text rather than -a git range, so no part of this act is offered as evidence of the review payload. - **A human-exception record this cycle owes goes in the commit its closing act uses**, so the two never land in different places. **No new revision of the artifact is made to close a Gate-A cycle**, a new revision being one no pass has run against — and committing already-reviewed text @@ -172,17 +182,19 @@ because a reason left out is a decision made by omission. A finding the clearly- surfaces that also carries either trigger takes the scope stop's answers at that same surface, so it is not asked twice; the two-tell stop surfaces tells and not a finding. -**Third, a pass that neither closes nor suspends continues** — the loop runs another pass on the +**Third, and the branch is qualified before it is stated.** **Where any unmet closure condition's +own source prescribes stop-and-surface** — a profile present but unresolvable, governing headers +that disagree, a `Story:` header that cannot be read — **the cycle stays open, that source decides +what must be repaired or answered, and no further pass runs while its block stands.** It is neither +a suspension nor a continue and needs no name and no procedure of its own: the source rule carries +both, and this ordering's part is to send the reader there rather than to run a pass over a cycle +another rule has stopped. **That case is settled first**, so what follows never reaches it. + +**Otherwise, a pass that neither closes nor suspends continues** — the loop runs another pass on the **current** artifact, revised where the severity and scope rules require a repair and unrevised where they do not. **An eligible pass with an unmet closure precondition lands here**: clean completion did not close it, and being eligible it cannot suspend, so the loop continues on whatever -the unmet precondition requires — most often a repair still owed from an earlier pass. -**Where that precondition's own source prescribes stop-and-surface instead** — a profile present -but unresolvable, governing headers that disagree, a `Story:` header that cannot be read — **the -cycle stays open, that source decides what must be repaired or answered, and no further pass runs -while its block stands.** It is neither a suspension nor a continue, and it needs no name and no -procedure of its own: the source rule already carries both, and the ordering's part is to send the -reader there rather than to run a pass over a cycle another rule has stopped. A +the unmet precondition requires — most often a repair still owed from an earlier pass. A below-floor clean pass lands here too, **only where no suspension applies to it**; where one does, the second branch has already taken it, because clean completion did not close the pass and only closing outranks a suspension. So does a pass whose only findings are Minors and Nits, which are From 36e5fe22a8c74809dca9ccec9e4c8264455007b2 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sat, 12 Sep 2026 12:38:16 +0200 Subject: [PATCH 047/181] =?UTF-8?q?docs(context):=20Gate-A=20pass=2025=20?= =?UTF-8?q?=E2=80=94=20B+M=202,=20cycle=20low,=20decision=20owed?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 10 findings, 0 Blockers, 2 Majors, 7 Minors, 1 Nit. Valid pass. One tell of five, so no mandatory stop. Blocker+Major at 2 and Majors at 2 are both the lowest of the cycle by a wide margin. Five of the ten are collected Minors being re-raised. One is on its fourth appearance across passes 22 to 25 and another on its third. Under the standing decision they get no repair round, so they will return at every future pass and keep the finding count inflated by a fixed amount. That is a property of the rule rather than a signal about the artifact. Both Majors are new ground. The first shows the pass-24 withdrawal leaking: the closing-time summary reads "any of them differing" over profile, cited set, fix set and artifact together, which is the decision taken for the artifact and wrong for the profile, where CLAUDE.md:760 makes any change cost a further pass regardless of the final value. The second shows §G's test cannot see a member that is wholly absent along with every sentence referring to it; the remedy asked for is to say so. Finding 7 was verified in the hook source rather than inferred. plugins/dev-workflow/hooks/codex-gate.sh:896 resets the Gate-A counter at skill boundaries, not on a commit, so §A's unqualified claim that a non-WIP commit discards Gate-A pass credit is false as written. No repair round was authorised beyond pass 25. Findings are open and the decision is Daniel's. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-25.md | 11 +++++ .../gate-a-spec-awsf1ec771-resume.md | 45 ++++++++++++++++++- 2 files changed, 55 insertions(+), 1 deletion(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-25.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-25.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-25.md new file mode 100644 index 0000000..fa577d9 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-25.md @@ -0,0 +1,11 @@ +MAJOR | medium | §A "Any of them differing at the closing act" | The artifact-only exception says current equality is deliberately enough and every other closure condition remains independent, but this later summary groups the profile and cited set with the artifact and makes a value still differing at the closing act the condition that costs the pass; both standing copies instead say "Any profile change costs at least one further pass" and require a new final pass whenever cited-set membership changes | A profile or cited set changed and then restored before closure can either require the standing further pass or satisfy the target's closing-time sameness summary, so the cycle can close over a live duty | State that current equality applies only to artifact content and that profile and cited-set transitions retain their source-prescribed further-pass semantics even when their final values match +MAJOR | high | §G "Membership is decided by a test a reader can apply to the text in front of them" | The test can classify only a live sentence, while the terminal rule also requires the reader to know that the project carries "some of them and not others"; if partial adoption omits an entire member together with the wording that would refer to it, a downstream reader holding only the shipped prompt has no sentence to test and no way to identify the absence | A wholly omitted contract member can be treated as a coherent adoption and a gate can run under the partial contract despite the paragraph's promised stop | Limit the terminal instruction to missing or inconsistent membership that is observable from the present text and state that whole-member omission is undecidable from that text alone, without adding a checker or closed list +MINOR | high | §A ownership boundary versus §§B and H | §A says it owns what a surfaced-finding hold is and what discharges it, yet §B separately says the membership answer ends the membership hold and §H says the hold stands until its answers are given; §A also repeats "repair or a validated dismissal" while §E calls itself the only resolve-duty statement | The target preserves parallel authorities for the same decisions, recreating the drift mechanism this consolidation is meant to remove and violating its own one-section assignment | Keep each definition at its promised owner and turn the other occurrences into references that do not repeat the hold or resolve transition +MINOR | high | §A and §C "That third condition is what makes a plateau rather than a finish" | §C leaves this rationale as a standalone sentence while §A reproduces it again as the opening of the supposedly word-for-word moved precedence sentence | One passage is assigned to two proposed sections, contrary to the target's section-assignment rule, and the duplicated rationale can drift independently | Keep the rationale only beside the clearly-stuck predicate in §C and move only the precedence clause to §A, or remove the duplicate from §C +MINOR | high | §A "No other pass outcome makes a cycle eligible to close" | Its exhaustive reason says every rejected pass leaves a repair, hold, question or floor shortfall, but an at-floor pass that re-raises an in-set effective-Major finding already validly dismissed is unclean while its resolve duty remains discharged and, before the clearly-stuck conditions mature, has no hold or question | The operative eligibility predicate still rejects the pass, but its rationale can make an agent demand a second dismissal or repair, contradicting the explicit idempotent-dismissal rule | Add fixed pass-level uncleanliness as the remaining reason a pass can be ineligible after every cycle-level duty and answer is settled +MINOR | high | §E "The curve must stay derivable from the findings files alone" | The standing curve also records model identifiers that findings files do not contain and explicitly permits `?` when a count cannot be recovered from those files; only the available Findings, Blockers and Majors series are derived from finding lines | The replacement overstates the curve's checkability and conflicts with the surviving unknown-count and model-source rules | Narrow the claim to the three numeric series when their validated findings files remain available +MINOR | high | §A "a non-`WIP` commit mid-cycle makes the hook read the cycle as closed and discards the passes" | The statement is unqualified across the Gate-A and Gate-B ordering, but the hook resets Gate-B review state on a non-WIP commit and does not reset the Gate-A pass counter on commit; Gate-A state is reset at its skill boundaries instead | A Gate-A reader is told that an accidental commit destroyed pass credit when the hook retains it, overstating what the mechanism observes and causing unnecessary reruns or false closure diagnostics | Scope this hook observation explicitly to Gate B and describe Gate-A logical closure without attributing it to the commit-reset behavior +MINOR | high | §G "A curve without a cycle field cannot be told from another cycle's" | Even where several curves are read together, standing Mechanics says missing attribution is a limitation rather than a disqualification because cycle kind and surrounding context can sometimes distinguish them; the nonce merely means a later reader usually does not need that context | The contract rationale upgrades a probabilistic, context-dependent attribution aid into an impossibility claim, the mechanism-overclaim class AGENTS.md warns against | Say the curves cannot be reliably distinguished in every multi-cycle context without the cycle field +MINOR | medium | §G "any record this section obliges a cycle to write" | Once installed, §G sits under `### Mechanics (reference)` inside §5, so "this section" can mean Mechanics or the whole Cross-Model Review section; those scopes need not contain the same record obligations, yet the phrase decides contract membership | Two downstream readers can classify a record-production sentence differently while both apply the promised text-only test | Name §5 explicitly as the scope of the record-obligation clause +NIT | high | §A "Below the floor the pass suspends, clean completion having not closed it" | The same section defines clean completion to include eligibility, so a nonzero clean pass below the floor is not a clean completion at all | The phrase gives clean completion two meanings inside the ordering and makes the precedence explanation needlessly difficult to apply literally | Say the clean pass failed eligibility because it was below the floor +END OF FINDINGS (10 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index 0486479..27436d6 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -51,7 +51,50 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 22 | 2ff9f24 | 6→**13** | 0→**1** | 5→**5** | yes | **TWO-TELL STOP — mandatory, surfaced to Daniel.** Findings more than doubled and the Blocker returned. Three findings (5, 8, 12) are defects in sentences the pass-21 round itself wrote; the Gate-A closing act has now taken a Blocker or Major at passes 20, 21 and 22, each out of the previous repair. All 13 held open; session 01a091bb-b1cd-7562-98d1-a50647a30370 | | 23 | d6052fe | 13→**9** | 1→**0** | 5→**7** | yes | **zero tells, no mandatory stop.** But findings 1, 2 and 3 are all in the closing-act paragraph the rollback rewrote the day before, and the paragraph has now produced a Blocker or Major at passes 20, 21, 22 and 23; session 01a094e9-76bd-7521-9c93-c11491c14202 | | 24 | 6913c35 | 9→**8** | 0→**1** | 7→**3** | yes | **B+M 4 and Majors 3, both the lowest of the cycle.** One tell. The Blocker is new ground: the closing act reads `HEAD` and the working artifact and never the **index**, which is what `git commit` commits — `AGENTS.md:93`'s own documented failure class; session 01a0950c-f003-79d3-985f-73e2885c9621 | -| 25 | — | — | — | — | not run | **awaiting Daniel's decision**; no automatic follow-up round was authorised | +| 25 | 31bbe8f | 8→**10** | 1→**0** | 3→**2** | yes | **B+M 2 and Majors 2, both by far the lowest of the cycle; zero Blockers.** Both Majors are new ground. **Five of the ten are re-raised collected Minors**, one on its fourth appearance; session 01a09522-65bc-7691-8adc-fb26e330810e | +| 26 | — | — | — | — | not run | **awaiting Daniel's decision**; no automatic follow-up round was authorised | + +## Pass-25 three-line report + +**Floor line:** derived floor **3**; risk **high**, security **none**; read fresh from +`docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md`. One cited story, level 2. + +- **Trend:** findings …, 9, 8, **10**. Blockers …, 1, 0, **0**. Majors …, 7, 3, **2**. + Blocker+Major …, 6, 5, 6, 7, 4, **2** — **the lowest of the cycle by a wide margin**; Majors at 2 + are also a cycle low, and this is the sixth zero-Blocker pass in twenty-five. +- **Cluster (pass 25):** product 9 of 10; prose about the artifact 1 (finding 4); the instrument 0. +- **require↔withdraw:** none. Finding 1 is adjacent and is not one: it does **not** ask the withdrawn + historical condition back for the artifact — it says the withdrawal leaked onto rules it was never + meant to touch. + +**Tells: one of five** — the finding count rose 8 → 10. Blockers fell to zero, Majors fell, the +cluster is product, no pair. **No mandatory stop.** + +**What moved the count, and it is not new defect surface.** **Five of the ten are collected Minors +being re-raised** — findings 3, 4, 5, 8 and 10. Finding 4 (one rationale installed in both §A and +§C) is on its **fourth** appearance across passes 22, 23, 24 and 25; finding 3 on its third. Under +the standing decision they get no repair round, so **they will return at every future pass and keep +the finding count — and its tell — inflated by a fixed amount.** That is a property of the rule, not +a signal about the artifact, and it is recorded here so the count is not misread later. + +**Both Majors are new ground and both are narrow.** +- **Finding 1 — the withdrawal leaked.** The closing-time summary reads *"Any of them **differing** + at the closing act"* over profile, cited set, fix set and artifact together. For the artifact that + is exactly the decision taken: current equality. **For the profile it is wrong** — `CLAUDE.md:760` + says *"Any profile change costs at least one further pass, in either direction"*, a **change**, + not a differing value, and `CLAUDE.md:770` says the same of cited-set membership. A profile + changed and restored before closure satisfies my summary and violates the standing rule. +- **Finding 2 — §G's test cannot see a wholly missing member.** It classifies sentences that are + present; where a partial adoption drops a member together with every sentence referring to it, a + reader holding only the shipped prompt has nothing to test and no way to notice the absence. The + remedy asked for is to **say so**, not to build anything. + +**Finding 7 verified in the hook source, not inferred.** §A says without qualification that a +non-`WIP` commit mid-cycle makes the hook read the cycle as closed and discards the passes counted +so far. `plugins/dev-workflow/hooks/codex-gate.sh:896` resets the **Gate-A** counter at skill +boundaries — `superpowers:brainstorming` and `superpowers:writing-plans` — **not on a commit**. The +claim holds for Gate B and is false as written for Gate A, which is the overclaim class `AGENTS.md` +names, found for the fifth pass running. ## Pass-24 dispositions — and a recommendation of the reviewer's own, withdrawn From 8b8e14605ad2c6961a363a4f131087c3fb2b0c66 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sat, 12 Sep 2026 12:45:13 +0200 Subject: [PATCH 048/181] docs(specs): apply Gate-A pass 25; dissolve the joint equality rule MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Two numeric corrections of mine first, both recounted from the pass files. The Blocker sequence for passes 23 to 25 is 0, 1, 0 and my trend line read 1, 0, 0; the tell reading is unaffected. And the re-raised Minor pile does not explain the rise, since a constant base cannot produce one: pass 24 carried 3 Minors, pass 25 carries 7, three of them new, and recurrence is not guaranteed either. The claim that the pile inflates the count by a fixed amount is withdrawn. Major 1 is repaired by dissolving the joint sentence rather than adding exceptions to it. What gates the closing act is read condition by condition: the artifact on current equality, the profile and cited set on their own sources, which answer a change and not a differing value, so one changed and then restored still costs the pass. The artifact's current-equality reading is stated as the artifact's alone. Major 2 names §G's limit without weakening its duty. The test classifies sentences that are present and does not establish completeness, since an adoption dropping a member with every sentence referring to it leaves nothing to mark the absence. That bounds detection and not the obligation: a partial adoption stops however it becomes known. The hook sentence is scoped to Gate B inside the paragraph Major 1 already opened, so it costs no Minor round. Verified at plugins/dev-workflow/hooks/codex-gate.sh:896. The neighbouring sentences were checked and one did not match: the ownership boundary claimed every closing-time sameness test is defined in the block, which is now true only of the artifact's. Corrected in the same round. Gate A stays open; pass 25 was not clean. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-resume.md | 54 +++++++++++++++---- ...-10-loop-rule-consolidation-target-text.md | 46 ++++++++++------ 2 files changed, 74 insertions(+), 26 deletions(-) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index 27436d6..62e5dc1 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -54,12 +54,39 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 25 | 31bbe8f | 8→**10** | 1→**0** | 3→**2** | yes | **B+M 2 and Majors 2, both by far the lowest of the cycle; zero Blockers.** Both Majors are new ground. **Five of the ten are re-raised collected Minors**, one on its fourth appearance; session 01a09522-65bc-7691-8adc-fb26e330810e | | 26 | — | — | — | — | not run | **awaiting Daniel's decision**; no automatic follow-up round was authorised | +## Pass-25 dispositions — two Majors, one mis-scoped hook sentence, two corrected numbers + +**Two numeric corrections of mine, both verified against the pass files before acting:** +1. The Blocker sequence for passes 23–25 is **0, 1, 0**; my trend line read 1, 0, 0. The tell + reading is unaffected — 1 → 0 is a fall either way. +2. **The re-raised Minor pile does not explain the rise, and a constant base could not.** Pass 24 + carried 3 Minors, pass 25 carries 7, of which **three are new**. Recurrence is also not + guaranteed. The claim that the pile keeps the count "inflated by a fixed amount" is withdrawn, + and the tell rule is neither explained away nor reinterpreted. + +| Finding | Disposition | +|---|---| +| 1 (Major) | **The joint equality sentence is dissolved rather than qualified.** What gates the closing act is now read **condition by condition**: the artifact on current equality; the **profile** and **cited set** on their own sources, which answer a **change** and not a differing value, so one changed and then restored **still costs the pass**; the fix set and the evidence entry likewise. The artifact's current-equality reading is stated as the artifact's alone. | +| 2 (Major) | **§G names the limit without weakening the duty.** The test classifies sentences that are present and does not establish completeness — an adoption that drops a member with every sentence referring to it leaves nothing to mark the absence. **That bounds detection, never the obligation:** a partial adoption is a stop **however it becomes known**, including from outside this text. | +| 7 (Minor) | **Repaired inside the paragraph Major 1 already opened**, so it costs no Minor round: the hook sentence is scoped to **Gate B**, and the text says it does not reach a Gate-A cycle's count, which the hook clears at the skill boundaries that start a new Gate-A cycle. Verified at `plugins/dev-workflow/hooks/codex-gate.sh:896`. | +| 3, 4, 5, 6, 8, 9, 10 | **Collected, no round.** | + +**The neighbouring sentences were checked and one did not match.** The ownership boundary claimed +that *"the closing-time sameness tests below are defined here … they borrow no rule and have no +other source"* — true of the artifact's test and **false of the profile's and the cited set's** +after this round, since those now defer to their sources. Corrected in the same round; it is exactly +the mismatch the instruction warned about. + +**No zero-findings target is implied by any of this.** A pass carrying no in-set Blocker or Major at +effective severity and no scope-stop trigger is clean with collected Minors still standing. + ## Pass-25 three-line report **Floor line:** derived floor **3**; risk **high**, security **none**; read fresh from `docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md`. One cited story, level 2. -- **Trend:** findings …, 9, 8, **10**. Blockers …, 1, 0, **0**. Majors …, 7, 3, **2**. +- **Trend:** findings …, 9, 8, **10**. Blockers …, **0, 1, 0** (corrected — the first version of this + line read 1, 0, 0 and was wrong; recounted from the pass files). Majors …, 7, 3, **2**. Blocker+Major …, 6, 5, 6, 7, 4, **2** — **the lowest of the cycle by a wide margin**; Majors at 2 are also a cycle low, and this is the sixth zero-Blocker pass in twenty-five. - **Cluster (pass 25):** product 9 of 10; prose about the artifact 1 (finding 4); the instrument 0. @@ -67,15 +94,24 @@ Nothing depends on it; the pass files and the repo are authoritative where this historical condition back for the artifact — it says the withdrawal leaked onto rules it was never meant to touch. -**Tells: one of five** — the finding count rose 8 → 10. Blockers fell to zero, Majors fell, the -cluster is product, no pair. **No mandatory stop.** +**Tells: one of five** — the finding count rose 8 → 10. Blockers fell 1 → 0, Majors fell, the +cluster is product, no pair. **No mandatory stop.** The tell stands as the rule states it. + +**Five of the ten are collected Minors being re-raised** — findings 3, 4, 5, 8 and 10. Finding 4 +(one rationale installed in both §A and §C) is on its **fourth** appearance across passes 22–25; +finding 3 on its third. + +**What that does and does not explain — corrected after the reviewer challenged it.** The first +version of this report said the re-raised pile keeps the count "inflated by a fixed amount" and +would return at every future pass. **Both halves are wrong.** A constant base cannot produce a +**rise**, and the rise is real: pass 24 carried 3 Minors, pass 25 carries 7, of which **three are +new** (§E's curve-derivability claim, the hook's scope, §G's "this section"). Nor is recurrence +guaranteed — a reviewer may simply not raise a collected finding again. **The tell rule is not +explained away and is not reinterpreted here.** -**What moved the count, and it is not new defect surface.** **Five of the ten are collected Minors -being re-raised** — findings 3, 4, 5, 8 and 10. Finding 4 (one rationale installed in both §A and -§C) is on its **fourth** appearance across passes 22, 23, 24 and 25; finding 3 on its third. Under -the standing decision they get no repair round, so **they will return at every future pass and keep -the finding count — and its tell — inflated by a fixed amount.** That is a property of the rule, not -a signal about the artifact, and it is recorded here so the count is not misread later. +**And no zero-findings target is being built.** A later pass with no in-set Blocker or Major at +effective severity and no scope-stop trigger **is clean with collected Minors still standing**; the +remaining closure conditions decide the rest. The Minor pile does not have to reach zero. **Both Majors are new ground and both are narrow.** - **Finding 1 — the withdrawal leaked.** The closing-time summary reads *"Any of them **differing** diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index fad3b2f..6e2ec70 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -54,10 +54,12 @@ evaluation order or as a thing beside it is a question of wording, not of author turns that wording into an argument about whether the paragraph has overrun its own boundary. **Everything else it names it cites**: the scope triggers, the assigned fix set, every severity rule, and every closure precondition **that has a source of its own** keep their one definition in the paragraph that owns them, and a -reader who finds one of *those* defined here has found a defect. **The closing-time sameness -tests below are defined here and are not an exception to that**: they borrow no rule and have no -other source, being part of the closure decision this paragraph owns rather than a precondition -stated elsewhere and read from here. +reader who finds one of *those* defined here has found a defect. **The artifact's closing-time +equality test below is defined here and is not an exception to that**: it borrows no rule and has +no other source, being part of the closure decision this paragraph owns rather than a precondition +stated elsewhere and read from here. **The other closing-time gates are not like it** — the +profile's and the cited set's are their own sources' rules, cited from here and not restated, which +is the sentence before this one applied rather than excepted. **What a pass is read from.** Every finding-derived predicate reads the validated findings file **or files** of the logical pass as **the concatenation of their finding lines after each file @@ -145,18 +147,24 @@ what; the safe git sequence for each belongs to the plan.** **A human-exception record this cycle owes goes in the commit its closing act uses**, so the two never land in different places. **No new revision of the artifact is made to close a Gate-A cycle**, a new revision being one no pass has run against — and committing already-reviewed text -that was never committed is not one. Nothing is -closed before that act, and -what changes in between still gates it: the **profile**, the **cited set**, the **assigned fix -set**, the **artifact measured against the final pass's request text**, and — in a Gate-B cycle — -the **evidence entry**. Any of them differing at the closing act makes that pass non-final and -owes another, which is the same answer a mid-pass change already gets. **Sameness is read on the -artifact and the duties, never on the branch tip**: writing the closing body is itself a commit, -so a commit that only records the closure changes neither, while an edit to the artifact or a -broadening of the set changes one and costs the pass. **A commit the hook reads as cycle-closing is a separate matter**: a non-`WIP` commit -mid-cycle makes the hook read the cycle as closed and discards the passes counted so far, which -is an observation about the counter — the cycle itself stays open until the conditions above -hold, so an accidental commit destroys the pass credit and closes nothing. A plateau or tells on +that was never committed is not one. **Nothing is closed before that act, and what still gates it in +between is read condition by condition rather than as one comparison.** The **artifact** gates on +the equality above and on nothing stronger. The **profile** and the **cited set** gate as their own +sources say, and those sources answer a **change** and not a differing value: a profile change +costs at least one further pass in either direction, and a cited-set membership change makes the +final clean pass run against the current set — **so one of those changed and then restored still +costs the pass**, which a single "do they differ" test would have let through. The **assigned fix +set** and, in a Gate-B cycle, the **evidence entry** likewise keep the rules their own sources +give them. **The artifact's current-equality reading belongs to the artifact and is not carried to +any of the others.** None of these is read on the branch tip: writing the closing body is itself a +commit and changes none of them. + +**A commit the hook reads as cycle-closing is a separate matter, and it is a Gate-B one.** A +non-`WIP` commit mid-cycle makes the hook drop its Gate-B review state and that gate's counter — +an observation about the counter, since the cycle itself stays open until the conditions above +hold, so an accidental commit destroys Gate-B pass credit and closes nothing. **It does not reach +a Gate-A cycle's count**, which the hook clears at the skill boundaries that start a new Gate-A +cycle rather than on any commit. A plateau or tells on **the pass that closes** go into **that pass's status report to the user** — the carrier the three-line duty already names, and no second report form is introduced — and never block it, because reporting "will not converge" on a converged loop is a false report; on an eligible pass @@ -622,7 +630,11 @@ carry rule naming records a project does not produce is inert, and a clean predi fix-set boundary it reads decides membership by accident. **A project whose text carries some of them and not others, or carries all of them in versions that disagree, stops and has a human complete, revert or reconcile the adoption before running a -gate under it.** +gate under it.** **The test classifies sentences that are present, and completeness is not among +what it establishes**: where an adoption drops a member together with every sentence that would +refer to it, what remains reads as coherent and nothing in it marks the absence. **That bounds +what this text lets a reader detect, never what the rule obliges** — a partial adoption is a stop +however it becomes known, and learning of it from outside this text is learning of it. ``` **What this is, stated so nothing reads it as more:** an **instruction to the agent**, over a From 56d8f5a608be974cd852eaec962611a09471a2c4 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sat, 12 Sep 2026 13:02:43 +0200 Subject: [PATCH 049/181] =?UTF-8?q?docs(context):=20Gate-A=20pass=2026=20?= =?UTF-8?q?=E2=80=94=20mandatory=20two-tell=20stop,=20findings=20held=20op?= =?UTF-8?q?en?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 14 findings, 1 Blocker, 5 Majors, 7 Minors, 1 Nit. Valid pass. Two tells of five, which is the threshold: the finding count rose 10 to 14 and the Blocker count failed to fall, 0 to 1. Stop-and-surface is mandatory. All 14 are held open and unrepaired pending the answer. Every Blocker and Major comes out of a repair made in passes 23, 24 or 25, each nameable. The Blocker is the artifact-to-request equality being written cycle-generally while Gate B passes a git range and has no value to test, which leaves an eligible Gate-B pass unable to close and, being eligible, unable to suspend. One shape runs through five of the six: a rule stated in §A that also lives at a source, or a source sentence §A's wording contradicts. That is the two-authorities mechanism pass 11 named. The collected Minor pointing at it has been re-raised at passes 23 through 26 and another at five consecutive passes, while the Majors it produces were repaired one at a time. Verified rather than inferred: the cycle-general scoping at target 117-120 and 149-153; §F's suspension-answer sentence at 540 against §A's membership answer at 259; and CLAUDE.md:354, which ships the optional dispositions companion the target says these rules do not ship. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-26.md | 15 ++++++ .../gate-a-spec-awsf1ec771-resume.md | 48 ++++++++++++++++++- 2 files changed, 62 insertions(+), 1 deletion(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-26.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-26.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-26.md new file mode 100644 index 0000000..4aa0924 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-26.md @@ -0,0 +1,15 @@ +BLOCKER | high | §A "The order is fixed" | The artifact-to-request equality is stated as a closure condition for every cycle, but its explanation and two cases are Gate-A-only: Gate B sends a git range identified by baseSha and headSha rather than artifact text in the review request, so Gate B has no value with which to evaluate this condition, while the standing Gate-B freshness rule instead concerns re-review after changes and the content-derived effective-index/worktree state | An otherwise eligible Gate-B pass cannot establish closure, and because eligibility also forbids suspension it can only continue indefinitely; alternatively a reader may silently ignore the universal condition and close under a different ordering | Scope the artifact/request equality and the two repository cases explicitly to Gate A, and cite Gate B's existing range, re-review, and content-derived validity rules as that gate's own closure inputs without adding a new fingerprint +MAJOR | high | §A "Second" and "Third" branches | The suspension branch is evaluated before the source-prescribed-block branch, yet the latter says its case is "settled first"; a pass can carry a scope or health suspension while a profile or governing header becomes unresolvable during the pass, so one branch suspends under the composition rule while the other says the source block is neither a suspension nor a continue and forbids another pass | After the suspension answers arrive, the composition rule can authorize the next pass while the profile or header source still forbids it, and the required source-block reason can be omitted from the surface | Put the source-prescribed-block check before suspension evaluation, or explicitly compose it with the surface while making clear that it blocks every next pass until its own source condition is repaired +MAJOR | high | §A "assigned fix set" closing gate | The text says a closing-time assigned-fix-set change is governed by that set's own source, but §B only defines the set before the pass and §A itself only says that a later broadening is read by the next pass; no source says what a narrowing, another kind of set change, or a change followed by restoration does at the closing act | A final pass can be accepted after the set changed without reviewing the set on which closure now relies, despite the settled requirement that this gate answers a change rather than final-value inequality | Put the complete change-sensitive closing rule at the assigned-fix-set source and leave §A as a pure citation, covering either direction and restoration without introducing a record or history mechanism +MAJOR | medium | design §2 "closing-time sameness tests" | The design still says the block-owned closing-time tests are plural, borrow no rule, and have no other source, while target §A now says only the artifact equality is block-owned and the profile and cited-set gates retain their own source rules | The plan receives incompatible authority boundaries and can rebuild the joint equality rule pass 25 deliberately dissolved or duplicate the source-owned gates in the ordering | Align the design with the target's singular artifact equality and separately identify the commit-carry duty, while stating that every other gate remains source-owned +MAJOR | high | §F "human-exception scope sentence" | The replacement says "The answer a suspension asks for" is continue or stop, but §A defines the scope stop as a suspension whose membership component asks accept or decline and whose question component asks for the user's decision; only the two health readings ask continue or stop | A scope-stopped finding can be given the wrong answer vocabulary, leaving its hold undischargeable or bypassing the accept/decline decision that determines fix-set membership | Scope continue or stop to the health suspensions and point scope-stop answers to the ordering without restating them +MAJOR | medium | §H "Surfacing does not close the cycle" | The replacement summarizes the resolve rule as meaning every surfaced finding is repaired or validly dismissed and closes with the unqualified rationale that every Blocker and Major resolves, but §E deliberately limits that duty to findings in the assigned fix set | An out-of-set Blocker or Major surfaced by a membership or question stop, including one declined and kept outside, can be treated as still owing repair or dismissal, contradicting D5 and the source-owned severity rule | Retain a pure pointer to Mechanics · Severity and qualify the rationale to in-set Blocker/Major findings, removing the unscoped repair-or-dismissal restatement +MINOR | high | §F "seven standing sentences" | §A says recognizing a finding across a lost session needs "a record these rules do not ship", but standing §5 explicitly ships the optional `-dispositions.md` companion and says it makes a dismissal durable so its verdict and reason outlive the session; this is an eighth live sentence the target currently makes false | Readers are told both that a durability record exists and that none ships, obscuring the actual residual, which is the lack of a mandatory identity, sameness, and recovery rule rather than the absence of any record | Narrow the target claim to say that the optional note is non-authoritative and that no mandatory recognition or recovery semantics ship; do not add a new record +MINOR | high | §A ownership boundary versus §§B, F, and H | §A says it owns what a surfaced-finding hold is and what discharges it and that source-owned preconditions are cited rather than restated, yet §B directly says the membership answer ends the membership hold, §H says the hold stands until its answers are given, §F repeats the Gate-A human-exception destination, and §A repeats the profile, cited-set, evidence, and resolve outcomes it says their sources own | The target preserves several parallel authorities for the decisions this consolidation exists to state once, so a later repair can update one copy and leave another contradictory | Keep each rule at its declared owner and turn the other occurrences into references that do not repeat a discharge, destination, or source-owned transition +MINOR | high | §§A and C "That third condition is what makes a plateau rather than a finish" | §C leaves this rationale as a standalone sentence while §A reproduces it again as the opening of the supposedly word-for-word moved precedence sentence | One passage is assigned to two proposed sections, contrary to the target's single-section rule, and the duplicated rationale can drift independently | Keep the rationale beside the clearly-stuck predicate in §C and move only the precedence clause, or remove the duplicate from §C +MINOR | high | §A "No other pass outcome makes a cycle eligible" | Its exhaustive reason says every rejected pass leaves a repair, hold, question, or floor shortfall, but an at-floor pass that re-raises an in-set effective-Major finding already validly dismissed is immutably unclean while its resolve duty is discharged and, before the three clearly-stuck conditions mature, has no hold or question | The operative predicate rejects the pass, but its rationale can make an agent demand a second dismissal or repair, contradicting the idempotent-dismissal rule | Add fixed pass-level uncleanliness as the remaining reason a pass can be ineligible after cycle-level duties and answers are settled +MINOR | high | §E "curve must stay derivable from the findings files alone" | The standing curve also records cycle identity, pass ranges, and model identifiers that findings files do not contain, and it explicitly permits `?` when a numeric count cannot be recovered from those files | The replacement overstates both the curve's provenance and its checkability, conflicting with the surviving unknown-count and model-source rules | Narrow the claim to the three numeric series when their validated findings files remain available +MINOR | high | §G "curve without a cycle field cannot be told from another cycle's" | Standing Mechanics says missing attribution is a limitation rather than a disqualification because cycle kind and surrounding context can sometimes distinguish records, and that a nonce only means a later reader usually does not need that context | The one-contract rationale upgrades a probabilistic, context-dependent attribution aid into an impossibility claim, repeating the mechanism-overclaim class AGENTS.md warns against | Say the curves cannot be reliably distinguished in every multi-cycle context without the cycle field +MINOR | medium | §G "any record this section obliges" | Once installed, §G sits inside `### Mechanics (reference)` under §5, so "this section" can mean Mechanics or the whole Cross-Model Review section even though those scopes need not oblige the same records | Two downstream readers can derive different one-contract memberships while both apply the promised text-only test | Name §5 or Mechanics explicitly, whichever scope is intended +NIT | high | §A "Below the floor the pass suspends, clean completion having not closed it" | The section defines clean completion to include eligibility, so a nonzero clean pass below the floor is not a clean completion rather than a completion that failed to close | The phrase gives clean completion two meanings inside the ordering and makes the precedence explanation harder to apply literally | Say the clean pass failed eligibility because it was below the floor +END OF FINDINGS (14 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index 62e5dc1..8461ae8 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -52,7 +52,53 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 23 | d6052fe | 13→**9** | 1→**0** | 5→**7** | yes | **zero tells, no mandatory stop.** But findings 1, 2 and 3 are all in the closing-act paragraph the rollback rewrote the day before, and the paragraph has now produced a Blocker or Major at passes 20, 21, 22 and 23; session 01a094e9-76bd-7521-9c93-c11491c14202 | | 24 | 6913c35 | 9→**8** | 0→**1** | 7→**3** | yes | **B+M 4 and Majors 3, both the lowest of the cycle.** One tell. The Blocker is new ground: the closing act reads `HEAD` and the working artifact and never the **index**, which is what `git commit` commits — `AGENTS.md:93`'s own documented failure class; session 01a0950c-f003-79d3-985f-73e2885c9621 | | 25 | 31bbe8f | 8→**10** | 1→**0** | 3→**2** | yes | **B+M 2 and Majors 2, both by far the lowest of the cycle; zero Blockers.** Both Majors are new ground. **Five of the ten are re-raised collected Minors**, one on its fourth appearance; session 01a09522-65bc-7691-8adc-fb26e330810e | -| 26 | — | — | — | — | not run | **awaiting Daniel's decision**; no automatic follow-up round was authorised | +| 26 | 8b8e146 | 10→**14** | 0→**1** | 2→**5** | yes | **TWO-TELL STOP — mandatory, surfaced to Daniel.** **All six Blocker/Majors trace to repairs made in passes 23, 24 and 25**, each nameable. All 14 held open; session 01a09539-5b8c-7293-af82-cc5f3bd2ba48 | +| 27 | — | — | — | — | not run | blocked on the two-tell answer | + +## Pass-26 three-line report — MANDATORY TWO-TELL STOP + +**Floor line:** derived floor **3**; risk **high**, security **none**; read fresh from +`docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md`. One cited story, level 2. + +- **Trend:** findings …, 8, 10, **14**. Blockers …, 1, 0, **1**. Majors …, 3, 2, **5**. + Blocker+Major …, 7, 4, 2, **6** — the pass-25 low of 2 tripled. +- **Cluster (pass 26):** product 10 of 14; prose about the design or this artifact 2 (4, 9); the + instrument 0. +- **require↔withdraw:** none. Finding 4 asks the **design** to follow the target text's singular + artifact equality, which is alignment rather than a demand for removed text. + +**Tells: two of five — the threshold. Stop-and-surface is mandatory, not discretionary.** The +finding count rose 10 → 14 and the Blocker count failed to fall, 0 → 1. + +**Every Blocker and Major comes out of a repair made in the last three rounds. All six, nameable:** + +| # | Sev | What it says | Repair it comes from | +|---|---|---|---| +| 1 | BLOCKER | the artifact/request equality is written cycle-generally, but **Gate B passes a git range, not artifact text**, so it has no value to test — leaving an eligible Gate-B pass unable to close and, being eligible, unable to suspend | pass 25 (the per-condition gate reading) | +| 2 | MAJOR | the source-block branch says its case is "settled first" while the **suspension branch is evaluated before it** | pass 24 (branch order) | +| 3 | MAJOR | each gate was sent to "its own source", but the **assigned fix set has no closing-time source rule** to be sent to | pass 25 (same repair) | +| 4 | MAJOR | design §2 still calls the block-owned closing tests plural and source-free | pass 25 (§A's neighbour was fixed, the design's was not) | +| 5 | MAJOR | "the answer a suspension asks for" is given as continue-or-stop, but a **membership stop asks accept or decline** | pass 23 (§F entry 4) | +| 6 | MAJOR | §H restates the resolve duty **unscoped**, while §E limits it to the assigned fix set | pass 24 (§H surfacing) | + +**Verified before this report, not inferred:** the equality condition and the closing-time gate list +are written cycle-generally (target 117–120 and 149–153, the list naming the evidence entry as the +Gate-B-only member and so implying the rest covers both); §F's sentence at target 540 against §A's +membership answer at target 259; and `CLAUDE.md:354`, which ships the optional +`-dispositions.md` companion — an **eighth** standing sentence the target contradicts by +saying a cross-session record is one "these rules do not ship" (finding 7). + +**One shape runs through findings 2, 3, 4, 5 and 6, and the Minors have been naming it for five +passes.** Each is **a rule stated in §A that also lives at a source, or a source sentence §A's +wording contradicts**. That is the two-authorities mechanism pass 11's finding 14 named. Finding 8 +has re-raised it as a collected Minor at passes 23, 24, 25 and 26; finding 9 is on its **fifth** +appearance. The Majors were repaired one at a time while the Minor pointing at their common cause +was collected each round under the standing rule. + +**The clearly-stuck reading stands at two of three** — regeneration is nameable for all six, and +B+M has never reached zero in twenty-six passes — but **coverage is not affirmable**: finding 1 +reaches Gate-B scoping that no earlier pass had read. Two of three is not the exit, and the two-tell +threshold stands alone anyway. ## Pass-25 dispositions — two Majors, one mis-scoped hook sentence, two corrected numbers From 5446a458c7d4d97405ec54b00fc2379711ec681b Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sat, 12 Sep 2026 14:12:15 +0200 Subject: [PATCH 050/181] docs(context): park Gate-A cycle awsf1ec771 open at pass 26 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Daniel's answer to the pass-26 two-tell stop is B: park the cycle open. Gate A stays open, the target text stays inactive, all 14 findings stay open and unrepaired. No further repair round and no pass 27. A restart needs an explicit later decision. Two overstatements in my pass-26 report are corrected, both traced with git log -S rather than argued. The claim that all six Blocker/Majors came out of the last three repair rounds is false. Three did, from rounds 24, 24 and 25. Three are older defects rediscovered at pass 26, from rounds 19, 20 and 22, and none came from round 23. Repair damage and rediscovered defects are separate categories and the record now keeps them apart. The claim that option C would remove five findings as a class is not established. Evaluation order is something §A must decide, the missing closing rule for a fix-set change cannot be replaced by citing a rule that does not exist, and the design/target divergence is its own problem. Precise citation answers two of the six. "§A cites throughout" would itself be a false rule, so C is a partial diagnosis rather than a bounded approach, and no analysis of it is commissioned. The Minor rule is also not the culprit: it forbids Minors their own rounds and does not forbid repairing a demonstrated common cause of several Majors. Repeated Minor hints do not establish that clearing them would have prevented the heavy defects. Nothing this cycle wrote governs anything. CLAUDE.md and the workflow-init template are untouched, so the Gate-B scoping Blocker sits in text that ships to nobody. The 26 passes and their findings are kept. What is withdrawn is the funding of this working method, not the work, and the prior investment is explicitly not a reason to continue. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-resume.md | 76 ++++++++++++++----- 1 file changed, 59 insertions(+), 17 deletions(-) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index 8461ae8..592d236 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -70,16 +70,24 @@ Nothing depends on it; the pass files and the repo are authoritative where this **Tells: two of five — the threshold. Stop-and-surface is mandatory, not discretionary.** The finding count rose 10 → 14 and the Blocker count failed to fall, 0 → 1. -**Every Blocker and Major comes out of a repair made in the last three rounds. All six, nameable:** - -| # | Sev | What it says | Repair it comes from | -|---|---|---|---| -| 1 | BLOCKER | the artifact/request equality is written cycle-generally, but **Gate B passes a git range, not artifact text**, so it has no value to test — leaving an eligible Gate-B pass unable to close and, being eligible, unable to suspend | pass 25 (the per-condition gate reading) | -| 2 | MAJOR | the source-block branch says its case is "settled first" while the **suspension branch is evaluated before it** | pass 24 (branch order) | -| 3 | MAJOR | each gate was sent to "its own source", but the **assigned fix set has no closing-time source rule** to be sent to | pass 25 (same repair) | -| 4 | MAJOR | design §2 still calls the block-owned closing tests plural and source-free | pass 25 (§A's neighbour was fixed, the design's was not) | -| 5 | MAJOR | "the answer a suspension asks for" is given as continue-or-stop, but a **membership stop asks accept or decline** | pass 23 (§F entry 4) | -| 6 | MAJOR | §H restates the resolve duty **unscoped**, while §E limits it to the assigned fix set | pass 24 (§H surfacing) | +**Provenance of the six Blocker/Majors — CORRECTED 2026-09-12 after the reviewer challenged it.** +The first version of this table said all six came out of the last three repair rounds. **That was +wrong**, and the two categories must stay apart: *damage made by a recent repair* and *an older +defect only now rediscovered*. Traced with `git log -S` against each sentence: + +| # | Sev | What it says | Entered at | Round | +|---|---|---|---|---| +| 1 | BLOCKER | the artifact/request equality is written cycle-generally, but **Gate B passes a git range, not artifact text**, so it has no value to test — leaving an eligible Gate-B pass unable to close and, being eligible, unable to suspend | `31bbe8f` | **24** | +| 2 | MAJOR | the source-block branch says its case is "settled first" while the **suspension branch is evaluated before it** | `31bbe8f` | **24** | +| 3 | MAJOR | each gate was sent to "its own source", but the **assigned fix set has no closing-time source rule** to be sent to | `8b8e146` | **25** | +| 4 | MAJOR | design §2 still calls the block-owned closing tests plural and source-free | `8dc22fb` | **20** | +| 5 | MAJOR | "the answer a suspension asks for" is given as continue-or-stop, but a **membership stop asks accept or decline** | `5174d7a` | **19** | +| 6 | MAJOR | §H restates the resolve duty **unscoped**, while §E limits it to the assigned fix set | (reworked) | **22** | + +**So: three of six from the last three rounds (24, 24, 25), and three rediscovered from rounds 19, +20 and 22. None from round 23.** The loop is still producing repair damage, and it is also still +finding old defects for the first time at pass 26 — which is the same fact that makes the +clearly-stuck coverage condition unaffirmable. **Verified before this report, not inferred:** the equality condition and the closing-time gate list are written cycle-generally (target 117–120 and 149–153, the list naming the evidence entry as the @@ -88,12 +96,26 @@ membership answer at target 259; and `CLAUDE.md:354`, which ships the optional `-dispositions.md` companion — an **eighth** standing sentence the target contradicts by saying a cross-session record is one "these rules do not ship" (finding 7). -**One shape runs through findings 2, 3, 4, 5 and 6, and the Minors have been naming it for five -passes.** Each is **a rule stated in §A that also lives at a source, or a source sentence §A's -wording contradicts**. That is the two-authorities mechanism pass 11's finding 14 named. Finding 8 -has re-raised it as a collected Minor at passes 23, 24, 25 and 26; finding 9 is on its **fifth** -appearance. The Majors were repaired one at a time while the Minor pointing at their common cause -was collected each round under the standing rule. +**The two-authorities reading explains part of the six, and my pass-26 report overstated it.** +It was written as "one shape runs through findings 2, 3, 4, 5 and 6" and as "option C removes five +findings as a class". **Neither is established**, and the reviewer's breakdown is the accurate one: + +| # | Would precise citation fix it? | +|---|---| +| 2 | **No** — evaluation order is something §A itself must decide. | +| 3 | **No** — the closing rule for a fix-set change **does not exist anywhere**, and a missing rule cannot be replaced by citing one. | +| 4 | **No** — it is a divergence between design and target text. | +| 5, 6 | **Yes** — precise references, and removing a wrong restatement, do answer these two. | + +**"§A cites throughout" would itself be a false rule**, since §A must keep defining the evaluation +order and its own closure decisions. **Option C is therefore not a bounded approach yet**, only a +partial diagnosis. + +**And the Minor rule is not the culprit either.** It forbids Minors their own repair rounds; it does +**not** forbid repairing a demonstrated common cause of several Majors. That a Minor was re-raised +at passes 23–26 (finding 8) and another at five consecutive passes (finding 9) does **not** +establish that clearing them would have prevented the heavy defects. Recorded because the pass-26 +report implied it did. **The clearly-stuck reading stands at two of three** — regeneration is nameable for all six, and B+M has never reached zero in twenty-six passes — but **coverage is not affirmable**: finding 1 @@ -1257,7 +1279,26 @@ here" blocks, which must stay in or the reviewer re-raises deferred material. survives a context clear on this machine and not a fresh clone. This resume note and every pass file are tracked from 2026-09-10 and do survive. -## STATE AT HANDOFF — 2026-09-11 midday +## PARKED 2026-09-12 — cycle open, not running. READ THIS FIRST. + +**Daniel's answer to the pass-26 two-tell stop: B — park the cycle open.** Gate A stays open, the +target text stays **inactive**, all 14 pass-26 findings stay open and unrepaired. **No further +repair round and no pass 27.** Restarted only by an explicit later decision. + +- **HEAD `56d8f5a`** on `loop-rule-consolidation`, tree clean, **26 passes run, none clean**. +- **Nothing is closed and nothing is claimed closed.** Surfacing credits no pass as clean; + `CLAUDE.md` and `plugins/dev-workflow/commands/workflow-init.md` are untouched, so **no rule this + cycle wrote governs anything**, and the Gate-B scoping Blocker sits in text that ships to nobody. +- **The 26 passes and their findings are kept**, not discarded. What is withdrawn is the funding of + this working method, not the work. +- **Option C is not decided and no analysis of it is commissioned.** Its claimed common cause and + effect are only partly substantiated — see the pass-26 section. A later restart needs **a + concretely bounded alternative to the repair-round strategy**, and producing one is not an open + task. +- **The prior investment is not a reason to continue.** Recorded because this cycle's reports twice + reached for it. + +## STATE AT HANDOFF — 2026-09-11 midday (superseded by the parking entry above) - **HEAD `ef9a504`** on `loop-rule-consolidation`, tree clean, 18 passes run, none clean. - **THE ARTIFACT CHANGED AT PASS 18.** Passes 1–17 reviewed the design spec; from pass 18 the @@ -1331,6 +1372,7 @@ nothing checks them. They are candidates captured in | pass 14 scope stop | **B — accept, bounded.** Extend the existing coherence *instruction* to the closure block and its coupled edits (now the one-contract paragraph, target text §G). No checker, no new mechanism, no record-durability work. | | pass 15 | **Bounded rollback**, not another twelve-repair round: totals, item-number lists and enumerated precondition lists removed; behaviour stays decided in the design. | | pass 17 two-tell stop | **Interrupt the repair mode.** Produce a non-active **target-text** version; Gate A stays open, no transition to Gate B, no clearly-stuck close claimed. | +| pass 26 two-tell stop | **B — park the cycle open.** Gate A open, target text inactive, all findings kept. No repair round, no pass 27. Option C undecided and uncommissioned; a restart needs a bounded alternative to the repair-round strategy. The 26 passes are explicitly **not** a reason to continue. | | pass 22 two-tell stop | **Continue, bounded.** Rollback of the source-revision line confirmed. One repair round for the seven pass-23 Majors under three precisions — the closing cases decide only *how* to close and never override an intervening artifact change; the stop-and-surface precondition keeps its blocking effect stated positively; §G's membership criterion admits no blanket exemption by section. Then pass 24, then a fresh decision. Recorded late: the answer was given 2026-09-12 after the agent found no confirmation on record. | ## For the execution phase, not needed yet From d971ae7dba518c6cedbf53b8ea14e5b134c56e25 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sun, 13 Sep 2026 10:23:40 +0200 Subject: [PATCH 051/181] docs(specs): split the ordering into gate-general and per-gate closure MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Applies the counter-draft Daniel authorized on 2026-09-13, answering pass 26's six Blocker/Majors. Target text §A becomes three paragraphs: A1 the gate-general ordering, A2 Gate A's own content condition and closing act, A3 Gate B's. The artifact/request equality was stated cycle-generally while its explanation and its two repository cases were Gate A's alone, leaving an eligible Gate-B pass with no value to test and, being eligible, unable to suspend (pass 26 finding 1). Shared closure preconditions stay shared and are named in A1 — the floor, the resolve duty, the hold, no-clean-credit, and the profile, cited-set and assigned-fix-set gates. Neither gate paragraph is an inventory of what its gate requires, so the fix-set rule reaches both. Other repairs from pass 26: - the source-block branch is read first and explicitly silences no suspension raised by the same pass (finding 2) - §B gains the closing-time rule for a change to the assigned fix set, which existed nowhere. Daniel's decision of 2026-09-12: a change costs at least one further pass, in either direction and whether or not later undone; the window opens where the set is fixed for the pass (finding 3) - design §2 stops calling the block-owned closing tests plural; the exception it named no longer exists (finding 4) - §F point 4 refers to the ordering's answer rules instead of naming a vocabulary that is only the health readings' (finding 5) - §H scopes the resolve restatement to the assigned fix set (finding 6) Branches are named rather than numbered, and every cross-reference names the branch: five references had gone stale, two of them outside §A. Collected, not repaired this round: findings 7 and 10 answered in passing while the paragraphs were reset; finding 8 remains only partly answered — A1 still carries the three sources' change consequence and §H still describes the hold's discharge. Cycle awsf1ec771 stays open; nothing is installed. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- ...26-09-10-loop-rule-consolidation-design.md | 25 +- ...-10-loop-rule-consolidation-target-text.md | 567 ++++++++++-------- 2 files changed, 332 insertions(+), 260 deletions(-) diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index 59b8a2b..b554da5 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -62,9 +62,21 @@ boundary is stated once, in the target text's §A opening, and is deliberately n two copies of it are what let them drift, which is pass 20 finding 3. What this spec records is the decision behind it: the block defines no trigger and no severity rule of its own, and every closure precondition **that has a source of its own** keeps its one definition there, changed **at that -source** where it had to change to agree with the ordering. **The closing-time sameness tests are -the exception in substance and not in principle**: they borrow no rule and have no other source, -being part of the closure decision the block owns. +source** where it had to change to agree with the ordering. + +**The ordering is split into three paragraphs on Daniel's decision of 2026-09-12, and that split is +the answer to pass 26's Blocker.** The ordering had stated one closure condition — the artifact's +equality with the text sent to the reviewer — cycle-generally, while its explanation and its two +repository cases were Gate-A's alone; Gate B passes a git range rather than artifact text, so an +otherwise eligible Gate-B pass had no value with which to evaluate it and, being eligible, could +not suspend either. **What is gate-general stays gate-general and what differs goes to the gate**: +the ordering keeps the evaluation of a pass, the classification of the duties, and the closure +conditions both gates share — the floor, the resolve duty, the hold, no-clean-credit, and the +profile, cited-set and assigned-fix-set gates; each gate states only **its own content condition +and its own closing act**, and neither gate's paragraph is an inventory of what that gate requires. +**No closing-time test is an exception to the block's citation rule any more**, because the one +that was is now Gate A's own condition stated at Gate A's paragraph. **The split buys ownership and +not brevity** — the three paragraphs together run slightly longer than the single block did. **The closure-ordering block is an addition beside the source edits**, not one of them. §4 lists the edits; **no total is stated here or there**, because the unit — one contiguous replacement at @@ -148,7 +160,7 @@ here. This table says what happens to each inventoried passage, so the map stays | Passage | This change | Target text | |---|---|---| | (a) the floor paragraphs | edited — closure sentences trimmed to a pointer, the no-restating prohibition scoped, the per-pass fix command pointed at Mechanics | §H | -| (b) what a loop absorbs | edited — owns both triggers and the fix set, so every qualification the ordering needs is made **here**, which is what keeps one definition per rule | §B | +| (b) what a loop absorbs | edited — owns both triggers and the fix set, so every qualification the ordering needs is made **here**, which is what keeps one definition per rule; **gains the closing-time rule for a change to the set**, which no inventoried condition carried because none existed (Daniel, 2026-09-12: a change costs at least one further pass, in either direction and whether or not it is later undone, the window opening where the set is fixed for the pass) | §B | | (c) recognizing clearly stuck | edited — keeps its three-condition reading, stops carrying evaluation order; its precedence sentence moves into the block unchanged, and the third condition itself is widened at this source to admit a re-raised validated dismissal | §C | | (d) from pass 4 onward | **no longer edited.** The unavailable-history block moved to the successor with **D10** | — | | (e) the five tells | edited — the threshold is read after clean completion, and a pointer says what its answer does | §D | @@ -222,8 +234,9 @@ complete or that the plan's fragments discriminate.** The enumeration moved to w exists; the completeness claim did not, because nothing supports it. **The counterfactual splits, and stating it as one understates what is owed.** For the **ordering -block** it is **ABSENT and is claimed as absent**: the parent carries no such block, so no old -wording of it can be shown to disappear and presence alone is the check. For **every replacement** +and the two gate-closure paragraphs beside it** it is **ABSENT and is claimed as absent**: the +parent carries none of the three, so no old wording of them can be shown to disappear and presence +alone is the check. For **every replacement** the parent carries the old wording and the change removes it, so each owes the old-wording-gone half of its pair. **The obligation reaches every passage the target text marks REPLACED, and no list of them is kept here** — a second enumeration beside the markers is the diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index 6e2ec70..dcdd838 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -37,193 +37,161 @@ per-condition disposition, parity diff and verification fragments (the plan, aga --- -## A. The closure ordering — NEW - -Sits immediately before "**What a loop absorbs, and what stops it**" in both copies. - -``` -**How a cycle ends — one ordering, stated here and referenced everywhere else.** A pass is read -in a fixed order, because every rule bearing on one decision — may this cycle close — otherwise -qualifies the others and the ranking survives only in a reader's head. **This paragraph owns how a -pass is evaluated and what follows from that**: the evaluation order, **what each predicate is -read from**, closure and its eligibility, the hold a surfaced finding places and what discharges -it, the composition of several suspensions, the pairs that cannot co-occur, and — as a stated -exception, because precedence is evaluation order — the clearly-stuck precedence sentence quoted -into it below. **No number is put on that list.** Whether the read model counts as part of the -evaluation order or as a thing beside it is a question of wording, not of authority, and a count -turns that wording into an argument about whether the paragraph has overrun its own boundary. **Everything else it names it cites**: -the scope triggers, the assigned fix set, every severity rule, and every closure precondition -**that has a source of its own** keep their one definition in the paragraph that owns them, and a -reader who finds one of *those* defined here has found a defect. **The artifact's closing-time -equality test below is defined here and is not an exception to that**: it borrows no rule and has -no other source, being part of the closure decision this paragraph owns rather than a precondition -stated elsewhere and read from here. **The other closing-time gates are not like it** — the -profile's and the cited set's are their own sources' rules, cited from here and not restated, which -is the sentence before this one applied rather than excepted. +## A. The closure ordering and the two gates' closure — NEW + +Three paragraphs, in this order, sitting immediately before "**What a loop absorbs, and what stops +it**" in both copies. **A1 is the ordering and is gate-general; A2 and A3 are each gate's own +closure.** The split is the answer to pass 26's Blocker: the ordering stated one closure condition +cycle-generally while its explanation and its two cases were Gate-A's, leaving an eligible Gate-B +pass with no value to test. + +### A1 — the ordering + +``` +**How a cycle ends — one ordering, stated here and referenced everywhere else.** A pass is read in +a fixed order, because every rule bearing on one decision — may this cycle close — otherwise +qualifies the others and the ranking survives only in a reader's head. + +**What this paragraph owns.** It owns **how a pass is evaluated and what follows from that**: the +evaluation order, what each predicate is read from, the eligibility test, the hold a surfaced +finding places and what discharges it, the composition of several suspensions, the pairs that +cannot co-occur, and — as a stated exception, because precedence is evaluation order — the +clearly-stuck precedence clause quoted into it below. **No number is put on that list.** Whether +the read model counts as part of the evaluation order or as a thing beside it is a question of +wording, not of authority. It also owns the **classification** of the standing duties below, and +**those duties bind both gates alike**: the derived floor and the Blocker/Major-resolve duty keep +their definitions at their own sources and are read here for every cycle whatever its gate, while +the hold and no-clean-credit are defined here, being properties of the evaluation itself. + +**What belongs to a gate is what the two gates do differently: the content condition each gate's +closure requires, and the closing act.** Each is stated in that gate's own paragraph below and read +from there, and **neither gate's paragraph restates the conditions this one gives every cycle**, so +neither is a complete inventory on its own. **Everything else this paragraph names it cites**: the +scope triggers, the assigned fix set, every severity rule and every closure precondition with a +source of its own keep their one definition in the paragraph that owns them, and a reader who finds +one of *those* defined here has found a defect. + +**The conditions every cycle has, whatever its gate.** The duties classified below; and the +**profile**, the **cited set** and the **assigned fix set**, each gating as its own source says and +cited here without restatement. **Those three sources answer a change and not a differing value**, +which is why one of them changed and then undone still costs a pass — the contrast that matters +against a gate's own content condition, which may be a comparison of current values and says so +where it is stated. **What a pass is read from.** Every finding-derived predicate reads the validated findings file -**or files** of the logical pass as **the concatenation of their finding lines after each file -has been validated separately** — a `full` Gate-B pass has two, one branch alone is already an +**or files** of the logical pass as **the concatenation of their finding lines after each file has +been validated separately** — a `full` Gate-B pass has two, one branch alone is already an incomplete pass, each file's terminator is not a finding line, and a branch whose body is `NO FINDINGS` contributes an empty sequence rather than a line. **Which severity field each predicate reads is settled in Mechanics · Severity**, which is where that split lives and is not -repeated here. Beyond the findings, closure reads the **derived floor**, the final-acceptance -preconditions the floor section states, the **resolve duty's standing over this cycle**, and any -**hold still standing**; **in a Gate-B cycle it also reads the evidence entry's revalidation -rule** — a changed entry meaning the clean pass no longer covers what is being committed — which -is a Gate-B precondition because the entry is about the diff being committed and a Gate-A cycle -produces none. The scope triggers read the **current assigned fix set** as the absorb paragraph -defines it, and the answers already given; the clearly-stuck reading adds its own coverage -judgement. **A line in one branch file and a line in the other are distinct findings for holds -and answers**, so a `full` pass asks twice rather than risk resuming over one it never asked -about. **Within one running cycle an answer binds to the finding or question as the pass that -raised it recorded them** — which is what an agent running the cycle can do with nothing written -down. Recognising the same finding or question across a lost session needs a record these rules -do not ship. - -**First, clean completion, and eligibility is its own test.** §5 uses *clean* in two senses and -now says which is which: a **clean findings file** is the `NO FINDINGS` signal the protocol -defines, and a **clean pass** is the predicate here, read on the logical pass with every required -branch file combined, so one branch's clean file never establishes a clean pass. **A pass is -clean** when its findings carry no in-set Blocker or Major at effective severity and **no +repeated here. Beyond the findings, closure reads the **derived floor**, the **resolve duty's +standing over this cycle**, any **hold still standing**, and **every closure condition this cycle +has** — those above and this cycle's gate's. The scope triggers read the **current assigned fix +set** as the absorb paragraph defines it, and the answers already given; the clearly-stuck reading +adds its own coverage judgement. **A line in one branch file and a line in the other are distinct +findings for holds and answers**, so a `full` pass asks twice rather than risk resuming over one it +never asked about. **Within one running cycle an answer binds to the finding or question as the +pass that raised it recorded them** — which is what an agent running the cycle can do with nothing +written down. Recognising the same finding or question **across a lost session** has no mandatory +identity, sameness or recovery rule in these rules; the optional `-dispositions.md` note is +advisory and authoritative for nothing. + +**The four branches are named, and a cross-reference anywhere in this section names the branch +rather than its position**, so reordering them breaks no reference. + +**The source-block branch, read first.** **Where any unmet closure condition's own source +prescribes stop-and-surface** — a profile present but unresolvable, governing headers that +disagree, a `Story:` header that cannot be read — **the cycle stays open, that source decides what +must be repaired or answered, and no further pass runs while its block stands.** It is neither a +suspension nor a continue and needs no name and no procedure of its own: the source rule carries +both, and this ordering's part is to send the reader there rather than to run a pass over a cycle +another rule has stopped. **It is read first and it silences nothing.** Where the same pass also +carries a suspension, that suspension is surfaced with its reasons and its questions exactly as the +suspension branch requires and its answers are collected; what the block adds is that **no next +pass runs until its own source condition is repaired**, whatever those answers were. + +**Then the clean-completion branch, and eligibility is its own test.** §5 uses *clean* in two +senses and now says which is which: a **clean findings file** is the `NO FINDINGS` signal the +protocol defines, and a **clean pass** is the predicate here, read on the logical pass with every +required branch file combined, so one branch's clean file never establishes a clean pass. **A pass +is clean** when its findings carry no in-set Blocker or Major at effective severity and **no scope-stop trigger** — the two the absorb paragraph defines, read there and not redefined here, -each already carrying the qualification **an answer given before that pass ran** puts on it. -**A pass's cleanliness is settled on what it found and on the answers standing when it ran**, and -a later answer never rewrites it: an answer discharges the holds it was asked for and leaves the -pass that raised them exactly as clean or unclean as it was, which is the same fact the duties -paragraph states of no-clean-credit. Reading a later answer back onto an earlier pass would let a -cycle close on a pass that was surfaced, answered and never re-run. A pass with -**zero** findings is clean whatever the floor, because a floor buys further looks at an artifact -that keeps yielding findings, and one yielding none has already given what those looks were for; -don't manufacture findings to pad. +each already carrying the qualification **an answer given before that pass ran** puts on it. **A +pass's cleanliness is settled on what it found and on the answers standing when it ran**, and a +later answer never rewrites it: an answer discharges the holds it was asked for and leaves the pass +that raised them exactly as clean or unclean as it was, which is the same fact the duties paragraph +states of no-clean-credit. Reading a later answer back onto an earlier pass would let a cycle close +on a pass that was surfaced, answered and never re-run. A pass with **zero** findings is clean +whatever the floor, because a floor buys further looks at an artifact that keeps yielding findings, +and one yielding none has already given what those looks were for; don't manufacture findings to +pad. **Eligibility is exactly this and nothing more: a clean pass at or above the derived floor, or a zero-finding pass.** It is a property of the pass. **Closure is eligibility plus every closure -precondition holding plus the cycle's closing act**, and the preconditions are properties of the -*cycle*, not of the pass — so an eligible pass whose preconditions are unmet closes nothing, and -is not thereby made unclean. Keeping the two apart is what lets the order decide the case where -they disagree, which the branches below do. - -**The closing act, by cycle kind.** A **Gate-B** cycle closes with the closing amend -Mechanics · Finishing the cycle describes. A **Gate-A** cycle has no WIP snapshot to replace, so -it closes by **writing the closing commit — or the closing message of one that already exists — -carrying the records this cycle already owes**: its provenance line and its per-pass curve, in the -forms Mechanics fixes, **neither of them altered**. **Closure introduces no new kind of record, and -it excuses none**: every other record this cycle owes, a human-exception record among them, is -owed and written exactly as before. Writing that body once every closure condition holds is the -closing act, and nothing before it closes anything. - -**The order is fixed.** **First** every closure condition is established, among them that **the -artifact as it now stands is identical to the text that went into the final pass's review -request**. **Only then** is the closing act performed, and the shape it takes is read off the -repository as it stands. **That reading decides how a cycle closes, never whether it may**: where -the artifact differs from that request text the condition has failed, the pass is not final, and -another is owed. - -**The condition is current equality, and deliberately nothing more.** It does **not** say the -artifact went untouched in between: text edited and then restored byte for byte satisfies it. That -is a decision and not an oversight — a content comparison cannot tell those two states apart, and a -condition nobody can check is a condition nobody applies. **Every other closure condition stands -independently and is not relaxed by this one.** It likewise says nothing about **what the reviewer -consumed**: Gate A hands the reviewer text rather than a git range, so no part of this act is -offered as evidence of the review payload. - -**The content must survive into the commit the act produces, and carrying it there is part of the -act.** The equality above is read on the artifact, while **the commit is written from the effective -index** — so an act that does not carry that same content through has checked the condition without -performing it. Both cases below owe the same thing: **the commit that closes the cycle contains, at -the artifact path, exactly the text the condition was read against.** The safe command sequence and -whatever demonstrates it belong to the plan; **the duty belongs here**, and it needs no new -fingerprint and no new record. - -**Two cases, told apart by the repository's current state rather than by which commit introduced -what; the safe git sequence for each belongs to the plan.** -- **`HEAD` already carries that text at the artifact path.** Close by **amending `HEAD`'s message** - to add the records, leaving that path as it stands. -- **`HEAD` does not, the reviewed text being still uncommitted.** **Commit it unchanged** and close - in that commit. This case had no answer before: an eligible pass over repairs nobody had - committed could neither close nor suspend. - -**A human-exception record this cycle owes goes in the commit its closing act uses**, so the two -never land in different places. **No new revision of the artifact is made to close a Gate-A -cycle**, a new revision being one no pass has run against — and committing already-reviewed text -that was never committed is not one. **Nothing is closed before that act, and what still gates it in -between is read condition by condition rather than as one comparison.** The **artifact** gates on -the equality above and on nothing stronger. The **profile** and the **cited set** gate as their own -sources say, and those sources answer a **change** and not a differing value: a profile change -costs at least one further pass in either direction, and a cited-set membership change makes the -final clean pass run against the current set — **so one of those changed and then restored still -costs the pass**, which a single "do they differ" test would have let through. The **assigned fix -set** and, in a Gate-B cycle, the **evidence entry** likewise keep the rules their own sources -give them. **The artifact's current-equality reading belongs to the artifact and is not carried to -any of the others.** None of these is read on the branch tip: writing the closing body is itself a -commit and changes none of them. - -**A commit the hook reads as cycle-closing is a separate matter, and it is a Gate-B one.** A -non-`WIP` commit mid-cycle makes the hook drop its Gate-B review state and that gate's counter — -an observation about the counter, since the cycle itself stays open until the conditions above -hold, so an accidental commit destroys Gate-B pass credit and closes nothing. **It does not reach -a Gate-A cycle's count**, which the hook clears at the skill boundaries that start a new Gate-A -cycle rather than on any commit. A plateau or tells on -**the pass that closes** go into **that pass's status report to the user** — the carrier the -three-line duty already names, and no second report form is introduced — and never block it, -because reporting "will not converge" on a converged loop is a false report; on an eligible pass -that does **not** close they go into the same place, a pass reporting what it read of the loop -whether or not it closes. **No other pass outcome makes a cycle eligible -to close**, because every other pass either leaves a required repair, a hold or a question -outstanding **or has not reached the floor** — and closing over any of those is the failure this -ordering exists to prevent. The floor is named separately because a below-floor pass whose only -findings are Minors leaves nothing outstanding and is still not eligible. The one termination -that is not a pass outcome is the Gate-B triviality skip, which runs no passes and is outside -this ordering. - -**Second, only a pass that is not a clean completion can suspend — and *clean completion* here -means the whole first branch, eligibility included.** So a clean pass **below** the floor is not -a clean completion and can suspend, while an **eligible** pass cannot, whatever its preconditions -do. That is what makes "clean completion outranks the two-tell stop" executable rather than -asserted, and cleanliness alone never decides it. Three suspensions, -by the names their paragraphs use and read by those paragraphs: the **scope stop**, raised by -either trigger above — a **membership stop** by the first, a **question stop** by the second; the -**clearly-stuck exit**; and the **two-tell stop**. A suspension waives nothing. Any non-empty set -of them can apply to one pass: **one surface, every reason reported, every question asked**, -because a reason left out is a decision made by omission. A finding the clearly-stuck reading -surfaces that also carries either trigger takes the scope stop's answers at that same surface, so -it is not asked twice; the two-tell stop surfaces tells and not a finding. - -**Third, and the branch is qualified before it is stated.** **Where any unmet closure condition's -own source prescribes stop-and-surface** — a profile present but unresolvable, governing headers -that disagree, a `Story:` header that cannot be read — **the cycle stays open, that source decides -what must be repaired or answered, and no further pass runs while its block stands.** It is neither -a suspension nor a continue and needs no name and no procedure of its own: the source rule carries -both, and this ordering's part is to send the reader there rather than to run a pass over a cycle -another rule has stopped. **That case is settled first**, so what follows never reaches it. - -**Otherwise, a pass that neither closes nor suspends continues** — the loop runs another pass on the -**current** artifact, revised where the severity and scope rules require a repair and unrevised -where they do not. **An eligible pass with an unmet closure precondition lands here**: clean -completion did not close it, and being eligible it cannot suspend, so the loop continues on whatever -the unmet precondition requires — most often a repair still owed from an earlier pass. A -below-floor clean pass lands here too, **only where no suspension applies to it**; where one does, -the second branch has already taken it, because clean completion did not close the pass and only -closing outranks a suspension. So does a pass whose only findings are Minors and Nits, which are -collected and never iterated and may leave nothing to revise. It is a branch and not an inference, -because "does not close" read alone says nothing about whether to run again. +condition of this cycle holding plus this cycle's gate's closing act**, and those conditions are +properties of the *cycle*, not of the pass — so an eligible pass whose conditions are unmet closes +nothing, and is not thereby made unclean. Keeping the two apart is what lets the order decide the +case where they disagree, which the branches do. **The order inside closure is fixed: every +condition is established first, and only then is the closing act performed.** + +**Closure introduces no new kind of record, and it excuses none**: every other record this cycle +owes, a human-exception record among them, is owed and written exactly as before, and **a +human-exception record this cycle owes goes in the commit its closing act uses**, so the two never +land in different places. **No closure condition is read on the branch tip**: writing the closing +body is itself a commit and changes none of them. + +A plateau or tells on **the pass that closes** go into **that pass's status report to the user** — +the carrier the three-line duty already names, and no second report form is introduced — and never +block it, because reporting "will not converge" on a converged loop is a false report; on an +eligible pass that does **not** close they go into the same place, a pass reporting what it read of +the loop whether or not it closes. **No other pass outcome makes a cycle eligible to close**, +because every other pass either leaves a required repair, a hold or a question outstanding, **or +has not reached the floor, or is itself unclean on its own findings** — and closing over any of +those is the failure this ordering exists to prevent. The floor is named separately because a +below-floor pass whose only findings are Minors leaves nothing outstanding and is still not +eligible; **pass-level uncleanliness is named separately** because an in-set finding already +validly dismissed and re-raised leaves the resolve duty discharged and still makes its pass +unclean. The one termination that is not a pass outcome is the Gate-B triviality skip, which runs +no passes and is outside this ordering. + +**Then the suspension branch, which only a pass that is not a clean completion reaches — and +*clean completion* means that whole branch, eligibility included.** So a clean pass **below** the +floor is not a clean completion and can suspend, while an **eligible** pass cannot, whatever its +conditions do. That is what makes "clean completion outranks the two-tell stop" executable rather +than asserted, and cleanliness alone never decides it. Three suspensions, by the names their +paragraphs use and read by those paragraphs: the **scope stop**, raised by either trigger above — a +**membership stop** by the first, a **question stop** by the second; the **clearly-stuck exit**; +and the **two-tell stop**. A suspension waives nothing. Any non-empty set of them can apply to one +pass: **one surface, every reason reported, every question asked**, because a reason left out is a +decision made by omission. A finding the clearly-stuck reading surfaces that also carries either +trigger takes the scope stop's answers at that same surface, so it is not asked twice; the two-tell +stop surfaces tells and not a finding. + +**Otherwise the continue branch: a pass that neither closes nor suspends continues** — the loop +runs another pass on the **current** artifact, revised where the severity and scope rules require a +repair and unrevised where they do not. **An eligible pass with an unmet closure condition lands +here**: clean completion did not close it, and being eligible it cannot suspend, so the loop +continues on whatever the unmet condition requires — most often a repair still owed from an earlier +pass. A below-floor clean pass lands here too, **only where no suspension applies to it**; where +one does, the suspension branch has already taken it, because clean completion did not close the +pass and only closing outranks a suspension. So does a pass whose only findings are Minors and +Nits, which are collected and never iterated and may leave nothing to revise. It is a branch and +not an inference, because "does not close" read alone says nothing about whether to run again. **The four standing duties, classified.** The **derived floor** is a **precondition on closure**: -it gates closing, discharged by the count of valid logical passes reaching it with the last of -them clean, or by the zero-finding exit. The **Blocker/Major-resolve duty** is a **precondition on +it gates closing, discharged by the count of valid logical passes reaching it with the last of them +clean, or by the zero-finding exit. The **Blocker/Major-resolve duty** is a **precondition on closure**: Mechanics · Severity, scoped to the assigned fix set, states what it demands and what -discharges it — a repair or a validated dismissal — and what a dismissal is, all at that source. -**It is discharged per finding and tracked across the cycle, never inferred from a later pass.** A -findings file establishes the **inventory** of what that pass found and not the resolution of -anything, so a later pass that does not mention an earlier in-set Blocker or Major says nothing -about whether it was repaired or dismissed; reading its absence as discharge would let an omission -close a cycle. **The duty is not a second test on whether a pass is clean, and the two are not run -together.** A pass is clean on its own findings. Stated as the case that separates them, because a -reader who conflates them decides it wrongly: **pass 1 raises an in-set Major; it is neither -repaired nor dismissed; pass 2 finds nothing.** Pass 2 **is** clean, and at or above the floor it -is **eligible** — and the cycle **still cannot close**, the duty being unmet; it continues on the -third branch until that Major is discharged. What the open Major does *not* do is make pass 2 +discharges it, at that source and not here. **It is discharged per finding and tracked across the +cycle, never inferred from a later pass.** A findings file establishes the **inventory** of what +that pass found and not the resolution of anything, so a later pass that does not mention an +earlier in-set Blocker or Major says nothing about whether it was resolved; reading its absence as +discharge would let an omission close a cycle. **The duty is not a second test on whether a pass is +clean, and the two are not run together.** A pass is clean on its own findings. Stated as the case +that separates them, because a reader who conflates them decides it wrongly: **pass 1 raises an +in-set Major; it is not resolved; pass 2 finds nothing.** Pass 2 **is** clean, and at or above the +floor it is **eligible** — and the cycle **still cannot close**, the duty being unmet; it takes the +continue branch until that Major is discharged. What the open Major does *not* do is make pass 2 unclean. The **hold** a surfaced finding places on closure **participates in the ordering**: it gates closing while it stands, and is discharged by the answers that surface requires. **It attaches to every surfaced finding, whichever suspension surfaced it** — clean completion creates @@ -247,75 +215,142 @@ regeneration chain that were repaired or dismissed are history the reading consu findings it re-surfaces, so no discharged finding takes a second hold. Each surfaced finding takes a hold like any other. **A re-raised valid dismissal stays discharged for the resolve duty** — the dismissal was the resolution and a reviewer repeating the finding does not undo it, so no second -dismissal is owed — and what the recurrence creates is the **clearly-stuck hold**, ended by -that reading's continue-or-stop answer. Any trigger the recurrence independently carries raises its -own stop as usual. **Where one finding is surfaced by both, it carries two hold components and -each is discharged by its own answer**: the **membership** -component ends on the membership answer **in either direction**, a decline releasing it exactly as -an accept does; the **clearly-stuck** component ends on the reading's continue-or-stop answer. Neither -answer discharges the other's component, and where the finding carries no scope-stop trigger the -continue-or-stop answer is the only one its surface asks for and discharges the only component -there is. **Resumption is still the composition rule's**, which waits for every outstanding -answer. At a **membership stop** the answer is **accept**, the finding joining the fix set where -Mechanics · Severity governs it, or **decline**, the finding staying outside and binding so for -the rest of this cycle. A later answer that contradicts a decline **does not reverse it**: the -decline **remains binding** and the contradiction is **surfaced to the user as information**, -changing neither membership, nor the cycle's state, nor any outstanding question — **whether the -loop resumes is decided by the composition rule below and by nothing here**, so a contradictory -answer is never itself a resumption and cannot step past a hold or a health question still -awaiting its own answer. Nothing here turns one answer into another, since that would let a -finding be moved out of the set and back into it to escape what it owes inside it. **There is no -withdrawal inside the cycle that declined**, a decline binding for the remainder of -its cycle and admitting no exception; a reconsideration is a later cycle's, where that decline has -no effect at all and the finding takes the ordinary route. Either answer is an -**explicit, attributable decision on that specific finding** — never silence, never a general -remark about scope, never inferred, because a fix set changed by inference is a fix set nobody -chose. **Membership is answered against the set as the absorb paragraph fixes it for the pass that -raised the question**: a later broadening is a new fact the **next** pass reads and never -discharges a standing hold, a hold discharged by a scope change being a hold nobody answered. At a -**question stop** the answer is the user's decision on the question and membership does not -change; an out-of-set finding that opened one is a membership stop as well. **Decline is available -only at a membership stop**, that being the only stop whose question is whether a finding belongs -to the set. The **clearly-stuck and two-tell readings** ask **continue or stop**. **Continue -consumes the reading that raised the suspension**: a further health suspension needs that reading -recomputed over a pass run after the answer, which is new data — so continue produces a distinct -next state, and the same reading cannot return the same stop unanswered. It permits an -**unrevised** artifact **only where no repair is owed**; where effective severity or scope requires -one, that repair comes before the post-answer pass, since a pass run over an unrepaired in-set -Blocker or Major spends a look on text the rules already say must change. **Stop parks the -cycle**: open, not running, spending no passes, restarted only by an explicit later continue — a -distinct state from the suspended-awaiting-answer one it was in before the answer. **That continue -restarts the cycle and never skips an answer**: where any question the suspension raised is still -outstanding, it returns the cycle to suspended-awaiting-answer, and only once every answer the -composition rule requires has been given does the next pass run. So a cycle parked with an -unanswered membership or question stop cannot be continued into a pass, and cannot sit parked with -no transition either — the continue is always available and always moves it. Nothing a parked -cycle wrote is a closing commit, and a parked cycle nobody restarts is a human's to resolve, -exactly as the nonce rules already say of open cycles. +dismissal is owed — and what the recurrence creates is the **clearly-stuck hold**, ended by that +reading's continue-or-stop answer. Any trigger the recurrence independently carries raises its own +stop as usual. **Where one finding is surfaced by both, it carries two hold components and each is +discharged by its own answer**: the **membership** component ends on the membership answer **in +either direction**, a decline releasing it exactly as an accept does; the **clearly-stuck** +component ends on the reading's continue-or-stop answer. Neither answer discharges the other's +component, and where the finding carries no scope-stop trigger the continue-or-stop answer is the +only one its surface asks for and discharges the only component there is. **Resumption is still the +composition rule's**, which waits for every outstanding answer. At a **membership stop** the answer +is **accept**, the finding joining the fix set where Mechanics · Severity governs it, or +**decline**, the finding staying outside and binding so for the rest of this cycle. A later answer +that contradicts a decline **does not reverse it**: the decline **remains binding** and the +contradiction is **surfaced to the user as information**, changing neither membership, nor the +cycle's state, nor any outstanding question — **whether the loop resumes is decided by the +composition rule below and by nothing here**, so a contradictory answer is never itself a +resumption and cannot step past a hold or a health question still awaiting its own answer. Nothing +here turns one answer into another, since that would let a finding be moved out of the set and back +into it to escape what it owes inside it. **There is no withdrawal inside the cycle that +declined**, a decline binding for the remainder of its cycle and admitting no exception; a +reconsideration is a later cycle's, where that decline has no effect at all and the finding takes +the ordinary route. Either answer is an **explicit, attributable decision on that specific +finding** — never silence, never a general remark about scope, never inferred, because a fix set +changed by inference is a fix set nobody chose. **Membership is answered against the set as the +absorb paragraph fixes it for the pass that raised the question**: a later broadening is a new fact +the **next** pass reads and never discharges a standing hold, a hold discharged by a scope change +being a hold nobody answered. At a **question stop** the answer is the user's decision on the +question and membership does not change; an out-of-set finding that opened one is a membership stop +as well. **Decline is available only at a membership stop**, that being the only stop whose +question is whether a finding belongs to the set. The **clearly-stuck and two-tell readings** ask +**continue or stop**. **Continue consumes the reading that raised the suspension**: a further +health suspension needs that reading recomputed over a pass run after the answer, which is new +data — so continue produces a distinct next state, and the same reading cannot return the same stop +unanswered. It permits an **unrevised** artifact **only where no repair is owed**; where effective +severity or scope requires one, that repair comes before the post-answer pass, since a pass run +over an unrepaired in-set Blocker or Major spends a look on text the rules already say must change. +**Stop parks the cycle**: open, not running, spending no passes, restarted only by an explicit +later continue — a distinct state from the suspended-awaiting-answer one it was in before the +answer. **That continue restarts the cycle and never skips an answer**: where any question the +suspension raised is still outstanding, it returns the cycle to suspended-awaiting-answer, and only +once every answer the composition rule requires has been given does the next pass run. So a cycle +parked with an unanswered membership or question stop cannot be continued into a pass, and cannot +sit parked with no transition either — the continue is always available and always moves it. +Nothing a parked cycle wrote is a closing commit, and a parked cycle nobody restarts is a human's +to resolve, exactly as the nonce rules already say of open cycles. **Composition, and what cannot happen.** Every **question** is answered on its own and the loop -resumes only when every answer resumes it — accept or decline at a membership stop, a decision at -a question stop, continue at the health readings; one stop answer parks the whole suspension, -because a loop resumed over an unanswered question decides it by running. **The clearly-stuck and -two-tell readings raise one question between them, not two**, both asking continue or stop, so one -answer carrying every reason ends both — an instance of the sentence before it, not an exception. -**Two pairings cannot occur**, and no rule ranks them: clean completion and a **scope stop**, since -that stop's triggers are the clean predicate's own second half, so a pass raising one is not clean; -and a zero-finding pass and any suspension, since it has nothing to surface, nothing regenerating, -no cluster and no require↔withdraw pair. **Clean completion and the clearly-stuck exit can**, and -the overlap is admitted rather than argued away: the two read different severity fields, as -Mechanics · Severity sets out, so an **in-set** Blocker or Major the ceiling demotes below Major can regenerate across passes on -a pass that is clean. **A declined finding is not a route into that reading**: the third condition -admits regeneration across repair attempts and a re-raised validated dismissal, and a decline is -neither — it is the user's decision that a *true* finding stays outside the set, and it binds -for the cycle. **The order decides it and no new rule is needed.** The clearly-stuck paragraph's own -precedence sentence is stated here rather than there, because precedence is evaluation order and -this paragraph is where evaluation order is stated once; its opening words point back to that -paragraph, which is where the reading itself lives. That third condition is what makes a plateau -rather than a finish, and it is why **a clean completion takes precedence over this exit**: a -Blocker/Major-free pass **at or above the floor** has satisfied the clean-final-pass rule — -collect the Minors and Nits and close — and reporting "will not converge" on a converged loop is a -false report. Below the floor the pass **suspends**, clean completion having not closed it. +resumes only when every answer resumes it — accept or decline at a membership stop, a decision at a +question stop, continue at the health readings; one stop answer parks the whole suspension, because +a loop resumed over an unanswered question decides it by running. **The clearly-stuck and two-tell +readings raise one question between them, not two**, both asking continue or stop, so one answer +carrying every reason ends both — an instance of the sentence before it, not an exception. **Two +pairings cannot occur**, and no rule ranks them: clean completion and a **scope stop**, since that +stop's triggers are the clean predicate's own second half, so a pass raising one is not clean; and +a zero-finding pass and any suspension, since it has nothing to surface, nothing regenerating, no +cluster and no require↔withdraw pair. **Clean completion and the clearly-stuck exit can**, and the +overlap is admitted rather than argued away: the two read different severity fields, as +Mechanics · Severity sets out, so an **in-set** Blocker or Major the ceiling demotes below Major +can regenerate across passes on a pass that is clean. **A declined finding is not a route into that +reading**: the third condition admits regeneration across repair attempts and a re-raised validated +dismissal, and a decline is neither — it is the user's decision that a *true* finding stays outside +the set, and it binds for the cycle. **The order decides it and no new rule is needed.** The +clearly-stuck paragraph's own precedence clause is stated here rather than there, because +precedence is evaluation order and this paragraph is where evaluation order is stated once; the +rationale that clause turns on stays beside the reading in that paragraph, which is where the +reading itself lives. **A clean completion takes precedence over this exit**: a Blocker/Major-free +pass **at or above the floor** has satisfied the clean-final-pass rule — collect the Minors and +Nits and close — and reporting "will not converge" on a converged loop is a false report. Below the +floor the pass **suspends**, the clean pass having failed eligibility. +``` + +### A2 — Gate A's closure + +``` +**Gate A's content condition, and its closing act.** These are what Gate A adds to the conditions +the ordering states for every cycle; that list is there and is not repeated here. + +**Its content condition: the artifact as it now stands is identical to the text that went into the +final pass's review request.** Gate A hands the reviewer text rather than a git range, which is why +this condition is Gate A's and is written nowhere else. It is **current equality and deliberately +nothing more**: it does **not** say the artifact went untouched in between, and text edited and +then restored byte for byte satisfies it — the one place in this cycle's conditions where an undone +change costs nothing, and stated here because the ordering's conditions answer a change instead. +That is a decision rather than an oversight: a content comparison cannot tell those two states +apart, and a condition nobody can check is a condition nobody applies. It likewise says nothing +about **what the reviewer consumed**: no part of this act is offered as evidence of the review +payload. + +**The act.** A Gate-A cycle has no WIP snapshot to replace, so it closes by **writing the closing +commit — or the closing message of one that already exists — carrying the records this cycle +already owes**: its provenance line and its per-pass curve, in the forms Mechanics fixes, neither +of them altered. **The content must survive into that commit, and carrying it there is part of the +act**: the equality above is read on the artifact, while the commit is written from the effective +index, so an act that does not carry that same content through has checked the condition without +performing it. **The commit that closes the cycle contains, at the artifact path, exactly the text +the condition was read against.** The safe command sequence and whatever demonstrates it belong to +the plan; **the duty belongs here**, and it needs no new fingerprint and no new record. + +**Two cases, told apart by the repository's current state rather than by which commit introduced +what; the safe git sequence for each belongs to the plan. That reading decides how a cycle closes, +never whether it may.** +- **`HEAD` already carries that text at the artifact path.** Close by **amending `HEAD`'s message** + to add the records, leaving that path as it stands. +- **`HEAD` does not, the reviewed text being still uncommitted.** **Commit it unchanged** and close + in that commit. This case had no answer before: an eligible pass over repairs nobody had + committed could neither close nor suspend. + +**No new revision of the artifact is made to close a Gate-A cycle**, a new revision being one no +pass has run against — and committing already-reviewed text that was never committed is not one. +``` + +### A3 — Gate B's closure + +``` +**Gate B's content condition, and its closing act.** These are what Gate B adds to the conditions +the ordering states for every cycle; that list is there and is not repeated here, so nothing below +is an inventory of what this gate requires. + +**Its content condition is not an artifact/request equality, and none is written for it.** Gate B +reviews a **diff** identified by `baseSha` and `headSha` rather than a text handed to the reviewer, +so there is no reviewed text to compare an artifact against. What holds that place is already in +this section and is cited rather than restated: the range those two names fix, **both branches +issued against the same commit**; **a re-review after every fix**, a fix changing the artifact so +the prior review no longer covers it; **a fix that changes specified behaviour updating the spec in +the same commit**, so the re-review covers both; the battery and the mode-derived evidence the +profiles section obliges before a call; and the **evidence entry**, revalidated as that section +says. + +**The act** is the closing amend Mechanics · Finishing the cycle describes, performed once the +ordering reaches it — an eligible pass with every closure condition holding, **never a clean pass +on its own**. + +**A commit the hook reads as cycle-closing is a Gate-B matter.** A non-`WIP` commit mid-cycle makes +the hook drop its Gate-B review state and that gate's counter — an observation about the counter, +since the cycle itself stays open until the conditions hold, so an accidental commit destroys +Gate-B pass credit and closes nothing. **It does not reach a Gate-A cycle's count**, which the hook +clears at the skill boundaries that start a new Gate-A cycle rather than on any commit. ``` --- @@ -333,7 +368,8 @@ differently about the same clause. **The replacement bytes for this passage appe nowhere else in this file** — other sections cite what this paragraph defines, as §A cites the fix set and the two triggers, and citing is not a second copy. -**Changed:** `b3`, `b7`, `b8`, `b11`, `b12`, `b13`, `b16`, `b17`–`b18`. **Carried:** every other +**Changed:** `b3`, `b7`, `b8`, `b11`, `b12`, `b13`, `b16`, `b17`–`b18`. **Added:** the closing-time +change rule, which no inventoried condition carried, because none existed. **Carried:** every other sentence. The rationales follow the text. ``` @@ -370,7 +406,13 @@ decided by default. Stopping this way is **not an exit from the gate**: it is a in the closure ordering's sense, the floor, the Blocker/Major filter and the clean-final-pass rule all stand, and what the answer does is stated there — what the stop prevents is a loop committing you to a design you never chose, which is a different failure from an unfinished -review. +review. **A change to this set costs the cycle at least one further pass.** The set a pass was +begun under is the set its closure would rely on, so the window opens **when that set is fixed +for the pass**, as this paragraph defines it, and runs to the closing act; a change anywhere in +that window costs a further pass, **in either direction and whether or not the change is later +undone**, the set having governed the pass differently while it stood. That further pass must +itself be clean and every other closure duty must be satisfied; it is one more pass, not a +licence to close on the next one. ``` *Why `b3`:* the four severity actions are restated beside a pointer to the section that defines @@ -398,6 +440,18 @@ asks only the structural question, and absorbs out-of-set work silently. *Why `b17`–`b18`:* the stop is named a suspension and defers to the ordering rather than restating what ends it. +*Why the added closing-time rule (pass 26 finding 3):* §A cited this paragraph as the source +governing a closing-time change to the set, and **no such rule existed anywhere** — this paragraph +fixed the set before a pass, and §A said only that a later broadening is read by the **next** pass. +Neither answers what a narrowing, another kind of change, or a change since undone does at the +closing act, so a final pass could be accepted after the set changed under it. **Daniel decided the +behaviour on 2026-09-12**, choosing a change-costs-a-pass rule over "only a broadening counts" and +over "nothing at closing time": a narrowing is not harmless for having been inside a superset, +since it changes which findings are in the set and what follows from that. **The window opens where +the set is fixed for the pass**, not where that pass's findings arrive, so a change made while the +pass runs is inside it. It is written here because this paragraph owns the set, which is what makes +§A's citation true. + --- ## C. Passage (c) — recognizing clearly stuck — REPLACED, from the third condition @@ -427,8 +481,8 @@ because precedence is evaluation order. **D3**. The **below-the-floor sentence does not move at all — it is replaced**: it says a Blocker/Major-free pass below the floor carrying a Minor "keeps looping", while the ordering splits that case, such a pass **suspending** where any suspension applies to it and **continuing** -where none does. Its two halves live in the ordering's second and third branches, and no copy of -the live wording survives beside them — carrying it word for word would install an unconditional +where none does. Its two halves live in the ordering's suspension and continue branches, and no +copy of the live wording survives beside them — carrying it word for word would install an unconditional continuation next to the conditional one and give the same pass two answers. --- @@ -537,15 +591,20 @@ across C 1014–1015 and W 1198–1199. ``` Those have their own terminal actions and this paragraph changes none of them: on a STOP you still stop, and **neither a human's general assent nor this record** lets an agent close or -continue a cycle. **The answer a suspension asks for is not assent of that kind**: continue and -stop are the answers the closure ordering prescribes, given on the question that suspension -raised, and what each produces is stated there. +continue a cycle. **The answer a suspension asks for is not assent of that kind**: it is the +answer the closure ordering prescribes for that suspension, given on the question that suspension +raised, and both which answer that is and what it produces are stated there. ``` *Why (pass 19 finding 4):* the live sentence says no human answer lets an agent continue a cycle, while the ordering makes **continue** the prescribed answer that restarts a parked one. Left as it stands, an agent following it refuses the exact transition the ordering requires. The distinction the repair draws is between a human waving a rule through — which this paragraph still forbids — and answering the question a suspension actually asked. +*And (pass 26 finding 5):* an earlier repair named **continue or stop** as the answer, which is the +two health readings' vocabulary and not a membership stop's accept-or-decline. The replacement +**refers** to the ordering rather than enumerating the three, because an enumeration here is the +second description of the answer rules that finding 5 was raised about — this paragraph is not +their definition site. **5. The `Finishing the cycle` lead-in** (Mechanics · `baseSha`). It wraps across C 827–828 and W 1011–1012. @@ -574,8 +633,8 @@ answer given since the last pass changed what the rules require of it, or becaus reaches what the last reading did not. ``` *Why (pass 21 finding 1):* the live wording says to re-run "over the revised artifact … because -the artifact changes between passes", which the ordering's third branch contradicts — that branch -continues on an **unrevised** artifact wherever no repair is owed. A pass whose only findings were +the artifact changes between passes", which the ordering's continue branch contradicts — that +branch continues on an **unrevised** artifact wherever no repair is owed. A pass whose only findings were Minors, or one run after a health answer with nothing left to repair, receives two instructions and can satisfy the live one only by manufacturing a change. **The replacement keeps the breadth demand**, which is what that sentence exists for and the reason it must not be narrowed @@ -659,15 +718,15 @@ ordering decides both and decided them differently. ``` **Surfacing does not close the cycle, and that is what makes this reachable.** You surface *with the cycle and the new hold still open* — the resolve rule stands over the finding exactly as -Mechanics · Severity states it, which means **repaired or validly dismissed**, so a recurrence of +Mechanics · Severity states it, **which scopes it to the assigned fix set**, so a recurrence of one already validly dismissed **stays resolved** and owes neither a second dismissal nor a repair; the hold stands until its answers are given, and **what the answer does is the closure ordering's**. **A pass is credited clean or not on its own findings**, as that ordering defines cleanliness; surfacing a finding that carries a scope-stop trigger is what withholds the credit, and no credit is withheld for surfacing alone. Reading this as "stop instead of fixing" would put the -exit in competition with the rule that every Blocker and Major resolves, and then nothing could -satisfy both. +exit in competition with the rule that every **in-set** Blocker and Major resolves, and then +nothing could satisfy both. ``` **`a13`** — scoped to its own paragraph, which is what it was written to police. *(It wraps From 4b00d44053518a81e36cbf54ecae917f148fad52 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sun, 13 Sep 2026 10:45:12 +0200 Subject: [PATCH 052/181] =?UTF-8?q?docs(context):=20Gate-A=20pass=2027=20?= =?UTF-8?q?=E2=80=94=20one=20tell,=20B+M=206=20to=204,=20findings=20held?= =?UTF-8?q?=20open?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit First pass on the split ordering. Valid: terminator correct, 8 finding lines, count matches, all severity tokens recognized. Trend: findings 10, 14, 8. Blockers 0, 1, 1. Majors 2, 5, 3. Blocker+Major 2, 6, 4. One tell of five — the Blocker count failed to fall — which is below the mandatory threshold. Clearly-stuck is not reachable: coverage is not affirmable while finding 4 reaches ground no earlier pass read. None of pass 26's six Blocker/Majors returns, and the new Blocker is on different ground. The split answered them; the text is not thereby correct. Provenance, traced with git log -S rather than assumed: finding 4 is this round's new §A3 text naming a real pre-existing gap between Gate B's reviewed range and the index its amend commits; findings 2 and 3 are older defects from rounds 3 and 4, untouched this round; finding 1's trap predates the round, and this round's pass-26-finding-10 repair made it visible by naming pass-level uncleanliness. Minor 6 is genuine damage: repairing pass-26 finding 9 falsified the surviving "word for word" claim in §C and design §5. Findings 5, 7 and 8 are collected Minors returning, knowingly left unrepaired. The pass is not clean, so the cycle continues by its own rules. Nothing runs: the authorization covered exactly this pass and is spent. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-27.md | 9 ++++ .../gate-a-spec-awsf1ec771-resume.md | 45 ++++++++++++++++++- 2 files changed, 53 insertions(+), 1 deletion(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-27.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-27.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-27.md new file mode 100644 index 0000000..7f4be03 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-27.md @@ -0,0 +1,9 @@ +BLOCKER | high | §A/§C "validly dismissed and re-raised" | A re-raised in-set effective-Major dismissal is explicitly still unclean, but the only new suspension route still requires both a six-pass plateau and affirmative sufficient coverage; with a known unreviewed area, a stable product-behaviour false positive yields only the non-falling-Blocker tell, so neither health suspension applies | The cycle can never close while the reviewer repeats the finding and can never suspend while coverage remains insufficient, violating acceptance criterion 4 | Make recurrence of a validated dismissal an independent suspension trigger, or otherwise prevent that already-discharged recurrence from permanently poisoning pass cleanliness +MAJOR | high | §A/§B "accept, the finding joining the fix set" | The accept transition says every accepted finding joins the fix set, but §B computes that set from governing scope plus accepted repair obligations only; accepting an out-of-set Minor or Nit creates no repair obligation because Severity says collect and never iterate | A later pass can recompute the accepted Minor or Nit as outside and raise the same membership stop again, so the accept answer is not idempotent and the promised set change is undefined | Add every finding explicitly accepted at a membership stop to the in-session fix-set computation independently of whether its severity creates a repair duty +MAJOR | high | §A/§H "only findings are Minors and Nits" | §A sends a pass whose only findings are Minors and Nits to continue and §H says a Minor-only pass is clean, but a Minor or Nit can be out of set or open a new structural question and therefore carry a scope-stop trigger, which the clean predicate says makes the pass unclean | Readers can bypass the mandatory scope suspension and continue or close over an unanswered membership or contract question | Qualify both summaries with the absence of every scope-stop trigger and route triggered Minor or Nit findings through the suspension branch +MAJOR | high | §A3 "Gate B's content condition" | Gate B reviews only the committed baseSha-to-headSha range, yet the closing amend commits the effective index; staged content already present before the final review is outside that range, leaves the hook fingerprint unchanged through review and close, and is not excluded by any cited condition | A clean Gate-B cycle can close with staged content that neither review branch received, even while the advisory fingerprint check reports no change | Require the effective index tree to equal the reviewed headSha tree before the final pass and closing amend, folding any difference into the WIP snapshot and re-reviewing it +MINOR | high | §A "What this paragraph owns" | The declared single-authority boundary is still violated: §A owns hold discharge while §B says the membership answer ends the membership hold and §H says the hold stands until its answers are given; §A also restates the three source-owned change consequences it says it only cites | The target retains parallel descriptions of the same decisions, contrary to its section-assignment rule and the consolidation's reason for existing | Keep the operative rule at its declared owner and turn every other occurrence into a non-normative reference that does not repeat the discharge or change consequence +MINOR | high | §C "moves into the block above word for word" | The standing precedence is one sentence wrapping across CLAUDE.md lines 234–238 and workflow-init.md lines 437–441, beginning "That third condition"; §A leaves that opening in §C and moves only the later clause, capitalized as a new sentence, while both §C and design §5 claim the sentence moves unchanged | The mechanical source-accounting claim is false and can make the plan verify or preserve bytes that are not the proposed text | Say that only the operative precedence clause moves and that it is capitalized as a standalone sentence; keep the rationale assigned to §C +MINOR | high | §E "curve must stay derivable from the findings files alone" | The categorical claim covers the curve, but standing Mechanics also sources cycle identity, pass ranges and model identifiers outside findings files and explicitly permits an unrecoverable numeric count to be written as `?` | The replacement overstates the curve's provenance and checkability and disagrees with live curve grammar that remains unchanged | Limit the claim to the three numeric series while their validated findings files are available +MINOR | medium | §G "any record this section obliges" | In the installed prompt §G sits inside `### Mechanics (reference)` under §5, so "this section" can denote Mechanics or all of §5 even though record duties also occur outside Mechanics | Downstream readers can classify different record-production and transport sentences while both appear to apply the promised text-only membership test | Name §5 explicitly as the intended scope of the record-obligation clause +END OF FINDINGS (8 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index 592d236..a9a9985 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -53,7 +53,50 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 24 | 6913c35 | 9→**8** | 0→**1** | 7→**3** | yes | **B+M 4 and Majors 3, both the lowest of the cycle.** One tell. The Blocker is new ground: the closing act reads `HEAD` and the working artifact and never the **index**, which is what `git commit` commits — `AGENTS.md:93`'s own documented failure class; session 01a0950c-f003-79d3-985f-73e2885c9621 | | 25 | 31bbe8f | 8→**10** | 1→**0** | 3→**2** | yes | **B+M 2 and Majors 2, both by far the lowest of the cycle; zero Blockers.** Both Majors are new ground. **Five of the ten are re-raised collected Minors**, one on its fourth appearance; session 01a09522-65bc-7691-8adc-fb26e330810e | | 26 | 8b8e146 | 10→**14** | 0→**1** | 2→**5** | yes | **TWO-TELL STOP — mandatory, surfaced to Daniel.** **All six Blocker/Majors trace to repairs made in passes 23, 24 and 25**, each nameable. All 14 held open; session 01a09539-5b8c-7293-af82-cc5f3bd2ba48 | -| 27 | — | — | — | — | not run | blocked on the two-tell answer | +| 27 | d971ae7 | 14→**8** | 1→**1** | 5→**3** | yes | **restarted by Daniel 2026-09-13 after the gate-split counter-draft was applied. B+M 6→4; findings the second-lowest of the cycle. ONE tell — no mandatory stop.** Three of the eight are collected Minors knowingly left unrepaired (5, 7, 8); session 01a099dd-b906-7970-8b39-8a3498510af4 | + +## Pass-27 report — ONE TELL, no mandatory stop, cycle open and NOT running + +**Restart:** Daniel authorized the restart on 2026-09-13 for one thing only — apply the gate-split +counter-draft to the inactive target text, align the design, run **exactly** pass 27, report, and +stop. **No automatic follow-up round, no activation, no Gate B.** That authorization is spent. + +**Floor line:** derived floor **3**; risk **high**, security **none**; read fresh from +`docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md`. One cited story, level 2. + +- **Trend:** findings …, 10, 14, **8**. Blockers …, 0, 1, **1**. Majors …, 2, 5, **3**. + Blocker+Major …, 2, 6, **4**. +- **Cluster (pass 27):** product behaviour 7 of 8; prose about the change 1 (finding 6); the + instrument 0. +- **require↔withdraw:** none. Findings 5, 7 and 8 are collected Minors returning — pass 26's 8, 11 + and 13, knowingly left unrepaired under the no-Minor-rounds instruction. A collected Minor + returning is not a withdrawal reversed. + +**Tells: one of five — below the threshold. No mandatory stop.** The finding count fell 14 → 8; +the Blocker count failed to fall, 1 → 1, which is the one tell. The clearly-stuck exit is **not** +reachable either: coverage is not affirmable, finding 4 reaching Gate-B index-versus-range ground +no earlier pass had read. + +**The pass is not clean** — one Blocker and three Majors, all in-set — **so the cycle continues**, +and the continue branch is where it sits. It is **not** parked: no suspension applies. What holds +it is the authorization, not the rules. + +**Provenance of the four Blocker/Majors, traced with `git log -S` before being reported.** + +| # | Sev | Entered at | Reading | +|---|---|---|---| +| 1 | BLOCKER | clause new in `d971ae7`; trap older | **Rediscovered, made findable by this round.** The clean predicate already made a re-raised validated dismissal unclean; this round's pass-26-finding-10 repair named pass-level uncleanliness out loud, and naming it exposed that such a cycle can neither close nor suspend. The defect is not new; its visibility is. | +| 2 | MAJOR | round 4, reworked at 11 and `0168f88` | **Rediscovered, older.** Untouched this round. | +| 3 | MAJOR | round 3, reworked at 9 and `0168f88` | **Rediscovered, older.** Untouched this round. | +| 4 | MAJOR | §A3 is new text in `d971ae7` | **This round's new text, real gap.** Gate B reviews a `baseSha`..`headSha` range while the closing amend commits the effective index; staged content present before the final review is in neither. §A3 is the first text to make that comparable. | + +**One Minor is genuine damage from this round:** finding 6. §A1 dropped the "That third condition…" +opening per pass-26 finding 9, which makes §C's and design §5's surviving "moves … word for word" +claim false. Repairing finding 9 broke that claim; nothing else did. + +**What the six pass-26 Blocker/Majors did.** None of them returns. Findings 1–6 of pass 26 are +absent from pass 27, and the pass-27 Blocker is on different ground. That is the strongest evidence +the split worked; it is **not** a claim that the text is correct. ## Pass-26 three-line report — MANDATORY TWO-TELL STOP From 233e915490a3970826247192e9045f3a4698b2d1 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sun, 13 Sep 2026 10:55:47 +0200 Subject: [PATCH 053/181] docs(specs): decide repeated-refuted findings; apply pass 27's three Majors MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Daniel's behaviour decision of 2026-09-13, and it changes the clean predicate: a finding this cycle has validly dismissed, repeated by a later pass with no new evidence and no change to the text the dismissal turned on, does not on its own make that later pass unclean. Narrow by construction — the repetition stays a finding in its file, every loop-health reading counts it, it remains the clearly-stuck re-raise condition, no earlier pass becomes clean in retrospect, and doubt about whether it is the same complaint resolves against the exclusion. No new suspension type and no record mechanism. The Blocker that prompted it is accepted on that core and not on its reasoning: it argued the cycle could never close, resting on a six-pass plateau the text sets as no threshold and on a coverage judgement that is a fact about now rather than forever. §C's third condition keeps the re-raised-dismissal clause and loses the reason that named the deadlock the clean predicate now answers. The three Majors: - an accept at a membership stop puts the finding in the set whatever its severity; membership and the repair duty are different things, and a later pass recomputing it as outside would make the accept decide nothing - a Minor-only pass continues or is clean only where it carries no scope-stop trigger, in both summaries that said otherwise - §A3 adds Gate B's missing condition: the effective index at the closing act carries nothing outside the reviewed range, a difference folding into the WIP snapshot and re-reviewed. The hook proves none of this and A3 says so in the terms AGENTS.md invariant 3 fixes Also repaired: pass 27 finding 6, damage from the previous round. The precedence sentence is split rather than moved whole, and §C and design §5 stop claiming it moves word for word. Collected, not repaired: findings 5, 7 and 8. The new clean rule governs the target rules only. Cycle awsf1ec771 continues under the rules it started with and may not use it to clear itself. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- ...26-09-10-loop-rule-consolidation-design.md | 32 ++++++++- ...-10-loop-rule-consolidation-target-text.md | 67 ++++++++++++++++--- 2 files changed, 86 insertions(+), 13 deletions(-) diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index b554da5..1ce0cd3 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -64,6 +64,22 @@ decision behind it: the block defines no trigger and no severity rule of its own precondition **that has a source of its own** keeps its one definition there, changed **at that source** where it had to change to agree with the ordering. +**One behaviour decision was added after the split, on Daniel's decision of 2026-09-13, and it +changes the clean predicate.** A finding this cycle has **validly dismissed** and a later pass +merely repeats, with no new evidence and no relevant change to the text the dismissal turned on, +**does not on its own make that later pass unclean**. Without it the standing duty *dismiss +validly, then run another pass* cannot finish: the reviewer would be the authority on whether its +own refuted claim had been dealt with, and a cycle could be held permanently unclean by a repeated +false positive. **It is deliberately narrow and it is not a waiver** — the repetition stays a +finding in its file, every loop-health reading counts it, it remains the clearly-stuck reading's +re-raise condition, no earlier pass becomes clean in retrospect, and doubt about whether it is the +same complaint is resolved against the exclusion. **No new suspension type and no record mechanism +were introduced for it**, which two earlier candidate answers would have required. **The Blocker +that prompted it is accepted on that core and not on its reasoning**: pass 27 argued the cycle +could *never* close, resting on a six-pass plateau the text does not set as a threshold and on a +coverage judgement that is a fact about now rather than forever. The repeated-refuted-complaint +problem stands on its own without either. + **The ordering is split into three paragraphs on Daniel's decision of 2026-09-12, and that split is the answer to pass 26's Blocker.** The ordering had stated one closure condition — the artifact's equality with the text sent to the reviewer — cycle-generally, while its explanation and its two @@ -78,6 +94,18 @@ and its own closing act**, and neither gate's paragraph is an inventory of what that was is now Gate A's own condition stated at Gate A's paragraph. **The split buys ownership and not brevity** — the three paragraphs together run slightly longer than the single block did. +**Writing Gate B's closure down exposed one condition nobody had stated** (pass 27 finding 4). +Gate B reviews a `baseSha`..`headSha` range while the closing amend commits the **effective +index**, so content staged before the final review is in the index and in no reviewer's payload. +A3 adds the one condition that closes it: the effective index at the closing act carries nothing +outside the reviewed range, a difference being folded into the `WIP:` snapshot and re-reviewed +under the re-review rule that already exists. **The hook proves none of this, and A3 says so in +the terms invariant 3 fixes**: its fingerprint includes an effective-index tree, so content staged +before the review call and still staged at the commit leaves it unmoved between the two +invocations — an unmoved fingerprint reports that nothing changed since it last looked, never that +a review covered what it is looking at. That distinction is read from `AGENTS.md` invariant 3, +which states the comparison; nothing is claimed here about the script beyond it. + **The closure-ordering block is an addition beside the source edits**, not one of them. §4 lists the edits; **no total is stated here or there**, because the unit — one contiguous replacement at one site — is not stable across revisions that merge or split a span, and a stated total then @@ -160,8 +188,8 @@ here. This table says what happens to each inventoried passage, so the map stays | Passage | This change | Target text | |---|---|---| | (a) the floor paragraphs | edited — closure sentences trimmed to a pointer, the no-restating prohibition scoped, the per-pass fix command pointed at Mechanics | §H | -| (b) what a loop absorbs | edited — owns both triggers and the fix set, so every qualification the ordering needs is made **here**, which is what keeps one definition per rule; **gains the closing-time rule for a change to the set**, which no inventoried condition carried because none existed (Daniel, 2026-09-12: a change costs at least one further pass, in either direction and whether or not it is later undone, the window opening where the set is fixed for the pass) | §B | -| (c) recognizing clearly stuck | edited — keeps its three-condition reading, stops carrying evaluation order; its precedence sentence moves into the block unchanged, and the third condition itself is widened at this source to admit a re-raised validated dismissal | §C | +| (b) what a loop absorbs | edited — owns both triggers and the fix set, so every qualification the ordering needs is made **here**, which is what keeps one definition per rule; **gains the closing-time rule for a change to the set**, which no inventoried condition carried because none existed (Daniel, 2026-09-12: a change costs at least one further pass, in either direction and whether or not it is later undone, the window opening where the set is fixed for the pass); **an accept at a membership stop puts the finding in the set whatever its severity** (pass 27 finding 2), membership and the repair duty being different things, so an accepted Minor is in the set although Severity asks no repair for it | §B | +| (c) recognizing clearly stuck | edited — keeps its three-condition reading, stops carrying evaluation order; **its precedence sentence is split**, the operative clause moving into the block unchanged and capitalized there while the plateau rationale stays at this source (pass 26 finding 9, pass 27 finding 6); the third condition is widened here to admit a re-raised validated dismissal, and its reason is restated because the deadlock it named is now answered by the clean predicate | §C | | (d) from pass 4 onward | **no longer edited.** The unavailable-history block moved to the successor with **D10** | — | | (e) the five tells | edited — the threshold is read after clean completion, and a pointer says what its answer does | §D | | (f) the two rules above do not compete | **unchanged.** "The two rules above" still names the absorb rule and the stuck reading; the block sits before both and adds no third rule between them | — | diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index dcdd838..8000e83 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -126,6 +126,28 @@ whatever the floor, because a floor buys further looks at an artifact that keeps and one yielding none has already given what those looks were for; don't manufacture findings to pad. +**A repetition of a finding this cycle has validly dismissed does not, on its own, make a pass +unclean.** Every part of this is required and the exclusion is narrow: the **dismissal was made in +an earlier pass of this cycle**, its stated reason is **still true of the artifact as it now +stands**, and the later finding **makes the same complaint and brings no new evidence** — no +observation the dismissal did not answer, and no change to the text the reason turned on. Where any +part fails — new evidence, relevant content changed, or genuine doubt that this finding is that +one — the finding is read afresh like any other, and **doubt never resolves in the exclusion's +favour**. Without it the standing duty *dismiss validly, then run another pass* cannot finish, +because a reviewer repeating its own refuted claim would decide whether that claim had been dealt +with. **It reaches only a dismissal this cycle made**, which is what an agent running the cycle +knows; nothing here ships a record, and recognising a dismissal across a lost session has no more +support than the paragraph above gives it. + +**It changes cleanliness and nothing else, which is what keeps it from being a waiver.** The +repetition is **still a finding**: it stands in its pass's findings file, and every loop-health +reading counts it exactly as it counts any other, so a loop spending passes on a point it keeps +refuting still shows up as one. It remains the clearly-stuck reading's re-raise condition. **No +earlier pass becomes clean in retrospect** — a pass's cleanliness is settled on what it found and +is never rewritten, which this exclusion leaves untouched: it decides the pass being read and no +other. And it is **not** a second dismissal; the resolve duty was discharged when the finding was +dismissed and there is nothing here to discharge again. + **Eligibility is exactly this and nothing more: a clean pass at or above the derived floor, or a zero-finding pass.** It is a property of the pass. **Closure is eligibility plus every closure condition of this cycle holding plus this cycle's gate's closing act**, and those conditions are @@ -175,7 +197,10 @@ continues on whatever the unmet condition requires — most often a repair still pass. A below-floor clean pass lands here too, **only where no suspension applies to it**; where one does, the suspension branch has already taken it, because clean completion did not close the pass and only closing outranks a suspension. So does a pass whose only findings are Minors and -Nits, which are collected and never iterated and may leave nothing to revise. It is a branch and +Nits **and which carries no scope-stop trigger** — those are collected and never iterated and may +leave nothing to revise, while a Minor or Nit that is out of set or opens a new structural question +carries a trigger like any other finding, is not clean, and has already been taken by the +suspension branch. It is a branch and not an inference, because "does not close" read alone says nothing about whether to run again. **The four standing duties, classified.** The **derived floor** is a **precondition on closure**: @@ -342,6 +367,18 @@ the same commit**, so the re-review covers both; the battery and the mode-derive profiles section obliges before a call; and the **evidence entry**, revalidated as that section says. +**What the range does not reach, and the one condition this gate adds for it.** The reviewed range +ends at a commit; **the closing amend commits the effective index**, and content staged before the +final review sits in the index without being in the range, so no reviewer saw it and no condition +above excludes it. **So: the effective index at the closing act carries nothing outside the range +the final pass reviewed.** A difference is not a failed review — it is unreviewed content: fold it +into the `WIP:` snapshot and re-review, which the re-review rule above already requires of any +change, and the pass that sees it becomes the candidate final one. **The hook's fingerprint +establishes none of this**: it is advisory, it compares its own inputs across its own invocations, +and content staged before the review call and still staged at the commit has not moved between +them — an unmoved fingerprint says nothing changed since it last looked, never that a review +covered what it is looking at. + **The act** is the closing amend Mechanics · Finishing the cycle describes, performed once the ordering reaches it — an eligible pass with every closure condition holding, **never a clean pass on its own**. @@ -380,8 +417,12 @@ severity exactly as Mechanics · Severity says. Ancestry decides where a finding never decides what you do with it, and it grants no Minor or Nit a repair round it would not otherwise get. **The assigned fix set is fixed before the pass you are answering: it is the union of the scope every approved story or plan governing this change assigns to this cycle, -plus repair obligations you already accepted in earlier passes, minus every finding this cycle -has declined.** A finding is in-set when repairing it stays inside **the assigned fix set as +plus every finding this cycle has accepted at a membership stop together with any repair +obligation accepted with it, minus every finding this cycle has declined.** **An accept puts the +finding in the set whatever its severity**: membership and the repair duty are different things, +so a Minor or Nit accepted into the set is in it though Mechanics · Severity asks no repair for +it, and a later pass that recomputed it as outside would raise the membership question a second +time and make the accept decide nothing. A finding is in-set when repairing it stays inside **the assigned fix set as just defined** — never merely because it arrived in the current pass, which would put every new finding in the set by definition and leave the boundary deciding nothing. Where membership is genuinely unclear treat the finding as **outside**, which costs a question and never a silent @@ -469,16 +510,20 @@ coverage is sufficient**, stated — a known materially unreviewed area forbids and disclosing it does not license it; and **Blocker or Major findings that keep regenerating across genuine repair attempts**, each round's fix producing the next — **or a finding the author has validly dismissed that the reviewer re-raises across passes**, the re-raise standing in for -the regenerating fix, since a dismissal gets no repair and produces none and a false positive that -returns every pass would otherwise leave the cycle unable to close and unable to suspend. That -third condition is what makes a plateau rather than a finish. **Where this reading and a clean +the regenerating fix, since a dismissal gets no repair and produces none, and a reviewer returning +to the same refuted point every pass says the same thing about the loop that a fix producing the +next finding says. That third condition is what makes a plateau rather than a finish. **Where this reading and a clean completion both apply, the closure ordering decides it** — the precedence sentence lives there, because precedence is evaluation order. ``` -*What follows in the live paragraph, and the two are not treated alike (pass 19 finding 8).* The -**precedence sentence** moves into the block above **word for word**, which is what satisfies -**D3**. The **below-the-floor sentence does not move at all — it is replaced**: it says a +*What follows in the live paragraph, and the two are not treated alike (pass 19 finding 8).* **Only the operative precedence clause moves into the block above**, from "a clean completion takes +precedence over this exit" to the end of the sentence, **capitalized there as a standalone +sentence**; its opening clause — the plateau rationale — **stays here**, which is pass 26 finding +9. So the words that move are unchanged and the sentence is **split, not moved whole**: saying it +moves word for word would be false of the sentence and would have the plan preserve bytes that are +not the proposed text. **D3 is satisfied by the operative clause surviving unchanged**, which is +what D3 is about. The **below-the-floor sentence does not move at all — it is replaced**: it says a Blocker/Major-free pass below the floor carrying a Minor "keeps looping", while the ordering splits that case, such a pass **suspending** where any suspension applies to it and **continuing** where none does. Its two halves live in the ordering's suspension and continue branches, and no @@ -754,8 +799,8 @@ closure ordering. Codex is advisory — validate before applying; dismissed find **The Gate-A clean-signal sentence** — the signal stops being the only route to clean. ``` Ask for one line per finding and a literal `NO FINDINGS` when a pass found none — that explicit -signal is what lets a pass be read as clean without inspecting it, and a pass carrying only -Minors is clean too and could never produce that file: +signal is what lets a pass be read as clean without inspecting it, and a pass carrying only Minors +**and no scope-stop trigger** is clean too and could never produce that file: ``` **The gate-prompt template's clean sentence**, in the block both gates paste. From 74e3ef04974495842ddb8b74ec644c7e89af00a8 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sun, 13 Sep 2026 11:09:26 +0200 Subject: [PATCH 054/181] =?UTF-8?q?docs(context):=20Gate-A=20pass=2028=20?= =?UTF-8?q?=E2=80=94=20mandatory=20two-tell=20stop,=20findings=20held=20op?= =?UTF-8?q?en?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Valid pass: terminator correct, 13 finding lines, count matches, all severity tokens recognized. Trend: findings 14, 8, 13. Blockers 1, 1, 4. Majors 5, 3, 2. Blocker+Major 6, 4, 6. Two tells of five — the count rose and the Blockers failed to fall — which is the mandatory threshold. Stop-and-surface, every finding open, no pass credited clean. Provenance traced with git log -S rather than assumed. Two of the six Blocker/Majors are damage this round made: the pass-27 round's pass-level uncleanliness example now contradicts the exclusion the pass-28 round added, and §A3's new rationale turns request targeting into observed coverage, which is the overclaim class AGENTS.md names as this repo's most persistent. Two are older text the new decision broke or left incomplete: the discharge sites state the recurrence rule without its three qualifications, and the unknown-start list shipped without a strict reading for the exclusion. Two are rediscovered: D3's preserved precedence rationale, and the Gate-A mirror of pass 27's index finding. Four of the six need no decision — they carry an already-made decision to the sites that state it. Two do: finding 1 would replace the rationale of the sentence D3 preserves verbatim, and finding 4 would add a Gate-A closure condition symmetric with the Gate-B one just decided. The run-to-completion authorization names a prescribed stop as a reporting point. This is one. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-28.md | 14 ++++++ .../gate-a-spec-awsf1ec771-resume.md | 49 +++++++++++++++++++ 2 files changed, 63 insertions(+) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-28.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-28.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-28.md new file mode 100644 index 0000000..e85e7a4 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-28.md @@ -0,0 +1,14 @@ +BLOCKER | high | §A1 "A clean completion takes precedence" | The retained rationale says a Blocker/Major-free pass at or above the floor should "collect the Minors and Nits and close", but the new clean predicate also requires no scope-stop trigger and closure also requires every cycle-level and gate-specific condition; the same section separately sends an eligible pass with an unmet condition to continue | An accepted-at-floor pass can close over a Minor or Nit carrying an unanswered scope question, an unresolved earlier Major, a standing hold, or a failed gate content condition, so the ordering gives the dangerous cases two outcomes | Preserve the settled precedence clause, but replace its rationale with the defined rule: close only when the pass is eligible, every closure condition holds, and the gate's closing act succeeds; otherwise take the specified suspension or continue branch +BLOCKER | high | §A1 "pass-level uncleanliness is named separately" | The example says an in-set finding already validly dismissed and re-raised "still makes its pass unclean", directly contradicting the new earlier rule that the same no-new-evidence, no-relevant-change repetition does not on its own make the later pass unclean | A later pass containing only that repetition is simultaneously clean and unclean, so its eligibility and whether it closes, suspends, or continues are undecidable | Replace this stale example with an ordinary in-set Blocker or Major whose repair after the pass discharges the resolve duty without rewriting that pass's unclean result +BLOCKER | high | §§A1, C and H "a re-raised valid dismissal stays discharged" | The no-new-evidence, same-complaint and no-relevant-text-change qualifications appear only in the cleanliness exclusion; the later resolve-duty text and H's replacement categorically say a recurrence stays resolved and owes no repair, while C admits any re-raised prior dismissal into its third condition | New evidence or a relevant artifact change can make the complaint true, yet the old dismissal can still be read as discharging the resolve duty; a later omitted-finding pass could then close over a true in-set Blocker or Major | Carry the same three qualifications into every discharge and clearly-stuck reference, and state that any failure of them creates an ordinary fresh finding whose current effective severity controls the resolve duty +BLOCKER | high | §A2 "No closure condition is read on the branch tip" | Gate A protects the reviewed artifact's text through the resulting commit, but protects no other closure input; an effective index can contain a staged version of a cited story profile or another governing scope source that differs from the working-tree version used for the final checks, and the amend or commit will publish that value despite the claim that the closing act changes no condition | A Gate-A cycle can close under one profile, cited-set-derived obligation, or assigned fix set while its closing commit installs another, bypassing the further-pass rule and invalidating the close | Require the resulting commit to preserve every source-owned closure input as it was established, or re-establish those inputs against the effective-index result before the act; keep only the safe Git sequence in the plan +MAJOR | high | §A3 "the range the final pass reviewed" | A3 repeatedly upgrades request targeting into observed coverage by saying staged content was seen by no reviewer, naming "the range the final pass reviewed", and making the next pass one that "sees it"; standing Mechanics instead says the kept baseSha/headSha values establish only that both calls were aimed at one range and are not evidence of what either branch reviewed | The new content condition can conservatively exclude index drift, but its explanation falsely claims the gate proves review consumption, violating AGENTS.md's gate-mechanism invariant and overstating what a closing commit demonstrates | Describe the range as the range both final-pass requests were aimed at and the staged difference as unavailable through that range, without claiming either branch actually consumed or saw it +MAJOR | high | §H "every suspension binding" | The standing unknown-start rule says each further rule this change ships adds its own strict reading to the list, but H adds only the suspensions and gives no unknown-start reading for the new repeated-dismissal cleanliness exclusion | When a cycle's starting rules cannot be established, one reader can apply the exclusion and close over a re-raised finding while another can withhold it as the stricter reading, so the fallback no longer decides the branch | Add the exclusion's strict reading explicitly, for example that it is unavailable unless the cycle can establish that its starting rules contained it, while leaving the deferred record and rollback material out +MINOR | high | §§A, B, F and H "What this paragraph owns" | The declared one-authority boundary is still false: A says it owns hold discharge, yet B says the membership answer ends the membership hold and H says the hold stands until its answers are given; A also fixes the human-exception destination that F assigns to the source paragraph and repeats source-owned profile, cited-set and fix-set change consequences | The target preserves parallel normative descriptions of the same transitions, recreating the exact drift mechanism the consolidation says it removes and violating its own one-passage-per-section rule | Keep each operative rule at its declared owner and make the other sites pure references that do not repeat a discharge, destination, or source-owned change consequence +MINOR | high | §E "The curve must stay derivable from the findings files alone" | The standing curve also carries cycle identity, pass ranges and model identifiers that findings files do not supply, and it explicitly permits a numeric series value of `?` when the count cannot be recovered from those files | The replacement overstates both the whole curve's provenance and its checkability and disagrees with the live grammar that remains beside it | Limit the derivability claim to the Findings, Blockers and Majors series when their validated findings files remain available +MINOR | medium | §G "any record this section obliges" | In the installed prompt G sits inside `### Mechanics (reference)` under §5, so "this section" can mean Mechanics or the whole Cross-Model Review section even though the phrase decides one-contract membership | Two downstream readers holding the same shipped text can classify different record-production and transport sentences while both appear to apply the promised membership test | Name §5 explicitly as the intended record-obligation scope +MINOR | high | §G "A curve without a cycle field cannot be told from another cycle's" | This categorical rationale contradicts standing Mechanics, which calls absent cycle attribution a limitation rather than a disqualification because kind and surrounding context can sometimes distinguish records and says the nonce only means a later reader usually does not need that context | The one-contract paragraph turns a probabilistic attribution aid into an impossibility claim, repeating the mechanism-overclaim class AGENTS.md specifically prohibits | Say curves cannot be reliably distinguished in every multi-cycle context without the cycle field +MINOR | high | §A3 "the re-review rule above already requires of any change" | The cited standing rule says only "Re-review after every fix"; it does not require re-review after arbitrary pre-existing or newly staged index content, so A3's own new instruction rather than the old rule is what covers a non-fix difference | The rationale assigns the new obligation to a source that does not contain it, leaving a future reader or edit able to remove A3's direct sentence while believing the coverage still exists elsewhere | State that A3 extends re-review to every effective-index difference from the targeted head, including differences that are not fixes, instead of attributing that breadth to the standing fix rule +MINOR | medium | standing §5 "the other closure and stop predicates" | The live sentence, which wraps across C 181–185 and W 388–392, presents assigned-fix-set membership, a new structural question, an accepted Blocker or Major and the tell thresholds as "the other closure and stop predicates" without marking the enumeration non-exhaustive; the target adds or makes explicit holds, no-clean-credit, gate content conditions and source blocks while leaving that list untouched | A reader entering through the floor-coherence paragraph can treat the four-item list as the closure surface and omit conditions the new ordering requires, while the one-contract rule then has two incompatible descriptions of membership | Make the list explicitly illustrative or replace it with a reference to the closure ordering's classifications +NIT | high | §A1 "A plateau or tells on the pass that closes" | The shared rationale says neither reading blocks closure because reporting "will not converge" on a converged loop is false, but only the clearly-stuck reading makes that claim; the two-tell rule reports the tells and asks whether to continue or stop, without concluding non-convergence | The settled precedence remains explicit, but its stated reason does not justify the two-tell half and fails the prompt-standard requirement that a constraint carry its actual why | Give the two-tell precedence its own accurate reason, such as that a pass satisfying every closure condition has no remaining cycle decision for the health question to suspend +END OF FINDINGS (13 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index a9a9985..5b5f01b 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -54,6 +54,55 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 25 | 31bbe8f | 8→**10** | 1→**0** | 3→**2** | yes | **B+M 2 and Majors 2, both by far the lowest of the cycle; zero Blockers.** Both Majors are new ground. **Five of the ten are re-raised collected Minors**, one on its fourth appearance; session 01a09522-65bc-7691-8adc-fb26e330810e | | 26 | 8b8e146 | 10→**14** | 0→**1** | 2→**5** | yes | **TWO-TELL STOP — mandatory, surfaced to Daniel.** **All six Blocker/Majors trace to repairs made in passes 23, 24 and 25**, each nameable. All 14 held open; session 01a09539-5b8c-7293-af82-cc5f3bd2ba48 | | 27 | d971ae7 | 14→**8** | 1→**1** | 5→**3** | yes | **restarted by Daniel 2026-09-13 after the gate-split counter-draft was applied. B+M 6→4; findings the second-lowest of the cycle. ONE tell — no mandatory stop.** Three of the eight are collected Minors knowingly left unrepaired (5, 7, 8); session 01a099dd-b906-7970-8b39-8a3498510af4 | +| 28 | 233e915 | 8→**13** | 1→**4** | 3→**2** | yes | **MANDATORY TWO-TELL STOP.** B+M 4→6. Four of the six trace to the pass-27/28 repair rounds (2, 5 damage; 3, 6 carry-through); 1 and 4 are rediscovered. All 13 held open; session 01a099fb-0706-7c53-ab6e-a0ccd8e191d3 | + +## Pass-28 report — MANDATORY TWO-TELL STOP + +**Authorization in force:** Daniel's run-to-completion commission of 2026-09-13 — decide the +repeated-refuted-finding rule, apply pass 27's three Majors, continue Gate A to a genuine clean +close, then plan → Gate A → implement both copies → version bump → evidence → Gate B, reporting +only at a prescribed stop, a new behaviour decision or a real obstacle. **A mandatory two-tell stop +is the first of those**, so the loop stops here with every finding open. + +**Floor line:** derived floor **3**; risk **high**, security **none**; read fresh from +`docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md`. One cited story, level 2. + +- **Trend:** findings …, 14, 8, **13**. Blockers …, 1, 1, **4**. Majors …, 5, 3, **2**. + Blocker+Major …, 6, 4, **6**. +- **Cluster (pass 28):** product behaviour 11 of 13; prose about the change 2 (5, 11 — both are + rationales overstating what a mechanism establishes); the instrument 0. +- **require↔withdraw:** none. + +**Tells: two of five — the threshold. Stop-and-surface is mandatory, not discretionary.** The +finding count rose 8 → 13 and the Blocker count failed to fall, 1 → 4. + +**Provenance of the six Blocker/Majors, traced with `git log -S` before being reported.** + +| # | Sev | Entered at | Round | Reading | +|---|---|---|---|---| +| 2 | BLOCKER | `d971ae7` | **27** | **My damage.** The pass-27 round wrote "pass-level uncleanliness is named separately", using a re-raised validated dismissal as its example; the pass-28 round then made exactly that case *not* unclean and left the example standing. The two sentences now contradict each other outright. | +| 3 | BLOCKER | `ef9a504` | 18 | **Older text, broken by this round.** §H and the resolve-duty text say categorically that a recurrence stays resolved. That was safe while any re-raise was unclean; the new exclusion's three qualifications were written into the clean predicate only, so a *qualified* recurrence and an *unqualified* one now read alike at the discharge sites. | +| 5 | MAJOR | new in `233e915` | **28** | **My damage, and it is the class `AGENTS.md` names as this repo's most persistent.** §A3's new rationale says staged content "was seen by no reviewer" and names "the range the final pass reviewed" — turning request targeting into observed coverage, which standing Mechanics explicitly refuses to claim. | +| 6 | MAJOR | 11 / `c8f96b8` | 11 | **Older list, incomplete for this round's decision.** The unknown-start rule says each further rule this change ships adds its own strict reading; the new exclusion shipped without one. | +| 1 | BLOCKER | round 1, reworked at 4 | 1 | **Rediscovered, oldest text in the cycle.** D3's preserved precedence rationale says "collect the Minors and Nits and close", which the strengthened clean predicate and the closure conditions now both qualify. | +| 4 | BLOCKER | `8b8e146` | 25 | **Rediscovered, and it is the Gate-A mirror of pass 27's finding 4.** §A2 protects the artifact's text through the closing commit and no other closure input; a staged profile or story differing from the working tree is published by the closing act. | + +**So: two of six are damage this round made (2, 5), two are older text this round's decision +broke or left incomplete (3, 6), and two are rediscovered defects (1, 4).** The loop is not +regenerating blindly — it is finding the carry-through of a decision made one round ago. + +**What needs a decision and what does not, stated so the stop is actionable.** +- **No decision needed** — 2, 3, 5, 6 and 11 carry an already-made decision through to the sites + that state it. That is exactly the `AGENTS.md` Don't about replacing a decision procedure + without accounting for its old conditions, applied to the new exclusion. +- **Decision needed — finding 1.** Its fix replaces the rationale of the sentence **D3 preserves + verbatim**. Touching it is a change to a settled decision. +- **Decision needed — finding 4.** Its fix adds a Gate-A closure condition symmetric with the + Gate-B one just decided. Same shape, but it is a new behaviour rule. + +**The five collected Minors return again** (7, 8, 9, 10, 12) — the ownership boundary, the curve's +derivability, §G's scope word, §G's attribution overclaim, and the standing four-item list. None +has been repaired in any round; they are not evidence of a plateau. ## Pass-27 report — ONE TELL, no mandatory stop, cycle open and NOT running From 36db7f087bdda1c2d3bbdbe248bc76019291d5bf Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sun, 13 Sep 2026 11:13:00 +0200 Subject: [PATCH 055/181] docs(specs): apply pass 28; carry the dismissal rule to every site that states it MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Both findings the pass-28 report escalated were tested against a concrete counter-case before any change was derived, per Daniel's process decision of 2026-09-13, and neither needed a behaviour decision. Finding 1 does not hold as stated: a Minor carrying a scope-stop trigger makes the pass unclean, so the precedence sentence's antecedent is false and it never reaches that case. What is real is the entry-point hazard §F exists for — the sentence carries an unqualified "close". D3's sentence is preserved verbatim and one sentence after it scopes the ranking to the clean-completion branch. Finding 4 does hold: a profile staged as trivial and restored to high in the worktree is read as high by the pass and published as trivial by the closing commit. But the profile rule already forbids the result, so the repair corrects the false claim that the closing act changes no condition and points at the rule that bites. No new condition. The rest: - the pass-27 round's pass-level-uncleanliness example named the very case the pass-28 round excluded; replaced with an in-set Major repaired after its pass - the exclusion's three qualifications now appear at every site that states the discharge — the suspension paragraph, §C's third condition, §H's surfacing sentence, and the composition paragraph's re-raise clause. Stated only at the clean predicate they let a recurrence that has become true read as discharged - §A3 stops claiming staged content was seen by no reviewer. The range is what the requests were aimed at; nothing here establishes what either branch consumed. §A3 also owns the widened re-review duty instead of attributing it to the standing after-every-fix rule - the unknown-start list gains the exclusion's strict reading Three recurring Minors are corrected in passing rather than given a round, because two of them are the overclaim class AGENTS.md prohibits and invariant compliance is not optional: the curve's derivability claim is limited to the three numeric series, §G names §5 rather than "this section", and §G stops calling missing cycle attribution an impossibility. Collected: the ownership boundary, and the standing four-item predicate list. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- ...26-09-10-loop-rule-consolidation-design.md | 17 +++- ...-10-loop-rule-consolidation-target-text.md | 91 +++++++++++++------ 2 files changed, 75 insertions(+), 33 deletions(-) diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index 1ce0cd3..bf32693 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -74,7 +74,13 @@ false positive. **It is deliberately narrow and it is not a waiver** — the rep finding in its file, every loop-health reading counts it, it remains the clearly-stuck reading's re-raise condition, no earlier pass becomes clean in retrospect, and doubt about whether it is the same complaint is resolved against the exclusion. **No new suspension type and no record mechanism -were introduced for it**, which two earlier candidate answers would have required. **The Blocker +were introduced for it**, which two earlier candidate answers would have required. **Its three +qualifications are carried to every site that states the discharge** — the ordering's suspension +paragraph, §C's third condition and §H's surfacing sentence — because stated only at the clean +predicate they would leave a recurrence that has *become* true reading as discharged (pass 28 +finding 3). **It also has an unknown-start strict reading**: unavailable where a cycle cannot +establish that its starting rules contained it, since the standing fallback says each rule this +change ships adds its own (pass 28 finding 6). **The Blocker that prompted it is accepted on that core and not on its reasoning**: pass 27 argued the cycle could *never* close, resting on a six-pass plateau the text does not set as a threshold and on a coverage judgement that is a fact about now rather than forever. The repeated-refuted-complaint @@ -98,8 +104,13 @@ not brevity** — the three paragraphs together run slightly longer than the sin Gate B reviews a `baseSha`..`headSha` range while the closing amend commits the **effective index**, so content staged before the final review is in the index and in no reviewer's payload. A3 adds the one condition that closes it: the effective index at the closing act carries nothing -outside the reviewed range, a difference being folded into the `WIP:` snapshot and re-reviewed -under the re-review rule that already exists. **The hook proves none of this, and A3 says so in +outside the range the final pass's requests were aimed at, a difference being folded into the +`WIP:` snapshot and re-reviewed. **A3 widens this gate's re-review duty to say so and does not +attribute the breadth to the standing rule**, which requires a re-review after every *fix* and an +index difference not being one (pass 28 finding 11). **Gate A's mirror of the same hazard is not a +second condition**: a staged edit to a cited story or a profile header is published by A2's closing +commit, and the source rules already answer it — a profile or cited-set change costs a further +pass — so A2 says where those rules bite at the closing act and adds nothing (pass 28 finding 4). **The hook proves none of this, and A3 says so in the terms invariant 3 fixes**: its fingerprint includes an effective-index tree, so content staged before the review call and still staged at the commit leaves it unmoved between the two invocations — an unmoved fingerprint reports that nothing changed since it last looked, never that diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index 8000e83..389c706 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -159,8 +159,15 @@ condition is established first, and only then is the closing act performed.** **Closure introduces no new kind of record, and it excuses none**: every other record this cycle owes, a human-exception record among them, is owed and written exactly as before, and **a human-exception record this cycle owes goes in the commit its closing act uses**, so the two never -land in different places. **No closure condition is read on the branch tip**: writing the closing -body is itself a commit and changes none of them. +land in different places. **No closure condition is read on the branch tip**, the conditions being +read where their sources say and not off the tip — but **the closing act must not publish a value +none of them was established under.** A commit is written from the effective index, so a staged +edit to a cited story, a profile header or any other source-owned input lands in the closing commit +even though the pass read the working tree; the source rules already forbid the result — a profile +change costs a further pass, a cited-set change makes the final clean pass run against the current +set — so what this says is **where those rules bite at the closing act**, and it adds no condition +of its own. Where the act would publish such a change, the change has happened and its own rule +applies: another pass is owed. A plateau or tells on **the pass that closes** go into **that pass's status report to the user** — the carrier the three-line duty already names, and no second report form is introduced — and never @@ -171,9 +178,9 @@ because every other pass either leaves a required repair, a hold or a question o has not reached the floor, or is itself unclean on its own findings** — and closing over any of those is the failure this ordering exists to prevent. The floor is named separately because a below-floor pass whose only findings are Minors leaves nothing outstanding and is still not -eligible; **pass-level uncleanliness is named separately** because an in-set finding already -validly dismissed and re-raised leaves the resolve duty discharged and still makes its pass -unclean. The one termination that is not a pass outcome is the Gate-B triviality skip, which runs +eligible; **pass-level uncleanliness is named separately** because an in-set Blocker or Major +repaired after the pass that raised it discharges the resolve duty without making that pass clean, +so a cycle can owe nothing and still hold no pass it may close on. The one termination that is not a pass outcome is the Gate-B triviality skip, which runs no passes and is outside this ordering. **Then the suspension branch, which only a pass that is not a clean completion reaches — and @@ -238,10 +245,15 @@ on until it is answered. The **clearly-stuck reading surfaces findings** — **t pass being read that satisfy its regeneration condition, and only those**; earlier members of a regeneration chain that were repaired or dismissed are history the reading consults and never findings it re-surfaces, so no discharged finding takes a second hold. Each surfaced finding takes -a hold like any other. **A re-raised valid dismissal stays discharged for the resolve duty** — the -dismissal was the resolution and a reviewer repeating the finding does not undo it, so no second -dismissal is owed — and what the recurrence creates is the **clearly-stuck hold**, ended by that -reading's continue-or-stop answer. Any trigger the recurrence independently carries raises its own +a hold like any other. **A re-raised valid dismissal stays discharged for the resolve duty on +exactly the terms the clean predicate sets out above** — the same complaint, no new evidence, no +change to the text the dismissal turned on, and the dismissal's reason still true of the artifact. +The dismissal was the resolution and a reviewer repeating it does not undo it, so no second +dismissal is owed. **Where any of those fails the recurrence is an ordinary fresh finding**, judged +at its current effective severity and owing a repair or a dismissal of its own; reading the old +dismissal as covering it would let a finding that has since become true close a cycle. What a +qualifying recurrence creates is the **clearly-stuck hold**, ended by that reading's +continue-or-stop answer. Any trigger the recurrence independently carries raises its own stop as usual. **Where one finding is surfaced by both, it carries two hold components and each is discharged by its own answer**: the **membership** component ends on the membership answer **in either direction**, a decline releasing it exactly as an accept does; the **clearly-stuck** @@ -298,8 +310,8 @@ cluster and no require↔withdraw pair. **Clean completion and the clearly-stuck overlap is admitted rather than argued away: the two read different severity fields, as Mechanics · Severity sets out, so an **in-set** Blocker or Major the ceiling demotes below Major can regenerate across passes on a pass that is clean. **A declined finding is not a route into that -reading**: the third condition admits regeneration across repair attempts and a re-raised validated -dismissal, and a decline is neither — it is the user's decision that a *true* finding stays outside +reading**: the third condition admits regeneration across repair attempts and a qualifying +re-raised validated dismissal, and a decline is neither — it is the user's decision that a *true* finding stays outside the set, and it binds for the cycle. **The order decides it and no new rule is needed.** The clearly-stuck paragraph's own precedence clause is stated here rather than there, because precedence is evaluation order and this paragraph is where evaluation order is stated once; the @@ -307,7 +319,11 @@ rationale that clause turns on stays beside the reading in that paragraph, which reading itself lives. **A clean completion takes precedence over this exit**: a Blocker/Major-free pass **at or above the floor** has satisfied the clean-final-pass rule — collect the Minors and Nits and close — and reporting "will not converge" on a converged loop is a false report. Below the -floor the pass **suspends**, the clean pass having failed eligibility. +floor the pass **suspends**, the clean pass having failed eligibility. **That sentence ranks two +readings and licenses no closure**, its "close" being the clean-completion branch's and carrying +every condition that branch carries: a Minor or Nit bearing a scope-stop trigger makes the pass +unclean, so the sentence does not reach it, and an undischarged duty, a standing hold or an unmet +gate condition leaves an eligible pass at the continue branch exactly as that branch says. ``` ### A2 — Gate A's closure @@ -369,15 +385,18 @@ says. **What the range does not reach, and the one condition this gate adds for it.** The reviewed range ends at a commit; **the closing amend commits the effective index**, and content staged before the -final review sits in the index without being in the range, so no reviewer saw it and no condition -above excludes it. **So: the effective index at the closing act carries nothing outside the range -the final pass reviewed.** A difference is not a failed review — it is unreviewed content: fold it -into the `WIP:` snapshot and re-review, which the re-review rule above already requires of any -change, and the pass that sees it becomes the candidate final one. **The hook's fingerprint -establishes none of this**: it is advisory, it compares its own inputs across its own invocations, -and content staged before the review call and still staged at the commit has not moved between -them — an unmoved fingerprint says nothing changed since it last looked, never that a review -covered what it is looking at. +final review sits in the index without being in the range, so it was never inside what the review +request selected and no condition above excludes it. **So: the effective index at the closing act +carries nothing outside the range the final pass's requests were aimed at.** A difference is not a +failed review — it is content the request could not reach: fold it into the `WIP:` snapshot and +re-review. **This gate's re-review duty is widened here to say so**, the standing rule requiring a +re-review after every *fix* and an index difference not being one; the pass run over the widened +snapshot becomes the candidate final one. **Nothing here establishes what either branch actually +consumed** — the reply reports no reviewed revision, which is why the kept `baseSha` and `headSha` +establish only that both calls were aimed at one range. **The hook's fingerprint establishes less +still**: it is advisory and compares its own inputs across its own invocations, and content staged +before the review call and still staged at the commit has not moved between them — an unmoved +fingerprint says nothing changed since it last looked, never that anything reviewed it. **The act** is the closing amend Mechanics · Finishing the cycle describes, performed once the ordering reaches it — an eligible pass with every closure condition holding, **never a clean pass @@ -509,7 +528,9 @@ visible across passes (six or more is where the field saw one); an **affirmative coverage is sufficient**, stated — a known materially unreviewed area forbids this exit outright, and disclosing it does not license it; and **Blocker or Major findings that keep regenerating across genuine repair attempts**, each round's fix producing the next — **or a finding the author -has validly dismissed that the reviewer re-raises across passes**, the re-raise standing in for +has validly dismissed that the reviewer re-raises across passes on the terms the closure ordering +sets**, a recurrence failing them being an ordinary fresh finding and not a re-raise at all, the +re-raise standing in for the regenerating fix, since a dismissal gets no repair and produces none, and a reviewer returning to the same refuted point every pass says the same thing about the loop that a fix producing the next finding says. That third condition is what makes a plateau rather than a finish. **Where this reading and a clean @@ -579,7 +600,10 @@ suspensions and what its answer produces are both stated. severity**, and that difference is the point rather than a discrepancy. It stays in the fix set either way; the ceiling moves what the cycle owes for it and never whether it is in. The line is **what the cycle owes versus what it observes about itself**, which is why no list of readings has to be kept complete - here. Two reasons for the split. The curve must stay derivable from the findings files alone — + here. Two reasons for the split. The curve's **three numeric series** must stay derivable from + the validated findings files alone wherever those files remain available — the rest of the curve + is not and does not claim to be, its cycle field, pass ranges and model identifiers coming from + elsewhere, and an unrecoverable count being written `?` exactly as the standing grammar allows — the finding total counts finding lines and the Blocker and Major series count the lines whose normalized severity is each, which is the only thing that makes a self-reported curve checkable; the subject clusters use no severity at all, being a judgement per finding that no count @@ -713,8 +737,9 @@ present. **Membership is decided by a test a reader can apply to the text in fro no list to consult, and the test reads what a rule states rather than what changing it would do: a live rule belongs to this contract when what it says determines or supplies an input the closure ordering reads, which branch a pass takes, what a hold is or what discharges it, -whether a cycle may close, or the production, identity or transport of any record this section -obliges a cycle to write.** **Read it on the sentence, never on the section the sentence sits in.** +whether a cycle may close, or the production, identity or transport of any record **§5** obliges a +cycle to write — §5 entire and not the Mechanics subsection this paragraph sits in, record duties +being stated in both.** **Read it on the sentence, never on the section the sentence sits in.** A sentence is a member when **it itself** fixes one of those things — what counts as a valid finding line, which files or records are owed, what ends a hold. It is not a member when it only shapes what a review produces, as the choice of reviewer, the lens set and the wording of a prompt @@ -727,8 +752,10 @@ any rule can be edited into deciding a branch and none decides one when edited c membership would follow the edit a reader pictured rather than the text in front of them. The last clause is why the squash carry belongs: it moves no pass and decides no branch, and a record that does not survive the merge is unreachable from the squash -commit and from `main`'s history. A curve without a cycle field cannot be told from another -cycle's where several are read together, a slot rule without a nonce cannot keep sibling cycles +commit and from `main`'s history. A curve without a cycle field cannot be reliably told from +another cycle's in every multi-cycle context — kind and surrounding context sometimes separate +them, which is why the standing rule calls missing attribution a limitation rather than a +disqualification — a slot rule without a nonce cannot keep sibling cycles apart — the bare names staying reserved for the legacy single-cycle case they already serve — a carry rule naming records a project does not produce is inert, and a clean predicate without the fix-set boundary it reads decides membership by accident. @@ -764,7 +791,9 @@ ordering decides both and decided them differently. **Surfacing does not close the cycle, and that is what makes this reachable.** You surface *with the cycle and the new hold still open* — the resolve rule stands over the finding exactly as Mechanics · Severity states it, **which scopes it to the assigned fix set**, so a recurrence of -one already validly dismissed **stays resolved** and owes neither a second dismissal nor a repair; +one already validly dismissed **stays resolved on the terms the closure ordering sets** and owes +neither a second dismissal nor a repair, while a recurrence failing any of them is an ordinary +fresh finding; the hold stands until its answers are given, and **what the answer does is the closure ordering's**. **A pass is credited clean or not on its own findings**, as that ordering defines cleanliness; @@ -825,8 +854,10 @@ rule in this section alone. dash-delimited list is given, the last item being the addition. ``` — at minimum floor 3, severity classified without the demotion, the provenance-line duty owed, -the curve duty owed, the nonce duties at their strictest, and **every suspension binding, since -starting rules that cannot be established cannot be read as having waived an open hold** — +the curve duty owed, the nonce duties at their strictest, **every suspension binding, since +starting rules that cannot be established cannot be read as having waived an open hold**, and **the +repeated-dismissal cleanliness exclusion unavailable, a cycle that cannot establish its starting +rules being unable to establish that they contained it** — ``` --- From 95439f86e611ceb084c04fde7cb704848aae76f8 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sun, 13 Sep 2026 11:25:54 +0200 Subject: [PATCH 056/181] docs(specs): apply pass 29; precedence turns on closing, not eligibility MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Zero tells: findings 13 to 10, Blockers 4 to 0. Pass 18 asked for clarity about which predicate did the work, not for the rule pass 29 narrows, so the pair is a follow-on and not a require-withdraw. Each Major was tested against a concrete case before any change was derived. Clean completion outranks a suspension by closing the cycle, not by being eligible. The earlier wording let a clean pass blocked by an unresolved prior Major walk past §D's mandatory two-tell stop and keep spending passes. This narrows a rule this change wrote and contradicts no settled decision: D2 and D3 forbid reporting "will not converge" on a converged loop, and a loop still owing a repair, an answer or a condition has not converged. The suspension branch, the continue branch, the precedence scoping clause and §D now say it the same way. Compliance repairs, no decision involved: - the unknown-start list gains every closure condition and pass-cost rule this change ships, the standing fallback requiring each to add its own reading - §A3's index condition names a tree-to-tree comparison against the explicit headSha; a range is not something an index can be inside, and a staged revert at a path the range touches would have read as covered - §G's membership test reaches rules deciding termination without a pass, the Gate-B triviality skip being the only such route and previously outside it One gap is named rather than closed, in §I and the design: a Gate-A closing act publishes a staged edit to a review input no source rule governs — a cited story's acceptance criteria, say. Profile values, cited-set membership and the assigned fix set are governed and answer themselves. Gate B's equivalent is answered because it was decided; Gate A's is a behaviour decision nobody has made, so the text states the residual instead of inventing a rule. The previous wording claimed the source rules already forbade every such value, which was false. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-29.md | 11 +++ ...26-09-10-loop-rule-consolidation-design.md | 18 ++++ ...-10-loop-rule-consolidation-target-text.md | 83 ++++++++++++------- 3 files changed, 82 insertions(+), 30 deletions(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-29.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-29.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-29.md new file mode 100644 index 0000000..b179d22 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-29.md @@ -0,0 +1,11 @@ +MAJOR | high | §§A1 and D "eligible pass cannot [suspend]" | A1 routes an eligible pass with an unmet closure condition directly to continue and says it cannot suspend whatever its conditions do, while §D still makes two tells a mandatory stop after the clean-completion branch and story D1-D3 say only clean completion outranks the health suspensions; an eligible pass that does not close is not a completion under A1's own definition that closure also needs every condition and the closing act | A clean current pass blocked by an unresolved prior Major, stale evidence or another unmet condition can ignore a mandatory two-tell or clearly-stuck suspension and keep spending passes without the user decision the standing rule requires | Limit health-suspension precedence to a pass that actually satisfies every closure condition and reaches the closing act; after the source-block check, let an eligible but non-closing pass take any applicable suspension before the continue branch +MAJOR | high | §A1 "a staged edit to a cited story" | The text claims the source rules already forbid every source-owned value the closing index could publish, but the cited source rules cover profile values, cited-set membership and assigned-fix-set scope, not a cited story's acceptance criteria or settled decisions; Gate A can read the working-tree story during its final pass while the index holds a different staged story body that changes none of those three predicates | The closing commit can publish governing decisions the clean pass never checked the artifact against, leaving a cycle marked closed even though its target text disagrees with the story in that same commit | Require the Gate-A closing act to preserve the complete cited-story and other review-input content under which the final pass ran, or make any differing effective-index source content current and run the further pass before closing +MAJOR | high | §H "unknown-start strict-reading list" | The standing fallback says each further rule this change ships adds its own strict reading to the list, but H adds only binding suspensions and an unavailable repeated-dismissal exclusion; the new Gate-A equality and commit-carry duty, Gate-B effective-index condition, and assigned-fix-set change-costs-a-pass rule have no stated unknown-start reading | A cycle whose starting rules cannot be established can apply the two listed additions yet omit a new closure condition or further-pass duty and close under a weaker hybrid of the old and new rules | Add the conservative unknown-start reading for every new closure condition and pass-cost rule this change ships, requiring rather than waiving each one when applicability cannot be established +MAJOR | high | §A3 "effective index ... carries nothing outside the range" | `baseSha..headSha` is a diff range while the effective index is a tree, and the condition never names the exact comparison that makes "outside" decidable; for example a staged revert at a path already changed in the range can be read as inside that range even though the closing amend would produce a tree different from `headSha` | Two readers can accept different closing indexes, allowing staged content absent from the final requests' selected head tree to land in the closing commit | State that the effective-index tree must equal the final pass's explicit `headSha` tree, with any inequality folded into the WIP snapshot and re-reviewed +MAJOR | high | §G "Membership is decided by a test" | The membership test covers ordering inputs, pass branches, holds, closure and records, but A1 explicitly places the Gate-B triviality skip outside the ordering as a no-pass termination; the sentence that defines the skip's eligibility therefore satisfies none of G's membership clauses even though it decides the only gate-off termination path | A downstream reader can treat a missing, stale or weaker triviality-skip predicate as outside the one-contract rule and run no Gate-B passes without triggering the required partial-adoption stop | Include rules that decide whether a cycle may terminate without a pass, specifically the Gate-B triviality skip, in the semantic membership test +MINOR | high | §§A, B, E, F and H "one definition" | The declared ownership boundary remains mechanically false: A says source-owned discharges are only cited but states when the resolve duty stays discharged, B states that a membership answer ends its hold, H states when the hold stands and ends, and both A and F assign the human-exception record to the closing-act commit while E calls itself the resolve duty's only statement | The proposed prompt retains parallel normative descriptions of the same transitions, recreating the drift mechanism the consolidation says it removes and violating the target's one-passage-per-section requirement | Keep each operative discharge, hold transition and record destination at its declared owner and make the other sites pure references that do not repeat the rule +MINOR | high | standing §5 "the other closure and stop predicates" | The live sentence at CLAUDE.md 181-185 and workflow-init.md 388-392 names assigned-fix-set membership, a new structural question, an accepted Blocker or Major and the tell thresholds as "the other closure and stop predicates" without saying the list is illustrative; the target leaves it unchanged while adding or making explicit source blocks, holds, no-clean-credit and gate-specific content conditions | A reader entering through the surviving floor-coherence paragraph can treat its four-item apposition as the closure surface and omit conditions the new ordering requires, contrary to the target's claim that standing text remains consistent | Mark the live list illustrative and point to the closure ordering for the complete classification +MINOR | high | §A1 "derived floor ... discharged" | The duties paragraph says the floor is discharged only by enough valid passes "with the last of them clean," even though the same section separates the numeric floor from pass cleanliness and defines eligibility as their conjunction; a dirty pass can satisfy the pass count without satisfying cleanliness | The floor and clean-pass predicates are merged again at the duty source, giving readers two accounts of what the floor itself gates and discharges and weakening acceptance criterion 2's visible classification | Say the floor is discharged by the required valid-pass count, with zero findings taking the stated early-exit alternative, and leave the clean-pass requirement solely in eligibility +MINOR | high | §H "the last item being the addition" | H's own description says the unknown-start list gains the suspensions and that the last item is the addition, but the shown replacement now contains two additions: every suspension binding and the repeated-dismissal exclusion unavailable | The section's accounting is mechanically false after the pass-28 edit and can make the plan or old-condition disposition treat one new clause as unassigned | State that the list gains both additions and name both in the introduction +NIT | high | §A1 "A plateau or tells on the pass that closes" | The shared reason says neither health reading blocks closure because reporting "will not converge" on a converged loop would be false, but only the clearly-stuck reading makes that report; the two-tell rule reports observed tells and asks continue or stop without claiming non-convergence | The precedence is operative, but its rationale does not justify the two-tell half and fails prompt-standard item 6's requirement that a constraint carry its actual why | Give the two-tell precedence its own reason, such as that a pass satisfying every closure condition leaves no cycle decision for that health question to suspend +END OF FINDINGS (10 total) diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index bf32693..89e632c 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -86,6 +86,24 @@ could *never* close, resting on a six-pass plateau the text does not set as a th coverage judgement that is a fact about now rather than forever. The repeated-refuted-complaint problem stands on its own without either. +**Clean completion outranks a suspension by closing the cycle, not by being eligible** (pass 29 +finding 1). The earlier wording barred every eligible pass from suspending, which let a clean pass +blocked by an unresolved prior Major, a standing hold or stale evidence walk past §D's **mandatory** +two-tell stop and keep spending passes. **This narrows a rule this change itself wrote, and +contradicts no settled decision**: D2 and D3 forbid reporting "will not converge" on a loop that +converged, and a loop still owing a repair, an answer or a closure condition has not converged. +Three other pass-29 repairs are compliance rather than decision — the unknown-start list gains the +remaining rules this change ships, §A3's index condition names a **tree-to-tree** comparison against +the explicit `headSha` rather than membership of a range, and §G's membership test reaches rules +deciding termination **without** a pass, the Gate-B triviality skip being the only such route. + +**One gap is named and not closed** (pass 29 finding 2, and §I carries it): a Gate-A closing act is +written from the effective index, so a staged edit to a review input **no source rule governs** — a +cited story's acceptance criteria, say — is published by the closing commit unchecked. Profile +values, cited-set membership and the assigned fix set are governed and answer themselves. Gate B's +equivalent is answered because its index condition was decided; **Gate A's is a behaviour decision +nobody has made**, so the text states the residual instead of inventing a rule for it. + **The ordering is split into three paragraphs on Daniel's decision of 2026-09-12, and that split is the answer to pass 26's Blocker.** The ordering had stated one closure condition — the artifact's equality with the text sent to the reviewer — cycle-generally, while its explanation and its two diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index 389c706..a9679b3 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -160,14 +160,15 @@ condition is established first, and only then is the closing act performed.** owes, a human-exception record among them, is owed and written exactly as before, and **a human-exception record this cycle owes goes in the commit its closing act uses**, so the two never land in different places. **No closure condition is read on the branch tip**, the conditions being -read where their sources say and not off the tip — but **the closing act must not publish a value -none of them was established under.** A commit is written from the effective index, so a staged -edit to a cited story, a profile header or any other source-owned input lands in the closing commit -even though the pass read the working tree; the source rules already forbid the result — a profile -change costs a further pass, a cited-set change makes the final clean pass run against the current -set — so what this says is **where those rules bite at the closing act**, and it adds no condition -of its own. Where the act would publish such a change, the change has happened and its own rule -applies: another pass is owed. +read where their sources say and not off the tip. **That is not a claim that the act publishes what +the pass read.** A commit is written from the effective index, so a staged edit the working tree +does not show lands in the closing commit; where the edit changes something a **source rule** +governs — a profile value, cited-set membership, the assigned fix set — **the change has happened +and that source's own rule applies**, so a further pass is owed and no condition is added here for +it. **Where it changes a review input no source rule governs** — a cited story's acceptance +criteria or settled decisions, say — **nothing here reaches it**, and that is stated as a residual +in §I rather than answered: the artifact's own equality condition covers the artifact, and the +inputs beside it have only the rules their sources give them. A plateau or tells on **the pass that closes** go into **that pass's status report to the user** — the carrier the three-line duty already names, and no second report form is introduced — and never @@ -183,11 +184,17 @@ repaired after the pass that raised it discharges the resolve duty without makin so a cycle can owe nothing and still hold no pass it may close on. The one termination that is not a pass outcome is the Gate-B triviality skip, which runs no passes and is outside this ordering. -**Then the suspension branch, which only a pass that is not a clean completion reaches — and -*clean completion* means that whole branch, eligibility included.** So a clean pass **below** the -floor is not a clean completion and can suspend, while an **eligible** pass cannot, whatever its -conditions do. That is what makes "clean completion outranks the two-tell stop" executable rather -than asserted, and cleanliness alone never decides it. Three suspensions, by the names their +**Then the suspension branch, which a pass reaches unless it closed the cycle.** **Clean completion +outranks a suspension by closing, not by being eligible**: a pass that took the branch above, met +every closure condition and had the closing act performed has ended the cycle, and a suspension has +nothing left to suspend. **A pass that did not close reaches this branch whatever its +cleanliness** — a clean pass below the floor, and equally an eligible pass the cycle's unmet +conditions kept from closing. That is what makes "clean completion outranks the two-tell stop" +executable rather than asserted, and cleanliness alone never decides it. **What D2 and D3 forbid is +reporting "will not converge" on a loop that converged, and a loop still owing a repair, an answer +or a closure condition has not converged** — so a mandatory two-tell stop and the clearly-stuck +reading stay reachable exactly where the loop is still running, which is the only place their +question means anything. Three suspensions, by the names their paragraphs use and read by those paragraphs: the **scope stop**, raised by either trigger above — a **membership stop** by the first, a **question stop** by the second; the **clearly-stuck exit**; and the **two-tell stop**. A suspension waives nothing. Any non-empty set of them can apply to one @@ -199,11 +206,11 @@ stop surfaces tells and not a finding. **Otherwise the continue branch: a pass that neither closes nor suspends continues** — the loop runs another pass on the **current** artifact, revised where the severity and scope rules require a repair and unrevised where they do not. **An eligible pass with an unmet closure condition lands -here**: clean completion did not close it, and being eligible it cannot suspend, so the loop -continues on whatever the unmet condition requires — most often a repair still owed from an earlier -pass. A below-floor clean pass lands here too, **only where no suspension applies to it**; where -one does, the suspension branch has already taken it, because clean completion did not close the -pass and only closing outranks a suspension. So does a pass whose only findings are Minors and +here**, and like every other non-closing pass **only where no suspension applies to it**: clean +completion did not close it, so the loop continues on whatever the unmet condition requires — most +often a repair still owed from an earlier pass. A below-floor clean pass lands here on the same +terms; where a suspension does apply, the suspension branch has already taken it, because only +closing outranks a suspension. So does a pass whose only findings are Minors and Nits **and which carries no scope-stop trigger** — those are collected and never iterated and may leave nothing to revise, while a Minor or Nit that is out of set or opens a new structural question carries a trigger like any other finding, is not clean, and has already been taken by the @@ -320,10 +327,11 @@ reading itself lives. **A clean completion takes precedence over this exit**: a pass **at or above the floor** has satisfied the clean-final-pass rule — collect the Minors and Nits and close — and reporting "will not converge" on a converged loop is a false report. Below the floor the pass **suspends**, the clean pass having failed eligibility. **That sentence ranks two -readings and licenses no closure**, its "close" being the clean-completion branch's and carrying -every condition that branch carries: a Minor or Nit bearing a scope-stop trigger makes the pass -unclean, so the sentence does not reach it, and an undischarged duty, a standing hold or an unmet -gate condition leaves an eligible pass at the continue branch exactly as that branch says. +readings and licenses no closure**, its "close" being the closure this ordering defines and +carrying every condition that closure carries: a Minor or Nit bearing a scope-stop trigger makes +the pass unclean, so the sentence does not reach it, and an undischarged duty, a standing hold or +an unmet gate condition means the pass does not close — leaving it on the suspension branch where +one applies and the continue branch where none does, exactly as those branches say. ``` ### A2 — Gate A's closure @@ -386,8 +394,10 @@ says. **What the range does not reach, and the one condition this gate adds for it.** The reviewed range ends at a commit; **the closing amend commits the effective index**, and content staged before the final review sits in the index without being in the range, so it was never inside what the review -request selected and no condition above excludes it. **So: the effective index at the closing act -carries nothing outside the range the final pass's requests were aimed at.** A difference is not a +request selected and no condition above excludes it. **So: the effective index tree at the closing act equals the tree of the explicit `headSha` the +final pass's requests named** — a tree compared with a tree, since a range is not a thing an index +can be inside and a staged revert at a path the range already touches would otherwise read as +covered. A difference is not a failed review — it is content the request could not reach: fold it into the `WIP:` snapshot and re-review. **This gate's re-review duty is widened here to say so**, the standing rule requiring a re-review after every *fix* and an index difference not being one; the pass run over the widened @@ -558,9 +568,9 @@ continuation next to the conditional one and give the same pass two answers. **`e7`, the threshold.** Gains one clause; the sentence is given entire. ``` **Any two present makes stop-and-surface mandatory, not discretionary** — read **after** the -clean-completion branch of the closure ordering, which outranks it — and you report the -tells and hand the decision to the user, and the "clearly stuck" reading above is not a -precondition for it. +clean-completion branch of the closure ordering, which outranks it **by closing the cycle and only +then** — and you report the tells and hand the decision to the user, and the "clearly stuck" +reading above is not a precondition for it. ``` **A pointer is added** at the end of the passage: @@ -737,7 +747,9 @@ present. **Membership is decided by a test a reader can apply to the text in fro no list to consult, and the test reads what a rule states rather than what changing it would do: a live rule belongs to this contract when what it says determines or supplies an input the closure ordering reads, which branch a pass takes, what a hold is or what discharges it, -whether a cycle may close, or the production, identity or transport of any record **§5** obliges a +whether a cycle may close **or may terminate without running a pass at all, the Gate-B triviality +skip being the one such route and its eligibility test therefore a member**, or the production, +identity or transport of any record **§5** obliges a cycle to write — §5 entire and not the Mechanics subsection this paragraph sits in, record duties being stated in both.** **Read it on the sentence, never on the section the sentence sits in.** A sentence is a member when **it itself** fixes one of those things — what counts as a valid @@ -855,9 +867,12 @@ dash-delimited list is given, the last item being the addition. ``` — at minimum floor 3, severity classified without the demotion, the provenance-line duty owed, the curve duty owed, the nonce duties at their strictest, **every suspension binding, since -starting rules that cannot be established cannot be read as having waived an open hold**, and **the +starting rules that cannot be established cannot be read as having waived an open hold**, **the repeated-dismissal cleanliness exclusion unavailable, a cycle that cannot establish its starting -rules being unable to establish that they contained it** — +rules being unable to establish that they contained it**, and **every closure condition and +pass-cost rule this change ships owed rather than waived — the gate content conditions and their +commit-carry duty, and the further pass an assigned-fix-set change costs — since a rule that cannot +be established as absent is cheaper to owe than to skip** — ``` --- @@ -872,3 +887,11 @@ rules being unable to establish that they contained it** — - **Partial adoption is instructed against, never detected.** §G says so in its own words. - **Which exit a cycle took is not observable from history.** The transport left with the record (successor story) and no story has taken it. An admitted gap, unowned. +- **A Gate-A closing act can publish a review input the final pass never read** (pass 29 finding + 2). The commit is written from the effective index, so a staged edit to a cited story's + acceptance criteria or settled decisions — anything no source rule governs, profile values, + cited-set membership and the assigned fix set being the ones that are governed — lands in the + closing commit unchecked. Gate B's equivalent **is** answered, its index condition being one of + this change's decisions; **Gate A's is not, and the asymmetry is deliberate only in the sense + that nobody has decided it.** Named here rather than answered, because answering it is a + behaviour decision this change has not been given. From e48259df5ab46ce0fcea6b17efa7b49429faf757 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sun, 13 Sep 2026 11:40:25 +0200 Subject: [PATCH 057/181] docs(specs): apply pass 30; a non-closing pass suspends before it continues MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Zero tells: findings 10 to 9, Blockers 0 to 0. A Blocker count already at zero cannot fall, and reading that as the "failing to fall" tell would flag the healthiest state the loop has; recorded so the reading is auditable rather than silent. Three of the four Majors are the carry-through of the pass-29 decision that a non-closing pass reaches the suspension branch. My own pre-review sweep covered the precedence wording and missed the unconditional continuation instructions, which is the sweep being too narrow rather than the review finding new ground. - §H's Gate-A cadence re-runs only where the ordering selects its continue branch; a suspension is answered first - §A3's fold-and-re-review is the continuation, not an instruction that outruns the ordering - the standing evidence-entry remedy is the EIGHTH falsified standing sentence and the sixth sharing §F's mechanism: an entry point carrying an unqualified fix-re-review-close. §F's two counts move with it The fourth is new ground: a closing act that fails leaves no defined next state, so a permission or signing error would spend a review pass and surface a loop-health problem the loop does not have. A failed act closes nothing, costs no pass, keeps every condition established and is retried once the concrete failure is repaired — available only while nothing a condition is read from was touched by the repair. This adds no mechanism; it says what "the act was not performed" already means. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-30.md | 10 +++++ ...26-09-10-loop-rule-consolidation-design.md | 9 +++++ ...-10-loop-rule-consolidation-target-text.md | 39 ++++++++++++++++--- 3 files changed, 53 insertions(+), 5 deletions(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-30.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-30.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-30.md new file mode 100644 index 0000000..2fa8459 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-30.md @@ -0,0 +1,10 @@ +MAJOR | high | §H "Each pass: validate, revise, re-run" | The replacement keeps an unconditional Gate-A re-run cadence even though §A1 now requires every non-closing pass to take an applicable suspension before the continue branch; a pass with two tells or a clearly-stuck reading is therefore told both to await the user's answer and to re-run | An agent entering through the Gate-A instructions can bypass a mandatory suspension and spend another pass without the decision the ordering requires | Limit the cadence to passes for which the closure ordering selects the continue branch, and point suspension outcomes back to that ordering +MAJOR | high | §A3 "fold it into the WIP snapshot and re-review" | The index-mismatch remediation is unconditional, but an eligible final pass whose index tree differs from its named headSha tree did not close and must first take any applicable clearly-stuck or two-tell suspension under §A1 | A Gate-B pass with an index mismatch and two tells has two incompatible next actions, so the direct re-review instruction can bypass the mandatory user stop | Qualify folding and re-reviewing as the post-suspension or no-suspension continuation, leaving the ordering to decide whether a user answer is owed first +MAJOR | high | standing §5 "If revalidation changes the entry" | The live evidence-entry sentence at CLAUDE.md 727–729 and workflow-init.md 913–915 says to fix, re-review and close when revalidation changes the entry, while the new ordering sends that same non-closing pass through any applicable health suspension before continuation; the target does not replace or qualify this source instruction | An eligible pass with stale evidence and two tells can skip the mandatory suspension by following the standing evidence path directly | Make the standing remediation conditional on the ordering selecting continuation, with any applicable suspension answered before the fix and re-review +MAJOR | medium | §A1 "only then is the closing act performed" | Closure requires the gate's closing act to succeed, but neither gate nor the four-branch ordering defines the next state when the commit or amend fails after every condition was established; because the pass did not close, the current text can only route it through an unrelated health suspension or into another full review pass | A permission, hook, signing or repository error can waste review passes indefinitely or surface a false loop-health problem while the actual closing failure remains untreated | Add a closing-act failure path that keeps the cycle open, surfaces the concrete command failure, and permits retrying the act after repair without another pass when every reviewed input and closure condition remains unchanged +MINOR | high | §§A, B, E, F and H "one definition" | The declared ownership boundary is still false: A says source-owned discharges are only cited but adds per-finding resolution and no-omission rules, B states when the membership hold ends, and both A and F assign a human-exception record to the closing-act commit while E calls itself the resolve duty's only statement | The proposed prompt retains parallel normative descriptions of the same transitions, recreating the drift mechanism the consolidation says it removes and violating the target's one-passage-per-section requirement | Keep each operative discharge, hold transition and record destination at its declared owner and make every other occurrence a pure reference +MINOR | high | standing §5 "the other closure and stop predicates" | The live sentence at CLAUDE.md 181–185 and workflow-init.md 388–392 presents assigned-fix-set membership, a new structural question, an accepted Blocker or Major and the tell thresholds as the other closure and stop predicates without marking the enumeration illustrative; the target leaves it unchanged while adding source blocks, holds, no-clean-credit and gate-specific content conditions | A reader entering through the standing floor-coherence paragraph can treat its four-item apposition as the closure surface and omit conditions the new ordering requires | Mark the standing list illustrative and point to the closure ordering for the complete classification +MINOR | high | §A1 "derived floor ... discharged" | The duties paragraph says the floor is discharged only by enough valid passes with the last one clean, although the same section defines the numeric floor and pass cleanliness as separate predicates combined by eligibility; a dirty pass can satisfy the required pass count without satisfying cleanliness | The text gives two accounts of what discharges the floor and weakens the story's requirement that every duty have one visible classification | Say the floor is discharged by the required valid-pass count, retain the zero-finding exception, and leave the clean-pass requirement solely in eligibility +MINOR | high | §H "the last item being the addition" | H says the unknown-start list gains the suspensions and that the last item is the addition, but its displayed replacement now adds three separate clauses: suspension binding, exclusion of repeated-dismissal cleanliness, and the new gate-condition and pass-cost strict reading | The section's own mechanical accounting is false and can cause the plan's old-condition disposition to leave two new clauses unassigned | State that the list gains all three additions and name each in the introduction +NIT | high | §A1 "reporting will not converge" | The shared precedence rationale says D2 and D3 forbid reporting that a converged loop will not converge, but only the clearly-stuck reading makes that report; the two-tell rule reports observed tells and asks continue or stop without concluding non-convergence | The precedence is operative, but the reason does not justify its two-tell half and fails prompt-standard item 6 | Give two-tell precedence its own reason, such as that a completed cycle leaves no remaining cycle decision for its continue-or-stop question to suspend +END OF FINDINGS (9 total) diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index 89e632c..1973e32 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -97,6 +97,14 @@ remaining rules this change ships, §A3's index condition names a **tree-to-tree the explicit `headSha` rather than membership of a range, and §G's membership test reaches rules deciding termination **without** a pass, the Gate-B triviality skip being the only such route. +**A closing act that does not complete closes nothing and costs no pass** (pass 30 finding 4). A +failed commit or amend leaves the cycle open with every condition still established and is +retried once the concrete failure is repaired; no branch is taken, because nothing the review +reads has changed. Retrying is available only while that holds — a repair touching anything a +condition is read from re-establishes that condition first. **This adds no mechanism**: it says +what "the act was not performed" already means, against a text that otherwise routes a `git` +error into a full review pass. + **One gap is named and not closed** (pass 29 finding 2, and §I carries it): a Gate-A closing act is written from the effective index, so a staged edit to a review input **no source rule governs** — a cited story's acceptance criteria, say — is published by the closing commit unchecked. Profile @@ -195,6 +203,7 @@ to point at and so a reader can see the shape of the change without reading the | (i) when these rules bind | §H | | the Gate-A section and the gate-prompt template | the two senses of *clean* are separated; the cadence makes revision conditional; the broad-prompt instruction stops assuming the artifact is revised between passes, keeping its breadth demand | | the profiles section, the lens paragraph | its unchanged-list is scoped to the lens sets | +| the profiles section, the evidence-entry revalidation remedy | its fix-re-review-close instruction becomes conditional on the ordering selecting continuation, a non-closing pass taking any applicable suspension first (pass 30 finding 3). The eighth falsified standing sentence, and the sixth sharing §F's mechanism | **Two sentences are deliberately not edited**, named so nobody looks for them: the "Copy every record into the squash body" sentence inside the human-exception block, and the "records every diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index a9679b3..5a5308d 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -156,6 +156,16 @@ nothing, and is not thereby made unclean. Keeping the two apart is what lets the case where they disagree, which the branches do. **The order inside closure is fixed: every condition is established first, and only then is the closing act performed.** +**An act that does not complete has not closed the cycle, and costs no pass.** Where the commit or +amend fails — a permission, signing, hook or repository error — **the cycle stays open with every +condition still established, the failure is surfaced as the concrete command failure it is, and +the act is retried once that is repaired.** No branch below is taken and no further review pass is +owed, because nothing the review reads has changed: not the artifact, not a finding, not an answer, +not a condition. **Retrying is only available while that stays true** — where the repair touches +anything a closure condition is read from, that condition is re-established first and the ordinary +rules decide what that costs. Reading a failed act as an ordinary non-closing pass would spend a +review pass on a `git` error and report a loop-health problem the loop does not have. + **Closure introduces no new kind of record, and it excuses none**: every other record this cycle owes, a human-exception record among them, is owed and written exactly as before, and **a human-exception record this cycle owes goes in the commit its closing act uses**, so the two never @@ -399,7 +409,9 @@ final pass's requests named** — a tree compared with a tree, since a range is can be inside and a staged revert at a path the range already touches would otherwise read as covered. A difference is not a failed review — it is content the request could not reach: fold it into the `WIP:` snapshot and -re-review. **This gate's re-review duty is widened here to say so**, the standing rule requiring a +re-review, **which is what the pass does once the closure ordering selects its continue branch**; +the mismatch means the pass did not close, so any suspension applying to it is answered first. +**This gate's re-review duty is widened here to say so**, the standing rule requiring a re-review after every *fix* and an index difference not being one; the pass run over the widened snapshot becomes the candidate final one. **Nothing here establishes what either branch actually consumed** — the reply reports no reviewed revision, which is why the kept `baseSha` and `headSha` @@ -629,10 +641,10 @@ demotes it — the two counts are meant to differ. --- -## F. The seven standing sentences this change falsifies — REPLACED +## F. The eight standing sentences this change falsifies — REPLACED -Each is a live sentence that the block makes wrong. All seven are **known contradictions** and -none is deferred. **Five of them share one mechanism** — an entry point other than the ordering +Each is a live sentence that the block makes wrong. All eight are **known contradictions** and +none is deferred. **Six of them share one mechanism** — an entry point other than the ordering carrying an unqualified instruction — which is why each is **replaced** rather than given an exception to point at. @@ -734,6 +746,21 @@ plan commit", which is the closing commit only on the first of the three closing other two the record and the closure would land in different commits. Naming the closing act instead keeps them together on all three without changing the record's form or force. +**8. The evidence-entry revalidation remedy** (the profiles section). It wraps across C 728–729 and +W 914–915. +``` +If revalidation changes the entry, the clean pass no longer covers what is being committed: the +pass did not close, so the closure ordering decides what happens next — fix, re-review and close +on the entry that pass validated where it selects the continue branch, and where it selects a +suspension, the answer comes first. +``` +*Why (pass 30 finding 3):* the live sentence is a complete instruction to whoever enters through +the profiles section — fix, re-review, close — and under the ordering a non-closing pass takes any +applicable suspension first. An eligible pass with stale evidence and two tells could follow this +path straight past a mandatory stop. It is the sixth sentence of this section's shared mechanism, +an entry point other than the ordering carrying an unqualified instruction, and it takes the same +repair: the **operation** stays here, the **permission** is the ordering's. + --- ## G. The one-contract paragraph — REPLACED @@ -852,7 +879,9 @@ A **clean findings file** is the single body line `NO FINDINGS` with `END OF FIN **The Gate-A cadence** — revision becomes conditional, since unconditional it tells a Minor-only pass to manufacture the repair the severity rule forbids. ``` -Each pass: validate, revise **where a repair is required**, re-run. +Each pass: validate, revise **where a repair is required**, and re-run **where the closure +ordering selects its continue branch** — where it selects a suspension instead, the answer comes +first and that ordering says what the answer produces. ``` **The lens paragraph's unchanged-list** — scoped to the lens sets, which is what that paragraph is From 45b7d366eea3a10d524ad11835b08589737a622a Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sun, 13 Sep 2026 11:55:27 +0200 Subject: [PATCH 058/181] docs(specs): apply pass 31; shrink the failed-act rule instead of extending it MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit One tell: findings 9 to 13. Blockers 0 to 0, which is not the falling tell — a zero cannot fall. Below the mandatory threshold. Three of the six Majors are damage from the pass-30 failed-act repair, which I adopted from the reviewer's suggested fix without testing the fix itself. It claimed every closure condition stayed established, which a pre-commit hook that modifies and stages content before failing falsifies, and it added a "no branch below is taken" clause that contradicted the pass-29 rule. Both are withdrawn. What stands is smaller: nothing is assumed about what a failed attempt left behind, every condition is re-established against the repository as it stands, the act is performed again where they hold, and anything the attempt or its repair moved is that condition's own rule to price. No branch, no mechanism. The other three: - §C carried the standing "a missing one means keep going" unqualified; the clearly-stuck reading failing decides only that this exit does not apply - §F item 8 commanded a later close after re-review without saying the new pass traverses the ordering like any other - §A3 named the act as the amend only, while standing Mechanics also prescribes a reset and single commit where several WIP snapshots piled up - §A3's index condition is read on the tree the act produces, not only on the index before it, since a hook can stage content during the commit Pre-review sweep, both failure classes this round cost: no unqualified continuation instruction and no guarantee claim about the act or the conditions remains in any prompt block. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-31.md | 14 +++++++ ...26-09-10-loop-rule-consolidation-design.md | 17 ++++---- ...-10-loop-rule-consolidation-target-text.md | 40 ++++++++++--------- 3 files changed, 46 insertions(+), 25 deletions(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-31.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-31.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-31.md new file mode 100644 index 0000000..e411150 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-31.md @@ -0,0 +1,14 @@ +MAJOR | high | §C "a missing one means keep going" | The replacement carries the standing unqualified continuation from CLAUDE.md 230 and workflow-init.md 433 even though failure of the clearly-stuck reading decides only that this suspension does not apply; the same pass may still have a scope stop or two-tell stop, or may close cleanly | An agent entering through the clearly-stuck paragraph can bypass another mandatory suspension or run another pass after the cycle should close | Replace "means keep going" with "means this reading does not apply" and send the pass back to the closure ordering +MAJOR | high | §F item 8 "fix, re-review and close" | The replacement conditions the sequence only on the current stale-evidence pass taking the continue branch, but still commands a later close after re-review without saying that the new pass must itself traverse the ordering; its own rationale claims that permission to close lives only in the ordering | A re-review that finds a Blocker, Major, scope trigger or health suspension can be followed by the source paragraph's direct close instruction | End the source remedy at fix and re-review, then state that the resulting pass is evaluated normally and closes only if the ordering reaches its closing act +MAJOR | high | §A1 "No branch below is taken" | The failed-closing-act rule says a pass whose act fails takes no later branch, while the suspension branch says without qualification that "A pass that did not close reaches this branch whatever its cleanliness" | A failed act on a pass carrying two tells or a clearly-stuck reading has two incompatible next states: immediate retry versus mandatory suspension | Make failed-act retry an explicit exception in the suspension branch and in the later claim about every eligible pass that does not close +MAJOR | high | §A1 "the cycle stays open with every condition still established" | That assertion is false for the named hook-failure case: a pre-commit hook can modify and stage content before failing, so the failed act itself can change the effective index, artifact or another closure input even when the later repair changes nothing | The agent can rely on the categorical no-change claim, retry without re-review and close over content the final pass did not cover | Allow the no-pass retry only after checking that neither the failed act nor its repair changed anything a condition or review reads; otherwise re-establish the affected conditions and ordinary pass duties +MAJOR | high | §A3 "effective index tree at the closing act equals" | The tree equality is established before `git commit --amend`, but a successful pre-commit hook can modify and stage content during that command; the target forbids reading a closure condition on the resulting branch tip and provides no other check of the committed tree | Gate B can close on a commit tree different from the reviewed `headSha` tree, making §I's claim that Gate B's equivalent is answered false | Make successful completion of the Gate-B act require the resulting commit tree to equal the final pass's explicit `headSha` tree, with a mismatch leaving the cycle open for the ordinary fold-and-re-review path +MAJOR | high | §A3 "The act is the closing amend" | The standing Mechanics sentence immediately after the replaced lead-in wraps across CLAUDE.md 829–830 and workflow-init.md 1013–1014 and says "If several WIP snapshots piled up, `git reset --soft ` first, then commit once"; that is a reset-plus-new-commit act, not the amend A3 declares, and A1's failure rule names only commit or amend rather than a failing reset | A multi-WIP cycle receives contradictory closing operations, can leave WIP commits in history, and has no defined retry state if the reset step fails | Define Gate B's act by reference to the complete Mechanics closing operation, preserving both the one-WIP amend and multi-WIP reset-plus-commit cases, and apply failed-act handling to every step +MINOR | high | §H "every closure condition and pass-cost rule" | The unknown-start replacement names positive pass obligations but gives no decidable strict reading for the new rule that a failed closing act costs no pass; "owed rather than waived" cannot say whether the zero-pass retry applies or whether strict fallback requires another pass | Two agents resuming an unknown-start cycle can charge different pass costs for the same failed act | State explicitly whether the failed-act no-pass retry is available under unknown-start fallback and carry the reason for the conservative choice +MINOR | high | §A "one definition in the paragraph that owns them" | The declared ownership boundary remains false: E calls itself the only resolve-duty statement while A adds per-finding tracking and no-omission discharge rules; A owns hold discharge while B and H also state when holds end; and A and F both assign the human-exception record to the closing-act commit | The proposed prompt retains parallel normative descriptions of the same transitions, recreating the drift mechanism the consolidation says it removes and violating the requested one-passage-per-section property | Keep each operative transition at its declared owner and turn every other occurrence into a pure reference without repeating the rule +MINOR | high | standing §5 "the other closure and stop predicates" | The live sentence wraps across CLAUDE.md 181–185 and workflow-init.md 388–392 and presents assigned-fix-set membership, a new structural question, an accepted Blocker or Major and the tell thresholds as "the other closure and stop predicates" without marking the list illustrative; the target leaves it intact while adding holds, no-clean-credit, source blocks and gate-specific content conditions | A reader entering through the standing floor-coherence paragraph can treat the four-item apposition as the closure surface and omit conditions the new ordering requires | Mark the standing list illustrative and point to the closure ordering for the complete classification +MINOR | high | §A1 "derived floor ... discharged" | The duties paragraph says the floor is discharged by enough valid passes "with the last of them clean", although the same section defines floor attainment and cleanliness as separate predicates combined only by eligibility | A dirty pass that reaches the numeric floor both does and does not discharge the floor, so the duty has two incompatible definitions | Say the floor is discharged by the required valid-pass count, retain the zero-finding exception, and leave the clean-pass requirement solely in eligibility +MINOR | high | §G "any record §5 obliges a cycle to write" | The opening declares "slot naming" part of the contract, but the operative membership test reaches record production, identity or transport only for records the cycle is obliged to write; standing §5 explicitly makes the resume note and dispositions companion optional even though their nonce-bearing slot rules are part of slot naming | A downstream reader applying only the shipped text cannot decide whether the optional-record slot and nonce sentences are contract members, so partial adoption can be accepted by one reader and stopped by another and can reintroduce sibling-file overwrites | State explicitly whether optional cycle-record naming and identity rules are members, and make the opening enumeration and semantic test agree +MINOR | high | §H "the last item being the addition" | The section says the unknown-start list gains the suspensions and that its last item is the addition, but the displayed replacement adds three distinct clauses: suspension binding, unavailability of the repeated-dismissal exclusion, and the closure-condition and pass-cost strict reading | The target's own mechanical accounting is false and can leave two new clauses unassigned when the plan builds its disposition | Say that the list gains all three clauses and name each in the introduction +NIT | high | §A1 "What D2 and D3 forbid" | The shared rationale says both health precedences prevent reporting "will not converge" on a converged loop, but only the clearly-stuck reading makes that report; the two-tell rule reports observed tells and asks continue or stop | The precedence works, but the two-tell half lacks its actual reason and fails prompt-standard item 6 | Give the two-tell precedence its own reason: once the cycle has closed there is no running-cycle decision left for its continue-or-stop question to suspend +END OF FINDINGS (13 total) diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index 1973e32..6e11edf 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -97,13 +97,16 @@ remaining rules this change ships, §A3's index condition names a **tree-to-tree the explicit `headSha` rather than membership of a range, and §G's membership test reaches rules deciding termination **without** a pass, the Gate-B triviality skip being the only such route. -**A closing act that does not complete closes nothing and costs no pass** (pass 30 finding 4). A -failed commit or amend leaves the cycle open with every condition still established and is -retried once the concrete failure is repaired; no branch is taken, because nothing the review -reads has changed. Retrying is available only while that holds — a repair touching anything a -condition is read from re-establishes that condition first. **This adds no mechanism**: it says -what "the act was not performed" already means, against a text that otherwise routes a `git` -error into a full review pass. +**A closing act that does not complete has not closed the cycle, and a failed command is not a +pass outcome** (pass 30 finding 4). **The first wording of this was wrong and pass 31 said why**: +it claimed every condition stayed established, which a pre-commit hook that modifies and stages +content before failing falsifies, and it added a "no branch is taken" clause contradicting the +pass-29 rule that a non-closing pass reaches the suspension branch. **The rule was shrunk rather +than extended**: nothing is assumed about what a failed attempt left behind, every closure +condition is re-established against the repository as it stands, the act is performed again where +they hold, and where the attempt or its repair moved anything a condition is read from that +condition's own rule decides the cost. **It introduces no branch and no mechanism.** Recorded +because adopting a reviewer's suggested fix wholesale is what produced three Majors here. **One gap is named and not closed** (pass 29 finding 2, and §I carries it): a Gate-A closing act is written from the effective index, so a staged edit to a review input **no source rule governs** — a diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index 5a5308d..5d20c57 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -156,15 +156,15 @@ nothing, and is not thereby made unclean. Keeping the two apart is what lets the case where they disagree, which the branches do. **The order inside closure is fixed: every condition is established first, and only then is the closing act performed.** -**An act that does not complete has not closed the cycle, and costs no pass.** Where the commit or -amend fails — a permission, signing, hook or repository error — **the cycle stays open with every -condition still established, the failure is surfaced as the concrete command failure it is, and -the act is retried once that is repaired.** No branch below is taken and no further review pass is -owed, because nothing the review reads has changed: not the artifact, not a finding, not an answer, -not a condition. **Retrying is only available while that stays true** — where the repair touches -anything a closure condition is read from, that condition is re-established first and the ordinary -rules decide what that costs. Reading a failed act as an ordinary non-closing pass would spend a -review pass on a `git` error and report a loop-health problem the loop does not have. +**An act that does not complete has not closed the cycle**, and a failed command is not a pass +outcome. **Nothing is assumed about what the attempt left behind** — a hook can modify and stage +content before failing, so the attempt itself can move what a condition is read from — and +therefore: surface the concrete command failure, and **re-establish every closure condition against +the repository as it now stands.** Where they all still hold, perform the act again; no review pass +is owed, because nothing the review reads has changed. Where the attempt or its repair moved +anything a condition is read from, **that condition has changed and its own rule decides what it +costs**, a further pass included. This introduces no branch: the pass is where the branches below +put it, and a retry is simply the act being performed once its conditions hold. **Closure introduces no new kind of record, and it excuses none**: every other record this cycle owes, a human-exception record among them, is owed and written exactly as before, and **a @@ -404,8 +404,10 @@ says. **What the range does not reach, and the one condition this gate adds for it.** The reviewed range ends at a commit; **the closing amend commits the effective index**, and content staged before the final review sits in the index without being in the range, so it was never inside what the review -request selected and no condition above excludes it. **So: the effective index tree at the closing act equals the tree of the explicit `headSha` the -final pass's requests named** — a tree compared with a tree, since a range is not a thing an index +request selected and no condition above excludes it. **So: the tree the closing act commits equals the tree of the explicit `headSha` the final pass's +requests named** — read on what the act produces and not only on the index before it, since a hook +running during the commit can stage content of its own; where the act produces a different tree it +has produced unreviewed content and has not closed the cycle, which is the mismatch case below — a tree compared with a tree, since a range is not a thing an index can be inside and a staged revert at a path the range already touches would otherwise read as covered. A difference is not a failed review — it is content the request could not reach: fold it into the `WIP:` snapshot and @@ -420,9 +422,10 @@ still**: it is advisory and compares its own inputs across its own invocations, before the review call and still staged at the commit has not moved between them — an unmoved fingerprint says nothing changed since it last looked, never that anything reviewed it. -**The act** is the closing amend Mechanics · Finishing the cycle describes, performed once the -ordering reaches it — an eligible pass with every closure condition holding, **never a clean pass -on its own**. +**The act** is the one Mechanics · Finishing the cycle describes — the closing amend, or, where +several `WIP:` snapshots piled up, the reset and single commit that section prescribes instead — +performed once the ordering reaches it: an eligible pass with every closure condition holding, +**never a clean pass on its own**. **A commit the hook reads as cycle-closing is a Gate-B matter.** A non-`WIP` commit mid-cycle makes the hook drop its Gate-B review state and that gate's counter — an observation about the counter, @@ -545,7 +548,7 @@ clause left standing alone — its second half, the precedence sentence, moves i the second is new. ``` -So this exit needs three things **together**, and a missing one means keep going: a plateau +So this exit needs three things **together**, and a missing one means only that *this* exit does not apply — what the pass does instead is the closure ordering's, another suspension or a close being open to it: a plateau visible across passes (six or more is where the field saw one); an **affirmative judgement that coverage is sufficient**, stated — a known materially unreviewed area forbids this exit outright, and disclosing it does not license it; and **Blocker or Major findings that keep regenerating @@ -750,9 +753,10 @@ instead keeps them together on all three without changing the record's form or f W 914–915. ``` If revalidation changes the entry, the clean pass no longer covers what is being committed: the -pass did not close, so the closure ordering decides what happens next — fix, re-review and close -on the entry that pass validated where it selects the continue branch, and where it selects a -suspension, the answer comes first. +pass did not close, so the closure ordering decides what happens next: where it selects the +continue branch, fix and re-review, and where it selects a suspension, the answer comes first. The +pass that follows is read by that ordering like any other and closes only if it reaches closure, +on the entry revalidated for it. ``` *Why (pass 30 finding 3):* the live sentence is a complete instruction to whoever enters through the profiles section — fix, re-review, close — and under the ordering a non-closing pass takes any From c7cb99372a358b7822917c209c426e0554fc18e6 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sun, 13 Sep 2026 12:11:14 +0200 Subject: [PATCH 059/181] =?UTF-8?q?docs(context):=20Gate-A=20pass=2032=20?= =?UTF-8?q?=E2=80=94=20mandatory=20two-tell=20stop,=20repair=20collisions?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Valid pass: terminator correct, 15 finding lines, count matches, all severity tokens recognized. Trend: findings 13, 15. Blockers 0, 1. Majors 6, 5. Two tells — the count rose and the Blocker returned — which is the mandatory threshold. All 15 findings open, no pass credited clean. Over passes 29 to 32 Blocker+Major went 5, 4, 6, 6 while findings went 10, 9, 13, 15. Each round verified its repairs against a concrete case first and the pre-review sweep was widened twice. The count still did not fall. Three of this pass's six Blocker/Majors are collisions between repairs made in the three preceding rounds: pass 31's "tree the act produces" against pass 29's "established before the act"; the same comparison against the older "no closure condition is read on the branch tip"; and pass 31's failed-act retry against the two pass-credit sentences in §F item 1 and §A3. The artifact is dense enough that a correct local repair reliably falsifies a sentence elsewhere. Clearly-stuck stands at two of three — plateau and nameable regeneration, but coverage is not affirmable while findings still reach new §G ground. The five settled behaviour decisions have held. No pass since 27 has challenged one, and the Gate-A/Gate-B split has produced no finding since then. The defects are in the seams between repairs. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-32.md | 16 ++++++++ .../gate-a-spec-awsf1ec771-resume.md | 38 +++++++++++++++++++ 2 files changed, 54 insertions(+) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-32.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-32.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-32.md new file mode 100644 index 0000000..a6d8dc7 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-32.md @@ -0,0 +1,16 @@ +BLOCKER | high | §A1 "every condition is established first" / §A3 "tree the closing act commits" | A1 requires every closure condition to be established before the closing act, but A3 defines Gate B's added condition on the tree the act produces and says a hook can change that tree during the command, so the condition cannot be established in A1's required order | An otherwise eligible Gate-B pass with no suspension cannot close because its last condition is unknowable before the act, and it can only continue into the same state | Treat the tree equality already defined in A3 as a success postcondition of Gate B's closing act and scope A1's pre-act ordering to conditions that can be established before the act; keep the existing mismatch path +MAJOR | high | §A1 "No closure condition is read on the branch tip" | A3's Gate-B condition compares the produced tree with the explicit `headSha` tree, while the surviving Mechanics text defines that `headSha` as the full object name `HEAD` resolved to at the call, which is the branch tip at that read point | The target simultaneously denies and requires a branch-tip-derived closure input, so a reader can discard the comparison or misstate what it proves | Delete the categorical sentence or narrow it to the post-act branch tip while expressly leaving A3's explicit-`headSha` comparison intact +MAJOR | high | §A1 "Those three sources answer a change" | The standing cited-set source re-reads the current set at each pass and before final acceptance and invalidates a header change during a pass, but it does not say that a transient cited-set change after the pass and later undone costs a pass; A1 nevertheless groups the cited set with the profile and assigned fix set as sources that answer any change rather than a differing value | A1 silently adds a pass-cost rule at the ordering instead of citing the source, contradicting its ownership claim and extending the settled changed-then-undone decision beyond the assigned fix set | Remove the cited set from the changed-not-different assertion and leave its current-set behavior at its standing source +MAJOR | high | §A1 failed act / §F item 1 and §A3 hook reset | A1 says an unchanged failed closing attempt owes no review pass, while F item 1 says a non-WIP attempt loses accumulated pass credit and A3 says the hook reset destroys Gate-B pass credit; the hook reset occurs on the command even when the commit fails, so the same failed attempt has two stated pass costs | An agent can either retry immediately or rerun the floor, and the retry path no longer has one determinate state | Call the lost state the hook's advisory counter state rather than logical pass credit, preserving A1's no-pass retry when no closure or review input moved +MAJOR | high | §G "wording of a prompt" | The membership test excludes "the wording of a prompt" as merely shaping findings, then says a sentence inside a prompt that fixes a valid input is a member; the standing gate-prompt sentences that define finding-line and `NO FINDINGS` validity satisfy both descriptions | A downstream reader cannot decide whether those prompt-format sentences belong to the one contract, defeating the authorised paragraph's required membership test | Narrow the exclusion to prompt wording that only frames review questions, leaving prompt sentences that define valid ordering inputs inside the contract +MAJOR | high | §G "slot naming" / obligatory-record test | The opening expressly includes slot naming in the contract, but the operative test reaches record production, identity or transport only for records §5 obliges a cycle to write; the surviving resume note and dispositions companions are optional while their nonce-bearing slot rules are part of slot naming | Two downstream readers can classify the optional slot and nonce sentences differently, so partial adoption can recreate sibling-file collisions without a determinate stop | Extend the semantic test to production, identity or transport of any §5 cycle record, whether optional or required, without adding a checker or durability mechanism +MINOR | high | standing §5 "the other closure and stop predicates" | The untouched floor-coherence paragraph presents assigned-fix-set membership, a new structural question, an accepted Blocker or Major and the tell thresholds as "the other closure and stop predicates" without marking the list illustrative, while the target adds holds, no-clean-credit, source blocks and gate content conditions | A reader entering through the standing paragraph can treat its four-item apposition as the closure surface and omit conditions the new ordering requires | Mark the standing list as illustrative and point to the closure ordering for the complete classification +MINOR | high | §A ownership boundary | A says the resolve duty keeps its definition only in Severity and that A owns hold discharge, yet A adds per-finding resolve tracking and no-omission discharge rules, B and H also say when holds end, and A and F both assign a human-exception record to the closing-act commit | The target retains parallel normative descriptions of the same transitions, violating its one-owner claim and recreating the drift mechanism the consolidation is meant to remove | Keep each operative rule at its declared owner and make the other occurrences pure references that do not repeat the transition +MINOR | high | §A1 "derived floor ... discharged" | The duties paragraph says the floor is discharged only by enough valid passes "with the last of them clean", although the same section defines floor attainment and cleanliness as separate predicates combined by eligibility | A dirty pass reaching the numeric floor both does and does not discharge the floor, leaving the duty classification internally inconsistent even though a later clean pass may mask it | Define floor discharge by the required valid-pass count, retain the zero-finding exception, and leave the clean-pass requirement solely in eligibility +MINOR | high | §H unknown-start pass costs | The strict-reading clause says every pass-cost rule this change ships is "owed rather than waived", while A1 says an unchanged failed closing act owes no review pass; H names only the assigned-fix-set pass cost and does not tell a reader whether strict fallback reverses A1's retry result | Agents resuming an unknown-start cycle can charge different review costs for the same unchanged failed act, reintroducing the wider failed-act rule withdrawn at pass 31 | State narrowly that unknown-start does not itself add a pass after a failed act and that A1 still re-establishes every current closure condition before retry +MINOR | high | §H "the last item being the addition" | The introduction says the unknown-start list gains the suspensions and that its last item is "the addition", but the displayed replacement adds three distinct clauses: suspension binding, unavailability of the repeated-dismissal exclusion, and the closure-condition/pass-cost strict reading | The target's own enumeration is mechanically false and can leave two new clauses unassigned when the plan performs condition disposition | Say the list gains all three clauses and name them in the introduction +NIT | high | §A1 "source-block branch" | The text names one of its four branches "the source-block branch" and requires cross-references to use branch names, then says that branch "needs no name" | The ownership description contradicts itself and makes references to this branch look forbidden and required at once | Replace "needs no name" with the narrower claim that the source condition needs no new subprocedure +NIT | high | §A1 "What D2 and D3 forbid" | The shared rationale says D2 and D3 both prevent reporting "will not converge" on a converged loop, but only the clearly-stuck reading makes that report; the two-tell rule reports tells and asks whether to continue or stop | The two-tell precedence carries a false reason and misses its actual reason under prompt-standard item 6 | Give the two-tell case its own reason: after closure there is no running-cycle decision left for its continue-or-stop question +NIT | high | target-text introduction "byte-identical" | The introduction says both prompt copies take every NEW and REPLACED section byte-identically, while §B is marked REPLACED and explicitly leaves the CLAUDE-only field-mint sentence outside its common replacement span | The file's installation metadata overstates parity and disagrees with its own declared retained divergence | Say every proposed replacement block is installed byte-identically, except content explicitly marked as copy-specific and carried unchanged +NIT | high | §H "§B ... is the only place this file states anything about it" | Sections A, F and H themselves discuss passage B's fix set, scope triggers and replacement status, so the claim is false even though §B is the only place that writes out the replacement bytes | The section-ownership metadata fails its own literal test and can make legitimate cross-references look like duplicate replacement text | Say §B is the only place that writes out passage B's replacement bytes +END OF FINDINGS (15 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index 5b5f01b..7764164 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -55,6 +55,44 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 26 | 8b8e146 | 10→**14** | 0→**1** | 2→**5** | yes | **TWO-TELL STOP — mandatory, surfaced to Daniel.** **All six Blocker/Majors trace to repairs made in passes 23, 24 and 25**, each nameable. All 14 held open; session 01a09539-5b8c-7293-af82-cc5f3bd2ba48 | | 27 | d971ae7 | 14→**8** | 1→**1** | 5→**3** | yes | **restarted by Daniel 2026-09-13 after the gate-split counter-draft was applied. B+M 6→4; findings the second-lowest of the cycle. ONE tell — no mandatory stop.** Three of the eight are collected Minors knowingly left unrepaired (5, 7, 8); session 01a099dd-b906-7970-8b39-8a3498510af4 | | 28 | 233e915 | 8→**13** | 1→**4** | 3→**2** | yes | **MANDATORY TWO-TELL STOP.** B+M 4→6. Four of the six trace to the pass-27/28 repair rounds (2, 5 damage; 3, 6 carry-through); 1 and 4 are rediscovered. All 13 held open; session 01a099fb-0706-7c53-ab6e-a0ccd8e191d3 | +| 29 | 36db7f0 | 13→**10** | 4→**0** | 2→**5** | yes | zero tells; all 5 Majors verified by counter-case; session 01a09a0a-c844-7792-b136-f779a4b82730 | +| 30 | 95439f8 | 10→**9** | 0→**0** | 5→**4** | yes | zero tells; 3 of 4 are carry-through of the pass-29 decision — my pre-review sweep was too narrow; session 01a09a16-99f4-7ca1-973e-2ac344f65a7c | +| 31 | e48259d | 9→**13** | 0→**0** | 4→**6** | yes | one tell; 3 of 6 are damage from the pass-30 failed-act repair, adopted from the reviewer's suggested fix without testing the fix. Rule shrunk, not extended; session 01a09a23-e45f-7f80-b6b0-d7dab19cac19 | +| 32 | 45b7d36 | 13→**15** | 0→**1** | 6→**5** | yes | **MANDATORY TWO-TELL STOP.** Findings rose and the Blocker returned. **Three of six B+M are collisions between my own repairs of passes 29–31.** All 15 held open; session 01a09a31-a7bb-7a03-9437-6c9396908fdb | + +## Pass-32 report — MANDATORY TWO-TELL STOP, and the repair strategy is the subject + +**Floor line:** derived floor **3**; risk **high**, security **none**; read fresh from +`docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md`. One cited story, level 2. + +- **Trend:** findings 9, 13, **15**. Blockers 0, 0, **1**. Majors 4, 6, **5**. + Blocker+Major 4, 6, **6**. +- **Cluster (pass 32):** product behaviour 11 of 15; prose about this file's own metadata 4 + (11, 12, 14, 15); the instrument 0. +- **require↔withdraw:** none, on the strict reading. But findings 2 and 4 are **my own repairs + contradicting each other**, which the five tells do not have a name for. + +**Tells: two of five — mandatory.** The finding count rose 13 → 15 and the Blocker count failed to +fall, 0 → 1. + +**The reading that matters, and it is not the tell count.** Over passes 29–32 the Blocker/Major +count went **5, 4, 6, 6** while findings went **10, 9, 13, 15**. Every round's repairs were +verified against a concrete case first, the pre-review sweep was widened twice, and the count still +did not fall. **Three of this pass's six B+M are collisions between repairs I made in the three +preceding rounds** — the pass-31 "tree the act produces" against the pass-29 "established before +the act" (Blocker 1); the same tree comparison against the older "no condition is read on the +branch tip" (finding 2); the pass-31 failed-act retry against §F item 1's and §A3's pass-credit +sentences (finding 4). **The artifact is now dense enough that a correct local repair reliably +falsifies a sentence elsewhere in it.** + +**Clearly-stuck stands at two of three.** Plateau: yes, across six passes. Regeneration across +genuine repair attempts: yes, and nameable per finding. **Coverage: not affirmable** — findings 5 +and 6 reach §G membership ground no earlier pass read. Two of three is not that exit, and the +two-tell threshold stands alone anyway. + +**What is not in doubt.** The five settled behaviour decisions have held: none of the last four +passes challenged one, and the Gate-A/Gate-B split has produced no finding since pass 27. The +defects are in the seams between repairs, not in the decisions. ## Pass-28 report — MANDATORY TWO-TELL STOP From 3ee9132ac7f57dcb6b2e906cd919b852e7396334 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sun, 13 Sep 2026 13:18:07 +0200 Subject: [PATCH 060/181] docs(specs): cut the Gate-B tree condition; apply the four surviving Majors MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Daniel answers the pass-32 mandatory stop with continue under a bounded scope cut, on 2026-09-13. The Gate-B tree-equality condition and its widened re-review duty leave this delivery. The gap they addressed is real and recorded in §I as a later task, not worked out: content staged before the review, or staged by a hook during the commit, reaches the closing commit through neither review branch. The price is stated rather than implied — this delivery does not close that gap. Gate B's existing review, re-review and evidence duties are unchanged and are not a substitute, and the gate hook is not one either. Two of pass 32's six Blocker/Majors go with it, both collisions the condition created: against the ordering's pre-act requirement, and against the branch-tip sentence. Faster convergence is plausible, not guaranteed. Three corrections to my own pass-32 report, made after the reviewer checked it against the files and recorded in the resume note: - three findings did not hang on the tree condition; two did. Finding 4 is §A1's failed-act rule against §F item 1's pass-credit sentence, and neither mentions the tree. It is expressly not resolved by the cut - §G was not new ground: pass 32 finding 6 is pass 31 finding 11 re-raised. Coverage stays unaffirmable, but not for the reason I gave - "this is no longer negligence" is withdrawn The four surviving Blocker/Majors, each tested against the standing conditions before the suggested fix was adopted: - the cited set does not price a change since undone; its source says the final clean pass runs against the current set, which an undone change satisfies. Only the profile and the assigned fix set answer a change as such - what a stray non-WIP commit costs is the hook's count, not the passes: a valid pass is established by its validated findings file, which no commit touches. Both sites now say so - §G's prompt-wording exclusion is narrowed to wording that only frames the review question; wording that fixes a valid input is a member - §G's record clause reaches any §5 cycle record, optional or required, the optional companions' slot rules carrying the nonce that keeps siblings apart Collected, not repaired: the ownership boundary, the standing four-item predicate list, §H's enumeration of its own additions, and three metadata NITs. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-resume.md | 16 ++++- ...26-09-10-loop-rule-consolidation-design.md | 45 +++++++------ ...-10-loop-rule-consolidation-target-text.md | 67 +++++++++---------- 3 files changed, 68 insertions(+), 60 deletions(-) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index 7764164..8269e20 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -86,9 +86,19 @@ sentences (finding 4). **The artifact is now dense enough that a correct local r falsifies a sentence elsewhere in it.** **Clearly-stuck stands at two of three.** Plateau: yes, across six passes. Regeneration across -genuine repair attempts: yes, and nameable per finding. **Coverage: not affirmable** — findings 5 -and 6 reach §G membership ground no earlier pass read. Two of three is not that exit, and the -two-tell threshold stands alone anyway. +genuine repair attempts: yes, and nameable per finding. **Coverage: not affirmable** — but **not +for the reason first given here.** *(Corrected 2026-09-13 after the reviewer checked it.)* This +section claimed findings 5 and 6 reached §G ground no earlier pass had read. **Finding 6 is pass +31's finding 11 re-raised** — the same optional-record complaint — so it is not new ground and +proves no earlier coverage gap. It proves no sufficiency either, which is why coverage stays +unaffirmable. Two of three is not that exit, and the two-tell threshold stands alone anyway. + +**Two further corrections to this report, made after the reviewer checked it against the files.** +**Three of six B+M do not hang on the Gate-B tree condition — two do.** Finding 4 is §A1's +failed-act rule (target 159) against §F item 1's "discards the passes you just accumulated" +(target 654); neither mentions the tree, and the conflict survives any cut of it. And **"this is +no longer negligence" is withdrawn**: the contradicting sentences stood in the same file, and +density explains the miss without making it unavoidable. **What is not in doubt.** The five settled behaviour decisions have held: none of the last four passes challenged one, and the Gate-A/Gate-B split has produced no finding since pass 27. The diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index 6e11edf..603c0e3 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -92,10 +92,10 @@ blocked by an unresolved prior Major, a standing hold or stale evidence walk pas two-tell stop and keep spending passes. **This narrows a rule this change itself wrote, and contradicts no settled decision**: D2 and D3 forbid reporting "will not converge" on a loop that converged, and a loop still owing a repair, an answer or a closure condition has not converged. -Three other pass-29 repairs are compliance rather than decision — the unknown-start list gains the -remaining rules this change ships, §A3's index condition names a **tree-to-tree** comparison against -the explicit `headSha` rather than membership of a range, and §G's membership test reaches rules -deciding termination **without** a pass, the Gate-B triviality skip being the only such route. +Two other pass-29 repairs are compliance rather than decision — the unknown-start list gains the +remaining rules this change ships, and §G's membership test reaches rules deciding termination +**without** a pass, the Gate-B triviality skip being the only such route. A third, sharpening the +index condition to a tree-to-tree comparison, **left with that condition** in the 2026-09-13 cut. **A closing act that does not complete has not closed the cycle, and a failed command is not a pass outcome** (pass 30 finding 4). **The first wording of this was wrong and pass 31 said why**: @@ -129,22 +129,27 @@ and its own closing act**, and neither gate's paragraph is an inventory of what that was is now Gate A's own condition stated at Gate A's paragraph. **The split buys ownership and not brevity** — the three paragraphs together run slightly longer than the single block did. -**Writing Gate B's closure down exposed one condition nobody had stated** (pass 27 finding 4). -Gate B reviews a `baseSha`..`headSha` range while the closing amend commits the **effective -index**, so content staged before the final review is in the index and in no reviewer's payload. -A3 adds the one condition that closes it: the effective index at the closing act carries nothing -outside the range the final pass's requests were aimed at, a difference being folded into the -`WIP:` snapshot and re-reviewed. **A3 widens this gate's re-review duty to say so and does not -attribute the breadth to the standing rule**, which requires a re-review after every *fix* and an -index difference not being one (pass 28 finding 11). **Gate A's mirror of the same hazard is not a -second condition**: a staged edit to a cited story or a profile header is published by A2's closing -commit, and the source rules already answer it — a profile or cited-set change costs a further -pass — so A2 says where those rules bite at the closing act and adds nothing (pass 28 finding 4). **The hook proves none of this, and A3 says so in -the terms invariant 3 fixes**: its fingerprint includes an effective-index tree, so content staged -before the review call and still staged at the commit leaves it unmoved between the two -invocations — an unmoved fingerprint reports that nothing changed since it last looked, never that -a review covered what it is looking at. That distinction is read from `AGENTS.md` invariant 3, -which states the comparison; nothing is claimed here about the script beyond it. +**The Gate-B tree-equality condition is deferred out of this change on Daniel's decision of +2026-09-13, as a bounded scope cut.** Writing Gate B's closure down had exposed a real gap (pass 27 +finding 4): the gate reviews a `baseSha`..`headSha` range while the closing amend commits the +**effective index**, so content staged before the review, or staged by a hook during the commit, +reaches the closing commit through neither branch. The condition written for it, and the widened +re-review duty it carried, **are removed from the target text**; the gap is recorded in §I as a +later task and is **not** worked out here. **The price is stated rather than implied**: this +delivery does not close that gap. Gate B's existing review, re-review and evidence duties stand +unchanged and are not a substitute, and **the gate hook is not one either** — its fingerprint is +advisory and compares its own inputs across its own invocations, so content staged before the +review call and still staged at the commit leaves it unmoved between the two. **Faster convergence +is plausible, not guaranteed.** What the cut removes is two of pass 32's six Blocker/Majors, both +collisions this condition created — against the ordering's pre-act requirement and against the +branch-tip sentence. **Pass 32's finding 4 is expressly not resolved by it**: §A1's failed-act rule +contradicts §F item 1's pass-credit sentence without either mentioning the tree. + +**Gate A's mirror of the same hazard is not a second condition**: a staged edit to a cited story or +a profile header is published by A2's closing commit, and the source rules answer what they govern — +a profile change costs a further pass — so A2 says where those rules bite at the closing act and +adds nothing (pass 28 finding 4). Where the staged edit touches a review input **no** source rule +governs, nothing reaches it, which §I carries as its own residual (pass 29 finding 2). **The closure-ordering block is an addition beside the source edits**, not one of them. §4 lists the edits; **no total is stated here or there**, because the unit — one contiguous replacement at diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index 5d20c57..33025b7 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -73,10 +73,11 @@ one of *those* defined here has found a defect. **The conditions every cycle has, whatever its gate.** The duties classified below; and the **profile**, the **cited set** and the **assigned fix set**, each gating as its own source says and -cited here without restatement. **Those three sources answer a change and not a differing value**, -which is why one of them changed and then undone still costs a pass — the contrast that matters -against a gate's own content condition, which may be a comparison of current values and says so -where it is stated. +cited here without restatement. **The profile and the assigned fix set answer a change and not a +differing value**, which is why one of them changed and then undone still costs a pass; **the +cited set answers as its own source says** — the final clean pass runs against the current set — +which a change since undone can already satisfy. The contrast that matters is against a gate's own +content condition, which may be a comparison of current values and says so where it is stated. **What a pass is read from.** Every finding-derived predicate reads the validated findings file **or files** of the logical pass as **the concatenation of their finding lines after each file has @@ -389,7 +390,8 @@ pass has run against — and committing already-reviewed text that was never com ``` **Gate B's content condition, and its closing act.** These are what Gate B adds to the conditions the ordering states for every cycle; that list is there and is not repeated here, so nothing below -is an inventory of what this gate requires. +is an inventory of what this gate requires. **This gate adds no content condition of its own**, +and the next paragraph says why. **Its content condition is not an artifact/request equality, and none is written for it.** Gate B reviews a **diff** identified by `baseSha` and `headSha` rather than a text handed to the reviewer, @@ -401,27 +403,6 @@ the same commit**, so the re-review covers both; the battery and the mode-derive profiles section obliges before a call; and the **evidence entry**, revalidated as that section says. -**What the range does not reach, and the one condition this gate adds for it.** The reviewed range -ends at a commit; **the closing amend commits the effective index**, and content staged before the -final review sits in the index without being in the range, so it was never inside what the review -request selected and no condition above excludes it. **So: the tree the closing act commits equals the tree of the explicit `headSha` the final pass's -requests named** — read on what the act produces and not only on the index before it, since a hook -running during the commit can stage content of its own; where the act produces a different tree it -has produced unreviewed content and has not closed the cycle, which is the mismatch case below — a tree compared with a tree, since a range is not a thing an index -can be inside and a staged revert at a path the range already touches would otherwise read as -covered. A difference is not a -failed review — it is content the request could not reach: fold it into the `WIP:` snapshot and -re-review, **which is what the pass does once the closure ordering selects its continue branch**; -the mismatch means the pass did not close, so any suspension applying to it is answered first. -**This gate's re-review duty is widened here to say so**, the standing rule requiring a -re-review after every *fix* and an index difference not being one; the pass run over the widened -snapshot becomes the candidate final one. **Nothing here establishes what either branch actually -consumed** — the reply reports no reviewed revision, which is why the kept `baseSha` and `headSha` -establish only that both calls were aimed at one range. **The hook's fingerprint establishes less -still**: it is advisory and compares its own inputs across its own invocations, and content staged -before the review call and still staged at the commit has not moved between them — an unmoved -fingerprint says nothing changed since it last looked, never that anything reviewed it. - **The act** is the one Mechanics · Finishing the cycle describes — the closing amend, or, where several `WIP:` snapshots piled up, the reset and single commit that section prescribes instead — performed once the ordering reaches it: an eligible pass with every closure condition holding, @@ -429,8 +410,9 @@ performed once the ordering reaches it: an eligible pass with every closure cond **A commit the hook reads as cycle-closing is a Gate-B matter.** A non-`WIP` commit mid-cycle makes the hook drop its Gate-B review state and that gate's counter — an observation about the counter, -since the cycle itself stays open until the conditions hold, so an accidental commit destroys -Gate-B pass credit and closes nothing. **It does not reach a Gate-A cycle's count**, which the hook +since the cycle itself stays open until the conditions hold, so an accidental commit resets what +the hook reports and closes nothing; **the passes themselves stand on their validated findings +files**. **It does not reach a Gate-A cycle's count**, which the hook clears at the skill boundaries that start a new Gate-A cycle rather than on any commit. ``` @@ -654,9 +636,11 @@ exception to point at. **1. The `WIP:` naming warning** (Mechanics · `baseSha`). ``` A pre-review snapshot named anything else reads to the hook as a real commit: the hook treats the -cycle as closed and **discards the passes you just accumulated**, while the cycle itself stays -open until the closure ordering's conditions hold. The cost of the mistake is the lost pass -credit, not a close nobody intended. +cycle as closed and **discards its count of the passes you just accumulated**, while the cycle +itself stays open until the closure ordering's conditions hold. **What is lost is the hook's +counter state and not the passes** — a valid pass is established by its validated findings file, +which no commit touches — so the cost is a reminder that now understates what you hold, not a +close nobody intended and not a review you have to run again. ``` *The sibling sentence in the profile-change paragraph needs no edit* — it already says such a commit "reads **to the hook** as the cycle closing", which claims no closure. @@ -780,14 +764,15 @@ a live rule belongs to this contract when what it says determines or supplies an closure ordering reads, which branch a pass takes, what a hold is or what discharges it, whether a cycle may close **or may terminate without running a pass at all, the Gate-B triviality skip being the one such route and its eligibility test therefore a member**, or the production, -identity or transport of any record **§5** obliges a -cycle to write — §5 entire and not the Mechanics subsection this paragraph sits in, record duties -being stated in both.** **Read it on the sentence, never on the section the sentence sits in.** +identity or transport of **any §5 cycle record, required or optional** — §5 entire and not the +Mechanics subsection this paragraph sits in, record duties being stated in both, and the optional +companions' slot rules carrying the nonce that keeps sibling cycles apart.** **Read it on the sentence, never on the section the sentence sits in.** A sentence is a member when **it itself** fixes one of those things — what counts as a valid finding line, which files or records are owed, what ends a hold. It is not a member when it only -shapes what a review produces, as the choice of reviewer, the lens set and the wording of a prompt -do: those change the findings without deciding what a finding *is* or what the ordering may do with -one. **No paragraph is exempt as a paragraph** — a sentence inside a routing or prompt paragraph +shapes what a review produces, as the choice of reviewer, the lens set and **prompt wording that only frames the +review question** do: those change the findings without deciding what a finding *is* or what the +ordering may do with one. **Prompt wording that fixes a valid input is a member**, the gate-prompt +sentences defining a finding line and the `NO FINDINGS` signal being exactly that. **No paragraph is exempt as a paragraph** — a sentence inside a routing or prompt paragraph that fixes a valid input or an owed file is a member, and a sentence anywhere that only influences the findings is not. The examples follow the test; they do not stand in for it. Asking instead what an imagined edit would do decides nothing, because @@ -920,6 +905,14 @@ be established as absent is cheaper to owe than to skip** — - **Partial adoption is instructed against, never detected.** §G says so in its own words. - **Which exit a cycle took is not observable from history.** The transport left with the record (successor story) and no story has taken it. An admitted gap, unowned. +- **Deferred out of this change on Daniel's decision of 2026-09-13: a Gate-B tree-equality + condition.** Its purpose was to close the gap between what the final review request selected and + what the closing commit carries — content staged before the review, or staged by a hook during + the commit, reaches the closing commit through neither review branch. **The gap is real and this + change does not close it.** Gate B's existing review, re-review and evidence duties are unchanged + and are not a substitute; **the gate hook is not one either**, its fingerprint being advisory and + comparing its own inputs across its own invocations. Named as a later task and deliberately not + worked out here. - **A Gate-A closing act can publish a review input the final pass never read** (pass 29 finding 2). The commit is written from the effective index, so a staged edit to a cited story's acceptance criteria or settled decisions — anything no source rule governs, profile values, From 6a2805e51cd58cc846cf26cbaa4c33db2ac8a736 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sun, 13 Sep 2026 13:30:26 +0200 Subject: [PATCH 061/181] =?UTF-8?q?docs(context):=20Gate-A=20pass=2033=20?= =?UTF-8?q?=E2=80=94=20mandatory=20two-tell=20stop,=20the=20cut=20did=20no?= =?UTF-8?q?t=20help?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Valid pass: terminator correct, 22 finding lines, count matches, all severity tokens recognized. Trend: findings 15, 22. Blockers 1, 1. Majors 5, 8. Two tells — the count rose and the Blocker failed to fall. Third mandatory stop in six passes. Across passes 28-33 Blocker+Major went 6, 5, 4, 6, 6, 9 and findings went 13, 10, 9, 13, 15, 22. The scope cut removed the two collisions predicted for it and the count still rose to its worst since pass 26. Three findings are references to the removed condition my sweep did not reach: §I still says Gate B's equivalent is answered by an index condition, design §2 says the same, and design §7's oracle still requires every precondition to hold through the closing commit. I grepped for the condition's own vocabulary and not for sentences relying on it — the third round running that a sweep was too narrow. Three more come out of this round's own repairs: the failed-act rule's "nothing the review reads has changed" against §I's admitted residuals, the same rule against the continue branch, and §F item 1's new "which no commit touches" against §A1's own statement that a commit attempt and its hooks may move what a condition is read from. Finding 7 is a ninth falsified standing sentence, at CLAUDE.md 752-753, verified against the file before being reported. Clearly-stuck stands at two of three: plateau and nameable regeneration, but coverage is not affirmable while design §7's oracle had not been read at this depth before. Every repair since pass 29 was verified against a concrete case, suggested fixes were tested before adoption and one withdrawn as too wide, the sweep was widened twice, and a scope cut removed a whole condition. Blocker+Major did not fall in any of them. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-33.md | 23 ++++++++++ .../gate-a-spec-awsf1ec771-resume.md | 42 +++++++++++++++++++ 2 files changed, 65 insertions(+) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-33.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-33.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-33.md new file mode 100644 index 0000000..511a388 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-33.md @@ -0,0 +1,23 @@ +BLOCKER | high | §§A1 and H "each gate's content condition" | A1 says each gate requires and states its own content condition, and H's unknown-start list owes the plural "gate content conditions and their commit-carry duty," while A3 says Gate B adds no content condition after the settled scope cut | An unknown-start Gate-B cycle following A1 and H owes an undefined condition that has no source, discharge, suspension or repair path, so it can neither close nor reach a determinate stop | Make both cross-gate statements singular and explicit about Gate A, and state that Gate B adds no content condition in this delivery +MAJOR | high | §A1 "no review pass is owed" | Re-establishing every closure condition does not establish the stated reason that "nothing the review reads has changed": §I admits a Gate-A review input outside those conditions, and the Gate-B scope cut admits closing content outside the final review range | A failed act can change one of those admitted uncovered inputs, leave every defined condition holding and be retried under a false claim that the final pass still covers the repository | Preserve the settled retry result but say only that failure by itself costs no pass; delete the inference about review inputs and leave each source rule and admitted residual at its stated width +MAJOR | high | §A1 failed closing act versus "Otherwise the continue branch" | The failed-act paragraph says to retry the act without another review pass when every condition still holds, but the continue branch says every pass that neither closes nor suspends runs another pass; an unchanged signing or permission failure satisfies both descriptions | The same failure has two incompatible next states, so an agent can waste a review pass or ignore the retry instruction | Clarify without adding a branch: after any applicable suspension is answered, retry the act when the re-established conditions still hold, and enter continue only when a source condition itself requires another pass +MAJOR | high | §C "another suspension or a close being open to it" | Failure of one clearly-stuck condition is sent back to the ordering but the sentence then enumerates only another suspension or a close, omitting the ordering's continue branch | A pass that is unclean or below the floor with no other suspension receives no stated outcome at this entry point and can be misread as closable or suspended | End the sentence after "what the pass does instead is the closure ordering's," or add continue to the non-exhaustive examples +MAJOR | high | §A3 "What holds that place" | A3 says the existing range, re-review and evidence duties hold the place of a Gate-B content condition, while §I and the settled scope cut say those duties are not a substitute for the removed tree-equality condition | A reader can treat the deferred staged-content gap as covered and close Gate B over content selected by neither final review branch | Say that Gate B adds no content condition and that the listed duties remain separately sourced duties; remove the claim that they hold the missing condition's place +MAJOR | high | §F item 1 "which no commit touches" | The replacement justifies retaining logical passes by claiming no commit touches their validated findings files, contradicting A1's explicit rule that a commit attempt and its hooks may modify what it leaves behind; the design also says this pass-32 conflict survived the tree-condition cut | A stray or failed non-WIP commit can alter or remove the evidence used to establish pass validity, so the text promises preserved passes in a state where that preservation is not known | State only the mechanically established effect that the hook resets its advisory counter, and remove the categorical findings-file and no-re-review claims +MAJOR | high | standing profiles paragraph "would discard the accumulated passes" | The live profile-change sentence, which wraps across CLAUDE.md 749–752 and workflow-init.md 935–938, still says a non-WIP commit "would discard the accumulated passes"; §F claims this sibling needs no edit, but only its preceding "reads to the hook" clause is qualified and the discarded object is still the logical passes | The installed prompt will simultaneously say the mistake loses only the hook count and loses the passes themselves, so agents can rerun or retain the floor on opposite readings | Replace the surviving object with the hook's accumulated pass count in both copies +MAJOR | high | §I Gate-A residual and design §2 "Gate B's equivalent is answered" | The final §I bullet and design lines 111–116 say Gate B's equivalent is answered by an index condition, contradicting A3, the preceding §I bullet and design lines 132–146, all of which say that condition was removed and the gap remains | The plan can treat the deferred gap as closed or reintroduce the withdrawn condition despite the explicit scope cut | Delete the stale Gate-B comparison from the Gate-A residual and update the earlier design paragraph to say both gates' named residuals remain unanswered at their stated widths +MAJOR | high | design §7 "every precondition ... through the closing commit" | The verification oracle requires every precondition to hold through the closing commit and attributes that window to the target block, but the target establishes conditions before the act and gives the through-act change window only to source rules that state one; Gate B's general post-act tree condition was removed | The plan is instructed to verify a stronger post-act closure rule than the shipped text contains, which can silently rebuild the removed condition or overclaim what the verification proves | Align the oracle with the target's source-specific read points and Gate A's explicit carry duty, without imposing through-commit persistence on every precondition +MINOR | high | §A2 "the one place ... where an undone change costs nothing" | A2 calls restored artifact equality the only condition where an undone change costs nothing, while A1 says a cited-set change later undone can already satisfy the standing current-set rule | Readers can charge an extra pass for a restored cited-set change even though its source says the final clean pass runs against the current set | Remove the uniqueness claim and state only that restored artifact bytes satisfy Gate A's equality condition +MINOR | high | standing §5 "the other closure and stop predicates" | The untouched sentence, wrapping across CLAUDE.md 181–185 and workflow-init.md 388–392, presents assigned-fix-set membership, a new question, an accepted Blocker or Major and tell thresholds as "the other closure and stop predicates" without marking the list illustrative; the target adds holds, no-clean-credit, source blocks and gate-specific closure rules | A reader entering through the surviving coherence paragraph can treat its four-item apposition as the closure surface and omit predicates the new ordering requires | Mark the live list illustrative and point to the closure ordering for the complete classification +MINOR | high | §§A1 and E resolve-duty ownership | E calls itself the resolve duty's only statement and A says the duty keeps its definition at Severity, yet A also defines per-finding discharge, cross-pass tracking and the rule that absence from a later findings file never discharges it | The promised one-owner structure still contains two normative definitions that can drift during partial adoption or later edits | Move the discharge and tracking clauses to Severity and leave A with the duty's classification and a reference +MINOR | high | §§A1, B and H hold discharge | A claims ownership of what discharges a hold, but B says the membership answer ends the membership hold and H says the hold stands until its answers are given | The same transition is specified in three proposed sections, violating the target's one-passage assignment and recreating the drift mechanism the consolidation is meant to remove | Keep the operative discharge rule in A and turn B and H into pure references to that rule +MINOR | high | §§A1 and F item 7 human-exception destination | A says a human-exception record goes in the commit used by the closing act, and F item 7 independently installs the same destination at the standing source | One record-placement rule has two authorities despite the requirement that no passage be described in more than one section | Keep the destination at the human-exception source and make A cite the records already owed without restating their destination +MINOR | high | §A1 "derived floor ... discharged" | The duties paragraph says the floor is discharged only by enough passes with the last one clean, although eligibility just above is explicitly the conjunction of floor attainment and pass cleanliness | A dirty pass at the numeric floor both satisfies and does not satisfy the floor, giving the standing duty two incompatible discharge definitions | Define floor discharge by the required valid-pass count, retain the zero-finding exit, and leave cleanliness solely in the eligibility predicate +MINOR | high | §H unknown-start "every ... pass-cost rule ... owed rather than waived" | The strict fallback says every pass-cost rule shipped here is owed, while A1 says an unchanged failed closing act owes no pass; H names only the assigned-fix-set cost and never says whether the fallback reverses the settled retry result | Two agents resuming an unknown-start cycle can charge different pass costs for the same unchanged failed act | Scope the fallback wording to the positive assigned-fix-set pass duty it actually adds, leaving A1's failed-act rule unchanged +MINOR | high | §H "the last item being the addition" | The introduction says the unknown-start list gains the suspensions and that its last item is "the addition," but the displayed replacement adds three distinct clauses: suspension binding, unavailability of the repeated-dismissal exclusion, and the closure-condition and pass-cost reading | The section's own enumeration is mechanically false and can make the plan leave two additions unassigned in the old-condition disposition | Say that the list gains all three clauses and name them in the introduction +NIT | high | §A1 "source-block branch" | The text names this branch and then says it "needs no name" | The section makes a mechanically false claim about its own four-branch structure | Remove "needs no name" while retaining that the source rule supplies the procedure +NIT | high | §A1 "a cross-reference anywhere ... names the branch" | The paragraph claims all cross-references use branch names, but the proposed text uses positional references including "branches below," "the branch above" and "either trigger above" | Reordering can break the very references the sentence says are stable | Replace the positional references with the clean-completion, suspension, continue or scope-trigger names +NIT | high | §§A1 and D two-tell precedence rationale | The shared reason says D2 and D3 prevent reporting "will not converge," but only clearly stuck makes that report; the two-tell stop reports observed tells and asks continue or stop | The precedence works, but the two-tell half lacks its actual reason under prompt-standard item 6 | Give two-tell precedence its own reason: a completed cycle leaves no running-cycle decision for its continue-or-stop question +NIT | high | §H "the only place this file states anything about" passage B | §H says §B is the only place the file states anything about passage B, although A cites and describes its fix set and triggers and §H itself discusses passage B; only the full replacement bytes are unique to §B | The section's assignment metadata is literally false and can mislead the plan's passage accounting | Say §B is the only place that writes out passage B's replacement bytes +MINOR | medium | §H Gate-A clean signal "without inspecting it" | The replacement says the NO FINDINGS signal lets a pass be read as clean without inspecting the file, while the standing protocol requires the file, body, terminator and count to be inspected and validated before any pass is accepted | An agent can read the sentence as permission to skip structural validation and accept a stale or malformed zero-finding file | Say the signal establishes zero findings without classifying finding lines, while file validation remains mandatory +END OF FINDINGS (22 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index 8269e20..45f56d3 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -59,6 +59,48 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 30 | 95439f8 | 10→**9** | 0→**0** | 5→**4** | yes | zero tells; 3 of 4 are carry-through of the pass-29 decision — my pre-review sweep was too narrow; session 01a09a16-99f4-7ca1-973e-2ac344f65a7c | | 31 | e48259d | 9→**13** | 0→**0** | 4→**6** | yes | one tell; 3 of 6 are damage from the pass-30 failed-act repair, adopted from the reviewer's suggested fix without testing the fix. Rule shrunk, not extended; session 01a09a23-e45f-7f80-b6b0-d7dab19cac19 | | 32 | 45b7d36 | 13→**15** | 0→**1** | 6→**5** | yes | **MANDATORY TWO-TELL STOP.** Findings rose and the Blocker returned. **Three of six B+M are collisions between my own repairs of passes 29–31.** All 15 held open; session 01a09a31-a7bb-7a03-9437-6c9396908fdb | +| 33 | 3ee9132 | 15→**22** | 1→**1** | 5→**8** | yes | **MANDATORY TWO-TELL STOP — third in six passes.** B+M 6→**9**, the worst since pass 26. **The scope cut did not reduce the count; it added cleanup debt.** Three findings are references to the cut condition my sweep did not reach. All 22 held open; session 01a09a7d-55f2-7513-b61b-09c9a34626d3 | + +## Pass-33 report — MANDATORY TWO-TELL STOP, and the scope cut did not help + +**Floor line:** derived floor **3**; risk **high**, security **none**; read fresh from +`docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md`. One cited story, level 2. + +- **Trend:** findings 15, **22**. Blockers 1, **1**. Majors 5, **8**. Blocker+Major 6, **9**. + Across passes 28–33: B+M **6, 5, 4, 6, 6, 9**; findings **13, 10, 9, 13, 15, 22**. +- **Cluster (pass 33):** the ordering and the gate paragraphs 13 of 22; this file's or the design's + own metadata 8; the verification instrument 1 (design §7's oracle). +- **require↔withdraw:** none. + +**Tells: two of five — mandatory. Third mandatory stop in six passes** (28, 32, 33). + +**The cut did not reduce the count.** It removed the two collisions it was predicted to remove, and +the pass still rose 15 → 22 with B+M at its worst since pass 26. **Three findings are references to +the removed condition that my sweep did not reach** — §I line 920 still says "Gate B's equivalent +**is** answered, its index condition being one of this change's decisions"; design §2 says the +same; design §7's oracle still requires every precondition to hold through the closing commit. +I grepped the target text for the condition's own vocabulary and did not grep for sentences that +merely *rely* on it, which is the third round running that a sweep was too narrow. + +**Three of the remaining Blocker/Majors come out of this round's own repairs.** The failed-act +rule's "nothing the review reads has changed" is falsified by §I's own admitted residuals +(finding 2); the same rule contradicts the continue branch (finding 3); and §F item 1's new "which +no commit touches" contradicts §A1's own statement that a commit attempt and its hooks may modify +what a condition is read from (finding 6). **Finding 7 is a ninth falsified standing sentence** — +the profiles paragraph at `CLAUDE.md` 752–753 still says a non-`WIP` commit "would discard the +accumulated passes", which this round's repair made false. Verified against the file before being +reported. + +**Clearly-stuck stands at two of three, again.** Plateau across six passes: yes. Regeneration +across genuine repair attempts, each round's fix producing the next: yes, and nameable per finding. +**Coverage: not affirmable** — design §7's oracle had not been read at this depth by any earlier +pass. Two of three is not that exit. + +**What five rounds of tightened method did not do.** Every repair since pass 29 was verified +against a concrete case first; suggested fixes were tested before adoption and one was withdrawn +for being wider than its defect; the pre-review sweep was widened twice; a scope cut removed a whole +condition. **The Blocker/Major count did not fall in any of them.** That is the finding this stop +carries, and it is about the approach rather than about any sentence in the artifact. ## Pass-32 report — MANDATORY TWO-TELL STOP, and the repair strategy is the subject From ea5e76c1d026bae72a13195208491642797685ea Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sun, 13 Sep 2026 14:47:23 +0200 Subject: [PATCH 062/181] docs(specs): finish the scope cut; drop the unsupported assurances MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Daniel answers the pass-33 mandatory stop with continue in the delivery scope already decided. No further cut, no redraft. Three corrections to my own pass-33 report, checked against the files and recorded in the resume note: - "Blocker+Major did not fall in any round" is false. P28 to P29 fell 6 to 5 and P29 to P30 fell 5 to 4. The supportable claim is no sustained convergence - findings 2 and 3 are not this round's damage. 3ee9132 did not touch the failed-act rule; it came from the pass-31 round. What 3ee9132 did newly add was the assurance "which no commit touches", which the repair commission never asked for and which finding 6 caught - "not locally repairable" is withdrawn. Line count does not establish it The cut is now finished consistently. It had been left half-done: §I said the gap is real and then said Gate B's equivalent is answered, §H still owed plural gate content conditions, §A3 said the existing duties hold the condition's place, and design §7's oracle still required preconditions to hold through the closing commit. All four are gone. Neither gate answers the hazard now, and both residuals are stated as residuals. The unsupported assurances are removed rather than qualified: no claim that a findings file is untouchable by any commit, and none that nothing the review reads has changed after a failed act. The hook reset is described only as what it reaches — its counter. The failed-act path is restated rather than excepted: the act is not a branch, the pass already took the clean-completion branch to reach closure, and a failed act returns to the closure step rather than to the branches. Where the attempt or its repair moved anything a condition is read from, that condition's own rule prices it and the cycle is back in the ordering with a pass owed. §C names all three continuations instead of two. §F gains a ninth falsified standing sentence — the profile-change paragraph's claim that a stray commit discards the accumulated passes, when what it reaches is the hook's count. Six of the nine still share the entry-point mechanism; this one does not. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-resume.md | 21 ++++-- ...26-09-10-loop-rule-consolidation-design.md | 18 +++-- ...-10-loop-rule-consolidation-target-text.md | 73 +++++++++++-------- 3 files changed, 70 insertions(+), 42 deletions(-) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index 45f56d3..ff4b0d2 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -96,11 +96,22 @@ across genuine repair attempts, each round's fix producing the next: yes, and na **Coverage: not affirmable** — design §7's oracle had not been read at this depth by any earlier pass. Two of three is not that exit. -**What five rounds of tightened method did not do.** Every repair since pass 29 was verified -against a concrete case first; suggested fixes were tested before adoption and one was withdrawn -for being wider than its defect; the pre-review sweep was widened twice; a scope cut removed a whole -condition. **The Blocker/Major count did not fall in any of them.** That is the finding this stop -carries, and it is about the approach rather than about any sentence in the artifact. +**What five rounds of tightened method did and did not do.** *(Corrected 2026-09-13 after the +reviewer checked it.)* This section claimed the Blocker/Major count "did not fall in any of them". +**That is false**: P28→P29 fell 6→5 and P29→P30 fell 5→4, by this table's own numbers. What the +evidence supports is the narrower claim — **no sustained convergence**, with the count returning to +6, 6 and then 9. + +**And the provenance given here for findings 2 and 3 was wrong.** They complain about the +failed-act rule, which `3ee9132` did not touch; it came from the pass-31 round. **What that commit +did newly add was the unsupported assurance "which no commit touches"** — a claim the repair +commission never asked for, which I wrote on my own initiative and which finding 6 then caught. +That is the accurate charge against the round. + +**"Not locally repairable" is withdrawn.** Line count does not establish it: the 923 lines carry +installation notes, rationales and history as well as rules, and the nine Blocker/Majors are +incomplete carry-throughs, over-wide claims and one contradictory path — none of which needs a new +product decision. ## Pass-32 report — MANDATORY TWO-TELL STOP, and the repair strategy is the subject diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index 603c0e3..6942a15 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -111,9 +111,9 @@ because adopting a reviewer's suggested fix wholesale is what produced three Maj **One gap is named and not closed** (pass 29 finding 2, and §I carries it): a Gate-A closing act is written from the effective index, so a staged edit to a review input **no source rule governs** — a cited story's acceptance criteria, say — is published by the closing commit unchecked. Profile -values, cited-set membership and the assigned fix set are governed and answer themselves. Gate B's -equivalent is answered because its index condition was decided; **Gate A's is a behaviour decision -nobody has made**, so the text states the residual instead of inventing a rule for it. +values, cited-set membership and the assigned fix set are governed and answer themselves. **Neither gate answers this after the +2026-09-13 cut**: Gate B's own version of the hazard is deferred with that condition, and Gate A's +has never been decided. The text states both residuals instead of inventing a rule for either. **The ordering is split into three paragraphs on Daniel's decision of 2026-09-12, and that split is the answer to pass 26's Blocker.** The ordering had stated one closure condition — the artifact's @@ -316,7 +316,9 @@ old-wording-gone half of its pair. **The obligation reaches every passage the ta REPLACED, and no list of them is kept here** — a second enumeration beside the markers is the bookkeeping that goes stale, which it did: the list this sentence used to carry omitted §G while §G was marked REPLACED. **The plan reads the markers off the target text**, where the concrete -replacements live. Nothing is +replacements live. **§F now states nine falsified standing sentences**, the ninth being the +profile-change paragraph's pass claim (pass 33 finding 7); six of the nine share the section's +entry-point mechanism and that one does not. Nothing is claimed as "contradictory" — the second of the two defects Gate B found in the `fic2` instrument. **The named verification of the risk path** (story AC 4) is a **next-state table**, written in the @@ -349,9 +351,11 @@ route the block states. **The closure conditions are read from the block and not here**: a re-enumeration is a second definition that drifts, and pass 15 found this list already missing two of them. Concretely the row must **enter closure from the clean-completion or zero-finding branch** — so a pass carrying a scope-stop trigger cannot close on the answer to that -trigger, no-clean-credit being the clean predicate's own second half — and every precondition the -block names must hold **through the closing commit**, not merely during the pass, which is the -window the block's own "in between" wording fixes. Naming only the distinct-state half would pass +trigger, no-clean-credit being the clean predicate's own second half — and every precondition the block names must hold +**when it is established, immediately before the closing act**, which is the window the block +fixes. **The oracle does not require a precondition to be re-read on what the act produces** — the +condition that would have demanded that left with the 2026-09-13 cut, and the gap it addressed is +recorded in the target text's §I rather than checked here. Naming only the distinct-state half would pass the exact no-progress defect AC 4 cites from the parent cycle. **The consumption clause is what keeps the oracle and the shipped text in agreement**: the target text's §A says continue consumes the reading that raised the suspension and a further health suspension diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index 33025b7..180aa61 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -63,9 +63,9 @@ wording, not of authority. It also owns the **classification** of the standing d their definitions at their own sources and are read here for every cycle whatever its gate, while the hold and no-clean-credit are defined here, being properties of the evaluation itself. -**What belongs to a gate is what the two gates do differently: the content condition each gate's -closure requires, and the closing act.** Each is stated in that gate's own paragraph below and read -from there, and **neither gate's paragraph restates the conditions this one gives every cycle**, so +**What belongs to a gate is what the two gates do differently: the content condition its closure +requires, where it has one, and the closing act.** **Gate A has such a condition and Gate B has +none**; each paragraph below states its own and is read from there, and **neither gate's paragraph restates the conditions this one gives every cycle**, so neither is a complete inventory on its own. **Everything else this paragraph names it cites**: the scope triggers, the assigned fix set, every severity rule and every closure precondition with a source of its own keep their one definition in the paragraph that owns them, and a reader who finds @@ -157,15 +157,15 @@ nothing, and is not thereby made unclean. Keeping the two apart is what lets the case where they disagree, which the branches do. **The order inside closure is fixed: every condition is established first, and only then is the closing act performed.** -**An act that does not complete has not closed the cycle**, and a failed command is not a pass -outcome. **Nothing is assumed about what the attempt left behind** — a hook can modify and stage -content before failing, so the attempt itself can move what a condition is read from — and -therefore: surface the concrete command failure, and **re-establish every closure condition against -the repository as it now stands.** Where they all still hold, perform the act again; no review pass -is owed, because nothing the review reads has changed. Where the attempt or its repair moved -anything a condition is read from, **that condition has changed and its own rule decides what it -costs**, a further pass included. This introduces no branch: the pass is where the branches below -put it, and a retry is simply the act being performed once its conditions hold. +**An act that does not complete has not closed the cycle**, and it is not a branch: the branches +below decide what a *pass* is, and this cycle's pass already took the clean-completion branch and +reached closure. **A failed act returns to the closure step it failed in, not to the branches.** +**Nothing is assumed about what the attempt left behind** — a hook can modify and stage content +before failing, so the attempt itself can move what a condition is read from — so surface the +concrete command failure and **re-establish every closure condition against the repository as it +now stands.** Where they all still hold, perform the act again. Where the attempt or its repair +moved anything a condition is read from, **that condition has changed and its own rule decides what +it costs**, a further pass included, and the cycle is back in the ordering with that pass owed. **Closure introduces no new kind of record, and it excuses none**: every other record this cycle owes, a human-exception record among them, is owed and written exactly as before, and **a @@ -395,8 +395,9 @@ and the next paragraph says why. **Its content condition is not an artifact/request equality, and none is written for it.** Gate B reviews a **diff** identified by `baseSha` and `headSha` rather than a text handed to the reviewer, -so there is no reviewed text to compare an artifact against. What holds that place is already in -this section and is cited rather than restated: the range those two names fix, **both branches +so there is no reviewed text to compare an artifact against, and **nothing is put in its place** — +the bullet in §I records what that leaves open. What this gate does have is already in this section +and is cited rather than restated: the range those two names fix, **both branches issued against the same commit**; **a re-review after every fix**, a fix changing the artifact so the prior review no longer covers it; **a fix that changes specified behaviour updating the spec in the same commit**, so the re-review covers both; the battery and the mode-derived evidence the @@ -411,8 +412,8 @@ performed once the ordering reaches it: an eligible pass with every closure cond **A commit the hook reads as cycle-closing is a Gate-B matter.** A non-`WIP` commit mid-cycle makes the hook drop its Gate-B review state and that gate's counter — an observation about the counter, since the cycle itself stays open until the conditions hold, so an accidental commit resets what -the hook reports and closes nothing; **the passes themselves stand on their validated findings -files**. **It does not reach a Gate-A cycle's count**, which the hook +the hook reports and closes nothing; **what a pass is stands on its validated findings file, not +on that counter**. **It does not reach a Gate-A cycle's count**, which the hook clears at the skill boundaries that start a new Gate-A cycle rather than on any commit. ``` @@ -530,7 +531,7 @@ clause left standing alone — its second half, the precedence sentence, moves i the second is new. ``` -So this exit needs three things **together**, and a missing one means only that *this* exit does not apply — what the pass does instead is the closure ordering's, another suspension or a close being open to it: a plateau +So this exit needs three things **together**, and a missing one means only that *this* exit does not apply — what the pass does instead is the closure ordering's — another suspension, a continue, or a close: a plateau visible across passes (six or more is where the field saw one); an **affirmative judgement that coverage is sufficient**, stated — a known materially unreviewed area forbids this exit outright, and disclosing it does not license it; and **Blocker or Major findings that keep regenerating @@ -626,9 +627,9 @@ demotes it — the two counts are meant to differ. --- -## F. The eight standing sentences this change falsifies — REPLACED +## F. The nine standing sentences this change falsifies — REPLACED -Each is a live sentence that the block makes wrong. All eight are **known contradictions** and +Each is a live sentence that the block makes wrong. All nine are **known contradictions** and none is deferred. **Six of them share one mechanism** — an entry point other than the ordering carrying an unqualified instruction — which is why each is **replaced** rather than given an exception to point at. @@ -637,13 +638,13 @@ exception to point at. ``` A pre-review snapshot named anything else reads to the hook as a real commit: the hook treats the cycle as closed and **discards its count of the passes you just accumulated**, while the cycle -itself stays open until the closure ordering's conditions hold. **What is lost is the hook's -counter state and not the passes** — a valid pass is established by its validated findings file, -which no commit touches — so the cost is a reminder that now understates what you hold, not a -close nobody intended and not a review you have to run again. +itself stays open until the closure ordering's conditions hold. **What the hook loses is its counter state**, and a +valid pass is established by its validated findings file rather than by that counter — so the cost +is a reminder that now understates what you hold, not a close nobody intended. ``` -*The sibling sentence in the profile-change paragraph needs no edit* — it already says such a -commit "reads **to the hook** as the cycle closing", which claims no closure. +*The sibling sentence in the profile-change paragraph claims no closure* — it says such a commit +"reads **to the hook** as the cycle closing" — **but its second half is falsified by this same +repair and is item 8 below.** **2. The Gate-B coverage instruction** (Gate B section). ``` @@ -733,7 +734,20 @@ plan commit", which is the closing commit only on the first of the three closing other two the record and the closure would land in different commits. Naming the closing act instead keeps them together on all three without changing the record's form or force. -**8. The evidence-entry revalidation remedy** (the profiles section). It wraps across C 728–729 and +**8. The profile-change paragraph's pass claim** (the profiles section). It wraps across C 752–753 +and W 938–939. +``` +Inside an active Gate-B cycle, fold the edit into the active `WIP:` snapshot by amend — a non-`WIP` +commit reads to the hook as the cycle closing and would discard **the hook's count of** the +accumulated passes. +``` +*Why (pass 33 finding 7):* the live clause says such a commit "would discard the accumulated +passes". A pass is established by its validated findings file; what the commit reaches is the +hook's counter. Left standing it tells an author that a stray commit destroyed review work it +cannot reach. **This one does not share the section's shared mechanism** — it is a false claim +about a mechanism rather than an entry point carrying an unqualified instruction. + +**9. The evidence-entry revalidation remedy** (the profiles section). It wraps across C 728–729 and W 914–915. ``` If revalidation changes the entry, the clean pass no longer covers what is being committed: the @@ -888,7 +902,7 @@ the curve duty owed, the nonce duties at their strictest, **every suspension bin starting rules that cannot be established cannot be read as having waived an open hold**, **the repeated-dismissal cleanliness exclusion unavailable, a cycle that cannot establish its starting rules being unable to establish that they contained it**, and **every closure condition and -pass-cost rule this change ships owed rather than waived — the gate content conditions and their +pass-cost rule this change ships owed rather than waived — Gate A's content condition and its commit-carry duty, and the further pass an assigned-fix-set change costs — since a rule that cannot be established as absent is cheaper to owe than to skip** — ``` @@ -917,7 +931,6 @@ be established as absent is cheaper to owe than to skip** — 2). The commit is written from the effective index, so a staged edit to a cited story's acceptance criteria or settled decisions — anything no source rule governs, profile values, cited-set membership and the assigned fix set being the ones that are governed — lands in the - closing commit unchecked. Gate B's equivalent **is** answered, its index condition being one of - this change's decisions; **Gate A's is not, and the asymmetry is deliberate only in the sense - that nobody has decided it.** Named here rather than answered, because answering it is a + closing commit unchecked. **Neither gate answers this**: Gate B's own version of the + hazard is the bullet above, deferred out of this delivery, and Gate A's has never been decided. Named here rather than answered, because answering it is a behaviour decision this change has not been given. From 8847414361d5081fe4381a51a6e595391c9e5fe8 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sun, 13 Sep 2026 15:04:15 +0200 Subject: [PATCH 063/181] docs(specs): apply pass 34; five Majors, each verified before repair MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Zero tells: findings 22 to 19, Blockers 1 to 0, Blocker+Major 9 to 5 — the largest fall of this stretch and the lowest since pass 29. Finishing the cut is what did it. - the cited set is described by what its source actually says: a governing header changed during a pass makes that pass not final, and the final clean pass runs against the current set. The pass-33 wording read it as a pure current-value test, which the standing floor paragraph contradicts. §A2 drops its uniqueness claim with it - the continue branch no longer reaches a source-blocked pass, which previously satisfied both branches and could run review passes an unresolvable profile forbids - §F item 9 stops implying the clean pass covered what the closing commit carries; after the cut, §I says it may not - §G's criterion is replaced, not qualified: a rule is a member when it defines an input's validity or how it is read. "Determines or supplies an input" literally included the reviewer choice and lens sets the next sentence excludes - §F item 5 now states the hook's actual behaviour, read at codex-gate.sh:876 rather than asserted: it clears Gate-B state on any non-WIP commit attempt, whether or not the command succeeded. A failed closing act therefore leaves the counter cleared and no commit made Design bookkeeping aligned: §1 no longer implies the target claims no total at all — it claims none over edit spans and does state nine falsified standing sentences — and §4 gains the profile-change site. Collected: the ownership boundary, the standing four-item predicate list, §H's enumeration of its own additions, and five metadata NITs. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-34.md | 20 +++++++++++++ .../gate-a-spec-awsf1ec771-resume.md | 24 +++++++++++++++ ...26-09-10-loop-rule-consolidation-design.md | 9 ++++-- ...-10-loop-rule-consolidation-target-text.md | 30 +++++++++++-------- 4 files changed, 67 insertions(+), 16 deletions(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-34.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-34.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-34.md new file mode 100644 index 0000000..fcc3f5b --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-34.md @@ -0,0 +1,20 @@ +MAJOR | high | §A1 "cited set" and standing §5 | A1 says a cited-set change that is later undone can already satisfy closure, but the standing sentence, wrapping across CLAUDE.md 116–119 and workflow-init.md 323–326, says "A header or profile that changed during that pass means the pass is not final"; §F does not replace it despite claiming all known contradictions are covered | The same restored header makes the pass final on one entry path and non-final on another, so a cycle can charge or skip the required further pass depending on which sentence the agent follows | Replace the standing sentence rather than append an exception: separate the event-based profile rule from the current-value cited-set rule, and include that replacement in the target text +MINOR | high | §§A1 and A2 "undone change" | A2 calls restored artifact equality "the one place" where an undone change costs nothing, while A1 says a cited-set change since undone can already satisfy its current-set condition | The uniqueness claim gives readers opposite pass-cost answers for a restored cited-set change | Delete the uniqueness claim and state only that byte-for-byte restoration satisfies Gate A's equality condition +MAJOR | high | §A1 source-block and continue branches | The source-block branch says no further pass runs while its source block stands, but the continue branch says every pass that neither closes nor suspends runs another pass and expressly includes an eligible pass with an unmet closure condition; a source-blocked pass with no suspension satisfies both | An agent can run review passes while an unresolvable profile, malformed governing header, or unreadable Story header still forbids them | Replace the continue catch-all so it applies only when no source block stands, preserving the existing source-block branch rather than adding another branch +MAJOR | high | §F item 9 "evidence-entry revalidation" | "The clean pass no longer covers what is being committed" implies that the unchanged entry meant the pass covered the closing commit, while A3 and §I expressly admit that the final Gate-B range can omit content carried by that commit after the tree-equality condition was removed | A reader can treat the deferred staged-content gap as covered by evidence revalidation and close over content selected by neither review branch | Replace the sentence with the narrower fact that the pass assessed a stale evidence entry and must be followed by the ordering; do not make any claim about coverage of the closing commit +MAJOR | high | §G "Membership is decided by a test" | The test includes any rule that "determines or supplies an input" the ordering reads, which literally includes reviewer choice and lens or review-prompt rules because they supply the findings, but the next sentences expressly exclude reviewer choice, lens sets, and prompt wording that frames the question | A downstream reader with only the shipped prompt cannot decide contract membership, so the partial-adoption stop has contradictory membership answers | Replace "determines or supplies an input" with a criterion limited to defining an input's validity or interpretation, and align the examples to that single criterion +MAJOR | high | §F item 5 "hook reads the amend" | The replacement says the hook reads the amend as "the real cycle-closing commit", but codex-gate.sh resets Gate-B state for every non-WIP git commit command regardless of whether the command succeeds; on the failed closing act A1 explicitly handles, no real commit exists | After a failed amend the agent can misread the hook reset as evidence of a completed commit or fail to understand why the advisory counter vanished before retry | Replace the mechanism claim with the exact behavior: the hook treats a non-WIP commit attempt as a Gate-B boundary and resets its state even when the command fails +MINOR | high | §A1 "derived floor" duty | The duty paragraph says the floor is discharged only by the valid-pass count "with the last of them clean", although the standing rule defines the floor as the pass count owed and A1 separately makes cleanliness part of eligibility | A dirty pass at the numeric floor simultaneously satisfies and fails the floor, making floor reports and later closure reasoning disagree | Define floor discharge by the required valid-pass count, retain the zero-finding exception, and leave last-pass cleanliness solely in eligibility +MINOR | high | §§A1 and E resolve-duty ownership | E calls itself "the only statement" of the resolve duty, but A1 also defines per-finding discharge, cross-pass tracking, the effect of omission from a later findings file, and the repeated-dismissal discharge | The promised one-owner structure has two normative definitions that can drift under implementation or partial adoption | Move the discharge and tracking rules to Mechanics · Severity and leave A1 with the duty's classification and a reference +MINOR | high | §§A1, B, and H hold discharge | A1 claims ownership of what discharges a hold, while B independently says the membership answer ends the membership hold and H says the hold stands until its answers are given | One transition is specified in three proposed sections, violating the target's one-passage assignment and recreating the drift mechanism the consolidation is meant to remove | Keep the operative discharge rule in A1 and replace the B and H formulations with references to it +MINOR | high | §§A1 and F item 7 human-exception destination | A1 says a human-exception record goes in the commit used by the closing act, and F item 7 independently installs the same destination at the human-exception source | The record-placement rule has two authorities despite the target's one-definition contract | Keep the destination at the human-exception source and replace A1's destination sentence with a citation to records already owed +MINOR | high | standing §5 "the other closure and stop predicates" | The untouched sentence, wrapping across CLAUDE.md 181–185 and workflow-init.md 388–392, presents assigned-fix-set membership, a new question, an accepted Blocker or Major, and tell thresholds as "the other closure and stop predicates" without marking the apposition illustrative; the target adds holds, no-clean-credit, source blocks, shared conditions, and a Gate-A condition | A reader entering through the surviving coherence paragraph can treat its four-item list as the closure surface and omit predicates the new ordering requires | Replace the live sentence so the list is explicitly illustrative and points to the closure ordering for the complete classification +MINOR | high | §H "Gate-A clean-signal sentence" | The replacement says NO FINDINGS lets a pass be read as clean "without inspecting it", while the standing acceptance protocol requires inspection of readability, body, terminator, count, and extra lines before any pass is accepted | An agent can skip structural validation and accept a stale, truncated, or malformed zero-finding file as clean | Say that the signal establishes zero findings after mandatory file validation, without requiring severity classification of finding lines +MINOR | high | §F and design §§1, 4, and 7 | The target states a total of nine falsified standing sentences and adds the profile-change replacement as item 8, while design §§1 and 4 say the target claims no total and the §4 edit-site map omits that profile-change site; design §7 then acknowledges the nine-item total | The implementation plan receives incompatible bookkeeping instructions and can omit the profile-change edit or preserve a count the design says must not exist | Add the profile-change site to the design's edit map and replace the stale no-total and nine-total claims with one consistent instruction derived from the concrete target list +MINOR | high | §H "unknown-start strict-reading list" | The introduction says the list "gains the suspensions" and calls its last item "the addition", but the displayed replacement adds three distinct clauses: suspension binding, unavailability of the repeated-dismissal exclusion, and the closure-condition and pass-cost reading | The section's own enumeration is mechanically false and can leave two additions unassigned in the old-condition disposition | Replace the introduction with a statement naming all three added clauses +NIT | high | §A1 "source-block branch" | The text names the source-block branch and then says it "needs no name" | The section makes a mechanically false claim about its own four-branch structure | Remove "needs no name" while retaining that the source rule supplies the procedure +NIT | high | §A1 "cross-reference anywhere" | A1 claims every cross-reference names a branch rather than its position, but the proposed text uses positional references including "branches below", "the branch above", and "either trigger above" | Reordering can break the references the sentence says are stable | Replace each positional reference with the clean-completion, suspension, continue, or scope-trigger name +NIT | high | §§A1 and D two-tell precedence rationale | A1 says D2 and D3 both forbid reporting "will not converge", but only the clearly-stuck exit makes that report; the two-tell stop reports observed tells and asks continue or stop | The precedence is operationally stated, but the two-tell rule lacks its actual reason under prompt-standard item 6 | Replace the shared rationale with separate reasons, using the false convergence report only for clearly stuck and the absence of a running-cycle decision after closure for two tells +NIT | high | §H "Passage (b) is not among them" | H says §B is "the only place this file states anything about" passage B, although A cites and describes B's fix set and triggers and H itself discusses the passage | The section's assignment metadata is literally false | Say §B is the only place that writes out passage B's replacement bytes +NIT | high | §A "Three paragraphs" | A says the installed addition is three paragraphs, but each of A1, A2, and A3 contains multiple blank-line-delimited Markdown paragraphs | The stated count does not match its own proposed text and can mislead the insertion plan about the replacement unit | Call A1, A2, and A3 three blocks or passages rather than three paragraphs +END OF FINDINGS (19 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index ff4b0d2..90970a6 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -60,6 +60,30 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 31 | e48259d | 9→**13** | 0→**0** | 4→**6** | yes | one tell; 3 of 6 are damage from the pass-30 failed-act repair, adopted from the reviewer's suggested fix without testing the fix. Rule shrunk, not extended; session 01a09a23-e45f-7f80-b6b0-d7dab19cac19 | | 32 | 45b7d36 | 13→**15** | 0→**1** | 6→**5** | yes | **MANDATORY TWO-TELL STOP.** Findings rose and the Blocker returned. **Three of six B+M are collisions between my own repairs of passes 29–31.** All 15 held open; session 01a09a31-a7bb-7a03-9437-6c9396908fdb | | 33 | 3ee9132 | 15→**22** | 1→**1** | 5→**8** | yes | **MANDATORY TWO-TELL STOP — third in six passes.** B+M 6→**9**, the worst since pass 26. **The scope cut did not reduce the count; it added cleanup debt.** Three findings are references to the cut condition my sweep did not reach. All 22 held open; session 01a09a7d-55f2-7513-b61b-09c9a34626d3 | +| 34 | ea5e76c | 22→**19** | 1→**0** | 8→**5** | yes | **zero tells.** B+M 9→**5**, the largest fall of this stretch and the lowest since pass 29 — finishing the cut is what did it. All 5 Majors verified before repair, one of them (§F item 5) against `codex-gate.sh:876` rather than asserted; session 01a09acf-1185-7883-b23f-3ee707d5ad58 | + +## Pass-34 report — zero tells, B+M 9 → 5 + +**Floor line:** derived floor **3**; risk **high**, security **none**; read fresh from +`docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md`. One cited story, level 2. + +- **Trend:** findings 22, **19**. Blockers 1, **0**. Majors 8, **5**. Blocker+Major 9, **5**. +- **Cluster (pass 34):** the ordering and the gate paragraphs 12 of 19; this file's or the design's + own metadata 7; the instrument 0. +- **require↔withdraw:** none. + +**Tells: zero of five.** The count fell and the Blocker went to zero. + +**What is now excluded that was not before** — the measure that matters more than the count. A +source-blocked pass can no longer be run through the continue branch. §G's membership test no +longer classifies reviewer choice and lens sets as contract members. A reader can no longer take +evidence revalidation as covering what the closing commit carries. And the two mechanism claims +that were false are gone: the hook clears its Gate-B state on any non-`WIP` commit **attempt**, +success or not — read at `codex-gate.sh:876`, not inferred — and a stray commit reaches that +counter and nothing else. + +**Still open and stated as such:** the ownership boundary (three sections describe one hold +discharge), the standing four-item predicate list, and §H's enumeration of its own additions. ## Pass-33 report — MANDATORY TWO-TELL STOP, and the scope cut did not help diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index 6942a15..1a44b01 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -29,9 +29,11 @@ cycle and which merely *suspend* it, in what order a pass is read so the ranking rather than asserted, what any set of suspensions at once does, and which of the four standing duties participate in that ordering versus gate it as preconditions. With it: the answer to what a severity demotion does to the loop-health counts, and the standing sentences the ordering -falsifies or leaves ambiguous if they are not edited at their source, which §4 lists row by row -**without claiming a total** — a count over spans that merge and split is bookkeeping the plan -re-derives against the files. +falsifies or leaves ambiguous if they are not edited at their source. **§4 lists the sites row by +row and claims no total over them**, a count over spans that merge and split being bookkeeping the +plan re-derives against the files. **The nine falsified standing sentences are a different count** +and the target text's §F states it, because those are individually enumerated sentences rather than +spans. **What does not:** the pass floor and severity semantics, which the parent shipped and this spec reads as given, and everything §9 lists as moved or parked. @@ -211,6 +213,7 @@ to point at and so a reader can see the shape of the change without reading the | (i) when these rules bind | §H | | the Gate-A section and the gate-prompt template | the two senses of *clean* are separated; the cadence makes revision conditional; the broad-prompt instruction stops assuming the artifact is revised between passes, keeping its breadth demand | | the profiles section, the lens paragraph | its unchanged-list is scoped to the lens sets | +| the profiles section, the profile-change paragraph | its claim that a stray non-`WIP` commit "would discard the accumulated passes" is narrowed to the hook's count of them (pass 33 finding 7); §F item 8 | | the profiles section, the evidence-entry revalidation remedy | its fix-re-review-close instruction becomes conditional on the ordering selecting continuation, a non-closing pass taking any applicable suspension first (pass 30 finding 3). The eighth falsified standing sentence, and the sixth sharing §F's mechanism | **Two sentences are deliberately not edited**, named so nobody looks for them: the "Copy every diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index 180aa61..a6ef17f 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -74,10 +74,11 @@ one of *those* defined here has found a defect. **The conditions every cycle has, whatever its gate.** The duties classified below; and the **profile**, the **cited set** and the **assigned fix set**, each gating as its own source says and cited here without restatement. **The profile and the assigned fix set answer a change and not a -differing value**, which is why one of them changed and then undone still costs a pass; **the -cited set answers as its own source says** — the final clean pass runs against the current set — -which a change since undone can already satisfy. The contrast that matters is against a gate's own -content condition, which may be a comparison of current values and says so where it is stated. +differing value**, which is why one of them changed and then undone still costs a pass. **The cited +set answers as its own source says**, and that source says two things: a governing header changed +**during** a pass makes that pass not final, and the final clean pass runs against the **current** +set. The contrast that matters is against a gate's own content condition, which may be a comparison +of current values and says so where it is stated. **What a pass is read from.** Every finding-derived predicate reads the validated findings file **or files** of the logical pass as **the concatenation of their finding lines after each file has @@ -214,7 +215,8 @@ decision made by omission. A finding the clearly-stuck reading surfaces that als trigger takes the scope stop's answers at that same surface, so it is not asked twice; the two-tell stop surfaces tells and not a finding. -**Otherwise the continue branch: a pass that neither closes nor suspends continues** — the loop +**Otherwise the continue branch, which no source block reaches: where none stands, a pass that +neither closes nor suspends continues** — the loop runs another pass on the **current** artifact, revised where the severity and scope rules require a repair and unrevised where they do not. **An eligible pass with an unmet closure condition lands here**, and like every other non-closing pass **only where no suspension applies to it**: clean @@ -355,8 +357,8 @@ the ordering states for every cycle; that list is there and is not repeated here final pass's review request.** Gate A hands the reviewer text rather than a git range, which is why this condition is Gate A's and is written nowhere else. It is **current equality and deliberately nothing more**: it does **not** say the artifact went untouched in between, and text edited and -then restored byte for byte satisfies it — the one place in this cycle's conditions where an undone -change costs nothing, and stated here because the ordering's conditions answer a change instead. +then restored byte for byte satisfies it — stated here because the ordering's conditions answer a change +instead, and said of this condition rather than as a claim about every other. That is a decision rather than an oversight: a content comparison cannot tell those two states apart, and a condition nobody can check is a condition nobody applies. It likewise says nothing about **what the reviewer consumed**: no part of this act is offered as evidence of the review @@ -690,9 +692,11 @@ W 1011–1012. ``` **Finishing the cycle:** once the closure ordering reaches a Gate-B cycle's closing act — an eligible pass with every closure precondition holding, never a clean pass on its own — close it -with `git commit --amend -m ""`; that replaces the WIP commit, and the hook reads -the amend as the real cycle-closing commit. This section states the operation and never whether -the cycle may close. +with `git commit --amend -m ""`; that replaces the WIP commit. **The hook treats any +non-`WIP` commit *attempt* as a Gate-B boundary and clears its state even where the command +fails**, so a failed closing act leaves that counter cleared and no commit made — a fact about the +counter and not about the cycle. This section states the operation and never whether the cycle may +close. ``` *Why (pass 20 finding 5):* the live sentence says "after the final clean pass, close it with `git commit --amend`", which is a complete instruction to whoever enters through Mechanics — and @@ -750,7 +754,7 @@ about a mechanism rather than an entry point carrying an unqualified instruction **9. The evidence-entry revalidation remedy** (the profiles section). It wraps across C 728–729 and W 914–915. ``` -If revalidation changes the entry, the clean pass no longer covers what is being committed: the +If revalidation changes the entry, the pass was read against an entry that no longer stands: the pass did not close, so the closure ordering decides what happens next: where it selects the continue branch, fix and re-review, and where it selects a suspension, the answer comes first. The pass that follows is read by that ordering like any other and closes only if it reaches closure, @@ -774,8 +778,8 @@ semantics **and the closure ordering together with every rule it reads** depend and the requirement is that the adopted definitions **agree**, not merely that all of them are present. **Membership is decided by a test a reader can apply to the text in front of them, with no list to consult, and the test reads what a rule states rather than what changing it would do: -a live rule belongs to this contract when what it says determines or supplies an input the -closure ordering reads, which branch a pass takes, what a hold is or what discharges it, +a live rule belongs to this contract when what it says **defines an input's validity or how it is +read**, which branch a pass takes, what a hold is or what discharges it, whether a cycle may close **or may terminate without running a pass at all, the Gate-B triviality skip being the one such route and its eligibility test therefore a member**, or the production, identity or transport of **any §5 cycle record, required or optional** — §5 entire and not the From 701aa5063f7b584f885f1e84126021d093c5faf2 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sun, 13 Sep 2026 15:22:25 +0200 Subject: [PATCH 064/181] docs(specs): apply pass 35; four Majors, all wordings of earlier repairs MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit One tell — the count rose 19 to 20. Blocker+Major falls a third time: 9, 5, 4. The cluster was an even split between product rules and this file's metadata, disclosed as borderline and read as not the prose tell, since all four Majors are product behaviour and the metadata half is collected NITs. - §C stops enumerating what follows a failed clearly-stuck reading and points at the ordering. The enumeration was the defect in both rounds it appeared: first it omitted the continue branch, then the source block - a non-qualifying recurrence of a dismissed finding is handled as any fresh finding is, by the severity and scope rules at their sources. Saying it owes a repair or a dismissal forced an out-of-set or Minor recurrence into a duty the resolve rule does not reach - a stray non-WIP commit costs the hook's counter AND leaves the WIP snapshot as an ancestor the closing amend does not replace, so the closing act takes the reset-and-single-commit shape. Pricing it as counter loss alone dropped the structural cost; Mechanics already says a WIP commit left in history is what the amend exists to prevent - §G's "input" is named as an input the closure ordering reads, so the lens and broad-prompt rules no longer satisfy the criterion the next sentence excludes All four were wordings of my own earlier repairs, two too wide and two too narrow. None needed a decision and none of the five settled decisions moved. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-35.md | 21 +++++++++++++++ .../gate-a-spec-awsf1ec771-resume.md | 26 +++++++++++++++++++ ...-10-loop-rule-consolidation-target-text.md | 20 ++++++++------ 3 files changed, 59 insertions(+), 8 deletions(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-35.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-35.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-35.md new file mode 100644 index 0000000..7a2080a --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-35.md @@ -0,0 +1,21 @@ +MAJOR | high | §C "missing one means" | The replacement enumerates only "another suspension, a continue, or a close" after the clearly-stuck reading fails, but A1's first branch also requires a source-blocked pass to remain blocked | A reader entering through §C can bypass an unresolvable profile, disagreeing governing headers, or an unreadable Story header and run or close when the source rule forbids another pass | Replace the enumeration with "what follows is the closure ordering's" so the source-block branch is preserved without a second branch list +MAJOR | high | §A1 "Where any of those fails" | A1 says every non-qualifying recurrence of a dismissed finding owes a repair or dismissal, while Mechanics · Severity requires that only in-set Blocker/Major findings resolve and passage B lets an out-of-set fresh finding take a membership stop | A doubtful-identity, out-of-set, Minor, or Nit recurrence can be forced into an unauthorized repair or dismissal instead of following the ordinary severity and scope rules | Replace the repair-or-dismissal claim with a reference saying the recurrence is handled as an ordinary fresh finding under Mechanics · Severity and the absorb paragraph +MAJOR | high | §A3 "non-WIP commit mid-cycle" | A3 prices an accidental non-WIP commit only as lost hook state, but a successful ordinary commit on top of the active WIP leaves that WIP as an ancestor; the cited closing amend replaces only the tip, and the standing reset alternative applies only when several WIP snapshots piled up | The cycle can close with a WIP commit still in history and with its final message or records split across commits, contrary to the standing reason for the closing act | Replace the blanket cost sentence and widen the existing reset-and-single-commit act to cover any case where the active WIP is no longer the sole tip commit to replace, including a stray non-WIP descendant +MAJOR | high | §G "defines an input's validity or how it is read" | The membership test does not say that "input" means an input to the closure ordering; literally, the Gate-A broad-prompt and lens rules define how the artifact is read, while the next sentences expressly classify review-framing prompt wording and lenses as non-members | A downstream reader with only the shipped prompt still gets both membership answers for the same sentence and cannot apply the partial-adoption stop consistently | Replace the limb with a criterion limited to defining whether a value presented to the closure ordering is valid or how that ordering interprets the value +MINOR | high | §A1 "derived floor" duty | A1 says the floor is discharged only when the last counted pass is clean, although the standing text defines the floor as the valid pass count owed and A1 separately makes cleanliness part of eligibility | A dirty pass at the numeric floor simultaneously meets and fails the floor, making floor status and closure reasoning disagree | Define floor discharge only by reaching the required valid-pass count, retain the zero-finding exception, and leave last-pass cleanliness solely in eligibility +MINOR | high | §§A1 and E "only statement" | E calls itself the only statement of the resolve duty, but A1 also normatively defines per-finding discharge, cross-pass tracking, the effect of omission from a later findings file, and repeated-dismissal discharge | The promised single-owner structure has two definitions that can drift under implementation or partial adoption | Move the discharge and tracking rules to Mechanics · Severity and leave A1 with only the duty's classification and a reference +MINOR | high | §§A1 and F item 7 "human-exception record" | A1 says a human-exception record goes in the commit used by the closing act, and F item 7 independently installs the same destination at the human-exception source | One record-placement rule has two authorities despite the target's one-definition contract | Keep the destination at the human-exception source and replace A1's destination clause with a citation to records already owed +MINOR | high | §H "without inspecting it" | The Gate-A clean-signal replacement says NO FINDINGS lets a pass be read as clean without inspection, while the standing acceptance protocol requires inspection of readability, exact body, terminator, count, and extra lines before any pass is accepted | An agent can accept a stale, truncated, or malformed zero-finding file without performing mandatory validation | Say the signal establishes zero findings after mandatory file validation, without requiring severity classification of finding lines +MINOR | high | §G and standing §5 "partial adoption" | G applies its stop-and-reconcile rule to all contract members including the floor, but the untouched "Downstream has no shipping commit" passage already defines the floor subset's coherence test and the same stop-and-reconcile action | The shipped prompt has two authorities for the same partial-adoption decision, recreating the duplication the one-contract paragraph claims to remove | Preserve the standing paragraph's floor-specific facts but replace its terminal instruction with a pointer to the one-contract rule, or consolidate the broader contract at that existing site +MINOR | high | §A3 "Gate-A cycle's count" | A3 says the hook clears Gate-A count only at skill boundaries that start a new Gate-A cycle, but codex-gate.sh also clears it on executing-plans and subagent-driven-development, which close the Gate-A plan cycle rather than start another one | Readers can infer that the advisory count survives a closing skill boundary when the hook has already deleted it | Replace the sentence with the exact behavior: commits never clear Gate-A count, while the hook clears it at the four listed opening and plan-execution skill boundaries +MINOR | high | §I "Which exit a cycle took" | The residual says no cycle exit is observable from history, but the standing commit grammar records a Gate-B triviality termination explicitly as "skipped" with its adjacent skip reason | The residual overstates the observability gap and contradicts the passless route that A1 and G expressly retain | Narrow the sentence to say that history does not reveal whether a running cycle took a scope, clearly-stuck, or two-tell suspension before clean completion +MINOR | high | §I "sites known at pass 17" | I says the target contains the sites known at pass 17, but F itself attributes added sites to passes 30 and 33 and other target rationales cite post-17 discoveries | The target gives the plan a false provenance boundary for its known edit set | Delete the pass-17 boundary or replace it with an accurate through-current-pass statement while retaining the no-completeness claim +NIT | high | §F item 5 line citation | The target says the live Finishing-the-cycle sentence wraps across CLAUDE.md 827–828 and workflow-init.md 1011–1012, but the sentence continues through lines 829 and 1013 respectively | The mechanical source citation truncates the sentence it claims to locate | Change the ranges to C 827–829 and W 1011–1013 +NIT | high | §F item 9 and design §4 | The target enumerates the evidence-entry revalidation remedy as falsified sentence 9, while the design's final edit-table row calls it "The eighth falsified standing sentence" immediately after assigning item 8 to the profile-change claim | The target and its decision document give incompatible mechanical numbering for the same replacement | Change the design row from "eighth" to "ninth" +NIT | high | §C and story D3 | Story decision D3 says the existing precedence sentence is preserved verbatim, while C explicitly splits that sentence, moves only its operative clause, and capitalizes the initial "a" to "A" while still claiming the clause survives unchanged | The target does not meet the settled text-preservation claim even though it retains the behavior | Replace the verbatim-preservation claim with the accurate narrower statement that the operative precedence rule is preserved, or retain the full source sentence byte-for-byte +NIT | high | §A1 "source-block branch" | The text names "The source-block branch" and then says it needs no name | The section makes a mechanically false claim about its own four-branch structure | Remove "needs no name" while retaining that the source rule supplies the procedure +NIT | high | §A1 "cross-reference anywhere" | A1 claims every cross-reference names a branch rather than its position, but the proposed text uses positional references including "branches below", "the branch above", and "either trigger above" | Reordering can break the references the sentence says are stable | Replace the positional references with the clean-completion, suspension, continue, or scope-trigger names +NIT | high | §A1 "What D2 and D3 forbid" | A1 says D2 and D3 both forbid reporting "will not converge", but only the clearly-stuck reading makes that report; the two-tell stop reports observed tells and asks continue or stop | The two-tell precedence rule is justified by a consequence it does not produce, leaving its actual rationale unstated under prompt-standard item 6 | Give the two decisions separate reasons, reserving the false-convergence rationale for clearly stuck and using the absence of a running-cycle decision after closure for two tells +NIT | high | §H "only place this file states anything" | H says B is the only place the file states anything about passage B, although A1 describes and cites B's assigned fix set and scope triggers and H's own sentence discusses the passage | The section's assignment metadata is literally false | Say B is the only place that writes out passage B's replacement bytes +NIT | high | §A "Three paragraphs" | A says the installed addition is three paragraphs, but A1, A2, and A3 each contain multiple blank-line-delimited Markdown paragraphs | The stated count does not match the proposed text and can mislead the insertion plan about the replacement unit | Call A1, A2, and A3 three blocks or passages +END OF FINDINGS (20 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index 90970a6..98465a2 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -61,6 +61,32 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 32 | 45b7d36 | 13→**15** | 0→**1** | 6→**5** | yes | **MANDATORY TWO-TELL STOP.** Findings rose and the Blocker returned. **Three of six B+M are collisions between my own repairs of passes 29–31.** All 15 held open; session 01a09a31-a7bb-7a03-9437-6c9396908fdb | | 33 | 3ee9132 | 15→**22** | 1→**1** | 5→**8** | yes | **MANDATORY TWO-TELL STOP — third in six passes.** B+M 6→**9**, the worst since pass 26. **The scope cut did not reduce the count; it added cleanup debt.** Three findings are references to the cut condition my sweep did not reach. All 22 held open; session 01a09a7d-55f2-7513-b61b-09c9a34626d3 | | 34 | ea5e76c | 22→**19** | 1→**0** | 8→**5** | yes | **zero tells.** B+M 9→**5**, the largest fall of this stretch and the lowest since pass 29 — finishing the cut is what did it. All 5 Majors verified before repair, one of them (§F item 5) against `codex-gate.sh:876` rather than asserted; session 01a09acf-1185-7883-b23f-3ee707d5ad58 | +| 35 | 8847414 | 19→**20** | 0→**0** | 5→**4** | yes | one tell (count rose by one). **B+M 5→4, third consecutive fall: 9, 5, 4.** All four Majors are over-narrow or over-wide wordings of my own earlier repairs; session 01a09ade-8e4c-7352-9e19-75be21c2178e | + +## Pass-35 report — one tell, B+M 9 → 5 → 4 + +**Floor line:** derived floor **3**; risk **high**, security **none**; read fresh from +`docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md`. One cited story, level 2. + +- **Trend:** findings 19, **20**. Blockers 0, **0**. Majors 5, **4**. Blocker+Major 5, **4**. +- **Cluster (pass 35):** an **even split** — 10 on the ordering and the gate blocks, 10 on this + file's or the design's own metadata. Disclosed because it is borderline: all four Majors are + product behaviour and the metadata half is mostly NITs collected across several passes, so this + is read as **not** the prose-cluster tell. The instrument: 0. +- **require↔withdraw:** none. Pass 35's §G finding is a follow-on to the replacement pass 34 asked + for, not a demand for text an earlier pass removed. + +**Tells: one of five** — the count rose by one. Below the mandatory threshold. + +**What is now excluded.** A reader entering at §C can no longer walk past a source block, because +§C stopped enumerating branches at all — the enumeration was the defect in both rounds it appeared. +A non-qualifying recurrence of a dismissed finding is no longer forced into a repair it may not +owe. A stray non-`WIP` commit is priced for what it actually costs: the hook's counter **and** a +`WIP:` snapshot left in history, which the closing amend would not replace. §G's "input" is the +ordering's input. + +**All four Majors were wordings of my own earlier repairs** — two too wide, two too narrow. None +required a new decision, and none of the five settled decisions was touched. ## Pass-34 report — zero tells, B+M 9 → 5 diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index a6ef17f..778a75d 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -269,9 +269,10 @@ a hold like any other. **A re-raised valid dismissal stays discharged for the re exactly the terms the clean predicate sets out above** — the same complaint, no new evidence, no change to the text the dismissal turned on, and the dismissal's reason still true of the artifact. The dismissal was the resolution and a reviewer repeating it does not undo it, so no second -dismissal is owed. **Where any of those fails the recurrence is an ordinary fresh finding**, judged -at its current effective severity and owing a repair or a dismissal of its own; reading the old -dismissal as covering it would let a finding that has since become true close a cycle. What a +dismissal is owed. **Where any of those fails the recurrence is an ordinary fresh finding** and is +handled as one — by the severity and scope rules at their own sources, which decide whether it is +in set and what it owes; reading the old dismissal as covering it would let a finding that has +since become true close a cycle. What a qualifying recurrence creates is the **clearly-stuck hold**, ended by that reading's continue-or-stop answer. Any trigger the recurrence independently carries raises its own stop as usual. **Where one finding is surfaced by both, it carries two hold components and each is @@ -414,8 +415,11 @@ performed once the ordering reaches it: an eligible pass with every closure cond **A commit the hook reads as cycle-closing is a Gate-B matter.** A non-`WIP` commit mid-cycle makes the hook drop its Gate-B review state and that gate's counter — an observation about the counter, since the cycle itself stays open until the conditions hold, so an accidental commit resets what -the hook reports and closes nothing; **what a pass is stands on its validated findings file, not -on that counter**. **It does not reach a Gate-A cycle's count**, which the hook +the hook reports and closes nothing; **what a pass is stands on its validated findings file, not on +that counter**. **It also leaves the `WIP:` snapshot as an ancestor**, which the closing amend +replaces nothing of — so the closing act takes the reset-and-single-commit shape Mechanics gives +for that case, a `WIP:` commit left in history being exactly what Mechanics says the amend exists +to prevent. **It does not reach a Gate-A cycle's count**, which the hook clears at the skill boundaries that start a new Gate-A cycle rather than on any commit. ``` @@ -533,7 +537,7 @@ clause left standing alone — its second half, the precedence sentence, moves i the second is new. ``` -So this exit needs three things **together**, and a missing one means only that *this* exit does not apply — what the pass does instead is the closure ordering's — another suspension, a continue, or a close: a plateau +So this exit needs three things **together**, and a missing one means only that *this* exit does not apply — what the pass does instead is the closure ordering's, read there in full: a plateau visible across passes (six or more is where the field saw one); an **affirmative judgement that coverage is sufficient**, stated — a known materially unreviewed area forbids this exit outright, and disclosing it does not license it; and **Blocker or Major findings that keep regenerating @@ -778,8 +782,8 @@ semantics **and the closure ordering together with every rule it reads** depend and the requirement is that the adopted definitions **agree**, not merely that all of them are present. **Membership is decided by a test a reader can apply to the text in front of them, with no list to consult, and the test reads what a rule states rather than what changing it would do: -a live rule belongs to this contract when what it says **defines an input's validity or how it is -read**, which branch a pass takes, what a hold is or what discharges it, +a live rule belongs to this contract when what it says **defines the validity of an input the +closure ordering reads, or how that input is read**, which branch a pass takes, what a hold is or what discharges it, whether a cycle may close **or may terminate without running a pass at all, the Gate-B triviality skip being the one such route and its eligibility test therefore a member**, or the production, identity or transport of **any §5 cycle record, required or optional** — §5 entire and not the From a37b5ab2eac5e52fab9d58d4fcf631f17e57569f Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sun, 13 Sep 2026 15:38:07 +0200 Subject: [PATCH 065/181] docs(specs): apply pass 36; three Majors, two of the same assurance class MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Zero tells. Findings 20 to 9, Blocker+Major 9, 5, 4, 3 across four passes. All three Majors were wordings from the pass-35 round: - "what a pass is stands on its validated findings file" is false of a full Gate-B pass, which also needs both branch files and equal kept baseSha and headSha. Replaced with what is true: the hook reset erases counter state and does not invalidate a pass that satisfied the validation rules where they are stated - §A3 named the reset-and-single-commit shape for a state standing Mechanics does not prescribe it for — it gives that shape only where several WIP snapshots piled up. §A3 now states the duty Mechanics already carries, no WIP commit left in history, and leaves the git sequence to the plan as every other closing sequence is - §F item 1 keeps the counter fact and points at the Gate-B closure paragraph for what the commit does to the repository, instead of claiming the counter was the whole cost Two of the three are the assurance class I have now re-introduced three times: a claim about what a mechanism guarantees, written where a duty or a pointer belonged. Swept the file afterwards; none remains. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-36.md | 10 +++++++ .../gate-a-spec-awsf1ec771-resume.md | 27 +++++++++++++++++++ ...-10-loop-rule-consolidation-target-text.md | 18 +++++++------ 3 files changed, 47 insertions(+), 8 deletions(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-36.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-36.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-36.md new file mode 100644 index 0000000..a014e42 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-36.md @@ -0,0 +1,10 @@ +MAJOR | high | §A3 "reset-and-single-commit shape Mechanics gives for that case" | The standing Mechanics text gives reset-and-single-commit only when "several WIP snapshots piled up"; it does not cover the newly named case of one active WIP followed by a stray non-WIP commit, so A3 cites an operation the standing entry point does not prescribe for that state | An agent entering through Mechanics can follow its default amend, replace only the stray tip and leave the WIP ancestor in history despite A3 requiring the reset shape | Replace the standing Mechanics reset sentence so its condition includes every state in which the active WIP is not the sole tip commit to replace, including a non-WIP descendant +MAJOR | high | §F item 1 "so the cost is a reminder" | This replacement says the cost of a successful non-WIP commit is only an understated hook reminder, while A3 now says the same commit also leaves the WIP snapshot as an ancestor and changes the required closing operation to reset-and-single-commit | The two entry points give different consequences for the requested stray-commit path, and the Mechanics reader can miss the repository repair and leave a WIP commit in history | Replace the exhaustive cost sentence with the narrow counter fact and point repository cleanup to Gate B's closing paragraph +MAJOR | high | §§A3 and F item 1 "what a pass is stands on its validated findings file" | A validated findings file does not by itself establish a valid full Gate-B logical pass: standing Mechanics also requires both branch files and exact equality of the kept `baseSha` and `headSha` values, treating a mismatch as incomplete | A reader recovering from the hook reset can count files from branches aimed at different ranges as a valid pass and close on review work §5 explicitly excludes | Say the reset erases advisory counter state but does not invalidate a pass that already satisfied every standing file-validation and branch-agreement condition; do not name the findings file as sufficient +MINOR | high | §A3 "skill boundaries that start a new Gate-A cycle" | The hook does not clear Gate-A count only at boundaries that start a cycle: `codex-gate.sh` lines 891–897 also clear it for `superpowers:executing-plans` and `superpowers:subagent-driven-development`, which close the Gate-A plan cycle | The proposed mechanism description is false and can make a reader expect Gate-A count to survive a plan-execution boundary after the hook deleted it | Replace the clause with the exact distinction: commits do not clear Gate-A count, while Gate-A opening and plan-execution skill boundaries do +MINOR | high | §I "Which exit a cycle took is not observable from history" | The claim is categorical, but standing Mechanics writes a skipped Gate-B cycle as `skipped (see skip reason)` with the adjacent reason in the commit body, so the passless triviality termination that A1 and G retain is observable from history | The residual overstates the observability gap and contradicts an existing cycle record, making downstream work appear to own a case already recorded | Restrict the residual to suspension history inside a running cycle and leave the existing skip record outside the claimed gap +MINOR | high | §C "D3 is satisfied" | Story decision D3 says the existing precedence sentence is preserved verbatim, while §C explicitly splits that sentence, moves only its operative clause and capitalizes its first word; semantic preservation does not satisfy the settled byte-preservation requirement | The target disagrees with a settled D1–D8 input while claiming agreement, so adoption would silently override a decision the review was told not to reopen | Preserve the full standing D3 sentence byte-for-byte at one authoritative site and reference it from the ordering without altering its bytes +NIT | high | §I "It is the sites known at pass 17" | This provenance claim is mechanically false because §F itself attributes added replacement sites to passes 30 and 33 | The plan receives an inaccurate boundary for the target's known edit set | Delete the pass-17 provenance boundary and retain only the supported no-completeness statement +NIT | high | §F item 5 line citation | The target says the live Finishing-the-cycle sentence wraps across C 827–828 and W 1011–1012, but the same sentence continues through C 829 and W 1013 | The mechanical source citation truncates the sentence it claims to locate | Correct the ranges to C 827–829 and W 1011–1013 +NIT | high | §F item 9 and design §4 | The target numbers the evidence-entry revalidation remedy as falsified sentence 9, while the design's corresponding final edit-table row calls it "The eighth falsified standing sentence" after assigning item 8 to the profile-change claim | The target text and its decision document give incompatible mechanical numbering for the same replacement | Change the design row from "eighth" to "ninth" +END OF FINDINGS (9 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index 98465a2..ca2af79 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -62,6 +62,33 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 33 | 3ee9132 | 15→**22** | 1→**1** | 5→**8** | yes | **MANDATORY TWO-TELL STOP — third in six passes.** B+M 6→**9**, the worst since pass 26. **The scope cut did not reduce the count; it added cleanup debt.** Three findings are references to the cut condition my sweep did not reach. All 22 held open; session 01a09a7d-55f2-7513-b61b-09c9a34626d3 | | 34 | ea5e76c | 22→**19** | 1→**0** | 8→**5** | yes | **zero tells.** B+M 9→**5**, the largest fall of this stretch and the lowest since pass 29 — finishing the cut is what did it. All 5 Majors verified before repair, one of them (§F item 5) against `codex-gate.sh:876` rather than asserted; session 01a09acf-1185-7883-b23f-3ee707d5ad58 | | 35 | 8847414 | 19→**20** | 0→**0** | 5→**4** | yes | one tell (count rose by one). **B+M 5→4, third consecutive fall: 9, 5, 4.** All four Majors are over-narrow or over-wide wordings of my own earlier repairs; session 01a09ade-8e4c-7352-9e19-75be21c2178e | +| 36 | 701aa50 | 20→**9** | 0→**0** | 4→**3** | yes | **zero tells. Findings more than halved; B+M 9, 5, 4, 3 across four passes.** All three Majors are pass-35 wordings of mine, two of them the same assurance class again; session 01a09aef-327b-7a11-8033-9e3584c47b27 | + +## Pass-36 report — zero tells, B+M 9 → 5 → 4 → 3 + +**Floor line:** derived floor **3**; risk **high**, security **none**; read fresh from +`docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md`. One cited story, level 2. + +- **Trend:** findings 20, **9**. Blockers 0, **0**. Majors 4, **3**. Blocker+Major 4, **3**. +- **Cluster (pass 36):** the ordering and the gate blocks 4 of 9; this file's own metadata 5; the + instrument 0. +- **require↔withdraw:** none. All three Majors are follow-ons to the pass-35 repairs, not demands + for text an earlier pass removed. + +**Tells: zero of five.** Findings more than halved and no count rose. + +**All three Majors were mine, and two are the same class I keep re-introducing** — an assurance +about what a mechanism guarantees. "What a pass is stands on its validated findings file" is false +of a `full` Gate-B pass, which also needs both branch files and equal kept `baseSha`/`headSha`. And +§A3 had named the reset-and-single-commit shape for a state standing Mechanics does not prescribe +it for. **Both are now stated as duties and pointers rather than as guarantees or command shapes**: +the reset erases counter state and does not invalidate a pass that met the validation rules stated +where they live, and the closing act owes what Mechanics already requires — no `WIP:` commit left +in history — with the sequence left to the plan like every other. + +**Four passes of falling counts, and the mechanism behind it is legible**: every Major since pass +34 has been a wording of an earlier repair, narrowing or widening, and none has touched a settled +decision or opened a behaviour question. ## Pass-35 report — one tell, B+M 9 → 5 → 4 diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index 778a75d..c3f7c24 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -415,11 +415,12 @@ performed once the ordering reaches it: an eligible pass with every closure cond **A commit the hook reads as cycle-closing is a Gate-B matter.** A non-`WIP` commit mid-cycle makes the hook drop its Gate-B review state and that gate's counter — an observation about the counter, since the cycle itself stays open until the conditions hold, so an accidental commit resets what -the hook reports and closes nothing; **what a pass is stands on its validated findings file, not on -that counter**. **It also leaves the `WIP:` snapshot as an ancestor**, which the closing amend -replaces nothing of — so the closing act takes the reset-and-single-commit shape Mechanics gives -for that case, a `WIP:` commit left in history being exactly what Mechanics says the amend exists -to prevent. **It does not reach a Gate-A cycle's count**, which the hook +the hook reports and closes nothing; **what the reset erases is counter state, and it does not +invalidate a pass that already satisfied the validation rules this section states** — which is +where what makes a pass valid stays. **It also leaves the `WIP:` snapshot as an ancestor**, which +an amend of the stray tip would not replace, so **the closing act still owes what Mechanics already +requires of it: no `WIP:` commit left in history.** Which git sequence reaches that from this state +belongs to the plan, as every other closing sequence does. **It does not reach a Gate-A cycle's count**, which the hook clears at the skill boundaries that start a new Gate-A cycle rather than on any commit. ``` @@ -644,9 +645,10 @@ exception to point at. ``` A pre-review snapshot named anything else reads to the hook as a real commit: the hook treats the cycle as closed and **discards its count of the passes you just accumulated**, while the cycle -itself stays open until the closure ordering's conditions hold. **What the hook loses is its counter state**, and a -valid pass is established by its validated findings file rather than by that counter — so the cost -is a reminder that now understates what you hold, not a close nobody intended. +itself stays open until the closure ordering's conditions hold. **What the hook loses is its counter state**, and that +counter is not what makes a pass valid — so the reminder now understates what you hold, and no +close was intended or made. **What such a commit does to the repository, and what the closing act +then owes, is in the Gate-B closure paragraph**, not here. ``` *The sibling sentence in the profile-change paragraph claims no closure* — it says such a commit "reads **to the hook** as the cycle closing" — **but its second half is falsified by this same From cc7f6dd42a0196eef6a470610994bbae1feb704a Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sun, 13 Sep 2026 15:51:27 +0200 Subject: [PATCH 066/181] docs(specs): apply pass 37; the unrepairable closing failure gets a terminal state MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit One tell — the Blocker count rose 0 to 1. Findings 9 to 3, the lowest of the cycle. Below the mandatory threshold. The Blocker is real and mine: decision 5 says a failed closing act is not a branch, so a signing or permission failure nobody can repair left the cycle unable to close and unable to suspend — exactly the shape acceptance criterion 4 exists to catch. Such a failure is now surfaced and the cycle left PARKED, which is the state a stop answer already produces. Named, not invented; no new branch and no new mechanism. - the suspension branch is reached by a non-closing pass where a suspension applies to it. "Whatever its cleanliness" was meant to drop the eligibility bar and read as claiming every non-closing pass, which collided with both the continue branch and the source block - only a stray commit that succeeds and does not amend leaves the WIP snapshot as an ancestor: a failed attempt adds no commit and an amend replaces the tip. The duty is scoped to that one shape Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-37.md | 4 ++++ ...9-10-loop-rule-consolidation-target-text.md | 18 +++++++++++++----- 2 files changed, 17 insertions(+), 5 deletions(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-37.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-37.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-37.md new file mode 100644 index 0000000..a400b1f --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-37.md @@ -0,0 +1,4 @@ +BLOCKER | high | §A1 "A failed act returns to the closure step" | After surfacing a failed closing command and re-establishing the conditions, the text unconditionally says to perform the act again when they still hold, but gives no terminal action when the same signing, permission, hook, or repository failure cannot be repaired; because the act failure is expressly not a branch, neither suspension nor continue can park that cycle | A persistent closing failure leaves the cycle unable to close and unable to suspend, so the agent can retry the same command indefinitely without any answer or state change | Keep failure handling inside the closure step but replace the unconditional retry with a duty to retry only after the concrete cause is resolved and to stop and surface when no repair is available +MAJOR | high | §A1 "Then the suspension branch" | The suspension branch says every pass that did not close reaches it regardless of cleanliness, while the continue branch says an eligible pass with an unmet closure condition and no suspension lands in continue, and the source-block branch separately keeps a source-blocked pass with no suspension blocked; the same no-suspension pass is therefore assigned both to the suspension branch and to its actual non-suspension outcome | A reader can invent an undefined suspension, bypass a source block, or fail to continue an eligible pass whose closure condition remains unmet | Replace both unconditional "reaches this branch" claims with a conditional statement that a non-closing pass takes the suspension branch only when at least one named suspension applies, leaving source-blocked passes at their source and all other non-closing passes to continue +MAJOR | high | §A3 "It also leaves the `WIP:` snapshot as an ancestor" | The ancestor claim is attached to a non-`WIP` commit boundary generally, but §F item 5 and codex-gate.sh lines 876–888 establish that the hook clears state even when the command fails, in which case no new commit exists, and a successful `git commit --amend` replaces the WIP tip rather than leaving it as an ancestor; only a successful non-amending commit added on top has the stated repository shape | After a failed or amending attempt, an agent can infer a stray descendant that does not exist and choose unnecessary reset-and-recommit cleanup instead of the failed-act retry path, risking misplaced staged content or commit records | Condition the ancestor sentence on a successful non-amending commit added on top of the active WIP, and leave failed attempts to §A1 and hook-counter behavior to the exact §F item 5 statement +END OF FINDINGS (3 total) diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index c3f7c24..8448f8a 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -167,6 +167,11 @@ concrete command failure and **re-establish every closure condition against the now stands.** Where they all still hold, perform the act again. Where the attempt or its repair moved anything a condition is read from, **that condition has changed and its own rule decides what it costs**, a further pass included, and the cycle is back in the ordering with that pass owed. +**Where the failure cannot be repaired at all** — a signing key nobody has, a permission nobody can +grant — **surface it and leave the cycle parked**: open, not running, spending no passes, restarted +by an explicit later continue, which is the state a stop answer already produces and is named here +rather than invented. A cycle that can neither close nor be parked is the outcome this sentence +exists to prevent. **Closure introduces no new kind of record, and it excuses none**: every other record this cycle owes, a human-exception record among them, is owed and written exactly as before, and **a @@ -200,8 +205,10 @@ no passes and is outside this ordering. outranks a suspension by closing, not by being eligible**: a pass that took the branch above, met every closure condition and had the closing act performed has ended the cycle, and a suspension has nothing left to suspend. **A pass that did not close reaches this branch whatever its -cleanliness** — a clean pass below the floor, and equally an eligible pass the cycle's unmet -conditions kept from closing. That is what makes "clean completion outranks the two-tell stop" +cleanliness, where a suspension applies to it** — a clean pass below the floor, and equally an +eligible pass the cycle's unmet conditions kept from closing. Cleanliness is what this branch stops +asking about; whether a suspension applies is still what puts a pass here, and where none does the +continue branch has it. That is what makes "clean completion outranks the two-tell stop" executable rather than asserted, and cleanliness alone never decides it. **What D2 and D3 forbid is reporting "will not converge" on a loop that converged, and a loop still owing a repair, an answer or a closure condition has not converged** — so a mandatory two-tell stop and the clearly-stuck @@ -417,9 +424,10 @@ the hook drop its Gate-B review state and that gate's counter — an observation since the cycle itself stays open until the conditions hold, so an accidental commit resets what the hook reports and closes nothing; **what the reset erases is counter state, and it does not invalidate a pass that already satisfied the validation rules this section states** — which is -where what makes a pass valid stays. **It also leaves the `WIP:` snapshot as an ancestor**, which -an amend of the stray tip would not replace, so **the closing act still owes what Mechanics already -requires of it: no `WIP:` commit left in history.** Which git sequence reaches that from this state +where what makes a pass valid stays. **A stray commit that succeeds and does not amend also leaves the +`WIP:` snapshot as an ancestor** — a failed attempt adds no commit, and an amend replaces the tip — +so in that one shape **the closing act still owes what Mechanics already requires of it: no `WIP:` +commit left in history.** Which git sequence reaches that from this state belongs to the plan, as every other closing sequence does. **It does not reach a Gate-A cycle's count**, which the hook clears at the skill boundaries that start a new Gate-A cycle rather than on any commit. ``` From 13005ded51fdb1964e415ceea82b4582d4ac178d Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sun, 13 Sep 2026 16:13:03 +0200 Subject: [PATCH 067/181] docs(specs): apply pass 38; dimensions follow the profile, components follow routes MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit One tell — the count rose 3 to 10. Blocker+Major flat at 3, 3, 3. Below the mandatory threshold. - §F item 6 froze the review dimensions across passes, while the profiles section requires lens sets to follow the current profile and cited set and recomputes every derived obligation after a change. The broad question stays fixed; the dimensions are recomputed and appended each pass - the both-routes sentence hard-coded membership plus clearly-stuck. A finding can carry a question-stop trigger alongside, or both scope triggers. It now states one hold component per route and one answer per component, which removes the enumeration rather than extending it - §G's membership test gained the limb it was missing: what a suspension asks and what state its answer produces. Continue-consumes-reading and stop-parks-cycle decide a post-pass state and matched no existing limb, so a downstream partial adoption could have dropped the only transition out of a two-tell stop while §G called that sentence a non-member §G has now taken a Major at passes 32, 34, 35 and 38, each naming a different limb and each fix correct for what it named. Recorded as a regeneration chain on the hardest paragraph, not treated as a stop signal at one tell. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-38.md | 11 ++++++++++ ...-10-loop-rule-consolidation-target-text.md | 22 +++++++++++-------- 2 files changed, 24 insertions(+), 9 deletions(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-38.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-38.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-38.md new file mode 100644 index 0000000..19e8d60 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-38.md @@ -0,0 +1,11 @@ +MAJOR | high | §F item 6 "dimensions stay the same every pass" | The replacement freezes the review dimensions across passes, while the standing Profiles text requires lens sets to follow the current profile and cited set and explicitly recomputes every derived obligation after a profile change | A Gate-A pass after a profile or cited-set change can omit newly owed risk or security lenses while still satisfying this instruction | Keep the broad base question stable, but require each pass to recompute and append the dimensions owed by the current profiles and cited set +MAJOR | high | §A1 "Where one finding is surfaced by both" | The sentence is stated for any finding surfaced by a scope stop and the clearly-stuck reading, but it hard-codes exactly two components, membership and clearly-stuck; a finding can instead carry a question-stop trigger, or both scope triggers, alongside clearly-stuck | An in-set finding opening a structural question can be given an inapplicable accept-or-decline decision and its actual question can remain unanswered, while a three-trigger finding loses one required answer | Narrow the sentence to a membership-stop-plus-clearly-stuck finding, or state that every active membership, question and clearly-stuck component is discharged by its own prescribed answer +MAJOR | high | §G "Membership is decided by a test" | The semantic membership test has no limb for rules that define what a suspension asks or what its answer does; the continue-consumes-reading and stop-parks-cycle sentences decide a post-pass state, not input validity, a pass branch, a hold for the two-tell case, closure, passless termination or a record | A downstream partial adoption can omit or mismatch the only transition out of a two-tell stop while §G classifies that sentence as outside the contract, leaving a suspended cycle without the resume behavior the contract is meant to keep coherent | Add the question and answer-produced state transition to the semantic membership test, without adding a checker or a second member list +MINOR | high | §A1 "cited here without restatement" | The paragraph claims the profile, cited set and assigned fix set are only cited, then restates that profile and fix-set changes rather than differing values cost a pass and repeats both cited-set finality rules | The target creates two authorities for three closure inputs despite its own ownership rule, so a later source edit or partial adoption can make the ordering disagree with the source it says it merely cites | Remove the repeated change semantics and leave direct pointers to the three source rules +MINOR | high | §A1 "derived floor" | The duty paragraph says the numeric floor is discharged only when the last counted pass is clean, although the standing floor is the valid-pass count owed and §A1 separately makes final-pass cleanliness part of eligibility | A dirty pass at the numeric floor both meets and fails the floor, making pass reports and closure diagnosis disagree even though the ultimate close remains blocked | Define floor discharge by the required valid-pass count alone, keep cleanliness in eligibility, and describe zero findings as bypassing the floor rather than discharging it +MINOR | high | §A3 "skill boundaries that start a new Gate-A cycle" | The hook clears Gate-A count at four skill boundaries, including `superpowers:executing-plans` and `superpowers:subagent-driven-development`, which close the Gate-A plan cycle rather than start a new one | The proposed mechanism description is false and can make a reader expect the count to survive a plan-execution boundary after the hook has deleted it | State that commits do not clear Gate-A count and that the hook clears it at the four named opening and plan-execution skill boundaries +MINOR | high | §F item 5 "any non-WIP commit attempt" | `codex-gate.sh` does not inspect the resulting commit message; `is_wip_commit` only searches the Bash command for an `-m` argument beginning with `wip`, so a `git commit --amend --no-edit` that preserves a `WIP:` message is still reset and an unrelated matching token can suppress a reset | The target gives readers a message-based model of a command-text matcher, so ordinary WIP housekeeping can unexpectedly erase counters and a matching command can preserve them across a real boundary | Replace the categorical non-WIP claim with a pointer to the hook's cycle-internal command matcher, or state its exact `-m WIP` limitation +MINOR | high | §G and standing "Downstream has no shipping commit" | §G puts the floor rules inside its broader contract and repeats the stop-and-have-a-human-complete-revert-or-reconcile action, while the untouched standing downstream paragraph already defines the floor subset and the same terminal action | The final prompt has two authorities for the same partial-adoption decision, contrary to §G's one-contract purpose and the target's no-duplicate-section requirement | Keep the standing floor-specific failure states but replace its terminal action with a pointer to §G's one-contract action +MINOR | high | §A1 "failure cannot be repaired" | The target adds parking plus an explicit later continue for an unrepairable closing-act failure, while design §2 still says the failed-act rule was "shrunk rather than extended," inventories only re-establish-retry-or-source-cost behavior, and says it introduces no mechanism | The final text and the decision document behind it disagree about the newly settled terminal path, so the plan and later reviewers can treat required behavior as an unsupported addition | Update design §2 to record that the existing parked state and explicit-continue transition are reused for an unrepairable closing-act failure +NIT | high | §A3 "Gate B's content condition, and its closing act" | The section announces that it contains Gate B's content condition and the next paragraph begins "Its content condition," but both then say Gate B has no content condition | The section's own description briefly asserts the existence of the thing its settled behavior correctly denies, adding avoidable ambiguity at the gate-specific split | Rename the lead to "Gate B's lack of a content condition, and its closing act" and state the absence directly +END OF FINDINGS (10 total) diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index 8448f8a..772b597 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -282,12 +282,13 @@ in set and what it owes; reading the old dismissal as covering it would let a fi since become true close a cycle. What a qualifying recurrence creates is the **clearly-stuck hold**, ended by that reading's continue-or-stop answer. Any trigger the recurrence independently carries raises its own -stop as usual. **Where one finding is surfaced by both, it carries two hold components and each is -discharged by its own answer**: the **membership** component ends on the membership answer **in -either direction**, a decline releasing it exactly as an accept does; the **clearly-stuck** -component ends on the reading's continue-or-stop answer. Neither answer discharges the other's -component, and where the finding carries no scope-stop trigger the continue-or-stop answer is the -only one its surface asks for and discharges the only component there is. **Resumption is still the +stop as usual. **Where one finding is surfaced by more than one route it carries a hold +component per route, and each is discharged by its own answer**: the **membership** component by +the membership answer **in either direction**, a decline releasing it exactly as an accept does; +the **question** component by the user's decision on that question; the **clearly-stuck** component +by that reading's continue-or-stop answer. **No answer discharges another route's component**, and +a finding surfaced by one route has one component, discharged by the one answer its surface asks +for. **Resumption is still the composition rule's**, which waits for every outstanding answer. At a **membership stop** the answer is **accept**, the finding joining the fix set where Mechanics · Severity governs it, or **decline**, the finding staying outside and binding so for the rest of this cycle. A later answer @@ -722,8 +723,10 @@ which is the split every sentence of this mechanism takes. **6. The Gate-A broad-prompt instruction** (Gate A section). It wraps across C 553–555 and W 745–747. ``` -Use ONE broad prompt: **its review question and dimensions stay the same every pass, while the -artifact text it carries is always the current one**. Re-running it over an **unrevised** artifact +Use ONE broad prompt: **its review question stays the same every pass, while the artifact text it +carries and the dimensions it asks for are always the current ones** — the lens sets the profiles +section derives are recomputed from the current profile and cited set each pass and appended, since +a profile or cited-set change changes what is owed. Re-running it over an **unrevised** artifact is legitimate wherever no repair is owed, an edit made to justify a pass being no reason to run one. Don't narrow per-dimension: new findings surface because the artifact changed, because an answer given since the last pass changed what the rules require of it, or because a broad prompt @@ -793,7 +796,8 @@ and the requirement is that the adopted definitions **agree**, not merely that a present. **Membership is decided by a test a reader can apply to the text in front of them, with no list to consult, and the test reads what a rule states rather than what changing it would do: a live rule belongs to this contract when what it says **defines the validity of an input the -closure ordering reads, or how that input is read**, which branch a pass takes, what a hold is or what discharges it, +closure ordering reads, or how that input is read**, which branch a pass takes, what a hold is or what discharges it, **what a suspension asks or what +state its answer produces**, whether a cycle may close **or may terminate without running a pass at all, the Gate-B triviality skip being the one such route and its eligibility test therefore a member**, or the production, identity or transport of **any §5 cycle record, required or optional** — §5 entire and not the From 12cf2473f107efa7ae1d47f6870c2cd66379f221 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sun, 13 Sep 2026 16:27:50 +0200 Subject: [PATCH 068/181] docs(specs): apply pass 39; name the three-site fan-out every new rule owes MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Zero tells. Findings 10 to 5; Blocker+Major 3, 3, 4. Three of the four Majors are the same shape: a rule added in the pass-37 or pass-38 round without the three further sites it owes. Every rule this change adds needs a limb in §G's membership test, a strict reading in §H's unknown-start list, and any §F pointer that states it at its own entry point. That is checkable before a pass rather than after one, and is now part of the pre-review sweep. - §G's limb covers what state an incomplete closing act produces, not only what a suspension answer produces. Parking an unrepairable failure matched no limb - §H's strict list gains the parked state, since the standing fallback requires each rule this change ships to add its own reading, and a cycle whose starting rules cannot be established is the last one to leave without a terminal transition - §F item 4's pointer is plural: the ordering prescribes one answer per surfacing route, so a single-answer pointer let a reader resume with another route's component unanswered - the source-block branch's list is marked as examples rather than the set. It named three profile and header failures while the standing evidence-gap rule prescribes stop-and-surface too. Enumerations have now cost four rounds across three sites; each is removed or marked illustrative Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-39.md | 6 ++++ .../gate-a-spec-awsf1ec771-resume.md | 30 +++++++++++++++++++ ...-10-loop-rule-consolidation-target-text.md | 17 ++++++----- 3 files changed, 46 insertions(+), 7 deletions(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-39.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-39.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-39.md new file mode 100644 index 0000000..b8c0690 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-39.md @@ -0,0 +1,6 @@ +MAJOR | high | §A1 "source-block branch" | The em-dash list after "Where any unmet closure condition's own source prescribes stop-and-surface" names only the three profile/header failures, while the unchanged standing sentence "An unobservable counterfactual is a blocking evidence gap ... stop and surface" occurs at CLAUDE.md 717–719 and workflow-init.md 903–905, wrapping in both copies | A Gate-B cycle with missing required evidence and no suspension can be read as taking the continue branch and spending more review passes even though its live evidence source forbids another pass until the gap is answered | Remove the enumeration and point the branch to every source-owned stop-and-surface condition +MAJOR | high | §F item 4 "the answer a suspension asks for" | The replacement says the ordering prescribes one answer "for that suspension" on "the question that suspension raised", but §A1 assigns answers per surfacing route and one scope-stop suspension can carry both membership and question routes | A reader entering through the human-exception paragraph can treat one answer as sufficient and resume while the other route's hold component remains unanswered | Use the plural pointer "the required answers are those the closure ordering prescribes on the questions raised" without enumerating them here +MAJOR | high | §G "Membership is decided by a test" | The new limb covers what state a suspension answer produces, but no limb covers what an incomplete closing act produces; §A1 expressly says that failure is not a branch, and parking an unrepairable failure is neither a suspension answer, closure permission, input rule, passless termination nor record rule | A downstream partial adoption can omit or mismatch the only terminal transition for an unrepairable closing failure while §G classifies the remaining contract as coherent, recreating a cycle unable to close or park | Add the missing narrow limb: what next state an incomplete closing act produces +MAJOR | high | §H "unknown-start strict-reading list" | The unchanged standing sentence "Each further rule this change ships adds its own strict reading to this list" occurs at CLAUDE.md 162–163 and workflow-init.md 369–370, wrapping in both copies, but §H gives no strict reading for §A1's new retry-or-park handling of a failed closing act | When a cycle's starting rules cannot be established, the text does not decide whether the new parked terminal state binds, so an unrepairable closing failure can fall back into the no-progress state this change otherwise closes | Add a narrow strict-reading clause making the failed-closing-act transition binding in the unknown-start case +MINOR | high | §F item 6 rationale "Unchanged now scopes to the question and the dimensions" | The replacement now requires dimensions to be recomputed from the current profile and cited set each pass, so the rationale's claim that the dimensions stay unchanged directly contradicts the proposed sentence above it | The plan or a later reviewer can preserve stale dimensions after a profile change and omit newly owed lenses despite installing the correct replacement bytes | Say that only the broad review question stays unchanged, while both the artifact payload and derived dimensions are current for each pass +END OF FINDINGS (5 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index ca2af79..d8b1b1b 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -63,6 +63,36 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 34 | ea5e76c | 22→**19** | 1→**0** | 8→**5** | yes | **zero tells.** B+M 9→**5**, the largest fall of this stretch and the lowest since pass 29 — finishing the cut is what did it. All 5 Majors verified before repair, one of them (§F item 5) against `codex-gate.sh:876` rather than asserted; session 01a09acf-1185-7883-b23f-3ee707d5ad58 | | 35 | 8847414 | 19→**20** | 0→**0** | 5→**4** | yes | one tell (count rose by one). **B+M 5→4, third consecutive fall: 9, 5, 4.** All four Majors are over-narrow or over-wide wordings of my own earlier repairs; session 01a09ade-8e4c-7352-9e19-75be21c2178e | | 36 | 701aa50 | 20→**9** | 0→**0** | 4→**3** | yes | **zero tells. Findings more than halved; B+M 9, 5, 4, 3 across four passes.** All three Majors are pass-35 wordings of mine, two of them the same assurance class again; session 01a09aef-327b-7a11-8033-9e3584c47b27 | +| 37 | a37b5ab | 9→**3** | 0→**1** | 3→**2** | yes | one tell (Blocker 0→1). **Findings 3 — lowest of the cycle.** The Blocker was real and mine: an unrepairable closing failure had no terminal state; now PARKED, the state a stop answer already produces; session 01a09afd-8b4c-7712-93c1-6a3afa109d87 | +| 38 | cc7f6dd | 3→**10** | 1→**0** | 2→**3** | yes | one tell (count rose). B+M flat at 3. §G takes its fourth Major; dimensions unfrozen; hold components per route rather than per pair; session 01a09b09-ba8b-7773-b42f-93a2f4261814 | +| 39 | 13005de | 10→**5** | 0→**0** | 3→**4** | yes | **zero tells.** Three of four Majors are the three-site fan-out of the pass-37/38 rules — §G limb, §H strict reading, §F pointer. Pattern named and now checked before each pass; session 01a09b1d-8cbc-7d43-8d2f-49ff2339c31f | + +## Passes 37–39 — the fan-out is the mechanism, and it is checkable + +**Floor line (unchanged across all three):** derived floor **3**; risk **high**, security **none**; +read fresh from `docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md`. One cited +story, level 2. + +- **Trend:** findings 9, 3, 10, **5**. Blockers 0, 1, 0, **0**. Majors 3, 2, 3, **4**. + Blocker+Major 3, 3, 3, **4**. +- **Cluster (pass 39):** the ordering and the gate blocks 4 of 5; metadata 1; the instrument 0. +- **require↔withdraw:** none across the three. +- **Tells:** pass 37 one (Blocker rose to 1), pass 38 one (count rose), pass 39 **zero**. + +**The mechanism behind the plateau at 3, named rather than described.** Every rule this change adds +owes **three further sites**: a limb in §G's membership test, a strict reading in §H's unknown-start +list, and any §F pointer that states the rule at its own entry point. Pass 39's four Majors were +three of those fan-outs plus one enumeration. **This is checkable before a pass rather than after +one**, and is now part of the pre-review sweep: after adding a rule, check §G, §H and §F. + +**The enumerations keep costing.** §C's branch list cost two rounds, the both-routes pair cost one, +and the source-block list cost pass 39 finding 1. Each is now either removed or explicitly marked +as examples rather than the set. + +**§G has taken a Major at passes 32, 34, 35, 38 and 39** — five rounds, each naming a different +limb, each fix correct for what it named. It is a semantic membership test over the whole section, +which is why it is the last paragraph to settle; it is recorded as a regeneration chain and has not +reached the two-tell threshold on its own. ## Pass-36 report — zero tells, B+M 9 → 5 → 4 → 3 diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index 772b597..3606ff4 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -103,7 +103,8 @@ rather than its position**, so reordering them breaks no reference. **The source-block branch, read first.** **Where any unmet closure condition's own source prescribes stop-and-surface** — a profile present but unresolvable, governing headers that -disagree, a `Story:` header that cannot be read — **the cycle stays open, that source decides what +disagree, a `Story:` header that cannot be read, an unobservable counterfactual, and **any other +source rule that prescribes it; the list is examples and not the set** — **the cycle stays open, that source decides what must be repaired or answered, and no further pass runs while its block stands.** It is neither a suspension nor a continue and needs no name and no procedure of its own: the source rule carries both, and this ordering's part is to send the reader there rather than to run a pass over a cycle @@ -687,9 +688,9 @@ across C 1014–1015 and W 1198–1199. ``` Those have their own terminal actions and this paragraph changes none of them: on a STOP you still stop, and **neither a human's general assent nor this record** lets an agent close or -continue a cycle. **The answer a suspension asks for is not assent of that kind**: it is the -answer the closure ordering prescribes for that suspension, given on the question that suspension -raised, and both which answer that is and what it produces are stated there. +continue a cycle. **The answers a suspension asks for are not assent of that kind**: they are the +answers the closure ordering prescribes, one per surfacing route, and both which answers those are +and what they produce are stated there. ``` *Why (pass 19 finding 4):* the live sentence says no human answer lets an agent continue a cycle, while the ordering makes **continue** the prescribed answer that restarts a parked one. Left as @@ -796,8 +797,8 @@ and the requirement is that the adopted definitions **agree**, not merely that a present. **Membership is decided by a test a reader can apply to the text in front of them, with no list to consult, and the test reads what a rule states rather than what changing it would do: a live rule belongs to this contract when what it says **defines the validity of an input the -closure ordering reads, or how that input is read**, which branch a pass takes, what a hold is or what discharges it, **what a suspension asks or what -state its answer produces**, +closure ordering reads, or how that input is read**, which branch a pass takes, what a hold is or what discharges it, **what a suspension asks, or what state its answer or an +incomplete closing act produces**, whether a cycle may close **or may terminate without running a pass at all, the Gate-B triviality skip being the one such route and its eligibility test therefore a member**, or the production, identity or transport of **any §5 cycle record, required or optional** — §5 entire and not the @@ -923,7 +924,9 @@ dash-delimited list is given, the last item being the addition. the curve duty owed, the nonce duties at their strictest, **every suspension binding, since starting rules that cannot be established cannot be read as having waived an open hold**, **the repeated-dismissal cleanliness exclusion unavailable, a cycle that cannot establish its starting -rules being unable to establish that they contained it**, and **every closure condition and +rules being unable to establish that they contained it**, **the parked state binding after a closing act that cannot +be repaired, a cycle whose starting rules cannot be established being the last one that should be +left with no terminal transition**, and **every closure condition and pass-cost rule this change ships owed rather than waived — Gate A's content condition and its commit-carry duty, and the further pass an assigned-fix-set change costs — since a rule that cannot be established as absent is cheaper to owe than to skip** — From d26d0979c2e68f93d5b14041196a454074213401 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sun, 13 Sep 2026 16:42:35 +0200 Subject: [PATCH 069/181] docs(specs): apply pass 40; the pre-pass fan-out check held MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Zero tells. Findings 5 to 3, Blocker+Major 4 to 3. The three-site check run before this pass produced no fan-out Majors, which is the first round where that was true. - §A3's branch-agreement pointer named only a shared commit, while the standing rule requires the kept baseSha AND headSha to be exactly equal before the branches are summed. A shared head over different bases is two ranges, not one pass, and could have fed closure - §G gains the limb for what a gate's closing act is. The test covered closure permission and the state an incomplete act produces, but not the operation itself, which A2 and A3 define and a downstream mismatch would break - §G gains the limb for which version of these rules governs a cycle. Its own opening declares the unknown-start activation semantics a member while the test had no category for a sentence deciding rule applicability Every member §G's opening declares now has a limb in its test: the nonce and slot naming through record identity, the provenance line and curve through record production, the carry rule through transport, the activation semantics through rule applicability, and the ordering through the branch, hold, suspension, closure, closing-act and passless-termination limbs. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .context/codex-reviews/gate-a-spec-awsf1ec771-pass-40.md | 4 ++++ .../2026-09-10-loop-rule-consolidation-target-text.md | 8 +++++--- 2 files changed, 9 insertions(+), 3 deletions(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-40.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-40.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-40.md new file mode 100644 index 0000000..b768f4a --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-40.md @@ -0,0 +1,4 @@ +MAJOR | high | §A3 "both branches issued against the same commit" | The compressed pointer drops a load-bearing part of the standing branch-agreement rule: the live sentence, wrapping across CLAUDE.md 963–965 and workflow-init.md 1147–1149, requires the retained `baseSha` and `headSha` values both to be exactly equal before the branches are summed, while A3 names only a shared commit | Two branches can share a head but use different bases, select different ranges, and still be read through A3 as one valid `full` pass that may feed closure | Remove this enumeration and point to the standing branch-agreement rule in full +MAJOR | high | §G "whether a cycle may close" | A membership limb is missing: the test covers closure permission and the state an incomplete closing act produces, but not what the successful closing act is; A2 expressly says its two cases decide how a cycle closes and never whether it may, and A3 likewise defines the amend/reset operation | A downstream adoption can retain or mismatch a gate's closing operation while §G classifies that sentence as outside the contract, allowing a cycle to leave a `WIP:` ancestor or carry content and records through the wrong commit shape without the required partial-adoption stop | Add the narrow limb "what a gate's closing act is" +MAJOR | high | §G "unknown-start activation semantics" | A membership limb is missing: the opening declares the unknown-start activation semantics a contract member, but the operative test has no category for a sentence that decides which version of the rules governs a cycle; the standing activation sentences decide rule applicability, not input validity or interpretation, a pass branch, a hold, a suspension, closure permission, passless termination, or record production, identity or transport | A downstream project can carry a mismatched activation boundary and apply new rules to an old running cycle or old rules after activation without §G requiring the adoption to stop and reconcile | Add the narrow limb "which version of these rules governs a cycle" +END OF FINDINGS (3 total) diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index 3606ff4..48c9916 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -409,8 +409,9 @@ and the next paragraph says why. reviews a **diff** identified by `baseSha` and `headSha` rather than a text handed to the reviewer, so there is no reviewed text to compare an artifact against, and **nothing is put in its place** — the bullet in §I records what that leaves open. What this gate does have is already in this section -and is cited rather than restated: the range those two names fix, **both branches -issued against the same commit**; **a re-review after every fix**, a fix changing the artifact so +and is cited rather than restated: the range those two names fix, **both branches issued +with `baseSha` and `headSha` exactly equal before they are summed**, a shared head over different +bases being two ranges and not one pass; **a re-review after every fix**, a fix changing the artifact so the prior review no longer covers it; **a fix that changes specified behaviour updating the spec in the same commit**, so the re-review covers both; the battery and the mode-derived evidence the profiles section obliges before a call; and the **evidence entry**, revalidated as that section @@ -798,7 +799,8 @@ present. **Membership is decided by a test a reader can apply to the text in fro no list to consult, and the test reads what a rule states rather than what changing it would do: a live rule belongs to this contract when what it says **defines the validity of an input the closure ordering reads, or how that input is read**, which branch a pass takes, what a hold is or what discharges it, **what a suspension asks, or what state its answer or an -incomplete closing act produces**, +incomplete closing act produces**, **what a gate's closing act is**, **which version of these rules +governs a cycle**, whether a cycle may close **or may terminate without running a pass at all, the Gate-B triviality skip being the one such route and its eligibility test therefore a member**, or the production, identity or transport of **any §5 cycle record, required or optional** — §5 entire and not the From 3b61fe3a4ad3804687252d41877ca29d79f8e492 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sun, 13 Sep 2026 16:56:26 +0200 Subject: [PATCH 070/181] docs(specs): apply pass 41; stop compressing cited rules, stop enumerating branches MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Zero tells. Findings 3 to 2, Blocker+Major 4, 3, 2 across three passes. Both Majors are the same class and both are mine: - §A3's pointer named only the equality of the kept baseSha and headSha and dropped the rest of the standing branch-agreement rule — that HEAD is resolved to its full 40-character object name before each call, passed explicitly rather than symbolically, and kept with that branch's result. Two calls can carry equal symbolic values and resolve different commits across an intervening WIP amend. The rule is now cited entire - §F item 9 said the ordering decides what follows and then enumerated only continue and suspension, while changed evidence can take the source-block branch. The enumeration is removed rather than extended, which is the sixth round an enumeration has cost and the standing preference for this file Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-41.md | 3 +++ ...09-10-loop-rule-consolidation-target-text.md | 17 +++++++++-------- 2 files changed, 12 insertions(+), 8 deletions(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-41.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-41.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-41.md new file mode 100644 index 0000000..f98751e --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-41.md @@ -0,0 +1,3 @@ +MAJOR | high | §A3 "`baseSha` and `headSha` exactly equal" | The compressed citation still drops the standing branch-agreement rule's requirement to resolve `HEAD` to its full 40-character object name before each call, pass that explicit value rather than symbolic `HEAD`, and retain it with the branch result; that sentence occurs in both copies and wraps across CLAUDE.md 960–965 and workflow-init.md 1144–1149 | Two calls can carry textually equal symbolic `headSha` values while resolving different commits after an intervening `WIP:` amend, yet A3 permits their files to be summed into a pass that can feed closure | Remove this enumeration and point to the standing branch-agreement rule in full +MAJOR | high | §F item 9 "where it selects the continue branch" | The replacement says the ordering decides what follows but then enumerates only continue and suspension; the standing unobservable-counterfactual rule, which occurs in both copies and wraps across CLAUDE.md 717–719 and workflow-init.md 903–905, can instead make changed evidence take A1's source-block branch | A Gate-B cycle whose evidence revalidation exposes an unrepairable counterfactual gap can follow this source paragraph into fix and re-review without first repairing or answering the source-owned block | Replace the two-case enumeration with a pointer to the closure ordering in full, retaining fix and re-review only as the duty when that ordering reaches continue +END OF FINDINGS (2 total) diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index 48c9916..818e888 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -409,9 +409,10 @@ and the next paragraph says why. reviews a **diff** identified by `baseSha` and `headSha` rather than a text handed to the reviewer, so there is no reviewed text to compare an artifact against, and **nothing is put in its place** — the bullet in §I records what that leaves open. What this gate does have is already in this section -and is cited rather than restated: the range those two names fix, **both branches issued -with `baseSha` and `headSha` exactly equal before they are summed**, a shared head over different -bases being two ranges and not one pass; **a re-review after every fix**, a fix changing the artifact so +and is cited rather than restated: the range those two names fix and **the branch-agreement rule +entire**, which says how `headSha` is resolved and passed, what is kept with each branch result and +what must be equal before two branches are summed — cited here and not compressed, a part of it +dropped being a range nobody checked; **a re-review after every fix**, a fix changing the artifact so the prior review no longer covers it; **a fix that changes specified behaviour updating the spec in the same commit**, so the re-review covers both; the battery and the mode-derived evidence the profiles section obliges before a call; and the **evidence entry**, revalidated as that section @@ -773,11 +774,11 @@ about a mechanism rather than an entry point carrying an unqualified instruction **9. The evidence-entry revalidation remedy** (the profiles section). It wraps across C 728–729 and W 914–915. ``` -If revalidation changes the entry, the pass was read against an entry that no longer stands: the -pass did not close, so the closure ordering decides what happens next: where it selects the -continue branch, fix and re-review, and where it selects a suspension, the answer comes first. The -pass that follows is read by that ordering like any other and closes only if it reaches closure, -on the entry revalidated for it. +If revalidation changes the entry, the pass was read against an entry that no longer stands: the pass +did not close, so **what happens next is the closure ordering's, read there in full** — this +paragraph states the fix and the re-review it owes and never which branch the pass takes. The pass +that follows is read by that ordering like any other and closes only if it reaches closure, on the +entry revalidated for it. ``` *Why (pass 30 finding 3):* the live sentence is a complete instruction to whoever enters through the profiles section — fix, re-review, close — and under the ordering a non-closing pass takes any From 95f8591935cd2cc9b26b57a45d3188940c8a45a4 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sun, 13 Sep 2026 17:14:25 +0200 Subject: [PATCH 071/181] =?UTF-8?q?docs(specs):=20apply=20pass=2042;=20?= =?UTF-8?q?=C2=A7F=20item=204=20stops=20counting=20answers?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Zero tells. One finding, Blocker+Major 2 to 1. §F item 4 had taken back the enumeration it was repaired to avoid: "one per surfacing route" is wrong where both health readings apply, since the ordering gives them one shared continue-or-stop question answered once. The paragraph now points at the ordering without counting anything, which is what its own rationale requires of it. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .context/codex-reviews/gate-a-spec-awsf1ec771-pass-42.md | 2 ++ .../specs/2026-09-10-loop-rule-consolidation-target-text.md | 4 ++-- 2 files changed, 4 insertions(+), 2 deletions(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-42.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-42.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-42.md new file mode 100644 index 0000000..716c33c --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-42.md @@ -0,0 +1,2 @@ +MAJOR | high | §F item 4 "one per surfacing route" | The replacement says a suspension has one answer per surfacing route, but §A1 says the clearly-stuck and two-tell routes raise one shared continue-or-stop question and one answer carrying both reasons ends both; its own rationale also says this paragraph must not enumerate the answer rules | When both health readings apply, a reader entering through the human-exception paragraph can wait for a second answer that does not exist or treat the two routes inconsistently, preventing the suspended cycle from resuming deterministically | Remove "one per surfacing route" and leave the pointer to the closure ordering entire +END OF FINDINGS (1 total) diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index 818e888..07cfa3d 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -691,8 +691,8 @@ across C 1014–1015 and W 1198–1199. Those have their own terminal actions and this paragraph changes none of them: on a STOP you still stop, and **neither a human's general assent nor this record** lets an agent close or continue a cycle. **The answers a suspension asks for are not assent of that kind**: they are the -answers the closure ordering prescribes, one per surfacing route, and both which answers those are -and what they produce are stated there. +answers the closure ordering prescribes, and both which answers those are and what they produce are +stated there. ``` *Why (pass 19 finding 4):* the live sentence says no human answer lets an agent continue a cycle, while the ordering makes **continue** the prescribed answer that restarts a parked one. Left as From 4a440075c27ed215dc9cf705b187acb14f330b38 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sun, 13 Sep 2026 17:37:11 +0200 Subject: [PATCH 072/181] docs(specs): apply pass 43; the branches are read once, before the act MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit One tell — the count rose 1 to 6. Blocker+Major 1 to 3. - the branch leads classified every pass that did not close, which collided with the failed-act rule keeping such a pass at the closure step. The branches are now stated as read ONCE, on the pass, before any closing act is attempted, so a pass that reached the act is not re-classified by what the act does. That removes the collision at its root rather than excepting it - accepting an out-of-set Minor changes the assigned fix set, which costs a further pass, while Severity said Minor and Nit never iterate. Both are true and the conflict was in attributing the cost: a Minor's severity buys no repair round and no further pass, and where accepting one costs a pass that cost is the set change's, stated at the absorb paragraph. The continue branch's own never-iterated wording is about repairs and needed no change - §F item 9 had lost the standing "fix, re-review" imperative to a self-referential claim that the paragraph states it. The imperative is back and the branch decision stays with the ordering Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-43.md | 7 +++++++ ...-09-10-loop-rule-consolidation-target-text.md | 16 ++++++++++------ 2 files changed, 17 insertions(+), 6 deletions(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-43.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-43.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-43.md new file mode 100644 index 0000000..deb5a86 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-43.md @@ -0,0 +1,7 @@ +MAJOR | high | §§A1 and D "failed closing act" | A1's specific rule says a pass whose closing act fails remains at the closure step for retry or park, but the later suspension and continue leads classify every pass that did not close, and §D says a suspension is outranked only when closing completes | A repairable failure can be sent to a health question or an unnecessary new pass instead of retrying, while an unrepairable failure has both park and branch outcomes | Replace the suspension and continue leads with rules for "a pass that reaches this step" and make §D defer to the ordering entire, so a pass that entered the failed-act handler stays there until retry or park +MAJOR | high | §§B and E "accepted Minor" | Accepting an out-of-set Minor changes the assigned fix set and §B therefore mandates at least one further pass, while §E and the retained deciding-severity sentence still say Minor and Nit findings are collected and "never iterate" | The same accepted Minor both requires and forbids another pass, so an agent must violate either the settled set-change rule or the standing severity instruction | Replace both unqualified "never iterate" phrases in Mechanics with "severity alone requires no repair or further pass", leaving the independently owed set-change pass intact +MAJOR | medium | §F item 9 "evidence-entry revalidation remedy" | The replacement removes the standing imperative "fix, re-review" and preserves it only through the self-referential claim that "this paragraph states the fix and the re-review it owes"; neither that sentence nor the unchanged preceding text identifies the repair action | A changed evidence entry can be routed through the ordering without a clear duty to repair the evidence mismatch before the next review, dropping a load-bearing part of the standing remedy | Replace the meta-claim with the direct duty to repair the cause and re-review against the resulting validated entry, then leave branch selection and closure to the ordering +MINOR | high | §§A3 and F item 1 "counter state" | The target says what the reset erases is counter state, but codex-gate.sh lines 886–888 delete the recorded fingerprint as well as both Gate-B count files; A3's preceding sentence itself separately names review state and the counter | A reader can believe the reviewed fingerprint survives a stray or failed commit attempt and misdiagnose the hook's subsequent "not run" or freshness result | Replace the exhaustive counter-state wording with "the recorded Gate-B fingerprint and counters", or point to §F item 5's broader state-clearing statement +NIT | high | §F item 4 line citation | The quoted human-exception sentence begins on CLAUDE.md line 1013 and workflow-init.md line 1197, then wraps through lines 1015 and 1199; the target cites only C 1014–1015 and W 1198–1199 | The mechanical citation omits the first line of the sentence it claims to locate | Change the ranges to C 1013–1015 and W 1197–1199 +NIT | high | §F item 8 line citation | The quoted profile-change sentence begins on CLAUDE.md line 751 and workflow-init.md line 937, then wraps through lines 753 and 939; the target cites only C 752–753 and W 938–939 | The mechanical citation omits the first line of the sentence it claims to locate | Change the ranges to C 751–753 and W 937–939 +END OF FINDINGS (6 total) diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index 07cfa3d..199a531 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -99,7 +99,9 @@ identity, sameness or recovery rule in these rules; the optional `-disposi advisory and authoritative for nothing. **The four branches are named, and a cross-reference anywhere in this section names the branch -rather than its position**, so reordering them breaks no reference. +rather than its position**, so reordering them breaks no reference. **They are read once, on the +pass, before any closing act is attempted** — so a pass that reached the act has already been +classified and is not classified again by what the act does. **The source-block branch, read first.** **Where any unmet closure condition's own source prescribes stop-and-surface** — a profile present but unresolvable, governing headers that @@ -608,8 +610,10 @@ suspensions and what its answer produces are both stated. carrying the one-line why this section already requires, that the finding is not true of the artifact. A dismissal does not rewrite the pass that found it and the later clean pass is still owed; **a dismissal is not a decline**, a dismissal saying the finding is false and a decline - being the user's decision that a **true** finding stays outside the set. Minor · Nit → collect, - never iterate. + being the user's decision that a **true** finding stays outside the set. Minor · Nit → collect; + **their severity buys no repair round and no further pass.** Where accepting one into the + assigned fix set costs a pass, that cost is the **set change's** and is stated at the absorb + paragraph, not this severity's. ``` **The handed-over question, replaced by its answer.** @@ -774,9 +778,9 @@ about a mechanism rather than an entry point carrying an unqualified instruction **9. The evidence-entry revalidation remedy** (the profiles section). It wraps across C 728–729 and W 914–915. ``` -If revalidation changes the entry, the pass was read against an entry that no longer stands: the pass -did not close, so **what happens next is the closure ordering's, read there in full** — this -paragraph states the fix and the re-review it owes and never which branch the pass takes. The pass +If revalidation changes the entry, the pass was read against an entry that no longer stands: **fix the +entry and re-review on it**, and **which branch the pass takes meanwhile is the closure ordering's, +read there in full** — this paragraph states the repair and never the branch. The pass that follows is read by that ordering like any other and closes only if it reaches closure, on the entry revalidated for it. ``` From 1c858acf20ee0481c23c598c6f9d6a3c17a22def Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sun, 13 Sep 2026 17:48:24 +0200 Subject: [PATCH 073/181] docs(specs): apply pass 44; reword the branch leads, add the tenth standing sentence MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit One tell — the count rose 6 to 7. Blocker+Major flat at 3. All three Majors say the pass-43 repairs did not go far enough. - the "read once" sentence added a claim without changing the leads it was about. The leads now classify without mentioning closing at all: the suspension branch is reached where the clean-completion branch did not take the pass and a suspension applies, and the continue branch where neither of the two above took it. Closing is downstream of classification, so a failed act cannot re-enter - §F item 5 commanded `git commit --amend` unconditionally while §A3 names a shape where the amend replaces the stray tip and leaves the WIP as an ancestor. It now performs the closing act as the ordering describes it, the amend being the ordinary shape and not the only one - a TENTH falsified standing sentence: Mechanics' severity fallback ends "collect, never iterate", verified at CLAUDE.md 788-789 and workflow-init.md 974-975. Accepting an out-of-set Minor changes the fix set and that change costs a pass — owed by the decision, not by the severity. §F's two counts move with it, and it is the seventh sentence sharing the section's mechanism Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-44.md | 8 +++++ ...-10-loop-rule-consolidation-target-text.md | 33 ++++++++++++++----- 2 files changed, 32 insertions(+), 9 deletions(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-44.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-44.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-44.md new file mode 100644 index 0000000..b541191 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-44.md @@ -0,0 +1,8 @@ +MAJOR | high | §§A1 and D "failed closing act" | The new one-time classification sentence says a pass that reached a closing act is not classified again and the failed-act rule returns it to the closure step, but the later suspension lead still classifies "a pass ... unless it closed", the continue lead still covers a pass that "neither closes nor suspends", and §D says clean completion outranks its stop only by actually closing | A repairable failed act can still be sent to a health suspension or another pass, while an unrepairable one has both the prescribed parked state and a second branch outcome | Replace the suspension and continue lead-ins and §D's order clause so they apply only to the single pre-act classification and leave every failed attempt in the failed-act handler until retry or park +MAJOR | high | §E and standing Mechanics "Minor or below" | §E now correctly says a Minor's severity alone buys no further pass, but the next live Mechanics sentence remains in both copies: "If you cannot name both, the finding is Minor or below: collect, never iterate"; that unqualified command conflicts with §B when accepting an out-of-set Minor changes the assigned fix set and therefore costs a further pass | An agent entering through the retained severity procedure can refuse the settled set-change pass, leaving the cycle short of a closure duty | Replace the retained "collect, never iterate" sentence with the same severity-scoped rule §E uses, so it forbids a severity-driven repair round without forbidding a pass owed by a set change +MAJOR | high | §§A3 and F item 5 "stray non-amending commit" | A3 says a successful non-amending commit on top of the active WIP leaves the WIP as an ancestor and still requires no WIP commit in history, while F item 5 unconditionally commands `git commit --amend` and says it replaces the WIP; in this named shape the amend replaces the stray tip, and the retained soft-reset rule applies only when several WIP snapshots piled up | Following the concrete Mechanics instruction leaves the WIP in history, so the purported closing act does not meet A3's own duty and the cycle has conflicting closure instructions | Make F item 5 conditional on the WIP being the tip and point the already-named ancestor shape to A3's duty, leaving its exact git sequence to the plan as required +MINOR | high | §F item 1 "counter state" | The exhaustive claim "What the hook loses is its counter state" omits the recorded Gate-B fingerprint: `codex-gate.sh` lines 886–888 delete `state_file` as well as `count_file` and `fresh_file`, and A3 itself distinguishes review state from the counter | A reader can wrongly believe the prior fingerprint survives the commit attempt and misdiagnose the hook's next freshness or not-run result | Replace the enumeration with a pointer to item 5's state-clearing statement, or name the recorded fingerprint and both counters +NIT | high | §F item 4 line citation | The human-exception sentence begins on CLAUDE.md line 1013 and workflow-init.md line 1197 with "Those have their", then wraps through lines 1015 and 1199; the stated ranges C 1014–1015 and W 1198–1199 omit that first wrapped line | The mechanical location claim does not contain the complete sentence it says it locates | Change the ranges to C 1013–1015 and W 1197–1199 +NIT | high | §F item 8 line citation | The profile-change sentence begins midway through CLAUDE.md line 751 and workflow-init.md line 937 with "Inside an active Gate-B cycle", then wraps through lines 753 and 939; the stated ranges C 752–753 and W 938–939 omit that first wrapped line | The mechanical location claim does not contain the complete sentence it says it locates | Change the ranges to C 751–753 and W 937–939 +NIT | high | §F item 5 quoted standing sentence | The rationale says the live sentence says "after the final clean pass, close it with `git commit --amend`", but that quoted text occurs in neither prompt copy even after joining wrapped lines: both sources say `git commit --amend -m ""` | The mechanical quote check fails and a reader cannot locate the purported quotation verbatim | Quote the complete command from the standing sentence or mark the omission with an ellipsis outside the code span +END OF FINDINGS (7 total) diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index 199a531..986dc53 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -204,8 +204,9 @@ repaired after the pass that raised it discharges the resolve duty without makin so a cycle can owe nothing and still hold no pass it may close on. The one termination that is not a pass outcome is the Gate-B triviality skip, which runs no passes and is outside this ordering. -**Then the suspension branch, which a pass reaches unless it closed the cycle.** **Clean completion -outranks a suspension by closing, not by being eligible**: a pass that took the branch above, met +**Then the suspension branch, which a pass reaches where the clean-completion branch did not take +it and a suspension applies to it.** **Clean completion outranks a suspension by taking the pass to +the closing act, not by eligibility alone**: a pass that took the branch above, met every closure condition and had the closing act performed has ended the cycle, and a suspension has nothing left to suspend. **A pass that did not close reaches this branch whatever its cleanliness, where a suspension applies to it** — a clean pass below the floor, and equally an @@ -225,8 +226,8 @@ decision made by omission. A finding the clearly-stuck reading surfaces that als trigger takes the scope stop's answers at that same surface, so it is not asked twice; the two-tell stop surfaces tells and not a finding. -**Otherwise the continue branch, which no source block reaches: where none stands, a pass that -neither closes nor suspends continues** — the loop +**Otherwise the continue branch, which no source block reaches: where none stands and neither of +the two branches above took the pass, it continues** — the loop runs another pass on the **current** artifact, revised where the severity and scope rules require a repair and unrevised where they do not. **An eligible pass with an unmet closure condition lands here**, and like every other non-closing pass **only where no suspension applies to it**: clean @@ -650,10 +651,10 @@ demotes it — the two counts are meant to differ. --- -## F. The nine standing sentences this change falsifies — REPLACED +## F. The ten standing sentences this change falsifies — REPLACED -Each is a live sentence that the block makes wrong. All nine are **known contradictions** and -none is deferred. **Six of them share one mechanism** — an entry point other than the ordering +Each is a live sentence that the block makes wrong. All ten are **known contradictions** and +none is deferred. **Seven of them share one mechanism** — an entry point other than the ordering carrying an unqualified instruction — which is why each is **replaced** rather than given an exception to point at. @@ -713,8 +714,10 @@ their definition site. W 1011–1012. ``` **Finishing the cycle:** once the closure ordering reaches a Gate-B cycle's closing act — an -eligible pass with every closure precondition holding, never a clean pass on its own — close it -with `git commit --amend -m ""`; that replaces the WIP commit. **The hook treats any +eligible pass with every closure precondition holding, never a clean pass on its own — perform that gate's +closing act as the ordering describes it — ordinarily `git commit --amend -m ""`, +which replaces the WIP commit, and the shape the ordering names instead wherever the amend would +leave a `WIP:` commit in history. **The hook treats any non-`WIP` commit *attempt* as a Gate-B boundary and clears its state even where the command fails**, so a failed closing act leaves that counter cleared and no commit made — a fact about the counter and not about the cycle. This section states the operation and never whether the cycle may @@ -775,6 +778,18 @@ hook's counter. Left standing it tells an author that a stray commit destroyed r cannot reach. **This one does not share the section's shared mechanism** — it is a false claim about a mechanism rather than an entry point carrying an unqualified instruction. +**9b. The severity-deciding fallback** (Mechanics · Severity). It wraps across C 788–789 and +W 974–975. +``` +If you cannot name both, the finding is Minor or below: collect; **its severity buys no repair +round and no further pass**, and any pass a later scope decision costs is that decision's. +``` +*Why (pass 44 finding 2):* the live clause ends "collect, never iterate", which is an unqualified +command at an entry point other than the ordering. Accepting an out-of-set Minor changes the +assigned fix set, and that change costs a further pass — owed by the set change, not by the +Minor's severity. Left standing, the two instructions decide the same pass in opposite directions. +It is the seventh sentence sharing this section's mechanism. + **9. The evidence-entry revalidation remedy** (the profiles section). It wraps across C 728–729 and W 914–915. ``` From 4b713dd5583a656638335e482cd9b56b422fc4f2 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sun, 13 Sep 2026 17:57:00 +0200 Subject: [PATCH 074/181] docs(specs): apply pass 45; classification stops referring to closing anywhere MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Zero tells. Findings flat at 7, Blocker+Major 3 to 1 — the lowest of the cycle. The one Major is the third round on the same subject: rewording the branch leads at pass 44 left the sentences inside them still tying classification to whether the pass closed. Swept all four this time rather than the one the finding named: - the suspension branch's body now reads "a pass the clean-completion branch did not take", not "a pass that did not close" - the continue branch reads "clean completion did not take it" - its precedence clause reads "only the clean-completion branch outranks a suspension" - §D reads "outranks it by taking the pass to the closing act and only then" Classification is now stated entirely in terms of which branch took the pass, never in terms of an act that happens after classification. Verified by sweep: no prompt block still ties a branch to closing. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-45.md | 8 ++++++++ ...26-09-10-loop-rule-consolidation-target-text.md | 14 +++++++------- 2 files changed, 15 insertions(+), 7 deletions(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-45.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-45.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-45.md new file mode 100644 index 0000000..661a154 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-45.md @@ -0,0 +1,8 @@ +MAJOR | high | §§A1 and D "failed closing act" | A1 says the four branches are read once before any closing act and that a failed act stays at the closure step, but the surviving sentence "A pass that did not close reaches this branch whatever its cleanliness" sends that same pass to suspension where one applies, while §D says clean completion outranks the two-tell stop only "by closing the cycle and only then" | A repairable act failure can detour to a health suspension instead of retrying, and an unrepairable failure has both the prescribed parked state and a suspension outcome | Restrict the suspension sentence to passes the clean-completion branch did not take, and replace §D's closing-dependent ranking with a pointer to the closure ordering in full +MINOR | high | §F item 1 "What the hook loses" | The proposed exhaustive claim says the hook loses "its counter state", but `codex-gate.sh` lines 886–888 delete `state_file` as well as `count_file` and `fresh_file`; A3 itself distinguishes the Gate-B review state from the counter | A reader can wrongly believe the recorded fingerprint survives a stray or failed commit attempt and misdiagnose the hook's next freshness or not-run result | Remove the enumeration and point to A3's state-clearing statement +MINOR | high | design §1 lines 34–35 versus target §F | The governing design says there are nine falsified standing sentences and that target §F states that count, while target §F now enumerates and states ten | The design and target give different implementation scope, so a plan reader can omit the newly added severity-fallback replacement | Remove the cross-file count from the design and cite target §F's complete enumeration +MINOR | high | design §7 lines 322–324 versus target §F | The verification section says target §F has nine falsified sentences of which six share the entry-point mechanism, while target §F states ten and seven respectively | Verification and its evidence entry can be planned against one fewer replacement and one fewer shared-mechanism case than the target requires | Replace both repeated counts with a pointer to target §F's enumeration +NIT | high | §F item 4 line citation | The quoted human-exception sentence begins on CLAUDE.md line 1013 and workflow-init.md line 1197 with "Those have their", then wraps through lines 1015 and 1199; the stated ranges C 1014–1015 and W 1198–1199 omit that first line | The mechanical location claim does not contain the complete sentence it says it locates | Change the ranges to C 1013–1015 and W 1197–1199 +NIT | high | §F item 8 line citation | The profile-change sentence begins midway through CLAUDE.md line 751 and workflow-init.md line 937 with "Inside an active Gate-B cycle", then wraps through lines 753 and 939; the stated ranges C 752–753 and W 938–939 omit that first line | The mechanical location claim does not contain the complete sentence it says it locates | Change the ranges to C 751–753 and W 937–939 +NIT | high | §F item 5 quoted standing sentence | The rationale says the live sentence says "after the final clean pass, close it with `git commit --amend`", but neither prompt copy contains that quoted text after joining its wrap; both say `git commit --amend -m ""` | The required mechanical quote check fails and a reader cannot locate the purported quotation verbatim | Quote the complete command from the standing sentence or mark the omitted arguments with an ellipsis outside the code span +END OF FINDINGS (7 total) diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index 986dc53..cef2449 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -208,9 +208,9 @@ no passes and is outside this ordering. it and a suspension applies to it.** **Clean completion outranks a suspension by taking the pass to the closing act, not by eligibility alone**: a pass that took the branch above, met every closure condition and had the closing act performed has ended the cycle, and a suspension has -nothing left to suspend. **A pass that did not close reaches this branch whatever its -cleanliness, where a suspension applies to it** — a clean pass below the floor, and equally an -eligible pass the cycle's unmet conditions kept from closing. Cleanliness is what this branch stops +nothing left to suspend. **A pass the clean-completion branch did not take reaches this branch +whatever its cleanliness, where a suspension applies to it** — a clean pass below the floor, and +equally an eligible pass whose unmet closure conditions kept that branch from taking it. Cleanliness is what this branch stops asking about; whether a suspension applies is still what puts a pass here, and where none does the continue branch has it. That is what makes "clean completion outranks the two-tell stop" executable rather than asserted, and cleanliness alone never decides it. **What D2 and D3 forbid is @@ -231,10 +231,10 @@ the two branches above took the pass, it continues** — the loop runs another pass on the **current** artifact, revised where the severity and scope rules require a repair and unrevised where they do not. **An eligible pass with an unmet closure condition lands here**, and like every other non-closing pass **only where no suspension applies to it**: clean -completion did not close it, so the loop continues on whatever the unmet condition requires — most +completion did not take it, so the loop continues on whatever the unmet condition requires — most often a repair still owed from an earlier pass. A below-floor clean pass lands here on the same terms; where a suspension does apply, the suspension branch has already taken it, because only -closing outranks a suspension. So does a pass whose only findings are Minors and +the clean-completion branch outranks a suspension. So does a pass whose only findings are Minors and Nits **and which carries no scope-stop trigger** — those are collected and never iterated and may leave nothing to revise, while a Minor or Nit that is out of set or opens a new structural question carries a trigger like any other finding, is not clean, and has already been taken by the @@ -588,8 +588,8 @@ continuation next to the conditional one and give the same pass two answers. **`e7`, the threshold.** Gains one clause; the sentence is given entire. ``` **Any two present makes stop-and-surface mandatory, not discretionary** — read **after** the -clean-completion branch of the closure ordering, which outranks it **by closing the cycle and only -then** — and you report the tells and hand the decision to the user, and the "clearly stuck" +clean-completion branch of the closure ordering, which outranks it **by taking the pass to the +closing act and only then** — and you report the tells and hand the decision to the user, and the "clearly stuck" reading above is not a precondition for it. ``` From 6c88db5a4f7248aa8fc7e166ca52cea52754986b Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sun, 13 Sep 2026 18:08:56 +0200 Subject: [PATCH 075/181] docs(specs): apply pass 46; source-block repair reopens the reading MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit TWO TELLS at pass 46 — the count rose 7 to 9 and the Blocker count failed to fall, 0 to 1. §5 makes stop-and-surface mandatory at two. Daniel answered that question in advance on 2026-09-13 at 17:02: no stops without an absolute block. The tells are surfaced here and in the resume note, the decision was already given, and what the rule forbids is honoured: no pass credited clean, every finding open, the cycle open. Recorded explicitly, because continuing past a mandatory stop on an agent's own judgement is a gate-off path §5 names. The Blocker is real: a source block repaired without moving any pass-cost value left a clean eligible pass unclassifiable — no new pass owed, no branch to the act, none to a suspension. The source-block branch now says the ordering is read again on that pass once its condition is repaired, no act having been attempted on it. The read-once rule is about a pass that reached the act, not about one a block held before it. An ELEVENTH falsified standing sentence: the evidence-entry revalidation trigger names only "the cycle-closing amend", verified at CLAUDE.md 726-727 and workflow-init.md 912-913, while §A3 closes through reset plus a single commit in two shapes where no amend occurs. Scoped to the amend, the final revalidation is owed on one path and skipped on the others. §F item 7 likewise stops naming the amend and names the commit the closing act produces, so a Gate-B human-exception record survives every closing shape. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-46.md | 10 +++++ .../gate-a-spec-awsf1ec771-resume.md | 38 +++++++++++++++++++ ...-10-loop-rule-consolidation-target-text.md | 27 ++++++++++--- 3 files changed, 70 insertions(+), 5 deletions(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-46.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-46.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-46.md new file mode 100644 index 0000000..3c9476c --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-46.md @@ -0,0 +1,10 @@ +BLOCKER | high | §A1 "source-block branch" | The four branches are read once on the pass, and the source-block branch says only that no further pass runs while the block stands; when a repair clears a block without changing any value governed by a pass-cost rule, the pass cannot be classified again, no new pass is owed, and no branch reaches the closing act | A clean eligible pass can remain open after its source block is repaired with no route to close or suspend | State that clearing a source block sends the cycle through a new pass before the ordering is read again +MAJOR | high | standing Profiles "before the cycle-closing amend" | CLAUDE.md 725–729 and workflow-init.md 911–915 require final evidence-entry revalidation only before a cycle-closing amend, while §A3 also closes through reset plus a new commit when WIP snapshots piled up or a stray non-amending commit left a WIP ancestor; §F replaces only the remedy after a changed entry and leaves this trigger standing | A non-amend Gate-B closing path can skip the final evidence revalidation and commit a stale or newly unrepairable evidence claim | Replace "before the cycle-closing amend" with a pointer requiring revalidation immediately before the Gate-B closing act in every shape +MAJOR | high | §F item 7 "restated by the closing amend" | The replacement preserves the Gate-B instruction to put a human-exception record in the WIP commit and restate it through a closing amend, but §A3 permits reset plus a new commit in two closing shapes where no closing amend occurs and reset does not carry commit-message records | The source instruction and §A1's closing-commit duty disagree, and following the source alone loses the exception record from the commit that closes the cycle | Replace the Gate-B clause with a duty to carry the WIP record into whichever commit the Gate-B closing act uses +MINOR | high | §F item 1 "What the hook loses" | The replacement says the hook loses its counter state, but codex-gate.sh 886–888 deletes the recorded fingerprint state file as well as the pass and freshness count files; §A3 itself distinguishes Gate-B review state from its counter | A reader can wrongly expect the prior fingerprint to survive a stray or failed commit attempt and misdiagnose the hook's next freshness or not-run result | Remove the enumeration and point to §A3's state-clearing statement, or name the recorded fingerprint and both counters +MINOR | high | design §1 lines 34–35 versus target §F | The decision document says there are nine falsified standing sentences and that target §F states that count, while target §F states and enumerates ten | The two governing inputs give different implementation scope, so a plan reader can omit the newly added severity-fallback replacement | Remove the cross-file count from the design and cite target §F's enumeration +MINOR | high | design §7 lines 322–324 versus target §F | The verification section says target §F has nine falsified sentences and six sharing the entry-point mechanism, while target §F states ten and seven respectively | The planned verification and evidence entry can cover one fewer replacement and one fewer shared-mechanism case than the target requires | Replace both counts with a pointer to target §F's enumeration +NIT | high | §F item 4 line citation | The quoted human-exception sentence begins on CLAUDE.md line 1013 and workflow-init.md line 1197 with "Those have their", then wraps through lines 1015 and 1199; the stated ranges C 1014–1015 and W 1198–1199 omit its first line | The required mechanical location check does not contain the complete sentence it says it locates | Change the ranges to C 1013–1015 and W 1197–1199 +NIT | high | §F item 8 line citation | The profile-change sentence begins midway through CLAUDE.md line 751 and workflow-init.md line 937 with "Inside an active Gate-B cycle", then wraps through lines 753 and 939; the stated ranges C 752–753 and W 938–939 omit its first line | The required mechanical location check does not contain the complete sentence it says it locates | Change the ranges to C 751–753 and W 937–939 +NIT | high | §F item 5 quoted standing sentence | The rationale says the live sentence says "after the final clean pass, close it with `git commit --amend`", but neither prompt copy contains that quotation after joining its wrap; both say `git commit --amend -m ""` | The required exact-quote check fails and a reader cannot locate the purported quotation verbatim | Quote the complete command from the standing sentence or mark the omitted arguments with an ellipsis outside the code span +END OF FINDINGS (9 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index d8b1b1b..b1fae95 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -66,6 +66,44 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 37 | a37b5ab | 9→**3** | 0→**1** | 3→**2** | yes | one tell (Blocker 0→1). **Findings 3 — lowest of the cycle.** The Blocker was real and mine: an unrepairable closing failure had no terminal state; now PARKED, the state a stop answer already produces; session 01a09afd-8b4c-7712-93c1-6a3afa109d87 | | 38 | cc7f6dd | 3→**10** | 1→**0** | 2→**3** | yes | one tell (count rose). B+M flat at 3. §G takes its fourth Major; dimensions unfrozen; hold components per route rather than per pair; session 01a09b09-ba8b-7773-b42f-93a2f4261814 | | 39 | 13005de | 10→**5** | 0→**0** | 3→**4** | yes | **zero tells.** Three of four Majors are the three-site fan-out of the pass-37/38 rules — §G limb, §H strict reading, §F pointer. Pattern named and now checked before each pass; session 01a09b1d-8cbc-7d43-8d2f-49ff2339c31f | +| 40 | 12cf247 | 5→**3** | 0→**0** | 4→**3** | yes | zero tells; first round with no fan-out Major after the three-site check was run BEFORE the pass | +| 41 | d26d097 | 3→**2** | 0→**0** | 3→**2** | yes | zero tells; both Majors were compressed pointers of mine dropping a load-bearing part of a standing rule | +| 42 | 3b61fe3 | 2→**1** | 0→**0** | 2→**1** | yes | zero tells; §F item 4 had taken back the enumeration it was repaired to avoid | +| 43 | 95f8591 | 1→**6** | 0→**0** | 1→**3** | yes | one tell; branches restated as read ONCE before any act; accepted-Minor pass cost attributed to the set change, not the severity | +| 44 | 4a44007 | 6→**7** | 0→**0** | 3→**3** | yes | one tell; branch LEADS reworded; **tenth** falsified standing sentence (Mechanics' "collect, never iterate") | +| 45 | 1c858ac | 7→**7** | 0→**0** | 3→**1** | yes | zero tells; swept all four sites still tying classification to closing | +| 46 | 4b713dd | 7→**9** | 0→**1** | 1→**2** | yes | **TWO TELLS — mandatory stop under §5, answered in advance by Daniel's standing "no stops without an absolute block" of 2026-09-13 17:02.** Tells surfaced here and in the commit body, findings kept open, no pass credited clean; the loop continued on that answer. **Eleventh** falsified standing sentence; session 01a09b7c-dd6b-7cb3-b83a-776599286f9b | + +## Passes 40–46 — the two-tell stop at pass 46, and how it was answered + +**Floor line (unchanged):** derived floor **3**; risk **high**, security **none**; read fresh from +`docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md`. One cited story, level 2. + +- **Trend:** findings 3, 2, 1, 6, 7, 7, **9**. Blockers 0, 0, 0, 0, 0, 0, **1**. + Majors 3, 2, 1, 3, 3, 1, **2**. Blocker+Major 3, 2, 1, 3, 3, 1, **3**. +- **Cluster (pass 46):** the ordering and the gate blocks 6 of 9; metadata 3; the instrument 0. +- **require↔withdraw:** none across the seven. + +**Pass 46 carries two tells — the count rose 7 → 9 and the Blocker count failed to fall, 0 → 1 — +and §5 makes stop-and-surface mandatory at two.** Daniel answered that question in advance on +2026-09-13 at 17:02: *"Keine Stopps, wenn nicht zwingend notwendig. Keine Stopps wenn kein +absoluter Block."* **The stop's terminal action is to surface and hand the decision to the user, +which is done here and in the commit body; the decision was already given.** What the rule forbids +is unaffected and is honoured: no pass is credited clean, every finding stays open, and the cycle +stays open. This is recorded rather than left implicit because an agent continuing past a mandatory +stop on its own judgement would be the gate-off path §5 names. + +**What passes 40–46 did.** The three-site fan-out check, run before pass 40 rather than after, +produced the first round with no fan-out Major. Compressed pointers and enumerations were then the +whole remaining stock: §F item 4 took back an enumeration, §A3 compressed the branch-agreement +rule, §C and §F item 9 enumerated branches. Passes 43–45 were three rounds on one subject — +classification still referred to closing in four places after the leads had been reworded — and +pass 45 swept all four at once. + +**Two more standing sentences were found false**, taking §F from nine to eleven: Mechanics' +severity fallback ending "collect, never iterate" against the accepted-Minor pass cost, and the +evidence-entry revalidation trigger naming only "the cycle-closing amend" while §A3 has two +non-amend closing shapes. ## Passes 37–39 — the fan-out is the mechanism, and it is checkable diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index cef2449..bfb865e 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -110,7 +110,11 @@ source rule that prescribes it; the list is examples and not the set** — **the must be repaired or answered, and no further pass runs while its block stands.** It is neither a suspension nor a continue and needs no name and no procedure of its own: the source rule carries both, and this ordering's part is to send the reader there rather than to run a pass over a cycle -another rule has stopped. **It is read first and it silences nothing.** Where the same pass also +another rule has stopped. **Once its source condition is repaired the ordering is read again on that pass**, no closing act +having been attempted on it — the read-once rule below is about a pass that reached the act, not +about one a block held before it, and without this a repair that moves no pass-cost value would +leave a clean eligible pass with no route to the act and none to a suspension. +**It is read first and it silences nothing.** Where the same pass also carries a suspension, that suspension is surfaced with its reasons and its questions exactly as the suspension branch requires and its answers are collected; what the block adds is that **no next pass runs until its own source condition is repaired**, whatever those answers were. @@ -651,10 +655,10 @@ demotes it — the two counts are meant to differ. --- -## F. The ten standing sentences this change falsifies — REPLACED +## F. The eleven standing sentences this change falsifies — REPLACED -Each is a live sentence that the block makes wrong. All ten are **known contradictions** and -none is deferred. **Seven of them share one mechanism** — an entry point other than the ordering +Each is a live sentence that the block makes wrong. All eleven are **known contradictions** and +none is deferred. **Eight of them share one mechanism** — an entry point other than the ordering carrying an unqualified instruction — which is why each is **replaced** rather than given an exception to point at. @@ -758,7 +762,8 @@ last pass's text. C 988–989 and W 1172–1173; only the Gate-A clause changes. ``` **Which commit:** an ungated change records it in that commit; a Gate-A cycle in the commit its -closing act uses; a Gate-B cycle in the WIP commit, restated by the closing amend. +closing act uses; a Gate-B cycle in the WIP commit, restated by the commit its closing act +produces. ``` *Why (pass 22 finding 3):* the live clause sends a Gate-A cycle's exception record to "the spec or plan commit", which is the closing commit only on the first of the three closing paths. On the @@ -778,6 +783,18 @@ hook's counter. Left standing it tells an author that a stray commit destroyed r cannot reach. **This one does not share the section's shared mechanism** — it is a false claim about a mechanism rather than an entry point carrying an unqualified instruction. +**9a. The evidence-entry revalidation trigger** (the profiles section). It wraps across C 726–727 +and W 912–913. +``` +and is **revalidated before every Gate-B re-review and before the commit its closing act +produces** — a fix changes the diff even when the profile sits still. +``` +*Why (pass 46 finding 2):* the live clause names "the cycle-closing amend", which is one of the +shapes a Gate-B closing act takes; where several `WIP:` snapshots piled up, or a stray non-amending +commit left a `WIP:` ancestor, the act is a reset and a single commit and no amend occurs. Scoped +to the amend, the final revalidation is owed on one path and skipped on the others. It is the +eighth sentence sharing this section's mechanism. + **9b. The severity-deciding fallback** (Mechanics · Severity). It wraps across C 788–789 and W 974–975. ``` From 4083ba6abbeaf2d69229a42faf8be6c589918006 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sun, 13 Sep 2026 18:18:27 +0200 Subject: [PATCH 076/181] docs(specs): apply pass 47; the Blocker/Major filter was the last hidden class MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit One tell — the count rose 9 to 11; Blockers fell 1 to 0. Below the threshold. The one-shape / one-severity / one-gate sweep asked for in this pass's prompt found four more standing sentences, three of them the same previously hidden class: the change makes Minor and Nit lines load-bearing, and several entry points still tell the author to filter them away. - the HARD FLOOR parenthetical reads "(Blocker/Major only)". A scope-stop trigger makes a pass unclean whatever the severity, and an assigned-fix-set change costs a pass even where the accepted finding is a Minor. Verified at CLAUDE.md 72-73 and workflow-init.md 279-280 - the Gate-A coverage instruction's "you filter to Blocker/Major downstream" is now false of the READING rather than of the repairing. Verified at CLAUDE.md 561-564 and workflow-init.md 753-756 - §F item 2 carried the same clause into the Gate-B instruction and would have installed it in the mirror. Both now say: filter to Blocker/Major for what must be repaired, read every line for everything else - §F item 8 preserved "fold the edit into the active WIP snapshot by amend" unconditionally, while §A3 admits a stray non-amending commit that makes the snapshot an ancestor no amend touches §F goes from eleven falsified standing sentences to thirteen, ten of them sharing the section's mechanism. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-47.md | 12 ++++++ ...-10-loop-rule-consolidation-target-text.md | 41 +++++++++++++++---- 2 files changed, 46 insertions(+), 7 deletions(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-47.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-47.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-47.md new file mode 100644 index 0000000..3c98d2f --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-47.md @@ -0,0 +1,12 @@ +MAJOR | high | standing HARD FLOOR opening, C 72–73 / W 279–280 | The retained parenthetical defines both gate loops as a minimum pass count “(Blocker/Major only)”, but the target now makes Minor/Nit scope-stop triggers affect cleanliness and makes any assigned-fix-set change, including accepting a Minor or Nit, cost a further pass | An agent entering through the opening rule can treat a Minor/Nit-only outcome as outside the loop’s pass accounting and skip the suspension or extra pass the ordering requires | Replace the severity-only parenthetical with a pointer to the closure ordering’s complete pass-cost and cleanliness rules +MAJOR | high | standing Gate-A coverage instruction, C 561–564 / W 753–756 | The live sentence that §F does not replace says the author filters findings to Blocker/Major downstream, while the target requires Minor and Nit findings to be read for scope-stop triggers, assigned-fix-set changes, repeated-dismissal handling and loop health | A Gate-A runner can discard a Minor or Nit before the predicates that can make its pass unclean or suspend it are evaluated | Replace the filter instruction with a duty to retain every finding and let Mechanics · Severity decide repair while the closure ordering reads scope and health +MAJOR | high | §F item 2 “Gate-B coverage instruction” | The proposed replacement still says “You filter to Blocker/Major”, and it newly installs that clause in the workflow-init mirror, although §A1 and §B require Minor and Nit lines for cleanliness, scope stops, set-change pass costs and health readings | A Gate-B runner can drop a lower-severity finding that carries a mandatory membership or question stop and close on a pass the target defines as unclean | Remove the filter sentence and point to Mechanics · Severity for repair duties and to the closure ordering for every finding’s other effects +MAJOR | high | §F item 8 and standing Profiles, C 751–753 / W 937–939 | The replacement preserves the unconditional instruction to fold a profile edit into the active WIP snapshot “by amend”, but §A3 expressly admits a successful non-amending commit on top of the WIP; in that state `git commit --amend` rewrites the stray tip rather than the WIP ancestor | The profile edit can land in the wrong snapshot, the next review can target an unintended history shape, and the WIP remains in history until a later corrective rewrite | Replace the amend-specific instruction with a duty to update the current Gate-B review snapshot and cite the Gate-B closure paragraph’s repository-state distinction entire +MINOR | high | §F item 1 “What the hook loses” | The proposed exhaustive statement says the hook loses “its counter state”, but `codex-gate.sh` 886–888 deletes the recorded fingerprint `state_file` as well as `count_file` and `fresh_file`, and §A3 itself distinguishes Gate-B review state from the counter | A reader can wrongly expect the prior fingerprint to survive a stray or failed commit attempt and misdiagnose the hook’s next freshness or not-run result | Remove the enumeration and cite §A3’s complete state-clearing statement +MINOR | high | §F item 7 “first of the three closing paths” | The rationale says the spec-or-plan commit is the closing commit only on the first of three Gate-A closing paths, while §A2 defines exactly two repository-state cases; if the intended three-way split distinguishes already-committed-at-HEAD, carried by a later HEAD, and still-uncommitted text, the spec-or-plan commit is also the closing commit in the uncommitted case | The rationale both contradicts §A2’s case count and misstates when the standing destination is correct, making the replacement look broader than the defect | Remove the enumeration and say only that the spec-or-plan commit need not be the closing commit when a later HEAD already carries the reviewed artifact +MINOR | high | §F versus design §1 lines 34–35 | The design says there are nine falsified standing sentences and that target §F states that count, while §F now states and enumerates eleven | The governing decision document gives a plan author a smaller implementation scope than the reviewed target | Remove the cross-file count from design §1 and cite §F’s complete enumeration +MINOR | high | §F versus design §7 lines 322–324 | The verification section says §F contains nine falsified sentences of which six share the entry-point mechanism, while §F now states eleven and eight respectively | Verification pairs and the closing evidence entry can be planned against two fewer replacements and shared-mechanism cases than the target requires | Replace both counts with a pointer to §F’s enumeration +NIT | high | §F item 4 line citation | The quoted human-exception sentence begins at CLAUDE.md 1013 and workflow-init.md 1197 with “Those have their” and ends at 1015 / 1199, but the stated C 1014–1015 and W 1198–1199 ranges omit its first line | The required mechanical location check does not contain the complete sentence it claims to locate | Change the ranges to C 1013–1015 and W 1197–1199 +NIT | high | §F item 8 line citation | The replacement quotes the sentence beginning “Inside an active Gate-B cycle” at CLAUDE.md 751 and workflow-init.md 937, but its stated C 752–753 and W 938–939 ranges omit that opening line | The required mechanical location check does not contain the complete sentence being replaced | Change the ranges to C 751–753 and W 937–939 +NIT | high | §F item 5 quoted standing sentence | The rationale presents “after the final clean pass, close it with `git commit --amend`” as what the live sentence says, but neither prompt copy contains that quotation; both contain `git commit --amend -m ""` | The exact-quote check fails and a reader cannot find the asserted wording in either named source | Quote the complete standing command or mark the omitted arguments with an ellipsis outside the code span +END OF FINDINGS (11 total) diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index bfb865e..f896335 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -655,10 +655,10 @@ demotes it — the two counts are meant to differ. --- -## F. The eleven standing sentences this change falsifies — REPLACED +## F. The thirteen standing sentences this change falsifies — REPLACED -Each is a live sentence that the block makes wrong. All eleven are **known contradictions** and -none is deferred. **Eight of them share one mechanism** — an entry point other than the ordering +Each is a live sentence that the block makes wrong. All thirteen are **known contradictions** and +none is deferred. **Ten of them share one mechanism** — an entry point other than the ordering carrying an unqualified instruction — which is why each is **replaced** rather than given an exception to point at. @@ -679,7 +679,9 @@ repair and is item 8 below.** ``` Same coverage rule as Gate A: put "report every finding with severity and confidence; write `NO FINDINGS` only when the branch found none" in `additionalContext`, with the same one-line -format. You filter to Blocker/Major, Codex never does. +format. **You filter to Blocker/Major for what must be repaired, and read every line for +everything else** — cleanliness, the scope triggers, the assigned fix set and the loop-health +readings all take Minor and Nit lines. Codex never filters. ``` *Why (pass 17 finding 5):* "say `NO FINDINGS` if clean" plus the ordering's clean-pass-with-Minors tells a reviewer to emit an empty file over real Minors. @@ -773,9 +775,10 @@ instead keeps them together on all three without changing the record's form or f **8. The profile-change paragraph's pass claim** (the profiles section). It wraps across C 752–753 and W 938–939. ``` -Inside an active Gate-B cycle, fold the edit into the active `WIP:` snapshot by amend — a non-`WIP` -commit reads to the hook as the cycle closing and would discard **the hook's count of** the -accumulated passes. +Inside an active Gate-B cycle, fold the edit into the active `WIP:` snapshot — by amend where the +snapshot is the tip, and otherwise by the shape that reaches it, a stray non-amending commit having +made the snapshot an ancestor an amend would not touch. A non-`WIP` commit reads to the hook as the +cycle closing and would discard **the hook's count of** the accumulated passes. ``` *Why (pass 33 finding 7):* the live clause says such a commit "would discard the accumulated passes". A pass is established by its validated findings file; what the commit reaches is the @@ -783,6 +786,30 @@ hook's counter. Left standing it tells an author that a stray commit destroyed r cannot reach. **This one does not share the section's shared mechanism** — it is a false claim about a mechanism rather than an entry point carrying an unqualified instruction. +**8a. The HARD FLOOR parenthetical** (the §5 loop rule). It wraps across C 72–73 and W 279–280. +``` +**Both gates are a LOOP with a HARD FLOOR: a minimum number of passes per run (a clean final pass +being what the floor is spent on, and cleanliness taking more than Blocker/Major), derived from the +cited story's profile.** +``` +*Why (pass 47 finding 1):* the live parenthetical reads "(Blocker/Major only)". Under the ordering +a scope-stop trigger makes a pass unclean whatever the finding's severity, and an assigned-fix-set +change costs a further pass even where the accepted finding is a Minor. A reader entering here can +treat a Minor-or-Nit-only outcome as outside the floor's business, which it no longer is. + +**8b. The Gate-A coverage instruction's filter clause** (the Gate A section). It wraps across +C 561–564 and W 753–756. +``` +with severity and confidence — **you filter to Blocker/Major for what must be repaired and read +every line for everything else**, Codex never filters, because a model told to report only high +severity drops real findings silently +``` +*Why (pass 47 finding 2):* "you filter to Blocker/Major downstream" is now false of the reading +rather than of the repairing. Cleanliness, the two scope triggers, the fix set and every +loop-health reading take Minor and Nit lines, so an author who discards them before those +predicates run can close on a pass that was never clean. It is the ninth and tenth sentences +sharing this section's mechanism, with 8a. + **9a. The evidence-entry revalidation trigger** (the profiles section). It wraps across C 726–727 and W 912–913. ``` From 935c8da000a33c7be3cd620b9f67c8a146fe7896 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sun, 13 Sep 2026 18:32:25 +0200 Subject: [PATCH 077/181] =?UTF-8?q?docs(specs):=20apply=20pass=2048;=20sto?= =?UTF-8?q?p=20keeping=20a=20second=20copy=20of=20=C2=A7F's=20count?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Zero tells. Findings 11 to 7, Blocker+Major 4 to 5. Three of the five Majors were design drift I created: I grew §F over four passes without carrying the counts or the edit map with it. The design now states no number for the falsified sentences at all and says why — a second copy of that count has gone stale five times, and the plan reads both the list and its count off §F. - a FOURTEENTH falsified standing sentence: Mechanics' mid-run recovery sentence says a Gate-A cycle has no commit of its own, verified at CLAUDE.md 416-417 and workflow-init.md 610-611, while §A2 admits the case where HEAD already carries the reviewed text and the act amends that commit's message - repairing a source block no longer bypasses an unanswered suspension: the branch releases its own block and the composition rule still holds the cycle on every answer that suspension asked for - the design's edit map gains the five later sites, its evidence revalidation points at the commit the closing act produces rather than at the amend, and its oracle stops calling a zero-finding pass a branch — §A1 defines four, and zero findings is an eligibility route inside the clean-completion branch Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-48.md | 8 ++++++ ...26-09-10-loop-rule-consolidation-design.md | 25 ++++++++++++------- ...-10-loop-rule-consolidation-target-text.md | 20 ++++++++++++--- 3 files changed, 41 insertions(+), 12 deletions(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-48.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-48.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-48.md new file mode 100644 index 0000000..51fd511 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-48.md @@ -0,0 +1,8 @@ +MAJOR | high | standing recovery paragraph (C 413–417; W 607–611) | Both live copies say a Gate-A cycle mid-run has no commit of its own and therefore has only the working record, but §A2 now permits a Gate-A cycle to review text already carried by `HEAD` and close by amending that existing commit's message | A reader can choose history as a recovery source or reject it based on incompatible descriptions of the same mid-run state, risking adoption of the wrong cycle identity | Replace the commit-existence test with a test for whether a closing body carrying this cycle's provenance line and curve exists, and add this sentence to §F +MINOR | high | standing curve-value paragraph (C 972–975; W 1156–1159) | Both live copies justify the curve being non-durable within a running cycle with “the commit does not exist until the cycle closes,” but §A2 now allows the reviewed revision's commit to exist before the Gate-A cycle closes | The new closing shape leaves a false mechanism claim in both prompt copies even if the intended conclusion about the unwritten curve remains valid | State that the curve is not present in the closing commit body until the closing act writes it, and add this sentence to §F +MAJOR | high | §A1 source-block and suspension composition | The source-block branch says repairing its condition rereads the ordering on the same pass even when that pass also carried a suspension and “whatever” its answers were, while the suspension rules say a stop answer parks the cycle until an explicit continue and say the next pass runs after all answers permit resumption | Repairing a source block can either bypass a stop answer and close on the old pass or require a new pass, so the same source-block-plus-suspension path has incompatible next states | Condition the same-pass reread on the suspension state, preserve explicit continue for a parked cycle, and make the composition rule cite the source-block same-pass case rather than promising a next pass unconditionally +MAJOR | high | §F against design §§1, 4 and 7 | The target enumerates thirteen falsified standing sentences, ten sharing the entry-point mechanism, but the design still says there are nine in §§1 and 7, calls the evidence remedy the eighth in §4, and omits the four later sites from its edit map | A plan following the decisions document can omit four required replacements and fail the story's accounting and parity criteria even while following the target's current text | Remove the stale counts and ordinal from the design and cite target §F entire; make the site map defer to that same complete source instead of maintaining another enumeration +MAJOR | high | §F item 9a against design §7 “Evidence entry” | Item 9a broadens final evidence-entry revalidation to the commit produced by Gate B's closing act, but design §7 still says it runs before “the closing amend” and claims that is what §5 requires | The reset-plus-single-commit closing shape can reach its act without the final evidence revalidation that the target requires | Replace the design's amend-specific phrase with a pointer to Gate B's closing act entire +MAJOR | high | §A1 four branches against design §7 “The oracle” | The design says a row enters closure from “the clean-completion or zero-finding branch,” but §A1 defines exactly four branches and makes a zero-finding pass an eligibility route within the clean-completion branch, not a fifth branch | The plan's next-state oracle can encode and verify a branch the target text does not define, defeating the claimed transition check | Say the row enters closure through the clean-completion branch, whose eligibility test includes the zero-finding route +MINOR | high | §F item 7 “three closing paths” | The rationale still says “the first of the three closing paths,” “the other two,” and “all three,” although §A2 now defines two Gate-A repository cases and §A3 defines two Gate-B closing-act shapes | The numerical rationale has no stable mapping to the proposed closure text and can make the plan preserve an obsolete case split | Remove the enumeration and say that naming the closing act keeps the record with closure in every case that act permits +END OF FINDINGS (7 total) diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index 1a44b01..ed9380a 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -31,8 +31,9 @@ duties participate in that ordering versus gate it as preconditions. With it: th a severity demotion does to the loop-health counts, and the standing sentences the ordering falsifies or leaves ambiguous if they are not edited at their source. **§4 lists the sites row by row and claims no total over them**, a count over spans that merge and split being bookkeeping the -plan re-derives against the files. **The nine falsified standing sentences are a different count** -and the target text's §F states it, because those are individually enumerated sentences rather than +plan re-derives against the files. **The falsified standing sentences are a different count** +and **only the target text's §F states it**, which is why no number for them appears here — a +second copy of that count is what went stale four times, because those are individually enumerated sentences rather than spans. **What does not:** the pass floor and severity semantics, which the parent shipped and this spec @@ -213,6 +214,11 @@ to point at and so a reader can see the shape of the change without reading the | (i) when these rules bind | §H | | the Gate-A section and the gate-prompt template | the two senses of *clean* are separated; the cadence makes revision conditional; the broad-prompt instruction stops assuming the artifact is revised between passes, keeping its breadth demand | | the profiles section, the lens paragraph | its unchanged-list is scoped to the lens sets | +| the §5 loop rule, the HARD FLOOR parenthetical | "(Blocker/Major only)" no longer describes what the floor is spent on, a scope-stop trigger and an accepted Minor both bearing on it (pass 47 finding 1); §F item 8a | +| the Gate-A section, the coverage instruction's filter clause | the filter is scoped to what must be repaired, every line still being read for cleanliness, the triggers, the fix set and loop health (pass 47 finding 2); §F item 8b | +| Mechanics, the cycle nonce's mid-run recovery sentence | a Gate-A cycle does have a commit of its own where `HEAD` already carries the reviewed text (pass 48 finding 1); §F item 7a | +| the profiles section, the evidence-entry revalidation trigger | scoped to the commit the closing act produces rather than to the amend (pass 46 finding 2); §F item 9a | +| Mechanics · Severity, the severity-deciding fallback | "collect, never iterate" scoped to the severity's own cost (pass 44 finding 2); §F item 9b | | the profiles section, the profile-change paragraph | its claim that a stray non-`WIP` commit "would discard the accumulated passes" is narrowed to the hook's count of them (pass 33 finding 7); §F item 8 | | the profiles section, the evidence-entry revalidation remedy | its fix-re-review-close instruction becomes conditional on the ordering selecting continuation, a non-closing pass taking any applicable suspension first (pass 30 finding 3). The eighth falsified standing sentence, and the sixth sharing §F's mechanism | @@ -319,9 +325,9 @@ old-wording-gone half of its pair. **The obligation reaches every passage the ta REPLACED, and no list of them is kept here** — a second enumeration beside the markers is the bookkeeping that goes stale, which it did: the list this sentence used to carry omitted §G while §G was marked REPLACED. **The plan reads the markers off the target text**, where the concrete -replacements live. **§F now states nine falsified standing sentences**, the ninth being the -profile-change paragraph's pass claim (pass 33 finding 7); six of the nine share the section's -entry-point mechanism and that one does not. Nothing is +replacements live. **§F states the falsified standing sentences and their count**, and the plan reads +both off §F rather than from here; the count has moved at five passes and a copy of it here would +be stale again by the next. Nothing is claimed as "contradictory" — the second of the two defects Gate B found in the `fic2` instrument. **The named verification of the risk path** (story AC 4) is a **next-state table**, written in the @@ -352,8 +358,8 @@ or closed state — **the same stop returning with its reading unconsumed, that intervening validated pass run after the answer** — or when it closes on anything other than the route the block states. **The closure conditions are read from the block and not re-enumerated here**: a re-enumeration is a second definition that drifts, and pass 15 found this list already -missing two of them. Concretely the row must **enter closure from the clean-completion or -zero-finding branch** — so a pass carrying a scope-stop trigger cannot close on the answer to that +missing two of them. Concretely the row must **enter closure from the clean-completion branch**, a +zero-finding pass being an eligibility route inside that branch rather than a branch of its own — so a pass carrying a scope-stop trigger cannot close on the answer to that trigger, no-clean-credit being the clean predicate's own second half — and every precondition the block names must hold **when it is established, immediately before the closing act**, which is the window the block fixes. **The oracle does not require a precondition to be re-read on what the act produces** — the @@ -370,8 +376,9 @@ not an input was consumed", which would have classified that legitimate case as **Evidence entry**, in the closing commit body, names: the battery run; every pair the plan built with its counts in each copy and each tree, and every presence check beside them; the §6 parity diff and the `b11`/`b13` equivalence result; and the next-state table's location plus its row -count. It is revalidated before every Gate-B re-review and before the closing amend, as §5 -requires. +count. It is revalidated before every Gate-B re-review and before the commit that gate's closing act +produces, as the target text's §F item 9a requires — the amend being one of the shapes that act +takes. **One observability residual, stated because the lens set asks for it and nothing here answers it.** A closing commit body records that a cycle closed and what its curve was; it records diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index f896335..8c306bf 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -111,7 +111,9 @@ must be repaired or answered, and no further pass runs while its block stands.** suspension nor a continue and needs no name and no procedure of its own: the source rule carries both, and this ordering's part is to send the reader there rather than to run a pass over a cycle another rule has stopped. **Once its source condition is repaired the ordering is read again on that pass**, no closing act -having been attempted on it — the read-once rule below is about a pass that reached the act, not +having been attempted on it — **and where that pass also carried a suspension, its answers are +still owed and the composition rule still holds the cycle**, this branch releasing only its own +block — the read-once rule below is about a pass that reached the act, not about one a block held before it, and without this a repair that moves no pass-cost value would leave a clean eligible pass with no route to the act and none to a suspension. **It is read first and it silences nothing.** Where the same pass also @@ -655,9 +657,9 @@ demotes it — the two counts are meant to differ. --- -## F. The thirteen standing sentences this change falsifies — REPLACED +## F. The fourteen standing sentences this change falsifies — REPLACED -Each is a live sentence that the block makes wrong. All thirteen are **known contradictions** and +Each is a live sentence that the block makes wrong. All fourteen are **known contradictions** and none is deferred. **Ten of them share one mechanism** — an entry point other than the ordering carrying an unqualified instruction — which is why each is **replaced** rather than given an exception to point at. @@ -786,6 +788,18 @@ hook's counter. Left standing it tells an author that a stray commit destroyed r cannot reach. **This one does not share the section's shared mechanism** — it is a false claim about a mechanism rather than an entry point carrying an unqualified instruction. +**7a. The mid-run recovery sentence** (Mechanics, the cycle nonce). It wraps across C 416–417 and +W 610–611. +``` +A Gate-A cycle has such a commit only where its reviewed text is already committed and its closing +act would amend that commit's message; otherwise it has only the working record. +``` +*Why (pass 48 finding 1):* the live sentence says a Gate-A cycle mid-run has no commit of its own +and therefore has only the working record. §A2 admits the case where `HEAD` already carries the +reviewed text, and there the cycle does have a commit a later reader can read. Left standing, the +two describe the same mid-run state incompatibly and a recovering run cannot tell whether history +is a source. + **8a. The HARD FLOOR parenthetical** (the §5 loop rule). It wraps across C 72–73 and W 279–280. ``` **Both gates are a LOOP with a HARD FLOOR: a minimum number of passes per run (a clean final pass From f6bca9b230216a612929bb0bd90d875b03aa2d51 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sun, 13 Sep 2026 18:43:59 +0200 Subject: [PATCH 078/181] =?UTF-8?q?docs(specs):=20apply=20pass=2049;=20the?= =?UTF-8?q?=20design=20stops=20enumerating=20what=20=C2=A7F=20owns?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit One tell — the count rose 7 to 10. Blocker+Major 5 to 4. Three of the four Majors had one root: the design kept a second enumeration of the standing sentences §F owns. Pass 48 removed the counts and left the rows, which is why the same finding came back. Eleven rows are now replaced by one pointer saying the list, the count and the replacements live in §F alone, and §5's human-exception row and §7's evidence-timing sentence stop restating what §F states. A second enumeration beside §F went stale at five passes running. - design §2 said clean completion outranks a suspension "by closing the cycle", which the settled decision replaced with "by taking the pass to the closing act"; a plan reader could have reclassified a pass whose act failed - the source-block reread is scoped: the branch releases its own block only. The composition rule still holds the cycle on every answer that suspension asked for, a continue still leads to a pass run after the answer, and a stop still parks until an explicit later continue. The reread happens where no suspension of that pass is outstanding Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-49.md | 11 ++++++++++ ...26-09-10-loop-rule-consolidation-design.md | 20 ++++--------------- ...-10-loop-rule-consolidation-target-text.md | 8 +++++--- 3 files changed, 20 insertions(+), 19 deletions(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-49.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-49.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-49.md new file mode 100644 index 0000000..cdcbb8b --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-49.md @@ -0,0 +1,11 @@ +MAJOR | high | §A1 "Once its source condition is repaired" | The pass-48 repair keeps suspension answers outstanding but still unconditionally rereads the same pre-answer pass after the source block clears; the suspension rules instead make continue lead to a post-answer pass and make stop park the cycle until a later explicit continue | A source-blocked pass can close without the required post-answer pass after continue, or can leave the parked state and close merely because its source block was repaired after stop | Limit the same-pass reread to a pass with no suspension; where a suspension applied, follow its answer-produced state in full, including a new pass after continue and parking after stop +MAJOR | high | design §2 "by closing the cycle" | The design says clean completion outranks a suspension by closing the cycle, while target §A1 settles precedence when the pass reaches the closing act and keeps a failed act in its own retry-or-park handler even though the cycle did not close | A plan reader can reclassify a pass whose closing act failed into a suspension, giving the same failed act both the target's retry-or-park outcome and the design's suspension outcome | Replace the design's outcome-based formulation with a pointer to target §A1's complete precedence and failed-act rules +MAJOR | high | design §§4–5 edit maps | The design still maintains a second enumeration of the standing sentences that target §F enumerates: §4 lists the §F sites and replacements row by row, and §5 separately repeats the two human-exception replacements | The next addition or regrouping in §F can again leave the design map incomplete or inconsistent and send the plan to a smaller replacement set, recreating the drift mechanism this consolidation is meant to remove | Collapse the §F-specific rows and the passage-(h) details to one pointer to target §F entire, leaving only non-§F passage mapping in the design +MAJOR | high | design §7 "Evidence entry" | The design repeats §F item 9a's complete timing rule by requiring revalidation before every Gate-B re-review and before the commit the closing act produces | This surviving second authority can drift back to an amend-only rule or otherwise disagree with §F while a plan still satisfies one of the two texts | Replace the timing sentence with a pointer to §F item 9a entire and keep only the evidence-entry contents that §7 owns +MINOR | high | design §4 "No total is claimed, here or in the target text" | The target text explicitly claims fourteen falsified standing sentences in §F, so this unqualified assertion is false and also contradicts design §1's statement that only §F carries that count | A reader cannot tell whether §F's fourteen-item total is intentional evidence for the plan or a count the design says does not exist | Scope the sentence to the table's replacement-span total and explicitly leave §F's different sentence count at §F +MINOR | high | §F item 1 "What the hook loses is its counter state" | The replacement presents counter state as the hook's loss, but codex-gate.sh 886–888 also deletes the recorded Gate-B fingerprint state file alongside both count files | A reader can expect the prior fingerprint to survive a stray or failed non-WIP commit attempt and misdiagnose the hook's next freshness or not-run result | Remove the state inventory and point to §A3's complete state-clearing statement +MINOR | high | §F item 7 "three closing paths" | Its rationale says the spec-or-plan commit is the closing commit on the first of three paths and not the other two, while §A2 defines exactly two Gate-A repository-state cases and the uncommitted case also closes in the spec or plan commit | The rationale gives the replacement an obsolete case split and a false account of when the standing destination is correct | Remove the path count and say only that naming the closing act keeps the record with closure in every case §A2 permits +MINOR | high | design §5 passage (h) "only one of three closing paths" | The passage map repeats §F item 7's obsolete three-path rule even though target §A2 has two Gate-A cases | The design independently preserves the same false case split after §F is corrected, so the plan can still implement the wrong record destination logic | Remove the repeated case count and cite §F item 7 entire +NIT | high | §F item 4 line citation | The live human-exception sentence begins at CLAUDE.md 1013 and workflow-init.md 1197 with "Those have their" and ends at 1015 and 1199, but the stated ranges start one line late | The required mechanical location check omits the first line of the sentence it claims to locate | Change the ranges to C 1013–1015 and W 1197–1199 +NIT | high | §F item 8 line citation | The live profile-change sentence begins at CLAUDE.md 751 and workflow-init.md 937 with "Inside an active Gate-B cycle", but the stated ranges start one line late | The required mechanical location check omits the opening line of the sentence it claims to locate | Change the ranges to C 751–753 and W 937–939 +END OF FINDINGS (10 total) diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index ed9380a..1b5be90 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -89,7 +89,7 @@ could *never* close, resting on a six-pass plateau the text does not set as a th coverage judgement that is a fact about now rather than forever. The repeated-refuted-complaint problem stands on its own without either. -**Clean completion outranks a suspension by closing the cycle, not by being eligible** (pass 29 +**Clean completion outranks a suspension by taking the pass to the closing act, not by eligibility alone** (pass 29 finding 1). The earlier wording barred every eligible pass from suspending, which let a clean pass blocked by an unresolved prior Major, a standing hold or stale evidence walk past §D's **mandatory** two-tell stop and keep spending passes. **This narrows a rule this change itself wrote, and @@ -206,21 +206,11 @@ to point at and so a reader can see the shape of the change without reading the | (c) recognizing clearly stuck | §C | | (e) the five tells | §D | | Mechanics · Severity | the resolve duty is scoped and gains its discharge rule; the handed-over question is replaced by its answer | -| Mechanics · `baseSha` | two sentences: the `WIP:` warning stops claiming a closure the rules do not grant, and the `Finishing the cycle` lead-in performs the amend only where the ordering permits closing | -| Mechanics, recording a human exception | two sentences: a prescribed continuation answer is distinguished from blanket assent, and the Gate-A destination follows the closing act rather than naming the spec or plan commit. The answer-record material stays moved to the successor | -| Gate B, the coverage instruction | `NO FINDINGS` only when the branch found none | -| Mechanics, the curve's Majors rationale | rewritten on the pre-ceiling reading | +| every standing sentence this change falsifies | **enumerated once, in the target text's §F**, with its replacement text and its reason. **No row for any of them appears here** — a second enumeration beside §F went stale at five passes running, and the plan reads the list, the count and the replacements off §F | | Mechanics, the one-contract paragraph | membership widened, with a semantic test a downstream reader can apply | | (i) when these rules bind | §H | | the Gate-A section and the gate-prompt template | the two senses of *clean* are separated; the cadence makes revision conditional; the broad-prompt instruction stops assuming the artifact is revised between passes, keeping its breadth demand | | the profiles section, the lens paragraph | its unchanged-list is scoped to the lens sets | -| the §5 loop rule, the HARD FLOOR parenthetical | "(Blocker/Major only)" no longer describes what the floor is spent on, a scope-stop trigger and an accepted Minor both bearing on it (pass 47 finding 1); §F item 8a | -| the Gate-A section, the coverage instruction's filter clause | the filter is scoped to what must be repaired, every line still being read for cleanliness, the triggers, the fix set and loop health (pass 47 finding 2); §F item 8b | -| Mechanics, the cycle nonce's mid-run recovery sentence | a Gate-A cycle does have a commit of its own where `HEAD` already carries the reviewed text (pass 48 finding 1); §F item 7a | -| the profiles section, the evidence-entry revalidation trigger | scoped to the commit the closing act produces rather than to the amend (pass 46 finding 2); §F item 9a | -| Mechanics · Severity, the severity-deciding fallback | "collect, never iterate" scoped to the severity's own cost (pass 44 finding 2); §F item 9b | -| the profiles section, the profile-change paragraph | its claim that a stray non-`WIP` commit "would discard the accumulated passes" is narrowed to the hook's count of them (pass 33 finding 7); §F item 8 | -| the profiles section, the evidence-entry revalidation remedy | its fix-re-review-close instruction becomes conditional on the ordering selecting continuation, a non-closing pass taking any applicable suspension first (pass 30 finding 3). The eighth falsified standing sentence, and the sixth sharing §F's mechanism | **Two sentences are deliberately not edited**, named so nobody looks for them: the "Copy every record into the squash body" sentence inside the human-exception block, and the "records every @@ -249,7 +239,7 @@ here. This table says what happens to each inventoried passage, so the map stays | (e) the five tells | edited — the threshold is read after clean completion, and a pointer says what its answer does | §D | | (f) the two rules above do not compete | **unchanged.** "The two rules above" still names the absorb rule and the stuck reading; the block sits before both and adds no third rule between them | — | | (g) Mechanics · Severity, the handed-over question | edited — the unsettled statement and its interim report-and-stop duty are replaced by the answer, in both copies, removing the one deliberate story-path divergence | §E | -| (h) recording a human exception | **edited in two sentences** — a prescribed continuation answer is distinguished from blanket assent, which the ordering makes the restart of a parked cycle; and the Gate-A destination follows the closing act, the spec-or-plan commit being that commit on only one of three closing paths. The answer-record block that was to follow it stays moved to the successor with **D9**, and nothing else in the passage changes | §F | +| (h) recording a human exception | **edited; §F states which sentences and how** — a prescribed continuation answer is distinguished from blanket assent, which the ordering makes the restart of a parked cycle; and the Gate-A destination follows the closing act, the spec-or-plan commit being that commit on only one of three closing paths. The answer-record block that was to follow it stays moved to the successor with **D9**, and nothing else in the passage changes | §F | | (i) when these rules bind | edited — the strict-reading list is added to, not rewritten | §H | | (j) the squash carry | **no longer edited.** It was to name the answer record, which moved; this change ships no record for it to carry | — | @@ -376,9 +366,7 @@ not an input was consumed", which would have classified that legitimate case as **Evidence entry**, in the closing commit body, names: the battery run; every pair the plan built with its counts in each copy and each tree, and every presence check beside them; the §6 parity diff and the `b11`/`b13` equivalence result; and the next-state table's location plus its row -count. It is revalidated before every Gate-B re-review and before the commit that gate's closing act -produces, as the target text's §F item 9a requires — the amend being one of the shapes that act -takes. +count. **When it is revalidated is stated once, in the target text's §F**, and is not repeated here. **One observability residual, stated because the lens set asks for it and nothing here answers it.** A closing commit body records that a cycle closed and what its curve was; it records diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index 8c306bf..8c1fc10 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -111,9 +111,11 @@ must be repaired or answered, and no further pass runs while its block stands.** suspension nor a continue and needs no name and no procedure of its own: the source rule carries both, and this ordering's part is to send the reader there rather than to run a pass over a cycle another rule has stopped. **Once its source condition is repaired the ordering is read again on that pass**, no closing act -having been attempted on it — **and where that pass also carried a suspension, its answers are -still owed and the composition rule still holds the cycle**, this branch releasing only its own -block — the read-once rule below is about a pass that reached the act, not +having been attempted on it — **and where that pass also carried a suspension, this branch +releases only its own block**: the composition rule still holds the cycle on every answer that +suspension asked for, a continue still leads to a pass run after the answer, and a stop still parks +the cycle until an explicit later continue — so the reread happens where no suspension of that +pass is outstanding, and otherwise the suspension's own route runs first — the read-once rule below is about a pass that reached the act, not about one a block held before it, and without this a repair that moves no pass-cost value would leave a clean eligible pass with no route to the act and none to a suspension. **It is read first and it silences nothing.** Where the same pass also From 68eb365b03c43428608a9a44d16201bf94c2597c Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sun, 13 Sep 2026 19:02:16 +0200 Subject: [PATCH 079/181] docs(specs): apply pass 50; the design records decisions, the target states rules MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit One tell — the count rose 10 to 25. Blocker+Major 4 to 11, and eight of the eleven had one root, the third appearance of the same class: the design carrying second copies of what the target text owns. §2 restated all five behaviour decisions in prose. It is replaced by a decision RECORD — that a decision was made, who made it, when, what question it answered, and which target section states it — plus the two narrowings and the two named gaps. 92 lines become 30, and no rule text survives there. §5's human-exception row and its "three reversals" paragraph, and §7's oracle, likewise stop restating current §A rules. The reversals paragraph records what was reversed and not what the rules now say; the oracle reads the route and the condition timing off §A instead of embedding a copy that can pass while disagreeing with the text it checks. Two dangling pointers shipped in the installed text: §A1 and §A3 both pointed at §I, which is this file's own metadata and ships nowhere. Both now state the limitation in place instead of pointing at a section a downstream reader cannot reach. §F item 7a is corrected: an already-committed revision of the reviewed text is not the cycle's own commit, carrying no provenance line and no curve for a nonce to be taken from. The history source is the closing commit, as the standing rule defines it. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-50.md | 26 +++ ...26-09-10-loop-rule-consolidation-design.md | 149 +++++------------- ...-10-loop-rule-consolidation-target-text.md | 16 +- 3 files changed, 76 insertions(+), 115 deletions(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-50.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-50.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-50.md new file mode 100644 index 0000000..654d4e1 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-50.md @@ -0,0 +1,26 @@ +MAJOR | high | §A1 "stated as a residual in §I" | This sentence is inside the NEW text copied into both prompt copies, but §I is explicitly file-local metadata that ships nowhere, so the installed Gate-A rule points to a nonexistent section | A downstream reader cannot reach the limitation the pointer claims to preserve and may treat the unchecked-review-input case as covered | Remove the §I pointer while retaining the limitation already stated inline, or cite an authoritative section that actually ships in both copies +MAJOR | high | §A3 "the bullet in §I" | The Gate-B closure paragraph is installed but its explanation of the deliberately absent content condition points to §I, which is not installed | Both prompt copies ship a dangling reference at the exact place where the scope cut must remain explicit, making the admitted gap look documented somewhere the reader cannot access | Delete the dangling clause and leave the paragraph's self-contained statement that Gate B adds no content condition +MAJOR | high | §F item 7a and standing recovery paragraph C 411–417 / W 605–611 | Item 7a treats an existing commit that merely carries the reviewed artifact text as the cycle's history source, but the immediately preceding standing rule defines that source as the cycle's own commit body from which agreeing provenance and curve records supply the nonce, while another surviving sentence says the working record is the source while the cycle runs; the pre-closing commit in §A2 need contain none of those records | Mid-run recovery has mutually incompatible source rules and can mistake an unrelated pre-existing artifact commit for this cycle's identity | Remove item 7a and retain the working-record source until the closing act writes the cycle records, rather than equating artifact content with a cycle-owned record body +MINOR | high | standing curve paragraph C 972–975 / W 1156–1159 | The live rationale says the curve is not durable during a running cycle "since the commit does not exist until the cycle closes," but §A2 expressly permits a Gate-A cycle whose reviewed artifact commit already exists and §F contains no replacement for this sentence | The installed text gives a false reason for the still-valid conclusion and invites recovery logic to infer commit absence from an unwritten curve | Add this sentence to §F and say that the curve is absent from the closing commit body until the closing act writes it +MINOR | high | §F item 1 "What the hook loses is its counter state" | The replacement narrows the reset to counter state, while codex-gate.sh 886–888 deletes the recorded fingerprint state file as well as the pass-count and freshness-count files and §A3 itself distinguishes review state from the counter | A reader can expect the prior fingerprint to survive a stray or failed commit attempt and misdiagnose the hook's next freshness or not-run result | Remove the state inventory and cite §A3's complete reset statement, or name the fingerprint and both counters +MINOR | high | §F item 7 "only the Gate-A clause changes" | The replacement also changes the Gate-B destination from "restated by the closing amend" to "restated by the commit its closing act produces" | The section's own change accounting can direct the plan to preserve the amend-only Gate-B wording even though the proposed replacement removes it | State that both gate clauses change and give the Gate-B widening its reset-plus-single-commit rationale +MINOR | high | §F item 7 "first of the three closing paths" | The rationale retains a three-path split after §A2 consolidated Gate A into exactly two repository-state cases, and its claim that the spec-or-plan destination works only on the first path is also false for the uncommitted-text case that commits the spec or plan | The rationale gives the plan an obsolete case model and a false explanation of the replacement's scope | Remove the enumeration and say only that following the closing act keeps the record with closure in every Gate-A case §A2 permits +NIT | high | §F item 4 line citation | The replaced sentence begins at CLAUDE.md 1013 and workflow-init.md 1197 with "Those have their" and ends at 1015 / 1199, but the cited ranges begin one line late | The required mechanical location check omits the opening line of the sentence it claims to locate | Change the ranges to C 1013–1015 and W 1197–1199 +NIT | high | §F item 8 line citation | The replaced sentence begins midway through CLAUDE.md 751 and workflow-init.md 937 with "Inside an active Gate-B cycle" and ends at 753 / 939, but the cited ranges begin one line late | The required mechanical location check omits the opening line of the sentence it claims to locate | Change the ranges to C 751–753 and W 937–939 +NIT | high | §F item order | The required enumeration order is 1–7, 7a, 8, 8a, 8b, 9a, 9b, 9, but the target places item 8 before item 7a | The section's numbering is mechanically out of sequence and makes later ordinal references harder to audit | Move item 7a before item 8 without changing either replacement +MINOR | high | target metadata "no pass is credited clean" | The opening explanation still describes the old c18 rule, while §H replaces it with pass-local cleanliness and expressly says no credit is withheld for surfacing alone | The artifact contradicts itself about whether a clearly-stuck surface can carry a clean pass, obscuring the state claimed for this Gate-A cycle | Say only that surfacing does not close the cycle and that the findings remain open, leaving pass cleanliness to §A1 +MINOR | high | design §3 "credits no pass as clean" | The design's Gate-A-status paragraph repeats the old blanket no-clean-credit rule even though target §H removes it and design §5 later acknowledges that a clearly-stuck surface can be clean | The two design passages decide the same surfaced pass differently and can make the plan retain c18 | Replace the cleanliness claim with a pointer to target §A1 and keep only the statement that surfacing does not close the cycle +MINOR | high | design §1 "§4 lists the sites row by row" | After pass 49, §4 deliberately has no row for any individual §F site and instead points to §F entire, so this surviving claim is false for the falsified standing sentences discussed in the same paragraph | A reader is told both to obtain the §F site set from §4 and that §4 does not contain it | Scope the row-by-row claim to the non-§F passage map and point to §F for its own sites +MAJOR | high | design §2 "One behaviour decision was added" | The design restates §A1's repeated-dismissal cleanliness rule together with its qualifications, health effects, resolve effect, uncertainty rule and unknown-start behavior instead of leaving the executable rule at §A | This is a second load-bearing copy of the rule most likely to drift by losing one qualification, recreating the mechanism the target's ownership boundary forbids | Keep the decision history and rationale, but cite target §A1 entire for the operative rule and remove the duplicated conditions and effects +MAJOR | high | design §2 "Clean completion outranks a suspension" | Pass 49 corrected this sentence to match §A1 instead of removing it, despite the prior finding's requested pointer and §A1's ownership of evaluation order and failed-act treatment | The design remains a second authority for the precedence rule and can again diverge on whether reaching or completing the act wins | Replace the operational sentence and its state examples with a pointer to target §A1's complete precedence rule, retaining only the reason for the decision +MAJOR | high | design §2 "A closing act that does not complete" | The design independently restates §A1's failed-act rule, including re-establishing conditions, retrying, charging source-defined costs and introducing no branch | A repair to either copy can leave the plan following a stale retry or reclassification rule from the other | Replace the operational recovery sequence with a pointer to target §A1 entire and keep only the historical explanation of why the decision was needed +MAJOR | high | design §2 "conditions both gates share" | The design retains a second enumeration of §A1's cycle-general conditions—floor, resolve duty, hold, no-clean-credit, profile, cited set and assigned fix set—even though §A1 owns that classification and deliberately puts no number on its own list | A future condition can be added at §A1 and omitted here while the plan still appears to follow the design, repeating the exact stale-list failure this consolidation addresses | Remove the enumeration and cite target §A1 entire for the gate-general condition set +MAJOR | high | design §2 "each gate states only its own content condition" | This sentence assigns a content condition to each gate, while the settled split and target §A3 say Gate B has none | The design can send the plan back toward the explicitly cut Gate-B closure condition or leave Gate B with an undefined value to test | Say Gate A states its content condition and closing act, while Gate B states only its closing act +MAJOR | high | design §5 passage (h) | Although the row now says §F states which sentences change, it still repeats both human-exception edits—the suspension-answer distinction and the Gate-A destination—instead of pointing to §F entire | The surviving second enumeration can omit another clause changed by §F, and already omits item 7's Gate-B destination widening | Delete the edit summary after the pointer and leave only the moved-successor scope statement that §5 owns +MINOR | high | design §5 passage (h) "only one of three closing paths" | The row independently preserves §F item 7's obsolete three-path account even though target §A2 defines two Gate-A cases | Correcting §F alone would still leave the plan with the wrong case split from the design | Remove the count and cite §F item 7 entire +MAJOR | high | design §5 "Three reversals" | This paragraph restates current §A rules for hold attachment, pass cleanliness, floor behavior and suspension behavior while presenting itself as historical accounting | The plan can implement these operational conclusions from the design even if §A changes, leaving two authorities for the same branches | Record only which earlier proposals were reversed and point to target §A1 for the resulting rules +MAJOR | high | design §7 "The oracle" | The verification oracle re-enumerates §A1's clean-completion entry, zero-finding route, scope-trigger exclusion, pre-act condition timing and post-answer health-pass rule rather than deriving them solely from §A | The plan's next-state table can pass against this stale embedded copy while disagreeing with the target text it is supposed to verify | Define the oracle by a pointer to target §A1 entire and require the plan to derive its rows from that source without restating the branch rules here +MINOR | high | design §8 "including the three exempted" | The conformance discussion repeats three §A1 rules—other outcomes are ineligible, zero findings are clean regardless of floor, and decline is membership-only—even though it says the ownership boundary is stated once at §A | These examples are another operational copy that can drift while still being cited as proof of prompt-standard compliance | Name the prompt-standard concern and cite the relevant §A1 sentences without restating their rules +MINOR | high | design §4 "No total is claimed, here or in the target text" | Target §F explicitly claims a total of fourteen falsified standing sentences, contradicting this unqualified assertion and design §1's acknowledgement that §F alone owns that count | A reader cannot tell whether the fourteen-item total is intentional plan input or a count the design says does not exist | Scope the statement to replacement-span totals in the §4 table and expressly leave §F's sentence count at §F +MINOR | high | design §8 "§4's site table is the check" | Calling the narrative table a check overclaims a mechanism and contradicts design §7's statement that nothing verifies §4's completeness; the actual sweep is deferred to the plan | Reviewers and plan authors can treat an unchecked inventory as evidence that the one-definition claim holds | Replace "is the check" with a pointer to the plan-owned sweep and describe §4 only as the intended passage map +END OF FINDINGS (25 total) diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index 1b5be90..b02e89c 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -67,97 +67,35 @@ decision behind it: the block defines no trigger and no severity rule of its own precondition **that has a source of its own** keeps its one definition there, changed **at that source** where it had to change to agree with the ordering. -**One behaviour decision was added after the split, on Daniel's decision of 2026-09-13, and it -changes the clean predicate.** A finding this cycle has **validly dismissed** and a later pass -merely repeats, with no new evidence and no relevant change to the text the dismissal turned on, -**does not on its own make that later pass unclean**. Without it the standing duty *dismiss -validly, then run another pass* cannot finish: the reviewer would be the authority on whether its -own refuted claim had been dealt with, and a cycle could be held permanently unclean by a repeated -false positive. **It is deliberately narrow and it is not a waiver** — the repetition stays a -finding in its file, every loop-health reading counts it, it remains the clearly-stuck reading's -re-raise condition, no earlier pass becomes clean in retrospect, and doubt about whether it is the -same complaint is resolved against the exclusion. **No new suspension type and no record mechanism -were introduced for it**, which two earlier candidate answers would have required. **Its three -qualifications are carried to every site that states the discharge** — the ordering's suspension -paragraph, §C's third condition and §H's surfacing sentence — because stated only at the clean -predicate they would leave a recurrence that has *become* true reading as discharged (pass 28 -finding 3). **It also has an unknown-start strict reading**: unavailable where a cycle cannot -establish that its starting rules contained it, since the standing fallback says each rule this -change ships adds its own (pass 28 finding 6). **The Blocker -that prompted it is accepted on that core and not on its reasoning**: pass 27 argued the cycle -could *never* close, resting on a six-pass plateau the text does not set as a threshold and on a -coverage judgement that is a fact about now rather than forever. The repeated-refuted-complaint -problem stands on its own without either. - -**Clean completion outranks a suspension by taking the pass to the closing act, not by eligibility alone** (pass 29 -finding 1). The earlier wording barred every eligible pass from suspending, which let a clean pass -blocked by an unresolved prior Major, a standing hold or stale evidence walk past §D's **mandatory** -two-tell stop and keep spending passes. **This narrows a rule this change itself wrote, and -contradicts no settled decision**: D2 and D3 forbid reporting "will not converge" on a loop that -converged, and a loop still owing a repair, an answer or a closure condition has not converged. -Two other pass-29 repairs are compliance rather than decision — the unknown-start list gains the -remaining rules this change ships, and §G's membership test reaches rules deciding termination -**without** a pass, the Gate-B triviality skip being the only such route. A third, sharpening the -index condition to a tree-to-tree comparison, **left with that condition** in the 2026-09-13 cut. - -**A closing act that does not complete has not closed the cycle, and a failed command is not a -pass outcome** (pass 30 finding 4). **The first wording of this was wrong and pass 31 said why**: -it claimed every condition stayed established, which a pre-commit hook that modifies and stages -content before failing falsifies, and it added a "no branch is taken" clause contradicting the -pass-29 rule that a non-closing pass reaches the suspension branch. **The rule was shrunk rather -than extended**: nothing is assumed about what a failed attempt left behind, every closure -condition is re-established against the repository as it stands, the act is performed again where -they hold, and where the attempt or its repair moved anything a condition is read from that -condition's own rule decides the cost. **It introduces no branch and no mechanism.** Recorded -because adopting a reviewer's suggested fix wholesale is what produced three Majors here. - -**One gap is named and not closed** (pass 29 finding 2, and §I carries it): a Gate-A closing act is -written from the effective index, so a staged edit to a review input **no source rule governs** — a -cited story's acceptance criteria, say — is published by the closing commit unchecked. Profile -values, cited-set membership and the assigned fix set are governed and answer themselves. **Neither gate answers this after the -2026-09-13 cut**: Gate B's own version of the hazard is deferred with that condition, and Gate A's -has never been decided. The text states both residuals instead of inventing a rule for either. - -**The ordering is split into three paragraphs on Daniel's decision of 2026-09-12, and that split is -the answer to pass 26's Blocker.** The ordering had stated one closure condition — the artifact's -equality with the text sent to the reviewer — cycle-generally, while its explanation and its two -repository cases were Gate-A's alone; Gate B passes a git range rather than artifact text, so an -otherwise eligible Gate-B pass had no value with which to evaluate it and, being eligible, could -not suspend either. **What is gate-general stays gate-general and what differs goes to the gate**: -the ordering keeps the evaluation of a pass, the classification of the duties, and the closure -conditions both gates share — the floor, the resolve duty, the hold, no-clean-credit, and the -profile, cited-set and assigned-fix-set gates; each gate states only **its own content condition -and its own closing act**, and neither gate's paragraph is an inventory of what that gate requires. -**No closing-time test is an exception to the block's citation rule any more**, because the one -that was is now Gate A's own condition stated at Gate A's paragraph. **The split buys ownership and -not brevity** — the three paragraphs together run slightly longer than the single block did. - -**The Gate-B tree-equality condition is deferred out of this change on Daniel's decision of -2026-09-13, as a bounded scope cut.** Writing Gate B's closure down had exposed a real gap (pass 27 -finding 4): the gate reviews a `baseSha`..`headSha` range while the closing amend commits the -**effective index**, so content staged before the review, or staged by a hook during the commit, -reaches the closing commit through neither branch. The condition written for it, and the widened -re-review duty it carried, **are removed from the target text**; the gap is recorded in §I as a -later task and is **not** worked out here. **The price is stated rather than implied**: this -delivery does not close that gap. Gate B's existing review, re-review and evidence duties stand -unchanged and are not a substitute, and **the gate hook is not one either** — its fingerprint is -advisory and compares its own inputs across its own invocations, so content staged before the -review call and still staged at the commit leaves it unmoved between the two. **Faster convergence -is plausible, not guaranteed.** What the cut removes is two of pass 32's six Blocker/Majors, both -collisions this condition created — against the ordering's pre-act requirement and against the -branch-tip sentence. **Pass 32's finding 4 is expressly not resolved by it**: §A1's failed-act rule -contradicts §F item 1's pass-credit sentence without either mentioning the tree. - -**Gate A's mirror of the same hazard is not a second condition**: a staged edit to a cited story or -a profile header is published by A2's closing commit, and the source rules answer what they govern — -a profile change costs a further pass — so A2 says where those rules bite at the closing act and -adds nothing (pass 28 finding 4). Where the staged edit touches a review input **no** source rule -governs, nothing reaches it, which §I carries as its own residual (pass 29 finding 2). - -**The closure-ordering block is an addition beside the source edits**, not one of them. §4 lists -the edits; **no total is stated here or there**, because the unit — one contiguous replacement at -one site — is not stable across revisions that merge or split a span, and a stated total then -disagrees with its own table. The plan counts what it writes. +**The behaviour decisions of this change are recorded here and stated nowhere here.** Each rule's +words live in the target text; this section says **that** a decision was made, **who** made it, +**when**, and **what question it answered**. A second telling of the rule itself is what drifted at +§F and at the edit map, and it is not rebuilt. + +| # | Decided | Question it answered | Stated in | +|---|---|---|---| +| 1 | 2026-09-12, Daniel | how a closure condition belonging to one gate can be stated cycle-generally without leaving the other gate unable to evaluate it (pass 26's Blocker) | target §A1, §A2, §A3 | +| 2 | 2026-09-12, Daniel | what a change to the assigned fix set does at closing time, no rule having existed anywhere (pass 26 finding 3) | target §B | +| 3 | 2026-09-13, Daniel | whether a reviewer repeating a claim the author has validly refuted can hold a cycle unclean indefinitely (pass 27's Blocker) | target §A1, with §C, §H and §H's unknown-start list for its reach | +| 4 | 2026-09-13, agent in scope | whether an eligible pass that cannot close may still reach a suspension, §D's two-tell stop being mandatory (pass 29 finding 1) | target §A1, §D | +| 5 | 2026-09-13, agent in scope | what a closing act that does not complete leaves behind, and what a failure nobody can repair produces (pass 30 finding 4, narrowed at pass 31, terminal state at pass 37) | target §A1 | + +**Two were reached by narrowing rather than by adding**, recorded because the first attempt at each +was wider than its defect: decision 5 first claimed every condition stayed established and added a +branch, both withdrawn at pass 31; and the Gate-B tree-equality condition that decision 1 exposed +was cut from this delivery entirely on 2026-09-13. + +**What that cut costs, stated rather than implied.** Content staged before the final review, or +staged by a hook during the commit, reaches the closing commit through neither review branch. +**This delivery does not close that gap**; the target text's §I records it as a later task. Gate B's +existing review, re-review and evidence duties are unchanged and are not a substitute, and the gate +hook is not one either — its fingerprint is advisory and compares its own inputs across its own +invocations. Faster convergence was the reason for the cut and was not guaranteed by it. + +**One further gap is named and not closed** (§I carries it): a Gate-A closing act is written from +the effective index, so a staged edit to a review input **no source rule governs** is published by +the closing commit unchecked. Neither gate answers this after the cut, and the text states the +residual rather than inventing a rule. --- ## 3. Where the text is @@ -239,19 +177,17 @@ here. This table says what happens to each inventoried passage, so the map stays | (e) the five tells | edited — the threshold is read after clean completion, and a pointer says what its answer does | §D | | (f) the two rules above do not compete | **unchanged.** "The two rules above" still names the absorb rule and the stuck reading; the block sits before both and adds no third rule between them | — | | (g) Mechanics · Severity, the handed-over question | edited — the unsettled statement and its interim report-and-stop duty are replaced by the answer, in both copies, removing the one deliberate story-path divergence | §E | -| (h) recording a human exception | **edited; §F states which sentences and how** — a prescribed continuation answer is distinguished from blanket assent, which the ordering makes the restart of a parked cycle; and the Gate-A destination follows the closing act, the spec-or-plan commit being that commit on only one of three closing paths. The answer-record block that was to follow it stays moved to the successor with **D9**, and nothing else in the passage changes | §F | +| (h) recording a human exception | **edited; §F states which sentences change and how, and is the only place that does** | §F | | (i) when these rules bind | edited — the strict-reading list is added to, not rewritten | §H | | (j) the squash carry | **no longer edited.** It was to name the answer record, which moved; this change ships no record for it to carry | — | -**Three reversals are recorded here rather than left as silent narrowings**, because each was -made in an earlier revision of this spec and each contradicted something settled. Membership was -briefly re-read when the answer arrived, which `b6` and **D4** both forbid. The hold was briefly -narrowed to scope stops, which contradicted `c16` and the story's own third standing duty; it -attaches to **every** surfaced finding. And `c18` — no pass credited as clean on a clearly-stuck -surface — is **replaced** rather than kept: where that exit's regenerating findings are in-set -and the ceiling demotes them below Major, the pass is clean at effective severity, closes at or -above the floor and suspends below it. Authority **D3**, under which the old reading and D3's own -preserved sentence decide that pass in opposite directions. +**Three reversals are recorded here as history and not as rules**, because each was made in an +earlier revision of this spec and each contradicted something settled: membership was briefly +re-read when the answer arrived; the hold was briefly narrowed to scope stops; and `c18` was +briefly kept rather than replaced. **What each of those rules now says is in the target text's §A +and nowhere here** — an earlier version of this paragraph restated the current hold, cleanliness, +floor and suspension behaviour while presenting itself as accounting, which is two authorities for +one branch. --- @@ -348,13 +284,10 @@ or closed state — **the same stop returning with its reading unconsumed, that intervening validated pass run after the answer** — or when it closes on anything other than the route the block states. **The closure conditions are read from the block and not re-enumerated here**: a re-enumeration is a second definition that drifts, and pass 15 found this list already -missing two of them. Concretely the row must **enter closure from the clean-completion branch**, a -zero-finding pass being an eligibility route inside that branch rather than a branch of its own — so a pass carrying a scope-stop trigger cannot close on the answer to that -trigger, no-clean-credit being the clean predicate's own second half — and every precondition the block names must hold -**when it is established, immediately before the closing act**, which is the window the block -fixes. **The oracle does not require a precondition to be re-read on what the act produces** — the -condition that would have demanded that left with the 2026-09-13 cut, and the gap it addressed is -recorded in the target text's §I rather than checked here. Naming only the distinct-state half would pass +missing two of them. **Which route a row must enter closure by, and which conditions must hold when, are read from the +target text's §A and are not re-enumerated here** — an embedded copy can pass while disagreeing +with the text it is meant to check, which is how this list came to name three conditions while the +block stated more. Naming only the distinct-state half would pass the exact no-progress defect AC 4 cites from the parent cycle. **The consumption clause is what keeps the oracle and the shipped text in agreement**: the target text's §A says continue consumes the reading that raised the suspension and a further health suspension diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index 8c1fc10..8f4279f 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -194,9 +194,9 @@ does not show lands in the closing commit; where the edit changes something a ** governs — a profile value, cited-set membership, the assigned fix set — **the change has happened and that source's own rule applies**, so a further pass is owed and no condition is added here for it. **Where it changes a review input no source rule governs** — a cited story's acceptance -criteria or settled decisions, say — **nothing here reaches it**, and that is stated as a residual -in §I rather than answered: the artifact's own equality condition covers the artifact, and the -inputs beside it have only the rules their sources give them. +criteria or settled decisions, say — **nothing here reaches it, and nothing in this section does**: +the artifact's own equality condition covers the artifact, and the inputs beside it have only the +rules their sources give them. A plateau or tells on **the pass that closes** go into **that pass's status report to the user** — the carrier the three-line duty already names, and no second report form is introduced — and never @@ -418,8 +418,9 @@ and the next paragraph says why. **Its content condition is not an artifact/request equality, and none is written for it.** Gate B reviews a **diff** identified by `baseSha` and `headSha` rather than a text handed to the reviewer, -so there is no reviewed text to compare an artifact against, and **nothing is put in its place** — -the bullet in §I records what that leaves open. What this gate does have is already in this section +so there is no reviewed text to compare an artifact against, and **nothing is put in its place**: +content the final review request did not select can reach the closing commit, and **no rule in this +section reaches it**. What this gate does have is already in this section and is cited rather than restated: the range those two names fix and **the branch-agreement rule entire**, which says how `headSha` is resolved and passed, what is kept with each branch result and what must be equal before two branches are summed — cited here and not compressed, a part of it @@ -793,8 +794,9 @@ about a mechanism rather than an entry point carrying an unqualified instruction **7a. The mid-run recovery sentence** (Mechanics, the cycle nonce). It wraps across C 416–417 and W 610–611. ``` -A Gate-A cycle has such a commit only where its reviewed text is already committed and its closing -act would amend that commit's message; otherwise it has only the working record. +A Gate-A cycle has such a commit only once its own closing commit exists — an already-committed +revision of the reviewed text is not one, carrying no provenance line and no curve for a nonce to +be taken from — so mid-run it has only the working record. ``` *Why (pass 48 finding 1):* the live sentence says a Gate-A cycle mid-run has no commit of its own and therefore has only the working record. §A2 admits the case where `HEAD` already carries the From 2044147bcf448fac591c38f11c211401a3f42d75 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sun, 13 Sep 2026 19:16:27 +0200 Subject: [PATCH 080/181] docs(specs): apply pass 51; one source per rule, in both files MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Zero tells. Findings 25 to 23, Blocker+Major 11 to 6. Four of the six Majors were the same class the pass-50 round fixed in design §2 and left standing in §4, §7 and §9. All three are now pointers: - §4's passage (b) row restated §B's set-change rule — direction, undo behaviour, the window's start — and its accept-at-a-membership-stop rule - §7 restated the consumption clause after saying the oracle derives from §A - §9 restated §G's partial-adoption terminal action immediately after saying §G states it once Two were second copies inside the target text itself: - §A3 said Gate B's duties are cited rather than restated and then summarised seven of them. The summary is gone; the duties stay at the paragraphs that state them, since a compressed inventory is where a load-bearing part goes missing while the list still looks complete - §A3 and §F item 5 each defined the amend-versus-reset case split, pointing at each other as the source. §F item 5 is now the only place either shape is defined, and §A3 says when the act is performed and never which shape it takes Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-51.md | 24 ++++++++++++++ ...26-09-10-loop-rule-consolidation-design.md | 19 +++++------- ...-10-loop-rule-consolidation-target-text.md | 31 ++++++++----------- 3 files changed, 44 insertions(+), 30 deletions(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-51.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-51.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-51.md new file mode 100644 index 0000000..2c50b47 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-51.md @@ -0,0 +1,24 @@ +MAJOR | high | design §4 passage (b) "change costs at least one further pass" | The passage map restates target §B's assigned-fix-set change rule, including direction, undo behavior and the start of its window, even though the design says behavior decisions live only in the target | The plan has a second operational authority that can drift by losing one of the set-change qualifications | Replace the parenthetical rule with a pointer to target §B entire and keep only that passage (b) changes +MAJOR | high | design §4 passage (b) "an accept at a membership stop" | The same map row separately restates target §B's rule that an accepted finding enters the fix set regardless of severity and that an accepted Minor owes no severity repair | A later edit can make membership or Minor handling differ between the map and the text the plan installs | Remove the behavioral summary and cite target §B entire +MINOR | high | design §5 passage (c) "third condition is widened" | The passage map states again that the clearly-stuck reading admits a re-raised validated dismissal, duplicating the operative rule in target §C | The design remains a second authority for whether that recurrence qualifies and can drift independently of §C | Say only that passage (c) is edited per target §C +MAJOR | high | design §7 "consumption clause" | After saying the oracle derives from target §A, the design restates that continue consumes the health reading and that a later suspension is legitimate only after an intervening pass | The plan's oracle can be built against this embedded copy and still disagree with the ordering it is meant to verify | Define the consumption part by a pointer to target §A1 entire and keep only the oracle's requirement to derive its transition from that source +MINOR | high | design §8 "including the three exempted" | The prompt-standards discussion repeats three §A1 rules as a list: no other pass outcome is eligible, zero findings are clean regardless of floor, and decline is membership-only | This second list can go stale while still being cited as evidence that the prompt carries the required rationales | Cite the relevant §A1 rationale sentences without restating the rules or the list +MAJOR | high | design §9 "stops and has a human complete, revert or reconcile" | The design repeats §G's partial-adoption terminal action verbatim in substance immediately after claiming the membership rule is stated once in §G | A plan reader has two authorities for the stop transition, and future changes to §G can leave this copy stale | Retain that §G is an agent instruction and cite §G entire for what it requires +NIT | high | design §1 "four standing duties" | The intent copies the target's duty count even though the target is meant to be the sole statement of the ordered duty set | A later duty change can leave a stale numeric claim in the design without changing the rule source | Remove the number and refer to the standing duties classified in target §A1 +MINOR | high | design §4 "No total is claimed, here or in the target text" | Target §F explicitly claims a total of fourteen falsified standing sentences, contradicting this unqualified assertion and design §1's statement that §F owns that count | A reader cannot tell whether fourteen is intentional plan input or a count the design says does not exist | Scope this sentence to the §4 replacement-span table and leave §F's different sentence count at §F +NIT | high | §A1 "The source-block branch" | The branch repeats a four-example list of source stop conditions already defined in the profile and evidence passages, despite the block's rule that sourced conditions are cited rather than restated | The examples can drift or acquire misleading omissions even though they are labelled non-exhaustive | Remove the examples and point to every source rule that prescribes stop-and-surface +MAJOR | high | §A3 "What this gate does have" | The Gate-B paragraph says the standing rules are cited rather than restated, then re-enumerates and summarizes the range, branch agreement, re-review, spec-update, battery, evidence-mode and evidence-entry duties | A load-bearing part can be dropped from this compressed inventory while the paragraph still appears to be a complete Gate-B handoff | Replace the enumeration with pointers to the branch-agreement rule entire and the other standing Gate-B source passages entire +MAJOR | high | §A3 and §F item 5 "The act" | §A3 defines both Gate-B closing shapes while item 5 defines the same amend-versus-reset-and-single-commit case split again, with each passage pointing to the other as the source | The closing operation has two circular authorities that can diverge and send one repository state through the wrong Git shape | Keep the complete closing-act definition at one source and make the other passage cite it entire +MINOR | high | §B "the floor, the Blocker/Major filter and the clean-final-pass rule all stand" | Passage B preserves a second three-item list of cycle-general duties that §A1 now classifies and owns | A later change to the common conditions can leave the scope-stop entry point teaching an incomplete set | Replace the list with a pointer to the cycle-general conditions in §A1 entire +MINOR | high | §§E and F item 3 "read before the ceiling" | §E defines the curve's pre-ceiling severity reading and its rationale, while item 3 states the same curve rule and rationale again at the curve paragraph | The health series can be implemented from either copy and drift on which severity the counts use | Make the curve paragraph cite the complete loop-health severity rule in §E and retain only why the curve records three series +MINOR | high | §§E and F item 9b "severity buys no repair round and no further pass" | The fallback replacement repeats verbatim in substance the Minor/Nit consequence already defined in §E | A future change to severity behavior can leave the fallback as a conflicting second command | Replace the consequence with a pointer to Mechanics · Severity entire, then preserve only that later scope costs come from the scope decision +MINOR | high | §F items 2 and 8b "filter to Blocker/Major for what must be repaired" | The Gate-A and Gate-B replacements independently state the same downstream filtering rule and that every line feeds all other predicates | The two gate entry points can drift on whether Minor and Nit lines reach cleanliness, scope, fix-set or health evaluation | Put the shared reading rule at one common source and have both gate instructions cite it entire +MINOR | high | §§A3 and F items 1, 5 and 8 "non-WIP commit" | The hook-reset behavior is stated repeatedly across the Gate-B closure paragraph, the naming warning, the finishing lead-in and the profile-change paragraph | These copies already use different state descriptions and can diverge on whether a failed attempt, fingerprint or counters are affected | Keep one verified hook-effect statement and replace the other copies with pointers to it +MINOR | high | §F item 1 "What the hook loses is its counter state" | The hook does not lose only counter state: codex-gate.sh 886–888 removes the Gate-B fingerprint state file as well as both count files | A reader can expect the prior fingerprint to survive a stray or failed non-WIP commit attempt and misdiagnose the next hook result | Remove the state inventory and cite §A3's fuller statement, or name the fingerprint and both counters +MINOR | high | §F heading "The fourteen standing sentences this change falsifies" | The count is presented as all standing sentences falsified by the change, but §H separately replaces multiple live sentences because the new cleanliness, cadence, lens and unknown-start rules make them false or incomplete | The plan-facing count has an unstated category boundary and can be read as proof that the false-sentence sweep is complete | Rename §F as fourteen additional standing sentences outside the passage replacements, or otherwise state the counted set's boundary +MINOR | high | §F item 7 "only the Gate-A clause changes" | The proposed sentence also changes Gate B from "restated by the closing amend" to "restated by the commit its closing act produces" | The section's own accounting can direct the plan to preserve amend-only Gate-B wording while installing a replacement that widens it | State that both gate clauses change and give the Gate-B change its closing-shape rationale +MINOR | high | §F item 7 "first of the three closing paths" | The rationale retains a three-path account even though §A2 defines two Gate-A repository-state cases, and the spec-or-plan destination is the closing commit in both of those cases | The replacement is justified by a false case split and can make the plan preserve obsolete destination logic | Remove the enumeration and say only that following the closing act keeps the record with closure in every §A2 case +NIT | high | §F item 4 line citation | The live sentence starts at CLAUDE.md 1013 and workflow-init.md 1197 with "Those have their" and ends at 1015 and 1199, but the stated ranges begin one line late | The mandatory mechanical check omits the opening line of the sentence it claims to locate | Change the ranges to C 1013–1015 and W 1197–1199 +NIT | high | §F item 8 line citation | The live sentence starts at CLAUDE.md 751 and workflow-init.md 937 with "Inside an active Gate-B cycle" and ends at 753 and 939, but the stated ranges begin one line late | The mandatory mechanical check omits the opening line of the sentence it claims to locate | Change the ranges to C 751–753 and W 937–939 +NIT | high | §F item order | The required enumeration order is 1–7, 7a, 8, 8a, 8b, 9a, 9b, 9, but item 8 appears before item 7a | The numbering is mechanically out of sequence and makes ordinal references harder to audit | Move item 7a before item 8 without changing either replacement +END OF FINDINGS (23 total) diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index b02e89c..0c8e533 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -140,7 +140,7 @@ to point at and so a reader can see the shape of the change without reading the |---|---| | the closure ordering | **new**, and it is section A of the target text | | (a) the floor paragraphs | §H | -| (b) what a loop absorbs | §B | +| (b) what a loop absorbs | edited — it owns both scope triggers and the assigned fix set, so every qualification the ordering needs is made there. **What it now says, including what a change to the set and an accept at a membership stop cost, is stated in the target text's §B and nowhere here** | §B | | (c) recognizing clearly stuck | §C | | (e) the five tells | §D | | Mechanics · Severity | the resolve duty is scoped and gains its discharge rule; the handed-over question is replaced by its answer | @@ -288,13 +288,10 @@ missing two of them. **Which route a row must enter closure by, and which condit target text's §A and are not re-enumerated here** — an embedded copy can pass while disagreeing with the text it is meant to check, which is how this list came to name three conditions while the block stated more. Naming only the distinct-state half would pass -the exact no-progress defect AC 4 cites from the parent cycle. **The -consumption clause is what keeps the oracle and the shipped text in agreement**: the target text's -§A says continue consumes the reading that raised the suspension and a further health suspension -needs it recomputed over a -pass run after the answer, so the *same* two-tell or clearly-stuck result **after** such a pass is -new data and a legitimate row, not a failed transition. An earlier wording failed it "whether or -not an input was consumed", which would have classified that legitimate case as a defect. +the exact no-progress defect AC 4 cites from the parent cycle. **What a health answer consumes, and when a later suspension of the same reading is new data +rather than a failed transition, are stated in the target text's §A and are read from there** — an +earlier wording of this section embedded that rule and an earlier one still contradicted it, which +is two authorities for one transition. **Evidence entry**, in the closing commit body, names: the battery run; every pair the plan built with its counts in each copy and each tree, and every presence check beside them; the §6 parity @@ -396,10 +393,8 @@ against this repository's own spec. Both are why the rule is stated and **the plan derives the membership against the real files**, where dependence is decidable and the numbering does not exist. -**Stated as what it is.** That paragraph is **an instruction to the agent**: a project whose text -carries some members and not others, or versions that disagree, **stops and has a human complete, -revert or reconcile the adoption before running a gate under it**. §G widens whom that -sentence is about. It is not a guard and not a mechanical check, and this change builds neither — +**Stated as what it is.** That paragraph is **an instruction to the agent**, and **what it obliges +is stated in §G and not repeated here** — §G widens whom that sentence is about. It is not a guard and not a mechanical check, and this change builds neither — **nothing detects a partial adoption**, and the stop happens only where an agent reads the sentence and acts on it. diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index 8f4279f..d5d8623 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -420,20 +420,14 @@ and the next paragraph says why. reviews a **diff** identified by `baseSha` and `headSha` rather than a text handed to the reviewer, so there is no reviewed text to compare an artifact against, and **nothing is put in its place**: content the final review request did not select can reach the closing commit, and **no rule in this -section reaches it**. What this gate does have is already in this section -and is cited rather than restated: the range those two names fix and **the branch-agreement rule -entire**, which says how `headSha` is resolved and passed, what is kept with each branch result and -what must be equal before two branches are summed — cited here and not compressed, a part of it -dropped being a range nobody checked; **a re-review after every fix**, a fix changing the artifact so -the prior review no longer covers it; **a fix that changes specified behaviour updating the spec in -the same commit**, so the re-review covers both; the battery and the mode-derived evidence the -profiles section obliges before a call; and the **evidence entry**, revalidated as that section -says. - -**The act** is the one Mechanics · Finishing the cycle describes — the closing amend, or, where -several `WIP:` snapshots piled up, the reset and single commit that section prescribes instead — -performed once the ordering reaches it: an eligible pass with every closure condition holding, -**never a clean pass on its own**. +section reaches it**. **What this gate does have is every Gate-B duty this section already +states, at the paragraphs that state them, and none of them is summarised here** — a compressed +inventory is where a load-bearing part goes missing while the list still looks complete. + +**The act** is the one Mechanics · Finishing the cycle describes, in whichever shape that section +gives the repository's current state. **This paragraph says when it is performed and never which +shape it takes**: once the ordering reaches it, on an eligible pass with every closure condition +holding, **never a clean pass on its own**. **A commit the hook reads as cycle-closing is a Gate-B matter.** A non-`WIP` commit mid-cycle makes the hook drop its Gate-B review state and that gate's counter — an observation about the counter, @@ -725,10 +719,11 @@ their definition site. W 1011–1012. ``` **Finishing the cycle:** once the closure ordering reaches a Gate-B cycle's closing act — an -eligible pass with every closure precondition holding, never a clean pass on its own — perform that gate's -closing act as the ordering describes it — ordinarily `git commit --amend -m ""`, -which replaces the WIP commit, and the shape the ordering names instead wherever the amend would -leave a `WIP:` commit in history. **The hook treats any +eligible pass with every closure precondition holding, never a clean pass on its own — close it with +`git commit --amend -m ""`, which replaces the WIP commit; **where a `WIP:` snapshot +would survive the amend** — several piled up, or a stray non-amending commit made one an +ancestor — **reset to the parent of the first and commit once instead**. This section is the only +place either shape is defined. **The hook treats any non-`WIP` commit *attempt* as a Gate-B boundary and clears its state even where the command fails**, so a failed closing act leaves that counter cleared and no commit made — a fact about the counter and not about the cycle. This section states the operation and never whether the cycle may From 06152d17f816a0943ec06c06bf8611cd74672a99 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sun, 13 Sep 2026 19:30:08 +0200 Subject: [PATCH 081/181] docs(specs): reconcile pass 51 by content; apply pass 52 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Daniel's advisory of 2026-09-13 19:22 is confirmed exactly, and the reconciliation is recorded in gate-a-spec-awsf1ec771-pass-51-dispositions.md. Pass 51 findings 1 and 2 were labelled design §4, passage (b). §4 line 143 already read "| (b) what a loop absorbs | §B |" before 2044147. That commit rewrote the clean pointer and left the offending row untouched at §5 line 174. Both findings were therefore still open. §4 is restored to its pointer; §5's row keeps the decision provenance and states no operative rule. A finding's section label is a claim about the artifact and gets the same verification as any other claim. Locating by quoted content would have cost one grep and was not done. Pass 52 located the same defect at §5 independently, which is the check that caught it. Findings 3, 4 and 5 verified resolved by grep against the current files. Finding 6 was partly open: §A3 and §F item 5 were repaired and §F item 9a's rationale still restated the same case split. Now a pointer. Pass 52, three further Majors beyond the two the reconciliation already closed: - §F item 8 required an edit to be folded into a WIP snapshot a stray commit had made an ancestor, while no section defines an operation that reaches one. It now states the duty — the edit must be in the content the next review reads — and says plainly that no operation here reaches such a snapshot - §F item 5 repeated the permission §A3 owns. It now gives only the operation - §A1 and §B both defined what accept and decline do to the fix set. §B is the single source, defining the set from those outcomes; §A1 names the two answers and points at it Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- ...-a-spec-awsf1ec771-pass-51-dispositions.md | 39 +++++++++++++++++++ .../gate-a-spec-awsf1ec771-pass-52.md | 30 ++++++++++++++ ...26-09-10-loop-rule-consolidation-design.md | 4 +- ...-10-loop-rule-consolidation-target-text.md | 21 +++++----- 4 files changed, 82 insertions(+), 12 deletions(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-51-dispositions.md create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-52.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-51-dispositions.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-51-dispositions.md new file mode 100644 index 0000000..daba732 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-51-dispositions.md @@ -0,0 +1,39 @@ +# Pass-51 dispositions — reconciled 2026-09-13 against the files, not the labels + +Reconciliation ordered by Daniel's advisory of 2026-09-13 19:22, which was raised against +`2044147` and is **confirmed**. Locations below were found by quoted content; every section label +and line number in the findings file was checked rather than trusted. + +| # | Sev | Label in the findings file | Actual location | Repair in `2044147` | Open after it | +|---|---|---|---|---|---| +| 1 | MAJOR | design §4, passage (b) | **design §5, line 174** | **wrong site.** §4 line 143 already read `\| (b) what a loop absorbs \| §B \|` before the commit; the repair rewrote that clean pointer and left §5's text untouched | **YES — now repaired** | +| 2 | MAJOR | design §4, passage (b) | **design §5, line 174** | same wrong site, same row | **YES — now repaired** | +| 3 | MAJOR | design §7 "consumption clause" | design §7, as labelled | replaced by a pointer to target §A | no — verified absent | +| 4 | MAJOR | design §9 partial-adoption terminal action | design §9, as labelled | replaced by a pointer to §G | no — verified absent | +| 5 | MAJOR | §A3 "What this gate does have" | target §A3, as labelled | the seven-duty summary deleted; duties stay at the paragraphs that state them | no — verified absent | +| 6 | MAJOR | §A3 and §F item 5 "The act" | target §A3 **and §F item 9a's rationale** | §A3 and item 5 repaired; **item 9a's rationale still restated the case split** | **partly — now repaired** | + +## What the mis-location cost, recorded because it is the lesson + +The findings file named §4 and I edited §4. **§4 was already correct.** The repair therefore +*added* a restatement to a clean pointer row while the offending row two sections down was never +touched — a repair that made one site worse and the named defect no worth of progress. **A finding's +section label is a claim about the artifact and gets the same verification as any other claim**; +locating by quoted content would have cost one grep. + +Finding 6 shows the second half of the same habit: the two sites the finding named were repaired +and a third site carrying the same duplication was not looked for. + +## Repairs now applied + +- **design §4** restored to `| (b) what a loop absorbs | §B |`, the pointer it was. +- **design §5's passage (b) row** keeps the decision provenance — decision 2 in §2, and pass 27 + finding 2 for the accept rule — and states no operative rule, pointing at target §B. +- **§F item 9a's rationale** points at item 5 instead of restating the amend-versus-reset split. + +Re-read after repair: §4 line 143 and §5 line 174 both verified in place; the precheck exits 0 on +both files; `grep` confirms no surviving copy of the consumption clause, the §9 terminal action or +the Gate-B duty summary. + +**This is local repair verification and not a clean gate pass.** Pass 52 ran against `2044147` and +its findings are open; the cycle continues under its own rules. diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-52.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-52.md new file mode 100644 index 0000000..0385e65 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-52.md @@ -0,0 +1,30 @@ +MAJOR | high | §F item 8 "otherwise by the shape that reaches it" | The profile-change replacement requires an edit to be folded into a WIP snapshot that a stray commit has made an ancestor, but neither this sentence nor its unnamed pointer defines the mid-cycle Git operation that reaches and rebuilds that snapshot; §F item 5 defines only the later closing act | The next Gate-B review can omit the current profile edit or an agent can perform the cycle-closing reset prematurely, leaving the active cycle without a reliable reviewed snapshot | Name the authoritative mid-cycle recovery operation in the installed text, or cite a complete standing procedure that actually defines it +MAJOR | high | §§A3 and F item 5 "once the ordering reaches" | §A3 says it alone states when Gate B's act is performed, but item 5 repeats the permission as an eligible pass with every closure precondition holding before giving the operation | Permission to close has two authorities despite item 5's claim that it states the operation and never whether the cycle may close, so one copy can lose a closure condition while still looking complete | Let §A3 own permission and begin item 5 with a pointer to §A3 entire before defining only the amend versus reset operation +MAJOR | high | design §5 passage (b) "change costs at least one further pass" | The passage map still restates target §B's complete assigned-fix-set change rule, including direction, undo behavior and the opening of its window, although design §4 now says the behavior is stated in target §B and nowhere in the design | The plan retains a second operational authority that can drift by losing one of the change qualifications | Replace the parenthetical behavior with a pointer to target §B entire +MAJOR | high | design §5 passage (b) "an accept at a membership stop" | The passage map separately restates target §B's rule that acceptance puts a finding in the set regardless of severity and that an accepted Minor owes no severity repair | Membership and Minor handling can diverge between the design map and the text the plan installs | Remove the behavioral summary and cite target §B entire +MINOR | high | design §5 passage (c) "third condition is widened" | The passage map states again that the clearly-stuck reading admits a re-raised validated dismissal, duplicating the operative rule in target §C | The design remains a second authority for whether that recurrence qualifies and can drift independently of §C | Say only that passage (c) is edited per target §C +MINOR | high | design §8 "including the three exempted" | The prompt-standards discussion repeats three §A1 rules as a list: no other pass outcome is eligible, zero findings are clean regardless of floor, and decline is membership-only | This second list can go stale while still being cited as evidence that the prompt carries the required rationales | Cite the relevant §A1 rationale sentences without restating the rules or the list +NIT | high | design §1 "four standing duties" | The intent copies target §A1's duty count even though target §A1 is the source that classifies the duty set | A later duty change can leave a stale numeric claim in the design | Remove the number and refer to the standing duties classified in target §A1 +NIT | high | design header, §5 and §9 "135 conditions" | The frozen inventory owns the total of 135, but the design copies that number at lines 12, 165 and 379 | Any correction to the frozen accounting would leave three stale numeric claims outside its source | Refer to every condition in the committed inventory without copying its total +NIT | high | design §5 "complete at ten" | The design copies the inventory's ten-passage count beside its own passage map | A passage-map change can leave the numeric completeness claim stale while the table still looks authoritative | Remove the count and require the plan to cover every passage in the frozen inventory +MINOR | high | target preface and design §3 "four parallel descriptions" | The target preface duplicates the design's history and counts: four parallel descriptions, seventeen passes, a Blocker flat at one for four passes, and three descriptions removed | The design and target metadata now provide two sources for the same change history and can drift independently | Keep this history in design §3 and replace the target preface account with a pointer +MINOR | high | design §4 "No total is claimed, here or in the target text" | Target §F explicitly claims a total of fourteen falsified standing sentences, contradicting this unqualified assertion and design §1's statement that §F owns that count | A reader cannot tell whether fourteen is intentional plan input or a count the design says does not exist | Scope the sentence to the §4 replacement-span table and leave §F's distinct sentence count at §F +NIT | high | §A1 "The source-block branch" | The branch repeats four examples of source stop conditions already defined in the profile and evidence passages, despite labeling the list non-exhaustive | The duplicate list can drift or acquire misleading omissions while appearing beside the governing branch | Remove the examples and point to every source rule that prescribes stop-and-surface +NIT | high | §A1 "a pass's cleanliness is settled" | The clean branch states that a later answer never rewrites a pass, then the dismissal paragraph states the same permanence rule again | Two statements inside the ordering can drift on whether an earlier pass changes retrospectively | Keep the rule in the clean predicate and make the dismissal consequence cite that sentence entire +MINOR | high | §A1 "A re-raised valid dismissal stays discharged" | The suspension paragraph repeats all four qualifying conditions already defined by the clean predicate for a validly dismissed recurrence | The two copies can disagree on new evidence, changed text, sameness or the continuing truth of the dismissal reason | Keep the complete qualification in the clean predicate and replace the later enumeration with a pointer to it entire +MINOR | high | §§A1 and B "out of set or opens a new structural question" | The continue branch redefines the two scope-stop triggers that §B owns, including the already-answered qualification only implicitly | Trigger membership can diverge between the branch and the assigned source, changing whether a Minor or Nit suspends | Refer to a scope-stop trigger as §B defines it and remove the trigger case split from the continue branch +MAJOR | high | §§A1 and B "accept" and "decline" membership effects | §A1 says acceptance joins the fix set and decline keeps a finding outside for the cycle, while §B independently defines the set as accepted findings minus declined findings and separately states the acceptance effect | A membership answer has two state-transition authorities, so a later edit can make the computed set disagree with the suspension transition | Keep the membership transition in §A1 and have §B define the set by reference to membership-stop outcomes, or choose §B as the complete source and cite it from §A1 +NIT | high | §A1 "Composition, and what cannot happen" | The composition paragraph repeats the complete answer list already given immediately above: accept or decline, a question decision, and continue | The duplicated case list can drift on which answer resumes each suspension | State composition over every answer required by the preceding paragraph without enumerating those answers again +MINOR | high | §B "the floor, the Blocker/Major filter and the clean-final-pass rule all stand" | Passage B preserves a second three-item list of cycle-general duties that §A1 now classifies and owns | A later change to the common conditions can leave the scope-stop entry point teaching an incomplete set | Replace the list with a pointer to the cycle-general conditions in §A1 entire +MINOR | high | §§E and F item 3 "read before the ceiling" | §E defines the curve's pre-ceiling severity reading and its rationale, while item 3 states the same curve rule and rationale again at the curve paragraph | The health series can be implemented from either copy and drift on which severity the counts use | Make the curve paragraph cite the complete loop-health severity rule in §E and retain only why the curve records its series +MINOR | high | §§E and F item 9b "severity buys no repair round and no further pass" | The fallback replacement repeats the Minor and Nit consequence already defined in §E | A future severity change can leave the fallback as a conflicting second command | Point to Mechanics · Severity entire and preserve only that any later pass cost comes from a scope decision +MINOR | high | §F items 2 and 8b "filter to Blocker/Major for what must be repaired" | The Gate-A and Gate-B replacements independently state the same downstream filtering rule and that every line feeds all other predicates | The gate entry points can drift on whether Minor and Nit lines reach cleanliness, scope, fix-set or health evaluation | Put the shared reading rule at one common source and have both gate instructions cite it entire +MINOR | high | §§A3 and F items 1 and 8 "non-WIP commit" | The hook reset is stated at three sites: §A3 says it drops Gate-B review state and the counter, item 1 says it discards the pass count, and item 8 says it discards the hook count | The copies already describe different portions of state and can drift on which state a stray commit loses | Keep one verified hook-effect statement and replace the other two with pointers to it +MINOR | high | §F item 1 "What the hook loses is its counter state" | The hook does not lose only counter state: codex-gate.sh lines 886–888 remove the Gate-B fingerprint file and both count files | A reader can expect the prior fingerprint to survive a stray or failed commit attempt and misdiagnose the next hook result | Remove the state inventory and cite §A3's fuller statement, or name the fingerprint and both counters +MINOR | high | §F heading "The fourteen standing sentences this change falsifies" | The count is presented as all standing sentences falsified by the change, but §H separately replaces live sentences because the new cleanliness, cadence, lens and unknown-start rules make them false or incomplete | The plan-facing count has no stated category boundary and can be mistaken for proof that the false-sentence sweep is complete | Name §F as fourteen additional standing sentences outside the passage replacements, or state the counted set's boundary +MINOR | high | §F item 7 "only the Gate-A clause changes" | The proposed sentence also changes Gate B from restatement by the closing amend to restatement by the commit its closing act produces | The section's own accounting can direct the plan to preserve amend-only Gate-B wording while installing a replacement that widens it | State that both gate clauses change and give the Gate-B edit its closing-shape rationale +MINOR | high | §F item 7 "first of the three closing paths" | The rationale retains a three-path account even though §A2 defines two Gate-A repository-state cases, and the old spec-or-plan destination is the closing commit in both cases | The replacement is justified by a false case split and can make the plan preserve obsolete destination logic | Remove the enumeration and say only that following the closing act keeps the record with closure in every §A2 case +NIT | high | §F item 4 line citation | The cited live sentence starts at CLAUDE.md 1013 and workflow-init.md 1197 with "Those have their" and ends at 1015 and 1199, but the stated ranges begin one line late | The mandatory mechanical check omits the opening line of the sentence it claims to locate | Change the ranges to C 1013–1015 and W 1197–1199 +NIT | high | §F item 8 line citation | The cited live sentence starts at CLAUDE.md 751 and workflow-init.md 937 with "Inside an active Gate-B cycle" and ends at 753 and 939, but the stated ranges begin one line late | The mandatory mechanical check omits the opening line of the sentence it claims to locate | Change the ranges to C 751–753 and W 937–939 +NIT | high | §F item order | The specified enumeration order is 1–7, 7a, 8, 8a, 8b, 9a, 9b, 9, but item 8 appears before item 7a | The numbering is mechanically out of sequence and makes ordinal references harder to audit | Move item 7a before item 8 without changing either replacement +END OF FINDINGS (29 total) diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index 0c8e533..751cd6f 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -140,7 +140,7 @@ to point at and so a reader can see the shape of the change without reading the |---|---| | the closure ordering | **new**, and it is section A of the target text | | (a) the floor paragraphs | §H | -| (b) what a loop absorbs | edited — it owns both scope triggers and the assigned fix set, so every qualification the ordering needs is made there. **What it now says, including what a change to the set and an accept at a membership stop cost, is stated in the target text's §B and nowhere here** | §B | +| (b) what a loop absorbs | §B | | (c) recognizing clearly stuck | §C | | (e) the five tells | §D | | Mechanics · Severity | the resolve duty is scoped and gains its discharge rule; the handed-over question is replaced by its answer | @@ -171,7 +171,7 @@ here. This table says what happens to each inventoried passage, so the map stays | Passage | This change | Target text | |---|---|---| | (a) the floor paragraphs | edited — closure sentences trimmed to a pointer, the no-restating prohibition scoped, the per-pass fix command pointed at Mechanics | §H | -| (b) what a loop absorbs | edited — owns both triggers and the fix set, so every qualification the ordering needs is made **here**, which is what keeps one definition per rule; **gains the closing-time rule for a change to the set**, which no inventoried condition carried because none existed (Daniel, 2026-09-12: a change costs at least one further pass, in either direction and whether or not it is later undone, the window opening where the set is fixed for the pass); **an accept at a membership stop puts the finding in the set whatever its severity** (pass 27 finding 2), membership and the repair duty being different things, so an accepted Minor is in the set although Severity asks no repair for it | §B | +| (b) what a loop absorbs | edited — it owns both scope triggers and the assigned fix set. **What it now says is stated in the target text's §B and nowhere here**; the decisions behind the two additions are recorded in §2 as decision 2 (a set change costs a further pass, Daniel 2026-09-12) and at pass 27 finding 2 (an accept enters the set whatever its severity) | §B | | (c) recognizing clearly stuck | edited — keeps its three-condition reading, stops carrying evaluation order; **its precedence sentence is split**, the operative clause moving into the block unchanged and capitalized there while the plateau rationale stays at this source (pass 26 finding 9, pass 27 finding 6); the third condition is widened here to admit a re-raised validated dismissal, and its reason is restated because the deadlock it named is now answered by the clean predicate | §C | | (d) from pass 4 onward | **no longer edited.** The unavailable-history block moved to the successor with **D10** | — | | (e) the five tells | edited — the threshold is read after clean completion, and a pointer says what its answer does | §D | diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index d5d8623..28ecc73 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -301,9 +301,9 @@ the **question** component by the user's decision on that question; the **clearl by that reading's continue-or-stop answer. **No answer discharges another route's component**, and a finding surfaced by one route has one component, discharged by the one answer its surface asks for. **Resumption is still the -composition rule's**, which waits for every outstanding answer. At a **membership stop** the answer -is **accept**, the finding joining the fix set where Mechanics · Severity governs it, or -**decline**, the finding staying outside and binding so for the rest of this cycle. A later answer +composition rule's**, which waits for every outstanding answer. At a **membership stop** the answer is **accept** or **decline**; +**what each does to the assigned fix set is the absorb paragraph's, which defines that set from +these outcomes**, and what an in-set finding then owes is Mechanics · Severity's. A later answer that contradicts a decline **does not reverse it**: the decline **remains binding** and the contradiction is **surfaced to the user as information**, changing neither membership, nor the cycle's state, nor any outstanding question — **whether the loop resumes is decided by the @@ -718,8 +718,8 @@ their definition site. **5. The `Finishing the cycle` lead-in** (Mechanics · `baseSha`). It wraps across C 827–828 and W 1011–1012. ``` -**Finishing the cycle:** once the closure ordering reaches a Gate-B cycle's closing act — an -eligible pass with every closure precondition holding, never a clean pass on its own — close it with +**Finishing the cycle:** **when a Gate-B cycle's closing act is performed is the closure +ordering's, stated there entire**; this section gives only the operation. Close it with `git commit --amend -m ""`, which replaces the WIP commit; **where a `WIP:` snapshot would survive the amend** — several piled up, or a stray non-amending commit made one an ancestor — **reset to the parent of the first and commit once instead**. This section is the only @@ -775,9 +775,10 @@ instead keeps them together on all three without changing the record's form or f **8. The profile-change paragraph's pass claim** (the profiles section). It wraps across C 752–753 and W 938–939. ``` -Inside an active Gate-B cycle, fold the edit into the active `WIP:` snapshot — by amend where the -snapshot is the tip, and otherwise by the shape that reaches it, a stray non-amending commit having -made the snapshot an ancestor an amend would not touch. A non-`WIP` commit reads to the hook as the +Inside an active Gate-B cycle, the edit must end up **in the content the next review reads** — +folded into the active `WIP:` snapshot by amend where that snapshot is the tip, and otherwise +reviewed as its own change, since no amend reaches a snapshot a stray commit has made an ancestor +and this section prescribes no operation that does. A non-`WIP` commit reads to the hook as the cycle closing and would discard **the hook's count of** the accumulated passes. ``` *Why (pass 33 finding 7):* the live clause says such a commit "would discard the accumulated @@ -830,8 +831,8 @@ and is **revalidated before every Gate-B re-review and before the commit its clo produces** — a fix changes the diff even when the profile sits still. ``` *Why (pass 46 finding 2):* the live clause names "the cycle-closing amend", which is one of the -shapes a Gate-B closing act takes; where several `WIP:` snapshots piled up, or a stray non-amending -commit left a `WIP:` ancestor, the act is a reset and a single commit and no amend occurs. Scoped +shapes a Gate-B closing act takes, item 5 above defining the other and being the only place either +is defined; where that other shape applies, no amend occurs. Scoped to the amend, the final revalidation is owed on one path and skipped on the others. It is the eighth sentence sharing this section's mechanism. From 81fffd193cb78be0c7d608395137ec05c4671024 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sun, 13 Sep 2026 19:35:54 +0200 Subject: [PATCH 082/181] docs(context): bring every handoff surface current for a context clear MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The pass table, the durable prompt and the session handoff had all fallen behind the work. - resume note: pass rows 50, 51, 52 and a 53 placeholder, plus one report covering the stretch — the second-copy class and the mis-located pass-51 repair - .context/gate-a-spec-prompt.md: the five settled decisions, the scope cut, the two residuals, the collected list, the two highest-yield sweeps, and the rule that a finding's section label is verified by quoted content. It was written at pass 18 and had none of this - .remember/remember.md: rewritten from "PARKED at pass 26" to the actual state, with Daniel's standing orders and the three mechanisms - the project memory file, same Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-resume.md | 28 +++++++++++++++++++ 1 file changed, 28 insertions(+) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index b1fae95..4ed9ec0 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -66,6 +66,10 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 37 | a37b5ab | 9→**3** | 0→**1** | 3→**2** | yes | one tell (Blocker 0→1). **Findings 3 — lowest of the cycle.** The Blocker was real and mine: an unrepairable closing failure had no terminal state; now PARKED, the state a stop answer already produces; session 01a09afd-8b4c-7712-93c1-6a3afa109d87 | | 38 | cc7f6dd | 3→**10** | 1→**0** | 2→**3** | yes | one tell (count rose). B+M flat at 3. §G takes its fourth Major; dimensions unfrozen; hold components per route rather than per pair; session 01a09b09-ba8b-7773-b42f-93a2f4261814 | | 39 | 13005de | 10→**5** | 0→**0** | 3→**4** | yes | **zero tells.** Three of four Majors are the three-site fan-out of the pass-37/38 rules — §G limb, §H strict reading, §F pointer. Pattern named and now checked before each pass; session 01a09b1d-8cbc-7d43-8d2f-49ff2339c31f | +| 50 | 68eb365 | 10→**25** | 0→**0** | 4→**11** | yes | one tell. 8 of 11 Majors = the design restating what the target owns, third appearance. design §2's prose of all five decisions replaced by a decision RECORD (92 lines → 30); §A1 and §A3 stopped pointing at §I, which ships nowhere | +| 51 | 2044147 | 25→**23** | 0→**0** | 11→**6** | yes | zero tells. 4 of 6 = the same class in design §§4, 7, 9; 2 inside the target (§A3's duty summary, the §A3/§F-5 circular case split) | +| 52 | 06152d1 | 23→**29** | 0→**0** | 6→**5** | yes | one tell. **Two of five were pass 51's findings 1–2, still open because the pass-51 repair went to the wrong section** — see the pass-51 dispositions file. Pass 52 located them correctly at design §5 | +| 53 | — | — | — | — | not run | next, against `06152d1` | | 40 | 12cf247 | 5→**3** | 0→**0** | 4→**3** | yes | zero tells; first round with no fan-out Major after the three-site check was run BEFORE the pass | | 41 | d26d097 | 3→**2** | 0→**0** | 3→**2** | yes | zero tells; both Majors were compressed pointers of mine dropping a load-bearing part of a standing rule | | 42 | 3b61fe3 | 2→**1** | 0→**0** | 2→**1** | yes | zero tells; §F item 4 had taken back the enumeration it was repaired to avoid | @@ -74,6 +78,30 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 45 | 1c858ac | 7→**7** | 0→**0** | 3→**1** | yes | zero tells; swept all four sites still tying classification to closing | | 46 | 4b713dd | 7→**9** | 0→**1** | 1→**2** | yes | **TWO TELLS — mandatory stop under §5, answered in advance by Daniel's standing "no stops without an absolute block" of 2026-09-13 17:02.** Tells surfaced here and in the commit body, findings kept open, no pass credited clean; the loop continued on that answer. **Eleventh** falsified standing sentence; session 01a09b7c-dd6b-7cb3-b83a-776599286f9b | +## Passes 50–52 — the second-copy class, and one mis-located repair + +**Floor line (unchanged):** derived floor **3**; risk **high**, security **none**; read fresh from +`docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md`. One cited story, level 2. + +- **Trend:** findings 23, 29 (from 25). Blockers 0, 0, 0. Majors 11, 6, **5**. + Blocker+Major 11, 6, **5**. +- **Cluster:** the ordering and the gate blocks, and the design's restatements of them. The + instrument: 0. +- **Tells:** pass 50 one, pass 51 zero, pass 52 one. None reached the threshold. + +**One class produced most of it: a rule stated twice.** Design §2 restated all five behaviour +decisions in prose and is now a **decision record** — that a decision was made, who, when, what +question it answered, which target section states it. §§4, 7 and 9 followed. Inside the target, +§A3's seven-duty summary went, and the §A3/§F-item-5 circular definition of the closing act +resolved to item 5 as the single source. + +**One repair went to the wrong place and is recorded in +`gate-a-spec-awsf1ec771-pass-51-dispositions.md`.** Pass 51's findings 1 and 2 were labelled design +§4; the text was at §5. §4 was already a pure `| §B |` pointer, so the repair rewrote a clean row +and left the offending one untouched. **A finding's section label is a claim about the artifact and +gets the same verification as any other claim** — locating by quoted content costs one grep. Daniel +raised it; pass 52 found the same thing independently. + ## Passes 40–46 — the two-tell stop at pass 46, and how it was answered **Floor line (unchanged):** derived floor **3**; risk **high**, security **none**; read fresh from From c73367f800861f7fe3eacfeb8290ab22e8533fb8 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sun, 13 Sep 2026 21:09:37 +0200 Subject: [PATCH 083/181] =?UTF-8?q?docs(specs):=20apply=20pass=2053;=20the?= =?UTF-8?q?=20=C2=A7F=20sweep=20reaches=20the=20shipped=20hook?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Major 1: §A1 assumes nothing about what a failed closing act left behind, while §A3 and §F item 5 asserted it adds no commit. Both assertions deleted; §A1's repository-state rule is the single source. Majors 19-21: three reminder strings in codex-gate.sh state rules this change supersedes — Gate A's "no content check behind it" honesty claim, and the clean-pass definition abbreviated to "no new Blocker/Major" in both satisfied branches. §F 14 -> 17 items; the Gate-B fingerprint overclaim in the same message stays parked. Narrow scope opening authorised by Daniel 2026-09-13: the three strings and the one exact-match expectation in codex-gate.test.sh are in scope; hook behaviour is not. Story §Out of scope and design §§3, 6, 7 carry the carve-out. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- ...-a-spec-awsf1ec771-pass-53-dispositions.md | 54 ++++++++++++++++++ .../gate-a-spec-awsf1ec771-pass-53.md | 22 ++++++++ .../gate-a-spec-awsf1ec771-resume.md | 29 +++++++++- ...26-09-10-loop-rule-consolidation-design.md | 22 ++++++-- ...-10-loop-rule-consolidation-target-text.md | 56 ++++++++++++++++--- ...026-08-29-loop-rule-consolidation-story.md | 9 ++- 6 files changed, 177 insertions(+), 15 deletions(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-53-dispositions.md create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-53.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-53-dispositions.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-53-dispositions.md new file mode 100644 index 0000000..a9ffee5 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-53-dispositions.md @@ -0,0 +1,54 @@ +# Pass 53 dispositions — cycle `awsf1ec771` + +Advisory. One line per finding: verdict + reason. 21 findings: 4 MAJOR, 14 MINOR, 3 NIT. +Reviewed at `81fffd1`. Blocker count 0. + +## Majors — all four accepted, three of them via a scope decision + +1. **ACCEPTED, repaired.** §A1 "Nothing is assumed about what the attempt left behind" versus + §A3's "a failed attempt adds no commit" and §F item 5's "no commit made". Verified by quoted + content at lines 175, 438 and 728. Both assertions deleted; §A1's repository-state rule is now + the single source, per the file's own cite-don't-restate rule. +19. **ACCEPTED, scope opened narrowly, repaired in §F.** The shipped string at + `codex-gate.sh:967` — "Gate A has no content check behind it — this floor is the only thing + keeping the spec review honest" — is falsified by §A2's content condition. Verified present + verbatim. Now §F item 10. +20. **ACCEPTED, scope opened narrowly, repaired in §F.** The clean-pass definition "no new + Blocker/Major" at `codex-gate.sh:973` and `:956` is an abbreviated second copy of a predicate + §A1 now states in full. Verified present verbatim in both. Now §F items 11 and 12. +21. **SPLIT.** The clean-definition half is item 20's, handled above. The fingerprint half — + "only the fresh pass(es) carry the same fingerprint as what you are committing" — is a + **pre-existing** overclaim, parked with the Gate-B tree-equality material by Daniel's decision + of 2026-09-13, and §I already records that the hook's fingerprint is not a substitute. Not + repaired, deliberately. Recorded in §F item 12 as explicitly untouched. + +## The scope decision behind 19, 20 and 21 + +The story and the design both put hook code out of scope. A reviewer opinion Daniel obtained +argued that prompt-standards item 7 (`docs/prompt-standards.md:53–55`, "if it supersedes one, +update the old text in the same change") reaches a hook reminder, a reminder being prompt text an +agent acts on, and that the blast radius was smaller than first estimated. Verified: three `note` +strings, one exact-match test expectation (`codex-gate.test.sh:1019`), no fixtures, and a version +bump already owed for the template. **Daniel authorised a narrow opening on 2026-09-13:** the +three strings and that one expectation are in scope; hook behaviour is not. Story §Out of scope +and design §§3, 6, 7 and the out-of-scope list now carry the carve-out. + +Two claims of mine in presenting that decision were unbacked and are corrected in the record: +"half a day" and "fixtures must change". Neither survived checking. + +## Minor and Nit — 17, collected, not iterated + +Per Mechanics · Severity. Four are worth naming because a later pass will meet them again: + +- **4 and 5** (the §F total): design §4 says "no total is claimed here or in the target text" + while §F's heading claims one. The heading now reads seventeen, so the contradiction stands + unchanged in substance. Collected. +- **14**: the target describes the non-`WIP` reset as counter state only, while + `codex-gate.sh:886–888` also clears the reviewed fingerprint state. A factual narrowing, not a + contradiction with the ordering. Collected. +- **17 and 18** (§F items 4 and 8 line citations): the stated start lines are one off against the + quoted text. Mechanical, cheap, and the plan sweeps citations anyway. Collected. + +The remaining twelve are second-copy observations — a rule stated in both §A1 and its owning +section. That is the mechanism the cycle has been converging on for twenty passes and each +instance is a judgement about which site should own the sentence, not a defect in the ordering. diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-53.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-53.md new file mode 100644 index 0000000..9726a47 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-53.md @@ -0,0 +1,22 @@ +MAJOR | high | §§A1, A3 and F item 5 "failed attempt" | §A1 says “Nothing is assumed about what the attempt left behind,” but §A3 says “a failed attempt adds no commit” and item 5 says a failed closing act leaves “no commit made” | Recovery can ignore a commit that was actually created before the surrounding command reported failure, leaving the cycle on the wrong tip or with a `WIP:` ancestor while claiming to retry the same act | Keep the repository-state rule only in §A1 and replace both no-commit assertions with a duty to inspect the current history before selecting the retry operation +MINOR | high | target preface lines 22–25 and design §3 lines 126–129 | Both metadata passages say the clearly-stuck rule credits no surfaced pass as clean, while proposed §H expressly replaces that blanket rule with pass-local cleanliness and says “no credit is withheld for surfacing alone” | The target and its decisions document preserve the very `c18` meaning the replacement removes, so the plan can retain or reintroduce the old no-clean-credit rule | Say only that surfacing does not close the cycle and leaves its findings and holds open; refer to §A1 for whether the pass itself is clean +MINOR | high | standing curve-value paragraph at CLAUDE.md 972–975 and workflow-init.md 1156–1159 | The live sentence says the curve is not durable during a running cycle “since the commit does not exist until the cycle closes,” but §A2 permits the reviewed artifact commit to exist before the Gate-A cycle closes and §F does not replace this sentence | The installed prompt would give a false mechanism reason for the still-valid fact that the curve is not yet durable | Add this live sentence to §F and state that the curve is absent from the closing commit body until the closing act writes it +MINOR | high | design §4 “No total is claimed” | The design says no total is claimed “here or in the target text,” while target §F explicitly claims fourteen falsified standing sentences and design §1 says that count lives only there | The plan cannot tell whether §F's total is intentional input or a count the design denies exists | Scope the sentence to the §4 replacement-span table and leave §F's distinct sentence count at §F +MINOR | high | §F “The fourteen standing sentences this change falsifies” | The heading presents fourteen as the full set of standing sentences falsified by the change, while §H separately replaces several other live sentences because the new cleanliness, cadence, lens and unknown-start rules make them false or incomplete | The plan-facing count has no category boundary and can be mistaken for evidence that the false-sentence sweep is complete | Name §F as the additional falsified sentences outside the passage replacements, or remove the totality claim and let the plan derive the set +MINOR | high | §A1 “A re-raised valid dismissal stays discharged” | The suspension-duty paragraph repeats the same four qualifications already owned by the clean predicate: same complaint, no new evidence, unchanged relevant text and a still-true dismissal reason | The two copies can drift on which recurrence remains dismissed, changing both cleanliness and the resolve duty | Keep the full test in the clean predicate and make the later resolve-duty sentence cite that predicate entire +MINOR | high | §§A1 and B “out of set or opens a new structural question” | The continue branch restates both scope-stop triggers that §B owns and carries the already-answered qualification only implicitly | A Minor or Nit can be classified as continuing at one entry point and suspending at the other after either copy changes | Say only “carries a scope-stop trigger as §B defines it” in the continue branch and leave both trigger definitions in §B +NIT | high | §A1 “Composition, and what cannot happen” | The composition paragraph repeats the complete answer list given immediately above: accept or decline, a question decision and continue | The duplicated list can drift on which answer resumes each suspension | Define composition over every answer required by the preceding paragraph without enumerating those answers again +MINOR | high | §B “the floor, the Blocker/Major filter and the clean-final-pass rule all stand” | Passage B preserves a second list of cycle-general duties even though §A1 owns their classification and expressly avoids complete inventories elsewhere | A later common-condition change can leave the scope-stop entry point teaching an incomplete closure rule | Replace the three-item list with a pointer to §A1's cycle-general conditions entire +MINOR | high | §§E and F item 3 “read before the ceiling” | §E defines the curve's pre-ceiling severity reading and its rationale, while item 3 defines the same reading again at the curve paragraph | The health series can be implemented from either authority and drift on which severity its counts use | Make the curve paragraph cite §E's complete observation rule and retain there only why the curve records all three series +MINOR | high | §§E and F item 9b “severity buys no repair round and no further pass” | Item 9b repeats the Minor/Nit consequence already defined in §E instead of pointing to the severity rule | A future severity change can leave the fallback issuing a conflicting pass instruction | Point item 9b to Mechanics · Severity entire and preserve only that a later pass cost can come from a scope decision +MINOR | high | §F items 2 and 8b “filter to Blocker/Major for what must be repaired” | The Gate-A and Gate-B replacements independently state the same downstream filter and the same every-line reading rule | The two gate entry points can drift on whether Minor and Nit findings reach cleanliness, scope, fix-set or health evaluation | Put the shared reading rule at one common source and have both gate instructions cite it entire +MINOR | high | §§A3 and F items 1 and 8 “non-WIP commit” | The hook-reset behavior is defined at three sites with different state descriptions: review state plus counter, counter state, and pass count | The copies can diverge on what a stray commit erases and send recovery to different states | Keep one verified hook-effect statement and replace the other two with pointers to it +MINOR | high | §§A3 and F item 1 “counter state” | Both passages narrow the reset to counter state even though `codex-gate.sh` lines 886–888 delete the reviewed fingerprint state file as well as the pass-count and freshness-count files; §A3's preceding sentence itself distinguishes review state from the counter | A reader can expect the reviewed fingerprint to survive and misdiagnose the next freshness or not-run result | Remove the state inventory and point to the hook, or name the fingerprint and both counters exactly once +MINOR | high | design §5 passage (c) “third condition is widened” | The passage map states again that a re-raised validated dismissal satisfies the clearly-stuck reading's third condition, duplicating the operative rule in target §C | The design remains a second authority for whether that recurrence qualifies | Say only that passage (c) is edited per target §C +MINOR | high | design §8 “including the three exempted” | The prompt-standards discussion repeats three §A1 rules as a list: no other pass outcome is eligible, zero findings are clean regardless of floor, and decline is membership-only | This second list can go stale while still being offered as evidence that the target carries the required rationales | Cite the relevant §A1 rationale sentences without restating their rules or count +NIT | high | §F item 4 line citation | The quoted live sentence begins at CLAUDE.md 1013 and workflow-init.md 1197 with “Those have their” and ends at 1015 and 1199, but the stated ranges begin one line late | The mandatory mechanical location check omits the opening line of the sentence it claims to locate | Change the ranges to C 1013–1015 and W 1197–1199 +NIT | high | §F item 8 line citation | The quoted live sentence begins at CLAUDE.md 751 and workflow-init.md 937 with “Inside an active Gate-B cycle” and ends at 753 and 939, but the stated ranges begin one line late | The mandatory mechanical location check omits the opening line of the sentence it claims to locate | Change the ranges to C 751–753 and W 937–939 +MAJOR | high | `codex-gate.sh` Gate-A below-floor reminder | The unchanged shipped reminder says “Gate A has no content check behind it — this floor is the only thing keeping the spec review honest,” while §A2 adds Gate A's artifact/request equality condition | Partial adoption or a reader following the reminder receives a direct contradiction about whether Gate A has a content condition and may ignore the new condition | Replace the reminder with the exact mechanism boundary: the hook checks only count, while §5's Gate-A content condition remains instruction-backed; update its fixture expectation with it +MAJOR | high | `codex-gate.sh` satisfied reminders | The unchanged Gate-A and Gate-B reminders define a clean final pass as “no new Blocker/Major,” while §A1 also requires no scope-stop trigger and permits a repeated valid dismissal to be excluded under its full test | A Minor or Nit carrying a membership or question trigger can be presented by the hook as clean, encouraging closure over an unanswered scope duty | Remove the abbreviated clean definition and point both reminders to §5's clean-pass predicate; update the hook tests with the prompt strings +MAJOR | high | `codex-gate.sh` Gate-B satisfied reminder | The unchanged reminder claims fresh passes “carry the same fingerprint as what you are committing,” while §A3 and §I expressly admit that staged content outside the reviewed range can reach the closing commit and that the hook fingerprint is not a substitute | The hook can emit false assurance that unreviewed staged content was covered, contradicting the settled Gate-B scope cut and AGENTS.md's mechanism-claim invariant | State only the exact comparison the hook performs—equality between its saved and current fingerprints—and remove any claim that it proves what the review or commit contains; update the test expectation +END OF FINDINGS (21 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index 4ed9ec0..b60a1ca 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -69,7 +69,8 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 50 | 68eb365 | 10→**25** | 0→**0** | 4→**11** | yes | one tell. 8 of 11 Majors = the design restating what the target owns, third appearance. design §2's prose of all five decisions replaced by a decision RECORD (92 lines → 30); §A1 and §A3 stopped pointing at §I, which ships nowhere | | 51 | 2044147 | 25→**23** | 0→**0** | 11→**6** | yes | zero tells. 4 of 6 = the same class in design §§4, 7, 9; 2 inside the target (§A3's duty summary, the §A3/§F-5 circular case split) | | 52 | 06152d1 | 23→**29** | 0→**0** | 6→**5** | yes | one tell. **Two of five were pass 51's findings 1–2, still open because the pass-51 repair went to the wrong section** — see the pass-51 dispositions file. Pass 52 located them correctly at design §5 | -| 53 | — | — | — | — | not run | next, against `06152d1` | +| 53 | 81fffd1 | 29→**21** | 0→**0** | 5→**4** | yes | zero tells. **Three of four Majors are the first §F sites outside the two prompt copies** — reminder strings in `codex-gate.sh`. Daniel authorised a NARROW SCOPE OPENING on 2026-09-13: the three strings plus one exact-match test expectation are in, hook behaviour stays out, the Gate-B fingerprint overclaim stays parked. §F 14 → 17. Major 1 was a genuine internal contradiction (§A1 vs §A3 + §F-5 on what a failed attempt leaves behind); both assertions deleted | +| 54 | — | — | — | — | not run | next, against the pass-53 repair commit | | 40 | 12cf247 | 5→**3** | 0→**0** | 4→**3** | yes | zero tells; first round with no fan-out Major after the three-site check was run BEFORE the pass | | 41 | d26d097 | 3→**2** | 0→**0** | 3→**2** | yes | zero tells; both Majors were compressed pointers of mine dropping a load-bearing part of a standing rule | | 42 | 3b61fe3 | 2→**1** | 0→**0** | 2→**1** | yes | zero tells; §F item 4 had taken back the enumeration it was repaired to avoid | @@ -78,6 +79,32 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 45 | 1c858ac | 7→**7** | 0→**0** | 3→**1** | yes | zero tells; swept all four sites still tying classification to closing | | 46 | 4b713dd | 7→**9** | 0→**1** | 1→**2** | yes | **TWO TELLS — mandatory stop under §5, answered in advance by Daniel's standing "no stops without an absolute block" of 2026-09-13 17:02.** Tells surfaced here and in the commit body, findings kept open, no pass credited clean; the loop continued on that answer. **Eleventh** falsified standing sentence; session 01a09b7c-dd6b-7cb3-b83a-776599286f9b | +## Pass-53 report — zero tells, and the sweep had been looking in too few files + +**Trend:** findings 25, 23, 29, **21** across 50–53; Blockers 0, 0, 0, **0**; Majors 11, 6, 5, **4**. +**Cluster:** product behaviour — three of four Majors are shipped hook strings the change falsifies; +the fourteen Minors are the familiar second-copy class. **Require↔withdraw:** none. +Two tells would be mandatory-stop; zero are present. + +**What this pass found that fifty-two did not.** §F's sweep had been run against `CLAUDE.md` §5 +and the `workflow-init.md` mirror only. `plugins/dev-workflow/hooks/codex-gate.sh` ships reminder +strings the agent acts on at exactly the moment it decides whether to proceed, and two of them +state rules this change supersedes: "Gate A has no content check behind it — this floor is the +only thing keeping the spec review honest" (`:967`), and the clean-pass definition "no new +Blocker/Major" (`:973` and `:956`). Both verified verbatim. + +**The boundary, and who drew it.** The story and design put hook code out of scope. Daniel took an +independent reviewer opinion, which argued prompt-standards item 7 reaches a hook reminder and that +the cost had been overstated. Checking bore that out: three `note` strings, one exact-match +expectation at `codex-gate.test.sh:1019`, no fixtures, and the version bump already owed for the +template. **Daniel authorised the narrow opening.** Hook behaviour — control flow, counters, +fingerprint computation, routing — stays out, and the Gate-B fingerprint overclaim stays parked +with the tree-equality material. + +**Carry into pass 54:** the prompt now records the opening so the boundary is not re-argued, and +sweep 1 names the hook as a third site. Expect §F citation Nits (items 4 and 8 are one line off) +and the standing second-copy class. + ## Passes 50–52 — the second-copy class, and one mis-located repair **Floor line (unchanged):** derived floor **3**; risk **high**, security **none**; read fresh from diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index 751cd6f..8daca3b 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -218,7 +218,10 @@ commit (`scripts/check-version-bump.sh main` needs the committed bump, §8). **The check — what it must establish, and where it is built.** Every edit that changes a standing meaning owes a **discriminating pair of counts**: one showing the new wording present, one -showing the old wording gone. Each half runs in **both copies** and against **both** the working +showing the old wording gone. Each half runs in **both copies** — and, for the three hook +reminder strings of §F items 10–12, in the hook's **single** copy, which has no mirror and owes +no parity check; the hook suite's exact-match assertion is that edit's second observation — and +against **both** the working tree and the parent tree, so every assertion is observed passing where the change exists and failing where it does not. A one-sided presence check is not enough: a copy carrying the new wording **and** the old one satisfies it, which is the two-instructions-that-disagree failure §4 @@ -341,9 +344,15 @@ repo's most persistent defect. The transport that could carry it left with the r `plugins/dev-workflow/commands/workflow-init.md` is under `plugins/`, so `plugins/dev-workflow/.claude-plugin/plugin.json` goes `0.11.0 → 0.12.0` with a `plugins/dev-workflow/CHANGELOG.md` entry: a minor bump, the template gaining a closure - ordering and the edited or extended sentences §4 lists. -- **Invariant 4 / the hook.** Untouched: `plugins/dev-workflow/hooks/codex-gate.sh` is not - edited, and the §5 heading it greps (`Cross-Model Review`) does not move. + ordering, the edited or extended sentences §4 lists, and the three hook reminder strings of + §F items 10–12. The hook edits add no bump the template did not already require. +- **Invariant 4 / the hook.** `plugins/dev-workflow/hooks/codex-gate.sh` is edited in exactly + three `note` **strings** (§F items 10–12) and nowhere else: no control flow, no counter, no + fingerprint computation, no routing, so the POSIX-`sh` and optional-`jq` obligations are not + reached. `plugins/dev-workflow/hooks/codex-gate.test.sh` changes in its one exact-match + expectation for the Gate-B satisfied message; the remaining hook assertions match loose + patterns these repairs leave standing. The §5 heading the hook greps (`Cross-Model Review`) + does not move. **Every path in this spec is written repository-relative and in full** — no ellipsis shorthand and no bare basename. An abbreviated citation fails the path-existence check a review pass runs @@ -415,6 +424,9 @@ not exist; §G names a sentence that does, and claims only what that sentence do - The pass-counter anomaly (`fic2` record). - The CodeRabbit plan-metadata contradiction (`fic2` record). - The fixture-per-predicate question (`fic2` record; story §2). -- Hook code under `plugins/dev-workflow/hooks/`. +- Hook **behaviour** under `plugins/dev-workflow/hooks/` — control flow, counters, fingerprint + computation, routing, event handling. The three contradictory reminder strings and the one + exact-match test expectation are in scope instead (§F items 10–12, story §Out of scope), and + the Gate-B fingerprint overclaim in the same message stays parked here. - `todos.md`: both-branches-misread-each-other; self-consuming-deletion (prompt-standards item 11); the three bot findings in resolved plans. diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index 28ecc73..d86f24e 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -435,7 +435,7 @@ since the cycle itself stays open until the conditions hold, so an accidental co the hook reports and closes nothing; **what the reset erases is counter state, and it does not invalidate a pass that already satisfied the validation rules this section states** — which is where what makes a pass valid stays. **A stray commit that succeeds and does not amend also leaves the -`WIP:` snapshot as an ancestor** — a failed attempt adds no commit, and an amend replaces the tip — +`WIP:` snapshot as an ancestor**, an amend replacing the tip instead — so in that one shape **the closing act still owes what Mechanics already requires of it: no `WIP:` commit left in history.** Which git sequence reaches that from this state belongs to the plan, as every other closing sequence does. **It does not reach a Gate-A cycle's count**, which the hook @@ -654,12 +654,13 @@ demotes it — the two counts are meant to differ. --- -## F. The fourteen standing sentences this change falsifies — REPLACED +## F. The seventeen standing sentences this change falsifies — REPLACED -Each is a live sentence that the block makes wrong. All fourteen are **known contradictions** and -none is deferred. **Ten of them share one mechanism** — an entry point other than the ordering +Each is a live sentence that the block makes wrong. All seventeen are **known contradictions** and +none is deferred. **Thirteen of them share one mechanism** — an entry point other than the ordering carrying an unqualified instruction — which is why each is **replaced** rather than given an -exception to point at. +exception to point at. **Fourteen live in the two prompt copies and cite a line in each; the last +three live in the shipped hook**, which has one copy and no template mirror. **1. The `WIP:` naming warning** (Mechanics · `baseSha`). ``` @@ -725,7 +726,7 @@ would survive the amend** — several piled up, or a stray non-amending commit m ancestor — **reset to the parent of the first and commit once instead**. This section is the only place either shape is defined. **The hook treats any non-`WIP` commit *attempt* as a Gate-B boundary and clears its state even where the command -fails**, so a failed closing act leaves that counter cleared and no commit made — a fact about the +fails**, so a failed closing act leaves that counter cleared — a fact about the counter and not about the cycle. This section states the operation and never whether the cycle may close. ``` @@ -864,6 +865,43 @@ path straight past a mandatory stop. It is the sixth sentence of this section's an entry point other than the ordering carrying an unqualified instruction, and it takes the same repair: the **operation** stays here, the **permission** is the ordering's. +**10. The Gate-A below-floor reminder's honesty claim** (the shipped hook, +`plugins/dev-workflow/hooks/codex-gate.sh` — one copy, no template mirror). +``` +Gate A has no content check in this hook; what the gate itself requires of the reviewed artifact is stated in $policy and is instruction-backed. +``` +*Why (pass 53 finding 19):* the live string says "Gate A has no content check behind it — this +floor is the only thing keeping the spec review honest", which the Gate-A content condition +falsifies: the floor is no longer the only thing, and a reader who believes it may treat the +condition as optional. The repair keeps the true half — **this hook** checks only the count — and +sends the reader to the gate's own text for what else is owed. It is the eleventh sentence sharing +this section's mechanism, and the first of three that live in the hook rather than in either +prompt copy. + +**11. The Gate-A satisfied reminder's clean definition** (the shipped hook, same file). +``` +Proceed only if your final pass was clean and every other closure condition holds, both as $policy defines them. +``` +*Why (pass 53 finding 20):* the live string defines a clean final pass as "no new Blocker/Major", +an abbreviated second copy of a definition the ordering now states in full — a pass carrying a +scope-stop trigger is not clean under it, whatever the severity of what triggered it. Left +standing, the hook presents such a pass as clean at exactly the moment an author is deciding +whether to proceed. The repair **cites** the definition rather than restating it, which is also +what keeps a later change to the definition from falsifying this string again. It is the twelfth +sentence sharing this section's mechanism. + +**12. The Gate-B satisfied reminder's clean definition** (the shipped hook, same file). **Its +fingerprint clause is untouched**: that overclaim predates this change, is parked with the +Gate-B tree-equality material, and no condition here reaches it. +``` +Per $policy, commit only if your final pass was clean and every other closure condition holds, both as it defines them. +``` +*Why (pass 53 finding 20):* the same abbreviated definition, in the branch an author reads +immediately before committing, and it takes the same repair. It is the thirteenth sentence sharing +this section's mechanism. **The exact-match expectation in +`plugins/dev-workflow/hooks/codex-gate.test.sh` pins this string in full and is replaced with +it** — the other hook assertions match on loose patterns this repair leaves standing. + --- ## G. The one-contract paragraph — REPLACED @@ -1018,8 +1056,10 @@ be established as absent is cheaper to owe than to skip** — - **Minor, collected and open (pass 17 finding 9 is resolved above; this is the remainder):** nothing in this file establishes that the edit set is complete. It is the sites known at pass - 17. **The plan sweeps both copies against the rule** — a live sentence the ordering falsifies - gets edited — and prints what it found; a spec cannot establish that claim against text the same + 17, plus the three hook reminder strings §F items 10–12 added at pass 53 — which is itself + evidence that the set was not complete. **The plan sweeps both prompt copies and the shipped + hook's reminder strings against the rule** — a live sentence the ordering falsifies gets + edited — and prints what it found; a spec cannot establish that claim against text the same change rewrites, which is the mechanism that produced a finding at passes 15, 16 and 17. - **Partial adoption is instructed against, never detected.** §G says so in its own words. - **Which exit a cycle took is not observable from history.** The transport left with the record diff --git a/docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md b/docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md index 03731fd..b4a6525 100644 --- a/docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md +++ b/docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md @@ -43,7 +43,14 @@ which are preconditions; and what a user's answer on a surfaced finding does in - **The pass floor and severity semantics** — the parent story ships those, and this story treats them as given rather than adjusting them. - **Reopening any decision in §4.** They are settled and paid for; the design starts from them. -- **Hook code** (anything under `plugins/dev-workflow/hooks/`), unchanged from the parent. +- **Hook behaviour** (anything under `plugins/dev-workflow/hooks/` that decides what the hook + does): its control flow, counters, fingerprint computation, routing and event handling, all + unchanged from the parent. **One narrow exception, authorised 2026-09-13:** the three reminder + **strings** this change makes contradictory, and the single exact-match expectation in + `plugins/dev-workflow/hooks/codex-gate.test.sh` that pins one of them, are in scope — listed as + items 10–12 of the target text's §F. Prompt-standards item 7 requires a superseded instruction + to be corrected in the same change, and a hook reminder is prompt text the agent acts on. The + Gate-B fingerprint overclaim in the same message stays parked. - **The pass-counter anomaly**, the CodeRabbit plan-metadata contradiction, and the fixture-per-predicate question — all still parked. From 1940181d9623a8d05c0c27d18344539f6586f3cc Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sun, 13 Sep 2026 21:21:13 +0200 Subject: [PATCH 084/181] docs(specs): apply pass 54; a fourth hook string, and the scope statement it caught MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Major 2: codex-gate.sh:908's WIP reminder ends "then make the real commit when your final pass is clean", making cleanliness the whole closing permission at the moment the ordering exists to govern. §F item 13; §F 17 -> 18. Inside the authorised class — Daniel authorised the reminder strings this change makes contradictory, and "three" in that authorisation was descriptive. Major 3: the design's opening still said no file under plugins/dev-workflow/hooks/ changes, three other scope statements having been updated. Major 1: §A1's continue branch re-enumerated §B's scope triggers, which would let a declined finding or an answered question be re-suspended. Replaced with a pointer to the absorb passage entire. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-54.md | 14 +++++++++ .../gate-a-spec-awsf1ec771-resume.md | 27 ++++++++++++++++- ...26-09-10-loop-rule-consolidation-design.md | 21 +++++++------ ...-10-loop-rule-consolidation-target-text.md | 30 ++++++++++++++----- ...026-08-29-loop-rule-consolidation-story.md | 4 +-- 5 files changed, 76 insertions(+), 20 deletions(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-54.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-54.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-54.md new file mode 100644 index 0000000..7b3c704 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-54.md @@ -0,0 +1,14 @@ +MAJOR | high | §A1 continue branch and §B scope triggers | The continue branch says a Minor or Nit that "is out of set or opens a new structural question carries a trigger like any other", but §B says an already-declined finding raises no membership trigger and an already-answered question raises no question trigger | A declined finding or answered question can be re-suspended and re-asked, contradicting D7 and making the later-pass handling non-idempotent | Replace the trigger enumeration with a pointer to the scope-stop triggers as §B defines them entire +MAJOR | high | `codex-gate.sh` WIP reminder | The live reminder says "then make the real commit when your final pass is clean", but §F does not replace it and §A1 makes the commit conditional on eligibility, every closure condition and the closing act rather than cleanliness alone | An author entering through the WIP reminder can commit over an unmet floor, standing hold, unresolved duty or gate condition | Add this reminder to §F and replace its commit permission with a pointer to the policy's complete closure ordering +MAJOR | high | design opening "No file under plugins/dev-workflow/hooks changes" | The design still says the change is confined to two mirrored prompt copies and that no hook file changes, while target §F items 10–12, story §2 and design §§7–9 require three hook strings and one hook-test expectation to change | The plan has two incompatible scope statements and can omit the newly-authorized hook repairs or treat them as out-of-scope edits | Update the design opening to name the two prompt copies, the three reminder strings and the one exact-match test expectation +MINOR | medium | §A1 "Which severity field each predicate reads" | The sentence claims Mechanics · Severity settles which severity field every finding-derived predicate reads, while §E says subject clusters use no severity and the scope triggers and assigned-fix-set membership are also not severity predicates | A reader can infer a severity dependency for predicates that intentionally inspect every finding independently of severity | Say that Mechanics settles whether and which severity a predicate reads, or scope the pointer to predicates that actually read severity +MINOR | high | §A1 clean predicate and §H Gate-A clean-signal sentence | §H states again that a pass carrying only Minors and no scope-stop trigger is clean, duplicating the clean predicate owned by §A1 | The Gate-A entry point can drift from the ordering on which nonzero-finding passes are clean | Keep §H's `NO FINDINGS` emission rule and cite §A1's clean-pass predicate entire for every other clean outcome +MINOR | high | §A1 "Composition, and what cannot happen" | The composition paragraph repeats the complete resuming-answer list just defined above it: accept or decline, a question decision and continue | The two adjacent authorities can diverge on which answer discharges or parks a suspension | Define composition over every answer required by the preceding paragraph without enumerating the answers again +MINOR | high | §B "the floor, the Blocker/Major filter and the clean-final-pass rule all stand" | The absorb passage preserves a second three-item account of duties that survive a scope stop even though §A1 owns the cycle-general duty classification and closure conditions | A later duty change can leave the scope-stop entry point teaching an incomplete closure rule | Replace the three-item enumeration with a pointer to §A1's cycle-general conditions and duties entire +MINOR | high | §§E and F item 3 "read before the ceiling" | §E defines the curve's pre-ceiling severity rule, while F item 3 defines that same rule again in the curve rationale | The curve can drift between two authorities on whether its severity series are counted before or after demotion | Make the curve paragraph cite §E's observation rule and retain only why all three series are recorded +MINOR | high | §§E and F item 9b "severity buys no repair round and no further pass" | F item 9b repeats the Minor-or-below pass consequence already defined in §E instead of pointing to the severity rule | A future severity edit can leave the fallback issuing a conflicting repair or pass instruction | Point item 9b to Mechanics · Severity entire and retain only the distinction that a later set-change pass cost comes from §B +MINOR | high | §F items 2 and 8b downstream filtering | The Gate-A and Gate-B replacements independently define the same rule to filter to Blocker/Major only for repair while reading every line for all other predicates | The two gate entry points can drift on whether Minor and Nit lines reach cleanliness, scope, fix-set and health evaluation | State the shared reading rule once in the common findings protocol and have both gate instructions cite it entire +MINOR | medium | §F item 6 and §H Gate-A cadence | Item 6 says an unrevised Gate-A rerun is legitimate whenever no repair is owed, while §H separately decides when rerunning is allowed and requires a suspension's answer first | The broad-prompt entry point can be read as permission to rerun before the ordering has released a suspension | Keep rerun timing in §H and make item 6 cite that cadence while retaining only the broad-question requirement +MINOR | high | design §5 passage (c) and target §C | The passage map restates that the clearly-stuck third condition is widened to admit a re-raised validated dismissal, duplicating the operative rule in target §C | The design remains a second authority for which recurrence satisfies the clearly-stuck reading | Reduce the passage-map row to a pointer to target §C and keep only its accounting role +MINOR | high | design §8 "including the three exempted" and target §A1 | The prompt-standards discussion repeats three §A1 rules as a numbered set: no other pass outcome is eligible, zero findings are clean regardless of floor and decline is membership-only | This second list can go stale while still being offered as evidence that the target carries all required rationales | Cite the relevant §A1 rationale sentences without restating their rules or their count +END OF FINDINGS (13 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index b60a1ca..53862d6 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -70,7 +70,8 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 51 | 2044147 | 25→**23** | 0→**0** | 11→**6** | yes | zero tells. 4 of 6 = the same class in design §§4, 7, 9; 2 inside the target (§A3's duty summary, the §A3/§F-5 circular case split) | | 52 | 06152d1 | 23→**29** | 0→**0** | 6→**5** | yes | one tell. **Two of five were pass 51's findings 1–2, still open because the pass-51 repair went to the wrong section** — see the pass-51 dispositions file. Pass 52 located them correctly at design §5 | | 53 | 81fffd1 | 29→**21** | 0→**0** | 5→**4** | yes | zero tells. **Three of four Majors are the first §F sites outside the two prompt copies** — reminder strings in `codex-gate.sh`. Daniel authorised a NARROW SCOPE OPENING on 2026-09-13: the three strings plus one exact-match test expectation are in, hook behaviour stays out, the Gate-B fingerprint overclaim stays parked. §F 14 → 17. Major 1 was a genuine internal contradiction (§A1 vs §A3 + §F-5 on what a failed attempt leaves behind); both assertions deleted | -| 54 | — | — | — | — | not run | next, against the pass-53 repair commit | +| 54 | c73367f | 21→**13** | 0→**0** | 4→**3** | yes | zero tells. All three Majors are pass-53 fallout and all are in-set: a **fourth** falsified hook string (the WIP reminder's "make the real commit when your final pass is clean") → §F item 13, 17→18; the design opening still said "No file under `plugins/dev-workflow/hooks/` changes"; §A1's continue branch re-enumerated §B's scope triggers instead of citing them | +| 55 | — | — | — | — | not run | next, against the pass-54 repair commit | | 40 | 12cf247 | 5→**3** | 0→**0** | 4→**3** | yes | zero tells; first round with no fan-out Major after the three-site check was run BEFORE the pass | | 41 | d26d097 | 3→**2** | 0→**0** | 3→**2** | yes | zero tells; both Majors were compressed pointers of mine dropping a load-bearing part of a standing rule | | 42 | 3b61fe3 | 2→**1** | 0→**0** | 2→**1** | yes | zero tells; §F item 4 had taken back the enumeration it was repaired to avoid | @@ -79,6 +80,30 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 45 | 1c858ac | 7→**7** | 0→**0** | 3→**1** | yes | zero tells; swept all four sites still tying classification to closing | | 46 | 4b713dd | 7→**9** | 0→**1** | 1→**2** | yes | **TWO TELLS — mandatory stop under §5, answered in advance by Daniel's standing "no stops without an absolute block" of 2026-09-13 17:02.** Tells surfaced here and in the commit body, findings kept open, no pass credited clean; the loop continued on that answer. **Eleventh** falsified standing sentence; session 01a09b7c-dd6b-7cb3-b83a-776599286f9b | +## Pass-54 report — zero tells, and the scope opening paid for itself immediately + +**Trend:** findings 23, 29, 21, **13** across 51–54; Blockers 0, 0, 0, **0**; Majors 6, 5, 4, **3**. +**Cluster:** product behaviour. **Require↔withdraw:** none. Zero tells. + +All three Majors were pass-53 fallout, and all three are inside the fix set as pass 53 left it: + +1. **A fourth falsified hook string**, which pass 53 did not find — the WIP-commit reminder's + "then make the real commit when your final pass is clean" (`codex-gate.sh:908`), which makes + cleanliness the whole permission at exactly the moment the ordering exists to govern. Now §F + item 13; §F 17 → 18. Treated as inside the authorised boundary: Daniel authorised the **class** + (reminder strings this change makes contradictory), and the count in that authorisation was + descriptive. Story, design and the pass prompt now say four. +2. **The design's own opening still said "No file under `plugins/dev-workflow/hooks/` changes"** — + the scope statement I updated in three places while missing the first one. It now names the two + prompt copies, the four strings and the one test expectation. +3. **§A1's continue branch re-enumerated §B's scope triggers**, which lets a declined finding or an + answered question be re-suspended. Replaced with a pointer to the absorb passage entire, plus + the standing "severity neither raises nor suppresses a trigger". + +**Reading:** finding count has halved twice (29 → 21 → 13) and Majors are falling. The one thing +worth watching is that pass 53's repair generated all of pass 54's Majors — a lineage, not yet a +plateau, and the three repairs here are pointer-and-scope work rather than new rules. + ## Pass-53 report — zero tells, and the sweep had been looking in too few files **Trend:** findings 25, 23, 29, **21** across 50–53; Blockers 0, 0, 0, **0**; Majors 11, 6, 5, **4**. diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index 8daca3b..c3da555 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -5,9 +5,12 @@ **Profile:** read from that header at every pass, never from here — it is the only writable copy, and a value copied here would be a remembered value. -Prompt-only, in two mirrored copies: `CLAUDE.md` §5 (**C** below) and the inline template in -`plugins/dev-workflow/commands/workflow-init.md` (**W** below). No file under -`plugins/dev-workflow/hooks/` changes. Condition ids `a1`…`j4` are defined in +Prompt-only, in two mirrored copies — `CLAUDE.md` §5 (**C** below) and the inline template in +`plugins/dev-workflow/commands/workflow-init.md` (**W** below) — **plus four reminder strings in +`plugins/dev-workflow/hooks/codex-gate.sh` and the one exact-match expectation in +`plugins/dev-workflow/hooks/codex-gate.test.sh` that pins one of them** (target text §F items +10–13, authorised 2026-09-13). **No hook behaviour changes**, and the hook strings have no mirror, +so they carry no parity obligation. Condition ids `a1`…`j4` are defined in `docs/superpowers/specs/2026-09-10-loop-rule-consolidation-condition-inventory.md` beside this file: 135 conditions quoted from `7c0d475`, so a reader can check an accounting rather than take it. @@ -218,8 +221,8 @@ commit (`scripts/check-version-bump.sh main` needs the committed bump, §8). **The check — what it must establish, and where it is built.** Every edit that changes a standing meaning owes a **discriminating pair of counts**: one showing the new wording present, one -showing the old wording gone. Each half runs in **both copies** — and, for the three hook -reminder strings of §F items 10–12, in the hook's **single** copy, which has no mirror and owes +showing the old wording gone. Each half runs in **both copies** — and, for the four hook +reminder strings of §F items 10–13, in the hook's **single** copy, which has no mirror and owes no parity check; the hook suite's exact-match assertion is that edit's second observation — and against **both** the working tree and the parent tree, so every assertion is observed passing where the change exists and @@ -344,10 +347,10 @@ repo's most persistent defect. The transport that could carry it left with the r `plugins/dev-workflow/commands/workflow-init.md` is under `plugins/`, so `plugins/dev-workflow/.claude-plugin/plugin.json` goes `0.11.0 → 0.12.0` with a `plugins/dev-workflow/CHANGELOG.md` entry: a minor bump, the template gaining a closure - ordering, the edited or extended sentences §4 lists, and the three hook reminder strings of - §F items 10–12. The hook edits add no bump the template did not already require. + ordering, the edited or extended sentences §4 lists, and the four hook reminder strings of + §F items 10–13. The hook edits add no bump the template did not already require. - **Invariant 4 / the hook.** `plugins/dev-workflow/hooks/codex-gate.sh` is edited in exactly - three `note` **strings** (§F items 10–12) and nowhere else: no control flow, no counter, no + four `note` **strings** (§F items 10–13) and nowhere else: no control flow, no counter, no fingerprint computation, no routing, so the POSIX-`sh` and optional-`jq` obligations are not reached. `plugins/dev-workflow/hooks/codex-gate.test.sh` changes in its one exact-match expectation for the Gate-B satisfied message; the remaining hook assertions match loose @@ -426,7 +429,7 @@ not exist; §G names a sentence that does, and claims only what that sentence do - The fixture-per-predicate question (`fic2` record; story §2). - Hook **behaviour** under `plugins/dev-workflow/hooks/` — control flow, counters, fingerprint computation, routing, event handling. The three contradictory reminder strings and the one - exact-match test expectation are in scope instead (§F items 10–12, story §Out of scope), and + exact-match test expectation are in scope instead (§F items 10–13, story §Out of scope), and the Gate-B fingerprint overclaim in the same message stays parked here. - `todos.md`: both-branches-misread-each-other; self-consuming-deletion (prompt-standards item 11); the three bot findings in resolved plans. diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index d86f24e..ad54ffa 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -244,9 +244,11 @@ often a repair still owed from an earlier pass. A below-floor clean pass lands h terms; where a suspension does apply, the suspension branch has already taken it, because only the clean-completion branch outranks a suspension. So does a pass whose only findings are Minors and Nits **and which carries no scope-stop trigger** — those are collected and never iterated and may -leave nothing to revise, while a Minor or Nit that is out of set or opens a new structural question -carries a trigger like any other finding, is not clean, and has already been taken by the -suspension branch. It is a branch and +leave nothing to revise, while a Minor or Nit **carrying a scope-stop trigger as the absorb +passage defines one** carries it like any other finding, is not clean, and has already been taken +by the suspension branch — **severity does not raise a trigger and does not suppress one**, and +which findings raise one is that passage's entire, an already-declined finding and an +already-answered question raising none. It is a branch and not an inference, because "does not close" read alone says nothing about whether to run again. **The four standing duties, classified.** The **derived floor** is a **precondition on closure**: @@ -654,13 +656,13 @@ demotes it — the two counts are meant to differ. --- -## F. The seventeen standing sentences this change falsifies — REPLACED +## F. The eighteen standing sentences this change falsifies — REPLACED -Each is a live sentence that the block makes wrong. All seventeen are **known contradictions** and -none is deferred. **Thirteen of them share one mechanism** — an entry point other than the ordering +Each is a live sentence that the block makes wrong. All eighteen are **known contradictions** and +none is deferred. **Fourteen of them share one mechanism** — an entry point other than the ordering carrying an unqualified instruction — which is why each is **replaced** rather than given an exception to point at. **Fourteen live in the two prompt copies and cite a line in each; the last -three live in the shipped hook**, which has one copy and no template mirror. +four live in the shipped hook**, which has one copy and no template mirror. **1. The `WIP:` naming warning** (Mechanics · `baseSha`). ``` @@ -902,6 +904,18 @@ this section's mechanism. **The exact-match expectation in `plugins/dev-workflow/hooks/codex-gate.test.sh` pins this string in full and is replaced with it** — the other hook assertions match on loose patterns this repair leaves standing. +**13. The WIP-commit reminder's closing permission** (the shipped hook, same file). +``` +Run the review against this commit; when the closing act may then be performed is $policy's closure ordering's, read there in full. +``` +*Why (pass 54 finding 2):* the live string ends "then make the real commit when your final pass is +clean", which makes cleanliness the whole permission — and it reaches the author at the one moment +the ordering exists to govern. Under the ordering a clean eligible pass closes nothing while a +precondition is unmet, so an author entering here can amend over a below-floor count, a standing +hold or an undischarged Major. It is the fourteenth sentence sharing this section's mechanism, and +the fourth found in the hook; the first three were found at pass 53 and this one at pass 54, which +is why §I records the edit set as established by sweeping rather than by this file. + --- ## G. The one-contract paragraph — REPLACED @@ -1056,7 +1070,7 @@ be established as absent is cheaper to owe than to skip** — - **Minor, collected and open (pass 17 finding 9 is resolved above; this is the remainder):** nothing in this file establishes that the edit set is complete. It is the sites known at pass - 17, plus the three hook reminder strings §F items 10–12 added at pass 53 — which is itself + 17, plus the four hook reminder strings §F items 10–13 added at passes 53 and 54 — which is itself evidence that the set was not complete. **The plan sweeps both prompt copies and the shipped hook's reminder strings against the rule** — a live sentence the ordering falsifies gets edited — and prints what it found; a spec cannot establish that claim against text the same diff --git a/docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md b/docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md index b4a6525..2d08417 100644 --- a/docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md +++ b/docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md @@ -45,10 +45,10 @@ which are preconditions; and what a user's answer on a surfaced finding does in - **Reopening any decision in §4.** They are settled and paid for; the design starts from them. - **Hook behaviour** (anything under `plugins/dev-workflow/hooks/` that decides what the hook does): its control flow, counters, fingerprint computation, routing and event handling, all - unchanged from the parent. **One narrow exception, authorised 2026-09-13:** the three reminder + unchanged from the parent. **One narrow exception, authorised 2026-09-13:** the four reminder **strings** this change makes contradictory, and the single exact-match expectation in `plugins/dev-workflow/hooks/codex-gate.test.sh` that pins one of them, are in scope — listed as - items 10–12 of the target text's §F. Prompt-standards item 7 requires a superseded instruction + items 10–13 of the target text's §F. Prompt-standards item 7 requires a superseded instruction to be corrected in the same change, and a hook reminder is prompt text the agent acts on. The Gate-B fingerprint overclaim in the same message stays parked. - **The pass-counter anomaly**, the CodeRabbit plan-metadata contradiction, and the From 6b8854fcf171f71c48ec81357c4f72542780a6b0 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sun, 13 Sep 2026 21:33:07 +0200 Subject: [PATCH 085/181] docs(specs): apply pass 55; the hook reminder class, swept to its edge MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Majors 3-6 plus pass 54's item 13 are one class: a gate reminder that ends by telling the agent what to do next states a permission the closure ordering owns. codex-gate.sh carries eight gate reminders; §F now covers seven, the docs-only notice stating no closure permission. §F 18 -> 22. The class is bounded and now swept, so the count stops moving unless the hook gains a reminder. Major 2: the standing Named residual still read "Hook text is out of scope here by decision", which the scope opening of 2026-09-13 falsifies. §F item 14. The residual it was written for stays out of scope. Major 1: §A1's source-block branch had a repaired block reread "that pass" where the block stood before any pass of the cycle existed — the floor derivation's blocks among them. Split into before-any-pass and on-a-pass-already-read. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-55.md | 19 ++++++ .../gate-a-spec-awsf1ec771-resume.md | 31 ++++++++- ...26-09-10-loop-rule-consolidation-design.md | 16 ++--- ...-10-loop-rule-consolidation-target-text.md | 64 ++++++++++++++++--- ...026-08-29-loop-rule-consolidation-story.md | 4 +- 5 files changed, 115 insertions(+), 19 deletions(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-55.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-55.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-55.md new file mode 100644 index 0000000..881f86a --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-55.md @@ -0,0 +1,19 @@ +MAJOR | high | §A1 "The source-block branch" | The branch includes a profile present but unresolvable, disagreeing governing headers and an unreadable `Story:` header, but the standing text detects each before a pass runs while §A1 says all branches are read "on the pass" and, after repair, to read the ordering again "on that pass" | At cycle start there is no logical pass or validated findings file to reread, so the repaired source block has no unambiguous route into the first review pass | Treat source-prescribed blocks as pre-pass gates: after repair run the pass, and reserve same-pass rereading for a block actually discovered while accepting an existing pass +MAJOR | high | `CLAUDE.md` §5 and `workflow-init.md` "Named residual" | Both standing copies still say "Hook text is out of scope here by decision"—the sentence wraps after "here" in each source—while §F, the design and the story now put four hook reminder strings and one exact-match expectation in scope | The installed text would contradict the change's settled scope and can make the plan omit authorized hook-string repairs | Replace the blanket sentence with the narrow residual that remains out of scope and point to §F's in-scope reminder-string exceptions +MAJOR | high | `codex-gate.sh` no-fingerprint reminder | The live string says "Run Gate B (mcp__codex__review) now" even though its own preceding alternatives include a pass whose fingerprint could not be written or read back, and that pass may have left a suspension awaiting answers | An agent can run a new pass over an unanswered scope or health suspension, violating §A1's composition rule | Preserve the fingerprint diagnosis but send pass timing to the closure ordering, allowing a rerun only after every outstanding suspension answer permits it +MAJOR | high | `codex-gate.sh` stale-fingerprint reminder | The live string unconditionally says "Run Gate B ... now" and later "run one more pass", while §A1 forbids another pass when the current pass is suspended or parked | A fingerprint warning can make the agent step past an unanswered scope, clearly-stuck or two-tell suspension instead of taking its required transition | Keep the diagnostic remedies, but make the next review conditional on the closure ordering releasing the cycle to a pass +MAJOR | high | `codex-gate.sh` Gate-B below-floor reminder | The live string orders "run more" and alternatively offers the skip rule, but §A1 sends a below-floor pass to a suspension when one applies and says the Gate-B triviality skip runs no passes | The reminder can bypass a mandatory suspension or retroactively turn an already-running Gate-B cycle into a skipped one | Replace the instruction with a pointer to the ordering; say another pass runs only on its continue route and the triviality skip is available only before a cycle runs a pass +MAJOR | high | `codex-gate.sh` Gate-A below-floor reminder | After §F item 10's proposed sentence replacement, the same live reminder still ends "Run more passes before executing" without accounting for a scope, clearly-stuck or two-tell suspension on the below-floor pass | An agent entering through the hook can spend another pass before answering the suspension that §A1 says must be handled first | Extend item 10 to replace the run-more command with a pointer to the ordering's suspension-or-continue decision +MINOR | high | §A1 "Which severity field each predicate reads" | The sentence says Mechanics · Severity settles which severity field each finding-derived predicate reads, but subject clustering, scope triggers and assigned-fix-set membership derive from findings without reading a severity field | Readers can infer a severity dependency for predicates that intentionally inspect findings independently of severity | Say Mechanics settles whether a predicate reads severity and, where it does, which severity field it reads +MINOR | high | §A1 continue branch and §B scope-trigger exceptions | After correctly declaring §B the entire source of which findings raise triggers, §A1 immediately restates that an already-declined finding and an already-answered question raise none | The new second copy can drift from §B and recreate the exact conflicting-trigger defect the pointer was added to remove | Delete the two-example apposition and leave the pointer to §B entire +MINOR | high | §A1 clean predicate and §H "Gate-A clean-signal sentence" | §H states again that a pass carrying only Minors and no scope-stop trigger is clean, duplicating the clean predicate owned by §A1 | The Gate-A entry point can drift from the ordering on which nonzero-finding passes are clean | Keep §H's `NO FINDINGS` emission rule and cite §A1's clean-pass predicate entire for every other clean outcome +MINOR | high | §A1 "Composition, and what cannot happen" | The composition paragraph repeats the complete resuming-answer list just defined in "What a suspension asks"—accept or decline, a question decision and continue | The adjacent authorities can diverge on which answer discharges or parks a suspension | Define composition over every answer required by the preceding paragraph without enumerating the answers again +MINOR | high | §B "the floor, the Blocker/Major filter and the clean-final-pass rule all stand" | The absorb passage preserves a second three-item account of duties surviving a scope stop even though §A1 owns the cycle-general duty classification and closure conditions | A later duty change can leave the scope-stop entry point teaching an incomplete closure rule | Replace the enumeration with a pointer to §A1's cycle-general conditions and duties entire +MINOR | high | §§E and F item 3 "read before the ceiling" | §E defines the curve's pre-ceiling severity rule, while F item 3 defines the same rule again in the curve rationale | The curve can drift between two authorities on whether severity series are counted before or after demotion | Make the curve paragraph cite §E's observation rule and retain only why all three series are recorded +MINOR | high | §§E and F item 9b "severity buys no repair round and no further pass" | F item 9b repeats the Minor-or-below pass consequence already defined in §E instead of pointing to the severity rule | A future severity edit can leave the fallback issuing a conflicting repair or pass instruction | Point item 9b to Mechanics · Severity entire and retain only that a later set-change pass cost comes from §B +MINOR | high | §F items 2 and 8b downstream filtering | The Gate-A and Gate-B replacements independently define the same rule to filter to Blocker/Major only for repair while reading every line for all other predicates | The two gate entry points can drift on whether Minor and Nit lines reach cleanliness, scope, fix-set and health evaluation | State the shared reading rule once in the common findings protocol and have both gate instructions cite it entire +MINOR | medium | §F item 6 and §H Gate-A cadence | Item 6 says an unrevised Gate-A rerun is legitimate wherever no repair is owed, while §H separately decides when rerunning is allowed and requires a suspension's answer first | The broad-prompt entry point can be read as permission to rerun before the ordering has released a suspension | Keep rerun timing in §H and make item 6 cite that cadence while retaining only the broad-question requirement +MINOR | high | design §5 passage (c) and target §C | The passage map restates that the clearly-stuck third condition admits a re-raised validated dismissal, duplicating the operative rule in target §C | The design remains a second authority for which recurrence satisfies the clearly-stuck reading | Reduce the passage-map row to a pointer to target §C and keep only its accounting role +MINOR | high | design §8 "including the three exempted" and target §A1 | The prompt-standards discussion repeats three §A1 rules as a numbered set: no other pass outcome is eligible, zero findings are clean regardless of floor and decline is membership-only | The second list can go stale while still being offered as evidence that the target carries every required rationale | Cite the relevant §A1 rationale sentences without restating their rules or count +MINOR | high | design §9 "three contradictory reminder strings" | The out-of-scope bullet says three reminder strings are in scope while the design opening, §§7–8, story §2 and target §F enumerate four as items 10–13 | The design's own stated count disagrees with its enumeration and can cause one hook edit to be omitted during planning | Change "three" to "four" +END OF FINDINGS (18 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index 53862d6..ba1baa2 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -71,7 +71,8 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 52 | 06152d1 | 23→**29** | 0→**0** | 6→**5** | yes | one tell. **Two of five were pass 51's findings 1–2, still open because the pass-51 repair went to the wrong section** — see the pass-51 dispositions file. Pass 52 located them correctly at design §5 | | 53 | 81fffd1 | 29→**21** | 0→**0** | 5→**4** | yes | zero tells. **Three of four Majors are the first §F sites outside the two prompt copies** — reminder strings in `codex-gate.sh`. Daniel authorised a NARROW SCOPE OPENING on 2026-09-13: the three strings plus one exact-match test expectation are in, hook behaviour stays out, the Gate-B fingerprint overclaim stays parked. §F 14 → 17. Major 1 was a genuine internal contradiction (§A1 vs §A3 + §F-5 on what a failed attempt leaves behind); both assertions deleted | | 54 | c73367f | 21→**13** | 0→**0** | 4→**3** | yes | zero tells. All three Majors are pass-53 fallout and all are in-set: a **fourth** falsified hook string (the WIP reminder's "make the real commit when your final pass is clean") → §F item 13, 17→18; the design opening still said "No file under `plugins/dev-workflow/hooks/` changes"; §A1's continue branch re-enumerated §B's scope triggers instead of citing them | -| 55 | — | — | — | — | not run | next, against the pass-54 repair commit | +| 55 | 1940181 | 13→**18** | 0→**0** | 3→**6** | yes | one tell (findings rose). **Four more hook strings**, same class as item 13 — every gate reminder that tells the agent what to do next, ignoring a suspension the ordering sends that pass to. §F 18 → 22, covering **7 of the hook's 8** gate reminders. Also: the standing Named residual still said "Hook text is out of scope here by decision", which our own scope opening falsifies → item 14. §A1's source-block branch said a repaired block rereads "that pass" where the block stood before any pass existed | +| 56 | — | — | — | — | not run | next, against the pass-55 repair commit | | 40 | 12cf247 | 5→**3** | 0→**0** | 4→**3** | yes | zero tells; first round with no fan-out Major after the three-site check was run BEFORE the pass | | 41 | d26d097 | 3→**2** | 0→**0** | 3→**2** | yes | zero tells; both Majors were compressed pointers of mine dropping a load-bearing part of a standing rule | | 42 | 3b61fe3 | 2→**1** | 0→**0** | 2→**1** | yes | zero tells; §F item 4 had taken back the enumeration it was repaired to avoid | @@ -80,6 +81,34 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 45 | 1c858ac | 7→**7** | 0→**0** | 3→**1** | yes | zero tells; swept all four sites still tying classification to closing | | 46 | 4b713dd | 7→**9** | 0→**1** | 1→**2** | yes | **TWO TELLS — mandatory stop under §5, answered in advance by Daniel's standing "no stops without an absolute block" of 2026-09-13 17:02.** Tells surfaced here and in the commit body, findings kept open, no pass credited clean; the loop continued on that answer. **Eleventh** falsified standing sentence; session 01a09b7c-dd6b-7cb3-b83a-776599286f9b | +## Pass-55 report — one tell, and the scope opening's true size + +**Trend:** findings 29, 21, 13, **18** across 52–55; Blockers 0, 0, 0, **0**; Majors 5, 4, 3, **6**. +**Cluster:** product behaviour — five of six Majors are shipped strings, one is the ordering itself. +**Require↔withdraw:** none. **One tell** (the finding count rose); two would be a mandatory stop. + +**The number Daniel authorised was three; the real number is seven.** Pass 53 found three falsified +hook strings, pass 54 a fourth, pass 55 four more. Before absorbing them I counted the ceiling: +`codex-gate.sh` carries **eight** gate reminders, and §F now covers **seven** — the docs-only +notice is the one left alone, because it states no closure permission. So this is a bounded class +that has now been swept, not an open-ended expansion; the growth stops here unless the hook gains +a reminder. + +**The class, stated once:** a gate reminder that ends by telling the agent what to do next — +"Run Gate B now", "run more passes before executing", "run one more pass", "proceed if the skip +rule applies" — states a permission the closure ordering owns, at the one moment the ordering +exists to govern. Each is replaced with its diagnostic kept and the next action pointed at the +ordering. Items 10–13 and 15–17. + +**Item 14 is the one worth noticing.** The standing §5 text says "Hook text is out of scope here by +decision" — a sentence our own scope opening falsified on 2026-09-13. The change had made its own +documentation wrong and no pass had looked there. The narrow residual it was written for (the hook +reporting its threshold as an obligation at a floor of 1) stays out of scope. + +**Reading:** the rise from 13 to 18 is the scope opening being swept to its edge, not a plateau — +Blockers are flat at zero and every Major was in-set and repairable. If pass 56 rises again with +the hook class closed, that is a different signal and two tells. + ## Pass-54 report — zero tells, and the scope opening paid for itself immediately **Trend:** findings 23, 29, 21, **13** across 51–54; Blockers 0, 0, 0, **0**; Majors 6, 5, 4, **3**. diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index c3da555..7bda305 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -6,10 +6,10 @@ copy, and a value copied here would be a remembered value. Prompt-only, in two mirrored copies — `CLAUDE.md` §5 (**C** below) and the inline template in -`plugins/dev-workflow/commands/workflow-init.md` (**W** below) — **plus four reminder strings in +`plugins/dev-workflow/commands/workflow-init.md` (**W** below) — **plus seven reminder strings in `plugins/dev-workflow/hooks/codex-gate.sh` and the one exact-match expectation in `plugins/dev-workflow/hooks/codex-gate.test.sh` that pins one of them** (target text §F items -10–13, authorised 2026-09-13). **No hook behaviour changes**, and the hook strings have no mirror, +10–17, authorised 2026-09-13). **No hook behaviour changes**, and the hook strings have no mirror, so they carry no parity obligation. Condition ids `a1`…`j4` are defined in `docs/superpowers/specs/2026-09-10-loop-rule-consolidation-condition-inventory.md` beside this file: 135 conditions quoted from `7c0d475`, so a reader can check an accounting rather than @@ -221,8 +221,8 @@ commit (`scripts/check-version-bump.sh main` needs the committed bump, §8). **The check — what it must establish, and where it is built.** Every edit that changes a standing meaning owes a **discriminating pair of counts**: one showing the new wording present, one -showing the old wording gone. Each half runs in **both copies** — and, for the four hook -reminder strings of §F items 10–13, in the hook's **single** copy, which has no mirror and owes +showing the old wording gone. Each half runs in **both copies** — and, for the seven hook +reminder strings of §F items 10–13 and 15–17, in the hook's **single** copy, which has no mirror and owes no parity check; the hook suite's exact-match assertion is that edit's second observation — and against **both** the working tree and the parent tree, so every assertion is observed passing where the change exists and @@ -347,10 +347,10 @@ repo's most persistent defect. The transport that could carry it left with the r `plugins/dev-workflow/commands/workflow-init.md` is under `plugins/`, so `plugins/dev-workflow/.claude-plugin/plugin.json` goes `0.11.0 → 0.12.0` with a `plugins/dev-workflow/CHANGELOG.md` entry: a minor bump, the template gaining a closure - ordering, the edited or extended sentences §4 lists, and the four hook reminder strings of - §F items 10–13. The hook edits add no bump the template did not already require. + ordering, the edited or extended sentences §4 lists, and the seven hook reminder strings of + §F items 10–13 and 15–17. The hook edits add no bump the template did not already require. - **Invariant 4 / the hook.** `plugins/dev-workflow/hooks/codex-gate.sh` is edited in exactly - four `note` **strings** (§F items 10–13) and nowhere else: no control flow, no counter, no + seven `note` **strings** (§F items 10–13 and 15–17) and nowhere else: no control flow, no counter, no fingerprint computation, no routing, so the POSIX-`sh` and optional-`jq` obligations are not reached. `plugins/dev-workflow/hooks/codex-gate.test.sh` changes in its one exact-match expectation for the Gate-B satisfied message; the remaining hook assertions match loose @@ -429,7 +429,7 @@ not exist; §G names a sentence that does, and claims only what that sentence do - The fixture-per-predicate question (`fic2` record; story §2). - Hook **behaviour** under `plugins/dev-workflow/hooks/` — control flow, counters, fingerprint computation, routing, event handling. The three contradictory reminder strings and the one - exact-match test expectation are in scope instead (§F items 10–13, story §Out of scope), and + exact-match test expectation are in scope instead (§F items 10–13 and 15–17, story §Out of scope), and the Gate-B fingerprint overclaim in the same message stays parked here. - `todos.md`: both-branches-misread-each-other; self-consuming-deletion (prompt-standards item 11); the three bot findings in resolved plans. diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index ad54ffa..d1aa876 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -110,7 +110,10 @@ source rule that prescribes it; the list is examples and not the set** — **the must be repaired or answered, and no further pass runs while its block stands.** It is neither a suspension nor a continue and needs no name and no procedure of its own: the source rule carries both, and this ordering's part is to send the reader there rather than to run a pass over a cycle -another rule has stopped. **Once its source condition is repaired the ordering is read again on that pass**, no closing act +another rule has stopped. **Once its source condition is repaired**, a block that stood **before any pass of this cycle was +read** — the ones its floor derivation and its cited set raise among them — leaves the cycle to run +its next pass, there being no pass to read again; a block that stood **on a pass already read** has +**that pass read again through the ordering**, no closing act having been attempted on it — **and where that pass also carried a suspension, this branch releases only its own block**: the composition rule still holds the cycle on every answer that suspension asked for, a continue still leads to a pass run after the answer, and a stop still parks @@ -656,13 +659,15 @@ demotes it — the two counts are meant to differ. --- -## F. The eighteen standing sentences this change falsifies — REPLACED +## F. The twenty-two standing sentences this change falsifies — REPLACED -Each is a live sentence that the block makes wrong. All eighteen are **known contradictions** and -none is deferred. **Fourteen of them share one mechanism** — an entry point other than the ordering +Each is a live sentence that the block makes wrong. All twenty-two are **known contradictions** and +none is deferred. **Seventeen of them share one mechanism** — an entry point other than the ordering carrying an unqualified instruction — which is why each is **replaced** rather than given an -exception to point at. **Fourteen live in the two prompt copies and cite a line in each; the last -four live in the shipped hook**, which has one copy and no template mirror. +exception to point at. **Fifteen live in the two prompt copies and cite a line in each; the last +seven live in the shipped hook**, which has one copy and no template mirror. The hook holds eight +gate reminders in all; the one this change leaves alone is the docs-only notice, which states no +closure permission. **1. The `WIP:` naming warning** (Mechanics · `baseSha`). ``` @@ -870,8 +875,11 @@ repair: the **operation** stays here, the **permission** is the ordering's. **10. The Gate-A below-floor reminder's honesty claim** (the shipped hook, `plugins/dev-workflow/hooks/codex-gate.sh` — one copy, no template mirror). ``` -Gate A has no content check in this hook; what the gate itself requires of the reviewed artifact is stated in $policy and is instruction-backed. +Gate A has no content check in this hook; what the gate itself requires of the reviewed artifact is stated in $policy and is instruction-backed. What this cycle does next is that policy's closure ordering's, a further pass being one of its answers and not the only one. ``` +*Why (pass 55 finding 6):* the same reminder's tail reads "Run more passes before executing", +which is the same unqualified next action as items 15–17 and is replaced with them; a below-floor +pass carrying a suspension owes that suspension's answer first. *Why (pass 53 finding 19):* the live string says "Gate A has no content check behind it — this floor is the only thing keeping the spec review honest", which the Gate-A content condition falsifies: the floor is no longer the only thing, and a reader who believes it may treat the @@ -916,6 +924,46 @@ hold or an undischarged Major. It is the fourteenth sentence sharing this sectio the fourth found in the hook; the first three were found at pass 53 and this one at pass 54, which is why §I records the edit set as established by sweeping rather than by this file. +**14. The Named residual's blanket exemption** (the §5 loop rule, the named residual). It wraps +after "here" in each copy: C 139–140 and W 346–347. +``` +**That particular overstatement is out of scope here by decision, and it is not a blanket exemption for hook text** — a reminder this change's own rules falsify is corrected in the same change, as the standing sentences above require. +``` +*Why (pass 55 finding 2):* the live sentence reads "Hook text is out of scope here by decision", +which the narrow scope opening of 2026-09-13 falsifies: seven reminder strings are edited by this +change. Left standing, the installed text tells a reader the opposite of what the change did, and +a plan following it omits the authorised repairs. The residual it was written for — the hook +reporting its own threshold as an obligation at a floor of 1 — is untouched and stays out of +scope. This one is **not** an entry point carrying an unqualified instruction; it is a false +statement about scope, so it does not join that count. + +**15. The no-fingerprint reminder's next action** (the shipped hook, same file). +``` +What this cycle does next is $policy's closure ordering's — a pass, an answer a suspension is waiting for, or a source block's repair — and this reminder decides none of it. +``` +*Why (pass 55 finding 3):* the live string says "Run Gate B (mcp__codex__review) now", which sends +the author into another pass whatever the cycle's state is — including a pass that carried a +suspension whose answers are still outstanding, which the composition rule holds the cycle on. The +diagnostic half, and the machinery checks after it, are kept: they are what the message is for. + +**16. The stale-fingerprint reminder's two instructions** (the shipped hook, same file). +``` +A fresh Gate-B pass is the complete remedy for the staging and post-upgrade cases too, and when this cycle may run one is $policy's closure ordering's; where it may, that pass records a usable fingerprint. +``` +*Why (pass 55 finding 4):* the live string carries two unqualified imperatives — "Run Gate B +(mcp__codex__review) now" and "then run one more pass to record a usable fingerprint" — in the +branch an author reads immediately before committing. Both step past a suspension or a source +block that the ordering says is answered first. The long diagnostic list between them is untouched. + +**17. The Gate-B below-floor reminder's instruction** (the shipped hook, same file). +``` +Per $policy the review is a LOOP with a hard minimum of $floor passes, and what this cycle does next — a further pass, an answer, or a repair — is that policy's closure ordering's; $policy's skip rule decides only whether a cycle runs at all, never whether one already running may stop short. +``` +*Why (pass 55 finding 5):* the live string says "run more … or proceed only if $policy's skip rule +applies to this change", which offers the triviality skip as an exit from a running cycle. The +skip runs **no** passes and is decided before the cycle starts, so a below-floor cycle cannot +reach it; and "run more" alone ignores a suspension the ordering sends that pass to. + --- ## G. The one-contract paragraph — REPLACED @@ -1070,7 +1118,7 @@ be established as absent is cheaper to owe than to skip** — - **Minor, collected and open (pass 17 finding 9 is resolved above; this is the remainder):** nothing in this file establishes that the edit set is complete. It is the sites known at pass - 17, plus the four hook reminder strings §F items 10–13 added at passes 53 and 54 — which is itself + 17, plus the seven hook reminder strings §F items 10–13 and 15–17 added at passes 53, 54 and 55 — which is itself evidence that the set was not complete. **The plan sweeps both prompt copies and the shipped hook's reminder strings against the rule** — a live sentence the ordering falsifies gets edited — and prints what it found; a spec cannot establish that claim against text the same diff --git a/docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md b/docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md index 2d08417..f8dc0ef 100644 --- a/docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md +++ b/docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md @@ -45,10 +45,10 @@ which are preconditions; and what a user's answer on a surfaced finding does in - **Reopening any decision in §4.** They are settled and paid for; the design starts from them. - **Hook behaviour** (anything under `plugins/dev-workflow/hooks/` that decides what the hook does): its control flow, counters, fingerprint computation, routing and event handling, all - unchanged from the parent. **One narrow exception, authorised 2026-09-13:** the four reminder + unchanged from the parent. **One narrow exception, authorised 2026-09-13:** the seven reminder **strings** this change makes contradictory, and the single exact-match expectation in `plugins/dev-workflow/hooks/codex-gate.test.sh` that pins one of them, are in scope — listed as - items 10–13 of the target text's §F. Prompt-standards item 7 requires a superseded instruction + items 10–13 and 15–17 of the target text's §F. Prompt-standards item 7 requires a superseded instruction to be corrected in the same change, and a hook reminder is prompt text the agent acts on. The Gate-B fingerprint overclaim in the same message stays parked. - **The pass-counter anomaly**, the CodeRabbit plan-metadata contradiction, and the From bb1ba031506860e77c549fa4303751868588edc6 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 08:50:39 +0200 Subject: [PATCH 086/181] docs(specs): apply pass 56; two count contradictions, eleven collected MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Pass 56 is the cycle's first Blocker- and Major-free pass. Two of its thirteen Minors were repaired because they are stated counts disagreeing with their own enumerations, which the pass prompt asks for as mechanical failures: design §9 still read "three contradictory reminder strings" against the seven its own cited item ranges enumerate — introduced by the pass-55 propagation, which missed this spelling. §F item 7's rationale said "the first of the three closing paths" while §A2 defines two cases; replaced with the condition rather than a count. The remaining eleven are the second-copy class and are collected, not iterated. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-56.md | 14 ++++++++ .../gate-a-spec-awsf1ec771-resume.md | 32 ++++++++++++++++++- ...26-09-10-loop-rule-consolidation-design.md | 2 +- ...-10-loop-rule-consolidation-target-text.md | 7 ++-- 4 files changed, 50 insertions(+), 5 deletions(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-56.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-56.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-56.md new file mode 100644 index 0000000..cc911d0 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-56.md @@ -0,0 +1,14 @@ +MINOR | high | §A1 "Which severity field each predicate reads" | The text says Mechanics · Severity settles which severity field every finding-derived predicate reads, but subject clustering, scope triggers, and assigned-fix-set membership derive from findings without reading severity | A reader can infer a severity dependency for predicates that intentionally inspect findings independently of severity | Say Mechanics settles whether a predicate reads severity and, where it does, which severity field it reads +MINOR | high | §A1 continue branch and §B scope-trigger exceptions | After declaring §B the complete source of which findings raise triggers, §A1 restates that an already-declined finding and an already-answered question raise none | The second copy can drift from §B and recreate conflicting trigger answers for a later pass | Delete the exception apposition and cite §B's trigger rules entire +MINOR | high | §A1 clean predicate and §H "Gate-A clean-signal sentence" | §H states again that a pass carrying only Minors and no scope-stop trigger is clean, duplicating the clean-pass predicate owned by §A1 | The Gate-A entry point can drift from the ordering on which nonzero-finding passes are clean | Keep §H's emission rule for `NO FINDINGS` and cite §A1's clean-pass predicate entire for every other clean outcome +MINOR | high | §A1 "Composition, and what cannot happen" | The composition paragraph repeats the complete resuming-answer list defined in "What a suspension asks"—accept or decline, a question decision, and continue | The adjacent authorities can diverge on which answer discharges or parks a suspension | Define composition over every answer required by the preceding paragraph without enumerating the answers again +MINOR | high | §B "the floor, the Blocker/Major filter and the clean-final-pass rule all stand" | The absorb passage keeps a second three-item account of duties surviving a scope stop even though §A1 owns the cycle-general duty classification and closure conditions | A later duty change can leave the scope-stop entry point teaching an incomplete closure rule | Replace the enumeration with a pointer to §A1's cycle-general conditions and duties entire +MINOR | high | §§E and F item 3 "read before the ceiling" | §E defines the curve's pre-ceiling severity rule while §F item 3 defines the same rule again in the curve rationale | The curve can drift between two authorities on whether its severity series are counted before or after demotion | Make the curve paragraph cite §E's observation rule and retain only why all three series are recorded +MINOR | high | §§E and F item 9b "severity buys no repair round and no further pass" | §F item 9b repeats the Minor-or-below pass consequence already defined in §E instead of pointing to the severity rule | A later severity edit can leave the fallback issuing a conflicting repair or pass instruction | Point item 9b to Mechanics · Severity entire and retain only that a later set-change pass cost comes from §B +MINOR | high | §F items 2 and 8b "filter to Blocker/Major for what must be repaired" | The Gate-A and Gate-B replacements independently define the same downstream filter and every-line reading rule | The two gate entry points can drift on whether Minor and Nit lines reach cleanliness, scope, fix-set, and health evaluation | State the shared reading rule once in the common findings protocol and have both gate instructions cite it entire +MINOR | medium | §F item 6 and §H "Gate-A cadence" | Item 6 says an unrevised Gate-A rerun is legitimate wherever no repair is owed while §H separately permits rerunning only when the ordering selects continue and requires a suspension's answer first | A reader can take item 6 as permission to rerun an unrevised artifact while a suspension is still awaiting its answer | Keep rerun timing in §H and make item 6 cite that cadence while retaining only the broad-question requirement +MINOR | high | design §5 passage (c) and target §C | The passage map restates that the clearly-stuck third condition admits a re-raised validated dismissal, duplicating the operative rule in target §C | The design remains a second authority for which recurrence satisfies the clearly-stuck reading | Reduce the passage-map row to a pointer to target §C and keep only its accounting role +MINOR | high | design §8 "including the three exempted" and target §A1 | The prompt-standards discussion repeats three §A1 rules as a numbered set: no other pass outcome is eligible, zero findings are clean regardless of floor, and decline is membership-only | The second list can go stale while still being offered as evidence that the target carries every required rationale | Cite the relevant §A1 rationale sentences without restating their rules or count +MINOR | high | design §9 "three contradictory reminder strings" | The out-of-scope bullet says three reminder strings are in scope while the cited ranges §F items 10–13 and 15–17 enumerate seven and every other scope statement now says seven | The design's own count disagrees with its enumeration and can cause four authorized hook edits to be omitted during planning | Replace "three" with "seven" +MINOR | high | §F item 7 "the first of the three closing paths" | The rationale says Gate A has three closing paths while §A2 explicitly defines "Two cases" and enumerates only `HEAD` already carrying the text or the reviewed text remaining uncommitted | The target gives the plan two incompatible path counts and can make it invent or overlook a closing operation | Remove the path count and say the old destination fails whenever the artifact's existing commit is not the closing commit +END OF FINDINGS (13 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index ba1baa2..0f9b4fd 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -72,7 +72,8 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 53 | 81fffd1 | 29→**21** | 0→**0** | 5→**4** | yes | zero tells. **Three of four Majors are the first §F sites outside the two prompt copies** — reminder strings in `codex-gate.sh`. Daniel authorised a NARROW SCOPE OPENING on 2026-09-13: the three strings plus one exact-match test expectation are in, hook behaviour stays out, the Gate-B fingerprint overclaim stays parked. §F 14 → 17. Major 1 was a genuine internal contradiction (§A1 vs §A3 + §F-5 on what a failed attempt leaves behind); both assertions deleted | | 54 | c73367f | 21→**13** | 0→**0** | 4→**3** | yes | zero tells. All three Majors are pass-53 fallout and all are in-set: a **fourth** falsified hook string (the WIP reminder's "make the real commit when your final pass is clean") → §F item 13, 17→18; the design opening still said "No file under `plugins/dev-workflow/hooks/` changes"; §A1's continue branch re-enumerated §B's scope triggers instead of citing them | | 55 | 1940181 | 13→**18** | 0→**0** | 3→**6** | yes | one tell (findings rose). **Four more hook strings**, same class as item 13 — every gate reminder that tells the agent what to do next, ignoring a suspension the ordering sends that pass to. §F 18 → 22, covering **7 of the hook's 8** gate reminders. Also: the standing Named residual still said "Hook text is out of scope here by decision", which our own scope opening falsifies → item 14. §A1's source-block branch said a repaired block rereads "that pass" where the block stood before any pass existed | -| 56 | — | — | — | — | not run | next, against the pass-55 repair commit | +| 56 | 6b8854f | 18→**13** | 0→**0** | 6→**0** | yes | **first Blocker- and Major-free pass of the cycle.** 11 of 13 are the second-copy class, collected. Two were count contradictions and were repaired, so 56 is not the final pass: design §9 still said "three contradictory reminder strings" against its own enumeration of seven (my pass-55 propagation missed it), and §F item 7's rationale said "the three closing paths" while §A2 defines two cases | +| 57 | — | — | — | — | not run | next, against the pass-56 repair commit — a clean 57 closes the cycle | | 40 | 12cf247 | 5→**3** | 0→**0** | 4→**3** | yes | zero tells; first round with no fan-out Major after the three-site check was run BEFORE the pass | | 41 | d26d097 | 3→**2** | 0→**0** | 3→**2** | yes | zero tells; both Majors were compressed pointers of mine dropping a load-bearing part of a standing rule | | 42 | 3b61fe3 | 2→**1** | 0→**0** | 2→**1** | yes | zero tells; §F item 4 had taken back the enumeration it was repaired to avoid | @@ -81,6 +82,35 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 45 | 1c858ac | 7→**7** | 0→**0** | 3→**1** | yes | zero tells; swept all four sites still tying classification to closing | | 46 | 4b713dd | 7→**9** | 0→**1** | 1→**2** | yes | **TWO TELLS — mandatory stop under §5, answered in advance by Daniel's standing "no stops without an absolute block" of 2026-09-13 17:02.** Tells surfaced here and in the commit body, findings kept open, no pass credited clean; the loop continued on that answer. **Eleventh** falsified standing sentence; session 01a09b7c-dd6b-7cb3-b83a-776599286f9b | +## Pass-56 report — the first clean pass, and why it is not the final one + +**Trend:** findings 21, 13, 18, **13** across 53–56; Blockers 0, 0, 0, **0**; Majors 4, 3, 6, **0**. +**Cluster:** prose about the rules — 11 of 13 are the second-copy class. **Require↔withdraw:** none. +**Zero tells.** The hook-string class stayed closed: pass 56 found none, which is what a swept +class looks like. + +**Zero Blockers and zero Majors.** Under the standing filter that is a clean pass, and at pass 56 +the floor has long been met. + +**It is still not the final pass, because two of the thirteen were repaired.** Both are count +contradictions rather than judgement calls, and a stated count disagreeing with its own +enumeration is a mechanical failure the pass prompt asks for by name: + +1. **design §9 still read "three contradictory reminder strings"** while every other scope + statement, and the item range cited in the same sentence, said seven. **I introduced this** in + the pass-55 propagation — the script replaced every spelling I had thought of and this bullet + used another. A plan reading it would have made three of seven authorised hook edits. +2. **§F item 7's rationale said "the first of the three closing paths"** while §A2 defines **two** + cases. Replaced with the condition itself rather than a count, per the file's own preference for + removing an enumeration over correcting it. + +**The other eleven are collected and not iterated**, per Mechanics · Severity. They are one +subject: a rule stated both in §A1 and in the section that owns it. Each is a judgement about +which site should carry the sentence, and the cycle has been converging on that class for twenty +passes without it ever producing a Blocker. + +**So pass 57 runs against the repair, and a clean 57 closes the cycle.** + ## Pass-55 report — one tell, and the scope opening's true size **Trend:** findings 29, 21, 13, **18** across 52–55; Blockers 0, 0, 0, **0**; Majors 5, 4, 3, **6**. diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index 7bda305..15984c2 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -428,7 +428,7 @@ not exist; §G names a sentence that does, and claims only what that sentence do - The CodeRabbit plan-metadata contradiction (`fic2` record). - The fixture-per-predicate question (`fic2` record; story §2). - Hook **behaviour** under `plugins/dev-workflow/hooks/` — control flow, counters, fingerprint - computation, routing, event handling. The three contradictory reminder strings and the one + computation, routing, event handling. The seven contradictory reminder strings and the one exact-match test expectation are in scope instead (§F items 10–13 and 15–17, story §Out of scope), and the Gate-B fingerprint overclaim in the same message stays parked here. - `todos.md`: both-branches-misread-each-other; self-consuming-deletion (prompt-standards item diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index d1aa876..9a0075b 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -776,9 +776,10 @@ closing act uses; a Gate-B cycle in the WIP commit, restated by the commit its c produces. ``` *Why (pass 22 finding 3):* the live clause sends a Gate-A cycle's exception record to "the spec or -plan commit", which is the closing commit only on the first of the three closing paths. On the -other two the record and the closure would land in different commits. Naming the closing act -instead keeps them together on all three without changing the record's form or force. +plan commit", which is the closing commit only where the artifact's existing commit **is** the +closing commit. Wherever it is not, the record and the closure would land in different commits. +Naming the closing act instead keeps them together in every case, without changing the record's +form or force. **8. The profile-change paragraph's pass claim** (the profiles section). It wraps across C 752–753 and W 938–939. From ebe3aa7cbe7bb7a1a76456bf19b9e91fc18d2b5f Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 09:04:53 +0200 Subject: [PATCH 087/181] docs(specs): apply pass 57; three pinned expectations, not one MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit codex-gate.test.sh pins three of the seven edited reminder strings with exact-match expectations (its three expected_ctx assignments), not one. My earlier check grepped for phrases I already knew were falsified, and two of the three expectations contain none of them. §F items 12, 15 and 16 now each carry the obligation and the §F opening states the count once; design, story and the pass prompt corrected with it. Findings 1-3 tightened the pass-55 hook replacements: item 13 still opened with an unqualified "Run the review against this commit"; item 15 enumerated the ordering's routes and omitted the closing act, which would tell a clean eligible Gate-B cycle with an unstorable fingerprint it had no commit route; and item 15's sentence tail "delete it and run a fresh pass" was left standing. The enumeration is removed rather than extended. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-57.md | 6 ++++ .../gate-a-spec-awsf1ec771-resume.md | 29 ++++++++++++++++++- ...26-09-10-loop-rule-consolidation-design.md | 11 +++---- ...-10-loop-rule-consolidation-target-text.md | 21 ++++++++++---- ...026-08-29-loop-rule-consolidation-story.md | 4 +-- 5 files changed, 57 insertions(+), 14 deletions(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-57.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-57.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-57.md new file mode 100644 index 0000000..4e8e88a --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-57.md @@ -0,0 +1,6 @@ +MAJOR | high | §F item 13 "Run the review against this commit" | The proposed WIP reminder retains this unqualified imperative even though §A1's source-block branch says a profile, header, or evidence block detected before the first pass forbids any pass until its source condition is repaired | An author following the reminder can spend a Gate-B pass while a source block stands, giving the same pre-pass state two incompatible next actions | Make the review instruction conditional on the closure ordering releasing the cycle to a pass, while retaining this commit as the range to use when it does +MAJOR | high | §F item 15 "a pass, an answer a suspension is waiting for, or a source block's repair" | This proposed enumeration omits the closing act, although an absent hook fingerprint is not a closure condition and §A3 deliberately gives Gate B no content condition, so a valid clean eligible cycle whose fingerprint merely could not be stored can proceed to its closing act | The hook reminder can make an otherwise closable Gate-B cycle appear to have no commit route and force an unowed pass | Remove the enumeration and cite the closure ordering entire so its closing-act route remains available +MAJOR | high | §F item 15 "delete it and run a fresh pass" | The live no-fingerprint reminder contains this second pass command later in the same sentence, and item 15 says the machinery checks after the first command are kept without replacing this one | Even after the new pointer says the reminder decides no next action, its retained tail can still run a pass past an outstanding suspension answer or source repair | Keep deletion as a diagnostic remedy but make any subsequent pass conditional on the closure ordering selecting one +MAJOR | high | §F item 16 "The stale-fingerprint reminder's two instructions" | `plugins/dev-workflow/hooks/codex-gate.test.sh` line 1006 repeats the full stale reminder in `expected_ctx`, and line 1008 compares it with exact equality, but the target names only item 12's exact-match expectation for replacement | Installing item 16's new wording while following the stated edit set leaves the hook regression suite failing | Add the stale `expected_ctx` exact fixture to item 16's replacement scope and update it to the complete resulting reminder +MAJOR | high | §F item 15 "The no-fingerprint reminder's next action" | `plugins/dev-workflow/hooks/codex-gate.test.sh` line 1029 repeats the full no-fingerprint reminder in `expected_ctx`, and line 1031 compares it with exact equality, but the target names only item 12's exact-match expectation for replacement | Installing item 15's new wording while following the stated edit set leaves the hook regression suite failing | Add the no-fingerprint `expected_ctx` exact fixture to item 15's replacement scope and update it to the complete resulting reminder +END OF FINDINGS (5 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index 0f9b4fd..9557143 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -73,7 +73,8 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 54 | c73367f | 21→**13** | 0→**0** | 4→**3** | yes | zero tells. All three Majors are pass-53 fallout and all are in-set: a **fourth** falsified hook string (the WIP reminder's "make the real commit when your final pass is clean") → §F item 13, 17→18; the design opening still said "No file under `plugins/dev-workflow/hooks/` changes"; §A1's continue branch re-enumerated §B's scope triggers instead of citing them | | 55 | 1940181 | 13→**18** | 0→**0** | 3→**6** | yes | one tell (findings rose). **Four more hook strings**, same class as item 13 — every gate reminder that tells the agent what to do next, ignoring a suspension the ordering sends that pass to. §F 18 → 22, covering **7 of the hook's 8** gate reminders. Also: the standing Named residual still said "Hook text is out of scope here by decision", which our own scope opening falsifies → item 14. §A1's source-block branch said a repaired block rereads "that pass" where the block stood before any pass existed | | 56 | 6b8854f | 18→**13** | 0→**0** | 6→**0** | yes | **first Blocker- and Major-free pass of the cycle.** 11 of 13 are the second-copy class, collected. Two were count contradictions and were repaired, so 56 is not the final pass: design §9 still said "three contradictory reminder strings" against its own enumeration of seven (my pass-55 propagation missed it), and §F item 7's rationale said "the three closing paths" while §A2 defines two cases | -| 57 | — | — | — | — | not run | next, against the pass-56 repair commit — a clean 57 closes the cycle | +| 57 | bb1ba03 | 13→**5** | 0→**0** | 0→**5** | yes | zero tells. All five are pass-53/55 fallout inside the hook items. **Two are a factual error of mine:** `codex-gate.test.sh` has **three** `expected_ctx` exact-match expectations (lines 1006, 1019, 1029), not one — my first grep searched only for phrases that occur in 1019. Items 12, 15 and 16 are all pinned. The other three tightened items 13 and 15, whose replacements still carried an unqualified imperative and an enumeration omitting the closing act | +| 58 | — | — | — | — | not run | next, against the pass-57 repair commit | | 40 | 12cf247 | 5→**3** | 0→**0** | 4→**3** | yes | zero tells; first round with no fan-out Major after the three-site check was run BEFORE the pass | | 41 | d26d097 | 3→**2** | 0→**0** | 3→**2** | yes | zero tells; both Majors were compressed pointers of mine dropping a load-bearing part of a standing rule | | 42 | 3b61fe3 | 2→**1** | 0→**0** | 2→**1** | yes | zero tells; §F item 4 had taken back the enumeration it was repaired to avoid | @@ -82,6 +83,32 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 45 | 1c858ac | 7→**7** | 0→**0** | 3→**1** | yes | zero tells; swept all four sites still tying classification to closing | | 46 | 4b713dd | 7→**9** | 0→**1** | 1→**2** | yes | **TWO TELLS — mandatory stop under §5, answered in advance by Daniel's standing "no stops without an absolute block" of 2026-09-13 17:02.** Tells surfaced here and in the commit body, findings kept open, no pass credited clean; the loop continued on that answer. **Eleventh** falsified standing sentence; session 01a09b7c-dd6b-7cb3-b83a-776599286f9b | +## Pass-57 report — zero tells, and a fact I had wrong twice + +**Trend:** findings 18, 13, 13, **5** across 54–57; Blockers 0, 0, 0, **0**; Majors 6, 0, 0... **5**. +**Cluster:** product behaviour, all five inside the hook items §F gained at passes 53–55. +**Require↔withdraw:** none. Zero tells — but note the Major count is not monotone, and the reason +is that pass 56's repair set was small while pass 57 read the hook items against the test suite +for the first time. + +**The error worth recording.** I told Daniel, and wrote into three files, that +`plugins/dev-workflow/hooks/codex-gate.test.sh` pins **one** of the edited reminder strings with an +exact-match expectation. It pins **three** — lines 1006, 1019 and 1029, all `expected_ctx`. My +check had been `grep` for the three phrases I already knew were falsified, and two of the three +expectations contain none of them. The lesson is the one AGENTS.md already states for gate claims +and which I did not apply to a test file: **name the exact thing the check performs** — here, +"every `expected_ctx` assignment" — rather than grepping for the strings I expected to find. +Items 12, 15 and 16 now all carry the obligation, and the §F opening states the count once. + +**The other three** tightened the hook replacements pass 55 had written: item 13 still opened +"Run the review against this commit" unqualified; item 15's replacement enumerated the ordering's +routes and **omitted the closing act**, which would tell a clean eligible Gate-B cycle with an +unstorable fingerprint that it had no commit route; and item 15 had left the same sentence's tail, +"delete it and run a fresh pass", standing. The enumeration was removed rather than extended. + +**Reading:** 5 findings from 13, all in-set, all repaired, no new subject. The hook class is still +closed — nothing in pass 57 named an eighth reminder. + ## Pass-56 report — the first clean pass, and why it is not the final one **Trend:** findings 21, 13, 18, **13** across 53–56; Blockers 0, 0, 0, **0**; Majors 4, 3, 6, **0**. diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index 15984c2..f4d5666 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -7,8 +7,8 @@ copy, and a value copied here would be a remembered value. Prompt-only, in two mirrored copies — `CLAUDE.md` §5 (**C** below) and the inline template in `plugins/dev-workflow/commands/workflow-init.md` (**W** below) — **plus seven reminder strings in -`plugins/dev-workflow/hooks/codex-gate.sh` and the one exact-match expectation in -`plugins/dev-workflow/hooks/codex-gate.test.sh` that pins one of them** (target text §F items +`plugins/dev-workflow/hooks/codex-gate.sh` and the three exact-match expectations in +`plugins/dev-workflow/hooks/codex-gate.test.sh` that pin three of them** (target text §F items 10–17, authorised 2026-09-13). **No hook behaviour changes**, and the hook strings have no mirror, so they carry no parity obligation. Condition ids `a1`…`j4` are defined in `docs/superpowers/specs/2026-09-10-loop-rule-consolidation-condition-inventory.md` beside this @@ -352,9 +352,10 @@ repo's most persistent defect. The transport that could carry it left with the r - **Invariant 4 / the hook.** `plugins/dev-workflow/hooks/codex-gate.sh` is edited in exactly seven `note` **strings** (§F items 10–13 and 15–17) and nowhere else: no control flow, no counter, no fingerprint computation, no routing, so the POSIX-`sh` and optional-`jq` obligations are not - reached. `plugins/dev-workflow/hooks/codex-gate.test.sh` changes in its one exact-match - expectation for the Gate-B satisfied message; the remaining hook assertions match loose - patterns these repairs leave standing. The §5 heading the hook greps (`Cross-Model Review`) + reached. `plugins/dev-workflow/hooks/codex-gate.test.sh` changes in all three of its + `expected_ctx` exact-match expectations — the Gate-B satisfied, stale-fingerprint and + no-fingerprint messages, §F items 12, 16 and 15 — each replaced with the complete resulting + message; the remaining hook assertions match loose patterns these repairs leave standing. The §5 heading the hook greps (`Cross-Model Review`) does not move. **Every path in this spec is written repository-relative and in full** — no ellipsis shorthand diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index 9a0075b..41838c3 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -667,7 +667,10 @@ carrying an unqualified instruction — which is why each is **replaced** rather exception to point at. **Fifteen live in the two prompt copies and cite a line in each; the last seven live in the shipped hook**, which has one copy and no template mirror. The hook holds eight gate reminders in all; the one this change leaves alone is the docs-only notice, which states no -closure permission. +closure permission. **Three of the seven are pinned by exact-match expectations in +`plugins/dev-workflow/hooks/codex-gate.test.sh`** — items 12, 15 and 16, at that file's three +`expected_ctx` assignments — and each is replaced there with its complete resulting message in the +same change. The remaining four are matched by loose patterns these repairs leave standing. **1. The `WIP:` naming warning** (Mechanics · `baseSha`). ``` @@ -909,13 +912,12 @@ Per $policy, commit only if your final pass was clean and every other closure co ``` *Why (pass 53 finding 20):* the same abbreviated definition, in the branch an author reads immediately before committing, and it takes the same repair. It is the thirteenth sentence sharing -this section's mechanism. **The exact-match expectation in -`plugins/dev-workflow/hooks/codex-gate.test.sh` pins this string in full and is replaced with -it** — the other hook assertions match on loose patterns this repair leaves standing. +this section's mechanism. This message is one of the three pinned by an exact-match expectation, +replaced there with the complete resulting message as the section opening requires. **13. The WIP-commit reminder's closing permission** (the shipped hook, same file). ``` -Run the review against this commit; when the closing act may then be performed is $policy's closure ordering's, read there in full. +Use this commit as the review range; whether this cycle runs a review now, and when its closing act may be performed, are both $policy's closure ordering's, read there in full. ``` *Why (pass 54 finding 2):* the live string ends "then make the real commit when your final pass is clean", which makes cleanliness the whole permission — and it reaches the author at the one moment @@ -940,12 +942,19 @@ statement about scope, so it does not join that count. **15. The no-fingerprint reminder's next action** (the shipped hook, same file). ``` -What this cycle does next is $policy's closure ordering's — a pass, an answer a suspension is waiting for, or a source block's repair — and this reminder decides none of it. +What this cycle does next is $policy's closure ordering's, read there entire, and this reminder decides none of it — including whether the state file's deletion below is followed by a pass. ``` *Why (pass 55 finding 3):* the live string says "Run Gate B (mcp__codex__review) now", which sends the author into another pass whatever the cycle's state is — including a pass that carried a suspension whose answers are still outstanding, which the composition rule holds the cycle on. The diagnostic half, and the machinery checks after it, are kept: they are what the message is for. +*And (pass 57 findings 2 and 3):* an earlier wording enumerated the ordering's routes as a pass, a +suspension's answer or a source block's repair, which **omits the closing act** — an absent +fingerprint is not a closure condition and Gate B has no content condition, so a clean eligible +cycle whose fingerprint merely could not be stored still has its commit route. The enumeration is +removed rather than extended, per this section's own preference. The same finding's second half +reaches the sentence's tail, "delete it and run a fresh pass": the deletion stays as a remedy, the +pass after it does not, and the new sentence says so where a reader meets it. **16. The stale-fingerprint reminder's two instructions** (the shipped hook, same file). ``` diff --git a/docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md b/docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md index f8dc0ef..428b301 100644 --- a/docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md +++ b/docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md @@ -46,8 +46,8 @@ which are preconditions; and what a user's answer on a surfaced finding does in - **Hook behaviour** (anything under `plugins/dev-workflow/hooks/` that decides what the hook does): its control flow, counters, fingerprint computation, routing and event handling, all unchanged from the parent. **One narrow exception, authorised 2026-09-13:** the seven reminder - **strings** this change makes contradictory, and the single exact-match expectation in - `plugins/dev-workflow/hooks/codex-gate.test.sh` that pins one of them, are in scope — listed as + **strings** this change makes contradictory, and the three exact-match expectations in + `plugins/dev-workflow/hooks/codex-gate.test.sh` that pin three of them, are in scope — listed as items 10–13 and 15–17 of the target text's §F. Prompt-standards item 7 requires a superseded instruction to be corrected in the same change, and a hook reminder is prompt text the agent acts on. The Gate-B fingerprint overclaim in the same message stays parked. From fecb3aaea4924565ebe2dd9fc6e122c198b4990f Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 09:21:35 +0200 Subject: [PATCH 088/181] docs(specs): apply pass 58; the hook messages' openings, one layer deeper MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Both fingerprint reminders still opened "STOP — Codex Gate B not satisfied", which states a gate verdict the hook cannot reach: its fingerprint is advisory and is not a closure condition, so a clean eligible cycle whose fingerprint merely could not be stored was told the gate was unsatisfied. Both openings now report the hook's own state and leave every verdict to the ordering — the AGENTS.md rule about naming the exact comparison, applied to the hook's words. Item 16 also claimed a permitted pass "records" a usable fingerprint, which no rule makes true where the store just failed; it now says attempts, and names the loop that would otherwise send an author round again. Item 17's enumeration omitted the closing act a zero-finding pass reaches below the floor, the one closing route that branch exists to describe; removed rather than extended. Design §9's bullet said "one exact-match test expectation" — the second time that same bullet has gone stale under a scripted propagation. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-58.md | 7 ++++ .../gate-a-spec-awsf1ec771-resume.md | 31 ++++++++++++++++- ...26-09-10-loop-rule-consolidation-design.md | 4 +-- ...-10-loop-rule-consolidation-target-text.md | 34 ++++++++++++++++--- 4 files changed, 68 insertions(+), 8 deletions(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-58.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-58.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-58.md new file mode 100644 index 0000000..d1e1231 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-58.md @@ -0,0 +1,7 @@ +MAJOR | high | §F item 15 "no-fingerprint reminder" | The replacement leaves the live opening "STOP — Codex Gate B not satisfied: no fingerprint is recorded for this cycle" untouched even though §A3 gives Gate B no content condition and item 15 says this reminder decides no next action | A clean eligible cycle whose fingerprint could not be stored receives a contradictory stop/unsatisfied instruction and can be driven into needless passes or an endless machinery loop | Replace the opening with a neutral hook-state diagnosis and leave closure permission entirely to the ordering pointer +MAJOR | high | §F item 16 "stale-fingerprint reminder" | The replacement leaves the live opening "STOP — Codex Gate B not satisfied" untouched although this branch expressly includes staging, upgrade, computation and storage states that need not violate any closure condition | An otherwise closable Gate-B cycle is still told that the gate is unsatisfied, contradicting §A3 and the replacement's deference to the ordering | Replace only the opening permission claim with a neutral hook-state warning, preserving the parked fingerprint diagnosis and leaving permission to the ordering +MAJOR | high | §F item 16 "records a usable fingerprint" | The proposed sentence says that, where a fresh pass may run, "that pass records a usable fingerprint", but the same branch includes a fingerprint that could not be stored and no rule makes storage succeed on the next pass | A permitted fresh pass can return to the same stale state, so the reminder overclaims its mechanism and can send the agent around another unproductive pass | Say that a fresh pass attempts to record a usable fingerprint, or condition the result explicitly on a successful store +MAJOR | high | §F item 17 "a further pass, an answer, or a repair" | The proposed below-floor reminder presents those three outcomes as the ordering's next-action list, but a zero-finding pass is eligible below the floor and can proceed to the closing act | The only below-floor closing route is omitted at the reminder an author sees before committing, so the hook text can prevent the early exit the ordering preserves | Remove the enumeration and cite the closure ordering entire, as item 15 already does +MINOR | high | §A1 "The profile and the assigned fix set answer a change" | §A1 defines that a change, including one later undone, costs a pass while §B and the standing profile-change paragraph own those same rules, contradicting §A1's claim that source-owned preconditions are only cited here | The installed prompt has multiple authorities for the same pass-cost semantics and a later repair can make them disagree | Delete the source-rule semantics from §A1 and cite the profile and assigned-fix-set sources entire +MINOR | high | design §9 "one exact-match test expectation" | The design's parked-scope bullet says one exact-match hook-test expectation is in scope, while design §1, design §8, target §F and the three `expected_ctx` assignments on disk all say three | A plan following §9 can update only one golden expectation, leaving two stale and failing the hook suite | Replace "one" with "three" in design §9 +END OF FINDINGS (6 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index 9557143..e810a59 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -74,7 +74,8 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 55 | 1940181 | 13→**18** | 0→**0** | 3→**6** | yes | one tell (findings rose). **Four more hook strings**, same class as item 13 — every gate reminder that tells the agent what to do next, ignoring a suspension the ordering sends that pass to. §F 18 → 22, covering **7 of the hook's 8** gate reminders. Also: the standing Named residual still said "Hook text is out of scope here by decision", which our own scope opening falsifies → item 14. §A1's source-block branch said a repaired block rereads "that pass" where the block stood before any pass existed | | 56 | 6b8854f | 18→**13** | 0→**0** | 6→**0** | yes | **first Blocker- and Major-free pass of the cycle.** 11 of 13 are the second-copy class, collected. Two were count contradictions and were repaired, so 56 is not the final pass: design §9 still said "three contradictory reminder strings" against its own enumeration of seven (my pass-55 propagation missed it), and §F item 7's rationale said "the three closing paths" while §A2 defines two cases | | 57 | bb1ba03 | 13→**5** | 0→**0** | 0→**5** | yes | zero tells. All five are pass-53/55 fallout inside the hook items. **Two are a factual error of mine:** `codex-gate.test.sh` has **three** `expected_ctx` exact-match expectations (lines 1006, 1019, 1029), not one — my first grep searched only for phrases that occur in 1019. Items 12, 15 and 16 are all pinned. The other three tightened items 13 and 15, whose replacements still carried an unqualified imperative and an enumeration omitting the closing act | -| 58 | — | — | — | — | not run | next, against the pass-57 repair commit | +| 58 | ebe3aa7 | 5→**6** | 0→**0** | 5→**4** | yes | one tell (findings rose 5→6). All four Majors in items 15–17, one layer deeper than pass 57 reached: both fingerprint reminders still **opened** "STOP — Codex Gate B not satisfied", a gate verdict the hook cannot reach; item 16 claimed a permitted pass "records" a usable fingerprint where the store can fail again; item 17's enumeration omitted the below-floor zero-finding closing route. Minor 6 was design §9's "one exact-match expectation" — the **second** time that bullet went stale under a propagation | +| 59 | — | — | — | — | not run | next, against the pass-58 repair commit | | 40 | 12cf247 | 5→**3** | 0→**0** | 4→**3** | yes | zero tells; first round with no fan-out Major after the three-site check was run BEFORE the pass | | 41 | d26d097 | 3→**2** | 0→**0** | 3→**2** | yes | zero tells; both Majors were compressed pointers of mine dropping a load-bearing part of a standing rule | | 42 | 3b61fe3 | 2→**1** | 0→**0** | 2→**1** | yes | zero tells; §F item 4 had taken back the enumeration it was repaired to avoid | @@ -83,6 +84,34 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 45 | 1c858ac | 7→**7** | 0→**0** | 3→**1** | yes | zero tells; swept all four sites still tying classification to closing | | 46 | 4b713dd | 7→**9** | 0→**1** | 1→**2** | yes | **TWO TELLS — mandatory stop under §5, answered in advance by Daniel's standing "no stops without an absolute block" of 2026-09-13 17:02.** Tells surfaced here and in the commit body, findings kept open, no pass credited clean; the loop continued on that answer. **Eleventh** falsified standing sentence; session 01a09b7c-dd6b-7cb3-b83a-776599286f9b | +## Pass-58 report — one tell, and the shape of the hook-item lineage + +**Trend:** findings 13, 5, **6** across 56–58; Blockers 0, 0, **0**; Majors 0, 5, **4**. +**Cluster:** product behaviour, all four Majors inside §F items 15–17. **Require↔withdraw:** none — +passes 57 and 58 each *removed* an enumeration pass 55 wrote; no pass demanded back what an earlier +one removed. **One tell:** the finding count rose 5 → 6. + +**The lineage, stated plainly, because it is the thing to watch.** Passes 53–55 added seven hook +items. Pass 57 found problems in them, pass 58 found more. Each round's fix produced the next +finding, which is the third condition of the clearly-stuck reading. + +**It is not a plateau, and here is the test that distinguishes them:** each pass reached an aspect +no earlier pass had examined, rather than re-litigating one. Pass 55 wrote the replacements; pass +57 was the first to read them against `codex-gate.test.sh` and found three pinned expectations +where the text claimed one; pass 58 was the first to read the *openings* of the messages rather +than the instructions inside them, and found "STOP — Codex Gate B not satisfied" stating a verdict +the hook cannot reach. That is coverage arriving in layers. + +**The clearly-stuck exit therefore does not apply**, and the reason is its second condition: I +cannot state an affirmative judgement that coverage is sufficient while each pass is still +reaching a first-time layer of the same seven items. When a pass reaches no new layer, that +judgement becomes available and this reading should be taken again. + +**Design §9 went stale a second time.** Pass 56 caught "three contradictory reminder strings" in +that bullet; pass 58 caught "one exact-match test expectation" in the same bullet. Both times a +scripted propagation had updated every other site. The bullet is worth reading by hand whenever +the hook scope moves. + ## Pass-57 report — zero tells, and a fact I had wrong twice **Trend:** findings 18, 13, 13, **5** across 54–57; Blockers 0, 0, 0, **0**; Majors 6, 0, 0... **5**. diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index f4d5666..8509ac8 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -429,8 +429,8 @@ not exist; §G names a sentence that does, and claims only what that sentence do - The CodeRabbit plan-metadata contradiction (`fic2` record). - The fixture-per-predicate question (`fic2` record; story §2). - Hook **behaviour** under `plugins/dev-workflow/hooks/` — control flow, counters, fingerprint - computation, routing, event handling. The seven contradictory reminder strings and the one - exact-match test expectation are in scope instead (§F items 10–13 and 15–17, story §Out of scope), and + computation, routing, event handling. The seven contradictory reminder strings and the three + exact-match test expectations are in scope instead (§F items 10–13 and 15–17, story §Out of scope), and the Gate-B fingerprint overclaim in the same message stays parked here. - `todos.md`: both-branches-misread-each-other; self-consuming-deletion (prompt-standards item 11); the three bot findings in resolved plans. diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index 41838c3..0e4968d 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -940,9 +940,11 @@ reporting its own threshold as an obligation at a floor of 1 — is untouched an scope. This one is **not** an entry point carrying an unqualified instruction; it is a false statement about scope, so it does not join that count. -**15. The no-fingerprint reminder's next action** (the shipped hook, same file). +**15. The no-fingerprint reminder's opening and its next action** (the shipped hook, same file). +Replaced span: the opening through "Run Gate B (mcp__codex__review) now;", the machinery checks +after it untouched. ``` -What this cycle does next is $policy's closure ordering's, read there entire, and this reminder decides none of it — including whether the state file's deletion below is followed by a pass. +Codex gate state: no fingerprint is recorded for this cycle — either no mcp__codex__review has run, or the last one's fingerprint could not be written or read back. Per $policy you MUST reach a minimum of $floor passes per cycle. What this cycle does next is $policy's closure ordering's, read there entire, and this reminder decides none of it — including whether the state file's deletion below is followed by a pass; ``` *Why (pass 55 finding 3):* the live string says "Run Gate B (mcp__codex__review) now", which sends the author into another pass whatever the cycle's state is — including a pass that carried a @@ -955,24 +957,46 @@ cycle whose fingerprint merely could not be stored still has its commit route. T removed rather than extended, per this section's own preference. The same finding's second half reaches the sentence's tail, "delete it and run a fresh pass": the deletion stays as a remedy, the pass after it does not, and the new sentence says so where a reader meets it. +*And (pass 58 finding 1):* the opening still read "STOP — Codex Gate B not satisfied", which +states a gate verdict the hook cannot reach — the fingerprint is advisory and is not a closure +condition, so a clean eligible cycle whose fingerprint merely could not be stored was being told +the gate was unsatisfied. The opening now reports the **hook's own state**, which is what it +observes, and leaves every verdict to the ordering. This is the AGENTS.md rule about naming the +exact comparison a mechanism performs, applied to the hook's own words. -**16. The stale-fingerprint reminder's two instructions** (the shipped hook, same file). +**16. The stale-fingerprint reminder's opening and its two instructions** (the shipped hook, same file). +**Two non-adjacent sentences of one message**, the diagnostic list between them untouched. The +opening: ``` -A fresh Gate-B pass is the complete remedy for the staging and post-upgrade cases too, and when this cycle may run one is $policy's closure ordering's; where it may, that pass records a usable fingerprint. +Codex gate state: the hook cannot confirm that the content you are about to commit is the content mcp__codex__review last saw ($passes recorded pass(es) this cycle). +``` +and, in place of the two imperatives: +``` +A fresh Gate-B pass is the complete remedy for the staging and post-upgrade cases too, and when this cycle may run one is $policy's closure ordering's; where it may, that pass attempts to record a usable fingerprint — a store that fails again returns this same state, which is a fault in the machinery and not a verdict on the cycle. ``` *Why (pass 55 finding 4):* the live string carries two unqualified imperatives — "Run Gate B (mcp__codex__review) now" and "then run one more pass to record a usable fingerprint" — in the branch an author reads immediately before committing. Both step past a suspension or a source block that the ordering says is answered first. The long diagnostic list between them is untouched. +*And (pass 58 findings 2 and 3):* the opening read "STOP — Codex Gate B not satisfied" over a +branch whose own text lists staging, a hook upgrade and a failed store — none of which violates a +closure condition — so it takes item 15's repair, reporting the hook's state and leaving the +verdict to the ordering. The second sentence also said a permitted pass "records a usable +fingerprint", which no rule makes true: the store that just failed can fail again. It now says +**attempts**, and names the loop that would otherwise send an author round again. **17. The Gate-B below-floor reminder's instruction** (the shipped hook, same file). ``` -Per $policy the review is a LOOP with a hard minimum of $floor passes, and what this cycle does next — a further pass, an answer, or a repair — is that policy's closure ordering's; $policy's skip rule decides only whether a cycle runs at all, never whether one already running may stop short. +Per $policy the review is a LOOP with a hard minimum of $floor passes, and what this cycle does next is that policy's closure ordering's, read there entire; $policy's skip rule decides only whether a cycle runs at all, never whether one already running may stop short. ``` *Why (pass 55 finding 5):* the live string says "run more … or proceed only if $policy's skip rule applies to this change", which offers the triviality skip as an exit from a running cycle. The skip runs **no** passes and is decided before the cycle starts, so a below-floor cycle cannot reach it; and "run more" alone ignores a suspension the ordering sends that pass to. +*And (pass 58 finding 4):* an earlier wording enumerated the ordering's routes as a pass, an answer +or a repair, which **omits the closing act a zero-finding pass reaches below the floor** — the one +closing route this branch exists to describe. Removed rather than extended, exactly as item 15's +was. --- From db9f70ec69c89ca4a9d2aef320d44abd67c25384 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 09:40:51 +0200 Subject: [PATCH 089/181] =?UTF-8?q?docs(specs):=20apply=20pass=2059;=20?= =?UTF-8?q?=C2=A7F's=20hook=20items=20carry=20complete=20messages=20now?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Two loop-health tells were present and the stop was mandatory: the finding count rose 5 -> 6 -> 9, and four of nine findings were about §F's prose describing its replacements rather than about the replacements. Surfaced; Daniel's standing answer applied and the loop continued. The tells pointed at the framing. §F's hook items were written as a span plus a sentence saying where it goes, and three passes running found a new gap in that: a tail standing behind a pointer, a fragment that is not independently grammatical, a diagnostic preservable only by inventing wording. Items 15 and 16 now give the complete resulting message for both the additionalContext and the systemMessage channel, so a plan installs what is written. Also: the parity rule claimed both prompt copies take every REPLACED section byte-identical, which is false of seven hook-only items; and the design opening wrote the hook items as the contiguous range 10-17, swallowing item 14, which is a prompt-copy sentence. Finding 7 is open and blocks pass 60: conflicting accept/decline on the same finding appearing in both full Gate-B branch files. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-59.md | 10 +++ .../gate-a-spec-awsf1ec771-resume.md | 39 ++++++++- ...26-09-10-loop-rule-consolidation-design.md | 6 +- ...-10-loop-rule-consolidation-target-text.md | 82 +++++++++++-------- 4 files changed, 98 insertions(+), 39 deletions(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-59.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-59.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-59.md new file mode 100644 index 0000000..fcf4ab8 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-59.md @@ -0,0 +1,10 @@ +MAJOR | high | §F item 15 "Replaced span" | The proposed replacement ends at the first "Run Gate B ... now;" and says the machinery after it is untouched, but the same live hook sentence later says "delete it and run a fresh pass" even though item 15's rationale says that pass command is removed | The installed reminder would first defer the next action to the ordering and then unconditionally order another pass, so it can still step past a suspension or source block | Replace the later tail explicitly with a deletion-only remedy, or give the complete resulting reminder string so no fresh-pass command survives outside the ordering pointer +MAJOR | high | §F item 16 "Two non-adjacent sentences" | Item 16 supplies one semicolon-linked sentence "in place of the two imperatives" while requiring the diagnostic sentence between those imperatives to remain untouched; splitting the proposal at the semicolon leaves the first fragment unterminated and the second beginning with dependent lowercase "where", while replacing the whole span deletes the promised diagnostic | The target does not determine one installable reminder or one complete `expected_ctx`, so the plan can preserve the diagnostic only by inventing wording | Give two independently grammatical replacement spans in their actual positions, or write the complete resulting reminder once +MAJOR | high | §F item 15 "no-fingerprint reminder" | The target changes the agent-facing opening to admit that a review may have run but its fingerprint could not be written or read back, yet the same live `note` call and its exact `expected_msg` fixture retain "⚠ Codex Gate B: no recorded review" | The operator and agent receive contradictory diagnoses, and an operator can order an unnecessary review after a valid pass whose state store failed | Replace the `systemMessage` with a neutral no-recorded-fingerprint diagnosis and update its exact fixture +MAJOR | high | §F item 16 "stale-fingerprint reminder" | The target says this branch is "a fault in the machinery and not a verdict on the cycle", but the live `systemMessage` and exact `expected_msg` fixture remain "⚠ Codex Gate B not satisfied (cannot confirm review)"; that same phrase is also what the stale-branch loose assertions match | The terse operator channel still states the gate verdict item 16 removes from the agent-facing opening, so partial adoption preserves the contradiction and the tests reward it | Replace the `systemMessage` with a neutral hook-state warning and update both its exact fixture and stale-branch assertions to test the observed fingerprint state instead of a gate verdict +MINOR | high | §F item 15 "either no review ... or ... could not be written or read back" | The two-cause opening omits the path §A3 itself adds: every non-`WIP` commit attempt makes the hook delete the Gate-B fingerprint and counters while the policy cycle and its already-valid passes remain open | The diagnostic can misclassify a reviewed cycle with deliberately cleared hook state as a missing or failed review-state write and send the reader to irrelevant permission checks | Name the non-`WIP` reset as a distinct cause with its Gate-B-closure pointer, or replace the closed cause enumeration with a diagnostic that does not claim why the fingerprint is absent +MINOR | high | §F item 16 "returns this same state" | A failed fingerprint store need not return the stale-fingerprint branch: opening or truncating the state file before a write failure can leave it empty or unreadable, which the hook routes to item 15's no-fingerprint branch | The replacement overclaims the hook's state transition and gives the reader a false expectation about which diagnostic repeats | Say that another failed store leaves the hook unable to confirm a fingerprint and may return either fingerprint-state diagnostic +MAJOR | medium | §A1 "A line in one branch file and a line in the other are distinct findings for holds and answers" | Two `full` Gate-B branches can emit the same out-of-set finding, and the target requires two independent answers but gives no rule for an accept on one copy and a decline on the other; §B then defines one assigned fix set from every accepted finding minus every declined finding even though both copies entail the same repair | A later occurrence can be simultaneously inside and outside the fix set, making its membership trigger, resolve duty and pass cleanliness undecidable and violating acceptance criterion 1's every-conflict requirement | Require one consistent membership decision for duplicate same-pass complaints, or surface conflicting answers until the user reconciles them before the next pass +MINOR | high | "How to read a section" "Both prompt copies take every NEW and REPLACED section byte-identical" | §F is marked REPLACED but contains seven hook-only replacements that §F itself says have one copy and no template mirror | The target's parity rule contradicts its hook scope and can make the plan apply or compare single-copy hook text as though it belonged in both prompt copies | Qualify the parity sentence to replacements whose destination is both prompt copies and exclude the hook-only items explicitly +MINOR | high | design opening "target text §F items 10–17" | The contiguous range includes item 14, which §F identifies as a sentence in both prompt copies, while the design, story and target elsewhere define the seven hook items as 10–13 and 15–17 | The design disagrees with the target's site map and can make the plan misclassify item 14 as a hook edit | Replace the range with the exact non-contiguous set "items 10–13 and 15–17" +END OF FINDINGS (9 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index e810a59..d953377 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -75,7 +75,8 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 56 | 6b8854f | 18→**13** | 0→**0** | 6→**0** | yes | **first Blocker- and Major-free pass of the cycle.** 11 of 13 are the second-copy class, collected. Two were count contradictions and were repaired, so 56 is not the final pass: design §9 still said "three contradictory reminder strings" against its own enumeration of seven (my pass-55 propagation missed it), and §F item 7's rationale said "the three closing paths" while §A2 defines two cases | | 57 | bb1ba03 | 13→**5** | 0→**0** | 0→**5** | yes | zero tells. All five are pass-53/55 fallout inside the hook items. **Two are a factual error of mine:** `codex-gate.test.sh` has **three** `expected_ctx` exact-match expectations (lines 1006, 1019, 1029), not one — my first grep searched only for phrases that occur in 1019. Items 12, 15 and 16 are all pinned. The other three tightened items 13 and 15, whose replacements still carried an unqualified imperative and an enumeration omitting the closing act | | 58 | ebe3aa7 | 5→**6** | 0→**0** | 5→**4** | yes | one tell (findings rose 5→6). All four Majors in items 15–17, one layer deeper than pass 57 reached: both fingerprint reminders still **opened** "STOP — Codex Gate B not satisfied", a gate verdict the hook cannot reach; item 16 claimed a permitted pass "records" a usable fingerprint where the store can fail again; item 17's enumeration omitted the below-floor zero-finding closing route. Minor 6 was design §9's "one exact-match expectation" — the **second** time that bullet went stale under a propagation | -| 59 | — | — | — | — | not run | next, against the pass-58 repair commit | +| 59 | fecb3aa | 6→**9** | 0→**0** | 4→**5** | yes | **MANDATORY TWO-TELL STOP** — findings rose 5→6→9, and four of nine were about §F's *prose describing* its replacements rather than the replacements. Surfaced, standing answer applied, loop continued. Structural repair: items 15 and 16 now give the **complete resulting message** for both channels instead of a span plus prose about where it goes. Finding 7 is a **new behaviour question** and is open with Daniel: conflicting accept/decline on the same finding in two `full` Gate-B branch files | +| 60 | — | — | — | — | not run | blocked on finding 7's answer; everything else from 59 is applied | | 40 | 12cf247 | 5→**3** | 0→**0** | 4→**3** | yes | zero tells; first round with no fan-out Major after the three-site check was run BEFORE the pass | | 41 | d26d097 | 3→**2** | 0→**0** | 3→**2** | yes | zero tells; both Majors were compressed pointers of mine dropping a load-bearing part of a standing rule | | 42 | 3b61fe3 | 2→**1** | 0→**0** | 2→**1** | yes | zero tells; §F item 4 had taken back the enumeration it was repaired to avoid | @@ -84,6 +85,42 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 45 | 1c858ac | 7→**7** | 0→**0** | 3→**1** | yes | zero tells; swept all four sites still tying classification to closing | | 46 | 4b713dd | 7→**9** | 0→**1** | 1→**2** | yes | **TWO TELLS — mandatory stop under §5, answered in advance by Daniel's standing "no stops without an absolute block" of 2026-09-13 17:02.** Tells surfaced here and in the commit body, findings kept open, no pass credited clean; the loop continued on that answer. **Eleventh** falsified standing sentence; session 01a09b7c-dd6b-7cb3-b83a-776599286f9b | +## Pass-59 report — MANDATORY TWO-TELL STOP, and the span-framing was the mechanism + +**Trend:** findings 5, 6, **9** across 57–59; Blockers 0, 0, **0**; Majors 5, 4, **5**. +**Cluster:** four of nine (findings 1, 2, 8, 9) are about **§F's prose describing** its own +replacements — the span framing, the parity sentence, the item range — rather than about the +replacement text. **Require↔withdraw:** none. + +**Two tells are present and the stop is mandatory, not discretionary:** the finding count rose +(5 → 6 → 9), and findings cluster on prose about the product rather than on the product. Surfaced +to Daniel; his standing answer of 2026-09-13 — surface in the record and continue — applies, as it +did at pass 46. **The loop continued on that answer.** + +**What the tells were pointing at, which is the useful part.** §F's hook items were written as a +**span plus a sentence saying where it goes**: "replaced span: the opening through X", "two +non-adjacent sentences of one message". Three passes running found a new gap in that framing — +a tail left standing behind a pointer (57, 59), a fragment that is not independently grammatical +(59), a diagnostic that could only be preserved by inventing wording (59). The framing was +generating the findings, not the text. + +**Repair: items 15 and 16 now carry the complete resulting message**, for both the +`additionalContext` and the `systemMessage` channel, with no prose about spans. A plan installs +what is written. This is the file's own preference — a duty or the thing itself over a description +of where the thing goes — applied to its own §F. + +**Two further sites the same pass exposed**, both mechanical: the "How to read a section" parity +rule said both prompt copies take every REPLACED section byte-identical, which is false of seven +hook-only items; and the design opening wrote the hook items as the contiguous range 10–17, which +swallows item 14, a prompt-copy sentence. Both corrected. + +**Open and blocking pass 60: finding 7.** In a `full` Gate-B pass the two branch files can carry +the same out-of-set finding. The target makes them distinct findings for holds and answers, so two +answers are required; §B builds one fix set from every acceptance minus every decline. An accept on +one copy and a decline on the other puts one repair simultaneously in and out of the set, leaving +its membership trigger, resolve duty and the pass's cleanliness undecidable. That is a behaviour +decision this change has not been given, and it is with Daniel. + ## Pass-58 report — one tell, and the shape of the hook-item lineage **Trend:** findings 13, 5, **6** across 56–58; Blockers 0, 0, **0**; Majors 0, 5, **4**. diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index 8509ac8..5b5adb4 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -7,9 +7,9 @@ copy, and a value copied here would be a remembered value. Prompt-only, in two mirrored copies — `CLAUDE.md` §5 (**C** below) and the inline template in `plugins/dev-workflow/commands/workflow-init.md` (**W** below) — **plus seven reminder strings in -`plugins/dev-workflow/hooks/codex-gate.sh` and the three exact-match expectations in -`plugins/dev-workflow/hooks/codex-gate.test.sh` that pin three of them** (target text §F items -10–17, authorised 2026-09-13). **No hook behaviour changes**, and the hook strings have no mirror, +`plugins/dev-workflow/hooks/codex-gate.sh` and their exact-match expectations in +`plugins/dev-workflow/hooks/codex-gate.test.sh` — three `expected_ctx` and two `expected_msg`** (target text §F items +10–13 and 15–17, authorised 2026-09-13). **No hook behaviour changes**, and the hook strings have no mirror, so they carry no parity obligation. Condition ids `a1`…`j4` are defined in `docs/superpowers/specs/2026-09-10-loop-rule-consolidation-condition-inventory.md` beside this file: 135 conditions quoted from `7c0d475`, so a reader can check an accounting rather than diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index 0e4968d..d477a2d 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -30,7 +30,9 @@ implementation diff later, and this file is not that diff. proposed text** — it is this file's own metadata and ships nowhere. A CARRIED passage is reproduced from the current state unchanged, because the block cites it and a reader has to see what it says; it is here to be read, not to be edited. Both prompt copies take every NEW and -REPLACED section **byte-identical**. +REPLACED section **byte-identical** — **except §F's hook items, whose destination is the shipped +hook and its test**; those have one copy, no template mirror and no parity obligation, and §F says +which they are. **What is deliberately not here:** the record-durability material (successor story), and the per-condition disposition, parity diff and verification fragments (the plan, against real files). @@ -665,12 +667,15 @@ Each is a live sentence that the block makes wrong. All twenty-two are **known c none is deferred. **Seventeen of them share one mechanism** — an entry point other than the ordering carrying an unqualified instruction — which is why each is **replaced** rather than given an exception to point at. **Fifteen live in the two prompt copies and cite a line in each; the last -seven live in the shipped hook**, which has one copy and no template mirror. The hook holds eight +seven live in the shipped hook**, which has one copy and no template mirror — so those seven are +outside the byte-identical parity rule the section opening states. The hook holds eight gate reminders in all; the one this change leaves alone is the docs-only notice, which states no closure permission. **Three of the seven are pinned by exact-match expectations in `plugins/dev-workflow/hooks/codex-gate.test.sh`** — items 12, 15 and 16, at that file's three -`expected_ctx` assignments — and each is replaced there with its complete resulting message in the -same change. The remaining four are matched by loose patterns these repairs leave standing. +`expected_ctx` assignments, **and items 15 and 16 also at the two `expected_msg` assignments beside +them**, since those items replace the terse operator line as well — each replaced there with its +complete resulting text in the same change. The remaining four are matched by loose patterns these +repairs leave standing, **except the stale-branch assertions item 16 moves with it**. **1. The `WIP:` naming warning** (Mechanics · `baseSha`). ``` @@ -940,50 +945,57 @@ reporting its own threshold as an obligation at a floor of 1 — is untouched an scope. This one is **not** an entry point carrying an unqualified instruction; it is a false statement about scope, so it does not join that count. -**15. The no-fingerprint reminder's opening and its next action** (the shipped hook, same file). -Replaced span: the opening through "Run Gate B (mcp__codex__review) now;", the machinery checks -after it untouched. +**15. The no-fingerprint reminder, both channels** (the shipped hook, same file). **The complete +resulting message is given, not a span**: a span plus a sentence about where it goes was found +three times running to leave a fragment ambiguous or a tail standing. ``` -Codex gate state: no fingerprint is recorded for this cycle — either no mcp__codex__review has run, or the last one's fingerprint could not be written or read back. Per $policy you MUST reach a minimum of $floor passes per cycle. What this cycle does next is $policy's closure ordering's, read there entire, and this reminder decides none of it — including whether the state file's deletion below is followed by a pass; +Codex gate state: no fingerprint is recorded for this cycle. The hook cannot tell why — no mcp__codex__review has run, the last one's fingerprint could not be written or read back, or a non-`WIP` commit attempt cleared it while the cycle itself stayed open. Per $policy you MUST reach a minimum of $floor passes per cycle. What this cycle does next is $policy's closure ordering's, read there entire, and this reminder decides none of it. If this repeats, check that .context/ and the state file inside it are readable and writable; if the file exists but is unreadable or empty, delete it — which restores no passes, and lets the next pass the ordering permits record a fingerprint. ``` -*Why (pass 55 finding 3):* the live string says "Run Gate B (mcp__codex__review) now", which sends +``` +⚠ Codex Gate B: no recorded fingerprint +``` +*Why (pass 55 finding 3):* the live string said "Run Gate B (mcp__codex__review) now", which sends the author into another pass whatever the cycle's state is — including a pass that carried a -suspension whose answers are still outstanding, which the composition rule holds the cycle on. The -diagnostic half, and the machinery checks after it, are kept: they are what the message is for. -*And (pass 57 findings 2 and 3):* an earlier wording enumerated the ordering's routes as a pass, a -suspension's answer or a source block's repair, which **omits the closing act** — an absent -fingerprint is not a closure condition and Gate B has no content condition, so a clean eligible -cycle whose fingerprint merely could not be stored still has its commit route. The enumeration is -removed rather than extended, per this section's own preference. The same finding's second half -reaches the sentence's tail, "delete it and run a fresh pass": the deletion stays as a remedy, the -pass after it does not, and the new sentence says so where a reader meets it. -*And (pass 58 finding 1):* the opening still read "STOP — Codex Gate B not satisfied", which -states a gate verdict the hook cannot reach — the fingerprint is advisory and is not a closure -condition, so a clean eligible cycle whose fingerprint merely could not be stored was being told -the gate was unsatisfied. The opening now reports the **hook's own state**, which is what it -observes, and leaves every verdict to the ordering. This is the AGENTS.md rule about naming the +suspension whose answers are still outstanding, which the composition rule holds the cycle on. +*And (pass 57 findings 2 and 3, pass 59 finding 1):* the routes the ordering offers were first +enumerated, which omitted the closing act, and then pointed at while the same sentence's tail, +"delete it and run a fresh pass", stayed standing behind the pointer. Both are gone because the +whole message is written out. +*And (pass 58 finding 1):* the opening read "STOP — Codex Gate B not satisfied", a gate verdict the +hook cannot reach — its fingerprint is advisory and is not a closure condition, so a clean eligible +cycle whose fingerprint merely could not be stored was being told the gate was unsatisfied. It now +reports the hook's own state, which is what it observes. That is AGENTS.md's rule about naming the exact comparison a mechanism performs, applied to the hook's own words. +*And (pass 59 findings 3 and 5):* the terse `systemMessage` still read "no recorded review", which +says a review did not happen where the hook knows only that no fingerprint is stored — an operator +could order a review over a valid pass whose store failed. Its exact fixture moves with it. The +cause list also omitted the one §A3 itself creates: a non-`WIP` commit attempt clears the +fingerprint while the cycle stays open, and a reader sent to permission checks for it would be +looking in the wrong place. -**16. The stale-fingerprint reminder's opening and its two instructions** (the shipped hook, same file). -**Two non-adjacent sentences of one message**, the diagnostic list between them untouched. The -opening: +**16. The stale-fingerprint reminder, both channels** (the shipped hook, same file). **The complete +resulting message, for the reason item 15 gives.** ``` -Codex gate state: the hook cannot confirm that the content you are about to commit is the content mcp__codex__review last saw ($passes recorded pass(es) this cycle). +Codex gate state: the hook cannot confirm that the content you are about to commit is the content mcp__codex__review last saw ($passes recorded pass(es) this cycle). Usually that means the working tree or the index changed since the review. It can also mean you only staged already-reviewed content — the bytes are fine, but the hook cannot tell staging from editing; that this hook was upgraded and the recorded fingerprint uses the older format (see CHANGELOG); or that the fresh fingerprint could not be computed or could not be stored. A fresh Gate-B pass is the complete remedy for the staging and post-upgrade cases too, and when this cycle may run one is $policy's closure ordering's, read there entire. If a fresh pass leaves this unchanged with nothing edited in between, the fault is in the machinery rather than the code: check that .context/ is writable, that TMPDIR is writable, that a checksum tool (shasum, sha1sum or cksum) runs, that git status works, and that the disk is not full — a store that fails again leaves the hook unable to confirm a fingerprint and may return either fingerprint-state diagnosis. Per $policy you MUST re-review after every fix. ``` -and, in place of the two imperatives: ``` -A fresh Gate-B pass is the complete remedy for the staging and post-upgrade cases too, and when this cycle may run one is $policy's closure ordering's; where it may, that pass attempts to record a usable fingerprint — a store that fails again returns this same state, which is a fault in the machinery and not a verdict on the cycle. +⚠ Codex Gate B: cannot confirm reviewed content ``` -*Why (pass 55 finding 4):* the live string carries two unqualified imperatives — "Run Gate B +*Why (pass 55 finding 4):* the live string carried two unqualified imperatives — "Run Gate B (mcp__codex__review) now" and "then run one more pass to record a usable fingerprint" — in the branch an author reads immediately before committing. Both step past a suspension or a source -block that the ordering says is answered first. The long diagnostic list between them is untouched. +block that the ordering says is answered first. *And (pass 58 findings 2 and 3):* the opening read "STOP — Codex Gate B not satisfied" over a branch whose own text lists staging, a hook upgrade and a failed store — none of which violates a -closure condition — so it takes item 15's repair, reporting the hook's state and leaving the -verdict to the ordering. The second sentence also said a permitted pass "records a usable -fingerprint", which no rule makes true: the store that just failed can fail again. It now says -**attempts**, and names the loop that would otherwise send an author round again. +closure condition — so it takes item 15's repair. The second sentence also said a permitted pass +"records a usable fingerprint", which no rule makes true where the store just failed. +*And (pass 59 findings 2, 4 and 6):* an earlier revision supplied one semicolon-linked sentence +"in place of the two imperatives" while requiring the diagnostic between them to stay, which no +plan could install without inventing wording — hence the complete message. The terse +`systemMessage` still read "not satisfied (cannot confirm review)", the gate verdict this item +removes from the agent-facing text, and its exact fixture and the stale-branch loose assertions +move with it. And a failed store can leave the state file empty or unreadable, which the hook +routes to item 15's branch, so the message no longer claims this diagnosis repeats. **17. The Gate-B below-floor reminder's instruction** (the shipped hook, same file). ``` From e911b39cd23626adea653c1f1272cdad98201edd Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 09:49:07 +0200 Subject: [PATCH 090/181] docs(specs): decision 6; a decline excludes its own finding and nothing else MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Daniel's answer to pass 59 finding 7, taken with narrow semantics. At the fix-set definition in §B: a decline excludes the finding it answers and does not cancel a repair obligation the approved scope or another accepted finding independently established. Where a full Gate-B pass's two branch files carry the same complaint the lines stay distinct findings, each owes its own explicit answer, each answer binds only its own line, and until both are given the unanswered one's membership hold stands. A declined finding stays declined for the cycle and discharging another finding's obligation does not return it. Deliberately no deduplication, no reconciliation stop and no general "acceptance overrides decline" rule. All four answer combinations were walked: accept/accept and accept/decline and decline/accept leave the repair owed through whichever finding was accepted; decline/decline leaves it owed only where the approved scope owns it independently. Recorded as design §2 decision 6. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../2026-09-10-loop-rule-consolidation-design.md | 1 + ...2026-09-10-loop-rule-consolidation-target-text.md | 12 +++++++++++- 2 files changed, 12 insertions(+), 1 deletion(-) diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index 5b5adb4..bb18bce 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -82,6 +82,7 @@ words live in the target text; this section says **that** a decision was made, * | 3 | 2026-09-13, Daniel | whether a reviewer repeating a claim the author has validly refuted can hold a cycle unclean indefinitely (pass 27's Blocker) | target §A1, with §C, §H and §H's unknown-start list for its reach | | 4 | 2026-09-13, agent in scope | whether an eligible pass that cannot close may still reach a suspension, §D's two-tell stop being mandatory (pass 29 finding 1) | target §A1, §D | | 5 | 2026-09-13, agent in scope | what a closing act that does not complete leaves behind, and what a failure nobody can repair produces (pass 30 finding 4, narrowed at pass 31, terminal state at pass 37) | target §A1 | +| 6 | 2026-09-14, Daniel | what an accept on one `full` Gate-B branch file and a decline on the other do to the assigned fix set, the two lines being distinct findings (pass 59 finding 7) | target §B, at the fix-set definition | **Two were reached by narrowing rather than by adding**, recorded because the first attempt at each was wider than its defect: decision 5 first claimed every condition stayed established and added a diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index d477a2d..1d3a6e8 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -477,7 +477,17 @@ never decides what you do with it, and it grants no Minor or Nit a repair round otherwise get. **The assigned fix set is fixed before the pass you are answering: it is the union of the scope every approved story or plan governing this change assigns to this cycle, plus every finding this cycle has accepted at a membership stop together with any repair -obligation accepted with it, minus every finding this cycle has declined.** **An accept puts the +obligation accepted with it, minus every finding this cycle has declined.** **A decline excludes +the finding it answers and nothing else**: it does not cancel a repair obligation that the +approved scope, or another accepted finding, has independently put in the set. So where a `full` +Gate-B pass's two branch files carry the same complaint, **each line is a finding of its own, owes +its own explicit answer, and each answer binds only its own line** — an acceptance puts its own +finding in, a decline takes only its own finding out, and neither reads the other; the set is +whatever the definition above then computes. **A declined finding stays declined for the cycle**, +and performing a repair to discharge a different finding's obligation neither reverses that +decision nor returns it to the set. **This adds no deduplication, no reconciliation stop and no +rule that an acceptance overrides a decline** — the two answers are about different findings, and +until both are given the unanswered one's membership hold stands. **An accept puts the finding in the set whatever its severity**: membership and the repair duty are different things, so a Minor or Nit accepted into the set is in it though Mechanics · Severity asks no repair for it, and a later pass that recomputed it as outside would raise the membership question a second From 5197b0c7fa0b90823fcb4250fbd99b2f1c46b406 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 10:10:36 +0200 Subject: [PATCH 091/181] docs(specs): apply pass 60; a backtick that would have run as a command MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Item 15's complete message carried a backtick around WIP. That message installs into a double-quoted POSIX-shell note argument, where a backtick is command substitution: the hook would have executed WIP, shipped the reminder with the word missing, failed ShellCheck, and the same hazard copied into expected_ctx would have made the suite pass on the corruption. Invisible while §F carried spans; visible one pass after the complete messages went in. §F now states the constraint for every hook replacement. Item 18: the standing work-loop sequence names a commit only after Gate B, while §A2 makes a completed closing act mandatory for a Gate-A cycle too. §F 22 -> 23. Item 11 let an author proceed on a clean pass plus the closure conditions, which permits execution after a closing act that never completed; it now requires the cycle to have closed. Item 16 called an unchanged fingerprint a machinery fault while the index moving changes it; it now checks worktree and index first. Three Minors and a Nit collected, not repaired. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-60.md | 9 +++++ .../gate-a-spec-awsf1ec771-resume.md | 35 +++++++++++++++- ...-10-loop-rule-consolidation-target-text.md | 40 ++++++++++++++----- 3 files changed, 74 insertions(+), 10 deletions(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-60.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-60.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-60.md new file mode 100644 index 0000000..0389f81 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-60.md @@ -0,0 +1,9 @@ +MAJOR | high | standing work-loop sequence "Gate A (spec) → plan ready" | The live sequence at CLAUDE.md:63 and workflow-init.md:262 moves directly from each Gate-A gate to the next phase and names `commit` only after Gate B, while §A2 makes a completed commit-or-amend closing act mandatory for both Gate-A cycles and §F does not replace this sentence | A reader can start planning or execution with the preceding Gate-A cycle still open and without its provenance line or curve | Put the Gate-A closing act after both Gate-A cycles in the sequence, citing §A2 rather than restating its operation +MAJOR | high | §F item 11 "Proceed only if your final pass was clean" | The proposed Gate-A satisfied reminder fires at the transition into execution but conditions proceeding only on a clean pass and the closure conditions; it never requires the closing act to have completed, although §A1 says an incomplete act leaves the cycle open | An agent entering through the hook can execute after an unattempted or failed Gate-A commit and bypass decision 5's terminal state | Replace the condition list with "Proceed only after the Gate-A cycle has closed under $policy's closure ordering" +MAJOR | high | §F item 15 "non-`WIP` commit attempt" | The complete message is destined for the existing double-quoted POSIX-shell `note` argument, where the proposed raw backticks perform command substitution; a literal installation runs `WIP`, emits `WIP: not found`, removes `WIP` from the reminder, and triggers ShellCheck SC2006, and the same quoting hazard reaches the exact `expected_ctx` assignment | The shipped hook text is corrupted at runtime and the required quality command fails even if the fixture is made to mirror the corruption | Write "non-WIP commit attempt" in the resulting message, avoiding shell-active delimiters in both the hook and its fixture +MINOR | high | §F item 15 "If this repeats, check that .context/ ..." | The message now names a non-`WIP` commit attempt as a distinct cause of the missing fingerprint, but its only repeat remedy unconditionally sends the reader to filesystem permissions and deletion; another non-`WIP` attempt reproduces the state with a healthy filesystem | The diagnostic directs the author away from the actual reset and violates prompt-standards item 10's cause-specific remedy rule | Route a known intervening non-`WIP` reset to §A3 and reserve the filesystem remedy for a repeat after a permitted pass with no intervening reset +MAJOR | high | §F item 16 "nothing edited in between" | The proposed stale-fingerprint reminder concludes that a repeat after a fresh pass is a machinery fault when "nothing [was] edited", but the same message distinguishes staging from editing and the hook fingerprint changes when the index is staged after that pass; another hook can also stage content during the next commit attempt | A normal index transition is misdiagnosed as broken machinery, so the reader can chase permissions and repeat reviews without checking the state that actually changed | Replace the categorical conclusion with a diagnostic that checks both the worktree and index for intervening changes before checking fingerprint computation and storage +MINOR | high | §A1 "distinct findings for holds and answers" / §B "each line is a finding of its own" / design §2 decision 6 | The branch-line identity and per-line answer-binding rule is stated in §A1, restated in §B, and asserted a third time by the design row's clause "the two lines being distinct findings", despite the design saying its table records only the question and the target saying §A1 owns the pass read model | Three authorities can drift on whether equal complaints in the two files are one or two findings, changing holds and assigned-fix-set membership together | Keep branch-line identity and answer binding in §A1, have §B cite that rule while defining only its set consequence, and phrase the design row as the question without asserting the answer +MINOR | high | §A1 "a decline binding for the remainder of its cycle" / §B "A declined finding stays declined for the cycle" | §B repeats the duration rule that §A1 already owns before adding the new consequence about a different finding's repair; the two sentences can later disagree on withdrawal or cross-cycle effect | The same decline can acquire different lifetime semantics depending on whether the reader enters through the ordering or the fix-set definition | Leave decline duration in §A1 and make §B state only that repairing another finding does not return this declined finding to the set +NIT | high | §B "an acceptance puts its own finding in, a decline takes only its own finding out" | The mixed-branch explanation repeats the immediately preceding rule that a decline excludes its finding and nothing else and the immediately following rule that an accept puts its finding in the set, creating two local copies of both membership effects | The extra enumeration adds another edit site for any future change to accept or decline semantics without adding a new case outcome | Remove the accept/decline enumeration and say that each answer binds only its own line and the definition above computes the resulting set +END OF FINDINGS (8 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index d953377..f5d4eba 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -76,7 +76,8 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 57 | bb1ba03 | 13→**5** | 0→**0** | 0→**5** | yes | zero tells. All five are pass-53/55 fallout inside the hook items. **Two are a factual error of mine:** `codex-gate.test.sh` has **three** `expected_ctx` exact-match expectations (lines 1006, 1019, 1029), not one — my first grep searched only for phrases that occur in 1019. Items 12, 15 and 16 are all pinned. The other three tightened items 13 and 15, whose replacements still carried an unqualified imperative and an enumeration omitting the closing act | | 58 | ebe3aa7 | 5→**6** | 0→**0** | 5→**4** | yes | one tell (findings rose 5→6). All four Majors in items 15–17, one layer deeper than pass 57 reached: both fingerprint reminders still **opened** "STOP — Codex Gate B not satisfied", a gate verdict the hook cannot reach; item 16 claimed a permitted pass "records" a usable fingerprint where the store can fail again; item 17's enumeration omitted the below-floor zero-finding closing route. Minor 6 was design §9's "one exact-match expectation" — the **second** time that bullet went stale under a propagation | | 59 | fecb3aa | 6→**9** | 0→**0** | 4→**5** | yes | **MANDATORY TWO-TELL STOP** — findings rose 5→6→9, and four of nine were about §F's *prose describing* its replacements rather than the replacements. Surfaced, standing answer applied, loop continued. Structural repair: items 15 and 16 now give the **complete resulting message** for both channels instead of a span plus prose about where it goes. Finding 7 is a **new behaviour question** and is open with Daniel: conflicting accept/decline on the same finding in two `full` Gate-B branch files | -| 60 | — | — | — | — | not run | blocked on finding 7's answer; everything else from 59 is applied | +| 60 | e911b39 | 9→**8** | 0→**0** | 5→**5** | yes | zero tells. **Finding 3 is the one that mattered:** item 15's complete message carried a backtick around `WIP`, which is command substitution inside the hook's double-quoted `note` argument — installed literally it would run `WIP`, corrupt the reminder and fail ShellCheck. Only visible because the message is now written out whole. §F now states the no-shell-active-characters constraint. Also: the standing work-loop sequence names a commit only after Gate B → item 18, §F 22 → 23; item 11 conditioned proceeding on a clean pass rather than on the cycle having closed; item 16 called an unchanged fingerprint a machinery fault while the index can move it. 3 Minors + 1 Nit **collected, not repaired** | +| 61 | — | — | — | — | not run | next, against the pass-60 repair commit. **If 61 has zero Blockers and zero Majors the cycle closes** — Minors are collected, per Mechanics · Severity | | 40 | 12cf247 | 5→**3** | 0→**0** | 4→**3** | yes | zero tells; first round with no fan-out Major after the three-site check was run BEFORE the pass | | 41 | d26d097 | 3→**2** | 0→**0** | 3→**2** | yes | zero tells; both Majors were compressed pointers of mine dropping a load-bearing part of a standing rule | | 42 | 3b61fe3 | 2→**1** | 0→**0** | 2→**1** | yes | zero tells; §F item 4 had taken back the enumeration it was repaired to avoid | @@ -85,6 +86,38 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 45 | 1c858ac | 7→**7** | 0→**0** | 3→**1** | yes | zero tells; swept all four sites still tying classification to closing | | 46 | 4b713dd | 7→**9** | 0→**1** | 1→**2** | yes | **TWO TELLS — mandatory stop under §5, answered in advance by Daniel's standing "no stops without an absolute block" of 2026-09-13 17:02.** Tells surfaced here and in the commit body, findings kept open, no pass credited clean; the loop continued on that answer. **Eleventh** falsified standing sentence; session 01a09b7c-dd6b-7cb3-b83a-776599286f9b | +## Pass-60 report — zero tells, and the span repair paying for itself + +**Trend:** findings 6, 9, **8** across 58–60; Blockers 0, 0, **0**; Majors 4, 5, **5**. +**Cluster:** product behaviour. **Require↔withdraw:** none. **Zero tells** — the finding count fell +and the prose-about-prose cluster that made pass 59 a two-tell stop did not return. + +**Finding 3 is why pass 59's repair was worth making.** Item 15's complete message contained +`` `WIP` ``. That message is destined for a **double-quoted POSIX-shell `note` argument**, where a +backtick is command substitution: installed literally the hook would execute `WIP`, print +`WIP: not found`, ship a reminder with the word missing, and fail ShellCheck — and the same hazard +would have been copied into the `expected_ctx` fixture, so the suite would have passed on the +corruption. **This was invisible while §F carried spans and prose about where they go.** Writing +the message out whole is what exposed it, one pass after the change. §F now carries the constraint +so the next writer does not reintroduce it. + +**Item 18, and the sweep reaching a fifth kind of site.** The standing work-loop sequence — +"spec ready → Gate A (spec) → plan ready → Gate A (plan) → execute → tests green → Gate B → +commit" — names a commit only after Gate B, while §A2 makes a completed closing act mandatory for +a Gate-A cycle too. A reader following it plans with the previous cycle still open. §F 22 → 23. + +**Two more Majors, both narrowing a permission:** item 11 let an author proceed on a clean pass +plus the closure conditions, which still permits execution after a closing act that never +completed; it now requires the cycle to have **closed**. And item 16 concluded that an unchanged +fingerprint after a fresh pass is a machinery fault, while the index moving — staging included, and +another hook can stage during the commit — changes it; the message now checks worktree and index +first. + +**Three Minors and a Nit were collected and not repaired.** Stated explicitly because pass 56's +two Minor repairs were not licensed by Mechanics · Severity and cost three passes; Daniel raised +that and it is correct. The discipline from here: a pass with zero Blockers and zero Majors closes +the cycle. + ## Pass-59 report — MANDATORY TWO-TELL STOP, and the span-framing was the mechanism **Trend:** findings 5, 6, **9** across 57–59; Blockers 0, 0, **0**; Majors 5, 4, **5**. diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index 1d3a6e8..1500b11 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -671,13 +671,13 @@ demotes it — the two counts are meant to differ. --- -## F. The twenty-two standing sentences this change falsifies — REPLACED +## F. The twenty-three standing sentences this change falsifies — REPLACED -Each is a live sentence that the block makes wrong. All twenty-two are **known contradictions** and -none is deferred. **Seventeen of them share one mechanism** — an entry point other than the ordering +Each is a live sentence that the block makes wrong. All twenty-three are **known contradictions** and +none is deferred. **Eighteen of them share one mechanism** — an entry point other than the ordering carrying an unqualified instruction — which is why each is **replaced** rather than given an -exception to point at. **Fifteen live in the two prompt copies and cite a line in each; the last -seven live in the shipped hook**, which has one copy and no template mirror — so those seven are +exception to point at. **Sixteen live in the two prompt copies and cite a line in each; seven +live in the shipped hook**, which has one copy and no template mirror — so those seven are outside the byte-identical parity rule the section opening states. The hook holds eight gate reminders in all; the one this change leaves alone is the docs-only notice, which states no closure permission. **Three of the seven are pinned by exact-match expectations in @@ -686,6 +686,11 @@ closure permission. **Three of the seven are pinned by exact-match expectations them**, since those items replace the terse operator line as well — each replaced there with its complete resulting text in the same change. The remaining four are matched by loose patterns these repairs leave standing, **except the stale-branch assertions item 16 moves with it**. +**Every hook replacement below is destined for a double-quoted POSIX-shell `note` argument**, so +its text carries no backtick, no `$(`, no backslash and no double quote; `$policy`, `$floor` and +`$passes` are the intended interpolations and stay. A backtick reached one of these blocks at pass +59 and would have run `WIP` as a command at install time, corrupting the shipped reminder and +failing ShellCheck. **1. The `WIP:` naming warning** (Mechanics · `baseSha`). ``` @@ -909,13 +914,17 @@ prompt copy. **11. The Gate-A satisfied reminder's clean definition** (the shipped hook, same file). ``` -Proceed only if your final pass was clean and every other closure condition holds, both as $policy defines them. +Proceed only once this Gate-A cycle has closed under $policy's closure ordering, read there entire. ``` *Why (pass 53 finding 20):* the live string defines a clean final pass as "no new Blocker/Major", an abbreviated second copy of a definition the ordering now states in full — a pass carrying a scope-stop trigger is not clean under it, whatever the severity of what triggered it. Left standing, the hook presents such a pass as clean at exactly the moment an author is deciding -whether to proceed. The repair **cites** the definition rather than restating it, which is also +whether to proceed. +*And (pass 60 finding 2):* an earlier replacement conditioned proceeding on a clean pass and the +closure conditions, which still lets an author execute after a Gate-A closing act that was never +attempted or did not complete — decision 5's terminal state. The condition is now **the cycle +having closed**, which is the one fact that covers all of them. The repair **cites** the definition rather than restating it, which is also what keeps a later change to the definition from falsifying this string again. It is the twelfth sentence sharing this section's mechanism. @@ -959,7 +968,7 @@ statement about scope, so it does not join that count. resulting message is given, not a span**: a span plus a sentence about where it goes was found three times running to leave a fragment ambiguous or a tail standing. ``` -Codex gate state: no fingerprint is recorded for this cycle. The hook cannot tell why — no mcp__codex__review has run, the last one's fingerprint could not be written or read back, or a non-`WIP` commit attempt cleared it while the cycle itself stayed open. Per $policy you MUST reach a minimum of $floor passes per cycle. What this cycle does next is $policy's closure ordering's, read there entire, and this reminder decides none of it. If this repeats, check that .context/ and the state file inside it are readable and writable; if the file exists but is unreadable or empty, delete it — which restores no passes, and lets the next pass the ordering permits record a fingerprint. +Codex gate state: no fingerprint is recorded for this cycle. The hook cannot tell why — no mcp__codex__review has run, the last one's fingerprint could not be written or read back, or a non-WIP commit attempt cleared it while the cycle itself stayed open. Per $policy you MUST reach a minimum of $floor passes per cycle. What this cycle does next is $policy's closure ordering's, read there entire, and this reminder decides none of it. If this repeats, check that .context/ and the state file inside it are readable and writable; if the file exists but is unreadable or empty, delete it — which restores no passes, and lets the next pass the ordering permits record a fingerprint. ``` ``` ⚠ Codex Gate B: no recorded fingerprint @@ -986,7 +995,7 @@ looking in the wrong place. **16. The stale-fingerprint reminder, both channels** (the shipped hook, same file). **The complete resulting message, for the reason item 15 gives.** ``` -Codex gate state: the hook cannot confirm that the content you are about to commit is the content mcp__codex__review last saw ($passes recorded pass(es) this cycle). Usually that means the working tree or the index changed since the review. It can also mean you only staged already-reviewed content — the bytes are fine, but the hook cannot tell staging from editing; that this hook was upgraded and the recorded fingerprint uses the older format (see CHANGELOG); or that the fresh fingerprint could not be computed or could not be stored. A fresh Gate-B pass is the complete remedy for the staging and post-upgrade cases too, and when this cycle may run one is $policy's closure ordering's, read there entire. If a fresh pass leaves this unchanged with nothing edited in between, the fault is in the machinery rather than the code: check that .context/ is writable, that TMPDIR is writable, that a checksum tool (shasum, sha1sum or cksum) runs, that git status works, and that the disk is not full — a store that fails again leaves the hook unable to confirm a fingerprint and may return either fingerprint-state diagnosis. Per $policy you MUST re-review after every fix. +Codex gate state: the hook cannot confirm that the content you are about to commit is the content mcp__codex__review last saw ($passes recorded pass(es) this cycle). Usually that means the working tree or the index changed since the review. It can also mean you only staged already-reviewed content — the bytes are fine, but the hook cannot tell staging from editing; that this hook was upgraded and the recorded fingerprint uses the older format (see CHANGELOG); or that the fresh fingerprint could not be computed or could not be stored. A fresh Gate-B pass is the complete remedy for the staging and post-upgrade cases too, and when this cycle may run one is $policy's closure ordering's, read there entire. If a fresh pass leaves this unchanged, check the worktree and the index first — the fingerprint moves when either does, staging included, and another hook can stage during the commit attempt. Where neither changed, the fault is in the machinery rather than the code: check that .context/ is writable, that TMPDIR is writable, that a checksum tool (shasum, sha1sum or cksum) runs, that git status works, and that the disk is not full — a store that fails again leaves the hook unable to confirm a fingerprint and may return either fingerprint-state diagnosis. Per $policy you MUST re-review after every fix. ``` ``` ⚠ Codex Gate B: cannot confirm reviewed content @@ -1020,6 +1029,19 @@ or a repair, which **omits the closing act a zero-finding pass reaches below the closing route this branch exists to describe. Removed rather than extended, exactly as item 15's was. + +**18. The work-loop sequence** (the §4 goal-driven-execution section). One line in each copy: +C 63 and W 262. +``` +The work loop includes the review gates: **spec ready → Gate A (spec) → Gate-A closing act → plan ready → Gate A (plan) → Gate-A closing act → execute → tests green → Gate B → Gate-B closing act** (see §5, which states when each act may be performed and what it is). +``` +*Why (pass 60 finding 1):* the live sequence moves straight from each Gate-A gate to the next +phase and names a commit only after Gate B, while §A2 makes a completed closing act mandatory for +a Gate-A cycle too. A reader following the live line starts planning or executing with the +preceding cycle still open, and with neither its provenance line nor its curve written. The +replacement names the act and cites §5 for when it may be performed, rather than restating the +operation. It is the eighteenth sentence sharing this section's mechanism. + --- ## G. The one-contract paragraph — REPLACED From 888d2d154fd09da03f37850f3d45e67d83fc70fc Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 10:24:14 +0200 Subject: [PATCH 092/181] docs(specs): apply pass 61; the operator line item 12 had left standing MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Items 15 and 16 replace both channels of their message; item 12 replaced only the long one, leaving "✓ Codex Gate B satisfied" on the terse operator line. That is a gate verdict the hook cannot establish — Gate B has no hook-backed content condition, and closure turns on duties, holds, the remaining conditions and a completed closing act, none of which the hook observes. It now reads "hook checks passed", and its exact fixture moves with it: expected_msg scope two to three. The counts in the line are kept; they are observations. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-61.md | 2 ++ .../gate-a-spec-awsf1ec771-resume.md | 19 ++++++++++++++++++- ...26-09-10-loop-rule-consolidation-design.md | 8 ++++---- ...-10-loop-rule-consolidation-target-text.md | 14 ++++++++++++-- 4 files changed, 36 insertions(+), 7 deletions(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-61.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-61.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-61.md new file mode 100644 index 0000000..ec15808 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-61.md @@ -0,0 +1,2 @@ +MAJOR | high | §F item 12 "✓ Codex Gate B satisfied" | The proposed replacement changes only the agent-facing `additionalContext`, leaving the live `systemMessage` and its exact fixture as `✓ Codex Gate B satisfied (...)`; that is a gate verdict the hook cannot establish, because §A3 gives Gate B no hook-backed content condition and §A1 makes closure depend on duties, holds, other conditions and a completed closing act the hook does not observe | An operator receives a green satisfied signal even when an undischarged Major, a standing hold, another unmet closure condition or an unperformed closing act keeps the cycle open | Replace the terse line with the hook state it actually observes, such as `Gate B hook checks passed`, and update its exact and loose test expectations +END OF FINDINGS (1 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index f5d4eba..82881fb 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -77,7 +77,8 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 58 | ebe3aa7 | 5→**6** | 0→**0** | 5→**4** | yes | one tell (findings rose 5→6). All four Majors in items 15–17, one layer deeper than pass 57 reached: both fingerprint reminders still **opened** "STOP — Codex Gate B not satisfied", a gate verdict the hook cannot reach; item 16 claimed a permitted pass "records" a usable fingerprint where the store can fail again; item 17's enumeration omitted the below-floor zero-finding closing route. Minor 6 was design §9's "one exact-match expectation" — the **second** time that bullet went stale under a propagation | | 59 | fecb3aa | 6→**9** | 0→**0** | 4→**5** | yes | **MANDATORY TWO-TELL STOP** — findings rose 5→6→9, and four of nine were about §F's *prose describing* its replacements rather than the replacements. Surfaced, standing answer applied, loop continued. Structural repair: items 15 and 16 now give the **complete resulting message** for both channels instead of a span plus prose about where it goes. Finding 7 is a **new behaviour question** and is open with Daniel: conflicting accept/decline on the same finding in two `full` Gate-B branch files | | 60 | e911b39 | 9→**8** | 0→**0** | 5→**5** | yes | zero tells. **Finding 3 is the one that mattered:** item 15's complete message carried a backtick around `WIP`, which is command substitution inside the hook's double-quoted `note` argument — installed literally it would run `WIP`, corrupt the reminder and fail ShellCheck. Only visible because the message is now written out whole. §F now states the no-shell-active-characters constraint. Also: the standing work-loop sequence names a commit only after Gate B → item 18, §F 22 → 23; item 11 conditioned proceeding on a clean pass rather than on the cycle having closed; item 16 called an unchanged fingerprint a machinery fault while the index can move it. 3 Minors + 1 Nit **collected, not repaired** | -| 61 | — | — | — | — | not run | next, against the pass-60 repair commit. **If 61 has zero Blockers and zero Majors the cycle closes** — Minors are collected, per Mechanics · Severity | +| 61 | 5197b0c | 8→**1** | 0→**0** | 5→**1** | yes | zero tells. **One finding.** Item 12 had only its long message repaired; the terse operator line "✓ Codex Gate B satisfied" stayed — a gate verdict the hook cannot establish, and the one systemMessage items 15 and 16 had already taught me to check. Now "hook checks passed". Its fixture moves with it: `expected_msg` scope 2 → 3 | +| 62 | — | — | — | — | not run | next, against the pass-61 repair commit. **Zero Blockers and zero Majors closes the cycle** — Minors collected, per Mechanics · Severity | | 40 | 12cf247 | 5→**3** | 0→**0** | 4→**3** | yes | zero tells; first round with no fan-out Major after the three-site check was run BEFORE the pass | | 41 | d26d097 | 3→**2** | 0→**0** | 3→**2** | yes | zero tells; both Majors were compressed pointers of mine dropping a load-bearing part of a standing rule | | 42 | 3b61fe3 | 2→**1** | 0→**0** | 2→**1** | yes | zero tells; §F item 4 had taken back the enumeration it was repaired to avoid | @@ -86,6 +87,22 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 45 | 1c858ac | 7→**7** | 0→**0** | 3→**1** | yes | zero tells; swept all four sites still tying classification to closing | | 46 | 4b713dd | 7→**9** | 0→**1** | 1→**2** | yes | **TWO TELLS — mandatory stop under §5, answered in advance by Daniel's standing "no stops without an absolute block" of 2026-09-13 17:02.** Tells surfaced here and in the commit body, findings kept open, no pass credited clean; the loop continued on that answer. **Eleventh** falsified standing sentence; session 01a09b7c-dd6b-7cb3-b83a-776599286f9b | +## Pass-61 report — one finding, and it was a gap I had left myself + +**Trend:** findings 9, 8, **1** across 59–61; Blockers 0, 0, **0**; Majors 5, 5, **1**. +**Cluster:** product behaviour, one finding. **Require↔withdraw:** none. **Zero tells.** + +Items 15 and 16 replace both channels of their message — the long `additionalContext` and the +terse `systemMessage` an operator sees. Item 12 replaced only the long one, leaving +"✓ Codex Gate B satisfied". That is a gate verdict the hook cannot establish: Gate B has no +hook-backed content condition, and closure turns on duties, holds, the remaining conditions and a +completed closing act, none of which the hook observes. A green "satisfied" could stand over an +undischarged Major. It now reads "hook checks passed" — what the hook knows — and its exact +fixture moves with it, taking the `expected_msg` scope from two to three. + +**The counts in that line are kept.** `$passes/$floor cycle, $fresh on current fingerprint` are +observations the hook does make; only the verdict was removed. + ## Pass-60 report — zero tells, and the span repair paying for itself **Trend:** findings 6, 9, **8** across 58–60; Blockers 0, 0, **0**; Majors 4, 5, **5**. diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index bb18bce..701eeb1 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -8,7 +8,7 @@ copy, and a value copied here would be a remembered value. Prompt-only, in two mirrored copies — `CLAUDE.md` §5 (**C** below) and the inline template in `plugins/dev-workflow/commands/workflow-init.md` (**W** below) — **plus seven reminder strings in `plugins/dev-workflow/hooks/codex-gate.sh` and their exact-match expectations in -`plugins/dev-workflow/hooks/codex-gate.test.sh` — three `expected_ctx` and two `expected_msg`** (target text §F items +`plugins/dev-workflow/hooks/codex-gate.test.sh` — three `expected_ctx` and three `expected_msg`** (target text §F items 10–13 and 15–17, authorised 2026-09-13). **No hook behaviour changes**, and the hook strings have no mirror, so they carry no parity obligation. Condition ids `a1`…`j4` are defined in `docs/superpowers/specs/2026-09-10-loop-rule-consolidation-condition-inventory.md` beside this @@ -354,9 +354,9 @@ repo's most persistent defect. The transport that could carry it left with the r seven `note` **strings** (§F items 10–13 and 15–17) and nowhere else: no control flow, no counter, no fingerprint computation, no routing, so the POSIX-`sh` and optional-`jq` obligations are not reached. `plugins/dev-workflow/hooks/codex-gate.test.sh` changes in all three of its - `expected_ctx` exact-match expectations — the Gate-B satisfied, stale-fingerprint and - no-fingerprint messages, §F items 12, 16 and 15 — each replaced with the complete resulting - message; the remaining hook assertions match loose patterns these repairs leave standing. The §5 heading the hook greps (`Cross-Model Review`) + `expected_ctx` exact-match expectations and all three `expected_msg` beside them — the Gate-B + satisfied, stale-fingerprint and no-fingerprint messages, §F items 12, 16 and 15 — each replaced + with the complete resulting text; the remaining hook assertions match loose patterns these repairs leave standing. The §5 heading the hook greps (`Cross-Model Review`) does not move. **Every path in this spec is written repository-relative and in full** — no ellipsis shorthand diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index 1500b11..345b581 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -682,8 +682,8 @@ outside the byte-identical parity rule the section opening states. The hook hold gate reminders in all; the one this change leaves alone is the docs-only notice, which states no closure permission. **Three of the seven are pinned by exact-match expectations in `plugins/dev-workflow/hooks/codex-gate.test.sh`** — items 12, 15 and 16, at that file's three -`expected_ctx` assignments, **and items 15 and 16 also at the two `expected_msg` assignments beside -them**, since those items replace the terse operator line as well — each replaced there with its +`expected_ctx` assignments, **and items 12, 15 and 16 also at the three `expected_msg` assignments +beside them**, since those items replace the terse operator line as well — each replaced there with its complete resulting text in the same change. The remaining four are matched by loose patterns these repairs leave standing, **except the stale-branch assertions item 16 moves with it**. **Every hook replacement below is destined for a double-quoted POSIX-shell `note` argument**, so @@ -934,10 +934,20 @@ Gate-B tree-equality material, and no condition here reaches it. ``` Per $policy, commit only if your final pass was clean and every other closure condition holds, both as it defines them. ``` +``` +✓ Codex Gate B hook checks passed ($passes/$floor cycle, $fresh on current fingerprint) +``` *Why (pass 53 finding 20):* the same abbreviated definition, in the branch an author reads immediately before committing, and it takes the same repair. It is the thirteenth sentence sharing this section's mechanism. This message is one of the three pinned by an exact-match expectation, replaced there with the complete resulting message as the section opening requires. +*And (pass 61 finding 1):* the terse `systemMessage` still read "✓ Codex Gate B satisfied", a gate +verdict the hook cannot establish — Gate B has no hook-backed content condition, and closure turns +on duties, holds, the remaining conditions and a completed closing act, none of which the hook +observes. An operator reading a green "satisfied" could see it over an undischarged Major or an +unperformed act. It now names **what the hook checked**, which is the whole of what it knows. Its +exact fixture moves with it. The counts either side of this message are kept: they are +observations, not verdicts. **13. The WIP-commit reminder's closing permission** (the shipped hook, same file). ``` From ef00b29eb47e1c1718e6458826681b7d09a86ef6 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 10:32:37 +0200 Subject: [PATCH 093/181] docs(specs): apply pass 62; item 12's long channel completes the transition MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Pass 61 converted item 12's systemMessage and the §F opening began promising a complete resulting text for items 12, 15 and 16, but item 12's additionalContext block still held only the replaced sentence. A plan reading the opening would have installed that sentence as the whole message and deleted the two before it — including the parked fingerprint overclaim, which is meant to survive untouched. The block now carries the complete message with those two sentences verbatim, and says that reproducing the overclaim is not endorsing it. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-62.md | 2 ++ .../gate-a-spec-awsf1ec771-resume.md | 20 ++++++++++++++++++- ...-10-loop-rule-consolidation-target-text.md | 10 ++++++---- 3 files changed, 27 insertions(+), 5 deletions(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-62.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-62.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-62.md new file mode 100644 index 0000000..736da0f --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-62.md @@ -0,0 +1,2 @@ +MAJOR | high | §F item 12 "Its fingerprint clause is untouched" | Item 12 says both channels are complete resulting text and that the fingerprint clause stays untouched, but its additionalContext block contains only "Per $policy, commit only if your final pass was clean and every other closure condition holds, both as it defines them."; the live reminder's two preceding sentences at plugins/dev-workflow/hooks/codex-gate.sh:956, beginning "Codex Gate B: $passes/$floor pass(es) this cycle" and "The floor counts the cycle", are absent. | Installing the block as the complete message silently deletes the parked fingerprint/count text, while treating it as a replacement span contradicts the full-message instruction and leaves the exact expected_ctx result to be invented. | Replace the block with the complete resulting additionalContext: preserve the live first two sentences verbatim and replace only the final clean-definition sentence; keep the supplied systemMessage. +END OF FINDINGS (1 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index 82881fb..b246362 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -78,7 +78,8 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 59 | fecb3aa | 6→**9** | 0→**0** | 4→**5** | yes | **MANDATORY TWO-TELL STOP** — findings rose 5→6→9, and four of nine were about §F's *prose describing* its replacements rather than the replacements. Surfaced, standing answer applied, loop continued. Structural repair: items 15 and 16 now give the **complete resulting message** for both channels instead of a span plus prose about where it goes. Finding 7 is a **new behaviour question** and is open with Daniel: conflicting accept/decline on the same finding in two `full` Gate-B branch files | | 60 | e911b39 | 9→**8** | 0→**0** | 5→**5** | yes | zero tells. **Finding 3 is the one that mattered:** item 15's complete message carried a backtick around `WIP`, which is command substitution inside the hook's double-quoted `note` argument — installed literally it would run `WIP`, corrupt the reminder and fail ShellCheck. Only visible because the message is now written out whole. §F now states the no-shell-active-characters constraint. Also: the standing work-loop sequence names a commit only after Gate B → item 18, §F 22 → 23; item 11 conditioned proceeding on a clean pass rather than on the cycle having closed; item 16 called an unchanged fingerprint a machinery fault while the index can move it. 3 Minors + 1 Nit **collected, not repaired** | | 61 | 5197b0c | 8→**1** | 0→**0** | 5→**1** | yes | zero tells. **One finding.** Item 12 had only its long message repaired; the terse operator line "✓ Codex Gate B satisfied" stayed — a gate verdict the hook cannot establish, and the one systemMessage items 15 and 16 had already taught me to check. Now "hook checks passed". Its fixture moves with it: `expected_msg` scope 2 → 3 | -| 62 | — | — | — | — | not run | next, against the pass-61 repair commit. **Zero Blockers and zero Majors closes the cycle** — Minors collected, per Mechanics · Severity | +| 62 | 888d2d1 | 1→**1** | 0→**0** | 1→**1** | yes | zero tells. One finding, the last thread of the same repair: item 12's `additionalContext` block held only the replaced sentence while the §F opening now promises a complete message for items 12, 15 and 16. Installing it literally would have deleted the two parked sentences. Now complete, the parked pair carried verbatim | +| 63 | — | — | — | — | not run | next, against the pass-62 repair commit. **Zero Blockers and zero Majors closes the cycle** | | 40 | 12cf247 | 5→**3** | 0→**0** | 4→**3** | yes | zero tells; first round with no fan-out Major after the three-site check was run BEFORE the pass | | 41 | d26d097 | 3→**2** | 0→**0** | 3→**2** | yes | zero tells; both Majors were compressed pointers of mine dropping a load-bearing part of a standing rule | | 42 | 3b61fe3 | 2→**1** | 0→**0** | 2→**1** | yes | zero tells; §F item 4 had taken back the enumeration it was repaired to avoid | @@ -87,6 +88,23 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 45 | 1c858ac | 7→**7** | 0→**0** | 3→**1** | yes | zero tells; swept all four sites still tying classification to closing | | 46 | 4b713dd | 7→**9** | 0→**1** | 1→**2** | yes | **TWO TELLS — mandatory stop under §5, answered in advance by Daniel's standing "no stops without an absolute block" of 2026-09-13 17:02.** Tells surfaced here and in the commit body, findings kept open, no pass credited clean; the loop continued on that answer. **Eleventh** falsified standing sentence; session 01a09b7c-dd6b-7cb3-b83a-776599286f9b | +## Pass-62 report — one finding, the last thread of the complete-message repair + +**Trend:** findings 8, 1, **1** across 60–62; Blockers 0, 0, **0**; Majors 5, 1, **1**. +**Cluster:** product behaviour. **Require↔withdraw:** none. **Zero tells.** + +Pass 61 took item 12's `systemMessage` to a complete replacement and the §F opening began +promising a complete resulting text for items 12, 15 and 16. Item 12's `additionalContext` block +had never been converted: it still held only the one sentence the repair replaces. A plan reading +the opening would have installed that sentence **as the whole message** and silently deleted the +two sentences before it — including the parked fingerprint overclaim, which is supposed to survive +untouched. The block now carries the complete message with those two sentences verbatim, and says +that reproducing the overclaim is not endorsing it. + +**This is the third and last item to make the span-to-complete-message transition**, which is why +the same defect appeared three passes running in three different items and has now run out of +places to be. + ## Pass-61 report — one finding, and it was a gap I had left myself **Trend:** findings 9, 8, **1** across 59–61; Blockers 0, 0, **0**; Majors 5, 5, **1**. diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index 345b581..a484361 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -928,11 +928,13 @@ having closed**, which is the one fact that covers all of them. The repair **cit what keeps a later change to the definition from falsifying this string again. It is the twelfth sentence sharing this section's mechanism. -**12. The Gate-B satisfied reminder's clean definition** (the shipped hook, same file). **Its -fingerprint clause is untouched**: that overclaim predates this change, is parked with the -Gate-B tree-equality material, and no condition here reaches it. +**12. The Gate-B satisfied reminder, both channels** (the shipped hook, same file). **The complete +resulting text, for the reason item 15 gives.** Its first two sentences are **carried verbatim**: +the fingerprint overclaim in the second predates this change, is parked with the Gate-B +tree-equality material, and no condition here reaches it — so it is reproduced rather than edited, +and reproducing it is not endorsing it. ``` -Per $policy, commit only if your final pass was clean and every other closure condition holds, both as it defines them. +Codex Gate B: $passes/$floor pass(es) this cycle, of which $fresh cover the CURRENT content fingerprint (unchanged since that review). The floor counts the cycle; only the fresh pass(es) carry the same fingerprint as what you are committing. Per $policy, commit only if your final pass was clean and every other closure condition holds, both as it defines them. ``` ``` ✓ Codex Gate B hook checks passed ($passes/$floor cycle, $fresh on current fingerprint) From 38042c8e6f4ea6018b1a889f46b13d6aeda540b7 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 10:45:35 +0200 Subject: [PATCH 094/181] docs(specs): apply pass 63; stop counting what this text does not survey MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Two loop-health tells were present and the stop was mandatory: findings rose 1 -> 1 -> 3 and two of three cluster on the test instrument. Surfaced; standing answer applied and the loop continued. §F has three times claimed which parts of codex-gate.test.sh a change reaches and been wrong every time: "the exact-match expectation" where there are three (pass 57), and "the remaining four are matched by loose patterns these repairs leave standing" where "Gate B satisfied" occurs 21 times and "STOP" 14 (both counted this pass). Each repair replaced one wrong count with another claim of the same kind, so the claim is now gone rather than corrected: the plan owes a sweep of that file for every assertion, label and comment naming a replaced string, and §F states that no count of them appears anywhere, with the two wrong counts as the reason. Item 15 also carried "you MUST reach a minimum of $floor passes" unconditionally, in a branch reachable after a zero-finding pass — eligible below the floor — whose first store failed. The floor now travels with the ordering pointer. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-63.md | 4 +++ .../gate-a-spec-awsf1ec771-resume.md | 33 ++++++++++++++++++- ...26-09-10-loop-rule-consolidation-design.md | 4 ++- ...-10-loop-rule-consolidation-target-text.md | 18 ++++++++-- 4 files changed, 54 insertions(+), 5 deletions(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-63.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-63.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-63.md new file mode 100644 index 0000000..c446374 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-63.md @@ -0,0 +1,4 @@ +MAJOR | high | §F item 15 "Per $policy you MUST reach a minimum of $floor passes per cycle" | The complete no-fingerprint reminder makes the floor unconditional even though §A1 makes a zero-finding pass eligible below the floor; this branch is reachable after such a pass when its first fingerprint store fails, its state cannot be read back, or a non-WIP closing attempt clears the state while the cycle stays open | The reminder gives a clean eligible cycle a MUST that contradicts the ordering and can force unowed passes instead of letting the cycle reach or retry its closing act | Remove the unconditional floor sentence and leave both floor eligibility and the next action to the pointer to the closure ordering entire +MAJOR | high | §F item 12 "✓ Codex Gate B hook checks passed" | The replacement removes the literal `Gate B satisfied`, but eighteen loose assertions in `plugins/dev-workflow/hooks/codex-gate.test.sh` still grep for that verdict; §F moves the three exact fixtures and item 16's stale-branch loose assertions but leaves these item-12 expectations live | Installing the specified message changes leaves the required hook suite failing across its satisfied-state behavior cases and preserves a test vocabulary that still calls the gate satisfied | Move every satisfied-branch loose assertion with item 12 and make each test the observed hook state, such as `hook checks passed`, updating verdict-bearing test labels and comments with it +MAJOR | high | §F item 15 "Codex gate state: no fingerprint is recorded" | The neutral opening removes the literal `STOP`, but the live marker-only assertion at `plugins/dev-workflow/hooks/codex-gate.test.sh:713` requires this no-fingerprint output to contain `STOP`, and §F names only item 15's exact fixtures for replacement | The required hook suite still fails after the specified fixture edits and continues to demand the gate verdict this item deliberately removes | Move that loose assertion with item 15 and test the observed no-fingerprint state plus delivery of the generic-policy reminder instead of `STOP` +END OF FINDINGS (3 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index b246362..aac2074 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -79,7 +79,8 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 60 | e911b39 | 9→**8** | 0→**0** | 5→**5** | yes | zero tells. **Finding 3 is the one that mattered:** item 15's complete message carried a backtick around `WIP`, which is command substitution inside the hook's double-quoted `note` argument — installed literally it would run `WIP`, corrupt the reminder and fail ShellCheck. Only visible because the message is now written out whole. §F now states the no-shell-active-characters constraint. Also: the standing work-loop sequence names a commit only after Gate B → item 18, §F 22 → 23; item 11 conditioned proceeding on a clean pass rather than on the cycle having closed; item 16 called an unchanged fingerprint a machinery fault while the index can move it. 3 Minors + 1 Nit **collected, not repaired** | | 61 | 5197b0c | 8→**1** | 0→**0** | 5→**1** | yes | zero tells. **One finding.** Item 12 had only its long message repaired; the terse operator line "✓ Codex Gate B satisfied" stayed — a gate verdict the hook cannot establish, and the one systemMessage items 15 and 16 had already taught me to check. Now "hook checks passed". Its fixture moves with it: `expected_msg` scope 2 → 3 | | 62 | 888d2d1 | 1→**1** | 0→**0** | 1→**1** | yes | zero tells. One finding, the last thread of the same repair: item 12's `additionalContext` block held only the replaced sentence while the §F opening now promises a complete message for items 12, 15 and 16. Installing it literally would have deleted the two parked sentences. Now complete, the parked pair carried verbatim | -| 63 | — | — | — | — | not run | next, against the pass-62 repair commit. **Zero Blockers and zero Majors closes the cycle** | +| 63 | ef00b29 | 1→**3** | 0→**0** | 1→**3** | yes | **two tells** (findings rose 1→1→3; two of three cluster on the test instrument). Surfaced, standing answer applied, loop continued. **Third wrong claim about which test assertions a change touches** — "Gate B satisfied" occurs 21 times and "STOP" 14, against §F's claim that loose patterns were left standing. Claim replaced by a **duty on the plan to sweep**, and §F now states that no count of them appears anywhere, with the two wrong counts as the reason. Item 15 also carried an unconditional floor MUST a zero-finding pass does not owe | +| 64 | — | — | — | — | not run | next, against the pass-63 repair commit. **Zero Blockers and zero Majors closes the cycle** | | 40 | 12cf247 | 5→**3** | 0→**0** | 4→**3** | yes | zero tells; first round with no fan-out Major after the three-site check was run BEFORE the pass | | 41 | d26d097 | 3→**2** | 0→**0** | 3→**2** | yes | zero tells; both Majors were compressed pointers of mine dropping a load-bearing part of a standing rule | | 42 | 3b61fe3 | 2→**1** | 0→**0** | 2→**1** | yes | zero tells; §F item 4 had taken back the enumeration it was repaired to avoid | @@ -88,6 +89,36 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 45 | 1c858ac | 7→**7** | 0→**0** | 3→**1** | yes | zero tells; swept all four sites still tying classification to closing | | 46 | 4b713dd | 7→**9** | 0→**1** | 1→**2** | yes | **TWO TELLS — mandatory stop under §5, answered in advance by Daniel's standing "no stops without an absolute block" of 2026-09-13 17:02.** Tells surfaced here and in the commit body, findings kept open, no pass credited clean; the loop continued on that answer. **Eleventh** falsified standing sentence; session 01a09b7c-dd6b-7cb3-b83a-776599286f9b | +## Pass-63 report — TWO TELLS, and the third wrong claim of one kind + +**Trend:** findings 1, 1, **3** across 61–63; Blockers 0, 0, **0**; Majors 1, 1, **3**. +**Cluster:** two of three findings are about `codex-gate.test.sh`, the instrument. +**Require↔withdraw:** none. **Two tells — the stop is mandatory**, surfaced, and the loop continued +on Daniel's standing answer of 2026-09-13. + +**The instrument cluster is legitimate severity, not noise.** CLAUDE.md's carve-out keeps an +instrument finding's severity where the instrument changes what a gate concludes; here the hook +suite would fail outright after installing the specified edits — a false red blocking a valid +change. Both were verified by count before repair. + +**The thing worth recording is that this is the third instance of one defect.** §F has three times +made a claim about which parts of `codex-gate.test.sh` a change reaches, and has been wrong every +time: "the exact-match expectation" where there are three (pass 57); "the remaining four are +matched by loose patterns these repairs leave standing" where `Gate B satisfied` occurs 21 times +and `STOP` 14 (pass 63, both counted). Each repair replaced one wrong count with another claim of +the same kind. + +**So the claim is gone rather than corrected.** §F now places a **duty on the plan** — sweep that +file for every assertion, label and comment naming a replaced string, move each with its item — and +states that **no count of them appears anywhere**, giving the two wrong counts as the reason. This +is the file's own stated preference (remove an enumeration rather than correct it; prefer a duty +over a claim about what a mechanism covers) applied to the place it kept failing. + +**Finding 1 was substantive and separate:** item 15's message carried "you MUST reach a minimum of +$floor passes" unconditionally, and that branch is reachable after a zero-finding pass — eligible +below the floor — whose first store failed. A clean eligible cycle was being handed a MUST the +ordering does not impose. The floor now travels with the pointer. + ## Pass-62 report — one finding, the last thread of the complete-message repair **Trend:** findings 8, 1, **1** across 60–62; Blockers 0, 0, **0**; Majors 5, 1, **1**. diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index 701eeb1..46c10aa 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -356,7 +356,9 @@ repo's most persistent defect. The transport that could carry it left with the r reached. `plugins/dev-workflow/hooks/codex-gate.test.sh` changes in all three of its `expected_ctx` exact-match expectations and all three `expected_msg` beside them — the Gate-B satisfied, stale-fingerprint and no-fingerprint messages, §F items 12, 16 and 15 — each replaced - with the complete resulting text; the remaining hook assertions match loose patterns these repairs leave standing. The §5 heading the hook greps (`Cross-Model Review`) + with the complete resulting text, **plus every other assertion, label or comment in that file + that tests or names a replaced string, which the plan finds by sweeping rather than from a list + here**; the remaining hook assertions match loose patterns these repairs leave standing. The §5 heading the hook greps (`Cross-Model Review`) does not move. **Every path in this spec is written repository-relative and in full** — no ellipsis shorthand diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index a484361..67ceecd 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -684,8 +684,15 @@ closure permission. **Three of the seven are pinned by exact-match expectations `plugins/dev-workflow/hooks/codex-gate.test.sh`** — items 12, 15 and 16, at that file's three `expected_ctx` assignments, **and items 12, 15 and 16 also at the three `expected_msg` assignments beside them**, since those items replace the terse operator line as well — each replaced there with its -complete resulting text in the same change. The remaining four are matched by loose patterns these -repairs leave standing, **except the stale-branch assertions item 16 moves with it**. +complete resulting text in the same change. +**Which other assertions move is a duty on the plan and is not enumerated here.** The plan sweeps +`plugins/dev-workflow/hooks/codex-gate.test.sh` for **every** assertion, label and comment that +tests or names a string any item below replaces — the loose `grep` assertions included — and moves +each with its item, testing the observed hook state rather than a gate verdict. **No count of them +appears anywhere**, and the reason is evidence rather than taste: this section twice stated one and +was twice wrong. It said one exact-match expectation where there are three (pass 57), and that the +remaining items' assertions could be left standing where "Gate B satisfied" occurs 21 times and +"STOP" 14 (pass 63). A count over a file this text does not survey is a claim it cannot keep. **Every hook replacement below is destined for a double-quoted POSIX-shell `note` argument**, so its text carries no backtick, no `$(`, no backslash and no double quote; `$policy`, `$floor` and `$passes` are the intended interpolations and stay. A backtick reached one of these blocks at pass @@ -980,7 +987,7 @@ statement about scope, so it does not join that count. resulting message is given, not a span**: a span plus a sentence about where it goes was found three times running to leave a fragment ambiguous or a tail standing. ``` -Codex gate state: no fingerprint is recorded for this cycle. The hook cannot tell why — no mcp__codex__review has run, the last one's fingerprint could not be written or read back, or a non-WIP commit attempt cleared it while the cycle itself stayed open. Per $policy you MUST reach a minimum of $floor passes per cycle. What this cycle does next is $policy's closure ordering's, read there entire, and this reminder decides none of it. If this repeats, check that .context/ and the state file inside it are readable and writable; if the file exists but is unreadable or empty, delete it — which restores no passes, and lets the next pass the ordering permits record a fingerprint. +Codex gate state: no fingerprint is recorded for this cycle. The hook cannot tell why — no mcp__codex__review has run, the last one's fingerprint could not be written or read back, or a non-WIP commit attempt cleared it while the cycle itself stayed open. What this cycle does next, the floor it owes included, is $policy's closure ordering's, read there entire, and this reminder decides none of it. If this repeats, check that .context/ and the state file inside it are readable and writable; if the file exists but is unreadable or empty, delete it — which restores no passes, and lets the next pass the ordering permits record a fingerprint. ``` ``` ⚠ Codex Gate B: no recorded fingerprint @@ -1003,6 +1010,11 @@ could order a review over a valid pass whose store failed. Its exact fixture mov cause list also omitted the one §A3 itself creates: a non-`WIP` commit attempt clears the fingerprint while the cycle stays open, and a reader sent to permission checks for it would be looking in the wrong place. +*And (pass 63 finding 1):* the message also carried "Per $policy you MUST reach a minimum of $floor +passes per cycle" unconditionally, which a zero-finding pass does not owe — and this branch is +reachable after exactly such a pass, when its first store fails or a non-WIP attempt clears the +state. A clean eligible cycle was being given a MUST the ordering does not impose. The floor now +travels with the pointer, where its own eligibility rules are. **16. The stale-fingerprint reminder, both channels** (the shipped hook, same file). **The complete resulting message, for the reason item 15 gives.** From 0593083e3473cd5c2f2bc30af4e2f3c4840c75b2 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 10:59:17 +0200 Subject: [PATCH 095/181] docs(specs): apply pass 64; the deleted claim's second copy in the design MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Pass 63 replaced §F's claim that the remaining hook assertions could be left standing with a sweep duty on the plan. The design carried the same sentence at a second site and only the first was updated. Deleted, so §F's duty is the single source. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-spec-awsf1ec771-pass-64.md | 2 ++ .../gate-a-spec-awsf1ec771-resume.md | 17 ++++++++++++++++- ...2026-09-10-loop-rule-consolidation-design.md | 2 +- 3 files changed, 19 insertions(+), 2 deletions(-) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-64.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-64.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-64.md new file mode 100644 index 0000000..8d0bcd5 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-64.md @@ -0,0 +1,2 @@ +MAJOR | high | §F "Which other assertions move is a duty on the plan" | The design's §8 repeats the superseded claim "the remaining hook assertions match loose patterns these repairs leave standing", even though item 12 removes `Gate B satisfied` from both channels and item 15 removes `STOP` while live loose assertions require those verdict strings | A plan can follow the design and leave verdict-based assertions and labels unchanged, causing the hook suite to fail and preserving tests that assert conclusions the hook no longer reports | Delete the design's "remaining hook assertions" clause so §F's sweep duty is the single source +END OF FINDINGS (1 total) diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md index aac2074..fda72bc 100644 --- a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md @@ -80,7 +80,8 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 61 | 5197b0c | 8→**1** | 0→**0** | 5→**1** | yes | zero tells. **One finding.** Item 12 had only its long message repaired; the terse operator line "✓ Codex Gate B satisfied" stayed — a gate verdict the hook cannot establish, and the one systemMessage items 15 and 16 had already taught me to check. Now "hook checks passed". Its fixture moves with it: `expected_msg` scope 2 → 3 | | 62 | 888d2d1 | 1→**1** | 0→**0** | 1→**1** | yes | zero tells. One finding, the last thread of the same repair: item 12's `additionalContext` block held only the replaced sentence while the §F opening now promises a complete message for items 12, 15 and 16. Installing it literally would have deleted the two parked sentences. Now complete, the parked pair carried verbatim | | 63 | ef00b29 | 1→**3** | 0→**0** | 1→**3** | yes | **two tells** (findings rose 1→1→3; two of three cluster on the test instrument). Surfaced, standing answer applied, loop continued. **Third wrong claim about which test assertions a change touches** — "Gate B satisfied" occurs 21 times and "STOP" 14, against §F's claim that loose patterns were left standing. Claim replaced by a **duty on the plan to sweep**, and §F now states that no count of them appears anywhere, with the two wrong counts as the reason. Item 15 also carried an unconditional floor MUST a zero-finding pass does not owe | -| 64 | — | — | — | — | not run | next, against the pass-63 repair commit. **Zero Blockers and zero Majors closes the cycle** | +| 64 | 38042c8 | 3→**1** | 0→**0** | 3→**1** | yes | zero tells. One finding: the claim pass 63 deleted from §F survived at a **second** design site — "the remaining hook assertions match loose patterns these repairs leave standing". Deleted; §F's sweep duty is the single source | +| 65 | — | — | — | — | not run | next, against the pass-64 repair commit. **Zero Blockers and zero Majors closes the cycle** | | 40 | 12cf247 | 5→**3** | 0→**0** | 4→**3** | yes | zero tells; first round with no fan-out Major after the three-site check was run BEFORE the pass | | 41 | d26d097 | 3→**2** | 0→**0** | 3→**2** | yes | zero tells; both Majors were compressed pointers of mine dropping a load-bearing part of a standing rule | | 42 | 3b61fe3 | 2→**1** | 0→**0** | 2→**1** | yes | zero tells; §F item 4 had taken back the enumeration it was repaired to avoid | @@ -89,6 +90,20 @@ Nothing depends on it; the pass files and the repo are authoritative where this | 45 | 1c858ac | 7→**7** | 0→**0** | 3→**1** | yes | zero tells; swept all four sites still tying classification to closing | | 46 | 4b713dd | 7→**9** | 0→**1** | 1→**2** | yes | **TWO TELLS — mandatory stop under §5, answered in advance by Daniel's standing "no stops without an absolute block" of 2026-09-13 17:02.** Tells surfaced here and in the commit body, findings kept open, no pass credited clean; the loop continued on that answer. **Eleventh** falsified standing sentence; session 01a09b7c-dd6b-7cb3-b83a-776599286f9b | +## Pass-64 report — one finding, the deleted claim's second copy + +**Trend:** findings 1, 3, **1** across 62–64; Blockers 0, 0, **0**; Majors 1, 3, **1**. +**Cluster:** product behaviour, one finding. **Require↔withdraw:** none. **Zero tells** — pass 63's +two did not return. + +Pass 63 removed §F's claim that the remaining hook assertions could be left standing, replacing it +with a sweep duty on the plan. The design carried the same sentence at a second site and I updated +only the first. Deleted, so §F's duty is the single source. + +**That is the cycle's oldest mechanism, caught in one pass rather than several:** a rule stated in +both the design and the target, repaired in one copy. It is worth noting that the repair-detection +loop is now finding each instance the pass after it is created rather than three passes later. + ## Pass-63 report — TWO TELLS, and the third wrong claim of one kind **Trend:** findings 1, 1, **3** across 61–63; Blockers 0, 0, **0**; Majors 1, 1, **3**. diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index 46c10aa..32e93bc 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -358,7 +358,7 @@ repo's most persistent defect. The transport that could carry it left with the r satisfied, stale-fingerprint and no-fingerprint messages, §F items 12, 16 and 15 — each replaced with the complete resulting text, **plus every other assertion, label or comment in that file that tests or names a replaced string, which the plan finds by sweeping rather than from a list - here**; the remaining hook assertions match loose patterns these repairs leave standing. The §5 heading the hook greps (`Cross-Model Review`) + here**. The §5 heading the hook greps (`Cross-Model Review`) does not move. **Every path in this spec is written repository-relative and in full** — no ellipsis shorthand From ba15e83088b7524744b9ed3ef3c177197466c987 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 11:11:06 +0200 Subject: [PATCH 096/181] docs(specs): close the Gate-A spec cycle; pass 65 returned NO FINDINGS The target text and its design are approved. Pass 65 read docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md at 0593083 against both prompt copies, both channels of every codex-gate.sh reminder, the expectations that pin them, the story, the successor story, the condition inventory and AGENTS.md, and returned a zero-finding file. The tree is unchanged since that read. cycle awsf1ec771; floor 3 per {docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md (level 2)}; hook reminder threshold absent cycle awsf1ec771; Gate-A spec (passes 1-65, codex): Findings 24,17,12,18,17,16,14,12,14,20,20,21,16,10,12,8,9,15,13,9,6,13,9,8,10,14,8,13,10,9,13,15,22,19,20,9,3,10,5,3,2,1,6,7,7,9,11,7,10,25,23,29,21,13,18,13,5,6,9,8,1,1,3,1,0. Blockers 5,2,4,1,0,4,0,1,2,2,1,1,2,1,1,1,1,0,1,1,0,1,0,1,0,1,1,4,0,0,0,1,1,0,0,0,1,0,0,0,0,0,0,0,0,1,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0. Majors 15,13,6,10,15,9,9,7,10,15,13,13,8,5,7,4,7,12,10,5,5,5,7,3,2,5,3,2,5,4,6,5,8,5,4,3,2,3,4,3,2,1,3,3,1,2,4,5,4,11,6,5,4,3,6,0,5,4,5,5,1,1,3,1,0. Counts for passes 47-49 were recovered from their findings files, the working record having no rows for them; every other pass is recorded from that record. No count is written as unknown. The working record is retired by rename rather than deletion: gate-a-spec-awsf1ec771-resume.md becomes gate-a-spec-awsf1ec771-history.md. The stated reason for retiring it is that a closed cycle's working record must not be adoptable as a later cycle's nonce candidate, and a name no recovery rule looks for satisfies that. Deleting it would have destroyed the only per-pass account of 65 passes, which the curve summarises and does not replace. Recorded here because Mechanics says delete. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- ...c-awsf1ec771-resume.md => gate-a-spec-awsf1ec771-history.md} | 0 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-65.md | 2 ++ 2 files changed, 2 insertions(+) rename .context/codex-reviews/{gate-a-spec-awsf1ec771-resume.md => gate-a-spec-awsf1ec771-history.md} (100%) create mode 100644 .context/codex-reviews/gate-a-spec-awsf1ec771-pass-65.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-resume.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-history.md similarity index 100% rename from .context/codex-reviews/gate-a-spec-awsf1ec771-resume.md rename to .context/codex-reviews/gate-a-spec-awsf1ec771-history.md diff --git a/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-65.md b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-65.md new file mode 100644 index 0000000..8d781f4 --- /dev/null +++ b/.context/codex-reviews/gate-a-spec-awsf1ec771-pass-65.md @@ -0,0 +1,2 @@ +NO FINDINGS +END OF FINDINGS (0 total) From 5e466f645f40ffd722328d38dc9eb54d4a19f538 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 11:13:45 +0200 Subject: [PATCH 097/181] docs(plans): the loop-rule-consolidation implementation plan MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Fifteen tasks plus a baseline task, against the target text approved at ba15e83. Carries the disposition for all 135 inventoried conditions by passage, which is story acceptance criterion 5, and the next-state table's specification, which is criterion 4. The plan cites the target text's fenced blocks rather than copying them. Copying would create the second-copy defect the 65-pass Gate-A cycle spent most of its findings on; the target text travels with the plan as its Spec, so a citation into it is not a placeholder. Self-review found §G had no task — the one-contract paragraph, which design §4 lists as its own site. Folded into Task 7, which now names nine sites. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../2026-09-14-loop-rule-consolidation.md | 953 ++++++++++++++++++ 1 file changed, 953 insertions(+) create mode 100644 docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md new file mode 100644 index 0000000..3fe4415 --- /dev/null +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -0,0 +1,953 @@ +# §5 loop-rule consolidation — Implementation Plan + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. + +**Goal:** Install one closure ordering into both §5 copies, replace the twenty-three standing sentences it falsifies across those copies and the shipped hook, and ship the result as plugin 0.12.0. + +**Architecture:** Every string this change installs is already written in final form in the target text. This plan does not restate any of it. Each task names the **site**, quotes the **anchor** it installs at, cites the **target-text section** whose fenced block is the bytes to install, and builds the **discriminating pair of counts** that shows the new wording present and the old wording gone. Copying the replacement text into this plan would create the second-copy defect the whole cycle fought; a citation into an approved artifact that travels with this plan is not a placeholder. + +**Tech Stack:** Markdown prompt text, POSIX `sh` (the hook), `grep`/`diff` for verification, `shellcheck`, the `claude` CLI. + +**Spec:** `docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md` (the text, in final form) and `docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md` (the decisions behind it). Both are approved: Gate-A cycle `awsf1ec771` closed at pass 65 with a zero-finding file, commit `ba15e83`. + +**Story:** `docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md` — read the profile from its header at every gate call; it is the only writable copy. Six acceptance criteria; §4 holds settled decisions D1–D8. + +--- + +## Global Constraints + +- **Both prompt copies take every NEW and REPLACED section byte-identical**, except §F's seven hook items, whose destination is the shipped hook and its test and which carry no parity obligation (target §"How to read a section", §F opening). +- **C** = `CLAUDE.md`. **W** = `plugins/dev-workflow/commands/workflow-init.md`. Every line number below is re-read at execution; the inventory's numbers cite `7c0d475` and have drifted. +- **Every hook replacement installs into a double-quoted POSIX-shell `note` argument** and therefore carries no backtick, no `$(`, no backslash and no double quote. `$policy`, `$floor`, `$passes`, `$passesA` and `$fresh` are the intended interpolations (target §F opening). +- **Install every replacement unwrapped** — as one line in the file — because a counted fragment must be single-line for `grep -F` to find it (design §7). +- **Invariant 5 (exact pinning)** and **invariant 12 (a plugin change requires a version bump)**: this change touches `plugins/`, so `plugins/dev-workflow/.claude-plugin/plugin.json` goes `0.11.0 → 0.12.0` with a CHANGELOG entry (design §8). +- **Invariant 11:** every prompt change passes all 12 items of `docs/prompt-standards.md`. Most at risk: item 6 (every constraint carries its reason in the same sentence) and item 8 (token-lean). +- **Invariant 4 / the hook:** no control flow, counter, fingerprint computation, routing or event handling changes. Only `note` strings and their expectations. +- **Gate-B cycle discipline:** snapshot commits are named `WIP: …`; the cycle closes by `git commit --amend`. A non-`WIP` commit mid-cycle resets the hook's counters. + +--- + +## File Structure + +| File | Responsibility in this change | +|---|---| +| `CLAUDE.md` | canonical §5 (and one §4 line). Receives §A–§E, §G, §H and §F's sixteen prompt-copy items. | +| `plugins/dev-workflow/commands/workflow-init.md` | the scaffolded mirror. Receives the same, byte-identical, minus the deliberate divergences the inventory records. | +| `plugins/dev-workflow/hooks/codex-gate.sh` | seven `note` strings, both channels each (§F items 10–13, 15–17). No behaviour change. | +| `plugins/dev-workflow/hooks/codex-gate.test.sh` | three `expected_ctx` and three `expected_msg` exact-match expectations, plus every other assertion, label or comment naming a replaced string. | +| `plugins/dev-workflow/.claude-plugin/plugin.json` | `version` `0.11.0 → 0.12.0`. | +| `plugins/dev-workflow/CHANGELOG.md` | the 0.12.0 entry, newest first. | +| this plan | the 135-condition disposition (below), the next-state table (Task 13), and the verification pairs each task builds. | + +--- + +## The condition disposition — all 135, by passage + +Story acceptance criterion 5 is satisfied here. Ids are `docs/superpowers/specs/2026-09-10-loop-rule-consolidation-condition-inventory.md`, snapshot at `7c0d475`. **Where the tree and the inventory disagree, the tree wins and the accounting is what needs correcting** — check each condition against the real file before marking it done. + +### Passage (a) — the floor paragraphs → target §H (Task 7) + +| Condition | Disposition | +|---|---| +| a1–a12, a14 | **kept**, untouched. The floor arithmetic, the hook-ratio rules and the two comparison points are outside this change. | +| a13 | **replaced** — scoped to its own paragraph. §H's `a13` block. | +| a15 | **carried** inside §H's `a16` block, which reproduces it so one contiguous string installs. | +| a16 | **replaced** — points at Mechanics · Severity instead of carrying an unscoped copy. §H's `a16` block. | +| a17, a18, a19 | **moved** — the floor paragraph stops stating the clean-final-pass rule and the early exit; both are stated once in §A. §H's `a17`–`a22` block. | +| a20 | **moved** to §A unchanged, beside the zero-finding rule it qualifies. | +| a21, a22 | **carried** unchanged, reproduced in §H's block for the same contiguity reason as a15. | + +### Passage (b) — what a loop absorbs → target §B (Task 3) + +§B states its own accounting and this table reproduces it: **Changed:** `b3`, `b7`, `b8`, `b11`, `b12`, `b13`, `b16`, `b17`, `b18`. **Added:** the closing-time change rule and decision 6's decline semantics, which no inventoried condition carried because none existed. **Carried:** `b1`, `b2`, `b4`, `b5`, `b6`, `b9`, `b10`, `b14`, `b15`. + +**Verify against the file, not against this table:** §B is written out whole and is the only place this change states passage (b). Read the installed passage and confirm each carried condition is present and each changed one is gone. + +### Passage (c) — recognizing clearly stuck → target §C (Task 4) and §H (Task 7) + +| Condition | Disposition | +|---|---| +| c1, c2, c3 | **kept** — the curve-reading sentences are untouched. | +| c4 | **replaced** — "a missing one means keep going" becomes "means only that *this* exit does not apply", the pass's actual next step being the ordering's. | +| c5, c6, c7 | **carried word for word** inside §C's block. | +| c8 | **changed** — gains the re-raised-dismissal clause. | +| c9 | **split.** The operative precedence clause moves into §A capitalized as a standalone sentence; the plateau rationale stays at this source. Not moved whole — §C says so explicitly. | +| c10, c11 | **moved** to §A, which states what a Blocker/Major-free pass at or above the floor does. | +| c12, c13 | **moved** to §A, beside a19. | +| c14 | **replaced, not moved.** The ordering splits the below-floor Minor case into suspend and continue; no copy of the live wording survives beside them. | +| c15, c16 | **replaced** — §H's `c18`-and-surfacing block. The hold is now over the cycle and the new hold. | +| c17 | **replaced** — the resolve rule now scopes to the assigned fix set and states what a validly dismissed recurrence owes. | +| c18 | **replaced** — the blanket no-clean-credit goes; a pass is credited on its own findings, and a scope-stop trigger is what withholds credit. | +| c19 | **replaced** — the one-answer resumption goes; what the answer does is the ordering's. | +| c20 | **carried**, with "Blocker and Major" narrowed to "**in-set** Blocker and Major". | + +### Passage (d) — from pass 4 onward (Task 0 check only) + +`d1`–`d7`: **all kept, untouched.** The unavailable-history block moved to the successor story with D10 and this change no longer edits this passage (design §5). **A diff touching C 255–261 or W 459–465 is a defect.** + +### Passage (e) — the five tells → target §D (Task 5) + +| Condition | Disposition | +|---|---| +| e1–e6 | **kept** — the five tells themselves are untouched. | +| e7 | **changed** — gains the read-after-clean-completion clause. The sentence is given entire in §D. | +| e8, e9, e10 | **carried** inside §D's block. | +| e11 | **kept** — the C-only rationale paragraph is untouched and stays C-only. | +| — | **added:** §D's pointer paragraph at the end of the passage. | + +### Passage (f) — the two rules above do not compete (Task 0 check only) + +`f1`–`f7`: **all kept, untouched.** "The two rules above" still names the absorb rule and the stuck reading; the block sits before both and adds no third rule between them (design §5). **A diff touching C 275–287 or W 473–483 is a defect, and so is inserting §A anywhere that would come between them.** + +### Passage (g) — Mechanics · Severity → target §E (Task 6) + +| Condition | Disposition | +|---|---| +| g1 | **replaced by the answer** — the demotion changes what a cycle must resolve, never what it observes. | +| g2 | **dropped** — the interim report-and-stop duty existed only until the question was settled, and it is settled. | +| g3 | **dropped with g2**, being that duty's justification. | +| g4 | **dropped (C only)** — the ownership sentence names the story this change discharges. **Removing it in C while leaving W's duty standing would desynchronise the copies in the opposite direction**, so g2/g3 must go from W in the same edit. This is the one deliberate story-path divergence and it disappears with this change. | + +### Passage (h) — recording a human exception → target §F (Tasks 8, 9) + +| Condition | Disposition | +|---|---| +| h1–h3, h5, h6, h8–h18, h20–h26 | **kept**, untouched. | +| h4 | **replaced** — §F item 7, the human-exception destination: a Gate-A cycle's record goes to the commit its closing act produces, not to "the spec or plan commit". | +| h7 | **kept.** | +| h19 | **replaced** — §F item 4, the scope sentence: "neither a human's **general** assent nor this record", plus the clause distinguishing the answers a suspension asks for from assent. | +| h13 | **kept.** Design §4 names this sentence as deliberately not edited: this change ships no record for the squash carry to carry. | + +### Passage (i) — when these rules bind → target §H (Task 7) + +`i1`, `i2`, `i3`, `i9`–`i11`, `i13`–`i16`: **kept**, untouched. +`i4`–`i8`: **kept**, and the dash-delimited list they sit in is **extended, not rewritten** — §H gives the whole list with the additions at the end. +`i12`: **discharged, and kept.** It is the extension point that licenses the addition; it stays because the next change needs it too. + +### Passage (j) — the squash carry (Task 0 check only) + +`j1`–`j4`: **all kept, untouched.** It was to name the answer record, which moved to the successor; this change ships no record for it (design §5). **A diff touching C 892 or W 1076 is a defect.** + +--- + +## Task 0: Establish the baseline and the do-not-touch set + +**Files:** none modified. + +**Interfaces:** +- Produces: `$BASE` (the parent commit every later task's counterfactual half runs against), and a recorded list of the five untouched passage ranges. + +- [ ] **Step 1: Confirm the approved artifacts and a clean tree** + +```bash +git log --oneline -1 # expect ba15e83 or later on loop-rule-consolidation +git status --porcelain # expect empty +BASE=$(git rev-parse HEAD) # every counterfactual count runs against this +echo "$BASE" +``` + +- [ ] **Step 2: Re-read the five untouched ranges and record their current line numbers** + +Passages (d), (f) and (j) are not edited by this change, and passage (a)'s arithmetic and passage (h)'s twenty-three kept conditions are mostly untouched. Record where they are now, because the inventory's numbers cite `7c0d475`: + +```bash +grep -n 'From pass 4 onward every pass report carries three lines' CLAUDE.md plugins/dev-workflow/commands/workflow-init.md +grep -n 'The two rules above do not compete' CLAUDE.md plugins/dev-workflow/commands/workflow-init.md +grep -n 'On squash-merge, copy every evidence entry' CLAUDE.md plugins/dev-workflow/commands/workflow-init.md +``` + +Expected: two hits each, one per copy. Write the six line numbers into a scratch note — Task 14 diffs against them. + +- [ ] **Step 3: Confirm the parity baseline of the inventoried ranges** + +```bash +diff <(sed -n '65,290p' CLAUDE.md) <(sed -n '264,489p' plugins/dev-workflow/commands/workflow-init.md) | head -40 +``` + +Expected: the deliberate divergences the inventory records (b3's cross-reference target, b's intensifier and field-mint parenthetical, e8's pronoun, e11, f5–f7's framing, g4). **Anything else is pre-existing drift — record it and raise it before editing**, because Task 14's parity diff cannot tell drift you introduced from drift you inherited. + +- [ ] **Step 4: Commit nothing** + +Task 0 produces a scratch note, not a commit. + +--- + +## Task 1: Install the closure ordering (§A) into both copies + +**Files:** +- Modify: `CLAUDE.md` — insert immediately before the line beginning `**What a loop absorbs, and what stops it` +- Modify: `plugins/dev-workflow/commands/workflow-init.md` — the same anchor +- Test: the counts in Step 4 + +**Interfaces:** +- Consumes: `$BASE` from Task 0. +- Produces: the installed block that every later task's pointers refer to. Later tasks cite it as "the closure ordering"; nothing depends on its line numbers. + +**The bytes:** target text §A, the three fenced blocks under `### A1 — the ordering`, `### A2` and `### A3`, in that order, as three paragraphs. §A's opening sentence states the placement: "sitting immediately before **What a loop absorbs, and what stops it** in both copies." + +**Counterfactual: ABSENT, and claimed as absent.** The parent carries none of the three paragraphs, so no old wording can be shown to disappear. **Presence alone is the check here**, and the evidence entry says so rather than implying a pair (design §7). + +- [ ] **Step 1: Find the anchor in both copies** + +```bash +grep -n '^\*\*What a loop absorbs, and what stops it' CLAUDE.md plugins/dev-workflow/commands/workflow-init.md +``` + +Expected: exactly one hit per file. + +- [ ] **Step 2: Verify the block is absent before installing** + +```bash +grep -cF 'How a cycle ends — one ordering, stated here and referenced everywhere else' CLAUDE.md plugins/dev-workflow/commands/workflow-init.md +``` + +Expected: `0` for both. If either is non-zero the block is already partly installed — stop and reconcile. + +- [ ] **Step 3: Install §A1, §A2 and §A3 in both copies** + +Insert the three blocks before the anchor line, blank-line separated, byte-identical in both files. **Do not reflow.** Install each paragraph unwrapped where the target text shows it unwrapped. + +**Placement constraint from the disposition table:** §A must not land between passage (b) and passage (c), because `f1` ("the two rules above") names those two and would then name the wrong pair. Inserting *before* (b) satisfies this. + +- [ ] **Step 4: Run the presence counts in both copies and both trees** + +```bash +for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do + printf '%s worktree: ' "$f" + grep -cF 'How a cycle ends — one ordering, stated here and referenced everywhere else' "$f" + printf '%s parent: ' "$f" + git show "$BASE:$f" | grep -cF 'How a cycle ends — one ordering, stated here and referenced everywhere else' +done +``` + +Expected: worktree `1`, parent `0`, for both files. A parent count above zero means `$BASE` is wrong. + +- [ ] **Step 5: Check parity of the installed block** + +```bash +diff <(sed -n '/^\*\*How a cycle ends/,/^\*\*What a loop absorbs/p' CLAUDE.md) \ + <(sed -n '/^\*\*How a cycle ends/,/^\*\*What a loop absorbs/p' plugins/dev-workflow/commands/workflow-init.md) +``` + +Expected: no output. + +- [ ] **Step 6: Commit** + +```bash +git add CLAUDE.md plugins/dev-workflow/commands/workflow-init.md +git commit -m "WIP: install the closure ordering into both §5 copies" +``` + +Named `WIP:` because Task 15 runs Gate B over the whole change and amends once. A non-`WIP` commit here would reset the hook's Gate-B counters mid-cycle. + +--- + +## Task 2: Verify the untouched passages are still untouched + +**Files:** none modified. + +**Interfaces:** +- Consumes: Task 0's recorded line numbers and `$BASE`. + +This task exists because `f1` and the (d)/(j) dispositions are falsifiable only by a diff, and the cheapest moment to catch an accidental edit is immediately after the insertion that could have caused one. + +- [ ] **Step 1: Diff each untouched passage against the parent** + +```bash +for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do + for anchor in 'From pass 4 onward every pass report carries three lines' \ + 'The two rules above do not compete' \ + 'On squash-merge, copy every evidence entry'; do + printf '%s | %s: ' "$f" "$anchor" + diff <(git show "$BASE:$f" | grep -A12 -F "$anchor") <(grep -A12 -F "$anchor" "$f") >/dev/null \ + && echo unchanged || echo CHANGED + done +done +``` + +Expected: `unchanged` six times. Any `CHANGED` is a defect — revert that hunk before continuing. + +- [ ] **Step 2: Confirm `f1` still names the right pair** + +Read the installed text around "The two rules above do not compete" and confirm the two rules immediately above it are still the absorb rule and the clearly-stuck reading, with §A before both rather than between them. + +- [ ] **Step 3: No commit** — this task verifies, it does not change. + +--- + +## Task 3: Replace passage (b) with §B in both copies + +**Files:** +- Modify: `CLAUDE.md` — the passage beginning `**What a loop absorbs, and what stops it` +- Modify: `plugins/dev-workflow/commands/workflow-init.md` — the same + +**The bytes:** target §B, which states it is "the only place this file states anything about" passage (b). **Preserve the two deliberate divergences** the inventory records and §B confirms: W says `the severity rule` where C says `Mechanics` (b3); and the field-mint parenthetical closes the paragraph in C and is absent from W. §B says it stops short of that parenthetical, which is left exactly as each copy has it. + +- [ ] **Step 1: Record the old-wording fragments** + +Two fragments that must reach zero, chosen because each is single-line in the file and is **not** preserved inside the replacement: + +```bash +grep -cF 'plus repair obligations you already accepted in earlier passes' CLAUDE.md +grep -cF 'it resumes the moment the user says whether the set now includes it' CLAUDE.md +``` + +Expected before the edit: `1` each. Re-check both against W with the same command. + +- [ ] **Step 2: Install §B's text over the passage in both copies** + +Replace from `**What a loop absorbs, and what stops it` through the sentence §B ends at, keeping each copy's own closing parenthetical. + +- [ ] **Step 3: Run the discriminating pair, both copies, both trees** + +```bash +NEW='union of the scope every approved story or plan governing this change assigns to this cycle' +OLD='plus repair obligations you already accepted in earlier passes' +for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do + printf '%s new/worktree=%s new/parent=%s old/worktree=%s old/parent=%s\n' "$f" \ + "$(grep -cF "$NEW" "$f")" "$(git show "$BASE:$f" | grep -cF "$NEW")" \ + "$(grep -cF "$OLD" "$f")" "$(git show "$BASE:$f" | grep -cF "$OLD")" +done +``` + +Expected per file: `new/worktree=1 new/parent=0 old/worktree=0 old/parent=1`. **All four values matter** — a copy carrying both the new and the old wording satisfies a one-sided presence check and is exactly the two-instructions-that-disagree failure this pair exists to catch. + +- [ ] **Step 4: Walk the carried conditions** + +Read the installed passage and confirm `b1`, `b2`, `b4`, `b5`, `b6`, `b9`, `b10`, `b14`, `b15` are each present, and that `b3`, `b7`, `b8`, `b11`, `b12`, `b13`, `b16`, `b17`, `b18` read as §B states rather than as the parent did. Nine plus nine; the disposition table above is the checklist. + +- [ ] **Step 5: Parity** + +```bash +diff <(sed -n '/^\*\*What a loop absorbs/,/^\*\*Recognizing "clearly stuck"/p' CLAUDE.md) \ + <(sed -n '/^\*\*What a loop absorbs/,/^\*\*Recognizing "clearly stuck"/p' plugins/dev-workflow/commands/workflow-init.md) +``` + +Expected: only the three recorded divergences. + +- [ ] **Step 6: Commit** + +```bash +git add CLAUDE.md plugins/dev-workflow/commands/workflow-init.md +git commit -m "WIP: replace passage (b) with the absorb paragraph that owns the fix set" +``` + +--- + +## Task 4: Replace passage (c)'s third condition and its two following sentences with §C + +**Files:** +- Modify: `CLAUDE.md` — the passage beginning `**Recognizing "clearly stuck"` +- Modify: `plugins/dev-workflow/commands/workflow-init.md` — the same + +**The bytes:** target §C's fenced block, which gives the three-condition sentence entire so nothing begins mid-clause. + +**The split, stated because getting it wrong is the failure §C names:** only the operative precedence clause — from "a clean completion takes precedence over this exit" to the end of that sentence — moves into §A, capitalized there. Its opening clause, the plateau rationale, **stays here**. The below-floor sentence (`c14`) **does not move; it is replaced**, and no copy of its live wording may survive beside the ordering's split. + +- [ ] **Step 1: Record the old-wording fragments** + +```bash +grep -cF 'a Blocker/Major-free pass below the floor' CLAUDE.md +grep -cF 'So this exit needs three things' CLAUDE.md +``` + +Expected: `1` each, in both copies. + +- [ ] **Step 2: Install §C's block** + +- [ ] **Step 3: Confirm the moved clause exists in §A and nowhere else** + +```bash +grep -cF 'clean completion' CLAUDE.md +grep -n 'takes precedence over this exit' CLAUDE.md +``` + +The precedence clause must appear **once**, inside the ordering. A second occurrence in passage (c) means the sentence was moved whole instead of split. + +- [ ] **Step 4: Run the discriminating pair, both copies, both trees** + +Same four-value shape as Task 3, with `OLD='a Blocker/Major-free pass below the floor'` and `NEW` a single-line fragment of §C's re-raised-dismissal clause taken from the installed file. + +Expected: `new/worktree=1 new/parent=0 old/worktree=0 old/parent=1`. + +- [ ] **Step 5: Walk the conditions** — `c1`–`c3` present unchanged, `c5`–`c7` word for word, `c8` carrying the new clause, `c9` split, `c10`–`c14` gone from here. + +- [ ] **Step 6: Commit** + +```bash +git add CLAUDE.md plugins/dev-workflow/commands/workflow-init.md +git commit -m "WIP: widen the clearly-stuck third condition and split its precedence sentence" +``` + +--- + +## Task 5: Replace `e7` and add the pointer (§D) in both copies + +**Files:** +- Modify: `CLAUDE.md` — the paragraph beginning `Those three lines expose` +- Modify: `plugins/dev-workflow/commands/workflow-init.md` — the same + +**The bytes:** target §D's two fenced blocks — the `e7` sentence entire, and the pointer paragraph added at the end of the passage. + +**Parity note:** W drops the pronoun (`report the tells`, not `you report the tells`). §D's block carries C's wording; **W keeps its own pronoun** unless the parity divergence list decides otherwise in Task 14. Record which you chose — this is one of the two divergences the inventory flags for passage (e), and the other (e11, the C-only rationale paragraph) is untouched. + +- [ ] **Step 1: Record the old wording** + +```bash +grep -cF 'Any two present makes stop-and-surface mandatory, not discretionary' CLAUDE.md +``` + +Expected `1` — but note this fragment **survives inside the replacement**, so it cannot be the old-wording half of the pair. Use instead: + +```bash +grep -cF 'and the "clearly stuck" reading above is not a precondition for it' CLAUDE.md +``` + +and pair it against the installed sentence's new clause. **This is the trap design §7 records:** four spec revisions named a fragment preserved inside its own replacement, whose old-wording-gone count could never reach zero. + +- [ ] **Step 2: Install §D's `e7` sentence and the pointer paragraph** + +- [ ] **Step 3: Run the discriminating pair**, both copies, both trees, with `NEW='read **after** the clean-completion branch of the closure ordering'` and `OLD` from Step 1. + +Expected: `new/worktree=1 new/parent=0 old/worktree=0 old/parent=1`. + +- [ ] **Step 4: Confirm `e1`–`e6` and `e8`–`e11` are untouched**, and that e11 is still C-only. + +- [ ] **Step 5: Commit** + +```bash +git add CLAUDE.md plugins/dev-workflow/commands/workflow-init.md +git commit -m "WIP: read the two-tell threshold after the clean-completion branch" +``` + +--- + +## Task 6: Replace Mechanics · Severity's resolve duty and the handed-over question with §E + +**Files:** +- Modify: `CLAUDE.md` — the `**Severity:**` bullet and the `How this demotion bears` paragraph +- Modify: `plugins/dev-workflow/commands/workflow-init.md` — the same + +**The bytes:** target §E's two fenced blocks. + +**This task removes the one deliberate story-path divergence.** `g4` — C's sentence naming `docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md` as owning the question — goes, because this change is that work and the question is answered. `g2` and `g3`, the interim report-and-stop duty and its justification, go from **both** copies in the same edit. **Removing g4 from C without removing g2/g3 from W desynchronises the copies in the opposite direction**, which is the failure the inventory flags by name. + +- [ ] **Step 1: Record the old wording in both copies** + +```bash +for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do + printf '%s g1/g2: %s g4: %s\n' "$f" \ + "$(grep -cF 'is not settled here, and this change does not settle it' "$f")" \ + "$(grep -cF '2026-08-29-loop-rule-consolidation-story.md' "$f")" +done +``` + +Expected: C `g1/g2: 1 g4: 1`; W `g1/g2: 1 g4: 0`. + +- [ ] **Step 2: Install §E's resolve-duty bullet and its answer paragraph in both copies** + +- [ ] **Step 3: Run the discriminating pair**, both copies, both trees, with `NEW='The demotion changes what a cycle must resolve, never what it observes'` and `OLD='is not settled here, and this change does not settle it'`. + +Expected: `new/worktree=1 new/parent=0 old/worktree=0 old/parent=1` in both. + +- [ ] **Step 4: Confirm g4 is gone from C** + +```bash +grep -c '2026-08-29-loop-rule-consolidation-story.md' CLAUDE.md +``` + +Expected: `0`. The story path may still appear in `docs/` — this check is scoped to `CLAUDE.md`. + +- [ ] **Step 5: Parity** — this passage should now be byte-identical, the one recorded divergence having been removed. Diff the two Severity bullets and expect no output. + +- [ ] **Step 6: Commit** + +```bash +git add CLAUDE.md plugins/dev-workflow/commands/workflow-init.md +git commit -m "WIP: answer the demotion question and scope the resolve duty to the fix set" +``` + +--- + +## Task 7: Install §G and §H's replacements in both copies + +**Files:** +- Modify: `CLAUDE.md`, `plugins/dev-workflow/commands/workflow-init.md` + +**The bytes:** target **§G**'s fenced block — the one-contract paragraph, whose membership is widened and which gains a semantic test a downstream reader can apply — and target **§H**'s blocks, in the order §H gives them: the `c18`-and-surfacing block; `a13`; `a16`; `a17`–`a22`; the Gate-A clean-signal sentence; the gate-prompt template's clean sentence; the Gate-A cadence; the lens paragraph's unchanged-list; and the unknown-start strict-reading list. **Passage (b) is deliberately not among them** — §B owns it. + +**§G is not in the a–j inventory** and therefore carries no condition ids; design §4 lists it as its own site. **§G is an instruction to the agent, not a checker** — install it as written and do not add a mechanical guard beside it. + +**`i12` is the licence for the strict-reading addition and stays in place.** The list is extended at the end, not rewritten; §H gives the whole dash-delimited list so one contiguous string installs. + +- [ ] **Step 1: Locate all eight sites in both copies** + +```bash +for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do + echo "== $f" + grep -n 'These rules and records are one contract' "$f" + grep -n 'Surfacing does not close the cycle' "$f" + grep -n 'Every other rule stated here about how a cycle closes' "$f" + grep -n 'Open a TodoWrite' "$f" + grep -n 'Your final pass must be clean' "$f" + grep -n 'literal `NO FINDINGS` when a pass is clean' "$f" + grep -n 'A clean pass is the single body line' "$f" + grep -n 'Each pass: validate, revise, re-run' "$f" + grep -n 'Lenses are \*\*different questions, not more passes\*\*' "$f" + grep -n 'at minimum floor 3, severity classified without the demotion' "$f" +done +``` + +Expected: one hit per pattern per file. A pattern with zero hits means the wording drifted since the inventory — find it before installing. + +- [ ] **Step 2: Install §G and all eight §H blocks, one site at a time, verifying each before moving to the next** + +- [ ] **Step 3: Run one discriminating pair per block**, both copies, both trees. Old-wording fragments, each single-line and none preserved inside its own replacement: + +| Block | OLD fragment | +|---|---| +| §G one-contract | `These rules and records are one contract` | +| `c18`/surfacing | `no pass is credited as clean` | +| `a13` | `Every other rule stated here about how a cycle closes` | +| `a16` | `fix Blocker/Major after each` | +| `a17`–`a22` | `Your final pass must be clean` | +| Gate-A clean signal | `a literal \`NO FINDINGS\` when a pass is clean` | +| template clean sentence | `A clean pass is the single body line` | +| Gate-A cadence | `Each pass: validate, revise, re-run` | +| lens unchanged-list | `The Blocker/Major filter, the file-first findings protocol` | +| strict-reading list | `the nonce duties at their strictest, the cycle is treated as post-rule` | + +Expected for each: `old/worktree=0 old/parent=1`, and the matching new fragment `1` / `0`. + +**The strict-reading list is add-only at its tail but replaces the dash-delimited run**, so it owes a full pair rather than presence alone. **The gate-prompt template's clean sentence is add-only** if the site carries no wording the change removes — classify it against the real file, per design §7, and check by presence alone if so. + +- [ ] **Step 4: Confirm `a21`, `a22`, `a15` survived** — they are carried inside blocks that install contiguously, so a mis-scoped replacement silently drops them. + +```bash +grep -cF 'Codex is advisory — validate before applying; dismissed finding → one-line why' CLAUDE.md +grep -cF 'Open a TodoWrite' CLAUDE.md +``` + +Expected: `1` each. + +- [ ] **Step 5: Parity** for all nine sites. + +- [ ] **Step 6: Commit** + +```bash +git add CLAUDE.md plugins/dev-workflow/commands/workflow-init.md +git commit -m "WIP: install the one-contract paragraph and the remaining prompt-copy replacements" +``` + +--- + +## Task 8: Replace §F's items 1–9 in both copies + +**Files:** +- Modify: `CLAUDE.md`, `plugins/dev-workflow/commands/workflow-init.md` + +**The bytes:** target §F items 1, 2, 3, 4, 5, 6, 7, 8, 7a, 8a, 8b, 9a, 9b and 9 — fourteen items, each with its own fenced replacement and its `C nnn` / `W nnn` citation. **Re-read every citation against the current file**: §F's own collected list records that items 4, 5 and 8 have line citations one off, and the numbers drifted further as this cycle edited the copies. + +**`h4` and `h19` are the human-exception conditions these items discharge** — item 7 is the destination, item 4 the scope sentence. + +- [ ] **Step 1: Re-derive every item's real location** + +For each item, take the quoted live sentence from §F and find it, rather than trusting the cited line: + +```bash +grep -n -F '' CLAUDE.md plugins/dev-workflow/commands/workflow-init.md +``` + +**Where a quoted sentence wraps across lines in the file, `grep -F` on the whole sentence returns nothing.** Search a single-line fragment of it instead and confirm by reading. Record which items wrap — Task 14's parity diff needs it. + +- [ ] **Step 2: Install all fourteen replacements** + +- [ ] **Step 3: Run one discriminating pair per item**, both copies, both trees, choosing each OLD fragment single-line and not preserved inside its replacement. + +Expected per item: `new/worktree=1 new/parent=0 old/worktree=0 old/parent=1`. + +- [ ] **Step 4: Count what was installed** + +```bash +# fourteen prompt-copy items from this task, each present once per copy +``` + +State the number you observed. **Do not carry a count from §F into a check** — §F states the count of falsified sentences and this task installs a subset of them; a count copied between the two is the stale-bookkeeping defect the design records at five passes running. + +- [ ] **Step 5: Parity** for all fourteen sites. + +- [ ] **Step 6: Commit** + +```bash +git add CLAUDE.md plugins/dev-workflow/commands/workflow-init.md +git commit -m "WIP: replace the fourteen falsified sentences in the two prompt copies" +``` + +--- + +## Task 9: Replace §F's items 14 and 18 + +**Files:** +- Modify: `CLAUDE.md` — the Named residual paragraph (§5), and the work-loop line (§4) +- Modify: `plugins/dev-workflow/commands/workflow-init.md` — the same two + +**The bytes:** target §F item 14 (the Named residual's blanket exemption) and item 18 (the work-loop sequence). + +**These two are separated from Task 8 because they were found last and because item 18 is the only edit outside §5.** A reviewer can reject this task while approving Task 8. + +- [ ] **Step 1: Locate both sites** + +```bash +grep -n 'Hook text is out of scope here' CLAUDE.md plugins/dev-workflow/commands/workflow-init.md +grep -n 'The work loop includes the review gates' CLAUDE.md plugins/dev-workflow/commands/workflow-init.md +``` + +Expected: one hit each per file. **Item 14's sentence wraps after "here"** — §F says so; grep a single-line fragment. + +- [ ] **Step 2: Install both replacements** + +- [ ] **Step 3: Discriminating pairs** + +```bash +OLD1='Hook text is out of scope here' +NEW1='not a blanket exemption for hook text' +OLD2='execute → tests green → Gate B → commit' +NEW2='Gate A (spec) → Gate-A closing act' +``` + +Expected for each, in both copies: `new/worktree=1 new/parent=0 old/worktree=0 old/parent=1`. + +- [ ] **Step 4: Confirm the residual the sentence was written for still stands** + +The hook reporting its own threshold as an obligation at a floor of 1 is **not** repaired by this change. Read the replaced paragraph and confirm it still says so. + +- [ ] **Step 5: Parity** for both sites. + +- [ ] **Step 6: Commit** + +```bash +git add CLAUDE.md plugins/dev-workflow/commands/workflow-init.md +git commit -m "WIP: correct the Named residual's blanket exemption and the work-loop sequence" +``` + +--- + +## Task 10: Replace the seven hook reminder strings + +**Files:** +- Modify: `plugins/dev-workflow/hooks/codex-gate.sh` + +**The bytes:** target §F items 10, 11, 12, 13, 15, 16 and 17. **Items 12, 15 and 16 give the complete resulting text for both channels**; items 10, 11, 13 and 17 give replacement sentences inside an otherwise unchanged message. + +**Shell constraint, and it is not advisory:** every string installs into a double-quoted `note` argument. No backtick, no `$(`, no backslash, no double quote. A backtick reached §F once and would have executed `WIP` at install time, shipped the reminder with the word missing, and failed ShellCheck — while the fixture copied from it would have made the suite pass on the corruption. + +- [ ] **Step 1: Verify the constraint before installing** + +```bash +# For each of the seven §F blocks, confirm the text you are about to install is clean: +printf '%s' "" | grep -nE '`|\$\(|\\\\|"' && echo HAZARD || echo clean +``` + +Expected: `clean` seven times. A hit here stops the task. + +- [ ] **Step 2: Locate the seven `note` calls** + +```bash +grep -n 'note "' plugins/dev-workflow/hooks/codex-gate.sh +``` + +Expected: thirteen hits. Seven are the gate reminders this task edits; four take a pre-built variable (`$FAILURE_CTX`, `$NORESULT_CTX`, `$BG_SHORT_CTX`, `$BG_LONG_CTX`), one is the tool-state echo, and one is the **docs-only notice, which is deliberately untouched** — it states no closure permission. + +- [ ] **Step 3: Install the seven replacements, one at a time** + +- [ ] **Step 4: Confirm no behaviour changed** + +```bash +git diff plugins/dev-workflow/hooks/codex-gate.sh | grep -E '^[-+]' | grep -vE '^[-+]\s*(note "|[A-Za-z ,.—;:()/$-]+")' | head +``` + +Expected: only the `---`/`+++` header lines. **Any changed line that is not inside a `note` string is out of scope** — no control flow, no counter, no fingerprint computation, no routing (invariant 4, design §8). + +- [ ] **Step 5: Discriminating pairs, worktree and parent** + +The hook has one copy and owes no parity check; **the hook suite's exact-match assertion is this edit's second observation** (design §7). Run the pair anyway: + +```bash +for pair in 'this floor is the only thing keeping the spec review honest|instruction-backed' \ + 'commit only if your final pass was clean — no new Blocker/Major|every other closure condition holds' \ + 'then make the real commit when your final pass is clean|Use this commit as the review range' \ + 'STOP — Codex Gate B not satisfied|Codex gate state:' ; do + OLD=${pair%%|*}; NEW=${pair##*|} + printf 'old/worktree=%s old/parent=%s new/worktree=%s new/parent=%s %s\n' \ + "$(grep -cF "$OLD" plugins/dev-workflow/hooks/codex-gate.sh)" \ + "$(git show "$BASE:plugins/dev-workflow/hooks/codex-gate.sh" | grep -cF "$OLD")" \ + "$(grep -cF "$NEW" plugins/dev-workflow/hooks/codex-gate.sh)" \ + "$(git show "$BASE:plugins/dev-workflow/hooks/codex-gate.sh" | grep -cF "$NEW")" "$OLD" +done +``` + +Expected: `old/worktree=0`, `old/parent` ≥ 1, `new/worktree` ≥ 1, `new/parent=0` for each. The `STOP` pair covers two messages, so its counts are 2 rather than 1 — **state the number you observed rather than asserting it**. + +- [ ] **Step 6: ShellCheck** + +```bash +shellcheck --shell=sh plugins/dev-workflow/hooks/codex-gate.sh +``` + +Expected: exit 0, no output. The suite will fail at this point because the expectations still pin the old strings — that is Task 11. + +- [ ] **Step 7: Commit** + +```bash +git add plugins/dev-workflow/hooks/codex-gate.sh +git commit -m "WIP: replace the seven gate reminders the ordering falsifies" +``` + +--- + +## Task 11: Sweep `codex-gate.test.sh` for every assertion naming a replaced string + +**Files:** +- Modify: `plugins/dev-workflow/hooks/codex-gate.test.sh` + +**There is no list of these assertions and building one here would repeat a defect.** §F states the duty and deliberately states no count: it twice named one and was twice wrong — "the exact-match expectation" where there are three, and "the remaining are matched by loose patterns these repairs leave standing" where `Gate B satisfied` occurs 21 times and `STOP` 14. **Sweep the file; do not work from a number.** + +- [ ] **Step 1: Find every site** + +```bash +grep -n 'expected_ctx=\|expected_msg=' plugins/dev-workflow/hooks/codex-gate.test.sh +grep -n 'Gate B satisfied\|Gate B not satisfied\|STOP\|no recorded review\|cannot confirm review\|only thing keeping\|no new Blocker/Major' plugins/dev-workflow/hooks/codex-gate.test.sh +``` + +**Record the counts you observe.** They will not match the numbers above if the file has changed; the numbers above are evidence for why no list is kept, not a target. + +- [ ] **Step 2: Update the three `expected_ctx` and three `expected_msg` assignments** + +Each takes the complete resulting text of its message from §F items 12, 15 and 16, with `$policy` rendered as the suite renders it (`this project's review policy`) and the counters as the fixture sets them. + +- [ ] **Step 3: Update every loose assertion, test label and comment that names a replaced string** + +Each must test **the observed hook state** rather than a gate verdict — `hook checks passed`, `no recorded fingerprint`, `cannot confirm reviewed content` — matching what Task 10 installed. **A label left saying "satisfied" is a test vocabulary that still calls the gate satisfied**, which is the claim this change removes. + +- [ ] **Step 4: Run the suite under both shells** + +```bash +HOOK_SH=sh sh plugins/dev-workflow/hooks/codex-gate.test.sh && \ +HOOK_SH=dash dash plugins/dev-workflow/hooks/codex-gate.test.sh +``` + +Expected: exit 0 from both. **`HOOK_SH` selects the shell the hook runs under; without it a `dash` invocation only exercises the harness.** Ubuntu's `/bin/sh` is dash, and dropping this second run is what let a dash-only defect ship once already. + +- [ ] **Step 5: Confirm no verdict vocabulary survives** + +```bash +grep -n 'Gate B satisfied\|Gate B not satisfied' plugins/dev-workflow/hooks/codex-gate.test.sh +``` + +Expected: no hits, or only hits you can justify one by one in the commit body. + +- [ ] **Step 6: ShellCheck the test file** + +```bash +shellcheck --shell=sh --exclude=SC2015 plugins/dev-workflow/hooks/codex-gate.test.sh +``` + +Expected: exit 0. **The `--exclude=SC2015` is a single-code exclusion**, not a blanket disable; every other rule still applies. + +- [ ] **Step 7: Commit** + +```bash +git add plugins/dev-workflow/hooks/codex-gate.test.sh +git commit -m "WIP: move every hook assertion that names a replaced reminder string" +``` + +--- + +## Task 12: The `b11`/`b13` equivalence check + +**Files:** none modified (unless the check fails). + +**Interfaces:** +- Consumes: the installed §A block and the installed passage (b). + +Design §6 requires a second check **within each copy**: that `b11` and `b13` as edited say what the ordering cites them as saying, **comparing the complete predicates and not a shared phrase** — including `b11`'s already-declined exception and `b13`'s already-answered qualification, in both directions. + +- [ ] **Step 1: Extract what the ordering cites** + +Read §A's continue branch and its scope-trigger references, and write down, in full, the predicate it attributes to the absorb paragraph. + +- [ ] **Step 2: Extract what the absorb paragraph states** + +Read the installed passage (b) and write down, in full, the predicate it defines for each of `b11` and `b13`. + +- [ ] **Step 3: Compare in both directions** + +A condition **in the block and not in the source** ships two triggers that disagree. A condition **in the source and not in the block** means the block cites a rule it has not read. Both are failures. + +- [ ] **Step 4: Repeat for the second copy** + +The two copies are byte-identical over this material, so a divergence here is a parity failure and belongs to Task 14. + +- [ ] **Step 5: Record the result** — it goes in the evidence entry verbatim. **No commit** unless the check failed and you repaired something. + +--- + +## Task 13: The next-state table and the per-condition closure checks + +**Files:** +- Modify: this plan (the table lives here; design §7) + +**What the table claims, at exactly this width:** it covers **answer-state transitions once the predicates producing them are established**. It does **not** establish how each predicate was derived — a wrongly derived predicate produces a row that passes — nor whether the rows cover every reachable combination of the clean, scope and health predicates. **The evidence entry states the claim at this width and no wider**; closing either gap is the parked fixture-per-predicate question, which this change does not reopen. + +**The oracle.** A row **fails** when its required answer does not produce a **distinct** resumable or closed state — the same stop returning with its reading unconsumed, that is, without an intervening validated pass run after the answer — or when it closes on anything other than the route the block states. **Read the closure conditions and the routes off the installed §A, not from this plan** — an embedded copy can pass while disagreeing with the text it checks. + +- [ ] **Step 1: Enumerate the rows** + +One row per (starting state, answer) pair the installed ordering admits. At minimum, and **this is a floor rather than the set**: a clean eligible pass with every condition met; a clean eligible pass with an unmet precondition; a clean pass below the floor; a zero-finding pass below the floor; a pass carrying a membership trigger, answered accept and answered decline; a pass carrying a new-question trigger, answered; a pass carrying both; a two-tell stop, answered; a clearly-stuck surface, answered; a source block raised before any pass was read; a source block raised on a pass already read; a closing act that does not complete and is repaired; a closing act that cannot be repaired; a `full` Gate-B pass with one branch clean and one not; the same complaint in both branch files under each of accept/accept, accept/decline, decline/accept and decline/decline. + +- [ ] **Step 2: Walk each row against the installed §A and record the next state** + +- [ ] **Step 3: Apply the oracle to each row** and mark pass or fail. + +- [ ] **Step 4: Write one named check per closure condition the block states** + +**Read the set off the block and write one check per condition.** **Fail this task where the block states a condition you have no check for** — an enumeration here is how design §7 came to name three conditions while the block stated more. + +- [ ] **Step 5: Write the separate named checks the table does not cover** + +Two, per design §7: that a logical pass was validated across every required branch file, and that every closure condition the block states held at the closing act. + +- [ ] **Step 6: Commit** + +```bash +git add docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +git commit -m "WIP: next-state table and per-condition closure checks" +``` + +--- + +## Task 14: The parity diff and the divergence list + +**Files:** possibly `CLAUDE.md` and `plugins/dev-workflow/commands/workflow-init.md`. + +Design §6: the two copies must agree on every rule this change ships. **The plan carries the divergence list** — which pre-existing wording differences are deliberate and stay, which are not and are aligned — **and performs the extraction and diff, passage by passage, against the real files.** + +**One divergence is decided by the design and is not a judgement call:** W's `b3` pointer names "the severity rule" on the inventory's reasoning that W has no Mechanics section, **which is false** — so **W takes C's wording** (design §6, target §B). + +- [ ] **Step 1: Extract and diff each changed passage** + +```bash +for anchor in 'How a cycle ends' 'What a loop absorbs' 'Recognizing "clearly stuck"' \ + 'Those three lines expose' '\*\*Severity:\*\*' 'Surfacing does not close' \ + 'When these rules bind' 'Named residual' 'The work loop includes'; do + echo "== $anchor" + diff <(grep -A25 -E "$anchor" CLAUDE.md) \ + <(grep -A25 -E "$anchor" plugins/dev-workflow/commands/workflow-init.md) +done +``` + +- [ ] **Step 2: Classify every difference the diff reports** + +Three buckets: **deliberate and stays** (the field-mint parenthetical, e11, f5–f7's evidence framing, e8's pronoun — each recorded in the inventory); **not deliberate, align it**; **introduced by this change, fix it**. Write the list into this plan. + +- [ ] **Step 3: Apply W's `b3` alignment** + +- [ ] **Step 4: Re-run the diff** and confirm only the deliberate divergences remain. + +- [ ] **Step 5: Commit** + +```bash +git add CLAUDE.md plugins/dev-workflow/commands/workflow-init.md docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +git commit -m "WIP: align the two copies and record the divergence list" +``` + +--- + +## Task 15: Version bump, CHANGELOG, battery, evidence, Gate B + +**Files:** +- Modify: `plugins/dev-workflow/.claude-plugin/plugin.json` +- Modify: `plugins/dev-workflow/CHANGELOG.md` + +**Mode:** read from the story header at execution. It was `battery+check+verification` at the time this plan was written; **read it fresh** — the header is the only writable copy and this plan carries the path, not the value. + +- [ ] **Step 1: Bump the version** + +```bash +grep -n '"version"' plugins/dev-workflow/.claude-plugin/plugin.json +``` + +`0.11.0 → 0.12.0` — a **minor** bump: the template gains a closure ordering, the edited sentences, and the seven hook reminder strings. The hook edits add no bump the template did not already require. + +- [ ] **Step 2: Add the CHANGELOG entry**, newest first, naming the ordering, the falsified-sentence replacements and the hook reminder strings. + +- [ ] **Step 3: Commit the bump into the WIP snapshot** + +```bash +git add plugins/dev-workflow/.claude-plugin/plugin.json plugins/dev-workflow/CHANGELOG.md +git commit -m "WIP: bump dev-workflow to 0.12.0" +``` + +`scripts/check-version-bump.sh` compares **commits**, so the bump must be committed before the battery runs. + +- [ ] **Step 4: Run the full quality battery** + +```bash +shellcheck --shell=sh plugins/dev-workflow/hooks/codex-gate.sh && \ +shellcheck --shell=sh --exclude=SC2015 plugins/dev-workflow/hooks/codex-gate.test.sh && \ +shellcheck --shell=sh scripts/check-invariants.sh && \ +shellcheck --shell=sh --exclude=SC2015 scripts/check-invariants.test.sh && \ +shellcheck --shell=sh scripts/check-version-bump.sh && \ +shellcheck --shell=sh scripts/check-version-bump.test.sh && \ +HOOK_SH=sh sh plugins/dev-workflow/hooks/codex-gate.test.sh && \ +HOOK_SH=dash dash plugins/dev-workflow/hooks/codex-gate.test.sh && \ +sh scripts/check-invariants.test.sh && sh scripts/check-invariants.sh && \ +sh scripts/check-version-bump.test.sh && sh scripts/check-version-bump.sh main && \ +claude plugin validate . --strict +``` + +Expected: exit 0. **`check-invariants.sh` includes the prompt-conformance checks** — a `Target model:` line naming one recognized model, a prose checklist-count claim matching the checklist, and the finding-severity vocabulary as a closed set in both prompt copies. Those three are a floor, not coverage; invariant 11's other eleven items are judged by a reader. + +- [ ] **Step 5: Write the evidence entry into the WIP commit body** + +It names: the battery run; **every pair this plan built, with its counts in each copy and each tree, and every presence check beside them**; the §6 parity diff and the `b11`/`b13` equivalence result; and the next-state table's location plus its row count. **State the §A presence checks as presence, not as pairs** — its counterfactual is absent and is claimed as absent. + +- [ ] **Step 6: Run Gate B** + +```bash +git rev-parse HEAD # the WIP commit +git rev-parse HEAD^ # baseSha +``` + +`mcp__codex__review` with `reviewType: full`, `baseSha` = the WIP commit's parent, `headSha` = the full 40-character object name `HEAD` resolves to **at that moment**, resolved once and kept with each branch's result. Carry the story path and the evidence entry quoted verbatim. Write findings to `.context/codex-reviews/gate-b---pass-

.md` — **draw a fresh nonce for this cycle**; it is a different cycle from `awsf1ec771`. + +**Standing lens, every call:** "which existing statements does this diff falsify?" and **name what this diff changes the size, value or position of** — a list, a count, a version, an identifier, a cited line — then grep for where each is described elsewhere. + +- [ ] **Step 7: Loop to a clean pass at or above the derived floor** + +Floor derives from the story profile: risk `high` → 2, security `none` → 0, max 2 ≠ 0 → **floor 3**. Re-derive it at each pass from the header. Fix Blocker/Major after each pass; re-review after every fix. **Revalidate the evidence entry before every re-review and before the closing amend.** + +**A fix that changes specified behaviour updates the spec in the same commit.** + +- [ ] **Step 8: Close the cycle** + +```bash +git reset --soft +git commit -m "" +``` + +The closing body carries: the validated evidence entry; the provenance line; the per-pass curve; and any human-exception record. **Amend rather than a follow-up commit** — a `WIP:` commit left in history defeats the convention, and a follow-up has nothing to commit when the review produced no fixes. + +--- + +## Self-Review + +**1. Spec coverage.** §A → Task 1. §B → Task 3. §C → Task 4. §D → Task 5. §E → Task 6. §F items 1–9 → Task 8; items 14, 18 → Task 9; items 10–13, 15–17 → Task 10 with its test sweep in Task 11. §G → Task 7, **added by this review**: the first draft gave the one-contract paragraph no task, though design §4 lists it as its own site and target §G carries its replacement. It is a prompt-copy replacement in both copies with the same shape as §H's blocks, owes the same discriminating pair with OLD `These rules and records are one contract`, and sits at C 879 / W 1063 as of this writing. §H → Task 7. §I ships nowhere and needs no task. Design §6 → Task 14. Design §7 → Tasks 13 and 15. Design §8 → Task 15's battery and the Global Constraints. Story AC 5 → the disposition tables. Story AC 4 → Task 13. + +**2. Placeholder scan.** The replacement text is cited rather than copied, deliberately and for the reason the Architecture note gives. Every verification step carries a runnable command and an expected value. Task 13's row list is explicitly a floor rather than a closed set, and says so. Task 11 deliberately carries no count, and says why. + +**3. Type consistency.** `$BASE` is set in Task 0 and used in Tasks 1–10. The four-value pair shape (`new/worktree`, `new/parent`, `old/worktree`, `old/parent`) is defined in Task 3 and referred to by name afterwards. Condition ids match the inventory throughout: a1–a22, b1–b18, c1–c20, d1–d7, e1–e11, f1–f7, g1–g4, h1–h26, i1–i16, j1–j4 — 135 total, every one dispositioned above. + +**One correction applied from this review:** §G was missing a task; it is now installed by Task 7, which names nine sites rather than eight. + +**One residual this plan does not close, stated rather than left to be found.** Nothing here establishes that the edit set is complete — it is the sites §F knows, and §F's own §I records that it cannot establish completeness either. Task 8's step 1 re-derives every citation against the real file, and Task 14's diff catches a copy that fell out of step; neither is a completeness check. The sweep for an affected site this text has not found is design §7's, owed by whoever executes, and a site found during execution is a finding against the spec rather than a gap in this plan. From 5871d0a26a76171408ff7a2e1b5d34dac1cd318a Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 11:13:55 +0200 Subject: [PATCH 098/181] docs(context): open Gate-A plan cycle om0bdd7udh Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-plan-om0bdd7udh-resume.md | 33 +++++++++++++++++++ 1 file changed, 33 insertions(+) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md new file mode 100644 index 0000000..a431c78 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md @@ -0,0 +1,33 @@ +# Gate-A (plan) working record — cycle `om0bdd7udh` + +Advisory, cycle-stable, per CLAUDE.md §5 optional companions. Retire at closure. +Nothing depends on it; the pass files and the repo are authoritative where this disagrees. + +- **Kind:** Gate-A plan +- **Nonce:** om0bdd7udh (drawn 2026-09-14 from /dev/urandom, 10 chars, no collision among open cycles — the only other cycle in this work, `awsf1ec771`, closed at `ba15e83`) +- **Artifact:** `docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md` +- **Branch:** loop-rule-consolidation +- **Story:** `docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md` — profile read from its header at each pass (high / none / battery+check+verification at pass 1) +- **Derived floor:** 3 (risk high → level 2; security none → 0; max 2 ≠ 0 → 3) +- **Hook knob:** absent (no floor-knob file in `.context/`) + +## What the spec cycle learned, carried here so this loop does not relearn it + +The Gate-A **spec** cycle ran 65 passes. Four mechanisms produced almost every finding; each is +worth checking before a pass rather than after. + +1. **Second copies.** A rule stated in both the design and the artifact. This plan cites the target + text rather than copying it, which is the structural answer — check that no task quietly + restates a rule instead of pointing at it. +2. **Counts over files the artifact does not survey.** §F claimed twice how many test assertions a + change reaches and was wrong twice. This plan carries **no count** of them and states a sweep + duty instead. A finding asking for a count back is asking for the removed defect. +3. **Enumerations that go stale.** Prefer removing one over correcting it. +4. **Spans plus prose about where they go.** Three passes found a new gap in that framing. The + hook items carry complete messages for exactly this reason. + +## Passes + +| Pass | Plan rev | Findings | Blockers | Majors | Valid | Notes | +|---|---|---|---|---|---|---| +| 1 | 5e466f6 | — | — | — | not run | first pass | From 9f13a2cda8c35b0233ce48f2658e7dd228b84477 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 11:31:12 +0200 Subject: [PATCH 099/181] docs(plans): apply Gate-A plan pass 1; one verified fragment table MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Blocker: with a WIP commit per task, baseSha = HEAD^ puts only the version bump in Gate B's range and the gate would close having reviewed no prompt or hook change. baseSha is now $BASE, the parent of the first WIP, persisted to a file in Task 0 because a shell variable does not survive between task invocations. The dominant class was verification fragments that count zero against the real files: nine of the plan's grep patterns wrap across a line break, quote text that does not exist, or name a fragment preserved inside its own replacement. That is the defect design §7 records from four consecutive spec revisions, reproduced here on the first try. Repaired structurally — the plan now carries one fragment table, every row checked against both copies and reporting its line, and every task cites a row instead of inventing a pattern. Two tasks installed text contrary to the approved target — W's b3 divergence, which design §6 decides against, and W's missing e8 pronoun. Both now install what the target says and Task 14 verifies rather than repairs. Also: per-paragraph presence checks for §A2 and §A3, which a single count on §A1 let be omitted; a hazard probe that reads blocks as inert data rather than executing them inside double quotes; a behaviour-change check by line number rather than by line shape, which admitted else and fi; pairs for hook items 11 and 17; zero tolerance for surviving gate-verdict vocabulary in the test file; and a final parent-tree diff of the five untouched ranges before Gate B. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-plan-om0bdd7udh-pass-1.md | 33 ++ .../gate-a-plan-om0bdd7udh-resume.md | 35 +- .../2026-09-14-loop-rule-consolidation.md | 473 ++++++++++++++---- 3 files changed, 445 insertions(+), 96 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-1.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-1.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-1.md new file mode 100644 index 0000000..8ff3074 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-1.md @@ -0,0 +1,33 @@ +BLOCKER | high | Task 15 step 6 | The plan creates separate WIP commits in Tasks 1 through 15, then sets `baseSha` to the final WIP commit's parent, so the review range contains only the version-bump commit rather than the whole implementation | Gate B can close without reviewing any prompt or hook change, leaving the repository unable to satisfy its required diff-review gate | Set `baseSha` to `$BASE`, the parent of the first WIP, and keep `headSha` at the current full HEAD object name +MAJOR | high | Task 0 interface | `$BASE` is assigned only as a shell-local variable in Task 0, but later task commands run in separate shell invocations and have no persisted value to consume | Every parent-tree count can execute with an empty revision and fail or produce unusable evidence | Persist the exact base in a scratch file and reload it in every task, or rederive the same immutable commit explicitly in each command block +MAJOR | high | Task 1 step 4 | The sole presence count names only A1's `How a cycle ends` heading and never observes the new A2 Gate-A closure paragraph | A2 can be omitted from both copies while every Task 1 count and parity check passes | Add a unique A2 fragment and require worktree 1 and parent 0 in both files +MAJOR | high | Task 1 step 4 | The sole presence count names only A1's `How a cycle ends` heading and never observes the new A3 Gate-B closure paragraph | A3 can be omitted from both copies while every Task 1 count and parity check passes | Add a unique A3 fragment and require worktree 1 and parent 0 in both files +MAJOR | high | Task 3 steps 1 and 3 | OLD `plus repair obligations you already accepted in earlier passes` wraps across C 201-202 and W 408-409 | Its parent count is 0 in both real files, so the required `old/parent=1` result is impossible and the passage replacement cannot pass its stated check | Use a genuinely single-line OLD fragment from b7 in each source layout +MAJOR | high | Task 3 steps 1 and 3 | The second recorded OLD, `it resumes the moment the user says whether the set now includes it`, wraps across C 207-208 and W 414-415 and is then omitted from the four-value pair entirely | The pre-edit expectation is false and the b12 immediate-resumption instruction can survive without a discriminating check failing | Choose a single-line b12 OLD fragment and give b12 its own four-value pair +MAJOR | high | Task 3 steps 2 and 5 | Task 3 says to preserve W's b3 `the severity rule` divergence and expects three divergences, while target §B and design §6 say that rationale is false and W must take C's `Mechanics · Severity` wording | Task 3 deliberately installs text contrary to the approved target and cannot be approved independently; Task 14 is then required to repair a defect this task introduces | Install the target's common b3 wording in Task 3 and make Task 14 verify it rather than change it +MAJOR | high | Task 4 step 4 | The pair leaves NEW as an executor-chosen fragment from the installed file and provides no concrete fragment or runnable four-value command | The reviewed plan has not built the discriminating pair design §7 assigns to it, and an executor can choose a shared or wrapped fragment that proves nothing | Name one exact single-line NEW fragment from target §C and spell out its four counts +MAJOR | high | Task 5 steps 1 and 3 | OLD `and the "clearly stuck" reading above is not a precondition for it` wraps across C 267-268 and W 471-472 | The parent count is 0 in both files, so the e7 replacement cannot produce the claimed pair | Use a single-line OLD fragment from e7 that exists once in each parent file +MAJOR | high | Task 5 steps 2 and 3 | §D's new `What the answer does` pointer is add-only, but the only check observes the separate e7 replacement | The pointer can be omitted from both copies while the pair, condition walk and parity instruction all pass | Add a presence-only worktree 1 and parent 0 check for a unique pointer fragment in both copies +MAJOR | high | Task 5 parity note | The task directs W to keep its old omitted pronoun, while target §D supplies C's complete replacement sentence and the plan's global constraint says every replaced prompt section is byte-identical | The implementation can knowingly diverge from the approved target and design §6 while Task 14 classifies the divergence as deliberate | Install target §D's complete sentence in both copies, including `you report`, and remove e8's pronoun from the surviving-divergence list +MAJOR | high | Task 6 step 5 | The parity check is scoped only to the two Severity bullets even though Task 6 also installs the long handed-over-question answer paragraph | The answer paragraph can differ between C and W while the NEW and OLD count pair and the stated parity check pass | Extract and diff both §E replacement blocks, including the complete answer paragraph +MINOR | high | Task 7 steps 1, 2 and 5 | The task enumerates ten sites, calls them eight in steps 1 and 2, then calls them nine in step 5 | An executor following the stated totals can omit one or two replacements or parity checks while believing the task complete | State nine §H blocks plus one §G block, ten sites total, consistently throughout the task +MAJOR | high | Task 7 steps 1 and 3 | `These rules and records are one contract` occurs nowhere in either source; the real line is `These records are one contract` at C 879 and W 1063 | The locator and §G OLD count both return 0 before editing, so the one-contract replacement has no working counterfactual | Replace the anchor and OLD fragment with the exact live wording +MAJOR | high | Task 7 steps 1 and 3 | `Your final pass must be clean` wraps across C 132-133 and W 339-340 | The locator and a17 OLD count return 0 in both correct parent files | Use a single-line a17 fragment such as `final pass must be clean — if the pass at the floor` +MAJOR | high | Task 7 steps 1 and 3 | `a literal `NO FINDINGS` when a pass is clean` is one line in C but wraps across W 756-757 | The claimed one-hit locator and `old/parent=1` pair fail for W even when its source is correct | Use per-copy single-line OLD fragments or extract and normalize the complete sentence before comparing +MAJOR | high | Task 7 steps 1 and 3 | `A clean pass is the single body line` is not a source line: C 328 ends with `A`, C 329 begins blockquoted lowercase `clean pass`, and W has the same shape at 522-523 | Both the locator and the template-clean OLD count are zero in the real files | Anchor on the exact single-line fragment `clean pass is the single body line` and count that fragment +MAJOR | high | Task 7 step 1 | `at minimum floor 3, severity classified without the demotion` wraps across C 155-156 and W 362-363 | The strict-reading locator returns zero in both files and falsely reports source drift | Use a single-line fragment from the live strict-reading list +MAJOR | high | Task 7 step 3 | The strict-reading OLD uses `the nonce duties at their strictest, the cycle is treated as post-rule`, but the source says `the nonce duties at their strictest — the cycle is treated as post-rule` | Its parent count is zero even after line wrapping is handled, so the pair is broken independently of the locator | Use the live em-dash fragment on C 157 and W 364 +MAJOR | high | Task 7 step 3 | The table provides only OLD fragments and says to use an unspecified matching NEW fragment for every §G and §H block | None of the ten pairs is fully reviewable or runnable, and an executor can choose a NEW fragment preserved in the parent or absent because of wrapping | Add an exact single-line NEW column and the four-value command for every block +MAJOR | high | Task 8 step 3 | All fourteen meaning-changing items defer both fragment selection and command construction to the executor | The plan does not build the per-edit discriminating pairs required by design §7, so this review cannot detect preserved OLD text, wrapped fragments or non-discriminating NEW text | Add a fragment table with one verified single-line OLD and NEW for each item and run the standard four counts from it +MAJOR | high | Task 10 step 1 | The hazard probe places the candidate block inside double quotes, so backticks and command substitutions execute before grep and intended variables such as `$policy` expand away; its backslash alternative also does not reliably test a single backslash | The check can execute the hazard it is meant to detect and can certify altered text rather than the literal string to be installed | Read each fenced block as inert data with Python or a single-quoted heredoc and test its literal characters without shell expansion +MAJOR | high | Task 10 step 4 | The diff filter admits any changed line made only of letters and allowed punctuation, including shell control words, and it filters the diff headers too despite claiming they remain as output | A control-flow edit such as changing a plain `else`, `return` or `fi` line can pass the no-behaviour-change check, while the documented expected output is mechanically wrong | Parse zero-context diff hunks and require every changed source line number to be one of the seven exact `note` calls; expect no output on success +MAJOR | high | Task 10 step 5 | No pair covers item 11, the Gate-A satisfied reminder changed to `Proceed only once this Gate-A cycle has closed` | That reminder can retain its old abbreviated clean definition while all four listed pairs pass | Add a dedicated item-11 OLD and NEW pair against hook worktree and parent +MAJOR | high | Task 10 step 5 | No pair covers item 17, the Gate-B below-floor reminder whose `run more` and skip-rule instruction is replaced | The invalid running-cycle skip exit can survive while all four listed pairs pass | Add a dedicated item-17 OLD and NEW pair against hook worktree and parent +MAJOR | high | Task 11 step 5 | The final vocabulary check allows `Gate B satisfied` and `Gate B not satisfied` hits when justified in the commit body, contradicting step 3's rule that even labels using `satisfied` must move to observed hook state | Old gate-verdict assertions or labels can survive and the task can still be marked complete | Require zero hits for both phrases in this test file; use different observed-state wording wherever a test still needs that case +MAJOR | high | Tasks 0, 2 and 14 | Task 0 says Task 14 will diff the untouched ranges against `$BASE`, but Task 14 performs only C-versus-W parity and never compares passages d, f or j, the untouched floor arithmetic, or h's kept conditions to the parent; Task 2 runs before all later edits | A later task can alter an untouched condition identically in both copies and every final check passes | Add a final parent-tree diff for all five recorded do-not-touch ranges after every text edit and before Gate B +MAJOR | high | Task 0 step 2 | The step promises a recorded list of five untouched passage ranges but locates only the three anchors for d, f and j | Its produced interface omits the untouched floor arithmetic and the kept human-exception conditions that later verification is supposed to consume | Record explicit bounded start and end anchors for all five ranges in both copies +MAJOR | high | Task 14 step 1 | The nine 25-line extraction windows omit changed sites including §G, the Gate-A and Gate-B prompt instructions, the human-exception edits, WIP and finishing-cycle text, evidence revalidation, and the curve rationale | The final §6 parity diff can pass while many rules shipped by this change differ between C and W | Drive the final parity check from every changed site named by target §§B-H, with exact bounded regions rather than nine broad windows +MAJOR | medium | Global constraints and Architecture | The plan calls each target fenced block the bytes to install and says not to reflow, but also orders every replacement to be installed unwrapped as one physical line even though the approved fenced blocks are extensively wrapped | Following one instruction violates the other, and following the unwrapped rule produces files whose bytes do not match the approved target text | State whether target newlines are normative; if they are, preserve them and choose fragments wholly within source lines, otherwise replace every byte-exact claim with an explicit whitespace-normalization rule +MINOR | high | Condition disposition c20 | c20 is labeled carried even though the plan immediately narrows `every Blocker and Major` to `every in-set Blocker and Major`, changing the condition's scope | The accounting overstates preservation and obscures a meaning-changing replacement in the exact class AGENTS.md requires the plan to expose | Mark c20 replaced or changed and point to the §H block that installs its narrowed form +MINOR | medium | Tasks 13 and 14 plan mutations | Neither the next-state table nor the divergence list has a named insertion or replacement region in this plan | Re-executing either task can duplicate or strand these verification artifacts, making the plan non-idempotent and its later location claims ambiguous | Add stable headings or placeholders and instruct each task to replace that exact region idempotently +END OF FINDINGS (32 total) diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md index a431c78..201a058 100644 --- a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md @@ -30,4 +30,37 @@ worth checking before a pass rather than after. | Pass | Plan rev | Findings | Blockers | Majors | Valid | Notes | |---|---|---|---|---|---|---| -| 1 | 5e466f6 | — | — | — | not run | first pass | +| 1 | 5871d0a | **32** | **1** | **28** | yes | first pass. The Blocker is real: with a WIP commit per task, `baseSha = HEAD^` would have put only the version bump in Gate B's range. **The dominant class is verification fragments that count zero** — the reviewer tested them against the real files and found five of Task 7's ten locators, both of Task 3's, Task 5's and Task 7's strict-reading OLD all wrapping across lines or quoting text that does not exist. Repaired by one **verified fragment table** for the whole plan instead of guessed fragments per task | +| 2 | — | — | — | — | not run | next, against the pass-1 repair commit | + +## Pass-1 report + +**Trend:** first pass, 32 findings, 1 Blocker, 28 Majors. **Cluster:** verification mechanics — +roughly half the findings are fragments that count zero against the real files. **Require↔withdraw:** +none. + +**What the reviewer did that made this pass worth its cost:** it ran every `grep` pattern the plan +contained against `CLAUDE.md` and the template. Nine of them return zero in a correct tree — five +wrap across a line break, one quotes `These rules and records are one contract` where the live text +says `These records are one contract`, one quotes an em-dash as a comma, and one names a fragment +preserved inside its own replacement. **This is the exact defect design §7 records from four +consecutive spec revisions**, reproduced by me on the first try despite the design warning about it +by name. + +**The repair is structural rather than nine corrections.** The plan now carries **one fragment +table**, every row verified with a checker script against both copies and reporting the line it sits +on, and every task cites a row rather than inventing a pattern. The three ways a fragment fails — +wrapped, preserved inside its replacement, not unique — are stated there once. The one thing the +table cannot pre-verify is the NEW half, which does not exist until a task installs it; that is +disclosed in the table rather than hidden, and each task's step fails if its chosen fragment is not +single-line and unique in the installed file. + +**Two findings were the plan installing text contrary to the approved target** — Task 3 preserving +W's `b3` divergence that design §6 decides against, and Task 5 keeping W's missing pronoun. Both now +install what the target says, and Task 14 verifies rather than repairs. + +**The three Minors were repaired rather than collected**, and the reason is that each was folded +into an edit a Major already required: the site count contradicting its own enumeration (8/9/10 in +one task), `c20` recorded as carried while the change narrows its scope — the dropped-condition +failure `AGENTS.md` names — and the two plan-mutating tasks having no stable region to replace. None +cost a pass of its own. diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index 3fe4415..9e0d101 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -4,7 +4,7 @@ **Goal:** Install one closure ordering into both §5 copies, replace the twenty-three standing sentences it falsifies across those copies and the shipped hook, and ship the result as plugin 0.12.0. -**Architecture:** Every string this change installs is already written in final form in the target text. This plan does not restate any of it. Each task names the **site**, quotes the **anchor** it installs at, cites the **target-text section** whose fenced block is the bytes to install, and builds the **discriminating pair of counts** that shows the new wording present and the old wording gone. Copying the replacement text into this plan would create the second-copy defect the whole cycle fought; a citation into an approved artifact that travels with this plan is not a placeholder. +**Architecture:** Every string this change installs is already written in final form in the target text. This plan does not restate any of it. Each task names the **site**, quotes the **anchor** it installs at, cites the **target-text section** whose fenced block supplies the wording, and runs the **discriminating pair of counts** — the new wording present, the old wording gone — from the one verified fragment table below. Copying the replacement text into this plan would create the second-copy defect the whole cycle fought; a citation into an approved artifact that travels with this plan is not a placeholder. **Tech Stack:** Markdown prompt text, POSIX `sh` (the hook), `grep`/`diff` for verification, `shellcheck`, the `claude` CLI. @@ -19,7 +19,7 @@ - **Both prompt copies take every NEW and REPLACED section byte-identical**, except §F's seven hook items, whose destination is the shipped hook and its test and which carry no parity obligation (target §"How to read a section", §F opening). - **C** = `CLAUDE.md`. **W** = `plugins/dev-workflow/commands/workflow-init.md`. Every line number below is re-read at execution; the inventory's numbers cite `7c0d475` and have drifted. - **Every hook replacement installs into a double-quoted POSIX-shell `note` argument** and therefore carries no backtick, no `$(`, no backslash and no double quote. `$policy`, `$floor`, `$passes`, `$passesA` and `$fresh` are the intended interpolations (target §F opening). -- **Install every replacement unwrapped** — as one line in the file — because a counted fragment must be single-line for `grep -F` to find it (design §7). +- **The target's fenced blocks are normative in their words, not in their line breaks.** The spec wraps for its own readability; each copy keeps its own wrapping style. What design §7 requires is narrower and is the rule here: **every fragment this plan counts must sit wholly within one line of the file it is grepped from**, and where installing a replacement would put a counted fragment across a wrap, that fragment's line is installed unwrapped. The fragment table below states, for every count, the line it must sit on. **Nothing in this plan claims the installed bytes equal the fenced block's bytes**, and no check asserts it. - **Invariant 5 (exact pinning)** and **invariant 12 (a plugin change requires a version bump)**: this change touches `plugins/`, so `plugins/dev-workflow/.claude-plugin/plugin.json` goes `0.11.0 → 0.12.0` with a CHANGELOG entry (design §8). - **Invariant 11:** every prompt change passes all 12 items of `docs/prompt-standards.md`. Most at risk: item 6 (every constraint carries its reason in the same sentence) and item 8 (token-lean). - **Invariant 4 / the hook:** no control flow, counter, fingerprint computation, routing or event handling changes. Only `note` strings and their expectations. @@ -27,6 +27,52 @@ --- +## Verification fragments — verified against the real files + +**Every fragment below was tested with `python3 /frag.py ""` against `CLAUDE.md` and the template at `5871d0a`, and each returned exactly one hit per copy on the line shown.** No task may invent a fragment; a task needing one not listed here adds it to this table and re-runs that check first. This table exists because the first draft scattered guessed fragments across ten tasks and pass 1 found that most of them wrapped across lines and counted zero in a correct tree — the exact defect design §7 records from four consecutive spec revisions. + +**Three ways a fragment fails, all of which this table's check catches:** it wraps across a line break, so `grep -F` counts zero in a correct file; it is **preserved inside its own replacement**, so its old-wording-gone count can never reach zero; or it is not unique, so a count of 1 proves nothing about which occurrence changed. + +| # | Edit | OLD fragment (single-line, not preserved in its replacement) | C | W | +|---|---|---|---|---| +| P1 | §A1/A2/A3 | *(none — counterfactual ABSENT, presence only)* | — | — | +| P2 | `b7`, the fix-set definition | `scope the approved story or plan assigns to this cycle, plus repair obligations you already` | 201 | 408 | +| P3 | `b12`, immediate resumption | `the moment the user says whether the set now includes it` | 208 | 415 | +| P4 | `c14`, below-floor Minor | `a Blocker/Major-free pass below the floor` | 239 | 442 | +| P5 | `e7`, the threshold — **C** | `the tells and hand the decision to the user, and the` | 267 | — | +| P5w | `e7`, the threshold — **W** | `tells and hand the decision to the user, and the` | — | 471 | +| P6 | `g1`/`g2`, the handed-over question | `is not settled here, and this change does not settle it` | 811 | 997 | +| P7 | §G, the one-contract paragraph | `These records are one contract` | 879 | 1063 | +| P8 | `c18`, no-clean-credit | `the resolve rule is not waived, no pass is credited as` | 243 | 446 | +| P9 | `a13`, the no-restating prohibition | `Every other rule stated here about how a cycle closes` | 129 | 336 | +| P10 | `a16`, the per-pass fix command | `fix Blocker/Major after each` | 132 | 339 | +| P11 | `a17`, the clean-final-pass rule | `final pass must be clean` | 133 | 340 | +| P12 | Gate-A clean signal | `when a pass is clean` | 565 | 757 | +| P13 | gate-prompt template clean sentence | `clean pass is the single body line` | 329 | 523 | +| P14 | Gate-A cadence | `Each pass: validate, revise, re-run` | 573 | 764 | +| P15 | lens unchanged-list | `The Blocker/Major filter, the file-first findings protocol` | 653 | 839 | +| P16 | strict-reading list | `nonce duties at their strictest` | 157 | 364 | +| P17 | §F item 14, Named residual | `Hook text is out of scope here` | 139 | 346 | +| P18 | §F item 18, work-loop line | `execute → tests green → Gate B → commit` | 63 | 262 | + +**`e7` is the one edit needing a per-copy fragment**, because W drops the pronoun: C reads `you report the tells`, W reads `report the tells`. That is the recorded `e8` divergence, and it **does not survive this change** — see Task 5. + +**The NEW fragment for every pair is taken from the installed line and checked the same way**, since the new wording does not exist until the task installs it. Each task's step says which sentence of its target block to take it from, and the step fails if the fragment it chooses is not single-line and unique in the installed file. **This is the one place the plan cannot pre-verify**, and it is disclosed rather than papered over. + +**The four-value command, defined once and cited by number afterwards:** + +```bash +pair() { # pair + printf '%s old/worktree=%s old/parent=%s new/worktree=%s new/parent=%s\n' "$3" \ + "$(grep -cF "$1" "$3")" "$(git show "$BASE:$3" | grep -cF "$1")" \ + "$(grep -cF "$2" "$3")" "$(git show "$BASE:$3" | grep -cF "$2")" +} +``` + +**A pair passes only on `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`.** All four values matter: a copy carrying the new wording **and** the old one satisfies a one-sided presence check and is exactly the two-instructions-that-disagree failure the pair exists to catch. + +--- + ## File Structure | File | Responsibility in this change | @@ -79,7 +125,7 @@ Story acceptance criterion 5 is satisfied here. Ids are `docs/superpowers/specs/ | c17 | **replaced** — the resolve rule now scopes to the assigned fix set and states what a validly dismissed recurrence owes. | | c18 | **replaced** — the blanket no-clean-credit goes; a pass is credited on its own findings, and a scope-stop trigger is what withholds credit. | | c19 | **replaced** — the one-answer resumption goes; what the answer does is the ordering's. | -| c20 | **carried**, with "Blocker and Major" narrowed to "**in-set** Blocker and Major". | +| c20 | **replaced.** The prohibition on the "stop instead of fixing" reading survives, but its scope narrows from "every Blocker and Major" to "every **in-set** Blocker and Major" — a meaning change, installed by §H's `c18`-and-surfacing block. **Marked replaced rather than carried**, because a narrowed condition recorded as preserved is exactly the dropped-condition failure AGENTS.md requires this accounting to expose. | ### Passage (d) — from pass 4 onward (Task 0 check only) @@ -140,23 +186,51 @@ Story acceptance criterion 5 is satisfied here. Ids are `docs/superpowers/specs/ - [ ] **Step 1: Confirm the approved artifacts and a clean tree** ```bash -git log --oneline -1 # expect ba15e83 or later on loop-rule-consolidation +git log --oneline -1 # expect 5871d0a or later on loop-rule-consolidation git status --porcelain # expect empty -BASE=$(git rev-parse HEAD) # every counterfactual count runs against this -echo "$BASE" +git rev-parse HEAD > .context/loop-rule-base +cat .context/loop-rule-base +``` + +**Persist it to a file, not to a shell variable.** Each task runs in its own shell invocation, so a +`BASE=` assignment in Task 0 is gone by Task 1 and every parent-tree count would run against an +empty revision — which fails loudly in `git show` but quietly in a `grep -c` pipeline. Every later +task begins with: + +```bash +BASE=$(cat .context/loop-rule-base) ``` +**This commit is also `baseSha` for Gate B.** It is the parent of the first WIP snapshot, and it is +the only value that puts the whole implementation inside the reviewed range. `.context/` is ignored +by the hook's fingerprint, so the file itself moves nothing. + - [ ] **Step 2: Re-read the five untouched ranges and record their current line numbers** Passages (d), (f) and (j) are not edited by this change, and passage (a)'s arithmetic and passage (h)'s twenty-three kept conditions are mostly untouched. Record where they are now, because the inventory's numbers cite `7c0d475`: +**Five ranges, each with a start and an end anchor** — the three whole passages, plus the floor +arithmetic and the kept human-exception conditions, which later verification consumes and which the +first draft omitted: + ```bash -grep -n 'From pass 4 onward every pass report carries three lines' CLAUDE.md plugins/dev-workflow/commands/workflow-init.md -grep -n 'The two rules above do not compete' CLAUDE.md plugins/dev-workflow/commands/workflow-init.md -grep -n 'On squash-merge, copy every evidence entry' CLAUDE.md plugins/dev-workflow/commands/workflow-init.md +for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do + echo "== $f" + grep -n 'From pass 4 onward every pass report carries three lines' "$f" # (d) start + grep -n 'Those three lines expose' "$f" # (d) end + grep -n 'The two rules above do not compete' "$f" # (f) start + grep -n 'Findings go to a FILE' "$f" # (f) end + grep -n 'On squash-merge, copy every evidence entry' "$f" # (j) single line + grep -n 'Both gates are a LOOP with a HARD FLOOR' "$f" # (a) arithmetic start + grep -n 'Nothing here writes the floor knob' "$f" # (a) arithmetic end + grep -n 'Recording a human exception' "$f" # (h) start + grep -n 'because writing it down makes it sound' "$f" # (h) end +done ``` -Expected: two hits each, one per copy. Write the six line numbers into a scratch note — Task 14 diffs against them. +Expected: one hit per pattern per file. Write all the line numbers into +`.context/loop-rule-untouched` — **Task 2 and the final check in Task 14 both read them**, so a +scratch note that dies with the session is not enough. - [ ] **Step 3: Confirm the parity baseline of the inventoried ranges** @@ -209,18 +283,27 @@ Insert the three blocks before the anchor line, blank-line separated, byte-ident **Placement constraint from the disposition table:** §A must not land between passage (b) and passage (c), because `f1` ("the two rules above") names those two and would then name the wrong pair. Inserting *before* (b) satisfies this. -- [ ] **Step 4: Run the presence counts in both copies and both trees** +- [ ] **Step 4: Run a presence count per paragraph, in both copies and both trees** + +**Three counts, not one.** A single count on §A1's opening lets §A2 or §A3 be omitted from both +copies while every count and the parity diff still pass — the paragraphs are installed together and +nothing else observes them. ```bash +BASE=$(cat .context/loop-rule-base) for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do - printf '%s worktree: ' "$f" - grep -cF 'How a cycle ends — one ordering, stated here and referenced everywhere else' "$f" - printf '%s parent: ' "$f" - git show "$BASE:$f" | grep -cF 'How a cycle ends — one ordering, stated here and referenced everywhere else' + for frag in 'How a cycle ends — one ordering' \ + '' \ + ''; do + printf '%s | worktree=%s parent=%s | %s\n' "$f" \ + "$(grep -cF "$frag" "$f")" "$(git show "$BASE:$f" | grep -cF "$frag")" "$frag" + done done ``` -Expected: worktree `1`, parent `0`, for both files. A parent count above zero means `$BASE` is wrong. +Expected: `worktree=1 parent=0` for all six. **Add the two §A2/§A3 fragments to the fragment table +once chosen**, with the check that verified each is single-line and unique — they are the two rows +that cannot be pre-verified because the text does not exist until this task installs it. - [ ] **Step 5: Check parity of the installed block** @@ -281,36 +364,56 @@ Read the installed text around "The two rules above do not compete" and confirm - Modify: `CLAUDE.md` — the passage beginning `**What a loop absorbs, and what stops it` - Modify: `plugins/dev-workflow/commands/workflow-init.md` — the same -**The bytes:** target §B, which states it is "the only place this file states anything about" passage (b). **Preserve the two deliberate divergences** the inventory records and §B confirms: W says `the severity rule` where C says `Mechanics` (b3); and the field-mint parenthetical closes the paragraph in C and is absent from W. §B says it stops short of that parenthetical, which is left exactly as each copy has it. +**The bytes:** target §B, which states it is "the only place this file states anything about" passage (b). -- [ ] **Step 1: Record the old-wording fragments** +**`b3` is aligned here, not in Task 14.** Design §6 decides it explicitly: W's pointer names "the +severity rule" on the inventory's reasoning that W has no Mechanics section, **which is false**, so +**W takes C's wording**. Installing W's old wording here and repairing it two tasks later would have +this task knowingly install text contrary to the approved target, and would make Task 14 repair a +defect this task introduced. + +**One divergence is preserved:** the field-mint parenthetical closes the paragraph in C and is +absent from W. §B says it stops short of that parenthetical, which is left exactly as each copy has +it. -Two fragments that must reach zero, chosen because each is single-line in the file and is **not** preserved inside the replacement: +- [ ] **Step 1: Confirm the two OLD fragments still count 1** + +Rows **P2** (`b7`, the fix-set definition) and **P3** (`b12`, immediate resumption) of the fragment +table. Both are verified single-line and unique; re-confirm before editing, since earlier tasks +have touched these files: ```bash -grep -cF 'plus repair obligations you already accepted in earlier passes' CLAUDE.md -grep -cF 'it resumes the moment the user says whether the set now includes it' CLAUDE.md +BASE=$(cat .context/loop-rule-base) +grep -cF 'scope the approved story or plan assigns to this cycle, plus repair obligations you already' CLAUDE.md +grep -cF 'the moment the user says whether the set now includes it' CLAUDE.md ``` -Expected before the edit: `1` each. Re-check both against W with the same command. +Expected: `1` each, and the same against W. + +**The first draft of this plan named `plus repair obligations you already accepted in earlier +passes` here, which wraps across C 201–202 and W 408–409 and counts zero in a correct file.** - [ ] **Step 2: Install §B's text over the passage in both copies** Replace from `**What a loop absorbs, and what stops it` through the sentence §B ends at, keeping each copy's own closing parenthetical. -- [ ] **Step 3: Run the discriminating pair, both copies, both trees** +- [ ] **Step 3: Run two discriminating pairs, both copies, both trees** + +`b7` and `b12` are separate meaning changes and each owes its own pair; one pair covering both +would let the surviving instruction pass behind the repaired one. ```bash -NEW='union of the scope every approved story or plan governing this change assigns to this cycle' -OLD='plus repair obligations you already accepted in earlier passes' +BASE=$(cat .context/loop-rule-base) +# pair() is defined once in the fragment-table section for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do - printf '%s new/worktree=%s new/parent=%s old/worktree=%s old/parent=%s\n' "$f" \ - "$(grep -cF "$NEW" "$f")" "$(git show "$BASE:$f" | grep -cF "$NEW")" \ - "$(grep -cF "$OLD" "$f")" "$(git show "$BASE:$f" | grep -cF "$OLD")" + pair 'scope the approved story or plan assigns to this cycle, plus repair obligations you already' \ + 'union of the scope every approved story or plan governing this change assigns to this cycle' "$f" + pair 'the moment the user says whether the set now includes it' \ + '' "$f" done ``` -Expected per file: `new/worktree=1 new/parent=0 old/worktree=0 old/parent=1`. **All four values matter** — a copy carrying both the new and the old wording satisfies a one-sided presence check and is exactly the two-instructions-that-disagree failure this pair exists to catch. +Expected for all four: `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`. - [ ] **Step 4: Walk the carried conditions** @@ -366,9 +469,19 @@ The precedence clause must appear **once**, inside the ordering. A second occurr - [ ] **Step 4: Run the discriminating pair, both copies, both trees** -Same four-value shape as Task 3, with `OLD='a Blocker/Major-free pass below the floor'` and `NEW` a single-line fragment of §C's re-raised-dismissal clause taken from the installed file. +Row **P4**. `NEW` is the single-line fragment `a recurrence failing them being an ordinary fresh +finding`, from §C's re-raised-dismissal clause — **install that clause's line unwrapped** so the +fragment sits wholly on one line, and confirm it is unique before counting. -Expected: `new/worktree=1 new/parent=0 old/worktree=0 old/parent=1`. +```bash +BASE=$(cat .context/loop-rule-base) +for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do + pair 'a Blocker/Major-free pass below the floor' \ + 'a recurrence failing them being an ordinary fresh finding' "$f" +done +``` + +Expected: `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`. - [ ] **Step 5: Walk the conditions** — `c1`–`c3` present unchanged, `c5`–`c7` word for word, `c8` carrying the new clause, `c9` split, `c10`–`c14` gone from here. @@ -389,27 +502,53 @@ git commit -m "WIP: widen the clearly-stuck third condition and split its preced **The bytes:** target §D's two fenced blocks — the `e7` sentence entire, and the pointer paragraph added at the end of the passage. -**Parity note:** W drops the pronoun (`report the tells`, not `you report the tells`). §D's block carries C's wording; **W keeps its own pronoun** unless the parity divergence list decides otherwise in Task 14. Record which you chose — this is one of the two divergences the inventory flags for passage (e), and the other (e11, the C-only rationale paragraph) is untouched. +**`e8` does not survive this change.** §D supplies the complete replacement sentence, which reads +`you report the tells`, and the Global Constraints make every replaced section byte-identical across +the copies. **Install §D's sentence as written in both copies** — W gains the pronoun — and remove +`e8` from the surviving-divergence list Task 14 carries. Keeping W's old pronoun would mean +knowingly installing text the approved target does not say, and would leave Task 14 classifying as +deliberate a divergence this task chose to create. -- [ ] **Step 1: Record the old wording** +**`e11`, the C-only rationale paragraph, is untouched and stays C-only.** -```bash -grep -cF 'Any two present makes stop-and-surface mandatory, not discretionary' CLAUDE.md -``` +- [ ] **Step 1: Confirm the per-copy OLD fragments** -Expected `1` — but note this fragment **survives inside the replacement**, so it cannot be the old-wording half of the pair. Use instead: +`Any two present makes stop-and-surface mandatory, not discretionary` **survives inside the +replacement**, so it can never be the old half — that is the trap design §7 records from four spec +revisions. Rows **P5** (C) and **P5w** (W) instead, and they differ because W drops the pronoun: ```bash -grep -cF 'and the "clearly stuck" reading above is not a precondition for it' CLAUDE.md +grep -cF 'the tells and hand the decision to the user, and the' CLAUDE.md # 1 +grep -cF 'tells and hand the decision to the user, and the' plugins/dev-workflow/commands/workflow-init.md # 1 ``` -and pair it against the installed sentence's new clause. **This is the trap design §7 records:** four spec revisions named a fragment preserved inside its own replacement, whose old-wording-gone count could never reach zero. +**The first draft named `and the "clearly stuck" reading above is not a precondition for it`, which +wraps across C 267–268 and W 471–472 and counts zero in both correct files.** - [ ] **Step 2: Install §D's `e7` sentence and the pointer paragraph** -- [ ] **Step 3: Run the discriminating pair**, both copies, both trees, with `NEW='read **after** the clean-completion branch of the closure ordering'` and `OLD` from Step 1. +- [ ] **Step 3: Run the discriminating pair, and a presence check for the pointer** + +The `e7` replacement and §D's added pointer paragraph are two separate observations. **The pointer +is add-only** — it replaces no wording — so it is checked by presence alone, and without that check +it can be omitted from both copies while the `e7` pair, the condition walk and the parity diff all +pass. -Expected: `new/worktree=1 new/parent=0 old/worktree=0 old/parent=1`. +```bash +BASE=$(cat .context/loop-rule-base) +pair 'the tells and hand the decision to the user, and the' \ + 'read **after** the clean-completion branch of the closure ordering' CLAUDE.md +pair 'tells and hand the decision to the user, and the' \ + 'read **after** the clean-completion branch of the closure ordering' plugins/dev-workflow/commands/workflow-init.md +for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do + printf 'pointer %s worktree=%s parent=%s\n' "$f" \ + "$(grep -cF 'where this stop'"'"'s place among the suspensions' "$f")" \ + "$(git show "$BASE:$f" | grep -cF 'where this stop'"'"'s place among the suspensions')" +done +``` + +Expected: both pairs `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`; both pointer checks +`worktree=1 parent=0`. - [ ] **Step 4: Confirm `e1`–`e6` and `e8`–`e11` are untouched**, and that e11 is still C-only. @@ -458,7 +597,18 @@ grep -c '2026-08-29-loop-rule-consolidation-story.md' CLAUDE.md Expected: `0`. The story path may still appear in `docs/` — this check is scoped to `CLAUDE.md`. -- [ ] **Step 5: Parity** — this passage should now be byte-identical, the one recorded divergence having been removed. Diff the two Severity bullets and expect no output. +- [ ] **Step 5: Parity over both §E blocks** + +**The bullet and the answer paragraph are two installs and the diff must cover both** — scoping the +check to the Severity bullet alone lets the long answer paragraph differ between the copies while +the pair and the stated parity check pass. + +```bash +diff <(sed -n '/^- \*\*Severity:\*\*/,/^- \*\*Tool routing:/p' CLAUDE.md) \ + <(sed -n '/^- \*\*Severity:\*\*/,/^- \*\*Tool routing:/p' plugins/dev-workflow/commands/workflow-init.md) +``` + +Expected: no output. This passage should now be byte-identical, `g4` having been removed and `g2`/`g3` having gone from both copies. - [ ] **Step 6: Commit** @@ -480,46 +630,77 @@ git commit -m "WIP: answer the demotion question and scope the resolve duty to t **`i12` is the licence for the strict-reading addition and stays in place.** The list is extended at the end, not rewritten; §H gives the whole dash-delimited list so one contiguous string installs. -- [ ] **Step 1: Locate all eight sites in both copies** +- [ ] **Step 1: Locate all ten sites in both copies** + +**Ten sites: one §G block and nine §H blocks.** Every locator below is a fragment-table row, so each +is verified single-line and unique. The first draft used five locators that count zero in a correct +file — `These rules and records are one contract` (the live text says `These records are one +contract`), `Your final pass must be clean` (wraps C 132–133), ``a literal `NO FINDINGS` when a pass +is clean`` (one line in C, wraps in W), `A clean pass is the single body line` (the line break falls +after `A`), and `at minimum floor 3, severity classified without the demotion` (wraps C 155–156). ```bash for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do echo "== $f" - grep -n 'These rules and records are one contract' "$f" - grep -n 'Surfacing does not close the cycle' "$f" - grep -n 'Every other rule stated here about how a cycle closes' "$f" - grep -n 'Open a TodoWrite' "$f" - grep -n 'Your final pass must be clean' "$f" - grep -n 'literal `NO FINDINGS` when a pass is clean' "$f" - grep -n 'A clean pass is the single body line' "$f" - grep -n 'Each pass: validate, revise, re-run' "$f" - grep -n 'Lenses are \*\*different questions, not more passes\*\*' "$f" - grep -n 'at minimum floor 3, severity classified without the demotion' "$f" + grep -cF 'These records are one contract' "$f" # P7 §G + grep -cF 'the resolve rule is not waived, no pass is credited as' "$f" # P8 c18 + grep -cF 'Every other rule stated here about how a cycle closes' "$f" # P9 a13 + grep -cF 'fix Blocker/Major after each' "$f" # P10 a16 + grep -cF 'final pass must be clean' "$f" # P11 a17 + grep -cF 'when a pass is clean' "$f" # P12 Gate-A clean signal + grep -cF 'clean pass is the single body line' "$f" # P13 template clean sentence + grep -cF 'Each pass: validate, revise, re-run' "$f" # P14 Gate-A cadence + grep -cF 'The Blocker/Major filter, the file-first findings protocol' "$f" # P15 lens list + grep -cF 'nonce duties at their strictest' "$f" # P16 strict-reading list done ``` -Expected: one hit per pattern per file. A pattern with zero hits means the wording drifted since the inventory — find it before installing. +Expected: `1` twenty times. **Any `0` means the wording drifted since this table was verified at +`5871d0a`** — re-derive that fragment and update the table before installing anything. -- [ ] **Step 2: Install §G and all eight §H blocks, one site at a time, verifying each before moving to the next** +- [ ] **Step 2: Install the §G block and all nine §H blocks — ten sites — one at a time, verifying each before moving to the next** -- [ ] **Step 3: Run one discriminating pair per block**, both copies, both trees. Old-wording fragments, each single-line and none preserved inside its own replacement: +- [ ] **Step 3: Run one discriminating pair per block**, both copies, both trees, one per row P7–P16. -| Block | OLD fragment | -|---|---| -| §G one-contract | `These rules and records are one contract` | -| `c18`/surfacing | `no pass is credited as clean` | -| `a13` | `Every other rule stated here about how a cycle closes` | -| `a16` | `fix Blocker/Major after each` | -| `a17`–`a22` | `Your final pass must be clean` | -| Gate-A clean signal | `a literal \`NO FINDINGS\` when a pass is clean` | -| template clean sentence | `A clean pass is the single body line` | -| Gate-A cadence | `Each pass: validate, revise, re-run` | -| lens unchanged-list | `The Blocker/Major filter, the file-first findings protocol` | -| strict-reading list | `the nonce duties at their strictest, the cycle is treated as post-rule` | +`NEW` for each is taken from the installed block; **check each chosen fragment is single-line and +unique in the installed file before counting it**, and add it to the fragment table. The suggested +source sentence per block: + +| Row | Block | NEW taken from | +|---|---|---| +| P7 | §G one-contract | the widened-membership clause of §G's block | +| P8 | `c18`/surfacing | `A pass is credited clean or not on its own findings` | +| P9 | `a13` | `Every other rule stated **in this paragraph**` | +| P10 | `a16` | `resolve Blocker/Major after each as` | +| P11 | `a17`–`a22` | `What a clean final pass and the zero-finding early exit mean for closing` | +| P12 | Gate-A clean signal | `and no scope-stop trigger** is clean too` | +| P13 | template clean sentence | `A **clean findings file** is the single body line` | +| P14 | Gate-A cadence | `revise **where a repair is required**` | +| P15 | lens unchanged-list | `**the lens sets** leave every other` | +| P16 | strict-reading list | `every suspension binding, since` | + +```bash +BASE=$(cat .context/loop-rule-base) +for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do + pair 'These records are one contract' '' "$f" + pair 'the resolve rule is not waived, no pass is credited as' '' "$f" + pair 'Every other rule stated here about how a cycle closes' '' "$f" + pair 'fix Blocker/Major after each' '' "$f" + pair 'final pass must be clean' '' "$f" + pair 'when a pass is clean' '' "$f" + pair 'clean pass is the single body line' '' "$f" + pair 'Each pass: validate, revise, re-run' '' "$f" + pair 'The Blocker/Major filter, the file-first findings protocol' '' "$f" + pair 'nonce duties at their strictest' '' "$f" +done +``` -Expected for each: `old/worktree=0 old/parent=1`, and the matching new fragment `1` / `0`. +Expected for all twenty: `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`. -**The strict-reading list is add-only at its tail but replaces the dash-delimited run**, so it owes a full pair rather than presence alone. **The gate-prompt template's clean sentence is add-only** if the site carries no wording the change removes — classify it against the real file, per design §7, and check by presence alone if so. +**Two classification notes, decided against the real files rather than asserted.** The +strict-reading list **replaces the dash-delimited run** even though its addition is at the tail, so +it owes a full pair. The gate-prompt template's clean sentence **replaces** `A clean pass is the +single body line …`, which is why P13 has an OLD at all; it is not add-only. - [ ] **Step 4: Confirm `a21`, `a22`, `a15` survived** — they are carried inside blocks that install contiguously, so a mis-scoped replacement silently drops them. @@ -530,7 +711,7 @@ grep -cF 'Open a TodoWrite' CLAUDE.md Expected: `1` each. -- [ ] **Step 5: Parity** for all nine sites. +- [ ] **Step 5: Parity** for all ten sites, each extracted by its own bounded region rather than a fixed line window. - [ ] **Step 6: Commit** @@ -562,9 +743,26 @@ grep -n -F '' CLAUDE.md plugins/dev-workflow/comma - [ ] **Step 2: Install all fourteen replacements** -- [ ] **Step 3: Run one discriminating pair per item**, both copies, both trees, choosing each OLD fragment single-line and not preserved inside its replacement. +- [ ] **Step 3: Build this task's fourteen fragment rows, then run fourteen pairs** -Expected per item: `new/worktree=1 new/parent=0 old/worktree=0 old/parent=1`. +**Do not choose fragments at run time.** Step 1 located each item's live sentence; for each, pick a +single-line OLD fragment from the located line, run it through the fragment check, and **append the +row to the plan's fragment table with the line numbers the check reported**. Only then run the +pairs. A fragment chosen and used in the same breath is how the first draft shipped five that count +zero. + +For each item, the OLD fragment must satisfy all three: single-line in both copies (or one fragment +per copy where the wrapping differs, as `e7` needed); unique in each; and **not preserved inside its +own replacement** — compare it against the §F block before accepting it. + +```bash +BASE=$(cat .context/loop-rule-base) +for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do + pair '' '' "$f" # ×14 +done +``` + +Expected for all twenty-eight: `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`. - [ ] **Step 4: Count what was installed** @@ -641,14 +839,28 @@ git commit -m "WIP: correct the Named residual's blanket exemption and the work- **Shell constraint, and it is not advisory:** every string installs into a double-quoted `note` argument. No backtick, no `$(`, no backslash, no double quote. A backtick reached §F once and would have executed `WIP` at install time, shipped the reminder with the word missing, and failed ShellCheck — while the fixture copied from it would have made the suite pass on the corruption. -- [ ] **Step 1: Verify the constraint before installing** +- [ ] **Step 1: Verify the constraint before installing, reading each block as inert data** -```bash -# For each of the seven §F blocks, confirm the text you are about to install is clean: -printf '%s' "" | grep -nE '`|\$\(|\\\\|"' && echo HAZARD || echo clean +**The probe must never put the candidate text inside double quotes.** A backtick or `$(` there is +executed by the probe itself, `$policy` expands away, and the check then certifies a string that is +not the one to be installed — the probe would run the hazard it exists to detect. + +Read the blocks straight from the spec file instead, and test the literal characters: + +```python +# python3 - <<'EOF' (single-quoted heredoc: the shell expands nothing) +import re, pathlib +spec = pathlib.Path("docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md").read_text() +sec = spec[spec.index("## F."):spec.index("## G.")] +for i, block in enumerate(re.findall(r"```\n(.*?)\n```", sec, re.S), 1): + bad = [c for c in ("`", "$(", "\\", '"') if c in block] + print(i, "HAZARD" if bad else "clean", bad, block[:60]) +# EOF ``` -Expected: `clean` seven times. A hit here stops the task. +Expected: every block destined for the hook prints `clean`. **Blocks for the prompt copies may +legitimately contain backticks and quotes** — they are markdown, not shell arguments — so read the +block index against §F's item numbers rather than requiring all of them clean. - [ ] **Step 2: Locate the seven `note` calls** @@ -660,13 +872,23 @@ Expected: thirteen hits. Seven are the gate reminders this task edits; four take - [ ] **Step 3: Install the seven replacements, one at a time** -- [ ] **Step 4: Confirm no behaviour changed** +- [ ] **Step 4: Confirm no behaviour changed, by line number rather than by shape** + +**A content filter cannot do this.** A pattern admitting "letters and allowed punctuation" admits +`else`, `fi` and `return` — a control-flow edit would pass the check it exists to fail. Compare the +**changed line numbers** against the seven `note` calls instead: ```bash -git diff plugins/dev-workflow/hooks/codex-gate.sh | grep -E '^[-+]' | grep -vE '^[-+]\s*(note "|[A-Za-z ,.—;:()/$-]+")' | head +BASE=$(cat .context/loop-rule-base) +git diff -U0 "$BASE" -- plugins/dev-workflow/hooks/codex-gate.sh \ + | grep -E '^@@' | sed -E 's/^@@ -([0-9]+).*/\1/' +grep -n 'note "' plugins/dev-workflow/hooks/codex-gate.sh ``` -Expected: only the `---`/`+++` header lines. **Any changed line that is not inside a `note` string is out of scope** — no control flow, no counter, no fingerprint computation, no routing (invariant 4, design §8). +**Every changed hunk must start inside one of the seven gate-reminder `note` calls this task +edits.** A hunk anywhere else is a behaviour change and is out of scope — no control flow, no +counter, no fingerprint computation, no routing (invariant 4, design §8). Read the two lists side by +side; do not automate the comparison with a pattern, which is what failed here the first time. - [ ] **Step 5: Discriminating pairs, worktree and parent** @@ -686,7 +908,16 @@ for pair in 'this floor is the only thing keeping the spec review honest|instruc done ``` -Expected: `old/worktree=0`, `old/parent` ≥ 1, `new/worktree` ≥ 1, `new/parent=0` for each. The `STOP` pair covers two messages, so its counts are 2 rather than 1 — **state the number you observed rather than asserting it**. +**Two more pairs, because four cover only five of the seven items.** Without them, item 11's old +abbreviated clean definition and item 17's running-cycle skip exit can both survive while every +listed pair passes: + +```bash + 'floor met by COUNT ONLY|Proceed only once this Gate-A cycle has closed' # item 11 + 'or proceed only if $policy|skip rule decides only whether a cycle runs at all' # item 17 +``` + +Expected: `old/worktree=0`, `old/parent` ≥ 1, `new/worktree` ≥ 1, `new/parent=0` for each of the six. The `STOP` pair covers two messages, so its counts are 2 rather than 1 — **state the number you observed rather than asserting it**. - [ ] **Step 6: ShellCheck** @@ -741,10 +972,12 @@ Expected: exit 0 from both. **`HOOK_SH` selects the shell the hook runs under; w - [ ] **Step 5: Confirm no verdict vocabulary survives** ```bash -grep -n 'Gate B satisfied\|Gate B not satisfied' plugins/dev-workflow/hooks/codex-gate.test.sh +grep -c 'Gate B satisfied\|Gate B not satisfied' plugins/dev-workflow/hooks/codex-gate.test.sh ``` -Expected: no hits, or only hits you can justify one by one in the commit body. +Expected: **`0`. Not "no hits you cannot justify"** — step 3 requires every label and comment to move +to observed hook state, and a justify-in-the-commit-body escape hatch is how the old gate-verdict +vocabulary survives a sweep. Where a test still needs that case, name it by what the hook observed. - [ ] **Step 6: ShellCheck the test file** @@ -801,6 +1034,11 @@ The two copies are byte-identical over this material, so a divergence here is a **The oracle.** A row **fails** when its required answer does not produce a **distinct** resumable or closed state — the same stop returning with its reading unconsumed, that is, without an intervening validated pass run after the answer — or when it closes on anything other than the route the block states. **Read the closure conditions and the routes off the installed §A, not from this plan** — an embedded copy can pass while disagreeing with the text it checks. +**Where it goes:** this plan, under the heading `## Next-state table (Task 13 output)` at the end of +the document. **Task 13 replaces that whole section idempotently** — re-running it must not append a +second table. If the heading is absent, create it; if present, replace everything under it up to the +next `## `. + - [ ] **Step 1: Enumerate the rows** One row per (starting state, answer) pair the installed ordering admits. At minimum, and **this is a floor rather than the set**: a clean eligible pass with every condition met; a clean eligible pass with an unmet precondition; a clean pass below the floor; a zero-finding pass below the floor; a pass carrying a membership trigger, answered accept and answered decline; a pass carrying a new-question trigger, answered; a pass carrying both; a two-tell stop, answered; a clearly-stuck surface, answered; a source block raised before any pass was read; a source block raised on a pass already read; a closing act that does not complete and is repaired; a closing act that cannot be repaired; a `full` Gate-B pass with one branch clean and one not; the same complaint in both branch files under each of accept/accept, accept/decline, decline/accept and decline/decline. @@ -834,18 +1072,30 @@ Design §6: the two copies must agree on every rule this change ships. **The pla **One divergence is decided by the design and is not a judgement call:** W's `b3` pointer names "the severity rule" on the inventory's reasoning that W has no Mechanics section, **which is false** — so **W takes C's wording** (design §6, target §B). -- [ ] **Step 1: Extract and diff each changed passage** +- [ ] **Step 1: Extract and diff every changed site, from the target text's own markers** + +**Nine fixed 25-line windows are not the site list.** They omit §G, both gate-prompt instructions, +the human-exception edits, the WIP and finishing-cycle text, the evidence-revalidation trigger and +the curve rationale — every one of which this change ships, and any of which could differ between +the copies while a nine-window diff passes. + +**Drive the list from the artifact:** every section the target text marks NEW or REPLACED, and every +item in §F whose destination is a prompt copy. For each, extract a **bounded** region — from its +first line to the first line of the next passage, not a fixed count — and diff the two copies: ```bash -for anchor in 'How a cycle ends' 'What a loop absorbs' 'Recognizing "clearly stuck"' \ - 'Those three lines expose' '\*\*Severity:\*\*' 'Surfacing does not close' \ - 'When these rules bind' 'Named residual' 'The work loop includes'; do - echo "== $anchor" - diff <(grep -A25 -E "$anchor" CLAUDE.md) \ - <(grep -A25 -E "$anchor" plugins/dev-workflow/commands/workflow-init.md) -done +# one bounded region per changed site; $start and $end are that site's own anchors +diff <(sed -n "/$start/,/$end/p" CLAUDE.md) \ + <(sed -n "/$start/,/$end/p" plugins/dev-workflow/commands/workflow-init.md) ``` +**Fail this step where a site named by §§A–H has no region in your list** — that is the same +completeness failure as Task 13's per-condition checks, and it is caught the same way: by reading +the set off the artifact rather than from a list kept here. + +**Where it goes:** this plan, under the heading `## Divergence list (Task 14 output)` at the end of +the document, replaced idempotently on re-run by the same rule Task 13 uses. + - [ ] **Step 2: Classify every difference the diff reports** Three buckets: **deliberate and stays** (the field-mint parenthetical, e11, f5–f7's evidence framing, e8's pronoun — each recorded in the inventory); **not deliberate, align it**; **introduced by this change, fix it**. Write the list into this plan. @@ -854,6 +1104,22 @@ Three buckets: **deliberate and stays** (the field-mint parenthetical, e11, f5 - [ ] **Step 4: Re-run the diff** and confirm only the deliberate divergences remain. +- [ ] **Step 4b: Diff all five untouched ranges against the parent, after every text edit** + +Task 2 ran this after Task 1 only, and Task 14's parity diff compares C against W rather than either +against the parent. **A later task can alter an untouched condition identically in both copies and +every other check still passes** — the parity diff sees no difference and no pair covers text no +task claims to change. + +```bash +BASE=$(cat .context/loop-rule-base) +# for each of the five ranges recorded in .context/loop-rule-untouched, in both copies: +diff <(git show "$BASE:$f" | sed -n "$start,$end p") <(sed -n "$start,$end p" "$f") +``` + +Expected: no output, ten times. **This is the last check before Gate B** and it is the only one that +would catch an identical accidental edit in both copies. + - [ ] **Step 5: Commit** ```bash @@ -915,11 +1181,16 @@ It names: the battery run; **every pair this plan built, with its counts in each - [ ] **Step 6: Run Gate B** ```bash -git rev-parse HEAD # the WIP commit -git rev-parse HEAD^ # baseSha +BASE=$(cat .context/loop-rule-base) # the parent of the FIRST WIP, from Task 0 +git rev-parse HEAD # the current WIP tip ``` -`mcp__codex__review` with `reviewType: full`, `baseSha` = the WIP commit's parent, `headSha` = the full 40-character object name `HEAD` resolves to **at that moment**, resolved once and kept with each branch's result. Carry the story path and the evidence entry quoted verbatim. Write findings to `.context/codex-reviews/gate-b---pass-

.md` — **draw a fresh nonce for this cycle**; it is a different cycle from `awsf1ec771`. +**`baseSha` is `$BASE`, never `HEAD^`.** This plan makes a WIP commit per task, so `HEAD^` is the +parent of the *last* one and the review range would hold the version bump alone — Gate B would +close having reviewed none of the prompt or hook changes. `$BASE` is the parent of the first WIP +and is the only value whose range contains the whole implementation. + +`mcp__codex__review` with `reviewType: full`, `baseSha` = `$BASE`, `headSha` = the full 40-character object name `HEAD` resolves to **at that moment**, resolved once and kept with each branch's result. Carry the story path and the evidence entry quoted verbatim. Write findings to `.context/codex-reviews/gate-b---pass-

.md` — **draw a fresh nonce for this cycle**; it is a different cycle from `awsf1ec771`. **Standing lens, every call:** "which existing statements does this diff falsify?" and **name what this diff changes the size, value or position of** — a list, a count, a version, an identifier, a cited line — then grep for where each is described elsewhere. @@ -951,3 +1222,15 @@ The closing body carries: the validated evidence entry; the provenance line; the **One correction applied from this review:** §G was missing a task; it is now installed by Task 7, which names nine sites rather than eight. **One residual this plan does not close, stated rather than left to be found.** Nothing here establishes that the edit set is complete — it is the sites §F knows, and §F's own §I records that it cannot establish completeness either. Task 8's step 1 re-derives every citation against the real file, and Task 14's diff catches a copy that fell out of step; neither is a completeness check. The sweep for an affected site this text has not found is design §7's, owed by whoever executes, and a site found during execution is a finding against the spec rather than a gap in this plan. + +--- + +## Next-state table (Task 13 output) + +*Empty until Task 13 runs. Task 13 replaces this entire section.* + +--- + +## Divergence list (Task 14 output) + +*Empty until Task 14 runs. Task 14 replaces this entire section.* From b767a2a1bb0d5d117c1f35db5badaa0ee03eedec Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 12:26:28 +0200 Subject: [PATCH 100/181] docs(plans): apply Gate-A plan pass 2; the third fragment condition, and runnable checks MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Pass 1's repair introduced one verified fragment table. Pass 2 found three of its rows preserved inside their own replacements — e7 in both copies and the strict-reading list — so each would return old/worktree=1 after a correct edit and no task could reach its required result. This plan states all three failure modes; the checker I wrote tested only two. The repair is the checker. It now tests the third condition against the target text's fenced blocks rather than the whole file, which is the narrower and correct test: a fragment quoted in an item's rationale is not preserved by its replacement. Re-running it over every row found exactly the three the reviewer named. The fourteen §F OLD fragments are derived and committed rather than deferred to the executor as placeholders. Second class: commands that cannot run. pair() defined once while each task runs in its own shell; three steps stating an expected four-value result with no command producing one; $BASE used unreloaded, where an empty value makes git show read the index rather than the parent. Task 0 writes .context/loop-rule-pair.sh, every counting block sources it, and a guard fails loudly on an empty $BASE. Third: the untouched-range checks used absolute line numbers recorded before Task 1's insertion, and two of the five ranges contained conditions this change replaces — a correct implementation would have failed its own check. Ranges are anchors now and the contaminated two are split around the sites they exclude. Also: Task 12b, the completeness sweep design §7 and target §I assign to the plan and which no task performed; per-edit pairs where one pair covered several changes; exact per-pair counts where a floor was used; and the closing commit, which wrote its evidence into a WIP body that git reset --soft discards. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-plan-om0bdd7udh-pass-2.md | 33 ++ .../gate-a-plan-om0bdd7udh-resume.md | 44 ++- .../2026-09-14-loop-rule-consolidation.md | 320 +++++++++++++++--- 3 files changed, 354 insertions(+), 43 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-2.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-2.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-2.md new file mode 100644 index 0000000..c8550ce --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-2.md @@ -0,0 +1,33 @@ +MAJOR | high | Verification fragment P5 / Task 5 step 3 (CLAUDE.md) | OLD `the tells and hand the decision to the user, and the` occurs at CLAUDE.md:267 and is preserved verbatim in target §D at line 612 | After the replacement the pair returns `old/worktree=1`, so Task 5 cannot reach its required four-value result for C | Choose a C-specific OLD fragment that occurs once in the live sentence and does not occur in target §D +MAJOR | high | Verification fragment P5w / Task 5 step 3 (workflow-init.md) | OLD `tells and hand the decision to the user, and the` occurs at W:471 and is preserved verbatim in target §D at line 612 | After the replacement the pair returns `old/worktree=1`, so Task 5 cannot reach its required four-value result for W | Choose a W-specific OLD fragment that occurs once in the live sentence and does not occur in target §D +MAJOR | high | Verification fragment P16 / Task 7 step 3 | OLD `nonce duties at their strictest` occurs once at C:157 and W:364 but is also carried verbatim in target §H at line 1206 | Both strict-reading pairs return `old/worktree=1`, contradicting the required zero and preventing Task 7 from passing | Use an OLD fragment from the replaced tail of the dash-delimited run rather than from i8, which is explicitly kept +MAJOR | high | Task 8 step 3 | The fourteen OLD halves remain the placeholder `` even though all fourteen old sentences already exist in the real files; no rows for them appear in the verified fragment table | Gate A cannot verify that any of these OLD fragments is single-line, unique, or absent from its own replacement, and all fourteen meaning-changing checks are deferred to implementation | Derive and commit the fourteen concrete OLD rows now; leave only the NEW positions deferred as the plan permits +MAJOR | high | Verification fragments / Tasks 3–8 | `pair()` is defined in a standalone shell snippet once, while the plan explicitly says task invocations use separate shells; none of the later command blocks defines or sources it | The commands in Tasks 3, 4, 5, 7 and 8 fail with `pair: command not found` even when the edits are correct | Define `pair()` in every command block that invokes it or put it in a persisted scratch helper that every task sources +MAJOR | high | Task 6 step 3 | The step supplies OLD and NEW prose but contains no command that invokes `pair()` or prints the four counts | The handed-over-question replacement can be marked checked without any worktree or parent observation | Add a runnable loop over C and W that reloads `$BASE`, invokes the pair, and asserts 0/1/1/0 +MAJOR | high | Task 9 step 3 | The code block only assigns `OLD1`, `NEW1`, `OLD2`, and `NEW2`; it never counts or compares them in either file or tree | Both replacements can be committed with no discriminating pair having run despite the stated expected result | Invoke the four-value check for both variable pairs against C and W after reloading `$BASE` +MAJOR | high | Task 3 steps 3–4 | The pairs cover only b7 and b12, while target §B separately changes b3, b8, b11, b13, b16, b17–b18 and adds decline and closing-time set-change rules | Those edits can be omitted or old instructions can survive while both pairs pass; a manual condition walk is not the discriminating pair design §7 requires for each meaning change | Add a pair for each independent replacement and a presence-only check for each add-only rule in §B +MAJOR | high | Task 4 step 4 | The sole pair combines OLD from c14 with NEW from c8, two different edits, and has no OLD/NEW observation for c4 or for c14's new conditional routing | The c8 addition and c14 deletion can make the pair pass while other meaning-changing parts of §C remain old or absent | Give each independent §C change its own discriminating pair, using presence only where the edit is genuinely add-only +MAJOR | high | Task 6 steps 2–3 | §E contains two separate fenced replacements, but the only described pair is for the handed-over-question paragraph and none observes the resolve-duty bullet | The old unscoped resolve duty can survive, or the new repair/dismissal distinction can be omitted, while Task 6's pair and parity check pass | Add a dedicated OLD/NEW pair for the resolve-duty block +MAJOR | high | Task 3 step 5 | The plan says the §B parity diff should show three recorded divergences, but Task 3 installs target §B's common byte-identical span, aligning b3, the intensifier, and the closing rationale; only C's retained field-mint parenthetical should differ | A correct implementation appears to fail, while stale W wording that target §B removes can be accepted as expected drift | Change the expected result to only the C-only field-mint parenthetical and identify any other difference as a failure +MAJOR | high | Task 5 step 4 | The step says e8 is untouched, although the same task explicitly requires W to gain `you` and says the e8 pronoun divergence does not survive | An executor cannot satisfy both instructions and may preserve the target-forbidden W divergence | Require e1–e6 and e9–e11 to remain untouched, and verify e8 is aligned to target §D +MAJOR | high | Task 14 step 2 | The surviving-divergence bucket still lists `e8's pronoun`, directly contradicting Task 5 and target §D, which install `you report` in both copies | Task 14 can classify a parity defect as deliberate and leave the two shipped copies inconsistent | Remove e8 from the deliberate-and-stays list +MAJOR | high | Condition disposition, passage (a) | a2 is marked kept and untouched, but target §F item 8a and Task 8 replace its literal `(Blocker/Major only)` parenthetical with a broader cleanliness statement | The 135-condition accounting records a replaced condition as preserved, reproducing the dropped-condition failure AGENTS.md forbids | Mark a2 replaced and cite §F item 8a; state which part of the old condition remains and which meaning changes +MAJOR | high | Task 0 step 2 / Task 14 step 4b, floor range | The recorded range from `Both gates are a LOOP with a HARD FLOOR` through `Nothing here writes the floor knob` includes a2, which Task 8 replaces, and a13, which Task 7 replaces | The final parent diff must report intended changes, so the claimed ten no-output checks cannot pass | Bound the untouched floor slices around the exact kept conditions, excluding a2 and a13, or compare those conditions individually +MAJOR | high | Task 0 step 2 / Task 14 step 4b, human-exception range | The range from `Recording a human exception` through `because writing it down makes it sound` contains h4 and h19, both intentionally replaced by Task 8 | A correct implementation necessarily fails the final untouched-range check | Record separate spans for the twenty-four kept h conditions and exclude the two replacement sites +MAJOR | high | Task 14 step 4b | Task 0 records absolute worktree line numbers before Task 1 inserts the large §A block, then Task 14 applies the same numbers to parent and worktree | Every recorded range after the insertion point slices different text in the two trees, producing false CHANGED results independently of content | Resolve each start and end anchor separately in the parent and current file at comparison time instead of reusing pre-insertion line numbers +MAJOR | high | Task 0 step 3 | The command compares only C:65–290 with W:264–489, omitting inventoried passages g, h and j and most other changed sites, yet its expected divergence list includes g4, which lies at C:815/W:997 and cannot appear | Pre-existing drift outside the early §5 span is not recorded, so Task 14 cannot distinguish inherited drift as promised | Run bounded baseline diffs for every inventoried and changed site, including g, h and j, and make the expected list match those regions +MAJOR | high | Task 0 step 3 | The baseline diff currently emits 90 lines, but the command pipes it through `head -40` | Fifty lines of baseline differences are hidden, so the instruction to raise anything beyond the deliberate divergences is not executable | Remove `head -40` or capture and inspect the complete diff +MAJOR | high | Task 10 step 4 | Checking only that each zero-context hunk starts on a `note` line does not constrain the rest of that hunk; a changed adjacent `else`, `return`, or `fi` is grouped into the same hunk and inherits the permitted start | Hook control flow can change while the no-behaviour check passes, violating the plan's invariant-4 claim | Compare parent and worktree after replacing only the seven exact note-line contents with neutral markers, or inspect and validate every changed old and new line in each hunk +MAJOR | high | Task 10 step 5 | Five single-message pairs have exact expected counts of 1/1, but the task weakens all six pairs to `old/parent >= 1` and `new/worktree >= 1` solely because the grouped STOP pair has count 2 | Duplicate installed wording or a non-unique fragment can pass for the five pairs that should be exact, contrary to the fragment-table rule | State exact counts per pair: 1 for the five single replacements and 2 only for the grouped items 15/16 pair +MAJOR | high | Design §3 / target §I duty | The approved inputs assign the plan a sweep of both prompt copies and hook reminders for affected live sentences not already found, but no task performs or records that sweep; line 1224 merely says it is owed by whoever executes | All listed tasks can finish while an additional standing instruction falsified by §A remains active, the exact completeness failure the design assigns to this plan | Add an explicit reader-led sweep task against the real files that records what was examined and turns every newly found site into a spec finding, without adding a mechanical guard +MINOR | high | Tasks 1, 7 and 8 commit steps | These tasks explicitly add fragments or rows to this plan, but their `git add` commands stage only C and W and their Files lists omit the plan | The reviewed fragment evidence remains dirty and is silently swept into Task 13's unrelated commit, so the task commits are not independent or reviewable as described | List the plan as modified and stage it in each task that changes the fragment table +MAJOR | high | Task 15 steps 5 and 8 | The evidence is written into a WIP commit body, then `git reset --soft` removes every WIP commit and `git commit -m ""` creates a new body without carrying that evidence | The closing commit loses the validated evidence entry, provenance, curve and exception records that the next sentence requires | Build the complete closing message in a file and use it for the squashed commit, preserving the revalidated body explicitly +MAJOR | high | Task 15 step 8 | `git reset --soft ` is an unresolved shell placeholder even though Task 0 persists the exact revision in `.context/loop-rule-base` | The closing command is not runnable as written and a mistaken substitution can squash the wrong range | Reload `$BASE` and run `git reset --soft "$BASE"` +MAJOR | high | Task 2 step 1 | The runnable block uses `$BASE` without the required `BASE=$(cat .context/loop-rule-base)` initialization; with an empty value, `git show ":$f"` reads the index rather than the recorded parent | The six checks can report unchanged while comparing the post-Task-1 index to the worktree and never test the parent | Reload and validate `$BASE` inside this command block before the loop +MINOR | high | Task 0 step 2 | The plan says passage h has twenty-three kept conditions, but its own disposition keeps every h condition except h4 and h19, which is twenty-four | The stated count disagrees with the 26-condition inventory and can cause one kept condition to be omitted from the intended check | Change twenty-three to twenty-four and make the kept-condition extraction enumerate all twenty-four +MINOR | high | Self-Review §1 | It quotes §G's OLD as `These rules and records are one contract`, but that string occurs nowhere; C:879 and W:1063 read `These records are one contract` | The self-review reintroduces the exact absent anchor it says Task 7 repaired and gives reviewers conflicting counterfactual text | Use the verified P7 OLD `These records are one contract` +MINOR | high | Self-Review correction count | It says Task 7 names nine sites rather than eight, while Task 7 consistently names one §G site plus nine §H sites, ten total | The plan's stated count fails its own enumeration and obscures whether the §G repair is included | Change the self-review to ten sites rather than nine +NIT | high | Verification-fragment introduction | It says every fragment returned exactly one hit per copy, although P1 has no fragment and P5/P5w are explicitly per-copy rows | The verification claim is literally broader than the evidence table | Say every OLD fragment returned one hit in each copy claimed by its row, with P1 excluded +MINOR | high | Self-Review §2 | It claims every verification step has a runnable command and expected value, but Task 6 step 3 has no command, Task 8 step 4 is only a comment, Task 9 step 3 only assigns variables, and Tasks 12–14 include reader checks without commands | The self-review overstates the plan's executability and can hide incomplete verification work | Narrow the claim to the checks that are intentionally mechanical, then add commands where a mechanical result is required +NIT | high | Task 8 step 4 | The `Count what was installed` code block contains only the comment `# fourteen prompt-copy items from this task, each present once per copy` | The step produces no observed number despite requiring the executor to state one | Add a command that totals the fourteen verified NEW counts already produced in step 3 +END OF FINDINGS (32 total) diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md index 201a058..05d6ebb 100644 --- a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md @@ -31,7 +31,8 @@ worth checking before a pass rather than after. | Pass | Plan rev | Findings | Blockers | Majors | Valid | Notes | |---|---|---|---|---|---|---| | 1 | 5871d0a | **32** | **1** | **28** | yes | first pass. The Blocker is real: with a WIP commit per task, `baseSha = HEAD^` would have put only the version bump in Gate B's range. **The dominant class is verification fragments that count zero** — the reviewer tested them against the real files and found five of Task 7's ten locators, both of Task 3's, Task 5's and Task 7's strict-reading OLD all wrapping across lines or quoting text that does not exist. Repaired by one **verified fragment table** for the whole plan instead of guessed fragments per task | -| 2 | — | — | — | — | not run | next, against the pass-1 repair commit | +| 2 | 9f13a2c | 32→**32** | 1→**0** | 28→**25** | yes | one tell (findings flat, and they cluster on the plan's own checks — the instrument). **Pass 1's repair reproduced the same defect at the next condition down:** three OLD fragments are *preserved inside their own replacements*, so their old-count can never reach zero. My checker tested single-line and unique but not that third condition, though this plan states all three. Checker fixed to test against the target's **fenced blocks**; three rows replaced; the fourteen §F OLD fragments derived and committed rather than deferred | +| 3 | — | — | — | — | not run | next, against the pass-2 repair commit | ## Pass-1 report @@ -64,3 +65,44 @@ into an edit a Major already required: the site count contradicting its own enum one task), `c20` recorded as carried while the change narrows its scope — the dropped-condition failure `AGENTS.md` names — and the two plan-mutating tasks having no stable region to replace. None cost a pass of its own. + + +## Pass-2 report + +**Trend:** findings 32, **32**; Blockers 1, **0**; Majors 28, **25**. **Cluster:** the plan's own +verification apparatus — fragments, pair commands, range checks. That is the **instrument**, so this +is one tell; the finding count is flat rather than rising, which is not. **Require↔withdraw:** none. + +**The finding worth the whole pass.** Pass 1's repair introduced one verified fragment table. Pass 2 +found three of its rows **preserved inside their own replacements** — `e7` in both copies and the +strict-reading list — so each would have returned `old/worktree=1` after a correct edit and no task +could have reached its required result. **This plan states all three failure modes in its own +fragment-table section**, and the checker I wrote tested only two of them. + +**The repair is the checker, not the three rows.** It now tests condition 3 against the target +text's **fenced blocks** rather than the whole file — the narrower test matters, because a fragment +quoted in an item's rationale is not preserved by its replacement, and the broad test rejected a +usable row (P17) that the reviewer correctly left alone. Re-running it over every existing row found +exactly the three the reviewer named and nothing else. + +**The second-largest class was commands that cannot run:** `pair()` defined once in a section while +the plan says each task runs in its own shell; three steps that state an expected four-value result +and contain no command that produces one; `$BASE` used without being reloaded, where an empty value +makes `git show` read the index rather than the parent and every count describe the wrong tree. +Task 0 now writes `.context/loop-rule-pair.sh`, every counting block sources it, and the source line +is followed by a guard that fails loudly on an empty `$BASE`. + +**Three findings were the plan's checks being un-runnable in a way that would have passed anyway** — +the untouched-range diffs used absolute line numbers recorded *before* Task 1's insertion, and two of +the five ranges contained conditions this change deliberately replaces, so a correct implementation +would have failed its own check. Ranges are anchors now, and the two contaminated ones are split +around the sites they must exclude. + +**One duty had no task at all.** Design §7 and target §I assign the plan a completeness sweep for +affected sites the spec has not found; the plan named it in a residual and gave it to "whoever +executes", which discharges nothing. It is **Task 12b** now — a reader-led sweep with a written +record, and explicitly not a mechanical guard. + +**And the closing commit would have destroyed its own evidence:** step 5 wrote the evidence entry +into a WIP body, step 8 squashed with `git reset --soft`, which keeps the tree and discards every WIP +message. The entry goes to a file now. diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index 9e0d101..2c6f77a 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -29,9 +29,11 @@ ## Verification fragments — verified against the real files -**Every fragment below was tested with `python3 /frag.py ""` against `CLAUDE.md` and the template at `5871d0a`, and each returned exactly one hit per copy on the line shown.** No task may invent a fragment; a task needing one not listed here adds it to this table and re-runs that check first. This table exists because the first draft scattered guessed fragments across ten tasks and pass 1 found that most of them wrapped across lines and counted zero in a correct tree — the exact defect design §7 records from four consecutive spec revisions. +**Every OLD fragment below was tested against all three conditions** — single-line, unique in each copy the row claims, and **absent from the target text's replacement blocks** — at `9f13a2c`. Each returned one hit in each copy its row names; **row P1 has no fragment**, and rows P5 and P5w are per-copy, so the claim is about the copies each row claims and not about both copies for every row. No task may invent a fragment; a task needing one not listed here adds it and runs the same three checks first. -**Three ways a fragment fails, all of which this table's check catches:** it wraps across a line break, so `grep -F` counts zero in a correct file; it is **preserved inside its own replacement**, so its old-wording-gone count can never reach zero; or it is not unique, so a count of 1 proves nothing about which occurrence changed. +**Three ways a fragment fails:** it wraps across a line break, so `grep -F` counts zero in a correct file; it is **preserved inside its own replacement**, so its old-wording-gone count can never reach zero; or it is not unique, so a count of 1 proves nothing about which occurrence changed. + +**This table exists because both of the first two drafts got this wrong, in a different one of those three ways each time.** Pass 1 found nine fragments that wrapped or quoted text that does not exist. Pass 2 found three that were preserved inside their own replacements — the condition the checker used for pass 1 did not test, though this plan had stated it. The third condition is now checked **against the target text's fenced blocks**, not against the whole file: a fragment quoted in an item's rationale is not preserved by its replacement, and testing the whole file rejects usable fragments. | # | Edit | OLD fragment (single-line, not preserved in its replacement) | C | W | |---|---|---|---|---| @@ -39,8 +41,8 @@ | P2 | `b7`, the fix-set definition | `scope the approved story or plan assigns to this cycle, plus repair obligations you already` | 201 | 408 | | P3 | `b12`, immediate resumption | `the moment the user says whether the set now includes it` | 208 | 415 | | P4 | `c14`, below-floor Minor | `a Blocker/Major-free pass below the floor` | 239 | 442 | -| P5 | `e7`, the threshold — **C** | `the tells and hand the decision to the user, and the` | 267 | — | -| P5w | `e7`, the threshold — **W** | `tells and hand the decision to the user, and the` | — | 471 | +| P5 | `e7`, the threshold — **C** | `not discretionary** — you report` | 266 | — | +| P5w | `e7`, the threshold — **W** | `not discretionary** — report the` | — | 470 | | P6 | `g1`/`g2`, the handed-over question | `is not settled here, and this change does not settle it` | 811 | 997 | | P7 | §G, the one-contract paragraph | `These records are one contract` | 879 | 1063 | | P8 | `c18`, no-clean-credit | `the resolve rule is not waived, no pass is credited as` | 243 | 446 | @@ -51,24 +53,59 @@ | P13 | gate-prompt template clean sentence | `clean pass is the single body line` | 329 | 523 | | P14 | Gate-A cadence | `Each pass: validate, revise, re-run` | 573 | 764 | | P15 | lens unchanged-list | `The Blocker/Major filter, the file-first findings protocol` | 653 | 839 | -| P16 | strict-reading list | `nonce duties at their strictest` | 157 | 364 | +| P16 | strict-reading list | `and the nonce duties at their strictest — the cycle` | 157 | 364 | | P17 | §F item 14, Named residual | `Hook text is out of scope here` | 139 | 346 | | P18 | §F item 18, work-loop line | `execute → tests green → Gate B → commit` | 63 | 262 | +**The fourteen §F prompt-copy items (Task 8), derived from each item's cited lines and checked the same three ways.** Pass 2 found these deferred to the executor as `` placeholders, which put fourteen meaning-changing checks outside Gate A's reach; they are concrete now. The NEW halves stay deferred, for the reason the paragraph below gives. + +| Row | §F item | OLD fragment | C | +|---|---|---|---| +| F1 | 1, the `WIP:` naming warning | `nor resets your pass counters. A pre-review snapshot named anything else reads as a` | 825 | +| F2 | 2, the Gate-B coverage instruction | `` `NO FINDINGS` if clean" in `additionalContext`, with the same one-line format. `` | 598 | +| F3 | 3, the curve's Majors rationale | `**Majors are recorded as well as Findings and Blockers**, because the severity rule moves the` | 938 | +| F4 | 4, the human-exception scope sentence | `own terminal actions and this paragraph changes none of them: on a STOP you still stop, and` | 1014 | +| F5 | 5, the `Finishing the cycle` lead-in | `**Finishing the cycle:** after the final clean pass, close it with` | 827 | +| F6 | 6, the Gate-A broad-prompt instruction | `reviews the TEXT you pass, not the git tree). Use ONE broad prompt, re-run it` | 553 | +| F7 | 7, the human-exception destination | `**Which commit:** an ungated change records it in that commit; a Gate-A cycle in the spec or` | 988 | +| F8 | 8, the profile-change pass claim | ``snapshot by amend — a non-`WIP` commit reads to the hook as the cycle closing and would`` | 752 | +| F9 | 7a, the mid-run recovery sentence | `taken from the provenance line and the curve, which must agree. A Gate-A cycle mid-run has no` | 416 | +| F10 | 8a, the HARD FLOOR parenthetical | `(Blocker/Major only), derived from the cited story's profile.**` | 73 | +| F11 | 8b, the Gate-A filter clause | `intent + artifact text + which invariants it touches. Ask for **every** finding` | 561 | +| F12 | 9a, the revalidation trigger | `and the named evidence but not the mode value**, and is **revalidated before every Gate-B` | 726 | +| F13 | 9b, the severity-deciding fallback | `decision that act takes differently if the text is wrong. Both are required. If you` | 788 | +| F14 | 9, the revalidation remedy | `profile sits still. If revalidation changes the entry, the clean pass no longer covers what` | 728 | + +**W line numbers are deliberately not carried for F1–F14.** Each fragment was verified unique in W as well as C, but the template's numbers drift with every earlier task and a stale number here would read as source drift. Locate each in W by the fragment. + **`e7` is the one edit needing a per-copy fragment**, because W drops the pronoun: C reads `you report the tells`, W reads `report the tells`. That is the recorded `e8` divergence, and it **does not survive this change** — see Task 5. **The NEW fragment for every pair is taken from the installed line and checked the same way**, since the new wording does not exist until the task installs it. Each task's step says which sentence of its target block to take it from, and the step fails if the fragment it chooses is not single-line and unique in the installed file. **This is the one place the plan cannot pre-verify**, and it is disclosed rather than papered over. -**The four-value command, defined once and cited by number afterwards:** +**The four-value command. Task 0 writes it to a file, and every task that uses it sources that file** — task command blocks run in separate shells, so a function defined here once would be `pair: command not found` everywhere it is called. + +Task 0 writes `.context/loop-rule-pair.sh`: ```bash +cat > .context/loop-rule-pair.sh <<'SH' +BASE=$(cat .context/loop-rule-base) pair() { # pair printf '%s old/worktree=%s old/parent=%s new/worktree=%s new/parent=%s\n' "$3" \ "$(grep -cF "$1" "$3")" "$(git show "$BASE:$3" | grep -cF "$1")" \ "$(grep -cF "$2" "$3")" "$(git show "$BASE:$3" | grep -cF "$2")" } +SH +``` + +and **every later command block that counts anything begins with**: + +```bash +. .context/loop-rule-pair.sh +test -n "$BASE" || { echo "BASE empty — Task 0 did not run"; exit 1; } ``` +**The `test` is not decoration.** With `$BASE` empty, `git show ":$f"` reads the *index* rather than the recorded parent and every parent count silently describes the wrong tree. + **A pair passes only on `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`.** All four values matter: a copy carrying the new wording **and** the old one satisfies a one-sided presence check and is exactly the two-instructions-that-disagree failure the pair exists to catch. --- @@ -95,7 +132,8 @@ Story acceptance criterion 5 is satisfied here. Ids are `docs/superpowers/specs/ | Condition | Disposition | |---|---| -| a1–a12, a14 | **kept**, untouched. The floor arithmetic, the hook-ratio rules and the two comparison points are outside this change. | +| a1, a3–a12, a14 | **kept**, untouched. The floor arithmetic, the hook-ratio rules and the two comparison points are outside this change. | +| a2 | **replaced.** §F item 8a rewrites the HARD FLOOR parenthetical `(Blocker/Major only)` — Task 8, row F10. The filter itself survives in the ordering, which states what a pass counting toward the floor must be; what goes is the parenthetical's claim that Blocker/Major is the *whole* of it. **Marked replaced rather than kept**, because a condition whose text the change in fact rewrites, recorded as preserved, is the dropped-condition failure `AGENTS.md` names. | | a13 | **replaced** — scoped to its own paragraph. §H's `a13` block. | | a15 | **carried** inside §H's `a16` block, which reproduces it so one contiguous string installs. | | a16 | **replaced** — points at Mechanics · Severity instead of carrying an unscoped copy. §H's `a16` block. | @@ -198,16 +236,21 @@ empty revision — which fails loudly in `git show` but quietly in a `grep -c` p task begins with: ```bash -BASE=$(cat .context/loop-rule-base) +. .context/loop-rule-pair.sh +test -n "$BASE" || { echo "BASE empty — Task 0 did not run"; exit 1; } ``` +**Task 0 also writes `.context/loop-rule-pair.sh`**, whose contents are given in the fragment-table +section. It sets `$BASE` and defines `pair()`, so both arrive together and neither can be used +without the other. + **This commit is also `baseSha` for Gate B.** It is the parent of the first WIP snapshot, and it is the only value that puts the whole implementation inside the reviewed range. `.context/` is ignored by the hook's fingerprint, so the file itself moves nothing. - [ ] **Step 2: Re-read the five untouched ranges and record their current line numbers** -Passages (d), (f) and (j) are not edited by this change, and passage (a)'s arithmetic and passage (h)'s twenty-three kept conditions are mostly untouched. Record where they are now, because the inventory's numbers cite `7c0d475`: +Passages (d), (f) and (j) are not edited by this change, and passage (a)'s arithmetic and passage (h)'s **twenty-four** kept conditions are untouched — twenty-six inventoried, less `h4` and `h19`, which Task 8 replaces. Record where they are now, because the inventory's numbers cite `7c0d475`: **Five ranges, each with a start and an end anchor** — the three whole passages, plus the floor arithmetic and the kept human-exception conditions, which later verification consumes and which the @@ -228,17 +271,57 @@ for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do done ``` -Expected: one hit per pattern per file. Write all the line numbers into -`.context/loop-rule-untouched` — **Task 2 and the final check in Task 14 both read them**, so a -scratch note that dies with the session is not enough. +Expected: one hit per pattern per file. Write the **anchor strings**, not line numbers, into +`.context/loop-rule-untouched` — **Task 2 and Task 14's final check both read them.** + +**Anchors, never absolute line numbers.** Task 1 inserts a large block, so every recorded number +after the insertion point addresses different text in the parent than in the worktree, and a +comparison using them reports CHANGED on identical content. Each check resolves its own start and +end anchor in each tree at comparison time. + +**Two of the five ranges must exclude the conditions this change replaces**, or a correct +implementation fails its own check: + +- **The floor arithmetic range must exclude `a2` and `a13`.** `a2` is the HARD FLOOR parenthetical + Task 8 replaces (row F10) and `a13` is the no-restating prohibition Task 7 replaces (row P9). + Record the arithmetic as **two spans** — from `Both gates are a LOOP with a HARD FLOOR` to just + before `(Blocker/Major only)`, and from just after `a13`'s sentence to `Nothing here writes the + floor knob`. +- **The human-exception range must exclude `h4` and `h19`**, which Task 8 replaces (rows F7 and F4). + Record the twenty-four kept conditions as the spans **between** those two sites. - [ ] **Step 3: Confirm the parity baseline of the inventoried ranges** +**Diff every inventoried and changed site, not one early window, and read the whole output.** +The first draft compared C 65–290 with W 264–489 and piped it through `head -40`. That window holds +none of passages (g), (h) or (j) — so `g4`, which sits at C 815 / W 997, could not appear in a diff +whose expected list named it — and the truncation hid about fifty of the roughly ninety lines the +comparison actually emits. + ```bash -diff <(sed -n '65,290p' CLAUDE.md) <(sed -n '264,489p' plugins/dev-workflow/commands/workflow-init.md) | head -40 +# one bounded region per site, resolved by anchor in each file, output in full +for site in 'Both gates are a LOOP:Nothing here writes the floor knob' \ + 'What a loop absorbs:Recognizing "clearly stuck"' \ + 'Recognizing "clearly stuck":Every pass report states' \ + 'From pass 4 onward:Those three lines expose' \ + 'Those three lines expose:The two rules above' \ + 'The two rules above:Findings go to a FILE' \ + '\*\*Severity:\*\*:\*\*Tool routing:' \ + 'Recording a human exception:because writing it down makes it sound' \ + 'On squash-merge:On squash-merge' \ + 'When these rules bind:Downstream has no shipping commit'; do + s=${site%%:*}; e=${site##*:} + echo "== $s" + diff <(sed -n "/$s/,/$e/p" CLAUDE.md) \ + <(sed -n "/$s/,/$e/p" plugins/dev-workflow/commands/workflow-init.md) +done | tee .context/loop-rule-baseline-diff.txt ``` -Expected: the deliberate divergences the inventory records (b3's cross-reference target, b's intensifier and field-mint parenthetical, e8's pronoun, e11, f5–f7's framing, g4). **Anything else is pre-existing drift — record it and raise it before editing**, because Task 14's parity diff cannot tell drift you introduced from drift you inherited. +Expected: the deliberate divergences the inventory records — `b3`'s cross-reference target, passage +(b)'s intensifier and field-mint parenthetical, `e8`'s pronoun, `e11`, `f5`–`f7`'s framing, and +`g4`. **Read every line of the output; do not truncate it.** Anything else is pre-existing drift — +record it in the file and raise it before editing, because Task 14's parity diff cannot tell drift +you introduced from drift you inherited. - [ ] **Step 4: Commit nothing** @@ -317,10 +400,17 @@ Expected: no output. - [ ] **Step 6: Commit** ```bash -git add CLAUDE.md plugins/dev-workflow/commands/workflow-init.md +git add CLAUDE.md plugins/dev-workflow/commands/workflow-init.md \ + docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md git commit -m "WIP: install the closure ordering into both §5 copies" ``` +**The plan is staged here because this task adds the §A2 and §A3 rows to the fragment table.** Every +task that adds a row stages the plan with its own edit; otherwise the reviewed fragment evidence +stays dirty and is swept into a later, unrelated commit, and the task commits are not the +independently reviewable units this plan claims they are. The same applies to Tasks 3, 4, 6, 7 and +8, each of which derives rows. + Named `WIP:` because Task 15 runs Gate B over the whole change and amends once. A non-`WIP` commit here would reset the hook's Gate-B counters mid-cycle. --- @@ -337,6 +427,8 @@ This task exists because `f1` and the (d)/(j) dispositions are falsifiable only - [ ] **Step 1: Diff each untouched passage against the parent** ```bash +. .context/loop-rule-pair.sh +test -n "$BASE" || { echo "BASE empty — Task 0 did not run"; exit 1; } for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do for anchor in 'From pass 4 onward every pass report carries three lines' \ 'The two rules above do not compete' \ @@ -415,6 +507,13 @@ done Expected for all four: `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`. +**Two pairs do not cover §B.** Target §B separately changes `b3`, `b8`, `b11`, `b13`, `b16`, +`b17`–`b18` and **adds** the closing-time set-change rule and decision 6's decline semantics. Each +independent replacement owes its own pair, and each add-only rule owes a presence check — a manual +condition walk is a reader's judgement, not the discriminating observation design §7 assigns here. +**Derive an OLD row for each changed condition and a NEW fragment for each added rule, add them to +the fragment table, and run the checks before Step 4.** + - [ ] **Step 4: Walk the carried conditions** Read the installed passage and confirm `b1`, `b2`, `b4`, `b5`, `b6`, `b9`, `b10`, `b14`, `b15` are each present, and that `b3`, `b7`, `b8`, `b11`, `b12`, `b13`, `b16`, `b17`, `b18` read as §B states rather than as the parent did. Nine plus nine; the disposition table above is the checklist. @@ -426,7 +525,10 @@ diff <(sed -n '/^\*\*What a loop absorbs/,/^\*\*Recognizing "clearly stuck"/p' C <(sed -n '/^\*\*What a loop absorbs/,/^\*\*Recognizing "clearly stuck"/p' plugins/dev-workflow/commands/workflow-init.md) ``` -Expected: only the three recorded divergences. +Expected: **only C's field-mint parenthetical.** Task 3 installs §B's common span, which aligns +`b3`, the intensifier and the closing rationale — so expecting "three recorded divergences" would +make a correct implementation look like a failure and would let stale W wording that §B removes pass +as inherited drift. **Any difference other than the parenthetical is a failure of this task.** - [ ] **Step 6: Commit** @@ -483,6 +585,11 @@ done Expected: `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`. +**One pair per edit, not one per task.** The pair above takes its OLD from `c14` and its NEW from +`c8` — two different changes — so `c8` can land while `c14` survives, or the reverse, and it still +reports a pass. `c4` has no observation at all. **Add a pair for `c4`'s replaced wording and one for +`c14`'s removal against §C's conditional routing, with rows in the fragment table**, before Step 5. + - [ ] **Step 5: Walk the conditions** — `c1`–`c3` present unchanged, `c5`–`c7` word for word, `c8` carrying the new clause, `c9` split, `c10`–`c14` gone from here. - [ ] **Step 6: Commit** @@ -550,7 +657,9 @@ done Expected: both pairs `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`; both pointer checks `worktree=1 parent=0`. -- [ ] **Step 4: Confirm `e1`–`e6` and `e8`–`e11` are untouched**, and that e11 is still C-only. +- [ ] **Step 4: Confirm `e1`–`e6` and `e9`–`e11` are untouched**, that `e11` is still C-only, and +that **`e8` is aligned rather than untouched** — this task gives W the pronoun, so listing `e8` +among the untouched conditions would contradict the task's own instruction. - [ ] **Step 5: Commit** @@ -571,7 +680,11 @@ git commit -m "WIP: read the two-tell threshold after the clean-completion branc **This task removes the one deliberate story-path divergence.** `g4` — C's sentence naming `docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md` as owning the question — goes, because this change is that work and the question is answered. `g2` and `g3`, the interim report-and-stop duty and its justification, go from **both** copies in the same edit. **Removing g4 from C without removing g2/g3 from W desynchronises the copies in the opposite direction**, which is the failure the inventory flags by name. -- [ ] **Step 1: Record the old wording in both copies** +- [ ] **Step 1: Record the old wording in both copies, and derive the resolve-duty OLD** + +§E replaces **two** blocks. The handed-over-question OLD is row P6; the resolve-duty bullet has no +row yet. Derive one from the live `**Severity:**` bullet, check it the three ways, and add it to the +fragment table before Step 3. ```bash for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do @@ -585,9 +698,24 @@ Expected: C `g1/g2: 1 g4: 1`; W `g1/g2: 1 g4: 0`. - [ ] **Step 2: Install §E's resolve-duty bullet and its answer paragraph in both copies** -- [ ] **Step 3: Run the discriminating pair**, both copies, both trees, with `NEW='The demotion changes what a cycle must resolve, never what it observes'` and `OLD='is not settled here, and this change does not settle it'`. +- [ ] **Step 3: Run two discriminating pairs, both copies, both trees** + +**§E has two fenced replacements and each owes its own pair.** With only the handed-over-question +pair, the old unscoped resolve duty can survive, or the new repair-versus-dismissal distinction be +omitted, while everything Task 6 checks passes. + +```bash +. .context/loop-rule-pair.sh +test -n "$BASE" || { echo "BASE empty — Task 0 did not run"; exit 1; } +for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do + pair 'is not settled here, and this change does not settle it' \ + 'The demotion changes what a cycle must resolve, never what it observes' "$f" + pair '' \ + 'for every finding in the assigned fix set as the absorb paragraph computes it' "$f" +done +``` -Expected: `new/worktree=1 new/parent=0 old/worktree=0 old/parent=1` in both. +Expected for all four: `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`. - [ ] **Step 4: Confirm g4 is gone from C** @@ -767,10 +895,19 @@ Expected for all twenty-eight: `old/worktree=0 old/parent=1 new/worktree=1 new/p - [ ] **Step 4: Count what was installed** ```bash -# fourteen prompt-copy items from this task, each present once per copy +# Each of step 3's fourteen pairs printed new/worktree; total them per copy. +. .context/loop-rule-pair.sh +for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do + n=0 + for new in '' '' '' '' '' '' '' \ + '' '' '' '' '' '' ''; do + n=$(( n + $(grep -cF "$new" "$f") )) + done + echo "$f installed=$n" +done ``` -State the number you observed. **Do not carry a count from §F into a check** — §F states the count of falsified sentences and this task installs a subset of them; a count copied between the two is the stale-bookkeeping defect the design records at five passes running. +Expected: `installed=14` per copy. State the number you observed. **Do not carry a count from §F into a check** — §F states the count of falsified sentences and this task installs a subset of them; a count copied between the two is the stale-bookkeeping defect the design records at five passes running. - [ ] **Step 5: Parity** for all fourteen sites. @@ -804,16 +941,21 @@ Expected: one hit each per file. **Item 14's sentence wraps after "here"** — - [ ] **Step 2: Install both replacements** -- [ ] **Step 3: Discriminating pairs** +- [ ] **Step 3: Run both discriminating pairs** + +Rows **P17** and **P18**. Assigning the variables is not running the check — the first draft stopped +at the assignment. ```bash -OLD1='Hook text is out of scope here' -NEW1='not a blanket exemption for hook text' -OLD2='execute → tests green → Gate B → commit' -NEW2='Gate A (spec) → Gate-A closing act' +. .context/loop-rule-pair.sh +test -n "$BASE" || { echo "BASE empty — Task 0 did not run"; exit 1; } +for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do + pair 'Hook text is out of scope here' 'not a blanket exemption for hook text' "$f" + pair 'execute → tests green → Gate B → commit' 'Gate A (spec) → Gate-A closing act' "$f" +done ``` -Expected for each, in both copies: `new/worktree=1 new/parent=0 old/worktree=0 old/parent=1`. +Expected for all four: `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`. - [ ] **Step 4: Confirm the residual the sentence was written for still stands** @@ -885,10 +1027,14 @@ git diff -U0 "$BASE" -- plugins/dev-workflow/hooks/codex-gate.sh \ grep -n 'note "' plugins/dev-workflow/hooks/codex-gate.sh ``` -**Every changed hunk must start inside one of the seven gate-reminder `note` calls this task -edits.** A hunk anywhere else is a behaviour change and is out of scope — no control flow, no -counter, no fingerprint computation, no routing (invariant 4, design §8). Read the two lists side by -side; do not automate the comparison with a pattern, which is what failed here the first time. +**A hunk's start is not enough.** A changed `else`, `fi` or `return` adjacent to an edited `note` +line lands in the same zero-context hunk and inherits its permitted start. **Read every removed and +added line of every hunk** and confirm each one is inside one of the seven gate-reminder `note` +strings this task edits. A changed line anywhere else is a behaviour change and is out of scope — +no control flow, no counter, no fingerprint computation, no routing (invariant 4, design §8). + +Read the diff yourself rather than filtering it. A pattern over line *shape* is what failed here the +first time — it admitted exactly the shell keywords it existed to catch. - [ ] **Step 5: Discriminating pairs, worktree and parent** @@ -917,7 +1063,19 @@ listed pair passes: 'or proceed only if $policy|skip rule decides only whether a cycle runs at all' # item 17 ``` -Expected: `old/worktree=0`, `old/parent` ≥ 1, `new/worktree` ≥ 1, `new/parent=0` for each of the six. The `STOP` pair covers two messages, so its counts are 2 rather than 1 — **state the number you observed rather than asserting it**. +**Exact counts per pair, not a floor.** Five of these replace one message each and must read +exactly `old/parent=1 new/worktree=1`; only the grouped `STOP` pair, which covers items 15 and 16, +reads `2`. Weakening all six to `≥ 1` because one of them is 2 lets a duplicated installation or a +non-unique fragment pass for the five that should be exact: + +| Pair | old/worktree | old/parent | new/worktree | new/parent | +|---|---|---|---|---| +| item 10, the honesty claim | 0 | 1 | 1 | 0 | +| item 11, the Gate-A clean definition | 0 | 1 | 1 | 0 | +| item 12, the Gate-B clean definition | 0 | 1 | 1 | 0 | +| item 13, the WIP reminder | 0 | 1 | 1 | 0 | +| item 17, the below-floor instruction | 0 | 1 | 1 | 0 | +| items 15+16, the two `STOP` openings | 0 | **2** | **2** | 0 | - [ ] **Step 6: ShellCheck** @@ -1025,6 +1183,60 @@ The two copies are byte-identical over this material, so a divergence here is a --- +## Task 12b: The completeness sweep design §3 and target §I assign to the plan + +**Files:** possibly `CLAUDE.md`, `plugins/dev-workflow/commands/workflow-init.md`, `plugins/dev-workflow/hooks/codex-gate.sh`; and this plan, which records what was examined. + +**This task exists because nothing else performs the sweep.** Design §7 leaves "the sweep for any +affected site this text has not found" to the plan, and target §I records that **nothing in the spec +establishes the edit set is complete** — it is the sites known when §F stopped growing. Without a +task, every other task can finish while a standing instruction the ordering falsifies is still live, +which is exactly the completeness failure the design hands to the plan. **The plan's own residual +paragraph says the sweep is "owed by whoever executes", which names a duty and discharges nothing.** + +**It is a reader-led sweep, not a mechanical guard.** No pattern decides whether a sentence the +ordering falsifies is still live; §G says so about its own subject, and inventing a checker here +would be the false precision this repo's invariants warn about. What is mechanical is the **record**: +what was read, and what was found. + +- [ ] **Step 1: Read the installed ordering once more, and list what it now decides** + +Closure, eligibility, cleanliness, the hold, composition, the two scope triggers, the suspensions and +their answers, the two gates' closing acts. **This list is the sweep's question set** — for each, +"does any other sentence in these three files still answer this?" + +- [ ] **Step 2: Read every section of both prompt copies that gives an instruction about a pass, a finding, a gate or a commit** + +Not a grep. §5 entire, §4's work-loop line, the Gate-A and Gate-B sections, the profiles section, and +Mechanics. **Record each section as read in `.context/loop-rule-sweep.md`**, with a line saying what +you were looking for and what you found. + +- [ ] **Step 3: Read both channels of all eight hook gate reminders** + +Seven are replaced by §F. **The eighth, the docs-only notice, is read too** — it is excluded because +it states no closure permission, and that exclusion is a claim this sweep is the place to confirm. + +- [ ] **Step 4: Classify anything found** + +A site the ordering falsifies that §F does not replace is **a finding against the approved spec, not +a gap in this plan**. Surface it: the spec's §F is the only enumeration of these sentences, and +adding one here would be the second copy that cycle spent sixty-five passes removing. **Where the +find is real, the spec's Gate-A cycle reopens for it.** + +- [ ] **Step 5: Record the result either way** + +`.context/loop-rule-sweep.md` states what was read and what was found, **including "nothing"**. A +sweep whose negative result is unrecorded cannot be told from a sweep that never ran. + +- [ ] **Step 6: Commit** + +```bash +git add .context/loop-rule-sweep.md docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +git commit -m "WIP: record the completeness sweep" +``` + +--- + ## Task 13: The next-state table and the per-condition closure checks **Files:** @@ -1098,7 +1310,14 @@ the document, replaced idempotently on re-run by the same rule Task 13 uses. - [ ] **Step 2: Classify every difference the diff reports** -Three buckets: **deliberate and stays** (the field-mint parenthetical, e11, f5–f7's evidence framing, e8's pronoun — each recorded in the inventory); **not deliberate, align it**; **introduced by this change, fix it**. Write the list into this plan. +Three buckets: **deliberate and stays** (the field-mint parenthetical, `e11`, `f5`–`f7`'s evidence +framing — each recorded in the inventory); **not deliberate, align it**; **introduced by this +change, fix it**. Write the list into this plan. + +**`e8` is not in the first bucket.** Task 5 installs §D's complete sentence in both copies, which +gives W the pronoun; classifying that divergence as deliberate here would let the two shipped copies +disagree on a sentence the approved target states once. **`b3` is not in it either** — Task 3 +aligns it, and this task verifies the alignment rather than performing it. - [ ] **Step 3: Apply W's `b3` alignment** @@ -1174,7 +1393,11 @@ claude plugin validate . --strict Expected: exit 0. **`check-invariants.sh` includes the prompt-conformance checks** — a `Target model:` line naming one recognized model, a prose checklist-count claim matching the checklist, and the finding-severity vocabulary as a closed set in both prompt copies. Those three are a floor, not coverage; invariant 11's other eleven items are judged by a reader. -- [ ] **Step 5: Write the evidence entry into the WIP commit body** +- [ ] **Step 5: Write the evidence entry into `.context/loop-rule-closing-msg`** + +**Not into a WIP commit body.** Step 8 squashes with `git reset --soft`, which keeps the tree and +discards every WIP message; evidence written only there would be destroyed by the close. Restate it +in the WIP body too if a mid-cycle reader would want it, but the file is the copy that survives. It names: the battery run; **every pair this plan built, with its counts in each copy and each tree, and every presence check beside them**; the §6 parity diff and the `b11`/`b13` equivalence result; and the next-state table's location plus its row count. **State the §A presence checks as presence, not as pairs** — its counterfactual is absent and is claimed as absent. @@ -1202,26 +1425,39 @@ Floor derives from the story profile: risk `high` → 2, security `none` → 0, - [ ] **Step 8: Close the cycle** +**Build the closing message in a file first.** `git reset --soft` discards every WIP commit *body*, +so an evidence entry written only into a WIP message is destroyed at exactly the moment the cycle +closes — which is what step 5 would otherwise have done. + ```bash -git reset --soft -git commit -m "" +. .context/loop-rule-pair.sh +test -n "$BASE" || { echo "BASE empty — Task 0 did not run"; exit 1; } +# .context/loop-rule-closing-msg holds the revalidated evidence entry, the provenance +# line, the per-pass curve and any human-exception record, written before this step. +git reset --soft "$BASE" +git commit -F .context/loop-rule-closing-msg ``` -The closing body carries: the validated evidence entry; the provenance line; the per-pass curve; and any human-exception record. **Amend rather than a follow-up commit** — a `WIP:` commit left in history defeats the convention, and a follow-up has nothing to commit when the review produced no fixes. +**`$BASE` is the recorded revision, not a placeholder to substitute by hand** — Task 0 persisted it +for this, and a mistaken substitution squashes the wrong range. + +The closing body carries: the validated evidence entry; the provenance line; the per-pass curve; and any human-exception record. **One commit rather than a follow-up** — a `WIP:` commit left in history defeats the convention, and a follow-up has nothing to commit when the review produced no fixes. --- ## Self-Review -**1. Spec coverage.** §A → Task 1. §B → Task 3. §C → Task 4. §D → Task 5. §E → Task 6. §F items 1–9 → Task 8; items 14, 18 → Task 9; items 10–13, 15–17 → Task 10 with its test sweep in Task 11. §G → Task 7, **added by this review**: the first draft gave the one-contract paragraph no task, though design §4 lists it as its own site and target §G carries its replacement. It is a prompt-copy replacement in both copies with the same shape as §H's blocks, owes the same discriminating pair with OLD `These rules and records are one contract`, and sits at C 879 / W 1063 as of this writing. §H → Task 7. §I ships nowhere and needs no task. Design §6 → Task 14. Design §7 → Tasks 13 and 15. Design §8 → Task 15's battery and the Global Constraints. Story AC 5 → the disposition tables. Story AC 4 → Task 13. +**1. Spec coverage.** §A → Task 1. §B → Task 3. §C → Task 4. §D → Task 5. §E → Task 6. §F items 1–9 → Task 8; items 14, 18 → Task 9; items 10–13, 15–17 → Task 10 with its test sweep in Task 11. §G → Task 7, **added by this review**: the first draft gave the one-contract paragraph no task, though design §4 lists it as its own site and target §G carries its replacement. It is a prompt-copy replacement in both copies with the same shape as §H's blocks, owes the same discriminating pair with row **P7**'s OLD `These records are one contract` — the live wording; `These rules and records are one contract` occurs nowhere — at C 879 / W 1063 as of this writing. §H → Task 7. §I ships nowhere and needs no task. Design §6 → Task 14. Design §7 → Tasks 13 and 15. Design §8 → Task 15's battery and the Global Constraints. Story AC 5 → the disposition tables. Story AC 4 → Task 13. + +**2. Placeholder scan.** The replacement text is cited rather than copied, deliberately and for the reason the Architecture note gives. Task 13's row list is explicitly a floor rather than a closed set, and says so. Task 11 deliberately carries no count, and says why. -**2. Placeholder scan.** The replacement text is cited rather than copied, deliberately and for the reason the Architecture note gives. Every verification step carries a runnable command and an expected value. Task 13's row list is explicitly a floor rather than a closed set, and says so. Task 11 deliberately carries no count, and says why. +**Which steps are mechanical and which are reader checks, stated rather than claimed uniformly.** Every OLD half of every pair is a concrete fragment with a runnable command and an exact expected result. **The NEW halves are `<...>` until their task installs the text**, which the fragment table discloses and each step requires to be verified before counting. **Tasks 12, 13 and 14 step 2 are reader checks by design** — a predicate comparison, a next-state walk and a divergence classification are judgements, and giving them commands would be the false-precision this repo's invariants warn about. An earlier revision of this section claimed every verification step had a runnable command, which was not true of them. **3. Type consistency.** `$BASE` is set in Task 0 and used in Tasks 1–10. The four-value pair shape (`new/worktree`, `new/parent`, `old/worktree`, `old/parent`) is defined in Task 3 and referred to by name afterwards. Condition ids match the inventory throughout: a1–a22, b1–b18, c1–c20, d1–d7, e1–e11, f1–f7, g1–g4, h1–h26, i1–i16, j1–j4 — 135 total, every one dispositioned above. -**One correction applied from this review:** §G was missing a task; it is now installed by Task 7, which names nine sites rather than eight. +**One correction applied from this review:** §G was missing a task; it is now installed by Task 7, which names **ten** sites — one §G block and nine §H blocks. -**One residual this plan does not close, stated rather than left to be found.** Nothing here establishes that the edit set is complete — it is the sites §F knows, and §F's own §I records that it cannot establish completeness either. Task 8's step 1 re-derives every citation against the real file, and Task 14's diff catches a copy that fell out of step; neither is a completeness check. The sweep for an affected site this text has not found is design §7's, owed by whoever executes, and a site found during execution is a finding against the spec rather than a gap in this plan. +**One residual this plan does not close, stated rather than left to be found.** Nothing here establishes that the edit set is complete — it is the sites §F knows, and §F's own §I records that it cannot establish completeness either. Task 8's step 1 re-derives every citation against the real file and Task 14's diff catches a copy that fell out of step; neither is a completeness check. **Task 12b performs the sweep and records what it read**, and a site it finds is a finding against the spec rather than a gap in this plan. **What stays open is that the sweep is a reader's judgement and nothing checks its coverage** — the record says what was examined, not that the examination was complete. --- From 6ace06f2c50ff36c8a5c59a3c84550538f6a10ff Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 12:45:39 +0200 Subject: [PATCH 101/181] docs(plans): apply Gate-A plan pass 3; fragments exist once, and are generated MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Pass 2 found three fragments preserved inside their own replacements. I repaired the three table rows. Pass 3 found the same three, because every task step also quoted its fragment inline — the plan held two copies of each and I had fixed one. That is the second-copy defect, in the document written to avoid it. The repair is that fragments now exist once. The table is the only authored copy and Task 0 generates .context/loop-rule-verify.sh from it, reading the rows out of the plan; every command block sources that file and refers to $P5_OLD, $F10_OLD and so on. No task step contains a fragment any more. The generator's known limit — backticks and single quotes inside a fragment — is stated with the two rows it affects, and a round-trip check runs before any task uses it. Nine findings were shell that cannot run: blocks assigning BASE but never sourcing the helper defining pair(); checks written as bare quoted strings after a loop; $start/$end/$f used as sed addresses with no step producing them; a colon-delimited site list whose anchors contain colons. Site and range lists are tab-separated files now, written by the step that needs them. Three were accounting: c15 marked replaced while §H reproduces it verbatim; the floor-arithmetic split dropping a3-a12 between its spans; the human-exception split naming no anchors. Two were duties with no home: no task applied the twelve prompt-standards items to the installed text, so Task 15 gains step 4b; and Task 12b's sweep record was to be committed into .context/, which .gitignore refuses. And Tasks 10 and 11 cannot be accepted independently — Task 10 leaves the hook suite red until Task 11 moves the expectations. Stated plainly rather than claimed away. Eight Minors collected. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-plan-om0bdd7udh-pass-3.md | 32 ++ .../gate-a-plan-om0bdd7udh-resume.md | 47 ++- .../2026-09-14-loop-rule-consolidation.md | 348 ++++++++++++------ 3 files changed, 319 insertions(+), 108 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-3.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-3.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-3.md new file mode 100644 index 0000000..0b92191 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-3.md @@ -0,0 +1,32 @@ +MAJOR | high | Task 4 step 1 | OLD fragment `So this exit needs three things` occurs once at C:229 and W:432 but is preserved in target §C's fenced replacement | It fails the required absent-from-own-replacement condition and cannot discriminate c4's change | Use a single-line unique fragment containing the removed `a missing one means keep going` wording and verify it against both files and the target fenced blocks +MAJOR | high | Task 5 step 3, CLAUDE.md | The pair still uses OLD `the tells and hand the decision to the user, and the`, which occurs at C:267 and is preserved in target §D's fenced replacement | The pair returns `old/worktree=1` after a correct installation and cannot reach its stated four-value result | Use verified row P5's C-specific OLD `not discretionary** — you report` +MAJOR | high | Task 5 step 3, workflow-init.md | The pair still uses OLD `tells and hand the decision to the user, and the`, which occurs at W:471 and is preserved in target §D's fenced replacement | The pair returns `old/worktree=1` after a correct installation and cannot reach its stated four-value result | Use verified row P5w's W-specific OLD `not discretionary** — report the` +MAJOR | high | Task 7 steps 1 and 3 | The task shortens row P16 to OLD `nonce duties at their strictest`, which occurs once at C:157 and W:364 but is preserved in target §H's fenced replacement | Both strict-reading pairs retain `old/worktree=1`, so Task 7 cannot pass even though table row P16 itself is valid | Use the complete verified P16 fragment `and the nonce duties at their strictest — the cycle` in both the locator and pair +MAJOR | high | Task 4 step 4 | The retained pair takes OLD from c14 and NEW `a recurrence failing them being an ordinary fresh finding` from c8, even while the next paragraph admits they are different edits | The evidence entry can cite a four-value result that proves neither edit as a discriminating pair and overclaims what was checked | Delete the cross-wired pair and run separate OLD/NEW pairs for c8 and c14, alongside the required c4 pair +MAJOR | high | Task 8 step 3 | The command still invokes `pair '' ''` and tells the executor to append fourteen OLD rows even though F1–F14 are already concrete in the fragment table | The reviewed rows are not consumed, the literal command cannot produce the expected counts, and following the prose duplicates the table rows | Drive the fourteen calls from F1–F14, derive and verify only each NEW half after installation, and update the existing rows rather than appending duplicates +MAJOR | high | Tasks 3, 4, 5, 7 and 8 pair blocks | Each block invokes `pair()` after only assigning `BASE`; none sources `.context/loop-rule-pair.sh`, despite the plan's separate-shell model and its claim that every such block sources the helper | All five checks fail with `pair: command not found`, so their expected results cannot be observed | Begin every block with `. .context/loop-rule-pair.sh` and the non-empty `BASE` guard +MAJOR | high | Task 1 step 4 | The parent-count block assigns `BASE=$(cat .context/loop-rule-base)` but never rejects an empty value | If the file is empty, `git show ":$f"` reads the index and can make the counterfactual appear valid without examining the recorded parent | Source the helper and run the stated `test -n "$BASE"` guard before counting +MAJOR | high | Task 10 step 5 | The four-pair hook block neither assigns `BASE` nor sources the helper in the shell that runs its `git show` commands | Under the plan's separate-shell rule, every parent count reads the index through an empty revision instead of the recorded baseline | Source `.context/loop-rule-pair.sh` and validate `BASE` at the start of the block +MAJOR | high | Task 10 step 5 | The item-11 and item-17 checks are shown as two standalone quoted strings after the four-item loop rather than as loop entries or commands | Executing that block attempts to run the quoted text as commands and never prints the two required four-value results | Add both OLD/NEW strings to the preceding loop's iteration list and produce their counts there +MINOR | high | Task 3 step 1 | The block runs both OLD counts only against `CLAUDE.md` while the expected result also claims the same counts against W | The pre-edit state of the template is asserted without a command producing it | Loop over C and W or add the two corresponding template commands +MINOR | high | Task 4 step 1 | The block runs both OLD-fragment counts only against `CLAUDE.md` while the expected result says `1` in both copies | The task can advance without observing either template locator | Run the same counts against W and record both results +MINOR | high | Task 7 step 5 | The task says to parity-check ten bounded sites but gives no extraction command, list, expected result, or place to record the output | Task 7 can be marked complete without its claimed parity observation, leaving review of its ten-site commit dependent on a later task | Provide the ten bounded extractions and exact expected differences, or explicitly make Task 14 the consuming check and remove this unperformed step +MINOR | high | Task 8 step 5 | The task says `Parity for all fourteen sites` without a command, region list, expected result, or recorded output | The fourteen-site commit has no task-local parity evidence and the step is non-falsifiable | Provide bounded C/W diffs for all fourteen installed replacements and require no unexplained output +MINOR | high | Task 9 step 5 | The task says `Parity for both sites` without any command or expected result | The two replacements can diverge while Task 9 is still marked done | Add bounded diffs for the Named residual and work-loop replacements and state the accepted output +MINOR | high | Tasks 3, 4, 6, 7 and 8 commit steps | These tasks add OLD or NEW fragment evidence to this plan, but their `git add` commands stage only the edited prompt files | The fragment-table edits remain dirty and are absorbed by a later unrelated commit, contradicting the plan's independently reviewable task units | Stage the plan in every commit whose task changes its fragment table +MAJOR | high | Task 0 step 2, floor untouched range | The prescribed two spans run from before a2 and from after a13, omitting a3–a12 at C:74–127 and W:281–334 even though the disposition marks them kept and untouched | The final parent check cannot detect accidental changes to most of the floor and cited-set conditions | Add the missing middle span from immediately after a2 through immediately before a13 and preserve all three spans in the recorded interface +MAJOR | high | Task 0 step 2, human-exception untouched range | Excluding h4 and h19 requires three concrete spans, but the task only says `the spans between those two sites` and defines no start/end anchors or serialization for them | The promised check of all twenty-four kept h conditions has no complete artifact that Task 14 can consume | Record explicit spans for h1–h3, h5–h18 and h20–h26 with stable anchors and a documented file format +MINOR | high | Task 2 step 1 | The task claims to consume the five recorded untouched ranges but hard-codes only passages d, f and j and never reads `.context/loop-rule-untouched` | The floor and human-exception conditions are not checked at this review boundary, contrary to the task's stated interface | Iterate every recorded span, including the floor and human-exception spans, or narrow the task's claim and interface to the three checks it actually performs +MAJOR | high | Task 0 step 3 | The colon-delimited site encoding cannot represent its own anchors: `\*\*Severity:\*\*:\*\*Tool routing:` yields an empty end after `${site##*:}`, and `On squash-merge:On squash-merge` uses the same unique line as both range endpoints | The Severity and squash baseline extractions run to an unintended later match or EOF, producing unrelated output instead of the inventoried sites | Store start and end anchors as two separately quoted fields with a delimiter absent from them, and give the squash line a distinct following end anchor or compare it as a single line +MINOR | high | Task 0 step 3 expected result | The raw C/W diffs also report the inventoried c-passage blank-line difference and d-passage line-wrap differences, but the closed expected list names only b, e, f and g wording divergences and calls anything else drift | A correct baseline is falsely classified as pre-existing drift | Include the known structural differences from the inventory in the expected result or compare normalized words where line breaks are intentionally non-normative +MAJOR | high | Task 14 step 1 | The purported mechanical parity check is one generic command using undefined `$start` and `$end`; no earlier step produces a changed-site region list for it to iterate | The task can record a divergence list without having diffed every target §§A–H site required by design §6 | Build and execute a concrete bounded-region list derived from every NEW/REPLACED marker and every prompt-copy §F item, then record every region checked +MAJOR | high | Task 14 step 4b | Task 0 records anchor strings, but this block uses undefined `$f`, `$start` and `$end` as numeric `sed` addresses and contains no loop over the recorded spans | It produces neither the promised ten comparisons nor valid parent/worktree slices, so identical accidental edits to kept conditions can survive to Gate B | Define a reader for the recorded anchor format, resolve each endpoint independently in parent and worktree, iterate every physical span in both files, and assert each diff is empty +MAJOR | high | Task 14 step 3 | Step 3 orders the executor to apply W's b3 alignment after Task 3 already installs the approved aligned wording, while the preceding paragraph says Task 14 verifies rather than performs that alignment | A correct Task 3 leaves Task 14 with an impossible or redundant mutation and obscures which task owns b3 | Replace Step 3 with an explicit verification that W already matches C and target §B, failing back to Task 3 if it does not +MAJOR | high | Condition disposition c15 | The plan marks c15 replaced, but target §H preserves verbatim `**Surfacing does not close the cycle, and that is what makes this reachable.**` | The 135-condition accounting records a carried condition as replaced and fails the per-condition accuracy duty in design §5 and story criterion 5 | Mark c15 carried unchanged inside the §H block and leave c16 as the condition whose hold wording is replaced +MAJOR | high | Task 12b step 6 | `.context/loop-rule-sweep.md` is ignored by `.gitignore`'s `.context/*` rule, so plain `git add` rejects it; the task also says the plan records the sweep but never copies the record into the plan | The completeness-sweep task cannot make its required commit and its evidence is absent from the final reviewed tree | Write the sweep record into a stable section of this plan and commit the plan, or deliberately change the repository's tracking policy and use an explicit supported path +MAJOR | high | Task 12b step 1 | The closed sweep question set omits decisions §A expressly owns, including source-block repair and reread routing, repeated-dismissal cleanliness, duty classification and discharge, and the parked state after an irreparable closing act | The reader can complete the recorded sweep without checking whether another live sentence still answers those parts of the ordering | Derive the question set from every ownership item in §A's opening and record each examined decision class, including these omitted ones +MAJOR | high | Tasks 10 and 11 | Task 10 deliberately commits hook text while stating the hook suite fails until Task 11, so the two commits are not independently acceptable and a reviewer cannot reject Task 11 while approving Task 10 | The plan's task boundaries create a known red intermediate state and separate each reminder from the assertions that specify it | Merge the hook-string and corresponding test updates into one task and commit only after both shell runs and ShellCheck pass +MAJOR | high | Task 15 step 7 | The Gate-B loop says to fix and re-review but never stages and commits or amends each fix before resolving the next `headSha` | A re-review can target the unchanged WIP tip while the repair remains only in the worktree, and the final squash can then publish a fix no clean pass reviewed | After every repair, rerun required checks, commit or amend the WIP snapshot, resolve the new full head object name, and use that committed head for both review branches +MAJOR | high | Task 15 steps 5 and 8 | Step 5 writes only the pre-review evidence entry, while the final Gate-B curve is not known until step 7 and no step writes the provenance line, final curve, or any exception record into `.context/loop-rule-closing-msg` before `git commit -F` | The closing commit can omit records the plan and installed rules require even though the comment beside the command claims the file already contains them | Add a post-cycle assembly step that writes the complete revalidated commit body after the clean pass, verifies every required record is present, and only then performs the squash commit +MAJOR | medium | Task 15 step 4 and invariant 11 | The plan notes that eleven prompt-standard items need reader judgement but contains no task that applies all twelve checklist predicates to the installed C, W and hook prompt text or records the result | The battery's three narrow prompt checks can pass while the implementation violates an unchecked prompt standard, leaving invariant 11 dependent on an unstated assumption about Gate B | Add an explicit reader check against all twelve items before Gate B and record the examined prompt sites and result without claiming mechanical coverage +END OF FINDINGS (31 total) diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md index 05d6ebb..545daaa 100644 --- a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md @@ -32,7 +32,8 @@ worth checking before a pass rather than after. |---|---|---|---|---|---|---| | 1 | 5871d0a | **32** | **1** | **28** | yes | first pass. The Blocker is real: with a WIP commit per task, `baseSha = HEAD^` would have put only the version bump in Gate B's range. **The dominant class is verification fragments that count zero** — the reviewer tested them against the real files and found five of Task 7's ten locators, both of Task 3's, Task 5's and Task 7's strict-reading OLD all wrapping across lines or quoting text that does not exist. Repaired by one **verified fragment table** for the whole plan instead of guessed fragments per task | | 2 | 9f13a2c | 32→**32** | 1→**0** | 28→**25** | yes | one tell (findings flat, and they cluster on the plan's own checks — the instrument). **Pass 1's repair reproduced the same defect at the next condition down:** three OLD fragments are *preserved inside their own replacements*, so their old-count can never reach zero. My checker tested single-line and unique but not that third condition, though this plan states all three. Checker fixed to test against the target's **fenced blocks**; three rows replaced; the fourteen §F OLD fragments derived and committed rather than deferred | -| 3 | — | — | — | — | not run | next, against the pass-2 repair commit | +| 3 | b767a2a | 32→**31** | 0→**0** | 25→**23** | yes | one tell (instrument cluster). **Pass 2's three fragments came back — because I repaired the table and left the same fragments quoted inline in the task steps.** Second-copy defect, in the plan written to avoid it. Structural repair: the table is the only authored copy and Task 0 **generates** the shell variables from it, so no step can restate a fragment. 8 Minors collected | +| 4 | — | — | — | — | not run | next, against the pass-3 repair commit | ## Pass-1 report @@ -106,3 +107,47 @@ record, and explicitly not a mechanical guard. **And the closing commit would have destroyed its own evidence:** step 5 wrote the evidence entry into a WIP body, step 8 squashed with `git reset --soft`, which keeps the tree and discards every WIP message. The entry goes to a file now. + + +## Pass-3 report + +**Trend:** findings 32, 32, **31**; Blockers 1, 0, **0**; Majors 28, 25, **23**. **Cluster:** the +plan's own verification apparatus, for the third pass running — the **instrument**, so one tell. +**Require↔withdraw:** none. + +**The finding that names the mechanism.** Pass 2 found three fragments preserved inside their own +replacements. I repaired the three **table rows**. Pass 3 found the same three, because every task +step also quoted its fragment **inline** — so the plan held two copies of each fragment and I had +fixed one. That is the second-copy defect, in the document written to avoid it, at the third +attempt. + +**The repair is that fragments now exist once.** The table is the only authored copy, and Task 0 +**generates** `.context/loop-rule-verify.sh` from it by reading the rows out of the plan; every +command block sources that file and refers to `$P5_OLD`, `$F10_OLD` and so on. **No task step +contains a fragment any more**, so the failure cannot recur in this shape. The generator's known +limit — fragments containing backticks or single quotes — is stated with the two rows it affects, +and a round-trip check runs before any task uses it. + +**Nine findings were shell that cannot run**: blocks assigning `BASE` but never sourcing the helper +that defines `pair()`; two checks written as bare quoted strings after a loop rather than as loop +entries; `$start`, `$end` and `$f` used as `sed` addresses with no step producing them; a +colon-delimited site list whose own anchors contain colons, so `**Severity:**` splits at the wrong +one. The site and range lists are tab-separated files now, written by the step that needs them. + +**Three were accounting.** `c15` was marked replaced while §H reproduces it verbatim — carried, and +corrected. The floor-arithmetic split dropped `a3`–`a12` between its two spans, so most of the floor +had no untouched check at all; it is three spans now. And the human-exception split named no anchors. + +**Two were duties with no home.** No task applies the twelve `prompt-standards.md` items to the +installed text — the battery's three checks are a floor, not coverage — so Task 15 gains step 4b. +And Task 12b's sweep record was to be committed into `.context/`, which `.gitignore` refuses; it +goes into the plan. + +**One boundary claim was false and is now stated plainly:** Tasks 10 and 11 cannot be accepted and +rejected independently, because Task 10 leaves the hook suite red until Task 11 moves the +expectations. They are one reviewable unit with two commits, and the plan says so rather than +claiming a boundary that is not there. + +**Eight Minors collected**, per Mechanics · Severity: locator commands run against C while claiming +a result for both copies; three parity steps with no command; commit steps not staging the plan; the +baseline expected-divergence list omitting the recorded blank-line and wrap differences. diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index 2c6f77a..b083677 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -82,31 +82,60 @@ **The NEW fragment for every pair is taken from the installed line and checked the same way**, since the new wording does not exist until the task installs it. Each task's step says which sentence of its target block to take it from, and the step fails if the fragment it chooses is not single-line and unique in the installed file. **This is the one place the plan cannot pre-verify**, and it is disclosed rather than papered over. -**The four-value command. Task 0 writes it to a file, and every task that uses it sources that file** — task command blocks run in separate shells, so a function defined here once would be `pair: command not found` everywhere it is called. +**The table is the only authored copy of every fragment, and Task 0 generates the shell variables from it.** Pass 3 found the three rows pass 2 had repaired still wrong — because the task *steps* repeated the same fragments inline, and only the table had been fixed. A plan that states a fragment twice has the second-copy defect it was written to avoid, so **no task step below quotes a fragment; each names a row id.** -Task 0 writes `.context/loop-rule-pair.sh`: +Task 0 writes `.context/loop-rule-verify.sh` by extracting the rows from this file: ```bash -cat > .context/loop-rule-pair.sh <<'SH' -BASE=$(cat .context/loop-rule-base) +{ + echo 'BASE=$(cat .context/loop-rule-base)' + echo 'test -n "$BASE" || { echo "BASE empty — Task 0 did not run"; exit 1; }' + # one `_OLD=''` line per table row, read out of this plan + awk -F'|' '/^\| (P|F)[0-9]+[a-z]? \|/ { + id=$2; frag=$4; + gsub(/^[ \t]+|[ \t]+$/, "", id); gsub(/^[ \t]+|[ \t]+$/, "", frag); + sub(/^`+/, "", frag); sub(/`+$/, "", frag); + gsub(/^[ \t]+|[ \t]+$/, "", frag); + if (frag ~ /^\*/) next; # P1 has no fragment + gsub(/'\''/, "'\''\\'\'''\''", frag); + printf "%s_OLD='\''%s'\''\n", id, frag }' \ + docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md + cat <<'SH' pair() { # pair printf '%s old/worktree=%s old/parent=%s new/worktree=%s new/parent=%s\n' "$3" \ "$(grep -cF "$1" "$3")" "$(git show "$BASE:$3" | grep -cF "$1")" \ "$(grep -cF "$2" "$3")" "$(git show "$BASE:$3" | grep -cF "$2")" } SH +} > .context/loop-rule-verify.sh ``` -and **every later command block that counts anything begins with**: +**Every command block below that counts anything begins with `. .context/loop-rule-verify.sh`** — task blocks run in their own shells, so a function or variable set elsewhere is `command not found` here, and the sourced file carries the empty-`$BASE` guard with it. + +- [ ] **Confirm the generator round-trips before any task uses it** ```bash -. .context/loop-rule-pair.sh -test -n "$BASE" || { echo "BASE empty — Task 0 did not run"; exit 1; } +. .context/loop-rule-verify.sh +for id in P2 P3 P4 P5 P5w P6 P7 P8 P9 P10 P11 P12 P13 P14 P15 P16 P17 P18 \ + F1 F2 F3 F4 F5 F6 F7 F8 F9 F10 F11 F12 F13 F14; do + eval "v=\$${id}_OLD" + printf '%-4s C=%s W=%s %s\n' "$id" \ + "$(grep -cF "$v" CLAUDE.md)" \ + "$(grep -cF "$v" plugins/dev-workflow/commands/workflow-init.md)" "${v:0:40}" +done ``` -**The `test` is not decoration.** With `$BASE` empty, `git show ":$f"` reads the *index* rather than the recorded parent and every parent count silently describes the wrong tree. +Expected: `C=1 W=1` for every row except **P5** (`C=1 W=0`) and **P5w** (`C=0 W=1`), which are the +per-copy `e7` rows. **A row printing `0` where its table entry claims a hit means the extraction +mangled it — fix the generator, not the table.** + +**Fragments containing a backtick or a single quote are the generator's known limit**, and two rows +have backticks inside them (F2, F8). Confirm those two round-trip by eye before trusting the loop. -**A pair passes only on `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`.** All four values matter: a copy carrying the new wording **and** the old one satisfies a one-sided presence check and is exactly the two-instructions-that-disagree failure the pair exists to catch. +**The four-value rule, stated once:** a pair passes only on `old/worktree=0 old/parent=1 +new/worktree=1 new/parent=0`. All four matter — a copy carrying the new wording **and** the old one +satisfies a one-sided presence check, which is the two-instructions-that-disagree failure the pair +exists to catch. --- @@ -159,7 +188,8 @@ Story acceptance criterion 5 is satisfied here. Ids are `docs/superpowers/specs/ | c10, c11 | **moved** to §A, which states what a Blocker/Major-free pass at or above the floor does. | | c12, c13 | **moved** to §A, beside a19. | | c14 | **replaced, not moved.** The ordering splits the below-floor Minor case into suspend and continue; no copy of the live wording survives beside them. | -| c15, c16 | **replaced** — §H's `c18`-and-surfacing block. The hold is now over the cycle and the new hold. | +| c15 | **carried.** §H's block reproduces `**Surfacing does not close the cycle, and that is what makes this reachable.**` verbatim; it is reproduced because the plan installs one contiguous string, not because it changes. | +| c16 | **replaced** — §H's block. The hold is now over the cycle and the new hold, not over the finding alone. | | c17 | **replaced** — the resolve rule now scopes to the assigned fix set and states what a validly dismissed recurrence owes. | | c18 | **replaced** — the blanket no-clean-credit goes; a pass is credited on its own findings, and a scope-stop trigger is what withholds credit. | | c19 | **replaced** — the one-answer resumption goes; what the answer does is the ordering's. | @@ -236,11 +266,10 @@ empty revision — which fails loudly in `git show` but quietly in a `grep -c` p task begins with: ```bash -. .context/loop-rule-pair.sh -test -n "$BASE" || { echo "BASE empty — Task 0 did not run"; exit 1; } +. .context/loop-rule-verify.sh ``` -**Task 0 also writes `.context/loop-rule-pair.sh`**, whose contents are given in the fragment-table +**Task 0 also writes `.context/loop-rule-verify.sh`**, whose generator is given in the fragment-table section. It sets `$BASE` and defines `pair()`, so both arrive together and neither can be used without the other. @@ -279,16 +308,22 @@ after the insertion point addresses different text in the parent than in the wor comparison using them reports CHANGED on identical content. Each check resolves its own start and end anchor in each tree at comparison time. -**Two of the five ranges must exclude the conditions this change replaces**, or a correct -implementation fails its own check: +**Two ranges must be split around the conditions this change replaces**, or a correct +implementation fails its own check — and each split must leave **every kept condition inside some +span**, which the first draft's version did not: -- **The floor arithmetic range must exclude `a2` and `a13`.** `a2` is the HARD FLOOR parenthetical - Task 8 replaces (row F10) and `a13` is the no-restating prohibition Task 7 replaces (row P9). - Record the arithmetic as **two spans** — from `Both gates are a LOOP with a HARD FLOOR` to just - before `(Blocker/Major only)`, and from just after `a13`'s sentence to `Nothing here writes the - floor knob`. -- **The human-exception range must exclude `h4` and `h19`**, which Task 8 replaces (rows F7 and F4). - Record the twenty-four kept conditions as the spans **between** those two sites. +- **The floor arithmetic is three spans**, because `a2` sits at its head and `a13` in its middle: + from `Both gates are a LOOP with a HARD FLOOR` to the line before row F10's fragment; from the + line after F10's to the line before row P9's; and from the line after P9's to `Nothing here writes + the floor knob`. **The middle span is the one the first draft dropped**, and it holds `a3`–`a12`. +- **The human exception is three spans**, around rows F7 (`h4`) and F4 (`h19`): from `Recording a + human exception` to before F7's line; between F7's and F4's; and from after F4's to `because + writing it down makes it sound`. All twenty-four kept conditions lie inside them. + +**Record the spans as one `startendfile` line each in `.context/loop-rule-untouched`**, the +anchors being literal strings. Tab-separated because the anchors contain colons — the first draft +used `:` as the delimiter, and `**Severity:**:**Tool routing:` splits at the wrong colon and yields +an empty end anchor. - [ ] **Step 3: Confirm the parity baseline of the inventoried ranges** @@ -299,22 +334,27 @@ whose expected list named it — and the truncation hid about fifty of the rough comparison actually emits. ```bash -# one bounded region per site, resolved by anchor in each file, output in full -for site in 'Both gates are a LOOP:Nothing here writes the floor knob' \ - 'What a loop absorbs:Recognizing "clearly stuck"' \ - 'Recognizing "clearly stuck":Every pass report states' \ - 'From pass 4 onward:Those three lines expose' \ - 'Those three lines expose:The two rules above' \ - 'The two rules above:Findings go to a FILE' \ - '\*\*Severity:\*\*:\*\*Tool routing:' \ - 'Recording a human exception:because writing it down makes it sound' \ - 'On squash-merge:On squash-merge' \ - 'When these rules bind:Downstream has no shipping commit'; do - s=${site%%:*}; e=${site##*:} +# Tab-separated start and end anchors: the anchors contain colons, so a +# colon delimiter splits '**Severity:**' at the wrong place and yields an +# empty end. A single-line site is given the same anchor twice and sed +# returns that one line. +printf '%s\n' \ + 'Both gates are a LOOP\tNothing here writes the floor knob' \ + 'What a loop absorbs\tRecognizing "clearly stuck"' \ + 'Recognizing "clearly stuck"\tEvery pass report states' \ + 'From pass 4 onward\tThose three lines expose' \ + 'Those three lines expose\tThe two rules above' \ + 'The two rules above\tFindings go to a FILE' \ + '\*\*Severity:\*\*\t\*\*Tool routing:' \ + 'Recording a human exception\tbecause writing it down makes it sound' \ + 'On squash-merge\tOn squash-merge' \ + 'When these rules bind\tDownstream has no shipping commit' \ + > .context/loop-rule-sites +while IFS=$(printf '\t') read -r s e; do echo "== $s" diff <(sed -n "/$s/,/$e/p" CLAUDE.md) \ <(sed -n "/$s/,/$e/p" plugins/dev-workflow/commands/workflow-init.md) -done | tee .context/loop-rule-baseline-diff.txt +done < .context/loop-rule-sites | tee .context/loop-rule-baseline-diff.txt ``` Expected: the deliberate divergences the inventory records — `b3`'s cross-reference target, passage @@ -373,7 +413,7 @@ copies while every count and the parity diff still pass — the paragraphs are i nothing else observes them. ```bash -BASE=$(cat .context/loop-rule-base) +. .context/loop-rule-verify.sh for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do for frag in 'How a cycle ends — one ordering' \ '' \ @@ -427,8 +467,8 @@ This task exists because `f1` and the (d)/(j) dispositions are falsifiable only - [ ] **Step 1: Diff each untouched passage against the parent** ```bash -. .context/loop-rule-pair.sh -test -n "$BASE" || { echo "BASE empty — Task 0 did not run"; exit 1; } +. .context/loop-rule-verify.sh +# read the five recorded ranges rather than hard-coding three of them for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do for anchor in 'From pass 4 onward every pass report carries three lines' \ 'The two rules above do not compete' \ @@ -475,12 +515,13 @@ table. Both are verified single-line and unique; re-confirm before editing, sinc have touched these files: ```bash -BASE=$(cat .context/loop-rule-base) -grep -cF 'scope the approved story or plan assigns to this cycle, plus repair obligations you already' CLAUDE.md -grep -cF 'the moment the user says whether the set now includes it' CLAUDE.md +. .context/loop-rule-verify.sh +for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do + printf '%s P2=%s P3=%s\n' "$f" "$(grep -cF "$P2_OLD" "$f")" "$(grep -cF "$P3_OLD" "$f")" +done ``` -Expected: `1` each, and the same against W. +Expected: `P2=1 P3=1` for **both** files — the first draft ran these against C only while claiming a result for both. **The first draft of this plan named `plus repair obligations you already accepted in earlier passes` here, which wraps across C 201–202 and W 408–409 and counts zero in a correct file.** @@ -552,11 +593,17 @@ git commit -m "WIP: replace passage (b) with the absorb paragraph that owns the - [ ] **Step 1: Record the old-wording fragments** ```bash -grep -cF 'a Blocker/Major-free pass below the floor' CLAUDE.md -grep -cF 'So this exit needs three things' CLAUDE.md +. .context/loop-rule-verify.sh +for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do + printf '%s P4=%s\n' "$f" "$(grep -cF "$P4_OLD" "$f")" +done ``` -Expected: `1` each, in both copies. +Expected: `P4=1` for both files. + +**`So this exit needs three things` is not usable and is not a row:** it occurs once in each copy but +is **preserved in target §C's fenced replacement**, so its old-count could never reach zero. Derive +`c4`'s OLD from the part of the sentence the replacement removes. - [ ] **Step 2: Install §C's block** @@ -575,20 +622,24 @@ Row **P4**. `NEW` is the single-line fragment `a recurrence failing them being a finding`, from §C's re-raised-dismissal clause — **install that clause's line unwrapped** so the fragment sits wholly on one line, and confirm it is unique before counting. +**Three pairs, one per edit.** An earlier draft ran a single pair taking its OLD from `c14` and its +NEW from `c8` — two different changes — so either could land while the other survived and it still +reported a pass, and `c4` had no observation at all. + ```bash -BASE=$(cat .context/loop-rule-base) +. .context/loop-rule-verify.sh for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do - pair 'a Blocker/Major-free pass below the floor' \ - 'a recurrence failing them being an ordinary fresh finding' "$f" + pair "$P4c4_OLD" '<§C c4 NEW — derive, verify, add as a row>' "$f" # c4 replaced + pair "$P4c8_OLD" 'a recurrence failing them being an ordinary fresh finding' "$f" # c8 widened + pair "$P4_OLD" '<§C c14 routing NEW — derive, verify, add as a row>' "$f" # c14 removed done ``` -Expected: `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`. +Expected for all six: `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`. -**One pair per edit, not one per task.** The pair above takes its OLD from `c14` and its NEW from -`c8` — two different changes — so `c8` can land while `c14` survives, or the reverse, and it still -reports a pass. `c4` has no observation at all. **Add a pair for `c4`'s replaced wording and one for -`c14`'s removal against §C's conditional routing, with rows in the fragment table**, before Step 5. +**`P4c4_OLD` and `P4c8_OLD` do not exist yet.** Derive each from the live passage, check it the three +ways, and **add it to the fragment table** — the generator reads the table, so a row that is not +there is not a variable. - [ ] **Step 5: Walk the conditions** — `c1`–`c3` present unchanged, `c5`–`c7` word for word, `c8` carrying the new clause, `c9` split, `c10`–`c14` gone from here. @@ -625,12 +676,18 @@ replacement**, so it can never be the old half — that is the trap design §7 r revisions. Rows **P5** (C) and **P5w** (W) instead, and they differ because W drops the pronoun: ```bash -grep -cF 'the tells and hand the decision to the user, and the' CLAUDE.md # 1 -grep -cF 'tells and hand the decision to the user, and the' plugins/dev-workflow/commands/workflow-init.md # 1 +. .context/loop-rule-verify.sh +printf 'P5 C=%s\n' "$(grep -cF "$P5_OLD" CLAUDE.md)" +printf 'P5w W=%s\n' "$(grep -cF "$P5w_OLD" plugins/dev-workflow/commands/workflow-init.md)" ``` -**The first draft named `and the "clearly stuck" reading above is not a precondition for it`, which -wraps across C 267–268 and W 471–472 and counts zero in both correct files.** +Expected: `1` each. + +**Two earlier drafts got this row wrong in two different ways.** The first named `and the "clearly +stuck" reading above is not a precondition for it`, which wraps across C 267–268 and W 471–472. The +second named `the tells and hand the decision to the user, and the`, which is **preserved in target +§D's replacement** and so could never reach zero. The rows now take the clause §D actually +removes. - [ ] **Step 2: Install §D's `e7` sentence and the pointer paragraph** @@ -705,16 +762,15 @@ pair, the old unscoped resolve duty can survive, or the new repair-versus-dismis omitted, while everything Task 6 checks passes. ```bash -. .context/loop-rule-pair.sh -test -n "$BASE" || { echo "BASE empty — Task 0 did not run"; exit 1; } +. .context/loop-rule-verify.sh for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do - pair 'is not settled here, and this change does not settle it' \ - 'The demotion changes what a cycle must resolve, never what it observes' "$f" - pair '' \ - 'for every finding in the assigned fix set as the absorb paragraph computes it' "$f" + pair "$P6_OLD" 'The demotion changes what a cycle must resolve, never what it observes' "$f" + pair "$P6r_OLD" 'for every finding in the assigned fix set as the absorb paragraph computes it' "$f" done ``` +**`P6r_OLD` is the resolve-duty row step 1 derives and adds to the table.** + Expected for all four: `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`. - [ ] **Step 4: Confirm g4 is gone from C** @@ -768,18 +824,13 @@ is clean`` (one line in C, wraps in W), `A clean pass is the single body line` ( after `A`), and `at minimum floor 3, severity classified without the demotion` (wraps C 155–156). ```bash +. .context/loop-rule-verify.sh for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do echo "== $f" - grep -cF 'These records are one contract' "$f" # P7 §G - grep -cF 'the resolve rule is not waived, no pass is credited as' "$f" # P8 c18 - grep -cF 'Every other rule stated here about how a cycle closes' "$f" # P9 a13 - grep -cF 'fix Blocker/Major after each' "$f" # P10 a16 - grep -cF 'final pass must be clean' "$f" # P11 a17 - grep -cF 'when a pass is clean' "$f" # P12 Gate-A clean signal - grep -cF 'clean pass is the single body line' "$f" # P13 template clean sentence - grep -cF 'Each pass: validate, revise, re-run' "$f" # P14 Gate-A cadence - grep -cF 'The Blocker/Major filter, the file-first findings protocol' "$f" # P15 lens list - grep -cF 'nonce duties at their strictest' "$f" # P16 strict-reading list + for id in P7 P8 P9 P10 P11 P12 P13 P14 P15 P16; do + eval "o=\$${id}_OLD" + printf '%-4s %s\n' "$id" "$(grep -cF "$o" "$f")" + done done ``` @@ -808,21 +859,19 @@ source sentence per block: | P16 | strict-reading list | `every suspension binding, since` | ```bash -BASE=$(cat .context/loop-rule-base) +. .context/loop-rule-verify.sh for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do - pair 'These records are one contract' '' "$f" - pair 'the resolve rule is not waived, no pass is credited as' '' "$f" - pair 'Every other rule stated here about how a cycle closes' '' "$f" - pair 'fix Blocker/Major after each' '' "$f" - pair 'final pass must be clean' '' "$f" - pair 'when a pass is clean' '' "$f" - pair 'clean pass is the single body line' '' "$f" - pair 'Each pass: validate, revise, re-run' '' "$f" - pair 'The Blocker/Major filter, the file-first findings protocol' '' "$f" - pair 'nonce duties at their strictest' '' "$f" + for id in P7 P8 P9 P10 P11 P12 P13 P14 P15 P16; do + eval "o=\$${id}_OLD"; eval "n=\$${id}_NEW" + pair "$o" "$n" "$f" + done done ``` +**`_NEW` is set by the executor after installing that block and verifying the chosen fragment +the same three ways** — append it to the generated file, or extend the table with a NEW column and +re-run the generator. + Expected for all twenty: `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`. **Two classification notes, decided against the real files rather than asserted.** The @@ -884,9 +933,12 @@ per copy where the wrapping differs, as `e7` needed); unique in each; and **not own replacement** — compare it against the §F block before accepting it. ```bash -BASE=$(cat .context/loop-rule-base) +. .context/loop-rule-verify.sh for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do - pair '' '' "$f" # ×14 + for id in F1 F2 F3 F4 F5 F6 F7 F8 F9 F10 F11 F12 F13 F14; do + eval "o=\$${id}_OLD"; eval "n=\$${id}_NEW" + pair "$o" "$n" "$f" + done done ``` @@ -896,7 +948,7 @@ Expected for all twenty-eight: `old/worktree=0 old/parent=1 new/worktree=1 new/p ```bash # Each of step 3's fourteen pairs printed new/worktree; total them per copy. -. .context/loop-rule-pair.sh +. .context/loop-rule-verify.sh for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do n=0 for new in '' '' '' '' '' '' '' \ @@ -947,11 +999,10 @@ Rows **P17** and **P18**. Assigning the variables is not running the check — t at the assignment. ```bash -. .context/loop-rule-pair.sh -test -n "$BASE" || { echo "BASE empty — Task 0 did not run"; exit 1; } +. .context/loop-rule-verify.sh for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do - pair 'Hook text is out of scope here' 'not a blanket exemption for hook text' "$f" - pair 'execute → tests green → Gate B → commit' 'Gate A (spec) → Gate-A closing act' "$f" + pair "$P17_OLD" 'not a blanket exemption for hook text' "$f" + pair "$P18_OLD" 'Gate A (spec) → Gate-A closing act' "$f" done ``` @@ -1083,7 +1134,15 @@ non-unique fragment pass for the five that should be exact: shellcheck --shell=sh plugins/dev-workflow/hooks/codex-gate.sh ``` -Expected: exit 0, no output. The suite will fail at this point because the expectations still pin the old strings — that is Task 11. +Expected: exit 0, no output. + +**The suite fails here, and that is a known red state between two commits.** Task 10 changes the +hook strings and Task 11 moves the expectations that pin them; a reviewer cannot accept Task 10 and +reject Task 11 without leaving the repo unable to pass its own battery. **They are therefore one +reviewable unit with two commits**, and neither is offered for independent acceptance — the plan's +task-independence claim does not extend to this pair, and saying so is cheaper than pretending to a +boundary that is not there. **Do not run the battery between them**; Task 11 step 4 is the first +point where green is expected. - [ ] **Step 7: Commit** @@ -1201,9 +1260,15 @@ what was read, and what was found. - [ ] **Step 1: Read the installed ordering once more, and list what it now decides** -Closure, eligibility, cleanliness, the hold, composition, the two scope triggers, the suspensions and -their answers, the two gates' closing acts. **This list is the sweep's question set** — for each, -"does any other sentence in these three files still answer this?" +Closure, eligibility, cleanliness, the hold and what discharges it, composition, the two scope +triggers, the suspensions and their answers, the two gates' closing acts, **the duty +classification**, **the source-block branch and its two reread routes**, **the repeated-dismissal +cleanliness exclusion**, and **the parked state after a closing act that cannot be repaired**. + +**Read the list off the installed §A rather than from here.** This enumeration is a floor and an +earlier draft's was short by four — a closed list in a plan is the bookkeeping that goes stale, and +the block itself is the only complete statement. For each item: "does any other sentence in these +three files still answer this?" - [ ] **Step 2: Read every section of both prompt copies that gives an instruction about a pass, a finding, a gate or a commit** @@ -1231,10 +1296,17 @@ sweep whose negative result is unrecorded cannot be told from a sweep that never - [ ] **Step 6: Commit** ```bash -git add .context/loop-rule-sweep.md docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +git add docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md git commit -m "WIP: record the completeness sweep" ``` +**The record goes in the plan, not in `.context/`.** `.gitignore` carries `.context/*` with only +`codex-gate.on` and `codex-reviews/` exempt, so `git add .context/loop-rule-sweep.md` is refused and +the task could not make its own commit. Write the sweep record under +`## Completeness sweep (Task 12b output)` at the end of this plan, replaced idempotently by the same +rule Tasks 13 and 14 use. The scratch files Task 0 writes stay in `.context/` and stay ignored — +they are working state, not a deliverable. + --- ## Task 13: The next-state table and the per-condition closure checks @@ -1296,11 +1368,21 @@ item in §F whose destination is a prompt copy. For each, extract a **bounded** first line to the first line of the next passage, not a fixed count — and diff the two copies: ```bash -# one bounded region per changed site; $start and $end are that site's own anchors -diff <(sed -n "/$start/,/$end/p" CLAUDE.md) \ - <(sed -n "/$start/,/$end/p" plugins/dev-workflow/commands/workflow-init.md) +# .context/loop-rule-changed-sites is written by this step: one +# startend line per section the target marks NEW or REPLACED and per +# §F item whose destination is a prompt copy. Build it by reading those +# markers off the target text, then: +while IFS=$(printf '\t') read -r s e; do + echo "== $s" + diff <(sed -n "/$s/,/$e/p" CLAUDE.md) \ + <(sed -n "/$s/,/$e/p" plugins/dev-workflow/commands/workflow-init.md) +done < .context/loop-rule-changed-sites ``` +**The list is a step output, not an assumed input.** An earlier draft gave one generic command with +undefined `$start` and `$end` and no step producing them, so the task could record a divergence list +without having diffed anything. + **Fail this step where a site named by §§A–H has no region in your list** — that is the same completeness failure as Task 13's per-condition checks, and it is caught the same way: by reading the set off the artifact rather than from a list kept here. @@ -1319,7 +1401,12 @@ gives W the pronoun; classifying that divergence as deliberate here would let th disagree on a sentence the approved target states once. **`b3` is not in it either** — Task 3 aligns it, and this task verifies the alignment rather than performing it. -- [ ] **Step 3: Apply W's `b3` alignment** +- [ ] **Step 3: Verify W's `b3` alignment, which Task 3 performed** + +Task 3 installs §B's common span, which aligns `b3`. **This step confirms it rather than repeating +it** — an earlier draft told this task to apply the alignment, which is either impossible or +redundant after a correct Task 3, and contradicted the paragraph above that assigns the alignment to +Task 3. - [ ] **Step 4: Re-run the diff** and confirm only the deliberate divergences remain. @@ -1332,11 +1419,16 @@ task claims to change. ```bash BASE=$(cat .context/loop-rule-base) -# for each of the five ranges recorded in .context/loop-rule-untouched, in both copies: -diff <(git show "$BASE:$f" | sed -n "$start,$end p") <(sed -n "$start,$end p" "$f") +while IFS=$(printf '\t') read -r s e f; do + echo "== $f :: $s" + diff <(git show "$BASE:$f" | sed -n "/$s/,/$e/p") <(sed -n "/$s/,/$e/p" "$f") +done < .context/loop-rule-untouched ``` -Expected: no output, ten times. **This is the last check before Gate B** and it is the only one that +**Anchors, resolved separately in each tree** — Task 1 inserts a large block, so a line number taken +from either tree addresses different text in the other, and an earlier draft's numeric `sed` +addresses would have reported CHANGED on identical content. Expected: no output for every recorded +span. **This is the last check before Gate B** and it is the only one that would catch an identical accidental edit in both copies. - [ ] **Step 5: Commit** @@ -1391,7 +1483,21 @@ sh scripts/check-version-bump.test.sh && sh scripts/check-version-bump.sh main & claude plugin validate . --strict ``` -Expected: exit 0. **`check-invariants.sh` includes the prompt-conformance checks** — a `Target model:` line naming one recognized model, a prose checklist-count claim matching the checklist, and the finding-severity vocabulary as a closed set in both prompt copies. Those three are a floor, not coverage; invariant 11's other eleven items are judged by a reader. +Expected: exit 0. + +- [ ] **Step 4b: Apply all twelve `docs/prompt-standards.md` items to the installed text** + +**Nothing mechanical does this and no other task claims it.** The battery's three narrow checks are +a floor — one `Target model:` spelling, one prose count claim, one severity vocabulary — and +invariant 11 requires all twelve items of every skill, command, hook message and scaffolded template +this change touches. **Read the installed §A–§H text in C, in W, and the seven hook strings, against +each of the twelve items, and record the result per item in this plan.** Items 6 (every constraint +carries its reason in the same sentence) and 8 (token-lean) are the ones design §8 names as most at +risk. + +**A reader check, deliberately** — no pattern decides whether a constraint carries its reason. + +**`check-invariants.sh` includes the prompt-conformance checks** — a `Target model:` line naming one recognized model, a prose checklist-count claim matching the checklist, and the finding-severity vocabulary as a closed set in both prompt copies. Those three are a floor, not coverage; invariant 11's other eleven items are judged by a reader. - [ ] **Step 5: Write the evidence entry into `.context/loop-rule-closing-msg`** @@ -1399,6 +1505,11 @@ Expected: exit 0. **`check-invariants.sh` includes the prompt-conformance checks discards every WIP message; evidence written only there would be destroyed by the close. Restate it in the WIP body too if a mid-cycle reader would want it, but the file is the copy that survives. +**The file is completed after step 7, not here.** The evidence entry can be drafted now, but the +**per-pass curve is not known until the Gate-B loop ends**, and the **provenance line** and any +**human-exception record** belong beside it. Step 7's last action is to append all three — this step +opens the file, step 7 closes it, and step 8 commits it. + It names: the battery run; **every pair this plan built, with its counts in each copy and each tree, and every presence check beside them**; the §6 parity diff and the `b11`/`b13` equivalence result; and the next-state table's location plus its row count. **State the §A presence checks as presence, not as pairs** — its counterfactual is absent and is claimed as absent. - [ ] **Step 6: Run Gate B** @@ -1419,7 +1530,19 @@ and is the only value whose range contains the whole implementation. - [ ] **Step 7: Loop to a clean pass at or above the derived floor** -Floor derives from the story profile: risk `high` → 2, security `none` → 0, max 2 ≠ 0 → **floor 3**. Re-derive it at each pass from the header. Fix Blocker/Major after each pass; re-review after every fix. **Revalidate the evidence entry before every re-review and before the closing amend.** +Floor derives from the story profile: risk `high` → 2, security `none` → 0, max 2 ≠ 0 → **floor 3**. Re-derive it at each pass from the header. + +**Each fix is committed before the next review is issued**, or the re-review targets the unchanged +WIP tip while the repair sits in the worktree — and the final squash then publishes a fix no pass +reviewed: + +```bash +git add -A && git commit -m "WIP: fix " +git rev-parse HEAD # resolve headSha fresh for the next call +``` + +Re-review after every fix. **Revalidate the evidence entry before every re-review and before the +closing commit.** **A fix that changes specified behaviour updates the spec in the same commit.** @@ -1430,8 +1553,7 @@ so an evidence entry written only into a WIP message is destroyed at exactly the closes — which is what step 5 would otherwise have done. ```bash -. .context/loop-rule-pair.sh -test -n "$BASE" || { echo "BASE empty — Task 0 did not run"; exit 1; } +. .context/loop-rule-verify.sh # .context/loop-rule-closing-msg holds the revalidated evidence entry, the provenance # line, the per-pass curve and any human-exception record, written before this step. git reset --soft "$BASE" @@ -1470,3 +1592,15 @@ The closing body carries: the validated evidence entry; the provenance line; the ## Divergence list (Task 14 output) *Empty until Task 14 runs. Task 14 replaces this entire section.* + +--- + +## Completeness sweep (Task 12b output) + +*Empty until Task 12b runs. Task 12b replaces this entire section.* + +--- + +## Prompt-standards result (Task 15 step 4b output) + +*Empty until Task 15 runs. It replaces this entire section, one line per checklist item.* From 58b3660d2c1b3a47a3e2f68ff9a95f0a130d8264 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 13:09:21 +0200 Subject: [PATCH 102/181] docs(plans): apply Gate-A plan pass 4; delete the generator, guard the empties MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit MANDATORY TWO-TELL STOP: Blockers rose 0 -> 10 and the findings cluster on the instrument for the fourth pass running. Surfaced; standing answer applied and the loop continued. Pass 3's fix for "fragments stated twice" was a generator that awk-parsed this plan's own markdown table into shell variables. It was broken three ways: it also matched Task 7's NEW-source table and overwrote nine _OLD variables; its sentinel test dropped every fragment beginning with **; and its id grammar could not express the rows Tasks 3, 4 and 6 add. Its failure mode is the dangerous one — an unset variable makes grep -cF "" match every line, so every pair reports a healthy-looking count against nothing. Deleted rather than debugged. A markdown parser is the wrong instrument for thirty-two lines. The helper is transcribed by hand and validated against the real files, and pair() now refuses an empty OLD or NEW. A fourth fragment was preserved inside its own replacement and invisible to my checker: F4 is one line in CLAUDE.md but the target's fenced block wraps it between "you" and "still". Re-running the check with line breaks normalized found exactly that row and no other. Also: two pair blocks still unsourced; two hook checks written as bare strings after a loop; Task 0 overwriting the recorded base on a re-run, which would put every earlier edit outside Gate B's range; plan records written after the last WIP commit and never staged, which reset --soft would have left in the worktree; h5 marked kept while §F item 7 changes it; and one pair per §F block treated as coverage where a block changes four conditions. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-plan-om0bdd7udh-pass-4.md | 21 ++ .../gate-a-plan-om0bdd7udh-resume.md | 51 ++++- .../2026-09-14-loop-rule-consolidation.md | 187 ++++++++++-------- 3 files changed, 173 insertions(+), 86 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-4.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-4.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-4.md new file mode 100644 index 0000000..212a343 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-4.md @@ -0,0 +1,21 @@ +BLOCKER | high | Task 0 fragment generator | The awk row matcher scans the entire plan, so it also matches Task 7's P7–P16 NEW-source table and emits later assignments that overwrite P7_OLD through P14_OLD and P16_OLD with NEW fragments. | The required round-trip immediately reports C=0 W=0 for nine rows, so Task 0 cannot complete and every later pair would read the wrong OLD value if the mismatch were overlooked. | Bound extraction to the verification-fragment tables, require their exact column schema, and reject duplicate ids before emitting assignments. +BLOCKER | high | Task 0 fragment generator | The P1 sentinel test skips every fragment beginning with an asterisk after Markdown delimiters are removed, which also drops the valid bold-leading F3, F5 and F7 fragments. | Those variables are unset; the round-trip counts an empty pattern against every line and cannot reach its expected result, blocking execution before Task 1. | Skip by id equal to P1 or by the exact no-fragment sentinel, never by the first character of fragment content. +MAJOR | high | Verification table row F4 | F4's OLD fragment is preserved word-for-word in target §F item 4; it only appears absent because the target wraps between “you” and “still.” | Installed unwrapped, the old count can never become zero; installed with that wrap, the pair passes while the old instruction remains semantically present, so the check does not discriminate the edit. | Replace F4 with a unique single-line clause the replacement actually changes, and test preservation after normalizing whitespace across fenced-block line breaks. +BLOCKER | high | Task 4 step 4 | The prescribed ids P4c4 and P4c8 cannot match the generator's id grammar, which permits only digits followed by at most one letter. | Even if the executor adds both rows and regenerates the helper, neither variable is emitted and all four c4/c8 pairs operate on empty OLD values. | Adopt ids accepted by a declared grammar or widen the parser to the actual ids, then include them in the round-trip verification. +BLOCKER | high | Tasks 3–7 fragment-table updates | The helper is generated only in Task 0, but later tasks add OLD rows and then source the already-generated file without any required regeneration step. | Newly required values such as P4c4_OLD, P4c8_OLD and P6r_OLD remain unavailable, so their pair commands cannot produce the stated four-value result. | Make generation a reusable checked command and require it immediately after every table edit before any new variable is consumed. +BLOCKER | high | Task 8 step 3 | F1–F14 already exist as OLD-only rows, yet the step says to build and append those fourteen rows and then consumes F1_NEW through F14_NEW, which neither the table nor the generator produces. | Following the step either creates duplicate authored rows or leaves every NEW variable empty; in both cases the fourteen pairs cannot be trusted or completed. | Keep one existing row per id, add a defined NEW field or a separate generated NEW map, regenerate once, and verify all twenty-eight variables before pairing. +MINOR | high | Task 0 round-trip check | The check proves only that an extracted value is some unique substring of C and W and prints only its first 40 characters; it never compares the generated value byte-for-byte with the table cell. F2's closing backtick around additionalContext is after character 40. | A truncated or otherwise mangled value that remains unique can report the expected counts and the instructed visual check cannot inspect the damaged suffix. | Compare each generated value to an independently parsed full table cell, or print and compare an unambiguous full-length representation plus length for every row. +BLOCKER | high | Tasks 3 step 3, 5 step 3 and 10 step 5 | The Task 3 and Task 5 pair blocks set BASE but never source the helper that defines pair, while Task 10's separately scoped pair block uses BASE without setting or sourcing it. | The first two blocks fail with pair not found, and Task 10's parent reads use an empty revision and return misleading zero counts through the grep pipelines. | Begin each block with the required helper source and remove the redundant local BASE assignment. +BLOCKER | high | Task 10 step 5 | The advertised item-11 and item-17 pairs are two standalone quoted strings after the loop, not entries in a loop or calls to any checker. | A shell tries to execute each whole string as a command, while the two meaning changes receive no four-value observation. | Put all six pair specifications in one loop or invoke a checked helper explicitly for each of the two omitted items. +BLOCKER | high | Task 5 step 3 | Both OLD halves are the stale fragments the surrounding prose says were rejected because target §D preserves them; the repaired P5_OLD and P5w_OLD table variables are not used. | With the target's shown wrapping the old/worktree counts stay 1 and the task cannot pass; with a different wrap they can become 0 without the words being removed. | Source the helper and pass P5_OLD and P5w_OLD as the OLD halves, deleting the inline stale copies. +MAJOR | high | Verification table and Tasks 3, 6, 9 plus Self-Review | The plan's “table is the only authored copy” invariant is false: P2 and P3 are repeated inline in Task 3, P6 in Task 6, P17 in Task 9, and P7 in Self-Review, with further table substrings repeated in Task 7's historical examples. | The plan still has parallel copies that can drift independently, the exact mechanism pass 3 was meant to remove; Task 5 already demonstrates the resulting stale-copy failure. | Replace every repeated OLD fragment with its generated variable or row id and keep historical explanations descriptive rather than quoting the fragment. +MAJOR | high | Task 0 step 3 | The baseline-site writer uses printf with a percent-s format and single-quoted backslash-t text, so it writes literal backslash-t characters rather than tab delimiters. | The tab-IFS reader receives the entire line as the start pattern and an empty end pattern, allowing the baseline diff to extract nothing instead of the ten intended regions. | Emit fields with an actual tab using a tabbed format string or explicit two-argument formatting, then validate that every row splits into exactly two non-empty fields before diffing. +BLOCKER | high | Task 0 step 1 | Re-running Task 0 after any partial implementation unconditionally replaces loop-rule-base with the current WIP tip. | Gate B and the final soft reset then omit all earlier WIP edits from their range, allowing prompt changes to escape review and leaving WIP commits in history. | Create the base file only when absent, validate an existing value as the parent of the first WIP, and stop on any attempted rebase of the recorded cycle. +MINOR | high | Task 2 step 1 | The task claims to consume Task 0's recorded five-range do-not-touch set, but its command ignores loop-rule-untouched and checks only hard-coded windows for passages d, f and j. | Accidental changes to the kept floor-arithmetic and human-exception spans are not caught at the immediate checkpoint and survive until Task 14, despite Task 2 being presented as the early guard. | Iterate the recorded anchor file here with the same independently resolved bounded-range comparison used by Task 14. +MAJOR | high | Condition disposition for h5 | h5 is marked kept and untouched, but target §F item 7 changes “restated by the closing amend” to “restated by the commit its closing act produces.” | The 135-condition accounting misclassifies a changed closing destination and Task 0's claimed untouched h spans cannot protect all twenty-four conditions it says they contain. | Mark h5 replaced, account for why its broader closing-act destination is required, and split the untouched spans around the complete item-7 sentence rather than only h4's fragment line. +MAJOR | high | Task 7 step 3 | One sampled pair per large block does not cover several independent meaning changes: the c18/surfacing block also changes c16, c17, c19 and c20, and P16 observes only the first of multiple newly added unknown-start readings. | An executor can omit the new hold route, answer route, in-set narrowing, repeated-dismissal rule, parked state or closure/pass-cost strict readings while every stated Task 7 pair and parity check passes. | Add a discriminating pair or add-only presence check for each independent changed clause, while continuing to source every OLD value from the single table. +BLOCKER | high | Task 15 step 4b | The prompt-standards reader check writes its per-item result into this plan after the last WIP commit, but no later step stages or commits that plan before Gate B or the final reset-and-commit. | Gate B reviews a head that excludes required invariant-11 evidence, and a clean cycle closes with the plan record left only as an unstaged worktree change. | Commit the updated plan as a WIP before resolving headSha and rerun any evidence affected by that commit. +MAJOR | high | Task 15 steps 6–8 | The final clean Gate-B branch files are created after the last WIP head and no subsequent step stages them; reset --soft stages committed WIP content only. | The repository's intentionally tracked codex-reviews record omits the closing pass, so the committed curve cannot be re-derived from the files it claims preserve those counts and the worktree remains dirty. | After validating the final pass, explicitly stage that cycle's two findings files before the closing commit without using an indiscriminate path. +MAJOR | high | Global Constraints, Gate-B cycle discipline | The global rule says the cycle closes by git commit --amend, while Task 15 correctly closes its many-WIP shape with reset --soft followed by a new commit. | A worker following the higher-level constraint can amend only the final WIP and leave the earlier WIP commits in history, violating the closing-act requirement the plan later enforces. | State the invariant as “use the repository-state-dependent closing act,” and name amend versus reset-and-commit only at the operation-defining task. +NIT | high | Task 7 step 1 | The table introduction says its fragments were verified at 9f13a2c, but Task 7 says the same table was verified at 5871d0a, a commit that predates the table's introduction. | A drift failure points the executor at the wrong provenance and makes the stated mechanical history internally inconsistent. | Use the actual verification commit consistently or remove the redundant commit claim from Task 7. +END OF FINDINGS (20 total) diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md index 545daaa..e73a10f 100644 --- a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md @@ -33,7 +33,8 @@ worth checking before a pass rather than after. | 1 | 5871d0a | **32** | **1** | **28** | yes | first pass. The Blocker is real: with a WIP commit per task, `baseSha = HEAD^` would have put only the version bump in Gate B's range. **The dominant class is verification fragments that count zero** — the reviewer tested them against the real files and found five of Task 7's ten locators, both of Task 3's, Task 5's and Task 7's strict-reading OLD all wrapping across lines or quoting text that does not exist. Repaired by one **verified fragment table** for the whole plan instead of guessed fragments per task | | 2 | 9f13a2c | 32→**32** | 1→**0** | 28→**25** | yes | one tell (findings flat, and they cluster on the plan's own checks — the instrument). **Pass 1's repair reproduced the same defect at the next condition down:** three OLD fragments are *preserved inside their own replacements*, so their old-count can never reach zero. My checker tested single-line and unique but not that third condition, though this plan states all three. Checker fixed to test against the target's **fenced blocks**; three rows replaced; the fourteen §F OLD fragments derived and committed rather than deferred | | 3 | b767a2a | 32→**31** | 0→**0** | 25→**23** | yes | one tell (instrument cluster). **Pass 2's three fragments came back — because I repaired the table and left the same fragments quoted inline in the task steps.** Second-copy defect, in the plan written to avoid it. Structural repair: the table is the only authored copy and Task 0 **generates** the shell variables from it, so no step can restate a fragment. 8 Minors collected | -| 4 | — | — | — | — | not run | next, against the pass-3 repair commit | +| 4 | 6ace06f | 31→**20** | 0→**10** | 23→**7** | yes | **MANDATORY TWO-TELL STOP** — Blockers rose 0→10, and the findings cluster on the instrument for the fourth pass running. Surfaced, standing answer applied, loop continued. **The awk generator pass 3 introduced was broken three ways**; it is deleted, not debugged — the helper is transcribed by hand and validated against the real files, with empty-string guards, because a silently unset variable makes `grep -cF ""` match every line. **A fourth fragment (F4) was preserved in its own replacement and invisible to my checker**, which compared without normalizing the block's line breaks | +| 5 | — | — | — | — | not run | next, against the pass-4 repair commit | ## Pass-1 report @@ -151,3 +152,51 @@ claiming a boundary that is not there. **Eight Minors collected**, per Mechanics · Severity: locator commands run against C while claiming a result for both copies; three parity steps with no command; commit steps not staging the plan; the baseline expected-divergence list omitting the recorded blank-line and wrap differences. + + +## Pass-4 report — MANDATORY TWO-TELL STOP + +**Trend:** findings 32, 32, 31, **20**; Blockers 1, 0, 0, **10**; Majors 28, 25, 23, **7**. +**Cluster:** the plan's verification apparatus, fourth pass running — the **instrument**. +**Require↔withdraw:** none. + +**Two tells: the Blocker count rose from zero to ten, and the instrument cluster persists.** The +stop is mandatory, not discretionary. Surfaced to Daniel; his standing answer of 2026-09-13 applies +and the loop continued on it. + +**The tells are reading something real, and it is my repair strategy rather than the plan.** Pass 3's +fix for "fragments stated twice" was a generator: Task 0 would `awk`-parse this plan's own markdown +table into shell variables. Pass 4 found it broken three ways at once — it also matched Task 7's +NEW-source table and overwrote nine `_OLD` variables; its sentinel test dropped every fragment +beginning with `**`; and its id grammar could not express the rows Tasks 3, 4 and 6 add. **Its +failure mode is the dangerous one**: an unset variable makes `grep -cF ""` match every line, so every +pair reports a healthy-looking count against nothing. + +**The generator is deleted rather than debugged.** A markdown parser is the wrong instrument for +thirty-two lines. The helper is **transcribed by hand** and **validated against the real files** — +the validation is what makes the transcription safe, and `pair()` now refuses an empty OLD or NEW, +which is the guard that would have caught the generator's failure had it existed. + +**A fourth fragment was preserved inside its own replacement**, and my checker could not see it: +`F4`'s text is one line in `CLAUDE.md` but the target's fenced block wraps it between `you` and +`still`, so a substring test found nothing. **Re-running the check with line breaks normalized found +exactly that one row and no other.** Three passes, three different ways for a fragment to be wrong, +and each time the checker learned the condition after the reviewer found it. + +**Four blockers were shell that cannot run** — two pair blocks still carrying the old +`BASE=$(cat …)` prefix without sourcing the helper that defines `pair`, and two hook checks written +as bare quoted strings after a loop rather than inside it. **Two more were state**: re-running Task 0 +would have overwritten the recorded base with the current WIP tip, putting every earlier edit outside +Gate B's range and outside the final reset; and the plan records written after the last WIP commit +were never staged, so `reset --soft` would have left them in the worktree and out of the closing +commit. + +**Two were accounting**, the fourth and fifth in four passes: `h5`, whose Gate-B destination §F item 7 +changes, was marked kept; and one pair per §F block is a sample rather than coverage where a block +changes four conditions. + +**What I would tell a reader of this record:** the product text has been stable since pass 1. Every +finding in four passes has been about the apparatus that checks it, and each of my repairs to that +apparatus has introduced a new defect in it. That is the signal the two tells are carrying, and the +answer taken here is to make the apparatus smaller — no parser, no generated state, hand-written +lines whose only guarantee is a check against the real files. diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index b083677..e4965e3 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -23,7 +23,7 @@ - **Invariant 5 (exact pinning)** and **invariant 12 (a plugin change requires a version bump)**: this change touches `plugins/`, so `plugins/dev-workflow/.claude-plugin/plugin.json` goes `0.11.0 → 0.12.0` with a CHANGELOG entry (design §8). - **Invariant 11:** every prompt change passes all 12 items of `docs/prompt-standards.md`. Most at risk: item 6 (every constraint carries its reason in the same sentence) and item 8 (token-lean). - **Invariant 4 / the hook:** no control flow, counter, fingerprint computation, routing or event handling changes. Only `note` strings and their expectations. -- **Gate-B cycle discipline:** snapshot commits are named `WIP: …`; the cycle closes by `git commit --amend`. A non-`WIP` commit mid-cycle resets the hook's counters. +- **Gate-B cycle discipline:** snapshot commits are named `WIP: …`, and a non-`WIP` commit mid-cycle resets the hook's counters. **This change closes with `git reset --soft "$BASE"` followed by one commit, not with `--amend`** — Mechanics prescribes the reset shape wherever several WIP snapshots piled up, and this plan makes one per task. Task 15 step 8 is the operation. --- @@ -64,7 +64,7 @@ | F1 | 1, the `WIP:` naming warning | `nor resets your pass counters. A pre-review snapshot named anything else reads as a` | 825 | | F2 | 2, the Gate-B coverage instruction | `` `NO FINDINGS` if clean" in `additionalContext`, with the same one-line format. `` | 598 | | F3 | 3, the curve's Majors rationale | `**Majors are recorded as well as Findings and Blockers**, because the severity rule moves the` | 938 | -| F4 | 4, the human-exception scope sentence | `own terminal actions and this paragraph changes none of them: on a STOP you still stop, and` | 1014 | +| F4 | 4, the human-exception scope sentence | `neither a human's assent nor this record` | 1015 | | F5 | 5, the `Finishing the cycle` lead-in | `**Finishing the cycle:** after the final clean pass, close it with` | 827 | | F6 | 6, the Gate-A broad-prompt instruction | `reviews the TEXT you pass, not the git tree). Use ONE broad prompt, re-run it` | 553 | | F7 | 7, the human-exception destination | `**Which commit:** an ungated change records it in that commit; a Gate-A cycle in the spec or` | 988 | @@ -82,60 +82,52 @@ **The NEW fragment for every pair is taken from the installed line and checked the same way**, since the new wording does not exist until the task installs it. Each task's step says which sentence of its target block to take it from, and the step fails if the fragment it chooses is not single-line and unique in the installed file. **This is the one place the plan cannot pre-verify**, and it is disclosed rather than papered over. -**The table is the only authored copy of every fragment, and Task 0 generates the shell variables from it.** Pass 3 found the three rows pass 2 had repaired still wrong — because the task *steps* repeated the same fragments inline, and only the table had been fixed. A plan that states a fragment twice has the second-copy defect it was written to avoid, so **no task step below quotes a fragment; each names a row id.** +**The table is the only authored copy of every fragment, and no task step below quotes one** — each names a row id. Pass 3 found the three rows pass 2 had repaired still wrong, because the steps repeated the fragments inline and only the table had been fixed; a plan stating a fragment twice has the second-copy defect it was written to avoid. -Task 0 writes `.context/loop-rule-verify.sh` by extracting the rows from this file: +**Task 0 transcribes the rows into `.context/loop-rule-verify.sh` by hand, and a round-trip check validates the transcription against the real files.** An earlier revision generated that file by `awk`-parsing this table out of the plan. Pass 4 found the parser broken three ways at once — it also matched Task 7's NEW-source table and overwrote nine `_OLD` variables; its "skip the sentinel row" test dropped every fragment beginning with `**`; and its id grammar could not express the rows Tasks 3, 4 and 6 add. **A markdown parser is the wrong instrument for thirty-two lines**, and its failure mode is the dangerous one: a silently empty variable makes `grep -cF ""` match every line and every pair report a passing-looking count against nothing. + +The transcription is safe because **nothing trusts it**. The check below counts each variable against the real files and rejects any that does not land exactly where its row says: ```bash -{ - echo 'BASE=$(cat .context/loop-rule-base)' - echo 'test -n "$BASE" || { echo "BASE empty — Task 0 did not run"; exit 1; }' - # one `_OLD=''` line per table row, read out of this plan - awk -F'|' '/^\| (P|F)[0-9]+[a-z]? \|/ { - id=$2; frag=$4; - gsub(/^[ \t]+|[ \t]+$/, "", id); gsub(/^[ \t]+|[ \t]+$/, "", frag); - sub(/^`+/, "", frag); sub(/`+$/, "", frag); - gsub(/^[ \t]+|[ \t]+$/, "", frag); - if (frag ~ /^\*/) next; # P1 has no fragment - gsub(/'\''/, "'\''\\'\'''\''", frag); - printf "%s_OLD='\''%s'\''\n", id, frag }' \ - docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md - cat <<'SH' +# .context/loop-rule-verify.sh — written by hand from the table, one line per row +BASE=$(cat .context/loop-rule-base) +test -n "$BASE" || { echo "BASE empty — Task 0 did not run"; exit 1; } +P2_OLD='scope the approved story or plan assigns to this cycle, plus repair obligations you already' +# … one line per row of the table above, P3 … P18 and F1 … F14 … pair() { # pair + test -n "$1" || { echo "pair: empty OLD — a variable is unset or mistyped"; return 1; } + test -n "$2" || { echo "pair: empty NEW"; return 1; } printf '%s old/worktree=%s old/parent=%s new/worktree=%s new/parent=%s\n' "$3" \ "$(grep -cF "$1" "$3")" "$(git show "$BASE:$3" | grep -cF "$1")" \ "$(grep -cF "$2" "$3")" "$(git show "$BASE:$3" | grep -cF "$2")" } -SH -} > .context/loop-rule-verify.sh ``` -**Every command block below that counts anything begins with `. .context/loop-rule-verify.sh`** — task blocks run in their own shells, so a function or variable set elsewhere is `command not found` here, and the sourced file carries the empty-`$BASE` guard with it. +**The two empty-string guards are the whole safety of the transcription.** Without them an unset variable counts every line in the file and reads as a healthy result. -- [ ] **Confirm the generator round-trips before any task uses it** +- [ ] **Validate the transcription before any task uses it** ```bash . .context/loop-rule-verify.sh for id in P2 P3 P4 P5 P5w P6 P7 P8 P9 P10 P11 P12 P13 P14 P15 P16 P17 P18 \ F1 F2 F3 F4 F5 F6 F7 F8 F9 F10 F11 F12 F13 F14; do eval "v=\$${id}_OLD" - printf '%-4s C=%s W=%s %s\n' "$id" \ + test -n "$v" || { echo "$id UNSET"; continue; } + printf '%-4s C=%s W=%s\n' "$id" \ "$(grep -cF "$v" CLAUDE.md)" \ - "$(grep -cF "$v" plugins/dev-workflow/commands/workflow-init.md)" "${v:0:40}" + "$(grep -cF "$v" plugins/dev-workflow/commands/workflow-init.md)" done ``` -Expected: `C=1 W=1` for every row except **P5** (`C=1 W=0`) and **P5w** (`C=0 W=1`), which are the -per-copy `e7` rows. **A row printing `0` where its table entry claims a hit means the extraction -mangled it — fix the generator, not the table.** +Expected: `C=1 W=1` for every row except **P5** (`C=1 W=0`) and **P5w** (`C=0 W=1`), the per-copy `e7` rows. **No `UNSET`, and no `0` where the row claims a hit.** A mismatch means the transcription is wrong — fix the file, not the table. + +**Rows Tasks 3, 4 and 6 add go into both places**: a row in this table, and a line in the helper, before the pair that uses them. **Re-run the validation loop after adding any row** — the helper is written once and never regenerates itself. -**Fragments containing a backtick or a single quote are the generator's known limit**, and two rows -have backticks inside them (F2, F8). Confirm those two round-trip by eye before trusting the loop. +**Every OLD row was checked three ways** — single-line in each copy it claims, unique there, and **absent from the target's fenced blocks compared with line breaks normalized**. The normalization matters: pass 4 found `F4`'s fragment preserved in its own replacement and invisible to a naive substring test, because the block wraps between `you` and `still`. Re-running the normalized check over all thirty-two rows found exactly that one and nothing else. -**The four-value rule, stated once:** a pair passes only on `old/worktree=0 old/parent=1 -new/worktree=1 new/parent=0`. All four matter — a copy carrying the new wording **and** the old one -satisfies a one-sided presence check, which is the two-instructions-that-disagree failure the pair -exists to catch. +**Three ways a fragment fails:** it wraps across a line break, so `grep -F` counts zero in a correct file; it is preserved inside its own replacement, so its old-count can never reach zero; or it is not unique. **Each of the first three passes found rows failing a different one of the three.** + +**The four-value rule, stated once:** a pair passes only on `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`. All four matter — a copy carrying the new wording **and** the old one satisfies a one-sided presence check, which is the two-instructions-that-disagree failure the pair exists to catch. --- @@ -226,7 +218,8 @@ Story acceptance criterion 5 is satisfied here. Ids are `docs/superpowers/specs/ | Condition | Disposition | |---|---| -| h1–h3, h5, h6, h8–h18, h20–h26 | **kept**, untouched. | +| h1–h3, h6, h8–h18, h20–h26 | **kept**, untouched. | +| h5 | **replaced** — §F item 7 changes the Gate-B destination from "restated by the closing amend" to "restated by the commit its closing act produces", the amend no longer being the only closing shape. Row F7. | | h4 | **replaced** — §F item 7, the human-exception destination: a Gate-A cycle's record goes to the commit its closing act produces, not to "the spec or plan commit". | | h7 | **kept.** | | h19 | **replaced** — §F item 4, the scope sentence: "neither a human's **general** assent nor this record", plus the clause distinguishing the answers a suspension asks for from assent. | @@ -254,12 +247,22 @@ Story acceptance criterion 5 is satisfied here. Ids are `docs/superpowers/specs/ - [ ] **Step 1: Confirm the approved artifacts and a clean tree** ```bash -git log --oneline -1 # expect 5871d0a or later on loop-rule-consolidation +git log --oneline -1 # expect 6ace06f or later on loop-rule-consolidation git status --porcelain # expect empty -git rev-parse HEAD > .context/loop-rule-base +if [ -s .context/loop-rule-base ]; then + echo "base already recorded: $(cat .context/loop-rule-base) — NOT overwriting" +else + git rev-parse HEAD > .context/loop-rule-base +fi cat .context/loop-rule-base ``` +**Never overwrite an existing base.** Re-running Task 0 after a partial implementation would record +the current WIP tip, and both Gate B's range and the final `reset --soft` would then start after +every edit made so far — prompt and hook changes would be squashed into the closing commit without +ever entering a review range. **If the recorded base is wrong, delete the file deliberately and say +why**; do not let a re-run decide it. + **Persist it to a file, not to a shell variable.** Each task runs in its own shell invocation, so a `BASE=` assignment in Task 0 is gone by Task 1 and every parent-tree count would run against an empty revision — which fails loudly in `git show` but quietly in a `grep -c` pipeline. Every later @@ -338,7 +341,7 @@ comparison actually emits. # colon delimiter splits '**Severity:**' at the wrong place and yields an # empty end. A single-line site is given the same anchor twice and sed # returns that one line. -printf '%s\n' \ +printf '%b\n' \ 'Both gates are a LOOP\tNothing here writes the floor knob' \ 'What a loop absorbs\tRecognizing "clearly stuck"' \ 'Recognizing "clearly stuck"\tEvery pass report states' \ @@ -536,16 +539,16 @@ Replace from `**What a loop absorbs, and what stops it` through the sentence §B would let the surviving instruction pass behind the repaired one. ```bash -BASE=$(cat .context/loop-rule-base) -# pair() is defined once in the fragment-table section +. .context/loop-rule-verify.sh for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do - pair 'scope the approved story or plan assigns to this cycle, plus repair obligations you already' \ - 'union of the scope every approved story or plan governing this change assigns to this cycle' "$f" - pair 'the moment the user says whether the set now includes it' \ - '' "$f" + pair "$P2_OLD" 'union of the scope every approved story or plan governing this change assigns to this cycle' "$f" + pair "$P3_OLD" "$P3_NEW" "$f" done ``` +**`P3_NEW` is the §B resumption sentence's fragment**, chosen after installing, verified the three +ways, and added to the helper before this step runs. + Expected for all four: `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`. **Two pairs do not cover §B.** Target §B separately changes `b3`, `b8`, `b11`, `b13`, `b16`, @@ -629,17 +632,18 @@ reported a pass, and `c4` had no observation at all. ```bash . .context/loop-rule-verify.sh for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do - pair "$P4c4_OLD" '<§C c4 NEW — derive, verify, add as a row>' "$f" # c4 replaced - pair "$P4c8_OLD" 'a recurrence failing them being an ordinary fresh finding' "$f" # c8 widened - pair "$P4_OLD" '<§C c14 routing NEW — derive, verify, add as a row>' "$f" # c14 removed + pair "$P19_OLD" '<§C c4 NEW — derive, verify, add as a row>' "$f" # c4 replaced + pair "$P20_OLD" 'a recurrence failing them being an ordinary fresh finding' "$f" # c8 widened + pair "$P4_OLD" '<§C c14 routing NEW — derive, verify, add as a row>' "$f" # c14 removed done ``` Expected for all six: `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`. -**`P4c4_OLD` and `P4c8_OLD` do not exist yet.** Derive each from the live passage, check it the three -ways, and **add it to the fragment table** — the generator reads the table, so a row that is not -there is not a variable. +**`P19_OLD` and `P20_OLD` do not exist yet.** Derive each from the live passage, check it the three +ways, **add a row to the fragment table and a line to `.context/loop-rule-verify.sh`**, then re-run +the validation loop. Ids continue the `P` series; a row id is whatever the table and the helper agree +on, so keep them plain. - [ ] **Step 5: Walk the conditions** — `c1`–`c3` present unchanged, `c5`–`c7` word for word, `c8` carrying the new clause, `c9` split, `c10`–`c14` gone from here. @@ -683,11 +687,10 @@ printf 'P5w W=%s\n' "$(grep -cF "$P5w_OLD" plugins/dev-workflow/commands/workflo Expected: `1` each. -**Two earlier drafts got this row wrong in two different ways.** The first named `and the "clearly -stuck" reading above is not a precondition for it`, which wraps across C 267–268 and W 471–472. The -second named `the tells and hand the decision to the user, and the`, which is **preserved in target -§D's replacement** and so could never reach zero. The rows now take the clause §D actually -removes. +**Two earlier drafts got this row wrong in two different ways**, which is why the step names ids and +not text. The first fragment wrapped across C 267–268 and W 471–472; the second was **preserved in +target §D's replacement** and could never reach zero. `P5_OLD` and `P5w_OLD` take the clause §D +actually removes. - [ ] **Step 2: Install §D's `e7` sentence and the pointer paragraph** @@ -699,15 +702,14 @@ it can be omitted from both copies while the `e7` pair, the condition walk and t pass. ```bash -BASE=$(cat .context/loop-rule-base) -pair 'the tells and hand the decision to the user, and the' \ - 'read **after** the clean-completion branch of the closure ordering' CLAUDE.md -pair 'tells and hand the decision to the user, and the' \ - 'read **after** the clean-completion branch of the closure ordering' plugins/dev-workflow/commands/workflow-init.md +. .context/loop-rule-verify.sh +NEW='read **after** the clean-completion branch of the closure ordering' +PTR='where this stop'"'"'s place among the suspensions' +pair "$P5_OLD" "$NEW" CLAUDE.md +pair "$P5w_OLD" "$NEW" plugins/dev-workflow/commands/workflow-init.md for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do printf 'pointer %s worktree=%s parent=%s\n' "$f" \ - "$(grep -cF 'where this stop'"'"'s place among the suspensions' "$f")" \ - "$(git show "$BASE:$f" | grep -cF 'where this stop'"'"'s place among the suspensions')" + "$(grep -cF "$PTR" "$f")" "$(git show "$BASE:$f" | grep -cF "$PTR")" done ``` @@ -769,7 +771,8 @@ for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do done ``` -**`P6r_OLD` is the resolve-duty row step 1 derives and adds to the table.** +**`P6r_OLD` is the resolve-duty row step 1 derives.** Add it to the table **and** to +`.context/loop-rule-verify.sh`, then re-run the validation loop before this step. Expected for all four: `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`. @@ -869,11 +872,17 @@ done ``` **`_NEW` is set by the executor after installing that block and verifying the chosen fragment -the same three ways** — append it to the generated file, or extend the table with a NEW column and -re-run the generator. +the same three ways** — append the assignment to `.context/loop-rule-verify.sh` and re-run the +validation loop before the pairs. Expected for all twenty: `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`. +**One pair per block is a sample, not coverage, and two blocks need more.** The `c18`-and-surfacing +block changes `c16`, `c17`, `c19` and `c20` besides the `c18` clause row P8 observes; the +strict-reading list adds several items where P16 observes the first. **Derive one OLD row per +independent meaning change in those two blocks**, add each to the table and the helper, and run a +pair for each — the other eight blocks change one thing each and one pair covers them. + **Two classification notes, decided against the real files rather than asserted.** The strict-reading list **replaces the dash-delimited run** even though its addition is at the tail, so it owes a full pair. The gate-prompt template's clean sentence **replaces** `A clean pass is the @@ -1092,26 +1101,22 @@ first time — it admitted exactly the shell keywords it existed to catch. The hook has one copy and owes no parity check; **the hook suite's exact-match assertion is this edit's second observation** (design §7). Run the pair anyway: ```bash -for pair in 'this floor is the only thing keeping the spec review honest|instruction-backed' \ - 'commit only if your final pass was clean — no new Blocker/Major|every other closure condition holds' \ - 'then make the real commit when your final pass is clean|Use this commit as the review range' \ - 'STOP — Codex Gate B not satisfied|Codex gate state:' ; do - OLD=${pair%%|*}; NEW=${pair##*|} - printf 'old/worktree=%s old/parent=%s new/worktree=%s new/parent=%s %s\n' \ - "$(grep -cF "$OLD" plugins/dev-workflow/hooks/codex-gate.sh)" \ - "$(git show "$BASE:plugins/dev-workflow/hooks/codex-gate.sh" | grep -cF "$OLD")" \ - "$(grep -cF "$NEW" plugins/dev-workflow/hooks/codex-gate.sh)" \ - "$(git show "$BASE:plugins/dev-workflow/hooks/codex-gate.sh" | grep -cF "$NEW")" "$OLD" -done -``` - -**Two more pairs, because four cover only five of the seven items.** Without them, item 11's old -abbreviated clean definition and item 17's running-cycle skip exit can both survive while every -listed pair passes: - -```bash - 'floor met by COUNT ONLY|Proceed only once this Gate-A cycle has closed' # item 11 - 'or proceed only if $policy|skip rule decides only whether a cycle runs at all' # item 17 +. .context/loop-rule-verify.sh +H=plugins/dev-workflow/hooks/codex-gate.sh +# Tab-separated OLDNEW: the fragments contain no tab, and | appears +# inside shell text. Six entries, all in the loop — an earlier draft left +# two of them as bare quoted strings after it, which a shell tries to run. +printf '%b\n' \ + 'this floor is the only thing keeping the spec review honest\tinstruction-backed' \ + 'floor met by COUNT ONLY\tProceed only once this Gate-A cycle has closed' \ + 'commit only if your final pass was clean — no new Blocker/Major\tevery other closure condition holds' \ + 'then make the real commit when your final pass is clean\tUse this commit as the review range' \ + 'or proceed only if $policy\tskip rule decides only whether a cycle runs at all' \ + 'STOP — Codex Gate B not satisfied\tCodex gate state:' \ + > .context/loop-rule-hookpairs +while IFS=$(printf '\t') read -r OLD NEW; do + pair "$OLD" "$NEW" "$H" +done < .context/loop-rule-hookpairs ``` **Exact counts per pair, not a floor.** Five of these replace one message each and must read @@ -1554,12 +1559,24 @@ closes — which is what step 5 would otherwise have done. ```bash . .context/loop-rule-verify.sh +# Everything that must be IN the squashed commit has to be committed before the +# reset: reset --soft stages only what the discarded commits already contained. +git add docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md \ + .context/codex-reviews/ +git status --porcelain # expect empty after this +git commit -m "WIP: plan records and Gate-B findings files" || true # .context/loop-rule-closing-msg holds the revalidated evidence entry, the provenance -# line, the per-pass curve and any human-exception record, written before this step. +# line, the per-pass curve and any human-exception record, completed at step 7. git reset --soft "$BASE" git commit -F .context/loop-rule-closing-msg ``` +**`reset --soft` stages committed content only.** The prompt-standards result (step 4b), the +completeness sweep, the next-state table and the divergence list all land in this plan, and +`.context/codex-reviews/` is tracked; **anything still uncommitted when the reset runs is left in the +worktree and is not in the closing commit** — and, for the plan records, was never in a Gate-B range +either. Commit them first, then reset. + **`$BASE` is the recorded revision, not a placeholder to substitute by hand** — Task 0 persisted it for this, and a mistaken substitution squashes the wrong range. From 916ede225f1cc290fb57b95e94aef0220144cd95 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 13:21:31 +0200 Subject: [PATCH 103/181] docs(plans): apply Gate-A plan pass 5; stop pre-writing the checks MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit MANDATORY TWO-TELL STOP: the instrument cluster for a fifth pass, plus a require↔withdraw pair — pass 4 demanded OLD rows for clauses pass 5 classifies as add-only. Surfaced; standing answer applied and the loop continued. Five passes have been saying one thing. Most of their findings are shell written in advance for text that does not exist yet: helpers defined in one shell and called in another, loop bodies outside their loops, sed ranges whose delimiters occur in their own data, sed -n "/x/,/x/p" for a single line, variables no step sets, a generated helper whose empty variable makes grep -cF "" match every line. Every one a defect in the apparatus, none in the change. Design §7 already said not to: the plan builds each pair against the real files and runs both directions THERE. The pre-written blocks are deleted and replaced by one procedure — take the OLD from its row, install, choose a NEW from the installed text, check it three ways, count four values, expect 0/1/1/0 — plus the two rules that were only implicit: guard every count against an empty pattern, and an add-only edit owes presence alone. Four findings were the condition table, now five for five: e8 marked carried while §D removes W's form; i4-i8 marked kept while §H reproduces them inside a replacement block; h5 needing its own row; and g2/g3 dropped rather than replaced, which owes an absence check and had none. Two were state bugs: .context/loop-rule-base accepted on being non-empty, so an abandoned run's value would put every edit outside Gate B's range; and the prompt-standards result written after the last WIP commit, where neither the review range nor reset --soft reaches it. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-plan-om0bdd7udh-pass-5.md | 17 + .../gate-a-plan-om0bdd7udh-resume.md | 40 ++- .../2026-09-14-loop-rule-consolidation.md | 314 ++++++------------ 3 files changed, 164 insertions(+), 207 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-5.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-5.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-5.md new file mode 100644 index 0000000..8baa98b --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-5.md @@ -0,0 +1,17 @@ +MAJOR | high | Verification-fragment validation loop | The loop only prints counts: an unset variable executes `continue`, no count is compared with its expected value, and its fixed id list omits every row later added by Tasks 1, 3, 4, 6, 7, and 8. Re-running it therefore exits successfully without examining later ids or rejecting an original mistype. | The plan's claimed round-trip validation can certify an incomplete or mistyped hand transcription, leaving later checks to fail ambiguously or, for rows never reached, to record evidence for fragments that were never validated. | Track a failure status for UNSET and unexpected C/W counts, return nonzero, and require every row-adding step to add its id to the validation list and validate any NEW assignment it creates. +BLOCKER | high | Task 0 step 1 and Task 15 step 8 | `.context/loop-rule-base` is never removed, and Task 0 accepts any pre-existing nonempty value without establishing that it is the parent of this execution's first WIP. After one completed or abandoned run, a later run preserves a stale but valid ancestor. | Gate B can review an unintended range and the final `reset --soft "$BASE"` can collapse unrelated intervening commits into the new closing commit. | Preserve the value only when the commits from it to HEAD are verified as this plan's active WIPs; otherwise stop and require deliberate reinitialization, and remove or retire the base file after a successful close. +BLOCKER | high | Task 0 step 2 and Task 14 step 4b | The proposed whole-line untouched spans cannot isolate every kept condition: `a14` begins on CLAUDE.md line 131 after changed `a13` text on the same line, and `h6` begins on CLAUDE.md line 989 and workflow-init.md line 1173 after changed `h4`/`h5` text on that same line. | Starting a span on those kept fragments makes `sed` include the changed prefix and fail a correct implementation; starting on the next line omits part of a kept condition, so the final check cannot support its completeness claim. | Record and compare normalized bounded substrings or individual kept-condition fragments for these boundaries instead of whole-line `sed` ranges. +BLOCKER | high | Task 0 step 3 and Task 14 step 4b | The plan says giving a single-line site the same start and end anchor makes `sed -n "/$s/,/$e/p"` return one line, but sed does not test the second regexp on the line that starts a regexp range. With the sole `On squash-merge` occurrence, extraction continues to EOF. | The baseline parity output includes the entire tail, and the final untouched check for passage (j) includes later intended human-exception edits, so a correct implementation reports a change and cannot satisfy the stated expectation. | Handle `s == e` with a single-address extraction or give passage (j) a distinct following end anchor in both trees. +MINOR | high | Task 2 step 1 | Despite the comment to read the five recorded ranges, the command never opens `.context/loop-rule-untouched`; it hard-codes only passages (d), (f), and (j) and uses arbitrary 12-line tails. | Accidental Task 1 damage to the floor-arithmetic or kept human-exception spans is reported as unchanged and is detected only much later by Task 14. | Iterate the recorded three-field span file with bounded extraction here, using the same corrected boundary method as the final check. +BLOCKER | high | Condition disposition, passage (e), `e8` | The table marks `e8` carried, but the inventory defines C as `you report` and W as `report`, while Task 5 explicitly installs `you report` into W and says the old W divergence does not survive. | Acceptance criterion 5's accounting is false for W and can direct a condition-preservation review to accept the very old wording Task 5 must remove. | Mark `e8` carried in C and changed/aligned in W, or mark the condition changed with the per-copy distinction. +MAJOR | medium | Condition disposition, passage (i), `i4`–`i8` | The table calls `i4`–`i8` kept even though Task 7 installs the target's whole dash-delimited replacement block containing them; this is the same contiguous-reproduction shape the table classifies as carried for `c15`, `a15`, `a21`, and `a22`. | The accounting hides that these five conditions must survive inside an edited span, and Task 7 has no explicit survival check for them, so they can be dropped while P16 observes only the extension. | Classify `i4`–`i8` as carried inside the replacement and add a reader or fragment check that each remains present. +MINOR | high | Task 0 step 2 | The plan states that passage (h) has twenty-four kept conditions, computed as twenty-six less `h4` and `h19`, but its own disposition also replaces `h5`; the actual kept total is twenty-three. | The stated completeness count disagrees with the enumeration and obscures whether the untouched spans account for every kept condition. | Change the count and deduction to twenty-three, excluding `h4`, `h5`, and `h19`. +MAJOR | high | Task 6 step 3 | P6's OLD fragment occurs in `g1`, while `g2`/`g3` are the separate interim report-and-stop duty and rationale that the disposition drops. No post-edit fragment or pair checks that duty is gone; the `g4` count and parity diff do not observe it. | An implementation can install the answer over `g1`, leave the obsolete stop duty in both copies, and satisfy both shown pairs, the C-only `g4` check, and parity. | Add a verified OLD row from the report-and-stop sentence and a four-value pair showing `g2`/`g3` gone in both copies. +MAJOR | high | Task 7 step 3, unknown-start strict-reading block | The plan recognizes several independent additions after P16 but instructs the executor to derive an OLD row for each. Those clauses are add-only and have no corresponding old wording; pairing them with unrelated text from the old list would mix edits. | The repeated-dismissal, parked-state, and remaining strict-reading additions can be omitted while P16 passes, or the executor is forced to invent nondiscriminating OLD halves. | Keep P16 for the changed list boundary and add a separately verified presence check for each independent add-only rule, following design section 7's add-only rule. +BLOCKER | high | Task 8 steps 3–4 | Task 0 creates only `F1_OLD` through `F14_OLD`; Task 8 never requires choosing or assigning `F1_NEW` through `F14_NEW` to the helper or re-running validation. Step 3 evaluates those unset names, so `pair()` rejects every empty NEW, while step 4 still iterates literal `` placeholders. | The task cannot produce its required twenty-eight passing pairs or installed=14 counts from a correct installation. | After installation, explicitly choose each NEW fragment, verify it against its own normalized target block and both installed copies, append each assignment to the helper, add the ids to a failing validator, and have both steps iterate those variables. +MAJOR | high | Task 8 step 3, §F item 7 | F7's OLD fragment observes only `h4`, ending in the old Gate-A destination, but item 7 also changes independent condition `h5` from `restated by the closing amend` to the commit produced by the closing act. One single-line pair cannot observe both changes, and no other check names the old h5 clause. | The obsolete amend-only Gate-B destination can survive in both copies while the F7 pair and parity pass. | Add a second OLD/NEW pair specific to h5's Gate-B destination in both copies. +BLOCKER | high | Task 10 step 5, item 11 pair | The OLD half `floor met by COUNT ONLY` is the unchanged count-only observation at the start of the Gate-A reminder, while the NEW half `Proceed only once this Gate-A cycle has closed` replaces the later permission sentence. Target §F says item 11 supplies a replacement sentence inside an otherwise unchanged message. | A correct installation retains the OLD phrase, so the required old/worktree=0 result cannot pass; deleting it to satisfy the pair removes approved wording and invalidates the test's still-valid count-only observation. | Use a unique OLD fragment from `Proceed only if your final pass was clean — no new Blocker/Major` and validate it against item 11's NEW sentence. +MAJOR | high | Task 10 step 5, items 10 and 17 | Item 10's pair checks removal of the honesty claim but not the separate old imperative `Run more passes before executing`; item 17 checks only `or proceed only if $policy` and not its old unconditional `run more` route. The target removes both unqualified next-action instructions. | Either stale run-more instruction can remain in its note while every shown pair has the expected values, leaving the hook contradictory to the closure ordering. | Add OLD absence observations for the two removed run-more clauses, paired with or accompanied by the corresponding verified NEW routing text. +MAJOR | high | Task 11 steps 1, 3, and 5 | The locator searches only `Gate B satisfied`, but the real test has `Gate A satisfied` in the assertions and labels at lines 504–506. Step 3 expressly says any label left saying `satisfied` is stale verdict vocabulary, and step 5 only proves the Gate-B spelling is gone. | The prescribed sweep can finish while the Gate-A test vocabulary still calls the gate satisfied, contrary to the task's own acceptance rule. | Include the Gate-A satisfied sites in the reader sweep and final zero-survival check, then rename them for the observed hook state such as floor met or cycle closed. +BLOCKER | high | Task 15 steps 4b, 6, and 8 | Step 4b writes the prompt-standards result into the plan after the last WIP commit, Step 6 resolves HEAD and runs Gate B without committing that edit, and Step 8 commits it only after the clean pass. | The closing commit contains a plan record that no Gate-B review range included, contradicting the plan's whole-implementation range claim and its own warning that fixes outside headSha are published unreviewed. | Commit the completed prompt-standards section in a WIP before resolving the initial Gate-B headSha, and re-resolve HEAD after any later plan-record repair before re-review. +END OF FINDINGS (16 total) diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md index e73a10f..0031178 100644 --- a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md @@ -34,7 +34,8 @@ worth checking before a pass rather than after. | 2 | 9f13a2c | 32→**32** | 1→**0** | 28→**25** | yes | one tell (findings flat, and they cluster on the plan's own checks — the instrument). **Pass 1's repair reproduced the same defect at the next condition down:** three OLD fragments are *preserved inside their own replacements*, so their old-count can never reach zero. My checker tested single-line and unique but not that third condition, though this plan states all three. Checker fixed to test against the target's **fenced blocks**; three rows replaced; the fourteen §F OLD fragments derived and committed rather than deferred | | 3 | b767a2a | 32→**31** | 0→**0** | 25→**23** | yes | one tell (instrument cluster). **Pass 2's three fragments came back — because I repaired the table and left the same fragments quoted inline in the task steps.** Second-copy defect, in the plan written to avoid it. Structural repair: the table is the only authored copy and Task 0 **generates** the shell variables from it, so no step can restate a fragment. 8 Minors collected | | 4 | 6ace06f | 31→**20** | 0→**10** | 23→**7** | yes | **MANDATORY TWO-TELL STOP** — Blockers rose 0→10, and the findings cluster on the instrument for the fourth pass running. Surfaced, standing answer applied, loop continued. **The awk generator pass 3 introduced was broken three ways**; it is deleted, not debugged — the helper is transcribed by hand and validated against the real files, with empty-string guards, because a silently unset variable makes `grep -cF ""` match every line. **A fourth fragment (F4) was preserved in its own replacement and invisible to my checker**, which compared without normalizing the block's line breaks | -| 5 | — | — | — | — | not run | next, against the pass-4 repair commit | +| 5 | 58b3660 | 20→**16** | 10→**7** | 7→**7** | yes | **MANDATORY TWO-TELL STOP** — instrument cluster for the fifth pass, plus a require↔withdraw pair (pass 4 demanded OLD rows for clauses pass 5 classifies as add-only). Surfaced, standing answer applied, loop continued. **The pre-written shell is deleted.** Design §7 says the plan builds each pair *against the real files*; five passes of findings were blocks written in advance for text that does not exist yet. One stated procedure replaces them | +| 6 | — | — | — | — | not run | next, against the pass-5 repair commit | ## Pass-1 report @@ -200,3 +201,40 @@ finding in four passes has been about the apparatus that checks it, and each of apparatus has introduced a new defect in it. That is the signal the two tells are carrying, and the answer taken here is to make the apparatus smaller — no parser, no generated state, hand-written lines whose only guarantee is a check against the real files. + + +## Pass-5 report — MANDATORY TWO-TELL STOP, and the over-specification named + +**Trend:** findings 32, 32, 31, 20, **16**; Blockers 1, 0, 0, 10, **7**; Majors 28, 25, 23, 7, **7**. +**Cluster:** the verification apparatus, fifth pass running. **Require↔withdraw:** pass 4 finding 16 +told me to derive an OLD row per addition in the strict-reading block; pass 5 finding 10 says those +clauses are add-only and have no old wording to remove. **Two tells, mandatory stop**, surfaced, loop +continued on the standing answer. + +**What five passes have actually been saying.** Findings 1, 3, 4, 9, 11, 13 and 16 of this pass, and +most of passes 2, 3 and 4, are one thing: **shell written in advance for text that does not exist +yet.** Helpers defined in one shell and called in another; loop bodies outside their loops; `sed` +ranges whose delimiters occur in their own data; `sed -n "/x/,/x/p"` for a single line, which runs to +the next match instead; variables no step sets; a generated helper whose empty variable makes +`grep -cF ""` match every line. **Every one was a defect in the apparatus and none in the change.** + +**Design §7 already said not to do this:** the plan *builds each pair against the real files and runs +both directions there*. **There** — with the installed text open. I had been pre-writing it. + +**So the pre-written blocks are gone**, replaced by one procedure stated once: take the OLD from its +row, install, choose a NEW from the installed text, check it the three ways, count four values, +expect `0/1/1/0`. Plus the two rules that were only ever implicit — **guard every count against an +empty pattern**, and **an add-only edit owes presence alone**, which is the withdraw half of this +pass's tell made into a rule. + +**Four findings were the condition table, which is now five for five.** `e8` was marked carried while +§D removes W's form of it; `i4`–`i8` were marked kept while §H reproduces them inside a replacement +block; `h5` needed its own row because §F item 7 changes two conditions; and `g2`/`g3` are **dropped** +rather than replaced, which owes an absence check and had none. **A dropped condition with no check +is the failure `AGENTS.md` names, and it took five passes to find the last of them.** + +**Two were genuine state bugs.** `.context/loop-rule-base` was accepted on being non-empty, so an +abandoned run's value would silently put every edit outside Gate B's range; it is now validated as an +ancestor of `HEAD` with only this run's `WIP:` commits between, and removed at close. And the +prompt-standards result was written into the plan after the last WIP commit, where neither Gate B's +range nor `reset --soft` would reach it. diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index e4965e3..cd0288a 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -67,7 +67,8 @@ | F4 | 4, the human-exception scope sentence | `neither a human's assent nor this record` | 1015 | | F5 | 5, the `Finishing the cycle` lead-in | `**Finishing the cycle:** after the final clean pass, close it with` | 827 | | F6 | 6, the Gate-A broad-prompt instruction | `reviews the TEXT you pass, not the git tree). Use ONE broad prompt, re-run it` | 553 | -| F7 | 7, the human-exception destination | `**Which commit:** an ungated change records it in that commit; a Gate-A cycle in the spec or` | 988 | +| F7 | 7, the human-exception destination (`h4`) | `**Which commit:** an ungated change records it in that commit; a Gate-A cycle in the spec or` | 988 | +| F7b | 7, the Gate-B destination (`h5`) | `restated by the closing amend` | 989 | | F8 | 8, the profile-change pass claim | ``snapshot by amend — a non-`WIP` commit reads to the hook as the cycle closing and would`` | 752 | | F9 | 7a, the mid-run recovery sentence | `taken from the provenance line and the curve, which must agree. A Gate-A cycle mid-run has no` | 416 | | F10 | 8a, the HARD FLOOR parenthetical | `(Blocker/Major only), derived from the cited story's profile.**` | 73 | @@ -82,52 +83,38 @@ **The NEW fragment for every pair is taken from the installed line and checked the same way**, since the new wording does not exist until the task installs it. Each task's step says which sentence of its target block to take it from, and the step fails if the fragment it chooses is not single-line and unique in the installed file. **This is the one place the plan cannot pre-verify**, and it is disclosed rather than papered over. -**The table is the only authored copy of every fragment, and no task step below quotes one** — each names a row id. Pass 3 found the three rows pass 2 had repaired still wrong, because the steps repeated the fragments inline and only the table had been fixed; a plan stating a fragment twice has the second-copy defect it was written to avoid. +**The table is the only authored copy of every fragment, and no task step below quotes one** — each names a row id. A plan stating a fragment twice has the second-copy defect it was written to avoid, which is what pass 3 found. -**Task 0 transcribes the rows into `.context/loop-rule-verify.sh` by hand, and a round-trip check validates the transcription against the real files.** An earlier revision generated that file by `awk`-parsing this table out of the plan. Pass 4 found the parser broken three ways at once — it also matched Task 7's NEW-source table and overwrote nine `_OLD` variables; its "skip the sentinel row" test dropped every fragment beginning with `**`; and its id grammar could not express the rows Tasks 3, 4 and 6 add. **A markdown parser is the wrong instrument for thirty-two lines**, and its failure mode is the dangerous one: a silently empty variable makes `grep -cF ""` match every line and every pair report a passing-looking count against nothing. +### The verification procedure, stated once and performed by the executor -The transcription is safe because **nothing trusts it**. The check below counts each variable against the real files and rejects any that does not land exactly where its row says: +**This plan does not pre-write the shell for each pair, and five Gate-A passes are the reason.** Design §7 assigns the plan to *build each pair against the real files and run both directions there* — **there**, where the installed text exists. Pre-writing commands for wording that does not exist yet produced, pass after pass, blocks that could not run: helpers defined in one shell and called in another, loop bodies outside their loops, `sed` ranges whose delimiters appeared in their own data, variables no step ever set. **Every one of those was a defect in the apparatus, never in the change.** What the plan owes is the *rule*, the *verified OLD fragments*, and the *expected result*; the executor writes the command in front of the files. -```bash -# .context/loop-rule-verify.sh — written by hand from the table, one line per row -BASE=$(cat .context/loop-rule-base) -test -n "$BASE" || { echo "BASE empty — Task 0 did not run"; exit 1; } -P2_OLD='scope the approved story or plan assigns to this cycle, plus repair obligations you already' -# … one line per row of the table above, P3 … P18 and F1 … F14 … -pair() { # pair - test -n "$1" || { echo "pair: empty OLD — a variable is unset or mistyped"; return 1; } - test -n "$2" || { echo "pair: empty NEW"; return 1; } - printf '%s old/worktree=%s old/parent=%s new/worktree=%s new/parent=%s\n' "$3" \ - "$(grep -cF "$1" "$3")" "$(git show "$BASE:$3" | grep -cF "$1")" \ - "$(grep -cF "$2" "$3")" "$(git show "$BASE:$3" | grep -cF "$2")" -} -``` +**For each meaning-changing edit, at the task that installs it:** -**The two empty-string guards are the whole safety of the transcription.** Without them an unset variable counts every line in the file and reads as a healthy result. +1. **Take the OLD fragment from its table row.** Confirm before editing that it counts **1** in each copy the row claims. +2. **Install the block.** +3. **Choose a NEW fragment from the installed text** and check it the same three ways: **single-line** in the file, **unique** there, and **absent from the parent tree**. +4. **Count four values** — OLD and NEW, each in the worktree and in `$BASE` — and record them. -- [ ] **Validate the transcription before any task uses it** +**A pair passes only on `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`.** All four matter: a copy carrying the new wording **and** the old one satisfies a one-sided presence check, which is the two-instructions-that-disagree failure the pair exists to catch. -```bash -. .context/loop-rule-verify.sh -for id in P2 P3 P4 P5 P5w P6 P7 P8 P9 P10 P11 P12 P13 P14 P15 P16 P17 P18 \ - F1 F2 F3 F4 F5 F6 F7 F8 F9 F10 F11 F12 F13 F14; do - eval "v=\$${id}_OLD" - test -n "$v" || { echo "$id UNSET"; continue; } - printf '%-4s C=%s W=%s\n' "$id" \ - "$(grep -cF "$v" CLAUDE.md)" \ - "$(grep -cF "$v" plugins/dev-workflow/commands/workflow-init.md)" -done -``` +**Guard every count against an empty pattern.** `grep -cF ""` matches every line, so a mistyped or unset fragment reports a healthy-looking number against nothing. **Check that each fragment is non-empty before counting with it** — this is the failure mode pass 4 found in a generated helper, and it is silent. -Expected: `C=1 W=1` for every row except **P5** (`C=1 W=0`) and **P5w** (`C=0 W=1`), the per-copy `e7` rows. **No `UNSET`, and no `0` where the row claims a hit.** A mismatch means the transcription is wrong — fix the file, not the table. +**An add-only edit owes presence alone**, because there is no old wording whose absence could be counted: `new/worktree=1 new/parent=0`, and no OLD half. **Which edits those are is decided against the real file** — an add-only edit is one whose site carries no wording the change removes. Design §7 declines to classify them and so does this plan; the executor does it with the file open. -**Rows Tasks 3, 4 and 6 add go into both places**: a row in this table, and a line in the helper, before the pair that uses them. **Re-run the validation loop after adding any row** — the helper is written once and never regenerates itself. +### Checking a fragment — the three conditions, and how each has failed -**Every OLD row was checked three ways** — single-line in each copy it claims, unique there, and **absent from the target's fenced blocks compared with line breaks normalized**. The normalization matters: pass 4 found `F4`'s fragment preserved in its own replacement and invisible to a naive substring test, because the block wraps between `you` and `still`. Re-running the normalized check over all thirty-two rows found exactly that one and nothing else. +| Condition | How it fails | Found at | +|---|---|---| +| **Single-line** in the file it is counted in | it wraps, so `grep -F` counts zero in a correct tree | pass 1, nine rows | +| **Unique** in that file | a count of 1 proves nothing about which occurrence changed | — | +| **Absent from its own replacement** | its old-count can never reach zero | pass 2, three rows | + +**The third check compares against the target's fenced blocks with line breaks normalized.** `F4` is one line in `CLAUDE.md` but the block wraps it between `you` and `still`; a substring test finds nothing and the row looks usable. Pass 4 found it, and re-running the normalized check over all thirty-two rows found that one and no other. -**Three ways a fragment fails:** it wraps across a line break, so `grep -F` counts zero in a correct file; it is preserved inside its own replacement, so its old-count can never reach zero; or it is not unique. **Each of the first three passes found rows failing a different one of the three.** +### The OLD fragments -**The four-value rule, stated once:** a pair passes only on `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`. All four matter — a copy carrying the new wording **and** the old one satisfies a one-sided presence check, which is the two-instructions-that-disagree failure the pair exists to catch. +**Every row below was checked all three ways at `58b3660`.** A row Tasks 3, 4, 6 and 7 add is checked the same way and appended here, so this table stays the one place they live. --- @@ -196,8 +183,9 @@ Story acceptance criterion 5 is satisfied here. Ids are `docs/superpowers/specs/ | Condition | Disposition | |---|---| | e1–e6 | **kept** — the five tells themselves are untouched. | +| e8 | **replaced.** The inventory records it as a parity divergence — C `you report`, W `report` — and §D's block supplies C's wording for **both** copies, so W's form goes. **Not carried**: the condition as inventoried names a difference this change removes. Task 5. | | e7 | **changed** — gains the read-after-clean-completion clause. The sentence is given entire in §D. | -| e8, e9, e10 | **carried** inside §D's block. | +| e9, e10 | **carried** inside §D's block. | | e11 | **kept** — the C-only rationale paragraph is untouched and stays C-only. | | — | **added:** §D's pointer paragraph at the end of the passage. | @@ -228,7 +216,7 @@ Story acceptance criterion 5 is satisfied here. Ids are `docs/superpowers/specs/ ### Passage (i) — when these rules bind → target §H (Task 7) `i1`, `i2`, `i3`, `i9`–`i11`, `i13`–`i16`: **kept**, untouched. -`i4`–`i8`: **kept**, and the dash-delimited list they sit in is **extended, not rewritten** — §H gives the whole list with the additions at the end. +`i4`–`i8`: **carried**, not merely kept. §H reproduces the whole dash-delimited run as one contiguous string so the plan installs it in a single edit, and the five conditions come through unchanged inside it. **Recorded as carried rather than kept** for the same reason `a15`, `a21`, `a22` and `c15` are: a condition reproduced inside a replacement block is not untouched text, and calling it untouched would put it outside the untouched-range checks that are supposed to protect it. `i12`: **discharged, and kept.** It is the extension point that licenses the addition; it stays because the next change needs it too. ### Passage (j) — the squash carry (Task 0 check only) @@ -257,20 +245,30 @@ fi cat .context/loop-rule-base ``` -**Never overwrite an existing base.** Re-running Task 0 after a partial implementation would record -the current WIP tip, and both Gate B's range and the final `reset --soft` would then start after -every edit made so far — prompt and hook changes would be squashed into the closing commit without -ever entering a review range. **If the recorded base is wrong, delete the file deliberately and say -why**; do not let a re-run decide it. +**Never overwrite an existing base, and never trust one you did not just write.** Re-running Task 0 +after a partial implementation would record the current WIP tip, and both Gate B's range and the +final `reset --soft` would then start after every edit made so far — prompt and hook changes would be +squashed into the closing commit without ever entering a review range. + +**A pre-existing value is not accepted on being non-empty.** It is valid only if it is an ancestor of +`HEAD` **and** every commit between it and `HEAD` is a `WIP:` commit of this execution. Check that +before proceeding: + +```bash +git log --oneline "$(cat .context/loop-rule-base)"..HEAD +``` + +Expected: nothing, or only `WIP:` commits of this run. **Anything else means the file is stale** — +left by an abandoned run, or by one whose work was already squashed. Delete it deliberately, record +why, and re-record from the true starting commit. **Task 15 step 8 removes the file after the +closing commit**, so a stale one is an abandoned run rather than a normal state. **Persist it to a file, not to a shell variable.** Each task runs in its own shell invocation, so a `BASE=` assignment in Task 0 is gone by Task 1 and every parent-tree count would run against an empty revision — which fails loudly in `git show` but quietly in a `grep -c` pipeline. Every later task begins with: -```bash -. .context/loop-rule-verify.sh -``` +*(Build the pair per the verification procedure; record the four values.)* **Task 0 also writes `.context/loop-rule-verify.sh`**, whose generator is given in the fragment-table section. It sets `$BASE` and defines `pair()`, so both arrive together and neither can be used @@ -324,9 +322,21 @@ span**, which the first draft's version did not: writing it down makes it sound`. All twenty-four kept conditions lie inside them. **Record the spans as one `startendfile` line each in `.context/loop-rule-untouched`**, the -anchors being literal strings. Tab-separated because the anchors contain colons — the first draft -used `:` as the delimiter, and `**Severity:**:**Tool routing:` splits at the wrong colon and yields -an empty end anchor. +anchors being literal strings. Tab-separated because the anchors contain colons — a `:` delimiter +splits `**Severity:**:**Tool routing:` at the wrong colon and yields an empty end anchor. + +**A line-span check cannot isolate every kept condition, and two of them prove it.** `a14` begins on +the same `CLAUDE.md` line as text `a13` changes, and `h6` shares a line with a changed neighbour; no +whole-line span contains one without the other. **For any kept condition that shares a line with a +changed one, check it by its own fragment instead** — count the condition's text before and after, +expecting `1` both times — and record which conditions are checked that way. **The span list is +therefore spans plus a short per-condition list**, and the two together must cover every kept +condition. + +**A single-line site cannot be expressed as `sed -n "/x/,/x/p"`.** `sed` does not test the end +address on the line that matched the start, so the range runs to the next match or to end of file. +**Give a single-line site a `grep -n` check rather than a `sed` range** — the squash-carry sentence +(`j1`–`j4`) is the one site of this shape. - [ ] **Step 3: Confirm the parity baseline of the inventoried ranges** @@ -415,17 +425,7 @@ Insert the three blocks before the anchor line, blank-line separated, byte-ident copies while every count and the parity diff still pass — the paragraphs are installed together and nothing else observes them. -```bash -. .context/loop-rule-verify.sh -for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do - for frag in 'How a cycle ends — one ordering' \ - '' \ - ''; do - printf '%s | worktree=%s parent=%s | %s\n' "$f" \ - "$(grep -cF "$frag" "$f")" "$(git show "$BASE:$f" | grep -cF "$frag")" "$frag" - done -done -``` +*(Build the pair per the verification procedure; record the four values.)* Expected: `worktree=1 parent=0` for all six. **Add the two §A2/§A3 fragments to the fragment table once chosen**, with the check that verified each is single-line and unique — they are the two rows @@ -469,19 +469,7 @@ This task exists because `f1` and the (d)/(j) dispositions are falsifiable only - [ ] **Step 1: Diff each untouched passage against the parent** -```bash -. .context/loop-rule-verify.sh -# read the five recorded ranges rather than hard-coding three of them -for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do - for anchor in 'From pass 4 onward every pass report carries three lines' \ - 'The two rules above do not compete' \ - 'On squash-merge, copy every evidence entry'; do - printf '%s | %s: ' "$f" "$anchor" - diff <(git show "$BASE:$f" | grep -A12 -F "$anchor") <(grep -A12 -F "$anchor" "$f") >/dev/null \ - && echo unchanged || echo CHANGED - done -done -``` +*(Build the pair per the verification procedure; record the four values.)* Expected: `unchanged` six times. Any `CHANGED` is a defect — revert that hunk before continuing. @@ -517,12 +505,7 @@ Rows **P2** (`b7`, the fix-set definition) and **P3** (`b12`, immediate resumpti table. Both are verified single-line and unique; re-confirm before editing, since earlier tasks have touched these files: -```bash -. .context/loop-rule-verify.sh -for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do - printf '%s P2=%s P3=%s\n' "$f" "$(grep -cF "$P2_OLD" "$f")" "$(grep -cF "$P3_OLD" "$f")" -done -``` +*(Build the pair per the verification procedure; record the four values.)* Expected: `P2=1 P3=1` for **both** files — the first draft ran these against C only while claiming a result for both. @@ -538,13 +521,7 @@ Replace from `**What a loop absorbs, and what stops it` through the sentence §B `b7` and `b12` are separate meaning changes and each owes its own pair; one pair covering both would let the surviving instruction pass behind the repaired one. -```bash -. .context/loop-rule-verify.sh -for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do - pair "$P2_OLD" 'union of the scope every approved story or plan governing this change assigns to this cycle' "$f" - pair "$P3_OLD" "$P3_NEW" "$f" -done -``` +*(Build the pair per the verification procedure; record the four values.)* **`P3_NEW` is the §B resumption sentence's fragment**, chosen after installing, verified the three ways, and added to the helper before this step runs. @@ -595,12 +572,7 @@ git commit -m "WIP: replace passage (b) with the absorb paragraph that owns the - [ ] **Step 1: Record the old-wording fragments** -```bash -. .context/loop-rule-verify.sh -for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do - printf '%s P4=%s\n' "$f" "$(grep -cF "$P4_OLD" "$f")" -done -``` +*(Build the pair per the verification procedure; record the four values.)* Expected: `P4=1` for both files. @@ -629,14 +601,7 @@ fragment sits wholly on one line, and confirm it is unique before counting. NEW from `c8` — two different changes — so either could land while the other survived and it still reported a pass, and `c4` had no observation at all. -```bash -. .context/loop-rule-verify.sh -for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do - pair "$P19_OLD" '<§C c4 NEW — derive, verify, add as a row>' "$f" # c4 replaced - pair "$P20_OLD" 'a recurrence failing them being an ordinary fresh finding' "$f" # c8 widened - pair "$P4_OLD" '<§C c14 routing NEW — derive, verify, add as a row>' "$f" # c14 removed -done -``` +*(Build the pair per the verification procedure; record the four values.)* Expected for all six: `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`. @@ -679,11 +644,7 @@ deliberate a divergence this task chose to create. replacement**, so it can never be the old half — that is the trap design §7 records from four spec revisions. Rows **P5** (C) and **P5w** (W) instead, and they differ because W drops the pronoun: -```bash -. .context/loop-rule-verify.sh -printf 'P5 C=%s\n' "$(grep -cF "$P5_OLD" CLAUDE.md)" -printf 'P5w W=%s\n' "$(grep -cF "$P5w_OLD" plugins/dev-workflow/commands/workflow-init.md)" -``` +*(Build the pair per the verification procedure; record the four values.)* Expected: `1` each. @@ -701,17 +662,7 @@ is add-only** — it replaces no wording — so it is checked by presence alone, it can be omitted from both copies while the `e7` pair, the condition walk and the parity diff all pass. -```bash -. .context/loop-rule-verify.sh -NEW='read **after** the clean-completion branch of the closure ordering' -PTR='where this stop'"'"'s place among the suspensions' -pair "$P5_OLD" "$NEW" CLAUDE.md -pair "$P5w_OLD" "$NEW" plugins/dev-workflow/commands/workflow-init.md -for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do - printf 'pointer %s worktree=%s parent=%s\n' "$f" \ - "$(grep -cF "$PTR" "$f")" "$(git show "$BASE:$f" | grep -cF "$PTR")" -done -``` +*(Build the pair per the verification procedure; record the four values.)* Expected: both pairs `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`; both pointer checks `worktree=1 parent=0`. @@ -763,16 +714,17 @@ Expected: C `g1/g2: 1 g4: 1`; W `g1/g2: 1 g4: 0`. pair, the old unscoped resolve duty can survive, or the new repair-versus-dismissal distinction be omitted, while everything Task 6 checks passes. -```bash -. .context/loop-rule-verify.sh -for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do - pair "$P6_OLD" 'The demotion changes what a cycle must resolve, never what it observes' "$f" - pair "$P6r_OLD" 'for every finding in the assigned fix set as the absorb paragraph computes it' "$f" -done -``` +*(Build the pair per the verification procedure; record the four values.)* + +**Three observations here, not two.** `P6` observes `g1`'s replacement and the resolve-duty row +observes the bullet, but **`g2` and `g3` are dropped rather than replaced** — the interim +report-and-stop duty and its rationale go, and nothing in §E takes their place. A dropped condition +owes an **absence check**, not a pair: count its text before and after, expecting `1` then `0`. +Without it the duty can survive beside the answer that makes it obsolete, which is the +two-instructions-that-disagree failure in its purest form. -**`P6r_OLD` is the resolve-duty row step 1 derives.** Add it to the table **and** to -`.context/loop-rule-verify.sh`, then re-run the validation loop before this step. +**Derive the resolve-duty OLD row and the `g2`/`g3` absence fragment**, check each the three ways, +and add both to the table before running the counts. Expected for all four: `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`. @@ -826,16 +778,7 @@ contract`), `Your final pass must be clean` (wraps C 132–133), ``a literal `NO is clean`` (one line in C, wraps in W), `A clean pass is the single body line` (the line break falls after `A`), and `at minimum floor 3, severity classified without the demotion` (wraps C 155–156). -```bash -. .context/loop-rule-verify.sh -for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do - echo "== $f" - for id in P7 P8 P9 P10 P11 P12 P13 P14 P15 P16; do - eval "o=\$${id}_OLD" - printf '%-4s %s\n' "$id" "$(grep -cF "$o" "$f")" - done -done -``` +*(Build the pair per the verification procedure; record the four values.)* Expected: `1` twenty times. **Any `0` means the wording drifted since this table was verified at `5871d0a`** — re-derive that fragment and update the table before installing anything. @@ -861,15 +804,7 @@ source sentence per block: | P15 | lens unchanged-list | `**the lens sets** leave every other` | | P16 | strict-reading list | `every suspension binding, since` | -```bash -. .context/loop-rule-verify.sh -for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do - for id in P7 P8 P9 P10 P11 P12 P13 P14 P15 P16; do - eval "o=\$${id}_OLD"; eval "n=\$${id}_NEW" - pair "$o" "$n" "$f" - done -done -``` +*(Build the pair per the verification procedure; record the four values.)* **`_NEW` is set by the executor after installing that block and verifying the chosen fragment the same three ways** — append the assignment to `.context/loop-rule-verify.sh` and re-run the @@ -941,32 +876,13 @@ For each item, the OLD fragment must satisfy all three: single-line in both copi per copy where the wrapping differs, as `e7` needed); unique in each; and **not preserved inside its own replacement** — compare it against the §F block before accepting it. -```bash -. .context/loop-rule-verify.sh -for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do - for id in F1 F2 F3 F4 F5 F6 F7 F8 F9 F10 F11 F12 F13 F14; do - eval "o=\$${id}_OLD"; eval "n=\$${id}_NEW" - pair "$o" "$n" "$f" - done -done -``` +*(Build the pair per the verification procedure; record the four values.)* Expected for all twenty-eight: `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`. - [ ] **Step 4: Count what was installed** -```bash -# Each of step 3's fourteen pairs printed new/worktree; total them per copy. -. .context/loop-rule-verify.sh -for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do - n=0 - for new in '' '' '' '' '' '' '' \ - '' '' '' '' '' '' ''; do - n=$(( n + $(grep -cF "$new" "$f") )) - done - echo "$f installed=$n" -done -``` +*(Build the pair per the verification procedure; record the four values.)* Expected: `installed=14` per copy. State the number you observed. **Do not carry a count from §F into a check** — §F states the count of falsified sentences and this task installs a subset of them; a count copied between the two is the stale-bookkeeping defect the design records at five passes running. @@ -1007,13 +923,7 @@ Expected: one hit each per file. **Item 14's sentence wraps after "here"** — Rows **P17** and **P18**. Assigning the variables is not running the check — the first draft stopped at the assignment. -```bash -. .context/loop-rule-verify.sh -for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do - pair "$P17_OLD" 'not a blanket exemption for hook text' "$f" - pair "$P18_OLD" 'Gate A (spec) → Gate-A closing act' "$f" -done -``` +*(Build the pair per the verification procedure; record the four values.)* Expected for all four: `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`. @@ -1100,24 +1010,14 @@ first time — it admitted exactly the shell keywords it existed to catch. The hook has one copy and owes no parity check; **the hook suite's exact-match assertion is this edit's second observation** (design §7). Run the pair anyway: -```bash -. .context/loop-rule-verify.sh -H=plugins/dev-workflow/hooks/codex-gate.sh -# Tab-separated OLDNEW: the fragments contain no tab, and | appears -# inside shell text. Six entries, all in the loop — an earlier draft left -# two of them as bare quoted strings after it, which a shell tries to run. -printf '%b\n' \ - 'this floor is the only thing keeping the spec review honest\tinstruction-backed' \ - 'floor met by COUNT ONLY\tProceed only once this Gate-A cycle has closed' \ - 'commit only if your final pass was clean — no new Blocker/Major\tevery other closure condition holds' \ - 'then make the real commit when your final pass is clean\tUse this commit as the review range' \ - 'or proceed only if $policy\tskip rule decides only whether a cycle runs at all' \ - 'STOP — Codex Gate B not satisfied\tCodex gate state:' \ - > .context/loop-rule-hookpairs -while IFS=$(printf '\t') read -r OLD NEW; do - pair "$OLD" "$NEW" "$H" -done < .context/loop-rule-hookpairs -``` +*(Build the pair per the verification procedure; record the four values.)* + +**Three of these pairs need a second look before they are run, and the executor derives the +fragments with the file open.** Item 11's removed text is the **clean definition** at the end of the +Gate-A satisfied message, not `floor met by COUNT ONLY`, which the change leaves standing — an OLD +half taken from unchanged text can never reach zero. Item 10 removes **two** things, the honesty +claim and the tail `Run more passes before executing`, and owes an observation for each. Item 17 +likewise removes both the skip-rule offer and `run more`. **Exact counts per pair, not a floor.** Five of these replace one message each and must read exactly `old/parent=1 new/worktree=1`; only the grouped `STOP` pair, which covers items 15 and 16, @@ -1169,9 +1069,13 @@ git commit -m "WIP: replace the seven gate reminders the ordering falsifies" ```bash grep -n 'expected_ctx=\|expected_msg=' plugins/dev-workflow/hooks/codex-gate.test.sh -grep -n 'Gate B satisfied\|Gate B not satisfied\|STOP\|no recorded review\|cannot confirm review\|only thing keeping\|no new Blocker/Major' plugins/dev-workflow/hooks/codex-gate.test.sh +grep -n 'Gate B satisfied\|Gate B not satisfied\|Gate A satisfied\|Gate A floor\|STOP\|no recorded review\|cannot confirm review\|only thing keeping\|no new Blocker/Major\|count only\|COUNT ONLY' plugins/dev-workflow/hooks/codex-gate.test.sh ``` +**`Gate A` as well as `Gate B`.** An earlier draft searched only the Gate-B phrases while step 3 +requires every label naming a gate verdict to move; the Gate-A assertions and labels around lines +504–506 would have been missed by the locator and then demanded by the rule. + **Record the counts you observe.** They will not match the numbers above if the file has changed; the numbers above are evidence for why no list is kept, not a target. - [ ] **Step 2: Update the three `expected_ctx` and three `expected_msg` assignments** @@ -1194,7 +1098,7 @@ Expected: exit 0 from both. **`HOOK_SH` selects the shell the hook runs under; w - [ ] **Step 5: Confirm no verdict vocabulary survives** ```bash -grep -c 'Gate B satisfied\|Gate B not satisfied' plugins/dev-workflow/hooks/codex-gate.test.sh +grep -c 'Gate B satisfied\|Gate B not satisfied\|Gate A satisfied' plugins/dev-workflow/hooks/codex-gate.test.sh ``` Expected: **`0`. Not "no hits you cannot justify"** — step 3 requires every label and comment to move @@ -1490,7 +1394,7 @@ claude plugin validate . --strict Expected: exit 0. -- [ ] **Step 4b: Apply all twelve `docs/prompt-standards.md` items to the installed text** +- [ ] **Step 4b: Apply all twelve `docs/prompt-standards.md` items to the installed text, and commit the result before Gate B** **Nothing mechanical does this and no other task claims it.** The battery's three narrow checks are a floor — one `Target model:` spelling, one prose count claim, one severity vocabulary — and @@ -1502,6 +1406,16 @@ risk. **A reader check, deliberately** — no pattern decides whether a constraint carries its reason. +**Commit the result before step 6.** Gate B reviews the range `$BASE..HEAD`; an edit to this plan +left in the worktree is in neither that range nor the final `reset --soft`, which stages only what +the discarded commits contained. The same applies to every record this plan collects — the sweep, the +next-state table, the divergence list: + +```bash +git add docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +git commit -m "WIP: plan records" +``` + **`check-invariants.sh` includes the prompt-conformance checks** — a `Target model:` line naming one recognized model, a prose checklist-count claim matching the checklist, and the finding-severity vocabulary as a closed set in both prompt copies. Those three are a floor, not coverage; invariant 11's other eleven items are judged by a reader. - [ ] **Step 5: Write the evidence entry into `.context/loop-rule-closing-msg`** @@ -1557,19 +1471,7 @@ closing commit.** so an evidence entry written only into a WIP message is destroyed at exactly the moment the cycle closes — which is what step 5 would otherwise have done. -```bash -. .context/loop-rule-verify.sh -# Everything that must be IN the squashed commit has to be committed before the -# reset: reset --soft stages only what the discarded commits already contained. -git add docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md \ - .context/codex-reviews/ -git status --porcelain # expect empty after this -git commit -m "WIP: plan records and Gate-B findings files" || true -# .context/loop-rule-closing-msg holds the revalidated evidence entry, the provenance -# line, the per-pass curve and any human-exception record, completed at step 7. -git reset --soft "$BASE" -git commit -F .context/loop-rule-closing-msg -``` +*(Build the pair per the verification procedure; record the four values.)* **`reset --soft` stages committed content only.** The prompt-standards result (step 4b), the completeness sweep, the next-state table and the divergence list all land in this plan, and From 33cdfa860032029c237fd202bc7e16d5dcefe09b Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 14:59:10 +0200 Subject: [PATCH 104/181] docs(plans): complete the pass-5 method change through its dependencies MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit An independent review found the pass-5 repair incomplete and partly damaging. Verified; none of it was a new requirement. My bulk edit matched the helper's filename as well as pair commands, so it deleted four blocks it should not have touched: Task 2's untouched-range diff, Task 8's installed-count step, Task 10's hook-pair list, and Task 15's git closing block — leaving "Build the pair … record the four values" where git reset --soft and git commit -F had been. A fragment count does not close a cycle. All four restored. Three references to the deleted helper survived: Task 0 still claimed to write .context/loop-rule-verify.sh, and Tasks 4 and 7 still told the executor to add lines to it and re-run a validation loop that no longer exists. One shell pattern pass 5 had rejected was still standing in Task 14 step 4b — the sed range whose behaviour with identical start and end anchors pass 5 named. Replaced with the anchor-resolved form plus per-condition checks for a14 and h6. Two claims in the pass-5 record were too broad and are corrected: not every finding in five passes was an apparatus defect — pass 1 found a baseSha that would have reviewed only the version bump, and two tasks installing text contrary to the approved target. And design §7 requires checks against the real files; it does not forbid pre-written shell. The new method is a defensible implementation of that requirement, not one the design had prescribed. Status: the verification strategy is simplified; whether it is consistently applied and effective is still being reviewed. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-plan-om0bdd7udh-resume.md | 52 ++++++++++-- .../2026-09-14-loop-rule-consolidation.md | 84 ++++++++++++------- 2 files changed, 100 insertions(+), 36 deletions(-) diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md index 0031178..0fb8d85 100644 --- a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md @@ -211,15 +211,22 @@ told me to derive an OLD row per addition in the strict-reading block; pass 5 fi clauses are add-only and have no old wording to remove. **Two tells, mandatory stop**, surfaced, loop continued on the standing answer. -**What five passes have actually been saying.** Findings 1, 3, 4, 9, 11, 13 and 16 of this pass, and -most of passes 2, 3 and 4, are one thing: **shell written in advance for text that does not exist +**What most of five passes have been saying.** Findings 1, 3, 4, 9, 11, 13 and 16 of this pass, and +much of passes 2, 3 and 4, are one thing: **shell written in advance for text that does not exist yet.** Helpers defined in one shell and called in another; loop bodies outside their loops; `sed` ranges whose delimiters occur in their own data; `sed -n "/x/,/x/p"` for a single line, which runs to the next match instead; variables no step sets; a generated helper whose empty variable makes -`grep -cF ""` match every line. **Every one was a defect in the apparatus and none in the change.** - -**Design §7 already said not to do this:** the plan *builds each pair against the real files and runs -both directions there*. **There** — with the installed text open. I had been pre-writing it. +`grep -cF ""` match every line. **Those were defects in the apparatus.** **It would be wrong to say none of the five passes found a +defect in the change itself** — pass 1 found a `baseSha` that would have put only the version bump in +Gate B's range, and two tasks installing text contrary to the approved target. Both would have +produced wrong behaviour, and neither is an apparatus problem. The apparatus class is the *majority* +of the findings, not all of them. + +**Design §7 requires the checks to run against the real files** — the plan *builds each pair against +the real files and runs both directions there*. **It does not forbid pre-written shell.** Writing the +commands at execution instead is a defensible way to meet that requirement and it removes a class of +error this cycle kept producing; it is not a method the design had already prescribed, and claiming +so would be reading a preference back into the text. **So the pre-written blocks are gone**, replaced by one procedure stated once: take the OLD from its row, install, choose a NEW from the installed text, check it the three ways, count four values, @@ -238,3 +245,36 @@ abandoned run's value would silently put every edit outside Gate B's range; it i ancestor of `HEAD` with only this run's `WIP:` commits between, and removed at close. And the prompt-standards result was written into the plan after the last WIP commit, where neither Gate B's range nor `reset --soft` would reach it. + + +## Pass-5 follow-up — the method change, traced through its dependencies + +An independent reviewer Daniel obtained found the pass-5 repair **incomplete, and partly damaging**. +Verified and corrected at the commit below; none of it was a new requirement. + +**My bulk edit destroyed four blocks it should not have touched.** The script that stripped +pre-written pair commands matched on the helper's filename as well, so it deleted **Task 2's +untouched-range diff**, **Task 8's installed-count step**, **Task 10's hook-pair list** and — worst — +**Task 15's git closing block**, leaving "Build the pair … record the four values" where +`git reset --soft` and `git commit -F` had been. A fragment count does not close a cycle. All four are +restored. + +**Three references to the deleted helper survived**: Task 0 still claimed to write +`.context/loop-rule-verify.sh`, and Tasks 4 and 7 still told the executor to add lines to it and +re-run a validation loop that no longer exists. Removed. + +**One shell pattern pass 5 had rejected was still standing** in Task 14 step 4b — the `sed` range +whose behaviour with identical start and end anchors pass 5 named. Replaced with the anchor-resolved +form plus the per-condition checks for `a14` and `h6`. + +**Two claims in the pass-5 report were too broad and are corrected above.** Not every finding in five +passes was an apparatus defect — pass 1 found a `baseSha` that would have reviewed only the version +bump, and two tasks installing text contrary to the approved target, both of which would have +produced wrong behaviour. And design §7 requires checks against the real files; it does not forbid +pre-written shell. The new method is a defensible implementation of that requirement, not one the +design had already prescribed. + +**The honest status is narrower than "cause understood and removed":** the verification strategy is +simplified, and **whether it is consistently applied and effective is still being reviewed.** The +change also moves responsibility for working check commands to execution time, which removes some +error sources and proves nothing about the checks being right. diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index cd0288a..23ba427 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -266,13 +266,15 @@ closing commit**, so a stale one is an abandoned run rather than a normal state. **Persist it to a file, not to a shell variable.** Each task runs in its own shell invocation, so a `BASE=` assignment in Task 0 is gone by Task 1 and every parent-tree count would run against an empty revision — which fails loudly in `git show` but quietly in a `grep -c` pipeline. Every later -task begins with: +task that counts anything begins by reading it back and refusing an empty value: -*(Build the pair per the verification procedure; record the four values.)* +```bash +BASE=$(cat .context/loop-rule-base) +test -n "$BASE" || { echo "BASE empty — Task 0 did not run"; exit 1; } +``` -**Task 0 also writes `.context/loop-rule-verify.sh`**, whose generator is given in the fragment-table -section. It sets `$BASE` and defines `pair()`, so both arrive together and neither can be used -without the other. +**The guard is not decoration.** With `$BASE` empty, `git show ":$f"` reads the *index* rather than +the recorded parent, and every parent count then describes the wrong tree without erroring. **This commit is also `baseSha` for Gate B.** It is the parent of the first WIP snapshot, and it is the only value that puts the whole implementation inside the reviewed range. `.context/` is ignored @@ -425,7 +427,7 @@ Insert the three blocks before the anchor line, blank-line separated, byte-ident copies while every count and the parity diff still pass — the paragraphs are installed together and nothing else observes them. -*(Build the pair per the verification procedure; record the four values.)* +*(Count each fragment in both copies and both trees, per the verification procedure; expect worktree 1, parent 0.)* Expected: `worktree=1 parent=0` for all six. **Add the two §A2/§A3 fragments to the fragment table once chosen**, with the check that verified each is single-line and unique — they are the two rows @@ -469,9 +471,13 @@ This task exists because `f1` and the (d)/(j) dispositions are falsifiable only - [ ] **Step 1: Diff each untouched passage against the parent** -*(Build the pair per the verification procedure; record the four values.)* +Read `.context/loop-rule-untouched`, and for each recorded span **resolve its start and end anchors +separately in each tree** — Task 1 inserts a large block, so a line number taken from one tree +addresses different text in the other. For each span, extract it from `$BASE:` and from the +worktree file and diff the two. -Expected: `unchanged` six times. Any `CHANGED` is a defect — revert that hunk before continuing. +Expected: no output for every recorded span. Any difference is a defect — revert that hunk before +continuing. - [ ] **Step 2: Confirm `f1` still names the right pair** @@ -505,7 +511,7 @@ Rows **P2** (`b7`, the fix-set definition) and **P3** (`b12`, immediate resumpti table. Both are verified single-line and unique; re-confirm before editing, since earlier tasks have touched these files: -*(Build the pair per the verification procedure; record the four values.)* +*(Count each named OLD fragment in the copies its row claims, per the verification procedure; expect 1 in each.)* Expected: `P2=1 P3=1` for **both** files — the first draft ran these against C only while claiming a result for both. @@ -572,7 +578,7 @@ git commit -m "WIP: replace passage (b) with the absorb paragraph that owns the - [ ] **Step 1: Record the old-wording fragments** -*(Build the pair per the verification procedure; record the four values.)* +*(Count each named OLD fragment in the copies its row claims, per the verification procedure; expect 1 in each.)* Expected: `P4=1` for both files. @@ -606,9 +612,8 @@ reported a pass, and `c4` had no observation at all. Expected for all six: `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`. **`P19_OLD` and `P20_OLD` do not exist yet.** Derive each from the live passage, check it the three -ways, **add a row to the fragment table and a line to `.context/loop-rule-verify.sh`**, then re-run -the validation loop. Ids continue the `P` series; a row id is whatever the table and the helper agree -on, so keep them plain. +ways, and **add a row to the fragment table** — that table is where every fragment lives, and a row +that is not there is a fragment stated somewhere else. Ids continue the `P` series. - [ ] **Step 5: Walk the conditions** — `c1`–`c3` present unchanged, `c5`–`c7` word for word, `c8` carrying the new clause, `c9` split, `c10`–`c14` gone from here. @@ -644,7 +649,7 @@ deliberate a divergence this task chose to create. replacement**, so it can never be the old half — that is the trap design §7 records from four spec revisions. Rows **P5** (C) and **P5w** (W) instead, and they differ because W drops the pronoun: -*(Build the pair per the verification procedure; record the four values.)* +*(Count each named OLD fragment in the copies its row claims, per the verification procedure; expect 1 in each.)* Expected: `1` each. @@ -778,7 +783,7 @@ contract`), `Your final pass must be clean` (wraps C 132–133), ``a literal `NO is clean`` (one line in C, wraps in W), `A clean pass is the single body line` (the line break falls after `A`), and `at minimum floor 3, severity classified without the demotion` (wraps C 155–156). -*(Build the pair per the verification procedure; record the four values.)* +*(Count each named OLD fragment in the copies its row claims, per the verification procedure; expect 1 in each.)* Expected: `1` twenty times. **Any `0` means the wording drifted since this table was verified at `5871d0a`** — re-derive that fragment and update the table before installing anything. @@ -807,8 +812,8 @@ source sentence per block: *(Build the pair per the verification procedure; record the four values.)* **`_NEW` is set by the executor after installing that block and verifying the chosen fragment -the same three ways** — append the assignment to `.context/loop-rule-verify.sh` and re-run the -validation loop before the pairs. +the same three ways** before it is counted with. A NEW fragment is not added to the table: the table +holds OLD fragments, which exist before the edit and can be checked in advance. Expected for all twenty: `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`. @@ -882,9 +887,9 @@ Expected for all twenty-eight: `old/worktree=0 old/parent=1 new/worktree=1 new/p - [ ] **Step 4: Count what was installed** -*(Build the pair per the verification procedure; record the four values.)* +Total the fourteen `new/worktree` values step 3 printed, per copy. -Expected: `installed=14` per copy. State the number you observed. **Do not carry a count from §F into a check** — §F states the count of falsified sentences and this task installs a subset of them; a count copied between the two is the stale-bookkeeping defect the design records at five passes running. +Expected: `14` per copy. State the number you observed. **Do not carry a count from §F into a check** — §F states the count of falsified sentences and this task installs a subset of them; a count copied between the two is the stale-bookkeeping defect the design records at five passes running. - [ ] **Step 5: Parity** for all fourteen sites. @@ -1008,9 +1013,10 @@ first time — it admitted exactly the shell keywords it existed to catch. - [ ] **Step 5: Discriminating pairs, worktree and parent** -The hook has one copy and owes no parity check; **the hook suite's exact-match assertion is this edit's second observation** (design §7). Run the pair anyway: - -*(Build the pair per the verification procedure; record the four values.)* +The hook has one copy and owes no parity check; **the hook suite's exact-match assertion is this +edit's second observation** (design §7). Run a pair anyway, one per replaced message, against +`plugins/dev-workflow/hooks/codex-gate.sh` — **seven items, and items 15 and 16 share one `STOP` +opening, so eight observations across seven pairs.** **Three of these pairs need a second look before they are run, and the executor derives the fragments with the file open.** Item 11's removed text is the **clean definition** at the end of the @@ -1326,13 +1332,17 @@ against the parent. **A later task can alter an untouched condition identically every other check still passes** — the parity diff sees no difference and no pair covers text no task claims to change. -```bash -BASE=$(cat .context/loop-rule-base) -while IFS=$(printf '\t') read -r s e f; do - echo "== $f :: $s" - diff <(git show "$BASE:$f" | sed -n "/$s/,/$e/p") <(sed -n "/$s/,/$e/p" "$f") -done < .context/loop-rule-untouched -``` +For each span in `.context/loop-rule-untouched`, resolve its start and end anchors **separately in +the parent and in the worktree**, extract the two slices, and diff them. + +**Not a `sed` range with the same anchor at both ends.** `sed` does not test the end address on the +line that matched the start, so a single-line site runs to the next match or to end of file — which +is why the squash-carry sentence is checked with `grep -n` instead. And not line numbers carried +from Task 0: Task 1's insertion shifts everything after it in one tree and not the other. + +**Plus the per-condition checks** for `a14`, `h6` and any other kept condition sharing a line with a +changed one: count the condition's own text in the parent and in the worktree, expecting `1` both +times. **Anchors, resolved separately in each tree** — Task 1 inserts a large block, so a line number taken from either tree addresses different text in the other, and an earlier draft's numeric `sed` @@ -1471,7 +1481,21 @@ closing commit.** so an evidence entry written only into a WIP message is destroyed at exactly the moment the cycle closes — which is what step 5 would otherwise have done. -*(Build the pair per the verification procedure; record the four values.)* +```bash +BASE=$(cat .context/loop-rule-base) +test -n "$BASE" || { echo "BASE empty — Task 0 did not run"; exit 1; } +# Everything that must be IN the squashed commit has to be committed before the +# reset: reset --soft stages only what the discarded commits already contained. +git add docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md .context/codex-reviews/ +git commit -m "WIP: plan records and Gate-B findings files" || true +git status --porcelain # expect empty +git reset --soft "$BASE" +git commit -F .context/loop-rule-closing-msg +rm -f .context/loop-rule-base +``` + +**`$BASE` is the recorded revision, not a placeholder to substitute by hand**, and the file is +removed afterwards so a later run cannot inherit a stale one. **`reset --soft` stages committed content only.** The prompt-standards result (step 4b), the completeness sweep, the next-state table and the divergence list all land in this plan, and From c21b45a630fb01477af4598cdc5e1b5258185c7d Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 15:07:49 +0200 Subject: [PATCH 105/181] docs(context): durable Gate-A plan prompt, and a resume block in the working record The plan-review prompt had been retyped inline each pass and would have died with the session. It is a file now, with the pass history and the settled/ collected blocks that keep a reviewer off answered ground. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-plan-om0bdd7udh-resume.md | 20 ++++++++++++++++++- 1 file changed, 19 insertions(+), 1 deletion(-) diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md index 0fb8d85..0582f93 100644 --- a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md @@ -11,6 +11,24 @@ Nothing depends on it; the pass files and the repo are authoritative where this - **Derived floor:** 3 (risk high → level 2; security none → 0; max 2 ≠ 0 → 3) - **Hook knob:** absent (no floor-knob file in `.context/`) +## Resume here + +**Next action:** Gate-A plan pass 6 against `33cdfa8`. The prompt is +`.context/gate-a-plan-prompt.md` — **durable and current**; substitute `__SHA__` and `__P__`, +precheck per its header, delete the target file, confirm it is gone. + +**After a clean close:** implement the plan — both prompt copies, the seven hook strings and their +test expectations, version bump 0.11.0 → 0.12.0 + CHANGELOG, the quality battery, the evidence +entry, then Gate B with a **fresh nonce** (this cycle's is Gate-A plan only). + +**Standing instructions from Daniel:** no stops unless an absolute block; a mandatory two-tell stop +is surfaced in the record and the loop continues on that standing answer. Escalate only new +behaviour decisions and real obstacles. Reports: result, verification, decision-relevant obstacles. + +**Two stories still await a profile confirmation:** +`docs/superpowers/stories/2026-09-10-record-durability-story.md` and +`docs/superpowers/stories/2026-09-10-harness-finding-termination-story.md`. + ## What the spec cycle learned, carried here so this loop does not relearn it The Gate-A **spec** cycle ran 65 passes. Four mechanisms produced almost every finding; each is @@ -35,7 +53,7 @@ worth checking before a pass rather than after. | 3 | b767a2a | 32→**31** | 0→**0** | 25→**23** | yes | one tell (instrument cluster). **Pass 2's three fragments came back — because I repaired the table and left the same fragments quoted inline in the task steps.** Second-copy defect, in the plan written to avoid it. Structural repair: the table is the only authored copy and Task 0 **generates** the shell variables from it, so no step can restate a fragment. 8 Minors collected | | 4 | 6ace06f | 31→**20** | 0→**10** | 23→**7** | yes | **MANDATORY TWO-TELL STOP** — Blockers rose 0→10, and the findings cluster on the instrument for the fourth pass running. Surfaced, standing answer applied, loop continued. **The awk generator pass 3 introduced was broken three ways**; it is deleted, not debugged — the helper is transcribed by hand and validated against the real files, with empty-string guards, because a silently unset variable makes `grep -cF ""` match every line. **A fourth fragment (F4) was preserved in its own replacement and invisible to my checker**, which compared without normalizing the block's line breaks | | 5 | 58b3660 | 20→**16** | 10→**7** | 7→**7** | yes | **MANDATORY TWO-TELL STOP** — instrument cluster for the fifth pass, plus a require↔withdraw pair (pass 4 demanded OLD rows for clauses pass 5 classifies as add-only). Surfaced, standing answer applied, loop continued. **The pre-written shell is deleted.** Design §7 says the plan builds each pair *against the real files*; five passes of findings were blocks written in advance for text that does not exist yet. One stated procedure replaces them | -| 6 | — | — | — | — | not run | next, against the pass-5 repair commit | +| 6 | — | — | — | — | **not run — NEXT** | against `33cdfa8`. Prompt: `.context/gate-a-plan-prompt.md`, substitute `__SHA__`=`33cdfa8` and `__P__`=`6` | ## Pass-1 report From 05ffbd46bc584acd8600963b2aa2f2924ce0f0b8 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 15:07:58 +0200 Subject: [PATCH 106/181] docs(context): note that the plan prompt is untracked and how to rebuild it Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md | 8 +++++++- 1 file changed, 7 insertions(+), 1 deletion(-) diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md index 0582f93..ed33539 100644 --- a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md @@ -14,9 +14,15 @@ Nothing depends on it; the pass files and the repo are authoritative where this ## Resume here **Next action:** Gate-A plan pass 6 against `33cdfa8`. The prompt is -`.context/gate-a-plan-prompt.md` — **durable and current**; substitute `__SHA__` and `__P__`, +`.context/gate-a-plan-prompt.md` — current at that commit; substitute `__SHA__` and `__P__`, precheck per its header, delete the target file, confirm it is gone. +**That prompt file is untracked.** `.gitignore` carries `.context/*` with only `codex-gate.on` and +`codex-reviews/` exempt, so it survives a context clear but not a `.context/` cleanup. **If it is +gone, rebuild it from this record** — the pass history below, the settled-and-not-open blocks and the +collected list are what it carries, and `.context/gate-a-spec-prompt.md` is the same shape for the +spec cycle. + **After a clean close:** implement the plan — both prompt copies, the seven hook strings and their test expectations, version bump 0.11.0 → 0.12.0 + CHANGELOG, the quality battery, the evidence entry, then Gate B with a **fresh nonce** (this cycle's is Gate-A plan only). From 34250beaa56a1677793d74fa2b71b1a57269d1eb Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 16:02:12 +0200 Subject: [PATCH 107/181] docs(plans): apply Gate-A plan pass 6; derive OLD fragments before installing MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Two Blockers and six Majors from pass 6, all validated against the real files. - Task 11's sweep locator matched only `Gate B not satisfied` (3 hits) and missed the 25 assertions greping the bare verdict word, which items 15/16 remove — the suite would have gone red and the battery could not pass. - Task 15 step 5 deferred the provenance line, the curve and the human-exception record to "step 7's last action", which step 7 did not contain. Step 7b is that action. - Passage (h) counted 24 kept conditions less h4 and h19; item 7 also replaces h5, so it is 23 less h4, h5 and h19 — and the untouched middle span opened after F7 (C 988) rather than after F7b (C 989), enclosing a line the change edits. - Five tasks derived OLD fragments after their own install step, where the old wording is gone and neither the three-condition check nor the pre-edit count can run. The rule is stated once in the procedure and each task now derives before installing. - Task 8 rebuilt fourteen rows the table already holds as fifteen, dropping F7b's independent Gate-B destination change. - §G changes two things; P7 observed only the widened membership, not the semantic test it gains. The test is add-only and owes presence. - The strict-reading tail's add-only clauses were told to produce OLD rows they cannot have; they owe presence alone. - Task 10 promised seven pairs and listed six, with item 10's second removal unaccounted; items 15 and 16 no longer share a non-discriminating count of 2. Two Minors and the Nit folded into edits a Major already required: the `clean completion` count with no expected result, Task 0/Task 2 saying "line numbers" where the body requires anchors, and h13 disposed twice. Two dead references to the helper pass 5 deleted are removed. Trend: findings 32, 32, 31, 20, 16, 11. Blockers 1, 0, 0, 10, 7, 2. Majors 28, 25, 23, 7, 7, 6. Cluster: the verification apparatus, sixth pass — one tell, so no mandatory stop. Require<->withdraw: none. Floor 3, from risk high (2) and security none (0), read at this pass from docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-plan-om0bdd7udh-pass-6.md | 12 + .../2026-09-14-loop-rule-consolidation.md | 253 +++++++++++++----- 2 files changed, 193 insertions(+), 72 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-6.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-6.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-6.md new file mode 100644 index 0000000..2f58f21 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-6.md @@ -0,0 +1,12 @@ +BLOCKER | high | Task 11 step 1 | The sweep locator searches `Gate B not satisfied` but not the bare `not satisfied` used by many real loose assertions, for example `grep -q 'not satisfied'` at `codex-gate.test.sh:189`; those assertions name the verdict wording items 15 and 16 remove. | Step 3 can miss them, after which the two-shell suite fails because the new hook messages say `cannot confirm` rather than `not satisfied`, leaving the repo unable to pass its quality battery. | Add a case-sensitive and uppercase-aware search for the bare old verdict wording, then read and update every match by observed hook state without introducing a fixed assertion count. +BLOCKER | high | Task 15 steps 5-8 | Step 5 says step 7's last action appends the final per-pass curve, provenance line and any human-exception record to `.context/loop-rule-closing-msg`, but step 7 contains no such action; it ends with the spec-update rule. | A literal execution reaches `git commit -F` with the draft created before Gate B and can publish a closing commit missing mandatory records, so the Gate-B cycle has not validly closed. | Add an explicit final substep after the clean Gate-B pass that appends all three record classes, validates the complete message, and only then allows step 8. +MAJOR | high | Task 0 step 2 and Task 14 step 4b | Passage (h) is said to have twenty-four kept conditions, computed as twenty-six less only `h4` and `h19`, although the disposition and target also replace `h5`; the live sentence wraps with `h4` on line 988 and `h5` as `restated by the closing amend` on line 989, and row F7b exists specifically for the latter. | The do-not-touch accounting is false and can either place changed `h5` inside an untouched span, making the final parent diff reject a correct item-7 installation, or omit it from both the changed and protected sets. | State twenty-three kept conditions, exclude `h4`, `h5` and `h19`, and define the human-exception spans around both F7 and F7b while retaining the separate `h6` check necessitated by the shared line. +MAJOR | high | Tasks 3 step 3, 4 step 4, 6 step 3, 7 step 3, and 10 step 5 | These tasks defer deriving one or more OLD fragments until after the task's install step has removed the old live wording, contrary to the global procedure's requirement to confirm `old/worktree=1` before editing; examples include Task 4's P19/P20, Task 6's `g2`/`g3` fragment, Task 7's extra `c18`-block rows and every Task 10 hook OLD. | At that state the executor can inspect only the parent or reconstruct text from prose, so current-file drift, a partial rerun, or a wrong locator is no longer caught before replacement and the claimed four-value pair has lost its pre-edit observation. | Move every OLD-fragment derivation, three-condition validation and current-count check ahead of its install step; leave only NEW-fragment selection and final four-value counting after installation. +MAJOR | high | Task 8 steps 3-4 | The verified table already contains fifteen OLD rows for the fourteen prompt-copy items, F1-F14 plus F7b because item 7 changes both `h4` and `h5`, but the task says to build fourteen rows, run fourteen pairs, expect twenty-eight pair instances and total fourteen NEW values. | Following the task either duplicates already-authored rows or drops F7b's independent Gate-B destination change; the latter lets `restated by the closing amend` survive while the stated counts still pass. | Consume the existing rows instead of rebuilding them, run a distinct observation for F7b as well as F7, and make the expected pair-instance and NEW-value accounting distinguish fourteen installed items from fifteen independent changed clauses. +MAJOR | high | Task 7 step 3, §G | P7's NEW is taken only from §G's widened-membership clause, although the task and target separately say §G also gains a semantic membership test a downstream reader can apply; the step then incorrectly says the other eight blocks change one thing each. | Both prompt copies can receive the widened clause while omitting the new semantic test, and P7, parity, the carried-condition checks and the later untouched diffs all still pass. | Add a separate presence observation for a verified single-line fragment from §G's semantic-test text, with worktree 1 and parent 0 in both copies. +MAJOR | high | Task 7 step 3, unknown-start block | The strict-reading tail adds several independent clauses with no predecessor, yet the task tells the executor to derive an OLD row and discriminating pair for each. | An add-only clause cannot produce `old/parent=1`; the executor must pair it with unrelated removed text or omit its check, allowing the repeated-dismissal, parked-state or remaining strict-reading addition to disappear unnoticed. | Keep P16 for the replaced list boundary and use a separate verified presence check for every independent add-only tail clause, as the global add-only procedure requires. +MAJOR | high | Task 10 step 5 | The prose promises eight observations across seven pairs, but the table enumerates only six pairs and collapses items 15 and 16 into one count-of-two row; it also names only the item-10 honesty claim despite the preceding paragraph requiring a separate observation of that message's removed `Run more passes before executing` tail. | A shared count of two is not discriminating per reminder and the closing evidence cannot truthfully claim the design-required pair for each of the seven hook edits; duplicated shared wording in one message can mask a stale second message, while item 10's old imperative is unaccounted for. | Give each hook reminder its own exact pair and add the explicitly required second removal observation for item 10, retaining exact per-message counts rather than one grouped count. +MINOR | high | Task 4 step 3 | The check `grep -cF 'clean completion' CLAUDE.md` names no expected result and is not the uniqueness check described below it: target §A1 contains four occurrences of that phrase, while only `takes precedence over this exit` is intended to occur once. | A correct install produces a surprising multi-hit count that the executor can misread as duplication or ignore, so the step does not supply an actionable oracle for its first command. | Remove the broad count or state its real subject and expected count; use the exact precedence-clause fragment in both copies for the once-only assertion. +MINOR | high | Task 0 step 2 and Task 2 Interfaces | The Task 0 heading says to record current line numbers and Task 2 says it consumes recorded line numbers, while the operative body forbids absolute line numbers and defines `.context/loop-rule-untouched` as anchor triples. | An executor following the task summary can create input Task 2 cannot safely resolve across the §A insertion, recreating the shifted-range false differences the body was meant to eliminate. | Change both summary statements to say anchor ranges and per-condition fragments, matching the defined scratch-file format. +NIT | high | Condition disposition, passage (h) | `h13` is disposed twice: once inside the `h8-h18` kept range and again in its own `h13` kept row. | The table has 136 disposition entries while claiming an all-135 accounting, even though its unique-id coverage is 135; the redundant row makes count audits need an undocumented deduplication rule. | Remove `h13` from the range or convert the later row into an unnumbered explanatory note. +END OF FINDINGS (11 total) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index 23ba427..7048342 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -57,7 +57,7 @@ | P17 | §F item 14, Named residual | `Hook text is out of scope here` | 139 | 346 | | P18 | §F item 18, work-loop line | `execute → tests green → Gate B → commit` | 63 | 262 | -**The fourteen §F prompt-copy items (Task 8), derived from each item's cited lines and checked the same three ways.** Pass 2 found these deferred to the executor as `` placeholders, which put fourteen meaning-changing checks outside Gate A's reach; they are concrete now. The NEW halves stay deferred, for the reason the paragraph below gives. +**The fourteen §F prompt-copy items (Task 8) — fifteen rows, because item 7 changes two clauses on two lines — derived from each item's cited lines and checked the same three ways.** Pass 2 found these deferred to the executor as `` placeholders, which put fourteen meaning-changing checks outside Gate A's reach; they are concrete now. The NEW halves stay deferred, for the reason the paragraph below gives. | Row | §F item | OLD fragment | C | |---|---|---|---| @@ -92,6 +92,7 @@ **For each meaning-changing edit, at the task that installs it:** 1. **Take the OLD fragment from its table row.** Confirm before editing that it counts **1** in each copy the row claims. + **Where the row does not exist yet, derive it here — before the install, never after.** Several tasks below add rows: after the edit the old wording is gone from the worktree, so the three-condition check cannot be run against the file it is about, the pre-edit count of 1 cannot be observed at all, and the executor is left reconstructing a fragment from the parent tree or from prose. A fragment derived that way can no longer catch the thing this step exists to catch — that the file drifted, or that a partial re-run already applied the edit. **Derive, validate the three conditions, count 1, append the row, and only then install.** 2. **Install the block.** 3. **Choose a NEW fragment from the installed text** and check it the same three ways: **single-line** in the file, **unique** there, and **absent from the parent tree**. 4. **Count four values** — OLD and NEW, each in the worktree and in `$BASE` — and record them. @@ -206,12 +207,12 @@ Story acceptance criterion 5 is satisfied here. Ids are `docs/superpowers/specs/ | Condition | Disposition | |---|---| -| h1–h3, h6, h8–h18, h20–h26 | **kept**, untouched. | -| h5 | **replaced** — §F item 7 changes the Gate-B destination from "restated by the closing amend" to "restated by the commit its closing act produces", the amend no longer being the only closing shape. Row F7. | +| h1–h3, h6, h8–h12, h14–h18, h20–h26 | **kept**, untouched. | +| h5 | **replaced** — §F item 7 changes the Gate-B destination from "restated by the closing amend" to "restated by the commit its closing act produces", the amend no longer being the only closing shape. Row **F7b**, which is its own row because item 7 changes `h4` and `h5` on two different lines and one fragment cannot observe both. | | h4 | **replaced** — §F item 7, the human-exception destination: a Gate-A cycle's record goes to the commit its closing act produces, not to "the spec or plan commit". | | h7 | **kept.** | | h19 | **replaced** — §F item 4, the scope sentence: "neither a human's **general** assent nor this record", plus the clause distinguishing the answers a suspension asks for from assent. | -| h13 | **kept.** Design §4 names this sentence as deliberately not edited: this change ships no record for the squash carry to carry. | +| h13 | **kept**, and lifted out of the `h8`–`h18` run above so no id is disposed twice. Design §4 names this sentence as deliberately not edited: this change ships no record for the squash carry to carry. | ### Passage (i) — when these rules bind → target §H (Task 7) @@ -230,7 +231,7 @@ Story acceptance criterion 5 is satisfied here. Ids are `docs/superpowers/specs/ **Files:** none modified. **Interfaces:** -- Produces: `$BASE` (the parent commit every later task's counterfactual half runs against), and a recorded list of the five untouched passage ranges. +- Produces: `$BASE` (the parent commit every later task's counterfactual half runs against), and the five untouched passage ranges recorded as **anchor spans plus a per-condition fragment list** — never as absolute line numbers, for the reason step 2 gives. - [ ] **Step 1: Confirm the approved artifacts and a clean tree** @@ -280,9 +281,9 @@ the recorded parent, and every parent count then describes the wrong tree withou the only value that puts the whole implementation inside the reviewed range. `.context/` is ignored by the hook's fingerprint, so the file itself moves nothing. -- [ ] **Step 2: Re-read the five untouched ranges and record their current line numbers** +- [ ] **Step 2: Re-read the five untouched ranges and record their current anchors** -Passages (d), (f) and (j) are not edited by this change, and passage (a)'s arithmetic and passage (h)'s **twenty-four** kept conditions are untouched — twenty-six inventoried, less `h4` and `h19`, which Task 8 replaces. Record where they are now, because the inventory's numbers cite `7c0d475`: +Passages (d), (f) and (j) are not edited by this change, and passage (a)'s arithmetic and passage (h)'s **twenty-three** kept conditions are untouched — twenty-six inventoried, less `h4`, `h5` and `h19`, which Task 8 replaces. **`h5` is in that list and an earlier draft left it out**, counting twenty-four kept: item 7 changes the Gate-B destination as well as the Gate-A one, which is why row F7b exists at all. Record where they are now, because the inventory's numbers cite `7c0d475`: **Five ranges, each with a start and an end anchor** — the three whole passages, plus the floor arithmetic and the kept human-exception conditions, which later verification consumes and which the @@ -319,9 +320,13 @@ span**, which the first draft's version did not: from `Both gates are a LOOP with a HARD FLOOR` to the line before row F10's fragment; from the line after F10's to the line before row P9's; and from the line after P9's to `Nothing here writes the floor knob`. **The middle span is the one the first draft dropped**, and it holds `a3`–`a12`. -- **The human exception is three spans**, around rows F7 (`h4`) and F4 (`h19`): from `Recording a - human exception` to before F7's line; between F7's and F4's; and from after F4's to `because - writing it down makes it sound`. All twenty-four kept conditions lie inside them. +- **The human exception is three spans**, around rows F7 (`h4`), **F7b (`h5`)** and F4 (`h19`): + from `Recording a human exception` to before F7's line; **from after F7b's line** to before + F4's; and from after F4's to `because writing it down makes it sound`. All twenty-three kept + conditions lie inside them. **F7 and F7b sit on adjacent lines** — C 988 and 989 as of this + writing — **so a middle span opening after F7 rather than after F7b would enclose a line item 7 + changes, and a correct implementation would fail its own untouched check.** Resolve both anchors + and open the middle span after the later of them. **Record the spans as one `startendfile` line each in `.context/loop-rule-untouched`**, the anchors being literal strings. Tab-separated because the anchors contain colons — a `:` delimiter @@ -465,7 +470,7 @@ Named `WIP:` because Task 15 runs Gate B over the whole change and amends once. **Files:** none modified. **Interfaces:** -- Consumes: Task 0's recorded line numbers and `$BASE`. +- Consumes: Task 0's recorded anchor spans and per-condition fragment list, and `$BASE`. This task exists because `f1` and the (d)/(j) dispositions are falsifiable only by a diff, and the cheapest moment to catch an accidental edit is immediately after the insertion that could have caused one. @@ -518,6 +523,17 @@ Expected: `P2=1 P3=1` for **both** files — the first draft ran these against C **The first draft of this plan named `plus repair obligations you already accepted in earlier passes` here, which wraps across C 201–202 and W 408–409 and counts zero in a correct file.** +- [ ] **Step 1b: Derive the rest of §B's OLD rows, still before installing** + +**Two pairs do not cover §B.** Target §B separately changes `b3`, `b8`, `b11`, `b13`, `b16`, +`b17`–`b18` and **adds** the closing-time set-change rule and decision 6's decline semantics. Each +independent replacement owes its own pair, and each add-only rule owes a presence check — a manual +condition walk is a reader's judgement, not the discriminating observation design §7 assigns here. +**Derive an OLD row for each changed condition now**, check each the three ways against the live +passage, confirm it counts 1 in each copy, and append it to the fragment table. Ids continue the +`P` series. **This is step 1b and not part of step 3 because step 2 removes the wording these rows +are taken from** — a row derived afterwards cannot be checked against the text it describes. + - [ ] **Step 2: Install §B's text over the passage in both copies** Replace from `**What a loop absorbs, and what stops it` through the sentence §B ends at, keeping each copy's own closing parenthetical. @@ -529,17 +545,15 @@ would let the surviving instruction pass behind the repaired one. *(Build the pair per the verification procedure; record the four values.)* -**`P3_NEW` is the §B resumption sentence's fragment**, chosen after installing, verified the three -ways, and added to the helper before this step runs. +**`P3_NEW` is the §B resumption sentence's fragment**, chosen after installing and verified the +three ways before this step counts with it. Expected for all four: `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`. -**Two pairs do not cover §B.** Target §B separately changes `b3`, `b8`, `b11`, `b13`, `b16`, -`b17`–`b18` and **adds** the closing-time set-change rule and decision 6's decline semantics. Each -independent replacement owes its own pair, and each add-only rule owes a presence check — a manual -condition walk is a reader's judgement, not the discriminating observation design §7 assigns here. -**Derive an OLD row for each changed condition and a NEW fragment for each added rule, add them to -the fragment table, and run the checks before Step 4.** +**Run a pair for every row step 1b derived**, and a **presence check** for each of §B's two added +rules — the closing-time set-change rule and decision 6's decline semantics — which replace no +wording and so owe `new/worktree=1 new/parent=0` and no OLD half, per the add-only rule. Choose +each NEW fragment from the installed text and verify it the three ways before counting it. - [ ] **Step 4: Walk the carried conditions** @@ -586,16 +600,29 @@ Expected: `P4=1` for both files. is **preserved in target §C's fenced replacement**, so its old-count could never reach zero. Derive `c4`'s OLD from the part of the sentence the replacement removes. +**`P19_OLD` and `P20_OLD` do not exist yet, and they are derived here rather than at step 4.** +Take each from the live passage, check it the three ways, confirm it counts 1 in each copy, and +**add a row to the fragment table** — that table is where every fragment lives, and a row that is +not there is a fragment stated somewhere else. Ids continue the `P` series. **Step 2 removes the +wording they are taken from**, so a derivation after it has no live text to check against and no +pre-edit count to observe. + - [ ] **Step 2: Install §C's block** - [ ] **Step 3: Confirm the moved clause exists in §A and nowhere else** ```bash -grep -cF 'clean completion' CLAUDE.md +grep -cF 'takes precedence over this exit' CLAUDE.md grep -n 'takes precedence over this exit' CLAUDE.md ``` -The precedence clause must appear **once**, inside the ordering. A second occurrence in passage (c) means the sentence was moved whole instead of split. +Expected: **exactly `1`**, and the hit inside the installed ordering. A second occurrence in +passage (c) means the sentence was moved whole instead of split. + +**The once-only assertion is on the precedence clause, not on `clean completion`.** An earlier +draft counted the bare phrase with no stated expectation; target §A1 uses it four times, so a +correct install produces a multi-hit count with nothing to compare it against — a check whose +result a reader can only shrug at. - [ ] **Step 4: Run the discriminating pair, both copies, both trees** @@ -697,9 +724,13 @@ git commit -m "WIP: read the two-tell threshold after the clean-completion branc - [ ] **Step 1: Record the old wording in both copies, and derive the resolve-duty OLD** -§E replaces **two** blocks. The handed-over-question OLD is row P6; the resolve-duty bullet has no -row yet. Derive one from the live `**Severity:**` bullet, check it the three ways, and add it to the -fragment table before Step 3. +§E replaces **two** blocks and **drops** a third thing. The handed-over-question OLD is row P6; +the resolve-duty bullet has no row yet, and neither does `g2`/`g3`'s absence fragment. **Derive +both here, before Step 2 installs over them** — the resolve-duty OLD from the live `**Severity:**` +bullet, the absence fragment from the live interim report-and-stop duty — check each the three +ways, confirm each counts 1 in each copy, and append both to the fragment table. **Not "before +Step 3":** Step 2 is the install, so a fragment derived after it is taken from text the edit has +already removed. ```bash for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do @@ -726,12 +757,11 @@ observes the bullet, but **`g2` and `g3` are dropped rather than replaced** — report-and-stop duty and its rationale go, and nothing in §E takes their place. A dropped condition owes an **absence check**, not a pair: count its text before and after, expecting `1` then `0`. Without it the duty can survive beside the answer that makes it obsolete, which is the -two-instructions-that-disagree failure in its purest form. - -**Derive the resolve-duty OLD row and the `g2`/`g3` absence fragment**, check each the three ways, -and add both to the table before running the counts. +two-instructions-that-disagree failure in its purest form. **Both fragments were derived and +validated at Step 1**; this step only runs them. -Expected for all four: `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`. +Expected for all four: `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`; and for the +`g2`/`g3` absence check, `worktree=0 parent=1` in each copy. - [ ] **Step 4: Confirm g4 is gone from C** @@ -788,6 +818,33 @@ after `A`), and `at minimum floor 3, severity classified without the demotion` ( Expected: `1` twenty times. **Any `0` means the wording drifted since this table was verified at `5871d0a`** — re-derive that fragment and update the table before installing anything. +- [ ] **Step 1b: Derive the rows the ten locators do not cover, still before installing** + +**One pair per block is a sample, not coverage, and three of the ten blocks change more than one +thing.** Derive each of the following from the **live** text now, check it the three ways, confirm +it counts 1 in each copy, and append it to the fragment table. **Step 2 installs over all of it**, +so a derivation afterwards has no live wording left to check against: + +- **The `c18`-and-surfacing block** changes `c16`, `c17`, `c19` and `c20` besides the `c18` clause + row P8 observes — **one OLD row per condition**, four in total. +- **§G changes two things, not one**: its membership is widened *and* it gains a semantic + membership test a downstream reader can apply, and target §G's own closing note names the test + as the point of the replacement. P7 observes the widened clause; **the test owes a second + observation**. It is **add-only** — the parent paragraph carries no membership test for it to + replace — so it owes presence alone, not a pair. +- **The strict-reading list** replaces the dash-delimited run, which P16 observes as one pair for + the whole run's boundary. Its **tail additions have no predecessor**: each added clause is + add-only and owes **presence alone**, `new/worktree=1 new/parent=0`, with no OLD half. **Do not + derive an OLD row per tail addition** — there is no removed wording for one to count, so the + executor would have to pair the clause with unrelated text and the observation would prove + nothing about the clause. One presence check per independent added clause. + +**The other seven blocks change one thing each and one pair covers them.** + +**One classification note, decided against the real file rather than asserted.** The gate-prompt +template's clean sentence **replaces** `A clean pass is the single body line …`, which is why P13 +has an OLD at all; it is not add-only. + - [ ] **Step 2: Install the §G block and all nine §H blocks — ten sites — one at a time, verifying each before moving to the next** - [ ] **Step 3: Run one discriminating pair per block**, both copies, both trees, one per row P7–P16. @@ -817,16 +874,11 @@ holds OLD fragments, which exist before the edit and can be checked in advance. Expected for all twenty: `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`. -**One pair per block is a sample, not coverage, and two blocks need more.** The `c18`-and-surfacing -block changes `c16`, `c17`, `c19` and `c20` besides the `c18` clause row P8 observes; the -strict-reading list adds several items where P16 observes the first. **Derive one OLD row per -independent meaning change in those two blocks**, add each to the table and the helper, and run a -pair for each — the other eight blocks change one thing each and one pair covers them. - -**Two classification notes, decided against the real files rather than asserted.** The -strict-reading list **replaces the dash-delimited run** even though its addition is at the tail, so -it owes a full pair. The gate-prompt template's clean sentence **replaces** `A clean pass is the -single body line …`, which is why P13 has an OLD at all; it is not add-only. +**Then run what step 1b added:** a pair for each of the `c18`-and-surfacing block's four further +conditions, and a **presence check** — `new/worktree=1 new/parent=0` in each copy — for §G's +semantic membership test and for every independent add-only clause in the strict-reading tail. +Choose each of those NEW fragments from the installed text and verify it the three ways before +counting it. - [ ] **Step 4: Confirm `a21`, `a22`, `a15` survived** — they are carried inside blocks that install contiguously, so a mis-scoped replacement silently drops them. @@ -855,7 +907,7 @@ git commit -m "WIP: install the one-contract paragraph and the remaining prompt- **The bytes:** target §F items 1, 2, 3, 4, 5, 6, 7, 8, 7a, 8a, 8b, 9a, 9b and 9 — fourteen items, each with its own fenced replacement and its `C nnn` / `W nnn` citation. **Re-read every citation against the current file**: §F's own collected list records that items 4, 5 and 8 have line citations one off, and the numbers drifted further as this cycle edited the copies. -**`h4` and `h19` are the human-exception conditions these items discharge** — item 7 is the destination, item 4 the scope sentence. +**`h4`, `h5` and `h19` are the human-exception conditions these items discharge** — item 7 is both destinations, `h4`'s and `h5`'s, item 4 the scope sentence. - [ ] **Step 1: Re-derive every item's real location** @@ -867,29 +919,38 @@ grep -n -F '' CLAUDE.md plugins/dev-workflow/comma **Where a quoted sentence wraps across lines in the file, `grep -F` on the whole sentence returns nothing.** Search a single-line fragment of it instead and confirm by reading. Record which items wrap — Task 14's parity diff needs it. +**Then confirm all fifteen of this task's fragment rows still count `1` in each copy** — `F1`–`F14` +plus `F7b`. **Here, not at step 3:** step 2 installs over every one of them, and after that the +pre-edit count of 1 can no longer be observed at all. A row that no longer counts 1 has drifted +since the table was verified — repair the row against the live line and update the table before +installing anything. + - [ ] **Step 2: Install all fourteen replacements** -- [ ] **Step 3: Build this task's fourteen fragment rows, then run fourteen pairs** +- [ ] **Step 3: Run the fifteen pairs the table already holds** -**Do not choose fragments at run time.** Step 1 located each item's live sentence; for each, pick a -single-line OLD fragment from the located line, run it through the fragment check, and **append the -row to the plan's fragment table with the line numbers the check reported**. Only then run the -pairs. A fragment chosen and used in the same breath is how the first draft shipped five that count -zero. +**Do not rebuild these rows.** The fragment table carries them already — `F1`–`F14` **plus `F7b`** +— each derived from its item's cited line and checked the three ways, and step 1 re-confirmed each +against the live file. Rebuilding them either duplicates authored rows or quietly re-derives one +differently, and the second-copy defect is what that produces. **Consume the rows; do not author +new ones here.** -For each item, the OLD fragment must satisfy all three: single-line in both copies (or one fragment -per copy where the wrapping differs, as `e7` needed); unique in each; and **not preserved inside its -own replacement** — compare it against the §F block before accepting it. +**Fifteen rows for fourteen items, because item 7 changes two clauses on two different lines** — +`h4`'s Gate-A destination at row F7 and `h5`'s Gate-B destination at row F7b. One fragment cannot +observe both, and running fourteen pairs would let `restated by the closing amend` survive while +every stated count still passed. *(Build the pair per the verification procedure; record the four values.)* -Expected for all twenty-eight: `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`. +Expected for all thirty pair instances — fifteen rows in each of the two copies: +`old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`. - [ ] **Step 4: Count what was installed** -Total the fourteen `new/worktree` values step 3 printed, per copy. +Total the fifteen `new/worktree` values step 3 printed, per copy. -Expected: `14` per copy. State the number you observed. **Do not carry a count from §F into a check** — §F states the count of falsified sentences and this task installs a subset of them; a count copied between the two is the stale-bookkeeping defect the design records at five passes running. +Expected: `15` per copy — **fifteen independent changed clauses, from fourteen installed items**. +State the number you observed. **Do not carry a count from §F into a check** — §F states the count of falsified sentences and this task installs a subset of them; a count copied between the two is the stale-bookkeeping defect the design records at five passes running. - [ ] **Step 5: Parity** for all fourteen sites. @@ -987,6 +1048,28 @@ grep -n 'note "' plugins/dev-workflow/hooks/codex-gate.sh Expected: thirteen hits. Seven are the gate reminders this task edits; four take a pre-built variable (`$FAILURE_CTX`, `$NORESULT_CTX`, `$BG_SHORT_CTX`, `$BG_LONG_CTX`), one is the tool-state echo, and one is the **docs-only notice, which is deliberately untouched** — it states no closure permission. +- [ ] **Step 2b: Derive the eight OLD fragments, before installing over them** + +Step 5 runs eight pairs. **Their OLD halves come from the live `note` strings, so they are derived +here** — after step 3 the old wording is gone from the only file that carries it, and a fragment +reconstructed from the parent tree can no longer catch a drifted string or a half-applied re-run. +For each, take a single-line fragment from the located `note` call, confirm it counts **1** in +`codex-gate.sh`, and confirm it is **not preserved inside its own §F replacement**. + +**Three of the eight need a second look while the file is open.** Item 11's removed text is the +**clean definition** at the end of the Gate-A satisfied message, not `floor met by COUNT ONLY`, +which the change leaves standing — an OLD half taken from unchanged text can never reach zero. +Item 10 removes **two separate sentences** of the Gate-A below-floor reminder — the honesty claim +and the tail `Run more passes before executing` — so it owes **two** fragments and two pairs. +Item 17 removes the skip-rule offer and `run more` as **one** clause, so one fragment covers it; +take that fragment from the skip-rule half, since `run more` alone also appears in item 10's tail. + +**Take items 15's and 16's fragments from each message's own body, never from the `STOP` opening +they share.** A fragment from the shared opening counts 2 and cannot tell the two messages apart, +so a stale second message passes behind a repaired first one. The two messages differ throughout — +item 15 is the no-fingerprint reminder, item 16 the stale-fingerprint one — and each has wording +unique to it. + - [ ] **Step 3: Install the seven replacements, one at a time** - [ ] **Step 4: Confirm no behaviour changed, by line number rather than by shape** @@ -1014,30 +1097,27 @@ first time — it admitted exactly the shell keywords it existed to catch. - [ ] **Step 5: Discriminating pairs, worktree and parent** The hook has one copy and owes no parity check; **the hook suite's exact-match assertion is this -edit's second observation** (design §7). Run a pair anyway, one per replaced message, against -`plugins/dev-workflow/hooks/codex-gate.sh` — **seven items, and items 15 and 16 share one `STOP` -opening, so eight observations across seven pairs.** - -**Three of these pairs need a second look before they are run, and the executor derives the -fragments with the file open.** Item 11's removed text is the **clean definition** at the end of the -Gate-A satisfied message, not `floor met by COUNT ONLY`, which the change leaves standing — an OLD -half taken from unchanged text can never reach zero. Item 10 removes **two** things, the honesty -claim and the tail `Run more passes before executing`, and owes an observation for each. Item 17 -likewise removes both the skip-rule offer and `run more`. - -**Exact counts per pair, not a floor.** Five of these replace one message each and must read -exactly `old/parent=1 new/worktree=1`; only the grouped `STOP` pair, which covers items 15 and 16, -reads `2`. Weakening all six to `≥ 1` because one of them is 2 lets a duplicated installation or a -non-unique fragment pass for the five that should be exact: +edit's second observation** (design §7). Run a pair anyway, against +`plugins/dev-workflow/hooks/codex-gate.sh`, using the fragments step 2b derived — **seven items and +eight observations, because item 10 removes two things, so eight pairs.** + +**Exact counts per pair, and every one of them is 1.** Each pair names one message, so a count of +2 anywhere means the fragment is not unique to the message it claims — a duplicated installation, +or a fragment taken from text two messages share. **An earlier draft grouped items 15 and 16 into +one pair reading `2`**, which cannot tell the two messages apart: a stale second message passes +behind a repaired first one, and the grouped `2` also removes the exactness from the row that was +supposed to carry it. Step 2b takes each fragment from its own message's body instead. | Pair | old/worktree | old/parent | new/worktree | new/parent | |---|---|---|---|---| | item 10, the honesty claim | 0 | 1 | 1 | 0 | +| item 10, the `Run more passes before executing` tail | 0 | 1 | 1 | 0 | | item 11, the Gate-A clean definition | 0 | 1 | 1 | 0 | | item 12, the Gate-B clean definition | 0 | 1 | 1 | 0 | | item 13, the WIP reminder | 0 | 1 | 1 | 0 | +| item 15, the no-fingerprint reminder | 0 | 1 | 1 | 0 | +| item 16, the stale-fingerprint reminder | 0 | 1 | 1 | 0 | | item 17, the below-floor instruction | 0 | 1 | 1 | 0 | -| items 15+16, the two `STOP` openings | 0 | **2** | **2** | 0 | - [ ] **Step 6: ShellCheck** @@ -1076,12 +1156,23 @@ git commit -m "WIP: replace the seven gate reminders the ordering falsifies" ```bash grep -n 'expected_ctx=\|expected_msg=' plugins/dev-workflow/hooks/codex-gate.test.sh grep -n 'Gate B satisfied\|Gate B not satisfied\|Gate A satisfied\|Gate A floor\|STOP\|no recorded review\|cannot confirm review\|only thing keeping\|no new Blocker/Major\|count only\|COUNT ONLY' plugins/dev-workflow/hooks/codex-gate.test.sh +grep -ni 'satisfied' plugins/dev-workflow/hooks/codex-gate.test.sh ``` **`Gate A` as well as `Gate B`.** An earlier draft searched only the Gate-B phrases while step 3 requires every label naming a gate verdict to move; the Gate-A assertions and labels around lines 504–506 would have been missed by the locator and then demanded by the rule. +**The third grep is the one that keeps the suite green, and it is case-insensitive on purpose.** +Most assertions in this file do not quote the qualified phrase at all: they grep the hook's output +for the bare verdict word — `grep -q 'not satisfied'` — and label the case `-> NOT satisfied` in +upper case. §F items 15 and 16 remove that word from both messages in favour of `cannot confirm` +and `Codex gate state:`, so **every one of those assertions goes red the moment Task 10 lands**, +and the repo cannot pass its own quality battery. A locator matching only `Gate B not satisfied` +sees three of them. **Read and update every hit of the bare word by the hook state the message now +reports**, exactly as step 3 requires — and do not turn the observed hit count into a target, for +the reason this task's opening gives. + **Record the counts you observe.** They will not match the numbers above if the file has changed; the numbers above are evidence for why no list is kept, not a target. - [ ] **Step 2: Update the three `expected_ctx` and three `expected_msg` assignments** @@ -1434,10 +1525,10 @@ git commit -m "WIP: plan records" discards every WIP message; evidence written only there would be destroyed by the close. Restate it in the WIP body too if a mid-cycle reader would want it, but the file is the copy that survives. -**The file is completed after step 7, not here.** The evidence entry can be drafted now, but the +**The file is completed at step 7b, not here.** The evidence entry can be drafted now, but the **per-pass curve is not known until the Gate-B loop ends**, and the **provenance line** and any -**human-exception record** belong beside it. Step 7's last action is to append all three — this step -opens the file, step 7 closes it, and step 8 commits it. +**human-exception record** belong beside it. **Step 7b appends all three and revalidates the +entry** — this step opens the file, step 7b closes it, and step 8 commits it. It names: the battery run; **every pair this plan built, with its counts in each copy and each tree, and every presence check beside them**; the §6 parity diff and the `b11`/`b13` equivalence result; and the next-state table's location plus its row count. **State the §A presence checks as presence, not as pairs** — its counterfactual is absent and is claimed as absent. @@ -1475,6 +1566,24 @@ closing commit.** **A fix that changes specified behaviour updates the spec in the same commit.** +- [ ] **Step 7b: After the clean pass, complete `.context/loop-rule-closing-msg` — this is the + action step 5 defers to, and step 8 has no other source for these records** + +Append to the file step 5 opened, then read the whole file back before step 8 runs: + +1. the **provenance line**, in the form Mechanics pins — the Gate-B cycle's nonce, the derived + floor, the cited set with each member's level, and the workspace knob; +2. the **per-pass curve** for this Gate-B cycle, `; Gate B (passes …): Findings … + Blockers … Majors …`, which is why this cannot be written at step 5: the counts do not exist + until the loop ends; +3. any **human-exception record**, and beside it the skip reason if a cycle was skipped; +4. the **revalidated evidence entry**, replacing step 5's draft if revalidation changed it. + +Confirm the file now carries all four before continuing. **A missing one is not recoverable after +step 8** — `git commit -F` publishes whatever the file holds, `reset --soft` has already discarded +every WIP body, and a closing commit without its provenance line or curve has not validly closed +the cycle. + - [ ] **Step 8: Close the cycle** **Build the closing message in a file first.** `git reset --soft` discards every WIP commit *body*, @@ -1512,7 +1621,7 @@ The closing body carries: the validated evidence entry; the provenance line; the ## Self-Review -**1. Spec coverage.** §A → Task 1. §B → Task 3. §C → Task 4. §D → Task 5. §E → Task 6. §F items 1–9 → Task 8; items 14, 18 → Task 9; items 10–13, 15–17 → Task 10 with its test sweep in Task 11. §G → Task 7, **added by this review**: the first draft gave the one-contract paragraph no task, though design §4 lists it as its own site and target §G carries its replacement. It is a prompt-copy replacement in both copies with the same shape as §H's blocks, owes the same discriminating pair with row **P7**'s OLD `These records are one contract` — the live wording; `These rules and records are one contract` occurs nowhere — at C 879 / W 1063 as of this writing. §H → Task 7. §I ships nowhere and needs no task. Design §6 → Task 14. Design §7 → Tasks 13 and 15. Design §8 → Task 15's battery and the Global Constraints. Story AC 5 → the disposition tables. Story AC 4 → Task 13. +**1. Spec coverage.** §A → Task 1. §B → Task 3. §C → Task 4. §D → Task 5. §E → Task 6. §F items 1–9 → Task 8; items 14, 18 → Task 9; items 10–13, 15–17 → Task 10 with its test sweep in Task 11. §G → Task 7, **added by this review**: the first draft gave the one-contract paragraph no task, though design §4 lists it as its own site and target §G carries its replacement. It is a prompt-copy replacement in both copies with the same shape as §H's blocks, owes a discriminating pair with row **P7**'s OLD `These records are one contract` — the live wording; `These rules and records are one contract` occurs nowhere — at C 879 / W 1063 as of this writing, **and a second, add-only presence check for the semantic membership test it gains**, which P7's pair does not observe. §H → Task 7. §I ships nowhere and needs no task. Design §6 → Task 14. Design §7 → Tasks 13 and 15. Design §8 → Task 15's battery and the Global Constraints. Story AC 5 → the disposition tables. Story AC 4 → Task 13. **2. Placeholder scan.** The replacement text is cited rather than copied, deliberately and for the reason the Architecture note gives. Task 13's row list is explicitly a floor rather than a closed set, and says so. Task 11 deliberately carries no count, and says why. From 54e793b662b3fe0d509230edfb9d5c79e640a928 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 16:17:54 +0200 Subject: [PATCH 108/181] docs(plans): apply Gate-A plan pass 7; stop pre-assigning fragment ids MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit MANDATORY TWO-TELL STOP — three tells present. Findings rose 11 -> 13, Blockers rose 2 -> 4, and the findings cluster on the verification apparatus for the seventh pass running. Surfaced to Daniel; his standing answer applies and the loop continues on it. Four Blockers and two Majors, all validated against the real files. - Task 0 step 1 required the recorded base to be an ancestor of HEAD and then tested it with `git log "$BASE"..HEAD`, which tests no such thing. A base from an abandoned branch passes, and `reset --soft` onto it drops every real commit since the fork. `git merge-base --is-ancestor` now runs first. - Task 0 step 3 put the single-line squash-carry site through `sed -n "/x/,/x/p"` with a comment claiming sed returns that one line — the failure step 2 of the same task states correctly. Extraction ran to EOF. It is a grep now, outside the range loop. - Passage (b) invented "decision 6's decline semantics" as a second §B added rule. Target §B adds one rule; the decline clause is §A's, covered by Task 1. Seventh condition-table misclassification in seven passes. - Task 4 was promised the ids P19/P20 while Task 3 was told to continue the same series — two fragments, one id. No task pre-assigns an id now; rows take the next free P at append and are named by the condition they observe. - Task 4 step 4 paired row P4's `c14` OLD with a `c8` NEW, the cross-edit pair its own next paragraph forbids. Three pairs, each with both halves from one edit, in a table. - Task 4 step 4 still derived rows after step 2 had installed over them, contradicting the pre-install derivation pass 6 added to step 1. - Task 7 step 3 said to add each NEW fragment to the fragment table and then said NEW fragments are not added to it. The table is OLD-only; chosen NEW fragments go in the task evidence. Task 1 step 4 carried the same contradiction and is corrected with it. - Task 15 step 8 deleted the recovery base file whether or not the closing commit succeeded — in the one state where the reset has already discarded every WIP commit and that file is the only record of where the cycle began. Minors and the Nit repaired rather than collected, each being a contradiction inside text pass 6 had just edited: Task 11 step 5's closing check narrower than its own locator; the stale "thirty-two rows" sweep claim, removed rather than renumbered; the span coverage claiming 23 where h6 is deliberately outside it; the Self-Review claiming a runnable command per OLD half; and Task 1 step 6 saying Task 15 amends where it resets and recommits. Trend: findings 32, 32, 31, 20, 16, 11, 13. Blockers 1, 0, 0, 10, 7, 2, 4. Majors 28, 25, 23, 7, 7, 6, 4. Floor 3, from risk high (2) and security none (0), read at this pass from docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-plan-om0bdd7udh-pass-7.md | 14 ++ .../2026-09-14-loop-rule-consolidation.md | 166 ++++++++++++------ 2 files changed, 127 insertions(+), 53 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-7.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-7.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-7.md new file mode 100644 index 0000000..b756ec9 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-7.md @@ -0,0 +1,14 @@ +BLOCKER | high | Task 0 step 1 | The plan states that a pre-existing `$BASE` must be an ancestor of `HEAD`, but `git log "$BASE"..HEAD` does not test ancestry; a commit on another branch can still yield a range containing only `WIP:` commits and pass the stated inspection. | Gate B can review the wrong range and Task 15 can soft-reset onto an unrelated commit, rewriting or absorbing work outside this execution. | Add an explicit failing ancestry check such as `git merge-base --is-ancestor "$BASE" HEAD` before inspecting the intervening commit subjects. +BLOCKER | high | Task 0 step 3 | The `On squash-merge` site is encoded with the same start and end anchor and passed to `sed -n "/$s/,/$e/p"`; despite the comment claiming one line, `sed` does not test the end address on the start line, and each real file contains this anchor only once, so extraction runs to EOF. | The baseline diff includes the entire unequal tail of both prompt files and cannot produce only the stated deliberate divergences, blocking a correct execution or misclassifying tail differences as inherited drift. | Extract the single-line squash-carry site with its own exact `grep` or one-line lookup, as Task 0 step 2 already requires, and exclude it from the range loop. +MAJOR | high | Passage (b) disposition and Task 3 steps 1b/3 | The plan invents “decision 6's decline semantics” as a second add-only rule in target §B, while approved target §B explicitly says its only Added item is the closing-time change rule; the decision-6 sentence “Decline is available only at a membership stop” is in §A. | Task 3 asks the executor to choose and certify a §B fragment for a rule that §B does not contain, so the presence check must either be impossible or observe an unrelated decline clause. | Remove the invented §B add-only item; cover decision 6 with Task 1's §A paragraph presence check and keep §B's added-rule check scoped to the closing-time change named by the target. +BLOCKER | high | Task 4 step 1 | Task 4 fixes its new rows as `P19_OLD` and `P20_OLD`, but Task 3 runs first and adds seven rows with ids continuing the existing P series after P18, so it already consumes P19 through P25 even if Task 1's two added rows use another naming scheme. | The fragment table gains duplicate ids and Task 4's later references cannot identify whether they mean a passage-(b) row or the intended `c4`/`c8` row. | Allocate Task 4's ids only after all earlier rows have been appended, or reference the rows by stable condition-qualified names instead of preassigning P19/P20. +BLOCKER | high | Task 4 step 4 | The step explicitly assigns row P4's OLD half from `c14` to a NEW fragment from `c8`, repeating the cross-edit pair the following paragraph says is invalid. | The three meaning changes are not given three unambiguous same-edit pairs, so old `c8` wording can coexist with its replacement or the new `c14` behavior can be absent without the intended discriminating observation. | Map one OLD and one NEW fragment to each of `c4`, `c8`, and `c14`; pair P4 with the `c14` replacement and pair the `c8` OLD row with the stated recurrence fragment. +BLOCKER | high | Task 4 step 4 | After step 2 has installed over the live passage, this step again says to derive `P19_OLD` and `P20_OLD` from that live passage, contradicting the valid pre-install derivation in step 1. | Following the later instruction makes the required pre-edit count impossible and forces reconstruction from the parent, losing the drift and partial-rerun detection the procedure requires. | Delete the post-install derivation paragraph and have step 4 consume only the rows derived and validated before step 2. +MAJOR | high | Task 7 step 3 | The first instruction says to add every chosen NEW fragment to the fragment table, while the next paragraph says a NEW fragment is not added because the table holds OLD fragments only. | The executor has two incompatible definitions of the table's contents, so verification evidence can be duplicated, omitted, or recorded in a form later tasks do not know how to consume. | Choose one storage rule and state it consistently; under the surrounding procedure, keep the table OLD-only and record chosen NEW fragments with their four counts in the task evidence. +MINOR | high | Task 11 step 5 | The final “no verdict vocabulary survives” check searches only three qualified phrases and does not repeat the case-insensitive bare `satisfied` sweep that steps 1 and 3 make part of the edit. | A missed continuation-line label such as `unborn repo hashes, self-matches, and can reach satisfied` can survive while the suite and step 5 both pass, leaving the test vocabulary claiming the old verdict. | Re-run the case-insensitive bare-verdict sweep after editing and record the disposition of every remaining hit, with an absence check for labels and comments that named the replaced hook verdicts. +MAJOR | high | Task 15 step 8 | `rm -f .context/loop-rule-base` runs after `git commit -F` without being conditional on that commit succeeding, and the command block enables no fail-fast behavior. | A hook, signing, identity, or other commit failure deletes the recovery base while leaving the repository soft-reset and without the closing commit, making recovery and a safe rerun materially harder. | Remove the base file only after explicitly verifying the closing commit succeeded, for example by chaining the removal to the commit's success and checking the resulting `HEAD` and worktree state. +MINOR | high | Verification fragments section | The current tables contain 34 rows and 33 actual fragments after adding F7b, but the normalized-replacement claim still says “all thirty-two rows”; both cited verification snapshots, `9f13a2c` and `58b3660`, predate F7b, with the former also predating every F row. | The plan falsely presents F7b as covered by its recorded three-condition sweeps, so its audit trail does not support the current table even though F7b does occur once per real copy and is absent from the normalized replacement text. | Re-run the three-condition verification over all 33 actual fragments, update the stated count, and cite a snapshot that contains F7b and the complete current tables. +NIT | high | Task 0 step 2 | The human-exception split says all twenty-three kept conditions lie inside the three spans, but `h6` shares the real line with changed `h5`/F7b and is intentionally excluded from those spans; the later per-condition paragraph correctly names `h6` separately. | The coverage description contradicts its own fallback and overstates what the span artifact proves. | State that twenty-two kept conditions are covered by the spans and `h6` is covered by its separate fragment check. +MINOR | high | Self-Review item 2 | It claims every OLD half is already a concrete fragment with a runnable command, although Tasks 1, 3, 4, 6, 7, and 10 deliberately derive rows or fragments during execution and the plan deliberately does not pre-write the per-task shell. | The self-review overstates the plan's current verification completeness and obscures which observations still depend on execution-time derivation. | Rewrite the claim to distinguish preverified table rows from execution-time OLD/NEW derivations and to describe the shared procedure and exact expected results without claiming prewritten commands. +MINOR | high | Task 1 step 6 | The rationale says Task 15 “amends once,” contradicting the Global Constraints and Task 15's explicit `reset --soft` followed by a new closing commit. | The plan gives two different closing mechanics, inviting an executor to use the obsolete amend path that can mishandle the accumulated WIP evidence. | Replace the amend statement with the actual reset-and-recommit closing shape used by Task 15. +END OF FINDINGS (13 total) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index 7048342..c860a8a 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -111,12 +111,23 @@ | **Unique** in that file | a count of 1 proves nothing about which occurrence changed | — | | **Absent from its own replacement** | its old-count can never reach zero | pass 2, three rows | -**The third check compares against the target's fenced blocks with line breaks normalized.** `F4` is one line in `CLAUDE.md` but the block wraps it between `you` and `still`; a substring test finds nothing and the row looks usable. Pass 4 found it, and re-running the normalized check over all thirty-two rows found that one and no other. +**The third check compares against the target's fenced blocks with line breaks normalized.** `F4` is one line in `CLAUDE.md` but the block wraps it between `you` and `still`; a substring test finds nothing and the row looks usable. Pass 4 found it, and re-running the normalized check over every row then in the table found that one and no other. + +**That sweep predates row F7b**, which pass 6 added, and the snapshots the two tables cite predate it too. **Re-run all three checks over every row the tables now hold before Task 0 finishes**, and record the revision you ran them at — a row presented as covered by a sweep that could not have seen it is a claim about evidence that does not exist. *(No count is stated here on purpose: a number in this sentence goes stale the next time a task appends a row, which is the enumeration failure this plan keeps finding in itself.)* ### The OLD fragments **Every row below was checked all three ways at `58b3660`.** A row Tasks 3, 4, 6 and 7 add is checked the same way and appended here, so this table stays the one place they live. +**No task pre-assigns a row id, and none refers to a row by a number this plan does not already +contain.** A task appending rows takes the **next free `P` id at the moment it appends**, and +names its own rows by the condition they observe — "the `c14` row", "the resolve-duty row" — not by +a number chosen in advance. Tasks 3, 4, 6 and 7 all append, and how many each adds is decided at +execution against the real files, so any number written here ahead of time is a guess that the +task running before it invalidates. An earlier draft promised Task 4 the ids `P19` and `P20` while +Task 3 was told to continue the same series, which hands two different fragments one id and leaves +every later reference ambiguous. + --- ## File Structure @@ -152,7 +163,9 @@ Story acceptance criterion 5 is satisfied here. Ids are `docs/superpowers/specs/ ### Passage (b) — what a loop absorbs → target §B (Task 3) -§B states its own accounting and this table reproduces it: **Changed:** `b3`, `b7`, `b8`, `b11`, `b12`, `b13`, `b16`, `b17`, `b18`. **Added:** the closing-time change rule and decision 6's decline semantics, which no inventoried condition carried because none existed. **Carried:** `b1`, `b2`, `b4`, `b5`, `b6`, `b9`, `b10`, `b14`, `b15`. +§B states its own accounting and this table reproduces it: **Changed:** `b3`, `b7`, `b8`, `b11`, `b12`, `b13`, `b16`, `b17`, `b18`. **Added:** the closing-time change rule, and that alone — no inventoried condition carried it because none existed. **Carried:** `b1`, `b2`, `b4`, `b5`, `b6`, `b9`, `b10`, `b14`, `b15`. + +**Decision 6's decline semantics are not a §B addition.** `**Decline is available only at a membership stop**` is a §A sentence, and Task 1's presence checks cover it. An earlier draft listed it here as a second added rule, which would have sent Task 3 looking for a §B fragment that does not exist — and the only way to satisfy that check is to certify an unrelated decline clause. **Verify against the file, not against this table:** §B is written out whole and is the only place this change states passage (b). Read the installed passage and confirm each carried condition is present and each changed one is gone. @@ -252,14 +265,22 @@ final `reset --soft` would then start after every edit made so far — prompt an squashed into the closing commit without ever entering a review range. **A pre-existing value is not accepted on being non-empty.** It is valid only if it is an ancestor of -`HEAD` **and** every commit between it and `HEAD` is a `WIP:` commit of this execution. Check that -before proceeding: +`HEAD` **and** every commit between it and `HEAD` is a `WIP:` commit of this execution. **Both +halves are checked, and the ancestry one first:** ```bash -git log --oneline "$(cat .context/loop-rule-base)"..HEAD +B=$(cat .context/loop-rule-base) +git merge-base --is-ancestor "$B" HEAD || { echo "recorded base is NOT an ancestor of HEAD — stale or from another branch"; exit 1; } +git log --oneline "$B"..HEAD ``` -Expected: nothing, or only `WIP:` commits of this run. **Anything else means the file is stale** — +**`git log "$B"..HEAD` does not test ancestry**, and reading it as if it did is how a base from an +abandoned branch passes: the range then lists what `HEAD` has and `$B` does not, which can be only +`WIP:` commits while `$B` sits on a branch of its own. `git reset --soft` onto it at Task 15 step 8 +would move `HEAD` to that unrelated commit and drop every real commit since the fork. `merge-base +--is-ancestor` is the test the sentence above actually names. + +Expected: the ancestry check exits 0, and the log shows nothing, or only `WIP:` commits of this run. **Anything else means the file is stale** — left by an abandoned run, or by one whose work was already squashed. Delete it deliberately, record why, and re-record from the true starting commit. **Task 15 step 8 removes the file after the closing commit**, so a stale one is an abandoned run rather than a normal state. @@ -322,8 +343,10 @@ span**, which the first draft's version did not: the floor knob`. **The middle span is the one the first draft dropped**, and it holds `a3`–`a12`. - **The human exception is three spans**, around rows F7 (`h4`), **F7b (`h5`)** and F4 (`h19`): from `Recording a human exception` to before F7's line; **from after F7b's line** to before - F4's; and from after F4's to `because writing it down makes it sound`. All twenty-three kept - conditions lie inside them. **F7 and F7b sit on adjacent lines** — C 988 and 989 as of this + F4's; and from after F4's to `because writing it down makes it sound`. **Twenty-two** of the + twenty-three kept conditions lie inside them; **`h6` is the exception** — it shares its line with + `h5`, which F7b changes, so no whole-line span can hold it and the per-condition fragment check + below is what covers it. **F7 and F7b sit on adjacent lines** — C 988 and 989 as of this writing — **so a middle span opening after F7 rather than after F7b would enclose a line item 7 changes, and a correct implementation would fail its own untouched check.** Resolve both anchors and open the middle span after the later of them. @@ -356,8 +379,8 @@ comparison actually emits. ```bash # Tab-separated start and end anchors: the anchors contain colons, so a # colon delimiter splits '**Severity:**' at the wrong place and yields an -# empty end. A single-line site is given the same anchor twice and sed -# returns that one line. +# empty end. Every entry here spans two DIFFERENT anchors — the single-line +# squash-carry site is handled below, not in this loop. printf '%b\n' \ 'Both gates are a LOOP\tNothing here writes the floor knob' \ 'What a loop absorbs\tRecognizing "clearly stuck"' \ @@ -367,7 +390,6 @@ printf '%b\n' \ 'The two rules above\tFindings go to a FILE' \ '\*\*Severity:\*\*\t\*\*Tool routing:' \ 'Recording a human exception\tbecause writing it down makes it sound' \ - 'On squash-merge\tOn squash-merge' \ 'When these rules bind\tDownstream has no shipping commit' \ > .context/loop-rule-sites while IFS=$(printf '\t') read -r s e; do @@ -375,8 +397,21 @@ while IFS=$(printf '\t') read -r s e; do diff <(sed -n "/$s/,/$e/p" CLAUDE.md) \ <(sed -n "/$s/,/$e/p" plugins/dev-workflow/commands/workflow-init.md) done < .context/loop-rule-sites | tee .context/loop-rule-baseline-diff.txt + +# The squash-carry sentence is ONE line and must not go through the loop. +echo "== On squash-merge" | tee -a .context/loop-rule-baseline-diff.txt +diff <(grep -F 'On squash-merge, copy every evidence entry' CLAUDE.md) \ + <(grep -F 'On squash-merge, copy every evidence entry' plugins/dev-workflow/commands/workflow-init.md) \ + | tee -a .context/loop-rule-baseline-diff.txt ``` +**The squash-carry site is extracted with `grep`, not as a range**, for the reason step 2 already +gives: `sed` does not test the end address on the line that matched the start, so +`sed -n "/x/,/x/p"` on a site occurring once runs **to end of file**. An earlier draft put it in +the loop with its own anchor at both ends and a comment claiming `sed` returns that one line — the +diff would then have carried the whole unequal tail of both files and no correct implementation +could have produced the expected divergence list. + Expected: the deliberate divergences the inventory records — `b3`'s cross-reference target, passage (b)'s intensifier and field-mint parenthetical, `e8`'s pronoun, `e11`, `f5`–`f7`'s framing, and `g4`. **Read every line of the output; do not truncate it.** Anything else is pre-existing drift — @@ -434,9 +469,10 @@ nothing else observes them. *(Count each fragment in both copies and both trees, per the verification procedure; expect worktree 1, parent 0.)* -Expected: `worktree=1 parent=0` for all six. **Add the two §A2/§A3 fragments to the fragment table -once chosen**, with the check that verified each is single-line and unique — they are the two rows -that cannot be pre-verified because the text does not exist until this task installs it. +Expected: `worktree=1 parent=0` for all six. **Record the three chosen fragments with their counts +in this task's evidence**, each with the check that verified it single-line and unique — **not in +the fragment table**, which holds OLD fragments only. §A has no OLD half at all: row P1 says so, +and these three exist only once this task installs them. - [ ] **Step 5: Check parity of the installed block** @@ -455,13 +491,16 @@ git add CLAUDE.md plugins/dev-workflow/commands/workflow-init.md \ git commit -m "WIP: install the closure ordering into both §5 copies" ``` -**The plan is staged here because this task adds the §A2 and §A3 rows to the fragment table.** Every -task that adds a row stages the plan with its own edit; otherwise the reviewed fragment evidence -stays dirty and is swept into a later, unrelated commit, and the task commits are not the -independently reviewable units this plan claims they are. The same applies to Tasks 3, 4, 6, 7 and -8, each of which derives rows. +**The plan is staged here because this task writes its fragment evidence into the plan.** Every +task that records a fragment — an appended OLD row, or a chosen NEW fragment with its counts — +stages the plan with its own edit; otherwise the reviewed fragment evidence stays dirty and is +swept into a later, unrelated commit, and the task commits are not the independently reviewable +units this plan claims they are. The same applies to Tasks 3, 4, 6, 7 and 8. -Named `WIP:` because Task 15 runs Gate B over the whole change and amends once. A non-`WIP` commit here would reset the hook's Gate-B counters mid-cycle. +Named `WIP:` because Task 15 runs Gate B over the whole change and closes it with **one +`git reset --soft "$BASE"` and a single commit**, per the Global Constraints and Task 15 step 8 — +not with an amend, which is the Mechanics shape for a cycle carrying one snapshot and this plan +makes one per task. A non-`WIP` commit here would reset the hook's Gate-B counters mid-cycle. --- @@ -526,13 +565,13 @@ passes` here, which wraps across C 201–202 and W 408–409 and counts zero in - [ ] **Step 1b: Derive the rest of §B's OLD rows, still before installing** **Two pairs do not cover §B.** Target §B separately changes `b3`, `b8`, `b11`, `b13`, `b16`, -`b17`–`b18` and **adds** the closing-time set-change rule and decision 6's decline semantics. Each -independent replacement owes its own pair, and each add-only rule owes a presence check — a manual -condition walk is a reader's judgement, not the discriminating observation design §7 assigns here. -**Derive an OLD row for each changed condition now**, check each the three ways against the live -passage, confirm it counts 1 in each copy, and append it to the fragment table. Ids continue the -`P` series. **This is step 1b and not part of step 3 because step 2 removes the wording these rows -are taken from** — a row derived afterwards cannot be checked against the text it describes. +`b17`–`b18` and **adds one rule**, the closing-time set-change rule. Each independent replacement +owes its own pair, and the added rule owes a presence check — a manual condition walk is a +reader's judgement, not the discriminating observation design §7 assigns here. **Derive an OLD row +for each changed condition now**, check each the three ways against the live passage, confirm it +counts 1 in each copy, and append it to the fragment table under the next free `P` id. **This is +step 1b and not part of step 3 because step 2 removes the wording these rows are taken from** — a +row derived afterwards cannot be checked against the text it describes. - [ ] **Step 2: Install §B's text over the passage in both copies** @@ -550,10 +589,10 @@ three ways before this step counts with it. Expected for all four: `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`. -**Run a pair for every row step 1b derived**, and a **presence check** for each of §B's two added -rules — the closing-time set-change rule and decision 6's decline semantics — which replace no -wording and so owe `new/worktree=1 new/parent=0` and no OLD half, per the add-only rule. Choose -each NEW fragment from the installed text and verify it the three ways before counting it. +**Run a pair for every row step 1b derived**, and a **presence check** for §B's one added rule — +the closing-time set-change rule — which replaces no wording and so owes `new/worktree=1 +new/parent=0` and no OLD half, per the add-only rule. Choose each NEW fragment from the installed +text and verify it the three ways before counting it. - [ ] **Step 4: Walk the carried conditions** @@ -600,12 +639,13 @@ Expected: `P4=1` for both files. is **preserved in target §C's fenced replacement**, so its old-count could never reach zero. Derive `c4`'s OLD from the part of the sentence the replacement removes. -**`P19_OLD` and `P20_OLD` do not exist yet, and they are derived here rather than at step 4.** -Take each from the live passage, check it the three ways, confirm it counts 1 in each copy, and -**add a row to the fragment table** — that table is where every fragment lives, and a row that is -not there is a fragment stated somewhere else. Ids continue the `P` series. **Step 2 removes the -wording they are taken from**, so a derivation after it has no live text to check against and no -pre-edit count to observe. +**`c4` and `c8` have no row yet, and both are derived here rather than at step 4.** Take each from +the live passage, check it the three ways, confirm it counts 1 in each copy, and **add a row to +the fragment table** under the next free `P` id — that table is where every fragment lives, and a +row that is not there is a fragment stated somewhere else. Refer to them afterwards as **the `c4` +row** and **the `c8` row**, never by a number picked here: Task 3 appends first and how many rows +it adds is decided at execution. **Step 2 removes the wording they are taken from**, so a +derivation after it has no live text to check against and no pre-edit count to observe. - [ ] **Step 2: Install §C's block** @@ -626,21 +666,25 @@ result a reader can only shrug at. - [ ] **Step 4: Run the discriminating pair, both copies, both trees** -Row **P4**. `NEW` is the single-line fragment `a recurrence failing them being an ordinary fresh -finding`, from §C's re-raised-dismissal clause — **install that clause's line unwrapped** so the -fragment sits wholly on one line, and confirm it is unique before counting. +**Three pairs, one per edit, and each pair's two halves come from the same edit.** An earlier draft +ran a single pair taking its OLD from `c14` and its NEW from `c8` — two different changes — so +either could land while the other survived and it still reported a pass, and `c4` had no +observation at all. **A later draft reintroduced exactly that pair** by naming row P4, whose OLD is +`c14`, and then taking its NEW from §C's re-raised-dismissal clause, which is `c8`. -**Three pairs, one per edit.** An earlier draft ran a single pair taking its OLD from `c14` and its -NEW from `c8` — two different changes — so either could land while the other survived and it still -reported a pass, and `c4` had no observation at all. +| Edit | OLD | NEW, taken from the installed §C block | +|---|---|---| +| `c4`, the widened third condition | the `c4` row (step 1) | the clause §C puts in place of "a missing one means keep going" | +| `c8`, the re-raised dismissal | the `c8` row (step 1) | `a recurrence failing them being an ordinary fresh finding` — **install that clause's line unwrapped** so the fragment sits wholly on one line | +| `c14`, the below-floor Minor | row **P4** | the ordering's replacement for the below-floor sentence | -*(Build the pair per the verification procedure; record the four values.)* +Each NEW is confirmed single-line and unique in the installed file before it is counted. -Expected for all six: `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`. +*(Build each pair per the verification procedure; record the four values.)* -**`P19_OLD` and `P20_OLD` do not exist yet.** Derive each from the live passage, check it the three -ways, and **add a row to the fragment table** — that table is where every fragment lives, and a row -that is not there is a fragment stated somewhere else. Ids continue the `P` series. +Expected for all six pair instances — three pairs in each of the two copies: +`old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`. **All three OLD rows were derived and +validated at step 1**; this step only runs them. - [ ] **Step 5: Walk the conditions** — `c1`–`c3` present unchanged, `c5`–`c7` word for word, `c8` carrying the new clause, `c9` split, `c10`–`c14` gone from here. @@ -850,8 +894,9 @@ has an OLD at all; it is not add-only. - [ ] **Step 3: Run one discriminating pair per block**, both copies, both trees, one per row P7–P16. `NEW` for each is taken from the installed block; **check each chosen fragment is single-line and -unique in the installed file before counting it**, and add it to the fragment table. The suggested -source sentence per block: +unique in the installed file before counting it**, and record it with its four counts in this +task's evidence — **not in the fragment table**, which holds OLD fragments only, for the reason +stated below the table. The suggested source sentence per block: | Row | Block | NEW taken from | |---|---|---| @@ -1196,9 +1241,16 @@ Expected: exit 0 from both. **`HOOK_SH` selects the shell the hook runs under; w ```bash grep -c 'Gate B satisfied\|Gate B not satisfied\|Gate A satisfied' plugins/dev-workflow/hooks/codex-gate.test.sh +grep -ni 'satisfied' plugins/dev-workflow/hooks/codex-gate.test.sh ``` -Expected: **`0`. Not "no hits you cannot justify"** — step 3 requires every label and comment to move +Expected: **`0` from the first**, and **every remaining hit of the second disposed of in writing** — +the bare case-insensitive sweep is repeated here because it is the one step 1 used to find the +sites, and a closing check narrower than the locator cannot confirm the sweep it closes. A label +wrapped onto a continuation line, or one naming the verdict without the gate, survives all three +qualified phrases while the suite passes. + +**Not "no hits you cannot justify"** — step 3 requires every label and comment to move to observed hook state, and a justify-in-the-commit-body escape hatch is how the old gate-verdict vocabulary survives a sweep. Where a test still needs that case, name it by what the hook observed. @@ -1599,10 +1651,18 @@ git add docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md .context/co git commit -m "WIP: plan records and Gate-B findings files" || true git status --porcelain # expect empty git reset --soft "$BASE" -git commit -F .context/loop-rule-closing-msg +git commit -F .context/loop-rule-closing-msg || { echo "closing commit FAILED — base file kept for recovery"; exit 1; } +git log --oneline -1 # expect the closing message, not a WIP +git status --porcelain # expect empty rm -f .context/loop-rule-base ``` +**The base file is removed only after the closing commit succeeded**, and the two checks above are +what "succeeded" means here. `git commit` can fail on a hook, a signing key or an unset identity, +and at that point the reset has already happened: the WIP commits are gone, the whole change is a +staged tree, and `.context/loop-rule-base` is the only record of where the cycle started. Deleting +it unconditionally destroys the one value a rerun needs, in the single state where it is needed. + **`$BASE` is the recorded revision, not a placeholder to substitute by hand**, and the file is removed afterwards so a later run cannot inherit a stale one. @@ -1625,7 +1685,7 @@ The closing body carries: the validated evidence entry; the provenance line; the **2. Placeholder scan.** The replacement text is cited rather than copied, deliberately and for the reason the Architecture note gives. Task 13's row list is explicitly a floor rather than a closed set, and says so. Task 11 deliberately carries no count, and says why. -**Which steps are mechanical and which are reader checks, stated rather than claimed uniformly.** Every OLD half of every pair is a concrete fragment with a runnable command and an exact expected result. **The NEW halves are `<...>` until their task installs the text**, which the fragment table discloses and each step requires to be verified before counting. **Tasks 12, 13 and 14 step 2 are reader checks by design** — a predicate comparison, a next-state walk and a divergence classification are judgements, and giving them commands would be the false-precision this repo's invariants warn about. An earlier revision of this section claimed every verification step had a runnable command, which was not true of them. +**Which steps are mechanical and which are reader checks, stated rather than claimed uniformly.** Every OLD half has an exact expected result and a procedure that produces it, but **not every one is a pre-verified table row**: the rows the tables carry were checked against the real files in advance, while Tasks 1, 3, 4, 6, 7 and 10 **derive their remaining OLD fragments at execution, before their install step**, against text this plan cannot quote without becoming a second copy of it. **The NEW halves are `<...>` until their task installs the text**, which the fragment table discloses and each step requires to be verified before counting. **And no per-task shell is pre-written at all** — the procedure is stated once and the executor writes the command in front of the files, so "a runnable command per step" is not what this plan claims. **Tasks 12, 13 and 14 step 2 are reader checks by design** — a predicate comparison, a next-state walk and a divergence classification are judgements, and giving them commands would be the false-precision this repo's invariants warn about. An earlier revision of this section claimed every verification step had a runnable command, which was not true of them. **3. Type consistency.** `$BASE` is set in Task 0 and used in Tasks 1–10. The four-value pair shape (`new/worktree`, `new/parent`, `old/worktree`, `old/parent`) is defined in Task 3 and referred to by name afterwards. Condition ids match the inventory throughout: a1–a22, b1–b18, c1–c20, d1–d7, e1–e11, f1–f7, g1–g4, h1–h26, i1–i16, j1–j4 — 135 total, every one dispositioned above. From 41d73a57fec6e653684542c9093e583c1052900f Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 16:18:10 +0200 Subject: [PATCH 109/181] docs(context): record passes 6 and 7 in the Gate-A plan working record Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../codex-reviews/gate-a-plan-om0bdd7udh-resume.md | 12 +++++++++--- 1 file changed, 9 insertions(+), 3 deletions(-) diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md index ed33539..f93bd9d 100644 --- a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md @@ -13,10 +13,14 @@ Nothing depends on it; the pass files and the repo are authoritative where this ## Resume here -**Next action:** Gate-A plan pass 6 against `33cdfa8`. The prompt is -`.context/gate-a-plan-prompt.md` — current at that commit; substitute `__SHA__` and `__P__`, +**Next action:** Gate-A plan pass 8 against `54e793b`. The prompt is +`.context/gate-a-plan-prompt.md`; substitute `__SHA__` and `__P__`, precheck per its header, delete the target file, confirm it is gone. +**Three rules pass 7 installed, each of which had been contradicted in two or three places at +once.** Check them every pass: a row a task must derive is derived **before** that task's install +step; **no task pre-assigns a fragment id**; the fragment table holds **OLD fragments only**. + **That prompt file is untracked.** `.gitignore` carries `.context/*` with only `codex-gate.on` and `codex-reviews/` exempt, so it survives a context clear but not a `.context/` cleanup. **If it is gone, rebuild it from this record** — the pass history below, the settled-and-not-open blocks and the @@ -59,7 +63,9 @@ worth checking before a pass rather than after. | 3 | b767a2a | 32→**31** | 0→**0** | 25→**23** | yes | one tell (instrument cluster). **Pass 2's three fragments came back — because I repaired the table and left the same fragments quoted inline in the task steps.** Second-copy defect, in the plan written to avoid it. Structural repair: the table is the only authored copy and Task 0 **generates** the shell variables from it, so no step can restate a fragment. 8 Minors collected | | 4 | 6ace06f | 31→**20** | 0→**10** | 23→**7** | yes | **MANDATORY TWO-TELL STOP** — Blockers rose 0→10, and the findings cluster on the instrument for the fourth pass running. Surfaced, standing answer applied, loop continued. **The awk generator pass 3 introduced was broken three ways**; it is deleted, not debugged — the helper is transcribed by hand and validated against the real files, with empty-string guards, because a silently unset variable makes `grep -cF ""` match every line. **A fourth fragment (F4) was preserved in its own replacement and invisible to my checker**, which compared without normalizing the block's line breaks | | 5 | 58b3660 | 20→**16** | 10→**7** | 7→**7** | yes | **MANDATORY TWO-TELL STOP** — instrument cluster for the fifth pass, plus a require↔withdraw pair (pass 4 demanded OLD rows for clauses pass 5 classifies as add-only). Surfaced, standing answer applied, loop continued. **The pre-written shell is deleted.** Design §7 says the plan builds each pair *against the real files*; five passes of findings were blocks written in advance for text that does not exist yet. One stated procedure replaces them | -| 6 | — | — | — | — | **not run — NEXT** | against `33cdfa8`. Prompt: `.context/gate-a-plan-prompt.md`, substitute `__SHA__`=`33cdfa8` and `__P__`=`6` | +| 6 | 33cdfa8 | 16→**11** | 7→**2** | 7→**6** | yes | one tell only (instrument cluster), no mandatory stop. **Task 11's sweep locator matched only `Gate B not satisfied` — 3 hits — and missed the 25 assertions greping the bare verdict word**, which §F items 15/16 remove: the suite would have gone red and the battery could not have passed. Task 15 step 5 deferred the provenance line, curve and human-exception record to an action step 7 did not contain. Passage (h) counted 24 kept less `h4`/`h19` where `h5` is replaced too, and the untouched middle span enclosed F7b's line. **Five tasks derived OLD fragments after their own install step.** Task 8 rebuilt fourteen rows the table already held as fifteen, dropping F7b. §G's semantic membership test had no observation; the strict-reading tail's add-only clauses were told to produce OLD rows they cannot have; Task 10 promised seven pairs and listed six | +| 7 | 34250be | 11→**13** | 2→**4** | 6→**4** | yes | **MANDATORY THREE-TELL STOP** — findings rose, Blockers rose, instrument cluster for the seventh pass. Surfaced, standing answer applied, loop continued. **Two of the four Blockers were pass 6's own repairs half-applied**: Task 4 kept its post-install derivation beside the new pre-install one, and Task 3's "ids continue the `P` series" collided with Task 4's pre-assigned `P19`/`P20`. The other two are five-pass survivors — `git log "$BASE"..HEAD` standing in for an ancestry test, and the single-line squash-carry site inside a `sed` range the same task forbids two paragraphs earlier. **Seventh condition-table misclassification in seven passes:** §B was credited with a second added rule that lives in §A | +| 8 | — | — | — | — | **not run — NEXT** | against `54e793b`. Prompt: `.context/gate-a-plan-prompt.md`, substitute `__SHA__`=`54e793b` and `__P__`=`8` | ## Pass-1 report From ebb371be9547cc9689535687fac2358ea2b62c34 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 16:44:07 +0200 Subject: [PATCH 110/181] docs(plans): apply Gate-A plan pass 8; observe the moved and carried conditions MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit MANDATORY TWO-TELL STOP — findings rose 13 -> 15 and the findings cluster on the verification apparatus for the eighth pass running. Blockers fell 4 -> 2. Surfaced to Daniel; his standing answer applies and the loop continued on it. Two Blockers and nine Majors, all validated against the real files. - Task 15 step 7 never re-ran the quality battery or the twelve-item reader check after a Gate-B fix. A repair to the hook, its test or either prompt copy could reach the closing commit on a tree that never passed the battery. - Task 15 step 8 suppressed the record commit with `|| true` and left the status and log lines as observations, so a failed commit, a dirty tree or a closing message still reading WIP proceeded through the soft reset and deleted the recovery base. Every check fails the script now, and the one harmless case `|| true` was hiding is tested for explicitly. - Two more condition-table misclassifications, the eighth pass in eight to find one: `a1` is carried inside §F item 8a's replacement block, not kept — and the floor-arithmetic untouched span opened on that very line, so a correct implementation would have failed its own check. `e10` is kept OUTSIDE §D, not carried inside it. - `c10`-`c13` and `a18`-`a20` are dispositioned moved and had only a reader walk; a move owes an absence check at the source. `c9`'s split owed both halves. Nine carried conditions owed preservation checks and three had them. - Task 10 derived eight hook OLD fragments and appended none, so a partial rerun could pick different ones and Task 15 could not audit them. - Task 14 treated §H as one section where it replaces nine noncontiguous sites; the unit is a destination block now. - Task 0 step 1 confirmed neither the branch nor the approved artifacts, and carried a stale "6ace06f or later". - Task 5's NEW fragment need not have contained the pronoun whose alignment it was the observation for. Minors repaired rather than collected, each a dangling reference in text pass 7 had just added: "this task's evidence" had no destination, which is now a defined output section; step 7b demanded all four records where the fourth is conditional; the evidence entry omitted the absence checks; and the version-bump step ran against a possibly stale base ref. Trend: findings 32, 32, 31, 20, 16, 11, 13, 15. Blockers 1, 0, 0, 10, 7, 2, 4, 2. Majors 28, 25, 23, 7, 7, 6, 4, 9. Floor 3, from risk high (2) and security none (0), read at this pass from docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-plan-om0bdd7udh-pass-8.md | 16 ++ .../2026-09-14-loop-rule-consolidation.md | 261 ++++++++++++++---- 2 files changed, 228 insertions(+), 49 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-8.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-8.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-8.md new file mode 100644 index 0000000..ea688be --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-8.md @@ -0,0 +1,16 @@ +MAJOR | high | Task 0 step 1 | The step titled “Confirm the approved artifacts” only prints the latest commit and accepts “6ace06f or later on loop-rule-consolidation”; it neither reports the branch nor establishes that the approved target commit and the reviewed plan revision are present in the current history. | A clean checkout on another branch or on a later history missing one approved artifact can record that checkout as `$BASE`, after which every install and the final soft reset operate on the wrong starting tree. | Verify the current branch and establish the approved target commit plus the Gate-A-reviewed plan revision against `HEAD` before writing the base file; make a mismatch stop the task. +MAJOR | high | Passage (a) disposition, `a1` | `a1` is marked kept and untouched, but Task 8 installs target §F item 8a’s complete fenced sentence, whose opening is `a1`; elsewhere the plan correctly calls unchanged conditions reproduced inside replacement blocks carried. | The 135-condition accounting is internally inconsistent and the untouched-range evidence claims a condition is outside an edited block when the approved replacement actually carries it. | Mark `a1` carried inside item 8a and make its preservation check follow the same carried-condition rule used for `a15`, `a21`, `a22`, `c15`, and `i4`–`i8`. +MAJOR | high | Passage (e) disposition and Task 5 step 4 | The disposition says both `e9` and `e10` are carried inside target §D, while the fenced `e7` replacement ends with `e9` and does not contain the following sentence `e10`; Task 5 then calls both conditions untouched. | Acceptance criterion 5 has two incompatible dispositions for `e9` and the wrong disposition for `e10`, so an executor cannot tell which text belongs to the replacement and which text must remain outside it. | Mark `e9` carried in the replacement and `e10` kept outside it, then make Task 5’s condition check use those same classifications. +MAJOR | high | Task 4 steps 1–5 | The task gives pairs to `c4`, `c8`, and `c14` and a post-install uniqueness check to `c9`, but `c10`–`c13` are independently dispositioned as moved and receive only a reader walk; this contradicts Task 3’s rule that a condition walk is not the discriminating observation design §7 requires. | The mechanical check can pass while one or more old closure instructions in `c10`–`c13` survive beside their ordering replacements, and the closing evidence can claim the required pairs without having them. | Before step 2, derive an OLD row for each moved condition not already observed, then pair each with the matching installed ordering clause; give `c9` the same parent/worktree observation rather than only a current-tree total. +MAJOR | high | Task 5 steps 1–4, `e8` | `e8` is correctly dispositioned as a meaning-changing W-only alignment, but the only pair is the `e7` threshold pair; its OLD boundary disappears when the new read-order clause is inserted and its freely chosen NEW fragment need not contain the added pronoun. | W can retain `report the tells` while the `e7` pair and pointer presence check pass, leaving the required counterfactual for the `e8` alignment to a later parity judgement rather than the check design §7 requires. | Add a W-specific OLD/NEW observation for `report the tells` becoming `you report the tells`, derived and confirmed before Task 5 installs the sentence. +MAJOR | high | Task 7 steps 1b–4, `a17`–`a22` block | The plan says the seven blocks outside the three expanded cases change one thing each, but this block moves or replaces `a17`, `a18`, `a19`, and `a20`; P11 observes only `a17`, while step 4 checks only the carried `a21` and `a22`. | Old loop-until-clean, zero-finding, or no-padding instructions can survive beside the new ordering while every Task 7 pair and carried-condition count still passes. | Treat `a18`, `a19`, and `a20` as separate changed conditions, derive their OLD rows before step 2, and run same-edit observations against the corresponding ordering text. +MAJOR | high | Task 7 steps 1b–4, carried conditions | The task’s survival check names `a15`, `a21`, and `a22` only; it never observes carried `c15` or `i4`–`i8`. P8 is taken from later surfacing clauses, and P16 observes only the changed boundary after `i8`, so neither proves the carried words remain. | Both copies can omit the surfacing premise or any interior strict-reading item and still satisfy every stated pair, add-only presence check, and parity comparison. | Add explicit preservation checks against the parent for `c15` and each of `i4`–`i8`, using per-condition fragments where one block-level fragment cannot isolate them. +MAJOR | high | Task 10 step 2b and Verification fragments procedure | The global procedure requires every execution-derived OLD fragment to be validated, appended under the next free `P` id, and kept in the OLD-only table, but Task 10 derives eight hook OLD fragments without appending any row; the table prose also names only Tasks 3, 4, 6, and 7 as appenders. | The hook pairs have no durable, single authored OLD values, so a partial rerun can choose different fragments and Task 15 cannot audit which exact counterfactual produced the recorded counts. | Require Task 10 to append its eight validated OLD rows before installation using the next free ids, and update the table’s appender list without preassigning ids. +MINOR | high | Tasks 1, 3, 4, 6, 7, and 8 fragment evidence | These tasks say to record chosen NEW fragments and counts “in this task’s evidence,” and Task 1 says that evidence edits this plan, but the plan defines no destination, schema, or idempotent replacement boundary for those records. | Executors can scatter or duplicate evidence on rerun, and Task 15 has no deterministic source from which to assemble every pair and presence count into the closing message. | Define one non-fragment-table output section keyed by task and condition, with an idempotent replacement rule for NEW fragments and observed counts. +MAJOR | high | Task 14 step 1 | The recipe requires one changed-site row per target section marked NEW or REPLACED, but target §H contains nine fenced replacements at nine noncontiguous live sites; one section-level start/end region cannot represent them, while the later instruction separately requires every site named by §§A–H to have a region. | A literal execution either cannot construct the §H row or emits one sampled region and leaves other changed prompt sites outside the final parity diff. | Define the unit as each destination fenced block or live site rather than each top-level target section, so all nine §H sites receive bounded regions without adding a fixed count. +MINOR | high | Task 15 step 5 | The closing evidence list requires every pair and every presence check, but omits the OLD-only absence checks used for dropped `g2`, `g3`, and `g4` wording. | The durable evidence can claim the verification set while leaving out the observations that prove obsolete stop and ownership instructions were removed. | Require the evidence entry to include every absence check and its parent/worktree counts alongside the pairs and presence checks. +BLOCKER | high | Task 15 step 7 | After Gate B finds a defect, the step only commits the fix, resolves a new head, revalidates the evidence entry, and re-reviews; it never explicitly reruns the quality battery or the prompt-standards reader check, and the latter is not part of the evidence entry step 5 defines. | A repair to the hook, test, prompt copies, or target text can make the final tree fail CI or invariant 11 after the only battery and twelve-item review have already passed, yet the plan can still reach the closing commit. | Before each re-review run the checks affected by the fix, and before closing rerun the full battery plus every reader check whose subject changed, recording fresh results for the new `HEAD`. +MINOR | high | Task 15 step 7b | Item 3 is conditional—carry any human-exception record and a skip reason only where one exists—yet the step then requires the file to “carry all four” numbered items before continuing. | An ordinary reviewed cycle with no exception and no skip cannot literally satisfy the check, so the executor must either invent a record, ignore the stated four-item oracle, or stop a valid close. | Require the three unconditional records and every applicable conditional record, or explicitly record `none` for item 3 and state that this satisfies the check. +BLOCKER | high | Task 15 step 8 | The pre-reset record commit suppresses every failure with `\|\| true`, and both `git status --porcelain` and the final `git log` are comments-only observations whose unexpected output does not stop `rm -f .context/loop-rule-base`. | A failed record commit, a dirty tree, or a closing message still beginning `WIP:` can proceed through the soft reset and delete the only recovery base after producing a closure state the plan itself says is invalid. | Distinguish the harmless no-changes case from a real commit failure, assert the required clean status and non-WIP closing commit, and remove the base only in the success branch after all assertions pass. +MINOR | high | Task 15 step 4 | The plan runs `scripts/check-version-bump.sh main` but never establishes the AGENTS.md precondition that the base ref is current. | A stale local `main` can yield a locally green “whole battery” for a version comparison different from the pull request’s actual base, deferring the failure to CI. | Confirm or update the intended base ref before the battery and record which current commit supplied the comparison. +END OF FINDINGS (15 total) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index c860a8a..5014b5f 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -122,11 +122,24 @@ **No task pre-assigns a row id, and none refers to a row by a number this plan does not already contain.** A task appending rows takes the **next free `P` id at the moment it appends**, and names its own rows by the condition they observe — "the `c14` row", "the resolve-duty row" — not by -a number chosen in advance. Tasks 3, 4, 6 and 7 all append, and how many each adds is decided at -execution against the real files, so any number written here ahead of time is a guess that the -task running before it invalidates. An earlier draft promised Task 4 the ids `P19` and `P20` while -Task 3 was told to continue the same series, which hands two different fragments one id and leaves -every later reference ambiguous. +a number chosen in advance. **Tasks 3, 4, 6, 7 and 10 all append**, and how many each adds is +decided at execution against the real files, so any number written here ahead of time is a guess +that the task running before it invalidates. An earlier draft promised Task 4 the ids `P19` and +`P20` while Task 3 was told to continue the same series, which hands two different fragments one id +and leaves every later reference ambiguous. + +**Task 10 appends too, and its rows are not optional.** Its eight OLD fragments come from the hook, +whose live wording exists in exactly one file and is gone after installation. A fragment derived, +used and never recorded leaves a partial rerun free to pick a different one, and leaves Task 15 +unable to say which counterfactual produced the counts it publishes. **The table is the one place +every OLD fragment lives, whatever file it came from.** + +**Where a NEW or presence fragment goes, since it does not go here.** The table is OLD-only. Each +task records its chosen NEW and presence fragments, with the four counts observed for each, under +`## Fragment evidence (per-task output)` at the end of this plan — **one subsection per task, +replaced idempotently on re-run by the same rule Tasks 13 and 14 use.** Task 15 step 5 assembles +the closing evidence entry from that section, so a fragment recorded anywhere else is a fragment +the closing commit cannot carry. --- @@ -152,7 +165,8 @@ Story acceptance criterion 5 is satisfied here. Ids are `docs/superpowers/specs/ | Condition | Disposition | |---|---| -| a1, a3–a12, a14 | **kept**, untouched. The floor arithmetic, the hook-ratio rules and the two comparison points are outside this change. | +| a3–a12, a14 | **kept**, untouched. The floor arithmetic, the hook-ratio rules and the two comparison points are outside this change. | +| a1 | **carried**, not kept. §F item 8a's fenced block opens with `**Both gates are a LOOP with a HARD FLOOR: a minimum number of passes per run` — `a1` itself — and reproduces it so one contiguous string installs. **Recorded as carried for the same reason as `a15`, `a21`, `a22`, `c15` and `i4`–`i8`**: a condition inside a replacement block is not untouched text, and calling it untouched would put it inside an untouched-range span that a correct implementation then fails. It owes a preservation check, not a span. | | a2 | **replaced.** §F item 8a rewrites the HARD FLOOR parenthetical `(Blocker/Major only)` — Task 8, row F10. The filter itself survives in the ordering, which states what a pass counting toward the floor must be; what goes is the parenthetical's claim that Blocker/Major is the *whole* of it. **Marked replaced rather than kept**, because a condition whose text the change in fact rewrites, recorded as preserved, is the dropped-condition failure `AGENTS.md` names. | | a13 | **replaced** — scoped to its own paragraph. §H's `a13` block. | | a15 | **carried** inside §H's `a16` block, which reproduces it so one contiguous string installs. | @@ -199,7 +213,8 @@ Story acceptance criterion 5 is satisfied here. Ids are `docs/superpowers/specs/ | e1–e6 | **kept** — the five tells themselves are untouched. | | e8 | **replaced.** The inventory records it as a parity divergence — C `you report`, W `report` — and §D's block supplies C's wording for **both** copies, so W's form goes. **Not carried**: the condition as inventoried names a difference this change removes. Task 5. | | e7 | **changed** — gains the read-after-clean-completion clause. The sentence is given entire in §D. | -| e9, e10 | **carried** inside §D's block. | +| e9 | **carried** inside §D's block — `the "clearly stuck" reading above is not a precondition for it` closes the replacement sentence and is reproduced in it. Owes a preservation check, not a span. | +| e10 | **kept**, untouched, and **outside** the replacement. `A loop can be worth stopping long before it plateaus.` is the sentence *after* §D's block; an earlier draft listed it as carried, which would have put a kept condition inside a replacement it never enters and left the real boundary of the edit unstated. | | e11 | **kept** — the C-only rationale paragraph is untouched and stays C-only. | | — | **added:** §D's pointer paragraph at the end of the passage. | @@ -249,7 +264,12 @@ Story acceptance criterion 5 is satisfied here. Ids are `docs/superpowers/specs/ - [ ] **Step 1: Confirm the approved artifacts and a clean tree** ```bash -git log --oneline -1 # expect 6ace06f or later on loop-rule-consolidation +git rev-parse --abbrev-ref HEAD # expect loop-rule-consolidation +git merge-base --is-ancestor ba15e83 HEAD || { echo "approved target text (ba15e83) is not in this history"; exit 1; } +test -f docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md || { echo "target text missing"; exit 1; } +test -f docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md || { echo "design missing"; exit 1; } +test -f docs/superpowers/specs/2026-09-10-loop-rule-consolidation-condition-inventory.md || { echo "inventory missing"; exit 1; } +git log --oneline -1 git status --porcelain # expect empty if [ -s .context/loop-rule-base ]; then echo "base already recorded: $(cat .context/loop-rule-base) — NOT overwriting" @@ -259,6 +279,14 @@ fi cat .context/loop-rule-base ``` +**The branch and the approved artifacts are established, not assumed.** An earlier draft printed +`git log --oneline -1` with the comment "expect 6ace06f or later on loop-rule-consolidation", +which names neither a branch the command reports nor a way to tell "later" from "a different +history". A checkout on another branch, or one whose history does not contain the approved target +text, would be recorded as `$BASE` and every install and the final soft reset would run from the +wrong starting tree. **A hard-coded commit is not repeated as a floor** — `ba15e83` appears once, +as the ancestry test for the approved text, because it is the commit that approval is *of*. + **Never overwrite an existing base, and never trust one you did not just write.** Re-running Task 0 after a partial implementation would record the current WIP tip, and both Gate B's range and the final `reset --soft` would then start after every edit made so far — prompt and hook changes would be @@ -318,7 +346,7 @@ for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do grep -n 'The two rules above do not compete' "$f" # (f) start grep -n 'Findings go to a FILE' "$f" # (f) end grep -n 'On squash-merge, copy every evidence entry' "$f" # (j) single line - grep -n 'Both gates are a LOOP with a HARD FLOOR' "$f" # (a) arithmetic start + grep -n 'Both gates are a LOOP with a HARD FLOOR' "$f" # (a) region marker — NOT a span start, see below grep -n 'Nothing here writes the floor knob' "$f" # (a) arithmetic end grep -n 'Recording a human exception' "$f" # (h) start grep -n 'because writing it down makes it sound' "$f" # (h) end @@ -337,10 +365,13 @@ end anchor in each tree at comparison time. implementation fails its own check — and each split must leave **every kept condition inside some span**, which the first draft's version did not: -- **The floor arithmetic is three spans**, because `a2` sits at its head and `a13` in its middle: - from `Both gates are a LOOP with a HARD FLOOR` to the line before row F10's fragment; from the - line after F10's to the line before row P9's; and from the line after P9's to `Nothing here writes - the floor knob`. **The middle span is the one the first draft dropped**, and it holds `a3`–`a12`. +- **The floor arithmetic is two spans, not three, and neither starts at `Both gates are a LOOP`**: + `a1` and `a2` are both inside item 8a's replacement block, which begins at that anchor, so a span + opening there opens on a changed line. Take the spans as: **from the line after row F10's + fragment** to the line before row P9's; and from the line after P9's to `Nothing here writes the + floor knob`. The first of them holds `a3`–`a12` — **the span an earlier draft dropped entirely** + — and the second holds `a14`. `a1` and `a2` are covered by their own observations, not by a span: + `a2` by row F10's pair, `a1` by the carried-condition preservation check. - **The human exception is three spans**, around rows F7 (`h4`), **F7b (`h5`)** and F4 (`h19`): from `Recording a human exception` to before F7's line; **from after F7b's line** to before F4's; and from after F4's to `because writing it down makes it sound`. **Twenty-two** of the @@ -639,9 +670,9 @@ Expected: `P4=1` for both files. is **preserved in target §C's fenced replacement**, so its old-count could never reach zero. Derive `c4`'s OLD from the part of the sentence the replacement removes. -**`c4` and `c8` have no row yet, and both are derived here rather than at step 4.** Take each from -the live passage, check it the three ways, confirm it counts 1 in each copy, and **add a row to -the fragment table** under the next free `P` id — that table is where every fragment lives, and a +**`c4`, `c8`, `c9` and the four moved conditions `c10`–`c13` have no row yet, and all of them are +derived here rather than at step 4.** Take each from the live passage, check it the three ways, +confirm it counts 1 in each copy, and **add a row to the fragment table** under the next free `P` id — that table is where every fragment lives, and a row that is not there is a fragment stated somewhere else. Refer to them afterwards as **the `c4` row** and **the `c8` row**, never by a number picked here: Task 3 appends first and how many rows it adds is decided at execution. **Step 2 removes the wording they are taken from**, so a @@ -680,13 +711,26 @@ observation at all. **A later draft reintroduced exactly that pair** by naming r Each NEW is confirmed single-line and unique in the installed file before it is counted. +**Plus an absence check per moved condition — `c10`, `c11`, `c12`, `c13` — and one for `c9`'s moved +clause.** A move removes wording here and adds it in §A, so the source removal is an **absence**: +count the condition's own text in each copy, expecting `parent=1 worktree=0`. Task 1's §A presence +checks are the other half. **A reader walk is not this observation** — Task 3 states the rule and +the same reason applies: without it an old closure instruction can survive in passage (c) beside +its §A replacement, and every pair, count and parity diff still passes. This is the +two-instructions-that-disagree failure, and `c10`–`c13` are four chances at it. + +**`c9` is split, so it owes both halves:** the moved precedence clause is absent here +(`parent=1 worktree=0`) and present in §A, while the plateau rationale stays — confirm the +rationale still counts `1` in each copy. + *(Build each pair per the verification procedure; record the four values.)* -Expected for all six pair instances — three pairs in each of the two copies: -`old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`. **All three OLD rows were derived and -validated at step 1**; this step only runs them. +Expected: for the three pairs, six pair instances reading +`old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`; for the five absence checks, +`parent=1 worktree=0` in each copy. **Every OLD row was derived and validated at step 1**; this +step only runs them. -- [ ] **Step 5: Walk the conditions** — `c1`–`c3` present unchanged, `c5`–`c7` word for word, `c8` carrying the new clause, `c9` split, `c10`–`c14` gone from here. +- [ ] **Step 5: Walk the conditions** — `c1`–`c3` present unchanged, `c5`–`c7` word for word, `c8` carrying the new clause, `c9` split, `c10`–`c14` gone from here. **This is a reader's confirmation on top of step 4's counts, not the observation for any of them** — `c9`–`c14` are each counted there, and a walk that found what the counts missed would mean a fragment was wrong rather than that the walk was the check. - [ ] **Step 6: Commit** @@ -738,14 +782,30 @@ is add-only** — it replaces no wording — so it is checked by presence alone, it can be omitted from both copies while the `e7` pair, the condition walk and the parity diff all pass. +**Take both copies' NEW fragment from the clause carrying the pronoun** — the installed +`you report the tells` — rather than from any other part of §D's sentence. P5w's OLD going to zero +shows W's pronoun-less form is gone; only a NEW fragment containing the pronoun shows W received +C's form rather than some third wording. **That is the `e8` alignment's own observation**, and +without it the alignment rests on Task 14's parity judgement instead of on a count. + *(Build the pair per the verification procedure; record the four values.)* Expected: both pairs `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`; both pointer checks `worktree=1 parent=0`. -- [ ] **Step 4: Confirm `e1`–`e6` and `e9`–`e11` are untouched**, that `e11` is still C-only, and -that **`e8` is aligned rather than untouched** — this task gives W the pronoun, so listing `e8` -among the untouched conditions would contradict the task's own instruction. +- [ ] **Step 4: Walk the conditions by their three different dispositions** + +They are not one class and the check differs per class: + +- **`e1`–`e6`, `e10`, `e11` are kept and outside the replacement** — confirm each is unchanged + against the parent, and that `e11` is still C-only. **`e10` belongs here, not with `e9`**: it is + the sentence after §D's block, so a check treating it as carried would look for it inside text + it never enters. +- **`e9` is carried inside §D's block** — confirm `the "clearly stuck" reading above is not a + precondition for it` is present in each copy after the install, expecting `1`. A carried + condition is reproduced rather than untouched, so a span cannot protect it. +- **`e8` is aligned rather than untouched** — this task gives W the pronoun, so listing `e8` among + the untouched conditions would contradict the task's own instruction. - [ ] **Step 5: Commit** @@ -871,6 +931,13 @@ so a derivation afterwards has no live wording left to check against: - **The `c18`-and-surfacing block** changes `c16`, `c17`, `c19` and `c20` besides the `c18` clause row P8 observes — **one OLD row per condition**, four in total. +- **The `a17`–`a22` block moves four conditions, not one.** P11 observes `a17`, the clean-final-pass + rule; `a18`, `a19` and `a20` are each independently dispositioned **moved** and have no + observation at all. **One OLD row each**, and each is an *absence* at this source — + `parent=1 worktree=0` — paired with the matching §A text Task 1 already installed, since a moved + condition leaves here and appears there. Without them the old loop-until-clean, zero-finding and + no-padding instructions can survive beside the ordering that replaces them while every pair and + both carried-condition counts still pass. - **§G changes two things, not one**: its membership is widened *and* it gains a semantic membership test a downstream reader can apply, and target §G's own closing note names the test as the point of the replacement. P7 observes the widened clause; **the test owes a second @@ -925,14 +992,27 @@ semantic membership test and for every independent add-only clause in the strict Choose each of those NEW fragments from the installed text and verify it the three ways before counting it. -- [ ] **Step 4: Confirm `a21`, `a22`, `a15` survived** — they are carried inside blocks that install contiguously, so a mis-scoped replacement silently drops them. +- [ ] **Step 4: Confirm every carried condition in this task's blocks survived** + +They are reproduced inside blocks that install contiguously, so a mis-scoped replacement silently +drops them — and because they are carried rather than kept, **no untouched-range span covers +them**, which is exactly why the disposition records them as carried. **Each owes a count of its +own text in each copy, expecting `1`.** + +The set is **`a15`, `a21`, `a22`, `c15`, and each of `i4`–`i8`** — nine conditions, not three. An +earlier draft checked three and left the rest to P8 and P16, which observe the *changed* clauses +around them and prove nothing about the reproduced words: both copies could omit the surfacing +premise, or any interior item of the strict-reading run, and every pair, presence check and parity +comparison would still pass. ```bash grep -cF 'Codex is advisory — validate before applying; dismissed finding → one-line why' CLAUDE.md grep -cF 'Open a TodoWrite' CLAUDE.md ``` -Expected: `1` each. +Expected: `1` each, and `1` for each of the remaining seven — take each fragment from the +condition's own text in the inventory, confirm it is single-line and unique, and run it in both +copies. - [ ] **Step 5: Parity** for all ten sites, each extracted by its own bounded region rather than a fixed line window. @@ -990,9 +1070,16 @@ every stated count still passed. Expected for all thirty pair instances — fifteen rows in each of the two copies: `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`. -- [ ] **Step 4: Count what was installed** +- [ ] **Step 4: Confirm `a1` survived, then count what was installed** + +**`a1` is carried inside item 8a's block**, which opens with it — `**Both gates are a LOOP with a +HARD FLOOR: a minimum number of passes per run` — and reproduces it so one contiguous string +installs. Carried, not kept: **no untouched-range span covers it**, and row F10's pair observes +`a2`, the parenthetical, not the opening it sits in. Count `a1`'s own text in each copy, expecting +`1`. Without it a mis-scoped item-8a replacement can drop the sentence's opening and every other +check in this task still passes. -Total the fifteen `new/worktree` values step 3 printed, per copy. +Then total the fifteen `new/worktree` values step 3 printed, per copy. Expected: `15` per copy — **fifteen independent changed clauses, from fourteen installed items**. State the number you observed. **Do not carry a count from §F into a check** — §F states the count of falsified sentences and this task installs a subset of them; a count copied between the two is the stale-bookkeeping defect the design records at five passes running. @@ -1099,7 +1186,10 @@ Step 5 runs eight pairs. **Their OLD halves come from the live `note` strings, s here** — after step 3 the old wording is gone from the only file that carries it, and a fragment reconstructed from the parent tree can no longer catch a drifted string or a half-applied re-run. For each, take a single-line fragment from the located `note` call, confirm it counts **1** in -`codex-gate.sh`, and confirm it is **not preserved inside its own §F replacement**. +`codex-gate.sh`, and confirm it is **not preserved inside its own §F replacement**. **Then append +all eight to the fragment table** under the next free `P` ids, naming each by the item it observes. +The hook's live wording exists in one file and is gone after step 3; a fragment used and never +recorded cannot be re-derived and cannot be audited. **Three of the eight need a second look while the file is open.** Item 11's removed text is the **clean definition** at the end of the Gate-A satisfied message, not `floor met by COUNT ONLY`, @@ -1421,15 +1511,22 @@ the human-exception edits, the WIP and finishing-cycle text, the evidence-revali the curve rationale — every one of which this change ships, and any of which could differ between the copies while a nine-window diff passes. -**Drive the list from the artifact:** every section the target text marks NEW or REPLACED, and every -item in §F whose destination is a prompt copy. For each, extract a **bounded** region — from its -first line to the first line of the next passage, not a fixed count — and diff the two copies: +**Drive the list from the artifact, and the unit is a destination site, not a target section.** +One row per **fenced block whose destination is a prompt copy** — so §H contributes **nine** rows, +one per live site it replaces, not one row for the section. Its nine blocks land at nine +noncontiguous places in each copy and no single start/end region spans them; a section-level row +would either be unbuildable or would sample one of the nine and leave the other eight out of the +final parity diff. §A, §B, §C, §E and §G are contiguous and contribute one row each; §D +contributes two; every §F item whose destination is a prompt copy contributes one. + +For each, extract a **bounded** region — from its first line to the first line of the next passage, +not a fixed count — and diff the two copies: ```bash # .context/loop-rule-changed-sites is written by this step: one -# startend line per section the target marks NEW or REPLACED and per -# §F item whose destination is a prompt copy. Build it by reading those -# markers off the target text, then: +# startend line per destination site — per fenced block the target +# marks NEW or REPLACED, and per §F item whose destination is a prompt +# copy. Build it by reading those markers off the target text, then: while IFS=$(printf '\t') read -r s e; do echo "== $s" diff <(sed -n "/$s/,/$e/p" CLAUDE.md) \ @@ -1443,7 +1540,8 @@ without having diffed anything. **Fail this step where a site named by §§A–H has no region in your list** — that is the same completeness failure as Task 13's per-condition checks, and it is caught the same way: by reading -the set off the artifact rather than from a list kept here. +the set off the artifact rather than from a list kept here. **Read it block by block**, which is +what makes §H's nine sites visible; reading it section by section is what hid eight of them. **Where it goes:** this plan, under the heading `## Divergence list (Task 14 output)` at the end of the document, replaced idempotently on re-run by the same rule Task 13 uses. @@ -1547,6 +1645,19 @@ claude plugin validate . --strict Expected: exit 0. +**`check-version-bump.sh main` has a precondition the rest of the battery does not**, and AGENTS.md +states it: it compares *commits* against a base ref, so a stale local `main` compares the bump +against a different base than the pull request will. **Fetch the base ref first and record which +commit supplied the comparison:** + +```bash +git fetch origin main +git rev-parse origin/main # record this — it is what the local run compared against +``` + +Run the battery's version-bump step against a current `main`. A green run against a stale one is +green about the wrong comparison and defers the failure to CI. + - [ ] **Step 4b: Apply all twelve `docs/prompt-standards.md` items to the installed text, and commit the result before Gate B** **Nothing mechanical does this and no other task claims it.** The battery's three narrow checks are @@ -1582,7 +1693,11 @@ in the WIP body too if a mid-cycle reader would want it, but the file is the cop **human-exception record** belong beside it. **Step 7b appends all three and revalidates the entry** — this step opens the file, step 7b closes it, and step 8 commits it. -It names: the battery run; **every pair this plan built, with its counts in each copy and each tree, and every presence check beside them**; the §6 parity diff and the `b11`/`b13` equivalence result; and the next-state table's location plus its row count. **State the §A presence checks as presence, not as pairs** — its counterfactual is absent and is claimed as absent. +It names: the battery run; **every pair this plan built, with its counts in each copy and each tree, every presence check beside them, and every absence check with its two counts**; the §6 parity diff and the `b11`/`b13` equivalence result; and the next-state table's location plus its row count. **Read them out of `## Fragment evidence (per-task output)`**, which is where every task records them. + +**The absence checks belong in the entry as much as the pairs do.** `g2`, `g3` and `g4` are dropped rather than replaced, and `c9`–`c13` and `a18`–`a20` are moved, so the observation that proves each is gone is an absence — `parent=1 worktree=0` — and nothing else in the entry carries it. An entry listing only pairs and presence checks claims the verification set while omitting the half that proves obsolete instructions were removed, which is the failure the absence checks exist for. + +**State the §A presence checks as presence, not as pairs** — its counterfactual is absent and is claimed as absent. - [ ] **Step 6: Run Gate B** @@ -1608,6 +1723,8 @@ Floor derives from the story profile: risk `high` → 2, security `none` → 0, WIP tip while the repair sits in the worktree — and the final squash then publishes a fix no pass reviewed: +**Before committing each fix, re-run what the fix could have broken**, then: + ```bash git add -A && git commit -m "WIP: fix " git rev-parse HEAD # resolve headSha fresh for the next call @@ -1618,6 +1735,20 @@ closing commit.** **A fix that changes specified behaviour updates the spec in the same commit.** +**The battery and the reader checks are not done when step 4 passed once.** A Gate-B fix can touch +the hook, its test, either prompt copy or the target text, and step 4's run and step 4b's +twelve-item review both describe the tree as it was *before* that fix. Without re-running them the +plan reaches its closing commit on a tree that never passed its own quality battery — a repo unable +to pass CI, or an invariant-11 violation, published by a cycle that closed clean. + +- **After each fix, re-run the checks whose subject it changed** — the hook or its test means the + suite under both shells; either prompt copy means `sh scripts/check-invariants.sh`; a shell file + means `shellcheck`. +- **Before the closing act, re-run the whole battery of step 4 against the current `HEAD`**, and + re-apply step 4b's twelve items to every artefact a fix touched, recording the fresh result in + the plan and committing it. **The results the closing commit carries are the ones from the final + tree**, not from the tree that first went green. + - [ ] **Step 7b: After the clean pass, complete `.context/loop-rule-closing-msg` — this is the action step 5 defers to, and step 8 has no other source for these records** @@ -1631,10 +1762,15 @@ Append to the file step 5 opened, then read the whole file back before step 8 ru 3. any **human-exception record**, and beside it the skip reason if a cycle was skipped; 4. the **revalidated evidence entry**, replacing step 5's draft if revalidation changed it. -Confirm the file now carries all four before continuing. **A missing one is not recoverable after -step 8** — `git commit -F` publishes whatever the file holds, `reset --soft` has already discarded -every WIP body, and a closing commit without its provenance line or curve has not validly closed -the cycle. +**Items 1, 2 and 4 are owed unconditionally; item 3 is owed only where such a record exists.** +Confirm the file carries the three, and either the applicable exception records or **the literal +line `Human exception: none`**, which is what satisfies this check for an ordinary cycle. Without +that line an executor reading "all four" must either invent a record or ignore the oracle, and a +valid close stops on a record nobody owed. + +**A missing one is not recoverable after step 8** — `git commit -F` publishes whatever the file +holds, `reset --soft` has already discarded every WIP body, and a closing commit without its +provenance line or curve has not validly closed the cycle. - [ ] **Step 8: Close the cycle** @@ -1648,20 +1784,36 @@ test -n "$BASE" || { echo "BASE empty — Task 0 did not run"; exit 1; } # Everything that must be IN the squashed commit has to be committed before the # reset: reset --soft stages only what the discarded commits already contained. git add docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md .context/codex-reviews/ -git commit -m "WIP: plan records and Gate-B findings files" || true -git status --porcelain # expect empty +if ! git diff --cached --quiet; then + git commit -m "WIP: plan records and Gate-B findings files" || { echo "record commit FAILED"; exit 1; } +fi +test -z "$(git status --porcelain)" || { echo "tree not clean before reset — aborting"; exit 1; } + git reset --soft "$BASE" git commit -F .context/loop-rule-closing-msg || { echo "closing commit FAILED — base file kept for recovery"; exit 1; } -git log --oneline -1 # expect the closing message, not a WIP -git status --porcelain # expect empty + +case "$(git log -1 --pretty=%s)" in + WIP:*|wip:*) echo "closing commit still reads WIP — cycle NOT closed"; exit 1 ;; +esac +test -z "$(git status --porcelain)" || { echo "worktree dirty after close — aborting before cleanup"; exit 1; } rm -f .context/loop-rule-base ``` -**The base file is removed only after the closing commit succeeded**, and the two checks above are -what "succeeded" means here. `git commit` can fail on a hook, a signing key or an unset identity, -and at that point the reset has already happened: the WIP commits are gone, the whole change is a -staged tree, and `.context/loop-rule-base` is the only record of where the cycle started. Deleting -it unconditionally destroys the one value a rerun needs, in the single state where it is needed. +**Every check in that block fails the script; none of them is a comment.** An earlier draft +suppressed the record commit with `|| true` and left the status and log lines as things to look at, +so a failed record commit, a dirty tree or a closing message still reading `WIP:` all proceeded +through the soft reset and deleted the recovery base — producing, silently, the exact state the +plan elsewhere calls an invalid close. + +**`|| true` is replaced by a test for the one harmless case it was hiding.** "Nothing to commit" +and "the commit failed" are different outcomes and only the first is fine, so the block asks +`git diff --cached --quiet` first and treats a real failure as a failure. + +**The base file is removed only in the success branch.** `git commit` can fail on a hook, a signing +key or an unset identity, and at that point the reset has already happened: the WIP commits are +gone, the whole change is a staged tree, and `.context/loop-rule-base` is the only record of where +the cycle started. Deleting it unconditionally destroys the one value a rerun needs, in the single +state where it is needed. **`$BASE` is the recorded revision, not a placeholder to substitute by hand**, and the file is removed afterwards so a later run cannot inherit a stale one. @@ -1716,3 +1868,14 @@ The closing body carries: the validated evidence entry; the provenance line; the ## Prompt-standards result (Task 15 step 4b output) *Empty until Task 15 runs. It replaces this entire section, one line per checklist item.* + +--- + +## Fragment evidence (per-task output) + +*Empty until the tasks run. One subsection per task — `### Task 1`, `### Task 3`, … — each holding +that task's chosen NEW and presence fragments with the four counts observed for each, and the +absence checks with their two. A task replaces its own subsection idempotently on re-run and +touches no other. **The OLD fragments live in the fragment table, never here**; this section is +where the halves that do not exist until installation are recorded, and it is what Task 15 step 5 +reads to assemble the closing evidence entry.* From ba614f68fe40aea7907141c544dfed304cb7bf06 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 17:03:27 +0200 Subject: [PATCH 111/181] docs(plans): apply Gate-A plan pass 9; state what each disposition owes, once MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit MANDATORY TWO-TELL STOP — Blockers flat at 2 (failing to fall) and the findings cluster on the verification apparatus for the ninth pass. Findings fell 15 -> 9. Surfaced to Daniel; standing answer applied, loop continued. Two Blockers and six Majors, all validated against the real files. The root repair is one table, not eight patches. Four consecutive passes found a condition whose check did not match its disposition, each in a different passage, because every task was inventing the rule for its own conditions. "What each disposition owes" now states it once -- kept, carried, replaced, moved, dropped, add-only -- and the tasks consult it. - The second floor-arithmetic span ended on C 131 / W 338, which carries both the tail of changed a13 and the start of kept a14: a correct implementation fails its own untouched check. The hand-written spans are gone; the plan states the derivation instead, because writing them out has now been wrong twice and neither error is visible without the file open. Task 2 also never ran the per-condition list it said it consumed. - Task 15 step 7's re-run set named only mechanical checks, so a Gate-B fix could invalidate Task 12's equivalence, Task 12b's sweep, Task 13's transitions or Task 14's parity and those records still reached the closing commit. It could also commit refreshed records after the clean pass, closing on a tree no pass reviewed. - Nine carried b conditions and carried c5-c7 had only reader walks. - c10-c13 and a17-a20 were given Task 1's three paragraph-level presence fragments as their destination half; those pass while any one moved predicate is missing from §A. Each moved condition gets its own §A presence check. - Task 6 derived one fragment for two separately dropped conditions -- it goes absent when either half goes -- and gave g4 no absence check at all. - Task 7 step 1b derived the a18-a20 absence rows and step 3 never ran them. - Task 12 persisted its equivalence result nowhere and committed nothing; it has an output section now. Task 15 step 5 read every reader record from the fragment-evidence section, which holds none of them. - The battery ran check-version-bump.sh against local main while the fetch updated origin/main, and recorded origin/main as the comparison it supplied. The fragment table's rule is sharpened from "OLD only" to the cut that actually holds: fragments that exist BEFORE the edit -- OLD halves, absence fragments, carried and line-sharing kept preservation fragments. Only post-install fragments go to the evidence section. Minor repaired: the Self-Review listed Task 1 among the OLD-deriving tasks where §A is add-only and P1 records no OLD half. Trend: findings 32, 32, 31, 20, 16, 11, 13, 15, 9. Blockers 1, 0, 0, 10, 7, 2, 4, 2, 2. Majors 28, 25, 23, 7, 7, 6, 4, 9, 6. Floor 3, from risk high (2) and security none (0), read at this pass from docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-plan-om0bdd7udh-pass-9.md | 10 + .../2026-09-14-loop-rule-consolidation.md | 267 +++++++++++++----- 2 files changed, 204 insertions(+), 73 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-9.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-9.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-9.md new file mode 100644 index 0000000..6319481 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-9.md @@ -0,0 +1,10 @@ +BLOCKER | high | Task 0 step 2 and Tasks 2/14 untouched checks | The prescribed second floor-arithmetic span starts after P9's line and ends on the line containing the a14 anchor, but the real C line 131 and W line 338 each contain both the tail of changed a13 and the start of kept a14; this contradicts the plan's later acknowledgement that no whole-line span can isolate a14, and Task 2 never executes the per-condition list it says it consumes | A correct a13 replacement either makes the untouched span fail, or an executor omits that span and leaves a14 unchecked until the end | Remove the impossible second span, record a14 only as a condition-specific fragment with parent=1 and worktree=1, define its scratch-record shape, and execute that check in both Task 2 and Task 14 +MAJOR | high | Task 3 step 4 and Task 4 step 5 | The nine carried b conditions and carried c5-c7 receive only reader walks, although the plan's governing rule requires a preservation count for every carried condition and explicitly says the Task 4 walk is not their observation | A mis-scoped contiguous replacement can omit carried wording from both copies while every changed-condition pair and parity check still passes | Before each install record a unique fragment for every carried condition, then count its own text after installation in both copies with the required preservation result +MAJOR | high | Task 1 step 4 with Task 4 step 4 and Task 7 step 3 | Task 1 records one freely chosen presence fragment per §A paragraph, but Tasks 4 and 7 treat those three paragraph samples as the destination half for the independently moved c10-c13 and a17-a20 conditions | Any one moved predicate can disappear from its source and be omitted from §A while the generic paragraph samples, source absences, parity checks, and later self-derived tables still pass | Add and record a condition-specific §A presence check for each of c10-c13 and a17-a20, paired with that condition's source absence check +MAJOR | high | Task 6 steps 1, 3 and 4 | The task derives one combined g2/g3 absence fragment even though g2 and g3 are separately dropped conditions, and handles g4 only with pre/post worktree path counts rather than an OLD-table row and a recorded parent/worktree absence check | A combined fragment becomes absent when either half is removed, so the other obsolete instruction may survive; g4 also has no auditable evidence for the Task 15 entry that explicitly requires it | Before step 2 derive separate OLD rows for g2, g3, and C-only g4, then run and record a distinct parent=1 worktree=0 absence check for each +MAJOR | high | Task 7 step 3 | Step 1b derives source-absence rows for moved a18, a19, and a20, but step 3's instruction to run what step 1b added names only the four c conditions and the add-only presence checks | The old loop-until-clean, zero-finding, or no-padding instruction can survive at its source without any post-install parent=1 worktree=0 observation detecting it | Add the three a18-a20 source-absence checks to step 3 and record their two counts in Task 7's fragment evidence +MAJOR | high | Task 12 step 5 and Task 15 step 5 | Task 12 tells the reader to write down the full b11/b13 predicates but persists only an unspecified result to a future evidence entry, modifies no file, and never writes the Fragment evidence section that Task 15 later names as its source; Task 15 likewise points to that section for parity and next-state outputs that actually live elsewhere | The comparison's subject and reasoning can be lost between tasks, and the closing evidence can be assembled from a location that contains none of the required reader-check record | Persist the full predicates examined and both-direction result in a named plan output before Gate B, commit it, and make Task 15 read each reader result from its actual named section +MAJOR | high | Task 15 step 4 | The displayed battery runs check-version-bump.sh against local main before the fetch, while the subsequent fetch updates origin/main rather than local main and the record falsely says origin/main supplied the comparison | The local check can be green against a stale base even though the pull-request base differs, deferring an invariant-12 failure to CI | Fetch first and pass origin/main or the recorded fetched commit directly to check-version-bump.sh, so the recorded revision is the exact argument the successful check used +BLOCKER | high | Task 15 step 7 | After a Gate-B fix the plan names only selected mechanical checks, and before closing it reruns the battery plus prompt standards but not the affected Task 12 equivalence check, Task 12b sweep, Task 13 transition and closure checks, Task 14 parity and untouched checks, or the affected pair, presence, and absence observations; it can also commit refreshed plan results after the clean pass without requiring that new HEAD to be reviewed | Gate B can close on stale reader evidence or on a final recorded tree that the clean pass never reviewed, so the cycle does not establish the story criteria or its own verification claims | Before every candidate re-review rerun and persist every mechanical and reader check whose subject changed, run the complete final set before the candidate final pass, commit those records, resolve the new HEAD, and allow only a clean response for that exact HEAD to close +MINOR | high | Self-Review item 2 | The self-review says Task 1 derives remaining OLD fragments even though §A is add-only, P1 explicitly has no OLD, and Task 1 correctly records presence fragments only | The summary contradicts the global add-only rule and can send an executor looking for or fabricating an OLD half that cannot pass | Remove Task 1 from the OLD-derivation list and state that it derives only NEW presence evidence +END OF FINDINGS (9 total) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index 5014b5f..a1f1001 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -134,7 +134,13 @@ used and never recorded leaves a partial rerun free to pick a different one, and unable to say which counterfactual produced the counts it publishes. **The table is the one place every OLD fragment lives, whatever file it came from.** -**Where a NEW or presence fragment goes, since it does not go here.** The table is OLD-only. Each +**What the table holds, stated as one cut: fragments that exist *before* the edit.** That is the +OLD half of every pair, every **absence** check's fragment, and every **carried** or +line-sharing **kept** condition's preservation fragment — all of them checkable in advance, which +is what the table is for. **What it does not hold is anything that exists only after installation** +— the NEW half of a pair, and an add-only edit's presence fragment. + +**Where those go, since they do not go here.** Each task records its chosen NEW and presence fragments, with the four counts observed for each, under `## Fragment evidence (per-task output)` at the end of this plan — **one subsection per task, replaced idempotently on re-run by the same rule Tasks 13 and 14 use.** Task 15 step 5 assembles @@ -161,6 +167,26 @@ the closing commit cannot carry. Story acceptance criterion 5 is satisfied here. Ids are `docs/superpowers/specs/2026-09-10-loop-rule-consolidation-condition-inventory.md`, snapshot at `7c0d475`. **Where the tree and the inventory disagree, the tree wins and the accounting is what needs correcting** — check each condition against the real file before marking it done. +### What each disposition owes, stated once + +**A disposition is not a label; it names the observation that condition is owed.** Every task below +consults this table rather than restating it, and **a condition whose check does not match its +disposition is a defect in one of the two** — four consecutive passes found one, each time in a +different passage, because each task was inventing the rule for its own conditions. + +| Disposition | The observation it owes | +|---|---| +| **kept** | inside an untouched span, **or**, where it shares a line with changed text, its own per-condition count: `parent=1 worktree=1` in each copy. Never both, and never neither. | +| **carried** | a **preservation count** after installation, `1` in each copy, from the condition's own text. **No untouched span covers a carried condition** — it sits inside a replacement block, which is the whole reason it is not recorded as kept. | +| **replaced** | a **discriminating pair**, `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`, **both halves from the same edit**. | +| **moved** | **two** observations: an **absence** at the source, `parent=1 worktree=0`, and a **condition-specific presence** at the destination, `worktree=1 parent=0`. A presence check on the destination *paragraph* is not the second half — it passes while any one moved predicate is missing from it. | +| **dropped** | an **absence check**, `parent=1 worktree=0`, **one per dropped condition**. Two dropped conditions sharing one fragment is one observation, and it goes absent when either half goes, leaving the other free to survive. | +| **add-only** | **presence alone**, `worktree=1 parent=0`. There is no old wording whose absence could be counted. | + +**A reader walk is never any of these.** Tasks walk their conditions as a reader's confirmation on +top of the counts; a walk that found what the counts missed means a fragment was wrong, not that +the walk was the check. + ### Passage (a) — the floor paragraphs → target §H (Task 7) | Condition | Disposition | @@ -365,22 +391,27 @@ end anchor in each tree at comparison time. implementation fails its own check — and each split must leave **every kept condition inside some span**, which the first draft's version did not: -- **The floor arithmetic is two spans, not three, and neither starts at `Both gates are a LOOP`**: - `a1` and `a2` are both inside item 8a's replacement block, which begins at that anchor, so a span - opening there opens on a changed line. Take the spans as: **from the line after row F10's - fragment** to the line before row P9's; and from the line after P9's to `Nothing here writes the - floor knob`. The first of them holds `a3`–`a12` — **the span an earlier draft dropped entirely** - — and the second holds `a14`. `a1` and `a2` are covered by their own observations, not by a span: - `a2` by row F10's pair, `a1` by the carried-condition preservation check. -- **The human exception is three spans**, around rows F7 (`h4`), **F7b (`h5`)** and F4 (`h19`): - from `Recording a human exception` to before F7's line; **from after F7b's line** to before - F4's; and from after F4's to `because writing it down makes it sound`. **Twenty-two** of the - twenty-three kept conditions lie inside them; **`h6` is the exception** — it shares its line with - `h5`, which F7b changes, so no whole-line span can hold it and the per-condition fragment check - below is what covers it. **F7 and F7b sit on adjacent lines** — C 988 and 989 as of this - writing — **so a middle span opening after F7 rather than after F7b would enclose a line item 7 - changes, and a correct implementation would fail its own untouched check.** Resolve both anchors - and open the middle span after the later of them. +**This plan states the derivation and not the spans**, because writing them out by hand has now +been wrong twice — once enclosing row F7b's line in the human-exception region, once opening on +`a1`'s line and ending on the line `a13` and `a14` share. Neither is visible without the file open, +and each makes a **correct** implementation fail its own check. **Derive them, with the files +open:** + +1. Resolve each of the five regions' start and end anchors. +2. **Collect every line in that region that a changed fragment sits on** — every fragment-table row + whose edit lands in the region, and every line an installed replacement will occupy. For the + floor arithmetic that is item 8a's block and row P9's line; for the human exception, rows F7, + F7b and F4. +3. **The spans are the gaps between those lines.** A region with none of them is one span; a region + with *n* is at most *n+1*. A span of zero lines is dropped, not recorded. +4. **Then assign every kept condition in that region to exactly one span, or to the per-condition + list.** A kept condition in neither is the failure this step exists to prevent; one in both is + an accounting error. + +**Expect the per-condition list to be non-empty, and do not treat its members as a list to +memorise.** `a1` and `a2` sit inside item 8a's block; `a13` ends on the line `a14` begins; `h5` +shares its line with `h6`; `h4` and `h5` are adjacent. Each is a reason a whole-line span cannot do +this alone, which is why the derivation replaces the enumeration rather than correcting it again. **Record the spans as one `startendfile` line each in `.context/loop-rule-untouched`**, the anchors being literal strings. Tab-separated because the anchors contain colons — a `:` delimiter @@ -554,6 +585,14 @@ worktree file and diff the two. Expected: no output for every recorded span. Any difference is a defect — revert that hunk before continuing. +**Then run Task 0's per-condition list, which is the other half of the same artifact.** Task 0 +records spans **plus** a per-condition fragment for every kept condition no span could hold, and +the two together are what covers the kept set — a step reading only the spans leaves exactly the +conditions that needed individual attention unchecked, which is the shape of the defect the list +exists for. For each entry count the condition's own text in the parent and in the worktree, +expecting `1` both times. **Task 14 step 4b runs both lists again after the last text edit**; +this run catches what Task 1's insertion could have broken, at the cheapest moment. + - [ ] **Step 2: Confirm `f1` still names the right pair** Read the installed text around "The two rules above do not compete" and confirm the two rules immediately above it are still the absorb rule and the clearly-stuck reading, with §A before both rather than between them. @@ -625,9 +664,16 @@ the closing-time set-change rule — which replaces no wording and so owes `new/ new/parent=0` and no OLD half, per the add-only rule. Choose each NEW fragment from the installed text and verify it the three ways before counting it. -- [ ] **Step 4: Walk the carried conditions** +- [ ] **Step 4: Count the carried conditions, then walk them** + +`b1`, `b2`, `b4`, `b5`, `b6`, `b9`, `b10`, `b14`, `b15` are **carried** — reproduced inside §B's +block, so no untouched span covers them and step 3's pairs observe only the changed clauses around +them. Each owes its **preservation count**: take a unique fragment from the condition's own text, +confirm it single-line and unique before step 2 installs, and count it in each copy afterwards, +expecting `1`. Without them a mis-scoped replacement can drop carried wording from both copies +while every pair, the presence check and the parity diff pass. -Read the installed passage and confirm `b1`, `b2`, `b4`, `b5`, `b6`, `b9`, `b10`, `b14`, `b15` are each present, and that `b3`, `b7`, `b8`, `b11`, `b12`, `b13`, `b16`, `b17`, `b18` read as §B states rather than as the parent did. Nine plus nine; the disposition table above is the checklist. +Then read the installed passage and confirm `b3`, `b7`, `b8`, `b11`, `b12`, `b13`, `b16`, `b17`, `b18` read as §B states rather than as the parent did. Nine plus nine; the disposition table above is the checklist, and the walk is the reader's confirmation on top of the counts. - [ ] **Step 5: Parity** @@ -670,8 +716,9 @@ Expected: `P4=1` for both files. is **preserved in target §C's fenced replacement**, so its old-count could never reach zero. Derive `c4`'s OLD from the part of the sentence the replacement removes. -**`c4`, `c8`, `c9` and the four moved conditions `c10`–`c13` have no row yet, and all of them are -derived here rather than at step 4.** Take each from the live passage, check it the three ways, +**`c4`, `c8`, `c9`, the four moved conditions `c10`–`c13` and the three carried ones `c5`–`c7` have +no row yet, and all of them are derived here rather than at step 4** — every one is pre-existing +text, and step 2 installs over the passage they live in. Take each from the live passage, check it the three ways, confirm it counts 1 in each copy, and **add a row to the fragment table** under the next free `P` id — that table is where every fragment lives, and a row that is not there is a fragment stated somewhere else. Refer to them afterwards as **the `c4` row** and **the `c8` row**, never by a number picked here: Task 3 appends first and how many rows @@ -711,14 +758,24 @@ observation at all. **A later draft reintroduced exactly that pair** by naming r Each NEW is confirmed single-line and unique in the installed file before it is counted. -**Plus an absence check per moved condition — `c10`, `c11`, `c12`, `c13` — and one for `c9`'s moved -clause.** A move removes wording here and adds it in §A, so the source removal is an **absence**: -count the condition's own text in each copy, expecting `parent=1 worktree=0`. Task 1's §A presence -checks are the other half. **A reader walk is not this observation** — Task 3 states the rule and -the same reason applies: without it an old closure instruction can survive in passage (c) beside -its §A replacement, and every pair, count and parity diff still passes. This is the +**Plus both halves of every move — `c10`, `c11`, `c12`, `c13`, and `c9`'s moved clause.** A move +removes wording here and adds it in §A, so it owes an **absence** at this source, `parent=1 +worktree=0`, **and a condition-specific presence in the installed §A**, `worktree=1 parent=0`. +**Task 1's three paragraph-level presence fragments are not that second half** — they pass while +any one moved predicate is missing from §A, so a predicate could vanish here and never arrive +there with every check green. Derive one §A fragment per moved condition and record it with the +absence. + +**A reader walk is neither half** — without the counts an old closure instruction can survive in +passage (c) beside its §A replacement, and every pair and parity diff still passes. This is the two-instructions-that-disagree failure, and `c10`–`c13` are four chances at it. +- [ ] **Step 4b: Count the carried conditions** + +`c5`, `c6`, `c7` are **carried word for word** inside §C's block, so no untouched span covers them +and no pair observes them. Each owes a preservation count of `1` in each copy after installation, +from a fragment verified before step 2. + **`c9` is split, so it owes both halves:** the moved precedence clause is absent here (`parent=1 worktree=0`) and present in §A, while the plateau rationale stays — confirm the rationale still counts `1` in each copy. @@ -828,13 +885,25 @@ git commit -m "WIP: read the two-tell threshold after the clean-completion branc - [ ] **Step 1: Record the old wording in both copies, and derive the resolve-duty OLD** -§E replaces **two** blocks and **drops** a third thing. The handed-over-question OLD is row P6; -the resolve-duty bullet has no row yet, and neither does `g2`/`g3`'s absence fragment. **Derive -both here, before Step 2 installs over them** — the resolve-duty OLD from the live `**Severity:**` -bullet, the absence fragment from the live interim report-and-stop duty — check each the three -ways, confirm each counts 1 in each copy, and append both to the fragment table. **Not "before -Step 3":** Step 2 is the install, so a fragment derived after it is taken from text the edit has -already removed. +§E replaces **two** blocks and **drops three conditions**. The handed-over-question OLD is row P6; +four fragments have no row yet. **Derive all four here, before Step 2 installs over them**, check +each the three ways, confirm each counts 1 in the copies its row claims, and append each to the +fragment table: + +- the **resolve-duty OLD**, from the live `**Severity:**` bullet; +- **`g2`'s own absence fragment**, from the interim report-and-stop duty; +- **`g3`'s own absence fragment**, from that duty's justification; +- **`g4`'s absence fragment (C only)**, from the sentence naming the story path. + +**One fragment per dropped condition, not one for `g2` and `g3` together.** A combined fragment +goes absent as soon as *either* half is removed, which leaves the other obsolete instruction free +to survive behind a passing check — and the whole point of these three is that an obsolete stop +duty must not outlive the answer that settles it. **`g4` owes a recorded absence check too**, not +only step 4's path grep: Task 15's evidence entry names every absence check with its two counts, +and a bare `grep -c` on the current tree supplies neither the parent count nor an auditable row. + +**Not "before Step 3":** Step 2 is the install, so a fragment derived after it is taken from text +the edit has already removed. ```bash for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do @@ -856,16 +925,18 @@ omitted, while everything Task 6 checks passes. *(Build the pair per the verification procedure; record the four values.)* -**Three observations here, not two.** `P6` observes `g1`'s replacement and the resolve-duty row -observes the bullet, but **`g2` and `g3` are dropped rather than replaced** — the interim -report-and-stop duty and its rationale go, and nothing in §E takes their place. A dropped condition -owes an **absence check**, not a pair: count its text before and after, expecting `1` then `0`. -Without it the duty can survive beside the answer that makes it obsolete, which is the -two-instructions-that-disagree failure in its purest form. **Both fragments were derived and -validated at Step 1**; this step only runs them. +**Five observations here, not two.** `P6` observes `g1`'s replacement and the resolve-duty row +observes the bullet, but **`g2`, `g3` and `g4` are dropped rather than replaced** — the interim +report-and-stop duty, its rationale and the ownership sentence go, and nothing in §E takes their +place. A dropped condition owes an **absence check**, not a pair: count its text before and after, +expecting `1` then `0`. Without it the duty can survive beside the answer that makes it obsolete, +which is the two-instructions-that-disagree failure in its purest form. **All four fragments were +derived and validated at Step 1**; this step only runs them. -Expected for all four: `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`; and for the -`g2`/`g3` absence check, `worktree=0 parent=1` in each copy. +Expected: for the two pairs, four pair instances reading +`old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`; for `g2`'s and `g3`'s absence checks, +`worktree=0 parent=1` in **each** copy; for `g4`'s, `worktree=0 parent=1` in **C only** — W never +carried it, which is the divergence this task removes. - [ ] **Step 4: Confirm g4 is gone from C** @@ -986,11 +1057,20 @@ holds OLD fragments, which exist before the edit and can be checked in advance. Expected for all twenty: `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`. -**Then run what step 1b added:** a pair for each of the `c18`-and-surfacing block's four further -conditions, and a **presence check** — `new/worktree=1 new/parent=0` in each copy — for §G's -semantic membership test and for every independent add-only clause in the strict-reading tail. -Choose each of those NEW fragments from the installed text and verify it the three ways before -counting it. +**Then run everything step 1b added, and that is three kinds, not one:** + +- **a pair** for each of the `c18`-and-surfacing block's four further conditions — `c16`, `c17`, + `c19`, `c20`; +- **an absence at this source plus a condition-specific presence in §A** for each of the three + moved conditions `a18`, `a19` and `a20` — `parent=1 worktree=0` here, `worktree=1 parent=0` + there. **An earlier draft derived these rows at step 1b and then never ran them**, so the old + loop-until-clean, zero-finding and no-padding instructions could each survive at their source + with nothing observing it; +- **a presence check** — `new/worktree=1 new/parent=0` in each copy — for §G's semantic membership + test and for every independent add-only clause in the strict-reading tail. + +Choose each post-install fragment from the installed text, verify it the three ways before counting +it, and record all of them in this task's fragment evidence. - [ ] **Step 4: Confirm every carried condition in this task's blocks survived** @@ -1386,7 +1466,22 @@ A condition **in the block and not in the source** ships two triggers that disag The two copies are byte-identical over this material, so a divergence here is a parity failure and belongs to Task 14. -- [ ] **Step 5: Record the result** — it goes in the evidence entry verbatim. **No commit** unless the check failed and you repaired something. +- [ ] **Step 5: Write the result into this plan, then commit it** + +**Write both predicates in full, as you extracted them, and the result in both directions**, under +`## b11/b13 equivalence (Task 12 output)` at the end of this plan — created if absent, replaced +whole if present, by the same idempotent rule Tasks 13 and 14 use. + +**"Record the result — it goes in the evidence entry" named no file and modified none.** Task 15 +step 5 assembles the closing entry from this plan's recorded outputs; a comparison whose subject +and reasoning were written down nowhere cannot be read back, cannot be re-checked after a Gate-B +fix changes either predicate, and reaches the closing commit as an assertion that the check +happened. + +```bash +git add docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +git commit -m "WIP: record the b11/b13 equivalence result" +``` --- @@ -1629,6 +1724,18 @@ git commit -m "WIP: bump dev-workflow to 0.12.0" - [ ] **Step 4: Run the full quality battery** +**First fetch the base ref, because one step in the battery has a precondition the others do not.** +`check-version-bump.sh` compares *commits* against a base ref, so a stale base compares the bump +against a different commit than the pull request will, and the run is green about the wrong +comparison. **Fetch, then pass the fetched ref to the checker itself** — an earlier draft fetched +`origin/main` while the battery went on passing the local `main`, and recorded `origin/main` as +what supplied a comparison it had not supplied: + +```bash +git fetch origin main +git rev-parse origin/main # record this — and it is the argument the check below receives +``` + ```bash shellcheck --shell=sh plugins/dev-workflow/hooks/codex-gate.sh && \ shellcheck --shell=sh --exclude=SC2015 plugins/dev-workflow/hooks/codex-gate.test.sh && \ @@ -1639,24 +1746,13 @@ shellcheck --shell=sh scripts/check-version-bump.test.sh && \ HOOK_SH=sh sh plugins/dev-workflow/hooks/codex-gate.test.sh && \ HOOK_SH=dash dash plugins/dev-workflow/hooks/codex-gate.test.sh && \ sh scripts/check-invariants.test.sh && sh scripts/check-invariants.sh && \ -sh scripts/check-version-bump.test.sh && sh scripts/check-version-bump.sh main && \ +sh scripts/check-version-bump.test.sh && sh scripts/check-version-bump.sh origin/main && \ claude plugin validate . --strict ``` -Expected: exit 0. - -**`check-version-bump.sh main` has a precondition the rest of the battery does not**, and AGENTS.md -states it: it compares *commits* against a base ref, so a stale local `main` compares the bump -against a different base than the pull request will. **Fetch the base ref first and record which -commit supplied the comparison:** - -```bash -git fetch origin main -git rev-parse origin/main # record this — it is what the local run compared against -``` - -Run the battery's version-bump step against a current `main`. A green run against a stale one is -green about the wrong comparison and defers the failure to CI. +Expected: exit 0. **`origin/main`, not `main`** — AGENTS.md's battery row writes `main` because a +human running it locally usually has one; here the fetched ref is the one just recorded, and the +recorded revision must be the exact argument the successful check received. - [ ] **Step 4b: Apply all twelve `docs/prompt-standards.md` items to the installed text, and commit the result before Gate B** @@ -1693,7 +1789,16 @@ in the WIP body too if a mid-cycle reader would want it, but the file is the cop **human-exception record** belong beside it. **Step 7b appends all three and revalidates the entry** — this step opens the file, step 7b closes it, and step 8 commits it. -It names: the battery run; **every pair this plan built, with its counts in each copy and each tree, every presence check beside them, and every absence check with its two counts**; the §6 parity diff and the `b11`/`b13` equivalence result; and the next-state table's location plus its row count. **Read them out of `## Fragment evidence (per-task output)`**, which is where every task records them. +It names: the battery run; **every pair this plan built, with its counts in each copy and each tree, every presence check beside them, and every absence check with its two counts**; the §6 parity diff and the `b11`/`b13` equivalence result; and the next-state table's location plus its row count. + +**Each of those is read from the section that actually holds it**, and they are not one place: +`## Fragment evidence (per-task output)` for the pairs, presence and absence counts; +`## b11/b13 equivalence (Task 12 output)` for the equivalence result; +`## Divergence list (Task 14 output)` for the parity outcome; +`## Next-state table (Task 13 output)` for the table and its row count; +`## Completeness sweep (Task 12b output)` and `## Prompt-standards result (Task 15 step 4b output)` +for those two reader records. **An entry assembled from one section would silently drop whatever +the other five hold.** **The absence checks belong in the entry as much as the pairs do.** `g2`, `g3` and `g4` are dropped rather than replaced, and `c9`–`c13` and `a18`–`a20` are moved, so the observation that proves each is gone is an absence — `parent=1 worktree=0` — and nothing else in the entry carries it. An entry listing only pairs and presence checks claims the verification set while omitting the half that proves obsolete instructions were removed, which is the failure the absence checks exist for. @@ -1741,13 +1846,22 @@ twelve-item review both describe the tree as it was *before* that fix. Without r plan reaches its closing commit on a tree that never passed its own quality battery — a repo unable to pass CI, or an invariant-11 violation, published by a cycle that closed clean. -- **After each fix, re-run the checks whose subject it changed** — the hook or its test means the - suite under both shells; either prompt copy means `sh scripts/check-invariants.sh`; a shell file - means `shellcheck`. -- **Before the closing act, re-run the whole battery of step 4 against the current `HEAD`**, and - re-apply step 4b's twelve items to every artefact a fix touched, recording the fresh result in - the plan and committing it. **The results the closing commit carries are the ones from the final - tree**, not from the tree that first went green. +- **After each fix, re-run every check whose subject it changed — mechanical and reader alike.** + The hook or its test means the suite under both shells; a shell file means `shellcheck`; either + prompt copy means `check-invariants.sh` **and** every observation over the text the fix touched — + its pair, presence, absence and preservation counts — **and** the reader checks whose subject + moved: Task 12's equivalence where `b11` or `b13` changed, Task 12b's sweep where the ordering or + a hook message changed, Task 13's transitions and closure checks where §A changed, Task 14's + parity and untouched checks wherever text moved at all. **Naming only the mechanical ones let a + fix invalidate a reader record that then reached the closing commit unchanged.** +- **Persist every re-run record and commit it before the re-review**, so the reviewed range holds + it. A refreshed record left in the worktree is in neither the Gate-B range nor `reset --soft`. +- **Before the candidate final pass, re-run the complete set** — the whole step-4 battery against + the current `HEAD`, step 4b's twelve items, and every reader check above — commit those records, + and resolve the new `HEAD`. **Only a clean response issued against that exact `HEAD` closes the + cycle.** A pass that was clean against an earlier tree, plus records committed afterwards, closes + on a tree no pass reviewed — which is the same defect as reviewing the wrong range, arrived at + from the other end. - [ ] **Step 7b: After the clean pass, complete `.context/loop-rule-closing-msg` — this is the action step 5 defers to, and step 8 has no other source for these records** @@ -1837,7 +1951,7 @@ The closing body carries: the validated evidence entry; the provenance line; the **2. Placeholder scan.** The replacement text is cited rather than copied, deliberately and for the reason the Architecture note gives. Task 13's row list is explicitly a floor rather than a closed set, and says so. Task 11 deliberately carries no count, and says why. -**Which steps are mechanical and which are reader checks, stated rather than claimed uniformly.** Every OLD half has an exact expected result and a procedure that produces it, but **not every one is a pre-verified table row**: the rows the tables carry were checked against the real files in advance, while Tasks 1, 3, 4, 6, 7 and 10 **derive their remaining OLD fragments at execution, before their install step**, against text this plan cannot quote without becoming a second copy of it. **The NEW halves are `<...>` until their task installs the text**, which the fragment table discloses and each step requires to be verified before counting. **And no per-task shell is pre-written at all** — the procedure is stated once and the executor writes the command in front of the files, so "a runnable command per step" is not what this plan claims. **Tasks 12, 13 and 14 step 2 are reader checks by design** — a predicate comparison, a next-state walk and a divergence classification are judgements, and giving them commands would be the false-precision this repo's invariants warn about. An earlier revision of this section claimed every verification step had a runnable command, which was not true of them. +**Which steps are mechanical and which are reader checks, stated rather than claimed uniformly.** Every OLD half has an exact expected result and a procedure that produces it, but **not every one is a pre-verified table row**: the rows the tables carry were checked against the real files in advance, while Tasks 3, 4, 6, 7 and 10 **derive their remaining pre-existing fragments at execution, before their install step**, against text this plan cannot quote without becoming a second copy of it. **Task 1 is not among them:** §A is add-only, row P1 records that it has no OLD half at all, and Task 1 derives presence fragments only — listing it would send an executor looking for a counterfactual that cannot exist. **The NEW halves are `<...>` until their task installs the text**, which the fragment table discloses and each step requires to be verified before counting. **And no per-task shell is pre-written at all** — the procedure is stated once and the executor writes the command in front of the files, so "a runnable command per step" is not what this plan claims. **Tasks 12, 13 and 14 step 2 are reader checks by design** — a predicate comparison, a next-state walk and a divergence classification are judgements, and giving them commands would be the false-precision this repo's invariants warn about. An earlier revision of this section claimed every verification step had a runnable command, which was not true of them. **3. Type consistency.** `$BASE` is set in Task 0 and used in Tasks 1–10. The four-value pair shape (`new/worktree`, `new/parent`, `old/worktree`, `old/parent`) is defined in Task 3 and referred to by name afterwards. Condition ids match the inventory throughout: a1–a22, b1–b18, c1–c20, d1–d7, e1–e11, f1–f7, g1–g4, h1–h26, i1–i16, j1–j4 — 135 total, every one dispositioned above. @@ -1847,6 +1961,13 @@ The closing body carries: the validated evidence entry; the provenance line; the --- +## b11/b13 equivalence (Task 12 output) + +*Empty until Task 12 runs. Task 12 replaces this entire section, carrying both predicates in full +as extracted and the result in both directions, per copy.* + +--- + ## Next-state table (Task 13 output) *Empty until Task 13 runs. Task 13 replaces this entire section.* From 55a27c92829bcd8762e108ffb72a5954b316fa88 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 17:21:37 +0200 Subject: [PATCH 112/181] docs(plans): apply Gate-A plan pass 10; tasks read the disposition table, not lists MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit MANDATORY TWO-TELL STOP — findings rose 9 -> 17 and the findings cluster on the verification apparatus for the tenth pass. Blockers fell 2 -> 1. Surfaced to Daniel; standing answer applied, loop continued. One Blocker, fourteen Majors, two Minors. Eleven of the seventeen are one family: pass 9 installed the disposition table, and the tasks had not been brought into line with it. The repair is not eleven patches. Root: "How a task discharges that table" states the per-task procedure, and NO TASK ENUMERATES ITS CONDITION IDS ANY MORE. Five consecutive passes found a condition missing its observation, and every one was in a task that had listed some of its ids and not the rest — an enumeration per task is a second copy of the disposition table and goes stale like every other enumeration this cycle has produced. Tasks 3, 4, 5, 7 and 8 now walk their passage's rows by class. Two consequences that had been got wrong and are now stated: a KEPT condition inside a passage a task replaces whole is protected by no untouched span, so it owes a per-condition count; and every pre-existing fragment — carried and kept preservation fragments included — is derived BEFORE the install. That closes, at the root: nine carried b conditions confirmed from inside a step that runs after the install; kept c1-c3, e1-e6, e10, e11, i1-i3 and i9-i16 with no observation at all; carried e9 whose clause wraps in both copies so a literal count returns zero in a correct file; carried a15/a21/a22/c15/i4-i8 chosen after installation; carried a1 the same. Individually: - Task 0 step 1 printed the branch and status under "expect" comments and asserted neither; both fail the script now, and the three approved inputs are compared by blob against ba15e83 rather than merely found. What cannot be established — that this is the Gate-A-closed plan revision — is said plainly. - .context/loop-rule-untouched had a schema for spans and none for the per-condition half its consumers must read. Both shapes are defined. - a17 is dispositioned moved and was observed only by P11, which pairs the old rule with §H's pointer — a different sentence. All four of a17-a20 get the moved pair, and the pointer is observed separately. - Task 5 gave e7 and e8 the same NEW fragment, so only the pronoun alignment was observed and e7's read-after-clean-completion clause was not. - c9 needed two rows: the moved clause's absence and the plateau rationale's preservation. One row cannot require a thing to vanish and to remain. - Task 15 step 7's "complete rerun" omitted every mechanical observation, so a clean pass could close on a HEAD whose fragment evidence came from an older tree. And step 8 staged the plan after the review; the final pass's own findings file is now the sole permitted post-review addition, asserted. - Task 14 applied its own block rule to §H and still gave §A and §E one row each, where §A has three blocks and §E two. Trend: findings 32, 32, 31, 20, 16, 11, 13, 15, 9, 17. Blockers 1, 0, 0, 10, 7, 2, 4, 2, 2, 1. Majors 28, 25, 23, 7, 7, 6, 4, 9, 6, 14. Floor 3, from risk high (2) and security none (0), read at this pass from docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-plan-om0bdd7udh-pass-10.md | 18 + .../2026-09-14-loop-rule-consolidation.md | 319 ++++++++++++------ 2 files changed, 241 insertions(+), 96 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-10.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-10.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-10.md new file mode 100644 index 0000000..5abed06 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-10.md @@ -0,0 +1,18 @@ +MAJOR | high | Task 0 step 1 | `git rev-parse --abbrev-ref HEAD` and `git status --porcelain` only print their results; neither command asserts the promised `loop-rule-consolidation` branch or clean tree before the base is recorded | A wrong-branch or dirty checkout can become `$BASE`, mixing unrelated work into the execution and the final soft reset | Test the branch value and empty status explicitly and exit before writing `.context/loop-rule-base` on either mismatch +MAJOR | high | Task 0 step 1 | The only approval assertion is that `ba15e83` is an ancestor; the step does not establish that the target/design blobs still equal their approved versions or that the plan revision being executed is the Gate-A-approved one | A later unapproved edit to an input or a stale/unreviewed plan can pass the baseline and supply different installation or execution instructions | Compare the input blobs to the approved commit and assert the eventual Gate-A closing revision of this plan before recording the base +MAJOR | high | Task 0 step 2 interface | Span records have the schema `startendfile`, but the per-condition half of `.context/loop-rule-untouched` has no schema for condition id, fragment, file/copy, or expected parent/worktree values | Tasks 2 and 14 cannot deterministically consume the very records that must protect line-sharing `a14`, `h6`, and any further derived member | Define a distinct parseable per-condition row shape and have both consumers validate `parent=1 worktree=1` for every recorded copy +MAJOR | high | Task 3 step 4 | The nine carried `b` conditions are first told to have their fragments confirmed before step 2 from inside step 4, after step 2 has replaced the passage; step 1b derives only changed-condition OLD rows | Their required pre-edit fragments cannot be validated against the live source or appended under the fragment-table cut, so a partial rerun or an omitted carried condition can pass | Derive, validate, count, and append one preservation row per carried `b` condition in step 1b before installing §B, then consume those rows in step 4 +MAJOR | high | Task 4 step 5 | Kept `c1`-`c3` receive only the reader walk even though no Task 0 untouched span covers passage (c); the live text is C 225-228 and W 428-431, with `Read the` wrapping before `**Blocker curve across passes**` in both copies | The task violates its own rule that a reader walk is never a kept-condition observation, so an accidental edit to the curve and coverage premises can survive every mechanical check | Record an untouched span for `c1`-`c3` or run per-condition parent/worktree preservation counts with verified single-line fragments +MAJOR | high | Task 4 steps 1 and 4b | The single `c9` row must observe removal of the moved precedence clause, while the separately kept plateau rationale gets only a post-install count; `That third condition is what makes a plateau rather than a finish` wraps at C 234-235 and W 437-438 | One row cannot prove both source absence and preservation, and choosing the quoted wrapped rationale after installation defeats the pre-edit fragment rule | Derive a second, single-line plateau-rationale preservation row before step 2 and count it `parent=1 worktree=1` in both copies +MAJOR | high | Task 5 step 4 | Kept `e1`-`e6`, `e10`, and C-only `e11` are only told to be confirmed unchanged in a reader walk; they are in no Task 0 untouched span and receive no disposition-required counts | The §D edit can alter or drop a tell, the long-before-plateau sentence, or the C-only rationale while the `e7` pair and pointer check pass | Add bounded untouched spans or per-condition parent/worktree preservation counts, including a C-only expectation for `e11` +MAJOR | high | Task 5 step 3 | Both `e7` pairs take NEW from `you report the tells`, which observes the `e8` pronoun alignment rather than `e7`'s new read-after-clean-completion clause; in C those words already exist across C 266-267 and become new only by reflow | The read-order change can be omitted while both OLD boundaries disappear and the pronoun fragment passes, and the two separately dispositioned changes do not each get a same-edit observation | Give `e7` a NEW fragment containing its read-after-clean-completion meaning and give W's `e8` alignment its own OLD/NEW observation +MAJOR | high | Task 5 step 4 | The quoted carried-`e9` text `the "clearly stuck" reading above is not a precondition for it` wraps at C 267-268 and W 471-472, and no preservation row is derived before step 2 | A literal single-line count returns zero in both correct source files, while choosing a different fragment after installation cannot satisfy the fragment table's pre-edit evidence rule | Before step 2 derive and append a unique single-line `e9` fragment, then record its post-install preservation count in each copy +MAJOR | high | Task 7 steps 1b and 3 | `a17` is dispositioned moved, but only `a18`-`a20` receive source absence plus condition-specific §A presence; P11 instead pairs the old clean-final-pass rule with the §H pointer `What a clean final pass and the zero-finding early exit mean for closing` | The actual `a17` predicate can disappear from its source without arriving in §A while the pointer pair and the paragraph-level §A samples pass | Include `a17` in the moved-condition source/destination checks and observe the §H pointer separately from the moved predicate +MAJOR | high | Task 7 step 4 | Preservation fragments for carried `a15`, `a21`, `a22`, `c15`, and `i4`-`i8` are first selected in step 4 after all ten blocks are installed, rather than derived and appended before step 2 as the fragment-table cut requires | The checks cannot detect source drift or a partial rerun and leave no pre-edit authored fragment for Task 15 to audit | Derive and append one condition-specific preservation row for all nine carried conditions in step 1b, then count those rows after installation +MAJOR | high | Task 7 kept-condition coverage | Kept `i1`-`i3`, `i9`-`i16` receive no untouched span or per-condition count; the real passage is C 153-167 and W 360-374, and `i3` shares the sentence whose dash-delimited list changes | The strict-reading edit can damage the activation, post-rule, nonce, extension-point, re-derivation, knob, or revert conditions while every P16 and carried-item check remains green | Derive untouched spans around the replaced list and a line-sharing preservation fragment where needed, then rerun them at Task 14 +MAJOR | high | Task 8 step 4 | Carried `a1` is first given a fragment/count after item 8a has already replaced its block, although the governing fragment-table cut explicitly includes every carried-condition preservation fragment before installation | A drifted or partially installed HARD FLOOR opening cannot be distinguished from the intended carried text and its evidence is absent from the OLD table | Derive, validate, count, and append the `a1` preservation row in step 1 before installing the fourteen items +MINOR | high | Task 7 step 1b | The prose says three of ten blocks change more than one thing and later says the other seven change one each, but its own bullets identify four: `c18`/surfacing, `a17`-`a22`, §G, and the strict-reading list | The stated arithmetic can lead an executor to omit one expanded observation while believing the ten-block accounting is complete | Change the count to four and the complement to six, preserving all four bullet sets +MINOR | high | Task 14 step 1 | The task requires one row per fenced destination block but then assigns one row each to §A and §E; mechanically, target §A has three fenced blocks at lines 53-375, 381-415, and 421-449, and §E has two at 628-637 and 642-666 | The parity-site inventory applies a section-level grouping to five blocks despite the plan's fixed destination-block unit, weakening its own completeness oracle | Emit one bounded row per A1, A2, A3, Severity, and demotion-answer destination block +BLOCKER | high | Task 15 step 7 final-pass preparation | The promised complete rerun before the candidate final pass enumerates the battery, prompt standards, and reader checks only; it omits the full mechanical set of discriminating pairs, add-only presence checks, moved/dropped absences, and kept/carried preservation counts | The clean pass can close on an exact HEAD whose fragment evidence was never re-established as a complete set, so the plan cannot substantiate its design §7 verification claim | Before committing records for the candidate final pass, rerun and persist every mechanical observation as well as every reader check, then review that resulting HEAD +MAJOR | high | Task 15 steps 7-8 | After requiring the clean response to target the exact closing HEAD, step 8 stages `.context/codex-reviews/` and may create a new WIP record commit containing the just-produced clean findings file before the soft reset | The final squashed tree differs from the HEAD that received the clean response, contradicting the exact-HEAD closure rule the task just installed | Make the generated clean-result artifact an explicit sole post-review exception with a check that no other tree delta exists, or keep it outside the squashed tree +END OF FINDINGS (17 total) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index a1f1001..2b675fa 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -187,6 +187,34 @@ different passage, because each task was inventing the rule for its own conditio top of the counts; a walk that found what the counts missed means a fragment was wrong, not that the walk was the check. +### How a task discharges that table — the procedure, so no task enumerates ids + +**No task below lists which of its conditions owe which check.** Five consecutive passes found a +condition missing its observation, and every one of them was in a task that had enumerated some of +its ids and not the rest. **An enumeration per task is a second copy of the disposition table**, +and it goes stale exactly the way every other enumeration in this cycle has. + +**Each editing task instead runs this, against the disposition table's rows for its own passage:** + +- [ ] **Before installing** — walk every condition in this task's passage and derive the + **pre-existing** fragment its disposition owes: the **OLD** half for a *replaced* one, the + **absence** fragment for a *moved* or *dropped* one, the **preservation** fragment for a + *carried* one, and, for a *kept* one, a preservation fragment **only where no untouched span can + hold it** — which is the case for every kept condition in a passage this task replaces whole, and + for any that shares a line with changed text. Check each the three ways, confirm its expected + pre-edit count, and append it to the fragment table under the next free `P` id. +- [ ] **After installing** — run each one to the result its class owes, and choose the + **post-install** halves: the **NEW** for each pair, and the destination **presence** for each + moved condition, taken from the installed text and verified the three ways before counting. + Record those in this task's fragment evidence. +- [ ] **Then walk the conditions as a reader**, which confirms the counts and replaces none of them. + +**Two consequences worth stating, because both have been got wrong.** A **kept** condition inside a +passage a task replaces whole is **not** protected by an untouched span — no span survives there — +so it owes a per-condition count like a carried one. And **every one of these fragments is +pre-existing**, so all of them are derived *before* the install, including the carried and kept +preservation fragments that an earlier draft chose afterwards. + ### Passage (a) — the floor paragraphs → target §H (Task 7) | Condition | Disposition | @@ -290,13 +318,15 @@ the walk was the check. - [ ] **Step 1: Confirm the approved artifacts and a clean tree** ```bash -git rev-parse --abbrev-ref HEAD # expect loop-rule-consolidation +test "$(git rev-parse --abbrev-ref HEAD)" = loop-rule-consolidation || { echo "wrong branch"; exit 1; } +test -z "$(git status --porcelain)" || { echo "tree not clean"; exit 1; } git merge-base --is-ancestor ba15e83 HEAD || { echo "approved target text (ba15e83) is not in this history"; exit 1; } -test -f docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md || { echo "target text missing"; exit 1; } -test -f docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md || { echo "design missing"; exit 1; } -test -f docs/superpowers/specs/2026-09-10-loop-rule-consolidation-condition-inventory.md || { echo "inventory missing"; exit 1; } -git log --oneline -1 -git status --porcelain # expect empty +# The three inputs must still be the versions that were approved, not merely present. +for f in target-text design condition-inventory; do + p="docs/superpowers/specs/2026-09-10-loop-rule-consolidation-$f.md" + test "$(git rev-parse "HEAD:$p")" = "$(git rev-parse "ba15e83:$p")" \ + || { echo "$p differs from its approved version at ba15e83"; exit 1; } +done if [ -s .context/loop-rule-base ]; then echo "base already recorded: $(cat .context/loop-rule-base) — NOT overwriting" else @@ -305,13 +335,22 @@ fi cat .context/loop-rule-base ``` -**The branch and the approved artifacts are established, not assumed.** An earlier draft printed -`git log --oneline -1` with the comment "expect 6ace06f or later on loop-rule-consolidation", -which names neither a branch the command reports nor a way to tell "later" from "a different -history". A checkout on another branch, or one whose history does not contain the approved target -text, would be recorded as `$BASE` and every install and the final soft reset would run from the -wrong starting tree. **A hard-coded commit is not repeated as a floor** — `ba15e83` appears once, -as the ancestry test for the approved text, because it is the commit that approval is *of*. +**Every line above fails the script; none of them prints for a human to notice.** An earlier draft +printed `git rev-parse --abbrev-ref HEAD` and `git status --porcelain` under comments saying what +to expect, which asserts nothing: a wrong-branch or dirty checkout would be recorded as `$BASE` and +every install and the final soft reset would run from the wrong starting tree. + +**The inputs are compared by blob, not merely found.** `ba15e83` being an ancestor says the +approved commit is in this history; it does not say the three files still hold what was approved, +and a later unapproved edit to the target text would supply different installation instructions to +every task. **`ba15e83` appears once, as the approval commit**, and the loop compares each input's +object id against its version there. + +**What this cannot establish, stated rather than implied:** that *this plan revision* is the one +Gate A closed on. A plan cannot name its own closing commit, and any sha written here would be from +before that commit exists. **The Gate-A findings files in `.context/codex-reviews/` are the record +of which revision was reviewed**; a reader checks the plan against them, and nothing mechanical +here does it. **Never overwrite an existing base, and never trust one you did not just write.** Re-running Task 0 after a partial implementation would record the current WIP tip, and both Gate B's range and the @@ -413,7 +452,21 @@ memorise.** `a1` and `a2` sit inside item 8a's block; `a13` ends on the line `a1 shares its line with `h6`; `h4` and `h5` are adjacent. Each is a reason a whole-line span cannot do this alone, which is why the derivation replaces the enumeration rather than correcting it again. -**Record the spans as one `startendfile` line each in `.context/loop-rule-untouched`**, the +**`.context/loop-rule-untouched` holds two record shapes, and both are parseable**, because Tasks 2 +and 14 consume them without a human in between. A file whose second half has no schema is a file +its consumers skip, which is what the per-condition list exists to prevent: + +``` +span +cond +``` + +Tab-separated, one record per line, the leading keyword distinguishing them. A kept condition's +expected pair is `11`; the shape carries the values rather than assuming them, so a moved or +dropped condition recorded here later needs no new format. **Both consumers validate every `cond` +row**, not only the `span` rows. + +**Record the spans as one `spanstartendfile` line each in `.context/loop-rule-untouched`**, the anchors being literal strings. Tab-separated because the anchors contain colons — a `:` delimiter splits `**Severity:**:**Tool routing:` at the wrong colon and yields an empty end anchor. @@ -632,16 +685,17 @@ Expected: `P2=1 P3=1` for **both** files — the first draft ran these against C **The first draft of this plan named `plus repair obligations you already accepted in earlier passes` here, which wraps across C 201–202 and W 408–409 and counts zero in a correct file.** -- [ ] **Step 1b: Derive the rest of §B's OLD rows, still before installing** +- [ ] **Step 1b: Run the pre-install half of the disposition procedure over passage (b)** -**Two pairs do not cover §B.** Target §B separately changes `b3`, `b8`, `b11`, `b13`, `b16`, -`b17`–`b18` and **adds one rule**, the closing-time set-change rule. Each independent replacement -owes its own pair, and the added rule owes a presence check — a manual condition walk is a -reader's judgement, not the discriminating observation design §7 assigns here. **Derive an OLD row -for each changed condition now**, check each the three ways against the live passage, confirm it -counts 1 in each copy, and append it to the fragment table under the next free `P` id. **This is -step 1b and not part of step 3 because step 2 removes the wording these rows are taken from** — a -row derived afterwards cannot be checked against the text it describes. +**Two pairs do not cover §B.** §B replaces the passage whole, so **every** condition in it owes an +observation — the changed ones a pair, the nine carried ones a preservation count, the added rule a +presence check — and **no untouched span protects anything here**, because none survives a +whole-passage replacement. Derive every pre-existing fragment now, per +`## How a task discharges that table`. + +**A manual condition walk is not the discriminating observation design §7 assigns here**, and §B's +nine carried conditions are the easiest ones to lose to it: a mis-scoped replacement drops carried +wording from both copies while every pair, the presence check and the parity diff pass. - [ ] **Step 2: Install §B's text over the passage in both copies** @@ -664,16 +718,11 @@ the closing-time set-change rule — which replaces no wording and so owes `new/ new/parent=0` and no OLD half, per the add-only rule. Choose each NEW fragment from the installed text and verify it the three ways before counting it. -- [ ] **Step 4: Count the carried conditions, then walk them** - -`b1`, `b2`, `b4`, `b5`, `b6`, `b9`, `b10`, `b14`, `b15` are **carried** — reproduced inside §B's -block, so no untouched span covers them and step 3's pairs observe only the changed clauses around -them. Each owes its **preservation count**: take a unique fragment from the condition's own text, -confirm it single-line and unique before step 2 installs, and count it in each copy afterwards, -expecting `1`. Without them a mis-scoped replacement can drop carried wording from both copies -while every pair, the presence check and the parity diff pass. +- [ ] **Step 4: Run the post-install half, then walk the passage** -Then read the installed passage and confirm `b3`, `b7`, `b8`, `b11`, `b12`, `b13`, `b16`, `b17`, `b18` read as §B states rather than as the parent did. Nine plus nine; the disposition table above is the checklist, and the walk is the reader's confirmation on top of the counts. +Every row step 1b appended, to the result its class owes; then read the installed passage and +confirm each condition reads as §B states rather than as the parent did. **The disposition table is +the checklist and the walk is the reader's confirmation on top of the counts.** - [ ] **Step 5: Parity** @@ -716,14 +765,22 @@ Expected: `P4=1` for both files. is **preserved in target §C's fenced replacement**, so its old-count could never reach zero. Derive `c4`'s OLD from the part of the sentence the replacement removes. -**`c4`, `c8`, `c9`, the four moved conditions `c10`–`c13` and the three carried ones `c5`–`c7` have -no row yet, and all of them are derived here rather than at step 4** — every one is pre-existing -text, and step 2 installs over the passage they live in. Take each from the live passage, check it the three ways, -confirm it counts 1 in each copy, and **add a row to the fragment table** under the next free `P` id — that table is where every fragment lives, and a -row that is not there is a fragment stated somewhere else. Refer to them afterwards as **the `c4` -row** and **the `c8` row**, never by a number picked here: Task 3 appends first and how many rows -it adds is decided at execution. **Step 2 removes the wording they are taken from**, so a -derivation after it has no live text to check against and no pre-edit count to observe. +**Run the pre-install half of the disposition procedure over passage (c)** — every condition in it, +by its class, derived from the live passage now and appended to the fragment table under the next +free `P` id. Refer to the rows afterwards by the condition they observe, never by a number picked +here: Task 3 appends first and how many rows it adds is decided at execution. + +**Two things about passage (c) that the class alone does not tell you:** + +- **`c1`–`c3` are kept and no untouched span reaches them.** Task 0's five regions do not include + passage (c), and §C replaces text inside it, so the curve-reading premises owe **per-condition + preservation counts** — `parent=1 worktree=1` in each copy — like a carried condition. Without + them an accidental edit to the curve or coverage sentences survives every mechanical check here. +- **`c9` is split and needs two rows, not one.** The moved precedence clause owes an **absence** + at this source; the plateau rationale that stays owes a **preservation count**. One row cannot + observe both — the clause must vanish and the rationale must not — and the rationale's own + sentence wraps in both copies, so its fragment has to be chosen single-line **before** step 2 + rather than picked out of the installed text afterwards. - [ ] **Step 2: Install §C's block** @@ -770,11 +827,11 @@ absence. passage (c) beside its §A replacement, and every pair and parity diff still passes. This is the two-instructions-that-disagree failure, and `c10`–`c13` are four chances at it. -- [ ] **Step 4b: Count the carried conditions** +- [ ] **Step 4b: Run every remaining row step 1 appended** -`c5`, `c6`, `c7` are **carried word for word** inside §C's block, so no untouched span covers them -and no pair observes them. Each owes a preservation count of `1` in each copy after installation, -from a fragment verified before step 2. +The carried `c5`–`c7`, the kept `c1`–`c3`, and `c9`'s plateau rationale — each to the result its +class owes. **None of them is observed by a pair**, which is why they are listed as a step rather +than left to the walk. **`c9` is split, so it owes both halves:** the moved precedence clause is absent here (`parent=1 worktree=0`) and present in §A, while the plateau rationale stays — confirm the @@ -830,6 +887,19 @@ not text. The first fragment wrapped across C 267–268 and W 471–472; the sec target §D's replacement** and could never reach zero. `P5_OLD` and `P5w_OLD` take the clause §D actually removes. +**Then run the pre-install half of the disposition procedure over passage (e)**, which P5 and P5w +do not cover: + +- **`e9` is carried** and its sentence **wraps** in both copies — C 267–268, W 471–472 — so a + literal count of the whole clause returns zero in a *correct* source file. Choose a single-line + fragment of it now, verify uniqueness, and append it; choosing one from the installed text + afterwards is the pre-edit rule broken in the one place the wrap makes it tempting. +- **`e1`–`e6`, `e10` and `e11` are kept, and no untouched span reaches them.** Passage (e) is not + one of Task 0's five regions and §D edits inside it, so each owes a **per-condition preservation + count**, `parent=1 worktree=1` — **`e11` in C only**, which is its recorded divergence. Without + them §D's edit can alter or drop a tell, the long-before-plateau sentence or the C-only + rationale while the `e7` pair and the pointer check both pass. + - [ ] **Step 2: Install §D's `e7` sentence and the pointer paragraph** - [ ] **Step 3: Run the discriminating pair, and a presence check for the pointer** @@ -839,11 +909,16 @@ is add-only** — it replaces no wording — so it is checked by presence alone, it can be omitted from both copies while the `e7` pair, the condition walk and the parity diff all pass. -**Take both copies' NEW fragment from the clause carrying the pronoun** — the installed -`you report the tells` — rather than from any other part of §D's sentence. P5w's OLD going to zero -shows W's pronoun-less form is gone; only a NEW fragment containing the pronoun shows W received -C's form rather than some third wording. **That is the `e8` alignment's own observation**, and -without it the alignment rests on Task 14's parity judgement instead of on a count. +**`e7` and `e8` are two dispositions and owe two different NEW fragments — an earlier draft gave +both pairs the same one and observed only the second.** + +- **`e7`'s NEW must carry the read-after-clean-completion meaning** — the clause §D adds. That is + what `e7` changed. A NEW taken from `you report the tells` observes nothing about it: those words + are already in C, and in C they become "new" only because the line reflows. +- **`e8`'s alignment is W's alone and owes its own observation**: `report the tells` becoming + `you report the tells`. P5w's OLD going to zero shows W's pronoun-less form is gone; a NEW + containing the pronoun shows W received C's form rather than some third wording. Without it the + alignment rests on Task 14's parity judgement instead of on a count. *(Build the pair per the verification procedure; record the four values.)* @@ -854,13 +929,13 @@ Expected: both pairs `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`; They are not one class and the check differs per class: -- **`e1`–`e6`, `e10`, `e11` are kept and outside the replacement** — confirm each is unchanged - against the parent, and that `e11` is still C-only. **`e10` belongs here, not with `e9`**: it is - the sentence after §D's block, so a check treating it as carried would look for it inside text - it never enters. -- **`e9` is carried inside §D's block** — confirm `the "clearly stuck" reading above is not a - precondition for it` is present in each copy after the install, expecting `1`. A carried - condition is reproduced rather than untouched, so a span cannot protect it. +Run every row step 1 appended, to the result its class owes, then read the passage. Three things +the class alone does not tell you: + +- **`e10` is kept and belongs with `e1`–`e6`, not with `e9`** — it is the sentence *after* §D's + block, so a check treating it as carried would look for it inside text it never enters. +- **`e9` is carried** — its count is `1` in each copy after the install, from the single-line + fragment step 1 appended rather than from the whole wrapped clause. - **`e8` is aligned rather than untouched** — this task gives W the pronoun, so listing `e8` among the untouched conditions would contradict the task's own instruction. @@ -995,20 +1070,32 @@ Expected: `1` twenty times. **Any `0` means the wording drifted since this table - [ ] **Step 1b: Derive the rows the ten locators do not cover, still before installing** -**One pair per block is a sample, not coverage, and three of the ten blocks change more than one -thing.** Derive each of the following from the **live** text now, check it the three ways, confirm -it counts 1 in each copy, and append it to the fragment table. **Step 2 installs over all of it**, -so a derivation afterwards has no live wording left to check against: +**Run the pre-install half of the disposition procedure over passages (a), (c) and (i) and over +§G** — every condition those ten blocks touch or reproduce, by its class, derived from the **live** +text now and appended to the fragment table. **Step 2 installs over all of it**, so a derivation +afterwards has no live wording left to check against — and that includes the **preservation +fragments for the nine carried conditions** `a15`, `a21`, `a22`, `c15` and `i4`–`i8`, which an +earlier draft chose from the installed text at step 4. + +**Passage (i)'s kept conditions need attention the class alone does not flag.** `i1`–`i3` and +`i9`–`i16` are kept, no untouched span reaches them — passage (i) is not one of Task 0's five +regions — and `i3` shares its sentence with the dash-delimited list this task replaces. Each owes a +per-condition preservation count, `parent=1 worktree=1`. + +**One pair per block is a sample, not coverage, and four of the ten blocks change more than one +thing:** - **The `c18`-and-surfacing block** changes `c16`, `c17`, `c19` and `c20` besides the `c18` clause row P8 observes — **one OLD row per condition**, four in total. -- **The `a17`–`a22` block moves four conditions, not one.** P11 observes `a17`, the clean-final-pass - rule; `a18`, `a19` and `a20` are each independently dispositioned **moved** and have no - observation at all. **One OLD row each**, and each is an *absence* at this source — - `parent=1 worktree=0` — paired with the matching §A text Task 1 already installed, since a moved - condition leaves here and appears there. Without them the old loop-until-clean, zero-finding and - no-padding instructions can survive beside the ordering that replaces them while every pair and - both carried-condition counts still pass. +- **The `a17`–`a22` block moves four conditions — `a17` included.** P11 pairs the old + clean-final-pass rule with §H's *pointer*, which is a different sentence from the predicate that + moved; so **all four of `a17`, `a18`, `a19`, `a20` owe the moved pair**: an *absence* at this + source, `parent=1 worktree=0`, **and a condition-specific presence in the installed §A**, + `worktree=1 parent=0`. Task 1's three paragraph-level fragments are not that second half. + **Observe §H's pointer separately** — P11's NEW is about the pointer arriving, not about the + predicate leaving. Without all of this the old loop-until-clean, zero-finding and no-padding + instructions can survive beside the ordering that replaces them while every pair and every + carried-condition count still passes. - **§G changes two things, not one**: its membership is widened *and* it gains a semantic membership test a downstream reader can apply, and target §G's own closing note names the test as the point of the replacement. P7 observes the widened clause; **the test owes a second @@ -1021,7 +1108,8 @@ so a derivation afterwards has no live wording left to check against: executor would have to pair the clause with unrelated text and the observation would prove nothing about the clause. One presence check per independent added clause. -**The other seven blocks change one thing each and one pair covers them.** +**The other six blocks change one thing each and one pair covers them** — six, because four of the +ten are named above. **One classification note, decided against the real file rather than asserted.** The gate-prompt template's clean sentence **replaces** `A clean pass is the single body line …`, which is why P13 @@ -1061,11 +1149,13 @@ Expected for all twenty: `old/worktree=0 old/parent=1 new/worktree=1 new/parent= - **a pair** for each of the `c18`-and-surfacing block's four further conditions — `c16`, `c17`, `c19`, `c20`; -- **an absence at this source plus a condition-specific presence in §A** for each of the three - moved conditions `a18`, `a19` and `a20` — `parent=1 worktree=0` here, `worktree=1 parent=0` - there. **An earlier draft derived these rows at step 1b and then never ran them**, so the old - loop-until-clean, zero-finding and no-padding instructions could each survive at their source - with nothing observing it; +- **an absence at this source plus a condition-specific presence in §A** for each of the four + moved conditions `a17`, `a18`, `a19` and `a20` — `parent=1 worktree=0` here, `worktree=1 + parent=0` there. **An earlier draft derived these rows at step 1b and then never ran them**, so + the old clean-final-pass, loop-until-clean, zero-finding and no-padding instructions could each + survive at their source with nothing observing it; +- **a preservation count** of `1` in each copy for each of the nine carried conditions and each + kept condition in passage (i), from the fragments step 1b appended; - **a presence check** — `new/worktree=1 new/parent=0` in each copy — for §G's semantic membership test and for every independent add-only clause in the strict-reading tail. @@ -1085,14 +1175,15 @@ around them and prove nothing about the reproduced words: both copies could omit premise, or any interior item of the strict-reading run, and every pair, presence check and parity comparison would still pass. +**All nine fragments were derived and appended at step 1b**, from the live text before step 2 +installed over it. Two of them, for orientation: + ```bash grep -cF 'Codex is advisory — validate before applying; dismissed finding → one-line why' CLAUDE.md grep -cF 'Open a TodoWrite' CLAUDE.md ``` -Expected: `1` each, and `1` for each of the remaining seven — take each fragment from the -condition's own text in the inventory, confirm it is single-line and unique, and run it in both -copies. +Expected: `1` each, and `1` in each copy for every one of the nine. - [ ] **Step 5: Parity** for all ten sites, each extracted by its own bounded region rather than a fixed line window. @@ -1130,6 +1221,13 @@ pre-edit count of 1 can no longer be observed at all. A row that no longer count since the table was verified — repair the row against the live line and update the table before installing anything. +**And derive `a1`'s preservation fragment here too.** `a1` is **carried** inside item 8a's block, +which reproduces it — so it is pre-existing text and the fragment-table cut puts it in the table, +before the install, like every other pre-existing fragment. An earlier draft chose it at step 4, +after item 8a had already replaced the block: at that point a drifted or half-installed HARD FLOOR +opening cannot be told from the intended carried text, and no authored fragment exists for Task 15 +to audit. + - [ ] **Step 2: Install all fourteen replacements** - [ ] **Step 3: Run the fifteen pairs the table already holds** @@ -1155,9 +1253,9 @@ Expected for all thirty pair instances — fifteen rows in each of the two copie **`a1` is carried inside item 8a's block**, which opens with it — `**Both gates are a LOOP with a HARD FLOOR: a minimum number of passes per run` — and reproduces it so one contiguous string installs. Carried, not kept: **no untouched-range span covers it**, and row F10's pair observes -`a2`, the parenthetical, not the opening it sits in. Count `a1`'s own text in each copy, expecting -`1`. Without it a mis-scoped item-8a replacement can drop the sentence's opening and every other -check in this task still passes. +`a2`, the parenthetical, not the opening it sits in. **Count the fragment step 1 appended**, in +each copy, expecting `1`. Without it a mis-scoped item-8a replacement can drop the sentence's +opening and every other check in this task still passes. Then total the fifteen `new/worktree` values step 3 printed, per copy. @@ -1611,8 +1709,14 @@ One row per **fenced block whose destination is a prompt copy** — so §H contr one per live site it replaces, not one row for the section. Its nine blocks land at nine noncontiguous places in each copy and no single start/end region spans them; a section-level row would either be unbuildable or would sample one of the nine and leave the other eight out of the -final parity diff. §A, §B, §C, §E and §G are contiguous and contribute one row each; §D -contributes two; every §F item whose destination is a prompt copy contributes one. +final parity diff. + +**Count the blocks, not the sections, everywhere** — including where a section's blocks happen to +be adjacent. **§A contributes three** (A1, A2, A3), **§D two**, **§E two** (the Severity bullet and +the demotion answer), **§H nine**; §B, §C and §G one each; every §F item whose destination is a +prompt copy, one. An earlier draft applied the block rule to §H and then gave §A and §E one row +each anyway, which is the section-level grouping the rule replaces, surviving in the two places it +looked harmless. For each, extract a **bounded** region — from its first line to the first line of the next passage, not a fixed count — and diff the two copies: @@ -1676,9 +1780,10 @@ line that matched the start, so a single-line site runs to the next match or to is why the squash-carry sentence is checked with `grep -n` instead. And not line numbers carried from Task 0: Task 1's insertion shifts everything after it in one tree and not the other. -**Plus the per-condition checks** for `a14`, `h6` and any other kept condition sharing a line with a -changed one: count the condition's own text in the parent and in the worktree, expecting `1` both -times. +**Plus every `cond` row in that file** — count the recorded fragment in the parent and in the +worktree and compare against the two expected values the row carries. **Read the set off the file, +not from a list here**: which kept conditions needed one is decided at Task 0 against the real +lines, and an enumeration in this step would be a second copy of it. **Anchors, resolved separately in each tree** — Task 1 inserts a large block, so a line number taken from either tree addresses different text in the other, and an earlier draft's numeric `sed` @@ -1856,12 +1961,23 @@ to pass CI, or an invariant-11 violation, published by a cycle that closed clean fix invalidate a reader record that then reached the closing commit unchanged.** - **Persist every re-run record and commit it before the re-review**, so the reviewed range holds it. A refreshed record left in the worktree is in neither the Gate-B range nor `reset --soft`. -- **Before the candidate final pass, re-run the complete set** — the whole step-4 battery against - the current `HEAD`, step 4b's twelve items, and every reader check above — commit those records, - and resolve the new `HEAD`. **Only a clean response issued against that exact `HEAD` closes the - cycle.** A pass that was clean against an earlier tree, plus records committed afterwards, closes - on a tree no pass reviewed — which is the same defect as reviewing the wrong range, arrived at - from the other end. +- **Before the candidate final pass, re-run the complete set, and "complete" means mechanical as + well as reader.** The whole step-4 battery against the current `HEAD`; step 4b's twelve items; + every reader check above; **and every mechanical observation this plan built** — each + discriminating pair, each add-only presence check, each moved and dropped absence, each carried + and kept preservation count — re-run and re-recorded in `## Fragment evidence (per-task output)`. + **An earlier draft named only the battery and the reader checks**, so the clean pass could close + on a `HEAD` whose fragment evidence had never been re-established as a set, and the plan's + design §7 claim would rest on counts taken from an older tree. +- **Commit those records, resolve the new `HEAD`, and only then issue the candidate final pass.** + **Only a clean response issued against that exact `HEAD` closes the cycle.** A pass that was + clean against an earlier tree, plus records committed afterwards, closes on a tree no pass + reviewed — which is the same defect as reviewing the wrong range, arrived at from the other end. +- **One thing cannot exist before that pass: the pass's own findings file.** It is therefore the + **sole permitted post-review addition**, and step 8 asserts that it is the only one — every other + path must already be in the reviewed `HEAD`. `.context/` moves none of the hook's fingerprint + inputs, so the file changes nothing the review looked at; what would be wrong is a *second* + delta riding along beside it. - [ ] **Step 7b: After the clean pass, complete `.context/loop-rule-closing-msg` — this is the action step 5 defers to, and step 8 has no other source for these records** @@ -1897,9 +2013,15 @@ BASE=$(cat .context/loop-rule-base) test -n "$BASE" || { echo "BASE empty — Task 0 did not run"; exit 1; } # Everything that must be IN the squashed commit has to be committed before the # reset: reset --soft stages only what the discarded commits already contained. -git add docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md .context/codex-reviews/ +# The ONLY thing allowed to differ from the reviewed HEAD is the final pass's own +# findings file, which could not exist when that pass was issued. Anything else +# means a change rode along that no pass reviewed. +git status --porcelain | grep -v '^.. \.context/codex-reviews/gate-b-' && { + echo "untracked or modified paths beyond the final findings files — NOT closing"; exit 1; } + +git add .context/codex-reviews/ if ! git diff --cached --quiet; then - git commit -m "WIP: plan records and Gate-B findings files" || { echo "record commit FAILED"; exit 1; } + git commit -m "WIP: Gate-B findings files" || { echo "record commit FAILED"; exit 1; } fi test -z "$(git status --porcelain)" || { echo "tree not clean before reset — aborting"; exit 1; } @@ -1933,10 +2055,15 @@ state where it is needed. removed afterwards so a later run cannot inherit a stale one. **`reset --soft` stages committed content only.** The prompt-standards result (step 4b), the -completeness sweep, the next-state table and the divergence list all land in this plan, and -`.context/codex-reviews/` is tracked; **anything still uncommitted when the reset runs is left in the -worktree and is not in the closing commit** — and, for the plan records, was never in a Gate-B range -either. Commit them first, then reset. +completeness sweep, the next-state table, the divergence list, the equivalence result and the +fragment evidence all land in this plan, and `.context/codex-reviews/` is tracked; **anything still +uncommitted when the reset runs is left in the worktree and is not in the closing commit** — and, +for the plan records, was never in a Gate-B range either. + +**Which is why step 7 commits them *before* the candidate final pass, and this step only adds the +findings file that pass produced.** The guard above is what keeps that distinction real: if any +other path is still dirty here, a record was written after the review rather than before it, and +the closing tree would differ from the one the clean pass read. **`$BASE` is the recorded revision, not a placeholder to substitute by hand** — Task 0 persisted it for this, and a mistaken substitution squashes the wrong range. From ff17edde9150c6bcbc314c4b455ac7749494731b Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 17:21:53 +0200 Subject: [PATCH 113/181] docs(context): record passes 8-10 in the Gate-A plan working record Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-plan-om0bdd7udh-resume.md | 25 +++++++++++++++---- 1 file changed, 20 insertions(+), 5 deletions(-) diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md index f93bd9d..df5e851 100644 --- a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md @@ -13,13 +13,25 @@ Nothing depends on it; the pass files and the repo are authoritative where this ## Resume here -**Next action:** Gate-A plan pass 8 against `54e793b`. The prompt is +**Next action:** Gate-A plan pass 11 against `55a27c9`. The prompt is `.context/gate-a-plan-prompt.md`; substitute `__SHA__` and `__P__`, precheck per its header, delete the target file, confirm it is gone. -**Three rules pass 7 installed, each of which had been contradicted in two or three places at -once.** Check them every pass: a row a task must derive is derived **before** that task's install -step; **no task pre-assigns a fragment id**; the fragment table holds **OLD fragments only**. +**The rules passes 6–10 installed, every one of which had been contradicted in two or three places +at once when found.** Check them every pass: + +- `## What each disposition owes, stated once` fixes the observation per class; `## How a task + discharges that table` is the per-task procedure. **No task enumerates its condition ids.** +- Every **pre-existing** fragment is derived **before** its task's install step — OLD halves, + absence fragments, and carried/kept preservation fragments alike. Only post-install fragments + go to `## Fragment evidence (per-task output)`. +- **No task pre-assigns a fragment id.** Tasks 3, 4, 6, 7 and 10 append. +- Untouched **spans are derived, not written out**. `.context/loop-rule-untouched` carries `span` + and `cond` records, both parseable, both consumed by Tasks 2 and 14. +- **Destination blocks, not target sections**, are the unit for the parity site list. +- Task 15 re-runs every affected check — mechanical *and* reader — after each Gate-B fix, and the + complete set before the candidate final pass; only a clean response against that exact `HEAD` + closes, and the final pass's own findings file is the sole permitted post-review addition. **That prompt file is untracked.** `.gitignore` carries `.context/*` with only `codex-gate.on` and `codex-reviews/` exempt, so it survives a context clear but not a `.context/` cleanup. **If it is @@ -65,7 +77,10 @@ worth checking before a pass rather than after. | 5 | 58b3660 | 20→**16** | 10→**7** | 7→**7** | yes | **MANDATORY TWO-TELL STOP** — instrument cluster for the fifth pass, plus a require↔withdraw pair (pass 4 demanded OLD rows for clauses pass 5 classifies as add-only). Surfaced, standing answer applied, loop continued. **The pre-written shell is deleted.** Design §7 says the plan builds each pair *against the real files*; five passes of findings were blocks written in advance for text that does not exist yet. One stated procedure replaces them | | 6 | 33cdfa8 | 16→**11** | 7→**2** | 7→**6** | yes | one tell only (instrument cluster), no mandatory stop. **Task 11's sweep locator matched only `Gate B not satisfied` — 3 hits — and missed the 25 assertions greping the bare verdict word**, which §F items 15/16 remove: the suite would have gone red and the battery could not have passed. Task 15 step 5 deferred the provenance line, curve and human-exception record to an action step 7 did not contain. Passage (h) counted 24 kept less `h4`/`h19` where `h5` is replaced too, and the untouched middle span enclosed F7b's line. **Five tasks derived OLD fragments after their own install step.** Task 8 rebuilt fourteen rows the table already held as fifteen, dropping F7b. §G's semantic membership test had no observation; the strict-reading tail's add-only clauses were told to produce OLD rows they cannot have; Task 10 promised seven pairs and listed six | | 7 | 34250be | 11→**13** | 2→**4** | 6→**4** | yes | **MANDATORY THREE-TELL STOP** — findings rose, Blockers rose, instrument cluster for the seventh pass. Surfaced, standing answer applied, loop continued. **Two of the four Blockers were pass 6's own repairs half-applied**: Task 4 kept its post-install derivation beside the new pre-install one, and Task 3's "ids continue the `P` series" collided with Task 4's pre-assigned `P19`/`P20`. The other two are five-pass survivors — `git log "$BASE"..HEAD` standing in for an ancestry test, and the single-line squash-carry site inside a `sed` range the same task forbids two paragraphs earlier. **Seventh condition-table misclassification in seven passes:** §B was credited with a second added rule that lives in §A | -| 8 | — | — | — | — | **not run — NEXT** | against `54e793b`. Prompt: `.context/gate-a-plan-prompt.md`, substitute `__SHA__`=`54e793b` and `__P__`=`8` | +| 8 | 54e793b | 13→**15** | 4→**2** | 4→**9** | yes | **MANDATORY TWO-TELL STOP** (findings rose, instrument cluster). Task 15 step 7 never re-ran the battery after a Gate-B fix; step 8 suppressed its record commit with `\|\| true`. Two more condition misclassifications: `a1` is carried inside §F item 8a's block **and the untouched span opened on that line**; `e10` is kept outside §D. `c10`–`c13` and `a18`–`a20` are moved and had only a reader walk; nine carried conditions owed preservation checks and three had them | +| 9 | ebb371b | 15→**9** | 2→**2** | 9→**6** | yes | **MANDATORY TWO-TELL STOP** (Blockers flat, instrument cluster). The floor span ended on the line carrying both changed `a13` and kept `a14`. **Root repair: `## What each disposition owes, stated once`** — kept/carried/replaced/moved/dropped/add-only. Spans are now derived, not written out. The fragment table's cut became "exists before the edit" | +| 10 | ba614f6 | 9→**17** | 2→**1** | 6→**14** | yes | **MANDATORY TWO-TELL STOP** (findings rose, instrument cluster). **Eleven of seventeen were one family: the tasks had not been brought into line with pass 9's table.** Root repair: **no task enumerates its condition ids** — `## How a task discharges that table` states the procedure and each task walks its passage's rows by class. A *kept* condition inside a wholly replaced passage has no span and owes a per-condition count; carried and kept preservation fragments are pre-existing and derived before the install | +| 11 | — | — | — | — | **not run — NEXT** | against `55a27c9`. Prompt: `.context/gate-a-plan-prompt.md`, substitute `__SHA__`=`55a27c9` and `__P__`=`11` | ## Pass-1 report From 9a79e01af0d6af3813f5c92a941f025991c57158 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 17:40:50 +0200 Subject: [PATCH 114/181] docs(plans): apply Gate-A plan pass 11; bound each task by its own blocks MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit MANDATORY TWO-TELL STOP — Blockers rose 1 -> 2 and the findings cluster on the verification apparatus for the eleventh pass. Findings fell 17 -> 11. Surfaced to Daniel; standing answer applied, loop continued. Two Blockers and seven Majors, all validated against the real files. Both Blockers and the first Major are pass 10's own repair, applied too widely. - Task 3 step 3 said "run a pair for every row step 1b derived", but step 1b now derives carried rows too, and a carried fragment has to still be there. A correct §B install could not satisfy it. Pairs are for replaced rows; the procedure now says each row runs to ITS OWN class's result. - The per-task walk was bounded by the PASSAGE, but passage (c) is split between Tasks 4 and 7 and passage (a) between Tasks 7 and 8. Task 4 would append rows it cannot discharge after its own install and both tasks would append a row for the same condition. The unit is the source BLOCK a task replaces. - Task 15 step 8's guard admitted any path starting gate-b- — including a modified or deleted findings file from another pass — and then staged the whole review directory. It now names the final pass's exact slot paths and requires the dirty set to be those, added, and nothing else. - Task 15 step 8 had no failure path. After a failed closing commit HEAD is at BASE with everything staged, and re-running hits the dirty-tree guard: the documented procedure was unrecoverable at its riskiest point. The pre-reset tip is recorded and restored on failure. Also: - c14 is replaced by two noncontiguous outcomes, suspend and continue; one NEW fragment observed one of them. - Task 14 interpolated DERIVED anchors into sed as basic regular expressions, with no escaping, no uniqueness test and no empty-extraction test — and two empty extractions diff equal, certifying parity for a block sed never found. - Task 14's site list lived only in ignored .context, so a missing destination row and a clean comparison left the same durable record. - Step 4b recorded a per-item result with no pass expectation and no repair path, so a recorded failure could be committed and the cycle continue. - The fragment-evidence section had one shape for pairs and named no preservation counts at all, so a required observation could be run and then vanish from the plan and the closing evidence. Four shapes now. - Step 5's moved-condition enumeration still said a18-a20 after a17 joined them; the enumeration is gone, per the no-enumeration rule. Trend: findings 32, 32, 31, 20, 16, 11, 13, 15, 9, 17, 11. Blockers 1, 0, 0, 10, 7, 2, 4, 2, 2, 1, 2. Majors 28, 25, 23, 7, 7, 6, 4, 9, 6, 14, 7. Floor 3, from risk high (2) and security none (0), read at this pass from docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-plan-om0bdd7udh-pass-11.md | 12 ++ .../2026-09-14-loop-rule-consolidation.md | 187 ++++++++++++++---- 2 files changed, 161 insertions(+), 38 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-11.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-11.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-11.md new file mode 100644 index 0000000..3553065 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-11.md @@ -0,0 +1,12 @@ +BLOCKER | high | Task 3 step 3 | The instruction to run a pair for every row step 1b derived includes the nine carried-condition rows that step 1b must append, even though a carried fragment must remain present in the worktree | A correct §B installation cannot satisfy old/worktree=0 for those rows, so the task either fails on correct text or requires the executor to ignore a direct instruction | Limit discriminating pairs to replaced rows and run carried rows as parent=1 worktree=1 preservation counts under the disposition procedure +MAJOR | high | Task 4 step 1 | The pre-install walk says to cover every condition in passage (c), whose inventory and disposition include the later Surfacing block, while Task 4 edits only the clearly-stuck block and Task 7 derives the Surfacing rows again | Task 4 appends rows it cannot discharge after its own install and Task 7 can append duplicate rows for the same conditions, breaking the fragment table's one-copy rule | Bound Task 4's walk by the actual source block it replaces, ending before the Surfacing paragraph, and leave the following replacement block to Task 7 without adding a per-task id list +MAJOR | high | Task 4 step 4 | The c14 replacement is assigned one NEW fragment even though the approved target explicitly replaces the old unconditional continuation with two noncontiguous outcomes: suspension when one applies and continuation when none does | Either outcome can be omitted while the single c14 pair, the three §A paragraph-presence samples and parity still pass | Observe both destination clauses separately, with the P4 absence paired to one condition-specific NEW and a second condition-specific add-only presence check for the other +MAJOR | high | Task 14 step 1 | The generated start and end anchors are interpolated directly as sed basic regular expressions without requiring escaping, uniqueness or nonempty extraction; target first lines commonly contain leading asterisks, slashes and other regex or delimiter characters | Both sed processes can error or select the wrong range and feed equal empty or unrelated output to diff, falsely certifying parity for an unchecked destination block | Require sed-safe escaped patterns, assert exactly one start and end match per copy and fail on an empty extraction before comparing the two regions +MAJOR | medium | Task 14 steps 1-2 | The derived destination-site list and raw parity results live only in ignored .context scratch state, while the committed Divergence list records differences but not which target blocks were actually examined | A missing destination row and a complete no-difference run produce the same durable record, so Task 15 cannot substantiate the block-by-block parity claim | Copy the derived site identifiers and each site's parity result into the Task 14 output section before committing, while continuing to derive the set from target markers +MAJOR | high | Task 15 step 4b | The prompt-standards review says to record a result per checklist item but never states that every item must pass before Gate B or gives the repair and rerun path for a failed item | An executor can commit a recorded failure and continue, or repair prompt text after the battery without re-running the checks that repair invalidated, violating invariant 11 | State an all-twelve-pass expected result, stop on any failure, repair it and rerun every affected mechanical and reader check plus the battery before committing the refreshed record +MINOR | medium | Task 15 step 4b output | One line per checklist item records no inventory of the C blocks, W blocks and seven hook strings actually examined | A twelve-line PASS record is indistinguishable from a review that skipped one copy or some hook messages | Record the derived subject set once in the output section and make each checklist result explicitly cover that set without adding a mechanical guard +MAJOR | high | Fragment evidence output and Task 15 step 5 | The only durable fragment-evidence schema names pairs, presence checks and absence checks but omits carried and kept preservation counts, and it ambiguously assigns four counts to standalone presence fragments; Task 15 later requires preservation counts to be re-recorded but reads this section as though it cannot contain them | Required preservation observations can be run and then disappear from both the plan record and closing evidence, leaving carried or kept text unsupported by the claimed complete verification set | Define separate durable record shapes for four-value pairs, two-value add-only presence, moved and dropped observations, and two-value preservation counts, then include every class when assembling the closing evidence +MINOR | high | Task 15 step 5 | The supposedly explanatory moved-condition list names a18-a20 but omits a17, although the disposition and Task 7 correctly classify all four as moved and explain that P11's pointer is not a17's destination observation | The closing-evidence instructions contradict the repaired disposition and can drop a17's condition-specific destination presence from the summary | Remove the stale id enumeration and derive every moved observation from the disposition table, or include a17 consistently if the explanatory list is retained +BLOCKER | high | Task 15 step 8 | The dirty-tree guard permits every status entry whose path merely starts with .context/codex-reviews/gate-b-, then stages the entire review directory, although the plan permits only the candidate final pass's own findings artifacts after the reviewed HEAD | Modified, deleted or unrelated findings files from another pass or cycle can ride into the closing commit without review while all closing checks pass | Derive the exact final-pass output paths from the issued review, require the status set and status kinds to match only those paths and stage only those exact files +MAJOR | high | Task 15 step 8 failure path | If the closing commit fails after the soft reset, HEAD is already at BASE with the whole implementation staged, but re-running step 8 immediately rejects those staged product paths at its initial dirty-tree guard | The promised recovery base does not make the documented close procedure recoverable after common identity, signing or hook failures, so the executor must invent a history-recovery sequence at the riskiest point | Save the pre-reset WIP tip and add an explicit failure branch that restores that tip before returning to the affected checks and Gate-B loop, retaining the recovery records until a later close succeeds +END OF FINDINGS (11 total) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index 2b675fa..51217d4 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -194,19 +194,29 @@ condition missing its observation, and every one of them was in a task that had its ids and not the rest. **An enumeration per task is a second copy of the disposition table**, and it goes stale exactly the way every other enumeration in this cycle has. -**Each editing task instead runs this, against the disposition table's rows for its own passage:** +**The unit is the source block a task replaces, not the passage it sits in.** A passage can be +edited by two tasks — passage (c) is split between Task 4's clearly-stuck block and Task 7's +`c18`-and-surfacing block, and passage (a) between Tasks 7 and 8 — so **a task walks the conditions +its own replaced blocks cover, and stops at their boundaries.** Walking the whole passage makes a +task append rows it cannot discharge after its own install, and makes two tasks append a row each +for the same condition, which is the fragment table's one-copy rule broken from a new direction. -- [ ] **Before installing** — walk every condition in this task's passage and derive the +**Each editing task runs this, against the disposition table's rows for its own blocks:** + +- [ ] **Before installing** — walk every condition those blocks cover and derive the **pre-existing** fragment its disposition owes: the **OLD** half for a *replaced* one, the **absence** fragment for a *moved* or *dropped* one, the **preservation** fragment for a *carried* one, and, for a *kept* one, a preservation fragment **only where no untouched span can hold it** — which is the case for every kept condition in a passage this task replaces whole, and for any that shares a line with changed text. Check each the three ways, confirm its expected pre-edit count, and append it to the fragment table under the next free `P` id. -- [ ] **After installing** — run each one to the result its class owes, and choose the - **post-install** halves: the **NEW** for each pair, and the destination **presence** for each - moved condition, taken from the installed text and verified the three ways before counting. - Record those in this task's fragment evidence. +- [ ] **After installing** — run each one **to the result its own class owes, which is not the same + result for all of them**: a *replaced* row to the four-value pair, a *moved* or *dropped* row to + `parent=1 worktree=0`, a *carried* or *kept* row to `parent=1 worktree=1`. **A step that says + "run a pair for every row" cannot be satisfied on correct text**, because a carried fragment has + to still be there. Choose the **post-install** halves — the **NEW** for each pair, the + destination **presence** for each moved condition — from the installed text, verified the three + ways before counting. Record those in this task's fragment evidence. - [ ] **Then walk the conditions as a reader**, which confirms the counts and replaces none of them. **Two consequences worth stating, because both have been got wrong.** A **kept** condition inside a @@ -713,10 +723,15 @@ three ways before this step counts with it. Expected for all four: `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`. -**Run a pair for every row step 1b derived**, and a **presence check** for §B's one added rule — -the closing-time set-change rule — which replaces no wording and so owes `new/worktree=1 -new/parent=0` and no OLD half, per the add-only rule. Choose each NEW fragment from the installed -text and verify it the three ways before counting it. +**Run a pair for every *replaced* row step 1b derived** — and only those. §B's carried rows are in +that set too and they owe `parent=1 worktree=1`, not a pair; step 4 runs them. A blanket "a pair +for every row" cannot be satisfied by a correct §B installation, because carried wording has to +still be there. + +Plus a **presence check** for §B's one added rule — the closing-time set-change rule — which +replaces no wording and so owes `new/worktree=1 new/parent=0` and no OLD half, per the add-only +rule. Choose each NEW fragment from the installed text and verify it the three ways before counting +it. - [ ] **Step 4: Run the post-install half, then walk the passage** @@ -765,10 +780,14 @@ Expected: `P4=1` for both files. is **preserved in target §C's fenced replacement**, so its old-count could never reach zero. Derive `c4`'s OLD from the part of the sentence the replacement removes. -**Run the pre-install half of the disposition procedure over passage (c)** — every condition in it, -by its class, derived from the live passage now and appended to the fragment table under the next -free `P` id. Refer to the rows afterwards by the condition they observe, never by a number picked -here: Task 3 appends first and how many rows it adds is decided at execution. +**Run the pre-install half of the disposition procedure over the block this task replaces** — the +clearly-stuck block, **ending before the Surfacing paragraph**, which is Task 7's +`c18`-and-surfacing block and whose conditions Task 7 derives. Passage (c) is edited by two tasks; +walking the whole passage here makes this task append rows it cannot discharge after its own +install, and makes both tasks append a row for the same condition. Derive each from the live text +now and append it under the next free `P` id, referring to the rows afterwards by the condition +they observe, never by a number picked here: Task 3 appends first and how many rows it adds is +decided at execution. **Two things about passage (c) that the class alone does not tell you:** @@ -811,10 +830,18 @@ observation at all. **A later draft reintroduced exactly that pair** by naming r |---|---|---| | `c4`, the widened third condition | the `c4` row (step 1) | the clause §C puts in place of "a missing one means keep going" | | `c8`, the re-raised dismissal | the `c8` row (step 1) | `a recurrence failing them being an ordinary fresh finding` — **install that clause's line unwrapped** so the fragment sits wholly on one line | -| `c14`, the below-floor Minor | row **P4** | the ordering's replacement for the below-floor sentence | +| `c14`, the below-floor Minor — **suspend** | row **P4** | the ordering's **suspension** outcome for a below-floor pass | +| `c14`, the below-floor Minor — **continue** | *(none — add-only)* | the ordering's **continuation** outcome where no suspension applies — presence alone, `worktree=1 parent=0` | Each NEW is confirmed single-line and unique in the installed file before it is counted. +**`c14` is replaced by two noncontiguous outcomes and owes two observations.** The old sentence +stated one unconditional continuation; the ordering splits it into a **suspension** where one +applies and a **continuation** where none does, and they do not sit together. One NEW fragment +observes one of them, so the other can be omitted while the pair, the three §A paragraph samples +and the parity diff all pass. P4's absence pairs with the suspension clause; the continuation +clause is add-only at its destination and owes presence. + **Plus both halves of every move — `c10`, `c11`, `c12`, `c13`, and `c9`'s moved clause.** A move removes wording here and adds it in §A, so it owes an **absence** at this source, `parent=1 worktree=0`, **and a condition-specific presence in the installed §A**, `worktree=1 parent=0`. @@ -1070,9 +1097,11 @@ Expected: `1` twenty times. **Any `0` means the wording drifted since this table - [ ] **Step 1b: Derive the rows the ten locators do not cover, still before installing** -**Run the pre-install half of the disposition procedure over passages (a), (c) and (i) and over -§G** — every condition those ten blocks touch or reproduce, by its class, derived from the **live** -text now and appended to the fragment table. **Step 2 installs over all of it**, so a derivation +**Run the pre-install half of the disposition procedure over this task's ten blocks** — every +condition *they* touch or reproduce, by its class, derived from the **live** text now and appended +to the fragment table. **Their boundaries are the bound**: passage (c) is shared with Task 4, which +owns the clearly-stuck block and stops before the Surfacing paragraph this task's `c18` block +begins at, and passage (a) is shared with Task 8, which owns item 8a's block. **Step 2 installs over all of it**, so a derivation afterwards has no live wording left to check against — and that includes the **preservation fragments for the nine carried conditions** `a15`, `a21`, `a22`, `c15` and `i4`–`i8`, which an earlier draft chose from the installed text at step 4. @@ -1728,11 +1757,29 @@ not a fixed count — and diff the two copies: # copy. Build it by reading those markers off the target text, then: while IFS=$(printf '\t') read -r s e; do echo "== $s" - diff <(sed -n "/$s/,/$e/p" CLAUDE.md) \ - <(sed -n "/$s/,/$e/p" plugins/dev-workflow/commands/workflow-init.md) + for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do + # Anchors are DERIVED from the target text, so they carry ** and / and . + # Unescaped they are a basic regular expression, not a literal. + test "$(grep -cF "$s" "$f")" = 1 || { echo "start anchor not unique in $f"; exit 1; } + test "$(grep -cF "$e" "$f")" = 1 || { echo "end anchor not unique in $f"; exit 1; } + done + se=$(printf '%s' "$s" | sed 's/[][\.*^$\/]/\\&/g') + ee=$(printf '%s' "$e" | sed 's/[][\.*^$\/]/\\&/g') + a=$(sed -n "/$se/,/$ee/p" CLAUDE.md) + b=$(sed -n "/$se/,/$ee/p" plugins/dev-workflow/commands/workflow-init.md) + test -n "$a" && test -n "$b" || { echo "empty extraction for '$s' — region not found"; exit 1; } + printf '%s\n' "$a" > .context/loop-rule-a.txt + printf '%s\n' "$b" > .context/loop-rule-b.txt + diff .context/loop-rule-a.txt .context/loop-rule-b.txt done < .context/loop-rule-changed-sites ``` +**Three assertions, and each one has already been the failure mode somewhere in this cycle.** +A derived anchor containing `**`, `/` or `.` is a **regular expression** to `sed`, not a literal, +so it can error or match the wrong line; a non-unique anchor selects a region that is not the one +named; and **two empty extractions diff equal**, which certifies parity for a block neither `sed` +ever found. Without the empty test that failure is silent and looks like success. + **The list is a step output, not an assumed input.** An earlier draft gave one generic command with undefined `$start` and `$end` and no step producing them, so the task could record a divergence list without having diffed anything. @@ -1745,6 +1792,14 @@ what makes §H's nine sites visible; reading it section by section is what hid e **Where it goes:** this plan, under the heading `## Divergence list (Task 14 output)` at the end of the document, replaced idempotently on re-run by the same rule Task 13 uses. +**It carries the site list as well as the differences — one line per destination block examined, +with its parity result, "no difference" included.** The derived list lives in `.context/`, which is +ignored, so nothing durable would otherwise say *which* blocks were compared. **A site missing from +the list and a site that compared equal produce the same record** if only differences are written +down, and Task 15's block-by-block parity claim would then rest on a record that cannot +distinguish them — the same failure Task 12b's "including nothing" rule exists for. Continue to +derive the set from the target's markers; copy the identifiers and results here. + - [ ] **Step 2: Classify every difference the diff reports** Three buckets: **deliberate and stays** (the field-mint parenthetical, `e11`, `f5`–`f7`'s evidence @@ -1871,6 +1926,19 @@ risk. **A reader check, deliberately** — no pattern decides whether a constraint carries its reason. +**Expected result: all twelve pass. A failure stops this step.** Invariant 11 is not "record the +score"; recording a failure and continuing ships a prompt change that violates it, which is the one +outcome this step exists to prevent. **On any failure: repair the text, then re-run every check the +repair invalidated** — the affected pairs, presence, absence and preservation counts, the parity +and untouched checks over the blocks touched, and the whole step-4 battery — and only then commit +the refreshed record. A repair made after the battery, with the battery not re-run, is the same +defect as a Gate-B fix with no re-run. + +**Record the subject set once, at the top of the output section**, and make each item's result +refer to it: the §A–§H blocks installed in C, the same in W, and the seven hook strings. +**Twelve `PASS` lines alone cannot be told from a review that skipped a copy or the hook** — the +subject list is what makes the twelve lines mean something. + **Commit the result before step 6.** Gate B reviews the range `$BASE..HEAD`; an edit to this plan left in the worktree is in neither that range nor the final `reset --soft`, which stages only what the discarded commits contained. The same applies to every record this plan collects — the sweep, the @@ -1905,7 +1973,13 @@ It names: the battery run; **every pair this plan built, with its counts in each for those two reader records. **An entry assembled from one section would silently drop whatever the other five hold.** -**The absence checks belong in the entry as much as the pairs do.** `g2`, `g3` and `g4` are dropped rather than replaced, and `c9`–`c13` and `a18`–`a20` are moved, so the observation that proves each is gone is an absence — `parent=1 worktree=0` — and nothing else in the entry carries it. An entry listing only pairs and presence checks claims the verification set while omitting the half that proves obsolete instructions were removed, which is the failure the absence checks exist for. +**All four record shapes belong in the entry, not only pairs and presence.** A **dropped** +condition's absence and a **moved** condition's absence are what prove an obsolete instruction was +removed; a **carried** or span-less **kept** condition's preservation count is what proves +reproduced text survived. An entry listing only pairs and presence claims the verification set +while omitting both. **Read the set off the disposition table and the fragment evidence, not from +an enumeration here** — an id list in this step was already stale once, naming `a18`–`a20` after +`a17` had joined them. **State the §A presence checks as presence, not as pairs** — its counterfactual is absent and is claimed as absent. @@ -2014,27 +2088,50 @@ test -n "$BASE" || { echo "BASE empty — Task 0 did not run"; exit 1; } # Everything that must be IN the squashed commit has to be committed before the # reset: reset --soft stages only what the discarded commits already contained. # The ONLY thing allowed to differ from the reviewed HEAD is the final pass's own -# findings file, which could not exist when that pass was issued. Anything else -# means a change rode along that no pass reviewed. -git status --porcelain | grep -v '^.. \.context/codex-reviews/gate-b-' && { - echo "untracked or modified paths beyond the final findings files — NOT closing"; exit 1; } - -git add .context/codex-reviews/ -if ! git diff --cached --quiet; then - git commit -m "WIP: Gate-B findings files" || { echo "record commit FAILED"; exit 1; } -fi +# findings files, which could not exist when that pass was issued. Name them +# EXACTLY — from the slot paths the final call was told to write — and require the +# dirty set to be those paths, added, and nothing else. +# NONCE and P are this cycle's nonce and its final pass number — the same two the +# final call's slot paths were built from. Set them to the values you used. +NONCE=; P= +FINAL=".context/codex-reviews/gate-b-spec-$NONCE-pass-$P.md .context/codex-reviews/gate-b-quality-$NONCE-pass-$P.md" +expected=$(printf '%s\n' $FINAL | sort) +actual=$(git status --porcelain | sed -n 's/^?? //p; s/^A //p' | sort) +test "$expected" = "$actual" || { + echo "dirty set is not exactly the final pass's findings files:"; git status --porcelain; exit 1; } + +# shellcheck disable=SC2086 +git add $FINAL +git commit -m "WIP: Gate-B findings files" || { echo "record commit FAILED"; exit 1; } test -z "$(git status --porcelain)" || { echo "tree not clean before reset — aborting"; exit 1; } +# Keep the pre-reset tip: after the reset it is the ONLY way back to the reviewed +# history if the closing commit fails. +git rev-parse HEAD > .context/loop-rule-wip-tip git reset --soft "$BASE" -git commit -F .context/loop-rule-closing-msg || { echo "closing commit FAILED — base file kept for recovery"; exit 1; } +git commit -F .context/loop-rule-closing-msg || { + echo "closing commit FAILED — restoring the reviewed tip" + git reset --hard "$(cat .context/loop-rule-wip-tip)" + echo "history restored; fix the cause, then re-run step 8. Base and tip files kept." + exit 1; } case "$(git log -1 --pretty=%s)" in WIP:*|wip:*) echo "closing commit still reads WIP — cycle NOT closed"; exit 1 ;; esac test -z "$(git status --porcelain)" || { echo "worktree dirty after close — aborting before cleanup"; exit 1; } -rm -f .context/loop-rule-base +rm -f .context/loop-rule-base .context/loop-rule-wip-tip ``` +**The failure branch is the point of `.context/loop-rule-wip-tip`.** A closing commit can fail on +an unset identity, a signing key or a hook, and by then the reset has already run: `HEAD` is at +`$BASE` with the whole implementation staged. **Re-running step 8 from that state hits its own +dirty-tree guard and refuses**, correctly — the guard cannot tell a staged implementation from a +record that rode along. `reset --hard` back to the recorded tip returns the repository to the +reviewed history, with every WIP commit and the reviewed `HEAD` intact, so the executor fixes the +cause and re-runs rather than inventing a recovery sequence at the riskiest moment. + +**Both files are removed only after a successful close**, so a later run cannot inherit either. + **Every check in that block fails the script; none of them is a comment.** An earlier draft suppressed the record commit with `|| true` and left the status and log lines as things to look at, so a failed record commit, a dirty tree or a closing message still reading `WIP:` all proceeded @@ -2121,9 +2218,23 @@ as extracted and the result in both directions, per copy.* ## Fragment evidence (per-task output) -*Empty until the tasks run. One subsection per task — `### Task 1`, `### Task 3`, … — each holding -that task's chosen NEW and presence fragments with the four counts observed for each, and the -absence checks with their two. A task replaces its own subsection idempotently on re-run and -touches no other. **The OLD fragments live in the fragment table, never here**; this section is -where the halves that do not exist until installation are recorded, and it is what Task 15 step 5 -reads to assemble the closing evidence entry.* +*Empty until the tasks run. One subsection per task — `### Task 1`, `### Task 3`, … — replaced +idempotently on re-run, touching no other. **The pre-existing fragments live in the fragment table, +never here**; this section records what each observation actually returned, and it is what Task 15 +step 5 reads to assemble the closing evidence entry.* + +***Four record shapes, because the classes do not return the same number of values.** Every +observation this plan makes is one of them, and a shape that fits only pairs is how a required +count gets run and then vanishes from both the plan and the closing evidence:* + +``` +pair old/worktree old/parent new/worktree new/parent +presence worktree parent # add-only, and a moved condition's destination +absence parent worktree # dropped, and a moved condition's source +preservation parent worktree # carried, and kept where no span holds it +``` + +*A **moved** condition therefore contributes two lines — one `absence` at its source, one +`presence` at its destination — and both are required for it to count as observed. **Every one of +the four classes goes into the closing evidence entry**; naming only pairs and presence leaves the +absences and preservations run but unrecorded.* From 1d5a892b282c1be1d3bd90b90f8aae3d8c474bc5 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 17:55:36 +0200 Subject: [PATCH 115/181] docs(plans): apply Gate-A plan pass 12; split the closing act into two invocations MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit One tell only (instrument cluster) — no mandatory stop. Findings fell 11 -> 6, Blockers 2 -> 1, Majors 7 -> 5. One Blocker and five Majors, all validated against the real files. - codex-gate.sh's is_wip_commit greps the WHOLE command string for -m ... wip (hooks/codex-gate.sh:763). Step 8 carried the WIP record commit and the real closing commit in one block, so the hook would have read the closing attempt as cycle-internal and carried this cycle's Gate-B count and fingerprint into the next one. Split into 8a and 8b as separate invocations, with neither the flag nor the word in 8b's command string. - The failure branch used reset --hard, which deletes the delta a failed closing act leaves — including anything a commit hook staged or modified. Target §A1 requires that delta to be inspected and the closure conditions re-established against it. It is a mixed reset now, and git status is the delta. - The recovery path could not be re-entered: after a failed close the findings files are already committed, so the dirty set is empty and the exact-dirty-set guard refuses. 8a has a retry branch that admits exactly that state. - Task 5 asked e8 for its own OLD half, which P5w already is. §D supplies ONE sentence for both copies, so in W the read-after clause and the pronoun install or fail together and no fragment can fail for one and not the other. The discrimination is in the NEW halves: two e7 pair instances plus a W-only presence check for the pronoun. - Task 12b wrote its sweep record to .context/loop-rule-sweep.md, which is ignored, while step 6 committed the plan and the paragraph below said the record belongs in the plan. A literal run committed an empty section. - Task 13's minimum rows named scenarios whose route the ordering does not determine — a clean pass with an unmet precondition, a clean pass below the floor. A row with more than one admissible route tests nothing. Rows now state every predicate their route depends on and split until determinate. Trend: findings 32, 32, 31, 20, 16, 11, 13, 15, 9, 17, 11, 6. Blockers 1, 0, 0, 10, 7, 2, 4, 2, 2, 1, 2, 1. Majors 28, 25, 23, 7, 7, 6, 4, 9, 6, 14, 7, 5. Floor 3, from risk high (2) and security none (0), read at this pass from docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-plan-om0bdd7udh-pass-12.md | 7 + .../2026-09-14-loop-rule-consolidation.md | 120 ++++++++++++------ 2 files changed, 86 insertions(+), 41 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-12.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-12.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-12.md new file mode 100644 index 0000000..e31a55e --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-12.md @@ -0,0 +1,7 @@ +MAJOR | high | Task 5 steps 1 and 3 | P5 and P5w are explicitly the C and W OLD halves for e7, but step 3 also assigns P5w to e8 and expects only two pairs; the plan therefore cannot both run e7 in both copies and give W-only e8 its own OLD/NEW observation | Either W's read-after-clean-completion change or its pronoun alignment lacks the disposition-required discriminating observation, and reusing P5w lets one OLD boundary stand for two edits | Run the e7 pair in C and W with P5 and P5w, then add a separately named W-only e8 observation and state the resulting three pair instances and expectations +MAJOR | high | Task 12b steps 2 and 5-6 | Steps 2 and 5 direct the sweep record to ignored .context/loop-rule-sweep.md, while step 6 commits only the plan and the following paragraph says the record must instead be written under the plan's Completeness sweep output section | A literal execution commits an empty plan section, leaves the real sweep evidence ignored, and gives Task 15 no durable record from which to assemble its closing evidence | Make steps 2 and 5 write or replace the plan output section; describe any .context file only as optional scratch that is copied into that section before the commit +MAJOR | medium | Task 13 step 1 | Several minimum rows do not establish all predicates needed to select a next state: a clean eligible pass with an unmet precondition can take the source-block, suspension, or continue route, and a clean below-floor pass can suspend or continue depending on the suspension predicates | The table can assign an arbitrary transition to an underspecified row or require the executor to invent missing assumptions, so its pass/fail oracle does not test the installed ordering for those named cases | Require every written row to state the source-block and suspension predicates it depends on, splitting each scenario family until the installed ordering yields one determinate route +MAJOR | high | Task 15 step 8 | The WIP findings-file commit and the real closing commit are in one Bash block, but codex-gate.sh classifies the entire Bash command with is_wip_commit and will match the earlier -m WIP text | When the block is executed as one tool call, the hook treats the real closing attempt as cycle-internal and preserves the prior Gate-B count and fingerprint into the next cycle | Put the WIP record commit and the reset-plus-closing commit in separate Bash invocations so the hook sees the closing commit in a command string containing no WIP commit +MAJOR | high | Task 15 step 8 recovery | After a failed closing commit the recovery restores the clean WIP tip where the final findings files are already committed, but a re-run immediately requires those same files to be dirty and unconditionally tries to commit them again; the prose later claiming a cached-diff no-op branch has no counterpart in the block | The documented fix-the-cause-then-re-run path stops at its own exact-dirty-set guard, so a recoverable identity, signing, or hook failure leaves no executable retry path | Admit a verified clean WIP tip whose last record commit contains exactly the two final-slot paths and skip the record commit in that state, while retaining the exact-path guard for the first run +BLOCKER | high | Task 15 step 8 closing-failure branch | A failed closing act can modify or stage tracked content in a commit hook, which target §A1 explicitly requires the executor to inspect and re-establish conditions against, but the branch immediately runs git reset --hard to the saved tip | The recovery can silently destroy the closing attempt's tracked side effects and then retry on a repository state different from the one the target says must be evaluated | Restore HEAD and the index without discarding the working tree, such as with a mixed reset, preserve and inspect the attempt delta, then re-run every closure condition affected before retrying +END OF FINDINGS (6 total) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index 51217d4..93e21f5 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -936,16 +936,26 @@ is add-only** — it replaces no wording — so it is checked by presence alone, it can be omitted from both copies while the `e7` pair, the condition walk and the parity diff all pass. -**`e7` and `e8` are two dispositions and owe two different NEW fragments — an earlier draft gave -both pairs the same one and observed only the second.** +**`e7` and `e8` are two dispositions but they cannot have two OLD halves, and saying why is the +point.** §D supplies **one** sentence for both copies, so in W there is no state where the +read-after clause arrives and the pronoun does not: the two changes install or fail together, and +any fragment that vanishes for one vanishes for the other. **A second OLD row would be a +counterfactual that cannot fail independently** — the thing this plan rejects everywhere else. -- **`e7`'s NEW must carry the read-after-clean-completion meaning** — the clause §D adds. That is - what `e7` changed. A NEW taken from `you report the tells` observes nothing about it: those words - are already in C, and in C they become "new" only because the line reflows. -- **`e8`'s alignment is W's alone and owes its own observation**: `report the tells` becoming - `you report the tells`. P5w's OLD going to zero shows W's pronoun-less form is gone; a NEW - containing the pronoun shows W received C's form rather than some third wording. Without it the - alignment rests on Task 14's parity judgement instead of on a count. +**So the discrimination lives in the NEW halves, and there are three of them:** + +| Observation | Copy | Shape | Fragment | +|---|---|---|---| +| `e7` pair | C | pair, OLD = **P5** | NEW from the **read-after-clean-completion clause** | +| `e7` pair | W | pair, OLD = **P5w** | NEW from the same clause | +| `e8` alignment | W only | **presence**, `worktree=1 parent=0` | the installed `you report the tells` | + +**Two pair instances and one presence check, not two pairs and a spare.** An earlier draft gave +both pairs a NEW taken from `you report the tells`, which observes the pronoun and says nothing +about `e7`: those words are already in C, and there they become "new" only because the line +reflows. A later draft asked `e8` for its own OLD, which P5w already is. **`e8`'s presence check is +what shows W received C's form rather than some third wording**, and without it the alignment rests +on Task 14's parity judgement instead of on a count. *(Build the pair per the verification procedure; record the four values.)* @@ -1643,8 +1653,9 @@ three files still answer this?" - [ ] **Step 2: Read every section of both prompt copies that gives an instruction about a pass, a finding, a gate or a commit** Not a grep. §5 entire, §4's work-loop line, the Gate-A and Gate-B sections, the profiles section, and -Mechanics. **Record each section as read in `.context/loop-rule-sweep.md`**, with a line saying what -you were looking for and what you found. +Mechanics. **Record each section as read under `## Completeness sweep (Task 12b output)` in this +plan**, with a line saying what you were looking for and what you found. A `.context/` scratch file +is fine while you work, but it is ignored and cannot be committed — the plan section is the record. - [ ] **Step 3: Read both channels of all eight hook gate reminders** @@ -1660,8 +1671,9 @@ find is real, the spec's Gate-A cycle reopens for it.** - [ ] **Step 5: Record the result either way** -`.context/loop-rule-sweep.md` states what was read and what was found, **including "nothing"**. A -sweep whose negative result is unrecorded cannot be told from a sweep that never ran. +`## Completeness sweep (Task 12b output)` states what was read and what was found, **including +"nothing"**. A sweep whose negative result is unrecorded cannot be told from a sweep that never +ran — and one recorded only in `.context/` is unrecorded as far as the commit is concerned. - [ ] **Step 6: Commit** @@ -1695,7 +1707,17 @@ next `## `. - [ ] **Step 1: Enumerate the rows** -One row per (starting state, answer) pair the installed ordering admits. At minimum, and **this is a floor rather than the set**: a clean eligible pass with every condition met; a clean eligible pass with an unmet precondition; a clean pass below the floor; a zero-finding pass below the floor; a pass carrying a membership trigger, answered accept and answered decline; a pass carrying a new-question trigger, answered; a pass carrying both; a two-tell stop, answered; a clearly-stuck surface, answered; a source block raised before any pass was read; a source block raised on a pass already read; a closing act that does not complete and is repaired; a closing act that cannot be repaired; a `full` Gate-B pass with one branch clean and one not; the same complaint in both branch files under each of accept/accept, accept/decline, decline/accept and decline/decline. +**Every row states the value of every predicate its route depends on, and is split until the +installed ordering yields exactly one next state.** A scenario naming only some of them is not a +row: "a clean eligible pass with an unmet precondition" can take the source-block route, a +suspension, or the continue branch depending on predicates it never states, and "a clean pass below +the floor" can suspend or continue on the same grounds. **A row with more than one admissible route +tests nothing** — the executor either picks one arbitrarily or invents the missing assumption, and +the oracle then passes on an answer the ordering never gave. Split the family until each member is +determinate, and say in the row which predicate values made it so. + +The list below is **scenario families, not rows** — one row per (starting state, answer) pair the +installed ordering admits, after splitting. At minimum, and **this is a floor rather than the set**: a clean eligible pass with every condition met; a clean eligible pass with an unmet precondition; a clean pass below the floor; a zero-finding pass below the floor; a pass carrying a membership trigger, answered accept and answered decline; a pass carrying a new-question trigger, answered; a pass carrying both; a two-tell stop, answered; a clearly-stuck surface, answered; a source block raised before any pass was read; a source block raised on a pass already read; a closing act that does not complete and is repaired; a closing act that cannot be repaired; a `full` Gate-B pass with one branch clean and one not; the same complaint in both branch files under each of accept/accept, accept/decline, decline/accept and decline/decline. - [ ] **Step 2: Walk each row against the installed §A and record the next state** @@ -2082,55 +2104,71 @@ provenance line or curve has not validly closed the cycle. so an evidence entry written only into a WIP message is destroyed at exactly the moment the cycle closes — which is what step 5 would otherwise have done. +**8a — record the final pass's findings files. Its own shell invocation, and that is not +cosmetic.** `codex-gate.sh`'s `is_wip_commit` matches `-m` followed by `wip` **anywhere in the +command string it is given**. A single block carrying this `-m "WIP: …"` and the closing +`git commit -F` would be classified cycle-internal in its entirety, and the hook would carry this +cycle's Gate-B count and fingerprint into the next one — a real closing commit read as a snapshot. +**Run 8a and 8b as separate Bash calls, and keep the word out of 8b's command string.** + ```bash BASE=$(cat .context/loop-rule-base) test -n "$BASE" || { echo "BASE empty — Task 0 did not run"; exit 1; } -# Everything that must be IN the squashed commit has to be committed before the -# reset: reset --soft stages only what the discarded commits already contained. -# The ONLY thing allowed to differ from the reviewed HEAD is the final pass's own -# findings files, which could not exist when that pass was issued. Name them -# EXACTLY — from the slot paths the final call was told to write — and require the -# dirty set to be those paths, added, and nothing else. # NONCE and P are this cycle's nonce and its final pass number — the same two the # final call's slot paths were built from. Set them to the values you used. NONCE=; P= FINAL=".context/codex-reviews/gate-b-spec-$NONCE-pass-$P.md .context/codex-reviews/gate-b-quality-$NONCE-pass-$P.md" expected=$(printf '%s\n' $FINAL | sort) actual=$(git status --porcelain | sed -n 's/^?? //p; s/^A //p' | sort) -test "$expected" = "$actual" || { - echo "dirty set is not exactly the final pass's findings files:"; git status --porcelain; exit 1; } - -# shellcheck disable=SC2086 -git add $FINAL -git commit -m "WIP: Gate-B findings files" || { echo "record commit FAILED"; exit 1; } -test -z "$(git status --porcelain)" || { echo "tree not clean before reset — aborting"; exit 1; } -# Keep the pre-reset tip: after the reset it is the ONLY way back to the reviewed -# history if the closing commit fails. +if [ -z "$actual" ] && git show --stat --name-only --pretty=format: HEAD | grep -q 'gate-b-'; then + echo "final findings files already recorded by a previous attempt — continue at 8b" +elif [ "$expected" = "$actual" ]; then + # shellcheck disable=SC2086 + git add $FINAL + git commit -m "WIP: Gate-B findings files" || { echo "record commit FAILED"; exit 1; } +else + echo "dirty set is not exactly the final pass's findings files:"; git status --porcelain; exit 1 +fi +test -z "$(git status --porcelain)" || { echo "tree not clean — aborting before 8b"; exit 1; } git rev-parse HEAD > .context/loop-rule-wip-tip +``` + +**The first branch is the retry path.** A closing commit that fails leaves the findings files +already committed, so on a second run the dirty set is empty — and an unconditional record commit, +or an exact-dirty-set guard with no such branch, refuses the very state the recovery produces. +The branch admits it only when `HEAD` is the record commit it would have made. + +**8b — reset and close. A separate invocation, with no `-m` and no `wip` in it.** + +```bash +BASE=$(cat .context/loop-rule-base); TIP=$(cat .context/loop-rule-wip-tip) git reset --soft "$BASE" git commit -F .context/loop-rule-closing-msg || { - echo "closing commit FAILED — restoring the reviewed tip" - git reset --hard "$(cat .context/loop-rule-wip-tip)" - echo "history restored; fix the cause, then re-run step 8. Base and tip files kept." + echo "closing act FAILED — restoring the reviewed tip without discarding its side effects" + git reset --mixed "$TIP" + git status --porcelain + echo "HEAD and index restored to the reviewed tip. Any files listed above were left by the" + echo "failed attempt — inspect them, then re-establish every closure condition before retrying 8a." exit 1; } case "$(git log -1 --pretty=%s)" in - WIP:*|wip:*) echo "closing commit still reads WIP — cycle NOT closed"; exit 1 ;; + [Ww][Ii][Pp]:*) echo "closing commit still reads as a snapshot — cycle NOT closed"; exit 1 ;; esac test -z "$(git status --porcelain)" || { echo "worktree dirty after close — aborting before cleanup"; exit 1; } rm -f .context/loop-rule-base .context/loop-rule-wip-tip ``` -**The failure branch is the point of `.context/loop-rule-wip-tip`.** A closing commit can fail on -an unset identity, a signing key or a hook, and by then the reset has already run: `HEAD` is at -`$BASE` with the whole implementation staged. **Re-running step 8 from that state hits its own -dirty-tree guard and refuses**, correctly — the guard cannot tell a staged implementation from a -record that rode along. `reset --hard` back to the recorded tip returns the repository to the -reviewed history, with every WIP commit and the reviewed `HEAD` intact, so the executor fixes the -cause and re-runs rather than inventing a recovery sequence at the riskiest moment. +**`--mixed`, never `--hard`.** A closing act can fail *after* a commit hook has modified or staged +tracked content, and `reset --hard` would delete exactly that — the delta the failure produced. +Target §A1 requires the executor to inspect what an incomplete closing act left and to re-establish +the closure conditions against it; a recovery that discards the evidence first makes that +impossible and retries against a repository state the target says must be evaluated. +`--mixed` puts `HEAD` and the index back at the reviewed tip and leaves the working tree alone, so +`git status` above *is* the attempt's delta. -**Both files are removed only after a successful close**, so a later run cannot inherit either. +**Both scratch files are removed only after a successful close**, so a later run cannot inherit +either, and the retry path above still has them. **Every check in that block fails the script; none of them is a comment.** An earlier draft suppressed the record commit with `|| true` and left the status and log lines as things to look at, From c1614ba66f29a47640c610cd86099c0e0d54c909 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 18:09:09 +0200 Subject: [PATCH 116/181] docs(plans): apply Gate-A plan pass 13; name the exact predicates the block tests MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit MANDATORY TWO-TELL STOP — Blockers flat at 1 (failing to fall) and the findings cluster on the verification apparatus. Findings fell 6 -> 4. Surfaced to Daniel; standing answer applied, loop continued. One Blocker, one Major, two Minors. - 8a's retry branch admitted any HEAD touching a path matching gate-b-. Step 7 commits earlier passes' findings files, so HEAD can carry one while the final pass's two are missing — a mistyped NONCE, a wrong P, a pass whose files were never written — and 8b would then close without the artifacts the closing commit exists to carry. It compares HEAD's exact changed-path set to $FINAL. - The span derivation collected the line each FRAGMENT sits on, not the line extent of each replacement BLOCK. A fragment is one line; its block spans several, and every line a replacement occupies is changed. a13's tail sits on the floor region's end-anchor line, so an anchor line is not exempt either. - The "no -m and no wip in 8b" rule outran what the hook tests. is_wip_commit matches -m followed by optional space/quote then wip; --mixed does not match, and a path containing wip does not match without a preceding -m. Verified against the regex. The rule now states that comparison; the scratch file is renamed to loop-rule-reviewed-tip anyway, so a reader need not check. - The step-8 explanation claimed a `git diff --cached --quiet` test the block stopped using in the pass-12 restructure. It names the two predicates 8a actually branches on. Trend: findings 32, 32, 31, 20, 16, 11, 13, 15, 9, 17, 11, 6, 4. Blockers 1, 0, 0, 10, 7, 2, 4, 2, 2, 1, 2, 1, 1. Majors 28, 25, 23, 7, 7, 6, 4, 9, 6, 14, 7, 5, 1. Floor 3, from risk high (2) and security none (0), read at this pass from docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-plan-om0bdd7udh-pass-13.md | 5 ++ .../2026-09-14-loop-rule-consolidation.md | 70 +++++++++++++------ 2 files changed, 52 insertions(+), 23 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-13.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-13.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-13.md new file mode 100644 index 0000000..40ba65c --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-13.md @@ -0,0 +1,5 @@ +MAJOR | high | Task 0 step 2 | The supposedly exhaustive floor-region changed-line example names only §F item 8a and P9, but P10's a16 replacement begins on the same physical line as the `Nothing here writes the floor knob` end anchor, so that line is also occupied by an installed replacement. | An executor following the stated `that is` list can record an untouched span across a line Task 7 later changes, making Task 2 or Task 14 reject a correct implementation. | Add P10's line to the floor-region changed-line derivation and make clear that an end-anchor line containing changed text cannot belong to an untouched span. +BLOCKER | high | Task 15 step 8a | The retry branch tests only that `actual` is empty and HEAD's changed-path listing contains any `gate-b-` name; it does not establish the following claim that HEAD is this final pass's two-file record commit, while step 7 can already have committed earlier Gate-B files. | If the final files are missing or NONCE/P is wrong on a retry, the branch can skip the record commit and let 8b close without the final pass artifacts. | Verify that HEAD is the expected record commit and that its exact changed-path set is the two expected final-pass files, failing otherwise. +MINOR | high | Task 15 step 8b | The plan twice requires the word `wip` to be absent from 8b's command string, but the supplied block reads and deletes `.context/loop-rule-wip-tip`. | The copy-pasteable command contradicts Pass 12's classifier-safety rule, so the executor cannot both follow the block and satisfy the review contract. | Rename the scratch file, for example to `.context/loop-rule-reviewed-tip`, consistently in 8a, 8b, and cleanup. +MINOR | high | Task 15 step 8 explanation | The paragraph says the block asks `git diff --cached --quiet` before distinguishing nothing-to-commit from failure, but neither 8a nor 8b executes that check; 8a instead uses a status-derived dirty set and a broad HEAD-path grep. | The recovery proof relies on a predicate the executable procedure does not have, obscuring the unsafe retry admission and violating the repo rule that prose name the exact comparison performed. | Either add and guard the claimed staged-diff check in the appropriate branch or rewrite the paragraph to state the actual status and HEAD-path predicates. +END OF FINDINGS (4 total) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index 93e21f5..54e6804 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -447,12 +447,17 @@ and each makes a **correct** implementation fail its own check. **Derive them, w open:** 1. Resolve each of the five regions' start and end anchors. -2. **Collect every line in that region that a changed fragment sits on** — every fragment-table row - whose edit lands in the region, and every line an installed replacement will occupy. For the - floor arithmetic that is item 8a's block and row P9's line; for the human exception, rows F7, - F7b and F4. -3. **The spans are the gaps between those lines.** A region with none of them is one span; a region - with *n* is at most *n+1*. A span of zero lines is dropped, not recorded. +2. **Collect the full line extent of every replacement that lands in that region — not the line its + fragment sits on.** A fragment is one line; the block it belongs to usually spans several, and + the lines a replacement occupies are all changed whether or not a fragment happens to sit on + them. Resolve each replacement's **first and last** live line and take the whole run. In the + floor region that is item 8a's block, `a13`'s paragraph and `a16`'s; in the human exception, + items 7 and 4. +3. **The spans are the gaps between those runs.** A region with no replacement in it is one span; + a region with *n* runs is at most *n+1*. A span of zero lines is dropped, not recorded. + **An anchor line is not exempt**: where a region's start or end anchor shares its line with + changed text — and `a13`'s tail sits on the same line as the floor region's end anchor — the + span begins or ends past it, and the kept condition on that line goes to the per-condition list. 4. **Then assign every kept condition in that region to exactly one span, or to the per-condition list.** A kept condition in neither is the failure this step exists to prevent; one in both is an accounting error. @@ -2105,11 +2110,17 @@ so an evidence entry written only into a WIP message is destroyed at exactly the closes — which is what step 5 would otherwise have done. **8a — record the final pass's findings files. Its own shell invocation, and that is not -cosmetic.** `codex-gate.sh`'s `is_wip_commit` matches `-m` followed by `wip` **anywhere in the -command string it is given**. A single block carrying this `-m "WIP: …"` and the closing -`git commit -F` would be classified cycle-internal in its entirety, and the hook would carry this -cycle's Gate-B count and fingerprint into the next one — a real closing commit read as a snapshot. -**Run 8a and 8b as separate Bash calls, and keep the word out of 8b's command string.** +cosmetic.** `codex-gate.sh`'s `is_wip_commit` (`hooks/codex-gate.sh:763`) tests the **whole command +string** it is given against `-m[[:space:]]*['"]?[[:space:]]*wip`, case-insensitively: **`-m` +immediately followed by optional whitespace, an optional quote, and `wip`.** A single block carrying +this `-m "WIP: …"` and the closing `git commit -F` matches, and is classified cycle-internal in its +entirety — the hook would then carry this cycle's Gate-B count and fingerprint into the next one, a +real closing commit read as a snapshot. **So run 8a and 8b as separate Bash calls.** + +**What 8b must avoid is that pattern, not the letters.** `--mixed` contains `-m` and matches +nothing, because `i` follows; a path containing `wip` matches nothing, because no `-m` precedes it. +Saying "no `-m` and no `wip` in 8b" would outrun the check the hook performs and make a correct +block look non-compliant. 8b carries no `-m` option at all, which is the property that matters. ```bash BASE=$(cat .context/loop-rule-base) @@ -2121,7 +2132,8 @@ FINAL=".context/codex-reviews/gate-b-spec-$NONCE-pass-$P.md .context/codex-revie expected=$(printf '%s\n' $FINAL | sort) actual=$(git status --porcelain | sed -n 's/^?? //p; s/^A //p' | sort) -if [ -z "$actual" ] && git show --stat --name-only --pretty=format: HEAD | grep -q 'gate-b-'; then +head_paths=$(git show --name-only --pretty=format: HEAD | sed '/^$/d' | sort) +if [ -z "$actual" ] && [ "$head_paths" = "$expected" ]; then echo "final findings files already recorded by a previous attempt — continue at 8b" elif [ "$expected" = "$actual" ]; then # shellcheck disable=SC2086 @@ -2131,18 +2143,25 @@ else echo "dirty set is not exactly the final pass's findings files:"; git status --porcelain; exit 1 fi test -z "$(git status --porcelain)" || { echo "tree not clean — aborting before 8b"; exit 1; } -git rev-parse HEAD > .context/loop-rule-wip-tip +git rev-parse HEAD > .context/loop-rule-reviewed-tip ``` -**The first branch is the retry path.** A closing commit that fails leaves the findings files -already committed, so on a second run the dirty set is empty — and an unconditional record commit, -or an exact-dirty-set guard with no such branch, refuses the very state the recovery produces. -The branch admits it only when `HEAD` is the record commit it would have made. +**The first branch is the retry path, and it compares `HEAD`'s exact changed-path set.** A closing +commit that fails leaves the findings files already committed, so on a second run the dirty set is +empty — and an unconditional record commit, or an exact-dirty-set guard with no such branch, +refuses the very state the recovery produces. -**8b — reset and close. A separate invocation, with no `-m` and no `wip` in it.** +**"`HEAD` touched some `gate-b-` file" is not that test, and would be unsafe.** Step 7 commits +earlier passes' findings files, so `HEAD` can carry one while the *final* pass's two are missing +entirely — from a mistyped `NONCE`, a wrong `P`, or a pass whose files were never written. The +branch would then skip the record commit and let 8b close without the artifacts the closing commit +exists to carry. **Requiring `HEAD`'s changed paths to equal `$FINAL` exactly** admits only the +commit 8a itself would have made. + +**8b — reset and close. A separate invocation, carrying no `-m` option at all.** ```bash -BASE=$(cat .context/loop-rule-base); TIP=$(cat .context/loop-rule-wip-tip) +BASE=$(cat .context/loop-rule-base); TIP=$(cat .context/loop-rule-reviewed-tip) git reset --soft "$BASE" git commit -F .context/loop-rule-closing-msg || { echo "closing act FAILED — restoring the reviewed tip without discarding its side effects" @@ -2156,7 +2175,7 @@ case "$(git log -1 --pretty=%s)" in [Ww][Ii][Pp]:*) echo "closing commit still reads as a snapshot — cycle NOT closed"; exit 1 ;; esac test -z "$(git status --porcelain)" || { echo "worktree dirty after close — aborting before cleanup"; exit 1; } -rm -f .context/loop-rule-base .context/loop-rule-wip-tip +rm -f .context/loop-rule-base .context/loop-rule-reviewed-tip ``` **`--mixed`, never `--hard`.** A closing act can fail *after* a commit hook has modified or staged @@ -2176,9 +2195,14 @@ so a failed record commit, a dirty tree or a closing message still reading `WIP: through the soft reset and deleted the recovery base — producing, silently, the exact state the plan elsewhere calls an invalid close. -**`|| true` is replaced by a test for the one harmless case it was hiding.** "Nothing to commit" -and "the commit failed" are different outcomes and only the first is fine, so the block asks -`git diff --cached --quiet` first and treats a real failure as a failure. +**`|| true` is replaced by branching on what the state actually is, before committing anything.** +"Nothing to commit" and "the commit failed" are different outcomes and only the first is fine. 8a +decides between them with two predicates it can name: the **status-derived dirty set** (`?? ` and +`A ` entries, sorted) and **`HEAD`'s changed-path set**. An empty dirty set with `HEAD` equal to +`$FINAL` is the harmless retry; the dirty set equal to `$FINAL` is the first run; anything else +stops. **A `git diff --cached --quiet` test is not among them** — an earlier draft's prose claimed +it after the block had stopped using it, which is the overclaim `AGENTS.md` names by requiring +prose to state the exact comparison performed. **The base file is removed only in the success branch.** `git commit` can fail on a hook, a signing key or an unset identity, and at that point the reset has already happened: the WIP commits are From 6a9cfc817f053f132417774a853a27dc88eda14d Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 18:18:21 +0200 Subject: [PATCH 117/181] docs(plans): apply Gate-A plan pass 14; close both recovery holes in the closing act MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit MANDATORY THREE-TELL STOP — findings rose 4 -> 5, Blockers rose 1 -> 2, and the findings cluster on the verification apparatus. Surfaced to Daniel; standing answer applied, loop continued. Two Blockers, two Majors, one Minor. - 8a built its "exact dirty set" from only '??' and 'A ' porcelain records, so a staged tracked modification beside the two findings files was invisible: the set compared equal, git commit committed the whole index and swept it in, the clean-tree test then passed, and 8b published content the reviewed HEAD never carried. Every porcelain record counts now, and the record commit's own changed-path set is checked after it lands. - 8b's post-commit subject and clean-tree checks were bare exits. A closing act that COMMITTED and then failed one of them left HEAD at the rejected commit with no way back — and 8a could not be re-entered, since that commit's changed-path set is the whole squash. Both now mixed-reset to the reviewed tip and report the delta. - The Gate-B loop admitted only "clean at or above the floor" and had no branch for a zero-finding pass below it, which installed §A defines as eligible and Task 13's table checks by name. The procedure would have overridden the rule it installs. - Tasks 3 and 4 consumed existing rows (P2, P3, P4) and then ran the global walk over the same conditions, which appends a second row for each. One finding applied without its stated cause: pass 14 finding 3 says the floor region's replacement list omits a17-a22. It does not — that block is at C 133 and the region ends at C 131. But the list it points at has now been wrong three times, in three different ways, so it is deleted rather than corrected a fourth time: which replacements intersect a region is resolved against the file. Trend: findings 32, 32, 31, 20, 16, 11, 13, 15, 9, 17, 11, 6, 4, 5. Blockers 1, 0, 0, 10, 7, 2, 4, 2, 2, 1, 2, 1, 1, 2. Majors 28, 25, 23, 7, 7, 6, 4, 9, 6, 14, 7, 5, 1, 2. Floor 3, from risk high (2) and security none (0), read at this pass from docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-plan-om0bdd7udh-pass-14.md | 6 ++ .../2026-09-14-loop-rule-consolidation.md | 66 ++++++++++++++++--- 2 files changed, 62 insertions(+), 10 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-14.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-14.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-14.md new file mode 100644 index 0000000..14ccf33 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-14.md @@ -0,0 +1,6 @@ +BLOCKER | high | Task 15 step 8a | The claimed exact dirty-set guard builds `actual` from only `??` and `A ` porcelain records, so it ignores tracked staged modifications, deletions, renames, copies, and conflicts. | If the two final findings files coexist with an unrelated staged tracked change, `expected = actual` passes, the record commit sweeps that change in, the following status check is clean, and 8b publishes content that was not in the reviewed HEAD. | Parse every porcelain record and its path, preferably with the NUL form, require the complete dirty path set to equal `$FINAL`, and verify the record commit's changed-path set equals `$FINAL` before writing the reviewed tip. +BLOCKER | high | Task 15 step 8b | Recovery exists only when `git commit -F` itself fails; if that commit succeeds and either the post-commit subject check or clean-tree check fails, the block exits with HEAD at the rejected closing commit and provides no restoration to `$TIP`. | A retry cannot enter 8a because the rejected closing commit's changed-path set is the whole squash rather than `$FINAL`, leaving the cycle unable to close through the plan and unable to return to its reviewed state, contrary to the incomplete-closing-act transition the target requires. | Put both post-commit checks on a failure path that mixed-resets to `$TIP`, preserves and reports side effects, and requires all closure conditions and the final review to be re-established before retrying 8a. +MAJOR | high | Task 0 step 2 | The floor-region replacement-run list says the replacements are item 8a, `a13`, and `a16`, but omits the separate `a17`–`a22` block that Task 7 replaces at the end of the same live floor paragraph. | Following the stated derivation records the nonzero gap after `a16` as untouched even though Task 7 changes it, so Task 14's mandatory parent diff rejects a correct implementation; treating the shorter list as authoritative also contradicts the full-block rule added in pass 13. | Include the full live line extent of the `a17`–`a22` replacement in the floor-region run set, merging overlapping extents where it shares a line with `a16`, before deriving gaps. +MAJOR | high | Task 15 step 7 | The Gate-B loop requires a clean pass at or above floor 3 and gives no branch for a zero-finding pass below the floor, although installed §A defines eligibility as a clean pass at or above the floor or a zero-finding pass and Task 13 explicitly checks the below-floor zero-finding case. | A first or second logical Gate-B pass whose required branch files both report `NO FINDINGS` is routed into more passes instead of the ordering's closing path, so the implementation procedure overrides the rule it has just installed. | Change the loop condition and final-pass procedure to close on either an eligible at-or-above-floor clean pass or a zero-finding logical pass, subject in both cases to every other closure condition and the closing act. +MINOR | high | Tasks 3 step 1b and 4 step 1 | Both tasks invoke the global pre-install procedure over every condition in their source block after already consuming existing OLD rows for conditions in that same block (`P2` and `P3` in Task 3, `P4` in Task 4); the global procedure says to derive and append every such pre-existing fragment. | A literal execution appends second OLD rows for `b7`, `b12`, and `c14`, then runs duplicate observations, violating the plan's one-authored-copy fragment rule and making the ledger depend on whether the executor silently inferred an exclusion. | State that the block walk reuses existing rows and appends only conditions not already covered, as Tasks 6 and 7 do, and run each existing or newly appended row once to its disposition's result. +END OF FINDINGS (5 total) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index 54e6804..298b1db 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -203,7 +203,11 @@ for the same condition, which is the fragment table's one-copy rule broken from **Each editing task runs this, against the disposition table's rows for its own blocks:** -- [ ] **Before installing** — walk every condition those blocks cover and derive the +- [ ] **Before installing** — walk every condition those blocks cover. **Where the fragment table + already has a row for it, reuse that row and confirm its pre-edit count; append nothing.** The + table is the one authored copy, and a walk that appends a second row for `b7`, `b12` or `c14` + would run two observations of one edit and leave the ledger depending on whether the executor + silently inferred an exclusion. **Otherwise** derive the **pre-existing** fragment its disposition owes: the **OLD** half for a *replaced* one, the **absence** fragment for a *moved* or *dropped* one, the **preservation** fragment for a *carried* one, and, for a *kept* one, a preservation fragment **only where no untouched span can @@ -447,12 +451,16 @@ and each makes a **correct** implementation fail its own check. **Derive them, w open:** 1. Resolve each of the five regions' start and end anchors. -2. **Collect the full line extent of every replacement that lands in that region — not the line its - fragment sits on.** A fragment is one line; the block it belongs to usually spans several, and - the lines a replacement occupies are all changed whether or not a fragment happens to sit on - them. Resolve each replacement's **first and last** live line and take the whole run. In the - floor region that is item 8a's block, `a13`'s paragraph and `a16`'s; in the human exception, - items 7 and 4. +2. **Collect the full line extent of every replacement whose live extent intersects that region — + not the line its fragment sits on.** A fragment is one line; the block it belongs to usually + spans several, and every line a replacement occupies is changed whether or not a fragment sits + on it. Resolve each replacement's **first and last** live line, take the whole run, and merge + overlapping runs. + **No example list is given here, and that is deliberate.** One was written three times and was + wrong all three — naming a block that sits outside the region, omitting one inside it, and + naming a fragment's line for a block's extent. Which replacements intersect a region is decided + by resolving both against the file, and any list written in advance is a fourth guess. Walk the + target's replacement blocks, resolve each one's extent, and keep the ones that overlap. 3. **The spans are the gaps between those runs.** A region with no replacement in it is one span; a region with *n* runs is at most *n+1*. A span of zero lines is dropped, not recorded. **An anchor line is not exempt**: where a region's start or end anchor shares its line with @@ -2030,6 +2038,14 @@ and is the only value whose range contains the whole implementation. Floor derives from the story profile: risk `high` → 2, security `none` → 0, max 2 ≠ 0 → **floor 3**. Re-derive it at each pass from the header. +**Two routes reach the closing act, and this loop must admit both**: a **clean pass at or above the +floor**, or a **zero-finding logical pass** at any pass number — the early exit below the floor. +Both are subject to every other closure condition. **An earlier draft named only the first**, so a +first or second Gate-B pass whose two branch files both read `NO FINDINGS` would have been routed +into further passes — the implementation procedure overriding the very rule §A installs, and the +one case Task 13's table checks by name. A zero-finding pass is `NO FINDINGS` in **every** required +branch file; one branch clean and the other not is not it. + **Each fix is committed before the next review is issued**, or the re-review targets the unchanged WIP tip while the repair sits in the worktree — and the final squash then publishes a fix no pass reviewed: @@ -2130,7 +2146,11 @@ test -n "$BASE" || { echo "BASE empty — Task 0 did not run"; exit 1; } NONCE=; P= FINAL=".context/codex-reviews/gate-b-spec-$NONCE-pass-$P.md .context/codex-reviews/gate-b-quality-$NONCE-pass-$P.md" expected=$(printf '%s\n' $FINAL | sort) -actual=$(git status --porcelain | sed -n 's/^?? //p; s/^A //p' | sort) +# EVERY porcelain record, whatever its status letters — staged modifications, +# deletions, renames and conflicts included. A filter that reads only '??' and +# 'A ' declares the set exact while a staged tracked change sits beside it, and +# `git commit` then commits the whole index. +actual=$(git status --porcelain -z | tr '\0' '\n' | sed -n 's/^.\{3\}//p' | sed '/^$/d' | sort) head_paths=$(git show --name-only --pretty=format: HEAD | sed '/^$/d' | sort) if [ -z "$actual" ] && [ "$head_paths" = "$expected" ]; then @@ -2139,6 +2159,8 @@ elif [ "$expected" = "$actual" ]; then # shellcheck disable=SC2086 git add $FINAL git commit -m "WIP: Gate-B findings files" || { echo "record commit FAILED"; exit 1; } + test "$(git show --name-only --pretty=format: HEAD | sed '/^$/d' | sort)" = "$expected" \ + || { echo "record commit changed paths beyond the final findings files"; exit 1; } else echo "dirty set is not exactly the final pass's findings files:"; git status --porcelain; exit 1 fi @@ -2151,6 +2173,13 @@ commit that fails leaves the findings files already committed, so on a second ru empty — and an unconditional record commit, or an exact-dirty-set guard with no such branch, refuses the very state the recovery produces. +**The dirty set is built from every porcelain record, not from two status codes.** An earlier draft +read only `?? ` and `A ` entries, so a **staged tracked modification** sitting beside the two +findings files was invisible: the set compared equal, `git commit` committed the whole index and +swept the change in, the following clean-tree test passed, and 8b published content the reviewed +`HEAD` never carried. **And the record commit's own changed-path set is checked afterwards**, +because the guard describes the working tree while the commit is what actually lands. + **"`HEAD` touched some `gate-b-` file" is not that test, and would be unsafe.** Step 7 commits earlier passes' findings files, so `HEAD` can carry one while the *final* pass's two are missing entirely — from a mistyped `NONCE`, a wrong `P`, or a pass whose files were never written. The @@ -2171,13 +2200,30 @@ git commit -F .context/loop-rule-closing-msg || { echo "failed attempt — inspect them, then re-establish every closure condition before retrying 8a." exit 1; } +bad="" case "$(git log -1 --pretty=%s)" in - [Ww][Ii][Pp]:*) echo "closing commit still reads as a snapshot — cycle NOT closed"; exit 1 ;; + [Ww][Ii][Pp]:*) bad="closing commit reads as a snapshot" ;; esac -test -z "$(git status --porcelain)" || { echo "worktree dirty after close — aborting before cleanup"; exit 1; } +test -z "$(git status --porcelain)" || bad="${bad:+$bad; }worktree dirty after the closing act" +if [ -n "$bad" ]; then + echo "closing act REJECTED: $bad — restoring the reviewed tip without discarding side effects" + git reset --mixed "$TIP" + git status --porcelain + echo "HEAD and index restored. Re-establish every closure condition, obtain a fresh clean" + echo "response against the new HEAD, then retry 8a." + exit 1 +fi + rm -f .context/loop-rule-base .context/loop-rule-reviewed-tip ``` +**Both post-commit checks restore the tip as well, and an earlier draft left them as bare exits.** +A closing act that *commits successfully* and then fails its subject or clean-tree check left +`HEAD` at the rejected commit with no way back — and 8a could not be re-entered from there either, +since that commit's changed-path set is the whole squash rather than `$FINAL`. The cycle was then +unable to close through the plan and unable to return to its reviewed state, which is exactly the +incomplete-closing-act transition target §A requires to be recoverable. + **`--mixed`, never `--hard`.** A closing act can fail *after* a commit hook has modified or staged tracked content, and `reset --hard` would delete exactly that — the delta the failure produced. Target §A1 requires the executor to inspect what an incomplete closing act left and to re-establish From c6d0773b8c51111a1def2b62d301339a06343244 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 18:29:16 +0200 Subject: [PATCH 118/181] docs(plans): apply Gate-A plan pass 15; "changed" is not a disposition MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit One tell only (instrument cluster) — no mandatory stop. Findings fell 5 -> 4 and Blockers fell 2 -> 0. Two Majors, two Minors. - Passage (b) classified nine conditions as "changed", and passages (c) and (e) used it for c8 and e7. "Changed" is not one of the six dispositions, so the procedure gives it no observation and a literal execution skips it — nine meaning changes in §B alone with no counterfactual. They are "replaced", and the disposition table now says the six words are the only ones. - 8a's post-commit changed-path check was a bare exit, leaving a rejected WIP commit at HEAD carrying an unreviewed path — which the retry branch then refuses, since its changed-path set is not $FINAL. It restores the pre-record tip, like every other rejection in step 8. - The battery recorded `git rev-parse origin/main` and then passed the mutable name to check-version-bump.sh, so the evidence could name one commit while the checker resolved another. Resolved once into $BASEREF, which is both the record and the argument. - The step-8 explanation still described the dirty set as "?? and A entries" after pass 14 made it consume every porcelain record — the same stale-prose shape pass 13 found, in the paragraph pass 13 repaired. Trend: findings 32, 32, 31, 20, 16, 11, 13, 15, 9, 17, 11, 6, 4, 5, 4. Blockers 1, 0, 0, 10, 7, 2, 4, 2, 2, 1, 2, 1, 1, 2, 0. Majors 28, 25, 23, 7, 7, 6, 4, 9, 6, 14, 7, 5, 1, 2, 2. Floor 3, from risk high (2) and security none (0), read at this pass from docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-plan-om0bdd7udh-pass-15.md | 5 ++ .../2026-09-14-loop-rule-consolidation.md | 49 ++++++++++++++----- 2 files changed, 42 insertions(+), 12 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-15.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-15.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-15.md new file mode 100644 index 0000000..b8ac9cb --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-15.md @@ -0,0 +1,5 @@ +MAJOR | high | Condition disposition passage (b) / Task 3 steps 1b and 3 | b3, b7, b8, b11, b12, b13, b16, b17 and b18 are classified as changed, but the only defined observation classes and derivation procedure use replaced, and step 3 runs additional pairs only for rows classified replaced | A literal execution can pair only the existing P2 and P3 rows and omit counterfactual observations for b3, b8, b11, b13, b16, b17 and b18, defeating design section 7 and leaving meaning changes out of the closing evidence | Classify every meaning-changing b condition as replaced, or define changed as an exact alias and use that term consistently in the procedure and Task 3 step 3 +MINOR | high | Task 15 step 4 | The plan records the object name from git rev-parse origin/main but invokes check-version-bump.sh with the mutable name origin/main while claiming the recorded revision is the exact argument the successful check receives | Another local update of the remote-tracking ref can make the evidence name one base while the checker resolves another, and the exact-argument claim is already false structurally | Store the resolved object name and pass that object name to check-version-bump.sh as well as recording it +MAJOR | high | Task 15 step 8a | If the post-commit changed-path check detects that the record commit gained another path, it exits without saving or restoring the pre-record tip | The rejected WIP commit remains at HEAD, the retry branch rejects its changed-path set, and the prescribed process cannot resume while an unreviewed side effect remains committed | Save the pre-record tip before the commit and, on a failed post-commit check, mixed-reset to it, expose the surviving side effects, and require closure conditions plus a fresh final pass to be re-established before retry +MINOR | high | Task 15 step 8 explanation | The explanation says the status-derived dirty set consists of ?? and A records even though the repaired block and its preceding explanation deliberately consume every porcelain status | The prose contradicts the implemented predicate and can cause a reviewer or executor to reinstate the exact staged-modification omission pass 14 repaired, violating AGENTS.md's exact-mechanism rule | Describe the predicate as every porcelain record with its status prefix removed, or delete the status-code parenthetical +END OF FINDINGS (4 total) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index 298b1db..9f9cd26 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -183,6 +183,12 @@ different passage, because each task was inventing the rule for its own conditio | **dropped** | an **absence check**, `parent=1 worktree=0`, **one per dropped condition**. Two dropped conditions sharing one fragment is one observation, and it goes absent when either half goes, leaving the other free to survive. | | **add-only** | **presence alone**, `worktree=1 parent=0`. There is no old wording whose absence could be counted. | +**These six words are the only dispositions, and the tables below use no others.** A condition +described as "changed" would name no class, so the procedure would give it no observation and a +literal execution would skip it — §B alone accounts for nine conditions that were labelled that +way. A condition whose sentence is given entire and differs from the parent is **replaced**, +however small the difference and whether the edit adds a clause or rewrites the sentence. + **A reader walk is never any of these.** Tasks walk their conditions as a reader's confirmation on top of the counts; a walk that found what the counts missed means a fragment was wrong, not that the walk was the check. @@ -245,7 +251,7 @@ preservation fragments that an earlier draft chose afterwards. ### Passage (b) — what a loop absorbs → target §B (Task 3) -§B states its own accounting and this table reproduces it: **Changed:** `b3`, `b7`, `b8`, `b11`, `b12`, `b13`, `b16`, `b17`, `b18`. **Added:** the closing-time change rule, and that alone — no inventoried condition carried it because none existed. **Carried:** `b1`, `b2`, `b4`, `b5`, `b6`, `b9`, `b10`, `b14`, `b15`. +§B states its own accounting and this table reproduces it: **Replaced:** `b3`, `b7`, `b8`, `b11`, `b12`, `b13`, `b16`, `b17`, `b18`. **Added:** the closing-time change rule, and that alone — no inventoried condition carried it because none existed. **Carried:** `b1`, `b2`, `b4`, `b5`, `b6`, `b9`, `b10`, `b14`, `b15`. **Decision 6's decline semantics are not a §B addition.** `**Decline is available only at a membership stop**` is a §A sentence, and Task 1's presence checks cover it. An earlier draft listed it here as a second added rule, which would have sent Task 3 looking for a §B fragment that does not exist — and the only way to satisfy that check is to certify an unrelated decline clause. @@ -258,7 +264,7 @@ preservation fragments that an earlier draft chose afterwards. | c1, c2, c3 | **kept** — the curve-reading sentences are untouched. | | c4 | **replaced** — "a missing one means keep going" becomes "means only that *this* exit does not apply", the pass's actual next step being the ordering's. | | c5, c6, c7 | **carried word for word** inside §C's block. | -| c8 | **changed** — gains the re-raised-dismissal clause. | +| c8 | **replaced** — gains the re-raised-dismissal clause. | | c9 | **split.** The operative precedence clause moves into §A capitalized as a standalone sentence; the plateau rationale stays at this source. Not moved whole — §C says so explicitly. | | c10, c11 | **moved** to §A, which states what a Blocker/Major-free pass at or above the floor does. | | c12, c13 | **moved** to §A, beside a19. | @@ -280,7 +286,7 @@ preservation fragments that an earlier draft chose afterwards. |---|---| | e1–e6 | **kept** — the five tells themselves are untouched. | | e8 | **replaced.** The inventory records it as a parity divergence — C `you report`, W `report` — and §D's block supplies C's wording for **both** copies, so W's form goes. **Not carried**: the condition as inventoried names a difference this change removes. Task 5. | -| e7 | **changed** — gains the read-after-clean-completion clause. The sentence is given entire in §D. | +| e7 | **replaced** — gains the read-after-clean-completion clause. The sentence is given entire in §D. | | e9 | **carried** inside §D's block — `the "clearly stuck" reading above is not a precondition for it` closes the replacement sentence and is reproduced in it. Owes a preservation check, not a span. | | e10 | **kept**, untouched, and **outside** the replacement. `A loop can be worth stopping long before it plateaus.` is the sentence *after* §D's block; an earlier draft listed it as carried, which would have put a kept condition inside a replacement it never enters and left the real boundary of the edit unstated. | | e11 | **kept** — the C-only rationale paragraph is untouched and stays C-only. | @@ -1928,9 +1934,15 @@ what supplied a comparison it had not supplied: ```bash git fetch origin main -git rev-parse origin/main # record this — and it is the argument the check below receives +BASEREF=$(git rev-parse origin/main) # resolve ONCE; record this object name +echo "$BASEREF" ``` +**Pass `$BASEREF` to `check-version-bump.sh`, not `origin/main`.** A remote-tracking ref is +mutable: another fetch between the record and the run makes the evidence name one commit while the +checker resolves another, and the claim that the recorded revision is the argument the check +received would be false. The object name is the argument. + ```bash shellcheck --shell=sh plugins/dev-workflow/hooks/codex-gate.sh && \ shellcheck --shell=sh --exclude=SC2015 plugins/dev-workflow/hooks/codex-gate.test.sh && \ @@ -1941,13 +1953,14 @@ shellcheck --shell=sh scripts/check-version-bump.test.sh && \ HOOK_SH=sh sh plugins/dev-workflow/hooks/codex-gate.test.sh && \ HOOK_SH=dash dash plugins/dev-workflow/hooks/codex-gate.test.sh && \ sh scripts/check-invariants.test.sh && sh scripts/check-invariants.sh && \ -sh scripts/check-version-bump.test.sh && sh scripts/check-version-bump.sh origin/main && \ +sh scripts/check-version-bump.test.sh && sh scripts/check-version-bump.sh "$BASEREF" && \ claude plugin validate . --strict ``` -Expected: exit 0. **`origin/main`, not `main`** — AGENTS.md's battery row writes `main` because a -human running it locally usually has one; here the fetched ref is the one just recorded, and the -recorded revision must be the exact argument the successful check received. +Expected: exit 0. **`$BASEREF`, not `main` and not `origin/main`** — AGENTS.md's battery row writes +`main` because a human running it locally usually has one; here the base is the fetched commit, +resolved once, so the recorded revision and the argument the successful check received are the same +value by construction. - [ ] **Step 4b: Apply all twelve `docs/prompt-standards.md` items to the installed text, and commit the result before Gate B** @@ -2156,11 +2169,17 @@ head_paths=$(git show --name-only --pretty=format: HEAD | sed '/^$/d' | sort) if [ -z "$actual" ] && [ "$head_paths" = "$expected" ]; then echo "final findings files already recorded by a previous attempt — continue at 8b" elif [ "$expected" = "$actual" ]; then + PRE=$(git rev-parse HEAD) # the reviewed tip, before the record commit # shellcheck disable=SC2086 git add $FINAL git commit -m "WIP: Gate-B findings files" || { echo "record commit FAILED"; exit 1; } - test "$(git show --name-only --pretty=format: HEAD | sed '/^$/d' | sort)" = "$expected" \ - || { echo "record commit changed paths beyond the final findings files"; exit 1; } + test "$(git show --name-only --pretty=format: HEAD | sed '/^$/d' | sort)" = "$expected" || { + echo "record commit changed paths beyond the final findings files — restoring the reviewed tip" + git reset --mixed "$PRE" + git status --porcelain + echo "HEAD and index restored. Inspect the paths listed above, re-establish every closure" + echo "condition, obtain a fresh clean response, then retry 8a." + exit 1; } else echo "dirty set is not exactly the final pass's findings files:"; git status --porcelain; exit 1 fi @@ -2180,6 +2199,11 @@ swept the change in, the following clean-tree test passed, and 8b published cont `HEAD` never carried. **And the record commit's own changed-path set is checked afterwards**, because the guard describes the working tree while the commit is what actually lands. +**That post-commit check restores the pre-record tip when it fails**, like every other rejection in +step 8. A bare exit there would leave a rejected WIP commit at `HEAD` carrying an unreviewed path — +and the retry branch would then refuse it, since its changed-path set is not `$FINAL`, so the +prescribed procedure could not resume while the side effect stayed committed. + **"`HEAD` touched some `gate-b-` file" is not that test, and would be unsafe.** Step 7 commits earlier passes' findings files, so `HEAD` can carry one while the *final* pass's two are missing entirely — from a mistyped `NONCE`, a wrong `P`, or a pass whose files were never written. The @@ -2243,8 +2267,9 @@ plan elsewhere calls an invalid close. **`|| true` is replaced by branching on what the state actually is, before committing anything.** "Nothing to commit" and "the commit failed" are different outcomes and only the first is fine. 8a -decides between them with two predicates it can name: the **status-derived dirty set** (`?? ` and -`A ` entries, sorted) and **`HEAD`'s changed-path set**. An empty dirty set with `HEAD` equal to +decides between them with two predicates it can name: the **status-derived dirty set** — every +`git status --porcelain -z` record with its three-character status prefix removed, sorted — and +**`HEAD`'s changed-path set**. An empty dirty set with `HEAD` equal to `$FINAL` is the harmless retry; the dirty set equal to `$FINAL` is the first run; anything else stops. **A `git diff --cached --quiet` test is not among them** — an earlier draft's prose claimed it after the block had stopped using it, which is the overclaim `AGENTS.md` names by requiring From d41ff5a352b94ec6d1b9aa723caedef866fb936d Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 18:40:26 +0200 Subject: [PATCH 119/181] docs(plans): apply Gate-A plan pass 16; a13 is two sentences, c9 is one clause MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit MANDATORY THREE-TELL STOP — findings rose 4 -> 5, Blockers rose 0 -> 2, and the findings cluster on the verification apparatus. Surfaced to Daniel; standing answer applied, loop continued. Two Blockers, two Majors, one Minor. - 8a's clean-tree check after the record commit was still a bare exit, the one post-commit rejection pass 15 did not reach. Both restore the pre-record tip. - $BASEREF was assigned in one fenced block and consumed in another. Each block is its own shell invocation — the same failure Task 0 solves for $BASE with a file — so the checker would have received an empty base argument. It goes to .context/loop-rule-baseref and the battery reads it back with a guard. - c9 was dispositioned "split", which is not one of the six, so the procedure had no result to run — the same shape as pass 15's "changed". The inventory defines c9 as the precedence clause ALONE, so c9 is moved; the plateau rationale carries no id, stays, and owes a preservation observation under its own name. One condition cannot be required to vanish and to remain. - The inventory defines a13 as TWO sentences and §H supplies one. Row P9 sits in the second, so nothing observed the first — "This replaces the pass-count number and nothing else" could survive beside an ordering that changes rather more than a number, with every stated count passing. It owes its own absence check, and the install covers the whole condition rather than P9's sentence. - Step 7's heading imposed the floor on a route the step body correctly exempts. Trend: findings 32, 32, 31, 20, 16, 11, 13, 15, 9, 17, 11, 6, 4, 5, 4, 5. Blockers 1, 0, 0, 10, 7, 2, 4, 2, 2, 1, 2, 1, 1, 2, 0, 2. Majors 28, 25, 23, 7, 7, 6, 4, 9, 6, 14, 7, 5, 1, 2, 2, 2. Floor 3, from risk high (2) and security none (0), read at this pass from docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-plan-om0bdd7udh-pass-16.md | 6 +++ .../2026-09-14-loop-rule-consolidation.md | 54 +++++++++++++------ 2 files changed, 44 insertions(+), 16 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-16.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-16.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-16.md new file mode 100644 index 0000000..3fe32fc --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-16.md @@ -0,0 +1,6 @@ +BLOCKER | high | Task 15 step 8a | After the record commit succeeds and its changed-path set passes, the clean-tree rejection exits without resetting to `$PRE`; this contradicts the step's rule that every post-commit rejection restores the reviewed tip. | A commit-hook side effect or a staged/unstaged split in a final findings file leaves `HEAD` on the record commit, and an extra dirty path then makes the exact-set retry branch impossible to enter. | On this rejection, run `git reset --mixed "$PRE"`, show the remaining status, and require closure conditions plus a fresh clean response before retrying 8a. +BLOCKER | high | Task 15 step 4 | `$BASEREF` is assigned in the first fenced shell block and consumed in a later fenced block, with no instruction or file preserving that shell-local value between executions; Task 0 explicitly persists `$BASE` for this same execution model. | When the blocks run as separate plan steps, the checker receives an empty base argument and the required quality battery cannot complete. | Resolve and consume `$BASEREF` in one shell block, or persist the resolved object name to a scratch file and read it back immediately before invoking `check-version-bump.sh`. +MAJOR | high | Passage (c) disposition and Task 4 steps 1 and 4b | `c9` is assigned the non-existent disposition `split` even though the plan says the six named dispositions are exhaustive; the inventory defines `c9` as only the operative precedence clause, while the plateau rationale is separate unnumbered source text. | The 135-condition accounting fails story criterion 5, and the generic procedure has no class result for `c9`, so execution invents two differently shaped records for one condition. | Mark `c9` moved and give it the required source absence plus §A presence; retain the plateau rationale with a separately named preservation observation that does not claim to be `c9`. +MAJOR | high | Task 7 steps 1b and 3, `a13` block | Row P9 observes only the changed second sentence of `a13`; no OLD fragment observes removal of `This replaces the pass-count number and nothing else`, although target §H deletes that now-false sentence and the inventory includes it in `a13`. | Both copies can retain the falsified sentence while P9, parity, and every stated per-condition count pass, leaving the new ordering beside text claiming the paragraph changed only a number. | Before installation, add and validate an absence fragment for that first sentence, then require `parent=1 worktree=0` in both copies and record it with the other fragment evidence. +MINOR | high | Task 15 step 7 heading | The heading says to loop to a clean pass at or above the floor, while the step body correctly admits a zero-finding logical pass below the floor as the other route to the closing act. | An executor following the step label can run extra passes and override the early-exit behavior the implementation is meant to install. | Rename the step to cover both closure-eligible routes without imposing the floor on a zero-finding pass. +END OF FINDINGS (5 total) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index 9f9cd26..74e8249 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -242,7 +242,7 @@ preservation fragments that an earlier draft chose afterwards. | a3–a12, a14 | **kept**, untouched. The floor arithmetic, the hook-ratio rules and the two comparison points are outside this change. | | a1 | **carried**, not kept. §F item 8a's fenced block opens with `**Both gates are a LOOP with a HARD FLOOR: a minimum number of passes per run` — `a1` itself — and reproduces it so one contiguous string installs. **Recorded as carried for the same reason as `a15`, `a21`, `a22`, `c15` and `i4`–`i8`**: a condition inside a replacement block is not untouched text, and calling it untouched would put it inside an untouched-range span that a correct implementation then fails. It owes a preservation check, not a span. | | a2 | **replaced.** §F item 8a rewrites the HARD FLOOR parenthetical `(Blocker/Major only)` — Task 8, row F10. The filter itself survives in the ordering, which states what a pass counting toward the floor must be; what goes is the parenthetical's claim that Blocker/Major is the *whole* of it. **Marked replaced rather than kept**, because a condition whose text the change in fact rewrites, recorded as preserved, is the dropped-condition failure `AGENTS.md` names. | -| a13 | **replaced** — scoped to its own paragraph. §H's `a13` block. | +| a13 | **replaced** — scoped to its own paragraph. §H's `a13` block. **The inventory defines `a13` as two sentences**, `This replaces the pass-count number and nothing else.` and `Every other rule stated here…`, and §H supplies **one** sentence for both. So the block replaces the whole condition, first sentence included, and **that first sentence owes an absence check of its own** — row P9's fragment sits in the second. §H's parenthetical line range is a locator, not the extent; installing over the second sentence alone leaves `This replaces the pass-count number` standing beside an ordering that changes more than a number, with every stated count still passing. | | a15 | **carried** inside §H's `a16` block, which reproduces it so one contiguous string installs. | | a16 | **replaced** — points at Mechanics · Severity instead of carrying an unscoped copy. §H's `a16` block. | | a17, a18, a19 | **moved** — the floor paragraph stops stating the clean-final-pass rule and the early exit; both are stated once in §A. §H's `a17`–`a22` block. | @@ -265,7 +265,8 @@ preservation fragments that an earlier draft chose afterwards. | c4 | **replaced** — "a missing one means keep going" becomes "means only that *this* exit does not apply", the pass's actual next step being the ordering's. | | c5, c6, c7 | **carried word for word** inside §C's block. | | c8 | **replaced** — gains the re-raised-dismissal clause. | -| c9 | **split.** The operative precedence clause moves into §A capitalized as a standalone sentence; the plateau rationale stays at this source. Not moved whole — §C says so explicitly. | +| c9 | **moved.** The inventory defines `c9` as the operative precedence clause alone — `**a clean completion takes precedence over this exit**` — and that clause moves into §A, capitalized as a standalone sentence. So it owes the moved pair: an absence here, a condition-specific presence in §A. **It was recorded as "split", which is not one of the six dispositions** and left the procedure with no result to run. | +| — | **the plateau rationale is not `c9`** and carries no id: it is unnumbered source text that **stays** at this site while `c9` leaves it. It owes a **preservation** observation under its own name, `parent=1 worktree=1`, and must not be recorded as part of `c9` — one condition cannot be required to vanish and to remain. | | c10, c11 | **moved** to §A, which states what a Blocker/Major-free pass at or above the floor does. | | c12, c13 | **moved** to §A, beside a19. | | c14 | **replaced, not moved.** The ordering splits the below-floor Minor case into suspend and continue; no copy of the live wording survives beside them. | @@ -814,11 +815,12 @@ decided at execution. passage (c), and §C replaces text inside it, so the curve-reading premises owe **per-condition preservation counts** — `parent=1 worktree=1` in each copy — like a carried condition. Without them an accidental edit to the curve or coverage sentences survives every mechanical check here. -- **`c9` is split and needs two rows, not one.** The moved precedence clause owes an **absence** - at this source; the plateau rationale that stays owes a **preservation count**. One row cannot - observe both — the clause must vanish and the rationale must not — and the rationale's own - sentence wraps in both copies, so its fragment has to be chosen single-line **before** step 2 - rather than picked out of the installed text afterwards. +- **`c9` and the plateau rationale are two subjects, not one condition with two halves.** `c9` is + the precedence clause and is **moved** — absence here, condition-specific presence in §A. The + plateau rationale carries no inventory id, **stays**, and owes a **preservation count** recorded + under its own name. One row cannot observe both, because one must vanish and the other must not. + The rationale's sentence wraps in both copies, so its fragment has to be chosen single-line + **before** step 2 rather than picked out of the installed text afterwards. - [ ] **Step 2: Install §C's block** @@ -1135,6 +1137,11 @@ afterwards has no live wording left to check against — and that includes the * fragments for the nine carried conditions** `a15`, `a21`, `a22`, `c15` and `i4`–`i8`, which an earlier draft chose from the installed text at step 4. +**`a13` is two sentences and row P9 sits in the second.** §H replaces the whole condition with one +sentence, so the first — `This replaces the pass-count number and nothing else.` — owes an +**absence** fragment of its own, `parent=1 worktree=0` in both copies. Derive it here, from the +live text, and install over the whole condition rather than over the sentence P9 names. + **Passage (i)'s kept conditions need attention the class alone does not flag.** `i1`–`i3` and `i9`–`i16` are kept, no untouched span reaches them — passage (i) is not one of Task 0's five regions — and `i3` shares its sentence with the dash-delimited list this task replaces. Each owes a @@ -1934,16 +1941,22 @@ what supplied a comparison it had not supplied: ```bash git fetch origin main -BASEREF=$(git rev-parse origin/main) # resolve ONCE; record this object name -echo "$BASEREF" +git rev-parse origin/main > .context/loop-rule-baseref # resolve ONCE; record this object name +cat .context/loop-rule-baseref ``` -**Pass `$BASEREF` to `check-version-bump.sh`, not `origin/main`.** A remote-tracking ref is +**Pass that object name to `check-version-bump.sh`, not `origin/main`.** A remote-tracking ref is mutable: another fetch between the record and the run makes the evidence name one commit while the checker resolves another, and the claim that the recorded revision is the argument the check received would be false. The object name is the argument. +**It goes to a file, not to a shell variable**, for the reason Task 0 gives about `$BASE`: each +fenced block below runs in its own shell invocation, so a `BASEREF=` assignment here is gone by the +time the battery runs and the checker would receive an empty argument. The battery reads it back. + ```bash +BASEREF=$(cat .context/loop-rule-baseref) +test -n "$BASEREF" || { echo "BASEREF empty — the fetch step did not run"; exit 1; } shellcheck --shell=sh plugins/dev-workflow/hooks/codex-gate.sh && \ shellcheck --shell=sh --exclude=SC2015 plugins/dev-workflow/hooks/codex-gate.test.sh && \ shellcheck --shell=sh scripts/check-invariants.sh && \ @@ -2047,7 +2060,7 @@ and is the only value whose range contains the whole implementation. **Standing lens, every call:** "which existing statements does this diff falsify?" and **name what this diff changes the size, value or position of** — a list, a count, a version, an identifier, a cited line — then grep for where each is described elsewhere. -- [ ] **Step 7: Loop to a clean pass at or above the derived floor** +- [ ] **Step 7: Loop until a pass is closure-eligible — clean at or above the derived floor, or zero-finding** Floor derives from the story profile: risk `high` → 2, security `none` → 0, max 2 ≠ 0 → **floor 3**. Re-derive it at each pass from the header. @@ -2183,7 +2196,13 @@ elif [ "$expected" = "$actual" ]; then else echo "dirty set is not exactly the final pass's findings files:"; git status --porcelain; exit 1 fi -test -z "$(git status --porcelain)" || { echo "tree not clean — aborting before 8b"; exit 1; } +test -z "$(git status --porcelain)" || { + echo "tree not clean after the record commit — restoring the reviewed tip" + test -n "${PRE:-}" && git reset --mixed "$PRE" + git status --porcelain + echo "Inspect the paths above, re-establish every closure condition, obtain a fresh clean" + echo "response, then retry 8a." + exit 1; } git rev-parse HEAD > .context/loop-rule-reviewed-tip ``` @@ -2199,10 +2218,13 @@ swept the change in, the following clean-tree test passed, and 8b published cont `HEAD` never carried. **And the record commit's own changed-path set is checked afterwards**, because the guard describes the working tree while the commit is what actually lands. -**That post-commit check restores the pre-record tip when it fails**, like every other rejection in -step 8. A bare exit there would leave a rejected WIP commit at `HEAD` carrying an unreviewed path — -and the retry branch would then refuse it, since its changed-path set is not `$FINAL`, so the -prescribed procedure could not resume while the side effect stayed committed. +**Both post-commit checks restore the pre-record tip when they fail**, like every other rejection in +step 8: the changed-path test **and** the clean-tree test after it. A bare exit at either would +leave a rejected WIP commit at `HEAD` carrying an unreviewed path, or an extra dirty path beside it +— and the retry branch would then refuse the state, since neither the dirty set nor the changed-path +set is `$FINAL`, so the prescribed procedure could not resume while the side effect stayed +committed. `${PRE:-}` is guarded because the retry branch reaches the clean-tree test without +setting it. **"`HEAD` touched some `gate-b-` file" is not that test, and would be unsafe.** Step 7 commits earlier passes' findings files, so `HEAD` can carry one while the *final* pass's two are missing From 67a50e024f16a19c3971b75b55127549c3b83647 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 18:57:52 +0200 Subject: [PATCH 120/181] docs(plans): apply Gate-A plan pass 17; the fragment test is class-specific MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit MANDATORY THREE-TELL STOP — findings rose 5 -> 9, Blockers flat at 2, and the findings cluster on the verification apparatus. Surfaced to Daniel; standing answer applied, loop continued. Two Blockers, three Majors, three Minors, one Nit. - The fragment admission test required every fragment to be ABSENT from its own replacement. Pass 9 widened the table to hold carried-condition preservation fragments, which must be PRESENT in the replacement — the block is what preserves them. The test never widened with the cut, so a correct implementation could not admit its own required rows: it either stops before installing §B or drops its carried coverage silently. The third condition is now stated per class, with the opposite requirement for each. - 8b never asserted HEAD still equals the tip 8a recorded. The two invocations are separate on purpose, so a commit can land between them without leaving the worktree dirty, and reset --soft would then squash it into the closing body with every check passing. It stops before the reset, which is the recoverable direction. - Task 0 admits re-entry after a partial implementation but steps 2 and 3 derived from the worktree, which by then is partly edited. They describe the tree at $BASE: on re-entry they are validated, not re-run, and where they must be built they are built from the $BASE blobs. - Step 7b appended to the closing message, so a second candidate close wrote a second provenance line and curve. It rebuilds the message whole and asserts exactly one of each. - Tasks 2 and 14 ran the cond rows and recorded nothing, while Task 15 reads every preservation from the fragment-evidence section. A zero-finding first Gate-B pass closes before step 7's re-record rule runs, so the omission would have shipped. - h3 is reproduced verbatim inside §F item 7's block — carried, not kept. Third condition to move class, so passage (h)'s kept count is deleted rather than corrected a third time. - Task 4 still said "c9 is split" twice after the disposition table stopped. - Task 7's "four of ten" arithmetic was wrong a third time, omitting the a13 block the paragraph above it had just added. Deleted, not corrected. - The cited path hooks/codex-gate.sh:763 does not resolve from the repo root. Trend: findings 32, 32, 31, 20, 16, 11, 13, 15, 9, 17, 11, 6, 4, 5, 4, 5, 9. Blockers 1, 0, 0, 10, 7, 2, 4, 2, 2, 1, 2, 1, 1, 2, 0, 2, 2. Majors 28, 25, 23, 7, 7, 6, 4, 9, 6, 14, 7, 5, 1, 2, 2, 2, 3. Floor 3, from risk high (2) and security none (0), read at this pass from docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-plan-om0bdd7udh-pass-17.md | 10 +++ .../2026-09-14-loop-rule-consolidation.md | 85 ++++++++++++++++--- 2 files changed, 83 insertions(+), 12 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-17.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-17.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-17.md new file mode 100644 index 0000000..b7d14b3 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-17.md @@ -0,0 +1,10 @@ +BLOCKER | high | Verification procedure and Tasks 3–8 pre-install derivation | The fragment table is expanded to carried and span-less-kept preservation fragments, but the shared pre-install procedure still requires every such fragment to be absent from its own replacement; a carried condition such as `b1` or `e9` must instead occur in that replacement to survive. | A correct implementation cannot satisfy the three-condition admission test for its required preservation rows, so execution either stops before installing §B or silently ignores the plan's fragment rule and loses auditable coverage of carried conditions. | Make fragment validation class-specific: OLD and source-absence fragments must be absent from the replacement, while carried preservation fragments must be present unchanged there; define the appropriate rule separately for kept fragments that merely share a changed line. +BLOCKER | high | Task 15 step 8b | The second closing invocation reads the saved `TIP` but never asserts that current `HEAD` still equals it before `git reset --soft "$BASE"`; the two invocations are deliberately separate, so a commit can land after 8a and before 8b without making the worktree dirty. | The reset then stages and squashes that intervening commit into the closing commit, while the subject and clean-tree checks both pass, defeating the stated rule that 8a's findings commit is the only commit admitted between the reviewed tip and the closing act. | Validate both scratch values, require `git rev-parse HEAD` to equal `TIP` before the reset, and on mismatch stop without moving `HEAD` and require the closure evidence to be re-established. +MAJOR | high | Task 0 steps 1–3 | Step 1 explicitly accepts re-entry after a partial implementation when the recorded base is an ancestor followed only by this run's WIP commits, but steps 2 and 3 then derive the untouched map and the “baseline” parity diff from the current worktree rather than from the recorded base; after any text task, OLD replacement extents are gone and aligned divergences no longer match the baseline expectation. | The advertised recovery path either overwrites its scratch evidence with a map of the partially edited tree or fails on differences the plan itself already installed, so an interrupted execution cannot safely restart through Task 0. | Tie the baseline artifacts to the recorded base and derive their source regions and parity from the `$BASE` blobs, or preserve and validate an existing base-keyed artifact instead of rebuilding it from a partial WIP tip. +MAJOR | high | Task 15 step 7b and step 8 recovery | Step 7b says to append the provenance line, curve and exception record to the existing closing-message file, but a failed closing act deliberately preserves that file and can require a fresh final pass whose curve and evidence replace the prior candidate's. | Re-entering step 7b after that recovery appends a second provenance line and curve, so the eventual commit body violates the standing one-of-each record grammar or carries a stale curve beside the current one. | Rebuild the complete closing message idempotently from the current records at every candidate close, or replace named sections in place and assert exactly one provenance line and one current cycle curve before 8b. +MAJOR | high | Task 14 step 4b and Task 15 steps 5–6 | The `cond` rows in `.context/loop-rule-untouched` are run in Tasks 2 and 14 but no step records their observed preservation results under `## Fragment evidence`; Task 15 nevertheless says every preservation is read from that section, and its complete re-recording rule appears only after the first Gate-B call. | Span-less kept conditions can have no durable preservation evidence in the reviewed `HEAD`, and a zero-finding first Gate-B pass can proceed directly to closing before the later rerun rule ever repairs the omission. | In Task 14 step 4b, record every `cond` result in Task 14's fragment-evidence subsection and commit it before Task 15 issues the first Gate-B pass; state where the untouched-span results are recorded as well. +MINOR | high | Passage (h) disposition and Task 0 step 2 | `h3` is classified among the “kept, untouched” conditions, but target §F item 7 supplies the complete replacement sentence beginning `Which commit:` and reproduces `an ungated change records it in that commit` inside that fenced replacement, which is exactly inventory condition `h3`. | The 135-condition accounting misstates the edit boundary and routes `h3` through incidental line-sharing protection instead of the carried-condition procedure; the stated count of twenty-three untouched `h` conditions is consequently also wrong. | Classify `h3` as carried, reduce the untouched `h` count accordingly, and have Task 8 derive and record its pre-install and post-install preservation count. +MINOR | high | Task 4 steps 4b–5 | The disposition table now correctly makes `c9` moved and explicitly says the adjacent plateau rationale has no id, but Task 4 still says “`c9` is split” twice; the inventory defines `c9` as the precedence clause alone, not the whole source sentence. | The task reintroduces a seventh disposition word and can lead an executor to attach the staying rationale to `c9`, even though the two subjects require opposite results. | Say that the source sentence is split, while `c9` is moved and the unnumbered plateau rationale is preserved separately. +MINOR | high | Task 7 step 1b | The task says four of its ten blocks change more than one thing and that the other six are covered by one pair each, but its immediately preceding `a13` instruction establishes a fifth multi-observation block: P9 covers the second sentence and a separate absence must cover the first. | The stated count disagrees with its own enumeration and tells a reader that one pair covers `a13`, undermining the additional absence check added in pass 16. | Count five multi-observation and five single-observation blocks, or delete the arithmetic and name only the required observations. +NIT | high | Task 15 step 8a introduction | The cited path `hooks/codex-gate.sh:763` does not exist from the reviewed repository root; the verified function is at `plugins/dev-workflow/hooks/codex-gate.sh:763`. | The required cited-path precheck fails and a reader following the citation is sent to a nonexistent file. | Use the full repository-relative plugin path. +END OF FINDINGS (9 total) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index 74e8249..0c7a9fb 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -31,7 +31,7 @@ **Every OLD fragment below was tested against all three conditions** — single-line, unique in each copy the row claims, and **absent from the target text's replacement blocks** — at `9f13a2c`. Each returned one hit in each copy its row names; **row P1 has no fragment**, and rows P5 and P5w are per-copy, so the claim is about the copies each row claims and not about both copies for every row. No task may invent a fragment; a task needing one not listed here adds it and runs the same three checks first. -**Three ways a fragment fails:** it wraps across a line break, so `grep -F` counts zero in a correct file; it is **preserved inside its own replacement**, so its old-wording-gone count can never reach zero; or it is not unique, so a count of 1 proves nothing about which occurrence changed. +**Three ways a fragment fails:** it wraps across a line break, so `grep -F` counts zero in a correct file; it stands in **the wrong relationship to its own replacement** for the result its class expects — an OLD half preserved inside the block can never reach zero, and a carried condition's fragment *absent* from the block can never be found after the install; or it is not unique, so a count of 1 proves nothing about which occurrence changed. **The third test is class-specific and the table below states it per class.** **This table exists because both of the first two drafts got this wrong, in a different one of those three ways each time.** Pass 1 found nine fragments that wrapped or quoted text that does not exist. Pass 2 found three that were preserved inside their own replacements — the condition the checker used for pass 1 did not test, though this plan had stated it. The third condition is now checked **against the target text's fenced blocks**, not against the whole file: a fragment quoted in an item's rationale is not preserved by its replacement, and testing the whole file rejects usable fragments. @@ -109,7 +109,21 @@ |---|---|---| | **Single-line** in the file it is counted in | it wraps, so `grep -F` counts zero in a correct tree | pass 1, nine rows | | **Unique** in that file | a count of 1 proves nothing about which occurrence changed | — | -| **Absent from its own replacement** | its old-count can never reach zero | pass 2, three rows | +| **The right relationship to its own replacement** — and that differs by class | a fragment required to disappear that survives, or one required to survive that is not there | pass 2, three rows; pass 17, the whole carried class | + +**The third condition is class-specific, and reading it as one rule breaks the carried class.** +A fragment whose count must reach **zero** has to be **absent** from the replacement; a fragment +whose count must stay **one** has to be **present** in it, unchanged. Those are opposite tests: + +| The row's class | Its fragment must be | Because its expected result is | +|---|---|---| +| **replaced** (OLD half), **dropped**, **moved** (source) | **absent** from the replacement block | `worktree=0` — a surviving fragment can never reach zero | +| **carried** | **present, unchanged**, in the replacement block | `worktree=1` — the block is what preserves it, so a fragment absent from the block cannot be found afterwards | +| **kept** sharing a line with changed text | **outside** every replacement block's extent | `worktree=1` — it is not reproduced by any block; it survives because nothing replaces it | + +**Applying the absent-from-its-replacement test to a carried row rejects every correct fragment**, +so a task either stops before installing or quietly drops its carried coverage. The table's cut +widened at pass 9 to hold these rows; this test did not widen with it. **The third check compares against the target's fenced blocks with line breaks normalized.** `F4` is one line in `CLAUDE.md` but the block wraps it between `you` and `still`; a substring test finds nothing and the row looks usable. Pass 4 found it, and re-running the normalized check over every row then in the table found that one and no other. @@ -310,7 +324,8 @@ preservation fragments that an earlier draft chose afterwards. | Condition | Disposition | |---|---| -| h1–h3, h6, h8–h12, h14–h18, h20–h26 | **kept**, untouched. | +| h1, h2, h6, h8–h12, h14–h18, h20–h26 | **kept**, untouched. | +| h3 | **carried**, not kept. §F item 7's fenced block reproduces `an ungated change records it in that commit` verbatim — `h3` itself — because the item replaces the whole `**Which commit:**` sentence and only the Gate-A clause changes. A condition inside a replacement block is not untouched text, so it owes a **preservation count** and is covered by no span. Task 8 derives its fragment before installing item 7. | | h5 | **replaced** — §F item 7 changes the Gate-B destination from "restated by the closing amend" to "restated by the commit its closing act produces", the amend no longer being the only closing shape. Row **F7b**, which is its own row because item 7 changes `h4` and `h5` on two different lines and one fragment cannot observe both. | | h4 | **replaced** — §F item 7, the human-exception destination: a Gate-A cycle's record goes to the commit its closing act produces, not to "the spec or plan commit". | | h7 | **kept.** | @@ -416,9 +431,22 @@ the recorded parent, and every parent count then describes the wrong tree withou the only value that puts the whole implementation inside the reviewed range. `.context/` is ignored by the hook's fingerprint, so the file itself moves nothing. +**On re-entry, steps 2 and 3 are not re-run — they are validated.** Step 1 admits a recorded base +followed only by this run's `WIP:` commits, which means text tasks may already have run. **Steps 2 +and 3 describe the tree as it was at `$BASE`**: after a text task, the OLD extents they map are +gone and the aligned divergences no longer match the baseline expectation, so re-deriving them from +the worktree either overwrites the evidence with a map of the partly edited tree or fails on +differences this plan itself installed. + +**So:** if `.context/loop-rule-untouched` and `.context/loop-rule-baseline-diff.txt` already exist +and the recorded base is valid, **keep them and confirm they are keyed to `$BASE`**; re-derive +neither. If they are missing while `$BASE` has `WIP:` commits after it, **derive both from the +`$BASE` blobs** — `git show "$BASE:"` — not from the worktree. On a clean first run the two +are the same thing, which is why this is stated once here rather than in each step. + - [ ] **Step 2: Re-read the five untouched ranges and record their current anchors** -Passages (d), (f) and (j) are not edited by this change, and passage (a)'s arithmetic and passage (h)'s **twenty-three** kept conditions are untouched — twenty-six inventoried, less `h4`, `h5` and `h19`, which Task 8 replaces. **`h5` is in that list and an earlier draft left it out**, counting twenty-four kept: item 7 changes the Gate-B destination as well as the Gate-A one, which is why row F7b exists at all. Record where they are now, because the inventory's numbers cite `7c0d475`: +Passages (d), (f) and (j) are not edited by this change, and passage (a)'s arithmetic and passage (h)'s kept conditions are untouched. **No count of them is stated here.** One was, and it was wrong twice — first omitting `h5`, which item 7 replaces alongside `h4`, then counting `h3` as kept where item 7's block reproduces it and it is carried. **The disposition table above is the one place the classes live**; a number repeated here is a second copy of it that goes stale the next time a condition moves class, which has now happened three times. Record where the regions are now, because the inventory's numbers cite `7c0d475`: **Five ranges, each with a start and an end anchor** — the three whole passages, plus the floor arithmetic and the kept human-exception conditions, which later verification consumes and which the @@ -881,7 +909,10 @@ The carried `c5`–`c7`, the kept `c1`–`c3`, and `c9`'s plateau rationale — class owes. **None of them is observed by a pair**, which is why they are listed as a step rather than left to the walk. -**`c9` is split, so it owes both halves:** the moved precedence clause is absent here +**The source sentence is split; `c9` is not.** `c9` is the precedence clause alone and is **moved**, +while the plateau rationale beside it carries no inventory id, **stays**, and is observed under its +own name. Calling `c9` "split" names a seventh disposition and invites an executor to attach the +staying rationale to a condition required to vanish. So: the moved precedence clause is absent here (`parent=1 worktree=0`) and present in §A, while the plateau rationale stays — confirm the rationale still counts `1` in each copy. @@ -892,7 +923,7 @@ Expected: for the three pairs, six pair instances reading `parent=1 worktree=0` in each copy. **Every OLD row was derived and validated at step 1**; this step only runs them. -- [ ] **Step 5: Walk the conditions** — `c1`–`c3` present unchanged, `c5`–`c7` word for word, `c8` carrying the new clause, `c9` split, `c10`–`c14` gone from here. **This is a reader's confirmation on top of step 4's counts, not the observation for any of them** — `c9`–`c14` are each counted there, and a walk that found what the counts missed would mean a fragment was wrong rather than that the walk was the check. +- [ ] **Step 5: Walk the conditions** — `c1`–`c3` present unchanged, `c5`–`c7` word for word, `c8` carrying the new clause, `c9` moved out and the plateau rationale still here, `c10`–`c14` gone from here. **This is a reader's confirmation on top of step 4's counts, not the observation for any of them** — `c9`–`c14` are each counted there, and a walk that found what the counts missed would mean a fragment was wrong rather than that the walk was the check. - [ ] **Step 6: Commit** @@ -1147,8 +1178,10 @@ live text, and install over the whole condition rather than over the sentence P9 regions — and `i3` shares its sentence with the dash-delimited list this task replaces. Each owes a per-condition preservation count, `parent=1 worktree=1`. -**One pair per block is a sample, not coverage, and four of the ten blocks change more than one -thing:** +**One pair per block is a sample, not coverage.** The blocks below each change more than one thing +and are named with what they owe; the rest are covered by one pair each. **No count is stated** — +the arithmetic here was wrong three times, most recently by omitting the `a13` block the paragraph +above had just added to the list. Take the set from the bullets, not from a number: - **The `c18`-and-surfacing block** changes `c16`, `c17`, `c19` and `c20` besides the `c18` clause row P8 observes — **one OLD row per condition**, four in total. @@ -1173,8 +1206,7 @@ thing:** executor would have to pair the clause with unrelated text and the observation would prove nothing about the clause. One presence check per independent added clause. -**The other six blocks change one thing each and one pair covers them** — six, because four of the -ten are named above. +**Every block not named above changes one thing and one pair covers it.** **One classification note, decided against the real file rather than asserted.** The gate-prompt template's clean sentence **replaces** `A clean pass is the single body line …`, which is why P13 @@ -1888,6 +1920,14 @@ worktree and compare against the two expected values the row carries. **Read the not from a list here**: which kept conditions needed one is decided at Task 0 against the real lines, and an enumeration in this step would be a second copy of it. +**Record every one of these results — the `cond` rows and the span diffs both — in this task's +`### Task 14` fragment-evidence subsection, and commit it in step 5.** Task 15 step 5 says every +preservation is read from that section, and nothing was writing them there: a span-less kept +condition would then have no durable evidence in the reviewed `HEAD` at all, and **a zero-finding +first Gate-B pass closes before step 7's re-record rule ever runs**, so the omission would ship. A +span diff's result is `no difference` and belongs beside the counts, for the same reason Task 12b +records "nothing". + **Anchors, resolved separately in each tree** — Task 1 inserts a large block, so a line number taken from either tree addresses different text in the other, and an earlier draft's numeric `sed` addresses would have reported CHANGED on identical content. Expected: no output for every recorded @@ -2125,7 +2165,14 @@ to pass CI, or an invariant-11 violation, published by a cycle that closed clean - [ ] **Step 7b: After the clean pass, complete `.context/loop-rule-closing-msg` — this is the action step 5 defers to, and step 8 has no other source for these records** -Append to the file step 5 opened, then read the whole file back before step 8 runs: +**Rebuild the file whole, do not append to it.** Step 5 opened it with a draft entry, and a failed +closing act preserves it while requiring a fresh final pass whose curve and evidence supersede the +previous candidate's. **Appending on re-entry writes a second provenance line and a second curve**, +which breaks the one-of-each grammar Mechanics pins, or leaves a stale curve standing beside the +current one. Write the complete message from the current records at every candidate close, then +assert **exactly one** provenance line and **exactly one** curve for this cycle before step 8. + +The message carries, in this order: 1. the **provenance line**, in the form Mechanics pins — the Gate-B cycle's nonce, the derived floor, the cited set with each member's level, and the workspace knob; @@ -2152,13 +2199,20 @@ so an evidence entry written only into a WIP message is destroyed at exactly the closes — which is what step 5 would otherwise have done. **8a — record the final pass's findings files. Its own shell invocation, and that is not -cosmetic.** `codex-gate.sh`'s `is_wip_commit` (`hooks/codex-gate.sh:763`) tests the **whole command +cosmetic.** `codex-gate.sh`'s `is_wip_commit` (`plugins/dev-workflow/hooks/codex-gate.sh:763`) tests the **whole command string** it is given against `-m[[:space:]]*['"]?[[:space:]]*wip`, case-insensitively: **`-m` immediately followed by optional whitespace, an optional quote, and `wip`.** A single block carrying this `-m "WIP: …"` and the closing `git commit -F` matches, and is classified cycle-internal in its entirety — the hook would then carry this cycle's Gate-B count and fingerprint into the next one, a real closing commit read as a snapshot. **So run 8a and 8b as separate Bash calls.** +**8b re-establishes that `HEAD` is still the tip 8a recorded, and stops without moving it if not.** +The split into two invocations is what makes this necessary: time passes between them, and a commit +landing in that window leaves the worktree clean, so every other check in 8b passes while +`reset --soft` stages and squashes that commit into the closing body — defeating the rule that 8a's +findings commit is the only commit admitted between the reviewed tip and the closing act. +**Stopping before the reset is the recoverable direction**; stopping after it is not. + **What 8b must avoid is that pattern, not the letters.** `--mixed` contains `-m` and matches nothing, because `i` follows; a path containing `wip` matches nothing, because no `-m` precedes it. Saying "no `-m` and no `wip` in 8b" would outrun the check the hook performs and make a correct @@ -2237,6 +2291,13 @@ commit 8a itself would have made. ```bash BASE=$(cat .context/loop-rule-base); TIP=$(cat .context/loop-rule-reviewed-tip) +test -n "$BASE" && test -n "$TIP" || { echo "BASE or TIP missing — 8a did not complete"; exit 1; } +# 8a and 8b are separate invocations on purpose, so time passes between them and a +# commit can land without leaving the worktree dirty. Nothing below would notice. +test "$(git rev-parse HEAD)" = "$TIP" || { + echo "HEAD has moved since 8a — it is $(git rev-parse HEAD), the reviewed tip is $TIP" + echo "NOT resetting. Whatever landed is unreviewed; re-establish the closure evidence first." + exit 1; } git reset --soft "$BASE" git commit -F .context/loop-rule-closing-msg || { echo "closing act FAILED — restoring the reviewed tip without discarding its side effects" From 6d276fafb659ede3bf22f13100ff8d8c6e3796f1 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 19:13:17 +0200 Subject: [PATCH 121/181] docs(plans): apply Gate-A plan pass 18; give spans a record shape and the map a base MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit One tell only (instrument cluster) — no mandatory stop. Findings fell 9 -> 4 and Blockers fell 2 -> 0. Three Majors and one Minor, all in text the last two passes added. - "Confirm they are keyed to $BASE" named a comparison the artifacts could not support: neither carried the base it was built from, so a stale map from an abandoned run was indistinguishable from this one's. Both now open with a `base` line and a re-entry compares it. - Task 14 was told to record its span-diff results in the fragment-evidence section, whose four shapes all require a fragment and counts. A span is two bounded extracts compared against each other and has neither, so the executor would have invented a fifth format or dropped the results. There is a fifth shape now, and Task 15 reads it. - Step 4b's subject set said "the seven hook strings". §F items 12, 15 and 16 replace both channels, so the change ships TEN prompt bodies — seven additionalContext and three systemMessage. Counting messages instead of bodies left three shipped operator prompts outside invariant 11's only gate. - The absence marker read `Human exception: none`, which is the reserved prefix of the fixed three-line record without its handle, date or two required lines. Every ordinary close would have written a malformed pseudo-record. It is `Human exceptions: none`. Trend: findings 32, 32, 31, 20, 16, 11, 13, 15, 9, 17, 11, 6, 4, 5, 4, 5, 9, 4. Blockers 1, 0, 0, 10, 7, 2, 4, 2, 2, 1, 2, 1, 1, 2, 0, 2, 2, 0. Majors 28, 25, 23, 7, 7, 6, 4, 9, 6, 14, 7, 5, 1, 2, 2, 2, 3, 3. Floor 3, from risk high (2) and security none (0), read at this pass from docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-plan-om0bdd7udh-pass-18.md | 5 ++ .../2026-09-14-loop-rule-consolidation.md | 53 ++++++++++++++----- 2 files changed, 46 insertions(+), 12 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-18.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-18.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-18.md new file mode 100644 index 0000000..d933799 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-18.md @@ -0,0 +1,5 @@ +MAJOR | high | Task 0 re-entry rule before Step 2 | `.context/loop-rule-untouched` and `.context/loop-rule-baseline-diff.txt` carry no `$BASE` field or sidecar, so "confirm they are keyed to `$BASE`" has no defined comparison and cannot distinguish prior-run scratch from this baseline | A re-entry can reuse a stale, incomplete span/condition map or inherited-divergence record, allowing untouched-condition or parity checks to certify the wrong parent | Persist the base object name with both artifacts and fail unless it equals `.context/loop-rule-base`, or rebuild and fully validate both from the recorded base blobs. +MAJOR | high | Task 14 step 4b / Fragment evidence schema | Step 4b requires every span-diff result in `### Task 14`, but the terminal schema says every observation must be one of four shapes and none can represent a span because those shapes require fragment/count fields rather than two bounded extracts | The executor must invent a fifth format or omit the span results, and Task 15's closing-evidence reader has no defined way to consume them | Keep the four condition shapes for `cond` rows and define a separate durable span-result table with file, anchors, parent, worktree and expected no difference that Task 15 explicitly reads. +MAJOR | medium | Task 15 step 4b | The recorded subject set calls the hook surface "the seven hook strings", although §F items 12, 15 and 16 replace both `additionalContext` and `systemMessage`, so the change has ten prompt bodies and the three terse operator messages are not named | A literal execution can apply the twelve prompt standards only to the seven long bodies and leave three shipped hook prompts outside invariant 11's reader gate | Name all seven `additionalContext` bodies and all three `systemMessage` bodies in the subject set and record all ten under each checklist item. +MINOR | high | Task 15 step 7b | The required literal `Human exception: none` uses the reserved prefix of the fixed three-line human-exception record while omitting its handle/date and two required lines, even though the same step says no exception record is owed in the ordinary case | Every ordinary close writes a malformed pseudo-record into the commit body, leaving a later reader unable to distinguish an absence marker from a corrupted exception record | Use a non-record label such as `Human exceptions: none`, or keep the absence assertion outside the closing body. +END OF FINDINGS (4 total) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index 0c7a9fb..f40abbe 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -438,11 +438,21 @@ gone and the aligned divergences no longer match the baseline expectation, so re the worktree either overwrites the evidence with a map of the partly edited tree or fails on differences this plan itself installed. -**So:** if `.context/loop-rule-untouched` and `.context/loop-rule-baseline-diff.txt` already exist -and the recorded base is valid, **keep them and confirm they are keyed to `$BASE`**; re-derive -neither. If they are missing while `$BASE` has `WIP:` commits after it, **derive both from the -`$BASE` blobs** — `git show "$BASE:"` — not from the worktree. On a clean first run the two -are the same thing, which is why this is stated once here rather than in each step. +**Each artifact carries the base it was built from, as its first line:** + +``` +base +``` + +Written by the step that creates it, and it is what "keyed to `$BASE`" means — without it there is +no comparison to make and a stale map from an abandoned run is indistinguishable from this one's. + +**So:** if `.context/loop-rule-untouched` and `.context/loop-rule-baseline-diff.txt` already exist, +**read their `base` line and require it to equal `.context/loop-rule-base`**; on a match keep them +and re-derive neither, on a mismatch or a missing line delete them and rebuild. If they are absent +while `$BASE` has `WIP:` commits after it, **derive both from the `$BASE` blobs** — +`git show "$BASE:"` — not from the worktree. On a clean first run the two sources are the +same thing, which is why this is stated once here rather than in each step. - [ ] **Step 2: Re-read the five untouched ranges and record their current anchors** @@ -515,11 +525,13 @@ and 14 consume them without a human in between. A file whose second half has no its consumers skip, which is what the per-condition list exists to prevent: ``` +base span cond ``` -Tab-separated, one record per line, the leading keyword distinguishing them. A kept condition's +Tab-separated, one record per line, the leading keyword distinguishing them. **The `base` line is +first and there is exactly one**, so a re-entry can tell this run's map from an abandoned run's. A kept condition's expected pair is `11`; the shape carries the values rather than assuming them, so a moved or dropped condition recorded here later needs no new format. **Both consumers validate every `cond` row**, not only the `span` rows. @@ -2036,8 +2048,13 @@ the refreshed record. A repair made after the battery, with the battery not re-r defect as a Gate-B fix with no re-run. **Record the subject set once, at the top of the output section**, and make each item's result -refer to it: the §A–§H blocks installed in C, the same in W, and the seven hook strings. -**Twelve `PASS` lines alone cannot be told from a review that skipped a copy or the hook** — the +refer to it: the §A–§H blocks installed in C, the same in W, and **ten hook prompt bodies, not +seven** — the seven `additionalContext` bodies items 10–13, 15–17 replace, **plus the three +`systemMessage` bodies** items 12, 15 and 16 replace alongside them. Those three are short +operator-facing prompts and they ship; counting the messages rather than the bodies leaves them +outside invariant 11's only reader gate. + +**Twelve `PASS` lines alone cannot be told from a review that skipped a copy or a channel** — the subject list is what makes the twelve lines mean something. **Commit the result before step 6.** Gate B reviews the range `$BASE..HEAD`; an edit to this plan @@ -2184,10 +2201,15 @@ The message carries, in this order: **Items 1, 2 and 4 are owed unconditionally; item 3 is owed only where such a record exists.** Confirm the file carries the three, and either the applicable exception records or **the literal -line `Human exception: none`**, which is what satisfies this check for an ordinary cycle. Without +line `Human exceptions: none`**, which is what satisfies this check for an ordinary cycle. Without that line an executor reading "all four" must either invent a record or ignore the oracle, and a valid close stops on a record nobody owed. +**Plural, and deliberately not `Human exception:`.** That singular prefix opens the fixed +three-line record Mechanics pins, and a line carrying it without its handle, date, `Not done:` and +`Accepted because:` is a **malformed record**, not an absence marker — every ordinary close would +write one, and a later reader could not tell it from an exception record that lost its body. + **A missing one is not recoverable after step 8** — `git commit -F` publishes whatever the file holds, `reset --soft` has already discarded every WIP body, and a closing commit without its provenance line or curve has not validly closed the cycle. @@ -2439,7 +2461,7 @@ idempotently on re-run, touching no other. **The pre-existing fragments live in never here**; this section records what each observation actually returned, and it is what Task 15 step 5 reads to assemble the closing evidence entry.* -***Four record shapes, because the classes do not return the same number of values.** Every +***Five record shapes, because the classes do not return the same number of values.** Every observation this plan makes is one of them, and a shape that fits only pairs is how a required count gets run and then vanishes from both the plan and the closing evidence:* @@ -2448,9 +2470,16 @@ pair old/worktree old/parent presence worktree parent # add-only, and a moved condition's destination absence parent worktree # dropped, and a moved condition's source preservation parent worktree # carried, and kept where no span holds it +span # an untouched range, parent vs worktree ``` +*The **fifth shape is for ranges, not conditions**: an untouched span is two bounded extracts +compared against each other, so it has no fragment and no counts and the four condition shapes +cannot express it. Its `` is `no difference` or the difference itself — recorded either +way, because a span that was never run and a span that compared equal are otherwise the same +record. Task 15 reads this shape alongside the other four.* + *A **moved** condition therefore contributes two lines — one `absence` at its source, one `presence` at its destination — and both are required for it to count as observed. **Every one of -the four classes goes into the closing evidence entry**; naming only pairs and presence leaves the -absences and preservations run but unrecorded.* +the five shapes goes into the closing evidence entry**; naming only pairs and presence leaves the +absences, preservations and span results run but unrecorded.* From 0b936eb35a0749223dd26cfd892336aba31f7d26 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 19:13:24 +0200 Subject: [PATCH 122/181] docs(plans): drop a stale shape count missed in pass 18 Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index f40abbe..cebca8d 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -2091,7 +2091,7 @@ It names: the battery run; **every pair this plan built, with its counts in each for those two reader records. **An entry assembled from one section would silently drop whatever the other five hold.** -**All four record shapes belong in the entry, not only pairs and presence.** A **dropped** +**Every record shape belongs in the entry, not only pairs and presence.** A **dropped** condition's absence and a **moved** condition's absence are what prove an obsolete instruction was removed; a **carried** or span-less **kept** condition's preservation count is what proves reproduced text survived. An entry listing only pairs and presence claims the verification set From 76cd2ce95456227ad3312193a5c8368b7963516b Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 19:26:14 +0200 Subject: [PATCH 123/181] docs(plans): apply Gate-A plan pass 19; every recording task stages the plan MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit One tell only (instrument cluster) — no mandatory stop. Findings fell 4 -> 3, Blockers held at 0. Two Majors and one Nit. - Nine editing tasks append fragment rows or record observations into this tracked plan and staged only the prompt or hook files. The evidence stayed dirty after the task commit, so Task 0's re-entry path — which requires a clean tree — was unusable after an interruption, and continuing swept the evidence into a later unrelated commit. The staging rule was also a stale task-number list naming five of the nine. It is stated by behaviour now: a task that appends a row or records an observation stages the plan. - Task 0 step 3 wrote .context/loop-rule-baseline-diff.txt without the base line pass 18's re-entry rule compares, so a correct first run produced an artifact its own next run had to reject. It also read the worktree unconditionally, which after any text task folds this plan's own edits into the inherited-drift record — the one distinction Task 14 depends on. The source is selected by the re-entry rule now, from the $BASE blobs once $BASE..HEAD is non-empty. - "ba15e83 appears once" was textually false; it has one ROLE, and the claim now says so. Trend: findings 32, 32, 31, 20, 16, 11, 13, 15, 9, 17, 11, 6, 4, 5, 4, 5, 9, 4, 3. Blockers 1, 0, 0, 10, 7, 2, 4, 2, 2, 1, 2, 1, 1, 2, 0, 2, 2, 0, 0. Majors 28, 25, 23, 7, 7, 6, 4, 9, 6, 14, 7, 5, 1, 2, 2, 2, 3, 3, 2. Floor 3, from risk high (2) and security none (0), read at this pass from docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-plan-om0bdd7udh-pass-19.md | 4 + .../2026-09-14-loop-rule-consolidation.md | 80 ++++++++++++++----- 2 files changed, 63 insertions(+), 21 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-19.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-19.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-19.md new file mode 100644 index 0000000..fb574c8 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-19.md @@ -0,0 +1,4 @@ +MAJOR | high | Task 0 step 3 | `.context/loop-rule-baseline-diff.txt` is created by the shown pipeline without the required first-line `base` record, so it starts with `== ...`; the same step also reads the worktree even though the re-entry contract requires `$BASE` blobs after WIP commits exist | The next re-entry necessarily treats a correctly produced first-run artifact as invalid, and following the shown rebuild command after partial implementation folds the plan's own edits into the supposed baseline, so inherited drift and introduced drift can no longer be distinguished reliably | Make step 3 prepend the validated base record and implement the stated source selection explicitly: retain a matching artifact, otherwise build every extraction and the squash-carry comparison from `git show "$BASE:"` when `$BASE..HEAD` is non-empty, using the worktree only on the clean first run +MAJOR | high | Tasks 3–10 commit steps | These editing tasks append OLD rows and/or write their observed NEW/presence/preservation results into this tracked plan, but their commit commands stage only the prompt or hook files; this contradicts Task 1's rule that each task stages the plan with its evidence, and Tasks 5, 9, and 10 are not even named in that rule despite recording evidence | After the first affected task, the intended plan edit remains dirty, so an interruption cannot use Task 0's mandatory clean-tree re-entry path; if execution continues, the evidence is swept into a later unrelated plan-record commit instead of the independently reviewable task snapshot the plan promises | Add this plan to every editing task's `git add` when that task appends a row or records fragment evidence (including Tasks 3–10 as applicable), and make the staging rule enumerate by behavior rather than a stale task-number list +NIT | high | Task 0 step 1 | The prose says ``ba15e83` appears once, as the approval commit`, but the literal object name appears seven times in this plan | The stated exact count is mechanically false and confuses a single semantic role with textual occurrence, weakening the plan's otherwise explicit count discipline | Replace the occurrence claim with a role claim such as ``ba15e83` has one role: the approval commit`, or remove the count +END OF FINDINGS (3 total) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index cebca8d..df1035d 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -379,8 +379,9 @@ every install and the final soft reset would run from the wrong starting tree. **The inputs are compared by blob, not merely found.** `ba15e83` being an ancestor says the approved commit is in this history; it does not say the three files still hold what was approved, and a later unapproved edit to the target text would supply different installation instructions to -every task. **`ba15e83` appears once, as the approval commit**, and the loop compares each input's -object id against its version there. +every task. **`ba15e83` has exactly one role here — the approval commit** — and the loop compares each input's +object id against its version there. (It is written more than once; the claim is about its role, +not its occurrences.) **What this cannot establish, stated rather than implied:** that *this plan revision* is the one Gate A closed on. A plan cannot name its own closing commit, and any sha written here would be from @@ -561,6 +562,27 @@ none of passages (g), (h) or (j) — so `g4`, which sits at C 815 / W 997, could whose expected list named it — and the truncation hid about fifty of the roughly ninety lines the comparison actually emits. +**Read both copies from the source the re-entry rule selects, not from the worktree unconditionally.** +On a clean first run they are the same; once `$BASE..HEAD` is non-empty the worktree carries this +plan's own edits, and comparing them would fold introduced drift into the inherited-drift record — +after which Task 14 can no longer tell the two apart, which is the whole purpose of this baseline. +**Set `C_SRC` and `W_SRC` first:** + +```bash +BASE=$(cat .context/loop-rule-base) +if [ -n "$(git log --oneline "$BASE"..HEAD)" ]; then + git show "$BASE:CLAUDE.md" > .context/loop-rule-c.src + git show "$BASE:plugins/dev-workflow/commands/workflow-init.md" > .context/loop-rule-w.src + C_SRC=.context/loop-rule-c.src; W_SRC=.context/loop-rule-w.src +else + C_SRC=CLAUDE.md; W_SRC=plugins/dev-workflow/commands/workflow-init.md +fi +printf 'base\t%s\n' "$BASE" > .context/loop-rule-baseline-diff.txt +``` + +**The `base` line is written first and by this step**, not assumed: the re-entry rule compares it, +so a first run that omits it produces an artifact its own next run must reject. + ```bash # Tab-separated start and end anchors: the anchors contain colons, so a # colon delimiter splits '**Severity:**' at the wrong place and yields an @@ -579,14 +601,14 @@ printf '%b\n' \ > .context/loop-rule-sites while IFS=$(printf '\t') read -r s e; do echo "== $s" - diff <(sed -n "/$s/,/$e/p" CLAUDE.md) \ - <(sed -n "/$s/,/$e/p" plugins/dev-workflow/commands/workflow-init.md) -done < .context/loop-rule-sites | tee .context/loop-rule-baseline-diff.txt + diff <(sed -n "/$s/,/$e/p" "$C_SRC") \ + <(sed -n "/$s/,/$e/p" "$W_SRC") +done < .context/loop-rule-sites | tee -a .context/loop-rule-baseline-diff.txt # The squash-carry sentence is ONE line and must not go through the loop. echo "== On squash-merge" | tee -a .context/loop-rule-baseline-diff.txt -diff <(grep -F 'On squash-merge, copy every evidence entry' CLAUDE.md) \ - <(grep -F 'On squash-merge, copy every evidence entry' plugins/dev-workflow/commands/workflow-init.md) \ +diff <(grep -F 'On squash-merge, copy every evidence entry' "$C_SRC") \ + <(grep -F 'On squash-merge, copy every evidence entry' "$W_SRC") \ | tee -a .context/loop-rule-baseline-diff.txt ``` @@ -676,11 +698,18 @@ git add CLAUDE.md plugins/dev-workflow/commands/workflow-init.md \ git commit -m "WIP: install the closure ordering into both §5 copies" ``` -**The plan is staged here because this task writes its fragment evidence into the plan.** Every -task that records a fragment — an appended OLD row, or a chosen NEW fragment with its counts — -stages the plan with its own edit; otherwise the reviewed fragment evidence stays dirty and is -swept into a later, unrelated commit, and the task commits are not the independently reviewable -units this plan claims they are. The same applies to Tasks 3, 4, 6, 7 and 8. +**The plan is staged here because this task writes its fragment evidence into the plan**, and the +rule is stated by behaviour rather than by a list of task numbers: **every task that appends a +fragment-table row or records an observation in `## Fragment evidence (per-task output)` stages +this plan in its own commit.** On the current shape that is every editing task, Tasks 1 and 3–11, +but the rule is the behaviour and a task that stops recording stops owing it. + +**A task-number list here was wrong twice** — it named Tasks 3, 4, 6, 7 and 8 while Tasks 5, 9, 10 +and 11 also record — and a stale list is the same defect as a stale count. **The cost of the +omission is concrete:** the evidence stays dirty after the task commits, so Task 0's re-entry path, +which requires a clean tree, cannot be used after an interruption; and if execution continues, that +task's evidence is swept into a later unrelated commit rather than the independently reviewable +snapshot this plan promises. Named `WIP:` because Task 15 runs Gate B over the whole change and closes it with **one `git reset --soft "$BASE"` and a single commit**, per the Global Constraints and Task 15 step 8 — @@ -814,7 +843,8 @@ as inherited drift. **Any difference other than the parenthetical is a failure o - [ ] **Step 6: Commit** ```bash -git add CLAUDE.md plugins/dev-workflow/commands/workflow-init.md +git add CLAUDE.md plugins/dev-workflow/commands/workflow-init.md \ + docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md git commit -m "WIP: replace passage (b) with the absorb paragraph that owns the fix set" ``` @@ -940,7 +970,8 @@ step only runs them. - [ ] **Step 6: Commit** ```bash -git add CLAUDE.md plugins/dev-workflow/commands/workflow-init.md +git add CLAUDE.md plugins/dev-workflow/commands/workflow-init.md \ + docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md git commit -m "WIP: widen the clearly-stuck third condition and split its precedence sentence" ``` @@ -1043,7 +1074,8 @@ the class alone does not tell you: - [ ] **Step 5: Commit** ```bash -git add CLAUDE.md plugins/dev-workflow/commands/workflow-init.md +git add CLAUDE.md plugins/dev-workflow/commands/workflow-init.md \ + docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md git commit -m "WIP: read the two-tell threshold after the clean-completion branch" ``` @@ -1138,7 +1170,8 @@ Expected: no output. This passage should now be byte-identical, `g4` having been - [ ] **Step 6: Commit** ```bash -git add CLAUDE.md plugins/dev-workflow/commands/workflow-init.md +git add CLAUDE.md plugins/dev-workflow/commands/workflow-init.md \ + docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md git commit -m "WIP: answer the demotion question and scope the resolve duty to the fix set" ``` @@ -1299,7 +1332,8 @@ Expected: `1` each, and `1` in each copy for every one of the nine. - [ ] **Step 6: Commit** ```bash -git add CLAUDE.md plugins/dev-workflow/commands/workflow-init.md +git add CLAUDE.md plugins/dev-workflow/commands/workflow-init.md \ + docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md git commit -m "WIP: install the one-contract paragraph and the remaining prompt-copy replacements" ``` @@ -1376,7 +1410,8 @@ State the number you observed. **Do not carry a count from §F into a check** - [ ] **Step 6: Commit** ```bash -git add CLAUDE.md plugins/dev-workflow/commands/workflow-init.md +git add CLAUDE.md plugins/dev-workflow/commands/workflow-init.md \ + docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md git commit -m "WIP: replace the fourteen falsified sentences in the two prompt copies" ``` @@ -1421,7 +1456,8 @@ The hook reporting its own threshold as an obligation at a floor of 1 is **not** - [ ] **Step 6: Commit** ```bash -git add CLAUDE.md plugins/dev-workflow/commands/workflow-init.md +git add CLAUDE.md plugins/dev-workflow/commands/workflow-init.md \ + docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md git commit -m "WIP: correct the Named residual's blanket exemption and the work-loop sequence" ``` @@ -1560,7 +1596,8 @@ point where green is expected. - [ ] **Step 7: Commit** ```bash -git add plugins/dev-workflow/hooks/codex-gate.sh +git add plugins/dev-workflow/hooks/codex-gate.sh \ + docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md git commit -m "WIP: replace the seven gate reminders the ordering falsifies" ``` @@ -1642,7 +1679,8 @@ Expected: exit 0. **The `--exclude=SC2015` is a single-code exclusion**, not a b - [ ] **Step 7: Commit** ```bash -git add plugins/dev-workflow/hooks/codex-gate.test.sh +git add plugins/dev-workflow/hooks/codex-gate.test.sh \ + docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md git commit -m "WIP: move every hook assertion that names a replaced reminder string" ``` From 9aa1178de665398a30814c9870005a319c503129 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 19:36:33 +0200 Subject: [PATCH 124/181] docs(plans): apply Gate-A plan pass 20; a fragment must lie inside what the edit removes MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit MANDATORY THREE-TELL STOP — findings rose 3 -> 4, Blockers rose 0 -> 2, and the findings cluster on the verification apparatus. Surfaced to Daniel; standing answer applied, loop continued. Two Blockers, two Majors. - Rows F11, F12 and F13 each end just BEFORE their item's first changed word, so a correct installation leaves them standing: old/worktree=1 for six pair instances, and Task 8 could not have met its own required result. Verified against both copies and against the target's blocks. The root is the test, not the three rows. "Absent from the replacement block" is necessary and NOT sufficient — a fragment in the unchanged text before an edit passes it and then survives, because the edit never reaches it. The rule is now that a disappearing class's fragment must lie WHOLLY INSIDE the live text its edit removes. The three rows are replaced with fragments that do, each verified single-line, unique in both copies, and absent from the blocks. - C_SRC and W_SRC were assigned in one fenced block and consumed in the next two — the same defect already fixed for $BASE and $BASEREF. An empty filename in a process substitution need not make diff fail, so the step would have certified a parity baseline over two empty extracts. One block, plus a test -r guard. - Task 0 compared the approved inputs at HEAD, while Task 15 step 7 permits a Gate-B fix to update the spec in a WIP snapshot. The advertised re-entry path would have rejected a state the plan itself creates. It compares at the recorded base. - The scratch artifacts were accepted on their base line alone, so an interruption after writing that line and before the cond rows left a same-base partial map that Tasks 2 and 14 consume as complete. They are written through a temp file and renamed, and a re-entry validates shapes and coverage rather than the first line. Trend: findings 32, 32, 31, 20, 16, 11, 13, 15, 9, 17, 11, 6, 4, 5, 4, 5, 9, 4, 3, 4. Blockers 1, 0, 0, 10, 7, 2, 4, 2, 2, 1, 2, 1, 1, 2, 0, 2, 2, 0, 0, 2. Majors 28, 25, 23, 7, 7, 6, 4, 9, 6, 14, 7, 5, 1, 2, 2, 2, 3, 3, 2, 2. Floor 3, from risk high (2) and security none (0), read at this pass from docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-plan-om0bdd7udh-pass-20.md | 5 ++ .../2026-09-14-loop-rule-consolidation.md | 72 ++++++++++++++----- 2 files changed, 58 insertions(+), 19 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-20.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-20.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-20.md new file mode 100644 index 0000000..d29705a --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-20.md @@ -0,0 +1,5 @@ +BLOCKER | high | Task 0 step 3 | `C_SRC` and `W_SRC` are assigned in one fenced shell block but consumed in the next two, even though this plan explicitly treats fenced blocks as separate shell invocations; neither value is persisted or reloaded. | The `sed` and `grep` process substitutions receive empty filenames, and their failures need not make `diff` fail, so Task 0 can append headings over two empty extracts and certify a false parity baseline instead of the approved-base text. | Put source selection and all extractions in one shell block, or persist and reload both paths with non-empty and readable-file guards and make every extraction failure fatal before accepting the diff. +BLOCKER | high | Task 8 step 3 | Rows F11, F12 and F13 choose OLD fragments that end before the first changed token: the installed item 8b still contains the entire F11 fragment, item 9a still contains the entire F12 fragment, and item 9b still contains the entire F13 fragment once unchanged surrounding text is joined to each replacement. | A correct installation leaves `old/worktree=1` for six pair instances, so Task 8 cannot meet its required `old/worktree=0` result and the implementation cannot reach Gate B without ignoring its verification contract. | Replace F11–F13 with unique single-line fragments that each include wording actually removed by its own item, then rerun the class-specific normalized replacement check in both prompt copies. +MAJOR | high | Task 0 step 1 and Task 15 step 7 | Task 0 always compares the target-text blob at `HEAD` with `ba15e83`, while Task 15 explicitly permits a Gate-B fix to update that target text in a committed WIP snapshot. | After an allowed spec-changing Gate-B repair, any interrupted run that follows the advertised Task 0 re-entry path is rejected as unapproved input even though its recorded base and WIP history are valid, so the recovery procedure cannot resume a state the plan itself creates. | On first entry compare `HEAD` with the approval commit; on re-entry compare the recorded `$BASE` blobs with it and separately validate that every later commit belongs to this WIP execution. +MAJOR | high | Task 0 re-entry rule before steps 2 and 3 | Existing baseline artifacts are accepted solely because their first `base` line matches; there is no completeness, schema or end-of-write validation before the plan says to keep them and re-derive neither. | An interruption after writing the base line but before all `span` and `cond` rows can leave a same-base partial `.context/loop-rule-untouched` file that Tasks 2 and 14 consume as complete, silently removing untouched-condition coverage; the baseline-diff artifact has the same partial-write ambiguity. | Write both artifacts through temporary files and rename only after successful completion, and on re-entry validate the exact required record shapes and coverage before reuse or rebuild them from the `$BASE` blobs. +END OF FINDINGS (4 total) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index df1035d..b0d9bd5 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -72,9 +72,9 @@ | F8 | 8, the profile-change pass claim | ``snapshot by amend — a non-`WIP` commit reads to the hook as the cycle closing and would`` | 752 | | F9 | 7a, the mid-run recovery sentence | `taken from the provenance line and the curve, which must agree. A Gate-A cycle mid-run has no` | 416 | | F10 | 8a, the HARD FLOOR parenthetical | `(Blocker/Major only), derived from the cited story's profile.**` | 73 | -| F11 | 8b, the Gate-A filter clause | `intent + artifact text + which invariants it touches. Ask for **every** finding` | 561 | -| F12 | 9a, the revalidation trigger | `and the named evidence but not the mode value**, and is **revalidated before every Gate-B` | 726 | -| F13 | 9b, the severity-deciding fallback | `decision that act takes differently if the text is wrong. Both are required. If you` | 788 | +| F11 | 8b, the Gate-A filter clause | `with severity and confidence — you filter to Blocker/Major downstream, Codex` | 562 | +| F12 | 9a, the revalidation trigger | `before the cycle-closing amend` | 727 | +| F13 | 9b, the severity-deciding fallback | `cannot name both, the finding is Minor or below: collect, never iterate.` | 789 | | F14 | 9, the revalidation remedy | `profile sits still. If revalidation changes the entry, the clean pass no longer covers what` | 728 | **W line numbers are deliberately not carried for F1–F14.** Each fragment was verified unique in W as well as C, but the template's numbers drift with every earlier task and a stale number here would read as source drift. Locate each in W by the fragment. @@ -117,7 +117,7 @@ whose count must stay **one** has to be **present** in it, unchanged. Those are | The row's class | Its fragment must be | Because its expected result is | |---|---|---| -| **replaced** (OLD half), **dropped**, **moved** (source) | **absent** from the replacement block | `worktree=0` — a surviving fragment can never reach zero | +| **replaced** (OLD half), **dropped**, **moved** (source) | **wholly inside the live text the edit removes** — which implies, but is stronger than, absent from the replacement block | `worktree=0` — a fragment the edit does not reach survives whatever the block says | | **carried** | **present, unchanged**, in the replacement block | `worktree=1` — the block is what preserves it, so a fragment absent from the block cannot be found afterwards | | **kept** sharing a line with changed text | **outside** every replacement block's extent | `worktree=1` — it is not reproduced by any block; it survives because nothing replaces it | @@ -125,6 +125,15 @@ whose count must stay **one** has to be **present** in it, unchanged. Those are so a task either stops before installing or quietly drops its carried coverage. The table's cut widened at pass 9 to hold these rows; this test did not widen with it. +**And for the disappearing classes, "absent from the replacement block" is the wrong test — it is +necessary and not sufficient.** A fragment sitting in the unchanged text *before* an edit passes it +and then survives the install, because the edit never reaches it. **Three rows failed exactly that +way** — F11 ended just before item 8b's first changed word, F12 just before item 9a's, F13 just +before item 9b's — and each would have returned `old/worktree=1` after a **correct** installation, +so Task 8 could not have met its own required result. **Resolve the removed span in the live file +and require the fragment to lie wholly inside it.** The narrower test is the only one that +distinguishes "this wording goes" from "this wording is near something that goes". + **The third check compares against the target's fenced blocks with line breaks normalized.** `F4` is one line in `CLAUDE.md` but the block wraps it between `you` and `still`; a substring test finds nothing and the row looks usable. Pass 4 found it, and re-running the normalized check over every row then in the table found that one and no other. **That sweep predates row F7b**, which pass 6 added, and the snapshots the two tables cite predate it too. **Re-run all three checks over every row the tables now hold before Task 0 finishes**, and record the revision you ran them at — a row presented as covered by a sweep that could not have seen it is a claim about evidence that does not exist. *(No count is stated here on purpose: a number in this sentence goes stale the next time a task appends a row, which is the enumeration failure this plan keeps finding in itself.)* @@ -358,10 +367,14 @@ test "$(git rev-parse --abbrev-ref HEAD)" = loop-rule-consolidation || { echo "w test -z "$(git status --porcelain)" || { echo "tree not clean"; exit 1; } git merge-base --is-ancestor ba15e83 HEAD || { echo "approved target text (ba15e83) is not in this history"; exit 1; } # The three inputs must still be the versions that were approved, not merely present. +# Compare the RECORDED BASE, not HEAD: Task 15 step 7 permits a Gate-B fix to update +# the spec in a committed WIP snapshot, so on re-entry HEAD legitimately differs while +# the base does not. On a first run the two are the same commit. +REF=$( [ -s .context/loop-rule-base ] && cat .context/loop-rule-base || git rev-parse HEAD ) for f in target-text design condition-inventory; do p="docs/superpowers/specs/2026-09-10-loop-rule-consolidation-$f.md" - test "$(git rev-parse "HEAD:$p")" = "$(git rev-parse "ba15e83:$p")" \ - || { echo "$p differs from its approved version at ba15e83"; exit 1; } + test "$(git rev-parse "$REF:$p")" = "$(git rev-parse "ba15e83:$p")" \ + || { echo "$p differs at $REF from its approved version at ba15e83"; exit 1; } done if [ -s .context/loop-rule-base ]; then echo "base already recorded: $(cat .context/loop-rule-base) — NOT overwriting" @@ -379,7 +392,14 @@ every install and the final soft reset would run from the wrong starting tree. **The inputs are compared by blob, not merely found.** `ba15e83` being an ancestor says the approved commit is in this history; it does not say the three files still hold what was approved, and a later unapproved edit to the target text would supply different installation instructions to -every task. **`ba15e83` has exactly one role here — the approval commit** — and the loop compares each input's +every task. + +**They are compared at the recorded base, not at `HEAD`.** Task 15 step 7 permits a Gate-B fix that +changes specified behaviour to update the spec in the same commit, so a valid run can reach a state +where `HEAD`'s target text deliberately differs from `ba15e83`. Comparing `HEAD` would then reject +the re-entry path this plan advertises, on a state the plan itself creates. **The base is what the +tasks were derived against**, and it is the right subject; the later commits are validated +separately, as this run's `WIP:` commits. **`ba15e83` has exactly one role here — the approval commit** — and the loop compares each input's object id against its version there. (It is written more than once; the claim is about its role, not its occurrences.) @@ -448,9 +468,19 @@ base Written by the step that creates it, and it is what "keyed to `$BASE`" means — without it there is no comparison to make and a stale map from an abandoned run is indistinguishable from this one's. +**Both are written through a temporary file and renamed only after the last record is written**, so +a run interrupted mid-write leaves no half-file. A `base` line arrives first in the finished +artifact and would otherwise arrive first on disk too — making a partial map, with its spans +written and its `cond` rows missing, indistinguishable from a complete one to a re-entry that +compares only that line. + **So:** if `.context/loop-rule-untouched` and `.context/loop-rule-baseline-diff.txt` already exist, -**read their `base` line and require it to equal `.context/loop-rule-base`**; on a match keep them -and re-derive neither, on a mismatch or a missing line delete them and rebuild. If they are absent +**read their `base` line, require it to equal `.context/loop-rule-base`, and validate the file +itself**: every line parses as one of the declared record shapes, and every kept condition in the +five regions appears in exactly one `span` or one `cond` row — the same coverage assertion Task 0 +step 2 makes when it builds the map. **On any failure, delete and rebuild rather than reuse**; a +same-base partial file is the one shape the base line cannot catch. On a full match keep them and +re-derive neither. If they are absent while `$BASE` has `WIP:` commits after it, **derive both from the `$BASE` blobs** — `git show "$BASE:"` — not from the worktree. On a clean first run the two sources are the same thing, which is why this is stated once here rather than in each step. @@ -566,7 +596,14 @@ comparison actually emits. On a clean first run they are the same; once `$BASE..HEAD` is non-empty the worktree carries this plan's own edits, and comparing them would fold introduced drift into the inherited-drift record — after which Task 14 can no longer tell the two apart, which is the whole purpose of this baseline. -**Set `C_SRC` and `W_SRC` first:** + +**Selection, the `base` line and every extraction are ONE shell block**, because each fenced block +is its own invocation: `C_SRC` set in one block and consumed in the next is empty by the time +`sed` and `grep` see it, and a process substitution reading an empty filename need not make `diff` +fail — so the step would append its headings over two empty extracts and certify a parity baseline +it never computed. **This is the same defect as `$BASE` and `$BASEREF`**, and those two are +persisted to files for exactly this reason; here one block is simpler than a third scratch file. +The `test -r` guard makes an unreadable source fatal rather than silent. ```bash BASE=$(cat .context/loop-rule-base) @@ -577,13 +614,9 @@ if [ -n "$(git log --oneline "$BASE"..HEAD)" ]; then else C_SRC=CLAUDE.md; W_SRC=plugins/dev-workflow/commands/workflow-init.md fi -printf 'base\t%s\n' "$BASE" > .context/loop-rule-baseline-diff.txt -``` - -**The `base` line is written first and by this step**, not assumed: the re-entry rule compares it, -so a first run that omits it produces an artifact its own next run must reject. +test -r "$C_SRC" && test -r "$W_SRC" || { echo "baseline source unreadable"; exit 1; } +printf 'base\t%s\n' "$BASE" > .context/loop-rule-baseline-diff.tmp # renamed at the end -```bash # Tab-separated start and end anchors: the anchors contain colons, so a # colon delimiter splits '**Severity:**' at the wrong place and yields an # empty end. Every entry here spans two DIFFERENT anchors — the single-line @@ -603,13 +636,14 @@ while IFS=$(printf '\t') read -r s e; do echo "== $s" diff <(sed -n "/$s/,/$e/p" "$C_SRC") \ <(sed -n "/$s/,/$e/p" "$W_SRC") -done < .context/loop-rule-sites | tee -a .context/loop-rule-baseline-diff.txt +done < .context/loop-rule-sites | tee -a .context/loop-rule-baseline-diff.tmp # The squash-carry sentence is ONE line and must not go through the loop. -echo "== On squash-merge" | tee -a .context/loop-rule-baseline-diff.txt +echo "== On squash-merge" | tee -a .context/loop-rule-baseline-diff.tmp diff <(grep -F 'On squash-merge, copy every evidence entry' "$C_SRC") \ <(grep -F 'On squash-merge, copy every evidence entry' "$W_SRC") \ - | tee -a .context/loop-rule-baseline-diff.txt + | tee -a .context/loop-rule-baseline-diff.tmp +mv .context/loop-rule-baseline-diff.tmp .context/loop-rule-baseline-diff.txt ``` **The squash-carry site is extracted with `grep`, not as a range**, for the reason step 2 already From 31119856d0d6b98127eb5028938b5685404f8389 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 19:37:10 +0200 Subject: [PATCH 125/181] docs(plans): state the fragment test as the post-edit region, and re-check all F rows MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Follow-up to pass 20, correcting the rule that pass applied. "Wholly inside the live text the edit removes" was too strong: it rejects F5, F7, F8 and F10, which run from unchanged wording into changed wording and are perfectly good — overlapping the removal by one word already makes the whole fragment unfindable afterwards. The accurate subject is the region's POST-EDIT text: unchanged prefix + replacement block + unchanged suffix. F11, F12 and F13 failed against that and passed against the block alone, which is the whole reason pass 20 found them and four earlier checks did not. Re-checked every §F row against the post-edit region built from its item's declared live range: F4, F5, F6, F7, F7b, F8, F9, F10, F11, F12, F13 and F14 all go to zero. F1, F2 and F3 could NOT be checked this way — their §F notes give no line range, so the region cannot be built mechanically — and the plan now says so and assigns them a reading check before Task 8 installs, rather than presenting twelve verified rows as fifteen. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../2026-09-14-loop-rule-consolidation.md | 27 ++++++++++++------- 1 file changed, 18 insertions(+), 9 deletions(-) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index b0d9bd5..7a25239 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -117,7 +117,7 @@ whose count must stay **one** has to be **present** in it, unchanged. Those are | The row's class | Its fragment must be | Because its expected result is | |---|---|---| -| **replaced** (OLD half), **dropped**, **moved** (source) | **wholly inside the live text the edit removes** — which implies, but is stronger than, absent from the replacement block | `worktree=0` — a fragment the edit does not reach survives whatever the block says | +| **replaced** (OLD half), **dropped**, **moved** (source) | **absent from the region's post-edit text** — the unchanged prefix, plus the replacement block, plus the unchanged suffix — not merely from the block | `worktree=0` — a fragment the edit does not reach survives whatever the block says | | **carried** | **present, unchanged**, in the replacement block | `worktree=1` — the block is what preserves it, so a fragment absent from the block cannot be found afterwards | | **kept** sharing a line with changed text | **outside** every replacement block's extent | `worktree=1` — it is not reproduced by any block; it survives because nothing replaces it | @@ -125,14 +125,23 @@ whose count must stay **one** has to be **present** in it, unchanged. Those are so a task either stops before installing or quietly drops its carried coverage. The table's cut widened at pass 9 to hold these rows; this test did not widen with it. -**And for the disappearing classes, "absent from the replacement block" is the wrong test — it is -necessary and not sufficient.** A fragment sitting in the unchanged text *before* an edit passes it -and then survives the install, because the edit never reaches it. **Three rows failed exactly that -way** — F11 ended just before item 8b's first changed word, F12 just before item 9a's, F13 just -before item 9b's — and each would have returned `old/worktree=1` after a **correct** installation, -so Task 8 could not have met its own required result. **Resolve the removed span in the live file -and require the fragment to lie wholly inside it.** The narrower test is the only one that -distinguishes "this wording goes" from "this wording is near something that goes". +**And for the disappearing classes, "absent from the replacement block" is the wrong subject.** Most +items replace only part of their live region, so what stands afterwards is **the unchanged prefix, +plus the block, plus the unchanged suffix** — and a fragment sitting in that prefix is absent from +the block, passes the test, and survives the install because the edit never reaches it. **Three +rows failed exactly that way**: F11 ended just before item 8b's first changed word, F12 just before +item 9a's, F13 just before item 9b's, so each would have returned `old/worktree=1` after a +**correct** installation and Task 8 could not have met its own required result. + +**So compare against the post-edit region, not the block.** Build it — prefix, block, suffix — and +require the fragment to be gone from it. **The fragment does not have to lie wholly inside the +removed text**, which would be too strong and would reject rows like F7 that run from unchanged +wording into changed wording: overlapping the removal by one word is enough to make the whole +fragment unfindable afterwards, and that is what the count measures. + +**All twelve §F rows whose items declare a live range were re-checked this way** at the commit that +records this. **F1, F2 and F3 were not**: their §F notes give no line range, so the region cannot be +built mechanically — **check those three by reading, before Task 8 installs.** **The third check compares against the target's fenced blocks with line breaks normalized.** `F4` is one line in `CLAUDE.md` but the block wraps it between `you` and `still`; a substring test finds nothing and the row looks usable. Pass 4 found it, and re-running the normalized check over every row then in the table found that one and no other. From df1e9f5cb7f2175268dfe8a848b17e98acb8033e Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 19:37:27 +0200 Subject: [PATCH 126/181] docs(context): record passes 11-20 in the Gate-A plan working record Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../codex-reviews/gate-a-plan-om0bdd7udh-resume.md | 14 ++++++++++++-- 1 file changed, 12 insertions(+), 2 deletions(-) diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md index df5e851..d269016 100644 --- a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md @@ -13,7 +13,7 @@ Nothing depends on it; the pass files and the repo are authoritative where this ## Resume here -**Next action:** Gate-A plan pass 11 against `55a27c9`. The prompt is +**Next action:** Gate-A plan pass 21 against `3111985`. The prompt is `.context/gate-a-plan-prompt.md`; substitute `__SHA__` and `__P__`, precheck per its header, delete the target file, confirm it is gone. @@ -80,7 +80,17 @@ worth checking before a pass rather than after. | 8 | 54e793b | 13→**15** | 4→**2** | 4→**9** | yes | **MANDATORY TWO-TELL STOP** (findings rose, instrument cluster). Task 15 step 7 never re-ran the battery after a Gate-B fix; step 8 suppressed its record commit with `\|\| true`. Two more condition misclassifications: `a1` is carried inside §F item 8a's block **and the untouched span opened on that line**; `e10` is kept outside §D. `c10`–`c13` and `a18`–`a20` are moved and had only a reader walk; nine carried conditions owed preservation checks and three had them | | 9 | ebb371b | 15→**9** | 2→**2** | 9→**6** | yes | **MANDATORY TWO-TELL STOP** (Blockers flat, instrument cluster). The floor span ended on the line carrying both changed `a13` and kept `a14`. **Root repair: `## What each disposition owes, stated once`** — kept/carried/replaced/moved/dropped/add-only. Spans are now derived, not written out. The fragment table's cut became "exists before the edit" | | 10 | ba614f6 | 9→**17** | 2→**1** | 6→**14** | yes | **MANDATORY TWO-TELL STOP** (findings rose, instrument cluster). **Eleven of seventeen were one family: the tasks had not been brought into line with pass 9's table.** Root repair: **no task enumerates its condition ids** — `## How a task discharges that table` states the procedure and each task walks its passage's rows by class. A *kept* condition inside a wholly replaced passage has no span and owes a per-condition count; carried and kept preservation fragments are pre-existing and derived before the install | -| 11 | — | — | — | — | **not run — NEXT** | against `55a27c9`. Prompt: `.context/gate-a-plan-prompt.md`, substitute `__SHA__`=`55a27c9` and `__P__`=`11` | +| 11 | 55a27c9 | 17→**11** | 1→**2** | 14→**7** | yes | two-tell stop. **The unit became the source block a task replaces, not the passage** — (c) is split between Tasks 4 and 7, (a) between 7 and 8. Each row runs to **its own class's** result, so no step says "a pair for every row" (a carried fragment must still be there). Five record shapes in the evidence section | +| 12 | 1d5a892 | 11→**6** | 2→**1** | 7→**5** | yes | one tell. **`is_wip_commit` greps the whole command string**, so step 8's WIP record commit and the real closing commit in one block would classify the close as cycle-internal. Split into 8a/8b, separate invocations. Failure branch is a **mixed** reset, never `--hard` | +| 13 | c1614ba | 6→**4** | 1→**1** | 5→**1** | yes | two-tell stop. The span derivation collected each **fragment's** line rather than each replacement **block's** extent. 8a's retry branch compares `HEAD`'s exact changed-path set | +| 14 | 6a9cfc8 | 4→**5** | 1→**2** | 1→**2** | yes | three-tell stop. The dirty-set guard read only `??`/`A ` records, so a staged tracked change was swept in. Both 8b post-commit checks restore the tip. The Gate-B loop admits the **zero-finding pass below the floor** | +| 15 | c6d0773 | 5→**4** | 2→**0** | 2→**2** | yes | one tell. **"changed" is not a disposition** — nine §B conditions had no observation class at all. Every step-8 rejection restores a tip. `$BASEREF` resolved once | +| 16 | d41ff5a | 4→**5** | 0→**2** | 2→**2** | yes | three-tell stop. **"split" is not a disposition either.** The inventory defines `c9` as one clause and `a13` as **two sentences** — §H supplies one, so nothing observed the first. Every cross-block value goes to a file | +| 17 | 67a50e0 | 5→**9** | 2→**2** | 2→**3** | yes | three-tell stop. **The fragment admission test was class-blind**: it demanded every fragment be absent from its replacement, where a *carried* one must be present in it — no correct implementation could admit its own rows. `h3` is carried, not kept | +| 18 | 6d276fa | 9→**4** | 2→**0** | 3→**3** | yes | one tell. A **fifth record shape, `span`**; both scratch artifacts open with a `base` line; the hook ships **ten** prompt bodies, not seven | +| 19 | 76cd2ce | 4→**3** | 0→**0** | 3→**2** | yes | one tell. Nine editing tasks record into the plan and staged only prompts — the staging rule was a stale task-number list. Task 0 step 3 selects its source by the re-entry rule | +| 20 | 9aa1178 | 3→**4** | 0→**2** | 2→**2** | yes | three-tell stop. **F11/F12/F13 each ended just before their item's first changed word** and would have survived a correct install. Root: the test's subject is the region's **post-edit text**, not the replacement block. All twelve F rows with a declared range re-checked and go to zero; F1–F3 have no range and owe a reading check | +| 21 | — | — | — | — | **not run — NEXT** | against `3111985`. Prompt: `.context/gate-a-plan-prompt.md`, substitute `__SHA__`=`3111985` and `__P__`=`21` | ## Pass-1 report From c51dce40ebe6447363b66cfa1669fceaa1d5012e Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Mon, 14 Sep 2026 19:47:28 +0200 Subject: [PATCH 127/181] docs(context): record Gate-A plan pass 21 findings (2 Blockers, 3 Majors) Not yet applied. Four of the five are repairs-of-repairs on Task 0's baseline extraction and Task 15's closing act, which raises a structural question that is the user's to answer. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-21.md | 6 ++++++ 1 file changed, 6 insertions(+) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-21.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-21.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-21.md new file mode 100644 index 0000000..ba8b845 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-21.md @@ -0,0 +1,6 @@ +BLOCKER | high | Task 15 steps 6–8a | The final review's resolved headSha is never persisted or compared with HEAD before 8a; the first-run branch sets PRE from whatever HEAD exists after the review, and the retry branch checks only the latest commit's changed-path set. | A commit landing after the clean response can become PRE, receive the findings-only commit on top, pass both changed-path checks, and be squashed into the close without any Gate-B pass reviewing it; an arbitrary findings-only HEAD can likewise hide an unreviewed parent on the retry route. | Persist the candidate pass's exact reviewed head before issuing it, require HEAD to equal that value before either 8a branch proceeds, and on retry require the findings-only commit's parent to equal that reviewed head before recording TIP. +BLOCKER | high | Task 15 step 7 | The loop commits only a fix before the next review and gives no transition for a non-closing pass that owes no repair, such as a Minor-only clean pass below the floor or an answered suspension that changes no artifact. | That pass's tracked findings files remain dirty; later pass files accumulate beside them, and 8a's exact dirty-set guard can never equal only the final pass's two files, so a valid route the installed ordering requires cannot close through this plan. | After every non-closing pass, commit that pass's findings files and refreshed records in a WIP snapshot even when no artifact repair exists, resolve the new head, and issue the next pass only against that exact head. +MAJOR | high | Task 0 re-entry rule | The rule requires both scratch artifacts to have every line parse as a declared record shape, but loop-rule-baseline-diff.txt is written as a base line followed by free-form headings and raw diff output and has no declared schema or completeness predicate. | A literal validation rejects a correctly generated baseline forever, while an executor who applies the declared base/span/cond shapes only to loop-rule-untouched still has no way to reject a same-base partial baseline; partial re-entry is either blocked or trusts incomplete parity evidence. | Give the baseline artifact its own parseable site/result schema and completeness check, or state and implement a separate validation that proves every expected baseline site was extracted and recorded before reuse. +MAJOR | high | Task 0 step 3 | The baseline extraction interpolates its anchors directly into sed ranges and never checks literal uniqueness or non-empty extracts; the while loop is piped to tee, so failures inside it also do not reliably prevent the final rename. | A missing, duplicated, or regex-active anchor can yield two empty extracts that compare equal, or a partial temporary result can be promoted as complete, causing Task 14 to treat introduced drift as inherited divergence. | Apply the same literal-count, escaped-anchor, and non-empty-extraction checks Task 14 uses to every baseline site, propagate loop failures explicitly, and rename the temporary file only after every extraction and the single-line site succeed. +MAJOR | high | Verification fragments paragraph before Task 0 | The plan orders a full three-condition recheck of every current fragment row and says to record the revision, but names no output location; Task 0 declares no modified files and ends with "Commit nothing." | The only sweep that includes F7b and the execution-time reading checks for F1–F3 leave no durable result, so an interruption or the closing evidence cannot distinguish a completed validation from one that never ran. | Add a keyed Task 0 fragment-sweep record with its subject, base revision, per-row result, and the F1–F3 reading result, then commit that record as the first WIP snapshot before later tasks consume the rows. +END OF FINDINGS (5 total) From 1d2f9dbaeec1a8923799681c6f479286708199d5 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Tue, 15 Sep 2026 10:42:06 +0200 Subject: [PATCH 128/181] docs(plans): replace Task 0 and Task 15's guarded shell with four procedures MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Method change, approved by Daniel after pass 21. NOT a reduction of what the close owes: the accounting table lists every one of the 41 existing guard conditions as kept, re-expressed, or proposed for deletion, per the AGENTS.md rule that exists because pass 5's bulk edit removed Task 15's closing block by accident. One condition is dropped and only one: "Task 0 commits nothing". It now commits the fragment sweep, because a validation whose result is not committed cannot be told from one that never ran. Six are corrected or extended, each against a pass-21 finding: - The reviewed HEAD and the restore point are two values. The first is what the pass read; the second is that plus the findings files. Treating them as one is how a commit landing after the response reaches the close. Every Gate-B call now records the head it is issued against, and 8a compares HEAD to it and the record commit's parent to it. - Every non-closing pass commits, repair or not. A Minor-only clean pass below the floor left its findings files dirty, after which the dirty-set check could never equal one pass's two files and a route the installed ordering requires could not close through this plan. - The failure procedure inspects HEAD, the index AND the working tree, and preserves and reports the delta. Restoring HEAD is not the whole duty and the approved §A says so; a HEAD-only rule never did. - The baseline artifact gains a `site` record per inventoried site, so the resume procedure has a completeness predicate for it and not only for the map. - The baseline extraction gets the literal-uniqueness, escaped-anchor and non-empty checks Task 14 already has, and is no longer piped to tee, so a partial temporary file cannot be promoted by the rename. - The fragment re-check has an output section and is committed. What actually changes in shape: one shell procedure no longer tries to transact every re-entry automatically. The four procedures — preparation, close, failure, resume — each say what is checked, how success is recognised, and what happens on deviation, and the executing agent reads the state and picks the operation. Concrete shell stays where it is simple and already verified. Two corrections to my own report, from Daniel: Task 0 and 15 are not the only tasks with shell — Task 14 keeps a substantial check block, untouched here. And another defect in a guard does not prove the guard was unnecessary; each closed a real gap, and pass 21's findings are real gaps too. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../2026-09-14-loop-rule-consolidation.md | 631 ++++++++++-------- 1 file changed, 343 insertions(+), 288 deletions(-) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index 7a25239..e4e99f1 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -181,6 +181,190 @@ the closing commit cannot carry. --- +## The four procedures — and what happened to every condition they replace + +**Why this section exists.** Tasks 0 and 15 accumulated guarded shell across ten review passes, and +by pass 21 thirty-nine of the previous fifty findings were in those two tasks, four of the last five +being repairs to guards the previous passes had added. **The reading is not that the guards were +unnecessary** — each closed a real execution gap, and pass 21's own findings are real gaps too. The +reading is that **one shell procedure was being asked to transact every re-entry automatically**, +and that is the thing being dropped. The obligations are not. + +**What replaces it:** four short procedures — **preparation, close, failure, resume** — each saying +only *what is checked*, *how success is recognised*, and *what happens on deviation*. The executing +agent observes the actual state and derives the operation. **Concrete shell stays where it is +simple and already verified**; what goes is the branching that tried to handle every state in one +block. + +### The accounting — every condition, kept / re-expressed / proposed for deletion + +Required before replacing a decision procedure (`AGENTS.md`, "Never replace a decision procedure +without accounting for its old conditions"). Pass 5 is why: a bulk edit that removed the +verification apparatus also removed Task 15's closing block, and nothing noticed until an +independent reader did. + +| # | Condition | Disposition | +|---|---|---| +| 1 | Branch is `loop-rule-consolidation` | **kept**, same shell — preparation | +| 2 | Tree clean before the base is recorded | **kept**, same shell — preparation | +| 3 | `ba15e83` is an ancestor of `HEAD` | **kept**, same shell — preparation | +| 4 | The three approved inputs' blobs equal their `ba15e83` versions, compared **at the recorded base** | **kept**, same shell — preparation | +| 5 | Never overwrite an existing base file | **kept** as an obligation; the shell branch becomes one line of the resume procedure | +| 6 | A pre-existing base is an ancestor of `HEAD` (`merge-base --is-ancestor`) | **kept**, same shell — resume | +| 7 | Only this run's `WIP:` commits lie between base and `HEAD` | **kept** as an obligation — resume; the reading is the agent's | +| 8 | `$BASE` persisted to a file; every consumer guards it non-empty | **kept**, same shell | +| 9 | The five regions are located by anchor, one hit per pattern per file | **kept**, same shell — preparation | +| 10 | Spans are **derived** from replacement extents, never hand-written | **kept** — the rule, unchanged | +| 11 | Every kept condition lands in exactly one `span` or one `cond` row | **kept** — the coverage assertion, unchanged | +| 12 | Anchors, never absolute line numbers; resolved separately per tree | **kept** — the rule, unchanged | +| 13 | A single-line site uses `grep`, never `sed -n "/x/,/x/p"` | **kept** — the rule, unchanged | +| 14 | Records are tab-separated `base` / `span` / `cond` shapes, `base` first | **kept**, and **extended**: the baseline artifact gains its own `site` shape (pass 21) | +| 15 | Artifacts written through a temp file, renamed after the last record | **kept**, same shell | +| 16 | On re-entry: base line matches, shapes parse, coverage complete — else rebuild | **re-expressed** as the resume procedure, and **corrected**: the completeness predicate now exists for both artifacts rather than only the map (pass 21) | +| 17 | The baseline source is `$BASE` blobs once `WIP:` commits exist | **kept**, same shell | +| 18 | Both baseline sources readable | **kept**, same shell | +| 19 | The baseline covers every inventoried site and is read whole | **kept**, and **strengthened**: per-site literal-uniqueness, escaped anchors and non-empty extraction, as Task 14 already requires (pass 21) | +| 20 | Task 0 commits nothing | **proposed for deletion, and deleted.** It now commits one record — the fragment sweep (pass 21). A validation whose result is not committed cannot be told from one that never ran, and every other reader check in this plan already commits its record | +| 21 | `baseSha` is `$BASE`, never `HEAD^` | **kept** — the rule, unchanged | +| 22 | `headSha` resolved to the full object name at that moment, kept with each branch's result | **kept**, and **extended**: it is persisted, because 8a needs it (pass 21) | +| 23 | A fresh nonce for the Gate-B cycle | **kept** — the rule, unchanged | +| 24 | Each fix is committed before the next review | **kept**, and **widened**: every non-closing pass commits, repair or not (pass 21) | +| 25 | Every affected check re-runs after each fix, mechanical and reader | **kept** — the rule, unchanged | +| 26 | Re-run records are committed before the re-review | **kept** — the rule, unchanged | +| 27 | The complete set re-runs before the candidate final pass | **kept** — the rule, unchanged | +| 28 | Only a clean response against that exact `HEAD` closes | **kept**, and **made checkable**: the reviewed head is recorded, so "that exact `HEAD`" has a value to compare against (pass 21) | +| 29 | Closure-eligible = clean at or above the floor **or** zero-finding | **kept** — the rule, unchanged | +| 30 | The final pass's own findings files are the sole permitted post-review addition | **kept** — the rule, unchanged | +| 31 | The closing message is rebuilt whole; exactly one provenance line and one curve | **kept** — the rule, unchanged | +| 32 | Every owed record is present before the close | **kept** — the rule, unchanged | +| 33 | The dirty set is exactly the final pass's findings files, read from every porcelain record | **kept** as the close procedure's check | +| 34 | The record commit's changed-path set equals those files | **kept** as the close procedure's check | +| 35 | The tree is clean after the record commit | **kept** as the close procedure's check | +| 36 | `HEAD` equals the reviewed tip before the reset | **kept** as the close procedure's check, and **corrected**: the reviewed *head* and the restore *point* are two values, not one (pass 21) | +| 37 | The closing commit's subject is not a snapshot | **kept** as the close procedure's check | +| 38 | The tree is clean after the close | **kept** as the close procedure's check | +| 39 | 8a and 8b are separate invocations; no `-m` in the closing one | **kept**, and it is why the close procedure names two invocations | +| 40 | Every rejection restores a tip by **mixed** reset, never `--hard` | **re-expressed** as the failure procedure, and **corrected**: restoring `HEAD` is not the whole duty — index and working tree are inspected, and the side effects are preserved and reported, which the approved §A requires and a `HEAD`-only rule never said | +| 41 | The scratch files are removed only after a successful close | **kept** as the close procedure's last step | + +**Nothing in that table is dropped except condition 20, and that one is dropped by being replaced +with a stronger obligation.** Six conditions are corrected or extended, each against a pass-21 +finding. The rest keep their force; what changes is that a person reads the state and picks the +operation, instead of one block trying to branch through every state in advance. + +### Preparation — before any task edits a file + +**What is checked.** The branch; a clean tree; that `ba15e83` is an ancestor; that the three +approved inputs still hold their approved blobs **at the recorded base**; that the recorded base, if +one exists, is an ancestor of `HEAD` with only this run's `WIP:` commits after it. + +```bash +test "$(git rev-parse --abbrev-ref HEAD)" = loop-rule-consolidation || { echo "wrong branch"; exit 1; } +test -z "$(git status --porcelain)" || { echo "tree not clean"; exit 1; } +git merge-base --is-ancestor ba15e83 HEAD || { echo "ba15e83 not in this history"; exit 1; } +REF=$( [ -s .context/loop-rule-base ] && cat .context/loop-rule-base || git rev-parse HEAD ) +for f in target-text design condition-inventory; do + p="docs/superpowers/specs/2026-09-10-loop-rule-consolidation-$f.md" + test "$(git rev-parse "$REF:$p")" = "$(git rev-parse "ba15e83:$p")" \ + || { echo "$p differs at $REF from its approved version"; exit 1; } +done +``` + +**How success is recognised.** Every line above exits 0, and `.context/loop-rule-base` holds a +40-character object name that is an ancestor of `HEAD`. + +**On deviation.** Stop. Each of these has a different fix and none of them is "retry": a wrong +branch is a checkout, a dirty tree is a decision about uncommitted work, a differing input blob is +an unapproved edit to a spec that the gates already closed. + +### Close — after a closure-eligible pass + +**What is checked**, in this order, and these are the conditions rather than a script: + +1. **`HEAD` equals the head the candidate pass was issued against** — the value recorded before that + call, in `.context/loop-rule-reviewed-head`. **That is the reviewed head.** It is not the same + value as the restore point below, and conflating them is how an unreviewed commit reaches the + close. +2. **The only thing dirty is the candidate pass's own findings files** — read from every porcelain + record, not from two status codes. Nothing else may be uncommitted: every record this plan + collects was committed before the pass was issued. +3. **Those files are committed** in their own invocation, and that commit **changes exactly those + paths**. Its parent is the reviewed head. Record the resulting commit as the **restore point**, + in `.context/loop-rule-reviewed-tip`. +4. **The tree is clean**, and `HEAD` is still the restore point, when the closing invocation begins. +5. **The closing message is complete** — rebuilt whole, exactly one provenance line, exactly one + curve, every owed evidence entry, and either the applicable human-exception records or + `Human exceptions: none`. +6. **After the closing commit**: its subject is not a snapshot, and the tree is clean. + +**How success is recognised.** A single commit at `HEAD` whose subject is the real message, whose +tree equals the restore point's tree, and whose parent is `$BASE`. Then, and only then, the scratch +files are removed. + +**Two invocations, not one.** `codex-gate.sh`'s `is_wip_commit` +(`plugins/dev-workflow/hooks/codex-gate.sh:763`) tests the **whole command string** against +`-m[[:space:]]*['"]?[[:space:]]*wip`. The findings-file commit carries that pattern; the closing +commit must not share a command string with it, or the hook reads the close as cycle-internal and +carries this cycle's count and fingerprint into the next. + +**On deviation at any of the six.** Stop **before** moving `HEAD`. Every one of these is checkable +while the repository is still in a state the plan understands, and that is the whole reason they +come before the reset. + +### Failure — the closing act did not complete + +**The approved §A requires the state a failed closing act leaves to be inspected**, and that is more +than `HEAD`. **Look at three things and say what each holds:** + +- **`HEAD`** — did the commit land? A failed `git commit` leaves it where it was; a hook that + rejected after committing does not. +- **The index** — a commit hook can stage content, and `reset --hard` would destroy exactly that. +- **The working tree** — the same hook can modify files without staging them. + +**How success is recognised.** The repository is back at the restore point **with the attempt's +delta preserved and reported** — `git status` names it — and the recovery base and reviewed-head +files still exist. **`git reset --mixed ` is the operation that does this**; `--hard` +is never it. + +**On deviation.** If the delta cannot be explained, stop and hand it to a person. **Do not retry a +close against a state you cannot account for** — the second attempt would carry whatever the first +one left. + +**What this procedure does not promise.** It does not transact the retry. After the state is +understood, the executor re-establishes every closure condition, obtains a fresh clean response +against the new head, and re-enters the close procedure from its first check. **That is a decision +sequence, not a rollback**, and trying to automate it is what grew the block this section replaces. + +### Resume — re-entering after an interruption + +**What is checked.** Whether the recorded base is this run's; whether the scratch artifacts belong +to it and are complete; and how far the implementation got. + +- **The base**, by the preparation procedure's checks. A base that is not an ancestor, or has + non-`WIP:` commits after it, is **stale** — delete it deliberately, record why, and re-record from + the true starting commit. **Never overwrite one you did not just write.** +- **Each scratch artifact**, by its own `base` line **and** its own completeness predicate: + - `.context/loop-rule-untouched` — every line parses as `base`, `span` or `cond`, and every kept + condition in the five regions appears in exactly one `span` or one `cond`. + - `.context/loop-rule-baseline-diff.txt` — every line parses as `base`, a `site` record, or diff + output belonging to the site above it, and **every inventoried site has a `site` record**. +- **The implementation**, by reading the `WIP:` commits between the base and `HEAD` and the plan's + own task checkboxes. + +**How success is recognised.** A base that passes preparation, artifacts whose `base` line matches +it and whose coverage is complete, and a task list whose ticked entries match the commits present. + +**On deviation.** An artifact that fails either test is **deleted and rebuilt from the `$BASE` +blobs** — never reused, and never repaired in place. A same-base partial file is the one shape a +`base` line alone cannot catch, which is why the completeness predicate exists. + +**Steps that describe the tree at `$BASE` are validated on re-entry, not re-run against the +worktree.** After a text task the worktree carries this plan's own edits, and rebuilding the +baseline from it would fold introduced drift into the inherited-drift record — the one distinction +Task 14 depends on. + +--- + ## File Structure | File | Responsibility in this change | @@ -369,85 +553,40 @@ preservation fragments that an earlier draft chose afterwards. **Interfaces:** - Produces: `$BASE` (the parent commit every later task's counterfactual half runs against), and the five untouched passage ranges recorded as **anchor spans plus a per-condition fragment list** — never as absolute line numbers, for the reason step 2 gives. -- [ ] **Step 1: Confirm the approved artifacts and a clean tree** +- [ ] **Step 1: Run the preparation procedure, then record the base** + +**`## The four procedures` · Preparation** holds the checks and the shell: branch, clean tree, +`ba15e83` an ancestor, and the three approved inputs compared by blob **at the recorded base** +rather than at `HEAD` — Task 15 step 7 permits a Gate-B fix to update the spec in a `WIP:` snapshot, +so on a re-entry `HEAD`'s target text legitimately differs while the base's does not. + +Then record the base, and **never overwrite one you did not just write**: ```bash -test "$(git rev-parse --abbrev-ref HEAD)" = loop-rule-consolidation || { echo "wrong branch"; exit 1; } -test -z "$(git status --porcelain)" || { echo "tree not clean"; exit 1; } -git merge-base --is-ancestor ba15e83 HEAD || { echo "approved target text (ba15e83) is not in this history"; exit 1; } -# The three inputs must still be the versions that were approved, not merely present. -# Compare the RECORDED BASE, not HEAD: Task 15 step 7 permits a Gate-B fix to update -# the spec in a committed WIP snapshot, so on re-entry HEAD legitimately differs while -# the base does not. On a first run the two are the same commit. -REF=$( [ -s .context/loop-rule-base ] && cat .context/loop-rule-base || git rev-parse HEAD ) -for f in target-text design condition-inventory; do - p="docs/superpowers/specs/2026-09-10-loop-rule-consolidation-$f.md" - test "$(git rev-parse "$REF:$p")" = "$(git rev-parse "ba15e83:$p")" \ - || { echo "$p differs at $REF from its approved version at ba15e83"; exit 1; } -done if [ -s .context/loop-rule-base ]; then - echo "base already recorded: $(cat .context/loop-rule-base) — NOT overwriting" + echo "base already recorded: $(cat .context/loop-rule-base) — validating, not overwriting" else git rev-parse HEAD > .context/loop-rule-base fi -cat .context/loop-rule-base -``` - -**Every line above fails the script; none of them prints for a human to notice.** An earlier draft -printed `git rev-parse --abbrev-ref HEAD` and `git status --porcelain` under comments saying what -to expect, which asserts nothing: a wrong-branch or dirty checkout would be recorded as `$BASE` and -every install and the final soft reset would run from the wrong starting tree. - -**The inputs are compared by blob, not merely found.** `ba15e83` being an ancestor says the -approved commit is in this history; it does not say the three files still hold what was approved, -and a later unapproved edit to the target text would supply different installation instructions to -every task. - -**They are compared at the recorded base, not at `HEAD`.** Task 15 step 7 permits a Gate-B fix that -changes specified behaviour to update the spec in the same commit, so a valid run can reach a state -where `HEAD`'s target text deliberately differs from `ba15e83`. Comparing `HEAD` would then reject -the re-entry path this plan advertises, on a state the plan itself creates. **The base is what the -tasks were derived against**, and it is the right subject; the later commits are validated -separately, as this run's `WIP:` commits. **`ba15e83` has exactly one role here — the approval commit** — and the loop compares each input's -object id against its version there. (It is written more than once; the claim is about its role, -not its occurrences.) - -**What this cannot establish, stated rather than implied:** that *this plan revision* is the one -Gate A closed on. A plan cannot name its own closing commit, and any sha written here would be from -before that commit exists. **The Gate-A findings files in `.context/codex-reviews/` are the record -of which revision was reviewed**; a reader checks the plan against them, and nothing mechanical -here does it. - -**Never overwrite an existing base, and never trust one you did not just write.** Re-running Task 0 -after a partial implementation would record the current WIP tip, and both Gate B's range and the -final `reset --soft` would then start after every edit made so far — prompt and hook changes would be -squashed into the closing commit without ever entering a review range. - -**A pre-existing value is not accepted on being non-empty.** It is valid only if it is an ancestor of -`HEAD` **and** every commit between it and `HEAD` is a `WIP:` commit of this execution. **Both -halves are checked, and the ancestry one first:** - -```bash -B=$(cat .context/loop-rule-base) -git merge-base --is-ancestor "$B" HEAD || { echo "recorded base is NOT an ancestor of HEAD — stale or from another branch"; exit 1; } -git log --oneline "$B"..HEAD +BASE=$(cat .context/loop-rule-base) +test -n "$BASE" || { echo "BASE empty"; exit 1; } +git merge-base --is-ancestor "$BASE" HEAD || { echo "recorded base is NOT an ancestor of HEAD — stale"; exit 1; } +git log --oneline "$BASE"..HEAD # expect nothing, or only this run's WIP: commits ``` -**`git log "$B"..HEAD` does not test ancestry**, and reading it as if it did is how a base from an -abandoned branch passes: the range then lists what `HEAD` has and `$B` does not, which can be only -`WIP:` commits while `$B` sits on a branch of its own. `git reset --soft` onto it at Task 15 step 8 -would move `HEAD` to that unrelated commit and drop every real commit since the fork. `merge-base ---is-ancestor` is the test the sentence above actually names. +**Re-running Task 0 after a partial implementation must not re-record the base.** It would capture +the current WIP tip, and both Gate B's range and the final reset would then start *after* every edit +made so far — prompt and hook changes squashed into the closing commit without entering a review +range. **A pre-existing value is validated, never trusted**: `git log "$BASE"..HEAD` does not test +ancestry, which is how a base from an abandoned branch passes; `merge-base --is-ancestor` is the +test the sentence names. Anything other than this run's `WIP:` commits in that log means the file is +stale — delete it deliberately, record why, and re-record from the true starting commit. **`## The +four procedures` · Resume** is what to do next in that case. -Expected: the ancestry check exits 0, and the log shows nothing, or only `WIP:` commits of this run. **Anything else means the file is stale** — -left by an abandoned run, or by one whose work was already squashed. Delete it deliberately, record -why, and re-record from the true starting commit. **Task 15 step 8 removes the file after the -closing commit**, so a stale one is an abandoned run rather than a normal state. - -**Persist it to a file, not to a shell variable.** Each task runs in its own shell invocation, so a -`BASE=` assignment in Task 0 is gone by Task 1 and every parent-tree count would run against an -empty revision — which fails loudly in `git show` but quietly in a `grep -c` pipeline. Every later -task that counts anything begins by reading it back and refusing an empty value: +**Persist it to a file, not to a shell variable.** Each fenced block runs in its own shell +invocation, so a `BASE=` assignment here is gone by the next task and every parent-tree count would +run against an empty revision — which fails loudly in `git show` but quietly in a `grep -c` +pipeline. Every later task that counts anything reads it back and refuses an empty value: ```bash BASE=$(cat .context/loop-rule-base) @@ -457,42 +596,9 @@ test -n "$BASE" || { echo "BASE empty — Task 0 did not run"; exit 1; } **The guard is not decoration.** With `$BASE` empty, `git show ":$f"` reads the *index* rather than the recorded parent, and every parent count then describes the wrong tree without erroring. -**This commit is also `baseSha` for Gate B.** It is the parent of the first WIP snapshot, and it is -the only value that puts the whole implementation inside the reviewed range. `.context/` is ignored -by the hook's fingerprint, so the file itself moves nothing. - -**On re-entry, steps 2 and 3 are not re-run — they are validated.** Step 1 admits a recorded base -followed only by this run's `WIP:` commits, which means text tasks may already have run. **Steps 2 -and 3 describe the tree as it was at `$BASE`**: after a text task, the OLD extents they map are -gone and the aligned divergences no longer match the baseline expectation, so re-deriving them from -the worktree either overwrites the evidence with a map of the partly edited tree or fails on -differences this plan itself installed. - -**Each artifact carries the base it was built from, as its first line:** - -``` -base -``` - -Written by the step that creates it, and it is what "keyed to `$BASE`" means — without it there is -no comparison to make and a stale map from an abandoned run is indistinguishable from this one's. - -**Both are written through a temporary file and renamed only after the last record is written**, so -a run interrupted mid-write leaves no half-file. A `base` line arrives first in the finished -artifact and would otherwise arrive first on disk too — making a partial map, with its spans -written and its `cond` rows missing, indistinguishable from a complete one to a re-entry that -compares only that line. - -**So:** if `.context/loop-rule-untouched` and `.context/loop-rule-baseline-diff.txt` already exist, -**read their `base` line, require it to equal `.context/loop-rule-base`, and validate the file -itself**: every line parses as one of the declared record shapes, and every kept condition in the -five regions appears in exactly one `span` or one `cond` row — the same coverage assertion Task 0 -step 2 makes when it builds the map. **On any failure, delete and rebuild rather than reuse**; a -same-base partial file is the one shape the base line cannot catch. On a full match keep them and -re-derive neither. If they are absent -while `$BASE` has `WIP:` commits after it, **derive both from the `$BASE` blobs** — -`git show "$BASE:"` — not from the worktree. On a clean first run the two sources are the -same thing, which is why this is stated once here rather than in each step. +**This commit is also `baseSha` for Gate B.** It is the parent of the first WIP snapshot, and the +only value that puts the whole implementation inside the reviewed range. `.context/` is ignored by +the hook's fingerprint, so the file itself moves nothing. - [ ] **Step 2: Re-read the five untouched ranges and record their current anchors** @@ -614,6 +720,17 @@ it never computed. **This is the same defect as `$BASE` and `$BASEREF`**, and th persisted to files for exactly this reason; here one block is simpler than a third scratch file. The `test -r` guard makes an unreadable source fatal rather than silent. +**The artifact has a schema, because the resume procedure validates it.** One `base` line, then one +`site` record per inventoried site followed by that site's diff output — +**every site gets a record, including one that compared equal**, which is what makes a site missing +from the run distinguishable from a site that matched. Without it the resume procedure has a +completeness predicate for the map and none for this file, and a same-base partial baseline reads as +complete. + +**And the loop is not piped to `tee`.** A failure inside a pipeline does not reliably stop the +script, so a partial temporary file could be promoted by the `mv`. Each site appends directly and +exits on failure; the rename is the last statement. + ```bash BASE=$(cat .context/loop-rule-base) if [ -n "$(git log --oneline "$BASE"..HEAD)" ]; then @@ -641,18 +758,30 @@ printf '%b\n' \ 'Recording a human exception\tbecause writing it down makes it sound' \ 'When these rules bind\tDownstream has no shipping commit' \ > .context/loop-rule-sites +# Each site: the same checks Task 14 applies — literal uniqueness per copy, an +# ESCAPED anchor (a derived anchor carries ** and / and . and is a regex to sed), +# and a non-empty extract. Two empty extracts diff equal and would certify a +# site sed never found. Failures exit; the rename happens only at the end. while IFS=$(printf '\t') read -r s e; do - echo "== $s" - diff <(sed -n "/$s/,/$e/p" "$C_SRC") \ - <(sed -n "/$s/,/$e/p" "$W_SRC") -done < .context/loop-rule-sites | tee -a .context/loop-rule-baseline-diff.tmp + for f in "$C_SRC" "$W_SRC"; do + test "$(grep -cF "$s" "$f")" = 1 || { echo "start anchor not unique in $f: $s"; exit 1; } + test "$(grep -cF "$e" "$f")" = 1 || { echo "end anchor not unique in $f: $e"; exit 1; } + done + se=$(printf '%s' "$s" | sed 's/[][\.*^$\/]/\\&/g') + ee=$(printf '%s' "$e" | sed 's/[][\.*^$\/]/\\&/g') + a=$(sed -n "/$se/,/$ee/p" "$C_SRC"); b=$(sed -n "/$se/,/$ee/p" "$W_SRC") + test -n "$a" && test -n "$b" || { echo "empty extraction for: $s"; exit 1; } + printf 'site\t%s\t%s\n' "$s" "$e" >> .context/loop-rule-baseline-diff.tmp + diff <(printf '%s\n' "$a") <(printf '%s\n' "$b") >> .context/loop-rule-baseline-diff.tmp +done < .context/loop-rule-sites || exit 1 # The squash-carry sentence is ONE line and must not go through the loop. -echo "== On squash-merge" | tee -a .context/loop-rule-baseline-diff.tmp +printf 'site\t%s\t%s\n' 'On squash-merge' 'On squash-merge' >> .context/loop-rule-baseline-diff.tmp diff <(grep -F 'On squash-merge, copy every evidence entry' "$C_SRC") \ <(grep -F 'On squash-merge, copy every evidence entry' "$W_SRC") \ - | tee -a .context/loop-rule-baseline-diff.tmp + >> .context/loop-rule-baseline-diff.tmp mv .context/loop-rule-baseline-diff.tmp .context/loop-rule-baseline-diff.txt +cat .context/loop-rule-baseline-diff.txt # READ IT WHOLE ``` **The squash-carry site is extracted with `grep`, not as a range**, for the reason step 2 already @@ -668,9 +797,29 @@ Expected: the deliberate divergences the inventory records — `b3`'s cross-refe record it in the file and raise it before editing, because Task 14's parity diff cannot tell drift you introduced from drift you inherited. -- [ ] **Step 4: Commit nothing** +- [ ] **Step 4: Re-check every fragment row, record the result, and commit it** -Task 0 produces a scratch note, not a commit. +**This is the sweep the fragment-table section orders, and it had no home.** Run all three +conditions over every row the tables now hold — single-line, unique in the copies the row claims, +and the class-specific relationship to its own replacement — and **write the result under +`## Fragment sweep (Task 0 output)`** at the end of this plan: the revision it ran at, one line per +row, and the reading result for `F1`, `F2` and `F3`, whose §F notes give no line range so their +region cannot be built mechanically. + +```bash +git add docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +git commit -m "WIP: fragment sweep at the recorded base" +``` + +**Task 0 used to commit nothing, and that is the one condition this change drops.** A validation +whose result is not committed cannot be told from one that never ran — the same rule Task 12b's +sweep and Task 15's prompt-standards result already follow — and every later task consumes these +rows on the strength of it. The commit is `WIP:` like every other snapshot in this cycle, so it is +inside Gate B's range and inside the final reset. + +**The scratch files stay in `.context/` and stay ignored.** `.gitignore` carries `.context/*`, so +`loop-rule-base`, `loop-rule-untouched`, `loop-rule-baseline-diff.txt` and the `.src` copies are +working state, not deliverables; only the sweep record is committed. --- @@ -2185,10 +2334,16 @@ an enumeration here** — an id list in this step was already stale once, naming - [ ] **Step 6: Run Gate B** ```bash -BASE=$(cat .context/loop-rule-base) # the parent of the FIRST WIP, from Task 0 -git rev-parse HEAD # the current WIP tip +BASE=$(cat .context/loop-rule-base) # the parent of the FIRST WIP, from Task 0 +git rev-parse HEAD > .context/loop-rule-reviewed-head # the head THIS call is issued against +cat .context/loop-rule-reviewed-head ``` +**Every Gate-B call records the head it is issued against, this first one included**, and step 7 +rewrites the file before each later call. It is what makes "a clean response against that exact +`HEAD`" a comparison rather than a claim: without it the close procedure has nothing to hold `HEAD` +up against, and a commit landing after the response is indistinguishable from the reviewed state. + **`baseSha` is `$BASE`, never `HEAD^`.** This plan makes a WIP commit per task, so `HEAD^` is the parent of the *last* one and the review range would hold the version bump alone — Gate B would close having reviewed none of the prompt or hook changes. `$BASE` is the parent of the first WIP @@ -2214,13 +2369,27 @@ branch file; one branch clean and the other not is not it. WIP tip while the repair sits in the worktree — and the final squash then publishes a fix no pass reviewed: -**Before committing each fix, re-run what the fix could have broken**, then: +**Every pass that does not close ends the same way, whether or not it produced a repair.** Commit +that pass's findings files and any refreshed records, resolve the new head, **record it**, and issue +the next pass against exactly that value: ```bash -git add -A && git commit -m "WIP: fix " -git rev-parse HEAD # resolve headSha fresh for the next call +# Before committing a fix, re-run what the fix could have broken. +git add -A && git commit -m "WIP: fix " # or: "WIP: pass records" where no repair was owed +git rev-parse HEAD > .context/loop-rule-reviewed-head # the head the NEXT call is issued against +cat .context/loop-rule-reviewed-head ``` +**A non-closing pass that owes no repair still commits.** A Minor-only clean pass below the floor, +or an answered suspension that changes no artifact, leaves its findings files tracked and dirty — +and the next pass's files pile up beside them, after which the close procedure's dirty-set check can +never equal one pass's two files and a route the installed ordering requires cannot close through +this plan. **There is no "nothing to commit" branch here**: the findings files are always something. + +**The recorded head is what makes "a clean response against that exact `HEAD`" checkable.** Resolved +and written before the call, it is the value the close procedure compares against — without it, +a commit landing after the response is indistinguishable from the reviewed state. + Re-review after every fix. **Revalidate the evidence entry before every re-review and before the closing commit.** @@ -2295,196 +2464,75 @@ write one, and a later reader could not tell it from an exception record that lo holds, `reset --soft` has already discarded every WIP body, and a closing commit without its provenance line or curve has not validly closed the cycle. -- [ ] **Step 8: Close the cycle** - -**Build the closing message in a file first.** `git reset --soft` discards every WIP commit *body*, -so an evidence entry written only into a WIP message is destroyed at exactly the moment the cycle -closes — which is what step 5 would otherwise have done. +- [ ] **Step 8: Close the cycle — the close procedure, in two invocations** -**8a — record the final pass's findings files. Its own shell invocation, and that is not -cosmetic.** `codex-gate.sh`'s `is_wip_commit` (`plugins/dev-workflow/hooks/codex-gate.sh:763`) tests the **whole command -string** it is given against `-m[[:space:]]*['"]?[[:space:]]*wip`, case-insensitively: **`-m` -immediately followed by optional whitespace, an optional quote, and `wip`.** A single block carrying -this `-m "WIP: …"` and the closing `git commit -F` matches, and is classified cycle-internal in its -entirety — the hook would then carry this cycle's Gate-B count and fingerprint into the next one, a -real closing commit read as a snapshot. **So run 8a and 8b as separate Bash calls.** +**Run `## The four procedures` · Close.** It states the six checks, what success looks like, and why +the two invocations are separate. What follows is the shell that is simple and already verified; +**the checks are the obligation and the shell is one way to run them** — where the observed state is +not one this shell expects, read the state and pick the operation, rather than extending the block. -**8b re-establishes that `HEAD` is still the tip 8a recorded, and stops without moving it if not.** -The split into two invocations is what makes this necessary: time passes between them, and a commit -landing in that window leaves the worktree clean, so every other check in 8b passes while -`reset --soft` stages and squashes that commit into the closing body — defeating the rule that 8a's -findings commit is the only commit admitted between the reviewed tip and the closing act. -**Stopping before the reset is the recoverable direction**; stopping after it is not. - -**What 8b must avoid is that pattern, not the letters.** `--mixed` contains `-m` and matches -nothing, because `i` follows; a path containing `wip` matches nothing, because no `-m` precedes it. -Saying "no `-m` and no `wip` in 8b" would outrun the check the hook performs and make a correct -block look non-compliant. 8b carries no `-m` option at all, which is the property that matters. +**8a — record the candidate pass's findings files.** ```bash BASE=$(cat .context/loop-rule-base) -test -n "$BASE" || { echo "BASE empty — Task 0 did not run"; exit 1; } -# NONCE and P are this cycle's nonce and its final pass number — the same two the -# final call's slot paths were built from. Set them to the values you used. -NONCE=; P= +HEADREV=$(cat .context/loop-rule-reviewed-head) +test -n "$BASE" && test -n "$HEADREV" || { echo "BASE or reviewed head missing"; exit 1; } +test "$(git rev-parse HEAD)" = "$HEADREV" || { echo "HEAD is not the head the candidate pass was issued against"; exit 1; } + +# Exactly this pass's two slot paths, from the values the call was built from. +NONCE=; P= FINAL=".context/codex-reviews/gate-b-spec-$NONCE-pass-$P.md .context/codex-reviews/gate-b-quality-$NONCE-pass-$P.md" expected=$(printf '%s\n' $FINAL | sort) -# EVERY porcelain record, whatever its status letters — staged modifications, -# deletions, renames and conflicts included. A filter that reads only '??' and -# 'A ' declares the set exact while a staged tracked change sits beside it, and -# `git commit` then commits the whole index. actual=$(git status --porcelain -z | tr '\0' '\n' | sed -n 's/^.\{3\}//p' | sed '/^$/d' | sort) +test "$expected" = "$actual" || { echo "dirty set is not exactly this pass's findings files:"; git status --porcelain; exit 1; } -head_paths=$(git show --name-only --pretty=format: HEAD | sed '/^$/d' | sort) -if [ -z "$actual" ] && [ "$head_paths" = "$expected" ]; then - echo "final findings files already recorded by a previous attempt — continue at 8b" -elif [ "$expected" = "$actual" ]; then - PRE=$(git rev-parse HEAD) # the reviewed tip, before the record commit - # shellcheck disable=SC2086 - git add $FINAL - git commit -m "WIP: Gate-B findings files" || { echo "record commit FAILED"; exit 1; } - test "$(git show --name-only --pretty=format: HEAD | sed '/^$/d' | sort)" = "$expected" || { - echo "record commit changed paths beyond the final findings files — restoring the reviewed tip" - git reset --mixed "$PRE" - git status --porcelain - echo "HEAD and index restored. Inspect the paths listed above, re-establish every closure" - echo "condition, obtain a fresh clean response, then retry 8a." - exit 1; } -else - echo "dirty set is not exactly the final pass's findings files:"; git status --porcelain; exit 1 -fi -test -z "$(git status --porcelain)" || { - echo "tree not clean after the record commit — restoring the reviewed tip" - test -n "${PRE:-}" && git reset --mixed "$PRE" - git status --porcelain - echo "Inspect the paths above, re-establish every closure condition, obtain a fresh clean" - echo "response, then retry 8a." - exit 1; } -git rev-parse HEAD > .context/loop-rule-reviewed-tip +# shellcheck disable=SC2086 +git add $FINAL +git commit -m "WIP: Gate-B findings files" || { echo "record commit FAILED"; exit 1; } +test "$(git show --name-only --pretty=format: HEAD | sed '/^$/d' | sort)" = "$expected" \ + || { echo "record commit changed paths beyond this pass's findings files"; exit 1; } +test "$(git rev-parse HEAD^)" = "$HEADREV" || { echo "record commit's parent is not the reviewed head"; exit 1; } +test -z "$(git status --porcelain)" || { echo "tree not clean after the record commit"; exit 1; } + +git rev-parse HEAD > .context/loop-rule-reviewed-tip # the RESTORE POINT, not the reviewed head ``` -**The first branch is the retry path, and it compares `HEAD`'s exact changed-path set.** A closing -commit that fails leaves the findings files already committed, so on a second run the dirty set is -empty — and an unconditional record commit, or an exact-dirty-set guard with no such branch, -refuses the very state the recovery produces. - -**The dirty set is built from every porcelain record, not from two status codes.** An earlier draft -read only `?? ` and `A ` entries, so a **staged tracked modification** sitting beside the two -findings files was invisible: the set compared equal, `git commit` committed the whole index and -swept the change in, the following clean-tree test passed, and 8b published content the reviewed -`HEAD` never carried. **And the record commit's own changed-path set is checked afterwards**, -because the guard describes the working tree while the commit is what actually lands. - -**Both post-commit checks restore the pre-record tip when they fail**, like every other rejection in -step 8: the changed-path test **and** the clean-tree test after it. A bare exit at either would -leave a rejected WIP commit at `HEAD` carrying an unreviewed path, or an extra dirty path beside it -— and the retry branch would then refuse the state, since neither the dirty set nor the changed-path -set is `$FINAL`, so the prescribed procedure could not resume while the side effect stayed -committed. `${PRE:-}` is guarded because the retry branch reaches the clean-tree test without -setting it. - -**"`HEAD` touched some `gate-b-` file" is not that test, and would be unsafe.** Step 7 commits -earlier passes' findings files, so `HEAD` can carry one while the *final* pass's two are missing -entirely — from a mistyped `NONCE`, a wrong `P`, or a pass whose files were never written. The -branch would then skip the record commit and let 8b close without the artifacts the closing commit -exists to carry. **Requiring `HEAD`'s changed paths to equal `$FINAL` exactly** admits only the -commit 8a itself would have made. +**The reviewed head and the restore point are two values.** The first is what the pass read; the +second is that plus the findings files. **A single value cannot be both**, and treating it as one is +how a commit that landed after the response reaches the close — which is why `HEAD^` is compared +above rather than assumed. **8b — reset and close. A separate invocation, carrying no `-m` option at all.** ```bash BASE=$(cat .context/loop-rule-base); TIP=$(cat .context/loop-rule-reviewed-tip) test -n "$BASE" && test -n "$TIP" || { echo "BASE or TIP missing — 8a did not complete"; exit 1; } -# 8a and 8b are separate invocations on purpose, so time passes between them and a -# commit can land without leaving the worktree dirty. Nothing below would notice. -test "$(git rev-parse HEAD)" = "$TIP" || { - echo "HEAD has moved since 8a — it is $(git rev-parse HEAD), the reviewed tip is $TIP" - echo "NOT resetting. Whatever landed is unreviewed; re-establish the closure evidence first." - exit 1; } +test "$(git rev-parse HEAD)" = "$TIP" || { echo "HEAD has moved since 8a — NOT resetting"; exit 1; } git reset --soft "$BASE" -git commit -F .context/loop-rule-closing-msg || { - echo "closing act FAILED — restoring the reviewed tip without discarding its side effects" - git reset --mixed "$TIP" - git status --porcelain - echo "HEAD and index restored to the reviewed tip. Any files listed above were left by the" - echo "failed attempt — inspect them, then re-establish every closure condition before retrying 8a." - exit 1; } - -bad="" -case "$(git log -1 --pretty=%s)" in - [Ww][Ii][Pp]:*) bad="closing commit reads as a snapshot" ;; -esac -test -z "$(git status --porcelain)" || bad="${bad:+$bad; }worktree dirty after the closing act" -if [ -n "$bad" ]; then - echo "closing act REJECTED: $bad — restoring the reviewed tip without discarding side effects" - git reset --mixed "$TIP" - git status --porcelain - echo "HEAD and index restored. Re-establish every closure condition, obtain a fresh clean" - echo "response against the new HEAD, then retry 8a." - exit 1 -fi +git commit -F .context/loop-rule-closing-msg +``` + +**Then check the result, and on any failure run `## The four procedures` · Failure** — which +inspects `HEAD`, the index **and** the working tree, preserves whatever the attempt left, and +restores by **mixed** reset to `$TIP`: -rm -f .context/loop-rule-base .context/loop-rule-reviewed-tip +```bash +git log -1 --pretty=%s # expect the real message, not a snapshot +git status --porcelain # expect empty +``` + +**Both clean, and only then:** + +```bash +rm -f .context/loop-rule-base .context/loop-rule-reviewed-tip .context/loop-rule-reviewed-head ``` -**Both post-commit checks restore the tip as well, and an earlier draft left them as bare exits.** -A closing act that *commits successfully* and then fails its subject or clean-tree check left -`HEAD` at the rejected commit with no way back — and 8a could not be re-entered from there either, -since that commit's changed-path set is the whole squash rather than `$FINAL`. The cycle was then -unable to close through the plan and unable to return to its reviewed state, which is exactly the -incomplete-closing-act transition target §A requires to be recoverable. - -**`--mixed`, never `--hard`.** A closing act can fail *after* a commit hook has modified or staged -tracked content, and `reset --hard` would delete exactly that — the delta the failure produced. -Target §A1 requires the executor to inspect what an incomplete closing act left and to re-establish -the closure conditions against it; a recovery that discards the evidence first makes that -impossible and retries against a repository state the target says must be evaluated. -`--mixed` puts `HEAD` and the index back at the reviewed tip and leaves the working tree alone, so -`git status` above *is* the attempt's delta. - -**Both scratch files are removed only after a successful close**, so a later run cannot inherit -either, and the retry path above still has them. - -**Every check in that block fails the script; none of them is a comment.** An earlier draft -suppressed the record commit with `|| true` and left the status and log lines as things to look at, -so a failed record commit, a dirty tree or a closing message still reading `WIP:` all proceeded -through the soft reset and deleted the recovery base — producing, silently, the exact state the -plan elsewhere calls an invalid close. - -**`|| true` is replaced by branching on what the state actually is, before committing anything.** -"Nothing to commit" and "the commit failed" are different outcomes and only the first is fine. 8a -decides between them with two predicates it can name: the **status-derived dirty set** — every -`git status --porcelain -z` record with its three-character status prefix removed, sorted — and -**`HEAD`'s changed-path set**. An empty dirty set with `HEAD` equal to -`$FINAL` is the harmless retry; the dirty set equal to `$FINAL` is the first run; anything else -stops. **A `git diff --cached --quiet` test is not among them** — an earlier draft's prose claimed -it after the block had stopped using it, which is the overclaim `AGENTS.md` names by requiring -prose to state the exact comparison performed. - -**The base file is removed only in the success branch.** `git commit` can fail on a hook, a signing -key or an unset identity, and at that point the reset has already happened: the WIP commits are -gone, the whole change is a staged tree, and `.context/loop-rule-base` is the only record of where -the cycle started. Deleting it unconditionally destroys the one value a rerun needs, in the single -state where it is needed. - -**`$BASE` is the recorded revision, not a placeholder to substitute by hand**, and the file is -removed afterwards so a later run cannot inherit a stale one. - -**`reset --soft` stages committed content only.** The prompt-standards result (step 4b), the -completeness sweep, the next-state table, the divergence list, the equivalence result and the -fragment evidence all land in this plan, and `.context/codex-reviews/` is tracked; **anything still -uncommitted when the reset runs is left in the worktree and is not in the closing commit** — and, -for the plan records, was never in a Gate-B range either. - -**Which is why step 7 commits them *before* the candidate final pass, and this step only adds the -findings file that pass produced.** The guard above is what keeps that distinction real: if any -other path is still dirty here, a record was written after the review rather than before it, and -the closing tree would differ from the one the clean pass read. - -**`$BASE` is the recorded revision, not a placeholder to substitute by hand** — Task 0 persisted it -for this, and a mistaken substitution squashes the wrong range. - -The closing body carries: the validated evidence entry; the provenance line; the per-pass curve; and any human-exception record. **One commit rather than a follow-up** — a `WIP:` commit left in history defeats the convention, and a follow-up has nothing to commit when the review produced no fixes. +**`reset --soft` stages committed content only.** The prompt-standards result, the completeness +sweep, the next-state table, the divergence list, the equivalence result and the fragment evidence +all land in this plan, and `.context/codex-reviews/` is tracked; **anything uncommitted when the +reset runs is left in the worktree and is not in the closing commit** — and, for the plan records, +was never in a Gate-B range either. Step 7 commits them before the candidate pass is issued, which +is what lets 8a's dirty-set check be exact. --- @@ -2535,6 +2583,13 @@ as extracted and the result in both directions, per copy.* --- +## Fragment sweep (Task 0 output) + +*Empty until Task 0 runs. Task 0 replaces this entire section: the revision the sweep ran at, one +line per fragment-table row with its three results, and the reading result for `F1`, `F2` and `F3`.* + +--- + ## Fragment evidence (per-task output) *Empty until the tasks run. One subsection per task — `### Task 1`, `### Task 3`, … — replaced From 9ced53c2f293671e09acbf84cf65f533a059bd7a Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Tue, 15 Sep 2026 11:02:24 +0200 Subject: [PATCH 129/181] docs(plans): apply Gate-A plan pass 22; the accounting table was itself wrong MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Targeted pass Daniel asked for after approving the method change: does the new flow close pass 21's five findings, and did the accounting drop anything silently. It found both — five Blockers, three Majors, one Minor, one Nit. The counts are not comparable to the previous curve: this pass carried a new charge, so it read text that had existed for one commit. THREE ACCOUNTING ROWS WERE FALSE WHEN WRITTEN. Rows 8, 15 and 40 claimed "kept, same shell" for obligations that were not present everywhere: - Row 8: three concrete blocks read $BASE with no non-empty guard — the baseline extraction, Task 10's hook diff, and the first Gate-B call, which would have issued Gate B with an empty baseSha. - Row 15: only the baseline artifact used temp-then-rename. The untouched map was written straight to the path Tasks 2 and 14 read, so an interruption left a same-base partial map at its consumer path — a second unlisted deletion. - Row 40: "mixed reset preserves the attempt's delta" is an overclaim. --mixed destroys an INDEX-ONLY change, which is exactly the hook side effect §A says must be inspected. The delta is captured to two patch files first. And an 8a rejection had no tip to restore to at all: the reviewed tip does not exist until 8a's commit lands, so 8a restores to the reviewed head. Other Blockers and Majors: - printf %b leaves \* as \*, so the Severity site's anchors were stored with backslashes that occur in neither copy and my new literal-uniqueness check would have stopped Task 0 deterministically on a correct tree. The sites are emitted as literals; the regex escaping stays on $se/$ee where it belongs. - Each site's diff status is now classified: 0 and 1 are results, above 1 is a failure, and a failed comparison can no longer be renamed as complete. - The failure procedure demanded a fresh clean response unconditionally. §A1 has THREE routes: retry the act where every condition still holds, owe a further pass only where the attempt moved something a condition reads, and PARK the cycle where the failure cannot be repaired at all. The third was missing entirely. - The disposition table's `kept` row had no route for a condition in a passage no untouched span reaches — which is passages (b), (c), (e) and (i) — so Tasks 4, 5 and 7 each patched it locally with an id list, against the rule that no task enumerates ids. The table has the third route now and the three lists are gone. - Global Constraints let each copy "keep its own wrapping style" while the byte-identical rule and Task 14's parity diff both compare C against W. - Every task that records into this plan now lists it under Files. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-plan-om0bdd7udh-pass-22.md | 11 + .../2026-09-14-loop-rule-consolidation.md | 208 ++++++++++++------ 2 files changed, 151 insertions(+), 68 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-22.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-22.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-22.md new file mode 100644 index 0000000..86138b6 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-22.md @@ -0,0 +1,11 @@ +BLOCKER | high | Task 0 step 3 | The Severity site is authored as `\*\*Severity:\*\*` and `\*\*Tool routing:` inside `printf %b`; `\*` is not converted to `*`, so the literal-uniqueness checks search for backslashes that do not exist in either source. | The baseline procedure deterministically stops at that site and Task 0 can never finish, so no implementation task is reachable. | Emit the two literal anchors without backslashes, using a field-wise tab-producing form, and reserve regex escaping for the later derived `se` and `ee` values. +MAJOR | high | Task 0 step 3 | A `diff` is the last command in each loop body, but the loop does not inspect its status per site; an I/O or invocation failure on any non-final iteration is overwritten by a later successful iteration, while status 1 is also the expected result for a real divergence. | A failed comparison can leave no diff payload for a recorded site and still allow the temporary baseline to be renamed as complete, so resume can trust parity evidence that was never produced. | Capture each `diff` status immediately, accept 0 and 1, fail on values above 1, and only write a completed site result after that classification. +BLOCKER | high | Accounting row 15 and Task 0 step 2 | The accounting says both artifacts retain the temp-file-then-rename condition, but only the baseline artifact has such a construction; the instructions write `loop-rule-untouched` directly and provide no temporary path or final rename. | This is a second unlisted deletion from the replaced Task 0 procedure, despite the table claiming condition 20 is the only deletion, and an interrupted map can exist at the final consumer path. | Require the untouched map to be built under a temporary name, complete its shape and coverage validation there, and rename it only as the final operation. +BLOCKER | high | Accounting row 8; Task 0 step 3, Task 10 step 4, and Task 15 step 6 | The accounting retains the obligation that every `$BASE` consumer rejects an empty value, but these concrete blocks read `loop-rule-base` and use the result without the non-empty guard; the generic paragraph narrows its promise to later tasks that count and therefore does not cover these consumers. | An empty or damaged base file can drive baseline extraction and hook-diff checks against the wrong Git operand and can issue Gate B with an empty or invalid `baseSha`, defeating the reviewed-range guarantee. | Put the non-empty guard in every concrete block immediately after reading the file, including the baseline, hook-only diff, and first Gate-B-call blocks. +BLOCKER | high | Task 15 step 8a and Failure procedure | The record commit can fail, or can land and then fail its changed-path or clean-tree check, before `loop-rule-reviewed-tip` is written; all three branches merely exit, while the Failure procedure can restore only a restore point that does not yet exist. | A rejected 8a can leave hook side effects, a bad WIP commit, or a dirty index with no prescribed recovery tip, so accounting row 40's promise that every rejection restores a tip was silently dropped and the close cannot be resumed safely. | Persist the pre-8a reviewed head as the recovery target before the commit, and route every 8a failure and postcondition rejection through recovery against that value while preserving and reporting side effects. +BLOCKER | high | Failure procedure | The procedure claims `git reset --mixed ` preserves the attempt's delta, but a hook can create an index-only change whose worktree copy still equals the restore point; mixed reset replaces that index entry and leaves no worktree delta for `git status` to report. | The recovery operation can destroy exactly the staged side effect the procedure says must be inspected and preserved, violating the approved incomplete-closing-act rule and creating data loss. | Snapshot or otherwise preserve the index delta before the mixed reset, inspect index and worktree separately, then restore the tip and report or reapply both classes of side effect without claiming mixed reset alone preserves them. +MAJOR | high | Failure procedure, final paragraph | The procedure unconditionally requires a fresh clean response after every failed closing act, although target §A1 says a failed act returns to the closure step and requires a further pass only when the attempt or its repair moved something a closure condition reads. | A transient commit failure that changes no relevant state is routed through an extra pass instead of retrying the already-authorized act, so the implementation procedure does not implement the approved failure transition. | Re-establish every closure condition first, retry the act directly when they still hold, and require a fresh pass only when the attempt or repair changed a condition under its own source rule. +MAJOR | high | How a task discharges that table; Tasks 4, 5, and 7 | The plan says no task enumerates condition ids and that the unit is the source block replaced, but these tasks reproduce condition-id sets and extend their walks to unchanged text outside their replacement blocks; several kept conditions then receive per-condition preservation counts even though the disposition table permits that route only when they share a line with changed text. | The 135-condition table is no longer the single authority and there is no execution that simultaneously follows the global procedure and the task instructions; later edits can stale the task-local lists or make an executor skip one source. | Put every block-specific exception in the disposition table, extend the untouched-span artifact to cover unchanged prefixes and suffixes where needed, and have each task invoke the table only for the exact source blocks it replaces. +MINOR | high | Global Constraints, fourth bullet | The plan allows each copy to keep its own wrapping style while also requiring every NEW and REPLACED prompt-copy block to be byte-identical and later expecting parity diffs with no output. | An executor can follow the wrapping permission, install identical words with different line breaks, and then fail the plan's own byte-parity checks despite having followed the earlier instruction. | State that line breaks may differ from the target-text artifact but the chosen wrapping for every common installed block must be identical between C and W. +NIT | high | Task file lists | Task 0 says no files are modified even though step 4 modifies and commits this plan, and Tasks 1, 10, 11, 12, 14, and 15 likewise omit this plan from their file lists despite recording or committing output into it. | Per-task scope metadata understates the files each task writes, making handoff and review boundaries less reliable. | Add this plan to every task's Modify list when that task records observations, rows, or result sections in it. +END OF FINDINGS (10 total) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index e4e99f1..aeacca2 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -19,7 +19,7 @@ - **Both prompt copies take every NEW and REPLACED section byte-identical**, except §F's seven hook items, whose destination is the shipped hook and its test and which carry no parity obligation (target §"How to read a section", §F opening). - **C** = `CLAUDE.md`. **W** = `plugins/dev-workflow/commands/workflow-init.md`. Every line number below is re-read at execution; the inventory's numbers cite `7c0d475` and have drifted. - **Every hook replacement installs into a double-quoted POSIX-shell `note` argument** and therefore carries no backtick, no `$(`, no backslash and no double quote. `$policy`, `$floor`, `$passes`, `$passesA` and `$fresh` are the intended interpolations (target §F opening). -- **The target's fenced blocks are normative in their words, not in their line breaks.** The spec wraps for its own readability; each copy keeps its own wrapping style. What design §7 requires is narrower and is the rule here: **every fragment this plan counts must sit wholly within one line of the file it is grepped from**, and where installing a replacement would put a counted fragment across a wrap, that fragment's line is installed unwrapped. The fragment table below states, for every count, the line it must sit on. **Nothing in this plan claims the installed bytes equal the fenced block's bytes**, and no check asserts it. +- **The target's fenced blocks are normative in their words, not in their line breaks.** The spec wraps for its own readability, and the installed wrapping need not match the spec's. **It must match between the two copies**, though: the byte-identical rule above and Task 14's parity diff both compare C against W, so "each copy keeps its own wrapping style" — as an earlier draft put it — licenses exactly the difference those checks then fail on. **Choose the wrapping once per installed block and use it in both.** What design §7 requires is narrower and is the rule here: **every fragment this plan counts must sit wholly within one line of the file it is grepped from**, and where installing a replacement would put a counted fragment across a wrap, that fragment's line is installed unwrapped. The fragment table below states, for every count, the line it must sit on. **Nothing in this plan claims the installed bytes equal the fenced block's bytes**, and no check asserts it. - **Invariant 5 (exact pinning)** and **invariant 12 (a plugin change requires a version bump)**: this change touches `plugins/`, so `plugins/dev-workflow/.claude-plugin/plugin.json` goes `0.11.0 → 0.12.0` with a CHANGELOG entry (design §8). - **Invariant 11:** every prompt change passes all 12 items of `docs/prompt-standards.md`. Most at risk: item 6 (every constraint carries its reason in the same sentence) and item 8 (token-lean). - **Invariant 4 / the hook:** no control flow, counter, fingerprint computation, routing or event handling changes. Only `note` strings and their expectations. @@ -212,14 +212,14 @@ independent reader did. | 5 | Never overwrite an existing base file | **kept** as an obligation; the shell branch becomes one line of the resume procedure | | 6 | A pre-existing base is an ancestor of `HEAD` (`merge-base --is-ancestor`) | **kept**, same shell — resume | | 7 | Only this run's `WIP:` commits lie between base and `HEAD` | **kept** as an obligation — resume; the reading is the agent's | -| 8 | `$BASE` persisted to a file; every consumer guards it non-empty | **kept**, same shell | +| 8 | `$BASE` persisted to a file; every consumer guards it non-empty | **kept**, and **repaired**: three concrete blocks read it without the guard while this row claimed otherwise — the baseline extraction, Task 10's hook diff, and the first Gate-B call. The guard is in each of them now (pass 22) | | 9 | The five regions are located by anchor, one hit per pattern per file | **kept**, same shell — preparation | | 10 | Spans are **derived** from replacement extents, never hand-written | **kept** — the rule, unchanged | | 11 | Every kept condition lands in exactly one `span` or one `cond` row | **kept** — the coverage assertion, unchanged | | 12 | Anchors, never absolute line numbers; resolved separately per tree | **kept** — the rule, unchanged | | 13 | A single-line site uses `grep`, never `sed -n "/x/,/x/p"` | **kept** — the rule, unchanged | | 14 | Records are tab-separated `base` / `span` / `cond` shapes, `base` first | **kept**, and **extended**: the baseline artifact gains its own `site` shape (pass 21) | -| 15 | Artifacts written through a temp file, renamed after the last record | **kept**, same shell | +| 15 | Artifacts written through a temp file, renamed after the last record | **kept**, and **repaired**: only the baseline artifact had the construction while this row claimed both did, so an interrupted map could sit at the path Tasks 2 and 14 read. The untouched map is built under a temp name and renamed after its coverage assertion (pass 22) | | 16 | On re-entry: base line matches, shapes parse, coverage complete — else rebuild | **re-expressed** as the resume procedure, and **corrected**: the completeness predicate now exists for both artifacts rather than only the map (pass 21) | | 17 | The baseline source is `$BASE` blobs once `WIP:` commits exist | **kept**, same shell | | 18 | Both baseline sources readable | **kept**, same shell | @@ -244,12 +244,15 @@ independent reader did. | 37 | The closing commit's subject is not a snapshot | **kept** as the close procedure's check | | 38 | The tree is clean after the close | **kept** as the close procedure's check | | 39 | 8a and 8b are separate invocations; no `-m` in the closing one | **kept**, and it is why the close procedure names two invocations | -| 40 | Every rejection restores a tip by **mixed** reset, never `--hard` | **re-expressed** as the failure procedure, and **corrected**: restoring `HEAD` is not the whole duty — index and working tree are inspected, and the side effects are preserved and reported, which the approved §A requires and a `HEAD`-only rule never said | +| 40 | Every rejection restores a tip by **mixed** reset, never `--hard` | **re-expressed** as the failure procedure, and **corrected twice**: restoring `HEAD` is not the whole duty — index and working tree are inspected and captured to patch files first, because **`--mixed` destroys an index-only change** and claiming it preserves the delta was an overclaim; and an 8a rejection restores to the reviewed *head*, the reviewed *tip* not existing yet (pass 22) | | 41 | The scratch files are removed only after a successful close | **kept** as the close procedure's last step | **Nothing in that table is dropped except condition 20, and that one is dropped by being replaced -with a stronger obligation.** Six conditions are corrected or extended, each against a pass-21 -finding. The rest keep their force; what changes is that a person reads the state and picks the +with a stronger obligation.** Six conditions are corrected or extended against pass-21 findings, and +**three rows were themselves wrong when first written** — 8, 15 and 40 claimed "kept, same shell" +for obligations that were not in fact present everywhere. Pass 22 was aimed at exactly that question +and found them; they are repaired rather than re-worded, and the fact that an accounting table can +itself be wrong is why the check was worth running. The rest keep their force; what changes is that a person reads the state and picks the operation, instead of one block trying to branch through every state in advance. ### Preparation — before any task edits a file @@ -314,26 +317,63 @@ come before the reset. ### Failure — the closing act did not complete **The approved §A requires the state a failed closing act leaves to be inspected**, and that is more -than `HEAD`. **Look at three things and say what each holds:** +than `HEAD`. **Look at three things and say what each holds, before moving anything:** - **`HEAD`** — did the commit land? A failed `git commit` leaves it where it was; a hook that rejected after committing does not. -- **The index** — a commit hook can stage content, and `reset --hard` would destroy exactly that. -- **The working tree** — the same hook can modify files without staging them. +- **The index** — a commit hook can stage content. `git diff --cached` names it. +- **The working tree** — the same hook can modify files without staging them. `git diff` names it. -**How success is recognised.** The repository is back at the restore point **with the attempt's -delta preserved and reported** — `git status` names it — and the recovery base and reviewed-head -files still exist. **`git reset --mixed ` is the operation that does this**; `--hard` -is never it. +**Capture the delta before restoring, because no reset preserves all of it.** `--hard` destroys +both. **`--mixed` destroys an index-only change** — one whose worktree copy still equals the +restore point — because it rewrites the index from the target commit and leaves nothing for +`git status` to report. Saying "`--mixed` preserves the delta" was an overclaim; what it preserves +is the worktree half. + +```bash +git stash create > /dev/null 2>&1 || true # no-op on an empty delta +git diff --cached > .context/loop-rule-failed-index.patch +git diff > .context/loop-rule-failed-worktree.patch +``` + +**Both patches are written even when empty**, so "nothing was left behind" is a recorded +observation rather than an absent file. Then restore, and report from the patches rather than from +`git status` alone. + +**The restore target is whichever tip exists.** After 8a's commit that is +`.context/loop-rule-reviewed-tip`; **before it** — a record commit that failed, or landed and then +failed its changed-path or clean-tree check — that file does not exist yet and the target is +`.context/loop-rule-reviewed-head`, which every Gate-B call wrote. **An 8a rejection is a rejection +like any other and owes the same restoration**, which an earlier draft's bare exits did not give it. + +**How success is recognised.** The repository is back at that tip, the two patch files describe what +the attempt left, and the recovery base and reviewed-head files still exist. **On deviation.** If the delta cannot be explained, stop and hand it to a person. **Do not retry a close against a state you cannot account for** — the second attempt would carry whatever the first one left. -**What this procedure does not promise.** It does not transact the retry. After the state is -understood, the executor re-establishes every closure condition, obtains a fresh clean response -against the new head, and re-enters the close procedure from its first check. **That is a decision -sequence, not a rollback**, and trying to automate it is what grew the block this section replaces. +**What happens next is §A's, and it has three routes, not one.** A failed act **returns to the +closure step it failed in**, not to the branches — this cycle's pass already took the +clean-completion branch. So: **re-establish every closure condition against the repository as it now +stands**, then read which route you are on. + +- **Every condition still holds** → **perform the act again.** No further pass is owed. An earlier + draft demanded a fresh clean response here unconditionally, which routes a transient identity or + signing failure — one that moved nothing any condition reads — through a whole extra pass the + approved text does not ask for. +- **The attempt or its repair moved something a condition is read from** → that condition has + changed, **and its own rule decides what it costs**, a further pass included. The cycle is back in + the ordering with that pass owed. The two patch files above are how you tell which case you are + in: an empty pair means nothing moved. +- **The failure cannot be repaired at all** — a signing key nobody has, a permission nobody can + grant — → **surface it and leave the cycle parked**: open, not running, spending no passes, + restarted by an explicit later continue. **This route was missing entirely**, and without it a + cycle that can neither close nor be parked is exactly the outcome §A names. + +**What this procedure does not promise.** It does not transact any of that. It restores a known +state and records what the attempt left; **the route is a reading**, and trying to automate the +reading is what grew the block this section replaces. ### Resume — re-entering after an interruption @@ -392,7 +432,7 @@ different passage, because each task was inventing the rule for its own conditio | Disposition | The observation it owes | |---|---| -| **kept** | inside an untouched span, **or**, where it shares a line with changed text, its own per-condition count: `parent=1 worktree=1` in each copy. Never both, and never neither. | +| **kept** | **exactly one of three routes**, never two and never none: inside an untouched span; **or** its own per-condition count, `parent=1 worktree=1` in each copy, where it shares a line with changed text; **or** that same per-condition count where **no untouched span reaches its passage at all** — Task 0 maps five regions, and passages (b), (c), (e) and (i) are not among them, so every kept condition there takes this third route. | | **carried** | a **preservation count** after installation, `1` in each copy, from the condition's own text. **No untouched span covers a carried condition** — it sits inside a replacement block, which is the whole reason it is not recorded as kept. | | **replaced** | a **discriminating pair**, `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`, **both halves from the same edit**. | | **moved** | **two** observations: an **absence** at the source, `parent=1 worktree=0`, and a **condition-specific presence** at the destination, `worktree=1 parent=0`. A presence check on the destination *paragraph* is not the second half — it passes while any one moved predicate is missing from it. | @@ -548,7 +588,8 @@ preservation fragments that an earlier draft chose afterwards. ## Task 0: Establish the baseline and the do-not-touch set -**Files:** none modified. +**Files:** +- Modify: `docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md` — the fragment sweep record (step 4) **Interfaces:** - Produces: `$BASE` (the parent commit every later task's counterfactual half runs against), and the five untouched passage ranges recorded as **anchor spans plus a per-condition fragment list** — never as absolute line numbers, for the reason step 2 gives. @@ -682,6 +723,13 @@ expected pair is `11`; the shape carries the values rather than assuming th dropped condition recorded here later needs no new format. **Both consumers validate every `cond` row**, not only the `span` rows. +**Build the map under `.context/loop-rule-untouched.tmp` and rename it only after the coverage +assertion passes.** The baseline artifact is written that way and accounting row 15 claims both are; +an earlier draft wrote this one straight to its consumer path, so an interruption between the `base` +line and the last `cond` row left a same-base partial map at exactly the path Tasks 2 and 14 read. +**The rename is the last operation**, after every `span` and `cond` row is written and every kept +condition is assigned — which is the same completeness predicate the resume procedure re-checks. + **Record the spans as one `spanstartendfile` line each in `.context/loop-rule-untouched`**, the anchors being literal strings. Tab-separated because the anchors contain colons — a `:` delimiter splits `**Severity:**:**Tool routing:` at the wrong colon and yields an empty end anchor. @@ -733,6 +781,7 @@ exits on failure; the rename is the last statement. ```bash BASE=$(cat .context/loop-rule-base) +test -n "$BASE" || { echo "BASE empty or unreadable — Task 0 did not run"; exit 1; } if [ -n "$(git log --oneline "$BASE"..HEAD)" ]; then git show "$BASE:CLAUDE.md" > .context/loop-rule-c.src git show "$BASE:plugins/dev-workflow/commands/workflow-init.md" > .context/loop-rule-w.src @@ -743,21 +792,37 @@ fi test -r "$C_SRC" && test -r "$W_SRC" || { echo "baseline source unreadable"; exit 1; } printf 'base\t%s\n' "$BASE" > .context/loop-rule-baseline-diff.tmp # renamed at the end -# Tab-separated start and end anchors: the anchors contain colons, so a -# colon delimiter splits '**Severity:**' at the wrong place and yields an -# empty end. Every entry here spans two DIFFERENT anchors — the single-line +# Tab-separated start and end anchors, written as LITERAL text — the escaping +# for sed happens later, on $se and $ee. Emitting '\*\*Severity:\*\*' here would +# store backslashes that occur in neither copy, and the literal-uniqueness +# check below would then fail at that site on a correct tree. +# Two fields per call so the tab is the format's, not the data's: the anchors +# contain colons, so a colon delimiter would split '**Severity:**' at the +# wrong one. Every entry spans two DIFFERENT anchors — the single-line # squash-carry site is handled below, not in this loop. -printf '%b\n' \ - 'Both gates are a LOOP\tNothing here writes the floor knob' \ - 'What a loop absorbs\tRecognizing "clearly stuck"' \ - 'Recognizing "clearly stuck"\tEvery pass report states' \ - 'From pass 4 onward\tThose three lines expose' \ - 'Those three lines expose\tThe two rules above' \ - 'The two rules above\tFindings go to a FILE' \ - '\*\*Severity:\*\*\t\*\*Tool routing:' \ - 'Recording a human exception\tbecause writing it down makes it sound' \ - 'When these rules bind\tDownstream has no shipping commit' \ - > .context/loop-rule-sites +: > .context/loop-rule-sites +while IFS= read -r s && IFS= read -r e; do + printf '%s\t%s\n' "$s" "$e" >> .context/loop-rule-sites +done <<'SITES' +Both gates are a LOOP +Nothing here writes the floor knob +What a loop absorbs +Recognizing "clearly stuck" +Recognizing "clearly stuck" +Every pass report states +From pass 4 onward +Those three lines expose +Those three lines expose +The two rules above +The two rules above +Findings go to a FILE +**Severity:** +**Tool routing: +Recording a human exception +because writing it down makes it sound +When these rules bind +Downstream has no shipping commit +SITES # Each site: the same checks Task 14 applies — literal uniqueness per copy, an # ESCAPED anchor (a derived anchor carries ** and / and . and is a regex to sed), # and a non-empty extract. Two empty extracts diff equal and would certify a @@ -828,6 +893,7 @@ working state, not deliverables; only the sweep record is committed. **Files:** - Modify: `CLAUDE.md` — insert immediately before the line beginning `**What a loop absorbs, and what stops it` - Modify: `plugins/dev-workflow/commands/workflow-init.md` — the same anchor +- Modify: `docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md` — this task's fragment evidence - Test: the counts in Step 4 **Interfaces:** @@ -912,7 +978,7 @@ makes one per task. A non-`WIP` commit here would reset the hook's Gate-B counte ## Task 2: Verify the untouched passages are still untouched -**Files:** none modified. +**Files:** none modified — this task verifies and records nothing of its own; Task 14 step 4b re-runs both lists after the last text edit and records the results. **Interfaces:** - Consumes: Task 0's recorded anchor spans and per-condition fragment list, and `$BASE`. @@ -950,7 +1016,7 @@ Read the installed text around "The two rules above do not compete" and confirm **Files:** - Modify: `CLAUDE.md` — the passage beginning `**What a loop absorbs, and what stops it` - Modify: `plugins/dev-workflow/commands/workflow-init.md` — the same - +- Modify: `docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md` — the rows this task appends and its fragment evidence **The bytes:** target §B, which states it is "the only place this file states anything about" passage (b). **`b3` is aligned here, not in Task 14.** Design §6 decides it explicitly: W's pointer names "the @@ -1047,7 +1113,7 @@ git commit -m "WIP: replace passage (b) with the absorb paragraph that owns the **Files:** - Modify: `CLAUDE.md` — the passage beginning `**Recognizing "clearly stuck"` - Modify: `plugins/dev-workflow/commands/workflow-init.md` — the same - +- Modify: `docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md` — the rows this task appends and its fragment evidence **The bytes:** target §C's fenced block, which gives the three-condition sentence entire so nothing begins mid-clause. **The split, stated because getting it wrong is the failure §C names:** only the operative precedence clause — from "a clean completion takes precedence over this exit" to the end of that sentence — moves into §A, capitalized there. Its opening clause, the plateau rationale, **stays here**. The below-floor sentence (`c14`) **does not move; it is replaced**, and no copy of its live wording may survive beside the ordering's split. @@ -1073,10 +1139,6 @@ decided at execution. **Two things about passage (c) that the class alone does not tell you:** -- **`c1`–`c3` are kept and no untouched span reaches them.** Task 0's five regions do not include - passage (c), and §C replaces text inside it, so the curve-reading premises owe **per-condition - preservation counts** — `parent=1 worktree=1` in each copy — like a carried condition. Without - them an accidental edit to the curve or coverage sentences survives every mechanical check here. - **`c9` and the plateau rationale are two subjects, not one condition with two halves.** `c9` is the precedence clause and is **moved** — absence here, condition-specific presence in §A. The plateau rationale carries no inventory id, **stays**, and owes a **preservation count** recorded @@ -1139,9 +1201,10 @@ two-instructions-that-disagree failure, and `c10`–`c13` are four chances at it - [ ] **Step 4b: Run every remaining row step 1 appended** -The carried `c5`–`c7`, the kept `c1`–`c3`, and `c9`'s plateau rationale — each to the result its -class owes. **None of them is observed by a pair**, which is why they are listed as a step rather -than left to the walk. +Every row step 1 appended that is not a pair — the carried conditions, the kept ones taking the +table's third route, and `c9`'s plateau rationale — each to the result its class owes. **None of +them is observed by a pair**, which is why they are a step of their own rather than left to the +walk. **Read the set off the disposition table's rows for this task's block**, not from a list. **The source sentence is split; `c9` is not.** `c9` is the precedence clause alone and is **moved**, while the plateau rationale beside it carries no inventory id, **stays**, and is observed under its @@ -1174,7 +1237,7 @@ git commit -m "WIP: widen the clearly-stuck third condition and split its preced **Files:** - Modify: `CLAUDE.md` — the paragraph beginning `Those three lines expose` - Modify: `plugins/dev-workflow/commands/workflow-init.md` — the same - +- Modify: `docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md` — the rows this task appends and its fragment evidence **The bytes:** target §D's two fenced blocks — the `e7` sentence entire, and the pointer paragraph added at the end of the passage. **`e8` does not survive this change.** §D supplies the complete replacement sentence, which reads @@ -1208,11 +1271,10 @@ do not cover: literal count of the whole clause returns zero in a *correct* source file. Choose a single-line fragment of it now, verify uniqueness, and append it; choosing one from the installed text afterwards is the pre-edit rule broken in the one place the wrap makes it tempting. -- **`e1`–`e6`, `e10` and `e11` are kept, and no untouched span reaches them.** Passage (e) is not - one of Task 0's five regions and §D edits inside it, so each owes a **per-condition preservation - count**, `parent=1 worktree=1` — **`e11` in C only**, which is its recorded divergence. Without - them §D's edit can alter or drop a tell, the long-before-plateau sentence or the C-only - rationale while the `e7` pair and the pointer check both pass. +- **`e11` is C-only**, which is its recorded divergence, so its preservation count is expected in + C and not in W. Every other kept condition in passage (e) takes the same route as any kept + condition in a passage no span reaches — the disposition table's third route — and this task + lists none of them by id. - [ ] **Step 2: Install §D's `e7` sentence and the pointer paragraph** @@ -1278,7 +1340,7 @@ git commit -m "WIP: read the two-tell threshold after the clean-completion branc **Files:** - Modify: `CLAUDE.md` — the `**Severity:**` bullet and the `How this demotion bears` paragraph - Modify: `plugins/dev-workflow/commands/workflow-init.md` — the same - +- Modify: `docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md` — the rows this task appends and its fragment evidence **The bytes:** target §E's two fenced blocks. **This task removes the one deliberate story-path divergence.** `g4` — C's sentence naming `docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md` as owning the question — goes, because this change is that work and the question is answered. `g2` and `g3`, the interim report-and-stop duty and its justification, go from **both** copies in the same edit. **Removing g4 from C without removing g2/g3 from W desynchronises the copies in the opposite direction**, which is the failure the inventory flags by name. @@ -1373,7 +1435,7 @@ git commit -m "WIP: answer the demotion question and scope the resolve duty to t **Files:** - Modify: `CLAUDE.md`, `plugins/dev-workflow/commands/workflow-init.md` - +- Modify: `docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md` — the rows this task appends and its fragment evidence **The bytes:** target **§G**'s fenced block — the one-contract paragraph, whose membership is widened and which gains a semantic test a downstream reader can apply — and target **§H**'s blocks, in the order §H gives them: the `c18`-and-surfacing block; `a13`; `a16`; `a17`–`a22`; the Gate-A clean-signal sentence; the gate-prompt template's clean sentence; the Gate-A cadence; the lens paragraph's unchanged-list; and the unknown-start strict-reading list. **Passage (b) is deliberately not among them** — §B owns it. **§G is not in the a–j inventory** and therefore carries no condition ids; design §4 lists it as its own site. **§G is an instruction to the agent, not a checker** — install it as written and do not add a mechanical guard beside it. @@ -1410,10 +1472,10 @@ sentence, so the first — `This replaces the pass-count number and nothing else **absence** fragment of its own, `parent=1 worktree=0` in both copies. Derive it here, from the live text, and install over the whole condition rather than over the sentence P9 names. -**Passage (i)'s kept conditions need attention the class alone does not flag.** `i1`–`i3` and -`i9`–`i16` are kept, no untouched span reaches them — passage (i) is not one of Task 0's five -regions — and `i3` shares its sentence with the dash-delimited list this task replaces. Each owes a -per-condition preservation count, `parent=1 worktree=1`. +**Passage (i) is not one of Task 0's five regions**, so its kept conditions take the disposition +table's third route — a per-condition preservation count — and `i3` additionally shares its sentence +with the dash-delimited list this task replaces. Both facts follow from the table; neither needs an +id list here. **One pair per block is a sample, not coverage.** The blocks below each change more than one thing and are named with what they owe; the rest are covered by one pair each. **No count is stated** — @@ -1503,11 +1565,11 @@ drops them — and because they are carried rather than kept, **no untouched-ran them**, which is exactly why the disposition records them as carried. **Each owes a count of its own text in each copy, expecting `1`.** -The set is **`a15`, `a21`, `a22`, `c15`, and each of `i4`–`i8`** — nine conditions, not three. An -earlier draft checked three and left the rest to P8 and P16, which observe the *changed* clauses -around them and prove nothing about the reproduced words: both copies could omit the surfacing -premise, or any interior item of the strict-reading run, and every pair, presence check and parity -comparison would still pass. +**The set is every condition the disposition table marks *carried* inside this task's ten blocks**, +and it is read from there rather than listed here. An earlier draft named three of them and left +the rest to P8 and P16, which observe the *changed* clauses around them and prove nothing about the +reproduced words: both copies could omit the surfacing premise, or any interior item of the +strict-reading run, and every pair, presence check and parity comparison would still pass. **All nine fragments were derived and appended at step 1b**, from the live text before step 2 installed over it. Two of them, for orientation: @@ -1535,7 +1597,7 @@ git commit -m "WIP: install the one-contract paragraph and the remaining prompt- **Files:** - Modify: `CLAUDE.md`, `plugins/dev-workflow/commands/workflow-init.md` - +- Modify: `docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md` — the rows this task appends and its fragment evidence **The bytes:** target §F items 1, 2, 3, 4, 5, 6, 7, 8, 7a, 8a, 8b, 9a, 9b and 9 — fourteen items, each with its own fenced replacement and its `C nnn` / `W nnn` citation. **Re-read every citation against the current file**: §F's own collected list records that items 4, 5 and 8 have line citations one off, and the numbers drifted further as this cycle edited the copies. **`h4`, `h5` and `h19` are the human-exception conditions these items discharge** — item 7 is both destinations, `h4`'s and `h5`'s, item 4 the scope sentence. @@ -1614,7 +1676,7 @@ git commit -m "WIP: replace the fourteen falsified sentences in the two prompt c **Files:** - Modify: `CLAUDE.md` — the Named residual paragraph (§5), and the work-loop line (§4) - Modify: `plugins/dev-workflow/commands/workflow-init.md` — the same two - +- Modify: `docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md` — the rows this task appends and its fragment evidence **The bytes:** target §F item 14 (the Named residual's blanket exemption) and item 18 (the work-loop sequence). **These two are separated from Task 8 because they were found last and because item 18 is the only edit outside §5.** A reviewer can reject this task while approving Task 8. @@ -1659,6 +1721,7 @@ git commit -m "WIP: correct the Named residual's blanket exemption and the work- **Files:** - Modify: `plugins/dev-workflow/hooks/codex-gate.sh` +- Modify: `docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md` — the eight appended rows and this task's fragment evidence **The bytes:** target §F items 10, 11, 12, 13, 15, 16 and 17. **Items 12, 15 and 16 give the complete resulting text for both channels**; items 10, 11, 13 and 17 give replacement sentences inside an otherwise unchanged message. @@ -1730,6 +1793,7 @@ unique to it. ```bash BASE=$(cat .context/loop-rule-base) +test -n "$BASE" || { echo "BASE empty or unreadable — Task 0 did not run"; exit 1; } git diff -U0 "$BASE" -- plugins/dev-workflow/hooks/codex-gate.sh \ | grep -E '^@@' | sed -E 's/^@@ -([0-9]+).*/\1/' grep -n 'note "' plugins/dev-workflow/hooks/codex-gate.sh @@ -1799,6 +1863,7 @@ git commit -m "WIP: replace the seven gate reminders the ordering falsifies" **Files:** - Modify: `plugins/dev-workflow/hooks/codex-gate.test.sh` +- Modify: `docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md` — this task's fragment evidence **There is no list of these assertions and building one here would repeat a defect.** §F states the duty and deliberately states no count: it twice named one and was twice wrong — "the exact-match expectation" where there are three, and "the remaining are matched by loose patterns these repairs leave standing" where `Gate B satisfied` occurs 21 times and `STOP` 14. **Sweep the file; do not work from a number.** @@ -1880,7 +1945,8 @@ git commit -m "WIP: move every hook assertion that names a replaced reminder str ## Task 12: The `b11`/`b13` equivalence check -**Files:** none modified (unless the check fails). +**Files:** +- Modify: `docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md` — the `b11`/`b13` equivalence result (step 5), and the copies only if the check fails. **Interfaces:** - Consumes: the installed §A block and the installed passage (b). @@ -2042,7 +2108,9 @@ git commit -m "WIP: next-state table and per-condition closure checks" ## Task 14: The parity diff and the divergence list -**Files:** possibly `CLAUDE.md` and `plugins/dev-workflow/commands/workflow-init.md`. +**Files:** +- Modify: `docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md` — the divergence list, the site results, and this task's fragment evidence +- Modify, where an alignment is needed: `CLAUDE.md`, `plugins/dev-workflow/commands/workflow-init.md` Design §6: the two copies must agree on every rule this change ships. **The plan carries the divergence list** — which pre-existing wording differences are deliberate and stay, which are not and are aligned — **and performs the extraction and diff, passage by passage, against the real files.** @@ -2190,6 +2258,7 @@ git commit -m "WIP: align the two copies and record the divergence list" **Files:** - Modify: `plugins/dev-workflow/.claude-plugin/plugin.json` - Modify: `plugins/dev-workflow/CHANGELOG.md` +- Modify: `docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md` — the prompt-standards result and every re-run record **Mode:** read from the story header at execution. It was `battery+check+verification` at the time this plan was written; **read it fresh** — the header is the only writable copy and this plan carries the path, not the value. @@ -2335,6 +2404,7 @@ an enumeration here** — an id list in this step was already stale once, naming ```bash BASE=$(cat .context/loop-rule-base) # the parent of the FIRST WIP, from Task 0 +test -n "$BASE" || { echo "BASE empty or unreadable — Task 0 did not run"; exit 1; } git rev-parse HEAD > .context/loop-rule-reviewed-head # the head THIS call is issued against cat .context/loop-rule-reviewed-head ``` @@ -2488,11 +2558,13 @@ test "$expected" = "$actual" || { echo "dirty set is not exactly this pass's fin # shellcheck disable=SC2086 git add $FINAL -git commit -m "WIP: Gate-B findings files" || { echo "record commit FAILED"; exit 1; } +# Any rejection below goes through the Failure procedure, restoring to $HEADREV — +# the reviewed-tip file does not exist yet at this point. +git commit -m "WIP: Gate-B findings files" || { echo "record commit FAILED — run Failure against $HEADREV"; exit 1; } test "$(git show --name-only --pretty=format: HEAD | sed '/^$/d' | sort)" = "$expected" \ - || { echo "record commit changed paths beyond this pass's findings files"; exit 1; } -test "$(git rev-parse HEAD^)" = "$HEADREV" || { echo "record commit's parent is not the reviewed head"; exit 1; } -test -z "$(git status --porcelain)" || { echo "tree not clean after the record commit"; exit 1; } + || { echo "record commit changed paths beyond this pass's findings files — run Failure against $HEADREV"; exit 1; } +test "$(git rev-parse HEAD^)" = "$HEADREV" || { echo "record commit's parent is not the reviewed head — run Failure against $HEADREV"; exit 1; } +test -z "$(git status --porcelain)" || { echo "tree not clean after the record commit — run Failure against $HEADREV"; exit 1; } git rev-parse HEAD > .context/loop-rule-reviewed-tip # the RESTORE POINT, not the reviewed head ``` From b8b433dc67f967dc54f5fd74c239843f385be381 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Tue, 15 Sep 2026 11:15:07 +0200 Subject: [PATCH 130/181] docs(plans): apply Gate-A plan pass 23; phase the restore target, widen the capture MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Three Blockers and two Majors. Findings fell 10 -> 5 and Blockers 5 -> 3 from the targeted pass. - The recorded base was accepted on being non-empty. A symbolic value such as HEAD passes the blob, ancestry and log checks and then RESOLVES DIFFERENTLY as WIP commits accrue, moving the reviewed range, every parent count and the final reset with the branch. It must be a full 40-hex id that resolves to itself as a commit. - The failure procedure chose its restore target by which tip file existed. After an 8b failure whose repair earns another pass, the previous candidate's reviewed-tip is still on disk when the new 8a starts, so an 8a failure would rewind past the newly reviewed repair. Step 7 deletes it when it writes a new reviewed-head, and the target is chosen by phase and validated before use. - The failure capture covered tracked content only, and then treated two empty patches as proof that nothing moved. The closing message, the recovery base and both tip files are IGNORED paths that the close reads, so a hook could rewrite a closure input with both patches empty. The capture now includes --untracked-files=all --ignored and checksums the four closure inputs. The `git stash create` line is dropped: it discarded its own object and did nothing. - Pass 22's diff-status classification was written as prose and not as code, in both Task 0 step 3 and Task 14 step 1. Status 0 and 1 are results; above 1 is a failure to compare, and the site record was being written before anything classified it. - The close procedure said all six conditions are checked before moving HEAD, while condition 3 makes a commit and condition 6 reads the state after one. Applied literally that leaves a rejected commit at HEAD or a live soft-reset index — the state a retry cannot start from. Preconditions and postconditions are split, and every post-move rejection routes through the failure procedure with its phase's target. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-plan-om0bdd7udh-pass-23.md | 6 ++ .../2026-09-14-loop-rule-consolidation.md | 80 ++++++++++++++----- 2 files changed, 67 insertions(+), 19 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-23.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-23.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-23.md new file mode 100644 index 0000000..aabcf29 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-23.md @@ -0,0 +1,6 @@ +BLOCKER | high | Preparation and Task 0 step 1 | A pre-existing `.context/loop-rule-base` is accepted when merely non-empty and usable as a revision even though success requires a full 40-character object name; a symbolic value such as `HEAD` passes the blob, ancestry, and log checks while resolving to a different commit as WIP commits accrue. | Parent counts, Gate B's `baseSha`, and the final soft reset can move with the branch, shrinking or emptying the reviewed range and excluding earlier implementation commits from the close. | Require exactly one full 40-hex commit id, peel it as a commit, and assert that its canonical resolved id equals the stored bytes before any consumer uses it. +BLOCKER | high | Failure procedure and Task 15 step 8a | The general restore rule chooses whichever `.context/loop-rule-reviewed-tip` exists, while step 8a says that file does not exist and restores to `$HEADREV`; after an earlier 8b failure and a repair that earns another pass, the old reviewed-tip still exists when the new 8a starts, so an 8a failure before the file is overwritten selects a stale restore point. | Recovery can rewind past the newly reviewed repair and its evidence, leaving the cycle on a prior candidate even though the current pass reviewed a later head. | Rotate or remove the prior tip before each new candidate, or persist an explicit phase-specific restore target and validate it against the current reviewed head instead of selecting by file existence. +BLOCKER | high | Failure procedure | The capture records only tracked index and worktree diffs, discards the object produced by `git stash create`, and then treats two empty patches as proof that nothing moved; untracked and ignored state is absent, including ignored closing-message and recovery files that the close reads. | A hook can change a closure input during a failed commit while both patches remain empty, causing the plan to retry or skip a pass on altered state and potentially publish malformed or stale evidence. | Snapshot and compare every closure input plus complete status including untracked paths before and after the act, preserve every observed delta, and never infer that nothing moved from the two tracked diffs alone. +MAJOR | high | Task 0 step 3 and Task 14 step 1 | The parity loops run `diff` without capturing and classifying its status; status 1 is an expected content difference, status greater than 1 is an execution or input failure, and later iterations or commands can overwrite either status after a site record has already been written. | Baseline or final parity evidence can certify a site whose comparison never completed, allowing prompt-copy drift to reach Gate B and closure. | Capture every diff status immediately, accept only 0 or 1, abort on any greater status, and append a site's evidence only after that classification; apply the same rule to Task 0's single-line squash comparison. +MAJOR | high | Close procedure | The deviation rule says every one of the six conditions is checked before moving `HEAD`, but condition 3 includes committing the findings and condition 6 is explicitly checked after the closing commit, so this instruction conflicts with the Failure procedure and Task 15's post-commit recovery route. | An executor following the local deviation rule can stop with a rejected record commit or closing commit still at `HEAD`, or with the soft-reset index still active, instead of restoring the known tip required for a valid retry or parked state. | Split pre-move preconditions from post-move postconditions and direct every failed post-move check explicitly through the Failure procedure with its validated phase-specific restore target. +END OF FINDINGS (5 total) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index aeacca2..eb8cad0 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -310,9 +310,19 @@ files are removed. commit must not share a command string with it, or the hook reads the close as cycle-internal and carries this cycle's count and fingerprint into the next. -**On deviation at any of the six.** Stop **before** moving `HEAD`. Every one of these is checkable -while the repository is still in a state the plan understands, and that is the whole reason they -come before the reset. +**Which of the six are preconditions and which are postconditions**, because they do not all come +before a move and an earlier draft said they did: + +- **1, 2 and 5 are preconditions** — checkable while nothing has moved. **On deviation, stop; no + restoration is owed**, because nothing was changed. +- **3, 4 and 6 straddle or follow a move** — 3 makes the record commit, 4 reads the state it left, + 6 reads the state after the closing commit. **On deviation, run `## The four procedures` · + Failure** against the phase's restore target: `loop-rule-reviewed-head` for a rejection at 3 or 4, + `loop-rule-reviewed-tip` for one at 6. + +**Stopping with a rejected commit at `HEAD`, or with the soft-reset index still live, is not an +outcome this procedure allows** — that is the state a retry cannot start from, and it is exactly +what "stop before moving `HEAD`" produced when applied to a check that runs after one. ### Failure — the closing act did not complete @@ -331,20 +341,39 @@ restore point — because it rewrites the index from the target commit and leave is the worktree half. ```bash -git stash create > /dev/null 2>&1 || true # no-op on an empty delta git diff --cached > .context/loop-rule-failed-index.patch git diff > .context/loop-rule-failed-worktree.patch +git status --porcelain --untracked-files=all --ignored > .context/loop-rule-failed-status.txt +for f in .context/loop-rule-closing-msg .context/loop-rule-base \ + .context/loop-rule-reviewed-head .context/loop-rule-reviewed-tip; do + test -e "$f" && cksum "$f" +done > .context/loop-rule-failed-inputs.txt ``` -**Both patches are written even when empty**, so "nothing was left behind" is a recorded -observation rather than an absent file. Then restore, and report from the patches rather than from -`git status` alone. - -**The restore target is whichever tip exists.** After 8a's commit that is -`.context/loop-rule-reviewed-tip`; **before it** — a record commit that failed, or landed and then -failed its changed-path or clean-tree check — that file does not exist yet and the target is -`.context/loop-rule-reviewed-head`, which every Gate-B call wrote. **An 8a rejection is a rejection -like any other and owes the same restoration**, which an earlier draft's bare exits did not give it. +**All four files are written even when empty**, so "nothing was left behind" is a recorded +observation rather than an absent file. + +**Two empty patches do not prove nothing moved, and an earlier draft said they did.** They cover +**tracked** content only. A hook can create or rewrite an **untracked or ignored** path — and the +closing message, the recovery base and both tip files are ignored, and the close reads every one of +them — so the status listing with `--untracked-files=all --ignored` and the checksums of the closure +inputs are what make "nothing moved" a statement about the things a condition is actually read +from. Compare them against the same four before the next attempt. + +**The restore target is chosen by phase, not by which file happens to exist.** After 8a's commit it +is `.context/loop-rule-reviewed-tip`; before it — a record commit that failed, or landed and then +failed a postcondition — it is `.context/loop-rule-reviewed-head`, which every Gate-B call wrote. +**An 8a rejection is a rejection like any other and owes the same restoration**, which an earlier +draft's bare exits did not give it. + +**"Whichever tip exists" was wrong, and the second candidate is why.** After an 8b failure whose +repair earns another pass, the *previous* candidate's `loop-rule-reviewed-tip` is still on disk when +the new 8a starts — so an 8a failure would select it and rewind past the newly reviewed repair and +its evidence. **Step 7 deletes `loop-rule-reviewed-tip` when it writes a new +`loop-rule-reviewed-head`**, so before 8a's commit the file is absent by construction rather than by +luck; and whichever target is chosen, **assert it is an ancestor of `HEAD` and that its own +`loop-rule-reviewed-head` matches the head this candidate was issued against** before restoring to +it. **How success is recognised.** The repository is back at that tip, the two patch files describe what the attempt left, and the recovery base and reviewed-head files still exist. @@ -610,7 +639,12 @@ else git rev-parse HEAD > .context/loop-rule-base fi BASE=$(cat .context/loop-rule-base) -test -n "$BASE" || { echo "BASE empty"; exit 1; } +# Not "non-empty": a symbolic value such as HEAD passes every check below and +# then RESOLVES DIFFERENTLY as WIP commits accrue, moving the reviewed range, +# the parent counts and the final reset with the branch. +case "$BASE" in [0-9a-f][0-9a-f][0-9a-f][0-9a-f][0-9a-f][0-9a-f][0-9a-f][0-9a-f]*) ;; *) echo "recorded base is not an object name: $BASE"; exit 1 ;; esac +test "${#BASE}" -eq 40 || { echo "recorded base is not a full 40-character object name"; exit 1; } +test "$(git rev-parse --verify "$BASE^{commit}")" = "$BASE" || { echo "recorded base does not resolve to itself as a commit"; exit 1; } git merge-base --is-ancestor "$BASE" HEAD || { echo "recorded base is NOT an ancestor of HEAD — stale"; exit 1; } git log --oneline "$BASE"..HEAD # expect nothing, or only this run's WIP: commits ``` @@ -836,15 +870,21 @@ while IFS=$(printf '\t') read -r s e; do ee=$(printf '%s' "$e" | sed 's/[][\.*^$\/]/\\&/g') a=$(sed -n "/$se/,/$ee/p" "$C_SRC"); b=$(sed -n "/$se/,/$ee/p" "$W_SRC") test -n "$a" && test -n "$b" || { echo "empty extraction for: $s"; exit 1; } + # diff exits 0 (same) or 1 (differs) — both are RESULTS. Above 1 is a failure + # to compare, and writing the site record before classifying would certify a + # comparison that never completed. + d=$(diff <(printf '%s\n' "$a") <(printf '%s\n' "$b")); st=$? + test $st -le 1 || { echo "diff failed (status $st) at site: $s"; exit 1; } printf 'site\t%s\t%s\n' "$s" "$e" >> .context/loop-rule-baseline-diff.tmp - diff <(printf '%s\n' "$a") <(printf '%s\n' "$b") >> .context/loop-rule-baseline-diff.tmp + printf '%s\n' "$d" >> .context/loop-rule-baseline-diff.tmp done < .context/loop-rule-sites || exit 1 # The squash-carry sentence is ONE line and must not go through the loop. +d=$(diff <(grep -F 'On squash-merge, copy every evidence entry' "$C_SRC") \ + <(grep -F 'On squash-merge, copy every evidence entry' "$W_SRC")); st=$? +test $st -le 1 || { echo "diff failed (status $st) at the squash-carry site"; exit 1; } printf 'site\t%s\t%s\n' 'On squash-merge' 'On squash-merge' >> .context/loop-rule-baseline-diff.tmp -diff <(grep -F 'On squash-merge, copy every evidence entry' "$C_SRC") \ - <(grep -F 'On squash-merge, copy every evidence entry' "$W_SRC") \ - >> .context/loop-rule-baseline-diff.tmp +printf '%s\n' "$d" >> .context/loop-rule-baseline-diff.tmp mv .context/loop-rule-baseline-diff.tmp .context/loop-rule-baseline-diff.txt cat .context/loop-rule-baseline-diff.txt # READ IT WHOLE ``` @@ -2160,7 +2200,8 @@ while IFS=$(printf '\t') read -r s e; do test -n "$a" && test -n "$b" || { echo "empty extraction for '$s' — region not found"; exit 1; } printf '%s\n' "$a" > .context/loop-rule-a.txt printf '%s\n' "$b" > .context/loop-rule-b.txt - diff .context/loop-rule-a.txt .context/loop-rule-b.txt + diff .context/loop-rule-a.txt .context/loop-rule-b.txt; st=$? + test $st -le 1 || { echo "diff failed (status $st) at site: $s"; exit 1; } done < .context/loop-rule-changed-sites ``` @@ -2446,6 +2487,7 @@ the next pass against exactly that value: ```bash # Before committing a fix, re-run what the fix could have broken. git add -A && git commit -m "WIP: fix " # or: "WIP: pass records" where no repair was owed +rm -f .context/loop-rule-reviewed-tip # the previous candidate's restore point git rev-parse HEAD > .context/loop-rule-reviewed-head # the head the NEXT call is issued against cat .context/loop-rule-reviewed-head ``` From d1ec3cc16aca536636c2a12cc3b299610e9b2b4e Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Tue, 15 Sep 2026 11:15:33 +0200 Subject: [PATCH 131/181] docs(context): record passes 21-23 and the method change in the working record Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md | 8 ++++++-- 1 file changed, 6 insertions(+), 2 deletions(-) diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md index d269016..df548ec 100644 --- a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md @@ -13,7 +13,7 @@ Nothing depends on it; the pass files and the repo are authoritative where this ## Resume here -**Next action:** Gate-A plan pass 21 against `3111985`. The prompt is +**Next action:** Gate-A plan pass 24 against `b8b433d`. The prompt is `.context/gate-a-plan-prompt.md`; substitute `__SHA__` and `__P__`, precheck per its header, delete the target file, confirm it is gone. @@ -90,7 +90,11 @@ worth checking before a pass rather than after. | 18 | 6d276fa | 9→**4** | 2→**0** | 3→**3** | yes | one tell. A **fifth record shape, `span`**; both scratch artifacts open with a `base` line; the hook ships **ten** prompt bodies, not seven | | 19 | 76cd2ce | 4→**3** | 0→**0** | 3→**2** | yes | one tell. Nine editing tasks record into the plan and staged only prompts — the staging rule was a stale task-number list. Task 0 step 3 selects its source by the re-entry rule | | 20 | 9aa1178 | 3→**4** | 0→**2** | 2→**2** | yes | three-tell stop. **F11/F12/F13 each ended just before their item's first changed word** and would have survived a correct install. Root: the test's subject is the region's **post-edit text**, not the replacement block. All twelve F rows with a declared range re-checked and go to zero; F1–F3 have no range and owe a reading check | -| 21 | — | — | — | — | **not run — NEXT** | against `3111985`. Prompt: `.context/gate-a-plan-prompt.md`, substitute `__SHA__`=`3111985` and `__P__`=`21` | +| 21 | 3111985 | 4→**5** | 2→**2** | 2→**3** | yes | three-tell stop. All five in Task 0 / Task 15. **This is where the loop was surfaced to Daniel**: 39 of the previous 49 findings sat in the two tasks carrying real shell, four of five being repairs to guards earlier passes had added | +| — | — | — | — | — | — | **METHOD CHANGE**, approved by Daniel. Task 0 and Task 15's guarded shell replaced by `## The four procedures` — preparation, close, failure, resume — plus an accounting table classifying all 41 prior guard conditions. **Not a reduction:** one condition deliberately dropped (Task 0 committing nothing), six corrected against pass-21 findings. Commit `1d2f9db` | +| 22 | 1d2f9db | **10** | **5** | **3** | yes | **targeted pass**, charged with: did anything get silently dropped, and are pass 21's five closed. **Three accounting rows were false when written** — 8, 15 and 40 claimed "kept, same shell" for obligations not present everywhere. Also `printf %b` storing `\*` in the site anchors; `--mixed` destroying an index-only change; an 8a rejection with no tip to restore to; §A1's failure transition having **three** routes where one was stated; the `kept` disposition having no route for a passage no span reaches | +| 23 | 9ced53c | 10→**5** | 5→**3** | 3→**2** | yes | **all five in Task 0 / Task 15 again.** A symbolic value in the base file; the restore target chosen by which tip file exists, which rewinds past a second candidate's repair; the failure capture covering tracked content only while the closure inputs are ignored paths; pass 22's diff-status classification written as prose and not as code; the close procedure's "check all six before moving `HEAD`" applied to two checks that run after a move | +| 24 | — | — | — | — | **not run — NEXT** | against `b8b433d`. Prompt: `.context/gate-a-plan-prompt.md`, substitute `__SHA__`=`b8b433d` and `__P__`=`24` | ## Pass-1 report From 892301eb63482a435a83126bdb2f19edf346aa26 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Tue, 15 Sep 2026 12:53:14 +0200 Subject: [PATCH 132/181] docs(plans): apply Gate-A plan pass 24; order the close checks, capture the landed commit MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Two Blockers, five Majors, one Minor. Findings 5 -> 8, Blockers 3 -> 2. Introduced by an earlier repair (pass 23's own split): - The close listed condition 5 (message complete) AFTER the record commit while classifying it a precondition owing no restoration. A literal executor could find an incomplete message with HEAD already moved and stop without restoring. The message check is condition 3 now, before anything moves. - The route decision said "an empty pair means nothing moved" three paragraphs after the same procedure said two empty patches prove nothing. The shortcut is gone, and the comparison is against a PRE-ACT capture that did not exist: status with untracked and ignored paths, plus checksums of the four closure inputs, taken immediately before the act. - Cleanup deleted three scratch files while success claimed "the scratch files are removed". Passes 16-23 added baseref, the pre/post captures, the .src copies and the site lists. All go now, and an `ls` asserts it. Previously untreated execution states: - After a commit LANDS and then fails a postcondition, both tracked patches are empty because the content is in the commit. Nothing recorded the rejected HEAD or restore-target..HEAD, so after the mixed restore there was no account of what the attempt published. Both are captured before restoring. - Task 0 derived line-sharing kept-condition fragments only into the ignored map, while the contract says every pre-existing fragment lives in the table and step 4's committed sweep runs over table rows. Those fragments were outside the sweep entirely. - Step 7b let revalidation replace the evidence entry after the candidate clean pass with no branch back to Gate B, so the closing commit could carry an entry no reviewer saw. A changed entry ends the candidate. Recurred despite an earlier repair: - Step 5 still said step 7b "appends all three" after pass 17 made 7b rebuild the message whole and forbid appending — the stale cross-reference class, and this time the stale referrer was the one the repair created. - Task 11 was told to record observed counts, which pass 19's staging change turned from a transient reading into a durable second numeric authority — exactly what §F refuses, having twice named a count and been wrong. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-plan-om0bdd7udh-pass-24.md | 9 ++ .../2026-09-14-loop-rule-consolidation.md | 100 ++++++++++++++---- 2 files changed, 91 insertions(+), 18 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-24.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-24.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-24.md new file mode 100644 index 0000000..8ca47c5 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-24.md @@ -0,0 +1,9 @@ +BLOCKER | high | Close, lines 285-321 | The ordered six-check sequence puts condition 5 after the findings-file commit, yet the classification calls 5 a precondition checked while nothing has moved and says no restoration is owed on its failure. | A literal executor can detect an incomplete closing message after HEAD moved, then stop without restoring, leaving a rejected record commit at HEAD and the cycle unable to retry from its reviewed state. | Move condition 5 before condition 3 in the authoritative order and keep 3, 4, and 6 as the only move/post-move checks. +BLOCKER | high | Failure, lines 343-397 | The route decision says an empty patch pair proves nothing moved, contradicting lines 356-361 and ignoring untracked or ignored status plus closure-input checksums; those checksums also have no pre-act baseline. | An attempt that mutates an ignored closure input can be classified as unchanged, retried without the pass that input's rule requires, and close on evidence the candidate pass did not cover. | Snapshot status and checksums before each closing act, compare all captured state after failure and repair, and remove the empty-pair shortcut. +MAJOR | high | Failure, lines 329-379 | The capture records only diffs against the current HEAD and never persists the rejected HEAD or a diff from the restore target to it; after a commit lands, both patches can be empty even though the act moved HEAD and published different content or metadata. | After the mixed restore, the plan lacks a durable record of the landed commit needed to explain or recover what the attempt changed, so its stated recovery-success predicate can be met by incomplete evidence. | Before restoring, persist the current HEAD and capture restore-target..HEAD plus index and worktree deltas; include those artifacts in success recognition. +MAJOR | high | Task 0 step 2 and the fragment-table contract, lines 150-180 and 744-765 | Task 0 derives line-sharing kept-condition fragments only into the ignored loop-rule-untouched file, although the contract says every pre-existing preservation fragment lives in the authoritative table and Task 0's committed sweep covers every table row. | Those fragments bypass the pre-install relationship sweep and can be reused after interruption based only on parse and coverage, allowing a wrong fragment to certify a kept condition and feed false-green evidence. | Append each derived cond fragment to the table before the Task 0 sweep and make the scratch map reference that row, or extend the durable sweep and resume validation to verify each cond fragment's single-line, uniqueness, and preservation relationship. +MAJOR | high | Task 11, lines 1904-1932 | The task orders observed grep counts to be recorded in the plan even though approved §F says no count of these assertions appears anywhere because such counts have twice become stale. | Executing the plan adds the exact second numeric authority the approved design rejects, so later file evolution can make the record false while the sweep remains correct. | Remove the count-recording instruction and the plan-fragment-evidence mutation for Task 11; record the examined sites and dispositions without an aggregate count if durable evidence is needed. +MAJOR | high | Task 15 step 5 versus step 7b, lines 2418-2421 and 2544-2552 | Step 5 says step 7b appends the closing records, while step 7b requires rebuilding the file whole and explicitly forbids append. | The executor receives mutually exclusive write semantics; following the earlier instruction on a retry duplicates provenance and curve lines and makes the closing message malformed. | Replace "appends all three" with "rebuilds the complete message whole" and describe the same four ordered record classes as step 7b. +MAJOR | high | Task 15 step 7 and step 7b, lines 2505-2506 and 2544-2562 | Step 7b allows the evidence entry to be replaced after the candidate clean pass but gives no branch back to Gate B when revalidation changes it. | The closing commit can carry an evidence entry different from the verbatim entry the final reviewer judged, violating the governing re-review rule and closing on unreviewed evidence. | Revalidate and finalize the evidence entry before recording the candidate HEAD and issuing the pass; if post-pass revalidation changes it, commit the change and run another candidate pass. +MINOR | high | Close success and Task 15 cleanup, lines 303-305 and 2638-2642 | Success says the scratch files are removed, but the cleanup deletes only base, reviewed-tip, and reviewed-head, leaving closing-msg, baseref, untouched maps, baseline sources, sites, and any failure snapshots. | The documented terminal state is false and later runs inherit stale cycle artifacts that must be detected or overwritten piecemeal. | Name the exact files intentionally retained and narrow the success claim, or delete every loop-rule scratch artifact only after success. +END OF FINDINGS (8 total) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index eb8cad0..fd61b60 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -291,13 +291,15 @@ an unapproved edit to a spec that the gates already closed. 2. **The only thing dirty is the candidate pass's own findings files** — read from every porcelain record, not from two status codes. Nothing else may be uncommitted: every record this plan collects was committed before the pass was issued. -3. **Those files are committed** in their own invocation, and that commit **changes exactly those +3. **The closing message is complete** — rebuilt whole, exactly one provenance line, exactly one + curve, every owed evidence entry, and either the applicable human-exception records or + `Human exceptions: none`. **This is checked here, before anything moves**, because an incomplete + message discovered after the record commit leaves `HEAD` moved for a reason no restoration was + owed for. +4. **Those files are committed** in their own invocation, and that commit **changes exactly those paths**. Its parent is the reviewed head. Record the resulting commit as the **restore point**, in `.context/loop-rule-reviewed-tip`. -4. **The tree is clean**, and `HEAD` is still the restore point, when the closing invocation begins. -5. **The closing message is complete** — rebuilt whole, exactly one provenance line, exactly one - curve, every owed evidence entry, and either the applicable human-exception records or - `Human exceptions: none`. +5. **The tree is clean**, and `HEAD` is still the restore point, when the closing invocation begins. 6. **After the closing commit**: its subject is not a snapshot, and the tree is clean. **How success is recognised.** A single commit at `HEAD` whose subject is the real message, whose @@ -313,11 +315,14 @@ carries this cycle's count and fingerprint into the next. **Which of the six are preconditions and which are postconditions**, because they do not all come before a move and an earlier draft said they did: -- **1, 2 and 5 are preconditions** — checkable while nothing has moved. **On deviation, stop; no - restoration is owed**, because nothing was changed. -- **3, 4 and 6 straddle or follow a move** — 3 makes the record commit, 4 reads the state it left, +- **1, 2 and 3 are preconditions** — checkable while nothing has moved. **On deviation, stop; no + restoration is owed**, because nothing was changed. **The message check is among them + deliberately**: an earlier draft classified it as a precondition while listing it *after* the + record commit, so a literal executor could find an incomplete message with `HEAD` already moved + and then stop without restoring. +- **4, 5 and 6 straddle or follow a move** — 4 makes the record commit, 5 reads the state it left, 6 reads the state after the closing commit. **On deviation, run `## The four procedures` · - Failure** against the phase's restore target: `loop-rule-reviewed-head` for a rejection at 3 or 4, + Failure** against the phase's restore target: `loop-rule-reviewed-head` for a rejection at 4 or 5, `loop-rule-reviewed-tip` for one at 6. **Stopping with a rejected commit at `HEAD`, or with the soft-reset index still live, is not an @@ -340,7 +345,24 @@ restore point — because it rewrites the index from the target commit and leave `git status` to report. Saying "`--mixed` preserves the delta" was an overclaim; what it preserves is the worktree half. +**There is a *before* half, and it runs as part of the close, not here.** A capture with nothing to +compare against says only what exists, never what changed. **Immediately before each closing act:** + ```bash +git status --porcelain --untracked-files=all --ignored > .context/loop-rule-pre-status.txt +for f in .context/loop-rule-closing-msg .context/loop-rule-base \ + .context/loop-rule-reviewed-head .context/loop-rule-reviewed-tip; do + test -e "$f" && cksum "$f" +done > .context/loop-rule-pre-inputs.txt +``` + +**And after a failed act, before restoring anything:** + +```bash +git rev-parse HEAD > .context/loop-rule-failed-head.txt # the rejected commit, if one landed +TGT=$(cat .context/loop-rule-reviewed-tip 2>/dev/null || cat .context/loop-rule-reviewed-head) +git log --oneline "$TGT"..HEAD > .context/loop-rule-failed-commits.txt +git diff "$TGT" HEAD > .context/loop-rule-failed-landed.patch git diff --cached > .context/loop-rule-failed-index.patch git diff > .context/loop-rule-failed-worktree.patch git status --porcelain --untracked-files=all --ignored > .context/loop-rule-failed-status.txt @@ -350,7 +372,13 @@ for f in .context/loop-rule-closing-msg .context/loop-rule-base \ done > .context/loop-rule-failed-inputs.txt ``` -**All four files are written even when empty**, so "nothing was left behind" is a recorded +**The landed patch is the half an earlier draft had no way to see.** After a commit lands and then +fails a postcondition, `git diff --cached` and `git diff` are both **empty** — the content is in the +commit — so the two tracked patches describe nothing while `HEAD` has moved and published different +content or a different message. `restore-target..HEAD` is what records it, and after the mixed +restore that record is the only account of what the attempt actually did. + +**Every one of these files is written even when empty**, so "nothing was left behind" is a recorded observation rather than an absent file. **Two empty patches do not prove nothing moved, and an earlier draft said they did.** They cover @@ -375,8 +403,10 @@ luck; and whichever target is chosen, **assert it is an ancestor of `HEAD` and t `loop-rule-reviewed-head` matches the head this candidate was issued against** before restoring to it. -**How success is recognised.** The repository is back at that tip, the two patch files describe what -the attempt left, and the recovery base and reviewed-head files still exist. +**How success is recognised.** The repository is back at that tip; the captured set — commit +listing, landed patch, index and worktree patches, status listing and input checksums — describes +what the attempt left, **each compared against the pre-act capture**; and the recovery base and +reviewed-head files still exist. **On deviation.** If the delta cannot be explained, stop and hand it to a person. **Do not retry a close against a state you cannot account for** — the second attempt would carry whatever the first @@ -394,7 +424,11 @@ stands**, then read which route you are on. - **The attempt or its repair moved something a condition is read from** → that condition has changed, **and its own rule decides what it costs**, a further pass included. The cycle is back in the ordering with that pass owed. The two patch files above are how you tell which case you are - in: an empty pair means nothing moved. + in — and **it is not the two tracked patches that decide it.** Compare the *after* capture against + the *before* one: the commit listing, the landed patch, the status listing with untracked and + ignored paths, and the four closure-input checksums. **An empty patch pair proves nothing**, since + a landed commit leaves both empty and a hook rewriting an ignored closure input leaves both empty + too. - **The failure cannot be repaired at all** — a signing key nobody has, a permission nobody can grant — → **surface it and leave the cycle parked**: open, not running, spending no passes, restarted by an explicit later continue. **This route was missing entirely**, and without it a @@ -757,6 +791,12 @@ expected pair is `11`; the shape carries the values rather than assuming th dropped condition recorded here later needs no new format. **Both consumers validate every `cond` row**, not only the `span` rows. +**And every `cond` fragment is appended to the fragment table as well**, under the next free `P` id, +with the map's row citing that id. The table is where every pre-existing fragment lives, and Task 0 +step 4's committed sweep runs the three conditions over **table rows** — a fragment that exists only +in this ignored scratch file is outside that sweep, so a wrong one could certify a kept condition +and be reused after an interruption on the strength of parsing and coverage alone. + **Build the map under `.context/loop-rule-untouched.tmp` and rename it only after the coverage assertion passes.** The baseline artifact is written that way and accounting row 15 claims both are; an earlier draft wrote this one straight to its consumer path, so an interruption between the `base` @@ -1929,7 +1969,7 @@ sees three of them. **Read and update every hit of the bare word by the hook sta reports**, exactly as step 3 requires — and do not turn the observed hit count into a target, for the reason this task's opening gives. -**Record the counts you observe.** They will not match the numbers above if the file has changed; the numbers above are evidence for why no list is kept, not a target. +**Read what you observe; do not record a count.** The numbers above are evidence for why no list is kept, not a target — and §F states no count of these assertions **anywhere**, having twice named one and been wrong. This task stages the plan like every other recording task, so a count written into its evidence would be exactly the second numeric authority the approved design rejects: durable, and false the next time the file grows an assertion. **Record the sites examined and what each became**, which stays true however many there are. - [ ] **Step 2: Update the three `expected_ctx` and three `expected_msg` assignments** @@ -2417,8 +2457,12 @@ in the WIP body too if a mid-cycle reader would want it, but the file is the cop **The file is completed at step 7b, not here.** The evidence entry can be drafted now, but the **per-pass curve is not known until the Gate-B loop ends**, and the **provenance line** and any -**human-exception record** belong beside it. **Step 7b appends all three and revalidates the -entry** — this step opens the file, step 7b closes it, and step 8 commits it. +**human-exception record** belong beside it. **Step 7b rebuilds the complete message whole** — +provenance line, curve, any human-exception record or `Human exceptions: none`, and the revalidated +evidence entry — **and does not append to this draft.** This step opens the file, step 7b replaces +it, and step 8 commits it. **"Appends" was the earlier wording and it is the opposite of what 7b +requires**: on a second candidate close, appending writes a second provenance line and a second +curve, which the one-of-each grammar refuses. It names: the battery run; **every pair this plan built, with its counts in each copy and each tree, every presence check beside them, and every absence check with its two counts**; the §6 parity diff and the `b11`/`b13` equivalence result; and the next-state table's location plus its row count. @@ -2559,7 +2603,12 @@ The message carries, in this order: Blockers … Majors …`, which is why this cannot be written at step 5: the counts do not exist until the loop ends; 3. any **human-exception record**, and beside it the skip reason if a cycle was skipped; -4. the **revalidated evidence entry**, replacing step 5's draft if revalidation changed it. +4. the **revalidated evidence entry**, replacing step 5's draft if revalidation changed it — and + **if it changed, this candidate is over.** The final reviewer judged the entry it was handed + verbatim; a different entry in the closing commit is evidence no pass covered. Commit the change, + resolve and record the new head, and issue another candidate pass. **Revalidate before recording + the candidate head and issuing the pass**, so that in the ordinary case this branch is never + reached. **Items 1, 2 and 4 are owed unconditionally; item 3 is owed only where such a record exists.** Confirm the file carries the three, and either the applicable exception records or **the literal @@ -2638,9 +2687,24 @@ git status --porcelain # expect empty **Both clean, and only then:** ```bash -rm -f .context/loop-rule-base .context/loop-rule-reviewed-tip .context/loop-rule-reviewed-head +rm -f .context/loop-rule-base .context/loop-rule-reviewed-tip .context/loop-rule-reviewed-head \ + .context/loop-rule-baseref .context/loop-rule-closing-msg \ + .context/loop-rule-untouched .context/loop-rule-baseline-diff.txt \ + .context/loop-rule-sites .context/loop-rule-changed-sites \ + .context/loop-rule-c.src .context/loop-rule-w.src \ + .context/loop-rule-a.txt .context/loop-rule-b.txt \ + .context/loop-rule-pre-status.txt .context/loop-rule-pre-inputs.txt \ + .context/loop-rule-failed-*.txt .context/loop-rule-failed-*.patch +ls .context/loop-rule-* 2>/dev/null && { echo "cycle scratch survives the close — list it above"; exit 1; } ``` +**Every `loop-rule-*` scratch file goes, and the `ls` is what makes "the scratch files are removed" +true rather than asserted.** An earlier draft deleted three of them and claimed the terminal state, +leaving a later run to inherit a closing message, a baseref, an untouched map and any failure +snapshots — each of which some check then has to detect or overwrite piecemeal. **The plan's records +are not among these**: they live in the plan and in `.context/codex-reviews/`, both tracked, both +already in the closing commit. + **`reset --soft` stages committed content only.** The prompt-standards result, the completeness sweep, the next-state table, the divergence list, the equivalence result and the fragment evidence all land in this plan, and `.context/codex-reviews/` is tracked; **anything uncommitted when the From f795bd2ad9a7cb3622389176572738dcee9bb713 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Tue, 15 Sep 2026 13:05:19 +0200 Subject: [PATCH 133/181] docs(plans): apply Gate-A plan pass 25; the capture gets a function and a home MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Two Blockers, two Majors, one Minor. Findings 8 -> 5, Blockers 2 -> 2. APPLIED BUT NOT YET REVIEWED — no pass has read these repairs. Introduced by pass 24's own repair: - `test -e "$f" && cksum "$f"` as a loop body makes the block's status that of its LAST iteration. Before 8a the reviewed tip is deliberately absent, so the pre-act capture returned 1 and a correct candidate could not enter the close; and a checksum that really failed on an earlier input was masked by a later success. Absence is recorded as `absent`, a failure returns 1 immediately. - Both captures wrote into ignored loop-rule-* files in the worktree being compared, so the two status listings differed because of the capture itself even on a no-op failure. They live under .context/loop-rule-capture/ now and both listings exclude it. Previously untreated execution state: - The pre-act capture snapshotted path statuses and four checksums, not the bytes. A hook that rewrites a staged findings file IN PLACE leaves its porcelain status unchanged, so the after-capture's index patch holds only the post-hook bytes and the reviewed artifact is gone with nothing to compare against. Every dirty path is checksummed before the act. Recurred despite an earlier repair — and the repair caused it: - Kept c1-c3 have no observation again. Pass 10 found exactly this and fixed it with a per-task note; pass 22 deleted that note along with the other id lists and left the disposition table's third route to cover it — but no task's walk is bounded to reach them. They sit before Task 4's block, passage (c) is in no untouched region, and Task 7 starts later, so they fall between two boundaries. The rule is now that the task opening a passage carries its kept prefix. - Task 11's sweep record had no shape. Pass 24 told it to record the sites examined rather than a count; the evidence section admits five shapes and none fits a site sweep. There is a sixth, and it carries no total, which is what §F refuses. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../gate-a-plan-om0bdd7udh-pass-25.md | 6 ++ .../2026-09-14-loop-rule-consolidation.md | 98 ++++++++++++++----- 2 files changed, 80 insertions(+), 24 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-25.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-25.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-25.md new file mode 100644 index 0000000..d36b537 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-25.md @@ -0,0 +1,6 @@ +BLOCKER | high | Failure procedure, pre-act and failed-input captures | Both checksum loops finish on `.context/loop-rule-reviewed-tip` using `test -e "$f" && cksum "$f"`; before 8a that tip is deliberately absent, so the pre-act block deterministically returns status 1, while a checksum failure on an earlier input can be masked by a later successful iteration. | A correct candidate cannot enter 8a through a successful before-capture, and even where the block continues its input evidence cannot distinguish an absent file from a failed checksum. | Use an explicit `if` per path that records either a checksum or an `absent` sentinel, fails immediately on a checksum error, and makes the completed loop return 0. +BLOCKER | high | Failure procedure, pre-act capture | The before half snapshots only path statuses and four closure-input checksums; it does not preserve the staged findings bytes or the pre-act index and worktree deltas that the after half claims to compare. A failing hook can rewrite and stage a findings file while leaving its porcelain status unchanged, after which the captured index patch contains only the post-hook bytes and the original review artifact is gone. | The recovery reader cannot determine whether a condition input changed and can retry or commit a mutated Gate-B findings artifact without the further pass that §A requires when an attempt moves such input. | Before each act, persist the pre-act index and worktree patches plus content and mode snapshots or checksums for every dirty path, especially the candidate findings files, and compare each after-artifact to that named baseline before selecting a recovery route. +MINOR | high | Failure procedure, status capture | Each status command redirects into an ignored `loop-rule-*` file before `git status` runs, and the after half creates several additional ignored failure artifacts before taking its status snapshot; therefore the two raw listings differ because of the capture procedure itself even when the closing act moved nothing, with no exclusion or normalized expected result stated. | A transient no-op failure cannot be classified from the prescribed status comparison alone, so an executor must invent an exception and may unnecessarily charge another pass or park the close. | Write the captures outside the compared worktree or remove one explicit set of procedure-owned paths from both snapshots, then state the expected normalized comparison. +MAJOR | high | Task 4 steps 1 and 5 | The disposition table requires kept `c1` through `c3` to receive per-condition preservation counts because no untouched span reaches passage (c), but Task 4's pre-install walk is bounded to the replacement beginning at `c4` and Task 7 starts at the later surfacing block; Task 4 step 5 merely reads `c1` through `c3` and explicitly says that walk is not an observation. | A mis-scoped edit can delete or alter those curve-reading conditions identically in both prompt copies while every pair, parity check, and recorded preservation check passes, so the promised 135-condition coverage is incomplete. | Add a bounded untouched span for the passage-(c) prefix or derive and record one preservation check per condition before Task 4 installs, and include those records in the Task 14 and Task 15 reruns. +MAJOR | high | Task 11 recording and Task 15 step 5 | Task 11 requires `sites examined and what each became` to be recorded in its fragment-evidence subsection, but the subsection declares that every observation must use one of five shapes and none represents a site-sweep result; Task 15's closing-entry inventory likewise consumes the five shapes and named reader sections but never names this sweep record. | The executor must either violate the evidence grammar or omit the durable proof that every stale test assertion, label, and comment was examined, so re-entry and the closing evidence cannot distinguish a complete sweep from a partial one. | Define a sixth idempotent sweep-record shape and include it in Task 15's assembly and complete reruns, or give Task 11 a dedicated output section with an explicit schema that Task 15 consumes. +END OF FINDINGS (5 total) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index fd61b60..b3de3ce 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -346,32 +346,66 @@ restore point — because it rewrites the index from the target commit and leave is the worktree half. **There is a *before* half, and it runs as part of the close, not here.** A capture with nothing to -compare against says only what exists, never what changed. **Immediately before each closing act:** +compare against says only what exists, never what changed. **Both halves write into +`.context/loop-rule-capture/`, and both listings exclude that directory** — otherwise the act of +capturing changes the thing being compared, and two listings differ over the procedure's own files +even when the closing act moved nothing. ```bash -git status --porcelain --untracked-files=all --ignored > .context/loop-rule-pre-status.txt -for f in .context/loop-rule-closing-msg .context/loop-rule-base \ - .context/loop-rule-reviewed-head .context/loop-rule-reviewed-tip; do - test -e "$f" && cksum "$f" -done > .context/loop-rule-pre-inputs.txt +capture() { # $1 = phase name: pre | failed + d=.context/loop-rule-capture; mkdir -p "$d" + git status --porcelain --untracked-files=all --ignored \ + | grep -v ' \.context/loop-rule-capture/' > "$d/$1-status.txt" + git diff --cached > "$d/$1-index.patch" + git diff > "$d/$1-worktree.patch" + # Every dirty path's bytes, so a hook that rewrites a staged findings file + # in place — leaving its porcelain status unchanged — is still detectable. + : > "$d/$1-dirty.txt" + git status --porcelain --untracked-files=all -z | tr '\0' '\n' | sed -n 's/^.\{3\}//p' | + while IFS= read -r f; do + case "$f" in .context/loop-rule-capture/*) continue ;; esac + if [ -f "$f" ]; then printf '%s\t%s\n' "$f" "$(cksum < "$f")" + else printf '%s\tabsent-or-not-a-regular-file\n' "$f"; fi + done >> "$d/$1-dirty.txt" + # Closure inputs, each recorded present-with-checksum or absent. An absent + # file is a RESULT, not a failure: before 8a the reviewed tip does not exist. + : > "$d/$1-inputs.txt" + for f in .context/loop-rule-closing-msg .context/loop-rule-base \ + .context/loop-rule-reviewed-head .context/loop-rule-reviewed-tip; do + if [ -e "$f" ]; then + c=$(cksum < "$f") || { echo "cksum failed on $f"; return 1; } + printf '%s\t%s\n' "$f" "$c" >> "$d/$1-inputs.txt" + else + printf '%s\tabsent\n' "$f" >> "$d/$1-inputs.txt" + fi + done + return 0 +} ``` +**Immediately before each closing act:** `capture pre || exit 1`. + **And after a failed act, before restoring anything:** ```bash -git rev-parse HEAD > .context/loop-rule-failed-head.txt # the rejected commit, if one landed +d=.context/loop-rule-capture +git rev-parse HEAD > "$d/failed-head.txt" # the rejected commit, if one landed TGT=$(cat .context/loop-rule-reviewed-tip 2>/dev/null || cat .context/loop-rule-reviewed-head) -git log --oneline "$TGT"..HEAD > .context/loop-rule-failed-commits.txt -git diff "$TGT" HEAD > .context/loop-rule-failed-landed.patch -git diff --cached > .context/loop-rule-failed-index.patch -git diff > .context/loop-rule-failed-worktree.patch -git status --porcelain --untracked-files=all --ignored > .context/loop-rule-failed-status.txt -for f in .context/loop-rule-closing-msg .context/loop-rule-base \ - .context/loop-rule-reviewed-head .context/loop-rule-reviewed-tip; do - test -e "$f" && cksum "$f" -done > .context/loop-rule-failed-inputs.txt +git log --oneline "$TGT"..HEAD > "$d/failed-commits.txt" +git diff "$TGT" HEAD > "$d/failed-landed.patch" +capture failed || exit 1 ``` +**Three things that loop gets right and an earlier draft did not.** `test -e "$f" && cksum "$f"` as +a loop body makes the block's status that of its **last** iteration — so before 8a, where the +reviewed tip is deliberately absent, the whole before-capture returned 1 and a correct candidate +could not enter the close; and a checksum that actually failed on an earlier input was masked by a +later success. Recording `absent` explicitly makes absence a result rather than an error. And +**checksumming every dirty path, not only the four inputs**, is what catches a hook that rewrites a +staged findings file in place: the porcelain status is unchanged, the after-capture's index patch +holds only the post-hook bytes, and without a before-image of those bytes the original review +artifact is simply gone. + **The landed patch is the half an earlier draft had no way to see.** After a commit lands and then fails a postcondition, `git diff --cached` and `git diff` are both **empty** — the content is in the commit — so the two tracked patches describe nothing while `HEAD` has moved and published different @@ -425,8 +459,9 @@ stands**, then read which route you are on. changed, **and its own rule decides what it costs**, a further pass included. The cycle is back in the ordering with that pass owed. The two patch files above are how you tell which case you are in — and **it is not the two tracked patches that decide it.** Compare the *after* capture against - the *before* one: the commit listing, the landed patch, the status listing with untracked and - ignored paths, and the four closure-input checksums. **An empty patch pair proves nothing**, since + the *before* one, file by file: the status listings, the index and worktree patches, the + per-dirty-path checksums, and the closure-input records. The commit listing and the landed patch + have no before-image by construction — a non-empty either one means the act moved `HEAD`. **An empty patch pair proves nothing**, since a landed commit leaves both empty and a hook rewriting an ignored closure input leaves both empty too. - **The failure cannot be repaired at all** — a signing key nobody has, a permission nobody can @@ -1210,7 +1245,16 @@ is **preserved in target §C's fenced replacement**, so its old-count could neve **Run the pre-install half of the disposition procedure over the block this task replaces** — the clearly-stuck block, **ending before the Surfacing paragraph**, which is Task 7's -`c18`-and-surfacing block and whose conditions Task 7 derives. Passage (c) is edited by two tasks; +`c18`-and-surfacing block and whose conditions Task 7 derives. + +**Plus the passage's kept prefix, which belongs to no block and would otherwise belong to no task.** +`c1`–`c3` sit before this task's replacement, passage (c) is not one of Task 0's five untouched +regions, and Task 7's blocks start later — so the curve-reading premises fall between two boundaries +and end up with no observation at all. **The task that opens a passage carries that passage's kept +prefix**, and this is that task: derive one preservation fragment per condition and run it +`parent=1 worktree=1` in each copy. **Step 5's walk is not that observation** and says so. Without +these a mis-scoped edit can alter the curve or coverage sentences identically in both copies while +every pair, preservation count and parity check passes. Passage (c) is edited by two tasks; walking the whole passage here makes this task append rows it cannot discharge after its own install, and makes both tasks append a row for the same condition. Derive each from the live text now and append it under the next free `P` id, referring to the rows afterwards by the condition @@ -2692,9 +2736,8 @@ rm -f .context/loop-rule-base .context/loop-rule-reviewed-tip .context/loop-rule .context/loop-rule-untouched .context/loop-rule-baseline-diff.txt \ .context/loop-rule-sites .context/loop-rule-changed-sites \ .context/loop-rule-c.src .context/loop-rule-w.src \ - .context/loop-rule-a.txt .context/loop-rule-b.txt \ - .context/loop-rule-pre-status.txt .context/loop-rule-pre-inputs.txt \ - .context/loop-rule-failed-*.txt .context/loop-rule-failed-*.patch + .context/loop-rule-a.txt .context/loop-rule-b.txt +rm -rf .context/loop-rule-capture ls .context/loop-rule-* 2>/dev/null && { echo "cycle scratch survives the close — list it above"; exit 1; } ``` @@ -2775,7 +2818,7 @@ idempotently on re-run, touching no other. **The pre-existing fragments live in never here**; this section records what each observation actually returned, and it is what Task 15 step 5 reads to assemble the closing evidence entry.* -***Five record shapes, because the classes do not return the same number of values.** Every +***Six record shapes, because the classes do not return the same number of values.** Every observation this plan makes is one of them, and a shape that fits only pairs is how a required count gets run and then vanishes from both the plan and the closing evidence:* @@ -2785,6 +2828,7 @@ presence worktree parent # add-only, and a absence parent worktree # dropped, and a moved condition's source preservation parent worktree # carried, and kept where no span holds it span # an untouched range, parent vs worktree +sweep # one examined site of a reader-led sweep ``` *The **fifth shape is for ranges, not conditions**: an untouched span is two bounded extracts @@ -2793,7 +2837,13 @@ cannot express it. Its `` is `no difference` or the difference itself way, because a span that was never run and a span that compared equal are otherwise the same record. Task 15 reads this shape alongside the other four.* +*The **sweep** shape is the one that carries no count: Task 11's sweep records **one line per site +examined** and what each became, because §F refuses a count of these assertions anywhere and a +number here would be the durable second authority it refuses. A complete sweep is distinguishable +from a partial one by the sites listed, not by a total. Task 15's closing entry reads this shape +with the other five.* + *A **moved** condition therefore contributes two lines — one `absence` at its source, one `presence` at its destination — and both are required for it to count as observed. **Every one of -the five shapes goes into the closing evidence entry**; naming only pairs and presence leaves the +the six shapes goes into the closing evidence entry**; naming only pairs and presence leaves the absences, preservations and span results run but unrecorded.* From a772083c40beb726ce5dc80098830d8c1fe14042 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Tue, 15 Sep 2026 13:05:43 +0200 Subject: [PATCH 134/181] docs(context): record passes 24-25 and the checkpoint in the working record Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md | 6 ++++-- 1 file changed, 4 insertions(+), 2 deletions(-) diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md index df548ec..a2b8c52 100644 --- a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md @@ -13,7 +13,7 @@ Nothing depends on it; the pass files and the repo are authoritative where this ## Resume here -**Next action:** Gate-A plan pass 24 against `b8b433d`. The prompt is +**Next action:** awaiting Daniel. The two-pass allowance (24, 25) is spent and the cycle did not close. Pass 25's repairs are applied at `f795bd2` and no pass has read them. The prompt is `.context/gate-a-plan-prompt.md`; substitute `__SHA__` and `__P__`, precheck per its header, delete the target file, confirm it is gone. @@ -94,7 +94,9 @@ worth checking before a pass rather than after. | — | — | — | — | — | — | **METHOD CHANGE**, approved by Daniel. Task 0 and Task 15's guarded shell replaced by `## The four procedures` — preparation, close, failure, resume — plus an accounting table classifying all 41 prior guard conditions. **Not a reduction:** one condition deliberately dropped (Task 0 committing nothing), six corrected against pass-21 findings. Commit `1d2f9db` | | 22 | 1d2f9db | **10** | **5** | **3** | yes | **targeted pass**, charged with: did anything get silently dropped, and are pass 21's five closed. **Three accounting rows were false when written** — 8, 15 and 40 claimed "kept, same shell" for obligations not present everywhere. Also `printf %b` storing `\*` in the site anchors; `--mixed` destroying an index-only change; an 8a rejection with no tip to restore to; §A1's failure transition having **three** routes where one was stated; the `kept` disposition having no route for a passage no span reaches | | 23 | 9ced53c | 10→**5** | 5→**3** | 3→**2** | yes | **all five in Task 0 / Task 15 again.** A symbolic value in the base file; the restore target chosen by which tip file exists, which rewinds past a second candidate's repair; the failure capture covering tracked content only while the closure inputs are ignored paths; pass 22's diff-status classification written as prose and not as code; the close procedure's "check all six before moving `HEAD`" applied to two checks that run after a move | -| 24 | — | — | — | — | **not run — NEXT** | against `b8b433d`. Prompt: `.context/gate-a-plan-prompt.md`, substitute `__SHA__`=`b8b433d` and `__P__`=`24` | +| 24 | b8b433d | 5→**8** | 3→**2** | 2→**5** | yes | close listed "message complete" *after* the record commit while calling it a precondition; the route said "an empty pair means nothing moved" three paragraphs after the same procedure said the opposite, and the checksums had no pre-act baseline; **after a commit lands and then fails a postcondition both tracked patches are empty** and nothing recorded the rejected `HEAD`; Task 0's `cond` fragments lived only in the ignored map, outside step 4's sweep; step 5 still said 7b "appends"; Task 11 recorded a count §F refuses | +| 25 | 892301e | 8→**5** | 2→**2** | 5→**2** | yes | `test -e && cksum` in a loop makes the block's status its *last* iteration, so the pre-act capture returned 1 before 8a and a correct candidate could not enter the close; the captures wrote into the worktree they compared; the pre-act capture held no **bytes**, so a hook rewriting a staged findings file in place was undetectable; **kept `c1`–`c3` lost their observation again** — pass 10 fixed it with a per-task note, pass 22 deleted that note with the other id lists, and no task's walk reaches them; Task 11's sweep record had no shape | +| — | — | — | — | — | — | **CHECKPOINT.** Daniel's allowance of two passes (24, 25) is spent. Pass 25's five repairs are **applied and not yet reviewed**. Next action is his call, not another pass | ## Pass-1 report From 05ec1b2026c3aae85450601f1f358d7d140d8d1e Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Tue, 15 Sep 2026 13:49:55 +0200 Subject: [PATCH 135/181] docs(plans): drop the plan-added automatic restoration for a bounded handoff MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Daniel's decision, validated against the repository before applying. VALIDATION. The approved spec (target-text.md:177) says: surface the concrete command failure and "re-establish every closure condition AGAINST THE REPOSITORY AS IT NOW STANDS", then retry, owe a pass, or park. It reads the state; it never asks for it to be rewound. The plan (line 329) additionally forbade stopping with a rejected commit at HEAD or a live soft-reset index, calling that "the state a retry cannot start from" — which is false against the approved text, and was this plan's own implementation choice. From it grew restore targets, phase selection, pre- and post-act captures and their own error handling, which passes 22-25 found defects in four times running. WHAT CHANGES. On a failed closing operation or a failed postcondition the plan now stops every further mutating action and reports: which step failed with the concrete command failure, whether a commit landed, what the tree holds, and which cycle values still exist. No reset, no second commit, no deletion of recovery inputs or review artifacts. It states plainly that stopping is no guarantee the failed operation destroyed nothing, and forbids reporting that nothing was lost. Resume is a decision on the actual state, after which §A's three routes apply unchanged, and a person's help replaces neither review nor evidence. WHAT DOES NOT CHANGE. The success path, the preconditions and postconditions, every review and evidence obligation, and §A's own failure transition. The reviewed-tip file survives as 8b's precondition value — HEAD must still equal it or 8b refuses to reset — and is renamed "closing tip" so no reader takes it for a restore target. Accounting row 40 is marked DELIBERATELY DROPPED, the second of two deletions the table records, with the reason: a guard the plan invented does not have to survive on the strength of existing, provided it takes no governing obligation with it. Not replaced by a generic backup mechanism. The seven findings tied to the removed machinery are reassessed individually in a new table. Four are moot. Three carried a duty that survives and is kept: the report names which cycle values exist rather than assuming them (23-3); it states whether a commit landed, because that is what resume needs first (24-3); and 25-2's undetectable in-place rewrite is answered by admission rather than by machinery. Every other finding from those passes stands and is owed. Not yet reviewed: no pass has read this change. Claude-Session: https://claude.ai/code/session_019keCkbwLtigAWW6wBi3QoA --- .../2026-09-14-loop-rule-consolidation.md | 271 +++++++----------- 1 file changed, 102 insertions(+), 169 deletions(-) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index b3de3ce..b4edbc3 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -240,21 +240,46 @@ independent reader did. | 33 | The dirty set is exactly the final pass's findings files, read from every porcelain record | **kept** as the close procedure's check | | 34 | The record commit's changed-path set equals those files | **kept** as the close procedure's check | | 35 | The tree is clean after the record commit | **kept** as the close procedure's check | -| 36 | `HEAD` equals the reviewed tip before the reset | **kept** as the close procedure's check, and **corrected**: the reviewed *head* and the restore *point* are two values, not one (pass 21) | +| 36 | `HEAD` equals the reviewed tip before the reset | **kept** as the close procedure's check, and **corrected**: the reviewed *head* and the closing *tip* are two values, not one (pass 21). With row 40 dropped the tip is **only** this precondition — nothing resets to it | | 37 | The closing commit's subject is not a snapshot | **kept** as the close procedure's check | | 38 | The tree is clean after the close | **kept** as the close procedure's check | | 39 | 8a and 8b are separate invocations; no `-m` in the closing one | **kept**, and it is why the close procedure names two invocations | -| 40 | Every rejection restores a tip by **mixed** reset, never `--hard` | **re-expressed** as the failure procedure, and **corrected twice**: restoring `HEAD` is not the whole duty — index and working tree are inspected and captured to patch files first, because **`--mixed` destroys an index-only change** and claiming it preserves the delta was an overclaim; and an 8a rejection restores to the reviewed *head*, the reviewed *tip* not existing yet (pass 22) | +| 40 | Every rejection restores a tip by **mixed** reset, never `--hard` | **DELIBERATELY DROPPED**, by Daniel's decision, and the second of two deletions this table records. **It was never a requirement of the approved spec** — §A says *"re-establish every closure condition against the repository as it now stands"*, which reads the state rather than rewinding it, and a rejected commit at `HEAD` is a legitimate starting point for that reading. It was this plan's own implementation choice, and it had grown restore targets, phase selection, pre- and post-act captures and their own error handling, which four consecutive passes then found defects in. **What replaces it is a bounded handoff**: stop every further mutating action, report the failed step and the observed state, change nothing else. **The obligations §A does impose are unchanged** — surface the failure, re-establish every closure condition against the state as it stands, and take one of its three routes. **Not replaced by a generic backup mechanism**, which would be the same growth under another name | | 41 | The scratch files are removed only after a successful close | **kept** as the close procedure's last step | -**Nothing in that table is dropped except condition 20, and that one is dropped by being replaced -with a stronger obligation.** Six conditions are corrected or extended against pass-21 findings, and +**Two conditions are dropped, 20 and 40, and each is named as a drop rather than lost.** Condition +20 is replaced by a stronger obligation. **Condition 40 is dropped outright** — a guard this plan +introduced itself, which the approved spec never asked for, removed by an explicit decision after +it became the largest single source of findings in the cycle. **A guard the plan invented does not +have to survive on the strength of existing**; what it may not do is take a governing obligation +with it, and §A's three routes, the closure conditions and every evidence duty stand exactly as +before. Six conditions are corrected or extended against pass-21 findings, and **three rows were themselves wrong when first written** — 8, 15 and 40 claimed "kept, same shell" for obligations that were not in fact present everywhere. Pass 22 was aimed at exactly that question and found them; they are repaired rather than re-worded, and the fact that an accounting table can itself be wrong is why the check was worth running. The rest keep their force; what changes is that a person reads the state and picks the operation, instead of one block trying to branch through every state in advance. +### The findings the dropped guard carried, reassessed one by one + +**Removing a mechanism does not settle the findings raised against it**, so each is judged on +whether it described a duty that survives the removal or only a defect in the removed machinery. + +| Finding | Was it about the guard? | Standing | +|---|---|---| +| **23-2** — the restore target was chosen by which tip file existed, rewinding past a second candidate's repair | yes, entirely | **moot.** Nothing restores. The tip survives as 8b's precondition value, and step 7 still clears it between candidates | +| **23-3** — the failure capture covered tracked content only, while the closure inputs are ignored paths | the capture, yes; **the duty, no** | **carried.** The report must name which cycle values still exist rather than assume them, and it says so by listing them | +| **24-2** — "an empty pair means nothing moved", and the checksums had no pre-act baseline | yes, entirely | **moot**, and the underlying error is answered differently: the procedure now refuses to claim anything was preserved at all | +| **24-3** — after a commit lands and then fails, both tracked patches are empty and nothing recorded the rejected `HEAD` | the capture, yes; **the duty, no** | **carried.** The report states whether a commit landed, with `git rev-parse HEAD` and the log, because that is the first thing the resume decision needs | +| **25-1** — `test -e && cksum` in a loop makes the block's status its last iteration | yes, entirely | **moot.** The loop is gone | +| **25-2** — the pre-act capture held no bytes, so a hook rewriting a staged file in place was undetectable | yes as stated; **the honesty duty survives** | **carried, and answered by admission rather than by machinery**: the procedure states plainly that stopping is no guarantee the failed operation destroyed nothing, and forbids reporting that nothing was lost | +| **25-3** — the captures wrote into the worktree they compared | yes, entirely | **moot.** There is nothing to compare | + +**Everything else these passes found stands and is owed** — the close-condition ordering, `c1`–`c3`'s +missing observation, Task 0's `cond` fragments bypassing the table, Task 11's count and its sweep +shape, step 5's stale "appends", the evidence-entry revalidation branch, and the cleanup list. None +of them touched this guard. + ### Preparation — before any task edits a file **What is checked.** The branch; a clean tree; that `ba15e83` is an ancestor; that the three @@ -286,7 +311,7 @@ an unapproved edit to a spec that the gates already closed. 1. **`HEAD` equals the head the candidate pass was issued against** — the value recorded before that call, in `.context/loop-rule-reviewed-head`. **That is the reviewed head.** It is not the same - value as the restore point below, and conflating them is how an unreviewed commit reaches the + value as the **closing tip** below, and conflating them is how an unreviewed commit reaches the close. 2. **The only thing dirty is the candidate pass's own findings files** — read from every porcelain record, not from two status codes. Nothing else may be uncommitted: every record this plan @@ -297,13 +322,15 @@ an unapproved edit to a spec that the gates already closed. message discovered after the record commit leaves `HEAD` moved for a reason no restoration was owed for. 4. **Those files are committed** in their own invocation, and that commit **changes exactly those - paths**. Its parent is the reviewed head. Record the resulting commit as the **restore point**, - in `.context/loop-rule-reviewed-tip`. -5. **The tree is clean**, and `HEAD` is still the restore point, when the closing invocation begins. + paths**. Its parent is the reviewed head. Record the resulting commit as the **closing tip**, in + `.context/loop-rule-reviewed-tip` — **a precondition value, not a restore target**: 8b refuses to + reset unless `HEAD` is still exactly it, which is how a commit landing between the two + invocations is caught. Nothing in this plan resets *to* it. +5. **The tree is clean**, and `HEAD` is still the closing tip, when the closing invocation begins. 6. **After the closing commit**: its subject is not a snapshot, and the tree is clean. **How success is recognised.** A single commit at `HEAD` whose subject is the real message, whose -tree equals the restore point's tree, and whose parent is `$BASE`. Then, and only then, the scratch +tree equals the closing tip's tree, and whose parent is `$BASE`. Then, and only then, the scratch files are removed. **Two invocations, not one.** `codex-gate.sh`'s `is_wip_commit` @@ -322,156 +349,60 @@ before a move and an earlier draft said they did: and then stop without restoring. - **4, 5 and 6 straddle or follow a move** — 4 makes the record commit, 5 reads the state it left, 6 reads the state after the closing commit. **On deviation, run `## The four procedures` · - Failure** against the phase's restore target: `loop-rule-reviewed-head` for a rejection at 4 or 5, - `loop-rule-reviewed-tip` for one at 6. + Failure**, which stops, reports the state and hands over. -**Stopping with a rejected commit at `HEAD`, or with the soft-reset index still live, is not an -outcome this procedure allows** — that is the state a retry cannot start from, and it is exactly -what "stop before moving `HEAD`" produced when applied to a check that runs after one. +**A rejected commit left at `HEAD`, or a live soft-reset index, is a state this plan stops in and +reports** — it is not an error condition the procedure has to undo. §A re-establishes the closure +conditions **against the repository as it now stands**, so that state is a legitimate starting +point for the resume decision, and an earlier draft's claim that "a retry cannot start from it" was +wrong about the approved text. ### Failure — the closing act did not complete -**The approved §A requires the state a failed closing act leaves to be inspected**, and that is more -than `HEAD`. **Look at three things and say what each holds, before moving anything:** - -- **`HEAD`** — did the commit land? A failed `git commit` leaves it where it was; a hook that - rejected after committing does not. -- **The index** — a commit hook can stage content. `git diff --cached` names it. -- **The working tree** — the same hook can modify files without staging them. `git diff` names it. - -**Capture the delta before restoring, because no reset preserves all of it.** `--hard` destroys -both. **`--mixed` destroys an index-only change** — one whose worktree copy still equals the -restore point — because it rewrites the index from the target commit and leaves nothing for -`git status` to report. Saying "`--mixed` preserves the delta" was an overclaim; what it preserves -is the worktree half. - -**There is a *before* half, and it runs as part of the close, not here.** A capture with nothing to -compare against says only what exists, never what changed. **Both halves write into -`.context/loop-rule-capture/`, and both listings exclude that directory** — otherwise the act of -capturing changes the thing being compared, and two listings differ over the procedure's own files -even when the closing act moved nothing. - -```bash -capture() { # $1 = phase name: pre | failed - d=.context/loop-rule-capture; mkdir -p "$d" - git status --porcelain --untracked-files=all --ignored \ - | grep -v ' \.context/loop-rule-capture/' > "$d/$1-status.txt" - git diff --cached > "$d/$1-index.patch" - git diff > "$d/$1-worktree.patch" - # Every dirty path's bytes, so a hook that rewrites a staged findings file - # in place — leaving its porcelain status unchanged — is still detectable. - : > "$d/$1-dirty.txt" - git status --porcelain --untracked-files=all -z | tr '\0' '\n' | sed -n 's/^.\{3\}//p' | - while IFS= read -r f; do - case "$f" in .context/loop-rule-capture/*) continue ;; esac - if [ -f "$f" ]; then printf '%s\t%s\n' "$f" "$(cksum < "$f")" - else printf '%s\tabsent-or-not-a-regular-file\n' "$f"; fi - done >> "$d/$1-dirty.txt" - # Closure inputs, each recorded present-with-checksum or absent. An absent - # file is a RESULT, not a failure: before 8a the reviewed tip does not exist. - : > "$d/$1-inputs.txt" - for f in .context/loop-rule-closing-msg .context/loop-rule-base \ - .context/loop-rule-reviewed-head .context/loop-rule-reviewed-tip; do - if [ -e "$f" ]; then - c=$(cksum < "$f") || { echo "cksum failed on $f"; return 1; } - printf '%s\t%s\n' "$f" "$c" >> "$d/$1-inputs.txt" - else - printf '%s\tabsent\n' "$f" >> "$d/$1-inputs.txt" - fi - done - return 0 -} -``` - -**Immediately before each closing act:** `capture pre || exit 1`. - -**And after a failed act, before restoring anything:** - -```bash -d=.context/loop-rule-capture -git rev-parse HEAD > "$d/failed-head.txt" # the rejected commit, if one landed -TGT=$(cat .context/loop-rule-reviewed-tip 2>/dev/null || cat .context/loop-rule-reviewed-head) -git log --oneline "$TGT"..HEAD > "$d/failed-commits.txt" -git diff "$TGT" HEAD > "$d/failed-landed.patch" -capture failed || exit 1 -``` - -**Three things that loop gets right and an earlier draft did not.** `test -e "$f" && cksum "$f"` as -a loop body makes the block's status that of its **last** iteration — so before 8a, where the -reviewed tip is deliberately absent, the whole before-capture returned 1 and a correct candidate -could not enter the close; and a checksum that actually failed on an earlier input was masked by a -later success. Recording `absent` explicitly makes absence a result rather than an error. And -**checksumming every dirty path, not only the four inputs**, is what catches a hook that rewrites a -staged findings file in place: the porcelain status is unchanged, the after-capture's index patch -holds only the post-hook bytes, and without a before-image of those bytes the original review -artifact is simply gone. - -**The landed patch is the half an earlier draft had no way to see.** After a commit lands and then -fails a postcondition, `git diff --cached` and `git diff` are both **empty** — the content is in the -commit — so the two tracked patches describe nothing while `HEAD` has moved and published different -content or a different message. `restore-target..HEAD` is what records it, and after the mixed -restore that record is the only account of what the attempt actually did. - -**Every one of these files is written even when empty**, so "nothing was left behind" is a recorded -observation rather than an absent file. - -**Two empty patches do not prove nothing moved, and an earlier draft said they did.** They cover -**tracked** content only. A hook can create or rewrite an **untracked or ignored** path — and the -closing message, the recovery base and both tip files are ignored, and the close reads every one of -them — so the status listing with `--untracked-files=all --ignored` and the checksums of the closure -inputs are what make "nothing moved" a statement about the things a condition is actually read -from. Compare them against the same four before the next attempt. - -**The restore target is chosen by phase, not by which file happens to exist.** After 8a's commit it -is `.context/loop-rule-reviewed-tip`; before it — a record commit that failed, or landed and then -failed a postcondition — it is `.context/loop-rule-reviewed-head`, which every Gate-B call wrote. -**An 8a rejection is a rejection like any other and owes the same restoration**, which an earlier -draft's bare exits did not give it. - -**"Whichever tip exists" was wrong, and the second candidate is why.** After an 8b failure whose -repair earns another pass, the *previous* candidate's `loop-rule-reviewed-tip` is still on disk when -the new 8a starts — so an 8a failure would select it and rewind past the newly reviewed repair and -its evidence. **Step 7 deletes `loop-rule-reviewed-tip` when it writes a new -`loop-rule-reviewed-head`**, so before 8a's commit the file is absent by construction rather than by -luck; and whichever target is chosen, **assert it is an ancestor of `HEAD` and that its own -`loop-rule-reviewed-head` matches the head this candidate was issued against** before restoring to -it. - -**How success is recognised.** The repository is back at that tip; the captured set — commit -listing, landed patch, index and worktree patches, status listing and input checksums — describes -what the attempt left, **each compared against the pre-act capture**; and the recovery base and -reviewed-head files still exist. - -**On deviation.** If the delta cannot be explained, stop and hand it to a person. **Do not retry a -close against a state you cannot account for** — the second attempt would carry whatever the first -one left. - -**What happens next is §A's, and it has three routes, not one.** A failed act **returns to the -closure step it failed in**, not to the branches — this cycle's pass already took the -clean-completion branch. So: **re-establish every closure condition against the repository as it now -stands**, then read which route you are on. - -- **Every condition still holds** → **perform the act again.** No further pass is owed. An earlier - draft demanded a fresh clean response here unconditionally, which routes a transient identity or - signing failure — one that moved nothing any condition reads — through a whole extra pass the - approved text does not ask for. -- **The attempt or its repair moved something a condition is read from** → that condition has - changed, **and its own rule decides what it costs**, a further pass included. The cycle is back in - the ordering with that pass owed. The two patch files above are how you tell which case you are - in — and **it is not the two tracked patches that decide it.** Compare the *after* capture against - the *before* one, file by file: the status listings, the index and worktree patches, the - per-dirty-path checksums, and the closure-input records. The commit listing and the landed patch - have no before-image by construction — a non-empty either one means the act moved `HEAD`. **An empty patch pair proves nothing**, since - a landed commit leaves both empty and a hook rewriting an ignored closure input leaves both empty - too. -- **The failure cannot be repaired at all** — a signing key nobody has, a permission nobody can - grant — → **surface it and leave the cycle parked**: open, not running, spending no passes, - restarted by an explicit later continue. **This route was missing entirely**, and without it a - cycle that can neither close nor be parked is exactly the outcome §A names. - -**What this procedure does not promise.** It does not transact any of that. It restores a known -state and records what the attempt left; **the route is a reading**, and trying to automate the -reading is what grew the block this section replaces. +**This procedure stops and hands over. It does not restore, retry or clean up.** That is a decision +about *this* plan, recorded in the accounting table as a deliberate drop, and the reason is that the +approved §A does not ask for automatic restoration: it says **"re-establish every closure condition +against the repository as it now stands"** — read the state, not rewind it. A one-off implementation +plan does not need to grow a general git-recovery mechanism to satisfy that. + +**On a failed closing operation, or a failed postcondition after one: stop every further mutating +action.** No reset. No second `git commit`. No deletion of recovery inputs, scratch values or review +artifacts. **Whatever the repository holds is what the resume decision is made from**, and moving it +first is what destroys the evidence that decision needs. + +**Then report, and report the state rather than a conclusion about it:** + +- **which step failed**, and the concrete command failure — the exit status and the message, quoted; +- **whether a commit landed**: `git rev-parse HEAD`, `git log --oneline -3`, and where a `reset + --soft` had already run, say so, because the index then holds the whole change; +- **what the tree holds**: `git status --porcelain --untracked-files=all`, `git diff --cached + --stat`, `git diff --stat`; +- **which cycle values still exist** — the recovery base, the reviewed head, the reviewed tip, the + closing message — by listing them, not by assuming. + +**What stopping does not do, said plainly.** It prevents *further* change; it is **no guarantee that +the failed operation itself destroyed nothing.** A hook that rejected after rewriting a staged file +has already done that, and this procedure cannot undo it or prove it did not happen. **Do not report +"nothing was lost".** Report what is there. + +**How success is recognised.** The failure is surfaced with the four observations above, nothing +further has been changed, and the cycle is waiting on a person. + +**Resuming is a decision about the actual state**, taken with that report in hand, and then §A's +rules apply unchanged: + +- **Every closure condition is re-established against the repository as it now stands.** Where they + all still hold, **perform the act again**. +- **Where the attempt or its repair moved anything a condition is read from**, that condition has + changed and **its own rule decides what it costs**, a further pass included, and the cycle is back + in the ordering with that pass owed. +- **Where the failure cannot be repaired at all** — a signing key nobody has, a permission nobody + can grant — **surface it and leave the cycle parked**: open, not running, spending no passes, + restarted by an explicit later continue. + +**A person's help replaces neither the review nor the evidence.** Whatever route resume takes, no +closing act happens until every prescribed condition is established again, and the evidence entry +and records the closing commit carries are the ones a pass actually validated. ### Resume — re-entering after an interruption @@ -2575,7 +2506,7 @@ the next pass against exactly that value: ```bash # Before committing a fix, re-run what the fix could have broken. git add -A && git commit -m "WIP: fix " # or: "WIP: pass records" where no repair was owed -rm -f .context/loop-rule-reviewed-tip # the previous candidate's restore point +rm -f .context/loop-rule-reviewed-tip # the previous candidate's closing tip git rev-parse HEAD > .context/loop-rule-reviewed-head # the head the NEXT call is issued against cat .context/loop-rule-reviewed-head ``` @@ -2693,18 +2624,18 @@ test "$expected" = "$actual" || { echo "dirty set is not exactly this pass's fin # shellcheck disable=SC2086 git add $FINAL -# Any rejection below goes through the Failure procedure, restoring to $HEADREV — -# the reviewed-tip file does not exist yet at this point. -git commit -m "WIP: Gate-B findings files" || { echo "record commit FAILED — run Failure against $HEADREV"; exit 1; } +# Any rejection below stops and goes through the Failure procedure, which reports +# the state and hands over. Nothing here resets, retries or deletes. +git commit -m "WIP: Gate-B findings files" || { echo "record commit FAILED — stop here and run Failure"; exit 1; } test "$(git show --name-only --pretty=format: HEAD | sed '/^$/d' | sort)" = "$expected" \ - || { echo "record commit changed paths beyond this pass's findings files — run Failure against $HEADREV"; exit 1; } -test "$(git rev-parse HEAD^)" = "$HEADREV" || { echo "record commit's parent is not the reviewed head — run Failure against $HEADREV"; exit 1; } -test -z "$(git status --porcelain)" || { echo "tree not clean after the record commit — run Failure against $HEADREV"; exit 1; } + || { echo "record commit changed paths beyond this pass's findings files — stop here and run Failure"; exit 1; } +test "$(git rev-parse HEAD^)" = "$HEADREV" || { echo "record commit's parent is not the reviewed head — stop here and run Failure"; exit 1; } +test -z "$(git status --porcelain)" || { echo "tree not clean after the record commit — stop here and run Failure"; exit 1; } git rev-parse HEAD > .context/loop-rule-reviewed-tip # the RESTORE POINT, not the reviewed head ``` -**The reviewed head and the restore point are two values.** The first is what the pass read; the +**The reviewed head and the closing tip are two values.** The first is what the pass read; the second is that plus the findings files. **A single value cannot be both**, and treating it as one is how a commit that landed after the response reaches the close — which is why `HEAD^` is compared above rather than assumed. @@ -2719,9 +2650,10 @@ git reset --soft "$BASE" git commit -F .context/loop-rule-closing-msg ``` -**Then check the result, and on any failure run `## The four procedures` · Failure** — which -inspects `HEAD`, the index **and** the working tree, preserves whatever the attempt left, and -restores by **mixed** reset to `$TIP`: +**Then check the result. On a failed `git commit`, or on either check below failing, run +`## The four procedures` · Failure** — which stops every further mutating action, reports the failed +step and the observed state, and hands over. **It does not reset, re-commit or clean up**, and the +`rm -f` below is therefore unreachable on that path: ```bash git log -1 --pretty=%s # expect the real message, not a snapshot @@ -2737,14 +2669,15 @@ rm -f .context/loop-rule-base .context/loop-rule-reviewed-tip .context/loop-rule .context/loop-rule-sites .context/loop-rule-changed-sites \ .context/loop-rule-c.src .context/loop-rule-w.src \ .context/loop-rule-a.txt .context/loop-rule-b.txt -rm -rf .context/loop-rule-capture ls .context/loop-rule-* 2>/dev/null && { echo "cycle scratch survives the close — list it above"; exit 1; } ``` **Every `loop-rule-*` scratch file goes, and the `ls` is what makes "the scratch files are removed" true rather than asserted.** An earlier draft deleted three of them and claimed the terminal state, -leaving a later run to inherit a closing message, a baseref, an untouched map and any failure -snapshots — each of which some check then has to detect or overwrite piecemeal. **The plan's records +leaving a later run to inherit a closing message, a baseref and an untouched map — each of which +some check then has to detect or overwrite piecemeal. **This runs only on a successful close**; a +failure leaves every one of these files exactly where it is, which is what the resume decision is +read from. **The plan's records are not among these**: they live in the plan and in `.context/codex-reviews/`, both tracked, both already in the closing commit. From f69db7de59c9d8fbb0c7e8317167159ef2961130 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Tue, 15 Sep 2026 14:04:01 +0200 Subject: [PATCH 136/181] docs(plans): apply Gate-A plan pass 26; Resume must recognise the handoff topologies MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit First pass on the revised plan, as the decision required. One Blocker, two Majors, one Minor, one Nit — all consequences of dropping automatic restoration, which is what this pass was aimed at. - Resume audited the implementation by reading WIP: commits between the base and HEAD, and treated any non-WIP commit as a stale base. The handoff deliberately leaves two other shapes: a rejected commit sitting above the cycle's, and HEAD AT the base with the whole implementation in the soft-reset index, where `git log "$BASE"..HEAD` is empty and the task checkboxes are the only record. Either would have had the executor delete and re-record the cycle base, losing the reviewed boundary. Resume recognises all three now, keeps the original base until a person picks a §A route, and accounting row 7 is scoped to the normal pre-close topology. - Step 7b said a failed closing act "preserves" the closing-message file and required a fresh final pass. Neither holds: the handoff states it cannot promise the failed operation left anything as it was, and §A retries the act where every closure condition still holds, owing a pass only where something a condition reads actually moved. The file is inspected and rebuilt on resume. - Step 4b's opening still said "the seven hook strings" after pass 18 corrected the subject set to ten prompt bodies — the stale-referrer class again, in the half of the step a reader reaches first. - The inline comment on loop-rule-reviewed-tip still called it "the RESTORE POINT", the most prominent description of a value that is now only 8b's precondition. - A stray emphasis delimiter on the record-shapes line. --- .../gate-a-plan-om0bdd7udh-pass-26.md | 6 +++ .../2026-09-14-loop-rule-consolidation.md | 48 ++++++++++++++----- 2 files changed, 43 insertions(+), 11 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-26.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-26.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-26.md new file mode 100644 index 0000000..2510222 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-26.md @@ -0,0 +1,6 @@ +BLOCKER | high | Accounting row 7 and Resume procedure | The WIP-only topology is still marked kept and Resume treats any non-WIP commit after the base as making that base stale, even though the new Failure procedure deliberately leaves a rejected closing commit at HEAD and can also leave HEAD at the base with the entire implementation only in the soft-reset index | A supported closing-failure state cannot pass Resume's implementation audit: the executor is told either to delete and re-record the original cycle base or sees no WIP commits at all, so re-entry can lose the reviewed cycle boundary or cannot establish how far implementation got | Make Resume recognize the reported post-8a and post-8b failure topologies explicitly, preserve the original base until the human chooses a section-A route, inspect HEAD plus index plus worktree plus cycle values, and scope accounting row 7 to the normal pre-close topology without reinstating automatic restoration +MAJOR | high | Task 15 step 7b | The step says a failed closing act preserves the closing-message file while requiring a fresh final pass, but the new Failure procedure expressly admits that the failed operation may already have destroyed or changed state, and approved section A1 says to retry the act without another pass when every closure condition still holds | An executor can trust a closing message the failed operation did not preserve and can force an unnecessary new review instead of taking the specified retry route, so the plan no longer implements the three failure routes it says remain unchanged | State that the closing message must be inspected and rebuilt from current records on resume, and require a new pass only when re-establishing the closure conditions shows that a condition's own rule costs one +MAJOR | high | Task 15 step 4b | The opening instruction scopes the twelve-item review to “the seven hook strings,” while the later subject-set instruction correctly requires ten hook prompt bodies because items 12, 15 and 16 also replace their systemMessage bodies | A literal execution can omit the three shipped operator-facing prompts from the only comprehensive prompt-standards review, violating invariant 11 while still recording twelve PASS lines | Use “ten hook prompt bodies” consistently in the opening instruction and identify them there as seven additionalContext bodies plus three systemMessage bodies +MINOR | high | Task 15 step 8a | The inline assignment comment still labels loop-rule-reviewed-tip as “the RESTORE POINT,” directly contradicting the dropped-restoration decision and the surrounding rule that this value is only a closing precondition | The most operationally prominent description of the value can mislead the bounded human handoff into treating it as an authorized reset target | Rename the comment to “the closing tip / 8b precondition, not a restore target” +NIT | high | Fragment evidence preface | The line introducing the six record shapes opens with three asterisks but closes with two | The record-format section renders with a stray or unmatched emphasis marker, making an already dense execution contract harder to read accurately | Change the opening delimiter from three asterisks to two +END OF FINDINGS (5 total) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index b4edbc3..1ec267d 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -211,7 +211,7 @@ independent reader did. | 4 | The three approved inputs' blobs equal their `ba15e83` versions, compared **at the recorded base** | **kept**, same shell — preparation | | 5 | Never overwrite an existing base file | **kept** as an obligation; the shell branch becomes one line of the resume procedure | | 6 | A pre-existing base is an ancestor of `HEAD` (`merge-base --is-ancestor`) | **kept**, same shell — resume | -| 7 | Only this run's `WIP:` commits lie between base and `HEAD` | **kept** as an obligation — resume; the reading is the agent's | +| 7 | Only this run's `WIP:` commits lie between base and `HEAD` | **kept, and scoped**: it is the *normal pre-close* topology. The handoff leaves two more — a rejected commit above the cycle's, and `HEAD` at the base with the implementation in the index — and Resume recognises both rather than reading them as a stale base | | 8 | `$BASE` persisted to a file; every consumer guards it non-empty | **kept**, and **repaired**: three concrete blocks read it without the guard while this row claimed otherwise — the baseline extraction, Task 10's hook diff, and the first Gate-B call. The guard is in each of them now (pass 22) | | 9 | The five regions are located by anchor, one hit per pattern per file | **kept**, same shell — preparation | | 10 | Spans are **derived** from replacement extents, never hand-written | **kept** — the rule, unchanged | @@ -417,8 +417,25 @@ to it and are complete; and how far the implementation got. condition in the five regions appears in exactly one `span` or one `cond`. - `.context/loop-rule-baseline-diff.txt` — every line parses as `base`, a `site` record, or diff output belonging to the site above it, and **every inventoried site has a `site` record**. -- **The implementation**, by reading the `WIP:` commits between the base and `HEAD` and the plan's - own task checkboxes. +- **The implementation**, by reading the commits between the base and `HEAD` and the plan's own task + checkboxes. + +**Three topologies are valid here, not one.** The `WIP:`-only shape is the normal pre-close one. The +handoff deliberately leaves two others, and Resume must recognise them rather than call the base +stale: + +- **After an 8a or 8b commit was rejected**: a non-`WIP` commit, or a `WIP:` findings commit, sits + at `HEAD` above the cycle's `WIP:` commits. **The base is not stale** — it is the same base, with + one more commit on top that a closure condition refused. +- **After `reset --soft` ran and the closing commit failed**: `HEAD` **is** the base and the entire + implementation is in the **index**, so `git log "$BASE"..HEAD` is empty and the task checkboxes + are the only record of how far the work got. **An empty log here is not an empty cycle.** + +**In both, keep the original base and change nothing until a person has chosen a §A route.** Read +`HEAD`, the index, the worktree and the cycle values together — `git status --porcelain +--untracked-files=all`, `git diff --cached --stat`, `git log --oneline -3` — and hand that reading +over. **The stale-base rule scopes to the normal topology**: a non-`WIP` commit means a stale base +only where no closing act was attempted, which the handoff report says. **How success is recognised.** A base that passes preparation, artifacts whose `base` line matches it and whose coverage is complete, and a task list whose ticked entries match the commits present. @@ -2387,7 +2404,8 @@ value by construction. **Nothing mechanical does this and no other task claims it.** The battery's three narrow checks are a floor — one `Target model:` spelling, one prose count claim, one severity vocabulary — and invariant 11 requires all twelve items of every skill, command, hook message and scaffolded template -this change touches. **Read the installed §A–§H text in C, in W, and the seven hook strings, against +this change touches. **Read the installed §A–§H text in C, in W, and the ten hook prompt bodies — +seven `additionalContext` and three `systemMessage` — against each of the twelve items, and record the result per item in this plan.** Items 6 (every constraint carries its reason in the same sentence) and 8 (token-lean) are the ones design §8 names as most at risk. @@ -2563,11 +2581,19 @@ to pass CI, or an invariant-11 violation, published by a cycle that closed clean - [ ] **Step 7b: After the clean pass, complete `.context/loop-rule-closing-msg` — this is the action step 5 defers to, and step 8 has no other source for these records** -**Rebuild the file whole, do not append to it.** Step 5 opened it with a draft entry, and a failed -closing act preserves it while requiring a fresh final pass whose curve and evidence supersede the -previous candidate's. **Appending on re-entry writes a second provenance line and a second curve**, -which breaks the one-of-each grammar Mechanics pins, or leaves a stale curve standing beside the -current one. Write the complete message from the current records at every candidate close, then +**Rebuild the file whole, do not append to it.** Step 5 opened it with a draft entry. **Appending on +re-entry writes a second provenance line and a second curve**, which breaks the one-of-each grammar +Mechanics pins, or leaves a stale curve standing beside the current one. + +**After a failed closing act, inspect this file before trusting it, and rebuild it from the current +records.** The handoff changes nothing, but it **does not claim the failed operation left the file +as it was** — a hook can rewrite anything before failing — so "the failure preserves it" is not a +statement this plan can make. + +**And a failed act does not by itself owe a further pass.** §A retries the act where every closure +condition still holds; a pass is owed only where the attempt or its repair moved something a +condition is read from, **and that condition's own rule is what decides.** An earlier draft required +a fresh final pass unconditionally here, which forces a review §A does not ask for. Write the complete message from the current records at every candidate close, then assert **exactly one** provenance line and **exactly one** curve for this cycle before step 8. The message carries, in this order: @@ -2632,7 +2658,7 @@ test "$(git show --name-only --pretty=format: HEAD | sed '/^$/d' | sort)" = "$ex test "$(git rev-parse HEAD^)" = "$HEADREV" || { echo "record commit's parent is not the reviewed head — stop here and run Failure"; exit 1; } test -z "$(git status --porcelain)" || { echo "tree not clean after the record commit — stop here and run Failure"; exit 1; } -git rev-parse HEAD > .context/loop-rule-reviewed-tip # the RESTORE POINT, not the reviewed head +git rev-parse HEAD > .context/loop-rule-reviewed-tip # the closing tip: 8b's precondition, NOT a reset target ``` **The reviewed head and the closing tip are two values.** The first is what the pass read; the @@ -2751,7 +2777,7 @@ idempotently on re-run, touching no other. **The pre-existing fragments live in never here**; this section records what each observation actually returned, and it is what Task 15 step 5 reads to assemble the closing evidence entry.* -***Six record shapes, because the classes do not return the same number of values.** Every +**Six record shapes, because the classes do not return the same number of values.** Every observation this plan makes is one of them, and a shape that fits only pairs is how a required count gets run and then vanishes from both the plan and the closing evidence:* From f2e5dda2f29cab8534350d2f5ff050a79b8a3356 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Tue, 15 Sep 2026 14:23:20 +0200 Subject: [PATCH 137/181] docs(plans): apply Gate-A plan pass 27; remove a closure condition the spec parks MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Three Blockers, one Major, two Nits. - The close's success criterion required the closing commit's tree to equal the closing tip's tree. Target §I PARKS a Gate-B tree-equality condition on Daniel's decision of 2026-09-13, and the shell performs no such comparison anyway — so the plan both added an out-of-scope closure condition and claimed a predicate it never established. Removed. The prose beside it was also wrong about git: `reset --soft` leaves the index untouched, so STAGED uncommitted content would land in the closing commit; only unstaged work stays out. That is what precondition 5's clean tree is for, and saying otherwise hid the very path §I parks. - Resume inferred a stale base from a non-WIP commit subject. Target §A3 names an accidental non-WIP commit mid-cycle as a reachable state: it resets what the hook counts, the cycle stays open, and the WIP snapshot can survive as an ancestor or be replaced by an amend. Deleting the base there drops earlier implementation out of Gate B's range. Staleness is decided by ancestry and by whether the history is this cycle's, never by a subject line. - 8b ran `git reset --soft` and `git commit` as consecutive bare commands, so a failed reset did not enter the handoff and the commit was attempted anyway — against the wrong HEAD, with the first failure's state already gone. - The closing message was validated before 8a and consumed by 8b without a second look, across an intervening commit whose hooks can rewrite anything, on a path under ignored .context that no clean-tree check can see. It is re-read and revalidated immediately before the commit. - Two counts: "three kinds" over a four-item list, and an unmatched emphasis delimiter pass 26 half-repaired. --- .../gate-a-plan-om0bdd7udh-pass-27.md | 7 +++ .../2026-09-14-loop-rule-consolidation.md | 49 ++++++++++++++----- 2 files changed, 45 insertions(+), 11 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-27.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-27.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-27.md new file mode 100644 index 0000000..3c48c11 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-27.md @@ -0,0 +1,7 @@ +BLOCKER | high | Resume procedure and Task 0 step 1 | Resume still declares a recorded base stale whenever a non-WIP commit follows it outside a reported closing failure, but approved target section A3 explicitly defines an accidental non-WIP commit mid-cycle as a reachable state that resets only hook state, leaves the cycle open, and can leave the WIP snapshot as an ancestor | A successful stray commit or amend can make the executor delete and re-record a base that is still the true boundary, losing earlier implementation from the Gate-B range and bypassing the required removal of surviving WIP history | Add the target-defined accidental non-WIP commit and amend shapes to Resume, preserve the original base while the state is inspected, and reserve stale-base replacement for evidence that the base belongs to another run rather than inferring staleness from the commit subject +BLOCKER | high | Close success criterion and Task 15 step 8 explanation | The plan requires the closing commit's tree to equal the closing tip's tree and says reset --soft stages committed content only, but approved target sections A3 and I explicitly add no Gate-B content condition and park this exact tree-equality decision; the shell also performs no tree comparison, and reset --soft actually leaves the index untouched, including any staged uncommitted content | The executor must either invent an out-of-scope closure check or declare success without establishing the plan's stated predicate, while the false index description hides the parked path by which staged content can enter the closing commit | Remove the tree-equality clause from success recognition, correct the reset prose to say that the index is preserved and only unstaged content stays out of the commit, and state only the target-authorized conditions the procedure really establishes without adding a tree comparison +BLOCKER | high | Task 15 step 8b | The reset and commit are consecutive bare commands, so a failed git reset --soft does not enter Failure and the next git commit is still attempted even though Failure requires every further mutating action to stop immediately after any failed closing operation | A reset error can be obscured by a second failure or followed by hooks and a commit against the wrong HEAD/index state, defeating the bounded handoff and leaving the operator without the state at the first failure | Check the reset exit status and invoke the Failure handoff immediately on failure before any commit attempt; do not restore or retry automatically +MAJOR | high | Close check 3 and Task 15 steps 7b–8b | The closing message is validated only before 8a, then 8a runs a commit and 8b later consumes the ignored .context/loop-rule-closing-msg without revalidation; the plan itself admits that hooks can rewrite anything, and the two invocations also permit an intervening edit that the porcelain clean-tree checks cannot see because .context is ignored | A successful 8a hook or concurrent edit can corrupt or replace the provenance, curve, exception marker, or evidence entry and the final commit can still pass every stated postcondition with an invalid closing record | Re-read and revalidate the current closing-message file after 8a and immediately before 8b uses it, handing off through Failure on any difference or malformed record rather than assuming the pre-8a check still holds +NIT | high | Task 7 step 3 | The sentence says step 1b added three kinds of observation, but the list contains four distinct kinds: pair, moved absence-plus-presence, preservation, and add-only presence | The local count contradicts its own enumeration and makes the observation contract harder to audit mechanically | Change three kinds to four kinds +NIT | high | Fragment evidence preface | Pass 26 removed the opening italic marker from the six-record-shapes paragraph but left the terminal single asterisk after closing evidence, so the paragraph now has an unmatched emphasis delimiter | The record-schema section renders with stray emphasis and fails the requested markup-balance check | Remove the trailing asterisk after closing evidence +END OF FINDINGS (6 total) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index 1ec267d..0530b9b 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -211,7 +211,7 @@ independent reader did. | 4 | The three approved inputs' blobs equal their `ba15e83` versions, compared **at the recorded base** | **kept**, same shell — preparation | | 5 | Never overwrite an existing base file | **kept** as an obligation; the shell branch becomes one line of the resume procedure | | 6 | A pre-existing base is an ancestor of `HEAD` (`merge-base --is-ancestor`) | **kept**, same shell — resume | -| 7 | Only this run's `WIP:` commits lie between base and `HEAD` | **kept, and scoped**: it is the *normal pre-close* topology. The handoff leaves two more — a rejected commit above the cycle's, and `HEAD` at the base with the implementation in the index — and Resume recognises both rather than reading them as a stale base | +| 7 | Only this run's `WIP:` commits lie between base and `HEAD` | **kept as an observation, dropped as a staleness test.** It is the *normal pre-close* topology. Three others are reachable and none makes the base stale: the handoff's rejected commit above the cycle's; `HEAD` at the base with the implementation in the index; and target §A3's accidental non-`WIP` commit or amend mid-cycle. **Staleness is decided by ancestry and by whether the history is this cycle's**, never by a commit subject | | 8 | `$BASE` persisted to a file; every consumer guards it non-empty | **kept**, and **repaired**: three concrete blocks read it without the guard while this row claimed otherwise — the baseline extraction, Task 10's hook diff, and the first Gate-B call. The guard is in each of them now (pass 22) | | 9 | The five regions are located by anchor, one hit per pattern per file | **kept**, same shell — preparation | | 10 | Spans are **derived** from replacement extents, never hand-written | **kept** — the rule, unchanged | @@ -320,7 +320,10 @@ an unapproved edit to a spec that the gates already closed. curve, every owed evidence entry, and either the applicable human-exception records or `Human exceptions: none`. **This is checked here, before anything moves**, because an incomplete message discovered after the record commit leaves `HEAD` moved for a reason no restoration was - owed for. + owed for. **And it is re-read and revalidated at 8b, immediately before the commit consumes it** + — 8a runs a commit in between, hooks can rewrite anything, and the file is under `.context/`, + which is ignored, so no clean-tree check between the two invocations can see it change. A + difference or a malformed record there is a Failure handoff, not a repair. 4. **Those files are committed** in their own invocation, and that commit **changes exactly those paths**. Its parent is the reviewed head. Record the resulting commit as the **closing tip**, in `.context/loop-rule-reviewed-tip` — **a precondition value, not a restore target**: 8b refuses to @@ -329,8 +332,11 @@ an unapproved edit to a spec that the gates already closed. 5. **The tree is clean**, and `HEAD` is still the closing tip, when the closing invocation begins. 6. **After the closing commit**: its subject is not a snapshot, and the tree is clean. -**How success is recognised.** A single commit at `HEAD` whose subject is the real message, whose -tree equals the closing tip's tree, and whose parent is `$BASE`. Then, and only then, the scratch +**How success is recognised.** A single commit at `HEAD` whose subject is the real message and whose +parent is `$BASE`. **No tree comparison** — target §I parks a Gate-B tree-equality condition on +Daniel's decision of 2026-09-13, and an earlier draft of this line added one anyway, which would +have made the executor either invent an out-of-scope closure check or declare success without +establishing its own stated predicate. Then, and only then, the scratch files are removed. **Two invocations, not one.** `codex-gate.sh`'s `is_wip_commit` @@ -409,9 +415,18 @@ and records the closing commit carries are the ones a pass actually validated. **What is checked.** Whether the recorded base is this run's; whether the scratch artifacts belong to it and are complete; and how far the implementation got. -- **The base**, by the preparation procedure's checks. A base that is not an ancestor, or has - non-`WIP:` commits after it, is **stale** — delete it deliberately, record why, and re-record from - the true starting commit. **Never overwrite one you did not just write.** +- **The base**, by the preparation procedure's checks. **A base that is not an ancestor of `HEAD` is + stale** — delete it deliberately, record why, and re-record from the true starting commit. + **Never overwrite one you did not just write.** + + **A non-`WIP` commit after the base does not make it stale, and inferring that from the subject + is wrong.** Target §A3 names an accidental non-`WIP` commit mid-cycle as a **reachable state**: it + resets what the hook counts, the cycle stays open until the conditions hold, and it can leave the + `WIP:` snapshot as an ancestor — or, as an amend, replace the tip. In both the base is still the + true boundary, and deleting it would drop earlier implementation out of Gate B's range and out of + the final reset. **Keep the base, inspect the state, and let §A3 decide what the stray commit + costs.** Replace a base only on evidence it belongs to **another run** — it is not an ancestor, or + its history has no part of this cycle in it. - **Each scratch artifact**, by its own `base` line **and** its own completeness predicate: - `.context/loop-rule-untouched` — every line parses as `base`, `span` or `cond`, and every kept condition in the five regions appears in exactly one `span` or one `cond`. @@ -1613,7 +1628,7 @@ holds OLD fragments, which exist before the edit and can be checked in advance. Expected for all twenty: `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`. -**Then run everything step 1b added, and that is three kinds, not one:** +**Then run everything step 1b added, and that is four kinds, not one:** - **a pair** for each of the `c18`-and-surfacing block's four further conditions — `c16`, `c17`, `c19`, `c20`; @@ -2672,7 +2687,13 @@ above rather than assumed. BASE=$(cat .context/loop-rule-base); TIP=$(cat .context/loop-rule-reviewed-tip) test -n "$BASE" && test -n "$TIP" || { echo "BASE or TIP missing — 8a did not complete"; exit 1; } test "$(git rev-parse HEAD)" = "$TIP" || { echo "HEAD has moved since 8a — NOT resetting"; exit 1; } -git reset --soft "$BASE" +git reset --soft "$BASE" || { echo "reset --soft FAILED — stop here and run Failure; do NOT commit"; exit 1; } +# Re-read the closing message here: it was validated before 8a, 8a ran a commit +# (hooks can rewrite anything), and .context is ignored, so no porcelain check +# between the two invocations can see it change. +test -s .context/loop-rule-closing-msg || { echo "closing message missing or empty — stop here and run Failure"; exit 1; } +# Revalidate its records — exactly one provenance line, exactly one curve for this +# cycle, every owed evidence entry, and the exception records or the plural marker. git commit -F .context/loop-rule-closing-msg ``` @@ -2707,7 +2728,13 @@ read from. **The plan's records are not among these**: they live in the plan and in `.context/codex-reviews/`, both tracked, both already in the closing commit. -**`reset --soft` stages committed content only.** The prompt-standards result, the completeness +**`reset --soft` moves `HEAD` and leaves the index exactly as it was** — it stages nothing and +unstages nothing, so the index still holds every `WIP:` commit's content, which is what makes the +single closing commit carry the whole change. **Anything *staged* and uncommitted at that moment is +in the index too and would land in the closing commit**; only *unstaged* work stays out. That is why +the close's precondition 5 requires a clean tree before 8b runs — **an earlier draft said the reset +"stages committed content only" and that unstaged-or-not, uncommitted content stays out, which is +wrong about the index and hides the path §I parks.** The prompt-standards result, the completeness sweep, the next-state table, the divergence list, the equivalence result and the fragment evidence all land in this plan, and `.context/codex-reviews/` is tracked; **anything uncommitted when the reset runs is left in the worktree and is not in the closing commit** — and, for the plan records, @@ -2775,7 +2802,7 @@ line per fragment-table row with its three results, and the reading result for ` *Empty until the tasks run. One subsection per task — `### Task 1`, `### Task 3`, … — replaced idempotently on re-run, touching no other. **The pre-existing fragments live in the fragment table, never here**; this section records what each observation actually returned, and it is what Task 15 -step 5 reads to assemble the closing evidence entry.* +step 5 reads to assemble the closing evidence entry. **Six record shapes, because the classes do not return the same number of values.** Every observation this plan makes is one of them, and a shape that fits only pairs is how a required From 58c49db1fd032c98a09c4df47ceb274856293359 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Tue, 15 Sep 2026 14:37:22 +0200 Subject: [PATCH 138/181] docs(plans): apply Gate-A plan pass 28; validate the commit that landed, not the file MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Two Blockers, two Majors, one Minor, one Nit. Both Blockers are pass 27's own fixes reaching one site out of several. - Pass 27 removed the subject-only stale-base rule from Resume's body and left it standing in Preparation's checks and Resume's closing sentence, against accounting row 7 and §A3. Staleness is decided by ancestry and by whether the history is this cycle's, in all three places now. - Pass 27's reset paragraph stated the index behaviour correctly and then, four lines later, repeated the old false claim that "anything uncommitted" stays out. An executor reading the second sentence can publish staged unreviewed content. The asymmetry is stated once: unstaged stays out, staged lands in. - Close defines success as a commit whose parent is $BASE, and nothing checked it. A HEAD movement between the soft reset and the commit yields a wrong-parent, unsquashed close that passed every stated postcondition. - The closing message was validated in the file; prepare-commit-msg and commit-msg hooks rewrite git's copy AFTER `-F` has read it, so the validated file proves nothing about what landed. The committed body is now diffed against it before cleanup. - Self-Review §3 gave the pair order as new-then-old, against the verification procedure, every task's expected result and the evidence schema. - The untouched map was described as two record shapes while enumerating three line forms, leaving a literal consumer unsure whether `base` is grammar. --- .../gate-a-plan-om0bdd7udh-pass-28.md | 7 ++++ .../2026-09-14-loop-rule-consolidation.md | 34 +++++++++++++------ 2 files changed, 30 insertions(+), 11 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-28.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-28.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-28.md new file mode 100644 index 0000000..8a26185 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-28.md @@ -0,0 +1,7 @@ +BLOCKER | high | Task 0 step 1 / Preparation / Resume | The plan still declares any non-`WIP:` commit after the recorded base stale in Preparation, Task 0 step 1, and Resume's final stale-base sentence, contradicting accounting row 7 and Resume's own rule that an accidental non-`WIP` commit or rejected close is reachable and does not stale the base. | A valid resumed cycle can delete and re-record its true base after implementation has begun, excluding earlier edits from Gate B and the final soft reset. | Remove every subject-only staleness rule; keep the original base when it is ancestral and belongs to this cycle, treat a non-`WIP` subject as a state to inspect under §A3 or the close handoff, and replace the base only on evidence it belongs to another run. +BLOCKER | high | Task 15 step 8 | The reset explanation first correctly says staged uncommitted content remains in the index and would land in the closing commit, then contradicts itself by saying "anything uncommitted" is left in the worktree and excluded. | An executor can rely on the false latter claim and publish staged, unreviewed content if closure condition 5 was not re-established or the index changes before the commit. | Replace the latter claim with "anything unstaged and uncommitted" and retain the explicit warning that every staged path lands in the closing commit. +MAJOR | high | Task 15 step 8b | Close defines success as a closing commit whose parent is `$BASE`, but the post-close commands check only its subject and a clean tree; nothing verifies `HEAD^` equals `$BASE` after `git commit`. | A concurrent or hook-induced `HEAD` movement between the soft reset and commit can produce a wrong-parent, unsquashed close that is accepted before the recovery inputs are deleted. | Before cleanup, assert that the resulting commit's parent is exactly `$BASE` and route any mismatch through Failure. +MAJOR | high | Task 15 step 8b | The mandatory closing message is validated only in `.context/loop-rule-closing-msg` before `git commit -F`; no postcondition validates the actual commit body, even though `prepare-commit-msg` or `commit-msg` hooks can rewrite Git's copied message after that check. | The cycle can close and delete its scratch state with provenance, curve, exception, or evidence records missing or altered in the commit while the existing subject and clean-tree checks still pass. | Before cleanup, read the resulting commit body and re-run the mandatory-record validation or compare it exactly with the validated source file; route any mismatch through Failure. +MINOR | high | Self-Review §3 | The claimed four-value pair shape is ordered `new/worktree`, `new/parent`, `old/worktree`, `old/parent`, while the verification procedure, Task 3 expectations, tables, and final evidence schema all define the order as OLD first and NEW second. | An executor following the self-review can reverse the durable evidence fields, making Task 15's assembled evidence ambiguous or mislabeled. | State the canonical order as `old/worktree`, `old/parent`, `new/worktree`, `new/parent` everywhere. +NIT | high | Task 0 step 2 | The plan says `.context/loop-rule-untouched` holds "two record shapes" and that both parse, but immediately enumerates three line shapes—`base`, `span`, and `cond`—and Resume validates all three. | The count is internally false and leaves a literal consumer unsure whether `base` is part of the record grammar or merely an out-of-band header. | Say "three record shapes", or explicitly distinguish two payload shapes plus one mandatory `base` header line. +END OF FINDINGS (6 total) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index 0530b9b..e7f7c96 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -284,7 +284,9 @@ of them touched this guard. **What is checked.** The branch; a clean tree; that `ba15e83` is an ancestor; that the three approved inputs still hold their approved blobs **at the recorded base**; that the recorded base, if -one exists, is an ancestor of `HEAD` with only this run's `WIP:` commits after it. +one exists, is an **ancestor** of `HEAD` and belongs to this cycle. **Ancestry and provenance decide +that, never a commit subject** — §A3's accidental non-`WIP` commit and the handoff's rejected close +both leave the base valid, and the commits after it are a state to inspect rather than a verdict. ```bash test "$(git rev-parse --abbrev-ref HEAD)" = loop-rule-consolidation || { echo "wrong branch"; exit 1; } @@ -449,8 +451,8 @@ stale: **In both, keep the original base and change nothing until a person has chosen a §A route.** Read `HEAD`, the index, the worktree and the cycle values together — `git status --porcelain --untracked-files=all`, `git diff --cached --stat`, `git log --oneline -3` — and hand that reading -over. **The stale-base rule scopes to the normal topology**: a non-`WIP` commit means a stale base -only where no closing act was attempted, which the handoff report says. +over. **There is no subject-based stale-base rule left**: replace the base only on evidence it +belongs to another run — it is not an ancestor, or its history contains no part of this cycle. **How success is recognised.** A base that passes preparation, artifacts whose `base` line matches it and whose coverage is complete, and a task list whose ticked entries match the commits present. @@ -773,7 +775,8 @@ memorise.** `a1` and `a2` sit inside item 8a's block; `a13` ends on the line `a1 shares its line with `h6`; `h4` and `h5` are adjacent. Each is a reason a whole-line span cannot do this alone, which is why the derivation replaces the enumeration rather than correcting it again. -**`.context/loop-rule-untouched` holds two record shapes, and both are parseable**, because Tasks 2 +**`.context/loop-rule-untouched` holds a mandatory `base` header line plus two payload shapes — three +line forms in all, and every one parseable**, because Tasks 2 and 14 consume them without a human in between. A file whose second half has no schema is a file its consumers skip, which is what the per-condition list exists to prevent: @@ -2703,8 +2706,15 @@ step and the observed state, and hands over. **It does not reset, re-commit or c `rm -f` below is therefore unreachable on that path: ```bash -git log -1 --pretty=%s # expect the real message, not a snapshot -git status --porcelain # expect empty +git log -1 --pretty=%s # expect the real message, not a snapshot +test "$(git rev-parse HEAD^)" = "$BASE" || { echo "closing commit's parent is not \$BASE — stop here and run Failure"; exit 1; } +git status --porcelain # expect empty +# The COMMIT BODY, not the source file: prepare-commit-msg and commit-msg hooks +# rewrite git's copy after -F has read it, so the validated file proves nothing +# about what landed. +git log -1 --pretty=%B > .context/loop-rule-landed-msg +diff .context/loop-rule-closing-msg .context/loop-rule-landed-msg \ + || { echo "the committed body differs from the validated message — stop here and run Failure"; exit 1; } ``` **Both clean, and only then:** @@ -2715,7 +2725,7 @@ rm -f .context/loop-rule-base .context/loop-rule-reviewed-tip .context/loop-rule .context/loop-rule-untouched .context/loop-rule-baseline-diff.txt \ .context/loop-rule-sites .context/loop-rule-changed-sites \ .context/loop-rule-c.src .context/loop-rule-w.src \ - .context/loop-rule-a.txt .context/loop-rule-b.txt + .context/loop-rule-a.txt .context/loop-rule-b.txt .context/loop-rule-landed-msg ls .context/loop-rule-* 2>/dev/null && { echo "cycle scratch survives the close — list it above"; exit 1; } ``` @@ -2736,9 +2746,11 @@ the close's precondition 5 requires a clean tree before 8b runs — **an earlier "stages committed content only" and that unstaged-or-not, uncommitted content stays out, which is wrong about the index and hides the path §I parks.** The prompt-standards result, the completeness sweep, the next-state table, the divergence list, the equivalence result and the fragment evidence -all land in this plan, and `.context/codex-reviews/` is tracked; **anything uncommitted when the -reset runs is left in the worktree and is not in the closing commit** — and, for the plan records, -was never in a Gate-B range either. Step 7 commits them before the candidate pass is issued, which +all land in this plan, and `.context/codex-reviews/` is tracked. **Anything uncommitted and +*unstaged* when the reset runs stays in the worktree and out of the closing commit; anything +uncommitted and *staged* lands in it.** That asymmetry is what precondition 5's clean tree +prevents; and the plan records, had they been left uncommitted, +would never have been in a Gate-B range either. Step 7 commits them before the candidate pass is issued, which is what lets 8a's dirty-set check be exact. --- @@ -2751,7 +2763,7 @@ is what lets 8a's dirty-set check be exact. **Which steps are mechanical and which are reader checks, stated rather than claimed uniformly.** Every OLD half has an exact expected result and a procedure that produces it, but **not every one is a pre-verified table row**: the rows the tables carry were checked against the real files in advance, while Tasks 3, 4, 6, 7 and 10 **derive their remaining pre-existing fragments at execution, before their install step**, against text this plan cannot quote without becoming a second copy of it. **Task 1 is not among them:** §A is add-only, row P1 records that it has no OLD half at all, and Task 1 derives presence fragments only — listing it would send an executor looking for a counterfactual that cannot exist. **The NEW halves are `<...>` until their task installs the text**, which the fragment table discloses and each step requires to be verified before counting. **And no per-task shell is pre-written at all** — the procedure is stated once and the executor writes the command in front of the files, so "a runnable command per step" is not what this plan claims. **Tasks 12, 13 and 14 step 2 are reader checks by design** — a predicate comparison, a next-state walk and a divergence classification are judgements, and giving them commands would be the false-precision this repo's invariants warn about. An earlier revision of this section claimed every verification step had a runnable command, which was not true of them. -**3. Type consistency.** `$BASE` is set in Task 0 and used in Tasks 1–10. The four-value pair shape (`new/worktree`, `new/parent`, `old/worktree`, `old/parent`) is defined in Task 3 and referred to by name afterwards. Condition ids match the inventory throughout: a1–a22, b1–b18, c1–c20, d1–d7, e1–e11, f1–f7, g1–g4, h1–h26, i1–i16, j1–j4 — 135 total, every one dispositioned above. +**3. Type consistency.** `$BASE` is set in Task 0 and used in Tasks 1–10. The four-value pair shape is **`old/worktree`, `old/parent`, `new/worktree`, `new/parent`** — the order the verification procedure, every task's expected result and the evidence schema use, and the one this section had reversed. Condition ids match the inventory throughout: a1–a22, b1–b18, c1–c20, d1–d7, e1–e11, f1–f7, g1–g4, h1–h26, i1–i16, j1–j4 — 135 total, every one dispositioned above. **One correction applied from this review:** §G was missing a task; it is now installed by Task 7, which names **ten** sites — one §G block and nine §H blocks. From f99607ba999f1f21bb87722bf2bf948609174125 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Tue, 15 Sep 2026 14:53:31 +0200 Subject: [PATCH 139/181] docs(plans): apply Gate-A plan pass 29; Resume gets its own validation MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit MANDATORY STOP — findings rose 6 -> 9 and Blockers rose 2 -> 5, with the findings clustered on the procedures. Surfaced; standing answer applied. Five Blockers, three Majors, one Minor. Three of the Blockers are one self-contradiction the handoff decision created and my passes 26-28 did not finish carrying through. - Resume declared the post-reset topology valid — HEAD at the base, the whole implementation in the index — and then validated the base through Preparation, which unconditionally requires a clean tree. That state could never pass, so the plan's own success path made a failed close unresumable unless the operator first mutated the evidence-bearing index. Preparation is now first-entry-only and refuses when a base already exists; Resume has its own checks and needs no clean tree. - Resume's stale-base predicate included "its history contains no part of this cycle". In the post-reset topology HEAD IS the base and every cycle commit is squashed into the index, so that predicate is necessarily true of a valid state. Replacement now needs affirmative evidence of another run. - The subject-based staleness rule was still standing in Task 0 step 1 — the third site, after passes 27 and 28 each found one. - Preparation's prose claimed a base predicate its shell never established: it assigned REF and compared three blobs, testing neither the file's existence, its shape, its ancestry nor its provenance. The claim is narrowed to what the shell does. - The post-close subject and clean-tree checks were `git log` and `git status` under expect comments, which return zero on a WIP subject and a dirty tree, and cleanup ran anyway. Both are predicates now. - The cond schema stored the fragment text while the rule beside it required the row to cite a table P id. Satisfying one broke the other. The map carries the id; the table keeps the only copy of the text. - Resume's delete-and-rebuild rule fired unconditionally, destroying the scratch state the handoff preserves, before anyone had chosen a §A route. - The cleanup's `ls` assertion returned nonzero exactly when cleanup succeeded. - Step 7 called the post-review addition a findings file, singular, where a full pass writes the spec and quality pair and step 8 requires both. --- .../gate-a-plan-om0bdd7udh-pass-29.md | 10 ++ .../2026-09-14-loop-rule-consolidation.md | 91 +++++++++++++------ 2 files changed, 71 insertions(+), 30 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-29.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-29.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-29.md new file mode 100644 index 0000000..0e4929c --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-29.md @@ -0,0 +1,10 @@ +BLOCKER | high | The four procedures · Preparation, lines 283–304 | The prose and success criterion say a pre-existing base is a 40-character ancestor of HEAD that belongs to this cycle, but the shell only assigns REF and compares three blobs; it neither requires the base file, validates its shape, tests its ancestry, nor establishes cycle provenance. A valid commit from another history carrying the same three blobs passes this procedure, and on a first run the procedure claims success while the base file still does not exist. | Resume and later counterfactuals can accept or claim a stale base, causing Gate B and the final soft reset to use the wrong boundary. | Make Preparation's executable checks establish the exact stated base predicate without using commit subjects, including the first-run versus existing-base case, or narrow its success contract and require the complete Task 0 validation everywhere Preparation is consumed. +BLOCKER | high | Preparation lines 291–304 and Resume lines 440–458 | Resume declares the post-reset state with HEAD at BASE and the whole implementation staged as valid, then recognizes success only when the base passes Preparation; Preparation unconditionally requires an empty porcelain status, so that valid staged-index topology can never pass. | A failed closing commit cannot be resumed through the plan's own success path, pressuring the operator either to mutate the evidence-bearing index before choosing a §A route or to remain permanently blocked. | Split initial clean-tree preparation from resume validation, and give the staged-index handoff topology a non-mutating validation path that preserves the original base and waits for the person's §A decision. +BLOCKER | high | Task 0 step 1, lines 683–693 | The step says anything other than this run's WIP commits after BASE makes the base stale and directs deletion, directly contradicting Resume and target §A3, which treat an accidental non-WIP commit and a rejected closing commit as reachable states that retain the same base. | Re-entry can delete the valid cycle boundary, excluding earlier implementation from the Gate-B range and from the final squash. | Remove the subject-based staleness rule here and route all non-normal histories to the state-reading Resume procedure; replace the base only on positive evidence that it belongs to another run. +BLOCKER | high | Resume, lines 420–455 | Resume twice permits replacing the base when its history contains no part of this cycle, but its own valid post-reset topology has HEAD equal to BASE and all cycle work only in the index, so the history necessarily contains no cycle commits. | The procedure can classify its expressly valid index-only handoff as another run and discard the true base, corrupting review-range and closing-boundary selection. | Remove absence of cycle commits from the stale-base predicate, or explicitly exempt and positively recognize the staged-index topology before any provenance decision; require affirmative evidence of another run before replacement. +BLOCKER | high | Close success criteria and Task 15 step 8 post-close checks, lines 335–342 and 2703–2729 | The closing subject and clean-tree conditions are only printed with expect comments: git log and git status return zero even when the subject is a WIP snapshot or the tree is dirty. The later parent and body comparisons do not establish either predicate, yet cleanup still runs. | The plan can declare an invalid close successful and delete every recovery input while uncommitted changes remain or the closing commit still has snapshot semantics. | Turn both observations into explicit predicates with the Failure handoff on mismatch, and permit cleanup only after the non-snapshot subject, BASE parent, exact landed body, and empty status have all been asserted. +MAJOR | high | Task 0 step 2 map schema, lines 783–799 | The cond schema stores the fragment text but the following rule requires the map row to cite the fragment table's P id; there is no field for that id and no lookup rule. Satisfying the schema duplicates the authored fragment, while satisfying the citation rule produces a row the stated consumers cannot count. | The committed sweep cannot be reliably tied to the ignored map rows, so an interrupted run can reuse an unverified or drifted condition fragment despite the pass-24 repair this structure is meant to provide. | Define one unambiguous record shape that includes a fragment-row identifier and specify how both consumers resolve it from the single authoritative table, then update parsing and completeness requirements to that shape. +MAJOR | high | Resume, lines 451–462 | The procedure says to change nothing in either failed-close topology until a person selects a §A route, but its unconditional deviation rule immediately deletes and rebuilds any incomplete scratch artifact. | Re-entry can destroy the very scratch-state evidence the Failure handoff deliberately preserves before the operator has decided whether to retry, owe another pass, or park. | While awaiting the §A decision, report an invalid artifact without mutating it; allow deletion and rebuild only after the selected route authorizes continuation and records why the old artifact is no longer evidence. +MAJOR | high | Task 15 step 8 cleanup, lines 2720–2730 | The final ls command returns nonzero when cleanup succeeds because no loop-rule path matches, making the successful branch's shell invocation itself fail even though the prose calls that absence success. | A valid close is surfaced as a command failure after its recovery files have already been deleted, creating a false Failure handoff with less state available to diagnose it. | Express the absence assertion with an if-style predicate whose no-match branch returns zero and whose match branch lists survivors and exits nonzero. +MINOR | high | Task 15 step 7, lines 2593–2597 | The final-pass rule calls the pass's own findings file the sole permitted addition, singular, while a full Gate-B pass requires two branch files and step 8's exact dirty set requires both. | A literal executor receives conflicting closure cardinality immediately before the exact-set check and may treat one missing branch artifact as eligible until step 8 rejects it. | Say findings files and state explicitly that the required spec and quality branch pair is the sole permitted post-review addition. +END OF FINDINGS (9 total) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index e7f7c96..9f63dfc 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -282,26 +282,30 @@ of them touched this guard. ### Preparation — before any task edits a file +**This procedure is for a FIRST entry, on a clean tree.** A re-entry runs **Resume** instead, which +requires no clean tree and mutates nothing — the handoff's valid topologies include `HEAD` at the +base with the whole implementation **staged**, and an unconditional clean-tree test would make that +state permanently unresumable through this plan's own success path. + **What is checked.** The branch; a clean tree; that `ba15e83` is an ancestor; that the three -approved inputs still hold their approved blobs **at the recorded base**; that the recorded base, if -one exists, is an **ancestor** of `HEAD` and belongs to this cycle. **Ancestry and provenance decide -that, never a commit subject** — §A3's accidental non-`WIP` commit and the handoff's rejected close -both leave the base valid, and the commits after it are a state to inspect rather than a verdict. +approved inputs still hold their approved blobs at the revision the tasks are derived against; and +that **no base file exists yet** — if one does, this is a re-entry and Resume owns it. ```bash test "$(git rev-parse --abbrev-ref HEAD)" = loop-rule-consolidation || { echo "wrong branch"; exit 1; } -test -z "$(git status --porcelain)" || { echo "tree not clean"; exit 1; } +test -z "$(git status --porcelain)" || { echo "tree not clean — if this is a re-entry, run Resume"; exit 1; } +test ! -e .context/loop-rule-base || { echo "a base is already recorded — this is a re-entry, run Resume"; exit 1; } git merge-base --is-ancestor ba15e83 HEAD || { echo "ba15e83 not in this history"; exit 1; } -REF=$( [ -s .context/loop-rule-base ] && cat .context/loop-rule-base || git rev-parse HEAD ) for f in target-text design condition-inventory; do p="docs/superpowers/specs/2026-09-10-loop-rule-consolidation-$f.md" - test "$(git rev-parse "$REF:$p")" = "$(git rev-parse "ba15e83:$p")" \ - || { echo "$p differs at $REF from its approved version"; exit 1; } + test "$(git rev-parse "HEAD:$p")" = "$(git rev-parse "ba15e83:$p")" \ + || { echo "$p differs at HEAD from its approved version"; exit 1; } done ``` -**How success is recognised.** Every line above exits 0, and `.context/loop-rule-base` holds a -40-character object name that is an ancestor of `HEAD`. +**How success is recognised.** Every line above exits 0. **No claim is made here about a recorded +base**, because on a first entry there is none — Task 0 step 1 records it immediately afterwards, +and validating it is Resume's job on every later entry. **On deviation.** Stop. Each of these has a different fix and none of them is "retry": a wrong branch is a checkout, a dirty tree is a decision about uncommitted work, a differing input blob is @@ -417,9 +421,17 @@ and records the closing commit carries are the ones a pass actually validated. **What is checked.** Whether the recorded base is this run's; whether the scratch artifacts belong to it and are complete; and how far the implementation got. -- **The base**, by the preparation procedure's checks. **A base that is not an ancestor of `HEAD` is - stale** — delete it deliberately, record why, and re-record from the true starting commit. - **Never overwrite one you did not just write.** +- **The base**, by its **own** checks, not Preparation's — Preparation is first-entry-only and + requires a clean tree, which the handoff's staged-index topology can never have. Read the file, + require a full 40-hex object name that resolves to itself as a commit, and require it to be an + **ancestor of `HEAD`**. **A base that is not an ancestor is stale.** **Never overwrite one you did + not just write.** + + **"Its history contains no part of this cycle" is NOT a staleness test** — in the post-reset + topology `HEAD` *is* the base and every cycle commit has been squashed away into the index, so + that predicate is necessarily true of a perfectly valid state. **Replace a base only on + affirmative evidence of another run**: it is not an ancestor of `HEAD`, or the recorded value + names a commit this branch never contained. **A non-`WIP` commit after the base does not make it stale, and inferring that from the subject is wrong.** Target §A3 names an accidental non-`WIP` commit mid-cycle as a **reachable state**: it @@ -451,15 +463,23 @@ stale: **In both, keep the original base and change nothing until a person has chosen a §A route.** Read `HEAD`, the index, the worktree and the cycle values together — `git status --porcelain --untracked-files=all`, `git diff --cached --stat`, `git log --oneline -3` — and hand that reading -over. **There is no subject-based stale-base rule left**: replace the base only on evidence it -belongs to another run — it is not an ancestor, or its history contains no part of this cycle. +over. **There is no subject-based stale-base rule left**, and no absence-of-cycle-commits rule +either: the index-only topology has neither. Replace the base only on the affirmative evidence +named above. **How success is recognised.** A base that passes preparation, artifacts whose `base` line matches it and whose coverage is complete, and a task list whose ticked entries match the commits present. -**On deviation.** An artifact that fails either test is **deleted and rebuilt from the `$BASE` -blobs** — never reused, and never repaired in place. A same-base partial file is the one shape a -`base` line alone cannot catch, which is why the completeness predicate exists. +**On deviation, and the answer differs by why you are here.** An artifact that fails either test is +**deleted and rebuilt from the `$BASE` blobs** — never reused, never repaired in place. A same-base +partial file is the one shape a `base` line alone cannot catch, which is why the completeness +predicate exists. + +**But not while a failed close is waiting on a person.** In either handoff topology the rule is +change nothing until a §A route is chosen, and an unconditional delete-and-rebuild here would +destroy the scratch state the handoff deliberately preserved — before anyone decided whether to +retry, owe a pass, or park. **Report the invalid artifact and leave it.** Rebuild it once the chosen +route authorizes continuing, and record why the old one is no longer evidence. **Steps that describe the tree at `$BASE` are validated on re-entry, not re-run against the worktree.** After a text task the worktree carries this plan's own edits, and rebuilding the @@ -680,7 +700,7 @@ case "$BASE" in [0-9a-f][0-9a-f][0-9a-f][0-9a-f][0-9a-f][0-9a-f][0-9a-f][0-9a-f] test "${#BASE}" -eq 40 || { echo "recorded base is not a full 40-character object name"; exit 1; } test "$(git rev-parse --verify "$BASE^{commit}")" = "$BASE" || { echo "recorded base does not resolve to itself as a commit"; exit 1; } git merge-base --is-ancestor "$BASE" HEAD || { echo "recorded base is NOT an ancestor of HEAD — stale"; exit 1; } -git log --oneline "$BASE"..HEAD # expect nothing, or only this run's WIP: commits +git log --oneline "$BASE"..HEAD # READ this; it does not decide staleness — see below ``` **Re-running Task 0 after a partial implementation must not re-record the base.** It would capture @@ -688,9 +708,11 @@ the current WIP tip, and both Gate B's range and the final reset would then star made so far — prompt and hook changes squashed into the closing commit without entering a review range. **A pre-existing value is validated, never trusted**: `git log "$BASE"..HEAD` does not test ancestry, which is how a base from an abandoned branch passes; `merge-base --is-ancestor` is the -test the sentence names. Anything other than this run's `WIP:` commits in that log means the file is -stale — delete it deliberately, record why, and re-record from the true starting commit. **`## The -four procedures` · Resume** is what to do next in that case. +test the sentence names. **What the log shows does not decide staleness** — §A3's accidental +non-`WIP` commit and the handoff's rejected close both leave this base valid, and the post-reset +topology shows an empty log with the whole cycle in the index. **Anything other than this run's +`WIP:` commits is a state to read, not a verdict: hand it to `## The four procedures` · Resume**, +which replaces a base only on affirmative evidence it belongs to another run. **Persist it to a file, not to a shell variable.** Each fenced block runs in its own shell invocation, so a `BASE=` assignment here is gone by the next task and every parent-tree count would @@ -783,7 +805,7 @@ its consumers skip, which is what the per-condition list exists to prevent: ``` base span -cond +cond

``` Tab-separated, one record per line, the leading keyword distinguishing them. **The `base` line is @@ -792,8 +814,11 @@ expected pair is `11`; the shape carries the values rather than assuming th dropped condition recorded here later needs no new format. **Both consumers validate every `cond` row**, not only the `span` rows. -**And every `cond` fragment is appended to the fragment table as well**, under the next free `P` id, -with the map's row citing that id. The table is where every pre-existing fragment lives, and Task 0 +**The fragment itself lives in the fragment table, and the map's row carries only its `P` id** — +that is the `

` field above, and both consumers resolve the text by looking the row up there. +**The map never stores the fragment text**, because two copies of an authored fragment is the +second-copy defect this cycle spent most of its findings on. Append the row to the table first, +under the next free `P` id, then write the `cond` line citing it. The table is where every pre-existing fragment lives, and Task 0 step 4's committed sweep runs the three conditions over **table rows** — a fragment that exists only in this ignored scratch file is outside that sweep, so a wrong one could certify a kept condition and be reused after an interruption on the strength of parsing and coverage alone. @@ -2590,8 +2615,10 @@ to pass CI, or an invariant-11 violation, published by a cycle that closed clean **Only a clean response issued against that exact `HEAD` closes the cycle.** A pass that was clean against an earlier tree, plus records committed afterwards, closes on a tree no pass reviewed — which is the same defect as reviewing the wrong range, arrived at from the other end. -- **One thing cannot exist before that pass: the pass's own findings file.** It is therefore the - **sole permitted post-review addition**, and step 8 asserts that it is the only one — every other +- **One thing cannot exist before that pass: the pass's own findings files.** A `full` Gate-B pass + writes **two** — the spec and quality branch files — and the close's dirty-set check requires both, + so the **pair** is the sole permitted post-review addition, and step 8 asserts that it is the only + one — every other path must already be in the reviewed `HEAD`. `.context/` moves none of the hook's fingerprint inputs, so the file changes nothing the review looked at; what would be wrong is a *second* delta riding along beside it. @@ -2706,9 +2733,11 @@ step and the observed state, and hands over. **It does not reset, re-commit or c `rm -f` below is therefore unreachable on that path: ```bash -git log -1 --pretty=%s # expect the real message, not a snapshot +case "$(git log -1 --pretty=%s)" in + [Ww][Ii][Pp]:*) echo "closing commit still reads as a snapshot — stop here and run Failure"; exit 1 ;; +esac test "$(git rev-parse HEAD^)" = "$BASE" || { echo "closing commit's parent is not \$BASE — stop here and run Failure"; exit 1; } -git status --porcelain # expect empty +test -z "$(git status --porcelain)" || { echo "tree dirty after the close — stop here and run Failure"; exit 1; } # The COMMIT BODY, not the source file: prepare-commit-msg and commit-msg hooks # rewrite git's copy after -F has read it, so the validated file proves nothing # about what landed. @@ -2726,7 +2755,9 @@ rm -f .context/loop-rule-base .context/loop-rule-reviewed-tip .context/loop-rule .context/loop-rule-sites .context/loop-rule-changed-sites \ .context/loop-rule-c.src .context/loop-rule-w.src \ .context/loop-rule-a.txt .context/loop-rule-b.txt .context/loop-rule-landed-msg -ls .context/loop-rule-* 2>/dev/null && { echo "cycle scratch survives the close — list it above"; exit 1; } +if ls .context/loop-rule-* >/dev/null 2>&1; then + echo "cycle scratch survives the close:"; ls .context/loop-rule-*; exit 1 +fi ``` **Every `loop-rule-*` scratch file goes, and the `ls` is what makes "the scratch files are removed" From 2b82702b367748628901d99bf5bdbbf52eb2726e Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Tue, 15 Sep 2026 15:09:53 +0200 Subject: [PATCH 140/181] docs(plans): apply Gate-A plan pass 30; rewrite Resume against all three topologies MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Four Blockers, three Majors. Five are my pass-29 fix applied inconsistently INSIDE the procedure it was fixing, so Resume is rewritten whole rather than patched again — patching sentence by sentence is what left the contradictions. - Resume said absence of cycle commits is not a staleness test and then kept the clause allowing replacement on exactly that. In the 8b topology HEAD IS the base and the chain is squashed into the index, so the clause condemns a state the plan calls valid. - Resume's success criterion still required a base that "passes Preparation", which refuses whenever a base exists and requires a clean tree. No valid re-entry could satisfy it. - Resume's progress check read the commit log only. After 8b the range is empty and the work is staged, so a literal reading treats a nearly finished cycle as untouched and invites every edit twice. It reads content now: `git diff --cached "$BASE"` for that topology, commits for the others. - The 8b topology was described as a commit above the WIP chain. `reset --soft` removes that chain from the ancestry: a landed closing commit is parented by BASE, and a failed one leaves HEAD at BASE. The three topologies are a table now, and every rule is read against all three. - Accounting row 4 said Preparation compares the approved inputs at the recorded base, which it cannot — it has none. Preparation compares at HEAD; Resume compares at BASE, and the split has a reason: a Gate-B fix may legitimately change HEAD's copy in a WIP snapshot. - Task 0 step 1 ran Preparation unconditionally and then branched on an existing base, but Preparation now exits on that file — the re-entry branch was unreachable and the plan was not idempotent at its first transition. - Three absence checks used bare `grep -c`, which prints 0 and EXITS 1 on no match, so the expected result surfaced as a failed shell step at Task 1 step 2, Task 6 step 4 and Task 11 step 5. --- .../gate-a-plan-om0bdd7udh-pass-30.md | 8 + .../2026-09-14-loop-rule-consolidation.md | 216 +++++++++--------- 2 files changed, 121 insertions(+), 103 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-30.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-30.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-30.md new file mode 100644 index 0000000..207619d --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-30.md @@ -0,0 +1,8 @@ +BLOCKER | high | The four procedures · Resume, lines 430–443 | Resume first states that absence of cycle commits is not a stale-base test, then repeats the opposite rule by allowing replacement when the base history has no part of the cycle. | In the expressly valid post-reset topology HEAD equals BASE and all implementation is staged, so this clause classifies the true base as another run and can discard the boundary used by Gate B and the final squash. | Delete the repeated history-has-no-cycle clause and keep only affirmative provenance evidence that is valid across all handoff topologies. +BLOCKER | high | The four procedures · Resume, lines 424–471 | Resume says its base uses its own checks because Preparation is first-entry-only, but its success criterion still requires a base that passes Preparation. | Every re-entry has a base file, and the staged-index handoff is dirty; both violate Preparation, so no valid re-entry can satisfy Resume's stated success condition. | Replace “passes preparation” with Resume's own full-object, self-resolving commit, ancestry, provenance and topology-aware checks. +BLOCKER | high | The four procedures · Resume, lines 449–471 | Resume validates implementation progress only from commits between BASE and HEAD plus task checkboxes, and recognizes success when ticked tasks match the commits present, even though its valid post-reset topology has an empty range and completed work only in the index. | A failed 8b close with the implementation staged cannot pass the progress check; following it literally either rejects a valid handoff or treats completed tasks as undone and risks duplicate edits. | Define progress validation over the actual topology: committed content for the WIP chain and the staged index content for the post-reset state, with checkboxes checked against that content rather than commits alone. +MAJOR | high | The four procedures · Resume, lines 452–461 | The first handoff bullet says a rejected 8b non-WIP commit sits above the cycle's WIP commits, but 8b soft-resets to BASE before committing. If the closing commit lands and a postcondition rejects it, its parent is BASE and the WIP commits are no longer in its ancestry; if the commit itself fails, HEAD is BASE with the index staged. | The topology description sends the implementation audit looking for history that cannot exist and can produce the wrong resume decision after a landed closing commit fails validation. | Split the 8a and 8b cases: retain the WIP-chain description only for a landed 8a record commit, and describe a landed 8b closing commit as one commit parented by BASE whose tree and closing records must be inspected. +BLOCKER | high | Accounting row 4 and Task 0 step 1, lines 211 and 680–685 | The accounting says the approved-input blobs are checked at the recorded base by Preparation, and Task 0 repeats that claim, but Preparation explicitly has no recorded base and compares HEAD; Resume contains no corresponding BASE-blob check. | The table claims a retained re-entry obligation that the procedures do not perform, while an executor is told that a Gate-B WIP edit to the spec is handled by a base comparison that exists nowhere on that path. | State the first-entry HEAD comparison accurately, add an explicit approved-input comparison at BASE to Resume if it is required to validate a pre-existing base, and update row 4 and Task 0 to name the procedure that actually performs each check. +MAJOR | high | Task 0 step 1, lines 680–704 | The step unconditionally says to run Preparation and then contains a branch for preserving an existing base, but Preparation now exits whenever that base exists. | The advertised partial-run re-entry branch is unreachable, so an interruption after the base is written cannot resume by following Task 0 and the plan is not idempotent at its first state transition. | Make the entry decision explicit before Preparation: existing base runs Resume and validation without Preparation; absent base runs Preparation and then writes the base once. +MAJOR | high | Task 1 step 2, Task 6 step 4, and Task 11 step 5 | These absence checks run bare grep count or grep listing commands while defining zero matches as success; grep prints zero but exits 1 on no match, and Task 11's final grep can likewise make the whole block return 1. | A correct installation is surfaced as a failed shell step at three verification points, so literal execution can stop even though the expected absence was established. | Use explicit predicates that accept grep status 1 as the expected no-match result, reject status 0 when absence is required, and still surface statuses above 1 as command errors. +END OF FINDINGS (7 total) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index 9f63dfc..5858bbf 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -208,7 +208,7 @@ independent reader did. | 1 | Branch is `loop-rule-consolidation` | **kept**, same shell — preparation | | 2 | Tree clean before the base is recorded | **kept**, same shell — preparation | | 3 | `ba15e83` is an ancestor of `HEAD` | **kept**, same shell — preparation | -| 4 | The three approved inputs' blobs equal their `ba15e83` versions, compared **at the recorded base** | **kept**, same shell — preparation | +| 4 | The three approved inputs' blobs equal their `ba15e83` versions | **kept, and split by entry**: Preparation compares them **at `HEAD`**, because a first entry has no recorded base; **Resume compares them at `$BASE`**, because a Gate-B fix may legitimately have changed `HEAD`'s copy in a `WIP:` snapshot. An earlier table row claimed Preparation did the base comparison, which it never could | | 5 | Never overwrite an existing base file | **kept** as an obligation; the shell branch becomes one line of the resume procedure | | 6 | A pre-existing base is an ancestor of `HEAD` (`merge-base --is-ancestor`) | **kept**, same shell — resume | | 7 | Only this run's `WIP:` commits lie between base and `HEAD` | **kept as an observation, dropped as a staleness test.** It is the *normal pre-close* topology. Three others are reachable and none makes the base stale: the handoff's rejected commit above the cycle's; `HEAD` at the base with the implementation in the index; and target §A3's accidental non-`WIP` commit or amend mid-cycle. **Staleness is decided by ancestry and by whether the history is this cycle's**, never by a commit subject | @@ -418,91 +418,88 @@ and records the closing commit carries are the ones a pass actually validated. ### Resume — re-entering after an interruption -**What is checked.** Whether the recorded base is this run's; whether the scratch artifacts belong -to it and are complete; and how far the implementation got. - -- **The base**, by its **own** checks, not Preparation's — Preparation is first-entry-only and - requires a clean tree, which the handoff's staged-index topology can never have. Read the file, - require a full 40-hex object name that resolves to itself as a commit, and require it to be an - **ancestor of `HEAD`**. **A base that is not an ancestor is stale.** **Never overwrite one you did - not just write.** - - **"Its history contains no part of this cycle" is NOT a staleness test** — in the post-reset - topology `HEAD` *is* the base and every cycle commit has been squashed away into the index, so - that predicate is necessarily true of a perfectly valid state. **Replace a base only on - affirmative evidence of another run**: it is not an ancestor of `HEAD`, or the recorded value - names a commit this branch never contained. - - **A non-`WIP` commit after the base does not make it stale, and inferring that from the subject - is wrong.** Target §A3 names an accidental non-`WIP` commit mid-cycle as a **reachable state**: it - resets what the hook counts, the cycle stays open until the conditions hold, and it can leave the - `WIP:` snapshot as an ancestor — or, as an amend, replace the tip. In both the base is still the - true boundary, and deleting it would drop earlier implementation out of Gate B's range and out of - the final reset. **Keep the base, inspect the state, and let §A3 decide what the stray commit - costs.** Replace a base only on evidence it belongs to **another run** — it is not an ancestor, or - its history has no part of this cycle in it. -- **Each scratch artifact**, by its own `base` line **and** its own completeness predicate: - - `.context/loop-rule-untouched` — every line parses as `base`, `span` or `cond`, and every kept - condition in the five regions appears in exactly one `span` or one `cond`. - - `.context/loop-rule-baseline-diff.txt` — every line parses as `base`, a `site` record, or diff - output belonging to the site above it, and **every inventoried site has a `site` record**. -- **The implementation**, by reading the commits between the base and `HEAD` and the plan's own task - checkboxes. - -**Three topologies are valid here, not one.** The `WIP:`-only shape is the normal pre-close one. The -handoff deliberately leaves two others, and Resume must recognise them rather than call the base -stale: - -- **After an 8a or 8b commit was rejected**: a non-`WIP` commit, or a `WIP:` findings commit, sits - at `HEAD` above the cycle's `WIP:` commits. **The base is not stale** — it is the same base, with - one more commit on top that a closure condition refused. -- **After `reset --soft` ran and the closing commit failed**: `HEAD` **is** the base and the entire - implementation is in the **index**, so `git log "$BASE"..HEAD` is empty and the task checkboxes - are the only record of how far the work got. **An empty log here is not an empty cycle.** - -**In both, keep the original base and change nothing until a person has chosen a §A route.** Read -`HEAD`, the index, the worktree and the cycle values together — `git status --porcelain ---untracked-files=all`, `git diff --cached --stat`, `git log --oneline -3` — and hand that reading -over. **There is no subject-based stale-base rule left**, and no absence-of-cycle-commits rule -either: the index-only topology has neither. Replace the base only on the affirmative evidence -named above. - -**How success is recognised.** A base that passes preparation, artifacts whose `base` line matches -it and whose coverage is complete, and a task list whose ticked entries match the commits present. - -**On deviation, and the answer differs by why you are here.** An artifact that fails either test is -**deleted and rebuilt from the `$BASE` blobs** — never reused, never repaired in place. A same-base -partial file is the one shape a `base` line alone cannot catch, which is why the completeness -predicate exists. - -**But not while a failed close is waiting on a person.** In either handoff topology the rule is -change nothing until a §A route is chosen, and an unconditional delete-and-rebuild here would -destroy the scratch state the handoff deliberately preserved — before anyone decided whether to -retry, owe a pass, or park. **Report the invalid artifact and leave it.** Rebuild it once the chosen -route authorizes continuing, and record why the old one is no longer evidence. +**Resume owns every entry after the first.** Preparation is first-entry-only and refuses when a base +file exists, so nothing here defers to it: the checks below are Resume's own, and **none of them +requires a clean tree** — two of the three valid topologies do not have one. -**Steps that describe the tree at `$BASE` are validated on re-entry, not re-run against the -worktree.** After a text task the worktree carries this plan's own edits, and rebuilding the -baseline from it would fold introduced drift into the inherited-drift record — the one distinction -Task 14 depends on. +**The three topologies, and the whole procedure reads against all three.** ---- +| Topology | `HEAD` | `$BASE..HEAD` | Where the work is | +|---|---|---|---| +| **Normal, mid-implementation** | the last `WIP:` snapshot | this run's `WIP:` commits, and possibly a §A3 stray commit or amend | committed | +| **8a rejected** | a `WIP:` findings commit, or a stray commit, **above** the cycle's `WIP:` chain | those commits | committed | +| **8b rejected** | either `$BASE` itself, if the closing commit never landed, **or one commit parented by `$BASE`**, if it landed and a postcondition refused it | **empty**, or that one commit — the `WIP:` chain is gone either way, squashed by `reset --soft` | the **index**, or that one commit's tree | -## File Structure +**The third row is the one every rule has to be re-read against.** `reset --soft` removes the +`WIP:` chain from the ancestry, so after 8b there is no chain to find; an empty `$BASE..HEAD` there +means the work is staged, not absent. -| File | Responsibility in this change | -|---|---| -| `CLAUDE.md` | canonical §5 (and one §4 line). Receives §A–§E, §G, §H and §F's sixteen prompt-copy items. | -| `plugins/dev-workflow/commands/workflow-init.md` | the scaffolded mirror. Receives the same, byte-identical, minus the deliberate divergences the inventory records. | -| `plugins/dev-workflow/hooks/codex-gate.sh` | seven `note` strings, both channels each (§F items 10–13, 15–17). No behaviour change. | -| `plugins/dev-workflow/hooks/codex-gate.test.sh` | three `expected_ctx` and three `expected_msg` exact-match expectations, plus every other assertion, label or comment naming a replaced string. | -| `plugins/dev-workflow/.claude-plugin/plugin.json` | `version` `0.11.0 → 0.12.0`. | -| `plugins/dev-workflow/CHANGELOG.md` | the 0.12.0 entry, newest first. | -| this plan | the 135-condition disposition (below), the next-state table (Task 13), and the verification pairs each task builds. | +- [ ] **Validate the base — Resume's own checks, not Preparation's** ---- +```bash +test -s .context/loop-rule-base || { echo "no base recorded — this is a first entry, run Preparation"; exit 1; } +BASE=$(cat .context/loop-rule-base) +test "${#BASE}" -eq 40 || { echo "base is not a full 40-character object name"; exit 1; } +case "$BASE" in *[!0-9a-f]*) echo "base is not an object name: $BASE"; exit 1 ;; esac +test "$(git rev-parse --verify "$BASE^{commit}")" = "$BASE" || { echo "base does not resolve to itself as a commit"; exit 1; } +git merge-base --is-ancestor "$BASE" HEAD || { echo "base is NOT an ancestor of HEAD"; exit 1; } +for f in target-text design condition-inventory; do + p="docs/superpowers/specs/2026-09-10-loop-rule-consolidation-$f.md" + test "$(git rev-parse "$BASE:$p")" = "$(git rev-parse "ba15e83:$p")" \ + || { echo "$p differs at the base from its approved version"; exit 1; } +done +``` + +**The approved-input comparison is at `$BASE` here, and at `HEAD` in Preparation.** They are +different revisions on purpose: Task 15 step 7 permits a Gate-B fix to update the spec in a `WIP:` +snapshot, so on a re-entry `HEAD`'s copy may legitimately differ while the revision the tasks were +derived against does not. + +**Replace a base only on affirmative evidence it belongs to another run** — it fails one of the +checks above. **Neither a commit subject nor an absence of cycle commits is evidence**: §A3's stray +non-`WIP` commit leaves the base valid, and the 8b topology has an empty range by construction, so +both tests would condemn states this plan calls valid. + +- [ ] **Validate the scratch artifacts** + +Each by its `base` line **and** its completeness predicate: + +- `.context/loop-rule-untouched` — every line parses as `base`, `span` or `cond`, and every kept + condition in the five regions appears in exactly one `span` or one `cond`. +- `.context/loop-rule-baseline-diff.txt` — every line parses as `base`, a `site` record, or diff + output belonging to the site above it, and **every inventoried site has a `site` record**. + +**On failure the answer depends on why you are here.** Outside a handoff: delete and rebuild from +the `$BASE` blobs — never reuse, never repair in place, because a same-base partial file is the one +shape a `base` line alone cannot catch. **Inside a handoff, while a failed close waits on a person: +report the invalid artifact and change nothing.** Rebuilding it would destroy the scratch state the +handoff preserved, before anyone chose a §A route. -## The condition disposition — all 135, by passage +- [ ] **Establish how far the implementation got — against the topology, not against the log** + +**Read the content, not only the commits.** In the first two topologies that is the commits between +`$BASE` and `HEAD`. **In the 8b topology it is `git diff --cached "$BASE"`** — the staged tree, plus +the landed closing commit's tree where one exists. A ticked checkbox is confirmed by the change +being *present in that content*, wherever the content lives. + +```bash +git log --oneline "$BASE"..HEAD # empty in the 8b topology; that is not an empty cycle +git diff --cached --stat "$BASE" # the staged implementation, where reset --soft left it +git status --porcelain --untracked-files=all +``` + +**A task whose checkbox is ticked but whose change is in neither place was not completed** — and one +whose change is present with the box unticked is completed. **Deciding from the commit log alone +reads the 8b topology as an untouched cycle and invites every edit to be made twice.** + +**How success is recognised.** The base passes Resume's own checks; the artifacts match it and are +complete, or are reported as invalid and left alone; and the plan's task list has been reconciled +against the content the topology actually holds. + +**Steps that describe the tree at `$BASE` are validated on re-entry, not re-run against the +worktree.** After a text task the worktree carries this plan's own edits, and rebuilding the +baseline from it would fold introduced drift into the inherited-drift record — the one distinction +Task 14 depends on. Story acceptance criterion 5 is satisfied here. Ids are `docs/superpowers/specs/2026-09-10-loop-rule-consolidation-condition-inventory.md`, snapshot at `7c0d475`. **Where the tree and the inventory disagree, the tree wins and the accounting is what needs correcting** — check each condition against the real file before marking it done. @@ -677,21 +674,25 @@ preservation fragments that an earlier draft chose afterwards. **Interfaces:** - Produces: `$BASE` (the parent commit every later task's counterfactual half runs against), and the five untouched passage ranges recorded as **anchor spans plus a per-condition fragment list** — never as absolute line numbers, for the reason step 2 gives. -- [ ] **Step 1: Run the preparation procedure, then record the base** - -**`## The four procedures` · Preparation** holds the checks and the shell: branch, clean tree, -`ba15e83` an ancestor, and the three approved inputs compared by blob **at the recorded base** -rather than at `HEAD` — Task 15 step 7 permits a Gate-B fix to update the spec in a `WIP:` snapshot, -so on a re-entry `HEAD`'s target text legitimately differs while the base's does not. +- [ ] **Step 1: Decide which entry this is, then run that procedure** -Then record the base, and **never overwrite one you did not just write**: +**The base file decides it, and nothing else does:** ```bash if [ -s .context/loop-rule-base ]; then - echo "base already recorded: $(cat .context/loop-rule-base) — validating, not overwriting" + echo "base recorded — this is a RE-ENTRY: run Resume, not Preparation" else - git rev-parse HEAD > .context/loop-rule-base + echo "no base — this is a FIRST ENTRY: run Preparation, then record the base below" fi +``` + +**First entry.** `## The four procedures` · Preparation holds the checks and the shell: branch, +clean tree, no base file, `ba15e83` an ancestor, and the three approved inputs compared by blob **at +`HEAD`** — which is the revision the tasks are about to be derived against. Then record the base, +**and never overwrite one you did not just write**: + +```bash +git rev-parse HEAD > .context/loop-rule-base BASE=$(cat .context/loop-rule-base) # Not "non-empty": a symbolic value such as HEAD passes every check below and # then RESOLVES DIFFERENTLY as WIP commits accrue, moving the reviewed range, @@ -699,20 +700,18 @@ BASE=$(cat .context/loop-rule-base) case "$BASE" in [0-9a-f][0-9a-f][0-9a-f][0-9a-f][0-9a-f][0-9a-f][0-9a-f][0-9a-f]*) ;; *) echo "recorded base is not an object name: $BASE"; exit 1 ;; esac test "${#BASE}" -eq 40 || { echo "recorded base is not a full 40-character object name"; exit 1; } test "$(git rev-parse --verify "$BASE^{commit}")" = "$BASE" || { echo "recorded base does not resolve to itself as a commit"; exit 1; } -git merge-base --is-ancestor "$BASE" HEAD || { echo "recorded base is NOT an ancestor of HEAD — stale"; exit 1; } -git log --oneline "$BASE"..HEAD # READ this; it does not decide staleness — see below ``` +**Re-entry.** `## The four procedures` · Resume validates the existing base, the scratch artifacts +and how far the implementation got — with its own checks, against all three topologies, and without +requiring a clean tree. **Do not run Preparation on a re-entry**: it refuses as soon as it sees the +base file, which is exactly what makes this branch reachable. + **Re-running Task 0 after a partial implementation must not re-record the base.** It would capture the current WIP tip, and both Gate B's range and the final reset would then start *after* every edit made so far — prompt and hook changes squashed into the closing commit without entering a review -range. **A pre-existing value is validated, never trusted**: `git log "$BASE"..HEAD` does not test -ancestry, which is how a base from an abandoned branch passes; `merge-base --is-ancestor` is the -test the sentence names. **What the log shows does not decide staleness** — §A3's accidental -non-`WIP` commit and the handoff's rejected close both leave this base valid, and the post-reset -topology shows an empty log with the whole cycle in the index. **Anything other than this run's -`WIP:` commits is a state to read, not a verdict: hand it to `## The four procedures` · Resume**, -which replaces a base only on affirmative evidence it belongs to another run. +range. The branch above is what prevents it: a recorded base sends this task to Resume, which +validates and never overwrites. **Persist it to a file, not to a shell variable.** Each fenced block runs in its own shell invocation, so a `BASE=` assignment here is gone by the next task and every parent-tree count would @@ -1021,10 +1020,17 @@ Expected: exactly one hit per file. - [ ] **Step 2: Verify the block is absent before installing** ```bash -grep -cF 'How a cycle ends — one ordering, stated here and referenced everywhere else' CLAUDE.md plugins/dev-workflow/commands/workflow-init.md +for f in CLAUDE.md plugins/dev-workflow/commands/workflow-init.md; do + n=$(grep -cF 'How a cycle ends — one ordering, stated here and referenced everywhere else' "$f"); st=$? + test $st -le 1 || { echo "grep failed (status $st) on $f"; exit 1; } + test "$n" = 0 || { echo "$f already carries the block ($n hits) — stop and reconcile"; exit 1; } +done ``` -Expected: `0` for both. If either is non-zero the block is already partly installed — stop and reconcile. +Expected: zero hits in both, and the block above says so as a **predicate**. **A bare `grep -c` +would not**: it prints `0` and **exits 1** on no match, so the expected result would surface as a +failed shell step. The same shape is used at Task 6 step 4 and Task 11 step 5, which are the other +two places this plan asserts an absence. - [ ] **Step 3: Install §A1, §A2 and §A3 in both copies** @@ -1518,7 +1524,9 @@ carried it, which is the divergence this task removes. - [ ] **Step 4: Confirm g4 is gone from C** ```bash -grep -c '2026-08-29-loop-rule-consolidation-story.md' CLAUDE.md +n=$(grep -c '2026-08-29-loop-rule-consolidation-story.md' CLAUDE.md); st=$? +test $st -le 1 || { echo "grep failed (status $st)"; exit 1; } +test "$n" = 0 || { echo "g4 still present in CLAUDE.md ($n hits)"; exit 1; } ``` Expected: `0`. The story path may still appear in `docs/` — this check is scoped to `CLAUDE.md`. @@ -2026,8 +2034,10 @@ Expected: exit 0 from both. **`HOOK_SH` selects the shell the hook runs under; w - [ ] **Step 5: Confirm no verdict vocabulary survives** ```bash -grep -c 'Gate B satisfied\|Gate B not satisfied\|Gate A satisfied' plugins/dev-workflow/hooks/codex-gate.test.sh -grep -ni 'satisfied' plugins/dev-workflow/hooks/codex-gate.test.sh +n=$(grep -c 'Gate B satisfied\|Gate B not satisfied\|Gate A satisfied' plugins/dev-workflow/hooks/codex-gate.test.sh); st=$? +test $st -le 1 || { echo "grep failed (status $st)"; exit 1; } +test "$n" = 0 || { echo "gate-verdict vocabulary survives ($n hits)"; exit 1; } +grep -ni 'satisfied' plugins/dev-workflow/hooks/codex-gate.test.sh || true # every hit disposed of in writing ``` Expected: **`0` from the first**, and **every remaining hit of the second disposed of in writing** — From 1e1638faa33cfa057ccc39cacfb03d3d6fa2a5c9 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Tue, 15 Sep 2026 15:25:38 +0200 Subject: [PATCH 141/181] docs(plans): apply Gate-A plan pass 31; run the loop through the ordering it installs MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Two Blockers, two Majors, one Nit. Down from seven and four. - Step 7 committed a non-closing pass and issued the next review without routing it through the ordering this change installs: no wait for a standing source block, no composition of open suspension answers, and no parked state for a stop answer — which acceptance criterion 4 requires to be distinct. The plan was executing a loop its own product forbids. - 8a checked that the record commit changed exactly the two findings paths and never looked at their content, while the plan elsewhere states that hooks rewrite staged files in place. A rewritten file keeps its pathname, so the path check, the parent check and the clean-tree check all pass and 8b closes on evidence different from the response whose eligibility was judged. The validated blob ids are pinned before staging and compared after the commit. - The topology table had one "8a rejected" row assuming the commit landed. 8a can also reject before it lands, or after it lands on the clean-tree check. Two rows now, and the progress audit reads HEAD, index and worktree CONTENT in every topology — `git status` names paths, and the delta that causes an 8a handoff is exactly the one a name cannot describe. - Two statements still said the handoff "preserved" scratch state and that a failure leaves every file "exactly where it is", against Failure's own admission that the failed operation may have rewritten anything. The plan performs no cleanup; it promises nothing about what survives, and Resume validates rather than trusts. - Task 4's NEW-column heading named §C where two rows take their fragment from the installed §A. --- .../gate-a-plan-om0bdd7udh-pass-31.md | 6 +++ .../2026-09-14-loop-rule-consolidation.md | 52 +++++++++++++++---- 2 files changed, 49 insertions(+), 9 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-31.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-31.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-31.md new file mode 100644 index 0000000..ab9bd60 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-31.md @@ -0,0 +1,6 @@ +BLOCKER | high | Task 15 step 7, lines 2573–2589 | The instruction “Every pass that does not close ends the same way” commits the pass and issues the next review without first routing the pass through the source-block and suspension branches; the later example mentions an answered suspension but never conditions the command on every required answer, and it gives a stop answer no parked path. | A Gate-B pass carrying an unrepaired source block, an outstanding membership or question hold, or a health stop can be followed immediately by another pass, so the execution procedure violates the ordering it installs and acceptance criterion 4's required distinct parked state. | After recording the pass, explicitly apply the installed ordering: wait for source repair and reread where a source block stands, collect every suspension answer and issue another pass only when their composition yields continue, park on stop until an explicit later continue, and enter Close only from the clean-completion branch. +BLOCKER | high | Close condition 4 and Task 15 step 8a, lines 333–338 and 2699–2716 | The record commit checks only that the changed-path set equals the two candidate findings paths; it never revalidates the committed blobs or proves they are the validated files the candidate pass produced, even though the plan expressly recognizes that commit hooks can rewrite staged files in place. | A hook can rewrite one findings file while keeping the same pathname and let the commit, path-set check, parent check, and clean-tree check all pass; 8b then closes on a malformed or non-clean logical pass and publishes evidence different from the response whose eligibility was evaluated. | Before 8a, persist the validated blob identities or equivalent exact content identities of both branch files; after the record commit, validate the committed blobs against those identities and rerun the findings-file structural and clean/eligibility checks, routing any difference through Failure without restoring anything. +MAJOR | high | Resume topology table and progress audit, lines 425–493 | The “8a rejected” row says the findings commit landed and all work is committed, but 8a can reject because the commit itself did not land or because its post-commit clean-tree check found an uncommitted delta; Resume then audits only commits for the first two topologies and only the staged tree for 8b, while `git status` reports names but does not inspect the uncommitted content. | A literal resume can classify an 8a failure as normal or reconcile task checkboxes against committed content while ignoring the very worktree or index delta that caused the handoff, leading to the wrong retry, repair, or further-pass decision. | Extend the topology table to cover 8a failures both before and after a commit lands, and reconcile progress and closure inputs against HEAD, the index, and the worktree contents for every topology, including the full staged and unstaged diffs rather than stats or status alone. +MAJOR | high | Resume lines 472–476 and Task 15 step 8 cleanup explanation lines 2773–2778 | Two surviving statements still say the handoff “preserved” scratch state and that a failure leaves every scratch file “exactly where it is”, contradicting Failure's explicit admission that the failed operation may already have rewritten or destroyed state and Task 15 step 7b's corrected rule not to claim preservation. | A resuming executor can trust a cycle value merely because cleanup was skipped, even though the failed hook or commit may have changed or removed it, undermining the state-based handoff this rewrite is meant to establish. | Say only that the plan performs no cleanup after failure; require Resume to enumerate and validate whatever scratch values actually survive, and remove every claim that the failed operation preserved them. +NIT | high | Task 4 step 4 table, lines 1295–1300 | The NEW-column heading says every value is taken from the installed §C block, but both `c14` rows explicitly take their NEW observations from the ordering in §A. | The row bodies are clear, but the false heading sends a literal executor to the wrong source before the rows correct it. | Rename the heading to cover the installed destination text, or qualify that only the `c4` and `c8` NEW fragments come from §C. +END OF FINDINGS (5 total) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index 5858bbf..511a640 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -427,7 +427,8 @@ requires a clean tree** — two of the three valid topologies do not have one. | Topology | `HEAD` | `$BASE..HEAD` | Where the work is | |---|---|---|---| | **Normal, mid-implementation** | the last `WIP:` snapshot | this run's `WIP:` commits, and possibly a §A3 stray commit or amend | committed | -| **8a rejected** | a `WIP:` findings commit, or a stray commit, **above** the cycle's `WIP:` chain | those commits | committed | +| **8a rejected, no commit landed** | the last `WIP:` snapshot, unchanged | this run's `WIP:` commits | committed, **plus whatever the failed attempt left in the index or worktree** | +| **8a rejected after its commit landed** | a `WIP:` findings commit **above** the cycle's `WIP:` chain | those commits | committed, **plus any delta the post-commit clean-tree check found** | | **8b rejected** | either `$BASE` itself, if the closing commit never landed, **or one commit parented by `$BASE`**, if it landed and a postcondition refused it | **empty**, or that one commit — the `WIP:` chain is gone either way, squashed by `reset --soft` | the **index**, or that one commit's tree | **The third row is the one every rule has to be re-read against.** `reset --soft` removes the @@ -472,8 +473,10 @@ Each by its `base` line **and** its completeness predicate: **On failure the answer depends on why you are here.** Outside a handoff: delete and rebuild from the `$BASE` blobs — never reuse, never repair in place, because a same-base partial file is the one shape a `base` line alone cannot catch. **Inside a handoff, while a failed close waits on a person: -report the invalid artifact and change nothing.** Rebuilding it would destroy the scratch state the -handoff preserved, before anyone chose a §A route. +report the invalid artifact and change nothing.** Rebuilding it would remove evidence before anyone +chose a §A route. **This plan performs no cleanup after a failure — it does not claim the failed +operation left anything intact.** Enumerate which scratch values actually survive and validate each; +a value is trustworthy because it passed a check, never because cleanup was skipped. - [ ] **Establish how far the implementation got — against the topology, not against the log** @@ -484,10 +487,16 @@ being *present in that content*, wherever the content lives. ```bash git log --oneline "$BASE"..HEAD # empty in the 8b topology; that is not an empty cycle -git diff --cached --stat "$BASE" # the staged implementation, where reset --soft left it +git diff --cached "$BASE" # the staged content — the FULL diff, not --stat +git diff # the unstaged content git status --porcelain --untracked-files=all ``` +**Read all three, in every topology.** `git status` names paths and says nothing about what is in +them, and `--stat` counts lines. **The delta that caused an 8a handoff is precisely the one a +path listing cannot describe** — a rewritten staged file keeps its name — so reconcile against +`HEAD`, the index and the worktree **contents** together, whichever topology you are in. + **A task whose checkbox is ticked but whose change is in neither place was not completed** — and one whose change is present with the box unticked is completed. **Deciding from the commit log alone reads the 8b topology as an untouched cycle and invites every edit to be made twice.** @@ -1292,7 +1301,7 @@ either could land while the other survived and it still reported a pass, and `c4 observation at all. **A later draft reintroduced exactly that pair** by naming row P4, whose OLD is `c14`, and then taking its NEW from §C's re-raised-dismissal clause, which is `c8`. -| Edit | OLD | NEW, taken from the installed §C block | +| Edit | OLD | NEW, from the installed destination text — §C here, §A where the row says so | |---|---|---| | `c4`, the widened third condition | the `c4` row (step 1) | the clause §C puts in place of "a missing one means keep going" | | `c8`, the re-raised dismissal | the `c8` row (step 1) | `a recurrence failing them being an ordinary fresh finding` — **install that clause's line unwrapped** so the fragment sits wholly on one line | @@ -2582,6 +2591,20 @@ git rev-parse HEAD > .context/loop-rule-reviewed-head # the head the NEXT call cat .context/loop-rule-reviewed-head ``` +**Committing the pass is not the same as being allowed to issue the next one.** After recording it, +**run the pass through the ordering this change installs** — the plan must not execute a loop its own +product forbids: + +- **A source block standing** — wait for the repair and reread by the route §A gives; no further pass + until that is done. +- **Any suspension open** — a membership stop, a new-question stop, a two-tell stop, a clearly-stuck + surface. **Collect every answer, compose them, and issue another pass only when the composition + yields continue.** +- **A stop answer** → **park**: open, not running, **spending no passes**, restarted only by an + explicit later continue. **There is no "commit and carry on" from a stop**, and acceptance + criterion 4 requires that parked state to be distinct. +- **Only the clean-completion branch enters Close.** + **A non-closing pass that owes no repair still commits.** A Minor-only clean pass below the floor, or an answered suspension that changes no artifact, leaves its findings files tracked and dirty — and the next pass's files pile up beside them, after which the close procedure's dirty-set check can @@ -2703,6 +2726,9 @@ expected=$(printf '%s\n' $FINAL | sort) actual=$(git status --porcelain -z | tr '\0' '\n' | sed -n 's/^.\{3\}//p' | sed '/^$/d' | sort) test "$expected" = "$actual" || { echo "dirty set is not exactly this pass's findings files:"; git status --porcelain; exit 1; } +# Pin the validated content BEFORE staging: a hook can rewrite a staged file in +# place, leaving its pathname — and therefore the changed-path check — unchanged. +for f in $FINAL; do git hash-object "$f"; done > .context/loop-rule-final-blobs # shellcheck disable=SC2086 git add $FINAL # Any rejection below stops and goes through the Failure procedure, which reports @@ -2711,6 +2737,12 @@ git commit -m "WIP: Gate-B findings files" || { echo "record commit FAILED — s test "$(git show --name-only --pretty=format: HEAD | sed '/^$/d' | sort)" = "$expected" \ || { echo "record commit changed paths beyond this pass's findings files — stop here and run Failure"; exit 1; } test "$(git rev-parse HEAD^)" = "$HEADREV" || { echo "record commit's parent is not the reviewed head — stop here and run Failure"; exit 1; } +# The committed blobs, not the paths: same names can hold different bytes. +for f in $FINAL; do git rev-parse "HEAD:$f"; done > .context/loop-rule-committed-blobs +diff .context/loop-rule-final-blobs .context/loop-rule-committed-blobs \ + || { echo "a findings file was rewritten between validation and commit — stop here and run Failure"; exit 1; } +# Then re-run the findings-file structural check on the committed content, and +# re-establish that this logical pass is still the eligible one it was judged as. test -z "$(git status --porcelain)" || { echo "tree not clean after the record commit — stop here and run Failure"; exit 1; } git rev-parse HEAD > .context/loop-rule-reviewed-tip # the closing tip: 8b's precondition, NOT a reset target @@ -2764,7 +2796,8 @@ rm -f .context/loop-rule-base .context/loop-rule-reviewed-tip .context/loop-rule .context/loop-rule-untouched .context/loop-rule-baseline-diff.txt \ .context/loop-rule-sites .context/loop-rule-changed-sites \ .context/loop-rule-c.src .context/loop-rule-w.src \ - .context/loop-rule-a.txt .context/loop-rule-b.txt .context/loop-rule-landed-msg + .context/loop-rule-a.txt .context/loop-rule-b.txt .context/loop-rule-landed-msg \ + .context/loop-rule-final-blobs .context/loop-rule-committed-blobs if ls .context/loop-rule-* >/dev/null 2>&1; then echo "cycle scratch survives the close:"; ls .context/loop-rule-*; exit 1 fi @@ -2773,9 +2806,10 @@ fi **Every `loop-rule-*` scratch file goes, and the `ls` is what makes "the scratch files are removed" true rather than asserted.** An earlier draft deleted three of them and claimed the terminal state, leaving a later run to inherit a closing message, a baseref and an untouched map — each of which -some check then has to detect or overwrite piecemeal. **This runs only on a successful close**; a -failure leaves every one of these files exactly where it is, which is what the resume decision is -read from. **The plan's records +some check then has to detect or overwrite piecemeal. **This runs only on a successful close**; after a +failure **the plan performs no cleanup at all**. That is not a promise the files are unchanged — a +hook that failed may have rewritten any of them — so Resume enumerates what survives and validates +it rather than trusting it. **The plan's records are not among these**: they live in the plan and in `.context/codex-reviews/`, both tracked, both already in the closing commit. From e429008eb6abd926264505baac8557c38e590433 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Tue, 15 Sep 2026 15:38:35 +0200 Subject: [PATCH 142/181] docs(plans): apply Gate-A plan pass 32; the ordering decides before a next call exists MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Two Blockers, one Major, one Minor. Findings 5 -> 4, Blockers 2 -> 2. - Pass 31 added the ordering bullets to step 7 and left them AFTER the block that already committed and recorded the next reviewed head. A literal run still issued a pass over a standing source block, an unanswered suspension or a parked cycle. Step 7 is two ordered steps now: record the pass, then read it against the ordering — and the reviewed-head file is written on the continue route alone, because writing it is what makes a next call possible. - Close condition 5 requires a clean tree AND the tip when the closing invocation begins; 8b checked only the tip, and the one clean-tree check sat in the earlier 8a invocation. A staged edit made between them — a rewritten findings file included — survives reset --soft, lands in the closing commit, and leaves the tree clean afterwards, so every postcondition passes while unreviewed content ships. 8b re-checks it before the reset. - Pass 31's blob-identity repair existed only in the example shell, not among the governing Close conditions, and step 8 explicitly tells an executor facing a different state to derive the operation from the conditions — so that route could legally omit it. Condition 4 now carries three bullets: the path set and parent, the committed blob identities, and a structural plus eligibility recheck of the committed findings files, each with its subject, base, expected result and Failure transition. - Splitting the 8a state gave four topology rows while the surrounding prose still said three, called 8b "the third row" and audited "the first two". The rows are named rather than numbered now. --- .../gate-a-plan-om0bdd7udh-pass-32.md | 5 ++ .../2026-09-14-loop-rule-consolidation.md | 89 +++++++++++++------ 2 files changed, 66 insertions(+), 28 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-32.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-32.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-32.md new file mode 100644 index 0000000..1b804d6 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-32.md @@ -0,0 +1,5 @@ +BLOCKER | high | Task 15 step 7, lines 2582–2606 | The unchanged lead-in still commands the executor to record a new head and “issue the next pass” before the newly added ordering bullets decide whether any next pass is permitted; the code block likewise writes `.context/loop-rule-reviewed-head` for that next call before source blocks, suspension composition, or park are evaluated. | A literal execution can issue another pass while a source repair or suspension answer is outstanding, or after a stop answer should have parked the cycle, so the plan still runs the loop in a way target §A and acceptance criterion 4 forbid. | Route the completed pass through the ordering before preparing any next call: source blocks wait and reread, suspensions collect and compose every answer, stop parks, clean enters Close, and only the continue result records a new reviewed head and issues another pass. +BLOCKER | high | Task 15 step 8b, lines 2758–2769 | Close condition 5 requires both a clean tree and `HEAD` equal to the closing tip when the separate closing invocation begins, but 8b checks only `HEAD`; the sole clean-tree check is line 2746 in the earlier 8a invocation. | A staged edit made between invocations — including a rewrite of either validated findings file — survives `reset --soft`, is swept into the closing commit, and leaves the post-commit tree clean, so every implemented postcondition can pass while unreviewed content ships. | Re-check the existing clean-tree precondition in 8b after the tip comparison and before `reset --soft`; on any dirty index or worktree, stop without mutating and follow the stated handoff route. +MAJOR | high | Close condition 4 and Task 15 step 8a, lines 333–339 and 2709–2746 | Pass 31's exact-content repair exists only in the example shell, not among the six governing Close conditions, and the following comment promises a findings-file structural and logical-eligibility recheck without naming its comparison or expected result or performing it. | Step 8 explicitly tells an executor facing a different observed state to derive an operation from the stated conditions, so that route can legally omit the blob-identity protection; even the normal route has no executable or reader instruction discharging the promised structural and eligibility recheck. | Promote committed-content identity plus findings-file structure and logical-pass eligibility into Close's stated obligations, update the six-condition accounting, and specify the subject, comparison base, expected result, and Failure transition for each check. +MINOR | high | Resume lines 421–498 and Task 0 step 1 lines 713–716 | Splitting the rejected-8a state produced four topology rows, but the surrounding procedure still says “three topologies”, “two of the three”, calls 8b “the third row”, describes committed progress in only “the first two topologies”, and tells Task 0 to validate “all three”. | The stale exhaustive numbering contradicts the new table and can make a literal resume group or omit one of the two distinct 8a failure states that pass 31 added specifically because they require different content reconciliation. | Change every count and ordinal to the four-row topology, preferably naming normal, 8a-before-commit, 8a-after-commit, and 8b directly instead of referring to row positions. +END OF FINDINGS (4 total) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index 511a640..b14b636 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -330,8 +330,21 @@ an unapproved edit to a spec that the gates already closed. — 8a runs a commit in between, hooks can rewrite anything, and the file is under `.context/`, which is ignored, so no clean-tree check between the two invocations can see it change. A difference or a malformed record there is a Failure handoff, not a repair. -4. **Those files are committed** in their own invocation, and that commit **changes exactly those - paths**. Its parent is the reviewed head. Record the resulting commit as the **closing tip**, in +4. **Those files are committed** in their own invocation, and that commit satisfies **three things, + not one**: + - it **changes exactly those paths**, and its parent is the reviewed head; + - **the committed blobs are the validated ones** — the object ids recorded before staging equal + the ids at `HEAD` afterwards. A path check cannot see this: a hook rewriting a staged file in + place leaves the pathname untouched, which the plan states elsewhere and must therefore guard + here; + - **the committed findings files still satisfy the findings-file structure and still make this + the eligible logical pass they were judged as** — every line before the terminator a finding + line, the terminator exact, the count matching, both branch files present, and the + clean-or-zero-finding reading unchanged. Subject: the committed content. Base: the protocol in + §5 and the eligibility this candidate was issued on. **Any difference at any of the three is a + Failure handoff, not a repair.** + + Record the resulting commit as the **closing tip**, in `.context/loop-rule-reviewed-tip` — **a precondition value, not a restore target**: 8b refuses to reset unless `HEAD` is still exactly it, which is how a commit landing between the two invocations is caught. Nothing in this plan resets *to* it. @@ -420,9 +433,9 @@ and records the closing commit carries are the ones a pass actually validated. **Resume owns every entry after the first.** Preparation is first-entry-only and refuses when a base file exists, so nothing here defers to it: the checks below are Resume's own, and **none of them -requires a clean tree** — two of the three valid topologies do not have one. +requires a clean tree** — three of the four valid topologies need not have one. -**The three topologies, and the whole procedure reads against all three.** +**Four topologies, named rather than numbered, and the whole procedure reads against all four.** | Topology | `HEAD` | `$BASE..HEAD` | Where the work is | |---|---|---|---| @@ -431,9 +444,11 @@ requires a clean tree** — two of the three valid topologies do not have one. | **8a rejected after its commit landed** | a `WIP:` findings commit **above** the cycle's `WIP:` chain | those commits | committed, **plus any delta the post-commit clean-tree check found** | | **8b rejected** | either `$BASE` itself, if the closing commit never landed, **or one commit parented by `$BASE`**, if it landed and a postcondition refused it | **empty**, or that one commit — the `WIP:` chain is gone either way, squashed by `reset --soft` | the **index**, or that one commit's tree | -**The third row is the one every rule has to be re-read against.** `reset --soft` removes the -`WIP:` chain from the ancestry, so after 8b there is no chain to find; an empty `$BASE..HEAD` there -means the work is staged, not absent. +**The 8b row is the one every rule has to be re-read against.** `reset --soft` removes the `WIP:` +chain from the ancestry, so after 8b there is no chain to find; an empty `$BASE..HEAD` there means +the work is staged, not absent. **And the two 8a rows differ from each other**: before its commit +the history is untouched and the delta is loose, after it the findings commit sits above the chain +— pass 31 split them because they need different content reconciliation. - [ ] **Validate the base — Resume's own checks, not Preparation's** @@ -480,9 +495,10 @@ a value is trustworthy because it passed a check, never because cleanup was skip - [ ] **Establish how far the implementation got — against the topology, not against the log** -**Read the content, not only the commits.** In the first two topologies that is the commits between -`$BASE` and `HEAD`. **In the 8b topology it is `git diff --cached "$BASE"`** — the staged tree, plus -the landed closing commit's tree where one exists. A ticked checkbox is confirmed by the change +**Read the content, not only the commits.** In the normal and the two 8a topologies that is the +commits between `$BASE` and `HEAD`, **plus whatever the failed attempt left loose**. **In the 8b +topology it is `git diff --cached "$BASE"`** — the staged tree, plus the landed closing commit's +tree where one exists. A ticked checkbox is confirmed by the change being *present in that content*, wherever the content lives. ```bash @@ -712,7 +728,7 @@ test "$(git rev-parse --verify "$BASE^{commit}")" = "$BASE" || { echo "recorded ``` **Re-entry.** `## The four procedures` · Resume validates the existing base, the scratch artifacts -and how far the implementation got — with its own checks, against all three topologies, and without +and how far the implementation got — with its own checks, against all four topologies, and without requiring a clean tree. **Do not run Preparation on a re-entry**: it refuses as soon as it sees the base file, which is exactly what makes this branch reachable. @@ -2579,31 +2595,40 @@ branch file; one branch clean and the other not is not it. WIP tip while the repair sits in the worktree — and the final squash then publishes a fix no pass reviewed: -**Every pass that does not close ends the same way, whether or not it produced a repair.** Commit -that pass's findings files and any refreshed records, resolve the new head, **record it**, and issue -the next pass against exactly that value: +**Every pass that does not close ends the same way, whether or not it produced a repair — and it +ends in two steps, in this order.** First record the pass. Then **run it through the ordering this +change installs, and only its continue result prepares another call.** The plan must not execute a +loop its own product forbids. + +**Step one — record the pass. This commits; it does not authorize anything.** ```bash # Before committing a fix, re-run what the fix could have broken. git add -A && git commit -m "WIP: fix " # or: "WIP: pass records" where no repair was owed -rm -f .context/loop-rule-reviewed-tip # the previous candidate's closing tip -git rev-parse HEAD > .context/loop-rule-reviewed-head # the head the NEXT call is issued against -cat .context/loop-rule-reviewed-head ``` -**Committing the pass is not the same as being allowed to issue the next one.** After recording it, -**run the pass through the ordering this change installs** — the plan must not execute a loop its own -product forbids: +**Step two — read the pass against the ordering, before any next call exists:** -- **A source block standing** — wait for the repair and reread by the route §A gives; no further pass - until that is done. +- **A source block standing** → **wait** for the repair and the reread by the route §A gives. No + further pass until that is done. - **Any suspension open** — a membership stop, a new-question stop, a two-tell stop, a clearly-stuck - surface. **Collect every answer, compose them, and issue another pass only when the composition - yields continue.** + surface — → **collect every answer and compose them.** Another pass only where the composition + yields **continue**. - **A stop answer** → **park**: open, not running, **spending no passes**, restarted only by an explicit later continue. **There is no "commit and carry on" from a stop**, and acceptance - criterion 4 requires that parked state to be distinct. -- **Only the clean-completion branch enters Close.** + criterion 4 requires that parked state to be distinct. **Nothing below runs on this route.** +- **The clean-completion branch** → **Close**, not another pass. +- **Continue** → and only then: + +```bash +rm -f .context/loop-rule-reviewed-tip # the previous candidate's closing tip +git rev-parse HEAD > .context/loop-rule-reviewed-head # the head the NEXT call is issued against +cat .context/loop-rule-reviewed-head +``` + +**The reviewed-head file is written on the continue route alone**, because writing it is what makes +a next call possible: recording it before the ordering has spoken is how a pass gets issued over a +standing source block, an unanswered suspension or a parked cycle. **A non-closing pass that owes no repair still commits.** A Minor-only clean pass below the floor, or an answered suspension that changes no artifact, leaves its findings files tracked and dirty — @@ -2741,8 +2766,11 @@ test "$(git rev-parse HEAD^)" = "$HEADREV" || { echo "record commit's parent is for f in $FINAL; do git rev-parse "HEAD:$f"; done > .context/loop-rule-committed-blobs diff .context/loop-rule-final-blobs .context/loop-rule-committed-blobs \ || { echo "a findings file was rewritten between validation and commit — stop here and run Failure"; exit 1; } -# Then re-run the findings-file structural check on the committed content, and -# re-establish that this logical pass is still the eligible one it was judged as. +# Then re-run the findings-file structural check on the committed content and +# re-establish this pass's eligibility, per close condition 4's third bullet. +# Expected: every line before the terminator is a finding line, the terminator +# is exact, the count matches, both branch files are present, and the pass reads +# clean-or-zero-finding exactly as it did when issued. Any difference: Failure. test -z "$(git status --porcelain)" || { echo "tree not clean after the record commit — stop here and run Failure"; exit 1; } git rev-parse HEAD > .context/loop-rule-reviewed-tip # the closing tip: 8b's precondition, NOT a reset target @@ -2759,6 +2787,11 @@ above rather than assumed. BASE=$(cat .context/loop-rule-base); TIP=$(cat .context/loop-rule-reviewed-tip) test -n "$BASE" && test -n "$TIP" || { echo "BASE or TIP missing — 8a did not complete"; exit 1; } test "$(git rev-parse HEAD)" = "$TIP" || { echo "HEAD has moved since 8a — NOT resetting"; exit 1; } +# Close condition 5 is BOTH: the tip AND a clean tree, checked in THIS invocation. +# A staged edit made between 8a and 8b — a rewritten findings file included — +# survives reset --soft, lands in the closing commit, and leaves the tree clean +# afterwards, so every postcondition passes while unreviewed content ships. +test -z "$(git status --porcelain)" || { echo "tree not clean at 8b — NOT resetting; stop here and run Failure"; exit 1; } git reset --soft "$BASE" || { echo "reset --soft FAILED — stop here and run Failure; do NOT commit"; exit 1; } # Re-read the closing message here: it was validated before 8a, 8a ran a commit # (hooks can rewrite anything), and .context is ignored, so no porcelain check From b395da05087666d21a666c426ced159cc58b987d Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Tue, 15 Sep 2026 15:54:48 +0200 Subject: [PATCH 143/181] docs(plans): apply Gate-A plan pass 33; the post-close block re-reads its own base MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Three Blockers, three Majors, one Nit. - Pass 32 split 8b's post-close checks into their own fence, where $BASE is empty — every fence is a separate shell invocation, which this plan states. A CORRECT closing commit therefore failed the parent check and entered Failure every time: the plan had no successful close path at all. The block re-reads and guards the base. - Step 7b's changed-evidence-entry branch committed and issued another candidate directly, stepping over the ordering step 7 had just been made to obey. A standing source block, an open suspension or a stop answer binds there too; only the continue result records a head and issues a call. - Landed-message identity existed only in the example shell. Step 8 tells an executor facing a different state to derive the operation from the governing conditions, so a guard outside them is optional by construction. Close condition 6 now names four postconditions — subject, parent, clean tree, landed body — and the success predicate carries them. - An EMPTY base file satisfied neither route: `-s` called it a first entry while Preparation refused because the path exists. The route is decided by the path existing; an empty or malformed file is Resume's, and its transition is stated. - Accounting row 7 described three alternative topologies where Resume has four, losing the loose-delta and landed-close reconciliation obligations. - Task 1 step 6 justified staging with "Task 0's re-entry path requires a clean tree". Resume requires no clean tree and three of its four topologies do not have one; that claim would have rejected the states Resume exists for. The real cost — evidence outside its own task's snapshot — is stated instead. - "Either check" and "Both clean" over a block that performs four. --- .../gate-a-plan-om0bdd7udh-pass-33.md | 8 +++ .../2026-09-14-loop-rule-consolidation.md | 67 +++++++++++++------ 2 files changed, 55 insertions(+), 20 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-33.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-33.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-33.md new file mode 100644 index 0000000..07d6014 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-33.md @@ -0,0 +1,8 @@ +BLOCKER | high | Task 15 step 8 post-close checks | The post-close fenced block compares `HEAD^` with `$BASE`, but `$BASE` is assigned only in the preceding 8b fence; the plan states that every fence is a separate shell invocation, so this block receives an empty value. | A correct closing commit always fails the parent check and enters Failure, so the plan has no successful close path. | Read and validate `.context/loop-rule-base` again at the start of the post-close block, or keep all of 8b in one fenced invocation. +BLOCKER | high | Task 15 step 7b, changed evidence-entry route | When revalidation changes the evidence entry, this branch says to commit it, record a new head, and issue another candidate pass directly instead of reading the non-closing pass through the source-block, suspension, composition, and parked-state ordering. | A mandatory two-tell or clearly-stuck suspension, a standing source block, or a stop answer can be bypassed by the plan's own Gate-B loop, contradicting target §A and story criterion 4. | After committing the changed entry, route the pass through step 7's ordering and permit a new reviewed-head value and review call only from its continue result. +BLOCKER | high | Close procedure condition 6 and success recognition | Identity between the landed commit body and the revalidated closing-message file exists only in the example shell at Task 15 step 8; it is absent from the governing six Close conditions and from the stated success predicate even though step 8 says alternate operations are derived from those conditions. | An executor adapting the sample to an observed state may accept a hook-rewritten commit body that dropped or changed the provenance line, curve, exception marker, or evidence entry. | Add landed-message identity as a governing postcondition and make any mismatch enter Failure; include it in the Close success predicate. +MAJOR | high | Task 0 step 1 entry selection and Resume base validation | An existing empty `.context/loop-rule-base` satisfies neither route: `-s` selects first entry, Preparation refuses because the path exists, and Resume also calls an empty file a first entry. | An interruption after redirection creates a state the plan can neither resume nor initialize, breaking its re-entry and idempotency design. | Route on path existence, give the invalid or empty existing-base state to one procedure, and state the inspected delete-and-recreate or human-handoff transition explicitly. +MAJOR | high | Accounting row 7 | The row says the normal topology has three alternatives but omits the 8a-rejected-no-commit state as a distinct dirty-content topology and the 8b-rejected-after-a-closing-commit-landed state, while treating the §A3 stray commit as separate although Resume nests it under normal. | The mandatory accounting disagrees with the four-row Resume model and can lose the loose-delta and landed-close reconciliation obligations in a later rewrite. | Make row 7 use the same four topologies and nesting as Resume, explicitly covering both 8a outcomes and both 8b outcomes. +MAJOR | high | Task 1 step 6 staging rationale | The plan says Task 0's re-entry path requires a clean tree, directly contradicting Preparation, Resume, and Task 0 step 1, which reserve clean-tree enforcement for first entry and require Resume to accept three dirty or staged topologies. | An executor can reject the exact interrupted states Resume was introduced to reconcile, or mutate them merely to obtain a clean tree before re-entry. | Remove the clean-tree claim and state the actual cost of not staging the plan: evidence is absent from its task's WIP snapshot and may be swept into a later commit. +NIT | high | Task 15 step 8 post-close prose | The prose says “either check below” and then “Both clean” although the following block performs four checks: subject, parent, clean tree, and landed-message identity. | The stale count obscures which failures enter the Failure procedure and violates the prompt's count-versus-enumeration check. | Say “any check below” and “All pass” or enumerate the four checks explicitly. +END OF FINDINGS (7 total) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index b14b636..ba70d40 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -211,7 +211,7 @@ independent reader did. | 4 | The three approved inputs' blobs equal their `ba15e83` versions | **kept, and split by entry**: Preparation compares them **at `HEAD`**, because a first entry has no recorded base; **Resume compares them at `$BASE`**, because a Gate-B fix may legitimately have changed `HEAD`'s copy in a `WIP:` snapshot. An earlier table row claimed Preparation did the base comparison, which it never could | | 5 | Never overwrite an existing base file | **kept** as an obligation; the shell branch becomes one line of the resume procedure | | 6 | A pre-existing base is an ancestor of `HEAD` (`merge-base --is-ancestor`) | **kept**, same shell — resume | -| 7 | Only this run's `WIP:` commits lie between base and `HEAD` | **kept as an observation, dropped as a staleness test.** It is the *normal pre-close* topology. Three others are reachable and none makes the base stale: the handoff's rejected commit above the cycle's; `HEAD` at the base with the implementation in the index; and target §A3's accidental non-`WIP` commit or amend mid-cycle. **Staleness is decided by ancestry and by whether the history is this cycle's**, never by a commit subject | +| 7 | Only this run's `WIP:` commits lie between base and `HEAD` | **kept as an observation, dropped as a staleness test.** Resume's four topologies are the model: *normal* (this run's `WIP:` commits, §A3's stray commit or amend nested inside it); *8a rejected before its commit landed* (history untouched, the delta loose in index or worktree); *8a rejected after it landed* (a findings commit above the chain, plus any delta); and *8b rejected* (`HEAD` at the base with the work staged, **or** one commit parented by the base where the closing commit landed and a postcondition refused it). **None makes the base stale**, and staleness is decided by ancestry and provenance, never by a commit subject | | 8 | `$BASE` persisted to a file; every consumer guards it non-empty | **kept**, and **repaired**: three concrete blocks read it without the guard while this row claimed otherwise — the baseline extraction, Task 10's hook diff, and the first Gate-B call. The guard is in each of them now (pass 22) | | 9 | The five regions are located by anchor, one hit per pattern per file | **kept**, same shell — preparation | | 10 | Spans are **derived** from replacement extents, never hand-written | **kept** — the rule, unchanged | @@ -349,10 +349,16 @@ an unapproved edit to a spec that the gates already closed. reset unless `HEAD` is still exactly it, which is how a commit landing between the two invocations is caught. Nothing in this plan resets *to* it. 5. **The tree is clean**, and `HEAD` is still the closing tip, when the closing invocation begins. -6. **After the closing commit**: its subject is not a snapshot, and the tree is clean. - -**How success is recognised.** A single commit at `HEAD` whose subject is the real message and whose -parent is `$BASE`. **No tree comparison** — target §I parks a Gate-B tree-equality condition on +6. **After the closing commit**, four things: its **subject is not a snapshot**; its **parent is + `$BASE`**; the **tree is clean**; and its **body is identical to the revalidated closing-message + file**. The last is not decoration — `prepare-commit-msg` and `commit-msg` hooks rewrite git's + copy *after* `-F` has read it, so a validated file proves nothing about what landed, and a + dropped provenance line, curve, exception marker or evidence entry would pass every other check. + **Any of the four failing is a Failure handoff.** + +**How success is recognised.** A single commit at `HEAD` whose subject is the real message, whose +parent is `$BASE`, **whose body is identical to the revalidated closing message**, with a clean tree +behind it. **No tree comparison** — target §I parks a Gate-B tree-equality condition on Daniel's decision of 2026-09-13, and an earlier draft of this line added one anyway, which would have made the executor either invent an out-of-scope closure check or declare success without establishing its own stated predicate. Then, and only then, the scratch @@ -453,7 +459,8 @@ the history is untouched and the delta is loose, after it the findings commit si - [ ] **Validate the base — Resume's own checks, not Preparation's** ```bash -test -s .context/loop-rule-base || { echo "no base recorded — this is a first entry, run Preparation"; exit 1; } +test -e .context/loop-rule-base || { echo "no base file — this is a first entry, run Preparation"; exit 1; } +test -s .context/loop-rule-base || { echo "base file is EMPTY — inspect, delete deliberately, record why, re-record from the true starting commit"; exit 1; } BASE=$(cat .context/loop-rule-base) test "${#BASE}" -eq 40 || { echo "base is not a full 40-character object name"; exit 1; } case "$BASE" in *[!0-9a-f]*) echo "base is not an object name: $BASE"; exit 1 ;; esac @@ -704,13 +711,21 @@ preservation fragments that an earlier draft chose afterwards. **The base file decides it, and nothing else does:** ```bash -if [ -s .context/loop-rule-base ]; then - echo "base recorded — this is a RE-ENTRY: run Resume, not Preparation" +if [ -e .context/loop-rule-base ]; then + echo "a base file exists — this is a RE-ENTRY: run Resume, not Preparation" else - echo "no base — this is a FIRST ENTRY: run Preparation, then record the base below" + echo "no base file — this is a FIRST ENTRY: run Preparation, then record the base below" fi ``` +**The route is decided by the path existing, not by it being non-empty.** An interruption between +the redirection and the write leaves an **empty** file, and an `-s` test calls that a first entry +while Preparation refuses because the path is there — a state the plan could neither initialize nor +resume. **An empty or malformed base file is Resume's**, and Resume fails it on the shape checks; the +transition from there is to **inspect it, delete it deliberately, record why, and re-record from the +true starting commit** — or, where what it should have held cannot be established, to hand it over +rather than guess. + **First entry.** `## The four procedures` · Preparation holds the checks and the shell: branch, clean tree, no base file, `ba15e83` an ancestor, and the three approved inputs compared by blob **at `HEAD`** — which is the revision the tasks are about to be derived against. Then record the base, @@ -1101,10 +1116,12 @@ but the rule is the behaviour and a task that stops recording stops owing it. **A task-number list here was wrong twice** — it named Tasks 3, 4, 6, 7 and 8 while Tasks 5, 9, 10 and 11 also record — and a stale list is the same defect as a stale count. **The cost of the -omission is concrete:** the evidence stays dirty after the task commits, so Task 0's re-entry path, -which requires a clean tree, cannot be used after an interruption; and if execution continues, that -task's evidence is swept into a later unrelated commit rather than the independently reviewable -snapshot this plan promises. +omission is concrete:** the evidence stays out of its own task's `WIP:` snapshot, so the reviewed +range does not hold it where the task claims, and if execution continues it is swept into a later +unrelated commit rather than the independently reviewable snapshot this plan promises. **Not a +clean-tree argument** — Resume requires no clean tree, and three of its four topologies do not have +one; an earlier draft said re-entry needs one, which would have rejected the very states Resume +exists to reconcile. Named `WIP:` because Task 15 runs Gate B over the whole change and closes it with **one `git reset --soft "$BASE"` and a single commit**, per the Global Constraints and Task 15 step 8 — @@ -2709,10 +2726,13 @@ The message carries, in this order: 3. any **human-exception record**, and beside it the skip reason if a cycle was skipped; 4. the **revalidated evidence entry**, replacing step 5's draft if revalidation changed it — and **if it changed, this candidate is over.** The final reviewer judged the entry it was handed - verbatim; a different entry in the closing commit is evidence no pass covered. Commit the change, - resolve and record the new head, and issue another candidate pass. **Revalidate before recording - the candidate head and issuing the pass**, so that in the ordinary case this branch is never - reached. + verbatim; a different entry in the closing commit is evidence no pass covered. **Commit the + change, then route the pass through step 7's ordering like any other non-closing pass** — a + standing source block, an open suspension or a stop answer binds here exactly as it does there, + and **only its continue result records a new reviewed head and issues another candidate.** An + earlier draft sent this branch straight to a new call, which is the plan's own Gate-B loop + stepping over the ordering it installs. **Revalidate before recording the candidate head and + issuing the pass**, so that in the ordinary case this branch is never reached. **Items 1, 2 and 4 are owed unconditionally; item 3 is owed only where such a record exists.** Confirm the file carries the three, and either the applicable exception records or **the literal @@ -2802,12 +2822,19 @@ test -s .context/loop-rule-closing-msg || { echo "closing message missing or emp git commit -F .context/loop-rule-closing-msg ``` -**Then check the result. On a failed `git commit`, or on either check below failing, run +**Then check the result. On a failed `git commit`, or on ANY of the four checks below failing, run `## The four procedures` · Failure** — which stops every further mutating action, reports the failed step and the observed state, and hands over. **It does not reset, re-commit or clean up**, and the -`rm -f` below is therefore unreachable on that path: +`rm -f` below is therefore unreachable on that path. + +**This block is its own shell invocation, so it re-reads `$BASE`.** An earlier draft compared +`HEAD^` against a variable assigned in the previous fence, where it is empty — which made a +**correct** closing commit fail the parent check and enter Failure every time, leaving the plan with +no successful close path at all: ```bash +BASE=$(cat .context/loop-rule-base) +test -n "$BASE" || { echo "BASE empty or unreadable — cannot verify the close"; exit 1; } case "$(git log -1 --pretty=%s)" in [Ww][Ii][Pp]:*) echo "closing commit still reads as a snapshot — stop here and run Failure"; exit 1 ;; esac @@ -2821,7 +2848,7 @@ diff .context/loop-rule-closing-msg .context/loop-rule-landed-msg \ || { echo "the committed body differs from the validated message — stop here and run Failure"; exit 1; } ``` -**Both clean, and only then:** +**All four pass — subject, parent, clean tree, landed body — and only then:** ```bash rm -f .context/loop-rule-base .context/loop-rule-reviewed-tip .context/loop-rule-reviewed-head \ From 8fc3c13412a73832b99dc740d5fa6c2dbe7b9333 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Tue, 15 Sep 2026 15:55:19 +0200 Subject: [PATCH 144/181] docs(context): record passes 26-33 and the stop in the working record --- .../codex-reviews/gate-a-plan-om0bdd7udh-resume.md | 11 ++++++++++- 1 file changed, 10 insertions(+), 1 deletion(-) diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md index a2b8c52..8dd81f0 100644 --- a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md @@ -13,7 +13,7 @@ Nothing depends on it; the pass files and the repo are authoritative where this ## Resume here -**Next action:** awaiting Daniel. The two-pass allowance (24, 25) is spent and the cycle did not close. Pass 25's repairs are applied at `f795bd2` and no pass has read them. The prompt is +**Next action:** awaiting Daniel. The handoff decision is applied and validated; passes 26–33 reviewed it. Pass 33's repairs are applied at `b395da0` and no pass has read them. The prompt is `.context/gate-a-plan-prompt.md`; substitute `__SHA__` and `__P__`, precheck per its header, delete the target file, confirm it is gone. @@ -94,6 +94,15 @@ worth checking before a pass rather than after. | — | — | — | — | — | — | **METHOD CHANGE**, approved by Daniel. Task 0 and Task 15's guarded shell replaced by `## The four procedures` — preparation, close, failure, resume — plus an accounting table classifying all 41 prior guard conditions. **Not a reduction:** one condition deliberately dropped (Task 0 committing nothing), six corrected against pass-21 findings. Commit `1d2f9db` | | 22 | 1d2f9db | **10** | **5** | **3** | yes | **targeted pass**, charged with: did anything get silently dropped, and are pass 21's five closed. **Three accounting rows were false when written** — 8, 15 and 40 claimed "kept, same shell" for obligations not present everywhere. Also `printf %b` storing `\*` in the site anchors; `--mixed` destroying an index-only change; an 8a rejection with no tip to restore to; §A1's failure transition having **three** routes where one was stated; the `kept` disposition having no route for a passage no span reaches | | 23 | 9ced53c | 10→**5** | 5→**3** | 3→**2** | yes | **all five in Task 0 / Task 15 again.** A symbolic value in the base file; the restore target chosen by which tip file exists, which rewinds past a second candidate's repair; the failure capture covering tracked content only while the closure inputs are ignored paths; pass 22's diff-status classification written as prose and not as code; the close procedure's "check all six before moving `HEAD`" applied to two checks that run after a move | +| 26 | 05ec1b2 | **5** | **1** | 2 | yes | first pass on the revised plan. Resume audited progress by `WIP:` commits and called any non-`WIP` commit a stale base — the handoff leaves two other shapes, one of them `HEAD` **at** the base with everything in the index | +| 27 | f69db7d | 5→**6** | 1→**3** | 2→**1** | yes | the close required the commit's **tree to equal the tip's**, which target §I **parks**; `reset --soft` leaves the index untouched, so the prose beside it was wrong about git; Resume still inferred staleness from a commit subject where §A3 names that state as reachable | +| 28 | f2e5dda | 6→**6** | 3→**2** | 1→**2** | yes | pass 27's stale-base fix reached one site of three; its reset paragraph stated the index behaviour correctly and repeated the false claim four lines later; nothing checked the closing commit's **parent**; the message was validated in the file, where `commit-msg` hooks rewrite git's copy after `-F` reads it | +| 29 | 58c49db | 6→**9** | 2→**5** | 2→**2** | yes | **mandatory stop.** Resume declared the post-reset topology valid and then validated the base through Preparation, which requires a clean tree — that state could never pass; its stale-base predicate included "history has no cycle commits", **necessarily true** of that same state | +| 30 | f99607b | 9→**7** | 5→**4** | 2→**3** | yes | five were pass 29's fix applied inconsistently **inside Resume itself**, so Resume was rewritten whole; three absence checks used bare `grep -c`, which **exits 1** on the no-match result they expect | +| 31 | 2b82702 | 7→**5** | 4→**2** | 3→**2** | yes | step 7 committed a non-closing pass and issued the next review **without routing it through the ordering this change installs**; 8a checked the findings files' paths and never their content | +| 32 | 1e1638f | 5→**4** | 2→**2** | 2→**1** | yes | pass 31's ordering bullets sat **after** the block that already recorded the next head; 8b checked the tip but not the clean tree; pass 31's blob guard lived only in the example shell, not among the governing conditions | +| 33 | e429008 | 4→**7** | 2→**3** | 1→**3** | yes | pass 32 split the post-close checks into their own fence where **`$BASE` is empty**, so a *correct* closing commit failed the parent check every time — the plan had no successful close path; an empty base file satisfied neither entry route | +| — | — | — | — | — | — | **STOP.** Eight passes since the handoff decision, no clean close. Reporting to Daniel; the decision authorized a review before implementation, not this many | | 24 | b8b433d | 5→**8** | 3→**2** | 2→**5** | yes | close listed "message complete" *after* the record commit while calling it a precondition; the route said "an empty pair means nothing moved" three paragraphs after the same procedure said the opposite, and the checksums had no pre-act baseline; **after a commit lands and then fails a postcondition both tracked patches are empty** and nothing recorded the rejected `HEAD`; Task 0's `cond` fragments lived only in the ignored map, outside step 4's sweep; step 5 still said 7b "appends"; Task 11 recorded a count §F refuses | | 25 | 892301e | 8→**5** | 2→**2** | 5→**2** | yes | `test -e && cksum` in a loop makes the block's status its *last* iteration, so the pre-act capture returned 1 before 8a and a correct candidate could not enter the close; the captures wrote into the worktree they compared; the pre-act capture held no **bytes**, so a hook rewriting a staged findings file in place was undetectable; **kept `c1`–`c3` lost their observation again** — pass 10 fixed it with a per-task note, pass 22 deleted that note with the other id lists, and no task's walk reaches them; Task 11's sweep record had no shape | | — | — | — | — | — | — | **CHECKPOINT.** Daniel's allowance of two passes (24, 25) is spent. Pass 25's five repairs are **applied and not yet reviewed**. Next action is his call, not another pass | From 18491647fd6535ea21b86d0f12a964b232ed77d9 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Tue, 15 Sep 2026 19:10:23 +0200 Subject: [PATCH 145/181] docs(plans): one home per closing condition; verify the close path end to end MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Bounded revision of the duplicated execution instructions around Task 0, Task 15 and "The four procedures", per Daniel's assignment. THE DUPLICATION. Each closing condition was stated three times: as a Close condition, again as a command in Task 15 step 8, and a third time as prose beside that command. Three statements drift, and the review history since pass 21 is largely that drift — a repair reaching one of the three and missing two. docs/prompt-standards.md item 8 is the rule: reference, do not duplicate. WHAT CHANGED. A new index, "Where each operational condition is defined", names one home per condition and is a lookup, not a procedure — it carries no commands and no expected results. Close keeps the six conditions and is where they are defined. Step 8 keeps the straight-line shell, with each block naming the condition it discharges. The third copy is gone. Task 0 step 1 and accounting row 7 now cite Preparation, Resume and Resume's topology table instead of re-describing them. NOTHING GOVERNING REMOVED. Two rationales that existed only in the deleted prose were moved into the conditions that own them: `reset --soft` leaving the index untouched into condition 5, and why the dirty set can be exact into condition 2. Checked that the is_wip_commit rule, the hook-rewrite rationale, the parked route, "re-establish every closure condition", the record grammar and the evidence duties all still have a home. ONE CORRECTION, FOUND BY RUNNING IT. Condition 6's landed-body check compared bytes, and `git log --pretty=%B` emits one trailing newline the source file does not carry — so the check rejected a CORRECT close and sent every successful cycle into the Failure handoff. The plan had no working close path. The comparison strips trailing blank lines on both sides; verified that it passes a correct close and still catches a rewritten, an added and a dropped line. VERIFIED IN A DISPOSABLE REPOSITORY, no permanent harness added: the whole of step 8 end to end — 8a's five conditions, 8b's reset and commit, condition 6's four postconditions, and the cleanup — leaving two commits and the work in the tree. Also verified that condition 2's dirty-set extraction catches a staged tracked modification, which was its point. The plan's shell needs `sh` or `bash`: `$FINAL` relies on word-splitting, which zsh does not do by default, and the normalized comparison uses process substitution. Observed — the first verification run was under zsh and failed on a path built from the unsplit string. Stated in the plan. --- .../2026-09-14-loop-rule-consolidation.md | 236 ++++++++++-------- 1 file changed, 130 insertions(+), 106 deletions(-) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index ba70d40..de5bdbd 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -196,6 +196,37 @@ agent observes the actual state and derives the operation. **Concrete shell stay simple and already verified**; what goes is the branching that tried to handle every state in one block. +### Where each operational condition is defined — one home each + +**This index exists because the same condition was being stated three times**: as a Close condition, +again as a command in Task 15 step 8, and a third time as prose beside that command. Three statements +drift, and the review history is mostly that drift — a repair reaching one of the three and missing +two. **`docs/prompt-standards.md` item 8 is the rule: reference, do not duplicate.** + +**Read this as a lookup, not as a procedure.** It carries no commands and no expected results; those +live in the home named. Where a task, a shell comment or an error string needs a condition, it +**cites the row** rather than restating it. + +| Operational condition | Defined in | Discharged by | +|---|---|---| +| Branch, clean tree, no base file yet, `ba15e83` ancestral, approved inputs unchanged at `HEAD` | **Preparation** | Task 0 step 1, first-entry branch | +| Base file shape, self-resolving commit, ancestral; approved inputs unchanged at `$BASE`; scratch-artifact validity; how far the implementation got | **Resume** | Task 0 step 1, re-entry branch | +| The four repository topologies a re-entry can meet | **Resume**, its table | accounting row 7; every state-reading rule | +| `HEAD` equals the reviewed head | **Close**, condition 1 | step 8a | +| The dirty set is exactly the candidate pass's findings files | **Close**, condition 2 | step 8a | +| The closing message is complete, and revalidated before it is consumed | **Close**, condition 3 | step 7b writes it; steps 8a and 8b check it | +| The record commit's paths, parent, blob identity, and the findings files' structure and eligibility | **Close**, condition 4 | step 8a | +| Clean tree and the closing tip when the closing invocation begins | **Close**, condition 5 | step 8b | +| After the closing commit: subject, parent, clean tree, landed body | **Close**, condition 6 | step 8b's post-close block | +| Two invocations, and why the hook requires them | **Close** | steps 8a and 8b | +| What happens on a failed closing operation or postcondition | **Failure** | every rejection in steps 8a and 8b | +| What a successful close removes | **Close**, its success recognition | step 8b's cleanup block | + +**Straight-line commands stay where the work happens.** Task 15 step 8 keeps the shell that runs +these checks — it is verified and it is what an executor types — and each block names the condition +it discharges instead of re-explaining it. **What was removed is the third copy**: the prose that +repeated a condition in words beside the command that already ran it. + ### The accounting — every condition, kept / re-expressed / proposed for deletion Required before replacing a decision procedure (`AGENTS.md`, "Never replace a decision procedure @@ -211,7 +242,7 @@ independent reader did. | 4 | The three approved inputs' blobs equal their `ba15e83` versions | **kept, and split by entry**: Preparation compares them **at `HEAD`**, because a first entry has no recorded base; **Resume compares them at `$BASE`**, because a Gate-B fix may legitimately have changed `HEAD`'s copy in a `WIP:` snapshot. An earlier table row claimed Preparation did the base comparison, which it never could | | 5 | Never overwrite an existing base file | **kept** as an obligation; the shell branch becomes one line of the resume procedure | | 6 | A pre-existing base is an ancestor of `HEAD` (`merge-base --is-ancestor`) | **kept**, same shell — resume | -| 7 | Only this run's `WIP:` commits lie between base and `HEAD` | **kept as an observation, dropped as a staleness test.** Resume's four topologies are the model: *normal* (this run's `WIP:` commits, §A3's stray commit or amend nested inside it); *8a rejected before its commit landed* (history untouched, the delta loose in index or worktree); *8a rejected after it landed* (a findings commit above the chain, plus any delta); and *8b rejected* (`HEAD` at the base with the work staged, **or** one commit parented by the base where the closing commit landed and a postcondition refused it). **None makes the base stale**, and staleness is decided by ancestry and provenance, never by a commit subject | +| 7 | Only this run's `WIP:` commits lie between base and `HEAD` | **kept as an observation, dropped as a staleness test.** The model is **Resume's topology table**, which is where the four shapes are defined — re-listing them here is the third copy this revision removes. **None of them makes the base stale**, and staleness is decided by ancestry and provenance, never by a commit subject | | 8 | `$BASE` persisted to a file; every consumer guards it non-empty | **kept**, and **repaired**: three concrete blocks read it without the guard while this row claimed otherwise — the baseline extraction, Task 10's hook diff, and the first Gate-B call. The guard is in each of them now (pass 22) | | 9 | The five regions are located by anchor, one hit per pattern per file | **kept**, same shell — preparation | | 10 | Spans are **derived** from replacement extents, never hand-written | **kept** — the rule, unchanged | @@ -247,6 +278,16 @@ independent reader did. | 40 | Every rejection restores a tip by **mixed** reset, never `--hard` | **DELIBERATELY DROPPED**, by Daniel's decision, and the second of two deletions this table records. **It was never a requirement of the approved spec** — §A says *"re-establish every closure condition against the repository as it now stands"*, which reads the state rather than rewinding it, and a rejected commit at `HEAD` is a legitimate starting point for that reading. It was this plan's own implementation choice, and it had grown restore targets, phase selection, pre- and post-act captures and their own error handling, which four consecutive passes then found defects in. **What replaces it is a bounded handoff**: stop every further mutating action, report the failed step and the observed state, change nothing else. **The obligations §A does impose are unchanged** — surface the failure, re-establish every closure condition against the state as it stands, and take one of its three routes. **Not replaced by a generic backup mechanism**, which would be the same growth under another name | | 41 | The scratch files are removed only after a successful close | **kept** as the close procedure's last step | +**This revision removed statements, not conditions — with one correction.** Every row below still +has exactly one home, named in the index above; what went is the **third copy**, the prose beside +Task 15 step 8's shell that repeated in words what the block already ran and the condition already +defined. Two rationales that existed only in that prose were moved into the conditions they belong +to: `reset --soft`'s index behaviour into condition 5, and why the dirty set can be exact into +condition 2. **The one correction is condition 6's landed-body check**, which compared bytes where +`git log --pretty=%B` adds a trailing newline the source file has none of — it rejected a *correct* +close, and the comparison now strips trailing blank lines on both sides. Verified in a disposable +repository, in both directions. + **Two conditions are dropped, 20 and 40, and each is named as a drop rather than lost.** Condition 20 is replaced by a stronger obligation. **Condition 40 is dropped outright** — a guard this plan introduced itself, which the approved spec never asked for, removed by an explicit decision after @@ -320,8 +361,11 @@ an unapproved edit to a spec that the gates already closed. value as the **closing tip** below, and conflating them is how an unreviewed commit reaches the close. 2. **The only thing dirty is the candidate pass's own findings files** — read from every porcelain - record, not from two status codes. Nothing else may be uncommitted: every record this plan - collects was committed before the pass was issued. + record, not from two status codes, because a staged tracked modification is invisible to a filter + that reads only `??` and `A `. **The set can be exact because step 7 commits every record this + plan collects before the candidate pass is issued**; anything else uncommitted here arrived after + the review and no pass has seen it. A full Gate-B pass writes **two** files, so the pair is what + is expected. 3. **The closing message is complete** — rebuilt whole, exactly one provenance line, exactly one curve, every owed evidence entry, and either the applicable human-exception records or `Human exceptions: none`. **This is checked here, before anything moves**, because an incomplete @@ -348,7 +392,14 @@ an unapproved edit to a spec that the gates already closed. `.context/loop-rule-reviewed-tip` — **a precondition value, not a restore target**: 8b refuses to reset unless `HEAD` is still exactly it, which is how a commit landing between the two invocations is caught. Nothing in this plan resets *to* it. -5. **The tree is clean**, and `HEAD` is still the closing tip, when the closing invocation begins. +5. **The tree is clean**, and `HEAD` is still the closing tip, when the closing invocation begins — + **both read in that invocation**, not carried from the previous one. **`git reset --soft` moves + `HEAD` and leaves the index exactly as it was**: it stages nothing and unstages nothing, so the + index still holds every `WIP:` commit's content, which is what makes one closing commit carry the + whole change. **Anything staged and uncommitted at that moment is in the index too and lands in + the closing commit; only unstaged work stays out.** This condition is what closes that asymmetry — + without it a staged edit made between the two invocations, a rewritten findings file included, + ships unreviewed and leaves a clean tree behind it. 6. **After the closing commit**, four things: its **subject is not a snapshot**; its **parent is `$BASE`**; the **tree is clean**; and its **body is identical to the revalidated closing-message file**. The last is not decoration — `prepare-commit-msg` and `commit-msg` hooks rewrite git's @@ -719,17 +770,12 @@ fi ``` **The route is decided by the path existing, not by it being non-empty.** An interruption between -the redirection and the write leaves an **empty** file, and an `-s` test calls that a first entry -while Preparation refuses because the path is there — a state the plan could neither initialize nor -resume. **An empty or malformed base file is Resume's**, and Resume fails it on the shape checks; the -transition from there is to **inspect it, delete it deliberately, record why, and re-record from the -true starting commit** — or, where what it should have held cannot be established, to hand it over -rather than guess. - -**First entry.** `## The four procedures` · Preparation holds the checks and the shell: branch, -clean tree, no base file, `ba15e83` an ancestor, and the three approved inputs compared by blob **at -`HEAD`** — which is the revision the tasks are about to be derived against. Then record the base, -**and never overwrite one you did not just write**: +the redirection and the write leaves an **empty** file, and an `-s` test would call that a first +entry while Preparation refuses because the path is there — a state neither route owns. **An empty +or malformed base file is Resume's**, which fails it on its shape checks and states what follows. + +**First entry — run `## The four procedures` · Preparation**, then record the base here. **Never +overwrite one you did not just write:** ```bash git rev-parse HEAD > .context/loop-rule-base @@ -742,16 +788,13 @@ test "${#BASE}" -eq 40 || { echo "recorded base is not a full 40-character objec test "$(git rev-parse --verify "$BASE^{commit}")" = "$BASE" || { echo "recorded base does not resolve to itself as a commit"; exit 1; } ``` -**Re-entry.** `## The four procedures` · Resume validates the existing base, the scratch artifacts -and how far the implementation got — with its own checks, against all four topologies, and without -requiring a clean tree. **Do not run Preparation on a re-entry**: it refuses as soon as it sees the -base file, which is exactly what makes this branch reachable. +**Re-entry — run `## The four procedures` · Resume.** It owns the existing base, the scratch +artifacts and how far the implementation got. -**Re-running Task 0 after a partial implementation must not re-record the base.** It would capture +**Why the branch exists at all:** re-recording the base after a partial implementation would capture the current WIP tip, and both Gate B's range and the final reset would then start *after* every edit -made so far — prompt and hook changes squashed into the closing commit without entering a review -range. The branch above is what prevents it: a recorded base sends this task to Resume, which -validates and never overwrites. +made so far — prompt and hook changes squashed into the closing commit without ever entering a +review range. **Persist it to a file, not to a shell variable.** Each fenced block runs in its own shell invocation, so a `BASE=` assignment here is gone by the next task and every parent-tree count would @@ -2751,104 +2794,106 @@ provenance line or curve has not validly closed the cycle. - [ ] **Step 8: Close the cycle — the close procedure, in two invocations** -**Run `## The four procedures` · Close.** It states the six checks, what success looks like, and why -the two invocations are separate. What follows is the shell that is simple and already verified; -**the checks are the obligation and the shell is one way to run them** — where the observed state is -not one this shell expects, read the state and pick the operation, rather than extending the block. +**`## The four procedures` · Close states the six conditions, what success looks like, and why the +two invocations are separate.** Nothing here restates them. What follows is the shell that discharges +them, each block naming its condition; **the conditions are the obligation and this shell is one way +to run them** — where the observed state is not one it expects, read the state and pick the +operation, rather than extending the block. -**8a — record the candidate pass's findings files.** +**8a — record the candidate pass's findings files. Discharges conditions 1, 2 and 4.** ```bash BASE=$(cat .context/loop-rule-base) HEADREV=$(cat .context/loop-rule-reviewed-head) test -n "$BASE" && test -n "$HEADREV" || { echo "BASE or reviewed head missing"; exit 1; } + +# Condition 1. test "$(git rev-parse HEAD)" = "$HEADREV" || { echo "HEAD is not the head the candidate pass was issued against"; exit 1; } -# Exactly this pass's two slot paths, from the values the call was built from. +# Condition 2. NONCE and P are this cycle's nonce and the candidate pass number, +# the same two the call's slot paths were built from. NONCE=; P= FINAL=".context/codex-reviews/gate-b-spec-$NONCE-pass-$P.md .context/codex-reviews/gate-b-quality-$NONCE-pass-$P.md" expected=$(printf '%s\n' $FINAL | sort) actual=$(git status --porcelain -z | tr '\0' '\n' | sed -n 's/^.\{3\}//p' | sed '/^$/d' | sort) test "$expected" = "$actual" || { echo "dirty set is not exactly this pass's findings files:"; git status --porcelain; exit 1; } -# Pin the validated content BEFORE staging: a hook can rewrite a staged file in -# place, leaving its pathname — and therefore the changed-path check — unchanged. +# Condition 4, first bullet: pin the validated blobs before staging. for f in $FINAL; do git hash-object "$f"; done > .context/loop-rule-final-blobs # shellcheck disable=SC2086 git add $FINAL -# Any rejection below stops and goes through the Failure procedure, which reports -# the state and hands over. Nothing here resets, retries or deletes. -git commit -m "WIP: Gate-B findings files" || { echo "record commit FAILED — stop here and run Failure"; exit 1; } +git commit -m "WIP: Gate-B findings files" || { echo "record commit FAILED — run Failure"; exit 1; } test "$(git show --name-only --pretty=format: HEAD | sed '/^$/d' | sort)" = "$expected" \ - || { echo "record commit changed paths beyond this pass's findings files — stop here and run Failure"; exit 1; } -test "$(git rev-parse HEAD^)" = "$HEADREV" || { echo "record commit's parent is not the reviewed head — stop here and run Failure"; exit 1; } -# The committed blobs, not the paths: same names can hold different bytes. + || { echo "record commit changed paths beyond this pass's findings files — run Failure"; exit 1; } +test "$(git rev-parse HEAD^)" = "$HEADREV" || { echo "record commit's parent is not the reviewed head — run Failure"; exit 1; } + +# Condition 4, second bullet. for f in $FINAL; do git rev-parse "HEAD:$f"; done > .context/loop-rule-committed-blobs diff .context/loop-rule-final-blobs .context/loop-rule-committed-blobs \ - || { echo "a findings file was rewritten between validation and commit — stop here and run Failure"; exit 1; } -# Then re-run the findings-file structural check on the committed content and -# re-establish this pass's eligibility, per close condition 4's third bullet. -# Expected: every line before the terminator is a finding line, the terminator -# is exact, the count matches, both branch files are present, and the pass reads -# clean-or-zero-finding exactly as it did when issued. Any difference: Failure. -test -z "$(git status --porcelain)" || { echo "tree not clean after the record commit — stop here and run Failure"; exit 1; } - -git rev-parse HEAD > .context/loop-rule-reviewed-tip # the closing tip: 8b's precondition, NOT a reset target -``` + || { echo "a findings file was rewritten between validation and commit — run Failure"; exit 1; } + +# Condition 4, third bullet: re-run the findings-file structural check on the +# COMMITTED content and re-establish this pass's eligibility. Reader check; the +# expected result is condition 4's. +test -z "$(git status --porcelain)" || { echo "tree not clean after the record commit — run Failure"; exit 1; } -**The reviewed head and the closing tip are two values.** The first is what the pass read; the -second is that plus the findings files. **A single value cannot be both**, and treating it as one is -how a commit that landed after the response reaches the close — which is why `HEAD^` is compared -above rather than assumed. +# Condition 4's tail: the closing tip. +git rev-parse HEAD > .context/loop-rule-reviewed-tip +``` -**8b — reset and close. A separate invocation, carrying no `-m` option at all.** +**8b — reset and close. A separate invocation, carrying no `-m` option at all. Discharges conditions +3, 5 and 6.** ```bash BASE=$(cat .context/loop-rule-base); TIP=$(cat .context/loop-rule-reviewed-tip) test -n "$BASE" && test -n "$TIP" || { echo "BASE or TIP missing — 8a did not complete"; exit 1; } + +# Condition 5: the tip AND a clean tree, both read in THIS invocation. test "$(git rev-parse HEAD)" = "$TIP" || { echo "HEAD has moved since 8a — NOT resetting"; exit 1; } -# Close condition 5 is BOTH: the tip AND a clean tree, checked in THIS invocation. -# A staged edit made between 8a and 8b — a rewritten findings file included — -# survives reset --soft, lands in the closing commit, and leaves the tree clean -# afterwards, so every postcondition passes while unreviewed content ships. -test -z "$(git status --porcelain)" || { echo "tree not clean at 8b — NOT resetting; stop here and run Failure"; exit 1; } -git reset --soft "$BASE" || { echo "reset --soft FAILED — stop here and run Failure; do NOT commit"; exit 1; } -# Re-read the closing message here: it was validated before 8a, 8a ran a commit -# (hooks can rewrite anything), and .context is ignored, so no porcelain check -# between the two invocations can see it change. -test -s .context/loop-rule-closing-msg || { echo "closing message missing or empty — stop here and run Failure"; exit 1; } -# Revalidate its records — exactly one provenance line, exactly one curve for this -# cycle, every owed evidence entry, and the exception records or the plural marker. +test -z "$(git status --porcelain)" || { echo "tree not clean at 8b — NOT resetting; run Failure"; exit 1; } + +git reset --soft "$BASE" || { echo "reset --soft FAILED — run Failure; do NOT commit"; exit 1; } + +# Condition 3's revalidation: re-read the message here, immediately before it is +# consumed. Reader check on its records; the expected result is condition 3's. +test -s .context/loop-rule-closing-msg || { echo "closing message missing or empty — run Failure"; exit 1; } git commit -F .context/loop-rule-closing-msg ``` -**Then check the result. On a failed `git commit`, or on ANY of the four checks below failing, run -`## The four procedures` · Failure** — which stops every further mutating action, reports the failed -step and the observed state, and hands over. **It does not reset, re-commit or clean up**, and the -`rm -f` below is therefore unreachable on that path. - -**This block is its own shell invocation, so it re-reads `$BASE`.** An earlier draft compared -`HEAD^` against a variable assigned in the previous fence, where it is empty — which made a -**correct** closing commit fail the parent check and enter Failure every time, leaving the plan with -no successful close path at all: +**Condition 6, in its own invocation — which is why it re-reads `$BASE`.** Every fenced block here is +a separate shell, so a variable assigned in 8b is empty in this one: ```bash BASE=$(cat .context/loop-rule-base) test -n "$BASE" || { echo "BASE empty or unreadable — cannot verify the close"; exit 1; } case "$(git log -1 --pretty=%s)" in - [Ww][Ii][Pp]:*) echo "closing commit still reads as a snapshot — stop here and run Failure"; exit 1 ;; + [Ww][Ii][Pp]:*) echo "closing commit still reads as a snapshot — run Failure"; exit 1 ;; esac -test "$(git rev-parse HEAD^)" = "$BASE" || { echo "closing commit's parent is not \$BASE — stop here and run Failure"; exit 1; } -test -z "$(git status --porcelain)" || { echo "tree dirty after the close — stop here and run Failure"; exit 1; } -# The COMMIT BODY, not the source file: prepare-commit-msg and commit-msg hooks -# rewrite git's copy after -F has read it, so the validated file proves nothing -# about what landed. +test "$(git rev-parse HEAD^)" = "$BASE" || { echo "closing commit's parent is not \$BASE — run Failure"; exit 1; } +test -z "$(git status --porcelain)" || { echo "tree dirty after the close — run Failure"; exit 1; } git log -1 --pretty=%B > .context/loop-rule-landed-msg -diff .context/loop-rule-closing-msg .context/loop-rule-landed-msg \ - || { echo "the committed body differs from the validated message — stop here and run Failure"; exit 1; } +# `%B` emits the body plus one trailing newline the source file does not carry, +# so a byte-for-byte diff rejects a CORRECT close. Compare with trailing blank +# lines stripped from both sides; a rewritten, added or dropped line still +# differs. Verified both ways in a disposable repository. +strip_trailing_blanks() { + awk '{ l[NR]=$0 } END { n=NR; while (n>0 && l[n]=="") n--; for(i=1;i<=n;i++) print l[i] }' "$1" +} +diff <(strip_trailing_blanks .context/loop-rule-closing-msg) \ + <(strip_trailing_blanks .context/loop-rule-landed-msg) \ + || { echo "the committed body differs from the validated message — run Failure"; exit 1; } ``` -**All four pass — subject, parent, clean tree, landed body — and only then:** +**Run step 8's blocks under `sh` or `bash`, not `zsh`.** `for f in $FINAL` and the `git add $FINAL` +beside it rely on the unquoted variable **word-splitting into two paths** — which is why the +`shellcheck disable=SC2086` is there — and `zsh` does not split unquoted parameters by default, so +the whole string is taken as one filename and every command fails on a path that does not exist. +**Observed, not assumed**: the first verification run of this block was made under `zsh` and failed +exactly that way. The process substitution above also needs `bash`; `sh` on a system where it is +`dash` has none, so use `bash` for that block or write the two normalized copies to temporary +files first. + +**All four pass, and only then the cleanup Close's success recognition names:** ```bash rm -f .context/loop-rule-base .context/loop-rule-reviewed-tip .context/loop-rule-reviewed-head \ @@ -2863,30 +2908,9 @@ if ls .context/loop-rule-* >/dev/null 2>&1; then fi ``` -**Every `loop-rule-*` scratch file goes, and the `ls` is what makes "the scratch files are removed" -true rather than asserted.** An earlier draft deleted three of them and claimed the terminal state, -leaving a later run to inherit a closing message, a baseref and an untouched map — each of which -some check then has to detect or overwrite piecemeal. **This runs only on a successful close**; after a -failure **the plan performs no cleanup at all**. That is not a promise the files are unchanged — a -hook that failed may have rewritten any of them — so Resume enumerates what survives and validates -it rather than trusting it. **The plan's records -are not among these**: they live in the plan and in `.context/codex-reviews/`, both tracked, both -already in the closing commit. - -**`reset --soft` moves `HEAD` and leaves the index exactly as it was** — it stages nothing and -unstages nothing, so the index still holds every `WIP:` commit's content, which is what makes the -single closing commit carry the whole change. **Anything *staged* and uncommitted at that moment is -in the index too and would land in the closing commit**; only *unstaged* work stays out. That is why -the close's precondition 5 requires a clean tree before 8b runs — **an earlier draft said the reset -"stages committed content only" and that unstaged-or-not, uncommitted content stays out, which is -wrong about the index and hides the path §I parks.** The prompt-standards result, the completeness -sweep, the next-state table, the divergence list, the equivalence result and the fragment evidence -all land in this plan, and `.context/codex-reviews/` is tracked. **Anything uncommitted and -*unstaged* when the reset runs stays in the worktree and out of the closing commit; anything -uncommitted and *staged* lands in it.** That asymmetry is what precondition 5's clean tree -prevents; and the plan records, had they been left uncommitted, -would never have been in a Gate-B range either. Step 7 commits them before the candidate pass is issued, which -is what lets 8a's dirty-set check be exact. +**The `ls` is what makes the removal checked rather than asserted**, and the list is every +`loop-rule-*` name this plan writes. **The plan's own records are not among them** — they live in +this plan and in `.context/codex-reviews/`, both tracked, both already inside the closing commit. --- From c7f4617a9fbde27ccb0391f0c5e43275de6b25f5 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Tue, 15 Sep 2026 19:23:14 +0200 Subject: [PATCH 146/181] docs(context): record the bounded revision and pass 34 in the working record --- .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-34.md | 7 +++++++ .context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md | 4 +++- 2 files changed, 10 insertions(+), 1 deletion(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-34.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-34.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-34.md new file mode 100644 index 0000000..2ea25f4 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-34.md @@ -0,0 +1,7 @@ +BLOCKER | high | Operational-condition index row for condition 3; Close condition 3; Task 15 steps 7b and 8a | The index says both 8a and 8b check the closing message, but 8a explicitly discharges only conditions 1, 2 and 4 and contains no condition-3 check. Step 7b's earlier check does not cover a re-entry between 7b and 8a. | On that supported re-entry path, 8a can make the findings record commit before the message-completeness precondition is re-established, leaving HEAD moved when 8b first discovers the defect and recreating the exact failure Close says condition 3 prevents. | Add a condition-3 reader revalidation at the start of 8a before any commit, retain the 8b revalidation against possible hook rewrites, and make the index and discharge labels name those two checks accurately. +MAJOR | high | Close condition 2 and Task 15 step 8a lines 2817-2819 | The governing condition requires every porcelain record, but the shell converts NUL-delimited status output to newline-delimited text and strips three characters from every resulting line; Git pathnames may contain newlines, and a rename's second NUL-delimited pathname has no status prefix. | A newline-bearing or rename pathname can be split, discarded, or mangled so the exact-dirty-set precondition passes or reports the wrong set; the record commit may mutate HEAD before later postconditions force a handoff. | Preserve NUL record framing and parse ordinary versus rename records explicitly, or replace this parser with checks whose path comparison remains exact for arbitrary Git pathnames; verify a newline pathname and a rename as rejection cases. +MINOR | high | Close condition 6 and Task 15 step 8 post-close block lines 2875-2884 | Condition 6 and its success text require an identical landed body, while the implemented oracle strips every trailing blank line from both inputs; the adjacent claim that any added or dropped line still differs is therefore false for trailing blank lines. | The sample can declare all four postconditions satisfied when the stated governing predicate is false, while an executor deriving a literal byte comparison from the condition can reject the same close the sample accepts. | Define the governing condition and success criterion as equality after trailing-blank normalization and narrow the rationale to nonblank record lines, or use a commit-message extraction and source normalization that establish the stated byte predicate. +MINOR | high | Task 15 step 8 heading and operational-condition index rows for condition 6 and cleanup | Step 8 is titled and justified as a two-invocation procedure, but it supplies four separate fenced shell invocations: 8a, 8b, condition 6's explicitly separate post-close check, and cleanup; the index nevertheless calls the latter two step-8b blocks. | The invocation boundary is part of the safety design, so contradictory counting and ownership can make an executor merge blocks, rely on variables that do not cross fences, or misidentify which invocation discharged a condition. | Describe the procedure as two commit invocations followed by separate post-close verification and cleanup invocations, and update the index to name those blocks directly, or intentionally merge the latter blocks into 8b and remove the cross-fence claims. +MINOR | high | Architecture paragraph, line 7 | The overview says each task runs a new-present and old-gone discriminating pair from one verified fragment table, but the plan's governing procedure requires add-only presence checks, carried and kept preservation checks, moved source and destination checks, dropped absences, spans and sweeps, and several rows are derived and appended during execution. | The plan's architectural contract contradicts the verification system that follows, inviting an implementer or reviewer to treat a pair as universal or assume every fragment was preverified. | Rewrite the overview to say each condition receives the observation its disposition requires, with pre-existing fragments kept in the table and post-install fragments and results recorded in the evidence section. +NIT | high | Task 15 step 8 cleanup lines 2898-2912 | The cleanup claims its explicit removal list contains every loop-rule scratch name the plan writes, but the plan also writes `.context/loop-rule-untouched.tmp` and `.context/loop-rule-baseline-diff.tmp`, neither of which is removed. | A stale temporary from an interrupted or repeated artifact build makes the wildcard verification fail only after the recovery base and other scratch values have already been deleted. | Add both temporary names to the removal list, or make Resume reject and reconcile them before closure, and narrow the completeness claim if transient names are intentionally excluded. +END OF FINDINGS (6 total) diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md index 8dd81f0..2db67c6 100644 --- a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md @@ -13,7 +13,7 @@ Nothing depends on it; the pass files and the repo are authoritative where this ## Resume here -**Next action:** awaiting Daniel. The handoff decision is applied and validated; passes 26–33 reviewed it. Pass 33's repairs are applied at `b395da0` and no pass has read them. The prompt is +**Next action:** awaiting Daniel. The bounded revision is at `1849164`; pass 34 reviewed it and is **unclean** — 1 Blocker, 1 Major, 3 Minors, 1 Nit, all validated, none repaired. The assignment's checkpoint forbids another repair/review round without his word. The prompt is `.context/gate-a-plan-prompt.md`; substitute `__SHA__` and `__P__`, precheck per its header, delete the target file, confirm it is gone. @@ -94,6 +94,8 @@ worth checking before a pass rather than after. | — | — | — | — | — | — | **METHOD CHANGE**, approved by Daniel. Task 0 and Task 15's guarded shell replaced by `## The four procedures` — preparation, close, failure, resume — plus an accounting table classifying all 41 prior guard conditions. **Not a reduction:** one condition deliberately dropped (Task 0 committing nothing), six corrected against pass-21 findings. Commit `1d2f9db` | | 22 | 1d2f9db | **10** | **5** | **3** | yes | **targeted pass**, charged with: did anything get silently dropped, and are pass 21's five closed. **Three accounting rows were false when written** — 8, 15 and 40 claimed "kept, same shell" for obligations not present everywhere. Also `printf %b` storing `\*` in the site anchors; `--mixed` destroying an index-only change; an 8a rejection with no tip to restore to; §A1's failure transition having **three** routes where one was stated; the `kept` disposition having no route for a passage no span reaches | | 23 | 9ced53c | 10→**5** | 5→**3** | 3→**2** | yes | **all five in Task 0 / Task 15 again.** A symbolic value in the base file; the restore target chosen by which tip file exists, which rewinds past a second candidate's repair; the failure capture covering tracked content only while the closure inputs are ignored paths; pass 22's diff-status classification written as prose and not as code; the close procedure's "check all six before moving `HEAD`" applied to two checks that run after a move | +| — | — | — | — | — | — | **BOUNDED REVISION**, Daniel's assignment after pass 33. The third statement of every closing condition removed; `### Where each operational condition is defined` gives one home each; step 8's shell kept and annotated with the condition it discharges. Condition 6's landed-body check **corrected** — it compared bytes where `%B` adds a trailing newline, so it rejected a *correct* close. Whole close path verified end to end in a disposable repo. Commit `1849164` | +| 34 | 1849164 | **6** | **1** | **1** | yes | **UNCLEAN — reported to Daniel, no repair round opened.** 8a discharges conditions 1, 2, 4 and never checks condition 3, while the index claims it does; the porcelain parser mangles a rename (`p.txt`) and a newline pathname (two fragments), verified by execution; condition 6 says "identical" where the corrected oracle normalizes trailing blanks; step 8 is titled "two invocations" and has four fenced blocks; the Architecture paragraph still says every task runs a discriminating pair; the two `.tmp` names are not in the cleanup list | | 26 | 05ec1b2 | **5** | **1** | 2 | yes | first pass on the revised plan. Resume audited progress by `WIP:` commits and called any non-`WIP` commit a stale base — the handoff leaves two other shapes, one of them `HEAD` **at** the base with everything in the index | | 27 | f69db7d | 5→**6** | 1→**3** | 2→**1** | yes | the close required the commit's **tree to equal the tip's**, which target §I **parks**; `reset --soft` leaves the index untouched, so the prose beside it was wrong about git; Resume still inferred staleness from a commit subject where §A3 names that state as reachable | | 28 | f2e5dda | 6→**6** | 3→**2** | 1→**2** | yes | pass 27's stale-base fix reached one site of three; its reset paragraph stated the index behaviour correctly and repeated the false claim four lines later; nothing checked the closing commit's **parent**; the message was validated in the file, where `commit-msg` hooks rewrite git's copy after `-F` reads it | From 4402260af6f48ad9949cd3c41935fda93795c160 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Tue, 15 Sep 2026 19:59:30 +0200 Subject: [PATCH 147/181] docs(context): make the Gate-A plan working record a standalone handoff MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The cycle is open and unclean at pass 34, and the next move is Daniel's, so the "Resume here" block is rewritten to carry a cold reader without this session: the state, the four open findings in cost order, what the bounded revision changed and why it must not be undone by accident, how to run a pass, and the rules every pass checks. Two things it now says that the old block did not. The checkpoint from Daniel's 2026-09-15 assignment supersedes the standing autonomy and is SPENT — it covered one revision and one pass, both delivered — so a resuming agent must not infer a repair round. And the plan's shell needs sh or bash, never zsh: $FINAL relies on word-splitting and the body comparison uses process substitution. --- .../gate-a-plan-om0bdd7udh-resume.md | 98 ++++++++++++++----- 1 file changed, 72 insertions(+), 26 deletions(-) diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md index 2db67c6..38ab8d9 100644 --- a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md @@ -13,40 +13,86 @@ Nothing depends on it; the pass files and the repo are authoritative where this ## Resume here -**Next action:** awaiting Daniel. The bounded revision is at `1849164`; pass 34 reviewed it and is **unclean** — 1 Blocker, 1 Major, 3 Minors, 1 Nit, all validated, none repaired. The assignment's checkpoint forbids another repair/review round without his word. The prompt is -`.context/gate-a-plan-prompt.md`; substitute `__SHA__` and `__P__`, -precheck per its header, delete the target file, confirm it is gone. - -**The rules passes 6–10 installed, every one of which had been contradicted in two or three places -at once when found.** Check them every pass: - -- `## What each disposition owes, stated once` fixes the observation per class; `## How a task - discharges that table` is the per-task procedure. **No task enumerates its condition ids.** -- Every **pre-existing** fragment is derived **before** its task's install step — OLD halves, - absence fragments, and carried/kept preservation fragments alike. Only post-install fragments - go to `## Fragment evidence (per-task output)`. -- **No task pre-assigns a fragment id.** Tasks 3, 4, 6, 7 and 10 append. -- Untouched **spans are derived, not written out**. `.context/loop-rule-untouched` carries `span` - and `cond` records, both parseable, both consumed by Tasks 2 and 14. -- **Destination blocks, not target sections**, are the unit for the parity site list. -- Task 15 re-runs every affected check — mechanical *and* reader — after each Gate-B fix, and the - complete set before the candidate final pass; only a clean response against that exact `HEAD` - closes, and the final pass's own findings file is the sole permitted post-review addition. +**State: `c7f4617`, tree clean, cycle OPEN and UNCLEAN.** The plan is reviewed at `1849164`; pass 34 +found 1 Blocker, 1 Major, 3 Minors and 1 Nit, **all validated, none repaired**. `.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-34.md` holds them. + +**Do not start a repair round on your own.** Daniel's assignment of 2026-09-15 ended with a +checkpoint that supersedes the standing autonomy: *"stop and report, whether clean or unclean. Do +not begin another repair/review round or implementation."* That checkpoint is spent — it covered one +revision and one pass, both done — so **the next move is Daniel's word, not an inference.** + +### The four open findings, in the order they cost most + +1. **BLOCKER — step 8a never checks Close condition 3.** It discharges 1, 2 and 4; the index claims + 8a and 8b both check the message. On a re-entry between 7b and 8a the record commit lands before + the message is re-established, and 8b is the first to notice — the exact failure condition 3 + exists to prevent. +2. **MAJOR — the porcelain parser mangles two real pathname shapes.** `git status --porcelain -z | + tr '\0' '\n' | sed 's/^.\{3\}//'` turns a rename into `p.txt` and a newline-bearing path into two + fragments. **Verified by execution, not by reading.** +3. **MINOR — condition 6 says the landed body is "identical"** where the corrected oracle strips + trailing blank lines. Introduced by the bounded revision itself. +4. **MINOR — step 8 is titled "two invocations" and has four fenced blocks**; **MINOR** — the + Architecture paragraph still says every task runs a discriminating pair, which the disposition + table replaced; **NIT** — `loop-rule-untouched.tmp` and `loop-rule-baseline-diff.tmp` are written + and never removed by the cleanup. + +### What the bounded revision did, so it is not undone by accident + +Each closing condition had been stated **three times** — a Close condition, a command in Task 15 +step 8, and prose beside that command. The third copy is gone. **`### Where each operational +condition is defined` is the index**: one home per condition, a lookup with no commands and no +expected results. Close defines the six closing conditions; step 8 keeps the shell with each block +naming the condition it discharges; Task 0 step 1 and accounting row 7 cite rather than re-describe. + +**One correction, found by running it:** condition 6 compared bytes, and `git log --pretty=%B` adds +one trailing newline the source file has none of — **it rejected a correct close, so the plan had no +working success path.** Fixed and checked in four directions. + +**Verified in a disposable repo:** the whole of step 8 end to end. **Unverified:** every reader check, +Failure and Resume — they need a real failure or a real interruption. + +**The plan's shell needs `sh` or `bash`, never `zsh`** — `$FINAL` relies on word-splitting and the +body comparison uses process substitution. + +### How to run a pass, if Daniel asks for one + +Prompt: `.context/gate-a-plan-prompt.md`. Substitute `__SHA__` and `__P__`, precheck per its header, +delete the target file and confirm it is gone. Codex reads the prompt from a file — write the +substituted text to a scratch path and tell it to read that path in full and follow it exactly. **That prompt file is untracked.** `.gitignore` carries `.context/*` with only `codex-gate.on` and `codex-reviews/` exempt, so it survives a context clear but not a `.context/` cleanup. **If it is -gone, rebuild it from this record** — the pass history below, the settled-and-not-open blocks and the -collected list are what it carries, and `.context/gate-a-spec-prompt.md` is the same shape for the -spec cycle. +gone, rebuild it from this record** — the pass history below and the settled-and-not-open blocks are +what it carries; `.context/gate-a-spec-prompt.md` is the same shape for the spec cycle. + +### Rules earlier passes installed — check each, every pass + +- `## What each disposition owes, stated once` fixes the observation per class; `## How a task + discharges that table` is the per-task procedure. **No task enumerates its condition ids**, and + the unit is the **source block a task replaces**, not the passage. +- Six disposition words and no others. Two non-dispositions were found this way (`changed`, `split`). +- A condition is checked against **the inventory's own definition of it** — `a13` is two sentences, + `c9` is one clause. +- The fragment test's third condition is **class-specific**: a disappearing fragment must be absent + from the region's **post-edit text**, a carried one **present** in the block. +- Every **pre-existing** fragment is derived **before** its task's install step. **No task + pre-assigns an id.** Every value crossing a fenced block goes to a file. +- Untouched **spans are derived, not written out**; `.context/loop-rule-untouched` carries `base`, + `span` and `cond` records, consumed by Tasks 2 and 14. +- **Destination blocks, not target sections**, are the parity site list's unit. +- **Resume's four topologies** are the model for every state-reading rule. +- **The plan must not execute a loop its own product forbids** — step 7 routes a non-closing pass + through the installed ordering before any next call exists. + +**The most reliable defect in this cycle: a repair reaching one site of several.** Passes 26–33 each +found the previous pass's fix applied in one place and missing in two. **When a rule changes, grep +for every statement of it before claiming the repair.** **After a clean close:** implement the plan — both prompt copies, the seven hook strings and their test expectations, version bump 0.11.0 → 0.12.0 + CHANGELOG, the quality battery, the evidence entry, then Gate B with a **fresh nonce** (this cycle's is Gate-A plan only). -**Standing instructions from Daniel:** no stops unless an absolute block; a mandatory two-tell stop -is surfaced in the record and the loop continues on that standing answer. Escalate only new -behaviour decisions and real obstacles. Reports: result, verification, decision-relevant obstacles. - **Two stories still await a profile confirmation:** `docs/superpowers/stories/2026-09-10-record-durability-story.md` and `docs/superpowers/stories/2026-09-10-harness-finding-termination-story.md`. From c55b1bb8424e91cba687030cc51c770ad056a2d9 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Tue, 15 Sep 2026 19:59:42 +0200 Subject: [PATCH 148/181] docs(context): anchor the handoff to the plan revision, not to HEAD --- .context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md | 6 ++++-- 1 file changed, 4 insertions(+), 2 deletions(-) diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md index 38ab8d9..b2ca0ad 100644 --- a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md @@ -13,8 +13,10 @@ Nothing depends on it; the pass files and the repo are authoritative where this ## Resume here -**State: `c7f4617`, tree clean, cycle OPEN and UNCLEAN.** The plan is reviewed at `1849164`; pass 34 -found 1 Blocker, 1 Major, 3 Minors and 1 Nit, **all validated, none repaired**. `.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-34.md` holds them. +**Cycle OPEN and UNCLEAN. The plan stands at `1849164` and that is the anchor** — later commits on +this branch are records, not plan edits, so check `git log --oneline -1 -- docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md` +rather than `HEAD`. Pass 34 reviewed `1849164` and found 1 Blocker, 1 Major, 3 Minors and 1 Nit, +**all validated, none repaired**. `.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-34.md` holds them. **Do not start a repair round on your own.** Daniel's assignment of 2026-09-15 ended with a checkpoint that supersedes the standing autonomy: *"stop and report, whether clean or unclean. Do From 2d79ac8270a1460cc551c16a73cf092c58a71813 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Tue, 15 Sep 2026 20:11:26 +0200 Subject: [PATCH 149/181] docs(plans): check condition 3 before 8a moves; ask git for the dirty set Pass-34 findings 1 and 2, plus the message-extraction mechanism behind finding 3. Findings 4-6 stay collected. Condition 3 had no check before the record commit: 8a discharged 1, 2 and 4 while the index claimed it checked the message too. On a re-entry between 7b and 8a the commit landed before the message was re-established, which is the failure condition 3 exists to prevent. 8a now runs condition 3's reader check as its last precondition; 8b keeps its revalidation against hook rewrites. The dirty-set check split `git status --porcelain -z` into pathnames, which is not decidable: a rename record carries a second, prefix-less path and a pathname may contain a newline. Observed in a disposable repository - an untracked file named with a leading newline passed the check invisibly, and a rename's original was reported as `cked.txt`. The check now asks git which paths outside the expected pair are dirty, via an exclusion pathspec, and names only this plan's own slot names. The record commit's path check moved to the same form. Condition 6 keeps byte equality and is now true as stated: the body is read with `--pretty=format:%B`, which emits the stored message alone, and the closing commit uses `--cleanup=verbatim`, so git stores the validated bytes rather than its own tidied copy. The trailing-blank normalizer is gone with its awk helper and its process substitution, so no block needs bash. Verified by execution, not by reading: the four step-8 blocks were extracted from this file verbatim and run under sh, dash and bash - 29 checks each, the ordinary close plus eleven rejection cases, including the newline path, the rename, the staged modification, a missing message before 8a, and hook rewrites that drop, add and change a record line. --- .../2026-09-14-loop-rule-consolidation.md | 88 ++++++++++++------- 1 file changed, 56 insertions(+), 32 deletions(-) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index de5bdbd..693b8c6 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -268,7 +268,7 @@ independent reader did. | 30 | The final pass's own findings files are the sole permitted post-review addition | **kept** — the rule, unchanged | | 31 | The closing message is rebuilt whole; exactly one provenance line and one curve | **kept** — the rule, unchanged | | 32 | Every owed record is present before the close | **kept** — the rule, unchanged | -| 33 | The dirty set is exactly the final pass's findings files, read from every porcelain record | **kept** as the close procedure's check | +| 33 | The dirty set is exactly the final pass's findings files, read from every porcelain record | **kept** as the close procedure's check, and **corrected**: it asks git which paths outside the expected pair are dirty instead of splitting status output into pathnames, which dropped a newline-bearing path invisibly and mangled a rename's second field (pass 34) | | 34 | The record commit's changed-path set equals those files | **kept** as the close procedure's check | | 35 | The tree is clean after the record commit | **kept** as the close procedure's check | | 36 | `HEAD` equals the reviewed tip before the reset | **kept** as the close procedure's check, and **corrected**: the reviewed *head* and the closing *tip* are two values, not one (pass 21). With row 40 dropped the tip is **only** this precondition — nothing resets to it | @@ -285,8 +285,11 @@ defined. Two rationales that existed only in that prose were moved into the cond to: `reset --soft`'s index behaviour into condition 5, and why the dirty set can be exact into condition 2. **The one correction is condition 6's landed-body check**, which compared bytes where `git log --pretty=%B` adds a trailing newline the source file has none of — it rejected a *correct* -close, and the comparison now strips trailing blank lines on both sides. Verified in a disposable -repository, in both directions. +close. **The predicate is unchanged and is now true as stated**: the body is extracted with +`--pretty=format:%B`, which emits the stored message alone, and the closing commit is made with +`--cleanup=verbatim`, so git stores the validated bytes rather than its own tidied copy. Verified in +a disposable repository, in both directions. An intermediate revision normalized trailing blank +lines on both sides instead; that made the check's own predicate false and is gone. **Two conditions are dropped, 20 and 40, and each is named as a drop rather than lost.** Condition 20 is replaced by a stronger obligation. **Condition 40 is dropped outright** — a guard this plan @@ -360,12 +363,17 @@ an unapproved edit to a spec that the gates already closed. call, in `.context/loop-rule-reviewed-head`. **That is the reviewed head.** It is not the same value as the **closing tip** below, and conflating them is how an unreviewed commit reaches the close. -2. **The only thing dirty is the candidate pass's own findings files** — read from every porcelain - record, not from two status codes, because a staged tracked modification is invisible to a filter - that reads only `??` and `A `. **The set can be exact because step 7 commits every record this - plan collects before the candidate pass is issued**; anything else uncommitted here arrived after - the review and no pass has seen it. A full Gate-B pass writes **two** files, so the pair is what - is expected. +2. **The only thing dirty is the candidate pass's own findings files** — every record counts, not + two status codes, because a staged tracked modification is invisible to a filter that reads only + `??` and `A `. **And the check does not split git's output into pathnames at all**: a pathname + may contain a newline and a rename record carries a second, prefix-less path, so a hand-rolled + split can mangle a name or drop a dirty path entirely — an extra file whose name begins with a + newline passed such a parser invisibly, observed in a disposable repository. **Ask git which + paths outside the expected pair are dirty**, naming only the two expected paths, which are this + plan's own slot names and carry no such bytes. **The set can be exact because step 7 commits + every record this plan collects before the candidate pass is issued**; anything else uncommitted + here arrived after the review and no pass has seen it. A full Gate-B pass writes **two** files, + so the pair is what is expected. 3. **The closing message is complete** — rebuilt whole, exactly one provenance line, exactly one curve, every owed evidence entry, and either the applicable human-exception records or `Human exceptions: none`. **This is checked here, before anything moves**, because an incomplete @@ -2800,7 +2808,8 @@ them, each block naming its condition; **the conditions are the obligation and t to run them** — where the observed state is not one it expects, read the state and pick the operation, rather than extending the block. -**8a — record the candidate pass's findings files. Discharges conditions 1, 2 and 4.** +**8a — record the candidate pass's findings files. Discharges conditions 1, 2, 3 (its pre-move +check) and 4.** ```bash BASE=$(cat .context/loop-rule-base) @@ -2811,21 +2820,36 @@ test -n "$BASE" && test -n "$HEADREV" || { echo "BASE or reviewed head missing"; test "$(git rev-parse HEAD)" = "$HEADREV" || { echo "HEAD is not the head the candidate pass was issued against"; exit 1; } # Condition 2. NONCE and P are this cycle's nonce and the candidate pass number, -# the same two the call's slot paths were built from. +# the same two the call's slot paths were built from. The exclusion pathspec is +# condition 2's "ask git which paths outside the expected pair are dirty". NONCE=; P= FINAL=".context/codex-reviews/gate-b-spec-$NONCE-pass-$P.md .context/codex-reviews/gate-b-quality-$NONCE-pass-$P.md" -expected=$(printf '%s\n' $FINAL | sort) -actual=$(git status --porcelain -z | tr '\0' '\n' | sed -n 's/^.\{3\}//p' | sed '/^$/d' | sort) -test "$expected" = "$actual" || { echo "dirty set is not exactly this pass's findings files:"; git status --porcelain; exit 1; } +set -- +for f in $FINAL; do set -- "$@" ":(exclude)$f"; done +for f in $FINAL; do + test -n "$(git status --porcelain --untracked-files=all -- "$f")" \ + || { echo "this pass's findings file is not dirty: $f"; exit 1; } +done +test -z "$(git status --porcelain -z --untracked-files=all -- "$@")" \ + || { echo "something outside this pass's findings files is dirty:"; git status --porcelain --untracked-files=all; exit 1; } + +# Condition 3, its pre-move check — the last precondition, so nothing has moved if +# it fails. Reader check on the message's records; the expected result is +# condition 3's. 8b revalidates it immediately before the commit consumes it. +test -s .context/loop-rule-closing-msg || { echo "closing message missing or empty"; exit 1; } # Condition 4, first bullet: pin the validated blobs before staging. for f in $FINAL; do git hash-object "$f"; done > .context/loop-rule-final-blobs # shellcheck disable=SC2086 git add $FINAL git commit -m "WIP: Gate-B findings files" || { echo "record commit FAILED — run Failure"; exit 1; } -test "$(git show --name-only --pretty=format: HEAD | sed '/^$/d' | sort)" = "$expected" \ - || { echo "record commit changed paths beyond this pass's findings files — run Failure"; exit 1; } test "$(git rev-parse HEAD^)" = "$HEADREV" || { echo "record commit's parent is not the reviewed head — run Failure"; exit 1; } +git diff --quiet "$HEADREV" HEAD -- "$@" \ + || { echo "record commit changed paths beyond this pass's findings files — run Failure"; exit 1; } +for f in $FINAL; do + git diff --quiet "$HEADREV" HEAD -- "$f" \ + && { echo "record commit did not carry $f — run Failure"; exit 1; } +done # Condition 4, second bullet. for f in $FINAL; do git rev-parse "HEAD:$f"; done > .context/loop-rule-committed-blobs @@ -2841,8 +2865,8 @@ test -z "$(git status --porcelain)" || { echo "tree not clean after the record c git rev-parse HEAD > .context/loop-rule-reviewed-tip ``` -**8b — reset and close. A separate invocation, carrying no `-m` option at all. Discharges conditions -3, 5 and 6.** +**8b — reset and close. A separate invocation, carrying no `-m` option at all. Discharges condition +3's revalidation and conditions 5 and 6.** ```bash BASE=$(cat .context/loop-rule-base); TIP=$(cat .context/loop-rule-reviewed-tip) @@ -2857,7 +2881,10 @@ git reset --soft "$BASE" || { echo "reset --soft FAILED — run Failure; do NOT # Condition 3's revalidation: re-read the message here, immediately before it is # consumed. Reader check on its records; the expected result is condition 3's. test -s .context/loop-rule-closing-msg || { echo "closing message missing or empty — run Failure"; exit 1; } -git commit -F .context/loop-rule-closing-msg +# `--cleanup=verbatim` so the stored body is the validated bytes: git's default +# cleanup for -F strips trailing whitespace and collapses blank runs, and +# condition 6 compares bytes. It carries no `-m`, so `is_wip_commit` still misses it. +git commit --cleanup=verbatim -F .context/loop-rule-closing-msg ``` **Condition 6, in its own invocation — which is why it re-reads `$BASE`.** Every fenced block here is @@ -2871,16 +2898,13 @@ case "$(git log -1 --pretty=%s)" in esac test "$(git rev-parse HEAD^)" = "$BASE" || { echo "closing commit's parent is not \$BASE — run Failure"; exit 1; } test -z "$(git status --porcelain)" || { echo "tree dirty after the close — run Failure"; exit 1; } -git log -1 --pretty=%B > .context/loop-rule-landed-msg -# `%B` emits the body plus one trailing newline the source file does not carry, -# so a byte-for-byte diff rejects a CORRECT close. Compare with trailing blank -# lines stripped from both sides; a rewritten, added or dropped line still -# differs. Verified both ways in a disposable repository. -strip_trailing_blanks() { - awk '{ l[NR]=$0 } END { n=NR; while (n>0 && l[n]=="") n--; for(i=1;i<=n;i++) print l[i] }' "$1" -} -diff <(strip_trailing_blanks .context/loop-rule-closing-msg) \ - <(strip_trailing_blanks .context/loop-rule-landed-msg) \ +# `--pretty=format:%B` emits the stored message alone; the `%B` spelling appends a +# trailing newline the source file has none of, which rejected a CORRECT close. +# 8b's `--cleanup=verbatim` is the other half — without it git stores its own +# tidied copy. Both observed in a disposable repository, so this diff is the +# byte equality condition 6 states, not a normalized stand-in for it. +git log -1 --pretty=format:%B > .context/loop-rule-landed-msg +diff .context/loop-rule-closing-msg .context/loop-rule-landed-msg \ || { echo "the committed body differs from the validated message — run Failure"; exit 1; } ``` @@ -2889,9 +2913,9 @@ beside it rely on the unquoted variable **word-splitting into two paths** — wh `shellcheck disable=SC2086` is there — and `zsh` does not split unquoted parameters by default, so the whole string is taken as one filename and every command fails on a path that does not exist. **Observed, not assumed**: the first verification run of this block was made under `zsh` and failed -exactly that way. The process substitution above also needs `bash`; `sh` on a system where it is -`dash` has none, so use `bash` for that block or write the two normalized copies to temporary -files first. +exactly that way. **Nothing in step 8 needs `bash` specifically** — an earlier revision compared the +two message copies through process substitution, which `dash` has none of; the comparison now reads +two ordinary files. **All four pass, and only then the cleanup Close's success recognition names:** From e4d2fba1589f09c042b55a822acf9336398b5d3c Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Tue, 15 Sep 2026 20:24:37 +0200 Subject: [PATCH 150/181] docs(context): record Gate-A plan pass 35 and the second bounded repair Pass 35 reviewed 2d79ac8: 1 Major, 0 Blockers, down from 6 and 1. Both pass-34 repairs held - nothing re-raised against 8a's condition-3 check, the dirty-set query or condition 6. The open Major is carried condition h3, whose preservation count passage (h) assigns to Task 8 while Task 8 owns only h4, h5 and h19. The working record also loses two claims that are no longer true: that condition 6 normalizes trailing blanks, and that Failure and Resume need a real incident to verify - controlled failures in a disposable repo are enough. Floor 3, from docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md (risk high, security none), read fresh at this pass. Cycle stays open and unclean; no repair round opened. --- .../gate-a-plan-om0bdd7udh-pass-35.md | 2 + .../gate-a-plan-om0bdd7udh-resume.md | 97 +++++++++++++------ 2 files changed, 70 insertions(+), 29 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-35.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-35.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-35.md new file mode 100644 index 0000000..1483112 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-35.md @@ -0,0 +1,2 @@ +MAJOR | high | Passage (h) disposition and Task 8 steps 1–4 | The disposition table says carried condition `h3` owes a preservation count and explicitly claims Task 8 derives its fragment before installing item 7, but Task 8 names only `h4`, `h5`, and `h19`, derives only `a1` beyond the existing F rows, and never runs or records an `h3` preservation observation. Row F7 cannot supply that check: its OLD spans `h3` into changed `h4` wording and is required to disappear, while its unconstrained NEW may omit `h3`. | Item 7 can drop `an ungated change records it in that commit` from both prompt copies while all fifteen pairs, the `a1` check, parity, the reader walk as currently scoped, and the battery pass, so the plan does not discharge its own 135-condition accounting or story criterion 5. | Add `h3` to Task 8's owned conditions; before step 2 derive and table its own single-line preservation fragment from the live sentence, then after installation require `parent=1 worktree=1` in both copies and record it in Task 8's fragment evidence and closing-entry reruns. +END OF FINDINGS (1 total) diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md index b2ca0ad..49862a4 100644 --- a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md @@ -13,33 +13,65 @@ Nothing depends on it; the pass files and the repo are authoritative where this ## Resume here -**Cycle OPEN and UNCLEAN. The plan stands at `1849164` and that is the anchor** — later commits on +**Cycle OPEN and UNCLEAN. The plan stands at `2d79ac8` and that is the anchor** — later commits on this branch are records, not plan edits, so check `git log --oneline -1 -- docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md` -rather than `HEAD`. Pass 34 reviewed `1849164` and found 1 Blocker, 1 Major, 3 Minors and 1 Nit, -**all validated, none repaired**. `.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-34.md` holds them. +rather than `HEAD`. Pass 35 reviewed `2d79ac8` and found **one Major and nothing else** +(0 Blockers). `.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-35.md` holds it. **Do not start a repair round on your own.** Daniel's assignment of 2026-09-15 ended with a -checkpoint that supersedes the standing autonomy: *"stop and report, whether clean or unclean. Do -not begin another repair/review round or implementation."* That checkpoint is spent — it covered one -revision and one pass, both done — so **the next move is Daniel's word, not an inference.** - -### The four open findings, in the order they cost most - -1. **BLOCKER — step 8a never checks Close condition 3.** It discharges 1, 2 and 4; the index claims - 8a and 8b both check the message. On a re-entry between 7b and 8a the record commit lands before - the message is re-established, and 8b is the first to notice — the exact failure condition 3 - exists to prevent. -2. **MAJOR — the porcelain parser mangles two real pathname shapes.** `git status --porcelain -z | - tr '\0' '\n' | sed 's/^.\{3\}//'` turns a rename into `p.txt` and a newline-bearing path into two - fragments. **Verified by execution, not by reading.** -3. **MINOR — condition 6 says the landed body is "identical"** where the corrected oracle strips - trailing blank lines. Introduced by the bounded revision itself. -4. **MINOR — step 8 is titled "two invocations" and has four fenced blocks**; **MINOR** — the - Architecture paragraph still says every task runs a discriminating pair, which the disposition - table replaced; **NIT** — `loop-rule-untouched.tmp` and `loop-rule-baseline-diff.tmp` are written - and never removed by the cleanup. - -### What the bounded revision did, so it is not undone by accident +checkpoint: *"After the repairs and verification, run exactly one complete Gate-A plan pass, then +stop and report regardless of outcome. Do not automatically repair its findings or start +implementation."* That checkpoint is spent — it covered one bounded repair and one pass, both +delivered — so **the next move is Daniel's word, not an inference.** + +### The one open finding from pass 35 + +**MAJOR — carried condition `h3` has no observation.** Passage (h)'s disposition table says `h3` is +carried, owes a preservation count, and that *"Task 8 derives its fragment before installing item +7"*. **Task 8 does not.** It names `h4`, `h5` and `h19` as its conditions, derives only `a1`'s +preservation fragment, and step 4 confirms only `a1`. Row F7 cannot stand in: its OLD spans `h3` +into changed `h4` wording and is required to reach zero, and its NEW is unconstrained. So item 7 +could drop `an ungated change records it in that commit` from both copies with all fifteen pairs, +the `a1` check, parity and the battery still green. **Validated against the plan, not taken on +trust.** It is the cycle's most frequent family — a kept-or-carried condition losing its +observation, as at passes 10, 15 and 25. + +### Still collected and deliberately unrepaired (pass 34's Minors and Nit) + +**MINOR** — step 8 is titled "two invocations" and carries four fenced blocks, and the index calls +the last two "8b blocks"; **MINOR** — the Architecture paragraph still says every task runs a +discriminating pair, which the disposition table replaced; **NIT** — +`.context/loop-rule-untouched.tmp` and `.context/loop-rule-baseline-diff.tmp` are written and never +removed by the cleanup. + +### What the second bounded repair did (pass-34 findings 1, 2 and the mechanism behind 3) + +**Condition 3 now has a pre-move check.** 8a runs the closing message's reader check as its **last +precondition**, so nothing has moved if it fails; 8b keeps its revalidation immediately before the +commit consumes the file. 8a's label and the index row agree with that now. + +**The dirty-set check no longer parses pathnames at all.** It asks git, through an exclusion +pathspec, which paths *outside* the expected pair are dirty, and names only this plan's own slot +names, which carry no awkward bytes. The record commit's changed-path check uses the same form. +**Why, observed rather than reasoned:** an untracked file named with a **leading** newline passed the +old parser completely invisibly, and a rename's original was reported as `cked.txt`. + +**Condition 6 keeps byte equality and is now true as stated.** The body is read with +`--pretty=format:%B` — the `%B` spelling appends a newline the source has none of — and the closing +commit carries `--cleanup=verbatim`, because git's default cleanup for `-F` strips trailing +whitespace and collapses blank runs, which would have re-created the same false rejection. The +trailing-blank normalizer, its `awk` helper and its process substitution are gone, so **no step-8 +block needs `bash`** any more. `--cleanup=verbatim` was checked against `is_wip_commit`: it does not +match. + +**Verified by extraction and execution, not by reading:** the four step-8 blocks were pulled from +the plan verbatim and run under `sh`, `dash` and `bash` — 29 checks each, all green. The ordinary +close plus eleven rejection cases: missing message before 8a, extra ordinary path, leading-newline +path, staged rename, staged modification, one findings file only, `HEAD` moved between 8a and 8b, +and hook rewrites that drop, add and change a record line. **A harness bug made the first run report +false "ok"s** (a relative path after `cd`) — read the failure column, not the tally. + +### What the first bounded revision did, so it is not undone by accident Each closing condition had been stated **three times** — a Close condition, a command in Task 15 step 8, and prose beside that command. The third copy is gone. **`### Where each operational @@ -49,13 +81,18 @@ naming the condition it discharges; Task 0 step 1 and accounting row 7 cite rath **One correction, found by running it:** condition 6 compared bytes, and `git log --pretty=%B` adds one trailing newline the source file has none of — **it rejected a correct close, so the plan had no -working success path.** Fixed and checked in four directions. +working success path.** Its first fix normalized trailing blanks, which pass 34 caught as making the +condition's own predicate false; the second repair replaced it with exact extraction (see above). -**Verified in a disposable repo:** the whole of step 8 end to end. **Unverified:** every reader check, -Failure and Resume — they need a real failure or a real interruption. +**Verified in a disposable repo:** the whole of step 8 end to end, twice — after each bounded +repair. **Unverified:** every reader check, and the Failure and Resume procedures. **They do not need +a real incident**, though: controlled failures and interruptions in the disposable repo are enough, +and that is the cheapest unclaimed verification left in this cycle. An earlier revision of this +record said they needed a real failure; that was wrong. -**The plan's shell needs `sh` or `bash`, never `zsh`** — `$FINAL` relies on word-splitting and the -body comparison uses process substitution. +**The plan's shell needs `sh`, `dash` or `bash`, never `zsh`** — `$FINAL` relies on word-splitting, +which `zsh` does not do for unquoted parameters. **`bash` is no longer required anywhere in step 8**; +the process substitution that once forced it is gone. ### How to run a pass, if Daniel asks for one @@ -144,6 +181,8 @@ worth checking before a pass rather than after. | 23 | 9ced53c | 10→**5** | 5→**3** | 3→**2** | yes | **all five in Task 0 / Task 15 again.** A symbolic value in the base file; the restore target chosen by which tip file exists, which rewinds past a second candidate's repair; the failure capture covering tracked content only while the closure inputs are ignored paths; pass 22's diff-status classification written as prose and not as code; the close procedure's "check all six before moving `HEAD`" applied to two checks that run after a move | | — | — | — | — | — | — | **BOUNDED REVISION**, Daniel's assignment after pass 33. The third statement of every closing condition removed; `### Where each operational condition is defined` gives one home each; step 8's shell kept and annotated with the condition it discharges. Condition 6's landed-body check **corrected** — it compared bytes where `%B` adds a trailing newline, so it rejected a *correct* close. Whole close path verified end to end in a disposable repo. Commit `1849164` | | 34 | 1849164 | **6** | **1** | **1** | yes | **UNCLEAN — reported to Daniel, no repair round opened.** 8a discharges conditions 1, 2, 4 and never checks condition 3, while the index claims it does; the porcelain parser mangles a rename (`p.txt`) and a newline pathname (two fragments), verified by execution; condition 6 says "identical" where the corrected oracle normalizes trailing blanks; step 8 is titled "two invocations" and has four fenced blocks; the Architecture paragraph still says every task runs a discriminating pair; the two `.tmp` names are not in the cleanup list | +| — | — | — | — | — | — | **SECOND BOUNDED REPAIR**, Daniel's assignment after pass 34, on his upstream reviewer's recommendation. Pass-34 findings 1 and 2 repaired plus the extraction mechanism behind 3; findings 4–6 left collected. Condition 3 gains a pre-move check in 8a; the dirty set is asked of git through an exclusion pathspec instead of parsed; condition 6 keeps byte equality via `--pretty=format:%B` + `--cleanup=verbatim`, and the normalizer, the `awk` helper and the process substitution are gone. Step 8 re-run end to end under `sh`, `dash` and `bash`, 29 checks each. Commit `2d79ac8` | +| 35 | 2d79ac8 | 6→**1** | 1→**0** | 1→**1** | yes | **UNCLEAN — reported to Daniel, no repair round opened.** One Major: carried `h3` owes a preservation count and passage (h) says Task 8 derives its fragment, but Task 8 owns only `h4`/`h5`/`h19` and derives only `a1`, so item 7 could drop `an ungated change records it in that commit` from both copies undetected. **Both pass-34 repairs held** — nothing re-raised against 8a, the dirty set or condition 6. One tell (instrument cluster); no mandatory stop | | 26 | 05ec1b2 | **5** | **1** | 2 | yes | first pass on the revised plan. Resume audited progress by `WIP:` commits and called any non-`WIP` commit a stale base — the handoff leaves two other shapes, one of them `HEAD` **at** the base with everything in the index | | 27 | f69db7d | 5→**6** | 1→**3** | 2→**1** | yes | the close required the commit's **tree to equal the tip's**, which target §I **parks**; `reset --soft` leaves the index untouched, so the prose beside it was wrong about git; Resume still inferred staleness from a commit subject where §A3 names that state as reachable | | 28 | f2e5dda | 6→**6** | 3→**2** | 1→**2** | yes | pass 27's stale-base fix reached one site of three; its reset paragraph stated the index behaviour correctly and repeated the false claim four lines later; nothing checked the closing commit's **parent**; the message was validated in the file, where `commit-msg` hooks rewrite git's copy after `-F` reads it | From 518121a39122c098a84d7d25fc3c3392c8d8f6cc Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Wed, 16 Sep 2026 09:25:01 +0200 Subject: [PATCH 151/181] docs(plans): let Task 8 discharge carried h3, via the procedure not a list Pass-35 Major. Passage (h) classifies h3 as carried, says it owes a preservation count, and assigns the derivation to Task 8. Task 8 did not do it: a line at the top named h4, h5 and h19 as "the conditions these items discharge" and omitted carried h3 and carried a1, and steps 1 and 4 handled only a1. F7 and F7b cannot stand in - both are OLD halves required to reach zero, and item 7's NEW text is not constrained to carry h3 - so item 7 could have dropped `an ungated change records it in that commit` from both prompt copies with every pair, the parity diff and the battery still green. The repair is to stop enumerating rather than to extend the enumeration. An enumeration inside a task is the second copy of the disposition table that `## How a task discharges that table` forbids, and this is the fourth time in this cycle that a carried or kept condition lost its only observation to one. Task 8 now cites the procedure; steps 1 and 4 walk every carried condition its blocks cover, with a1 and h3 kept as worked examples of what is invisible without the block open, not as the set owed. Two directly affected references corrected while pinning the result: the disposition table's carried row said a preservation count is `1` in each copy where the procedure and the evidence shape both say `parent=1 worktree=1`, and Task 7's step restated the same result as `1` alone. Verified on temporary copies, the live prompt copies untouched: item 7's install built from the target block in both copies, the h3 count is parent=1 worktree=1 on the intended text and fails on h3 removed from C alone, from W alone, and from both. F7's OLD reaches its required 0 in the damaged install as well as the intended one, which is the demonstration that it observes nothing about h3. --- .../2026-09-14-loop-rule-consolidation.md | 61 +++++++++++++------ 1 file changed, 44 insertions(+), 17 deletions(-) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index 693b8c6..bdf4e44 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -604,7 +604,7 @@ different passage, because each task was inventing the rule for its own conditio | Disposition | The observation it owes | |---|---| | **kept** | **exactly one of three routes**, never two and never none: inside an untouched span; **or** its own per-condition count, `parent=1 worktree=1` in each copy, where it shares a line with changed text; **or** that same per-condition count where **no untouched span reaches its passage at all** — Task 0 maps five regions, and passages (b), (c), (e) and (i) are not among them, so every kept condition there takes this third route. | -| **carried** | a **preservation count** after installation, `1` in each copy, from the condition's own text. **No untouched span covers a carried condition** — it sits inside a replacement block, which is the whole reason it is not recorded as kept. | +| **carried** | a **preservation count** after installation, `parent=1 worktree=1` in each copy, from the condition's own text — the same two values the *kept* route above owes, and stated the same way here because an earlier wording said `1` and left which count ambiguous. **No untouched span covers a carried condition** — it sits inside a replacement block, which is the whole reason it is not recorded as kept. | | **replaced** | a **discriminating pair**, `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`, **both halves from the same edit**. | | **moved** | **two** observations: an **absence** at the source, `parent=1 worktree=0`, and a **condition-specific presence** at the destination, `worktree=1 parent=0`. A presence check on the destination *paragraph* is not the second half — it passes while any one moved predicate is missing from it. | | **dropped** | an **absence check**, `parent=1 worktree=0`, **one per dropped condition**. Two dropped conditions sharing one fragment is one observation, and it goes absent when either half goes, leaving the other free to survive. | @@ -1766,8 +1766,9 @@ Expected for all twenty: `old/worktree=0 old/parent=1 new/worktree=1 new/parent= parent=0` there. **An earlier draft derived these rows at step 1b and then never ran them**, so the old clean-final-pass, loop-until-clean, zero-finding and no-padding instructions could each survive at their source with nothing observing it; -- **a preservation count** of `1` in each copy for each of the nine carried conditions and each - kept condition in passage (i), from the fragments step 1b appended; +- **a preservation count** to `parent=1 worktree=1` in each copy for each of the nine carried + conditions and each kept condition in passage (i), from the fragments step 1b appended — the + result the disposition table owes, which an earlier wording gave here as `1` alone; - **a presence check** — `new/worktree=1 new/parent=0` in each copy — for §G's semantic membership test and for every independent add-only clause in the strict-reading tail. @@ -1816,7 +1817,13 @@ git commit -m "WIP: install the one-contract paragraph and the remaining prompt- - Modify: `docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md` — the rows this task appends and its fragment evidence **The bytes:** target §F items 1, 2, 3, 4, 5, 6, 7, 8, 7a, 8a, 8b, 9a, 9b and 9 — fourteen items, each with its own fenced replacement and its `C nnn` / `W nnn` citation. **Re-read every citation against the current file**: §F's own collected list records that items 4, 5 and 8 have line citations one off, and the numbers drifted further as this cycle edited the copies. -**`h4`, `h5` and `h19` are the human-exception conditions these items discharge** — item 7 is both destinations, `h4`'s and `h5`'s, item 4 the scope sentence. +**Which conditions this task discharges is decided by `## How a task discharges that table`, run +against the disposition rows for the blocks these fourteen items replace — not by a list here.** +This line carried one: it named `h4`, `h5` and `h19` and silently omitted **carried `h3` and carried +`a1`**, and that is how `h3` reached pass 35 with no observation at all. An enumeration inside a +task is the second copy of the disposition table, which this plan forbids two sections up and has +now been bitten by from inside the section that forbids it. As orientation, not as the set owed: +item 7 carries both the `h4` and `h5` destinations, and item 4 the `h19` scope sentence. - [ ] **Step 1: Re-derive every item's real location** @@ -1834,12 +1841,27 @@ pre-edit count of 1 can no longer be observed at all. A row that no longer count since the table was verified — repair the row against the live line and update the table before installing anything. -**And derive `a1`'s preservation fragment here too.** `a1` is **carried** inside item 8a's block, -which reproduces it — so it is pre-existing text and the fragment-table cut puts it in the table, -before the install, like every other pre-existing fragment. An earlier draft chose it at step 4, -after item 8a had already replaced the block: at that point a drifted or half-installed HARD FLOOR -opening cannot be told from the intended carried text, and no authored fragment exists for Task 15 -to audit. +**And derive the preservation fragment of every carried condition these blocks cover, here** — they +are pre-existing text, so the fragment-table cut puts them in the table before the install like +every other pre-existing fragment, under the next free `P` id. `## How a task discharges that +table` is what says which conditions those are. **Two of them are invisible without the block open, +and both have been missed:** + +- **`a1`**, carried inside item 8a's block, which opens with it — `**Both gates are a LOOP with a + HARD FLOOR: a minimum number of passes per run`. Row F10's pair observes `a2`, the parenthetical, + not the opening it sits in. +- **`h3`**, carried inside item 7's block, which reproduces `an ungated change records it in that + commit` verbatim because the item replaces the whole `**Which commit:**` sentence and only the + Gate-A clause changes. **Neither F7 nor F7b can stand in for it:** both are OLD halves required to + reach **zero**, F7 runs from `h3`'s own wording into the changed `h4` wording, and item 7's NEW + text is not constrained to carry `h3` at all. So item 7 could drop that clause from both copies + with all fifteen pairs, the parity diff and the battery still green — pass 35, and the fourth time + in this cycle a carried or kept condition lost its only observation. + +An earlier draft chose `a1`'s fragment at step 4, after item 8a had already replaced the block: at +that point a drifted or half-installed HARD FLOOR opening cannot be told from the intended carried +text, and no authored fragment exists for Task 15 to audit. **The same applies to `h3` and to every +other fragment on this list** — derive before installing, never after. - [ ] **Step 2: Install all fourteen replacements** @@ -1861,14 +1883,19 @@ every stated count still passed. Expected for all thirty pair instances — fifteen rows in each of the two copies: `old/worktree=0 old/parent=1 new/worktree=1 new/parent=0`. -- [ ] **Step 4: Confirm `a1` survived, then count what was installed** +- [ ] **Step 4: Confirm the carried conditions survived, then count what was installed** -**`a1` is carried inside item 8a's block**, which opens with it — `**Both gates are a LOOP with a -HARD FLOOR: a minimum number of passes per run` — and reproduces it so one contiguous string -installs. Carried, not kept: **no untouched-range span covers it**, and row F10's pair observes -`a2`, the parenthetical, not the opening it sits in. **Count the fragment step 1 appended**, in -each copy, expecting `1`. Without it a mis-scoped item-8a replacement can drop the sentence's -opening and every other check in this task still passes. +**Count every preservation fragment step 1 appended** — in each copy, to the result its class owes: +`parent=1 worktree=1`, per `## What each disposition owes, stated once`. **Both values matter.** The +worktree half is what proves the block preserved the condition; the parent half is what proves the +fragment was ever in the file it claims to observe, so a mistyped one reports `parent=0 worktree=0` +instead of a healthy-looking `1`. Record each as a `preservation` line in this task's fragment +evidence — which is what carries it into Task 15's re-run set and the closing evidence entry. + +**No untouched-range span covers any of them**, which is why each owes its own count. Without these +counts a mis-scoped replacement passes every other check in this task: item 8a can drop the HARD +FLOOR sentence's opening, and item 7 can drop `an ungated change records it in that commit`, while +all fifteen pairs, the parity diff and the battery stay green. Then total the fifteen `new/worktree` values step 3 printed, per copy. From 1dc79905be43d9ce0981ccd90aadd39d3ef131de Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Wed, 16 Sep 2026 09:41:08 +0200 Subject: [PATCH 152/181] docs(context): record Gate-A plan pass 36 and the third bounded repair Pass 36 reviewed 518121a: 1 Major, 1 Minor, 0 Blockers. The h3 repair held - nothing re-raised against Task 8. Open Major: Resume's topology table names four topologies and says the whole procedure reads against all four, but has no row for the state Close condition 5 exists to catch - HEAD above the 8a findings commit after a commit lands between the two invocations. The condition-1 rejection reads as the "normal, mid-implementation" row, which is how an unreviewed commit gets treated as ordinary cycle progress. Open Minor, collected: Task 4 step 5's reader walk omits replaced c4. Same family as pass 35's Major - a task enumerating its own conditions - so the working record now carries that as a per-pass check. Two tells present (findings rose 1 to 2, instrument cluster for the third pass), so the stop is mandatory as well as instructed. Whether a Blocker count of 0 to 0 counts as "failing to fall" is not settled here and nothing turns on it. Floor 3, from docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md (risk high, security none), read fresh at this pass. Cycle stays open and unclean. --- .../gate-a-plan-om0bdd7udh-pass-36.md | 3 + .../gate-a-plan-om0bdd7udh-resume.md | 59 ++++++++++++------- 2 files changed, 41 insertions(+), 21 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-36.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-36.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-36.md new file mode 100644 index 0000000..320e0ae --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-36.md @@ -0,0 +1,3 @@ +MAJOR | high | Resume topology table and Close conditions 1 and 5 | Resume claims four exhaustive re-entry topologies, but the close deliberately rejects when `HEAD` has moved before 8a or between 8a and 8b. In those reachable states `HEAD` can be a new commit above the reviewed head or above the 8a findings tip; it is neither the last WIP snapshot, the unchanged pre-8a tip, the 8a findings commit itself, nor `$BASE` or one commit parented by `$BASE` after the soft reset, so no table row describes it. | The bounded Failure handoff can return a repository state that Resume says it owns but cannot classify against the model every state-reading rule is required to use; an executor can reconcile the wrong content or treat the extra commit as ordinary cycle progress even though no candidate pass reviewed it. | Add or generalize a pre-reset HEAD-moved topology covering both condition-1 and condition-5 rejection, then audit base validity, artifact handling, progress reconciliation and the next section-A route against that state without reinstating restoration. +MINOR | high | Task 4 step 5 | The closed reader-walk list names `c1`–`c3`, `c5`–`c8`, `c9`, and `c10`–`c14` but omits changed condition `c4`, even though the governing task procedure requires a reader walk over every condition the block covers. | The `c4` OLD/NEW sample can pass when the widened third-condition sentence is otherwise malformed or incomplete, while parity merely reproduces the same defect in both copies; the semantic confirmation this plan requires never reads that condition. | Replace the local condition enumeration with a citation to the governing per-block walk and require that walk to include every disposition row for Task 4's block, using `c4` only as an orientation example if one is needed. +END OF FINDINGS (2 total) diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md index 49862a4..f70a997 100644 --- a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md @@ -13,28 +13,38 @@ Nothing depends on it; the pass files and the repo are authoritative where this ## Resume here -**Cycle OPEN and UNCLEAN. The plan stands at `2d79ac8` and that is the anchor** — later commits on +**Cycle OPEN and UNCLEAN. The plan stands at `518121a` and that is the anchor** — later commits on this branch are records, not plan edits, so check `git log --oneline -1 -- docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md` -rather than `HEAD`. Pass 35 reviewed `2d79ac8` and found **one Major and nothing else** -(0 Blockers). `.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-35.md` holds it. - -**Do not start a repair round on your own.** Daniel's assignment of 2026-09-15 ended with a -checkpoint: *"After the repairs and verification, run exactly one complete Gate-A plan pass, then -stop and report regardless of outcome. Do not automatically repair its findings or start -implementation."* That checkpoint is spent — it covered one bounded repair and one pass, both -delivered — so **the next move is Daniel's word, not an inference.** - -### The one open finding from pass 35 - -**MAJOR — carried condition `h3` has no observation.** Passage (h)'s disposition table says `h3` is -carried, owes a preservation count, and that *"Task 8 derives its fragment before installing item -7"*. **Task 8 does not.** It names `h4`, `h5` and `h19` as its conditions, derives only `a1`'s -preservation fragment, and step 4 confirms only `a1`. Row F7 cannot stand in: its OLD spans `h3` -into changed `h4` wording and is required to reach zero, and its NEW is unconstrained. So item 7 -could drop `an ungated change records it in that commit` from both copies with all fifteen pairs, -the `a1` check, parity and the battery still green. **Validated against the plan, not taken on -trust.** It is the cycle's most frequent family — a kept-or-carried condition losing its -observation, as at passes 10, 15 and 25. +rather than `HEAD`. Pass 36 reviewed `518121a` and found **one Major and one Minor, 0 Blockers**. +`.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-36.md` holds them. + +**Do not start a repair round on your own.** Daniel's assignment of 2026-09-16 ended with a +checkpoint: *"After the repair and focused verification, run exactly one complete Gate-A plan +review, then stop and report. Do not automatically begin another repair round or implementation."* +That checkpoint is spent — it covered one bounded repair and one pass, both delivered — so **the +next move is Daniel's word, not an inference.** + +**Two tells present, so this stop is mandatory as well as instructed**: findings rose 1 → 2, and the +findings cluster on the instrument for the third pass running. A third reading is arguable — the +Blocker count "failed to fall" at 0 → 0, which is the literal test and nonsense at zero. **That +reading is not settled here and nothing turns on it**: the stop happens either way. + +### The two open findings from pass 36 + +1. **MAJOR — Resume's topology table is not exhaustive, and claims to be.** It names four + topologies and says *"the whole procedure reads against all four"*; the working record calls them + *"the model for every state-reading rule"*. But Close condition 5 exists precisely to catch **a + commit landing between 8a and 8b**, and in that rejected state `HEAD` sits **above** the 8a + findings commit — which is not the last `WIP:` snapshot, not the findings commit itself, and not + `$BASE` or its child after the reset. **No row describes it.** The condition-1 rejection has a + milder version of the same problem: a stray commit above the reviewed head reads as row 1, + "Normal, mid-implementation", which is how an unreviewed commit gets treated as ordinary cycle + progress. **Validated against the plan, not taken on trust.** +2. **MINOR — Task 4 step 5's reader walk omits `c4`.** It walks `c1`–`c3`, `c5`–`c7`, `c8`, `c9`, + `c10`–`c14`; passage (c) carries a `c4` row and `c4` is **replaced**. The mechanical pair still + covers it; the semantic confirmation never reads it. **Same defect family as pass 35's Major, in + a different task** — a task enumerating its own conditions and dropping one. Collected per the + Minor rule, not repaired. ### Still collected and deliberately unrepaired (pass 34's Minors and Nit) @@ -123,6 +133,11 @@ what it carries; `.context/gate-a-spec-prompt.md` is the same shape for the spec - **Resume's four topologies** are the model for every state-reading rule. - **The plan must not execute a loop its own product forbids** — step 7 routes a non-closing pass through the installed ordering before any next call exists. +- **A task that enumerates its own conditions is the second copy of the disposition table**, and it + goes stale silently. Task 8's list dropped carried `h3` (pass 35, MAJOR); Task 4 step 5's reader + walk drops `c4` (pass 36, MINOR). **Grep every task for id lists and check each against its + passage's rows** — it has now produced a finding in two consecutive passes, and the rule against + it is already written in `## How a task discharges that table`. **The most reliable defect in this cycle: a repair reaching one site of several.** Passes 26–33 each found the previous pass's fix applied in one place and missing in two. **When a rule changes, grep @@ -183,6 +198,8 @@ worth checking before a pass rather than after. | 34 | 1849164 | **6** | **1** | **1** | yes | **UNCLEAN — reported to Daniel, no repair round opened.** 8a discharges conditions 1, 2, 4 and never checks condition 3, while the index claims it does; the porcelain parser mangles a rename (`p.txt`) and a newline pathname (two fragments), verified by execution; condition 6 says "identical" where the corrected oracle normalizes trailing blanks; step 8 is titled "two invocations" and has four fenced blocks; the Architecture paragraph still says every task runs a discriminating pair; the two `.tmp` names are not in the cleanup list | | — | — | — | — | — | — | **SECOND BOUNDED REPAIR**, Daniel's assignment after pass 34, on his upstream reviewer's recommendation. Pass-34 findings 1 and 2 repaired plus the extraction mechanism behind 3; findings 4–6 left collected. Condition 3 gains a pre-move check in 8a; the dirty set is asked of git through an exclusion pathspec instead of parsed; condition 6 keeps byte equality via `--pretty=format:%B` + `--cleanup=verbatim`, and the normalizer, the `awk` helper and the process substitution are gone. Step 8 re-run end to end under `sh`, `dash` and `bash`, 29 checks each. Commit `2d79ac8` | | 35 | 2d79ac8 | 6→**1** | 1→**0** | 1→**1** | yes | **UNCLEAN — reported to Daniel, no repair round opened.** One Major: carried `h3` owes a preservation count and passage (h) says Task 8 derives its fragment, but Task 8 owns only `h4`/`h5`/`h19` and derives only `a1`, so item 7 could drop `an ungated change records it in that commit` from both copies undetected. **Both pass-34 repairs held** — nothing re-raised against 8a, the dirty set or condition 6. One tell (instrument cluster); no mandatory stop | +| — | — | — | — | — | — | **THIRD BOUNDED REPAIR**, Daniel's assignment after pass 35, scoped to that Major alone. **The fix was to stop enumerating, not to extend the enumeration**: Task 8's top line cited `## How a task discharges that table` instead of naming `h4`/`h5`/`h19`, step 1 derives the preservation fragment of every carried condition its blocks cover, step 4 counts each to `parent=1 worktree=1` and records a `preservation` line. Two directly affected references pinned to the same result — the disposition table's carried row and Task 7's step, both of which said `1` alone. Verified on temporary copies: the `h3` count passes on the intended install and fails on `h3` removed from C alone, W alone and both, and F7's OLD reaches 0 either way. Commit `518121a` | +| 36 | 518121a | 1→**2** | 0→**0** | 1→**1** | yes | **UNCLEAN — reported to Daniel, no repair round opened.** **The h3 repair held**; nothing re-raised against Task 8. MAJOR: Resume's topology table claims four exhaustive topologies and has **no row for the state condition 5 exists to catch** — `HEAD` above the 8a findings commit after a between-invocations commit; the condition-1 rejection reads as row 1, "normal". MINOR: Task 4 step 5's reader walk omits `c4` — the **same family** as pass 35's Major, a task enumerating its own conditions. Two tells (findings rose, instrument cluster) → mandatory stop, which the checkpoint already required | | 26 | 05ec1b2 | **5** | **1** | 2 | yes | first pass on the revised plan. Resume audited progress by `WIP:` commits and called any non-`WIP` commit a stale base — the handoff leaves two other shapes, one of them `HEAD` **at** the base with everything in the index | | 27 | f69db7d | 5→**6** | 1→**3** | 2→**1** | yes | the close required the commit's **tree to equal the tip's**, which target §I **parks**; `reset --soft` leaves the index untouched, so the prose beside it was wrong about git; Resume still inferred staleness from a commit subject where §A3 names that state as reachable | | 28 | f2e5dda | 6→**6** | 3→**2** | 1→**2** | yes | pass 27's stale-base fix reached one site of three; its reset paragraph stated the index behaviour correctly and repeated the false claim four lines later; nothing checked the closing commit's **parent**; the message was validated in the file, where `commit-msg` hooks rewrite git's copy after `-F` reads it | From d7f2af82f5f6262ed5e87581fc79df758cef7f5d Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Wed, 16 Sep 2026 09:59:16 +0200 Subject: [PATCH 153/181] docs(plans): a topology for a moved HEAD; Task 4 walks the table, not a list Pass 36's Major and Minor. Resume claimed four exhaustive topologies and had no row for the state Close condition 5 exists to catch: a commit landing between 8a and 8b leaves HEAD above the findings commit, which is neither the WIP tip, nor the findings commit, nor $BASE or its child after the reset. Condition 1's rejection had the milder version - a commit above the reviewed head reads as row 1, "normal, mid-implementation". One row now covers both, keyed on why the commit is there rather than on its shape, and the two rejections are distinguished where it matters: condition 1 is a precondition and a plain stop, condition 5 rejects after 8a's commit so a handoff is in progress and the scratch-artifact rule's inside-a-handoff branch applies. The progress reconciliation gains the distinction the gap was made of: present is not reviewed. Content above the reviewed head is not adopted - a ticked checkbox does not make it reviewed, and Close condition 1 decides what it costs - and not removed either, because this plan restores nothing and the commit is evidence a person has not yet chosen a section-A route on. Base validity is unaffected and now says so: the extra commit sits above the base, so ancestry holds. Task 4 step 5's reader walk was a closed list that omitted replaced c4, which has a pair at step 4 and was therefore counted and never read. It now walks every disposition row its block covers, the way step 4b already said to, with the per-condition expectations kept as orientation and c4's added. Bounded search of the remaining tasks: Tasks 3, 5 and 7 already read their sets off the disposition table; Task 8 was repaired last round. Task 7's id lists and Task 4 step 4's move list were checked against their passages and are complete, so neither was touched. This is not a claim about the whole file. Also corrected, and disclosed as a miss in the previous round's sweep: four more sites stated a carried condition's result as `1` where the disposition table says `parent=1 worktree=1` - Task 4 step 4b, Task 5 step 4, Task 7 step 4 and its expected line. The earlier grep keyed on "preservation count" and these do not use the phrase. Two dependent count claims the new topology row falsified are fixed with it: the operational-condition index and the WIP-naming paragraph both said "four topologies". Verified by execution under sh, dash and bash, 18 checks each: both HEAD-movement cases reject before any reset, the WIP chain and 8a's findings commit survive case B, the closing message is untouched, the base still validates through Resume's own checks, and the extra commit is visible to the progress reconciliation in both. Task 4's step 5 was diffed against 518121a: every id and every per-condition phrase the old list carried is retained, c4 is added, and the set now equals what the disposition table assigns. --- .../2026-09-14-loop-rule-consolidation.md | 69 +++++++++++++++---- 1 file changed, 55 insertions(+), 14 deletions(-) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index bdf4e44..e74f0f9 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -211,7 +211,7 @@ live in the home named. Where a task, a shell comment or an error string needs a |---|---|---| | Branch, clean tree, no base file yet, `ba15e83` ancestral, approved inputs unchanged at `HEAD` | **Preparation** | Task 0 step 1, first-entry branch | | Base file shape, self-resolving commit, ancestral; approved inputs unchanged at `$BASE`; scratch-artifact validity; how far the implementation got | **Resume** | Task 0 step 1, re-entry branch | -| The four repository topologies a re-entry can meet | **Resume**, its table | accounting row 7; every state-reading rule | +| The repository topologies a re-entry can meet | **Resume**, its table | accounting row 7; every state-reading rule | | `HEAD` equals the reviewed head | **Close**, condition 1 | step 8a | | The dirty set is exactly the candidate pass's findings files | **Close**, condition 2 | step 8a | | The closing message is complete, and revalidated before it is consumed | **Close**, condition 3 | step 7b writes it; steps 8a and 8b check it | @@ -498,15 +498,17 @@ and records the closing commit carries are the ones a pass actually validated. **Resume owns every entry after the first.** Preparation is first-entry-only and refuses when a base file exists, so nothing here defers to it: the checks below are Resume's own, and **none of them -requires a clean tree** — three of the four valid topologies need not have one. +requires a clean tree** — every topology below can legitimately lack one. **No count of them is +stated here**: one was, and adding the moved-`HEAD` row below falsified it. -**Four topologies, named rather than numbered, and the whole procedure reads against all four.** +**The topologies, named rather than numbered, and the whole procedure reads against all of them.** | Topology | `HEAD` | `$BASE..HEAD` | Where the work is | |---|---|---|---| | **Normal, mid-implementation** | the last `WIP:` snapshot | this run's `WIP:` commits, and possibly a §A3 stray commit or amend | committed | | **8a rejected, no commit landed** | the last `WIP:` snapshot, unchanged | this run's `WIP:` commits | committed, **plus whatever the failed attempt left in the index or worktree** | | **8a rejected after its commit landed** | a `WIP:` findings commit **above** the cycle's `WIP:` chain | those commits | committed, **plus any delta the post-commit clean-tree check found** | +| **A closure condition rejected because `HEAD` moved** — no reset has run | a commit **above** the value the condition expected: above the **reviewed head** at condition 1, or above **8a's findings commit** at condition 5 | this run's commits **plus that extra commit** | committed, **plus whatever is loose** — and the extra commit's content is **present but unreviewed** | | **8b rejected** | either `$BASE` itself, if the closing commit never landed, **or one commit parented by `$BASE`**, if it landed and a postcondition refused it | **empty**, or that one commit — the `WIP:` chain is gone either way, squashed by `reset --soft` | the **index**, or that one commit's tree | **The 8b row is the one every rule has to be re-read against.** `reset --soft` removes the `WIP:` @@ -515,6 +517,19 @@ the work is staged, not absent. **And the two 8a rows differ from each other**: the history is untouched and the delta is loose, after it the findings commit sits above the chain — pass 31 split them because they need different content reconciliation. +**The moved-`HEAD` row is the one that reads as ordinary progress and is not.** Its shape — a commit +above the chain — *is* the normal row's shape, which is why pass 36 found the state described by no +row at all and classifiable only as normal. **What separates them is why the commit is there**: +condition 1 or condition 5 rejected *on it*, so it arrived after the head the candidate pass was +issued against and **no pass has seen it**. Row 1's "possibly a §A3 stray commit" is the same object +without that signal. + +**The two rejections that reach this row differ in one way the rules below turn on.** Condition 1 is +a **precondition**, so nothing this plan does has moved: that is a plain stop, and no handoff is in +progress. **Condition 5 rejects after 8a's record commit has landed**, so the Failure handoff *is* +in progress — which is what selects the `inside a handoff` branch of the scratch-artifact rule +below, report and change nothing, rather than delete and rebuild. + - [ ] **Validate the base — Resume's own checks, not Preparation's** ```bash @@ -540,7 +555,10 @@ derived against does not. **Replace a base only on affirmative evidence it belongs to another run** — it fails one of the checks above. **Neither a commit subject nor an absence of cycle commits is evidence**: §A3's stray non-`WIP` commit leaves the base valid, and the 8b topology has an empty range by construction, so -both tests would condemn states this plan calls valid. +both tests would condemn states this plan calls valid. **Nor is an extra commit above the reviewed +head**: it sits *above* the base, so ancestry still holds and the approved inputs at `$BASE` are +untouched — the moved-`HEAD` topology leaves the base valid, and what that commit costs is Close +condition 1's to decide, not the base checks'. - [ ] **Validate the scratch artifacts** @@ -561,10 +579,10 @@ a value is trustworthy because it passed a check, never because cleanup was skip - [ ] **Establish how far the implementation got — against the topology, not against the log** -**Read the content, not only the commits.** In the normal and the two 8a topologies that is the -commits between `$BASE` and `HEAD`, **plus whatever the failed attempt left loose**. **In the 8b -topology it is `git diff --cached "$BASE"`** — the staged tree, plus the landed closing commit's -tree where one exists. A ticked checkbox is confirmed by the change +**Read the content, not only the commits.** In the normal, the two 8a and the moved-`HEAD` +topologies that is the commits between `$BASE` and `HEAD`, **plus whatever the failed attempt left +loose**. **In the 8b topology it is `git diff --cached "$BASE"`** — the staged tree, plus the landed +closing commit's tree where one exists. A ticked checkbox is confirmed by the change being *present in that content*, wherever the content lives. ```bash @@ -583,6 +601,17 @@ path listing cannot describe** — a rewritten staged file keeps its name — so whose change is present with the box unticked is completed. **Deciding from the commit log alone reads the 8b topology as an untouched cycle and invites every edit to be made twice.** +**Present is not reviewed, and the moved-`HEAD` topology is where the two come apart.** Reconciling +the task list tells you what the repository *holds*; it says nothing about what a pass has *seen*. +Content in a commit above the reviewed head arrived after the candidate pass was issued, so **no +pass has reviewed it**, and a ticked checkbox does not make it reviewed. **Do not adopt it** — +carrying it into a close is the defect Close condition 1 exists to stop, and that condition's own +rule is what decides the cost: only a clean response issued against that exact `HEAD` closes the +cycle. **And do not remove it** — this plan restores nothing, and a commit deleted here is evidence +a person has not yet chosen a §A route on. **Report it as present and unreviewed**, name it in the +Failure report's "whether a commit landed" line, and leave the three routes in `## The four +procedures` · Failure to decide. + **How success is recognised.** The base passes Resume's own checks; the artifacts match it and are complete, or are reported as invalid and left alone; and the plan's task list has been reconciled against the content the topology actually holds. @@ -1170,7 +1199,7 @@ and 11 also record — and a stale list is the same defect as a stale count. **T omission is concrete:** the evidence stays out of its own task's `WIP:` snapshot, so the reviewed range does not hold it where the task claims, and if execution continues it is swept into a later unrelated commit rather than the independently reviewable snapshot this plan promises. **Not a -clean-tree argument** — Resume requires no clean tree, and three of its four topologies do not have +clean-tree argument** — Resume requires no clean tree, and none of its topologies is required to have one; an earlier draft said re-entry needs one, which would have rejected the very states Resume exists to reconcile. @@ -1425,7 +1454,7 @@ while the plateau rationale beside it carries no inventory id, **stays**, and is own name. Calling `c9` "split" names a seventh disposition and invites an executor to attach the staying rationale to a condition required to vanish. So: the moved precedence clause is absent here (`parent=1 worktree=0`) and present in §A, while the plateau rationale stays — confirm the -rationale still counts `1` in each copy. +rationale still counts `parent=1 worktree=1` in each copy. *(Build each pair per the verification procedure; record the four values.)* @@ -1434,7 +1463,18 @@ Expected: for the three pairs, six pair instances reading `parent=1 worktree=0` in each copy. **Every OLD row was derived and validated at step 1**; this step only runs them. -- [ ] **Step 5: Walk the conditions** — `c1`–`c3` present unchanged, `c5`–`c7` word for word, `c8` carrying the new clause, `c9` moved out and the plateau rationale still here, `c10`–`c14` gone from here. **This is a reader's confirmation on top of step 4's counts, not the observation for any of them** — `c9`–`c14` are each counted there, and a walk that found what the counts missed would mean a fragment was wrong rather than that the walk was the check. +- [ ] **Step 5: Walk the conditions** — **every disposition row this task's block covers, read off + the table the way step 4b already says to**, not off a list here. The list this step used to carry + omitted `c4`, which has a pair at step 4 and was therefore counted and never read (pass 36); a + closed list in a task is the second copy of the disposition table, and this is the second one it + has cost this cycle. **This is a reader's confirmation on top of step 4's counts, not the + observation for any of them** — each of these is counted there, and a walk that found what the + counts missed would mean a fragment was wrong rather than that the walk was the check. + +For orientation, and not as the set owed: `c1`–`c3` present unchanged; **`c4` carrying the narrowed +reading** — a missing condition now means only that *this* exit does not apply; `c5`–`c7` word for +word; `c8` carrying the re-raised-dismissal clause; `c9` moved out with the plateau rationale still +here; `c10`–`c14` gone from here. - [ ] **Step 6: Commit** @@ -1534,7 +1574,7 @@ the class alone does not tell you: - **`e10` is kept and belongs with `e1`–`e6`, not with `e9`** — it is the sentence *after* §D's block, so a check treating it as carried would look for it inside text it never enters. -- **`e9` is carried** — its count is `1` in each copy after the install, from the single-line +- **`e9` is carried** — its count is `parent=1 worktree=1` in each copy after the install, from the single-line fragment step 1 appended rather than from the whole wrapped clause. - **`e8` is aligned rather than untouched** — this task gives W the pronoun, so listing `e8` among the untouched conditions would contradict the task's own instruction. @@ -1780,7 +1820,8 @@ it, and record all of them in this task's fragment evidence. They are reproduced inside blocks that install contiguously, so a mis-scoped replacement silently drops them — and because they are carried rather than kept, **no untouched-range span covers them**, which is exactly why the disposition records them as carried. **Each owes a count of its -own text in each copy, expecting `1`.** +own text in each copy, to `parent=1 worktree=1`** — the result the disposition table owes a +carried condition. **The set is every condition the disposition table marks *carried* inside this task's ten blocks**, and it is read from there rather than listed here. An earlier draft named three of them and left @@ -1796,7 +1837,7 @@ grep -cF 'Codex is advisory — validate before applying; dismissed finding → grep -cF 'Open a TodoWrite' CLAUDE.md ``` -Expected: `1` each, and `1` in each copy for every one of the nine. +Expected: `1` each in the worktree, and `parent=1 worktree=1` in each copy for every one of the nine. - [ ] **Step 5: Parity** for all ten sites, each extracted by its own bounded region rather than a fixed line window. From b231fc2023b7d7b2a1c2f35de04207b096cc5d0e Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Wed, 16 Sep 2026 10:09:04 +0200 Subject: [PATCH 154/181] docs(context): record Gate-A plan pass 37 and the fourth bounded repair Pass 37 reviewed d7f2af8: 1 Major, 1 Minor, 1 Nit, 0 Blockers. The open Major descends from pass 36's fix. The moved-HEAD topology row was keyed on the shape of the state - a commit above the expected head - where Close conditions 1 and 5 both test plain object-id inequality. An amend, a rebase or a reset to an earlier commit still above $BASE trips the same guard and matches no row, so a rejected, rewritten, unreviewed state can still read as normal progress. The row has to be defined from the predicate. The Minor was created by the same repair: accounting row 7 says Resume's table defines "four shapes" and it now has five rows. The stale-count sweep keyed on the word "topologies"; this site says "shapes". The Nit is pre-existing - Self-Review's reader-check sentence omits Task 12b and Task 15 step 4b. Both collected per the Minor/Nit rule. The working record now carries the regeneration signature plainly: findings 6, 1, 2, 3 and Majors exactly 1 for four consecutive passes, with one established lineage. That is not the clearly-stuck exit, which needs a plateau across roughly six passes and an affirmative coverage judgement as well, but it is the shape and the record says so rather than leaving it to be noticed. Also recorded as a per-pass lesson: grep for the claim, not the word. Two sweeps in two rounds each missed a site that used a synonym. Floor 3, from docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md (risk high, security none), read fresh at this pass. Cycle stays open and unclean. --- .../gate-a-plan-om0bdd7udh-pass-37.md | 4 + .../gate-a-plan-om0bdd7udh-resume.md | 75 ++++++++++++------- 2 files changed, 50 insertions(+), 29 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-37.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-37.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-37.md new file mode 100644 index 0000000..cffe7c3 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-37.md @@ -0,0 +1,4 @@ +MAJOR | high | Resume topology table and progress reconciliation | The moved-HEAD row assumes every Close condition-1 or condition-5 mismatch is a new descendant commit above the expected head, but both guards test only object-id inequality; an amend, rebase, or reset to an earlier commit that remains above $BASE also trips the guard while matching neither the row's HEAD/range description nor its "extra commit" reconciliation. | Resume can classify a rejected, unreviewed rewritten state as normal progress or apply the wrong handoff and scratch-artifact branch, so the claimed exhaustive re-entry procedure has no defined route for a reachable state. | Define the topology from the actual HEAD != expected predicate, split it by ancestry/content location where the operation differs, and make the progress and handoff rules cover every resulting relation rather than only an added descendant commit. +MINOR | high | The accounting, row 7 | Row 7 still says Resume's topology table defines "four shapes", but the table now has five rows after the moved-HEAD addition, contradicting the later no-count policy and the pass-37 repair note. | The accounting retains a stale numeric authority, so an audit can treat one reachable row as surplus or omit it when checking state-reading rules. | Remove the count and cite Resume's topology table without restating its size. +NIT | high | Self-Review section 2 | The sentence classifying reader checks names only Tasks 12, 13, and 14 step 2, omitting Task 12b's reader-led completeness sweep and Task 15 step 4b's deliberately reader-only prompt-standards review. | The self-review gives an incomplete account of which verification duties require human judgement, even though the task bodies classify both omitted checks correctly. | Add Task 12b and Task 15 step 4b to the sentence, or replace the enumeration with references to each task's own classification. +END OF FINDINGS (3 total) diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md index f70a997..81e99d9 100644 --- a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md @@ -13,38 +13,47 @@ Nothing depends on it; the pass files and the repo are authoritative where this ## Resume here -**Cycle OPEN and UNCLEAN. The plan stands at `518121a` and that is the anchor** — later commits on +**Cycle OPEN and UNCLEAN. The plan stands at `d7f2af8` and that is the anchor** — later commits on this branch are records, not plan edits, so check `git log --oneline -1 -- docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md` -rather than `HEAD`. Pass 36 reviewed `518121a` and found **one Major and one Minor, 0 Blockers**. -`.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-36.md` holds them. +rather than `HEAD`. Pass 37 reviewed `d7f2af8` and found **one Major, one Minor and one Nit, +0 Blockers**. `.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-37.md` holds them. **Do not start a repair round on your own.** Daniel's assignment of 2026-09-16 ended with a -checkpoint: *"After the repair and focused verification, run exactly one complete Gate-A plan -review, then stop and report. Do not automatically begin another repair round or implementation."* +checkpoint: *"Run exactly one complete Gate-A plan pass after the bounded revision and verification, +then stop and report regardless of outcome. Do not begin another repair round or implementation."* That checkpoint is spent — it covered one bounded repair and one pass, both delivered — so **the next move is Daniel's word, not an inference.** -**Two tells present, so this stop is mandatory as well as instructed**: findings rose 1 → 2, and the -findings cluster on the instrument for the third pass running. A third reading is arguable — the -Blocker count "failed to fall" at 0 → 0, which is the literal test and nonsense at zero. **That -reading is not settled here and nothing turns on it**: the stop happens either way. - -### The two open findings from pass 36 - -1. **MAJOR — Resume's topology table is not exhaustive, and claims to be.** It names four - topologies and says *"the whole procedure reads against all four"*; the working record calls them - *"the model for every state-reading rule"*. But Close condition 5 exists precisely to catch **a - commit landing between 8a and 8b**, and in that rejected state `HEAD` sits **above** the 8a - findings commit — which is not the last `WIP:` snapshot, not the findings commit itself, and not - `$BASE` or its child after the reset. **No row describes it.** The condition-1 rejection has a - milder version of the same problem: a stray commit above the reviewed head reads as row 1, - "Normal, mid-implementation", which is how an unreviewed commit gets treated as ordinary cycle - progress. **Validated against the plan, not taken on trust.** -2. **MINOR — Task 4 step 5's reader walk omits `c4`.** It walks `c1`–`c3`, `c5`–`c7`, `c8`, `c9`, - `c10`–`c14`; passage (c) carries a `c4` row and `c4` is **replaced**. The mechanical pair still - covers it; the semantic confirmation never reads it. **Same defect family as pass 35's Major, in - a different task** — a task enumerating its own conditions and dropping one. Collected per the - Minor rule, not repaired. +**Two tells present, so this stop is mandatory as well as instructed**: findings rose 1 → 2 → 3, +three passes running, and they cluster on the instrument. The Blocker count "failed to fall" at +0 → 0 is the literal third test and nonsense at zero; **that reading is still not settled here and +nothing turns on it.** + +**Read this before deciding the next round — the regeneration signature is now visible.** Findings +6 → 1 → 2 → 3; Blockers 1 → 0 → 0 → 0; **Majors exactly 1 in each of the last four passes**. And +pass 37's Major is **descended from pass 36's fix**: the moved-`HEAD` row was keyed on the *shape* +of the state ("a commit above the expected head") where both guards test **object-id inequality**, +so a rewrite that trips the same guard matches no row. That is one established lineage, a genuine +repair producing the next finding. **It is not yet the "clearly stuck" exit**, which needs three +things together — a plateau across roughly six passes, an affirmative coverage judgement, and +Blocker/Majors regenerating across repairs. We have the third and a quarter of the first. + +### The three open findings from pass 37 + +1. **MAJOR — the moved-`HEAD` topology covers one shape of several.** Conditions 1 and 5 are plain + inequalities (`test "$(git rev-parse HEAD)" = "$HEADREV"`, and the same against `$TIP`). The new + row describes only **a descendant commit above** the expected head. An **amend**, a **rebase**, + or a **reset to an earlier commit still above `$BASE`** trips the identical guard and matches + neither the row's `HEAD`/range description nor its "extra commit" reconciliation — so a rejected, + rewritten, unreviewed state can still read as normal progress. **The row has to be defined from + the predicate, not from a shape**, and split where the operation differs. Validated against the + plan. +2. **MINOR — accounting row 7 still says Resume's table defines "four shapes".** It now has five + rows. **Created by the previous repair**: the sweep for stale counts keyed on the word + *topologies* and this site says *shapes*. One-word fix; collected per the Minor rule. +3. **NIT — Self-Review's reader-check sentence names Tasks 12, 13 and 14 step 2**, omitting Task + 12b's reader-led sweep and Task 15 step 4b's prompt-standards review. The gate prompt's own + settled block lists all five. Pre-existing, not caused by this round. Collected. ### Still collected and deliberately unrepaired (pass 34's Minors and Nit) @@ -135,9 +144,15 @@ what it carries; `.context/gate-a-spec-prompt.md` is the same shape for the spec through the installed ordering before any next call exists. - **A task that enumerates its own conditions is the second copy of the disposition table**, and it goes stale silently. Task 8's list dropped carried `h3` (pass 35, MAJOR); Task 4 step 5's reader - walk drops `c4` (pass 36, MINOR). **Grep every task for id lists and check each against its - passage's rows** — it has now produced a finding in two consecutive passes, and the rule against - it is already written in `## How a task discharges that table`. + walk dropped `c4` (pass 36, MINOR). Both repaired. **Check each remaining task list against its + passage's rows — but a matching id is a search hit, not a defect.** Tasks 3, 5 and 7 already read + their sets off the table; Task 7's lists and Task 4 step 4's move list are complete and were left + alone deliberately. Replacing them blind would remove requirements. +- **When you state a rule in one place, grep for the CLAIM, not the word you used.** Two sweeps in + two rounds each reached most sites and missed the rest, both times because the missed site used a + synonym: `1` instead of `preservation count`, `four shapes` instead of `four topologies`. This is + `AGENTS.md`'s "search for the claim" rule, and the cycle keeps re-learning it. **A count anywhere + near a table you changed is the first thing to check.** **The most reliable defect in this cycle: a repair reaching one site of several.** Passes 26–33 each found the previous pass's fix applied in one place and missing in two. **When a rule changes, grep @@ -200,6 +215,8 @@ worth checking before a pass rather than after. | 35 | 2d79ac8 | 6→**1** | 1→**0** | 1→**1** | yes | **UNCLEAN — reported to Daniel, no repair round opened.** One Major: carried `h3` owes a preservation count and passage (h) says Task 8 derives its fragment, but Task 8 owns only `h4`/`h5`/`h19` and derives only `a1`, so item 7 could drop `an ungated change records it in that commit` from both copies undetected. **Both pass-34 repairs held** — nothing re-raised against 8a, the dirty set or condition 6. One tell (instrument cluster); no mandatory stop | | — | — | — | — | — | — | **THIRD BOUNDED REPAIR**, Daniel's assignment after pass 35, scoped to that Major alone. **The fix was to stop enumerating, not to extend the enumeration**: Task 8's top line cited `## How a task discharges that table` instead of naming `h4`/`h5`/`h19`, step 1 derives the preservation fragment of every carried condition its blocks cover, step 4 counts each to `parent=1 worktree=1` and records a `preservation` line. Two directly affected references pinned to the same result — the disposition table's carried row and Task 7's step, both of which said `1` alone. Verified on temporary copies: the `h3` count passes on the intended install and fails on `h3` removed from C alone, W alone and both, and F7's OLD reaches 0 either way. Commit `518121a` | | 36 | 518121a | 1→**2** | 0→**0** | 1→**1** | yes | **UNCLEAN — reported to Daniel, no repair round opened.** **The h3 repair held**; nothing re-raised against Task 8. MAJOR: Resume's topology table claims four exhaustive topologies and has **no row for the state condition 5 exists to catch** — `HEAD` above the 8a findings commit after a between-invocations commit; the condition-1 rejection reads as row 1, "normal". MINOR: Task 4 step 5's reader walk omits `c4` — the **same family** as pass 35's Major, a task enumerating its own conditions. Two tells (findings rose, instrument cluster) → mandatory stop, which the checkpoint already required | +| — | — | — | — | — | — | **FOURTH BOUNDED REPAIR**, Daniel's assignment after pass 36, on his reviewer's narrowed recommendation: repair the Major and Task 4's `c4`, plus a **bounded** search for the same coverage-list defect — **not** a blanket replacement of every id list, because a grep finds ids and cannot tell a constraining list from an orientation note. Resume gains one moved-`HEAD` row; the progress reconciliation gains *present is not reviewed*; Task 4 step 5 walks the table. Search result: Tasks 3, 5, 7 already cite the table, Task 7's and Task 4 step 4's lists checked complete and left alone. Four bare-`1` results and two stale counts corrected as disclosed misses of the previous sweep. Commit `d7f2af8` | +| 37 | d7f2af8 | 2→**3** | 0→**0** | 1→**1** | yes | **UNCLEAN — reported to Daniel, no repair round opened.** MAJOR: the new moved-`HEAD` row was keyed on the **shape** of the state where both guards test **object-id inequality**, so an amend, rebase or reset still above `$BASE` matches no row — **pass 36's fix produced it**, one established regeneration lineage. MINOR: accounting row 7 still says "four shapes", falsified by the fifth row — the stale-count sweep keyed on *topologies*, this site says *shapes*. NIT: Self-Review's reader-check sentence omits Task 12b and Task 15 step 4b. Two tells (findings rose a third pass, instrument cluster) → mandatory stop | | 26 | 05ec1b2 | **5** | **1** | 2 | yes | first pass on the revised plan. Resume audited progress by `WIP:` commits and called any non-`WIP` commit a stale base — the handoff leaves two other shapes, one of them `HEAD` **at** the base with everything in the index | | 27 | f69db7d | 5→**6** | 1→**3** | 2→**1** | yes | the close required the commit's **tree to equal the tip's**, which target §I **parks**; `reset --soft` leaves the index untouched, so the prose beside it was wrong about git; Resume still inferred staleness from a commit subject where §A3 names that state as reachable | | 28 | f2e5dda | 6→**6** | 3→**2** | 1→**2** | yes | pass 27's stale-base fix reached one site of three; its reset paragraph stated the index behaviour correctly and repeated the false claim four lines later; nothing checked the closing commit's **parent**; the message was validated in the file, where `commit-msg` hooks rewrite git's copy after `-F` reads it | From b283250c77fe1e4184099d2b26c5ab6e141df915 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Wed, 16 Sep 2026 11:26:55 +0200 Subject: [PATCH 155/181] docs(plans): Resume reads the repository; it stops classifying it Pass 37's Major, repaired by dropping the demand rather than by describing more states. Daniel's decision. Conditions 1 and 5 test object-id equality. The topology row added after pass 36 described only a further commit above the expected head, so an amend, a rebase or a move backwards tripped the same guard and matched no row. The previous two repairs both tried to describe the state space more precisely and each produced the next finding. The demand was this plan's own. Approved section A says to re-establish every closure condition against the repository as it now stands - it reads the conditions, it never asks for the history to be catalogued. So Resume no longer has to fit the repository to a named shape before it may proceed. The table stays as explicitly non-exhaustive illustration, and no permission, no base-validity conclusion, no recovery operation and no continuation route depends on a state matching a row; a state that fits none is an ordinary input, not an error. What is preserved: both equality checks, the precondition-stop versus Failure-handoff distinction, the scratch-artifact rules and their inside-a-handoff branch, the bounded handoff, no cleanup after failure, base validity checked rather than inferred, and section A's three routes. What is dropped is named in accounting row 7, and named as a requirement rather than as one of the forty-one conditions, so row 40's count of dropped conditions stays true. Two rules strengthened where the old text had leaned on shape. The progress reconciliation reads all four sources every time instead of routing on a guessed topology - routing on a guess is how a source gets skipped. And present-is-not-reviewed now covers any mismatch in any direction, with one addition: do not rewrite the recorded head or tip to make the equality pass, which would close on an unreviewed tree by discarding the only record of what was reviewed. Verified by execution under sh, dash and bash, 64 checks each: descendant, amended, rebased and backward mismatches, each before 8a and after 8a. Every one is rejected before any reset, the plan moves nothing, the recorded head and tip are not overwritten, the four reconciliation sources still read, the closing message is intact, and base validity is checked and holds. Representative mismatches only - no claim of exhaustive history coverage, which is the requirement this revision drops. Every bolded duty in the old Resume section was accounted for against the new one. --- .../2026-09-14-loop-rule-consolidation.md | 141 ++++++++++-------- 1 file changed, 80 insertions(+), 61 deletions(-) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index e74f0f9..9a4d95c 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -211,7 +211,7 @@ live in the home named. Where a task, a shell comment or an error string needs a |---|---|---| | Branch, clean tree, no base file yet, `ba15e83` ancestral, approved inputs unchanged at `HEAD` | **Preparation** | Task 0 step 1, first-entry branch | | Base file shape, self-resolving commit, ancestral; approved inputs unchanged at `$BASE`; scratch-artifact validity; how far the implementation got | **Resume** | Task 0 step 1, re-entry branch | -| The repository topologies a re-entry can meet | **Resume**, its table | accounting row 7; every state-reading rule | +| What a re-entry reads, and that no rule depends on the state fitting a named shape | **Resume** — its table is illustration, not a classification | accounting row 7; every state-reading rule | | `HEAD` equals the reviewed head | **Close**, condition 1 | step 8a | | The dirty set is exactly the candidate pass's findings files | **Close**, condition 2 | step 8a | | The closing message is complete, and revalidated before it is consumed | **Close**, condition 3 | step 7b writes it; steps 8a and 8b check it | @@ -242,7 +242,7 @@ independent reader did. | 4 | The three approved inputs' blobs equal their `ba15e83` versions | **kept, and split by entry**: Preparation compares them **at `HEAD`**, because a first entry has no recorded base; **Resume compares them at `$BASE`**, because a Gate-B fix may legitimately have changed `HEAD`'s copy in a `WIP:` snapshot. An earlier table row claimed Preparation did the base comparison, which it never could | | 5 | Never overwrite an existing base file | **kept** as an obligation; the shell branch becomes one line of the resume procedure | | 6 | A pre-existing base is an ancestor of `HEAD` (`merge-base --is-ancestor`) | **kept**, same shell — resume | -| 7 | Only this run's `WIP:` commits lie between base and `HEAD` | **kept as an observation, dropped as a staleness test.** The model is **Resume's topology table**, which is where the four shapes are defined — re-listing them here is the third copy this revision removes. **None of them makes the base stale**, and staleness is decided by ancestry and provenance, never by a commit subject | +| 7 | Only this run's `WIP:` commits lie between base and `HEAD` | **kept as an observation, dropped as a staleness test**, and **a second requirement is dropped here by decision**: Resume no longer has to fit the repository to a named shape before it may proceed. That demand was this plan's own — the approved §A re-establishes the conditions *against the repository as it now stands* — and it failed twice in two passes, once on a state no row described and once on a row describing one shape of several. Resume's table stays as **illustration, deliberately not exhaustive**, and no permission, base conclusion or recovery operation turns on it. **This drop is a requirement, not one of the forty-one conditions** — rows 20 and 40 are still the only *conditions* this table deletes, and row 40's count is about those. **Staleness is still decided by ancestry and provenance, never by a commit subject and never by an unfamiliar shape** | | 8 | `$BASE` persisted to a file; every consumer guards it non-empty | **kept**, and **repaired**: three concrete blocks read it without the guard while this row claimed otherwise — the baseline extraction, Task 10's hook diff, and the first Gate-B call. The guard is in each of them now (pass 22) | | 9 | The five regions are located by anchor, one hit per pattern per file | **kept**, same shell — preparation | | 10 | Spans are **derived** from replacement extents, never hand-written | **kept** — the rule, unchanged | @@ -327,7 +327,7 @@ of them touched this guard. ### Preparation — before any task edits a file **This procedure is for a FIRST entry, on a clean tree.** A re-entry runs **Resume** instead, which -requires no clean tree and mutates nothing — the handoff's valid topologies include `HEAD` at the +requires no clean tree and mutates nothing — a handoff can leave `HEAD` at the base with the whole implementation **staged**, and an unconditional clean-tree test would make that state permanently unresumable through this plan's own success path. @@ -498,37 +498,53 @@ and records the closing commit carries are the ones a pass actually validated. **Resume owns every entry after the first.** Preparation is first-entry-only and refuses when a base file exists, so nothing here defers to it: the checks below are Resume's own, and **none of them -requires a clean tree** — every topology below can legitimately lack one. **No count of them is -stated here**: one was, and adding the moved-`HEAD` row below falsified it. - -**The topologies, named rather than numbered, and the whole procedure reads against all of them.** - -| Topology | `HEAD` | `$BASE..HEAD` | Where the work is | +requires a clean tree** — a re-entry can legitimately arrive with the work staged, with it loose, or +with both, and an unconditional clean-tree test would refuse the very states Resume exists to +reconcile. + +**Resume reads the repository; it does not classify it first.** Every check below runs against the +**actual** `HEAD`, index, worktree and surviving cycle records, and **none of them is gated on the +state matching a named shape.** The approved §A is what this follows — *"re-establish every closure +condition against the repository as it now stands"* — and it asks for the conditions to be read, +never for the history to be catalogued. **The table below is illustration and is explicitly not +exhaustive**: no permission, no base-validity conclusion, no recovery operation and no continuation +route may depend on the state fitting a row, and **a state that fits none is an ordinary input to +Resume, not an error**. + +**This is a requirement being dropped, and it is named rather than lost.** Earlier revisions said +the whole procedure reads *against the rows*, which made a complete classification of git histories +a precondition for resuming at all. **That was this plan's invention, never §A's**, and it failed +twice in two passes in the same way: pass 36 found a reachable state no row described, and pass 37 +found the row added for it describing only one shape of the several that reach it. **What survives +unchanged is every governing duty** — the conditions, their object-id equality checks, the +precondition/handoff distinction, the scratch rules, the bounded handoff, and §A's three routes. + +| Illustration, not a classification | `HEAD` | `$BASE..HEAD` | Where the work is | |---|---|---|---| | **Normal, mid-implementation** | the last `WIP:` snapshot | this run's `WIP:` commits, and possibly a §A3 stray commit or amend | committed | | **8a rejected, no commit landed** | the last `WIP:` snapshot, unchanged | this run's `WIP:` commits | committed, **plus whatever the failed attempt left in the index or worktree** | -| **8a rejected after its commit landed** | a `WIP:` findings commit **above** the cycle's `WIP:` chain | those commits | committed, **plus any delta the post-commit clean-tree check found** | -| **A closure condition rejected because `HEAD` moved** — no reset has run | a commit **above** the value the condition expected: above the **reviewed head** at condition 1, or above **8a's findings commit** at condition 5 | this run's commits **plus that extra commit** | committed, **plus whatever is loose** — and the extra commit's content is **present but unreviewed** | +| **8a rejected after its commit landed** | a `WIP:` findings commit above the cycle's `WIP:` chain | those commits | committed, **plus any delta the post-commit clean-tree check found** | +| **A closure condition rejected on `HEAD`** — no reset has run | **not the object id the condition expected** — the reviewed head at condition 1, 8a's findings commit at condition 5. **The direction is not part of the test**: a further commit, an amend, a rebase and a move backwards all fail the same equality | whatever the actual history holds | wherever the actual content is — and **anything the mismatch introduced is present but unreviewed** | | **8b rejected** | either `$BASE` itself, if the closing commit never landed, **or one commit parented by `$BASE`**, if it landed and a postcondition refused it | **empty**, or that one commit — the `WIP:` chain is gone either way, squashed by `reset --soft` | the **index**, or that one commit's tree | -**The 8b row is the one every rule has to be re-read against.** `reset --soft` removes the `WIP:` -chain from the ancestry, so after 8b there is no chain to find; an empty `$BASE..HEAD` there means -the work is staged, not absent. **And the two 8a rows differ from each other**: before its commit -the history is untouched and the delta is loose, after it the findings commit sits above the chain -— pass 31 split them because they need different content reconciliation. - -**The moved-`HEAD` row is the one that reads as ordinary progress and is not.** Its shape — a commit -above the chain — *is* the normal row's shape, which is why pass 36 found the state described by no -row at all and classifiable only as normal. **What separates them is why the commit is there**: -condition 1 or condition 5 rejected *on it*, so it arrived after the head the candidate pass was -issued against and **no pass has seen it**. Row 1's "possibly a §A3 stray commit" is the same object -without that signal. - -**The two rejections that reach this row differ in one way the rules below turn on.** Condition 1 is -a **precondition**, so nothing this plan does has moved: that is a plain stop, and no handoff is in -progress. **Condition 5 rejects after 8a's record commit has landed**, so the Failure handoff *is* -in progress — which is what selects the `inside a handoff` branch of the scratch-artifact rule -below, report and change nothing, rather than delete and rebuild. +**The 8b illustration is the one worth reading before any rule below.** `reset --soft` removes the +`WIP:` chain from the ancestry, so after 8b there is no chain to find; an empty `$BASE..HEAD` there +means the work is staged, not absent. **And the two 8a illustrations differ from each other**: +before its commit the history is untouched and the delta is loose, after it the findings commit sits +above the chain — pass 31 separated them because they need different content reconciliation. + +**Conditions 1 and 5 are equality tests and stay that way.** `test "$(git rev-parse HEAD)" = +"$HEADREV"` and the same against `$TIP` ask one question — is `HEAD` the object the cycle recorded — +and **any answer of no stops the closing sequence**, whatever git operation produced it. **Do not +rewrite the recorded value to make the test pass**: those files are what make "the head the +candidate pass was issued against" a fact rather than a claim, and editing one turns an unreviewed +tree into a closable one with nothing left to notice. + +**The two rejections differ in one way the rules below turn on.** Condition 1 is a **precondition**, +so nothing this plan does has moved: that is a plain stop, and no handoff is in progress. **Condition +5 rejects after 8a's record commit has landed**, so the Failure handoff *is* in progress — which is +what selects the `inside a handoff` branch of the scratch-artifact rule below, report and change +nothing, rather than delete and rebuild. - [ ] **Validate the base — Resume's own checks, not Preparation's** @@ -554,11 +570,13 @@ derived against does not. **Replace a base only on affirmative evidence it belongs to another run** — it fails one of the checks above. **Neither a commit subject nor an absence of cycle commits is evidence**: §A3's stray -non-`WIP` commit leaves the base valid, and the 8b topology has an empty range by construction, so -both tests would condemn states this plan calls valid. **Nor is an extra commit above the reviewed -head**: it sits *above* the base, so ancestry still holds and the approved inputs at `$BASE` are -untouched — the moved-`HEAD` topology leaves the base valid, and what that commit costs is Close -condition 1's to decide, not the base checks'. +non-`WIP` commit leaves the base valid, and after a `reset --soft` the range is empty by +construction, so both tests would condemn states this plan calls valid. **Nor is a `HEAD` mismatch, in any direction**: +base validity is **established by the checks above and never inferred** — from a mismatch, from a +shape, or from the state not resembling anything named here. Run them and read the answer. A +mismatch costs whatever Close condition 1 says it costs, which is not the base checks' question, and +where the base does fail one of them the rule is the one already stated: **affirmative evidence, then +replace — never because a re-entry looked unfamiliar.** - [ ] **Validate the scratch artifacts** @@ -577,44 +595,46 @@ chose a §A route. **This plan performs no cleanup after a failure — it does n operation left anything intact.** Enumerate which scratch values actually survive and validate each; a value is trustworthy because it passed a check, never because cleanup was skipped. -- [ ] **Establish how far the implementation got — against the topology, not against the log** +- [ ] **Establish how far the implementation got — from the content, not from the log** -**Read the content, not only the commits.** In the normal, the two 8a and the moved-`HEAD` -topologies that is the commits between `$BASE` and `HEAD`, **plus whatever the failed attempt left -loose**. **In the 8b topology it is `git diff --cached "$BASE"`** — the staged tree, plus the landed -closing commit's tree where one exists. A ticked checkbox is confirmed by the change -being *present in that content*, wherever the content lives. +**Read all four sources, every time, and do not decide first which of them matters.** Which one +holds the work varies — after `reset --soft` it is the index and `$BASE..HEAD` is empty; after an 8a +handoff part of it is loose — and **routing on a guessed shape is how a source gets skipped**. So +read them all and reconcile against what they actually contain: ```bash -git log --oneline "$BASE"..HEAD # empty in the 8b topology; that is not an empty cycle +git log --oneline "$BASE"..HEAD # can legitimately be empty; that is not an empty cycle git diff --cached "$BASE" # the staged content — the FULL diff, not --stat git diff # the unstaged content git status --porcelain --untracked-files=all ``` -**Read all three, in every topology.** `git status` names paths and says nothing about what is in -them, and `--stat` counts lines. **The delta that caused an 8a handoff is precisely the one a -path listing cannot describe** — a rewritten staged file keeps its name — so reconcile against -`HEAD`, the index and the worktree **contents** together, whichever topology you are in. +A ticked checkbox is confirmed by the change being *present in that content*, wherever it lives. +`git status` names paths and says nothing about what is in them, and `--stat` counts lines. **The +delta that caused an 8a handoff is precisely the one a path listing cannot describe** — a rewritten +staged file keeps its name — so reconcile against `HEAD`, the index and the worktree **contents** +together. -**A task whose checkbox is ticked but whose change is in neither place was not completed** — and one +**A task whose checkbox is ticked but whose change is in none of them was not completed** — and one whose change is present with the box unticked is completed. **Deciding from the commit log alone -reads the 8b topology as an untouched cycle and invites every edit to be made twice.** - -**Present is not reviewed, and the moved-`HEAD` topology is where the two come apart.** Reconciling -the task list tells you what the repository *holds*; it says nothing about what a pass has *seen*. -Content in a commit above the reviewed head arrived after the candidate pass was issued, so **no -pass has reviewed it**, and a ticked checkbox does not make it reviewed. **Do not adopt it** — -carrying it into a close is the defect Close condition 1 exists to stop, and that condition's own -rule is what decides the cost: only a clean response issued against that exact `HEAD` closes the -cycle. **And do not remove it** — this plan restores nothing, and a commit deleted here is evidence -a person has not yet chosen a §A route on. **Report it as present and unreviewed**, name it in the -Failure report's "whether a commit landed" line, and leave the three routes in `## The four -procedures` · Failure to decide. +reads a post-reset repository as an untouched cycle and invites every edit to be made twice.** + +**Present is not reviewed, and a `HEAD` mismatch is where the two come apart.** Reconciling the task +list tells you what the repository *holds*; it says nothing about what a pass has *seen*. Whenever +`HEAD` is not the object id the cycle recorded, **something reached this repository that no pass +reviewed** — a further commit, an amended one, a rebased one, or a tree the history moved back to — +and a ticked checkbox does not make any of it reviewed. **Do not adopt it**: carrying it into a +close is the defect Close condition 1 exists to stop, and that condition's own rule decides the +cost — only a clean response issued against that exact `HEAD` closes the cycle. **Do not remove it**: +this plan restores nothing, and what you delete here is evidence a person has not yet chosen a §A +route on. **And do not rewrite `.context/loop-rule-reviewed-head` or `-reviewed-tip` to match**, +which would make the equality pass by discarding the only record of what was reviewed. **Report the +mismatch and the state**, name it in the Failure report's "whether a commit landed" line, and leave +`## The four procedures` · Failure and §A's three routes to decide. **How success is recognised.** The base passes Resume's own checks; the artifacts match it and are complete, or are reported as invalid and left alone; and the plan's task list has been reconciled -against the content the topology actually holds. +against the content the four sources actually hold. **Fitting a named shape is not among them.** **Steps that describe the tree at `$BASE` are validated on re-entry, not re-run against the worktree.** After a text task the worktree carries this plan's own edits, and rebuilding the @@ -1199,8 +1219,7 @@ and 11 also record — and a stale list is the same defect as a stale count. **T omission is concrete:** the evidence stays out of its own task's `WIP:` snapshot, so the reviewed range does not hold it where the task claims, and if execution continues it is swept into a later unrelated commit rather than the independently reviewable snapshot this plan promises. **Not a -clean-tree argument** — Resume requires no clean tree, and none of its topologies is required to have -one; an earlier draft said re-entry needs one, which would have rejected the very states Resume +clean-tree argument** — Resume requires no clean tree in any state it can meet; an earlier draft said re-entry needs one, which would have rejected the very states Resume exists to reconcile. Named `WIP:` because Task 15 runs Gate B over the whole change and closes it with **one From ab3cb8eef7694d0b7f008b5a836fa945f44a79ce Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Wed, 16 Sep 2026 11:40:15 +0200 Subject: [PATCH 156/181] docs(context): record Gate-A plan pass 38 and the classification decision Pass 38 reviewed b283250: 1 Blocker, 2 Majors, 0 Minors. The Blocker is the one that matters. Resume never checks the branch. Preparation does, and it is first-entry-only; the operational index and accounting row 1 both assign the condition to Preparation alone. The base file is ignored, so it survives a checkout, and another branch descended from $BASE passes every Resume check - after which the WIP commits, the soft reset and the closing commit land on the wrong branch while the plan reports a valid re-entry. Pre-existing, and thirty-seven passes never looked at it. That is a coverage signal, not a convergence signal, and the record now says so: no affirmative coverage judgement can be given for this artifact, which is one of the three things the clearly-stuck exit requires. It is also not a waiting period - the two-tell stop is already enough to decide on. The first Major is this round's own regression: "present is not reviewed" routes every mismatch to Failure, contradicting the precondition-versus- handoff split four paragraphs earlier, which would select the inside-a- handoff scratch branch for a condition-1 stop that owed a rebuild. The second is pre-existing: condition 6 diffs the landed body against the same mutable file the commit read, after hooks have run, while the plan elsewhere states hooks can rewrite anything and that the file is ignored. Repair-produces-the-next-finding is now three rounds running. Findings 6,1,2,3,3; Blockers 1,0,0,0,1; Majors 1,1,1,1,2. A new per-pass check is recorded: walk Preparation's conditions against Resume's one by one, because Preparation is first-entry-only and every condition it owns is one Resume may not have. Floor 3, from docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md (risk high, security none), read fresh at this pass. Cycle stays open and unclean. --- .../gate-a-plan-om0bdd7udh-pass-38.md | 4 + .../gate-a-plan-om0bdd7udh-resume.md | 93 +++++++++++-------- 2 files changed, 59 insertions(+), 38 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-38.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-38.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-38.md new file mode 100644 index 0000000..ab9d316 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-38.md @@ -0,0 +1,4 @@ +BLOCKER | high | Task 0 step 1 / Resume base validation | The re-entry route never checks that the current branch is `loop-rule-consolidation`; the ignored base file survives a branch checkout, and another branch descended from `$BASE` can pass every Resume check even though Preparation and accounting condition 1 require this branch. | Later WIP commits and the final soft reset and closing commit can be made on the wrong branch, rewriting that branch while the plan reports a valid re-entry. | Add the branch equality check to Resume's own conditions and to the operational-condition index/accounting, and run it before any re-entry mutation. +MAJOR | high | Resume, “Present is not reviewed” | The universal mismatch rule sends every `HEAD` mismatch to the Failure report and §A routes, contradicting the preceding governing distinction that condition 1 is a plain precondition stop with no handoff while only condition 5 starts a Failure handoff. | A condition-1 mismatch can be treated as an active handoff, selecting the “inside a handoff” scratch-artifact branch and suppressing a rebuild even though no closing operation moved anything. | Route the mismatch by the equality check that rejected it: report condition 1 as the plain precondition stop, and send condition 5 through Failure; retain the no-adopt, no-remove and no-rewrite rules for both. +MAJOR | medium | Task 15 step 8, conditions 3 and 6 | The closing-message bytes revalidated before `git commit` are not pinned; condition 6 compares the landed body with `.context/loop-rule-closing-msg` only after hooks have run, even though the plan elsewhere admits a hook can rewrite any file during the attempt. | A hook that changes both Git's message copy and the ignored closing-message file can make the postcondition pass while the commit body differs from the message the reader actually validated. | Capture the validated message bytes or digest before invoking `git commit` and compare the landed body to that captured value, ideally in the same 8b invocation, rather than using the mutable source path as the post-commit oracle. +END OF FINDINGS (3 total) diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md index 81e99d9..3cf4f51 100644 --- a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md @@ -13,47 +13,58 @@ Nothing depends on it; the pass files and the repo are authoritative where this ## Resume here -**Cycle OPEN and UNCLEAN. The plan stands at `d7f2af8` and that is the anchor** — later commits on +**Cycle OPEN and UNCLEAN. The plan stands at `b283250` and that is the anchor** — later commits on this branch are records, not plan edits, so check `git log --oneline -1 -- docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md` -rather than `HEAD`. Pass 37 reviewed `d7f2af8` and found **one Major, one Minor and one Nit, -0 Blockers**. `.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-37.md` holds them. +rather than `HEAD`. Pass 38 reviewed `b283250` and found **one Blocker and two Majors**. +`.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-38.md` holds them. **Do not start a repair round on your own.** Daniel's assignment of 2026-09-16 ended with a -checkpoint: *"Run exactly one complete Gate-A plan pass after the bounded revision and verification, -then stop and report regardless of outcome. Do not begin another repair round or implementation."* -That checkpoint is spent — it covered one bounded repair and one pass, both delivered — so **the -next move is Daniel's word, not an inference.** - -**Two tells present, so this stop is mandatory as well as instructed**: findings rose 1 → 2 → 3, -three passes running, and they cluster on the instrument. The Blocker count "failed to fall" at -0 → 0 is the literal third test and nonsense at zero; **that reading is still not settled here and -nothing turns on it.** - -**Read this before deciding the next round — the regeneration signature is now visible.** Findings -6 → 1 → 2 → 3; Blockers 1 → 0 → 0 → 0; **Majors exactly 1 in each of the last four passes**. And -pass 37's Major is **descended from pass 36's fix**: the moved-`HEAD` row was keyed on the *shape* -of the state ("a commit above the expected head") where both guards test **object-id inequality**, -so a rewrite that trips the same guard matches no row. That is one established lineage, a genuine -repair producing the next finding. **It is not yet the "clearly stuck" exit**, which needs three -things together — a plateau across roughly six passes, an affirmative coverage judgement, and -Blocker/Majors regenerating across repairs. We have the third and a quarter of the first. - -### The three open findings from pass 37 - -1. **MAJOR — the moved-`HEAD` topology covers one shape of several.** Conditions 1 and 5 are plain - inequalities (`test "$(git rev-parse HEAD)" = "$HEADREV"`, and the same against `$TIP`). The new - row describes only **a descendant commit above** the expected head. An **amend**, a **rebase**, - or a **reset to an earlier commit still above `$BASE`** trips the identical guard and matches - neither the row's `HEAD`/range description nor its "extra commit" reconciliation — so a rejected, - rewritten, unreviewed state can still read as normal progress. **The row has to be defined from - the predicate, not from a shape**, and split where the operation differs. Validated against the - plan. -2. **MINOR — accounting row 7 still says Resume's table defines "four shapes".** It now has five - rows. **Created by the previous repair**: the sweep for stale counts keyed on the word - *topologies* and this site says *shapes*. One-word fix; collected per the Minor rule. -3. **NIT — Self-Review's reader-check sentence names Tasks 12, 13 and 14 step 2**, omitting Task - 12b's reader-led sweep and Task 15 step 4b's prompt-standards review. The gate prompt's own - settled block lists all five. Pre-existing, not caused by this round. Collected. +checkpoint: *"After this bounded revision and verification, run exactly one complete Gate-A plan +pass and report back. Do not begin another repair round or implementation."* That checkpoint is +spent — one revision, one pass, both delivered — so **the next move is Daniel's word, not an +inference.** + +**Two tells, so this stop is mandatory as well as instructed**: the **Blocker count rose** 0 → 1 +after four passes at zero, and the findings cluster on the instrument for the fifth pass running. +Findings are flat at 3, which is neither rising nor falling and is not counted as a tell here. + +**Read this before deciding the next round. Two things, and the second is the more serious.** + +**One: the repair keeps producing the next finding — three rounds in a row now.** Pass 36's fix +produced pass 37's Major; pass 37's fix produced pass 38's first Major, which is a **regression +introduced by this round's own generalisation**. Findings 6 → 1 → 2 → 3 → 3; Blockers 1 → 0 → 0 → 0 +→ **1**; Majors 1 → 1 → 1 → 1 → **2**. + +**Two: pass 38 found a Blocker in an area thirty-seven passes never looked at.** The branch +condition has been in the accounting table since the beginning, assigned to Preparation, and nobody +noticed Resume never inherited it. **That is the coverage signal, not the convergence signal** — +`CLAUDE.md` says plainly that a low Blocker count can sit beside an entirely unreviewed subsystem, +and this is what that looks like from the inside. **No affirmative coverage judgement can be given +for this artifact right now**, which is one of the three things the "clearly stuck" exit needs, so +that exit is unavailable — and it is not a waiting period either: the two-tell stop is already +enough to decide on, and the six-pass figure is a field observation, not a threshold to sit out. + +### The three open findings from pass 38 + +1. **BLOCKER — Resume never checks the branch.** Preparation checks it and is **first-entry-only**; + the operational index and accounting row 1 both assign the condition to Preparation alone, and + Resume's own block checks base shape, self-resolution, ancestry and approved inputs — **and not + the branch**. `.context/loop-rule-base` is ignored, so it survives a checkout: another branch + descended from `$BASE` passes every Resume check. Then the WIP commits, the soft reset and the + closing commit all land on the wrong branch while the plan reports a valid re-entry. Real, + reachable, and **pre-existing** — not caused by this round. +2. **MAJOR — this round's own regression.** "Present is not reviewed" ends by sending **every** + mismatch to the Failure report and §A's routes, four paragraphs after the governing text says + condition 1 is a **plain precondition stop with no handoff** and only condition 5 starts one. + A condition-1 mismatch read as an active handoff selects the `inside a handoff` scratch branch + and suppresses a rebuild that was owed. **Route by the check that rejected**; keep no-adopt, + no-remove and no-rewrite for both. +3. **MAJOR — condition 6's oracle is the mutable file it is supposed to audit.** 8b commits with + `-F .context/loop-rule-closing-msg` and the post-close block diffs the landed body against **that + same path**, read after the hooks have run. The plan itself states that hooks can rewrite + anything and that the file is ignored, so no clean-tree check sees it change — so a hook editing + both git's copy and the source file makes the postcondition pass on a body nobody validated. + **Pin the bytes or a digest before `git commit`** and compare against that. ### Still collected and deliberately unrepaired (pass 34's Minors and Nit) @@ -148,6 +159,10 @@ what it carries; `.context/gate-a-spec-prompt.md` is the same shape for the spec passage's rows — but a matching id is a search hit, not a defect.** Tasks 3, 5 and 7 already read their sets off the table; Task 7's lists and Task 4 step 4's move list are complete and were left alone deliberately. Replacing them blind would remove requirements. +- **Preparation is first-entry-only, so every condition it owns is one Resume may not have.** The + branch check was assigned to Preparation alone and Resume never inherited it — thirty-seven passes + missed it (pass 38, BLOCKER). **Walk Preparation's conditions against Resume's, one by one**, and + do not assume the split is deliberate because it is written down. - **When you state a rule in one place, grep for the CLAIM, not the word you used.** Two sweeps in two rounds each reached most sites and missed the rest, both times because the missed site used a synonym: `1` instead of `preservation count`, `four shapes` instead of `four topologies`. This is @@ -217,6 +232,8 @@ worth checking before a pass rather than after. | 36 | 518121a | 1→**2** | 0→**0** | 1→**1** | yes | **UNCLEAN — reported to Daniel, no repair round opened.** **The h3 repair held**; nothing re-raised against Task 8. MAJOR: Resume's topology table claims four exhaustive topologies and has **no row for the state condition 5 exists to catch** — `HEAD` above the 8a findings commit after a between-invocations commit; the condition-1 rejection reads as row 1, "normal". MINOR: Task 4 step 5's reader walk omits `c4` — the **same family** as pass 35's Major, a task enumerating its own conditions. Two tells (findings rose, instrument cluster) → mandatory stop, which the checkpoint already required | | — | — | — | — | — | — | **FOURTH BOUNDED REPAIR**, Daniel's assignment after pass 36, on his reviewer's narrowed recommendation: repair the Major and Task 4's `c4`, plus a **bounded** search for the same coverage-list defect — **not** a blanket replacement of every id list, because a grep finds ids and cannot tell a constraining list from an orientation note. Resume gains one moved-`HEAD` row; the progress reconciliation gains *present is not reviewed*; Task 4 step 5 walks the table. Search result: Tasks 3, 5, 7 already cite the table, Task 7's and Task 4 step 4's lists checked complete and left alone. Four bare-`1` results and two stale counts corrected as disclosed misses of the previous sweep. Commit `d7f2af8` | | 37 | d7f2af8 | 2→**3** | 0→**0** | 1→**1** | yes | **UNCLEAN — reported to Daniel, no repair round opened.** MAJOR: the new moved-`HEAD` row was keyed on the **shape** of the state where both guards test **object-id inequality**, so an amend, rebase or reset still above `$BASE` matches no row — **pass 36's fix produced it**, one established regeneration lineage. MINOR: accounting row 7 still says "four shapes", falsified by the fifth row — the stale-count sweep keyed on *topologies*, this site says *shapes*. NIT: Self-Review's reader-check sentence omits Task 12b and Task 15 step 4b. Two tells (findings rose a third pass, instrument cluster) → mandatory stop | +| — | — | — | — | — | — | **FIFTH BOUNDED REVISION — a DECISION, not another description.** Daniel: **the plan does not classify the repository.** §A re-establishes the conditions *against the repository as it now stands*; the demand that Resume fit the state to a named shape was plan-invented and is **dropped**, recorded in accounting row 7 as a *requirement* (so row 40's count of dropped *conditions* stays true). The table is explicitly non-exhaustive illustration; nothing depends on fitting a row. Preserved: both equality checks, the precondition/handoff split, the scratch rules, the bounded handoff, base validity checked not inferred, §A's three routes. Strengthened: all four sources read every time, and no rewriting the recorded head or tip to make equality pass. 64 checks × 3 shells — descendant, amended, rebased, backward, each before and after 8a. Commit `b283250` | +| 38 | b283250 | 3→**3** | 0→**1** | 1→**2** | yes | **UNCLEAN — reported to Daniel, no repair round opened.** **BLOCKER: Resume never checks the branch** — Preparation does and is first-entry-only, the ignored base file survives a checkout, and another branch descended from `$BASE` passes everything; pre-existing, and **thirty-seven passes never looked**. MAJOR: this round's own regression — "present is not reviewed" routes *every* mismatch to Failure, contradicting the precondition/handoff split four paragraphs earlier. MAJOR: condition 6 diffs the landed body against the **same mutable file** the commit read, after hooks ran. Two tells (Blockers rose, instrument cluster) → mandatory stop | | 26 | 05ec1b2 | **5** | **1** | 2 | yes | first pass on the revised plan. Resume audited progress by `WIP:` commits and called any non-`WIP` commit a stale base — the handoff leaves two other shapes, one of them `HEAD` **at** the base with everything in the index | | 27 | f69db7d | 5→**6** | 1→**3** | 2→**1** | yes | the close required the commit's **tree to equal the tip's**, which target §I **parks**; `reset --soft` leaves the index untouched, so the prose beside it was wrong about git; Resume still inferred staleness from a commit subject where §A3 names that state as reachable | | 28 | f2e5dda | 6→**6** | 3→**2** | 1→**2** | yes | pass 27's stale-base fix reached one site of three; its reset paragraph stated the index behaviour correctly and repeated the false claim four lines later; nothing checked the closing commit's **parent**; the message was validated in the file, where `commit-msg` hooks rewrite git's copy after `-F` reads it | From df9123a68cd85e4b54e1c0a2923e4ac94402d98c Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Wed, 16 Sep 2026 12:00:08 +0200 Subject: [PATCH 157/181] docs(plans): Resume checks the branch; route by the check that rejected Pass 38's three findings, after the read-only reconciliation of the accounting table found no further operational gap. Blocker. Resume never checked the branch. Preparation did, and it is first-entry-only, so the operational index and accounting row 1 both assigned the condition to Preparation alone. .context/ is ignored, so the base file survives a checkout: on another branch descended from $BASE every re-entry check passed, and the WIP commits, the soft reset and the closing commit would all have landed on the wrong branch. Resume now checks the branch first, before anything reads or writes cycle state, and row 1 is split by entry the way row 4 already was. A detached HEAD prints HEAD and is refused for the same reason. Major, this round's own regression. "Present is not reviewed" sent every mismatch to Failure, four paragraphs after the governing text says condition 1 is a plain precondition stop with no handoff and only condition 5 starts one. The mismatch is now routed by the check that rejected it: condition 1 is a plain stop and takes the scratch rule's ordinary delete-and-rebuild branch; condition 5 goes to Failure and takes report-and-change-nothing. No-adopt, no-remove and no-rewrite hold for both. Major, pre-existing. Condition 6 compared the landed body against the same mutable file the commit read, after hooks had run, while the plan states elsewhere that hooks can rewrite anything and that the file is ignored. 8b now pins the validated bytes before the commit and condition 6 compares against that copy. The residual is stated in the condition: a hook that also rewrites the pinned copy defeats it, and nothing here detects that. The reconciliation itself produced no permanent structure, as instructed. Of the forty-one rows: one gap, which is the Blocker above; rows 2 and 9 are first-entry-only by design and moving row 2 to Resume would refuse valid re-entries; rows 7, 20 and 40 are deliberately dropped; rows 4, 5, 6, 8, 11, 14, 16 and 19 are checked at re-entry today; the rest bear on later steps. Row 3 is satisfied at re-entry by derivation rather than by a check - ba15e83 is an ancestor of $BASE and Resume checks $BASE is an ancestor of HEAD - and that derivation depends on $BASE having been recorded by Preparation. Verified by execution under sh, dash and bash: 13 new checks plus 111 re-run, all green. The wrong branch is refused before any mutation and a valid re-entry still proceeds; a detached HEAD is refused; conditions 1 and 5 still reject in the shell; a commit-msg hook rewriting both git's copy and the source file is now caught, with a control showing the old oracle would have passed that same case; an untouched message still passes. The routing half is reader text and was asserted as text. --- .../2026-09-14-loop-rule-consolidation.md | 52 ++++++++++++++----- 1 file changed, 40 insertions(+), 12 deletions(-) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index 9a4d95c..38ffdab 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -210,7 +210,7 @@ live in the home named. Where a task, a shell comment or an error string needs a | Operational condition | Defined in | Discharged by | |---|---|---| | Branch, clean tree, no base file yet, `ba15e83` ancestral, approved inputs unchanged at `HEAD` | **Preparation** | Task 0 step 1, first-entry branch | -| Base file shape, self-resolving commit, ancestral; approved inputs unchanged at `$BASE`; scratch-artifact validity; how far the implementation got | **Resume** | Task 0 step 1, re-entry branch | +| Branch; base file shape, self-resolving commit, ancestral; approved inputs unchanged at `$BASE`; scratch-artifact validity; how far the implementation got | **Resume** | Task 0 step 1, re-entry branch | | What a re-entry reads, and that no rule depends on the state fitting a named shape | **Resume** — its table is illustration, not a classification | accounting row 7; every state-reading rule | | `HEAD` equals the reviewed head | **Close**, condition 1 | step 8a | | The dirty set is exactly the candidate pass's findings files | **Close**, condition 2 | step 8a | @@ -236,7 +236,7 @@ independent reader did. | # | Condition | Disposition | |---|---|---| -| 1 | Branch is `loop-rule-consolidation` | **kept**, same shell — preparation | +| 1 | Branch is `loop-rule-consolidation` | **kept**, and **split by entry, like row 4**: Preparation checks it at a first entry and **Resume checks it again**, first, before anything reads or writes cycle state. An earlier revision assigned it to Preparation alone, which is first-entry-only — and `.context/` is ignored, so the base file survives a checkout and another branch descended from `$BASE` passed every re-entry check (pass 38, Blocker) | | 2 | Tree clean before the base is recorded | **kept**, same shell — preparation | | 3 | `ba15e83` is an ancestor of `HEAD` | **kept**, same shell — preparation | | 4 | The three approved inputs' blobs equal their `ba15e83` versions | **kept, and split by entry**: Preparation compares them **at `HEAD`**, because a first entry has no recorded base; **Resume compares them at `$BASE`**, because a Gate-B fix may legitimately have changed `HEAD`'s copy in a `WIP:` snapshot. An earlier table row claimed Preparation did the base comparison, which it never could | @@ -409,14 +409,18 @@ an unapproved edit to a spec that the gates already closed. without it a staged edit made between the two invocations, a rewritten findings file included, ships unreviewed and leaves a clean tree behind it. 6. **After the closing commit**, four things: its **subject is not a snapshot**; its **parent is - `$BASE`**; the **tree is clean**; and its **body is identical to the revalidated closing-message - file**. The last is not decoration — `prepare-commit-msg` and `commit-msg` hooks rewrite git's - copy *after* `-F` has read it, so a validated file proves nothing about what landed, and a - dropped provenance line, curve, exception marker or evidence entry would pass every other check. - **Any of the four failing is a Failure handoff.** + `$BASE`**; the **tree is clean**; and its **body is identical to the bytes 8b pinned when it + revalidated them** — **not to the source file, which is the wrong oracle**: the same hooks that + rewrite git's copy *after* `-F` has read it can rewrite that ignored file too, and a comparison + against it would then pass on a body nobody validated. The check is not decoration either way — + a dropped provenance line, curve, exception marker or evidence entry would pass every other one. + **What this does not close**, stated rather than left to be found: a hook that also rewrites the + pinned copy defeats it, and nothing here detects that. What it closes is the ordinary case, where + the oracle was the very file the commit read. **Any of the four failing is a Failure handoff.** **How success is recognised.** A single commit at `HEAD` whose subject is the real message, whose -parent is `$BASE`, **whose body is identical to the revalidated closing message**, with a clean tree +parent is `$BASE`, **whose body is identical to the pinned bytes of the revalidated closing +message**, with a clean tree behind it. **No tree comparison** — target §I parks a Gate-B tree-equality condition on Daniel's decision of 2026-09-13, and an earlier draft of this line added one anyway, which would have made the executor either invent an out-of-scope closure check or declare success without @@ -549,6 +553,14 @@ nothing, rather than delete and rebuild. - [ ] **Validate the base — Resume's own checks, not Preparation's** ```bash +# The branch, FIRST and before anything reads or writes cycle state. `.context/` is +# ignored, so the base file survives a checkout: on another branch descended from +# $BASE every check below passes and the WIP commits, the soft reset and the closing +# commit all land on the wrong branch. A detached HEAD prints `HEAD` and is refused +# for the same reason. +test "$(git rev-parse --abbrev-ref HEAD)" = loop-rule-consolidation \ + || { echo "not on loop-rule-consolidation (on: $(git rev-parse --abbrev-ref HEAD)) — refusing to re-enter"; exit 1; } + test -e .context/loop-rule-base || { echo "no base file — this is a first entry, run Preparation"; exit 1; } test -s .context/loop-rule-base || { echo "base file is EMPTY — inspect, delete deliberately, record why, re-record from the true starting commit"; exit 1; } BASE=$(cat .context/loop-rule-base) @@ -628,9 +640,18 @@ close is the defect Close condition 1 exists to stop, and that condition's own r cost — only a clean response issued against that exact `HEAD` closes the cycle. **Do not remove it**: this plan restores nothing, and what you delete here is evidence a person has not yet chosen a §A route on. **And do not rewrite `.context/loop-rule-reviewed-head` or `-reviewed-tip` to match**, -which would make the equality pass by discarding the only record of what was reviewed. **Report the -mismatch and the state**, name it in the Failure report's "whether a commit landed" line, and leave -`## The four procedures` · Failure and §A's three routes to decide. +which would make the equality pass by discarding the only record of what was reviewed. + +**Route the mismatch by the check that rejected it, not by the fact that it was a mismatch.** The +no-adopt, no-remove and no-rewrite rules above hold for both; **where they go does not.** A +**condition 1** mismatch is a precondition rejection — nothing this plan did has moved, so it is a +**plain stop**: report the mismatch and the state, and **no handoff is in progress**, which is what +lets the scratch-artifact rule above take its ordinary *delete and rebuild* branch. A **condition 5** +mismatch rejects after 8a's record commit has landed, so it goes to `## The four procedures` · +Failure, is named in that report's "whether a commit landed" line, and **is** inside a handoff — so +the scratch rule's *report and change nothing* branch applies. Either way §A's three routes decide +what happens next. **An earlier revision sent every mismatch to Failure**, which would have put a +condition-1 stop inside a handoff it is not in and suppressed a rebuild that was owed. **How success is recognised.** The base passes Resume's own checks; the artifacts match it and are complete, or are reported as invalid and left alone; and the plan's task list has been reconciled @@ -2968,6 +2989,11 @@ git reset --soft "$BASE" || { echo "reset --soft FAILED — run Failure; do NOT # Condition 3's revalidation: re-read the message here, immediately before it is # consumed. Reader check on its records; the expected result is condition 3's. test -s .context/loop-rule-closing-msg || { echo "closing message missing or empty — run Failure"; exit 1; } +# Pin the bytes this invocation validated, BEFORE the commit runs. Condition 6's +# oracle must not be the same mutable path the commit reads: a `commit-msg` hook +# that rewrites git's copy AND this ignored source file would otherwise make the +# post-close comparison pass on a body nobody validated. +cp .context/loop-rule-closing-msg .context/loop-rule-validated-msg # `--cleanup=verbatim` so the stored body is the validated bytes: git's default # cleanup for -F strips trailing whitespace and collapses blank runs, and # condition 6 compares bytes. It carries no `-m`, so `is_wip_commit` still misses it. @@ -2990,8 +3016,9 @@ test -z "$(git status --porcelain)" || { echo "tree dirty after the close — ru # 8b's `--cleanup=verbatim` is the other half — without it git stores its own # tidied copy. Both observed in a disposable repository, so this diff is the # byte equality condition 6 states, not a normalized stand-in for it. +test -s .context/loop-rule-validated-msg || { echo "no pinned message — 8b did not complete; run Failure"; exit 1; } git log -1 --pretty=format:%B > .context/loop-rule-landed-msg -diff .context/loop-rule-closing-msg .context/loop-rule-landed-msg \ +diff .context/loop-rule-validated-msg .context/loop-rule-landed-msg \ || { echo "the committed body differs from the validated message — run Failure"; exit 1; } ``` @@ -3013,6 +3040,7 @@ rm -f .context/loop-rule-base .context/loop-rule-reviewed-tip .context/loop-rule .context/loop-rule-sites .context/loop-rule-changed-sites \ .context/loop-rule-c.src .context/loop-rule-w.src \ .context/loop-rule-a.txt .context/loop-rule-b.txt .context/loop-rule-landed-msg \ + .context/loop-rule-validated-msg \ .context/loop-rule-final-blobs .context/loop-rule-committed-blobs if ls .context/loop-rule-* >/dev/null 2>&1; then echo "cycle scratch survives the close:"; ls .context/loop-rule-*; exit 1 From c1fcdd1ded1860998641808b203aafaf728549e0 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Wed, 16 Sep 2026 12:13:50 +0200 Subject: [PATCH 158/181] docs(context): record Gate-A plan pass 39 and the sixth bounded revision Pass 39 reviewed df9123a: 1 Blocker, 1 Major, 0 Minors. All three pass-38 repairs held. Nothing was re-raised against the branch check, the mismatch routing or the pinned message oracle, which breaks the repair-produces-the-next-finding chain for the first time in four rounds. Blocker. Task 15 step 7 commits the repair in step one and reads the pass through the installed ordering in step two, so a membership stop, a new-question stop or a stop answer arrives after the fix is already in the artifact - and this plan restores nothing, so no route removes it. The step's own header says "this commits; it does not authorize anything", which names the tension without resolving it. It is the loop its own section-A product forbids. Major. Step 4b repairs failed prompt-standards text and then commits only the plan record, leaving the repaired prompt copies and hook bodies outside $BASE..HEAD, where Gate B never sees them and the close's exact-dirty-set check then refuses them. Second consecutive pass finding a Blocker in an area no earlier pass flagged. That is a statement about these two passes, not a prediction about how many such areas remain. The accounting-table reconciliation run this round did what it claimed and no more: it walked the forty-one rows for re-entry relevance and could not have found either pass-39 finding, which are about step sequencing and staging. It is not a coverage certificate. A new per-pass lesson is recorded: ask the accounting table a scoped question, because "which row is unchecked at re-entry" would have moved row 2's clean-tree demand onto Resume and refused valid re-entries. Floor 3, from docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md (risk high, security none), read fresh at this pass. Cycle stays open and unclean. --- .../gate-a-plan-om0bdd7udh-pass-39.md | 3 + .../gate-a-plan-om0bdd7udh-resume.md | 104 +++++++++--------- 2 files changed, 58 insertions(+), 49 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-39.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-39.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-39.md new file mode 100644 index 0000000..398a65e --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-39.md @@ -0,0 +1,3 @@ +BLOCKER | high | Task 15 step 7 | The non-closing-pass sequence permits a finding repair to be applied and committed in step one before step two reads that pass through the installed ordering. | A membership or question suspension can therefore be answered only after its finding has already been adopted into the artifact, and a later decline or stop has no route that removes that unauthorized repair; the plan's own Gate-B loop can ship work outside the assigned fix set. | Commit only the pass record before routing, read the pass through the ordering, and apply and commit a repair only after the resulting answers and scope rules authorize it, before recording the next reviewed head. +MAJOR | high | Task 15 step 4b | The stated repair path tells the executor to repair failed prompt-standard text and then runs a commit command that stages only the plan record. | The repaired prompt files remain outside $BASE..HEAD when Gate B is issued, so the reviewer does not see them and the later exact-dirty-set check dead-ends the close; committing or staging them after review instead invalidates the reviewed-head condition. | Stage and commit every repaired artifact together with the refreshed plan record in a WIP snapshot, then rerun the affected checks and battery against that committed HEAD before issuing Gate B. +END OF FINDINGS (2 total) diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md index 3cf4f51..e643a32 100644 --- a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md @@ -13,58 +13,56 @@ Nothing depends on it; the pass files and the repo are authoritative where this ## Resume here -**Cycle OPEN and UNCLEAN. The plan stands at `b283250` and that is the anchor** — later commits on +**Cycle OPEN and UNCLEAN. The plan stands at `df9123a` and that is the anchor** — later commits on this branch are records, not plan edits, so check `git log --oneline -1 -- docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md` -rather than `HEAD`. Pass 38 reviewed `b283250` and found **one Blocker and two Majors**. -`.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-38.md` holds them. +rather than `HEAD`. Pass 39 reviewed `df9123a` and found **one Blocker and one Major**. +`.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-39.md` holds them. **Do not start a repair round on your own.** Daniel's assignment of 2026-09-16 ended with a -checkpoint: *"After this bounded revision and verification, run exactly one complete Gate-A plan -pass and report back. Do not begin another repair round or implementation."* That checkpoint is -spent — one revision, one pass, both delivered — so **the next move is Daniel's word, not an -inference.** - -**Two tells, so this stop is mandatory as well as instructed**: the **Blocker count rose** 0 → 1 -after four passes at zero, and the findings cluster on the instrument for the fifth pass running. -Findings are flat at 3, which is neither rising nor falling and is not counted as a tell here. - -**Read this before deciding the next round. Two things, and the second is the more serious.** - -**One: the repair keeps producing the next finding — three rounds in a row now.** Pass 36's fix -produced pass 37's Major; pass 37's fix produced pass 38's first Major, which is a **regression -introduced by this round's own generalisation**. Findings 6 → 1 → 2 → 3 → 3; Blockers 1 → 0 → 0 → 0 -→ **1**; Majors 1 → 1 → 1 → 1 → **2**. - -**Two: pass 38 found a Blocker in an area thirty-seven passes never looked at.** The branch -condition has been in the accounting table since the beginning, assigned to Preparation, and nobody -noticed Resume never inherited it. **That is the coverage signal, not the convergence signal** — -`CLAUDE.md` says plainly that a low Blocker count can sit beside an entirely unreviewed subsystem, -and this is what that looks like from the inside. **No affirmative coverage judgement can be given -for this artifact right now**, which is one of the three things the "clearly stuck" exit needs, so -that exit is unavailable — and it is not a waiting period either: the two-tell stop is already -enough to decide on, and the six-pass figure is a field observation, not a threshold to sit out. - -### The three open findings from pass 38 - -1. **BLOCKER — Resume never checks the branch.** Preparation checks it and is **first-entry-only**; - the operational index and accounting row 1 both assign the condition to Preparation alone, and - Resume's own block checks base shape, self-resolution, ancestry and approved inputs — **and not - the branch**. `.context/loop-rule-base` is ignored, so it survives a checkout: another branch - descended from `$BASE` passes every Resume check. Then the WIP commits, the soft reset and the - closing commit all land on the wrong branch while the plan reports a valid re-entry. Real, - reachable, and **pre-existing** — not caused by this round. -2. **MAJOR — this round's own regression.** "Present is not reviewed" ends by sending **every** - mismatch to the Failure report and §A's routes, four paragraphs after the governing text says - condition 1 is a **plain precondition stop with no handoff** and only condition 5 starts one. - A condition-1 mismatch read as an active handoff selects the `inside a handoff` scratch branch - and suppresses a rebuild that was owed. **Route by the check that rejected**; keep no-adopt, - no-remove and no-rewrite for both. -3. **MAJOR — condition 6's oracle is the mutable file it is supposed to audit.** 8b commits with - `-F .context/loop-rule-closing-msg` and the post-close block diffs the landed body against **that - same path**, read after the hooks have run. The plan itself states that hooks can rewrite - anything and that the file is ignored, so no clean-tree check sees it change — so a hook editing - both git's copy and the source file makes the postcondition pass on a body nobody validated. - **Pin the bytes or a digest before `git commit`** and compare against that. +checkpoint: *"Danach genau ein vollständiger Gate-A-Plan-Pass und Bericht. Keine automatische +Folgerunde oder Implementierung."* That checkpoint is spent — one reconciliation, one bounded +repair, one pass, all delivered — so **the next move is Daniel's word, not an inference.** + +**Two tells, so this stop is mandatory as well as instructed**: the **Blocker count failed to fall**, +1 → 1 — and this time that is the real test rather than the nonsense-at-zero reading, because there +is a Blocker and it did not go away — and the findings cluster on the instrument for the sixth pass +running. Findings fell 3 → 2 and Majors fell 2 → 1, so neither is a tell. + +**Read this before deciding the next round. Three things.** + +**One, and it is good news: all three pass-38 repairs held.** Nothing was re-raised against the +branch check, the mismatch routing or the pinned message oracle. **That breaks the +repair-produces-the-next-finding chain for the first time in four rounds** — passes 36, 37 and 38 +each found a defect descended from the previous fix; pass 39 found none. + +**Two: a second Blocker, in a second area no earlier pass flagged.** Pass 38's was Resume's missing +branch check; pass 39's is Task 15 step 7's ordering. **Two consecutive passes have each found a +Blocker in a previously unflagged area** — that is the evidence, and it is a statement about what +these two passes did, not a prediction about how many more such areas exist. + +**Three: the accounting-table reconciliation did what it claimed and no more.** It walked the +forty-one rows for re-entry relevance and found the one gap it was aimed at. **It could not have +found either pass-39 finding**, which are about Task 15's step *sequencing* and *staging*, not about +conditions at re-entry. **It is not a coverage certificate for the artifact** and must not be read +as one. + +### The two open findings from pass 39 + +1. **BLOCKER — Task 15 step 7 commits the repair before the ordering has spoken.** Step one is + `git add -A && git commit -m "WIP: fix "`, and step two *then* reads the pass through + the installed ordering — where a membership stop, a new-question stop or a stop answer may say + the finding was never in the assigned fix set. By then the repair is committed, and **this plan + restores nothing**, so no route removes it. The step's own header says "this commits; it does not + authorize anything", which names the tension without resolving it. **The plan's own Gate-B loop + can ship work outside the assigned fix set** — the failure its own §A product forbids, and the + thing this step was written to prevent. Validated against the plan. +2. **MAJOR — Task 15 step 4b repairs the prompt text and commits only the plan.** On a + prompt-standards failure it says to repair the text, re-run the invalidated checks, and commit — + but its command is `git add docs/superpowers/plans/…` alone. The repaired `CLAUDE.md`, + `workflow-init.md` and hook bodies stay in the worktree: **outside `$BASE..HEAD`, so Gate B never + sees them**, and then the close's exact-dirty-set check refuses a dirty prompt copy, so the cycle + dead-ends. Staging them after the review instead breaks the reviewed-head condition. Validated + against the plan. ### Still collected and deliberately unrepaired (pass 34's Minors and Nit) @@ -159,6 +157,12 @@ what it carries; `.context/gate-a-spec-prompt.md` is the same shape for the spec passage's rows — but a matching id is a search hit, not a defect.** Tasks 3, 5 and 7 already read their sets off the table; Task 7's lists and Task 4 step 4's move list are complete and were left alone deliberately. Replacing them blind would remove requirements. +- **Ask a scoped question of the accounting table, not a wide one.** "Which row is unchecked at + re-entry" is the wrong question — rows 2 and 9 are first-entry-only *by design*, and moving row + 2's clean-tree demand to Resume would refuse valid re-entries. The question that works: **which + conditions must still hold at re-entry, how is each one's validity established, and does that + happen before the first action depending on it?** Run once, as a report; do not turn it into a + permanent structure in the plan. - **Preparation is first-entry-only, so every condition it owns is one Resume may not have.** The branch check was assigned to Preparation alone and Resume never inherited it — thirty-seven passes missed it (pass 38, BLOCKER). **Walk Preparation's conditions against Resume's, one by one**, and @@ -234,6 +238,8 @@ worth checking before a pass rather than after. | 37 | d7f2af8 | 2→**3** | 0→**0** | 1→**1** | yes | **UNCLEAN — reported to Daniel, no repair round opened.** MAJOR: the new moved-`HEAD` row was keyed on the **shape** of the state where both guards test **object-id inequality**, so an amend, rebase or reset still above `$BASE` matches no row — **pass 36's fix produced it**, one established regeneration lineage. MINOR: accounting row 7 still says "four shapes", falsified by the fifth row — the stale-count sweep keyed on *topologies*, this site says *shapes*. NIT: Self-Review's reader-check sentence omits Task 12b and Task 15 step 4b. Two tells (findings rose a third pass, instrument cluster) → mandatory stop | | — | — | — | — | — | — | **FIFTH BOUNDED REVISION — a DECISION, not another description.** Daniel: **the plan does not classify the repository.** §A re-establishes the conditions *against the repository as it now stands*; the demand that Resume fit the state to a named shape was plan-invented and is **dropped**, recorded in accounting row 7 as a *requirement* (so row 40's count of dropped *conditions* stays true). The table is explicitly non-exhaustive illustration; nothing depends on fitting a row. Preserved: both equality checks, the precondition/handoff split, the scratch rules, the bounded handoff, base validity checked not inferred, §A's three routes. Strengthened: all four sources read every time, and no rewriting the recorded head or tip to make equality pass. 64 checks × 3 shells — descendant, amended, rebased, backward, each before and after 8a. Commit `b283250` | | 38 | b283250 | 3→**3** | 0→**1** | 1→**2** | yes | **UNCLEAN — reported to Daniel, no repair round opened.** **BLOCKER: Resume never checks the branch** — Preparation does and is first-entry-only, the ignored base file survives a checkout, and another branch descended from `$BASE` passes everything; pre-existing, and **thirty-seven passes never looked**. MAJOR: this round's own regression — "present is not reviewed" routes *every* mismatch to Failure, contradicting the precondition/handoff split four paragraphs earlier. MAJOR: condition 6 diffs the landed body against the **same mutable file** the commit read, after hooks ran. Two tells (Blockers rose, instrument cluster) → mandatory stop | +| — | — | — | — | — | — | **SIXTH BOUNDED REVISION**, Daniel's assignment after pass 38, narrowed by his reviewer: repair the three findings, **and first run a read-only reconciliation of the accounting table** — but with the right question (*which conditions must still hold at re-entry, how is validity established, and does it happen before the first action that depends on it*), not "which row is unchecked", since rows 2 and 9 are first-entry-only by design. Reconciliation found **no further operational gap**, so the package was not widened. Repairs: Resume checks the branch first (row 1 split by entry like row 4); the mismatch routes by the check that rejected it; 8b pins the validated message bytes and condition 6 compares against that copy. 13 new checks + 111 re-run × 3 shells. Commit `df9123a` | +| 39 | df9123a | 3→**2** | 1→**1** | 2→**1** | yes | **UNCLEAN — reported to Daniel, no repair round opened.** **All three pass-38 repairs held** — first round in four with no descendant finding. **BLOCKER: Task 15 step 7 commits the repair before the ordering has spoken**, so a membership or stop answer arrives after the fix is already in the artifact and nothing removes it — the loop its own §A product forbids. MAJOR: step 4b repairs prompt text and commits only the plan, leaving the repair outside `$BASE..HEAD`. **Second consecutive pass finding a Blocker in a previously unflagged area.** Two tells (Blockers flat at 1, instrument cluster) → mandatory stop | | 26 | 05ec1b2 | **5** | **1** | 2 | yes | first pass on the revised plan. Resume audited progress by `WIP:` commits and called any non-`WIP` commit a stale base — the handoff leaves two other shapes, one of them `HEAD` **at** the base with everything in the index | | 27 | f69db7d | 5→**6** | 1→**3** | 2→**1** | yes | the close required the commit's **tree to equal the tip's**, which target §I **parks**; `reset --soft` leaves the index untouched, so the prose beside it was wrong about git; Resume still inferred staleness from a commit subject where §A3 names that state as reachable | | 28 | f2e5dda | 6→**6** | 3→**2** | 1→**2** | yes | pass 27's stale-base fix reached one site of three; its reset paragraph stated the index behaviour correctly and repeated the false claim four lines later; nothing checked the closing commit's **parent**; the message was validated in the file, where `commit-msg` hooks rewrite git's copy after `-F` reads it | From d4a87c6e1c5c43e15ca8242fc3df0b54131fa028 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Wed, 16 Sep 2026 12:28:43 +0200 Subject: [PATCH 159/181] docs(assessment): record the OpenWolf context and memory assessment NOT part of the loop-rule-consolidation Gate-A cycle. A separate task, authorized by Daniel on 2026-09-16, committed here on its own so it stays identifiable and is not mixed into the plan repairs that follow. Three files, exactly as they stood in the worktree: - docs/openwolf-assessment.md, new - the evidence, alternatives and proposed evaluation criteria; - todos.md, one parked trigger-gated backlog entry pointing at it; - the loop-rule-consolidation plan, four lines of informational reference that state in their own text that they add no task, prerequisite or closure condition to the plan. Gate classification, surfaced rather than decided here: the two docs/**.md paths are prose and N/A per CLAUDE.md. `todos.md` is not in that rule's enumerated prose list, and it is not a prompt either. Whoever owns this task decides whether it owed a gate; nothing here claims one was run. --- docs/openwolf-assessment.md | 301 ++++++++++++++++++ .../2026-09-14-loop-rule-consolidation.md | 4 + todos.md | 7 + 3 files changed, 312 insertions(+) create mode 100644 docs/openwolf-assessment.md diff --git a/docs/openwolf-assessment.md b/docs/openwolf-assessment.md new file mode 100644 index 0000000..62e7f92 --- /dev/null +++ b/docs/openwolf-assessment.md @@ -0,0 +1,301 @@ +# OpenWolf assessment — context, memory and the dev-workflow kit + +**Assessment date:** 2026-09-16. **Status:** advisory documentation draft; no +integration or pilot approved. This is a project-local assessment, not a shipped +prompt, execution plan, or source of permanent workflow rules. + +## Authorization and source snapshot + +Daniel requested an evidence-based comparison, then asked to preserve the +recommendation and its trade-offs alongside the current plan. He subsequently +authorized the advisor to make these documentation edits directly. That authorizes +this document, an informational plan link, and a parked backlog reference. It does +not authorize installation, a pilot, product implementation, another repair round, +or a new gate call. The recommendations below remain the advisor's recommendations; +recording them does not turn them into approved implementation decisions. + +The original comparison inspected local HEAD `df9123a`. Before writing this record, +the advisor rechecked the clean worktree on `loop-rule-consolidation` at +`c1fcdd1ded1860998641808b203aafaf728549e0`. The latest committed plan revision was +`df9123a68cd85e4b54e1c0a2923e4ac94402d98c`. [Pass 39][kit-pass39] reports one Blocker +and one Major; the [working record][kit-resume] leaves cycle `om0bdd7udh` open and +unclean, awaiting Daniel's decision. This is a dated observation, not a second +maintained cycle-status record. The actual repository and gate artifacts govern +later work. The added plan reference was not part of the revision reviewed in pass 39. + +External sources: [OpenWolf's website][ow-site], [repository][ow-repo], and source at +commit `04b75ca9c40d10c345ae4f5157e033d7397f1b7a`, assessed alongside published version +2.5.2. Code links below pin that commit rather than following `main`. Local source +links refer to paths inspected at `c1fcdd1`; later edits to those paths do not +retroactively update this assessment. + +**Evidence boundary:** the advisor inspected source, documentation and tests, and +confirmed the successful [published CI run][ow-ci]. The advisor did not install +OpenWolf, execute its suite, or reproduce its native-agent sessions. Test counts +and live-session results below are attributed to the release report. Static +observations and their implications are distinguished from executed checks. + +## Recommendation and alternatives + +The advisor recommends using selected OpenWolf mechanisms as design references, +and considering a separately authorized external evaluation. Full incorporation +or installation with unchanged defaults is not recommended for the current kit. +OpenWolf manages context transport and visibility; our independent reviews, +closure conditions, quality command and hardening procedure retain their roles. + +| Candidate | Existing solution | Expected benefit | Cost or limitation | Advisor recommendation | Evidence / what could change the recommendation | +|---|---|---|---|---|---| +| Entire OpenWolf runtime | Prompt-based workflow plus one shipped POSIX hook | Automated context lifecycle across supported agents | Node runtime, dependencies, adapters, instruction conflicts and update policy | Do not incorporate wholesale | [Architecture][kit-agents], [package][ow-package]; reconsider only with demonstrated benefit and an explicit integration design | +| Evidence-linked handover | Optional resume companions and cycle identity | Faster reconstruction with inspectable sources and explicit gaps | Evidence may be incomplete; exported fields do not all enter active receiver state | Highest-value design inspiration | [Handover][ow-handover]; a pilot must preserve constraints and unresolved work across interruption | +| Memory and bug retrieval | Hardening ledger and recurrence procedure | Find relevant prior mistakes without scanning the entire history | Learned text can acquire unintended authority; a second ledger can diverge | Retrieve existing evidence; retain existing rule and ledger ownership | [Protocol][ow-protocol], [hardening][kit-hardening]; reconsider only with a clear authority boundary | +| Another code index | Codebase-Memory-MCP available in Daniel's environment | Compact navigation in other environments | Duplicate indexing and maintenance; limited extraction is not complete semantic analysis | No second index in the kit now | [Scanner][ow-scanner], [symbols][ow-symbols]; a concrete discovery gap could justify evaluation | +| Automatic output reduction | File-first findings and targeted reads | Less context spent on large outputs | Removed text can contain the condition or finding being reviewed | Exclude automatic evidence truncation from a proposed review-path pilot | [Governor][ow-governor], [output hook][ow-output]; benefit requires preserved review completeness | +| Archive, journal and hook health | Cycle-specific paths; locally tracked review files | Recoverable older context and visible runtime failures | Technical persistence does not establish semantic correctness or gate validity | Reuse ideas when separately commissioned | [Archive][ow-archive], [journal][ow-journal], [heartbeat][ow-shared]; a concrete loss or diagnostic need would motivate work | +| Usage reporting | Review curves and ledger; no comparable provider-usage collector in the shipped hook | Observe actual resource use separately from estimates | Harness coverage, attribution, collection overhead and no counterfactual baseline | Best candidate for an optional external evaluation | [Usage][ow-usage], [estimates][ow-estimates]; require measured net benefit at comparable quality | + +This is not evidence that context loss caused the current long plan cycle. The +latest findings concern sequencing and staging. Better context transport might +help an agent work on them; it does not establish correct rule interactions or +adequate review coverage. [Pass 39][kit-pass39] supports that narrower diagnosis. + +## Findings and trade-offs + +### Handover: references help, summaries do not confer approval + +**Source facts.** OpenWolf exports source-linked events and agent-authored +checkpoints, identifies repository/worktree and destination agent, records gaps +and omitted events, and labels the packet as untrusted historical evidence. +Import checks sources and repository state; validation must be rerun against the +receiving worktree. Retrieval and active-state injection can expose a bounded +selection rather than an entire transcript. [Sources][ow-sources], +[handover service][ow-handover], [active state][ow-state]. + +**Static limitation.** Export includes checkpoint `constraints` and `completed`. +The `importPacket` patch merges objective, next action and unresolved items, but +does not merge those two fields into the receiver's active state. They remain in +the stored packet. Packet availability therefore does not prove that every +restriction reaches the receiving agent's active context. This was observed in +source, not reproduced in a running harness. [Handover service][ow-handover]. + +**Interpretation.** The useful pattern is a short map to evidence, with provenance +and omissions, rather than another authoritative account of the project. Our +[optional companions and nonce rules][kit-claude] already cover parts of this +problem. The [shipped hook registration][kit-hooks] has no SessionStart, +PreCompact or Stop checkpoint lifecycle comparable to OpenWolf's. + +**Recommendation.** Preserve original sources and distinguish user constraints, +agent conclusions and historical test results. Shared context may help independent +reviewers locate evidence; an author's conclusions must not become their accepted +answer. A restored packet cannot restore gate approval by itself. + +### Memory authority: implementation and installed instructions differ + +**Source facts.** OpenWolf's protected-memory path requires a protected verifier +and an independently provisioned approval snapshot. An ordinary user-writable +installation does not satisfy that authority check. [Trusted memory][ow-trust]. +However, the shipped protocol tells the agent to use STATUS.md instead of +reconstructing context from plans/code, respect cerebrum entries, and avoid +rereading unchanged files. Those instructions are broader than treating memory as +untrusted evidence. Initialization rewrites OpenWolf's own protocol and Claude +rule file. [Protocol][ow-protocol], [initialization][ow-init]. + +**Interpretation.** Guarding automatic instruction reinjection does not eliminate +the authority conflict in separately installed prompts. A mistaken summary or +agent-authored learning could steer work away from current authoritative text. +This is a documented instruction conflict, not a demonstrated exploit. + +**Recommendation.** Memories may suggest sources and hypotheses. Permanent rules +still need deliberate review in their established homes. A safe evaluation must +inspect installed instructions, not just configuration switches. Custom edits +also incur maintenance if a later initialization overwrites them; our own +[workflow-init][kit-init] handles differing project files through an explicit +comparison and decision. + +### Existing knowledge: avoid a parallel index or ledger + +**Source facts.** OpenWolf builds file descriptions, bounded symbol outlines and +import-based navigation. Its scanner distinguishes incomplete refreshes from a +complete scan and preserves unseen entries during partial work. +[Scanner][ow-scanner], [symbol extraction][ow-symbols]. Codebase-Memory-MCP was +available and used in this advisory session. That is an environment capability, +not a feature installed by the kit's [MCP configuration][kit-mcp]. + +OpenWolf also records bugs for retrieval. Our [hardening procedure][kit-hardening] +instead checks recurrence, proposes a fitting stronger rung, verifies changes +and records the hardening. Neither the presence of a bug entry nor retrieval of +one performs those steps. [Bug tracker][ow-bugs]. + +**Recommendation.** Reuse existing discovery and make prior ledger evidence easier +to find before adding another index or authoritative bug collection. Durable +delivery of a fixed finding to the ledger is already owned by [Finding A][story-a]; +its identity, deduplication and consumption questions remain real design work. + +### Output reduction and repository snapshots: retain our review semantics + +**Source facts.** OpenWolf's default governor replaces selected large file-print, +grep and git-show outputs; test/build families default to suggestions. Condensation +can remove diff hunks or retain only selected parts of longer output. The output +hook checks whether the original survived caching, but can still emit condensed +text when preservation failed, without claiming a recoverable copy. Actual +replacement depends on harness support. [Configuration][ow-config], +[governor][ow-governor], [output hook][ow-output]. + +**Interpretation and recommendation.** This may help navigation, but an omitted +middle section can contain a necessary finding or predicate. Keep the +[file-first findings protocol][kit-claude] and exact-source access; evaluate query +hints rather than automatic truncation on the review path. The runtime's default +duplicate-read mode is a warning, which is narrower than the blanket reread +instruction in its protocol. [Read hook][ow-read], [protocol][ow-protocol]. + +**Separate static finding.** OpenWolf's repository snapshot combines HEAD-related +diff data, dirty-path discovery and contents read from disk. Staged paths are +enumerated, but staged blob contents are not independently hashed as an index +tree. Different staged contents at the same dirty path, with identical HEAD and +worktree, can therefore yield the same snapshot. This inference was not exercised +in a disposable repository. [Repository snapshot][ow-repository]. + +Our [tree_hash implementation][kit-hook] includes a tree of the effective index +(honoring GIT_INDEX_FILE), a worktree tree and tracked diff data, excluding +`.context/`. Retain that comparison for the hook's invalidation decision. It still +does not prove that a reviewer read a particular revision or that the review was +complete; those claims exceed the comparison. + +### Persistence and health: useful engineering patterns with bounded claims + +**Source facts.** OpenWolf archives older memory blocks by content hash, verifies +the archive before replacing active text with a pointer, and checks content and +marker when restoring. Its event journal persists events and uses event identities +when reconciling them. Hook heartbeats record success, error and consecutive +failure information even though hooks exit without blocking the agent. +[Archive][ow-archive], [journal][ow-journal], [heartbeat][ow-shared]. + +**Interpretation.** These are useful patterns for retaining original evidence and +making failed automation visible. They do not establish that a finding was +semantically processed exactly once, or that a gate passed. OpenWolf's memory +archive does not automatically archive our `.context/codex-reviews` artifacts. + +Our nonce paths and locally tracked review directory already mitigate some loss. +The [.gitignore][kit-ignore] explicitly documents that tracking review files is a +local divergence from the shipped template. A historical story saying all such +files are ignored is not the present local behavior. Further persistence work +belongs to an explicitly scoped decision; [record durability][story-durability] +does not silently acquire runtime machinery or a changed trigger from this memo. + +### Measurement, maintenance and remaining evidence + +**Source facts.** OpenWolf separates transcript-derived provider usage, including +coverage/availability, from estimated token effects of its interventions. +[Usage][ow-usage], [estimate accounting][ow-estimates]. Our review curves are +explicitly author-written and unchecked in [CLAUDE.md][kit-claude]; they are not +provider telemetry. [P8][story-p8] permits read-only ledger/git analysis and +explicitly excludes new instrumentation. + +**Interpretation.** Recorded consumption alone cannot establish how much the same +task would have consumed without OpenWolf. API price estimates are not necessarily +subscription expenditure; token usage is not remaining context capacity. An +evaluation needs a baseline and comparable task quality, including the context +cost of memory injection and collection. + +**Adoption costs.** OpenWolf adds a Node 20+ runtime and dependencies to a kit whose +shipped executable is a POSIX hook. Its default compatible-update policy can select +new runtimes between sessions, despite pinning within a session. That is not our +deliberate exact-version update policy. OpenWolf declares AGPL-3.0-only; this kit +uses MIT. Direct code reuse needs a separate licensing assessment; this memo makes +no legal compatibility determination. [Package][ow-package], [updates][ow-updates], +[our invariants][kit-agents], [our license][kit-license]. + +**Reported, not reproduced by the advisor.** The 2.5.2 release report describes +322 tests in 72 suites, package/install checks and native Codex recovery after +resume/compaction. It also states that the full live Claude/Codex round trip and +paired long-session quality/token-efficiency evaluation remain outstanding; +Claude model access prevented a successful coding turn. Automatic handover import +stays off. The advisor confirmed the referenced CI run's successful status, not +the general effectiveness of these features in this workflow. [Release report][ow-release]. + +## Proposed evaluation — parked, not executable + +**Activation requires Daniel's explicit authorization of a bounded evaluation.** +Finishing loop-rule-consolidation alone does not activate it. The eventual scope, +project and evaluation budget remain undecided. These are candidate criteria for +that decision, not additional acceptance criteria for the current plan: + +- Use an isolated test project and an exact runtime version, with automatic updates + off and installed instructions inspected for conflicts with project authority. +- Disable automatic output replacement; retain complete review evidence and the + existing gates. If configuration cannot isolate those behaviors, report the + adaptation cost before choosing an adapter or fork. +- Compare against existing resume artifacts and graph-assisted discovery. Keep + task, model/harness and validation comparable and record differences explicitly. +- Exercise interruption/compaction: check objective, constraints and unresolved + findings against their sources; detect stale or missing sources without reviving + earlier gate approval. Check journal replay separately from semantic consumption. +- Observe total recorded usage, retrieval effort and review completeness. Include + memory overhead and unavailable measurements; do not rename estimated savings + as measured improvement. +- Report lost constraints, concealed evidence, unintended instruction changes or + falsely restored approval as reasons to stop the trial and reassess. Reduced + token use alone is insufficient if quality declines. + +The evaluation could support optional external use, an independently designed +small feature, or a decision to adopt nothing. No outcome is assumed here. + +## Existing work and review boundary + +| Existing owner | Relationship | Boundary retained | +|---|---|---| +| [Finding A][story-a] | Fixed findings surviving handoff before ledger consideration | Preserve its exact scope and design questions; do not introduce a second route | +| [Record durability][story-durability] | Evidence surviving session and Git boundaries | Existing trigger, proposed profile and exclusions remain; this memo does not evaluate whether its trigger fired | +| [P8][story-p8] | Analysis of existing ledger/git records | New transcript collection is additional scope, not an already-authorized P8 implementation | +| [Current consolidation plan][kit-plan] | Discovery link to this assessment | No new task, prerequisite, closure condition or repair authorization | + +The documentation request preserves a decision basis; it does not approve the +assessment as a spec or settle pass 39. Existing review requirements remain. +The memo and plan-link amendment are not covered by the old pass merely because +their purpose is informational. Until committed through the applicable workflow, +these files are worktree drafts rather than durable Git history. + +## Sources + +Local references are relative to this document. External code references pin the +assessed revision. The website is background context, not authority for a code claim. + +[kit-agents]: ../AGENTS.md +[kit-claude]: ../CLAUDE.md +[kit-plan]: superpowers/plans/2026-09-14-loop-rule-consolidation.md +[kit-resume]: ../.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md +[kit-pass39]: ../.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-39.md +[kit-hooks]: ../plugins/dev-workflow/hooks/hooks.json +[kit-hook]: ../plugins/dev-workflow/hooks/codex-gate.sh +[kit-init]: ../plugins/dev-workflow/commands/workflow-init.md +[kit-hardening]: ../plugins/dev-workflow/skills/harden-finding/SKILL.md +[kit-mcp]: ../.mcp.json +[kit-ignore]: ../.gitignore +[kit-license]: ../LICENSE +[story-a]: superpowers/stories/2026-08-04-ledger-route-without-pull-requests-story.md +[story-durability]: superpowers/stories/2026-09-10-record-durability-story.md +[story-p8]: superpowers/stories/2026-08-04-passive-metrics-over-the-ledger-story.md +[ow-site]: https://openwolf.com/ +[ow-repo]: https://github.com/cytostack/openwolf +[ow-ci]: https://github.com/cytostack/openwolf/actions/runs/34894539164 +[ow-package]: https://github.com/cytostack/openwolf/blob/04b75ca9c40d10c345ae4f5157e033d7397f1b7a/package.json +[ow-handover]: https://github.com/cytostack/openwolf/blob/04b75ca9c40d10c345ae4f5157e033d7397f1b7a/src/handoff/service.ts +[ow-sources]: https://github.com/cytostack/openwolf/blob/04b75ca9c40d10c345ae4f5157e033d7397f1b7a/src/handoff/sources.ts +[ow-state]: https://github.com/cytostack/openwolf/blob/04b75ca9c40d10c345ae4f5157e033d7397f1b7a/src/hooks/handoff-state.ts +[ow-trust]: https://github.com/cytostack/openwolf/blob/04b75ca9c40d10c345ae4f5157e033d7397f1b7a/src/hooks/trusted-memory.ts +[ow-protocol]: https://github.com/cytostack/openwolf/blob/04b75ca9c40d10c345ae4f5157e033d7397f1b7a/src/templates/OPENWOLF.md +[ow-init]: https://github.com/cytostack/openwolf/blob/04b75ca9c40d10c345ae4f5157e033d7397f1b7a/src/cli/init.ts +[ow-scanner]: https://github.com/cytostack/openwolf/blob/04b75ca9c40d10c345ae4f5157e033d7397f1b7a/src/scanner/anatomy-scanner.ts +[ow-symbols]: https://github.com/cytostack/openwolf/blob/04b75ca9c40d10c345ae4f5157e033d7397f1b7a/src/anatomy/ts-symbol-extractor.ts +[ow-bugs]: https://github.com/cytostack/openwolf/blob/04b75ca9c40d10c345ae4f5157e033d7397f1b7a/src/buglog/bug-tracker.ts +[ow-config]: https://github.com/cytostack/openwolf/blob/04b75ca9c40d10c345ae4f5157e033d7397f1b7a/src/templates/config.json +[ow-governor]: https://github.com/cytostack/openwolf/blob/04b75ca9c40d10c345ae4f5157e033d7397f1b7a/src/hooks/bash-output-governor.ts +[ow-output]: https://github.com/cytostack/openwolf/blob/04b75ca9c40d10c345ae4f5157e033d7397f1b7a/src/hooks/post-bash.ts +[ow-read]: https://github.com/cytostack/openwolf/blob/04b75ca9c40d10c345ae4f5157e033d7397f1b7a/src/hooks/pre-read.ts +[ow-repository]: https://github.com/cytostack/openwolf/blob/04b75ca9c40d10c345ae4f5157e033d7397f1b7a/src/handoff/repository.ts +[ow-archive]: https://github.com/cytostack/openwolf/blob/04b75ca9c40d10c345ae4f5157e033d7397f1b7a/src/hooks/memory-archive.ts +[ow-journal]: https://github.com/cytostack/openwolf/blob/04b75ca9c40d10c345ae4f5157e033d7397f1b7a/src/hooks/event-journal.ts +[ow-shared]: https://github.com/cytostack/openwolf/blob/04b75ca9c40d10c345ae4f5157e033d7397f1b7a/src/hooks/shared.ts +[ow-usage]: https://github.com/cytostack/openwolf/blob/04b75ca9c40d10c345ae4f5157e033d7397f1b7a/src/tracker/usage.ts +[ow-estimates]: https://github.com/cytostack/openwolf/blob/04b75ca9c40d10c345ae4f5157e033d7397f1b7a/src/hooks/ledger-math.ts +[ow-updates]: https://github.com/cytostack/openwolf/blob/04b75ca9c40d10c345ae4f5157e033d7397f1b7a/src/hooks/runtime-updates.ts +[ow-release]: https://github.com/cytostack/openwolf/blob/04b75ca9c40d10c345ae4f5157e033d7397f1b7a/docs/release-2.5.2.md diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index 38ffdab..00e8857 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -12,6 +12,10 @@ **Story:** `docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md` — read the profile from its header at every gate call; it is the only writable copy. Six acceptance criteria; §4 holds settled decisions D1–D8. +**Related assessment (informational):** [OpenWolf context and memory assessment](../../openwolf-assessment.md) +records the advisor's evidence and trade-offs for a possible future evaluation requiring Daniel's separate authorization. +It adds no task, prerequisite or closure condition to this plan. + --- ## Global Constraints diff --git a/todos.md b/todos.md index cd92822..4aff5b1 100644 --- a/todos.md +++ b/todos.md @@ -24,6 +24,13 @@ driven by recurrence rather than by enthusiasm. ### Parked (trigger-gated) +- [ ] **OpenWolf: possible bounded context/memory evaluation.** The + [assessment](docs/openwolf-assessment.md) records the evidence, alternatives, + trade-offs and proposed evaluation criteria. Documentation authorized by + Daniel on 2026-09-16; installation, pilot and integration remain undecided. + *Trigger: Daniel explicitly authorizes a bounded evaluation. Completion of + loop-rule-consolidation alone does not activate it.* Existing Finding A, + record-durability and P8 scopes and triggers remain unchanged. - [ ] **Locator: TWO quadratic paths — `skipval`'s container walk and the record accumulator.** `substr(s,i,1)` is O(len) per call in BWK awk, so a large VALID sibling container before `tool_response` is quadratic: 3.2 s at 200 KB, 11.5 s at 400 KB, in one synchronous hook invocation. From c5d39be58d4f3a93d32eff6a4785579c1e863de2 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Wed, 16 Sep 2026 12:31:46 +0200 Subject: [PATCH 160/181] docs(plans): record, then read the ordering, then repair; 4b stages the repair Pass 39's Blocker and Major. Bounded to those two; no reconciliation of the rest of Task 15 was run and none is claimed. Blocker. Step 7 committed the fix in step one and read the pass through the installed ordering in step two, so a membership stop, a new-question stop or a stop answer arrived after the repair was already in the artifact, and this plan restores nothing. The step is three steps now: record the pass, read it against the ordering, then repair only on a route that authorizes one. Deferring the commit alone would not have fixed it, so step one says not to apply the repair in the worktree either, and stages the records by name instead of git add -A, which would sweep exactly that edit. Two route rules are preserved rather than flattened, both read from the approved text first. The source-block branch keeps its own repair path - section A sends the reader to the source and does not hold that repair behind a continue - so no blanket "every change only after continue" rule was created. And stop parks the open cycle and prescribes no rollback, so parking is not a retroactive revocation of a repair an earlier route had authorized. Where continue is reached with a repair owed, the repair comes before the post-answer pass, which is section A's rule and not this plan's. The escalation boundary is stated where it belongs: if a repair was applied before its route authorized it, report the concrete state and stop. Do not design a way to take it back. Major. Step 4b repaired prompt text and staged only the plan, leaving the repair outside $BASE..HEAD, where Gate B never saw it and the close's exact-dirty-set check then refused it. It now stages the repaired artifacts by name, from the set this step judges - the two prompt copies, the hook and its test - together with the refreshed record, and says why a blanket git add -A is the wrong fix. One directly dependent reference moved with them: step 7b's evidence-entry branch said "commit the change, then route", the same inversion. Verified by execution under sh, dash and bash, 17 checks each: step one leaves a premature repair loose and uncommitted while committing the records; a pass owing no repair still keeps them; step three commits the named repair and records the reviewed head at that commit, not the one before it; step 4b puts a repaired prompt file into the range, with a control showing the old command left it dirty and outside. Unrelated dirty work stayed out of every one of those commits. The four earlier suites re-run green. The route rules themselves are reader instructions and were asserted as text, 16 checks, not executed. --- .../2026-09-14-loop-rule-consolidation.md | 89 ++++++++++++++----- 1 file changed, 67 insertions(+), 22 deletions(-) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index 00e8857..345ee1f 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -2688,16 +2688,31 @@ outside invariant 11's only reader gate. **Twelve `PASS` lines alone cannot be told from a review that skipped a copy or a channel** — the subject list is what makes the twelve lines mean something. -**Commit the result before step 6.** Gate B reviews the range `$BASE..HEAD`; an edit to this plan -left in the worktree is in neither that range nor the final `reset --soft`, which stages only what -the discarded commits contained. The same applies to every record this plan collects — the sweep, the -next-state table, the divergence list: +**Commit the result before step 6 — and on a failure, commit the repaired files with it.** Gate B +reviews the range `$BASE..HEAD`; an edit left in the worktree is in neither that range nor the final +`reset --soft`, which stages only what the discarded commits contained. **An earlier revision staged +only this plan**, so a repair made under the failure branch above stayed in the worktree: Gate B +never saw it, and the close's exact-dirty-set check then refuses a dirty prompt copy, so the cycle +dead-ends. Staging it *after* the review instead breaks the reviewed-head condition — there is no +later moment that works. + +**Stage the repaired artifacts by name, and only those.** `git add -A` would carry unrelated work in +the tree into a range a reviewer is about to read as this change. The repairable set here is the +installed text this step judges: the two prompt copies, the hook and its test. **Only the ones the +repair actually touched go in**, together with this plan's refreshed record — and the re-run records +must describe *that* commit, which is why the re-runs above come first. ```bash -git add docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md -git commit -m "WIP: plan records" +# add only what the repair touched, from: CLAUDE.md, +# plugins/dev-workflow/commands/workflow-init.md, +# plugins/dev-workflow/hooks/codex-gate.sh, plugins/dev-workflow/hooks/codex-gate.test.sh +git add docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +git commit -m "WIP: plan records" # or: "WIP: prompt-standards repair + plan records" ``` +**The same applies to every record this plan collects** — the sweep, the next-state table, the +divergence list. + **`check-invariants.sh` includes the prompt-conformance checks** — a `Target model:` line naming one recognized model, a prose checklist-count claim matching the checklist, and the finding-severity vocabulary as a closed set in both prompt copies. Those three are a floor, not coverage; invariant 11's other eleven items are judged by a reader. - [ ] **Step 5: Write the evidence entry into `.context/loop-rule-closing-msg`** @@ -2776,39 +2791,68 @@ WIP tip while the repair sits in the worktree — and the final squash then publ reviewed: **Every pass that does not close ends the same way, whether or not it produced a repair — and it -ends in two steps, in this order.** First record the pass. Then **run it through the ordering this -change installs, and only its continue result prepares another call.** The plan must not execute a -loop its own product forbids. +ends in three steps, in this order.** Record the pass. Read it through the ordering this change +installs. **Then, and only on a route that authorizes it, repair.** The plan must not execute a loop +its own product forbids. -**Step one — record the pass. This commits; it does not authorize anything.** +**Step one — record the pass, and nothing else. This commits; it does not authorize anything.** + +**Do not apply the repair yet — not in the worktree either.** Deferring only the *commit* changes +nothing: this plan restores nothing, so an edit made before its route authorized it is just as +unauthorized, and a later decline has nothing that removes it. **Stage the records by name**; `git +add -A` here would sweep exactly the edit this step exists to keep out, along with any unrelated +work in the tree. ```bash -# Before committing a fix, re-run what the fix could have broken. -git add -A && git commit -m "WIP: fix " # or: "WIP: pass records" where no repair was owed +git add .context/codex-reviews/ docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +git commit -m "WIP: pass records" ``` **Step two — read the pass against the ordering, before any next call exists:** -- **A source block standing** → **wait** for the repair and the reread by the route §A gives. No - further pass until that is done. +- **A source block standing** → **the source rule decides what must be repaired or answered, and its + repair happens on that route** — §A sends the reader to the source and does not hold that repair + behind a continue. Then §A's release rule decides whether this pass is read again. **No further + pass while the block stands.** - **Any suspension open** — a membership stop, a new-question stop, a two-tell stop, a clearly-stuck - surface — → **collect every answer and compose them.** Another pass only where the composition - yields **continue**. + surface — → **collect every answer and compose them.** **No repair to a finding whose membership + is unanswered**: §A answers membership against the fix set *as it stood for the pass that raised + the question*, so repairing first decides the question the stop exists to ask. Another pass only + where the composition yields **continue**. - **A stop answer** → **park**: open, not running, **spending no passes**, restarted only by an explicit later continue. **There is no "commit and carry on" from a stop**, and acceptance criterion 4 requires that parked state to be distinct. **Nothing below runs on this route.** + **Parking is not a retroactive revocation**: §A defines stop as parking the open cycle and + prescribes no rollback, so a repair some earlier route had already authorized stays where it is. - **The clean-completion branch** → **Close**, not another pass. -- **Continue** → and only then: +- **Continue** → step three. + +**Step three — repair where the route authorized one, then prepare the call.** Continue permits an +**unrevised** artifact only where no repair is owed; **where one is owed it comes before the +post-answer pass**, which is §A's rule and not this plan's. So on this route: apply the repair now, +re-run every check it invalidated, commit both, and only then record the head. ```bash +# Stage the repaired artifacts BY NAME, plus the refreshed records. Not `git add -A`: +# it would carry unrelated work in the tree into a range a pass is about to review. +git add .context/codex-reviews/ \ + docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +git commit -m "WIP: fix " + rm -f .context/loop-rule-reviewed-tip # the previous candidate's closing tip git rev-parse HEAD > .context/loop-rule-reviewed-head # the head the NEXT call is issued against cat .context/loop-rule-reviewed-head ``` -**The reviewed-head file is written on the continue route alone**, because writing it is what makes -a next call possible: recording it before the ordering has spoken is how a pass gets issued over a -standing source block, an unanswered suspension or a parked cycle. +**The reviewed-head file is written on this route alone, and after the repair commit**, because +writing it is what makes a next call possible: recording it before the ordering has spoken is how a +pass gets issued over a standing source block, an unanswered suspension or a parked cycle — and +recording it before the repair lands would aim that call at a tree the repair is not in. + +**If a repair was already applied before its route authorized it, stop and report the concrete +state.** Do not design a way to take it back: this plan restores nothing, §A prescribes no rollback, +and inventing one here would be a new rule nobody approved. Report what is in the worktree, what is +committed, and which route had not yet spoken. **A non-closing pass that owes no repair still commits.** A Minor-only clean pass below the floor, or an answered suspension that changes no artifact, leaves its findings files tracked and dirty — @@ -2889,8 +2933,9 @@ The message carries, in this order: 3. any **human-exception record**, and beside it the skip reason if a cycle was skipped; 4. the **revalidated evidence entry**, replacing step 5's draft if revalidation changed it — and **if it changed, this candidate is over.** The final reviewer judged the entry it was handed - verbatim; a different entry in the closing commit is evidence no pass covered. **Commit the - change, then route the pass through step 7's ordering like any other non-closing pass** — a + verbatim; a different entry in the closing commit is evidence no pass covered. **Take this pass + through step 7's three steps like any other non-closing pass** — record it, read it against the + ordering, and repair only on a route that authorizes one; a standing source block, an open suspension or a stop answer binds here exactly as it does there, and **only its continue result records a new reviewed head and issues another candidate.** An earlier draft sent this branch straight to a new call, which is the plan's own Gate-B loop From 0a15f6d0eca7d3f027ad1bb1875897b8b07d35b6 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Wed, 16 Sep 2026 12:42:13 +0200 Subject: [PATCH 161/181] docs(context): record Gate-A plan pass 40 and the seventh bounded revision Pass 40 reviewed c5d39be: 1 Blocker, 1 Major, 0 Minors. Blocker. Task 0 step 3 branches on whether $BASE..HEAD is empty to decide whether to read the parity baseline from the $BASE blobs or from the worktree. After an 8b rejection that range is empty while the whole implementation is staged and present in the worktree - a topology Resume names as valid - so the else branch records the implemented copies as inherited drift, Task 14 can no longer tell introduced drift from inherited, and the cycle can close on false parity evidence. The same claim is stated correctly three times in Resume, where passes 29-30 repaired it, and is still used as a decision twice in Task 0 step 3. One fix, one site, never swept - the class AGENTS.md names, and greppable. Major, and it is this round's own repair: step 7 steps one and three say "stage the records by name" and then pass git add .context/codex-reviews/, a directory pathspec that stages everything under it. Another cycle's findings files are swept into the WIP commit and then the closing squash, before the next call, so the exact-dirty-set guard cannot see it. My verification missed it because it tested unrelated work using a file at the repo root, never a second findings file inside that directory. Demonstrated afterwards by execution. Three consecutive passes, each with exactly one Blocker, each in an area no earlier pass had flagged: Resume's branch check, Task 15 step 7, Task 0 step 3. That is what these three passes did, not a prediction about what remains. Commit d4a87c6 in this history is a separate authorized task - the OpenWolf assessment - committed on its own so it stays identifiable, with its gate classification surfaced rather than decided. A new per-pass lesson is recorded: git add

is not staging by name, and a sweep test belongs inside the directory. Floor 3, from docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md (risk high, security none), read fresh at this pass. Cycle stays open and unclean. --- .../gate-a-plan-om0bdd7udh-pass-40.md | 3 + .../gate-a-plan-om0bdd7udh-resume.md | 107 ++++++++++-------- 2 files changed, 63 insertions(+), 47 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-40.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-40.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-40.md new file mode 100644 index 0000000..3230166 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-40.md @@ -0,0 +1,3 @@ +BLOCKER | high | Task 0 step 3 | The baseline-source branch treats an empty `$BASE..HEAD` log as proof that the worktree is still the clean first-entry source, even though Resume explicitly says an 8b rejection after `reset --soft` has that same empty range while the full implementation is staged and present in the worktree. | If the baseline artifact must be rebuilt after that permitted topology leaves the handoff, this step records the implemented copies as inherited drift, so Task 14 can accept a parity or untouched-text defect introduced by this change and the cycle can close on false evidence. | Build the parity baseline from the recorded `$BASE` blobs on every entry (the first entry records clean `HEAD` as `$BASE` anyway), or persist an entry-mode predicate that distinguishes first entry without inferring it from commit-log shape. +MAJOR | high | Task 15 step 7 steps one and three | Both commands claim to stage records by name but pass the directory path `.context/codex-reviews/` to `git add`, which recursively stages every changed or untracked review artifact under that tracked directory. | Findings from another cycle or unrelated review work can be swept into a WIP commit and ultimately the closing squash; because they are committed before the next Gate-B call, the later exact-dirty-set guard cannot detect the scope leak. | Stage the current pass's two explicit slot paths in step one, and in step three list only the exact refreshed record paths plus the named repair files; do not use the review directory as a pathspec. +END OF FINDINGS (2 total) diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md index e643a32..ecadb71 100644 --- a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md @@ -13,56 +13,63 @@ Nothing depends on it; the pass files and the repo are authoritative where this ## Resume here -**Cycle OPEN and UNCLEAN. The plan stands at `df9123a` and that is the anchor** — later commits on +**Cycle OPEN and UNCLEAN. The plan stands at `c5d39be` and that is the anchor** — later commits on this branch are records, not plan edits, so check `git log --oneline -1 -- docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md` -rather than `HEAD`. Pass 39 reviewed `df9123a` and found **one Blocker and one Major**. -`.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-39.md` holds them. +rather than `HEAD`. Pass 40 reviewed `c5d39be` and found **one Blocker and one Major**. +`.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-40.md` holds them. + +**One unrelated commit sits in this history and is NOT part of this cycle.** `d4a87c6` records an +OpenWolf assessment — a separate task authorized 2026-09-16 — as `docs/openwolf-assessment.md`, a +parked `todos.md` entry, and **four informational lines near the top of this plan** which state in +their own text that they add no task, prerequisite or closure condition. It was committed on its own +so it stays identifiable. **Its gate classification was surfaced, not decided**: the two `docs/**.md` +paths are prose and N/A, `todos.md` is in neither CLAUDE.md's prose list nor the prompt paths, and +whoever owns that task decides whether it owed a gate. **Do not start a repair round on your own.** Daniel's assignment of 2026-09-16 ended with a -checkpoint: *"Danach genau ein vollständiger Gate-A-Plan-Pass und Bericht. Keine automatische -Folgerunde oder Implementierung."* That checkpoint is spent — one reconciliation, one bounded -repair, one pass, all delivered — so **the next move is Daniel's word, not an inference.** - -**Two tells, so this stop is mandatory as well as instructed**: the **Blocker count failed to fall**, -1 → 1 — and this time that is the real test rather than the nonsense-at-zero reading, because there -is a Blocker and it did not go away — and the findings cluster on the instrument for the sixth pass -running. Findings fell 3 → 2 and Majors fell 2 → 1, so neither is a tell. - -**Read this before deciding the next round. Three things.** - -**One, and it is good news: all three pass-38 repairs held.** Nothing was re-raised against the -branch check, the mismatch routing or the pinned message oracle. **That breaks the -repair-produces-the-next-finding chain for the first time in four rounds** — passes 36, 37 and 38 -each found a defect descended from the previous fix; pass 39 found none. - -**Two: a second Blocker, in a second area no earlier pass flagged.** Pass 38's was Resume's missing -branch check; pass 39's is Task 15 step 7's ordering. **Two consecutive passes have each found a -Blocker in a previously unflagged area** — that is the evidence, and it is a statement about what -these two passes did, not a prediction about how many more such areas exist. - -**Three: the accounting-table reconciliation did what it claimed and no more.** It walked the -forty-one rows for re-entry relevance and found the one gap it was aimed at. **It could not have -found either pass-39 finding**, which are about Task 15's step *sequencing* and *staging*, not about -conditions at re-entry. **It is not a coverage certificate for the artifact** and must not be read -as one. - -### The two open findings from pass 39 - -1. **BLOCKER — Task 15 step 7 commits the repair before the ordering has spoken.** Step one is - `git add -A && git commit -m "WIP: fix "`, and step two *then* reads the pass through - the installed ordering — where a membership stop, a new-question stop or a stop answer may say - the finding was never in the assigned fix set. By then the repair is committed, and **this plan - restores nothing**, so no route removes it. The step's own header says "this commits; it does not - authorize anything", which names the tension without resolving it. **The plan's own Gate-B loop - can ship work outside the assigned fix set** — the failure its own §A product forbids, and the - thing this step was written to prevent. Validated against the plan. -2. **MAJOR — Task 15 step 4b repairs the prompt text and commits only the plan.** On a - prompt-standards failure it says to repair the text, re-run the invalidated checks, and commit — - but its command is `git add docs/superpowers/plans/…` alone. The repaired `CLAUDE.md`, - `workflow-init.md` and hook bodies stay in the worktree: **outside `$BASE..HEAD`, so Gate B never - sees them**, and then the close's exact-dirty-set check refuses a dirty prompt copy, so the cycle - dead-ends. Staging them after the review instead breaks the reviewed-head condition. Validated - against the plan. +checkpoint: *"genau ein vollständiger Gate-A-Pass und Bericht. Keine automatische Folgerunde, keine +Produktimplementierung und keine OpenWolf-Evaluation."* Spent — one bounded repair, one pass, both +delivered — so **the next move is Daniel's word, not an inference.** + +**Two tells, so this stop is mandatory as well as instructed**: the **Blocker count failed to fall** +for the third pass running (1 → 1 → 1), and the findings cluster on the instrument for the seventh. +Findings are flat at 2 and Majors flat at 1, so neither is a tell. + +**Read this before deciding the next round. Three things, stated as what was observed.** + +**One: three consecutive passes, each with exactly one Blocker, each in a different area no earlier +pass had flagged.** Pass 38 — Resume's missing branch check. Pass 39 — Task 15 step 7's ordering. +Pass 40 — Task 0 step 3's baseline-source branch. That is a statement about these three passes, not +a prediction about how many such areas remain. + +**Two: the repair-produces-the-next-finding chain resumed.** It broke once, at pass 39. Pass 40's +Major is **this round's own repair**: I wrote "stage the records by name" and then passed +`git add .context/codex-reviews/`, a *directory* pathspec that stages everything under it +recursively. **And my verification missed it** — I tested that unrelated work stayed out using a file +at the repo root, never a second findings file inside that directory. Demonstrated afterwards: a +stray `gate-a-plan-OTHERCYCLE-pass-9.md` is swept straight into the commit. + +**Three, and this one is actionable: pass 40's Blocker is a known claim that was repaired in one +place and never swept for.** *"An empty `$BASE..HEAD` is not proof of anything"* is stated three +times in Resume (lines 539, 617, 622 — repaired at passes 29–30) and **still used as a decision** +twice in Task 0 step 3 (1023, 1049). The plan contradicts itself in its own words. This is +`AGENTS.md`'s named class — **search for the claim, not the phrase** — and it is greppable. + +### The two open findings from pass 40 + +1. **BLOCKER — Task 0 step 3 infers "first entry" from an empty commit range.** Its branch is + `if [ -n "$(git log --oneline "$BASE"..HEAD)" ]` → read the `$BASE` blobs, else read the + worktree. **After an 8b rejection the range is empty and the whole implementation is staged and + present in the worktree** — a topology Resume names as valid — so the `else` branch records the + *implemented* copies as the parity baseline, i.e. as inherited drift. Task 14 then cannot tell + introduced drift from inherited, which is the baseline's whole purpose, and **the cycle can close + on false parity evidence.** Validated against the plan. +2. **MAJOR — this round's own repair. `git add .context/codex-reviews/` is not "by name".** Both + step 7 step one and step three pass the directory, so another cycle's findings files, or an + earlier pass's, are swept into the WIP commit and then into the closing squash — and because they + are committed *before* the next call, the exact-dirty-set guard cannot see the leak. **Fix: the + two explicit slot paths in step one; the exact refreshed record paths plus the named repair files + in step three.** Demonstrated by execution. ### Still collected and deliberately unrepaired (pass 34's Minors and Nit) @@ -157,6 +164,10 @@ what it carries; `.context/gate-a-spec-prompt.md` is the same shape for the spec passage's rows — but a matching id is a search hit, not a defect.** Tasks 3, 5 and 7 already read their sets off the table; Task 7's lists and Task 4 step 4's move list are complete and were left alone deliberately. Replacing them blind would remove requirements. +- **`git add ` is not "staging by name".** A directory pathspec stages everything changed or + untracked beneath it. Pass 40 found this in a repair whose own sentence said "by name". **When a + step names what it stages, list paths — and test the sweep from INSIDE the directory**, not with a + file at the repo root, which is what let it through. - **Ask a scoped question of the accounting table, not a wide one.** "Which row is unchecked at re-entry" is the wrong question — rows 2 and 9 are first-entry-only *by design*, and moving row 2's clean-tree demand to Resume would refuse valid re-entries. The question that works: **which @@ -240,6 +251,8 @@ worth checking before a pass rather than after. | 38 | b283250 | 3→**3** | 0→**1** | 1→**2** | yes | **UNCLEAN — reported to Daniel, no repair round opened.** **BLOCKER: Resume never checks the branch** — Preparation does and is first-entry-only, the ignored base file survives a checkout, and another branch descended from `$BASE` passes everything; pre-existing, and **thirty-seven passes never looked**. MAJOR: this round's own regression — "present is not reviewed" routes *every* mismatch to Failure, contradicting the precondition/handoff split four paragraphs earlier. MAJOR: condition 6 diffs the landed body against the **same mutable file** the commit read, after hooks ran. Two tells (Blockers rose, instrument cluster) → mandatory stop | | — | — | — | — | — | — | **SIXTH BOUNDED REVISION**, Daniel's assignment after pass 38, narrowed by his reviewer: repair the three findings, **and first run a read-only reconciliation of the accounting table** — but with the right question (*which conditions must still hold at re-entry, how is validity established, and does it happen before the first action that depends on it*), not "which row is unchecked", since rows 2 and 9 are first-entry-only by design. Reconciliation found **no further operational gap**, so the package was not widened. Repairs: Resume checks the branch first (row 1 split by entry like row 4); the mismatch routes by the check that rejected it; 8b pins the validated message bytes and condition 6 compares against that copy. 13 new checks + 111 re-run × 3 shells. Commit `df9123a` | | 39 | df9123a | 3→**2** | 1→**1** | 2→**1** | yes | **UNCLEAN — reported to Daniel, no repair round opened.** **All three pass-38 repairs held** — first round in four with no descendant finding. **BLOCKER: Task 15 step 7 commits the repair before the ordering has spoken**, so a membership or stop answer arrives after the fix is already in the artifact and nothing removes it — the loop its own §A product forbids. MAJOR: step 4b repairs prompt text and commits only the plan, leaving the repair outside `$BASE..HEAD`. **Second consecutive pass finding a Blocker in a previously unflagged area.** Two tells (Blockers flat at 1, instrument cluster) → mandatory stop | +| — | — | — | — | — | — | **SEVENTH BOUNDED REVISION**, Daniel's assignment after pass 39, bounded by his reviewer to the two findings with **no Task-15-wide reconciliation**. Step 7 becomes three steps — record, read the ordering, then repair only on an authorizing route — and step one forbids applying the repair in the worktree too, since deferring the commit alone changes nothing where nothing restores. **Two route rules deliberately preserved, read from the approved text first:** the source-block branch keeps its own repair path (no blanket "only after continue"), and stop **parks** without rollback, so it is no retroactive revocation. Step 4b stages the repaired artifacts by name. 17 checks × 3 shells + 16 text assertions. Commit `c5d39be` | +| 40 | c5d39be | 2→**2** | 1→**1** | 1→**1** | yes | **UNCLEAN — reported to Daniel, no repair round opened.** **BLOCKER: Task 0 step 3 infers "first entry" from an empty `$BASE..HEAD`** — after an 8b rejection that range is empty while the implementation is staged, so the baseline records the *implemented* copies as inherited drift and the cycle can close on false parity evidence. **Same claim Resume states correctly three times and passes 29–30 repaired there; never swept for here.** MAJOR: **this round's own repair** — `git add .context/codex-reviews/` is a directory pathspec, not "by name", and sweeps another cycle's findings files in; my verification tested only outside that directory. Two tells (Blockers flat at 1 for a third pass, instrument cluster) → mandatory stop | | 26 | 05ec1b2 | **5** | **1** | 2 | yes | first pass on the revised plan. Resume audited progress by `WIP:` commits and called any non-`WIP` commit a stale base — the handoff leaves two other shapes, one of them `HEAD` **at** the base with everything in the index | | 27 | f69db7d | 5→**6** | 1→**3** | 2→**1** | yes | the close required the commit's **tree to equal the tip's**, which target §I **parks**; `reset --soft` leaves the index untouched, so the prose beside it was wrong about git; Resume still inferred staleness from a commit subject where §A3 names that state as reachable | | 28 | f2e5dda | 6→**6** | 3→**2** | 1→**2** | yes | pass 27's stale-base fix reached one site of three; its reset paragraph stated the index behaviour correctly and repeated the false claim four lines later; nothing checked the closing commit's **parent**; the message was validated in the file, where `commit-msg` hooks rewrite git's copy after `-F` reads it | From 4c9ed3cb23c974a0319d085bb781963921f9eb37 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Wed, 16 Sep 2026 14:25:57 +0200 Subject: [PATCH 162/181] docs(plans): the baseline always reads $BASE; stage slot paths, not a directory Pass 40's Blocker and Major, plus the scoped search the Blocker earned. Blocker. Task 0 step 3 selected the parity-baseline source by whether $BASE..HEAD was empty, reading the worktree when it was. After a reset --soft that range is empty while the whole implementation sits staged in the index and present in the worktree - a state Resume names as valid - so the baseline recorded the implemented copies as the original, Task 14 could no longer tell introduced drift from inherited, and the cycle could close on false parity evidence. It now reads the $BASE blobs unconditionally. That is the two sites agreeing rather than a new rule: Resume already required the rebuild to read those blobs. No entry-mode flag is needed either, because a first entry recorded a clean HEAD as $BASE. The scoped search, run as asked and reported as scoped: every place in the plan that infers worktree or index content, implementation progress, or a baseline's source from a commit range being empty or non-empty. Two hits, both in Task 0 step 3 - the branch and the prose justifying it - both repaired. Resume's five range statements are the correct direction and were left alone: they say the empty range proves nothing. Not every statement about a range is wrong; it says which commits are reachable and nothing about uncommitted content. Two found under that question, not a guaranteed total. Major, from the previous round's own repair. Step 7 steps one and three said "by name" and passed .context/codex-reviews/, a directory pathspec that stages every changed or untracked file beneath it, so another cycle's findings or an earlier pass's rode into the WIP commit and then the closing squash. Step one now names the two slot paths the call was built from; step three names the repaired files and this plan, and does not re-stage the findings that step one already committed. One residual is disclosed rather than guarded, and confirmed by execution: naming paths controls what the step adds to the index, but git commit still commits whatever the index already held, so a foreign path staged before the step runs is carried in and no check in this plan catches it. Verified by execution under sh, dash and bash, 15 checks each. A clean first entry takes the original as baseline. A re-entry with HEAD equal to $BASE and the implementation staged still takes the original, and a deliberately introduced deviation is not accepted as inherited drift - with the old selection run on the same state as a control, reproducing the Blocker. Another cycle's findings and an earlier pass's both stay out of step one's commit while this pass's two files are in, and the foreign files remain untracked and visible. The already-staged case is recorded with its outcome. The five earlier suites re-run green. --- .../2026-09-14-loop-rule-consolidation.md | 63 +++++++++++++------ 1 file changed, 43 insertions(+), 20 deletions(-) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index 345ee1f..c211471 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -1019,10 +1019,17 @@ none of passages (g), (h) or (j) — so `g4`, which sits at C 815 / W 997, could whose expected list named it — and the truncation hid about fifty of the roughly ninety lines the comparison actually emits. -**Read both copies from the source the re-entry rule selects, not from the worktree unconditionally.** -On a clean first run they are the same; once `$BASE..HEAD` is non-empty the worktree carries this -plan's own edits, and comparing them would fold introduced drift into the inherited-drift record — -after which Task 14 can no longer tell the two apart, which is the whole purpose of this baseline. +**Read both copies from the recorded `$BASE`, always — never from the worktree.** The worktree can +carry this plan's own edits, and folding those into the inherited-drift record leaves Task 14 unable +to tell introduced drift from inherited, which is the whole purpose of this baseline. **An earlier +revision selected the source by whether `$BASE..HEAD` was empty**, reading the worktree when it was. +That inference is unsound and this plan says so three times elsewhere: after a `reset --soft` the +range is empty **while the entire implementation sits staged in the index and present in the +worktree** — a state Resume names as valid — so the baseline would have recorded the *implemented* +copies as the original. **A commit range says which commits are reachable; it says nothing about +uncommitted content.** Resume already required the rebuild to read the `$BASE` blobs, so this is the +two sites agreeing rather than a new rule. **No entry-mode flag is needed either**: a first entry +recorded a clean `HEAD` as `$BASE`, so `$BASE` is the right source on every entry. **Selection, the `base` line and every extraction are ONE shell block**, because each fenced block is its own invocation: `C_SRC` set in one block and consumed in the next is empty by the time @@ -1046,14 +1053,15 @@ exits on failure; the rename is the last statement. ```bash BASE=$(cat .context/loop-rule-base) test -n "$BASE" || { echo "BASE empty or unreadable — Task 0 did not run"; exit 1; } -if [ -n "$(git log --oneline "$BASE"..HEAD)" ]; then - git show "$BASE:CLAUDE.md" > .context/loop-rule-c.src - git show "$BASE:plugins/dev-workflow/commands/workflow-init.md" > .context/loop-rule-w.src - C_SRC=.context/loop-rule-c.src; W_SRC=.context/loop-rule-w.src -else - C_SRC=CLAUDE.md; W_SRC=plugins/dev-workflow/commands/workflow-init.md -fi -test -r "$C_SRC" && test -r "$W_SRC" || { echo "baseline source unreadable"; exit 1; } +# Always the $BASE blobs. Never the worktree, and never a branch on the commit +# range: after a reset --soft that range is empty while the implementation is +# staged, so "empty range" would select the implemented copies as the original. +git show "$BASE:CLAUDE.md" > .context/loop-rule-c.src \ + || { echo "cannot read CLAUDE.md at $BASE"; exit 1; } +git show "$BASE:plugins/dev-workflow/commands/workflow-init.md" > .context/loop-rule-w.src \ + || { echo "cannot read workflow-init.md at $BASE"; exit 1; } +C_SRC=.context/loop-rule-c.src; W_SRC=.context/loop-rule-w.src +test -s "$C_SRC" && test -s "$W_SRC" || { echo "baseline source empty"; exit 1; } printf 'base\t%s\n' "$BASE" > .context/loop-rule-baseline-diff.tmp # renamed at the end # Tab-separated start and end anchors, written as LITERAL text — the escaping @@ -2799,15 +2807,29 @@ its own product forbids. **Do not apply the repair yet — not in the worktree either.** Deferring only the *commit* changes nothing: this plan restores nothing, so an edit made before its route authorized it is just as -unauthorized, and a later decline has nothing that removes it. **Stage the records by name**; `git -add -A` here would sweep exactly the edit this step exists to keep out, along with any unrelated -work in the tree. +unauthorized, and a later decline has nothing that removes it. + +**Stage the two slot paths this pass was called with, spelled out.** Not `git add -A`, which sweeps +exactly the edit this step exists to keep out — and **not the review directory either**: `git add +.context/codex-reviews/` is a directory pathspec that stages every changed or untracked file beneath +it, so another cycle's findings, or an earlier pass's, ride into this commit and then into the +closing squash, where the exact-dirty-set guard can no longer see them because they were committed +before the next call. **`NONCE` and `P` are this cycle's nonce and this pass's number** — the same +two the call's slot paths were built from. ```bash -git add .context/codex-reviews/ docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md -git commit -m "WIP: pass records" +NONCE=; P= +git add ".context/codex-reviews/gate-b-spec-$NONCE-pass-$P.md" \ + ".context/codex-reviews/gate-b-quality-$NONCE-pass-$P.md" \ + docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +git commit -m "WIP: pass $P records" ``` +**What this does not cover, disclosed rather than guarded:** naming the paths controls what this step +*adds* to the index; `git commit` still commits whatever the index already held. **A foreign path +staged before this step runs is carried in, and no check in this plan catches it** — the close's +dirty-set and changed-path checks read 8a's commit, not these. Nothing here fixes that. + **Step two — read the pass against the ordering, before any next call exists:** - **A source block standing** → **the source rule decides what must be repaired or answered, and its @@ -2833,9 +2855,10 @@ post-answer pass**, which is §A's rule and not this plan's. So on this route: a re-run every check it invalidated, commit both, and only then record the head. ```bash -# Stage the repaired artifacts BY NAME, plus the refreshed records. Not `git add -A`: -# it would carry unrelated work in the tree into a range a pass is about to review. -git add .context/codex-reviews/ \ +# The repaired files, spelled out, plus this plan — which carries the refreshed +# re-run records. The findings files went in at step one and are not re-staged. +# Neither `git add -A` nor the review directory: both carry in work no pass asked for. +git add \ docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md git commit -m "WIP: fix " From eee9d4eed128fdf94719f03426effb9107fd33f7 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Wed, 16 Sep 2026 14:32:57 +0200 Subject: [PATCH 163/181] docs(context): record Gate-A plan pass 41 and the eighth bounded revision Pass 41 reviewed 4c9ed3c: 1 Major, 0 Blockers, 0 Minors. Best pass since 35 and the first zero-Blocker pass since 37. Both pass-40 repairs went un-re-raised, and the scoped search for the empty-range fallacy produced no further hit beyond the two it repaired. The one Major is again this round's own repair. Step 7 step three runs git commit unguarded and then, unconditionally, deletes the prior tip and writes the reviewed head from git rev-parse HEAD. A failed commit is masked by the succeeding rev-parse: the next Gate-B call is issued against the old head while the repair sits staged, the review never sees it, and the cycle stops later on the close's dirty-set check. Demonstrated with a rejecting pre-commit hook - the commit does not land, the recorded head equals the old head, the repair stays staged. 8a two hundred lines away guards its commit exactly this way. Step one has the same unguarded shape with a smaller blast radius. Only one tell this pass - the instrument cluster - so the stop is instructed rather than mandated, the first time in six passes that is true. Two things kept in view. The repair still tends to produce the next finding, at passes 37, 40 and 41, but the class is narrowing: a design inversion, then a wrong pathspec, now a missing guard. And the last three Blockers each sat in an area no earlier pass had flagged, while pass 41 found none. One correction to the previous record: I had overstated the doubt about the OpenWolf filing's gate. CLAUDE.md's prose enumeration does not name todos.md, but the implemented classification does - is_docs_only accepts any *.md outside the prompt paths, a shipped test pins a root-level NOTES.md as docs, and running it on the three real commit paths returns yes. Reproduced here. That is not a claim that a gate ran. Floor 3, from docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md (risk high, security none), read fresh at this pass. Cycle stays open and unclean. --- .../gate-a-plan-om0bdd7udh-pass-41.md | 2 + .../gate-a-plan-om0bdd7udh-resume.md | 95 +++++++++---------- 2 files changed, 49 insertions(+), 48 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-41.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-41.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-41.md new file mode 100644 index 0000000..ae48436 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-41.md @@ -0,0 +1,2 @@ +MAJOR | high | Task 15 step 7, step three | The repair-commit command is followed unconditionally by deleting the prior tip and recording a new reviewed head; without `set -e` or an explicit status guard, a failed `git commit` is masked by the succeeding `git rev-parse` | The next Gate-B pass can be issued against the old commit while the repair remains staged or loose, so the review omits the repair and the cycle later stops on the close dirty-set check | Guard the repair commit and exit on failure; only clear the tip and write `loop-rule-reviewed-head` after a successful commit +END OF FINDINGS (1 total) diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md index ecadb71..579a3a3 100644 --- a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md @@ -13,63 +13,56 @@ Nothing depends on it; the pass files and the repo are authoritative where this ## Resume here -**Cycle OPEN and UNCLEAN. The plan stands at `c5d39be` and that is the anchor** — later commits on +**Cycle OPEN and UNCLEAN. The plan stands at `4c9ed3c` and that is the anchor** — later commits on this branch are records, not plan edits, so check `git log --oneline -1 -- docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md` -rather than `HEAD`. Pass 40 reviewed `c5d39be` and found **one Blocker and one Major**. -`.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-40.md` holds them. +rather than `HEAD`. Pass 41 reviewed `4c9ed3c` and found **one Major and nothing else — 0 Blockers**. +`.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-41.md` holds it. **One unrelated commit sits in this history and is NOT part of this cycle.** `d4a87c6` records an OpenWolf assessment — a separate task authorized 2026-09-16 — as `docs/openwolf-assessment.md`, a parked `todos.md` entry, and **four informational lines near the top of this plan** which state in their own text that they add no task, prerequisite or closure condition. It was committed on its own -so it stays identifiable. **Its gate classification was surfaced, not decided**: the two `docs/**.md` -paths are prose and N/A, `todos.md` is in neither CLAUDE.md's prose list nor the prompt paths, and -whoever owns that task decides whether it owed a gate. +so it stays identifiable, and pass 40 onward reviewed the plan including those lines. **Its gate +classification is settled and I had overstated the doubt:** CLAUDE.md's prose enumeration does not +name `todos.md`, but the **implemented** classification does — `is_docs_only` in +`plugins/dev-workflow/hooks/codex-gate.sh` accepts any `*.md` outside the prompt paths, and a +shipped test at `codex-gate.test.sh:446` pins a root-level `NOTES.md` as docs. Run on the three real +commit paths it returns **`is_docs_only: yes`**. Reproduced here. That is not a claim that a gate +ran; it is why no Gate-B round is owed for that filing on the strength of the filename alone. **Do not start a repair round on your own.** Daniel's assignment of 2026-09-16 ended with a checkpoint: *"genau ein vollständiger Gate-A-Pass und Bericht. Keine automatische Folgerunde, keine -Produktimplementierung und keine OpenWolf-Evaluation."* Spent — one bounded repair, one pass, both -delivered — so **the next move is Daniel's word, not an inference.** - -**Two tells, so this stop is mandatory as well as instructed**: the **Blocker count failed to fall** -for the third pass running (1 → 1 → 1), and the findings cluster on the instrument for the seventh. -Findings are flat at 2 and Majors flat at 1, so neither is a tell. - -**Read this before deciding the next round. Three things, stated as what was observed.** - -**One: three consecutive passes, each with exactly one Blocker, each in a different area no earlier -pass had flagged.** Pass 38 — Resume's missing branch check. Pass 39 — Task 15 step 7's ordering. -Pass 40 — Task 0 step 3's baseline-source branch. That is a statement about these three passes, not -a prediction about how many such areas remain. - -**Two: the repair-produces-the-next-finding chain resumed.** It broke once, at pass 39. Pass 40's -Major is **this round's own repair**: I wrote "stage the records by name" and then passed -`git add .context/codex-reviews/`, a *directory* pathspec that stages everything under it -recursively. **And my verification missed it** — I tested that unrelated work stayed out using a file -at the repo root, never a second findings file inside that directory. Demonstrated afterwards: a -stray `gate-a-plan-OTHERCYCLE-pass-9.md` is swept straight into the commit. - -**Three, and this one is actionable: pass 40's Blocker is a known claim that was repaired in one -place and never swept for.** *"An empty `$BASE..HEAD` is not proof of anything"* is stated three -times in Resume (lines 539, 617, 622 — repaired at passes 29–30) and **still used as a decision** -twice in Task 0 step 3 (1023, 1049). The plan contradicts itself in its own words. This is -`AGENTS.md`'s named class — **search for the claim, not the phrase** — and it is greppable. - -### The two open findings from pass 40 - -1. **BLOCKER — Task 0 step 3 infers "first entry" from an empty commit range.** Its branch is - `if [ -n "$(git log --oneline "$BASE"..HEAD)" ]` → read the `$BASE` blobs, else read the - worktree. **After an 8b rejection the range is empty and the whole implementation is staged and - present in the worktree** — a topology Resume names as valid — so the `else` branch records the - *implemented* copies as the parity baseline, i.e. as inherited drift. Task 14 then cannot tell - introduced drift from inherited, which is the baseline's whole purpose, and **the cycle can close - on false parity evidence.** Validated against the plan. -2. **MAJOR — this round's own repair. `git add .context/codex-reviews/` is not "by name".** Both - step 7 step one and step three pass the directory, so another cycle's findings files, or an - earlier pass's, are swept into the WIP commit and then into the closing squash — and because they - are committed *before* the next call, the exact-dirty-set guard cannot see the leak. **Fix: the - two explicit slot paths in step one; the exact refreshed record paths plus the named repair files - in step three.** Demonstrated by execution. +Produktimplementierung und keine OpenWolf-Evaluation."* Spent — one bounded repair, one scoped +search, one pass, all delivered — so **the next move is Daniel's word, not an inference.** + +**Only ONE tell this time — the instrument cluster, eighth pass running.** Findings fell 2 → 1, +**Blockers fell 1 → 0**, Majors flat at 1. **So this stop is instructed, not mandated** — the first +time in six passes that the tells alone would not have required it. + +**Where this stands after 41 passes.** Findings 3 → 3 → 2 → 2 → **1**; Blockers 0 → 1 → 1 → 1 → **0**; +Majors 1 → 2 → 1 → 1 → **1**. **Pass 41 is the best since pass 35** and the first zero-Blocker pass +since 37. Two things kept in view rather than glossed: + +**One: the repair still tends to produce the next finding** — passes 37, 40 and 41 each found a +defect in the previous round's own fix. **But the class is narrowing.** Pass 37's was a design +inversion, pass 40's a wrong pathspec, pass 41's a missing `||` guard — mechanical and local. + +**Two: the last three Blockers each sat in an area no earlier pass had flagged** (Resume's branch +check, Task 15 step 7, Task 0 step 3), and pass 41 found none. That is what these passes did; it +predicts nothing about what remains. + +### The one open finding from pass 41 + +**MAJOR — step 7 step three's repair commit is unguarded, and the head is recorded anyway.** The +block runs `git commit -m "WIP: fix "` and then, unconditionally, deletes the prior tip and +writes `.context/loop-rule-reviewed-head` from `git rev-parse HEAD`. **A failed commit is masked by +the succeeding `rev-parse`**: the next Gate-B call is issued against the *old* head while the repair +sits staged or loose, the review never sees it, and the cycle stops later on the close's dirty-set +check. **Demonstrated by execution** — with a rejecting `pre-commit` hook the commit does not land, +the recorded head equals the old head, and the repair stays staged. **This is this round's own +repair**, and 8a two hundred lines away guards its commit exactly this way, so the fix is that +spelling. **Step one has the same unguarded shape** with a smaller blast radius — it is followed by +no head write — and is worth carrying with it. ### Still collected and deliberately unrepaired (pass 34's Minors and Nit) @@ -164,6 +157,10 @@ what it carries; `.context/gate-a-spec-prompt.md` is the same shape for the spec passage's rows — but a matching id is a search hit, not a defect.** Tasks 3, 5 and 7 already read their sets off the table; Task 7's lists and Task 4 step 4's move list are complete and were left alone deliberately. Replacing them blind would remove requirements. +- **Guard every commit, then act on its result.** Pass 41: step three committed unguarded and wrote + the reviewed head from `git rev-parse HEAD` regardless, so a failed commit recorded the old head. + **8a already had the right spelling** — `|| { echo …; exit 1; }` — two hundred lines away. **When a + block commits and then records something derived from `HEAD`, the commit needs a guard.** - **`git add ` is not "staging by name".** A directory pathspec stages everything changed or untracked beneath it. Pass 40 found this in a repair whose own sentence said "by name". **When a step names what it stages, list paths — and test the sweep from INSIDE the directory**, not with a @@ -253,6 +250,8 @@ worth checking before a pass rather than after. | 39 | df9123a | 3→**2** | 1→**1** | 2→**1** | yes | **UNCLEAN — reported to Daniel, no repair round opened.** **All three pass-38 repairs held** — first round in four with no descendant finding. **BLOCKER: Task 15 step 7 commits the repair before the ordering has spoken**, so a membership or stop answer arrives after the fix is already in the artifact and nothing removes it — the loop its own §A product forbids. MAJOR: step 4b repairs prompt text and commits only the plan, leaving the repair outside `$BASE..HEAD`. **Second consecutive pass finding a Blocker in a previously unflagged area.** Two tells (Blockers flat at 1, instrument cluster) → mandatory stop | | — | — | — | — | — | — | **SEVENTH BOUNDED REVISION**, Daniel's assignment after pass 39, bounded by his reviewer to the two findings with **no Task-15-wide reconciliation**. Step 7 becomes three steps — record, read the ordering, then repair only on an authorizing route — and step one forbids applying the repair in the worktree too, since deferring the commit alone changes nothing where nothing restores. **Two route rules deliberately preserved, read from the approved text first:** the source-block branch keeps its own repair path (no blanket "only after continue"), and stop **parks** without rollback, so it is no retroactive revocation. Step 4b stages the repaired artifacts by name. 17 checks × 3 shells + 16 text assertions. Commit `c5d39be` | | 40 | c5d39be | 2→**2** | 1→**1** | 1→**1** | yes | **UNCLEAN — reported to Daniel, no repair round opened.** **BLOCKER: Task 0 step 3 infers "first entry" from an empty `$BASE..HEAD`** — after an 8b rejection that range is empty while the implementation is staged, so the baseline records the *implemented* copies as inherited drift and the cycle can close on false parity evidence. **Same claim Resume states correctly three times and passes 29–30 repaired there; never swept for here.** MAJOR: **this round's own repair** — `git add .context/codex-reviews/` is a directory pathspec, not "by name", and sweeps another cycle's findings files in; my verification tested only outside that directory. Two tells (Blockers flat at 1 for a third pass, instrument cluster) → mandatory stop | +| — | — | — | — | — | — | **EIGHTH BOUNDED REVISION**, Daniel's assignment after pass 40, bounded by his reviewer to the two findings **plus a scoped search for the one proven fallacy** — not a Task-15-wide audit. Task 0 step 3 reads the `$BASE` blobs unconditionally (Resume already required it; no entry-mode flag). Step one names the two slot paths; step three names the repaired files and the plan. **Search question, his wording, narrower than mine:** which places infer *worktree/index content, progress, or a baseline's source* from a range being empty or non-empty — **two hits, both in Task 0 step 3**, Resume's correct statements left alone. One residual disclosed, not guarded: `git commit` still commits an already-staged foreign path. 15 checks × 3 shells. Commit `4c9ed3c` | +| 41 | 4c9ed3c | 2→**1** | 1→**0** | 1→**1** | yes | **UNCLEAN — reported to Daniel, no repair round opened. Best pass since 35, first zero-Blocker pass since 37.** MAJOR, **this round's own repair**: step 7 step three's `git commit` is unguarded and the head is written unconditionally after it, so a failed commit records the **old** head and the next call reviews a tree without the repair — demonstrated with a rejecting `pre-commit` hook. 8a guards its commit exactly this way; step one has the same shape with a smaller blast radius. **Only one tell** (instrument cluster) — the stop is instructed, not mandated | | 26 | 05ec1b2 | **5** | **1** | 2 | yes | first pass on the revised plan. Resume audited progress by `WIP:` commits and called any non-`WIP` commit a stale base — the handoff leaves two other shapes, one of them `HEAD` **at** the base with everything in the index | | 27 | f69db7d | 5→**6** | 1→**3** | 2→**1** | yes | the close required the commit's **tree to equal the tip's**, which target §I **parks**; `reset --soft` leaves the index untouched, so the prose beside it was wrong about git; Resume still inferred staleness from a commit subject where §A3 names that state as reachable | | 28 | f2e5dda | 6→**6** | 3→**2** | 1→**2** | yes | pass 27's stale-base fix reached one site of three; its reset paragraph stated the index behaviour correctly and repeated the false claim four lines later; nothing checked the closing commit's **parent**; the message was validated in the file, where `commit-msg` hooks rewrite git's copy after `-F` reads it | From 0bb1d5d5e6898b60560a0aea479bc1e0a565efd0 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Wed, 16 Sep 2026 15:48:25 +0200 Subject: [PATCH 164/181] docs(plans): guard step 7's repair commit, and stop step one on a failed one Pass 41's Major, plus the adjacent step-one change as a separate bounded repair. Nothing else in the plan was changed. Step three ran git commit unguarded and then, unconditionally, deleted the prior reviewed tip and wrote the reviewed head from git rev-parse HEAD. A failed commit was masked by that rev-parse: the old head was recorded, the repair stayed staged, and the next call would have reviewed a tree it was not in. The commit is guarded now and exits before the tip is touched, so a failure leaves both records exactly as they were. 8a guards its record commit the same way and for the same reason. Step one is a different shape and is recorded as such: its commit is already the last command of its block, so the block's exit status is the commit's and nothing follows to mask it. The guard is added because what follows is prose - a reader who saw the failure scroll past can still walk into step two, and the ordering must not be read over a pass that was never recorded. Inspection, reported with its limits. Twenty-one of the plan's fifty-four fenced blocks contain a git commit. Eighteen end with the commit and have no following command, so no status can be masked within the block. One is 8a, already guarded. One is step three, the Major. One is 8b, whose commit is also terminal and whose next invocation - condition 6 - independently re-establishes that it landed: subject not WIP, parent equal to $BASE, clean tree, body matching the pinned bytes. That is not an instance of this dependency. The inspection is per block; cross-block sequencing was judged only where the next block re-establishes the state, and 8b to condition 6 is the one such pair. No third affected location, so no scope expansion. Verified by execution under sh, dash and bash, 21 checks each, in an isolated fixture. The old block reproduces the defect: it exits 0 on a rejected commit, records the old head as the reviewed head, and deletes the prior tip anyway. The repaired block stops, leaves the reviewed head and the prior tip byte-identical, and on a successful commit advances the head to the new commit, clears the tip and leaves the repair in HEAD. Step one stops on a rejected commit with no cycle record touched, and commits the records on a successful one. The six earlier suites re-run green. --- .../2026-09-14-loop-rule-consolidation.md | 21 +++++++++++++++++-- 1 file changed, 19 insertions(+), 2 deletions(-) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index c211471..07aa72b 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -2809,6 +2809,12 @@ its own product forbids. nothing: this plan restores nothing, so an edit made before its route authorized it is just as unauthorized, and a later decline has nothing that removes it. +**A failed records commit stops here; step two does not run.** This block *does* end with its +commit, so its exit status is the block's — unlike step three, it has no following command to mask +one. The guard is there because **what follows is prose, not a command**: a reader who saw the +failure scroll past can still walk into step two, and the ordering must not be read over a pass that +was never recorded. + **Stage the two slot paths this pass was called with, spelled out.** Not `git add -A`, which sweeps exactly the edit this step exists to keep out — and **not the review directory either**: `git add .context/codex-reviews/` is a directory pathspec that stages every changed or untracked file beneath @@ -2822,7 +2828,8 @@ NONCE=; P= git add ".context/codex-reviews/gate-b-spec-$NONCE-pass-$P.md" \ ".context/codex-reviews/gate-b-quality-$NONCE-pass-$P.md" \ docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md -git commit -m "WIP: pass $P records" +git commit -m "WIP: pass $P records" \ + || { echo "records commit FAILED — stop here; step two does not run"; exit 1; } ``` **What this does not cover, disclosed rather than guarded:** naming the paths controls what this step @@ -2854,13 +2861,23 @@ dirty-set and changed-path checks read 8a's commit, not these. Nothing here fixe post-answer pass**, which is §A's rule and not this plan's. So on this route: apply the repair now, re-run every check it invalidated, commit both, and only then record the head. +**The commit is guarded because this block does not end with it.** Three commands follow, and the +first thing they do is read `HEAD` — so an unguarded failure is masked by the `rev-parse` after it, +the *old* head is recorded, and the next call is issued against a tree the repair never reached. **On +a failed commit nothing moves**: the previous candidate's tip stays, the reviewed head stays, and no +call is issued. **8a guards its record commit the same way**, and for the same reason. + ```bash # The repaired files, spelled out, plus this plan — which carries the refreshed # re-run records. The findings files went in at step one and are not re-staged. # Neither `git add -A` nor the review directory: both carry in work no pass asked for. git add \ docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md -git commit -m "WIP: fix " +# Guard the commit. Everything below derives from HEAD, so an unguarded failure +# is masked by the `git rev-parse` that follows it: the old head gets recorded, +# the repair stays staged, and the next call reviews a tree it is not in. +git commit -m "WIP: fix " \ + || { echo "repair commit FAILED — tip and reviewed head left as they are; no call may be issued"; exit 1; } rm -f .context/loop-rule-reviewed-tip # the previous candidate's closing tip git rev-parse HEAD > .context/loop-rule-reviewed-head # the head the NEXT call is issued against From f4e04626333254a4c6dcbae01ca8e3b108cf78ea Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Wed, 16 Sep 2026 15:59:09 +0200 Subject: [PATCH 165/181] docs(context): record Gate-A plan pass 42 and the ninth bounded revision Pass 42 reviewed 0bb1d5d: 2 Blockers, 1 Major. Findings rose 1 to 3 and Blockers rose 0 to 2, so three tells stand and the stop is mandatory as well as instructed. Two of the three are defects the eighth revision introduced, and the record says so. When git add -A was replaced with named paths, the plan stayed in step one's list even though step one produces no plan output - so a post-review edit to the plan is committed before step two reads the ordering, which is pass 39's violation reopened through a different file. And splitting git add ... && git commit onto separate lines dropped the && that had guarded the staging, so a failed add now falls through to a commit that can still succeed on content the index already held. The shape that was replaced was safer in both respects than the narrowing that replaced it. One is pre-existing and was outside the declared limit of the inspection this round ran. Both reviewed-head writes - step 6 and step 7 step three - run git rev-parse HEAD into the file and then cat it without checking the write. Demonstrated by execution: with the target unwritable the redirect fails, cat prints the stale earlier value, and the block exits 0, so a Gate-B call is issued against a head that is not HEAD. Step 6's block contains no commit, and the assignment scoped the inspection to commit-containing blocks; the limit was declared and this finding went through it. Step 4b's unguarded add is pre-existing too. Two per-pass lessons are recorded: guard the staging and not only the commit, and treat a narrowing as a change that can lose a guarantee the wider form carried. Floor 3, from docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md (risk high, security none), read fresh at this pass. Cycle stays open and unclean. --- .../gate-a-plan-om0bdd7udh-pass-42.md | 4 + .../gate-a-plan-om0bdd7udh-resume.md | 100 ++++++++++-------- 2 files changed, 60 insertions(+), 44 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-42.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-42.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-42.md new file mode 100644 index 0000000..3392958 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-42.md @@ -0,0 +1,4 @@ +BLOCKER | high | Task 15 step 7, step one | The step says it records the pass and nothing else and forbids any repair before the ordering authorizes one, but its `git add` also stages this plan even though step one produces no plan output; any premature repair or other post-review edit to the plan is therefore committed before step two, and the later instruction to stop when a repair was applied too early comes after the mutation. | A membership, source-block, or stop answer can arrive only after the repair is already in history, recreating the pass-39 ordering violation and leaving a parked cycle carrying an unauthorized change. | Remove the plan from step one's path list and require it unchanged before the records commit; if step one is meant to write a plan record, define and isolate that exact record before staging it. +BLOCKER | high | Task 15 steps 4b and 7, the staging/commit blocks | The new guards cover `git commit` only; each preceding `git add` is unguarded, so a staging failure falls through to the commit, and the plan expressly admits that the index may already contain other staged content. A commit of that existing content can succeed and bypass the guard even though the named findings, repair, or refreshed records were not staged. | Step one can read an unrecorded or incomplete pass, and step three can record a new reviewed head and issue the next review without the intended repair or evidence in that head; step 4b can likewise put Gate B over a range that omits its repair. | Guard every staging command before its commit, and where downstream state depends on the named files, verify those files are actually in the resulting commit before proceeding. +MAJOR | high | Task 15 steps 6 and 7 step three | Both reviewed-head updates use `git rev-parse HEAD > .context/loop-rule-reviewed-head` followed by `cat` without checking the write; a failed `rev-parse` or redirection can be masked by a successful `cat`, including a stale existing file when the path cannot be replaced. | A Gate-B call can be issued with no trustworthy persisted head, wasting the pass or making the later exact-head condition reject a review whose recorded input was stale; the procedure has advanced despite failing the condition that makes the call attributable. | Resolve and validate the full object id, write it through a checked temporary file, rename only on success, and do not issue the call unless the persisted value equals the current `HEAD`. +END OF FINDINGS (3 total) diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md index 579a3a3..d5e96d5 100644 --- a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md @@ -13,56 +13,60 @@ Nothing depends on it; the pass files and the repo are authoritative where this ## Resume here -**Cycle OPEN and UNCLEAN. The plan stands at `4c9ed3c` and that is the anchor** — later commits on +**Cycle OPEN and UNCLEAN. The plan stands at `0bb1d5d` and that is the anchor** — later commits on this branch are records, not plan edits, so check `git log --oneline -1 -- docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md` -rather than `HEAD`. Pass 41 reviewed `4c9ed3c` and found **one Major and nothing else — 0 Blockers**. -`.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-41.md` holds it. +rather than `HEAD`. Pass 42 reviewed `0bb1d5d` and found **two Blockers and one Major**. +`.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-42.md` holds them. **One unrelated commit sits in this history and is NOT part of this cycle.** `d4a87c6` records an OpenWolf assessment — a separate task authorized 2026-09-16 — as `docs/openwolf-assessment.md`, a -parked `todos.md` entry, and **four informational lines near the top of this plan** which state in -their own text that they add no task, prerequisite or closure condition. It was committed on its own -so it stays identifiable, and pass 40 onward reviewed the plan including those lines. **Its gate -classification is settled and I had overstated the doubt:** CLAUDE.md's prose enumeration does not -name `todos.md`, but the **implemented** classification does — `is_docs_only` in -`plugins/dev-workflow/hooks/codex-gate.sh` accepts any `*.md` outside the prompt paths, and a -shipped test at `codex-gate.test.sh:446` pins a root-level `NOTES.md` as docs. Run on the three real -commit paths it returns **`is_docs_only: yes`**. Reproduced here. That is not a claim that a gate -ran; it is why no Gate-B round is owed for that filing on the strength of the filename alone. +parked `todos.md` entry, and four informational lines near the top of this plan which state in their +own text that they add no task, prerequisite or closure condition. Its gate classification is +settled: `is_docs_only` in `plugins/dev-workflow/hooks/codex-gate.sh` returns **yes** for the three +real commit paths, and a shipped test pins a root-level `.md` as docs. Not a claim that a gate ran. **Do not start a repair round on your own.** Daniel's assignment of 2026-09-16 ended with a -checkpoint: *"genau ein vollständiger Gate-A-Pass und Bericht. Keine automatische Folgerunde, keine -Produktimplementierung und keine OpenWolf-Evaluation."* Spent — one bounded repair, one scoped -search, one pass, all delivered — so **the next move is Daniel's word, not an inference.** - -**Only ONE tell this time — the instrument cluster, eighth pass running.** Findings fell 2 → 1, -**Blockers fell 1 → 0**, Majors flat at 1. **So this stop is instructed, not mandated** — the first -time in six passes that the tells alone would not have required it. - -**Where this stands after 41 passes.** Findings 3 → 3 → 2 → 2 → **1**; Blockers 0 → 1 → 1 → 1 → **0**; -Majors 1 → 2 → 1 → 1 → **1**. **Pass 41 is the best since pass 35** and the first zero-Blocker pass -since 37. Two things kept in view rather than glossed: - -**One: the repair still tends to produce the next finding** — passes 37, 40 and 41 each found a -defect in the previous round's own fix. **But the class is narrowing.** Pass 37's was a design -inversion, pass 40's a wrong pathspec, pass 41's a missing `||` guard — mechanical and local. - -**Two: the last three Blockers each sat in an area no earlier pass had flagged** (Resume's branch -check, Task 15 step 7, Task 0 step 3), and pass 41 found none. That is what these passes did; it -predicts nothing about what remains. - -### The one open finding from pass 41 - -**MAJOR — step 7 step three's repair commit is unguarded, and the head is recorded anyway.** The -block runs `git commit -m "WIP: fix "` and then, unconditionally, deletes the prior tip and -writes `.context/loop-rule-reviewed-head` from `git rev-parse HEAD`. **A failed commit is masked by -the succeeding `rev-parse`**: the next Gate-B call is issued against the *old* head while the repair -sits staged or loose, the review never sees it, and the cycle stops later on the close's dirty-set -check. **Demonstrated by execution** — with a rejecting `pre-commit` hook the commit does not land, -the recorded head equals the old head, and the repair stays staged. **This is this round's own -repair**, and 8a two hundred lines away guards its commit exactly this way, so the fix is that -spelling. **Step one has the same unguarded shape** with a smaller blast radius — it is followed by -no head write — and is worth carrying with it. +checkpoint: *"Validate and report that pass, then stop regardless of its outcome. No automatic repair +round, product implementation, or OpenWolf evaluation."* Spent — one bounded revision, one scoped +inspection, one pass, all delivered. + +**Three tells, so this stop is mandatory as well as instructed**: findings rose 1 → 3, **Blockers +rose 0 → 2**, and the findings cluster on the instrument for the ninth pass running. + +**Two of the three are mine, and the record says which.** The distinction matters more than the +count here. + +- **Newly introduced, by the eighth revision (mine).** Blocker 1 and half of Blocker 2. When I + replaced `git add -A` with named paths I **kept the plan in step one's list**, and I **split + `git add … && git commit` onto separate lines**, dropping the `&&` that had guarded the staging. + The original shape was safer in both respects than the narrowing I put in its place. +- **Newly discovered, pre-existing.** The Major at **step 6**, and step 4b's unguarded `git add`. + **My inspection could not have found step 6's**: the assignment scoped it to *commit-containing* + blocks, step 6's block contains none, and I declared that limit. The limit was real and this + finding walked straight through it. + +### The three open findings from pass 42 + +1. **BLOCKER — step one stages this plan, while forbidding any premature repair.** The block says + "record the pass, and nothing else" and "do not apply the repair yet — not in the worktree + either", then stages + `docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md`. **Step one produces no plan + output** — the re-run records are written at step three — so anything dirty in the plan is a + post-review edit, and it is committed **before** step two reads the ordering. A membership, + source-block or stop answer then arrives after the change is already in history, which is the + pass-39 violation reopened through a different file. **Introduced by me.** +2. **BLOCKER — the new guards cover `git commit` only; every `git add` before them is unguarded.** + A failed staging falls through to a commit that can still succeed on content the index already + held — and this plan **expressly discloses** that the index may hold foreign staged content. Step + one then records a pass without its findings files; step three records a new reviewed head with + no repair in it; step 4b puts Gate B over a range missing its repair. **Step one and step three + are mine** — the original `&&` guarded the add. **Step 4b's is pre-existing.** +3. **MAJOR — both reviewed-head writes are unchecked, and `cat` masks the failure.** Step 6 (line + 2767) and step 7 step three (2883) both run `git rev-parse HEAD > .context/loop-rule-reviewed-head` + followed by `cat`. **Demonstrated by execution**: with the target unwritable the redirect fails, + `cat` prints the **stale** earlier value, and the block exits **0** — so a Gate-B call is issued + against a head that is not `HEAD`, and the close's exact-head condition later rejects a review + whose recorded input was wrong. **Pre-existing at step 6; the same shape at step three.** ### Still collected and deliberately unrepaired (pass 34's Minors and Nit) @@ -157,6 +161,12 @@ what it carries; `.context/gate-a-spec-prompt.md` is the same shape for the spec passage's rows — but a matching id is a search hit, not a defect.** Tasks 3, 5 and 7 already read their sets off the table; Task 7's lists and Task 4 step 4's move list are complete and were left alone deliberately. Replacing them blind would remove requirements. +- **Guard the staging too, not only the commit.** Pass 42: `git add` on its own line can fail while + the following guarded `git commit` still succeeds on content the index already held. **The original + `git add … && git commit` was safer than the named-path rewrite that replaced it** — when you + narrow a command, carry its guarantees across. +- **A narrowing is a change, and it can lose something.** Two of pass 42's three findings came from + the eighth revision's own narrowing, not from the text it replaced. - **Guard every commit, then act on its result.** Pass 41: step three committed unguarded and wrote the reviewed head from `git rev-parse HEAD` regardless, so a failed commit recorded the old head. **8a already had the right spelling** — `|| { echo …; exit 1; }` — two hundred lines away. **When a @@ -252,6 +262,8 @@ worth checking before a pass rather than after. | 40 | c5d39be | 2→**2** | 1→**1** | 1→**1** | yes | **UNCLEAN — reported to Daniel, no repair round opened.** **BLOCKER: Task 0 step 3 infers "first entry" from an empty `$BASE..HEAD`** — after an 8b rejection that range is empty while the implementation is staged, so the baseline records the *implemented* copies as inherited drift and the cycle can close on false parity evidence. **Same claim Resume states correctly three times and passes 29–30 repaired there; never swept for here.** MAJOR: **this round's own repair** — `git add .context/codex-reviews/` is a directory pathspec, not "by name", and sweeps another cycle's findings files in; my verification tested only outside that directory. Two tells (Blockers flat at 1 for a third pass, instrument cluster) → mandatory stop | | — | — | — | — | — | — | **EIGHTH BOUNDED REVISION**, Daniel's assignment after pass 40, bounded by his reviewer to the two findings **plus a scoped search for the one proven fallacy** — not a Task-15-wide audit. Task 0 step 3 reads the `$BASE` blobs unconditionally (Resume already required it; no entry-mode flag). Step one names the two slot paths; step three names the repaired files and the plan. **Search question, his wording, narrower than mine:** which places infer *worktree/index content, progress, or a baseline's source* from a range being empty or non-empty — **two hits, both in Task 0 step 3**, Resume's correct statements left alone. One residual disclosed, not guarded: `git commit` still commits an already-staged foreign path. 15 checks × 3 shells. Commit `4c9ed3c` | | 41 | 4c9ed3c | 2→**1** | 1→**0** | 1→**1** | yes | **UNCLEAN — reported to Daniel, no repair round opened. Best pass since 35, first zero-Blocker pass since 37.** MAJOR, **this round's own repair**: step 7 step three's `git commit` is unguarded and the head is written unconditionally after it, so a failed commit records the **old** head and the next call reviews a tree without the repair — demonstrated with a rejecting `pre-commit` hook. 8a guards its commit exactly this way; step one has the same shape with a smaller blast radius. **Only one tell** (instrument cluster) — the stop is instructed, not mandated | +| — | — | — | — | — | — | **NINTH BOUNDED REVISION**, Daniel's assignment after pass 41: guard step three's repair commit (exiting **before** the tip is touched), and guard step one's records commit as a **separate** bounded repair of a different shape — its commit is already terminal, so the guard exists because what follows is *prose*. **Inspection, scoped to commit-containing blocks and reported with that limit:** 21 of 54 fenced blocks commit; 18 end with the commit; 8a is guarded; 8b is terminal and condition 6 re-establishes it independently; step three was the only masking sequence. 21 checks × 3 shells. Commit `0bb1d5d` | +| 42 | 0bb1d5d | 1→**3** | 0→**2** | 1→**1** | yes | **UNCLEAN — reported to Daniel, no repair round opened.** **BLOCKER: step one stages the plan** while forbidding a premature repair, so a post-review plan edit is committed before the ordering is read — pass 39's violation through another file. **BLOCKER: every `git add` before the new guards is unguarded**, so a failed staging falls through to a commit that succeeds on already-staged content. MAJOR: both reviewed-head writes are unchecked and `cat` masks a failed write — **demonstrated**, the block exits 0 on a stale value. **Two are mine** (the narrowing dropped the original `&&` and kept the plan in the list); **step 6's Major and step 4b's add are pre-existing**, and step 6 lay outside my inspection's declared limit. Three tells → mandatory stop | | 26 | 05ec1b2 | **5** | **1** | 2 | yes | first pass on the revised plan. Resume audited progress by `WIP:` commits and called any non-`WIP` commit a stale base — the handoff leaves two other shapes, one of them `HEAD` **at** the base with everything in the index | | 27 | f69db7d | 5→**6** | 1→**3** | 2→**1** | yes | the close required the commit's **tree to equal the tip's**, which target §I **parks**; `reset --soft` leaves the index untouched, so the prose beside it was wrong about git; Resume still inferred staleness from a commit subject where §A3 names that state as reachable | | 28 | f2e5dda | 6→**6** | 3→**2** | 1→**2** | yes | pass 27's stale-base fix reached one site of three; its reset paragraph stated the index behaviour correctly and repeated the false claim four lines later; nothing checked the closing commit's **parent**; the message was validated in the file, where `commit-msg` hooks rewrite git's copy after `-F` reads it | From 8361f0acf7dc02c66a459860bed6559b9743f81d Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Wed, 16 Sep 2026 17:11:04 +0200 Subject: [PATCH 166/181] docs(plans): guard the staging, pin the index, check the reviewed-head write Pass 42's two Blockers and one Major, bounded to the four blocks they name: Task 15 step 4b, step 6, and step 7's step one and step three. Nothing else in the plan was changed, and no plan-wide search was run. Provenance, corrected. The working record dated both Blockers to the eighth revision. They arrived at c5d39be, the seventh: it already carried the plan path in step one's staging list and already split git add from git commit. The eighth revision (4c9ed3c) replaced a directory pathspec with two named slot paths and inherited both. The earlier one-line form guarded the commit against a failed staging, and it would equally have swept a premature plan repair into the commit, so it was safer in one respect, not two. Step one no longer stages the plan, and requires it unchanged first. The step produces no plan output - the re-run records are written at step three - so a dirty plan there is a post-review edit, and committing it would put a change in history before step two has read the ordering. Removing the path from git add does not settle that on its own: git commit commits the index, so an edit already staged would ride in regardless. The precondition asks HEAD against the index, then the index against the worktree. A difference and a failed comparison both stop, and the message names neither cause, because the answer to both is to report the concrete state. Every git add is guarded, and each commit is checked against what it was given. A failed staging otherwise falls through to a commit that can still succeed on content the index already held: step one would then record a pass without its findings files, step three would record a new reviewed head with no repair in it, and step 4b would put Gate B over a range missing its repair. The check pins the index with git write-tree before the commit and compares that pin against the resulting commit tree. Successful staging does not establish presence. git add stages a removal as readily as a content change, so a tracked findings file absent from the worktree is staged as gone, and the pin and the commit then agree without it. Step one therefore checks both slot paths in the pin, and checks them as blobs: an existence test alone accepts a directory standing at the path, which is one of the causes CLAUDE.md already tells a reader to diagnose separately. Steps three and 4b carry no such check, because an authorized repair may delete a file. Both reviewed-head writes are guarded and the cat is gone. With the target unwritable the redirect fails, the cat printed the stale earlier value, and the block exited 0 - so a call could be issued against a head that is not HEAD, and the close's exact-head condition would later reject a review whose recorded input was wrong. Step 6's block contains no commit and so lay outside the ninth revision's inspection, which was scoped to commit-containing blocks and declared that limit; step three's block does contain one and was inside it. The limit explains one of the two sites, not both. Stated limits, so they are not read as more than they are. The pin is the index as staged at that moment, not the content the pass acceptance validated; nothing between that reading and the staging observes a change. The tree comparison is index-wide: it stops on any divergence between the pin and the commit tree, including on paths the step never named, and it is not a foreign-index check, since content staged beforehand stands on both sides. Step one has no re-run of the findings-file structural check on committed content; Close condition 4 does that for 8a's record commit alone. Verified by execution, not by reading. The four blocks were extracted from this plan verbatim, their authoring placeholders filled, and run in a disposable repository: 20 checks under sh, dash and bash, all green. The cases: both success paths; the plan staged-modified with the worktree identical to HEAD, which a worktree-only test cannot see; the plan dirty in the worktree; a findings file never written; a tracked findings file deleted before staging; a directory standing at a findings path; a commit rejected by a pre-commit hook; a hook changing the index after the pin, at step one and at step 4b; a repair that deletes a file, which must pass; a staging failure at step three and at 4b; and a failed reviewed-head write at step three and at step 6. Three failures in the first run were all defects in the harness, and each is recorded rather than quietly fixed: git rm had already staged the deletion the block was supposed to stage; a deleted tracked path was expected to fail git add when it is staged as a removal instead; and a hook wrote under .git/, where git add -A reaches nothing. Shell checks: sh -n, dash -n, bash -n and shellcheck --shell=sh clean on every changed block. Named and deliberately not repaired, per the bounded assignment: 8a's per-file presence check is fail-open on a git error; the rm -f before step three's head write is unguarded; other redirect writes in this plan carry the same unchecked shape. The working record's provenance lines are also still wrong and are advisory, not the plan. --- .../2026-09-14-loop-rule-consolidation.md | 129 +++++++++++++++--- 1 file changed, 109 insertions(+), 20 deletions(-) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index 07aa72b..29a234a 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -2710,12 +2710,25 @@ installed text this step judges: the two prompt copies, the hook and its test. * repair actually touched go in**, together with this plan's refreshed record — and the re-run records must describe *that* commit, which is why the re-runs above come first. +**Staging and commit are both guarded, and the commit is checked against what it was given**, because +step 6 issues Gate B over the range this commit ends: a failed `git add` otherwise falls through to a +commit that can still succeed on content the index already held, and the review then runs over a range +missing its repair. The pin is the same as step 7 step three's — **the whole index, unrelated paths +included**, demanding no presence, since a repair here may delete a file too. + ```bash # add only what the repair touched, from: CLAUDE.md, # plugins/dev-workflow/commands/workflow-init.md, # plugins/dev-workflow/hooks/codex-gate.sh, plugins/dev-workflow/hooks/codex-gate.test.sh -git add docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md -git commit -m "WIP: plan records" # or: "WIP: prompt-standards repair + plan records" +# Where a repair was made, the message is: "WIP: prompt-standards repair + plan records" +git add docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md \ + || { echo "staging FAILED — step 6 must not issue Gate B over this range"; exit 1; } +ITREE=$(git write-tree) \ + || { echo "cannot pin the staged index; do not proceed to step 6"; exit 1; } +git commit -m "WIP: plan records" \ + || { echo "commit FAILED — step 6 must not issue Gate B over this range"; exit 1; } +test "$(git rev-parse "HEAD^{tree}")" = "$ITREE" \ + || { echo "the commit's tree is not the index that was pinned; do not proceed to step 6"; exit 1; } ``` **The same applies to every record this plan collects** — the sweep, the next-state table, the @@ -2764,8 +2777,10 @@ an enumeration here** — an id list in this step was already stale once, naming ```bash BASE=$(cat .context/loop-rule-base) # the parent of the FIRST WIP, from Task 0 test -n "$BASE" || { echo "BASE empty or unreadable — Task 0 did not run"; exit 1; } -git rev-parse HEAD > .context/loop-rule-reviewed-head # the head THIS call is issued against -cat .context/loop-rule-reviewed-head +# The head THIS call is issued against. Guarded, and no `cat`: the redirect can fail +# while a following `cat` prints a STALE value and the block still exits 0. +git rev-parse HEAD > .context/loop-rule-reviewed-head \ + || { echo "recording the reviewed head FAILED — no call may be issued"; exit 1; } ``` **Every Gate-B call records the head it is issued against, this first one included**, and step 7 @@ -2809,11 +2824,35 @@ its own product forbids. nothing: this plan restores nothing, so an edit made before its route authorized it is just as unauthorized, and a later decline has nothing that removes it. -**A failed records commit stops here; step two does not run.** This block *does* end with its -commit, so its exit status is the block's — unlike step three, it has no following command to mask -one. The guard is there because **what follows is prose, not a command**: a reader who saw the -failure scroll past can still walk into step two, and the ordering must not be read over a pass that -was never recorded. +**A failed precondition, staging, commit or content check stops here; step two does not run.** No +command in this block is the last thing that happens — the commit is followed by a tree comparison, +and that by prose. A reader who saw a failure scroll past can still walk into step two, and the +ordering must not be read over a pass that was never recorded, nor over one whose findings files the +commit does not carry. + +**The plan is not staged here, and must be unchanged before this commit.** Step one produces no plan +output — the re-run records are written at step three — so a dirty plan here is a post-review edit, +and committing it puts a change in history before step two has read the ordering. Removing the path +from `git add` does not settle it: **`git commit` commits the index**, so an edit already staged +rides in regardless. The check therefore asks HEAD against the index first, then the index against +the worktree; a difference and a failed comparison both stop, and neither says what caused it, so +the answer to both is to **report the concrete state** — what is in the worktree, what is committed, +and which route had not yet spoken. This plan restores nothing and invents no rollback. The check +covers **this one path**; the residual below is unchanged. + +**A successful `git add` does not establish presence.** It stages a removal as readily as a content +change, so a tracked findings file that is absent from the worktree is staged as *gone* — and the +pinned index and the resulting commit then agree, both without it. **Tree equality preserves +presence only where presence was established in the pin**, which is why the two paths are checked +there, and checked as **blobs**: an existence test alone accepts a directory standing at the path, +one of the causes §5 already tells a reader to diagnose separately. Step one deletes nothing; an +authorized repair may, which is why step three and step 4b carry no such check. + +**What the pin is, and what it is not.** It is the index **as staged at that moment**, taken after +`git add` and before `git commit` — not the content the pass acceptance validated. Nothing between +that reading and this staging observes a change, so a findings file rewritten in between is pinned +as staged and passes here. Close condition 4 re-runs the structural check on committed content for +8a's record commit; **step one has no such re-run and this check does not stand in for one.** **Stage the two slot paths this pass was called with, spelled out.** Not `git add -A`, which sweeps exactly the edit this step exists to keep out — and **not the review directory either**: `git add @@ -2825,11 +2864,40 @@ two the call's slot paths were built from. ```bash NONCE=; P= -git add ".context/codex-reviews/gate-b-spec-$NONCE-pass-$P.md" \ - ".context/codex-reviews/gate-b-quality-$NONCE-pass-$P.md" \ - docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +SPEC=".context/codex-reviews/gate-b-spec-$NONCE-pass-$P.md" +QUAL=".context/codex-reviews/gate-b-quality-$NONCE-pass-$P.md" +PLAN=docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md + +# The plan is not staged here and must be unchanged: HEAD against the index (what the +# commit takes), then the index against the worktree. A difference and a failed +# comparison both stop, and neither says what caused it — report the state. +git diff --quiet --cached HEAD -- "$PLAN" \ + || { echo "plan is not identical between HEAD and the index, or the comparison failed — stop and report the state"; exit 1; } +git diff --quiet -- "$PLAN" \ + || { echo "plan is not identical between the index and the worktree, or the comparison failed — stop and report the state"; exit 1; } + +git add "$SPEC" "$QUAL" \ + || { echo "staging FAILED — the pass is not recorded; step two does not run"; exit 1; } + +# Pinned after staging, before the commit: the WHOLE index as staged at that moment, +# unrelated paths included. +ITREE=$(git write-tree) \ + || { echo "cannot pin the staged index; step two does not run"; exit 1; } + +# A successful add does not mean these paths are present — it stages a removal too. +# Require a blob at each: an existence test would accept a directory at the path. +test "$(git cat-file -t "$ITREE:$SPEC" 2>/dev/null)" = blob \ + || { echo "$SPEC is not a blob in the staged index; the pass is not recorded, step two does not run"; exit 1; } +test "$(git cat-file -t "$ITREE:$QUAL" 2>/dev/null)" = blob \ + || { echo "$QUAL is not a blob in the staged index; the pass is not recorded, step two does not run"; exit 1; } + git commit -m "WIP: pass $P records" \ || { echo "records commit FAILED — stop here; step two does not run"; exit 1; } + +# Any divergence between the pinned index and the resulting commit tree stops here, +# whatever produced it. +test "$(git rev-parse "HEAD^{tree}")" = "$ITREE" \ + || { echo "the commit's tree is not the index that was pinned; step two does not run"; exit 1; } ``` **What this does not cover, disclosed rather than guarded:** naming the paths controls what this step @@ -2861,27 +2929,48 @@ dirty-set and changed-path checks read 8a's commit, not these. Nothing here fixe post-answer pass**, which is §A's rule and not this plan's. So on this route: apply the repair now, re-run every check it invalidated, commit both, and only then record the head. -**The commit is guarded because this block does not end with it.** Three commands follow, and the -first thing they do is read `HEAD` — so an unguarded failure is masked by the `rev-parse` after it, -the *old* head is recorded, and the next call is issued against a tree the repair never reached. **On -a failed commit nothing moves**: the previous candidate's tip stays, the reviewed head stays, and no -call is issued. **8a guards its record commit the same way**, and for the same reason. +**The commit is guarded because this block does not end with it.** Commands follow, and among the +first things they do is read `HEAD` — so an unguarded failure is masked by the `rev-parse` after it, +the *old* head is recorded, and the next call is issued against a tree the repair never reached. **A +failed commit stops the commands below it**: the previous candidate's tip is not removed, the +reviewed head is not rewritten, and no call is issued. It is **not** a claim that the failed attempt +left the tree as it was — a hook can rewrite anything before failing, which this plan already says of +the closing message file. **8a guards its record commit the same way**, and for the same reason. + +**The staging is guarded too, and the commit is checked against what it was given.** A failed `git +add` otherwise falls through to a commit that can still succeed on content the index already held. +The pin is `git write-tree` — **the whole index, unrelated paths included**, so any divergence between +it and the resulting commit tree stops this step, whatever produced it. That is broader than this +step's obligation and deliberately fail-closed. It is **not** a foreign-index check: content staged +before this step stands on both sides and passes. And it demands no presence, which is what step one +needs and this step must not have — **an authorized repair may delete a file**, and a staged deletion +is in the pin and in the commit alike. ```bash # The repaired files, spelled out, plus this plan — which carries the refreshed # re-run records. The findings files went in at step one and are not re-staged. # Neither `git add -A` nor the review directory: both carry in work no pass asked for. git add \ - docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md + docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md \ + || { echo "staging FAILED — nothing committed, no call may be issued"; exit 1; } + +# Pinned after staging, before the commit: the whole index as staged at that moment. +ITREE=$(git write-tree) \ + || { echo "cannot pin the staged index; no call may be issued"; exit 1; } + # Guard the commit. Everything below derives from HEAD, so an unguarded failure # is masked by the `git rev-parse` that follows it: the old head gets recorded, # the repair stays staged, and the next call reviews a tree it is not in. git commit -m "WIP: fix " \ || { echo "repair commit FAILED — tip and reviewed head left as they are; no call may be issued"; exit 1; } +test "$(git rev-parse "HEAD^{tree}")" = "$ITREE" \ + || { echo "the commit's tree is not the index that was pinned; no call may be issued"; exit 1; } rm -f .context/loop-rule-reviewed-tip # the previous candidate's closing tip -git rev-parse HEAD > .context/loop-rule-reviewed-head # the head the NEXT call is issued against -cat .context/loop-rule-reviewed-head +# The head the NEXT call is issued against. Guarded, and no `cat`: the redirect can +# fail while a following `cat` prints the STALE value and the block exits 0. +git rev-parse HEAD > .context/loop-rule-reviewed-head \ + || { echo "recording the reviewed head FAILED — no call may be issued"; exit 1; } ``` **The reviewed-head file is written on this route alone, and after the repair commit**, because From ff708a17bc59463a8f999a9fd1361908b7c7d5cd Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Wed, 16 Sep 2026 17:20:58 +0200 Subject: [PATCH 167/181] docs(context): record Gate-A plan pass 43 and the tenth bounded revision Pass 43 reviewed the plan at 8361f0a and is valid: the findings file ends with the exact terminator, carries exactly five finding lines and nothing else. Two Blockers, two Majors, one Minor. All three pass-42 repairs held. Nothing was re-raised against the four blocks the tenth revision touched - Task 15 step 4b, step 6, and step 7's step one and step three. Every Blocker and Major is the same class at another location: an unguarded command whose status a following command masks. 8b's cp pinning the revalidated closing message; condition 6's git log redirect, masked by the diff after it; step 4's git fetch before the version-bump check; and every remaining git add before a git commit across Tasks 0-14 and step 8a. The last of those is the plan-wide sweep the assignment excluded by name. The only finding inside the assigned fix set is a Minor: the new blob check accepts a symlink, which git stores as a blob with mode 120000, so the indexed mode would have to be read as well. Collected, not iterated. Loop stopped rather than widened. The fix set was fixed before this pass, and a finding whose repair leaves it stops the loop even where it opens no new question. Three tells make the stop mandatory as well as instructed: findings rose 3 to 5, Blockers are flat at 2, and the findings cluster on the instrument for the tenth pass running. Floor 3, derived from docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md, read fresh at this pass: risk high (2), security none (0), max 2, not 0. --- .../gate-a-plan-om0bdd7udh-pass-43.md | 6 ++++ .../gate-a-plan-om0bdd7udh-resume.md | 31 +++++++++++++++++-- 2 files changed, 34 insertions(+), 3 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-43.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-43.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-43.md new file mode 100644 index 0000000..544119e --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-43.md @@ -0,0 +1,6 @@ +BLOCKER | high | Task 15 step 8, 8b | The copy that is supposed to pin the revalidated closing-message bytes is unguarded, so a failed `cp` is immediately followed by the closing `git commit`; an older `.context/loop-rule-validated-msg` can remain at the oracle path. | A retry can commit after the pin failed, and a hook-rewritten body can then be compared with stale bytes rather than the message this invocation validated, allowing condition 6 and cleanup to accept a close whose committed body was never validated. | Guard the `cp` and exit to Failure before `git commit` on any error; create the pin under a temporary name and rename it only after the copy succeeds so a failed refresh cannot leave a usable stale oracle. +BLOCKER | high | Task 15 step 8, condition 6 block | `git log -1 --pretty=format:%B > .context/loop-rule-landed-msg` is unguarded and the following `diff` masks its status; if opening the destination fails, an older landed-message file remains available to the comparison. | When a retry uses the same validated message but a hook changes the new commit body, the stale file can still equal `loop-rule-validated-msg`, so the postcondition passes and cleanup destroys the evidence even though the actual commit body differs. | Guard the extraction and stop in Failure before running `diff`; write to a fresh temporary path and rename it only after `git log` succeeds, or compare the command output without a reusable file. +MAJOR | high | Task 15 step 4 | The base refresh block does not guard `git fetch origin main`; a failed fetch is followed by a successful `git rev-parse origin/main` and `cat`, so the block can exit 0 using an old remote-tracking value while the prose claims the base was fetched. | The version-bump check can be recorded green against a stale base rather than the pull request's current base, defeating the reason this step fetches and pins the revision. | Stop immediately when `git fetch` fails, and guard the `rev-parse` write before allowing the battery to read the recorded object id. +MAJOR | high | Tasks 0-14 commit blocks and Task 15 steps 3 and 8a | The new revision guards staging in only four Task 15 blocks; the other `git add` followed by `git commit` sequences, including 8a's findings-record commit, still let a failed staging command fall through to a commit that can succeed from content already in the index. | A task or closing phase can report a successful snapshot while omitting the files it meant to stage, or can keep mutating after the first failed closing operation instead of entering the bounded handoff; this is especially reachable on Resume, which deliberately permits a live staged index. | Guard every staging operation before its commit and, where later steps rely on the commit's exact inputs, pin the resulting index and compare the commit tree just as the repaired Task 15 blocks do. +MINOR | medium | Task 15 step 7, step one | The new presence test accepts any Git `blob`, but an indexed symlink is also represented as a blob with mode `120000`; the worktree path can therefore be read as a valid findings file while the commit stores only the symlink target string. | Step two may route on the target file's valid contents while the durable pass record committed and later squashed is malformed, so the presence check does not establish the file shape its rationale claims. | Check the indexed mode as well as the object type and accept only regular-file modes for both slot paths before committing them. +END OF FINDINGS (5 total) diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md index d5e96d5..7c1803b 100644 --- a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md @@ -13,10 +13,29 @@ Nothing depends on it; the pass files and the repo are authoritative where this ## Resume here -**Cycle OPEN and UNCLEAN. The plan stands at `0bb1d5d` and that is the anchor** — later commits on +**Cycle OPEN and UNCLEAN. The plan stands at `8361f0a` and that is the anchor** — later commits on this branch are records, not plan edits, so check `git log --oneline -1 -- docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md` -rather than `HEAD`. Pass 42 reviewed `0bb1d5d` and found **two Blockers and one Major**. -`.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-42.md` holds them. +rather than `HEAD`. Pass 43 reviewed `8361f0a` and found **two Blockers, two Majors and one Minor**. +`.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-43.md` holds them. + +**Nothing pass 43 raised is inside the assigned fix set, and that is the whole finding.** The tenth +revision repaired pass 42's three findings in the four blocks they name, and **pass 43 re-raised none +of them**. Its two Blockers and two Majors are the *same class* — an unguarded command whose status a +following command masks — at **other** locations: 8b's `cp` pinning the validated message, condition +6's `git log … > file` followed by `diff`, step 4's `git fetch`, and every remaining `git add` +before a `git commit` across Tasks 0–14 and step 8a. Its one in-set finding is a **Minor**: the new +blob check accepts a symlink, since git stores one as a blob with mode `120000`. + +**So the loop stops here rather than absorbing them.** §5: the fix set was fixed before this pass — +pass 42's three findings and the four named blocks — and a finding whose repair leaves that set stops +the loop even when it opens no new question. Finding 4 is explicitly the plan-wide sweep Daniel's +assignment excluded. **Three tells** also make the stop mandatory: findings rose 3 → 5, Blockers are +flat at 2, and the findings cluster on the instrument for the tenth pass running. + +**The open question is one Daniel decides, not a repair to start:** whether to widen the fix set to +the same class plan-wide — which is finding 4, and is the audit the last four assignments each +declined — or to close only the in-set Minor, or to leave the class where it is. The Minor alone does +not iterate. **One unrelated commit sits in this history and is NOT part of this cycle.** `d4a87c6` records an OpenWolf assessment — a separate task authorized 2026-09-16 — as `docs/openwolf-assessment.md`, a @@ -70,6 +89,10 @@ count here. ### Still collected and deliberately unrepaired (pass 34's Minors and Nit) +**MINOR (pass 43, the only in-set finding)** — step one's new blob check accepts a **symlink**: git +stores one as a blob with mode `120000`, so the check would have to read the indexed mode and admit +regular-file modes only. Collected, not iterated, per Mechanics · Severity. + **MINOR** — step 8 is titled "two invocations" and carries four fenced blocks, and the index calls the last two "8b blocks"; **MINOR** — the Architecture paragraph still says every task runs a discriminating pair, which the disposition table replaced; **NIT** — @@ -264,6 +287,8 @@ worth checking before a pass rather than after. | 41 | 4c9ed3c | 2→**1** | 1→**0** | 1→**1** | yes | **UNCLEAN — reported to Daniel, no repair round opened. Best pass since 35, first zero-Blocker pass since 37.** MAJOR, **this round's own repair**: step 7 step three's `git commit` is unguarded and the head is written unconditionally after it, so a failed commit records the **old** head and the next call reviews a tree without the repair — demonstrated with a rejecting `pre-commit` hook. 8a guards its commit exactly this way; step one has the same shape with a smaller blast radius. **Only one tell** (instrument cluster) — the stop is instructed, not mandated | | — | — | — | — | — | — | **NINTH BOUNDED REVISION**, Daniel's assignment after pass 41: guard step three's repair commit (exiting **before** the tip is touched), and guard step one's records commit as a **separate** bounded repair of a different shape — its commit is already terminal, so the guard exists because what follows is *prose*. **Inspection, scoped to commit-containing blocks and reported with that limit:** 21 of 54 fenced blocks commit; 18 end with the commit; 8a is guarded; 8b is terminal and condition 6 re-establishes it independently; step three was the only masking sequence. 21 checks × 3 shells. Commit `0bb1d5d` | | 42 | 0bb1d5d | 1→**3** | 0→**2** | 1→**1** | yes | **UNCLEAN — reported to Daniel, no repair round opened.** **BLOCKER: step one stages the plan** while forbidding a premature repair, so a post-review plan edit is committed before the ordering is read — pass 39's violation through another file. **BLOCKER: every `git add` before the new guards is unguarded**, so a failed staging falls through to a commit that succeeds on already-staged content. MAJOR: both reviewed-head writes are unchecked and `cat` masks a failed write — **demonstrated**, the block exits 0 on a stale value. **Two are mine** (the narrowing dropped the original `&&` and kept the plan in the list); **step 6's Major and step 4b's add are pre-existing**, and step 6 lay outside my inspection's declared limit. Three tells → mandatory stop | +| — | — | — | — | — | — | **TENTH BOUNDED REVISION**, Daniel's assignment after pass 42, narrowed by his reviewer to the three findings and the four blocks they name — **no plan-wide search**. Step one drops the plan from its staging list and requires it unchanged first (HEAD↔index, then index↔worktree), since `git commit` commits the index. Every `git add` guarded; each commit checked against a `git write-tree` pin. **Step one additionally requires both slot paths to be blobs in the pin** — `git add` stages a *removal* as readily as a change, so a tracked findings file deleted before staging is staged as gone and pin and commit then agree without it; an existence test alone would also accept a directory at the path. Steps three and 4b take no presence check: an authorized repair may delete a file. Both reviewed-head writes guarded, both `cat`s gone. 20 checks × 3 shells in a disposable repo, all green; **three failures in the first run were all harness bugs** and are recorded as such. Commit `8361f0a` | +| 43 | 8361f0a | 3→**5** | 2→**2** | 1→**2** | yes | **UNCLEAN — reported to Daniel, no repair round opened.** **All three pass-42 repairs held; nothing was re-raised against the four repaired blocks.** Every Blocker and Major is the **same class at another location**: 8b's unguarded `cp` of the validated message, condition 6's unguarded `git log … > file` masked by the following `diff`, step 4's unguarded `git fetch`, and every remaining unguarded `git add` across Tasks 0–14 and 8a — **that last one is the plan-wide sweep the assignment excluded**. The only in-set finding is a MINOR: the new blob check accepts a **symlink** (blob, mode `120000`), so the indexed mode would have to be checked too. Three tells → mandatory stop | | 26 | 05ec1b2 | **5** | **1** | 2 | yes | first pass on the revised plan. Resume audited progress by `WIP:` commits and called any non-`WIP` commit a stale base — the handoff leaves two other shapes, one of them `HEAD` **at** the base with everything in the index | | 27 | f69db7d | 5→**6** | 1→**3** | 2→**1** | yes | the close required the commit's **tree to equal the tip's**, which target §I **parks**; `reset --soft` leaves the index untouched, so the prose beside it was wrong about git; Resume still inferred staleness from a commit subject where §A3 names that state as reachable | | 28 | f2e5dda | 6→**6** | 3→**2** | 1→**2** | yes | pass 27's stale-base fix reached one site of three; its reset paragraph stated the index behaviour correctly and repeated the false claim four lines later; nothing checked the closing commit's **parent**; the message was validated in the file, where `commit-msg` hooks rewrite git's copy after `-F` reads it | From 4752a355769d4a99fc450bb7f7a51a6b2158abe4 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Wed, 16 Sep 2026 19:13:15 +0200 Subject: [PATCH 168/181] docs(plans): guard every staging, and stop on a failed pin, extraction or fetch Pass 43's two Blockers and two Majors, bounded to the locations the findings name. The Minor stays collected. No other command-failure class was searched for and no other reported location was repaired. Finding 1, Task 15 step 8b. The cp that pins the revalidated closing message ran unguarded with the closing commit on the next line. It is guarded now, and any pin an earlier attempt left is removed first, so a failed refresh leaves condition 6 with no oracle rather than a stale one its test -s would accept. 8b's commit stays terminal and unguarded, as the ninth revision's inspection established: condition 6 re-establishes independently that it landed. Finding 2, condition 6. The extraction of the committed body redirected without a guard and the diff after it masked the status. Same shape, same repair: remove the previous extraction, then guard this one. Without it a stale landed message could equal the validated one and the close would be accepted on bytes nobody read out of that commit. Finding 3, Task 15 step 4. git fetch origin main ran unguarded before git rev-parse origin/main, which resolves the stale remote-tracking ref happily, so the block could record a base the pull request does not have and exit 0. Both the fetch and the recording are guarded, and the cat that would have printed an older recorded value is gone. Finding 4, seventeen locations, enumerated rather than described. Task 0's fragment sweep; the eight prompt-copy installs in Tasks 1, 3, 4, 5, 6, 7, 8 and 9; Task 10's hook install and Task 11's test sweep; Task 12's equivalence record, Task 12b's sweep record and Task 13's table; Task 14's alignment commit; Task 15 step 3's version bump; and Task 15 step 8a's findings record. Every git add in them is guarded. Guards everywhere, content comparisons only where a consumer reads that commit. A new Global Constraint states the rule and names the three places that have such a consumer: step 7 step one, whose record step two then reads, and step 7 step three and step 4b, whose commits bound the range the next Gate-B call reviews. The per-task snapshots have none - every check this plan runs afterwards reads the worktree, and Gate B reads the accumulated $BASE..HEAD range rather than any one commit - so they are guarded and not pinned. 8a needs no addition: Close condition 4 already pins its blobs before staging and compares them against the committed tree. Verified by execution. The four affected blocks were extracted from this plan verbatim and run in a disposable repository: 18 checks under sh, dash and bash, all green. A stale pin replaced on the success path; a failed pin with a stale file present and with none, neither producing a closing commit; a stale landed message equal to the validated one with a failed extraction, which is the case that would otherwise pass; a committed body that differs, whose existing route is unchanged; a failed fetch with the remote-tracking ref still resolvable, recording no base ref; a failed base-ref write; and a failed staging with foreign content already in the index, which commits nothing. Two harness errors are recorded rather than quietly fixed. A writable file in a read-only directory is not a failed write - cp and a redirect both truncate it and succeed - so the first model of the failure was wrong and the real case needs the destination itself unwritable. And 8b resets to $BASE before it pins, so HEAD moving is expected; the assertion that matters is that no closing commit was created. Unchanged and not repaired: four fenced blocks fail sh -n and dash -n through process substitution in their parity diffs. Identical before this change, and outside the assignment. --- .../2026-09-14-loop-rule-consolidation.md | 79 +++++++++++++------ 1 file changed, 57 insertions(+), 22 deletions(-) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index 29a234a..193b7d8 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -28,6 +28,7 @@ It adds no task, prerequisite or closure condition to this plan. - **Invariant 11:** every prompt change passes all 12 items of `docs/prompt-standards.md`. Most at risk: item 6 (every constraint carries its reason in the same sentence) and item 8 (token-lean). - **Invariant 4 / the hook:** no control flow, counter, fingerprint computation, routing or event handling changes. Only `note` strings and their expectations. - **Gate-B cycle discipline:** snapshot commits are named `WIP: …`, and a non-`WIP` commit mid-cycle resets the hook's counters. **This change closes with `git reset --soft "$BASE"` followed by one commit, not with `--amend`** — Mechanics prescribes the reset shape wherever several WIP snapshots piled up, and this plan makes one per task. Task 15 step 8 is the operation. +- **Every `git add` in this plan is guarded, and a guard is where most of them stop.** A failed staging otherwise falls through to the `git commit` on the next line, which **can still succeed on whatever the index already held** — this plan states in Task 15 step 7 that the index may carry foreign staged content, and Resume deliberately admits a live staged index — so the block reports a snapshot it did not take. **A content comparison after the commit is added only where a later step reads that commit rather than the worktree**, and that is three places: step 7's step one, whose record step two then reads; step 7's step three and step 4b, whose commits bound the range the next Gate-B call reviews. The per-task snapshots have no such consumer — every check this plan runs afterwards reads the worktree, and Gate B reads the accumulated `$BASE..HEAD` range rather than any one commit — so they are guarded and not pinned. **8a is the exception that needs no addition**: Close condition 4 already pins its blobs before staging and compares them against the committed tree. --- @@ -1150,7 +1151,8 @@ row, and the reading result for `F1`, `F2` and `F3`, whose §F notes give no lin region cannot be built mechanically. ```bash -git add docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +git add docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md \ + || { echo "staging FAILED — nothing is committed here; stop and fix the staging"; exit 1; } git commit -m "WIP: fragment sweep at the recorded base" ``` @@ -1237,7 +1239,8 @@ Expected: no output. ```bash git add CLAUDE.md plugins/dev-workflow/commands/workflow-init.md \ - docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md + docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md \ + || { echo "staging FAILED — nothing is committed here; stop and fix the staging"; exit 1; } git commit -m "WIP: install the closure ordering into both §5 copies" ``` @@ -1388,7 +1391,8 @@ as inherited drift. **Any difference other than the parenthetical is a failure o ```bash git add CLAUDE.md plugins/dev-workflow/commands/workflow-init.md \ - docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md + docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md \ + || { echo "staging FAILED — nothing is committed here; stop and fix the staging"; exit 1; } git commit -m "WIP: replace passage (b) with the absorb paragraph that owns the fix set" ``` @@ -1532,7 +1536,8 @@ here; `c10`–`c14` gone from here. ```bash git add CLAUDE.md plugins/dev-workflow/commands/workflow-init.md \ - docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md + docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md \ + || { echo "staging FAILED — nothing is committed here; stop and fix the staging"; exit 1; } git commit -m "WIP: widen the clearly-stuck third condition and split its precedence sentence" ``` @@ -1635,7 +1640,8 @@ the class alone does not tell you: ```bash git add CLAUDE.md plugins/dev-workflow/commands/workflow-init.md \ - docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md + docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md \ + || { echo "staging FAILED — nothing is committed here; stop and fix the staging"; exit 1; } git commit -m "WIP: read the two-tell threshold after the clean-completion branch" ``` @@ -1733,7 +1739,8 @@ Expected: no output. This passage should now be byte-identical, `g4` having been ```bash git add CLAUDE.md plugins/dev-workflow/commands/workflow-init.md \ - docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md + docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md \ + || { echo "staging FAILED — nothing is committed here; stop and fix the staging"; exit 1; } git commit -m "WIP: answer the demotion question and scope the resolve duty to the fix set" ``` @@ -1897,7 +1904,8 @@ Expected: `1` each in the worktree, and `parent=1 worktree=1` in each copy for e ```bash git add CLAUDE.md plugins/dev-workflow/commands/workflow-init.md \ - docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md + docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md \ + || { echo "staging FAILED — nothing is committed here; stop and fix the staging"; exit 1; } git commit -m "WIP: install the one-contract paragraph and the remaining prompt-copy replacements" ``` @@ -2001,7 +2009,8 @@ State the number you observed. **Do not carry a count from §F into a check** ```bash git add CLAUDE.md plugins/dev-workflow/commands/workflow-init.md \ - docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md + docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md \ + || { echo "staging FAILED — nothing is committed here; stop and fix the staging"; exit 1; } git commit -m "WIP: replace the fourteen falsified sentences in the two prompt copies" ``` @@ -2047,7 +2056,8 @@ The hook reporting its own threshold as an obligation at a floor of 1 is **not** ```bash git add CLAUDE.md plugins/dev-workflow/commands/workflow-init.md \ - docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md + docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md \ + || { echo "staging FAILED — nothing is committed here; stop and fix the staging"; exit 1; } git commit -m "WIP: correct the Named residual's blanket exemption and the work-loop sequence" ``` @@ -2189,7 +2199,8 @@ point where green is expected. ```bash git add plugins/dev-workflow/hooks/codex-gate.sh \ - docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md + docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md \ + || { echo "staging FAILED — nothing is committed here; stop and fix the staging"; exit 1; } git commit -m "WIP: replace the seven gate reminders the ordering falsifies" ``` @@ -2275,7 +2286,8 @@ Expected: exit 0. **The `--exclude=SC2015` is a single-code exclusion**, not a b ```bash git add plugins/dev-workflow/hooks/codex-gate.test.sh \ - docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md + docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md \ + || { echo "staging FAILED — nothing is committed here; stop and fix the staging"; exit 1; } git commit -m "WIP: move every hook assertion that names a replaced reminder string" ``` @@ -2320,7 +2332,8 @@ fix changes either predicate, and reaches the closing commit as an assertion tha happened. ```bash -git add docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +git add docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md \ + || { echo "staging FAILED — nothing is committed here; stop and fix the staging"; exit 1; } git commit -m "WIP: record the b11/b13 equivalence result" ``` @@ -2382,7 +2395,8 @@ ran — and one recorded only in `.context/` is unrecorded as far as the commit - [ ] **Step 6: Commit** ```bash -git add docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +git add docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md \ + || { echo "staging FAILED — nothing is committed here; stop and fix the staging"; exit 1; } git commit -m "WIP: record the completeness sweep" ``` @@ -2438,7 +2452,8 @@ Two, per design §7: that a logical pass was validated across every required bra - [ ] **Step 6: Commit** ```bash -git add docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +git add docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md \ + || { echo "staging FAILED — nothing is committed here; stop and fix the staging"; exit 1; } git commit -m "WIP: next-state table and per-condition closure checks" ``` @@ -2586,7 +2601,8 @@ would catch an identical accidental edit in both copies. - [ ] **Step 5: Commit** ```bash -git add CLAUDE.md plugins/dev-workflow/commands/workflow-init.md docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +git add CLAUDE.md plugins/dev-workflow/commands/workflow-init.md docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md \ + || { echo "staging FAILED — nothing is committed here; stop and fix the staging"; exit 1; } git commit -m "WIP: align the two copies and record the divergence list" ``` @@ -2614,7 +2630,8 @@ grep -n '"version"' plugins/dev-workflow/.claude-plugin/plugin.json - [ ] **Step 3: Commit the bump into the WIP snapshot** ```bash -git add plugins/dev-workflow/.claude-plugin/plugin.json plugins/dev-workflow/CHANGELOG.md +git add plugins/dev-workflow/.claude-plugin/plugin.json plugins/dev-workflow/CHANGELOG.md \ + || { echo "staging FAILED — nothing is committed here; stop and fix the staging"; exit 1; } git commit -m "WIP: bump dev-workflow to 0.12.0" ``` @@ -2630,9 +2647,14 @@ comparison. **Fetch, then pass the fetched ref to the checker itself** — an ea what supplied a comparison it had not supplied: ```bash -git fetch origin main -git rev-parse origin/main > .context/loop-rule-baseref # resolve ONCE; record this object name -cat .context/loop-rule-baseref +# Both guarded, and no `cat`. A failed fetch leaves whatever an earlier one wrote in +# `origin/main`, and `git rev-parse` resolves that stale ref happily — the block would +# exit 0 having recorded a base the pull request does not have. A failed write is the +# same shape: the `cat` after it printed the older recorded value. +git fetch origin main \ + || { echo "fetch FAILED — origin/main is whatever an earlier fetch left; do not run the battery"; exit 1; } +git rev-parse origin/main > .context/loop-rule-baseref \ + || { echo "recording the base ref FAILED — do not run the battery"; exit 1; } ``` **Pass that object name to `check-version-bump.sh`, not `origin/main`.** A remote-tracking ref is @@ -3127,7 +3149,8 @@ test -s .context/loop-rule-closing-msg || { echo "closing message missing or emp # Condition 4, first bullet: pin the validated blobs before staging. for f in $FINAL; do git hash-object "$f"; done > .context/loop-rule-final-blobs # shellcheck disable=SC2086 -git add $FINAL +git add $FINAL \ + || { echo "staging FAILED — run Failure; do NOT commit"; exit 1; } git commit -m "WIP: Gate-B findings files" || { echo "record commit FAILED — run Failure"; exit 1; } test "$(git rev-parse HEAD^)" = "$HEADREV" || { echo "record commit's parent is not the reviewed head — run Failure"; exit 1; } git diff --quiet "$HEADREV" HEAD -- "$@" \ @@ -3171,7 +3194,13 @@ test -s .context/loop-rule-closing-msg || { echo "closing message missing or emp # oracle must not be the same mutable path the commit reads: a `commit-msg` hook # that rewrites git's copy AND this ignored source file would otherwise make the # post-close comparison pass on a body nobody validated. -cp .context/loop-rule-closing-msg .context/loop-rule-validated-msg +# Remove any pin an earlier attempt left BEFORE writing this one, so a failed refresh +# leaves condition 6 with no oracle rather than a stale one its `test -s` would accept +# — and guard the copy, because the commit is the very next command and a failed pin +# must not fall through into it. +rm -f .context/loop-rule-validated-msg +cp .context/loop-rule-closing-msg .context/loop-rule-validated-msg \ + || { echo "pinning the validated closing message FAILED — run Failure; do NOT commit"; exit 1; } # `--cleanup=verbatim` so the stored body is the validated bytes: git's default # cleanup for -F strips trailing whitespace and collapses blank runs, and # condition 6 compares bytes. It carries no `-m`, so `is_wip_commit` still misses it. @@ -3195,7 +3224,13 @@ test -z "$(git status --porcelain)" || { echo "tree dirty after the close — ru # tidied copy. Both observed in a disposable repository, so this diff is the # byte equality condition 6 states, not a normalized stand-in for it. test -s .context/loop-rule-validated-msg || { echo "no pinned message — 8b did not complete; run Failure"; exit 1; } -git log -1 --pretty=format:%B > .context/loop-rule-landed-msg +# Same shape as 8b's pin: remove the previous extraction first, then guard this one. +# A failed redirect otherwise leaves an earlier body at the path and the `diff` below +# masks the failure — it can even pass, when the earlier body was the same validated +# message, so the close is accepted on bytes nobody read out of THIS commit. +rm -f .context/loop-rule-landed-msg +git log -1 --pretty=format:%B > .context/loop-rule-landed-msg \ + || { echo "extracting the committed body FAILED — run Failure; the comparison has no input"; exit 1; } diff .context/loop-rule-validated-msg .context/loop-rule-landed-msg \ || { echo "the committed body differs from the validated message — run Failure"; exit 1; } ``` From cdb6a236ff41041210b17587ecf32120f0c4e991 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Wed, 16 Sep 2026 19:22:57 +0200 Subject: [PATCH 169/181] docs(context): record Gate-A plan pass 44 and the eleventh bounded revision Pass 44 reviewed the plan at 4752a35 and is valid: the terminator is exact, the file carries nine finding lines and nothing else. Two Blockers, seven Majors. Nothing was re-raised against the eleventh revision. Neither the four repaired blocks nor the seventeen staging guards drew a finding, and the same was true of pass 43 against the tenth revision. The class widened again instead, from a status that is masked to a status that is discarded. test -z around a git status substitution swallows git's exit code, so a failed status reads as a clean tree - that is Preparation and closure conditions 2, 5 and 6, and it is one of the Blockers. Two git rev-parse substitutions inside one equality test both resolve empty and compare equal, in Preparation's and Resume's approved-input checks. Two unchecked grep substitutions compare two empty extractions as equal in Task 0 step 3, the other Blocker. A pipe without pipefail in Task 10 step 4. Unguarded scratch writes and a rename masked by cat in Task 0 step 3 and Task 14 step 1. An unguarded hash-object pin loop in 8a. A git log inside a case substitution in condition 6's subject check. This is not a plateau and the record should not read as one. The clearly-stuck condition requires Blockers and Majors regenerating from the repairs, and none of these do: they are pre-existing material in blocks nobody has swept for this class. The plan carries 54 fenced blocks. Three tells make the stop mandatory as well as instructed: findings rose 5 to 9, Blockers are flat at 2 for a third pass, and the findings cluster on the instrument for the eleventh. Two findings are escalation triggers rather than work to start. Finding 2 would change how closure conditions 2, 5 and 6 spell their clean-tree and dirty-set tests; finding 7 changes 8a's pre-move prerequisite ordering. The assignment names a change to existing closure conditions as a stop. Floor 3, derived from docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md, read fresh at this pass: risk high (2), security none (0), max 2, not 0. --- .../gate-a-plan-om0bdd7udh-pass-44.md | 10 ++++ .../gate-a-plan-om0bdd7udh-resume.md | 59 ++++++++++++------- 2 files changed, 47 insertions(+), 22 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-44.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-44.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-44.md new file mode 100644 index 0000000..47ef37a --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-44.md @@ -0,0 +1,10 @@ +BLOCKER | high | Task 0 step 3, squash-carry baseline | Accounting row 19 promises literal uniqueness and a non-empty extraction for every baseline site, but the single-line squash site diffs two unchecked grep substitutions and then writes a site record; a missing anchor in both copies yields two empty inputs that compare equal, and duplicate hits also pass | The baseline can certify parity for the j1–j4 site without extracting the intended line, making the accounting row false and leaving an inventoried site unproved | Assert exactly one literal hit in each base blob, capture each extracted line under a guarded command, require both non-empty, and only then diff and write the site record +BLOCKER | high | Preparation and Task 15 step 8 conditions 2, 5, and 6 | Each clean or empty-set predicate uses test -z around git status command substitution, which discards git status's exit status; a failed status command produces empty stdout and is therefore accepted as clean | Preparation can record a base without establishing a clean tree, and the close can commit, reset, accept its postconditions, and delete recovery state without establishing the exact dirty set or either clean-tree condition | Capture each git status result with an explicit guarded assignment or file write, fail on a nonzero status, and test emptiness only after success +MAJOR | high | Resume, Establish how far the implementation got | The four-source reconciliation runs git log, staged diff, unstaged diff, and status as consecutive unguarded commands, so a failure in any of the first three is masked by a later success even though Resume says all four are always read | Resume can omit committed, indexed, or worktree content, mis-mark tasks, and repeat edits or route a close using an incomplete view of the interrupted state | Guard each source read independently and stop the reconciliation unless all four commands succeeded +MAJOR | high | Task 10 step 4 | The zero-context hook diff is piped through grep and sed without pipefail or explicit status capture, and a following note-call grep becomes the block's final status | A failed git diff can be reported as an empty changed-line list while the task proceeds, so a hook control-flow or routing edit can escape the invariant-4 review | Write the diff to a temporary file under a guard, inspect that file for both hunk locations and every changed line, and guard the later locator separately +MAJOR | high | Task 0 step 3, baseline artifact construction | The base and site records, diff output appends, and final rename are unguarded; the trailing cat masks a failed rename, while failed writes can still leave or promote a partial or stale baseline | The task can exit successfully with a baseline that was not the one just computed, and first-entry execution does not run Resume's later completeness check before consumers use it | Guard every artifact write and append, guard the rename, and display the destination only after the guarded rename succeeds +MAJOR | high | Task 14 step 1 | Both extracted regions are written to reusable scratch files with unchecked printf redirections, while diff status 1 is deliberately accepted as an ordinary difference | A failed or partial write can compare stale or truncated content and be recorded as a real parity result, allowing the durable site list and divergence classification to describe bytes that were never extracted in this run | Remove or replace each scratch file through guarded writes to fresh temporary paths, verify both writes, then run and classify the diff +MAJOR | high | Task 15 step 8a, condition 4 first bullet | The hash-object loop that is supposed to pin both validated findings blobs is unguarded; an early hash failure is masked when the last iteration succeeds, and staging plus the record commit then run before the incomplete pin is detected by the later comparison | A pre-move prerequisite can fail yet the procedure still mutates HEAD, turning a failure that should stop before staging into a Failure handoff after a commit lands | Guard every hash-object call and the output write as one prerequisite, verify that both expected blob records exist, and do not stage until the pin succeeds +MAJOR | medium | Preparation and Resume approved-input checks | Each approved blob comparison embeds two unguarded git rev-parse calls inside one equality test; if both lookups fail, both substitutions are empty and the equality succeeds | The procedures can accept the target text, design, or condition inventory without proving either side resolves to the approved blob, invalidating the source basis for every task | Resolve and guard each object id separately, require both non-empty, then compare the two recorded ids +MAJOR | medium | Task 15 step 8, condition 6 subject check | The closing-subject predicate runs git log inside a case substitution without checking its status, so a failed subject read becomes an empty non-WIP subject and execution continues | The close can skip one of condition 6's four required postconditions and proceed through body comparison and cleanup without establishing that the closing commit is not a snapshot | Read the subject into a variable under an explicit guard, require the read to succeed, and only then apply the WIP pattern check +END OF FINDINGS (9 total) diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md index 7c1803b..2b479f9 100644 --- a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md @@ -13,29 +13,42 @@ Nothing depends on it; the pass files and the repo are authoritative where this ## Resume here -**Cycle OPEN and UNCLEAN. The plan stands at `8361f0a` and that is the anchor** — later commits on +**Cycle OPEN and UNCLEAN. The plan stands at `4752a35` and that is the anchor** — later commits on this branch are records, not plan edits, so check `git log --oneline -1 -- docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md` -rather than `HEAD`. Pass 43 reviewed `8361f0a` and found **two Blockers, two Majors and one Minor**. -`.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-43.md` holds them. - -**Nothing pass 43 raised is inside the assigned fix set, and that is the whole finding.** The tenth -revision repaired pass 42's three findings in the four blocks they name, and **pass 43 re-raised none -of them**. Its two Blockers and two Majors are the *same class* — an unguarded command whose status a -following command masks — at **other** locations: 8b's `cp` pinning the validated message, condition -6's `git log … > file` followed by `diff`, step 4's `git fetch`, and every remaining `git add` -before a `git commit` across Tasks 0–14 and step 8a. Its one in-set finding is a **Minor**: the new -blob check accepts a symlink, since git stores one as a blob with mode `120000`. - -**So the loop stops here rather than absorbing them.** §5: the fix set was fixed before this pass — -pass 42's three findings and the four named blocks — and a finding whose repair leaves that set stops -the loop even when it opens no new question. Finding 4 is explicitly the plan-wide sweep Daniel's -assignment excluded. **Three tells** also make the stop mandatory: findings rose 3 → 5, Blockers are -flat at 2, and the findings cluster on the instrument for the tenth pass running. - -**The open question is one Daniel decides, not a repair to start:** whether to widen the fix set to -the same class plan-wide — which is finding 4, and is the audit the last four assignments each -declined — or to close only the in-set Minor, or to leave the class where it is. The Minor alone does -not iterate. +rather than `HEAD`. Pass 44 reviewed `4752a35` and found **two Blockers and seven Majors**. +`.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-44.md` holds them. + +**Two rounds in a row, the repairs held and the reviewer widened the class instead.** Pass 43 +re-raised nothing against the tenth revision's four blocks; pass 44 re-raises nothing against the +eleventh's four findings and seventeen staging guards. What changes each round is the **class**: + +- pass 42 → an unguarded `git add` falling through to its `git commit`; +- pass 43 → the same shape at four other commands (`cp`, a redirect, `git fetch`, every remaining + `git add`); +- pass 44 → **status discarded rather than masked**: `test -z "$(git status …)"` swallows git's exit + code and reads a failed command as a clean tree; two `$(git rev-parse …)` inside one equality test + both resolve empty and compare equal; a pipe with no `pipefail`; unguarded scratch writes and an + unguarded `hash-object` loop. + +**This is not a plateau, and saying so matters.** The "clearly stuck" reading needs Blockers and +Majors **regenerating from the repairs**, and none of pass 44's do — they are newly discovered +pre-existing material in areas nobody has ever swept for this class. The plan carries **54 fenced +blocks**; the class is real wherever a status is read and discarded, and each round reaches further +into them. + +**Three tells stand anyway, so the stop is mandatory**: findings rose 5 → 9, Blockers are flat at 2 +for a third pass, and the findings cluster on the instrument for the eleventh. + +**Two of pass 44's findings are escalation triggers, not work to start.** Finding 2 would change how +**closure conditions 2, 5 and 6** spell their clean-tree and dirty-set tests, and finding 7 changes +8a's pre-move prerequisite ordering. Daniel's assignment names exactly that as a stop: *"Stop before +expanding scope if another location, a new contract decision, or a change to existing closure +conditions becomes necessary."* + +**The open question is his, not a repair to start:** whether to sweep all 54 blocks for discarded +status in one deliberate pass, or to keep taking the class a named handful at a time, or to draw the +line here and accept it as a disclosed residual of the plan. **The symlink Minor from pass 43 is +still collected and still unrepaired.** **One unrelated commit sits in this history and is NOT part of this cycle.** `d4a87c6` records an OpenWolf assessment — a separate task authorized 2026-09-16 — as `docs/openwolf-assessment.md`, a @@ -289,6 +302,8 @@ worth checking before a pass rather than after. | 42 | 0bb1d5d | 1→**3** | 0→**2** | 1→**1** | yes | **UNCLEAN — reported to Daniel, no repair round opened.** **BLOCKER: step one stages the plan** while forbidding a premature repair, so a post-review plan edit is committed before the ordering is read — pass 39's violation through another file. **BLOCKER: every `git add` before the new guards is unguarded**, so a failed staging falls through to a commit that succeeds on already-staged content. MAJOR: both reviewed-head writes are unchecked and `cat` masks a failed write — **demonstrated**, the block exits 0 on a stale value. **Two are mine** (the narrowing dropped the original `&&` and kept the plan in the list); **step 6's Major and step 4b's add are pre-existing**, and step 6 lay outside my inspection's declared limit. Three tells → mandatory stop | | — | — | — | — | — | — | **TENTH BOUNDED REVISION**, Daniel's assignment after pass 42, narrowed by his reviewer to the three findings and the four blocks they name — **no plan-wide search**. Step one drops the plan from its staging list and requires it unchanged first (HEAD↔index, then index↔worktree), since `git commit` commits the index. Every `git add` guarded; each commit checked against a `git write-tree` pin. **Step one additionally requires both slot paths to be blobs in the pin** — `git add` stages a *removal* as readily as a change, so a tracked findings file deleted before staging is staged as gone and pin and commit then agree without it; an existence test alone would also accept a directory at the path. Steps three and 4b take no presence check: an authorized repair may delete a file. Both reviewed-head writes guarded, both `cat`s gone. 20 checks × 3 shells in a disposable repo, all green; **three failures in the first run were all harness bugs** and are recorded as such. Commit `8361f0a` | | 43 | 8361f0a | 3→**5** | 2→**2** | 1→**2** | yes | **UNCLEAN — reported to Daniel, no repair round opened.** **All three pass-42 repairs held; nothing was re-raised against the four repaired blocks.** Every Blocker and Major is the **same class at another location**: 8b's unguarded `cp` of the validated message, condition 6's unguarded `git log … > file` masked by the following `diff`, step 4's unguarded `git fetch`, and every remaining unguarded `git add` across Tasks 0–14 and 8a — **that last one is the plan-wide sweep the assignment excluded**. The only in-set finding is a MINOR: the new blob check accepts a **symlink** (blob, mode `120000`), so the indexed mode would have to be checked too. Three tells → mandatory stop | +| — | — | — | — | — | — | **ELEVENTH BOUNDED REVISION**, Daniel's assignment after pass 43, scoped to its four Blocker/Major findings; the Minor stays collected. 8b removes any earlier pin **before** writing this one and copies under a guard, so a failed refresh leaves condition 6 **no** oracle rather than a stale one; 8b's commit stays terminal, condition 6 re-establishing it independently. Condition 6's extraction takes the same shape. Step 4 guards the fetch **and** the recording — `git rev-parse origin/main` resolves the stale ref happily after a failed fetch — and its `cat` is gone. **Seventeen staging locations guarded, enumerated not described.** A new Global Constraint states the rule: **guards everywhere, content comparison only where a consumer reads that commit** — three places have one, the per-task snapshots have none, and 8a needed no addition because condition 4 already pins its blobs. 18 checks × 3 shells, all green. **Two harness errors recorded**: a writable file in a read-only directory is *not* a failed write, and 8b resets before it pins. Commit `4752a35` | +| 44 | 4752a35 | 5→**9** | 2→**2** | 2→**7** | yes | **UNCLEAN — reported to Daniel, no repair round opened.** **Nothing re-raised against the eleventh revision.** The class widened again — from *masked* status to **discarded** status: `test -z "$(git status …)"` swallows git's exit code and reads a failed command as a clean tree (Preparation and closure conditions 2, 5, 6 — BLOCKER); two `$(git rev-parse …)` in one equality test both resolve empty and compare equal (Preparation and Resume approved-input checks); an unchecked `grep` substitution pair that compares two empty extractions as equal (Task 0 step 3 — BLOCKER); a pipe with no `pipefail` (Task 10 step 4); unguarded scratch writes and a rename masked by `cat` (Task 0 step 3, Task 14 step 1); an unguarded `hash-object` pin loop (8a condition 4); a `git log` inside a `case` substitution (condition 6's subject check). **Not a plateau** — none regenerates from a repair. **Two are escalation triggers**: finding 2 changes how closure conditions spell their tests, finding 7 changes 8a's pre-move ordering. Three tells → mandatory stop | | 26 | 05ec1b2 | **5** | **1** | 2 | yes | first pass on the revised plan. Resume audited progress by `WIP:` commits and called any non-`WIP` commit a stale base — the handoff leaves two other shapes, one of them `HEAD` **at** the base with everything in the index | | 27 | f69db7d | 5→**6** | 1→**3** | 2→**1** | yes | the close required the commit's **tree to equal the tip's**, which target §I **parks**; `reset --soft` leaves the index untouched, so the prose beside it was wrong about git; Resume still inferred staleness from a commit subject where §A3 names that state as reachable | | 28 | f2e5dda | 6→**6** | 3→**2** | 1→**2** | yes | pass 27's stale-base fix reached one site of three; its reset paragraph stated the index behaviour correctly and repeated the false claim four lines later; nothing checked the closing commit's **parent**; the message was validated in the file, where `commit-msg` hooks rewrite git's copy after `-F` reads it | From fc12f25358b9a8df9bf81eefd3f8acfcff3449b7 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Wed, 16 Sep 2026 19:58:16 +0200 Subject: [PATCH 170/181] docs(plans): stop reading failed operations as valid results Pass 44's nine findings, plus one bounded sweep of every executable block for the same mechanism. This does NOT close the defect class: it covers the mechanisms named in the assignment at the locations this sweep inspected, and a block can discard a status in a way none of those patterns describes. Inspection inventory. All 54 fenced bash blocks were read, in four parallel read-only passes plus a mechanical scan for command substitutions inside tests, pipelines, loops, writes, appends, renames and sequential commands. Repaired, by mechanism. Command substitution whose status was discarded: Preparation's clean-tree check; 8a's exact-dirty-set check and its post-commit clean-tree check; 8b's pre-reset clean-tree check; condition 6's subject read and its post-close clean-tree check. Each becomes a guarded assignment followed by the same test against the captured value. No predicate changed - only how it is established, and every existing failure message is preserved. Two substitutions inside one equality test, where both failing compares equal: Preparation's approved-input comparison at HEAD, and Resume's at $BASE. Each side now resolves under its own guard before the comparison runs. Sequential reads whose failure a later success masked: Resume's four-source reconciliation. The step's rule is that all four ARE read, so each is guarded and a failure stops the reconciliation. Pipelines: Task 10 step 4 captured the diff before filtering it, because the pipeline's status was sed's and sed succeeds on empty input. pipefail is not available - these blocks run under sh and dash too. Loops where only the last iteration's status survived, with an unguarded redirect: 8a's validated-blob pin and its committed-blob read. Both are guarded per iteration and on the redirect, and the in-loop messages go to stderr because the compound's stdout is the artifact. Inverted polarity: 8a's per-file carry check used `git diff --quiet ... && { error }`, so exit 2 and above - an execution failure - took the same path as "carried". It is a case on the status now: 0 not carried, 1 carried, 2+ stop. Unguarded writes, appends and renames: Task 0 step 3's baseline artifact, site list and per-site records, and its rename, which a following cat would otherwise have masked by displaying a previous run's file; Task 14 step 1's two reused scratch extracts; step 7 step three's removal of the previous candidate's tip; and 8a's closing-tip write. Extractions compared without being established: Task 0 step 3's squash-carry site diffed two unchecked greps, so a missing anchor in both copies compared equal and the site record certified a line sed never found - accounting row 19 promises uniqueness and a non-empty extraction for every site, and it now gets both. The same shape appears in two parity checks whose success signal is no output; both now require a non-empty extraction first. A third parity check expects specific non-empty text and is not affected. Justified non-changes. Sixteen task blocks end with their git commit, so the commit's status is the block's and nothing follows to mask it - the plan's own stated rule, and guarding them would contradict it. grep exit 1 and diff exit 1 are legitimate results throughout and are left alone; only status 2 and above is an execution failure, which the affected sites already test for. BASE=$(cat ...) and its siblings discard a status but are closed by the test -n that follows. test "$(git rev-parse ...)" = "$X" against an already non-empty value fails closed on a failed lookup. The cleanup's ls glob is unguarded only where .context/ cannot be read, and the alternative needs non-POSIX find -maxdepth; left as is. One correction. The claim that removing the pin before copying leaves condition 6 with no oracle is wrong, and was mine. Where the directory is unwritable the rm fails too and the earlier pin survives. The guard is what does the work: it stops before the commit, so nothing closes against a pin this invocation did not write. The prose now says that and reports the surviving state without promising cleanup. Verified by execution. Fourteen fixtures under sh and bash, thirteen under dash, all green, with the four affected blocks extracted from this plan verbatim and a stub git failing one named subcommand at a time. Observed: a failed status does not read as a clean tree, at Preparation, 8a, 8b and condition 6; two failed rev-parse lookups stop instead of comparing equal; a failed diff in the pipeline stops instead of printing an empty changed-line list; a failed hash-object and a failed committed-blob read each stop the pin; a failed git log is not read as a non-WIP subject; a missing squash-carry anchor stops; and every success path still passes. Coverage limits, stated rather than implied. The dash skip is Task 0 step 3, which uses process substitution and is bash-only - pre-existing and unchanged. Four blocks still fail sh -n and dash -n for that same reason, identically to before this change. Untested: the rename failure and the append failures in Task 0 step 3, the Task 14 scratch-write failure, the Resume four-source guards and step 7's tip removal - each is the same guard shape as one that was exercised, which is a reason to expect them to hold, not evidence that they do. Escalated and deliberately not done: condition 6 accepts an empty subject read successfully, and requiring a non-empty one would be a new closure condition; $FINAL emptiness is unguarded across 8a, and making it a stop adds a precondition that does not exist today; Task 14 step 1's while loop returns 0 when its site list is empty, and requiring at least one site is likewise a new condition. The pass-43 symlink Minor stays collected. --- .../2026-09-14-loop-rule-consolidation.md | 200 ++++++++++++++---- 1 file changed, 160 insertions(+), 40 deletions(-) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index 193b7d8..8342b92 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -342,12 +342,22 @@ that **no base file exists yet** — if one does, this is a re-entry and Resume ```bash test "$(git rev-parse --abbrev-ref HEAD)" = loop-rule-consolidation || { echo "wrong branch"; exit 1; } -test -z "$(git status --porcelain)" || { echo "tree not clean — if this is a re-entry, run Resume"; exit 1; } +# Two steps, because a FAILED `git status` prints nothing and `test -z ""` would read +# that as a clean tree. The predicate is unchanged; only its establishment is. +TREESTATE=$(git status --porcelain) \ + || { echo "reading the tree state FAILED — cleanliness is unestablished; stop"; exit 1; } +test -z "$TREESTATE" || { echo "tree not clean — if this is a re-entry, run Resume"; exit 1; } test ! -e .context/loop-rule-base || { echo "a base is already recorded — this is a re-entry, run Resume"; exit 1; } git merge-base --is-ancestor ba15e83 HEAD || { echo "ba15e83 not in this history"; exit 1; } for f in target-text design condition-inventory; do p="docs/superpowers/specs/2026-09-10-loop-rule-consolidation-$f.md" - test "$(git rev-parse "HEAD:$p")" = "$(git rev-parse "ba15e83:$p")" \ + # Resolved separately: with both substitutions inline, a path that resolves at + # NEITHER revision gives "" = "" and the comparison PASSES on two failed lookups. + HEAD_ID=$(git rev-parse "HEAD:$p") \ + || { echo "$p does not resolve at HEAD — the comparison is unestablished; stop"; exit 1; } + APPROVED_ID=$(git rev-parse "ba15e83:$p") \ + || { echo "$p does not resolve at ba15e83 — the comparison is unestablished; stop"; exit 1; } + test "$HEAD_ID" = "$APPROVED_ID" \ || { echo "$p differs at HEAD from its approved version"; exit 1; } done ``` @@ -575,7 +585,13 @@ test "$(git rev-parse --verify "$BASE^{commit}")" = "$BASE" || { echo "base does git merge-base --is-ancestor "$BASE" HEAD || { echo "base is NOT an ancestor of HEAD"; exit 1; } for f in target-text design condition-inventory; do p="docs/superpowers/specs/2026-09-10-loop-rule-consolidation-$f.md" - test "$(git rev-parse "$BASE:$p")" = "$(git rev-parse "ba15e83:$p")" \ + # Resolved separately, for the reason Preparation's copy gives: two inline + # substitutions that both fail compare equal and the check passes on nothing. + BASE_ID=$(git rev-parse "$BASE:$p") \ + || { echo "$p does not resolve at \$BASE — the comparison is unestablished; stop"; exit 1; } + APPROVED_ID=$(git rev-parse "ba15e83:$p") \ + || { echo "$p does not resolve at ba15e83 — the comparison is unestablished; stop"; exit 1; } + test "$BASE_ID" = "$APPROVED_ID" \ || { echo "$p differs at the base from its approved version"; exit 1; } done ``` @@ -620,10 +636,17 @@ handoff part of it is loose — and **routing on a guessed shape is how a source read them all and reconcile against what they actually contain: ```bash -git log --oneline "$BASE"..HEAD # can legitimately be empty; that is not an empty cycle -git diff --cached "$BASE" # the staged content — the FULL diff, not --stat -git diff # the unstaged content -git status --porcelain --untracked-files=all +# Guarded one by one, because this step's rule is that all four ARE read: unguarded, +# a failed read is indistinguishable from a source that holds nothing, and the next +# command's success carries the block to an exit 0 on an incomplete view. +git log --oneline "$BASE"..HEAD \ + || { echo "reading the commit range FAILED — the reconciliation is incomplete; stop"; exit 1; } +git diff --cached "$BASE" \ + || { echo "reading the staged content FAILED — the reconciliation is incomplete; stop"; exit 1; } +git diff \ + || { echo "reading the unstaged content FAILED — the reconciliation is incomplete; stop"; exit 1; } +git status --porcelain --untracked-files=all \ + || { echo "reading the path listing FAILED — the reconciliation is incomplete; stop"; exit 1; } ``` A ticked checkbox is confirmed by the change being *present in that content*, wherever it lives. @@ -1063,7 +1086,10 @@ git show "$BASE:plugins/dev-workflow/commands/workflow-init.md" > .context/loop- || { echo "cannot read workflow-init.md at $BASE"; exit 1; } C_SRC=.context/loop-rule-c.src; W_SRC=.context/loop-rule-w.src test -s "$C_SRC" && test -s "$W_SRC" || { echo "baseline source empty"; exit 1; } -printf 'base\t%s\n' "$BASE" > .context/loop-rule-baseline-diff.tmp # renamed at the end +# Renamed at the end. Every write to this artifact is guarded: an unguarded one lets +# the block build a partial baseline and still reach the rename. +printf 'base\t%s\n' "$BASE" > .context/loop-rule-baseline-diff.tmp \ + || { echo "cannot start the baseline artifact"; exit 1; } # Tab-separated start and end anchors, written as LITERAL text — the escaping # for sed happens later, on $se and $ee. Emitting '\*\*Severity:\*\*' here would @@ -1073,9 +1099,10 @@ printf 'base\t%s\n' "$BASE" > .context/loop-rule-baseline-diff.tmp # renamed a # contain colons, so a colon delimiter would split '**Severity:**' at the # wrong one. Every entry spans two DIFFERENT anchors — the single-line # squash-carry site is handled below, not in this loop. -: > .context/loop-rule-sites +: > .context/loop-rule-sites || { echo "cannot create the site list"; exit 1; } while IFS= read -r s && IFS= read -r e; do - printf '%s\t%s\n' "$s" "$e" >> .context/loop-rule-sites + printf '%s\t%s\n' "$s" "$e" >> .context/loop-rule-sites \ + || { echo "cannot append the site list"; exit 1; } done <<'SITES' Both gates are a LOOP Nothing here writes the floor knob @@ -1114,17 +1141,36 @@ while IFS=$(printf '\t') read -r s e; do # comparison that never completed. d=$(diff <(printf '%s\n' "$a") <(printf '%s\n' "$b")); st=$? test $st -le 1 || { echo "diff failed (status $st) at site: $s"; exit 1; } - printf 'site\t%s\t%s\n' "$s" "$e" >> .context/loop-rule-baseline-diff.tmp - printf '%s\n' "$d" >> .context/loop-rule-baseline-diff.tmp + printf 'site\t%s\t%s\n' "$s" "$e" >> .context/loop-rule-baseline-diff.tmp \ + || { echo "cannot record the site header for: $s"; exit 1; } + printf '%s\n' "$d" >> .context/loop-rule-baseline-diff.tmp \ + || { echo "cannot record the site diff for: $s"; exit 1; } done < .context/loop-rule-sites || exit 1 # The squash-carry sentence is ONE line and must not go through the loop. -d=$(diff <(grep -F 'On squash-merge, copy every evidence entry' "$C_SRC") \ - <(grep -F 'On squash-merge, copy every evidence entry' "$W_SRC")); st=$? +# It owes the SAME checks the loop applies, and accounting row 19 promises them for +# every site: unguarded, a missing anchor in both copies yields two empty extractions +# that diff equal, and the site record certifies a line sed never found. Duplicate +# hits would pass the same way. +SQ='On squash-merge, copy every evidence entry' +for f in "$C_SRC" "$W_SRC"; do + n=$(grep -cF "$SQ" "$f"); st=$? + test $st -le 1 || { echo "grep failed (status $st) at the squash-carry site in $f"; exit 1; } + test "$n" = 1 || { echo "squash-carry anchor is not unique in $f ($n hits)"; exit 1; } +done +ca=$(grep -F "$SQ" "$C_SRC") || { echo "cannot extract the squash-carry line from $C_SRC"; exit 1; } +cb=$(grep -F "$SQ" "$W_SRC") || { echo "cannot extract the squash-carry line from $W_SRC"; exit 1; } +test -n "$ca" && test -n "$cb" || { echo "empty extraction at the squash-carry site"; exit 1; } +d=$(diff <(printf '%s\n' "$ca") <(printf '%s\n' "$cb")); st=$? test $st -le 1 || { echo "diff failed (status $st) at the squash-carry site"; exit 1; } -printf 'site\t%s\t%s\n' 'On squash-merge' 'On squash-merge' >> .context/loop-rule-baseline-diff.tmp -printf '%s\n' "$d" >> .context/loop-rule-baseline-diff.tmp -mv .context/loop-rule-baseline-diff.tmp .context/loop-rule-baseline-diff.txt +printf 'site\t%s\t%s\n' 'On squash-merge' 'On squash-merge' >> .context/loop-rule-baseline-diff.tmp \ + || { echo "cannot record the squash-carry site header"; exit 1; } +printf '%s\n' "$d" >> .context/loop-rule-baseline-diff.tmp \ + || { echo "cannot record the squash-carry site diff"; exit 1; } +# Guarded, because the `cat` below would otherwise display a PREVIOUS run's .txt as +# though this sweep had produced it. +mv .context/loop-rule-baseline-diff.tmp .context/loop-rule-baseline-diff.txt \ + || { echo "rename FAILED — the baseline was NOT updated; stop"; exit 1; } cat .context/loop-rule-baseline-diff.txt # READ IT WHOLE ``` @@ -1229,8 +1275,14 @@ and these three exist only once this task installs them. - [ ] **Step 5: Check parity of the installed block** ```bash -diff <(sed -n '/^\*\*How a cycle ends/,/^\*\*What a loop absorbs/p' CLAUDE.md) \ - <(sed -n '/^\*\*How a cycle ends/,/^\*\*What a loop absorbs/p' plugins/dev-workflow/commands/workflow-init.md) +# Extract first and require both non-empty. This step's success signal is NO output, +# and two anchors that match nothing also produce no output — so an unchecked pair +# reads a parity check that never ran as parity confirmed. Task 0 step 3 applies the +# same rule to its sites. +a=$(sed -n '/^\*\*How a cycle ends/,/^\*\*What a loop absorbs/p' CLAUDE.md) +b=$(sed -n '/^\*\*How a cycle ends/,/^\*\*What a loop absorbs/p' plugins/dev-workflow/commands/workflow-init.md) +test -n "$a" && test -n "$b" || { echo "empty extraction — the installed §A block was not found in both copies"; exit 1; } +diff <(printf '%s\n' "$a") <(printf '%s\n' "$b") ``` Expected: no output. @@ -1729,8 +1781,12 @@ check to the Severity bullet alone lets the long answer paragraph differ between the pair and the stated parity check pass. ```bash -diff <(sed -n '/^- \*\*Severity:\*\*/,/^- \*\*Tool routing:/p' CLAUDE.md) \ - <(sed -n '/^- \*\*Severity:\*\*/,/^- \*\*Tool routing:/p' plugins/dev-workflow/commands/workflow-init.md) +# Same rule as the §A parity check: no output is this step's success signal, so two +# anchors that match nothing would read as byte-identical. +a=$(sed -n '/^- \*\*Severity:\*\*/,/^- \*\*Tool routing:/p' CLAUDE.md) +b=$(sed -n '/^- \*\*Severity:\*\*/,/^- \*\*Tool routing:/p' plugins/dev-workflow/commands/workflow-init.md) +test -n "$a" && test -n "$b" || { echo "empty extraction — the Severity passage was not found in both copies"; exit 1; } +diff <(printf '%s\n' "$a") <(printf '%s\n' "$b") ``` Expected: no output. This passage should now be byte-identical, `g4` having been removed and `g2`/`g3` having gone from both copies. @@ -2140,8 +2196,13 @@ unique to it. ```bash BASE=$(cat .context/loop-rule-base) test -n "$BASE" || { echo "BASE empty or unreadable — Task 0 did not run"; exit 1; } -git diff -U0 "$BASE" -- plugins/dev-workflow/hooks/codex-gate.sh \ - | grep -E '^@@' | sed -E 's/^@@ -([0-9]+).*/\1/' +# The diff is captured before it is filtered. In a pipeline the status is the LAST +# element's, and `sed` succeeds on empty input, so a failed `git diff` would print an +# empty list of changed lines — indistinguishable from a hook nobody edited. +# `pipefail` is not available here: these blocks run under sh and dash too. +DIFFOUT=$(git diff -U0 "$BASE" -- plugins/dev-workflow/hooks/codex-gate.sh) \ + || { echo "git diff FAILED — the changed-line list is unestablished; stop"; exit 1; } +printf '%s\n' "$DIFFOUT" | grep -E '^@@' | sed -E 's/^@@ -([0-9]+).*/\1/' grep -n 'note "' plugins/dev-workflow/hooks/codex-gate.sh ``` @@ -2511,8 +2572,13 @@ while IFS=$(printf '\t') read -r s e; do a=$(sed -n "/$se/,/$ee/p" CLAUDE.md) b=$(sed -n "/$se/,/$ee/p" plugins/dev-workflow/commands/workflow-init.md) test -n "$a" && test -n "$b" || { echo "empty extraction for '$s' — region not found"; exit 1; } - printf '%s\n' "$a" > .context/loop-rule-a.txt - printf '%s\n' "$b" > .context/loop-rule-b.txt + # Guarded: these two paths are REUSED on every iteration, so a failed write leaves + # the previous site's text standing and the diff below classifies bytes this + # iteration never extracted. + printf '%s\n' "$a" > .context/loop-rule-a.txt \ + || { echo "cannot write the C-side extract for: $s"; exit 1; } + printf '%s\n' "$b" > .context/loop-rule-b.txt \ + || { echo "cannot write the W-side extract for: $s"; exit 1; } diff .context/loop-rule-a.txt .context/loop-rule-b.txt; st=$? test $st -le 1 || { echo "diff failed (status $st) at site: $s"; exit 1; } done < .context/loop-rule-changed-sites @@ -2988,7 +3054,11 @@ git commit -m "WIP: fix " \ test "$(git rev-parse "HEAD^{tree}")" = "$ITREE" \ || { echo "the commit's tree is not the index that was pinned; no call may be issued"; exit 1; } -rm -f .context/loop-rule-reviewed-tip # the previous candidate's closing tip +# The previous candidate's closing tip. Guarded: where the path cannot be removed the +# stale tip survives, and 8b's condition 5 would compare `HEAD` against a value this +# run never produced. +rm -f .context/loop-rule-reviewed-tip \ + || { echo "removing the previous candidate's tip FAILED — it survives; no call may be issued"; exit 1; } # The head the NEXT call is issued against. Guarded, and no `cat`: the redirect can # fail while a following `cat` prints the STALE value and the block exits 0. git rev-parse HEAD > .context/loop-rule-reviewed-head \ @@ -3138,7 +3208,12 @@ for f in $FINAL; do test -n "$(git status --porcelain --untracked-files=all -- "$f")" \ || { echo "this pass's findings file is not dirty: $f"; exit 1; } done -test -z "$(git status --porcelain -z --untracked-files=all -- "$@")" \ +# Two steps: a FAILED `git status` prints nothing, and inline this reads as "nothing +# outside the findings files is dirty" — condition 2 satisfied by a comparison that +# never ran, and it is the only check between a foreign delta and the closing squash. +OUTSIDE=$(git status --porcelain -z --untracked-files=all -- "$@") \ + || { echo "reading the dirty set outside this pass's findings files FAILED"; exit 1; } +test -z "$OUTSIDE" \ || { echo "something outside this pass's findings files is dirty:"; git status --porcelain --untracked-files=all; exit 1; } # Condition 3, its pre-move check — the last precondition, so nothing has moved if @@ -3147,7 +3222,15 @@ test -z "$(git status --porcelain -z --untracked-files=all -- "$@")" \ test -s .context/loop-rule-closing-msg || { echo "closing message missing or empty"; exit 1; } # Condition 4, first bullet: pin the validated blobs before staging. -for f in $FINAL; do git hash-object "$f"; done > .context/loop-rule-final-blobs +# Guarded per iteration AND on the redirect. Only the last iteration's status survives +# a bare loop, and a failed redirect leaves an EARLIER attempt's pin standing — which, +# paired with the same failure below, makes the comparison compare two stale files and +# pass. The `>&2` matters: this compound's stdout IS the pin file. +for f in $FINAL; do + git hash-object "$f" \ + || { echo "pinning the validated blob for $f FAILED — nothing is staged" >&2; exit 1; } +done > .context/loop-rule-final-blobs \ + || { echo "writing the validated-blob pin FAILED — nothing is staged"; exit 1; } # shellcheck disable=SC2086 git add $FINAL \ || { echo "staging FAILED — run Failure; do NOT commit"; exit 1; } @@ -3155,23 +3238,45 @@ git commit -m "WIP: Gate-B findings files" || { echo "record commit FAILED — r test "$(git rev-parse HEAD^)" = "$HEADREV" || { echo "record commit's parent is not the reviewed head — run Failure"; exit 1; } git diff --quiet "$HEADREV" HEAD -- "$@" \ || { echo "record commit changed paths beyond this pass's findings files — run Failure"; exit 1; } +# The polarity is inverted here: exit 0 means "no difference", i.e. NOT carried. With +# `&&` an execution error (status 2 or above) takes the same path as "carried" and the +# close proceeds on a comparison that could not run. Status 1 — a real difference — is +# the legitimate result this check wants. for f in $FINAL; do - git diff --quiet "$HEADREV" HEAD -- "$f" \ - && { echo "record commit did not carry $f — run Failure"; exit 1; } + git diff --quiet "$HEADREV" HEAD -- "$f" + case $? in + 0) echo "record commit did not carry $f — run Failure"; exit 1 ;; + 1) : ;; + *) echo "comparing $f between the reviewed head and the record commit FAILED — run Failure"; exit 1 ;; + esac done # Condition 4, second bullet. -for f in $FINAL; do git rev-parse "HEAD:$f"; done > .context/loop-rule-committed-blobs +# Guarded the same way as the pin above, and for the same paired-stale reason. Note +# that on an unresolvable rev `git rev-parse` echoes its argument to stdout and exits +# 128, so without the guard this file can hold a junk line and the loop's status is gone. +for f in $FINAL; do + git rev-parse "HEAD:$f" \ + || { echo "reading the committed blob for $f FAILED — run Failure" >&2; exit 1; } +done > .context/loop-rule-committed-blobs \ + || { echo "writing the committed-blob list FAILED — run Failure"; exit 1; } diff .context/loop-rule-final-blobs .context/loop-rule-committed-blobs \ || { echo "a findings file was rewritten between validation and commit — run Failure"; exit 1; } # Condition 4, third bullet: re-run the findings-file structural check on the # COMMITTED content and re-establish this pass's eligibility. Reader check; the # expected result is condition 4's. -test -z "$(git status --porcelain)" || { echo "tree not clean after the record commit — run Failure"; exit 1; } +TREESTATE=$(git status --porcelain) \ + || { echo "reading the tree state after the record commit FAILED — run Failure"; exit 1; } +test -z "$TREESTATE" || { echo "tree not clean after the record commit — run Failure"; exit 1; } # Condition 4's tail: the closing tip. -git rev-parse HEAD > .context/loop-rule-reviewed-tip +# Guarded even though it ends the block: a failed redirect leaves whatever an earlier +# attempt wrote, and 8b's condition 5 compares HEAD against that. (A failed rev-parse +# after a successful truncate leaves the file empty, which 8b's `test -n "$TIP"` does +# catch — the redirect failure is the half it cannot.) +git rev-parse HEAD > .context/loop-rule-reviewed-tip \ + || { echo "recording the closing tip FAILED — run Failure; 8b must not run"; exit 1; } ``` **8b — reset and close. A separate invocation, carrying no `-m` option at all. Discharges condition @@ -3183,7 +3288,12 @@ test -n "$BASE" && test -n "$TIP" || { echo "BASE or TIP missing — 8a did not # Condition 5: the tip AND a clean tree, both read in THIS invocation. test "$(git rev-parse HEAD)" = "$TIP" || { echo "HEAD has moved since 8a — NOT resetting"; exit 1; } -test -z "$(git status --porcelain)" || { echo "tree not clean at 8b — NOT resetting; run Failure"; exit 1; } +# Two steps, and this is the most expensive instance: the next command is the +# `reset --soft`. Both branches keep both markers — "NOT resetting", because nothing +# has moved in this invocation, and "run Failure", because 8a's record commit has. +TREESTATE=$(git status --porcelain) \ + || { echo "reading the tree state at 8b FAILED — NOT resetting; run Failure"; exit 1; } +test -z "$TREESTATE" || { echo "tree not clean at 8b — NOT resetting; run Failure"; exit 1; } git reset --soft "$BASE" || { echo "reset --soft FAILED — run Failure; do NOT commit"; exit 1; } @@ -3194,10 +3304,12 @@ test -s .context/loop-rule-closing-msg || { echo "closing message missing or emp # oracle must not be the same mutable path the commit reads: a `commit-msg` hook # that rewrites git's copy AND this ignored source file would otherwise make the # post-close comparison pass on a body nobody validated. -# Remove any pin an earlier attempt left BEFORE writing this one, so a failed refresh -# leaves condition 6 with no oracle rather than a stale one its `test -s` would accept -# — and guard the copy, because the commit is the very next command and a failed pin -# must not fall through into it. +# Remove any pin an earlier attempt left, then copy under a guard — but the GUARD is +# what does the work. It stops before the commit, so nothing closes against a pin this +# invocation did not write. The `rm -f` promises no cleanup: where `.context/` is +# unwritable it fails too and the earlier pin survives, non-empty enough for condition +# 6's `test -s`. It is simply never reached, because the guard exits first and leaves +# both files on disk to be read. Observed in a disposable repository, not reasoned. rm -f .context/loop-rule-validated-msg cp .context/loop-rule-closing-msg .context/loop-rule-validated-msg \ || { echo "pinning the validated closing message FAILED — run Failure; do NOT commit"; exit 1; } @@ -3213,11 +3325,19 @@ a separate shell, so a variable assigned in 8b is empty in this one: ```bash BASE=$(cat .context/loop-rule-base) test -n "$BASE" || { echo "BASE empty or unreadable — cannot verify the close"; exit 1; } -case "$(git log -1 --pretty=%s)" in +# Read under a guard: a failed `git log` yields an empty subject, which matches no +# pattern, so the WIP arm never fires and the check is satisfied by a command that +# failed. No `test -n` is added — a genuinely empty subject read successfully passes +# today, and requiring one would be a new condition rather than a repair to this one. +SUBJECT=$(git log -1 --pretty=%s) \ + || { echo "reading the closing commit's subject FAILED — run Failure"; exit 1; } +case "$SUBJECT" in [Ww][Ii][Pp]:*) echo "closing commit still reads as a snapshot — run Failure"; exit 1 ;; esac test "$(git rev-parse HEAD^)" = "$BASE" || { echo "closing commit's parent is not \$BASE — run Failure"; exit 1; } -test -z "$(git status --porcelain)" || { echo "tree dirty after the close — run Failure"; exit 1; } +TREESTATE=$(git status --porcelain) \ + || { echo "reading the tree state after the close FAILED — run Failure"; exit 1; } +test -z "$TREESTATE" || { echo "tree dirty after the close — run Failure"; exit 1; } # `--pretty=format:%B` emits the stored message alone; the `%B` spelling appends a # trailing newline the source file has none of, which rejected a CORRECT close. # 8b's `--cleanup=verbatim` is the other half — without it git stores its own From f9bd330b793ea04e883dfcb07d3ca9eff031d6b8 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Wed, 16 Sep 2026 20:29:03 +0200 Subject: [PATCH 171/181] docs(plans): Resume checks ba15e83 against the base; Task 10 shows its diff Pass 45's two findings. Its Blocker is confirmed by direct observation, not accepted on assertion: ba15e83^ carries all three approved-input blobs identically to ba15e83, and Resume checked only that $BASE is an ancestor of HEAD. A base below the approved closure therefore passed every Resume check, and Gate B and the final soft reset would have pulled intervening history into the reviewed range and then squashed it into the closing commit. Accounting row 3 is split by entry, the way row 4 already is. The scope question - repair or new condition - went to Codex as a sparring partner and his answer was validated against the plan's own text: row 4 splits the approved-input comparison by entry for exactly this reason, Preparation being first-entry-only, so applying the same split to row 3 establishes an existing condition at re-entry rather than adding one. No human decision was owed. Task 10 step 4 now prints the captured diff before filtering it. The prose requires the executor to read every removed and added line for invariant 4, and the block printed only zero-context hunk headers. Pre-existing, confirmed against 4752a35: the pipeline consumed the body there too. Two claims of my own were wrong and are corrected. The tip-removal comment said a surviving stale tip is what 8b's condition 5 would read; it is not, because 8a rewrites the tip before 8b runs. The real risk is the next call being issued with two candidates' markers standing together, and the comment says that now. And the rev-parse rationale was false: without --verify, a failed git rev-parse : prints its unresolved ARGUMENT and exits 128, so two failed lookups compare equal only where the two revision expressions are equal. Observed: git rev-parse HEAD:no/such/file.md -> status 128, stdout "HEAD:no/such/file.md" git rev-parse ba15e83:no/such/file.md -> status 128, stdout "ba15e83:no/such/file.md" The separate guarded assignments remain the right implementation, because success must depend on both exit statuses; only the stated reason changes. Pass 44's finding 8 carried the same imprecision. Verified by execution: 16 fixtures under sh and bash, 15 under dash, all green, with the affected blocks extracted verbatim and a stub git failing one named subcommand at a time - including a base refused because ba15e83 is not its ancestor, and the success path. The dash skip is unchanged: Task 0 step 3 uses process substitution and is bash-only, pre-existing. No new syntax failure: the sh -n / dash -n failure set is identical to fc12f25's. The sparring exchange is kept at .context/codex-reviews/sparring-pass45.md. The pass-43 symlink Minor stays collected. --- .../gate-a-plan-om0bdd7udh-pass-45.md | 3 + .../gate-a-plan-om0bdd7udh-resume.md | 11 +- .context/codex-reviews/sparring-pass45.md | 104 ++++++++++++++++++ .../2026-09-14-loop-rule-consolidation.md | 27 +++-- 4 files changed, 136 insertions(+), 9 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-45.md create mode 100644 .context/codex-reviews/sparring-pass45.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-45.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-45.md new file mode 100644 index 0000000..492c557 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-45.md @@ -0,0 +1,3 @@ +BLOCKER | high | Resume — Validate the base; accounting row 3 | Resume proves that the recorded BASE is an ancestor of HEAD and that its three approved-input blobs equal ba15e83, but it never proves that ba15e83 is an ancestor of BASE. The plan assigns that ancestry condition only to first-entry Preparation even though the ignored base file selects Resume and can survive from another run. This repository makes the gap reachable: ba15e83's parent has the same three approved-input blobs as ba15e83, so that older commit satisfies every Resume base check. | Gate B and the final soft reset can use a base before the approved-spec closure, pulling intervening history into the reviewed range and then squashing it into the closing commit. | In Resume, guard git merge-base --is-ancestor ba15e83 "$BASE" before accepting the base, and split accounting row 3 by entry just as row 4 is split. +MAJOR | high | Task 10 step 4 | The guarded DIFFOUT capture is followed only by a pipeline that prints hunk headers; the block never emits or persists the removed and added lines. The prose then requires the executor to read every changed line and says to read the diff itself, but the captured diff body is no longer available after that fenced invocation. | An adjacent control-flow, counter, fingerprint, or routing edit can share an allowed zero-context hunk start and remain invisible to the required invariant-4 review. | After the guarded capture, print or persist the complete DIFFOUT and require the executor to inspect every added and removed line before accepting the seven note-only edits; keep the header list only as an aid. +END OF FINDINGS (2 total) diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md index 2b479f9..75eb248 100644 --- a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md @@ -13,7 +13,13 @@ Nothing depends on it; the pass files and the repo are authoritative where this ## Resume here -**Cycle OPEN and UNCLEAN. The plan stands at `4752a35` and that is the anchor** — later commits on +**READ THIS FIRST — the state below describes pass 44 and is superseded by the pass-45 rows in the +table.** The plan now stands at the thirteenth revision's commit; pass 45 found **2** findings (1 +Blocker, 1 Major), both repaired, and the next pass is **46**. Codex is being used as a **sparring +partner** on scope questions: put the question, require evidence, validate his answer, apply what +holds. That is how row 3's split was settled without spending a human decision. + +**Cycle OPEN and UNCLEAN. The plan stood at `4752a35` for pass 44 and that was the anchor** — later commits on this branch are records, not plan edits, so check `git log --oneline -1 -- docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md` rather than `HEAD`. Pass 44 reviewed `4752a35` and found **two Blockers and seven Majors**. `.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-44.md` holds them. @@ -304,6 +310,9 @@ worth checking before a pass rather than after. | 43 | 8361f0a | 3→**5** | 2→**2** | 1→**2** | yes | **UNCLEAN — reported to Daniel, no repair round opened.** **All three pass-42 repairs held; nothing was re-raised against the four repaired blocks.** Every Blocker and Major is the **same class at another location**: 8b's unguarded `cp` of the validated message, condition 6's unguarded `git log … > file` masked by the following `diff`, step 4's unguarded `git fetch`, and every remaining unguarded `git add` across Tasks 0–14 and 8a — **that last one is the plan-wide sweep the assignment excluded**. The only in-set finding is a MINOR: the new blob check accepts a **symlink** (blob, mode `120000`), so the indexed mode would have to be checked too. Three tells → mandatory stop | | — | — | — | — | — | — | **ELEVENTH BOUNDED REVISION**, Daniel's assignment after pass 43, scoped to its four Blocker/Major findings; the Minor stays collected. 8b removes any earlier pin **before** writing this one and copies under a guard, so a failed refresh leaves condition 6 **no** oracle rather than a stale one; 8b's commit stays terminal, condition 6 re-establishing it independently. Condition 6's extraction takes the same shape. Step 4 guards the fetch **and** the recording — `git rev-parse origin/main` resolves the stale ref happily after a failed fetch — and its `cat` is gone. **Seventeen staging locations guarded, enumerated not described.** A new Global Constraint states the rule: **guards everywhere, content comparison only where a consumer reads that commit** — three places have one, the per-task snapshots have none, and 8a needed no addition because condition 4 already pins its blobs. 18 checks × 3 shells, all green. **Two harness errors recorded**: a writable file in a read-only directory is *not* a failed write, and 8b resets before it pins. Commit `4752a35` | | 44 | 4752a35 | 5→**9** | 2→**2** | 2→**7** | yes | **UNCLEAN — reported to Daniel, no repair round opened.** **Nothing re-raised against the eleventh revision.** The class widened again — from *masked* status to **discarded** status: `test -z "$(git status …)"` swallows git's exit code and reads a failed command as a clean tree (Preparation and closure conditions 2, 5, 6 — BLOCKER); two `$(git rev-parse …)` in one equality test both resolve empty and compare equal (Preparation and Resume approved-input checks); an unchecked `grep` substitution pair that compares two empty extractions as equal (Task 0 step 3 — BLOCKER); a pipe with no `pipefail` (Task 10 step 4); unguarded scratch writes and a rename masked by `cat` (Task 0 step 3, Task 14 step 1); an unguarded `hash-object` pin loop (8a condition 4); a `git log` inside a `case` substitution (condition 6's subject check). **Not a plateau** — none regenerates from a repair. **Two are escalation triggers**: finding 2 changes how closure conditions spell their tests, finding 7 changes 8a's pre-move ordering. Three tells → mandatory stop | +| — | — | — | — | — | — | **TWELFTH BOUNDED REVISION**, Daniel's assignment after pass 44: all nine findings **plus one bounded sweep** of every executable block for the same mechanism — explicitly **not** a claim that the class is closed. All 54 blocks inspected (four parallel read-only passes plus a mechanical scan). Repaired: six discarded-substitution-status sites, each now a guarded assignment then the same test against the captured value; both two-substitution equality tests; Resume's four-source reconciliation; Task 10's pipeline; 8a's two loops (per iteration **and** on the redirect, messages to stderr because the compound's stdout is the artifact); 8a's inverted-polarity carry check, now a `case` on the status; and the unguarded writes, appends and renames in Task 0 step 3, Task 14 step 1, step 7 step three and 8a. **Justified non-changes:** sixteen terminal `git commit`s (the plan's own rule), `grep`/`diff` exit 1 as legitimate results, `$(cat …)` closed by `test -n`, the cleanup's `ls` glob. 14 fixtures × `sh`/`bash`, 13 × `dash`, all green, using a **stub `git` failing one named subcommand at a time**. Commit `fc12f25` | +| 45 | fc12f25 | 9→**2** | 2→**1** | 7→**1** | yes | **Best pass since 41.** Nothing re-raised against the twelfth revision. **BLOCKER: Resume never establishes that `ba15e83` is an ancestor of `$BASE`** — row 3 was Preparation-only, and **`ba15e83^` carries all three approved-input blobs identically** (verified directly), so a base below the approved closure passed every Resume check. **MAJOR: Task 10 step 4 never shows the diff body** the prose requires the executor to read — **pre-existing**, confirmed against `4752a35` | +| — | — | — | — | — | — | **THIRTEENTH BOUNDED REVISION**, both pass-45 findings, after putting the scope question to Codex as a sparring partner and validating his answer. **Splitting accounting row 3 by entry is a repair, not a new condition** — row 4 is already split the same way for the same reason, so no human decision was owed. Resume gains `merge-base --is-ancestor ba15e83 "$BASE"`; Task 10 prints the captured diff before filtering it. **Two of my own claims were wrong and are corrected**: the tip-removal comment (8a rewrites the tip before 8b, so a stale tip is not what condition 5 reads — the real risk is two candidates' markers standing together), and the rev-parse rationale — **without `--verify` a failed `git rev-parse :` prints its ARGUMENT, not nothing**, so two failures compare equal only where the revisions do. Observed, not reasoned. 16 fixtures × `sh`/`bash`, 15 × `dash`, all green | | 26 | 05ec1b2 | **5** | **1** | 2 | yes | first pass on the revised plan. Resume audited progress by `WIP:` commits and called any non-`WIP` commit a stale base — the handoff leaves two other shapes, one of them `HEAD` **at** the base with everything in the index | | 27 | f69db7d | 5→**6** | 1→**3** | 2→**1** | yes | the close required the commit's **tree to equal the tip's**, which target §I **parks**; `reset --soft` leaves the index untouched, so the prose beside it was wrong about git; Resume still inferred staleness from a commit subject where §A3 names that state as reachable | | 28 | f2e5dda | 6→**6** | 3→**2** | 1→**2** | yes | pass 27's stale-base fix reached one site of three; its reset paragraph stated the index behaviour correctly and repeated the false claim four lines later; nothing checked the closing commit's **parent**; the message was validated in the file, where `commit-msg` hooks rewrite git's copy after `-F` reads it | diff --git a/.context/codex-reviews/sparring-pass45.md b/.context/codex-reviews/sparring-pass45.md new file mode 100644 index 0000000..bb9b067 --- /dev/null +++ b/.context/codex-reviews/sparring-pass45.md @@ -0,0 +1,104 @@ +## ANCESTRY REPAIR + +Replace Resume's entire **Validate the base** shell block with this block. The added ancestry check is between resolving `$BASE` and accepting `$BASE` as an ancestor of `HEAD`; every existing message and the existing stop behavior remain unchanged. This is a Resume precondition check: it exits before Resume accepts the base and does not create a Failure handoff. If Resume was entered during an existing Failure handoff, that handoff remains in force and this block changes nothing. + +```sh +# The branch, FIRST and before anything reads or writes cycle state. `.context/` is +# ignored, so the base file survives a checkout: on another branch descended from +# $BASE every check below passes and the WIP commits, the soft reset and the closing +# commit all land on the wrong branch. A detached HEAD prints `HEAD` and is refused +# for the same reason. +test "$(git rev-parse --abbrev-ref HEAD)" = loop-rule-consolidation \ + || { echo "not on loop-rule-consolidation (on: $(git rev-parse --abbrev-ref HEAD)) — refusing to re-enter"; exit 1; } + +test -e .context/loop-rule-base || { echo "no base file — this is a first entry, run Preparation"; exit 1; } +test -s .context/loop-rule-base || { echo "base file is EMPTY — inspect, delete deliberately, record why, re-record from the true starting commit"; exit 1; } +BASE=$(cat .context/loop-rule-base) +test "${#BASE}" -eq 40 || { echo "base is not a full 40-character object name"; exit 1; } +case "$BASE" in *[!0-9a-f]*) echo "base is not an object name: $BASE"; exit 1 ;; esac +test "$(git rev-parse --verify "$BASE^{commit}")" = "$BASE" || { echo "base does not resolve to itself as a commit"; exit 1; } +git merge-base --is-ancestor ba15e83 "$BASE" || { echo "ba15e83 is NOT an ancestor of the base"; exit 1; } +git merge-base --is-ancestor "$BASE" HEAD || { echo "base is NOT an ancestor of HEAD"; exit 1; } +for f in target-text design condition-inventory; do + p="docs/superpowers/specs/2026-09-10-loop-rule-consolidation-$f.md" + # Resolve separately and check each status: `git rev-parse` can print its unresolved + # argument even when it fails, and equal revision expressions can therefore make two + # failed lookups compare equal. + BASE_ID=$(git rev-parse "$BASE:$p") \ + || { echo "$p does not resolve at \$BASE — the comparison is unestablished; stop"; exit 1; } + APPROVED_ID=$(git rev-parse "ba15e83:$p") \ + || { echo "$p does not resolve at ba15e83 — the comparison is unestablished; stop"; exit 1; } + test "$BASE_ID" = "$APPROVED_ID" \ + || { echo "$p differs at the base from its approved version"; exit 1; } +done +``` + +Replace accounting row 3 with exactly: + +```markdown +| 3 | `ba15e83` is an ancestor of the cycle's starting revision | **kept, and split by entry, like row 4**: Preparation checks it at `HEAD`, because a first entry has no recorded base; **Resume checks it at `$BASE`**, because that is the persisted starting revision Gate B diffs from and the close resets to. An earlier revision assigned it to Preparation alone, so a base older than `ba15e83` could pass Resume where the three approved-input blobs happened to match | +``` + +## TASK 10 + +This is a pre-existing gap. At `4752a35`, before the twelfth revision, the executor saw only the zero-context hunk headers produced by `grep` and `sed`, followed by the `note` locations; the complete removed and added lines were not printed or persisted by the fenced block. The prose told the executor to read them, but supplied no body to read after the block ran. I verified that with: + +```sh +git show 4752a35:docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md | nl -ba | sed -n '2120,2185p' +``` + +Use this minimal replacement block. It keeps the guarded read, prints the complete diff for the required inspection, then retains the hunk-start list as an aid: + +```sh +BASE=$(cat .context/loop-rule-base) +test -n "$BASE" || { echo "BASE empty or unreadable — Task 0 did not run"; exit 1; } +# Capture before filtering so a failed read cannot become an empty hunk list. +DIFFOUT=$(git diff -U0 "$BASE" -- plugins/dev-workflow/hooks/codex-gate.sh) \ + || { echo "git diff FAILED — the changed-line list is unestablished; stop"; exit 1; } +printf '%s\n' "$DIFFOUT" # READ EVERY REMOVED AND ADDED LINE +printf '%s\n' "$DIFFOUT" | grep -E '^@@' | sed -E 's/^@@ -([0-9]+).*/\1/' +grep -n 'note "' plugins/dev-workflow/hooks/codex-gate.sh +``` + +## THE OVERSTATEMENT + +At `docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:3057`–`3059` in `fc12f25`, the offending sentence is: + +> The previous candidate's closing tip. Guarded: where the path cannot be removed the stale tip survives, and 8b's condition 5 would compare `HEAD` against a value this run never produced. + +The last clause is false. A successful candidate close must run 8a before 8b, and 8a rewrites `.context/loop-rule-reviewed-tip` with the new record commit after all its preceding checks. The stale value left by this failed removal therefore cannot be the value condition 5 reads in a correctly followed closing sequence. The real risk is issuing the next review call with the old candidate's tip marker still present beside the new candidate's reviewed-head marker. + +Replace the comment with: + +```sh +# The previous candidate's closing tip. Guarded: without the guard, a failed removal +# leaves the stale tip beside the next candidate's reviewed head, so the next call is +# issued with two candidates' markers present. +rm -f .context/loop-rule-reviewed-tip \ + || { echo "removing the previous candidate's tip FAILED — it survives; no call may be issued"; exit 1; } +``` + +## CONFIRMATIONS + +The `case $?` rewrite of 8a's per-file carry check is correct: status 0 means the file was not carried and must fail, status 1 is the required difference, and status 2 or above is a comparison failure that must not be accepted as a difference. + +Both `>&2` uses inside the `for ... done > file` loops are correct: the loop's stdout is the blob-list artifact, so each failure diagnostic must bypass that redirection or it would be written into the artifact instead of shown to the executor. + +## ANYTHING ELSE WRONG + +The new explanations at `docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:354`–`355` and `:588`–`589` are factually wrong when they say two failed `git rev-parse` substitutions produce empty strings. In this repository's Git, a failed `git rev-parse ":"` writes the unresolved argument to stdout and exits 128. Two failures compare equal only when the two revision expressions are themselves equal, such as `HEAD = ba15e83` or `$BASE = ba15e83`; when the revisions differ, the emitted strings differ. The separate guarded assignments are still the right implementation because success must depend on both exit statuses. + +Use these comments: + +```sh +# Resolve separately and check each status: `git rev-parse` can print its unresolved +# argument even when it fails, and equal revision expressions can therefore make two +# failed lookups compare equal. +``` + +and in Resume: + +```sh +# Resolve separately, for the reason Preparation's copy gives: a failed lookup can +# still print its unresolved argument, so equality alone does not establish success. +``` diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index 8342b92..4679363 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -243,7 +243,7 @@ independent reader did. |---|---|---| | 1 | Branch is `loop-rule-consolidation` | **kept**, and **split by entry, like row 4**: Preparation checks it at a first entry and **Resume checks it again**, first, before anything reads or writes cycle state. An earlier revision assigned it to Preparation alone, which is first-entry-only — and `.context/` is ignored, so the base file survives a checkout and another branch descended from `$BASE` passed every re-entry check (pass 38, Blocker) | | 2 | Tree clean before the base is recorded | **kept**, same shell — preparation | -| 3 | `ba15e83` is an ancestor of `HEAD` | **kept**, same shell — preparation | +| 3 | `ba15e83` is an ancestor of the cycle's starting revision | **kept, and split by entry, like row 4**: Preparation checks it at `HEAD`, because a first entry has no recorded base; **Resume checks it at `$BASE`**, because that is the persisted starting revision Gate B diffs from and the close resets to. An earlier revision assigned it to Preparation alone, which is first-entry-only — so a base **below** `ba15e83` passed Resume wherever the three approved-input blobs happened to match, and in this repository `ba15e83^` carries all three identically (pass 45, Blocker) | | 4 | The three approved inputs' blobs equal their `ba15e83` versions | **kept, and split by entry**: Preparation compares them **at `HEAD`**, because a first entry has no recorded base; **Resume compares them at `$BASE`**, because a Gate-B fix may legitimately have changed `HEAD`'s copy in a `WIP:` snapshot. An earlier table row claimed Preparation did the base comparison, which it never could | | 5 | Never overwrite an existing base file | **kept** as an obligation; the shell branch becomes one line of the resume procedure | | 6 | A pre-existing base is an ancestor of `HEAD` (`merge-base --is-ancestor`) | **kept**, same shell — resume | @@ -351,8 +351,11 @@ test ! -e .context/loop-rule-base || { echo "a base is already recorded — this git merge-base --is-ancestor ba15e83 HEAD || { echo "ba15e83 not in this history"; exit 1; } for f in target-text design condition-inventory; do p="docs/superpowers/specs/2026-09-10-loop-rule-consolidation-$f.md" - # Resolved separately: with both substitutions inline, a path that resolves at - # NEITHER revision gives "" = "" and the comparison PASSES on two failed lookups. + # Resolve separately and check each status. `git rev-parse` without `--verify` + # prints its UNRESOLVED argument and exits 128, so two failed lookups compare equal + # exactly when the two revision expressions are equal — and either way the result + # says nothing, because success must depend on both exit statuses. (Observed: + # `git rev-parse HEAD:missing` prints `HEAD:missing`, status 128.) HEAD_ID=$(git rev-parse "HEAD:$p") \ || { echo "$p does not resolve at HEAD — the comparison is unestablished; stop"; exit 1; } APPROVED_ID=$(git rev-parse "ba15e83:$p") \ @@ -582,11 +585,17 @@ BASE=$(cat .context/loop-rule-base) test "${#BASE}" -eq 40 || { echo "base is not a full 40-character object name"; exit 1; } case "$BASE" in *[!0-9a-f]*) echo "base is not an object name: $BASE"; exit 1 ;; esac test "$(git rev-parse --verify "$BASE^{commit}")" = "$BASE" || { echo "base does not resolve to itself as a commit"; exit 1; } +# Row 3 at `$BASE`, the way row 4 is already split by entry. Preparation checks this +# against `HEAD` and is first-entry-only, so nothing established it for a base that +# arrived with the ignored base file — and a commit BELOW `ba15e83` can carry the same +# three approved-input blobs and pass every other check here. Observed in this +# repository: `ba15e83^` has identical blobs for all three. +git merge-base --is-ancestor ba15e83 "$BASE" || { echo "ba15e83 is NOT an ancestor of the base"; exit 1; } git merge-base --is-ancestor "$BASE" HEAD || { echo "base is NOT an ancestor of HEAD"; exit 1; } for f in target-text design condition-inventory; do p="docs/superpowers/specs/2026-09-10-loop-rule-consolidation-$f.md" - # Resolved separately, for the reason Preparation's copy gives: two inline - # substitutions that both fail compare equal and the check passes on nothing. + # Resolve separately, for the reason Preparation's copy gives: a failed lookup can + # still print its unresolved argument, so equality alone does not establish success. BASE_ID=$(git rev-parse "$BASE:$p") \ || { echo "$p does not resolve at \$BASE — the comparison is unestablished; stop"; exit 1; } APPROVED_ID=$(git rev-parse "ba15e83:$p") \ @@ -2202,6 +2211,7 @@ test -n "$BASE" || { echo "BASE empty or unreadable — Task 0 did not run"; exi # `pipefail` is not available here: these blocks run under sh and dash too. DIFFOUT=$(git diff -U0 "$BASE" -- plugins/dev-workflow/hooks/codex-gate.sh) \ || { echo "git diff FAILED — the changed-line list is unestablished; stop"; exit 1; } +printf '%s\n' "$DIFFOUT" # READ EVERY REMOVED AND ADDED LINE printf '%s\n' "$DIFFOUT" | grep -E '^@@' | sed -E 's/^@@ -([0-9]+).*/\1/' grep -n 'note "' plugins/dev-workflow/hooks/codex-gate.sh ``` @@ -3054,9 +3064,10 @@ git commit -m "WIP: fix " \ test "$(git rev-parse "HEAD^{tree}")" = "$ITREE" \ || { echo "the commit's tree is not the index that was pinned; no call may be issued"; exit 1; } -# The previous candidate's closing tip. Guarded: where the path cannot be removed the -# stale tip survives, and 8b's condition 5 would compare `HEAD` against a value this -# run never produced. +# The previous candidate's closing tip. Guarded: without the guard a failed removal +# leaves the stale tip standing beside the next candidate's reviewed head, so the next +# call is issued with two candidates' markers present. (It is NOT the value condition 5 +# reads in a correct sequence — 8a rewrites the tip before 8b runs.) rm -f .context/loop-rule-reviewed-tip \ || { echo "removing the previous candidate's tip FAILED — it survives; no call may be issued"; exit 1; } # The head the NEXT call is issued against. Guarded, and no `cat`: the redirect can From 221819c63f9313705758ede678a957c425fec594 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Wed, 16 Sep 2026 20:48:29 +0200 Subject: [PATCH 172/181] docs(plans): step two reads the committed findings, not the worktree copies MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Pass 46's single Blocker, taken in part and declined in part, with the reason for each. The scope question - repair to an existing condition, or a new obligation - went to two independent readings and they disagreed, both quoting the same target §A sentence: "Every finding-derived predicate reads the validated findings file or files of the logical pass ... after each file has been validated separately." One reading took that as a one-time ritual before the predicate reads, which would make this a new obligation. The other took it as an input invariant on what the predicate actually reads, which makes it a repair. The second reading was taken, for two reasons. Its supporting evidence was executed rather than argued: a pre-commit hook that rewrites only the worktree copy leaves every step-one check satisfied - staged index, pin and commit tree all agree - while the file a reader would open no longer holds what the pass was accepted on. And the first reading's supporting claim, that findings files are gitignored so nothing finding-derived is ever committed, is false in this repository: .gitignore exempts .context/codex-reviews/ and this cycle commits its findings files. Demonstrated under sh in a disposable repository, with a hook rewriting only the worktree slot: step one exits 0, git show HEAD: returns the validated finding line, and the worktree holds "REWRITTEN IN WORKTREE ONLY". The harness repeated that case under dash and bash too, but its own reset logic between iterations was broken and those two runs are not evidence. The observation is about git rather than about the shell. Taken: step two reads git show "HEAD:$SPEC" and "HEAD:$QUAL" rather than the worktree copies. The commit is immutable once step one's tree comparison has passed, so that is the one source a later rewrite cannot reach. Declined, with the reason: re-running the findings-file structural and branch-eligibility checks on the committed content. That is Close condition 4's third bullet, a duty the approved text places on the closing commit, and step one's record commit is not the closing commit. Object identity is what step one can establish. Named as a residual rather than repaired: nothing here establishes that the committed bytes are the ones pass acceptance validated. Step one pins what is on disk when step one runs, so an edit between accepting the result and staging it is pinned as staged and passes every check - which this plan already says of its pin. Closing that would mean carrying the accepted blob ids out of the acceptance step in a file, and neither the approved spec nor CLAUDE.md §5 carries that duty. The prose says so in those terms rather than claiming a tighter guarantee than the mechanism gives. The sparring exchange is kept at .context/codex-reviews/sparring-pass46.md. --- .../gate-a-plan-om0bdd7udh-pass-46.md | 2 + .../gate-a-plan-om0bdd7udh-resume.md | 2 + .context/codex-reviews/sparring-pass46.md | 123 ++++++++++++++++++ .../2026-09-14-loop-rule-consolidation.md | 19 +++ 4 files changed, 146 insertions(+) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-46.md create mode 100644 .context/codex-reviews/sparring-pass46.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-46.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-46.md new file mode 100644 index 0000000..f920f0b --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-46.md @@ -0,0 +1,2 @@ +BLOCKER | high | Task 15 step 7, step one | The record commit pins the index only after staging and explicitly does not re-run the findings-file validation or bind the committed blobs to the files supplied by the Gate-B result; step two then reads the pass after that commit. | An intervening edit before staging, or a commit hook that rewrites only the worktree copies, can make the ordering authorize a repair, continue, suspension, or park from findings content that was never validated, contrary to target §A's requirement that every finding-derived predicate read separately validated findings files. | Validate and hash both slot files immediately after the result is accepted, compare those hashes with the committed HEAD blobs after the record commit, re-run the structural and branch-eligibility checks on those committed blobs, and require step two to read that committed content. +END OF FINDINGS (1 total) diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md index 75eb248..615320c 100644 --- a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md @@ -313,6 +313,8 @@ worth checking before a pass rather than after. | — | — | — | — | — | — | **TWELFTH BOUNDED REVISION**, Daniel's assignment after pass 44: all nine findings **plus one bounded sweep** of every executable block for the same mechanism — explicitly **not** a claim that the class is closed. All 54 blocks inspected (four parallel read-only passes plus a mechanical scan). Repaired: six discarded-substitution-status sites, each now a guarded assignment then the same test against the captured value; both two-substitution equality tests; Resume's four-source reconciliation; Task 10's pipeline; 8a's two loops (per iteration **and** on the redirect, messages to stderr because the compound's stdout is the artifact); 8a's inverted-polarity carry check, now a `case` on the status; and the unguarded writes, appends and renames in Task 0 step 3, Task 14 step 1, step 7 step three and 8a. **Justified non-changes:** sixteen terminal `git commit`s (the plan's own rule), `grep`/`diff` exit 1 as legitimate results, `$(cat …)` closed by `test -n`, the cleanup's `ls` glob. 14 fixtures × `sh`/`bash`, 13 × `dash`, all green, using a **stub `git` failing one named subcommand at a time**. Commit `fc12f25` | | 45 | fc12f25 | 9→**2** | 2→**1** | 7→**1** | yes | **Best pass since 41.** Nothing re-raised against the twelfth revision. **BLOCKER: Resume never establishes that `ba15e83` is an ancestor of `$BASE`** — row 3 was Preparation-only, and **`ba15e83^` carries all three approved-input blobs identically** (verified directly), so a base below the approved closure passed every Resume check. **MAJOR: Task 10 step 4 never shows the diff body** the prose requires the executor to read — **pre-existing**, confirmed against `4752a35` | | — | — | — | — | — | — | **THIRTEENTH BOUNDED REVISION**, both pass-45 findings, after putting the scope question to Codex as a sparring partner and validating his answer. **Splitting accounting row 3 by entry is a repair, not a new condition** — row 4 is already split the same way for the same reason, so no human decision was owed. Resume gains `merge-base --is-ancestor ba15e83 "$BASE"`; Task 10 prints the captured diff before filtering it. **Two of my own claims were wrong and are corrected**: the tip-removal comment (8a rewrites the tip before 8b, so a stale tip is not what condition 5 reads — the real risk is two candidates' markers standing together), and the rev-parse rationale — **without `--verify` a failed `git rev-parse :` prints its ARGUMENT, not nothing**, so two failures compare equal only where the revisions do. Observed, not reasoned. 16 fixtures × `sh`/`bash`, 15 × `dash`, all green | +| 46 | f9bd330 | 2→**1** | 1→**1** | 1→**0** | yes | **BLOCKER: step one commits the pass records without binding them to what was validated, and step two then read the WORKTREE copies** — so a hook rewriting only the worktree leaves every step-one check satisfied while the file a reader opens no longer holds what the pass was accepted on. It points at the limit the twelfth revision **disclosed in its own prose**, which is why it was not created by the pass-45 repair | +| — | — | — | — | — | — | **FOURTEENTH BOUNDED REVISION.** **The scope question was put to two independent readings and they DISAGREED** — both quoting the same §A sentence (`target-text.md:85-89`). A subagent read it as "validate once, before the predicate reads" → new obligation; Codex read it as an **input invariant** on what the predicate actually reads → repair, and ran fixtures showing the worktree-rewrite hole is real. **Codex's reading was taken**, because the subagent's supporting claim — that findings files are gitignored so nothing finding-derived is committed — is **false in this repo**: `.gitignore` exempts `.context/codex-reviews/` and this cycle commits its findings files. **Repair taken:** step two reads `git show "HEAD:$SPEC"` / `HEAD:$QUAL`. **Declined with reason:** re-running the structural and eligibility checks on committed content — that is Close condition 4's *closing*-commit duty and the approved text does not place it on a record commit. **Residual named, not repaired:** nothing establishes that the committed bytes are the ones *pass acceptance* validated; an edit between acceptance and staging is pinned as staged. Closing that needs the accepted blob ids carried out of the acceptance step in a file, a duty **no approved text carries**. Demonstrated under `sh`: with a worktree-only rewriting hook, step one exits 0, `HEAD` holds the validated line and the worktree holds `REWRITTEN IN WORKTREE ONLY` | | 26 | 05ec1b2 | **5** | **1** | 2 | yes | first pass on the revised plan. Resume audited progress by `WIP:` commits and called any non-`WIP` commit a stale base — the handoff leaves two other shapes, one of them `HEAD` **at** the base with everything in the index | | 27 | f69db7d | 5→**6** | 1→**3** | 2→**1** | yes | the close required the commit's **tree to equal the tip's**, which target §I **parks**; `reset --soft` leaves the index untouched, so the prose beside it was wrong about git; Resume still inferred staleness from a commit subject where §A3 names that state as reachable | | 28 | f2e5dda | 6→**6** | 3→**2** | 1→**2** | yes | pass 27's stale-base fix reached one site of three; its reset paragraph stated the index behaviour correctly and repeated the false claim four lines later; nothing checked the closing commit's **parent**; the message was validated in the file, where `commit-msg` hooks rewrite git's copy after `-F` reads it | diff --git a/.context/codex-reviews/sparring-pass46.md b/.context/codex-reviews/sparring-pass46.md new file mode 100644 index 0000000..8b54249 --- /dev/null +++ b/.context/codex-reviews/sparring-pass46.md @@ -0,0 +1,123 @@ +# Sparring recommendation — Gate-A plan pass 46 + +## 1. PREMISE CHECK + +**Observed.** The asserted requirement exists, almost verbatim. Target §A says: + +> “Every finding-derived predicate reads the validated findings file **or files** of the logical pass as **the concatenation of their finding lines after each file has been validated separately**.” + +That is `docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md:85-89`. The same paragraph says closure also reads every closure condition the cycle has (`...target-text.md:91-93`). Target §A separately says Gate B has no gate-specific content condition (`...target-text.md:68-70`), but that does not remove A1's gate-general validated-file read rule. + +`CLAUDE.md` defines what that validation means. A pass is accepted only when each required file is readable, has the exact terminator, contains exactly the declared number of finding lines and nothing else, and, for full Gate B, both branch files pass those checks (`CLAUDE.md:478-486`). It then defines structural failure and severity parsing, including wrong field count, empty severity, malformed lines, bad terminator, and count mismatch (`CLAUDE.md:488-497`). + +**Conclusion.** The reviewer did **not** construct the premise. Every predicate derived from the findings must read files that were separately validated. What the two texts do **not** require is the reviewer's entire proposed mechanism: they do not say every intermediate record commit must duplicate Close condition 4's post-commit structural and eligibility procedure. They state the input invariant; the plan may establish it by binding the later read to the already validated blobs. + +## 2. SCOPE VERDICT + +**Recommendation: treat the core finding as a repair to an existing condition. Do not adopt the reviewer's full prescribed remedy.** + +**Observed.** The plan itself admits the gap. It says the current `git write-tree` pin is “not the content the pass acceptance validated,” that a rewrite before staging is pinned and passes, and that step one has no structural re-run (`docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:2949-2953`). Step two then makes finding-derived decisions: source-block repair, suspension and membership handling, park, close, or continue (`...loop-rule-consolidation.md:3006-3023`). The current shell only proves that the commit tree equals the index written after staging (`...loop-rule-consolidation.md:2963-2999`). It does not prove that step two's input is the accepted snapshot. + +**Reasoning.** Once step two derives a route from the findings, §A's existing validated-input rule applies. Disclosure accurately describes noncompliance; it does not turn that noncompliance into an allowed limit. Changing the acceptance operation so that it pins the accepted blob IDs, checking those IDs against `HEAD` after the record commit, and making step two read `HEAD:` changes **how the existing condition is established**. It adds no success predicate or closure condition. + +The asymmetry has two parts: + +- The lack of **any binding** in step one is an omission. The “different consumers” explanation cannot justify it because step two is itself a consumer of finding-derived predicates. +- The lack of a **second structural and eligibility re-run** can remain deliberate. Close condition 4 expressly requires three cumulative facts for the closing commit: path and parent, identity with the validated blobs, and a new structural/eligibility check on committed content (`...loop-rule-consolidation.md:403-415`). Step one's record commit is not the closing commit. If object-ID equality establishes that `HEAD:` is exactly the already validated input and step two reads that immutable committed content, §A and `CLAUDE.md` are satisfied without copying Close condition 4's third bullet. + +## 3. IF REPAIR + +Make the two `*_VALIDATED` assignments below the final act of the existing §5 acceptance operation. In the sentence introducing step two, state: **“Step two reads the two committed blobs with `git show "HEAD:$SPEC"` and `git show "HEAD:$QUAL"`; it does not read either worktree path.”** That sentence is necessary: a worktree-only hook rewrite is harmless only if the consumer actually reads `HEAD`. + +Replace the current step-one shell with this block. It keeps the plan precondition, guarded `git add`, whole-index `git write-tree` pin, both blob-type checks, tree comparison, all existing failure messages, one shell block, and no `set -e` or `pipefail`: + +```bash +NONCE=; P= +SPEC=".context/codex-reviews/gate-b-spec-$NONCE-pass-$P.md" +QUAL=".context/codex-reviews/gate-b-quality-$NONCE-pass-$P.md" +PLAN=docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md + +# These assignments are the final act of accepting the result under CLAUDE.md §5: +# they pin the exact bytes that were separately validated, before anything is staged. +SPEC_VALIDATED=$(git hash-object "$SPEC") \ + || { echo "pinning the validated blob for $SPEC FAILED — the pass is not accepted; step two does not run"; exit 1; } +QUAL_VALIDATED=$(git hash-object "$QUAL") \ + || { echo "pinning the validated blob for $QUAL FAILED — the pass is not accepted; step two does not run"; exit 1; } + +# The plan is not staged here and must be unchanged: HEAD against the index (what the +# commit takes), then the index against the worktree. A difference and a failed +# comparison both stop, and neither says what caused it — report the state. +git diff --quiet --cached HEAD -- "$PLAN" \ + || { echo "plan is not identical between HEAD and the index, or the comparison failed — stop and report the state"; exit 1; } +git diff --quiet -- "$PLAN" \ + || { echo "plan is not identical between the index and the worktree, or the comparison failed — stop and report the state"; exit 1; } + +git add "$SPEC" "$QUAL" \ + || { echo "staging FAILED — the pass is not recorded; step two does not run"; exit 1; } + +# Pinned after staging, before the commit: the WHOLE index as staged at that moment, +# unrelated paths included. +ITREE=$(git write-tree) \ + || { echo "cannot pin the staged index; step two does not run"; exit 1; } + +# A successful add does not mean these paths are present — it stages a removal too. +# Require a blob at each: an existence test would accept a directory at the path. +test "$(git cat-file -t "$ITREE:$SPEC" 2>/dev/null)" = blob \ + || { echo "$SPEC is not a blob in the staged index; the pass is not recorded, step two does not run"; exit 1; } +test "$(git cat-file -t "$ITREE:$QUAL" 2>/dev/null)" = blob \ + || { echo "$QUAL is not a blob in the staged index; the pass is not recorded, step two does not run"; exit 1; } + +git commit -m "WIP: pass $P records" \ + || { echo "records commit FAILED — stop here; step two does not run"; exit 1; } + +# Any divergence between the pinned index and the resulting commit tree stops here, +# whatever produced it. +test "$(git rev-parse "HEAD^{tree}")" = "$ITREE" \ + || { echo "the commit's tree is not the index that was pinned; step two does not run"; exit 1; } + +# Bind what step two will read to the exact blobs accepted above. Guard the reads: +# failed rev-parse commands must not leave two junk or empty values to compare equal. +SPEC_COMMITTED=$(git rev-parse "HEAD:$SPEC") \ + || { echo "reading the committed blob for $SPEC FAILED — step two does not run"; exit 1; } +QUAL_COMMITTED=$(git rev-parse "HEAD:$QUAL") \ + || { echo "reading the committed blob for $QUAL FAILED — step two does not run"; exit 1; } +test "$SPEC_COMMITTED" = "$SPEC_VALIDATED" \ + || { echo "$SPEC is not the validated findings blob; step two does not run"; exit 1; } +test "$QUAL_COMMITTED" = "$QUAL_VALIDATED" \ + || { echo "$QUAL is not the validated findings blob; step two does not run"; exit 1; } +``` + +This is deliberately smaller than Close condition 4: it adds identity binding and fixes the consumer. It does not add that closure condition's independent post-commit structural/eligibility predicate. + +**Verification performed outside the repository.** I substituted `NONCE='test'; P=46` and ran `sh -n`, `dash -n`, `bash -n`, and `shellcheck --shell=sh` on `/tmp/sparring-pass46-step7.sh`; all exited 0. In disposable `/tmp` repositories: + +- a pre-commit hook that rewrote only the worktree slot let the block exit 0, while `HEAD` retained `END OF FINDINGS (1 total)` and the worktree contained `rewritten only in worktree`; this confirms why step two must read `HEAD`; +- an edit injected after the validated hashes but before `git add` exited 1 with `.context/codex-reviews/gate-b-spec-test-pass-46.md is not the validated findings blob; step two does not run`; +- a pre-commit hook that rewrote and staged the slot exited 1 with `the commit's tree is not the index that was pinned; step two does not run`. + +These executions test the mechanics. **Reasoning:** object-ID equality means the committed bytes are the already accepted bytes; therefore re-running the same structural predicate cannot change the result unless the acceptance-to-hash binding was not actually performed as specified. + +## 4. IF NEW OBLIGATION + +This is not my verdict. A human choosing the reviewer's stronger remedy would be agreeing that every nonclosing pass-record commit must independently re-run the full findings-file structure and branch-eligibility checks on `HEAD`, even when blob identity already proves that `HEAD` contains the accepted snapshot; neither target §A nor `CLAUDE.md` currently states that extra procedural duty. The cheapest defensible alternatives are the identity binding plus committed-content read in §3, or “disclose the limit and do nothing”; the latter is the current state and knowingly leaves the existing §A condition unenforced when the bytes change. + +## 5. COST CHECK + +**Observed history.** The recent record shows a high repair-regression rate followed by three stabilizing rounds: + +- Pass 40 found that the preceding repair's directory pathspec swept another cycle's files (`.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md:304`). +- Pass 41 called its only finding “this round's own repair” (`...resume.md:306`). +- Pass 42 says two of three findings were introduced by the preceding narrowing (`...resume.md:308`). +- Pass 43 says the three pass-42 findings were repaired, but its Minor found that the newly added blob check accepted a symlink; the check's addition is recorded immediately before it (`...resume.md:309-310`). +- Passes 44 and 45 say nothing was re-raised against the preceding revisions (`...resume.md:312-314`). +- The thirteenth revision preceding pass 46 changed Resume ancestry and Task 10 diff display (`...resume.md:315`). Pass 46 instead points to the older, expressly disclosed step-one limit (`.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-46.md:1` and plan `:2949-2953`), so it was not created by the pass-45 repair. + +That is four repair-created findings in the first four of the last seven repair-to-review transitions, followed by three transitions with none. The headline trend is also 9 → 2 → 1 at passes 44-46 (`...resume.md:312`, `...resume.md:314`, and `...pass-46.md:1-2`). + +**Reasoned estimate: about one chance in three, plausibly 25-40%, that the next pass raises a finding created by this repair.** The raw recent rate, four of seven, argues against a low-single-digit estimate; the last three clean descendant checks, the narrow object-ID addition, and the `/tmp` shell and hook checks justify discounting the raw 57%. The remaining likely attack surface is semantic rather than shell portability: whether “the final act of acceptance” binds the human validation tightly enough to the two hashes, and whether step two is stated unambiguously enough to read `HEAD` rather than the worktree. + +## Evidence and repository state + +- `git rev-parse --verify HEAD` returned `f9bd330b793ea04e883dfcb07d3ca9eff031d6b8`. +- `git diff --exit-code f9bd330 -- CLAUDE.md docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md .context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md` exited 0, so the cited tracked texts are exactly the versions at the named commit. +- Before this requested report was written, `git status --short` showed only the supplied pass-46 findings file as untracked. No tracked file was edited, staged, or committed. diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index 4679363..8ce525b 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -3003,6 +3003,25 @@ test "$(git rev-parse "HEAD^{tree}")" = "$ITREE" \ staged before this step runs is carried in, and no check in this plan catches it** — the close's dirty-set and changed-path checks read 8a's commit, not these. Nothing here fixes that. +**Step two reads the COMMITTED findings, not the worktree copies** — `git show "HEAD:$SPEC"` and +`git show "HEAD:$QUAL"`, with `SPEC` and `QUAL` as step one spelled them. Target §A requires every +finding-derived predicate to read the **validated** findings file, and the worktree copy is the one +thing that can still change after validation: a `commit-msg` or `pre-commit` hook that rewrites only +the worktree leaves step one's checks entirely satisfied — the staged index, the pin and the commit +tree all agree — while the file a reader would open no longer holds what the pass was accepted on. +Reading `HEAD` closes that, because the commit is immutable once step one's tree comparison has +passed. + +**What this does not establish, disclosed rather than guarded.** It does not prove those committed +bytes are the ones **pass acceptance** validated. Step one pins what is on disk when step one runs; +an edit between accepting the result and staging it is pinned as staged and passes every check here, +which this plan already says of its pin. Closing *that* gap would mean carrying the accepted blob +ids out of the acceptance step in a file, and **no such duty exists in the approved spec or in §5** — +it is named here as the residual it is, not repaired. **The reviewer's further demand — re-running +the structural and branch-eligibility checks on the committed content — is declined for the same +reason**: object identity is what step one can establish, and Close condition 4's third bullet is a +*closing*-commit duty the approved text does not place on a non-closing record commit. + **Step two — read the pass against the ordering, before any next call exists:** - **A source block standing** → **the source rule decides what must be repaired or answered, and its From c2e49326338200248b68d185c42636eee91f6ca1 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Wed, 16 Sep 2026 20:56:44 +0200 Subject: [PATCH 173/181] docs(plans): read the findings by object id; guard 8b's pin removal Pass 47's two findings. Both are defects in the two repairs immediately before them, and both are confirmed by execution rather than accepted on assertion. The Blocker is an overclaim of mine. The previous revision made step two read git show "HEAD:" and said the source was immutable once step one's tree comparison had passed. That is true of the commit object and false of HEAD, which is a movable ref: a commit, amend, reset, rebase or checkout between the record and the read redirects both reads, and a move between the two reads can take the two branch files from different commits. Step one now persists the record commit's full object id to .context/loop-rule-records-commit, step two reads both slots from that exact object, and a moved HEAD is a reason to stop and report rather than to resolve the files again. The new scratch path is in the cleanup list. The Major is real and was demonstrated. With 8b's removal unguarded, a DIRECTORY at the pin path let the block reach the closing commit: rm -f fails on a directory, cp source dir then succeeds by writing beneath it, both commands report success, and condition 6 discovers the missing oracle only after history has moved. Observed under sh, dash and bash - the old shape printed REACHED THE COMMIT with the path still a directory, and the repaired shape stops at the removal. Guarding the removal is what closes that case, and the prose says so rather than crediting the test -f that follows it: no demonstrated path reaches that check, because the removal guard fires first. It stays as a residual check on the destination's shape, described as one. Verified: the sh -n / dash -n failure set is unchanged at seven, all pre-existing process substitution; step two's new block is shellcheck clean. --- .../gate-a-plan-om0bdd7udh-pass-47.md | 3 + .../2026-09-14-loop-rule-consolidation.md | 72 ++++++++++++++----- 2 files changed, 59 insertions(+), 16 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-47.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-47.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-47.md new file mode 100644 index 0000000..28c492b --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-47.md @@ -0,0 +1,3 @@ +BLOCKER | high | Task 15 step 7, step two | The repair reads `git show "HEAD:$SPEC"` and `git show "HEAD:$QUAL"`, but `HEAD` is a movable ref and step one persists neither the record commit's object id nor an identity that step two can use. A commit, amend, reset, rebase, or checkout between the tree comparison and step two can redirect both reads, and a move between the two commands can take the branch files from different commits; the statement that the source is immutable is true of the commit object, not of `HEAD`. | The ordering can authorize a repair, continue, suspension, park, or close from findings other than the two blobs step one committed, violating target §A's requirement that every finding-derived predicate read the separately validated files of one logical pass. | After step one's guarded commit and tree comparison, persist its full object id and have step two read and guard both slot paths from that same exact object; stop and report if `HEAD` moved before routing rather than resolving the files through `HEAD` again. +MAJOR | high | Task 15 step 8, 8b closing-message pin | `rm -f .context/loop-rule-validated-msg` is unchecked, and the guarded `cp` does not establish that the destination path itself became the pinned file: if that path is a directory, `rm -f` fails but `cp source directory` succeeds by writing beneath it, so the following closing commit runs despite the prose claiming the guard stops before any commit without a refreshed oracle. | The close mutates history and only discovers the missing oracle in condition 6 after the closing commit, forcing a Failure handoff and leaving a rejected closing commit for a scratch-path shape that should have failed before the act. | Guard the removal and require the destination to be the regular file just written before running `git commit`, or copy to a newly created regular temporary file and atomically rename it into place under guarded operations. +END OF FINDINGS (2 total) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index 8ce525b..49bf4a3 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -2996,6 +2996,14 @@ git commit -m "WIP: pass $P records" \ # whatever produced it. test "$(git rev-parse "HEAD^{tree}")" = "$ITREE" \ || { echo "the commit's tree is not the index that was pinned; step two does not run"; exit 1; } + +# Persist THIS commit's object id, because step two must read the two slots from this +# exact commit. `HEAD` is a movable ref: a commit, amend, reset, rebase or checkout +# between here and there redirects both reads, and a move BETWEEN step two's two reads +# can take the branch files from different commits. The commit object is immutable; +# `HEAD` is not, and only the object id carries that. +git rev-parse HEAD > .context/loop-rule-records-commit \ + || { echo "recording the records commit's object id FAILED; step two does not run"; exit 1; } ``` **What this does not cover, disclosed rather than guarded:** naming the paths controls what this step @@ -3003,14 +3011,34 @@ test "$(git rev-parse "HEAD^{tree}")" = "$ITREE" \ staged before this step runs is carried in, and no check in this plan catches it** — the close's dirty-set and changed-path checks read 8a's commit, not these. Nothing here fixes that. -**Step two reads the COMMITTED findings, not the worktree copies** — `git show "HEAD:$SPEC"` and -`git show "HEAD:$QUAL"`, with `SPEC` and `QUAL` as step one spelled them. Target §A requires every -finding-derived predicate to read the **validated** findings file, and the worktree copy is the one -thing that can still change after validation: a `commit-msg` or `pre-commit` hook that rewrites only -the worktree leaves step one's checks entirely satisfied — the staged index, the pin and the commit -tree all agree — while the file a reader would open no longer holds what the pass was accepted on. -Reading `HEAD` closes that, because the commit is immutable once step one's tree comparison has -passed. +**Step two reads the findings out of step one's RECORD COMMIT, by object id — not the worktree +copies, and not through `HEAD`.** Target §A requires every finding-derived predicate to read the +**validated** findings file, and the worktree copy is the one thing that can still change after +validation: a `commit-msg` or `pre-commit` hook that rewrites only the worktree leaves step one's +checks entirely satisfied — the staged index, the pin and the commit tree all agree — while the file +a reader would open no longer holds what the pass was accepted on. **A commit object is immutable; +`HEAD` is a movable ref and is not**, so reading `HEAD:` would be redirected by any commit, amend, +reset, rebase or checkout in between, and a move between the two reads could take the two branch +files from different commits. The object id step one persisted is what makes the source one commit: + +```bash +NONCE=; P= +SPEC=".context/codex-reviews/gate-b-spec-$NONCE-pass-$P.md" +QUAL=".context/codex-reviews/gate-b-quality-$NONCE-pass-$P.md" +RECORDS=$(cat .context/loop-rule-records-commit) \ + || { echo "cannot read the records commit id — step one did not complete"; exit 1; } +test "${#RECORDS}" -eq 40 || { echo "records commit id is not a full object name"; exit 1; } +# Both reads come from this one object, so neither can be redirected and the two +# branch files cannot come from different commits. +git show "$RECORDS:$SPEC" \ + || { echo "cannot read the committed spec findings from the records commit"; exit 1; } +git show "$RECORDS:$QUAL" \ + || { echo "cannot read the committed quality findings from the records commit"; exit 1; } +# `HEAD` having moved is not a reason to re-resolve the files — it is a reason to stop +# and report, because something reached this repository between the record and the read. +test "$(git rev-parse HEAD)" = "$RECORDS" \ + || { echo "HEAD moved after the records commit — stop and report the concrete state"; exit 1; } +``` **What this does not establish, disclosed rather than guarded.** It does not prove those committed bytes are the ones **pass acceptance** validated. Step one pins what is on disk when step one runs; @@ -3334,15 +3362,26 @@ test -s .context/loop-rule-closing-msg || { echo "closing message missing or emp # oracle must not be the same mutable path the commit reads: a `commit-msg` hook # that rewrites git's copy AND this ignored source file would otherwise make the # post-close comparison pass on a body nobody validated. -# Remove any pin an earlier attempt left, then copy under a guard — but the GUARD is -# what does the work. It stops before the commit, so nothing closes against a pin this -# invocation did not write. The `rm -f` promises no cleanup: where `.context/` is -# unwritable it fails too and the earlier pin survives, non-empty enough for condition -# 6's `test -s`. It is simply never reached, because the guard exits first and leaves -# both files on disk to be read. Observed in a disposable repository, not reasoned. -rm -f .context/loop-rule-validated-msg +# Guarding the REMOVAL is what closes the demonstrated case. Unguarded, a DIRECTORY at +# this path let the block reach the commit: `rm -f` fails on a directory, `cp source +# dir` then succeeds by writing beneath it, both commands report success and condition +# 6 only finds the missing oracle after history has moved. Observed under sh, dash and +# bash — the old shape printed REACHED THE COMMIT with the path still a directory. +# The copy is guarded because it is the write, and a surviving earlier pin would be +# non-empty enough for condition 6's `test -s`. The `test -f` after them is a residual +# check on the destination's shape; no demonstrated path reaches it, since the removal +# guard fires first. None of this promises cleanup — a failed step leaves whatever is +# on disk and stops before the commit, which is the whole of what it guarantees. +rm -f .context/loop-rule-validated-msg \ + || { echo "removing the previous pin FAILED — run Failure; do NOT commit"; exit 1; } cp .context/loop-rule-closing-msg .context/loop-rule-validated-msg \ || { echo "pinning the validated closing message FAILED — run Failure; do NOT commit"; exit 1; } +# A guarded `cp` is not enough on its own: where a DIRECTORY sits at the destination, +# `rm -f` fails and `cp source dir` succeeds by writing beneath it, so both commands +# report success, the commit runs, and condition 6 only discovers the missing oracle +# after history has moved. Require the destination to be the regular file just written. +test -f .context/loop-rule-validated-msg \ + || { echo "the pin path is not a regular file — run Failure; do NOT commit"; exit 1; } # `--cleanup=verbatim` so the stored body is the validated bytes: git's default # cleanup for -F strips trailing whitespace and collapses blank runs, and # condition 6 compares bytes. It carries no `-m`, so `is_wip_commit` still misses it. @@ -3404,7 +3443,8 @@ rm -f .context/loop-rule-base .context/loop-rule-reviewed-tip .context/loop-rule .context/loop-rule-c.src .context/loop-rule-w.src \ .context/loop-rule-a.txt .context/loop-rule-b.txt .context/loop-rule-landed-msg \ .context/loop-rule-validated-msg \ - .context/loop-rule-final-blobs .context/loop-rule-committed-blobs + .context/loop-rule-final-blobs .context/loop-rule-committed-blobs \ + .context/loop-rule-records-commit if ls .context/loop-rule-* >/dev/null 2>&1; then echo "cycle scratch survives the close:"; ls .context/loop-rule-*; exit 1 fi From 16f8167665d3c400aa7170bb42d233cd971ed837 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Wed, 16 Sep 2026 20:57:01 +0200 Subject: [PATCH 174/181] docs(context): record Gate-A plan passes 46 and 47 Both rounds' findings, their repairs, and the two claims of mine they corrected. Pass 46's scope question went to two independent readings that disagreed; the record names which was taken and why the other's supporting claim was false in this repository. Pass 47 caught the HEAD-versus-commit overclaim that the pass-46 repair introduced. --- .context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md index 615320c..feba097 100644 --- a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md @@ -315,6 +315,8 @@ worth checking before a pass rather than after. | — | — | — | — | — | — | **THIRTEENTH BOUNDED REVISION**, both pass-45 findings, after putting the scope question to Codex as a sparring partner and validating his answer. **Splitting accounting row 3 by entry is a repair, not a new condition** — row 4 is already split the same way for the same reason, so no human decision was owed. Resume gains `merge-base --is-ancestor ba15e83 "$BASE"`; Task 10 prints the captured diff before filtering it. **Two of my own claims were wrong and are corrected**: the tip-removal comment (8a rewrites the tip before 8b, so a stale tip is not what condition 5 reads — the real risk is two candidates' markers standing together), and the rev-parse rationale — **without `--verify` a failed `git rev-parse :` prints its ARGUMENT, not nothing**, so two failures compare equal only where the revisions do. Observed, not reasoned. 16 fixtures × `sh`/`bash`, 15 × `dash`, all green | | 46 | f9bd330 | 2→**1** | 1→**1** | 1→**0** | yes | **BLOCKER: step one commits the pass records without binding them to what was validated, and step two then read the WORKTREE copies** — so a hook rewriting only the worktree leaves every step-one check satisfied while the file a reader opens no longer holds what the pass was accepted on. It points at the limit the twelfth revision **disclosed in its own prose**, which is why it was not created by the pass-45 repair | | — | — | — | — | — | — | **FOURTEENTH BOUNDED REVISION.** **The scope question was put to two independent readings and they DISAGREED** — both quoting the same §A sentence (`target-text.md:85-89`). A subagent read it as "validate once, before the predicate reads" → new obligation; Codex read it as an **input invariant** on what the predicate actually reads → repair, and ran fixtures showing the worktree-rewrite hole is real. **Codex's reading was taken**, because the subagent's supporting claim — that findings files are gitignored so nothing finding-derived is committed — is **false in this repo**: `.gitignore` exempts `.context/codex-reviews/` and this cycle commits its findings files. **Repair taken:** step two reads `git show "HEAD:$SPEC"` / `HEAD:$QUAL`. **Declined with reason:** re-running the structural and eligibility checks on committed content — that is Close condition 4's *closing*-commit duty and the approved text does not place it on a record commit. **Residual named, not repaired:** nothing establishes that the committed bytes are the ones *pass acceptance* validated; an edit between acceptance and staging is pinned as staged. Closing that needs the accepted blob ids carried out of the acceptance step in a file, a duty **no approved text carries**. Demonstrated under `sh`: with a worktree-only rewriting hook, step one exits 0, `HEAD` holds the validated line and the worktree holds `REWRITTEN IN WORKTREE ONLY` | +| 47 | 221819c | 1→**2** | 1→**1** | 0→**1** | yes | **Both findings are defects in the two revisions immediately before them, both demonstrated.** **BLOCKER — my own overclaim**: step two read `git show "HEAD:"` and the prose called the source immutable. That is true of the **commit object** and false of **`HEAD`**, a movable ref — any commit, amend, reset, rebase or checkout in between redirects both reads, and a move *between* the two reads can take the branch files from different commits. **MAJOR**: 8b's unguarded `rm -f` let a **directory** at the pin path reach the closing commit — `rm -f` fails on a directory, `cp source dir` then succeeds by writing beneath it. Observed under all three shells: the old shape printed `REACHED THE COMMIT` with the path still a directory | +| — | — | — | — | — | — | **FIFTEENTH BOUNDED REVISION.** Step one persists the record commit's object id to `.context/loop-rule-records-commit` (added to the cleanup list); **step two reads both slots from that exact object**, and a moved `HEAD` is a reason to **stop and report**, never to re-resolve. 8b's removal is guarded — **that is the line that closes the demonstrated case**, and the prose credits it rather than the `test -f` after it, which **no demonstrated path reaches**. Commit `c2e4932` | | 26 | 05ec1b2 | **5** | **1** | 2 | yes | first pass on the revised plan. Resume audited progress by `WIP:` commits and called any non-`WIP` commit a stale base — the handoff leaves two other shapes, one of them `HEAD` **at** the base with everything in the index | | 27 | f69db7d | 5→**6** | 1→**3** | 2→**1** | yes | the close required the commit's **tree to equal the tip's**, which target §I **parks**; `reset --soft` leaves the index untouched, so the prose beside it was wrong about git; Resume still inferred staleness from a commit subject where §A3 names that state as reachable | | 28 | f2e5dda | 6→**6** | 3→**2** | 1→**2** | yes | pass 27's stale-base fix reached one site of three; its reset paragraph stated the index behaviour correctly and repeated the false claim four lines later; nothing checked the closing commit's **parent**; the message was validated in the file, where `commit-msg` hooks rewrite git's copy after `-F` reads it | From fe6bd6d5965011c8f7d45161da259da74238a935 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Thu, 17 Sep 2026 09:06:52 +0200 Subject: [PATCH 175/181] docs(context): stop cycle om0bdd7udh at pass 47, decided through a gate MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The question "run pass 48 or stop" was put to Codex as a gate rather than decided by instinct, with the pass data and the governing §5 text. Its assessment is kept at .context/codex-reviews/gate-loop-health-pass47.md. Three tells are established at pass 47: findings rose 1 to 2, Blockers are flat at 1 for a third pass, and both findings cluster on the instrument. §5 makes that surface mandatory and does not make clearly-stuck a precondition for it. The clearly-stuck exit is not available, and the record should not read as though it were. None of its three conjuncts is met: the Blocker curve across passes 40 to 47 is 1,0,2,2,2,1,1,1 - low and oscillating rather than plateaued; no affirmative coverage judgement exists, the twelfth revision saying in its own body that its sweep does not close the defect class; and the regeneration is isolated lineages rather than each round's fix producing the next, since passes 44 and 45 re-raised nothing and pass 46 pointed at a disclosed limit. One attribution in the pass-47 row was wrong and is corrected. Only the Blocker descends from the pass-46 repair. The Major's line was introduced at 4752a35 - git log -S'rm -f .context/loop-rule-validated-msg' names that commit and no other, and 221819c does not touch that line at all. The instinct this replaces is recorded so it is not re-adopted: "run one more pass and stop if it produces another self-created finding". Self-created ancestry decides fix-set membership, not loop health, and pass 47 already supplies the case that rule was waiting for. The cycle stays open and unclean. The floor of 3 is long satisfied; what is missing is a clean pass, and surfacing never closes a cycle. --- .../gate-a-plan-om0bdd7udh-resume.md | 24 +++++++- .../codex-reviews/gate-loop-health-pass47.md | 60 +++++++++++++++++++ 2 files changed, 82 insertions(+), 2 deletions(-) create mode 100644 .context/codex-reviews/gate-loop-health-pass47.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md index feba097..bd5b8a3 100644 --- a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md @@ -13,7 +13,27 @@ Nothing depends on it; the pass files and the repo are authoritative where this ## Resume here -**READ THIS FIRST — the state below describes pass 44 and is superseded by the pass-45 rows in the +**STOPPED AT PASS 47 ON A MANDATORY TWO-TELL SURFACE, AND THE STOP WAS DECIDED THROUGH A GATE, NOT BY +INSTINCT.** The question "run pass 48 or stop" was put to Codex with the pass data and the governing +§5 text; its assessment is `.context/codex-reviews/gate-loop-health-pass47.md`. **Three tells are +established** — findings rose 1 → 2, Blockers flat at 1 for a third pass, instrument cluster — and §5 +makes that surface mandatory, independently of "clearly stuck". **The clearly-stuck exit is NOT +available**: none of its three conjuncts is met (the Blocker curve is low and oscillating rather than +plateaued, no affirmative coverage judgement exists — the twelfth revision says in its own body that +its sweep does not close the class — and the regeneration is isolated lineages rather than each +round's fix producing the next). + +**What the human decides:** release the suspension and authorize pass 48 against `c2e4932`, or park +the cycle. **Neither closes it.** The floor of 3 is long satisfied; what is missing is a **clean** +pass, and surfacing never closes a cycle. + +**The instinct that was rejected, and why**, so it is not re-adopted: *"run one more pass, and stop if +it produces another self-created finding"*. Self-created ancestry decides **fix-set membership**, not +loop health, and §5 says a small correction-of-a-correction is not by itself evidence of a plateau. +Pass 47 **already** supplies the self-created case that rule was waiting for, so pass 48 would gather +no missing decision fact — it would spend a pass after the rule had already required the checkpoint. + +**READ THIS NEXT — the state below describes pass 44 and is superseded by the pass-45 rows in the table.** The plan now stands at the thirteenth revision's commit; pass 45 found **2** findings (1 Blocker, 1 Major), both repaired, and the next pass is **46**. Codex is being used as a **sparring partner** on scope questions: put the question, require evidence, validate his answer, apply what @@ -315,7 +335,7 @@ worth checking before a pass rather than after. | — | — | — | — | — | — | **THIRTEENTH BOUNDED REVISION**, both pass-45 findings, after putting the scope question to Codex as a sparring partner and validating his answer. **Splitting accounting row 3 by entry is a repair, not a new condition** — row 4 is already split the same way for the same reason, so no human decision was owed. Resume gains `merge-base --is-ancestor ba15e83 "$BASE"`; Task 10 prints the captured diff before filtering it. **Two of my own claims were wrong and are corrected**: the tip-removal comment (8a rewrites the tip before 8b, so a stale tip is not what condition 5 reads — the real risk is two candidates' markers standing together), and the rev-parse rationale — **without `--verify` a failed `git rev-parse :` prints its ARGUMENT, not nothing**, so two failures compare equal only where the revisions do. Observed, not reasoned. 16 fixtures × `sh`/`bash`, 15 × `dash`, all green | | 46 | f9bd330 | 2→**1** | 1→**1** | 1→**0** | yes | **BLOCKER: step one commits the pass records without binding them to what was validated, and step two then read the WORKTREE copies** — so a hook rewriting only the worktree leaves every step-one check satisfied while the file a reader opens no longer holds what the pass was accepted on. It points at the limit the twelfth revision **disclosed in its own prose**, which is why it was not created by the pass-45 repair | | — | — | — | — | — | — | **FOURTEENTH BOUNDED REVISION.** **The scope question was put to two independent readings and they DISAGREED** — both quoting the same §A sentence (`target-text.md:85-89`). A subagent read it as "validate once, before the predicate reads" → new obligation; Codex read it as an **input invariant** on what the predicate actually reads → repair, and ran fixtures showing the worktree-rewrite hole is real. **Codex's reading was taken**, because the subagent's supporting claim — that findings files are gitignored so nothing finding-derived is committed — is **false in this repo**: `.gitignore` exempts `.context/codex-reviews/` and this cycle commits its findings files. **Repair taken:** step two reads `git show "HEAD:$SPEC"` / `HEAD:$QUAL`. **Declined with reason:** re-running the structural and eligibility checks on committed content — that is Close condition 4's *closing*-commit duty and the approved text does not place it on a record commit. **Residual named, not repaired:** nothing establishes that the committed bytes are the ones *pass acceptance* validated; an edit between acceptance and staging is pinned as staged. Closing that needs the accepted blob ids carried out of the acceptance step in a file, a duty **no approved text carries**. Demonstrated under `sh`: with a worktree-only rewriting hook, step one exits 0, `HEAD` holds the validated line and the worktree holds `REWRITTEN IN WORKTREE ONLY` | -| 47 | 221819c | 1→**2** | 1→**1** | 0→**1** | yes | **Both findings are defects in the two revisions immediately before them, both demonstrated.** **BLOCKER — my own overclaim**: step two read `git show "HEAD:"` and the prose called the source immutable. That is true of the **commit object** and false of **`HEAD`**, a movable ref — any commit, amend, reset, rebase or checkout in between redirects both reads, and a move *between* the two reads can take the branch files from different commits. **MAJOR**: 8b's unguarded `rm -f` let a **directory** at the pin path reach the closing commit — `rm -f` fails on a directory, `cp source dir` then succeeds by writing beneath it. Observed under all three shells: the old shape printed `REACHED THE COMMIT` with the path still a directory | +| 47 | 221819c | 1→**2** | 1→**1** | 0→**1** | yes | **One finding descends from the revision immediately before it, the other does not — an earlier row here claimed both did and was wrong.** `git log -S'rm -f .context/loop-rule-validated-msg'` names `4752a35` as the commit that introduced the Major's line, and `221819c` does not touch it at all. Both demonstrated. **BLOCKER — my own overclaim**: step two read `git show "HEAD:"` and the prose called the source immutable. That is true of the **commit object** and false of **`HEAD`**, a movable ref — any commit, amend, reset, rebase or checkout in between redirects both reads, and a move *between* the two reads can take the branch files from different commits. **MAJOR**: 8b's unguarded `rm -f` let a **directory** at the pin path reach the closing commit — `rm -f` fails on a directory, `cp source dir` then succeeds by writing beneath it. Observed under all three shells: the old shape printed `REACHED THE COMMIT` with the path still a directory | | — | — | — | — | — | — | **FIFTEENTH BOUNDED REVISION.** Step one persists the record commit's object id to `.context/loop-rule-records-commit` (added to the cleanup list); **step two reads both slots from that exact object**, and a moved `HEAD` is a reason to **stop and report**, never to re-resolve. 8b's removal is guarded — **that is the line that closes the demonstrated case**, and the prose credits it rather than the `test -f` after it, which **no demonstrated path reaches**. Commit `c2e4932` | | 26 | 05ec1b2 | **5** | **1** | 2 | yes | first pass on the revised plan. Resume audited progress by `WIP:` commits and called any non-`WIP` commit a stale base — the handoff leaves two other shapes, one of them `HEAD` **at** the base with everything in the index | | 27 | f69db7d | 5→**6** | 1→**3** | 2→**1** | yes | the close required the commit's **tree to equal the tip's**, which target §I **parks**; `reset --soft` leaves the index untouched, so the prose beside it was wrong about git; Resume still inferred staleness from a commit subject where §A3 names that state as reachable | diff --git a/.context/codex-reviews/gate-loop-health-pass47.md b/.context/codex-reviews/gate-loop-health-pass47.md new file mode 100644 index 0000000..077ad09 --- /dev/null +++ b/.context/codex-reviews/gate-loop-health-pass47.md @@ -0,0 +1,60 @@ +# Gate-A plan cycle `om0bdd7udh`: loop-health assessment after pass 47 + +## 1. FACT CHECK + +- **Pass 44 → 47 counts: correct.** I counted the finding records themselves, not the summary prose. The command + + ```text + for n in 44 45 46 47; do f=.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-$n.md; awk -F' *\| *' -v n="$n" 'BEGIN{b=0;m=0;t=0} /^(BLOCKER|MAJOR|MINOR|NIT) *\|/{t++; if($1=="BLOCKER")b++; if($1=="MAJOR")m++} END{printf "pass %d: findings=%d blockers=%d majors=%d\n",n,t,b,m}' "$f"; done + ``` + + returned: + + ```text + pass 44: findings=9 blockers=2 majors=7 + pass 45: findings=2 blockers=1 majors=1 + pass 46: findings=1 blockers=1 majors=0 + pass 47: findings=2 blockers=1 majors=1 + ``` + + The source files independently show those exact totals and severities: `.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-44.md:1-10`, `.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-45.md:1-3`, `.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-46.md:1-2`, and `.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-47.md:1-3`. The working table agrees at `.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md:312-318`. + +- **Floor and pass count: correct.** The plan cites one story in its governing header (`docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:11-13`); that story currently says `Risk: high` and `Security: none` (`docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md:3-7`). The working record states the resulting calculation as “risk high → level 2; security none → 0; max 2 ≠ 0 → 3” (`.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md:6-12`). My table-count command returned `numbered table rows=47 highest_pass=47`, so the floor is long satisfied. It does **not** close the cycle: the governing sentence is, “Your final pass must be clean — if the pass at the floor still finds Blocker/Major, keep going until clean or clearly stuck → then STOP and surface to the user” (`CLAUDE.md:132-136`); pass 47 contains one of each (`.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-47.md:1-3`). + +- **Repair ancestry through pass 46: your reading is correct.** Pass 40 calls its Major “this round's own repair” (`.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md:304`); pass 41 says the same of its Major (`.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md:306`); pass 42 attributes its first Blocker and part of its second to its own narrowing while identifying the other sites as pre-existing (`.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md:77-107`, especially `.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md:95-107`, and the compressed row at `.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md:308`); and pass 43's Minor is against the “new presence test” (`.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-43.md:5`, with the row at `.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md:310`). Pass 44 says nothing was re-raised against the eleventh revision (`.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md:312`), pass 45 says nothing was re-raised against the twelfth (`.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md:314`), and pass 46 says its limit was already disclosed and therefore was not created by the pass-45 repair (`.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md:316`). + +- **One attribution is wrong: only one of pass 47's findings is a defect in the pass-46 repair.** The Blocker is the direct descendant of pass 46's `git show "HEAD:"` repair (`.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-47.md:1`; commit `221819c`, inspected with `git show --unified=20 221819c -- docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md`, adds only that committed-findings mechanism). The Major about unguarded `rm -f .context/loop-rule-validated-msg` is older: `git log -S'rm -f .context/loop-rule-validated-msg' --format='introduced-or-removed %h %s' -- docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md` returned only `4752a35 docs(plans): guard every staging, and stop on a failed pin, extraction or fetch`; the diff for `4752a35` shows that exact unguarded removal being added before the guarded copy. The table says only that the findings are defects in “the two revisions immediately before them” (`.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md:318`); that phrase is temporally ambiguous and does not establish that both descend from pass 46, while the `-S` provenance check establishes that the Major does not. + +- **Four explicitly self-attributed factual corrections: correct, with that scope.** `git show -s --format='%H%n%B' fc12f25` records one claim “wrong, and was mine” (removing the old pin does not ensure no oracle); the same command on `f9bd330` records “Two claims of my own” (the stale-tip/condition-5 explanation and failed `rev-parse` output); and on `c2e4932` records one “overclaim of mine” (`HEAD` called immutable). The execution status is also supported rather than merely asserted: the working record calls the first two pass-45 corrections “Observed, not reasoned” (`.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md:315`), and pass 47's row says both findings were demonstrated (`.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md:318`); the three commit bodies provide the concrete observed states and outputs. This count excludes the two harness-model errors recorded in `4752a35`, because that body labels them harness errors rather than self-attributed factual claims. + +## 2. THE TELLS, APPLIED + +The governing text requires each pass report to expose trend, cluster, and any require↔withdraw pair (`CLAUDE.md:255-261`), then defines five tells and says: **“Any two present makes stop-and-surface mandatory, not discretionary”**; it also says clearly-stuck is **not** a precondition (`CLAUDE.md:263-268`). Applied at pass 47: + +1. **Finding count rising — present.** Findings went `1 → 2` from pass 46 to 47 (`.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md:316-318`; direct file count above). +2. **Blocker count failing to fall — present.** Blockers are `2, 1, 1, 1` across passes 44–47, including `1 → 1` at 46→47 (`.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md:312-318`; direct file count above). This is a tell even though it is not a six-pass plateau; §5 deliberately makes the two-tell stop independent of clearly-stuck (`CLAUDE.md:263-268`). +3. **Instrument cluster — present.** Both findings are in the gate/closure apparatus: Task 15 step 7's findings-record read and Task 15 step 8's closing-message pin (`.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-47.md:1-2`). **Reasoning, not direct observation:** I classify those as the “test instrument,” not product behaviour, because both govern whether the review/close procedure reads and pins its evidence; neither changes the plugin behaviour that the plan ultimately installs. The plan's product goal is the two §5 prompt copies plus shipped-hook strings (`docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:5-9`), which is distinct from Task 15's gate machinery. +4. **Prose-about-product-or-instrument cluster — not counted separately.** Both findings expose inaccurate prose as well as executable defects—the Blocker corrects the “immutable” claim and the Major corrects what the guard was said to guarantee (`.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-47.md:1-2`). **Reasoning:** the findings' operative failures are the movable ref and unguarded removal, so counting the same two findings again as a prose cluster would double-classify an instrument cluster; the mandatory result does not depend on doing so. +5. **Require↔withdraw pair — not established.** Pass 47 asks for an exact object id and a guarded removal (`.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-47.md:1-2`); neither line identifies a requirement that an earlier pass withdrew. **Reasoning:** absence cannot be proven by a keyword search alone, but the two complete finding lines state no restoration of removed behaviour, and the working record identifies no such pair for pass 47 (`.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md:318`). + +Therefore **three tells are established**, and §5 makes stop-and-surface mandatory now (`CLAUDE.md:263-268`). The fact that both findings are absorbable repairs does not cancel that result: “A finding that corrects the correction you just made and stays inside the assigned fix set is inside this loop's scope” (`CLAUDE.md:195-203`), while “The two rules above do not compete” (`CLAUDE.md:275-278`). + +## 3. “CLEARLY STUCK”, APPLIED + +§5's governing text is: “this exit needs three things together, and a missing one means keep going”: a cross-pass plateau, “an affirmative judgement that coverage is sufficient,” and “Blocker or Major findings that keep regenerating across genuine repair attempts, each round's fix producing the next” (`CLAUDE.md:225-235`). On this cycle: + +1. **Plateau — not met.** The recent Blocker curve is low and oscillating, not plateaued: passes 40–47 are `1, 0, 2, 2, 2, 1, 1, 1` (`.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md:304-318`). The last four are `2,1,1,1`; only the last three are flat. **Reasoning:** that is not the six-or-more-pass plateau §5 says was seen in the field, and §5 expressly warns that one low count is a snapshot rather than a plateau (`CLAUDE.md:225-231`). +2. **Affirmative coverage-sufficient judgement — not met.** The twelfth revision explicitly says its sweep “does NOT close the defect class” and lists untested cases; `git show -s --format='%H%n%B' fc12f25` produced those exact statements. The working table likewise says the sweep was “explicitly not a claim that the class is closed” (`.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md:313`). **Reasoning:** a bounded sweep plus named limits is not the affirmative sufficiency judgement §5 requires, so this predicate fails even without deciding whether every named untested case is “materially unreviewed” (`CLAUDE.md:228-233`). +3. **Blocker/Major regeneration across genuine repair attempts, each round producing the next — not met as a sustained chain.** There are real local descendants: passes 40 and 41 each found their own repair defect, two pass-42 findings came from its narrowing, and pass 47's Blocker came from pass 46 (`.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md:304-308,318`). But pass 44 and pass 45 explicitly re-raised nothing against their preceding revisions (`.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md:312,314`), pass 46 was not caused by pass 45 (`.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md:316`), and pass 47's Major descends from `4752a35`, not pass 46 (the `git log -S` output in §1). **Reasoning:** isolated regeneration lineages do not satisfy “each round's fix producing the next” across a plateau (`CLAUDE.md:233-235`). + +So the **clearly-stuck exit is unavailable**: none of its three conjuncts is fully met. That does not license pass 48, because the independent two-tell rule already requires a surface (`CLAUDE.md:263-268`). Also, surfacing does not close the cycle or waive an open Blocker/Major; the cycle resumes only on the user's decision (`CLAUDE.md:242-246`). + +## 4. THE ACTUAL RECOMMENDATION + +**Stop the loop now and hand cycle `om0bdd7udh` to a human; do not run pass 48 without that decision.** This recommendation follows the mandatory three-tell result, not a claim that the cycle is clearly stuck (`CLAUDE.md:263-268`; §2 above). The current plan revision is `c2e49326338200248b68d185c42636eee91f6ca1`; `git log -1 --format='%H %s' -- docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md` returned that object and `docs(plans): read the findings by object id; guard 8b's pin removal`. + +The human is being asked to decide whether to release the mandatory two-tell suspension and authorize pass 48 against `c2e4932`, or park the still-open Gate-A cycle (`CLAUDE.md:263-268`). A continue decision would not waive the clean-final-pass rule: because the floor is already met, the cycle closes only when a new pass has no Blocker/Major, with Minors/Nits merely collected (`CLAUDE.md:133-136,235-241`). A park decision leaves the cycle open and unclean rather than accepting pass 47 as final, because surfacing never closes the cycle (`CLAUDE.md:242-246`). + +## 5. THE THING I AM PROBABLY GETTING WRONG + +The flaw in “run one more pass, and if it produces another self-created finding, stop” is that **self-created ancestry is neither the governing threshold nor still prospective**. §5 uses ancestry to decide whether a finding belongs inside the assigned fix set (`CLAUDE.md:195-215`) and explicitly says a small correction-of-a-correction is not by itself proof of a plateau (`CLAUDE.md:275-278`); loop-health stopping is instead mandatory when any two tells are present (`CLAUDE.md:263-268`). Pass 47 already supplies three tells and already supplies the immediate self-created case your instinct proposes waiting for—the Blocker is the defect in pass 46's `HEAD:` repair (`.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-47.md:1`; `.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md:318`)—so pass 48 would not gather a missing decision fact; it would spend another pass after the rule has already required the human checkpoint. From 1ba45beaa3c079e192d17df8ba474586a693a84a Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Thu, 17 Sep 2026 09:45:39 +0200 Subject: [PATCH 176/181] docs(plans): resolve a commit once and use that id; wire the marker into Resume Pass 48's three Blockers, all of them one defect: the plan resolved HEAD a second time as the identity of a commit whose properties it had already checked. Repairing them one site at a time is the failure this cycle's own record calls its most reliable - a repair reaching one site of several - so this swept every use instead. Three defect sites, two named by the reviewer and one not. Step 7 step one checked HEAD^{tree} and then resolved HEAD again to persist the id, so a move in between could pin one commit's tree and record another. Step 8a resolved HEAD at each condition-4 check and again to write the closing tip, so a moved commit could become the tip having satisfied none of the parent, path, blob-identity or eligibility checks - and that one has a shipped payer: condition 5 accepts the substituted tip, reset --soft folds it into the closing commit, and content the candidate pass never reviewed is published. The third site is mine and was not reported: step 7 step three checked the repair commit's tree and then resolved HEAD again for the reviewed head, so the next call could be issued against a commit whose content was never checked. Each now resolves once, immediately after the commit lands, and uses that object id for every check and every record. 8a's condition-4 chain runs against that id throughout - parent, changed paths, per-file carry, committed blobs - and writes the same id as the tip. Checked and deliberately unchanged: step 4b resolves HEAD once, for its tree check, and persists no identity afterwards, so a move makes the comparison fail rather than pass. Conditions 1, 5 and 6 read HEAD on purpose, to detect a move, and stay as they are. The third finding is an omission of mine from the revision before this one. I added .context/loop-rule-records-commit as cycle state and gave it no recovery rule, which this plan requires of every state file. Resume now validates it - full object name, resolving to a commit, ba15e83 an ancestor of it and it an ancestor of HEAD, both slot paths present as blobs in that commit rather than in the worktree - and reports a recorded pass as unrouted unless the reconciliation shows step two's outcome in the content. Where it cannot be shown, that is a precondition stop: report and take no mutation, issue no call. Re-running step two is the ordinary continuation and needs nothing new. The marker is retired by the close's cleanup and by nothing else, so an interrupted cycle keeps it precisely so this check can find it. Verified: the sh -n / dash -n failure set is unchanged at seven, all pre-existing process substitution; the three changed blocks are shellcheck clean. --- .../gate-a-plan-om0bdd7udh-pass-48.md | 4 + .../2026-09-14-loop-rule-consolidation.md | 80 ++++++++++++++----- 2 files changed, 63 insertions(+), 21 deletions(-) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-48.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-48.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-48.md new file mode 100644 index 0000000..c8f5f5a --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-48.md @@ -0,0 +1,4 @@ +BLOCKER | high | Task 15 step 7 step one | The block validates `HEAD^{tree}` against `$ITREE` and only afterwards resolves `HEAD` again into `.context/loop-rule-records-commit`; a commit, amend, reset, rebase or checkout between those two commands makes the persisted object a different commit from the one whose tree passed the pin, recreating the movable-`HEAD` defect the revision says it closes. | Step two can read both branch files from an unvalidated moved commit and its final `HEAD = $RECORDS` check still passes, so the ordering can repair, suspend or park on findings other than the pass that was recorded. | Resolve the record commit oid once immediately after `git commit`, guard it, compare that exact object's tree to `$ITREE`, persist that same oid, and never re-resolve `HEAD` as the identity of the validated commit. +BLOCKER | high | Task 15 step 8a, condition 4 tail | After validating the findings commit through `HEAD`, 8a resolves `HEAD` again when it writes `.context/loop-rule-reviewed-tip`; if `HEAD` moves after the last condition-4 check but before that write, the moved commit becomes the closing tip without satisfying the parent, path, blob-identity or findings-eligibility checks. | Condition 5 then accepts the substituted tip, `reset --soft` folds it into the closing commit, and condition 6 can pass because the deliberately parked tree-equality check is the only later content check, silently publishing content the candidate pass never reviewed. | Capture the findings commit oid once when the commit lands, run every condition-4 check against that immutable oid, write only that oid as the tip, and let condition 5 reject any subsequent `HEAD` move. +BLOCKER | high | Resume progress reconciliation and Task 15 step 7 | The new `.context/loop-rule-records-commit` marker is added to cleanup but not to Resume's required sources or transition rules; an interruption after step one leaves the pass records committed and the marker present, yet Resume can declare reconciliation successful without establishing whether step two consumed that pass through the ordering. | Re-entry can apply a repair or issue another pass while a source block, suspension or stop answer from the committed pass is still standing, so the plan can execute the exact ordering violation its product forbids. | Make the record-commit marker and its two slot paths part of Resume's read-every-time reconciliation, validate them against the immutable commit, require any pending recorded pass to complete step two before mutation or another call, and retire or advance the marker only after that read is complete. +END OF FINDINGS (3 total) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index 49bf4a3..bc68575 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -628,6 +628,22 @@ Each by its `base` line **and** its completeness predicate: condition in the five regions appears in exactly one `span` or one `cond`. - `.context/loop-rule-baseline-diff.txt` — every line parses as `base`, a `site` record, or diff output belonging to the site above it, and **every inventoried site has a `site` record**. +- `.context/loop-rule-records-commit` — **present means a pass was recorded and step two may not + have read it.** Step 7 step one writes it, step two consumes it, and nothing else touches it, so on + re-entry it is the one artifact that says *a pass is recorded and the ordering may not have + spoken*. Validate it: a full 40-character object name, resolving to a commit, with `ba15e83` as an + ancestor and the commit itself an ancestor of `HEAD`; and its two slot paths present as blobs + **in that commit**, not in the worktree. **This step establishes the marker; it decides nothing.** + +**A valid marker is a finding of the reconciliation, not an instruction.** It says the recorded pass +exists and reports it as **unrouted unless the reconciliation below shows step two's outcome in the +content** — a repair committed on an authorizing route, a park, a stop answer. **Where it cannot be +shown, this is a precondition stop: report the marker, the commit it names and what the four sources +hold, and take no mutation and issue no call.** Rerunning step two over a recorded pass is the +ordinary continuation and needs nothing new; **what must not happen is a repair or a next call while +the ordering has not spoken**, which is the loop §A forbids and the one this plan must not itself +execute. The marker is **retired by the close's cleanup**, with the other scratch values, and by +nothing else — an interrupted cycle keeps it precisely so this check can find it. **On failure the answer depends on why you are here.** Outside a handoff: delete and rebuild from the `$BASE` blobs — never reuse, never repair in place, because a same-base partial file is the one @@ -2992,17 +3008,24 @@ test "$(git cat-file -t "$ITREE:$QUAL" 2>/dev/null)" = blob \ git commit -m "WIP: pass $P records" \ || { echo "records commit FAILED — stop here; step two does not run"; exit 1; } -# Any divergence between the pinned index and the resulting commit tree stops here, -# whatever produced it. -test "$(git rev-parse "HEAD^{tree}")" = "$ITREE" \ +# Resolve the record commit ONCE, immediately after it lands, and use that object id for +# everything below. `HEAD` is a movable ref: resolving it a second time — for the tree +# check, and again to persist the identity — lets a commit, amend, reset, rebase or +# checkout in between hand those two commands DIFFERENT commits, which recreates the +# very defect this records the object id to close. The commit object is immutable; +# `HEAD` is not, and only the captured id carries that. +RECORD=$(git rev-parse HEAD) \ + || { echo "cannot resolve the records commit; step two does not run"; exit 1; } + +# Any divergence between the pinned index and THAT commit's tree stops here, whatever +# produced it. +test "$(git rev-parse "$RECORD^{tree}")" = "$ITREE" \ || { echo "the commit's tree is not the index that was pinned; step two does not run"; exit 1; } -# Persist THIS commit's object id, because step two must read the two slots from this -# exact commit. `HEAD` is a movable ref: a commit, amend, reset, rebase or checkout -# between here and there redirects both reads, and a move BETWEEN step two's two reads -# can take the branch files from different commits. The commit object is immutable; -# `HEAD` is not, and only the object id carries that. -git rev-parse HEAD > .context/loop-rule-records-commit \ +# Persist the same id, because step two must read both slots from this exact commit — +# and a move between its two reads could otherwise take the branch files from different +# commits. +printf '%s\n' "$RECORD" > .context/loop-rule-records-commit \ || { echo "recording the records commit's object id FAILED; step two does not run"; exit 1; } ``` @@ -3108,7 +3131,13 @@ ITREE=$(git write-tree) \ # the repair stays staged, and the next call reviews a tree it is not in. git commit -m "WIP: fix " \ || { echo "repair commit FAILED — tip and reviewed head left as they are; no call may be issued"; exit 1; } -test "$(git rev-parse "HEAD^{tree}")" = "$ITREE" \ +# Resolve the repair commit ONCE and use that id for both the tree check and the head +# record. Resolving `HEAD` twice lets a move in between check one commit's tree and +# record another as the head the next call is issued against — the same defect step one +# records an object id to close, and not one the reviewer named here. +REPAIR=$(git rev-parse HEAD) \ + || { echo "cannot resolve the repair commit; no call may be issued"; exit 1; } +test "$(git rev-parse "$REPAIR^{tree}")" = "$ITREE" \ || { echo "the commit's tree is not the index that was pinned; no call may be issued"; exit 1; } # The previous candidate's closing tip. Guarded: without the guard a failed removal @@ -3117,9 +3146,10 @@ test "$(git rev-parse "HEAD^{tree}")" = "$ITREE" \ # reads in a correct sequence — 8a rewrites the tip before 8b runs.) rm -f .context/loop-rule-reviewed-tip \ || { echo "removing the previous candidate's tip FAILED — it survives; no call may be issued"; exit 1; } -# The head the NEXT call is issued against. Guarded, and no `cat`: the redirect can -# fail while a following `cat` prints the STALE value and the block exits 0. -git rev-parse HEAD > .context/loop-rule-reviewed-head \ +# The head the NEXT call is issued against — the SAME id whose tree was checked above, +# never a fresh resolution. Guarded, and no `cat`: the redirect can fail while a +# following `cat` prints the STALE value and the block exits 0. +printf '%s\n' "$REPAIR" > .context/loop-rule-reviewed-head \ || { echo "recording the reviewed head FAILED — no call may be issued"; exit 1; } ``` @@ -3293,15 +3323,22 @@ done > .context/loop-rule-final-blobs \ git add $FINAL \ || { echo "staging FAILED — run Failure; do NOT commit"; exit 1; } git commit -m "WIP: Gate-B findings files" || { echo "record commit FAILED — run Failure"; exit 1; } -test "$(git rev-parse HEAD^)" = "$HEADREV" || { echo "record commit's parent is not the reviewed head — run Failure"; exit 1; } -git diff --quiet "$HEADREV" HEAD -- "$@" \ +# Resolve the record commit ONCE, here, and run every condition-4 check against that +# object id. `HEAD` is a movable ref: re-resolving it at each check — and again at the +# tail, where the tip is written — lets a move land a commit as the closing tip that +# satisfied none of the parent, path, blob-identity or eligibility checks. Condition 5 +# would then accept it, `reset --soft` would fold it into the closing commit, and the +# only later content check is the one this plan deliberately parks. +REC=$(git rev-parse HEAD) || { echo "cannot resolve the record commit — run Failure"; exit 1; } +test "$(git rev-parse "$REC^")" = "$HEADREV" || { echo "record commit's parent is not the reviewed head — run Failure"; exit 1; } +git diff --quiet "$HEADREV" "$REC" -- "$@" \ || { echo "record commit changed paths beyond this pass's findings files — run Failure"; exit 1; } # The polarity is inverted here: exit 0 means "no difference", i.e. NOT carried. With # `&&` an execution error (status 2 or above) takes the same path as "carried" and the # close proceeds on a comparison that could not run. Status 1 — a real difference — is # the legitimate result this check wants. for f in $FINAL; do - git diff --quiet "$HEADREV" HEAD -- "$f" + git diff --quiet "$HEADREV" "$REC" -- "$f" case $? in 0) echo "record commit did not carry $f — run Failure"; exit 1 ;; 1) : ;; @@ -3314,7 +3351,7 @@ done # that on an unresolvable rev `git rev-parse` echoes its argument to stdout and exits # 128, so without the guard this file can hold a junk line and the loop's status is gone. for f in $FINAL; do - git rev-parse "HEAD:$f" \ + git rev-parse "$REC:$f" \ || { echo "reading the committed blob for $f FAILED — run Failure" >&2; exit 1; } done > .context/loop-rule-committed-blobs \ || { echo "writing the committed-blob list FAILED — run Failure"; exit 1; } @@ -3328,12 +3365,13 @@ TREESTATE=$(git status --porcelain) \ || { echo "reading the tree state after the record commit FAILED — run Failure"; exit 1; } test -z "$TREESTATE" || { echo "tree not clean after the record commit — run Failure"; exit 1; } -# Condition 4's tail: the closing tip. -# Guarded even though it ends the block: a failed redirect leaves whatever an earlier -# attempt wrote, and 8b's condition 5 compares HEAD against that. (A failed rev-parse +# Condition 4's tail: the closing tip — the SAME object every check above ran against, +# never a fresh resolution, so a `HEAD` that moved after those checks cannot become the +# tip. Guarded even though it ends the block: a failed redirect leaves whatever an +# earlier attempt wrote, and 8b's condition 5 compares against that. (A failed write # after a successful truncate leaves the file empty, which 8b's `test -n "$TIP"` does # catch — the redirect failure is the half it cannot.) -git rev-parse HEAD > .context/loop-rule-reviewed-tip \ +printf '%s\n' "$REC" > .context/loop-rule-reviewed-tip \ || { echo "recording the closing tip FAILED — run Failure; 8b must not run"; exit 1; } ``` From 5fb3b942f1907f5d6dd25b32d15b497f855e05f8 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Thu, 17 Sep 2026 09:52:58 +0200 Subject: [PATCH 177/181] docs(context): record pass 48 and write the handoff for a fresh session The working record now opens with a handoff block: state, the one decision waiting on Daniel, the standing orders in force, and the counted efficiency finding behind the focused-versus-full choice - roughly 4 or 5 of 25 findings across passes 42 to 48 had a payer outside this repository. Pass 48's three Blockers were one defect and are recorded as one, with the third site the reviewer did not report and the two sites checked and deliberately left alone. --- .../gate-a-plan-om0bdd7udh-resume.md | 49 +++++++++++++++++++ 1 file changed, 49 insertions(+) diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md index bd5b8a3..7224808 100644 --- a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md @@ -13,6 +13,53 @@ Nothing depends on it; the pass files and the repo are authoritative where this ## Resume here +## HANDOFF — 2026-09-17, written for a session that has just lost its context + +**State.** Plan anchor `1ba45be`; cycle `om0bdd7udh` **OPEN and UNCLEAN at pass 48**. The plan's +sixteenth revision repaired all three of pass 48's Blockers. **Nothing is implemented yet** — +`CLAUDE.md`, `plugins/dev-workflow/commands/workflow-init.md` and the hook are untouched. + +**Not yours to touch:** `docs/superpowers/specs/2026-08-30-dark-factory-vision.md`, `todos.md` and +`docs/field-reports/2026-09-17-sfx-review-loop-economics.md` are Daniel's own edits in flight. Leave +them dirty; do not commit, revert or review them. + +**The open decision, and it is one word from Daniel.** Three tells stand at pass 48 (findings 2 → 3, +Blockers 1 → 3, instrument cluster), so the surface is mandatory. The choice he was given: + +- **"fokussiert weiter"** → run pass 49 with the prompt narrowed to the four areas that can reach a + shipped product (below); +- **"voll weiter"** → run pass 49 over the whole plan, as passes 1–48 did. + +**Why that choice exists — the efficiency finding, counted rather than felt.** Of roughly 25 findings +across passes 42–48, **about 4 or 5 had a payer outside this repository**; the rest cost the one +person executing the plan a bad session. **That is ~80% bookkeeping** — but the other 20% are +serious: unreviewed content published through `reset --soft`, a plugin shipping without its version +bump, a hook change slipping past the invariant-4 review. The loop is **misallocated, not worthless**. +Of the plan's **54 shell blocks**, only four areas decide what is published: **8a/8b** (the closing +commit), **steps 3–4** (version bump and battery, invariant 12), **step 6** (the Gate-B range), and +**Tasks 10/11** (the hook change and its invariant-4 review). + +**Standing orders now in force** (also in the session memory): +- **Subagents, parallel, by default.** Fan out read-only inspection; keep writing serial — one writer. +- **Model per job**: haiku for mechanical scans, sonnet for ordinary judgement, opus only where a + wrong call is expensive (the close, closure conditions, anything a gate depends on). +- **A stop goes through a gate, never into a menu for Daniel.** Take the report's own output, make it + the gate question, put it to Codex, validate the answer, follow it. Only a decision the rules + reserve for a human reaches him — and then as one decision with the gate's reasoning attached. +- **Codex is the sparring partner**: he proposes with evidence, I validate, I apply what holds and say + what I rejected. +- **The utility test on every finding**: name the payoff — **bugfix, security-fix or performance** — + and name **who pays** if it is left: a user of the shipped product, or the executor at a terminal. + No nameable payoff → collect, never iterate. **It waives no floor, no mandatory tell and no closure + condition** — `todos.md` "Review-loop usefulness" (Dark Factory vision §§4/7/11) is the formal + version and activates no thresholds. + +**After a clean close**, the real work starts: both prompt copies, the seven hook strings and their +test expectations, version bump 0.11.0 → 0.12.0 + CHANGELOG, the quality battery, the evidence entry, +then **Gate B with a fresh nonce** — this cycle's is Gate-A plan only. + +--- + **STOPPED AT PASS 47 ON A MANDATORY TWO-TELL SURFACE, AND THE STOP WAS DECIDED THROUGH A GATE, NOT BY INSTINCT.** The question "run pass 48 or stop" was put to Codex with the pass data and the governing §5 text; its assessment is `.context/codex-reviews/gate-loop-health-pass47.md`. **Three tells are @@ -337,6 +384,8 @@ worth checking before a pass rather than after. | — | — | — | — | — | — | **FOURTEENTH BOUNDED REVISION.** **The scope question was put to two independent readings and they DISAGREED** — both quoting the same §A sentence (`target-text.md:85-89`). A subagent read it as "validate once, before the predicate reads" → new obligation; Codex read it as an **input invariant** on what the predicate actually reads → repair, and ran fixtures showing the worktree-rewrite hole is real. **Codex's reading was taken**, because the subagent's supporting claim — that findings files are gitignored so nothing finding-derived is committed — is **false in this repo**: `.gitignore` exempts `.context/codex-reviews/` and this cycle commits its findings files. **Repair taken:** step two reads `git show "HEAD:$SPEC"` / `HEAD:$QUAL`. **Declined with reason:** re-running the structural and eligibility checks on committed content — that is Close condition 4's *closing*-commit duty and the approved text does not place it on a record commit. **Residual named, not repaired:** nothing establishes that the committed bytes are the ones *pass acceptance* validated; an edit between acceptance and staging is pinned as staged. Closing that needs the accepted blob ids carried out of the acceptance step in a file, a duty **no approved text carries**. Demonstrated under `sh`: with a worktree-only rewriting hook, step one exits 0, `HEAD` holds the validated line and the worktree holds `REWRITTEN IN WORKTREE ONLY` | | 47 | 221819c | 1→**2** | 1→**1** | 0→**1** | yes | **One finding descends from the revision immediately before it, the other does not — an earlier row here claimed both did and was wrong.** `git log -S'rm -f .context/loop-rule-validated-msg'` names `4752a35` as the commit that introduced the Major's line, and `221819c` does not touch it at all. Both demonstrated. **BLOCKER — my own overclaim**: step two read `git show "HEAD:"` and the prose called the source immutable. That is true of the **commit object** and false of **`HEAD`**, a movable ref — any commit, amend, reset, rebase or checkout in between redirects both reads, and a move *between* the two reads can take the branch files from different commits. **MAJOR**: 8b's unguarded `rm -f` let a **directory** at the pin path reach the closing commit — `rm -f` fails on a directory, `cp source dir` then succeeds by writing beneath it. Observed under all three shells: the old shape printed `REACHED THE COMMIT` with the path still a directory | | — | — | — | — | — | — | **FIFTEENTH BOUNDED REVISION.** Step one persists the record commit's object id to `.context/loop-rule-records-commit` (added to the cleanup list); **step two reads both slots from that exact object**, and a moved `HEAD` is a reason to **stop and report**, never to re-resolve. 8b's removal is guarded — **that is the line that closes the demonstrated case**, and the prose credits it rather than the `test -f` after it, which **no demonstrated path reaches**. Commit `c2e4932` | +| 48 | c2e4932 | 2→**3** | 1→**3** | 1→**0** | yes | **All three Blockers are ONE defect**: the plan resolved `HEAD` a second time as the identity of a commit whose properties it had already checked. Step 7 step one (tree checked on one commit, id persisted from another); **8a — the one with a shipped payer**: a moved commit could become the closing tip having satisfied **none** of condition 4's checks, and `reset --soft` folds it into the closing commit, publishing content no pass reviewed; and the `.context/loop-rule-records-commit` marker added one revision earlier with **no Resume rule**, which this plan requires of every state file | +| — | — | — | — | — | — | **SIXTEENTH BOUNDED REVISION.** **Swept rather than patched per site** — repairing one of several sites is this cycle's most reliable defect. Each site now resolves the commit **once**, immediately after it lands, and uses that object id for every check and every record; 8a's whole condition-4 chain runs against it and writes it as the tip. **A third site the reviewer did not report** was found and fixed: step 7 step three checked the repair commit's tree and then re-resolved `HEAD` for the reviewed head. **Checked and deliberately unchanged:** step 4b resolves `HEAD` once and persists no identity, so a move makes its comparison fail rather than pass; conditions 1, 5 and 6 read `HEAD` on purpose, to detect a move. Resume now validates the marker and reports a recorded pass as **unrouted** unless the reconciliation shows step two's outcome — a precondition stop, no mutation, no call. Commit `1ba45be` | | 26 | 05ec1b2 | **5** | **1** | 2 | yes | first pass on the revised plan. Resume audited progress by `WIP:` commits and called any non-`WIP` commit a stale base — the handoff leaves two other shapes, one of them `HEAD` **at** the base with everything in the index | | 27 | f69db7d | 5→**6** | 1→**3** | 2→**1** | yes | the close required the commit's **tree to equal the tip's**, which target §I **parks**; `reset --soft` leaves the index untouched, so the prose beside it was wrong about git; Resume still inferred staleness from a commit subject where §A3 names that state as reachable | | 28 | f2e5dda | 6→**6** | 3→**2** | 1→**2** | yes | pass 27's stale-base fix reached one site of three; its reset paragraph stated the index behaviour correctly and repeated the false claim four lines later; nothing checked the closing commit's **parent**; the message was validated in the file, where `commit-msg` hooks rewrite git's copy after `-F` reads it | From 266ecf5b59df9f230a874dd81ae294ac06f63be4 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sat, 26 Sep 2026 09:21:42 +0200 Subject: [PATCH 178/181] docs(plans): close the Gate-A plan cycle; pass 56 returned NO FINDINGS MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The loop-rule consolidation implementation plan is approved. Pass 56 read docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md at blob 79649d37e89ebd07b60e53d6be5f44a4ccb1cbb5 and returned a zero-finding file. The plan is committed here unchanged from that read. Revisions 24-27 after the pass-52 stop, each decided by Daniel on 2026-09-25/26: - option B: Task 15's battery runs in a disposable clone of the recorded candidate; the candidate id is resolved once and recorded in the same block; - D1: Gate-B findings files are committed once, at the close, not per pass; the loss exposure before the close is accepted and stated in the plan; - option 1: this change's Gate-B cycle follows CLAUDE.md §5 as it stands at $BASE; the installed ordering is the product under test, not the cycle's router; - the per-pass records commit, its marker, 8a, the closing tip and the two-invocation close are gone; step 8 is one closing block plus the postcondition block. The plan is 379 lines shorter than at pass 53. cycle om0bdd7udh; floor 3 per {docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md (level 2)}; hook reminder threshold absent cycle om0bdd7udh; Gate-A plan (passes 1-56, codex): Findings 32,32,31,20,16,11,13,15,9,17,11,6,4,5,4,5,9,4,3,4,5,10,5,8,5,5,6,6,9,7,5,4,7,6,1,2,3,3,2,2,1,3,5,9,2,1,2,3,6,4,6,4,7,7,3,0. Blockers 1,0,0,10,7,2,5,2,2,1,2,1,1,2,0,2,2,0,0,2,2,5,3,2,2,1,3,2,5,4,2,2,3,1,0,0,0,1,1,1,0,2,2,2,1,1,1,3,5,2,2,2,2,0,0,0. Majors 28,25,23,7,7,6,3,9,6,14,7,5,1,2,2,2,3,3,2,2,3,3,2,5,2,2,1,2,3,3,2,1,3,1,1,1,1,2,1,1,1,1,2,7,1,0,1,0,1,1,2,0,2,4,2,0. Every count is recomputed from the 56 findings files, each of which validates (terminator present, line count matching). No count is written as unknown. "codex" names the reviewer as the previous cycle's record (ba15e83) did; the session logs of passes 53-56 show model gpt-6-astra, and earlier passes were not re-derived. Collected and not repaired, per the Minor/Nit rule: pass 52's Task 0 baseline coverage Minor; pass 53's two Minors and Nit on 4b's reset rationale, 4c's merge message and 4c's fixture count; pass 54's Minor and two Nits on stale tip/condition wording and counts; pass 55's Task 7 checklist Minor. The working record is retired by rename, as ba15e83 did: gate-a-plan-om0bdd7udh-resume.md becomes gate-a-plan-om0bdd7udh-history.md, so no later cycle can adopt it, and the per-pass account survives. --- ...e.md => gate-a-plan-om0bdd7udh-history.md} | 880 ++++++++- ...-a-plan-om0bdd7udh-pass-49-dispositions.md | 43 + .../gate-a-plan-om0bdd7udh-pass-49.md | 7 + .../gate-a-plan-om0bdd7udh-pass-50.md | 5 + ...-a-plan-om0bdd7udh-pass-51-dispositions.md | 25 + .../gate-a-plan-om0bdd7udh-pass-51.md | 7 + ...-a-plan-om0bdd7udh-pass-52-dispositions.md | 146 ++ .../gate-a-plan-om0bdd7udh-pass-52.md | 5 + ...-a-plan-om0bdd7udh-pass-53-dispositions.md | 52 + .../gate-a-plan-om0bdd7udh-pass-53.md | 8 + ...-a-plan-om0bdd7udh-pass-54-dispositions.md | 52 + .../gate-a-plan-om0bdd7udh-pass-54.md | 8 + ...-a-plan-om0bdd7udh-pass-55-dispositions.md | 42 + .../gate-a-plan-om0bdd7udh-pass-55.md | 4 + ...-a-plan-om0bdd7udh-pass-56-dispositions.md | 28 + .../gate-a-plan-om0bdd7udh-pass-56.md | 2 + .../2026-09-14-loop-rule-consolidation.md | 1626 +++++++++-------- 17 files changed, 2209 insertions(+), 731 deletions(-) rename .context/codex-reviews/{gate-a-plan-om0bdd7udh-resume.md => gate-a-plan-om0bdd7udh-history.md} (54%) create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-49-dispositions.md create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-49.md create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-50.md create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-51-dispositions.md create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-51.md create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-52-dispositions.md create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-52.md create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-53-dispositions.md create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-53.md create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-54-dispositions.md create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-54.md create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-55-dispositions.md create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-55.md create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-56-dispositions.md create mode 100644 .context/codex-reviews/gate-a-plan-om0bdd7udh-pass-56.md diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-history.md similarity index 54% rename from .context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md rename to .context/codex-reviews/gate-a-plan-om0bdd7udh-history.md index 7224808..40854e3 100644 --- a/.context/codex-reviews/gate-a-plan-om0bdd7udh-resume.md +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-history.md @@ -13,31 +13,317 @@ Nothing depends on it; the pass files and the repo are authoritative where this ## Resume here +### PASS 56 — RUN, VALID, CLEAN (0 findings). Closure-eligible. NOT closed: the closing commit is not authorized. + +Blob `79649d37e89ebd07b60e53d6be5f44a4ccb1cbb5`, HEAD `5fb3b94`, session +`01a0dc89-184b-7d10-86d4-d0d659130526`. The file is exactly `NO FINDINGS` / `END OF FINDINGS (0 total)`. +Floor 3; pass 56 is above it. **Next: Daniel decides the Gate-A closing act.** Per §5 as this cycle +started, that is committing the reviewed plan unchanged together with its provenance line and its +passes 1–56 curve. The curve still has to be assembled from the findings files. + +### 2026-09-26 — revision 27 applied, pass 56 authorized (one pass only) + +**Release.** Daniel released the sparring brief `.context/sparring/2026-09-26-loop-pass55-prompt.md`: +repair the two pass-55 Majors, verify, run exactly one pass 56, then stop. + +**Revision 27** (blob `f34d032…` → `79649d37e89ebd07b60e53d6be5f44a4ccb1cbb5`, 3784 lines): +- a step-7 re-review route for a clean or zero-finding pass whose evidence changed, or which Failure + charges; Close now requires that no re-review is owed; +- 7b points at that route; +- the no-repair branch's dirty check is an exact-path block admitting only this plan, this cycle's + findings slots, their per-slot dispositions notes and the working record. + +**Executed:** the new block under sh and dash, 5 cases, as expected; the 12 close cases still hold. +**Walkthrough only:** the re-review and Failure-charged routes. The pass-55 Minor stays collected. + +### PASS 55 — RUN, VALID, UNCLEAN (0 Blockers, 2 Majors, 1 Minor). Cycle OPEN. No pass 56 authorized. + +Blob `f34d032…`, session `01a0da21-a62f-7d83-bb43-d0a00cecad82`. Major 1: the no-repair branch's +"anything else dirty" stop catches the cycle's own uncommitted findings files. Major 2: step 7 has no +route for a clean pass that owes re-review because evidence changed. Both came from revision 26, both +are in-set and small. The Minor (Task 7 checklist) is collected. One tell. Details: +`gate-a-plan-om0bdd7udh-pass-55-dispositions.md`. Next move is Daniel's. + +### 2026-09-25 late — option 1 decided, revision 26 applied, pass 55 authorized (one pass only) + +**Decision.** Daniel sent the sparring assessment and brief `2026-09-25-194705-loop-pass54-*`. Both +the agent and the reviewer recommended **option 1**: this change's Gate-B cycle follows **§5 at +`$BASE`** (`CLAUDE.md:153`); the installed ordering does not steer it and stays the product under test. +Daniel's message was taken as the release. + +**Revision 26** (blob `2d7eeff…` → `f34d03267b9a4387b8121b76b773c6eb7281d71a`, 3747 lines): +- step 7 routing rewritten to §5-at-`$BASE` routes, with an accounting table; +- Failure, Resume, Close and 7b relabelled as the plan's own procedures; +- the no-repair branch reconciled with the complete set (row 27); +- the step-4b block names repair files, including spec and version files; +- the working record is kept until the close has succeeded, and retired last. + +**Pass-54 Majors:** 1 → option 1 applied; 2, 3, 4 → fixed. Minor and Nits stay collected. +**Executed:** close composition test, 12 cases under sh and dash; the step-4b block with a spec repair +and with a records-only change. Prechecks: 54 fences, sh -n 7 / dash -n 8, same classes. + +### PASS 54 — RUN, VALID, UNCLEAN (0 Blockers, 4 Majors). Cycle OPEN. Stop: finding 1 is a contract question. + +Blob `2d7eeff…`, session `01a0d9e6-2fca-73a3-b15f-a58682ac11a8`. 7 findings: 0 Blockers, 4 Majors, +1 Minor, 2 Nits. Majors 2–4 are in-set corrections of revision 25; Major 1 asks which rule set +governs this change's Gate-B cycle (§5 at `$BASE` versus the installed ordering, which disagree on +tells at a closing pass). Details: `gate-a-plan-om0bdd7udh-pass-54-dispositions.md`. Next move is +Daniel's. + +### 2026-09-25 evening — D1 accepted, revision 25 applied, pass 54 authorized (one pass only) + +**Decision.** Pass 53's mandatory stop (tells 1, 2, 3) was surfaced. Daniel released the sparring +brief `.context/sparring/2026-09-25-183846-loop-d1-prompt.md`, accepting **D1**: Gate-B findings +files are committed once, at the close, not per pass. The durability loss before the close is +accepted and stated in the plan (Close, *D1*). Proposal: +`.context/loop-rule-task15-simplification-proposal-2026-09-25.md`, with the reviewer's three +corrections: no `.context/loop-rule-*` exemption; staged findings files are pinned too; exact slot +paths instead of a glob, so a dispositions note is not taken as findings. + +**Revision 25:** plan blob `769fbc6…` → `2d7eeff004758108222f9f921c78f72c4dacdd28` (3706 lines, 457 +fewer). Step 7, Close, step 8, Resume, the index, the accounting rows and line 31 were changed; 4b's +failure route was changed too. Details are in `.context/gate-a-plan-pass-54-instruction.md`, section +"REVISION TWENTY-FIVE". + +**Pass-53 findings → dispositions:** 1 → fixed (4b failure route); 2 → fixed (one sequence, +explicit no-repair branch); 3 → fixed (`$BASE` loaded in the block); 4 → dissolved (no records +commit, so `HEAD` stays the reviewed head). Minors and Nit stay collected. + +**Executed:** step 8's blocks, extracted from the plan, run under sh+dash in disposable repos, 10 +cases, all as expected (`scratchpad/close-compose-test.py`). Prechecks: 54 fences balanced, sh -n 7 / +dash -n 8, same classes. **Not executed:** the plan, and the reader re-check. + +### PASS 53 — RUN, VALID, UNCLEAN. Cycle OPEN. Mandatory stop (tells 1, 2, 3). No repair, no pass 54. + +Blob `769fbc61553fcc8b1d6069c3343282d78028e35c`, HEAD `5fb3b94`, session +`01a0d94e-ecfb-7390-8f8d-b299712092dc`. 7 findings: 2 Blockers, 2 Majors, 2 Minors, 1 Nit. Blocker 1 +(4b failure route runs the candidate battery before the candidate exists) was **created by revision +24**; Blocker 2 and both Majors are pre-existing. Details and loop health: +`gate-a-plan-om0bdd7udh-pass-53-dispositions.md`. Next move is Daniel's. + +### 2026-09-25 — Daniel decided pass 52's scope question: option B, bounded. Revision 24 made. Pass 53 authorized (one pass only). + +**Decision.** Daniel released the sparring reviewer's bounded brief +(`.context/sparring/2026-09-25-154933-loop-pass52-decision-prompt.md`) after the coding agent first +proposed option A (bind via CI). The reviewer showed A wrong on three verified points: +`process-pr-review.md:183` requires the local battery AND CI; `CLAUDE.md:708` owes the battery before +Gate B; `ci.yml` checks out the PR merge result, not the branch head. **B: the battery runs on the +recorded candidate in a disposable clone.** The mandatory stop from pass 52's tells was surfaced to +Daniel and answered by this decision; it is not a waiver, and no gate duty changed. + +**Revision 24** (plan blob `15a9b1c…` → `769fbc61553fcc8b1d6069c3343282d78028e35c`): +- Blocker 1: 4b's commit block resolves the commit once (`CAND`), checks `$CAND^{tree}`, writes `$CAND` + to `.context/loop-rule-reviewed-head`; the separate capture block is deleted. +- Blocker 2: step 4's battery runs in a `--shared` clone detached at the recorded candidate; exit + 0/1/2; limits and old-condition accounting written after the block; step 7's "battery is bound" + bullet and the line-31 constraint aligned. +- Evidence (scratchpad, disposable): Blocker 1 old shape recorded a different commit after an + intervening commit, new shape recorded the checked id. Blocker 2 control passed (exit 0) on + `5fb3b94`; defective candidate + worktree-only revert: old worktree shellcheck exit 0, new block exit 1. + Plan-extracted battery block byte-equal to the tested block; both blocks `sh -n`/`dash -n` clean. + Plan: 59 fences balanced, `sh -n` 8 / `dash -n` 9 failures — same classes as before. +- Minor and Nit from pass 52 stay collected. + +**Scope of the release:** at most pass 53, then report and stop. No pass 54, no automatic repair, no +commit/push/rebase. Instruction: `.context/gate-a-plan-pass-53-instruction.md`. + +### PASS 52 — RUN, VALID, UNCLEAN. Cycle stays OPEN. No repair authorized, none made. + +Reviewed the repaired worktree plan: blob **`15a9b1c2ff4beb3c6dfa6ae21e00550210d4b969`**, `HEAD` +`5fb3b942f1907f5d6dd25b32d15b497f855e05f8`, branch `loop-rule-consolidation`, diff +675/−134 against +`HEAD`. **Identity checked before and after the call — the blob did not move.** Tool: +`mcp__codex__exec`, session `01a0b0ac-5d48-75d1-86b8-e196edda1f40`. The pass reviewed the +**twenty-second and twenty-third revisions**, which pass 51 had not seen. + +**Result VALID** — terminator `END OF FINDINGS (4 total)` exact, 4 body lines, every line a finding +line, 6 fields each, no blanks, count matches. **4 findings: 2 Blockers, 0 Majors, 1 Minor, 1 Nit.** +Per-finding verdicts in `.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-52-dispositions.md`; +**both Blockers confirmed against the plan text, the Nit confirmed by direct read, the Minor confirmed +in kind with its named sites unverified.** + +**Floor 3**, read fresh from the story header at this pass: **Risk `high`** (level 2) · **Security +`none`** (0) · max 2 ≠ 0 → 3. One cited story, from the plan's `**Story:**` header. Floor long +satisfied; what is missing remains a **clean** pass. + +**Method change, declared.** The instruction was delivered by pointing Codex at +`.context/gate-a-plan-pass-52-instruction.md` (1006 lines) instead of inlining 100 KB in the tool +argument — reproducing that verbatim into a parameter risks silent drift, the file on disk does not. +The findings-protocol contract was restated inline. **That Codex read the file in full is not +established**; the reply and the findings are consistent with it, which is evidence and not proof. + +**The two Blockers, in one line each.** Step 4b's commit block pins `$ITREE` and checks +`HEAD^{tree}` against it in one fenced block (plan 2954–2967), then a **separate** block (2987–2993) +resolves `HEAD` **again** to write `.context/loop-rule-reviewed-head` — and the plan itself states at +line 3027 that each fenced block is its own shell invocation, so nothing binds the recorded candidate +to the commit that was checked. **That is the resolve-once class of passes 47–49, reintroduced by the +twenty-second revision's own new capture block.** And step 4's battery runs over the **mutable working +tree** while step 6 checks only that `HEAD` has not moved, so a dirty helper change can make the +battery green for bytes absent from the candidate. + +**The second Blocker is the blocking decision and it is NOT the loop's to absorb.** The plan +**already discloses** that window in its own text (2977–2985) as "a limit rather than guarded". +Disclosure does not discharge a Blocker — this cycle settled that at the twentieth revision. But the +proposed fix is a **new mechanism** (a detached disposable checkout for the battery, or an +index+worktree equality re-check on both sides of it), and this cycle holds **two live rulings that +point opposite ways**: it declined new preconditions at pass 44 and at pass 49's finding 3, and it +ruled at revisions twenty and twenty-one that "a different checking mechanism is not by itself a new +requirement" — which is the ruling that put step 4c in the plan. **A new structural question stops +the loop and goes to Daniel** (§5: novelty wins over ancestry). Finding 1, by contrast, sits inside +the assigned fix set and is an ordinary repair once authorized. + +**Loop health: at least two of five tells, so stop-and-surface is mandatory** — independently of the +instruction to stop. Findings 42–52: **3, 5, 9, 2, 1, 2, 3, 6, 4, 6, 4**; Blockers **2, 2, 2, 1, 1, +1, 3, 5, 2, 2, 2**; Majors **1, 2, 7, 1, 0, 1, 0, 1, 1, 2, 0**. (1) Finding count **falling**, 6 → 4 +— not present. (2) Blocker count **flat at 2 for the third pass — failing to fall**. (3) **Instrument +cluster, total**: all four are the plan's own execution machinery; zero touch the §5 target text. +(4) A **partial prose cluster** — findings 3 and 4 are both prose promising what the command beside it +does not do; reported, not resolved, and it changes nothing. (5) **No require↔withdraw pair**. + +**The "clearly stuck" exit is still NOT available.** No plateau at six or more — the Blocker curve +over the last six is 1, 3, 5, 2, 2, 2. **No affirmative coverage judgement is possible** — finding 1 +is a defect the twenty-second revision itself created. Regeneration is present; that is one conjunct +of three, and one is not the exit. + +**Mechanical prechecks on this blob, new this session and not comparable to earlier reported +baselines** (different extractor): **60 fenced blocks, all balanced, none unclosed**; `sh -n` fails on +**8**, `dash -n` on **9**, every one an intended `<…>` placeholder or `bash`-only process +substitution — **no new syntax regression**; every cited repo path resolves; `loop-rule-baseline-diff` +`.tmp`→`.txt` is an atomic rename, not an inconsistency. + +**Nothing was committed and nothing was repaired.** `docs/superpowers/specs/2026-08-30-dark-factory-vision.md`, +`todos.md` and `docs/field-reports/2026-09-17-sfx-review-loop-economics.md` remain Daniel's edits in +flight — untouched. + +**Next:** pass 53 is **not** authorized. The next step is Daniel's decision on finding 2's scope +question. `.context/gate-a-plan-prompt.md` now carries pass 52's history entry, the twenty-second and +twenty-third revision descriptions, the updated collected list and the identity-not-cleanliness +precondition; substitute `__SHA__`/`__P__` to build the next instruction file. + ## HANDOFF — 2026-09-17, written for a session that has just lost its context -**State.** Plan anchor `1ba45be`; cycle `om0bdd7udh` **OPEN and UNCLEAN at pass 48**. The plan's -sixteenth revision repaired all three of pass 48's Blockers. **Nothing is implemented yet** — -`CLAUDE.md`, `plugins/dev-workflow/commands/workflow-init.md` and the hook are untouched. +**State.** Plan anchor `1ba45be`; cycle `om0bdd7udh` **OPEN and UNCLEAN at pass 49**. The plan's +sixteenth revision repaired all three of pass 48's Blockers; **pass 49 found six more — 5 Blockers +and 1 Major — and none is repaired.** **Nothing is implemented yet** — `CLAUDE.md`, +`plugins/dev-workflow/commands/workflow-init.md` and the hook are untouched. + +**Pass 49 ran full, on Daniel's reviewer's written authorization of 2026-09-17** (exactly one full +Gate-A plan pass, four areas prioritized, no part of the artifact excluded; no repair round, no +implementation, no pass 50). Reviewed revision: the plan at `1ba45be`, blob +`bc685751b971c6708fc7abf133d4e73e530ec9a5`, **identical at `1ba45be`, at `HEAD` `5fb3b94` and in the +worktree** — the intervening commit touches only this record. Repo read at `HEAD` `5fb3b94`. Result +**VALID** under the findings protocol: terminator `END OF FINDINGS (6 total)` exact, 6 body lines, +every line a finding line, 6 fields each, no blanks. File +`.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-49.md` (untracked on disk, not committed). + +**All six were validated against the plan text before being reported** — five read-only subagents, +one per Blocker, plus execution for the Major. Verdicts: findings 1, 2, 4 **confirmed**; findings 3 +and 5 **partially confirmed**; finding 6 **confirmed by execution**. Details in the pass-49 rows +below. + +**The one structural fact pass 49 establishes: the resolve-once sweep is not complete.** Five of the +six are the same defect class the sixteenth revision swept — a movable ref re-resolved as the +*identity* of an object whose properties were already checked — at five sites that sweep did not +reach. Two were simply missed (Preparation/Task 0 step 1; step 6's call issuance, which is prose +describing an MCP parameter and therefore **structurally invisible** to a shellcheck-verified sweep). +One was **explicitly excluded with a rationale that does not hold for it**: `1ba45be` wrote +"conditions 1, 5 and 6 read `HEAD` on purpose, to detect a move, and stay as they are" — true of 1 +and 5, which each compare one read against one recorded baseline, but condition 6 has **no baseline +at all** and composes subject, parent, tree and body from four separate reads. **Not yours to touch:** `docs/superpowers/specs/2026-08-30-dark-factory-vision.md`, `todos.md` and `docs/field-reports/2026-09-17-sfx-review-loop-economics.md` are Daniel's own edits in flight. Leave them dirty; do not commit, revert or review them. -**The open decision, and it is one word from Daniel.** Three tells stand at pass 48 (findings 2 → 3, -Blockers 1 → 3, instrument cluster), so the surface is mandatory. The choice he was given: - -- **"fokussiert weiter"** → run pass 49 with the prompt narrowed to the four areas that can reach a - shipped product (below); -- **"voll weiter"** → run pass 49 over the whole plan, as passes 1–48 did. - -**Why that choice exists — the efficiency finding, counted rather than felt.** Of roughly 25 findings -across passes 42–48, **about 4 or 5 had a payer outside this repository**; the rest cost the one -person executing the plan a bad session. **That is ~80% bookkeeping** — but the other 20% are -serious: unreviewed content published through `reset --soft`, a plugin shipping without its version -bump, a hook change slipping past the invariant-4 review. The loop is **misallocated, not worthless**. -Of the plan's **54 shell blocks**, only four areas decide what is published: **8a/8b** (the closing +**That decision is spent: Daniel chose "voll", pass 49 ran full, and its authorization is exhausted.** +No repair round, no pass 50, no implementation is authorized by it. + +**The six open findings, with the validation verdict on each.** All must resolve before this cycle +can close — §5's Blocker/Major rule is not waived by the utility test. + +1. **BLOCKER — Preparation and Task 0 step 1. CONFIRMED.** Preparation checks branch, ancestry and + the three approved-input blobs through six separate live `HEAD` resolutions; Task 0 step 1 then + writes `.context/loop-rule-base` from a **seventh**, independent one, so the recorded base is + bound to no commit whose properties were checked. Each fenced block is its own shell invocation, + so the two can be arbitrarily far apart. **Not in either of `1ba45be`'s lists** — neither repaired + nor deliberately left; simply missed. No existing guard, no disclosure. +2. **BLOCKER — Task 15 step 6, call issuance. CONFIRMED.** Step 6 records the reviewed head to + `.context/loop-rule-reviewed-head`, then the call's `headSha` is specified in **prose** as "the + full 40-character object name `HEAD` resolves to at that moment" — never wired to the file just + written. Close condition 1 later compares `HEAD` against the **file**, never against what was + actually sent, and the review tool reports no reviewed revision, so the divergence is capturable + nowhere. **CLAUDE.md's own `headSha` rule is satisfied**; the stronger demand comes from the + plan's own logic (self-review item 28). **The sweep could not have caught this** — it is not shell. +3. **BLOCKER — Task 15 step 8b, Close condition 5. PARTIALLY CONFIRMED, and the weakest of the six.** + Condition 5 already closes the **wide** window: anything committed between 8a and 8b fails the + `HEAD == TIP` check, anything merely staged fails the clean-tree check. What remains is a + **sub-second TOCTOU inside 8b's own script**, between its own checks and its own `reset --soft`, + requiring a **second concurrent actor** on a normally sequential single-operator session. Git + offers no expected-old-object guard on `reset`; `update-ref` appears nowhere in the plan. "Silently + publish" is true of the script but **not** of the documentation: target text §I names this gap + loudly as parked on Daniel's decision of 2026-09-13. +4. **BLOCKER — Task 15 step 8, Close condition 6 and cleanup. CONFIRMED.** After the closing commit + lands, subject, parent, tree state and body are read through **four** independent live `HEAD` + resolutions with **no captured commit id**; the parent check asserts only `HEAD^ == $BASE`, true of + *any* commit parented by the base. Cleanup then deletes every recovery artifact. **Explicitly + excluded by `1ba45be` on a rationale valid for conditions 1 and 5 and not for 6** — those compare + one read against one recorded baseline; condition 6 has none and composes four properties. +5. **BLOCKER — the records-commit marker's recovery rule. PARTIALLY CONFIRMED, two halves.** + *Confirmed:* the five checks `1ba45be` added (40-char id · resolves to a commit · `ba15e83` an + ancestor · itself an ancestor of `HEAD` · both slot paths blobs **in that commit**) are satisfied + **trivially by any earlier already-routed records commit** on a linear branch — nothing pins "the + current pass", and no scan for a later competing records commit exists. Also confirmed: the generic + "delete and rebuild from the `$BASE` blobs" rule is a **category mismatch** — this marker is not a + function of `$BASE`'s content. *Blunted:* for the case the plan **does** name — valid marker, outcome + not shown — there is a real executable precondition stop (report, no mutation, no call). The + scenario nobody checks for is a stale-but-valid marker coexisting with a later unmarked records + commit at `HEAD`. +6. **MAJOR — Task 7, the worked carried-fragment example. CONFIRMED BY EXECUTION.** The plan runs + `grep -cF 'Codex is advisory — validate before applying; dismissed finding → one-line why'` and + states "Expected: `1` each in the worktree, and `parent=1 worktree=1` in each copy". In **both** + prompt copies the text wraps between `Codex is` and `advisory` (`CLAUDE.md:135-136`, + `plugins/dev-workflow/commands/workflow-init.md:342-343`), so the literal count is **0**. A + **correct** source text fails its own preservation check — a false red, which the severity + procedure's symmetric instrument carve-out keeps at Major. + +**The 80/20 claim is withdrawn as evidence.** Daniel's reviewer ruled it **unverified classification, +not a measurement**, and it **must not determine exclusions or severity**. What *is* counted, from the +pass files rather than from memory: passes 42–48 hold **exactly 25 findings — 12 Blockers, 12 Majors, +1 Minor** (`grep -cE '^(BLOCKER|MAJOR|MINOR|NIT) \|'` per file). The payer split is a judgement made +per finding under the utility test, and for **all six of pass 49's the payer is the one person +executing this plan at a terminal** — no user of the shipped plugin hits any of them. The closest to +a shipped consequence is finding 2: if it fires, *this change's own* edits to `CLAUDE.md`, +`workflow-init.md` and the hook could land without full Gate-B coverage. **The utility test waives no +floor, no mandatory tell and no closure condition**, so none of this makes the cycle closable. + +**Of the plan's 54 shell blocks, four areas decide what is published** — **8a/8b** (the closing commit), **steps 3–4** (version bump and battery, invariant 12), **step 6** (the Gate-B range), and -**Tasks 10/11** (the hook change and its invariant-4 review). +**Tasks 10/11** (the hook change and its invariant-4 review). Pass 49 was prioritized on these and +excluded nothing; five of its six findings landed inside them. + +**Loop health at pass 49 — three of the five tells stand, so the stop is mandatory, not +discretionary.** Trend across passes 42–49: findings **3, 5, 9, 2, 1, 2, 3, 6**; Blockers **2, 2, 2, +1, 1, 1, 3, 5**; Majors **1, 2, 7, 1, 0, 1, 0, 1**. (1) Finding count **rising**, 3 → 6. (2) Blocker +count **failing to fall**, 3 → 5. (3) **Instrument cluster** for the twelfth pass — all six are the +plan's own execution machinery; **zero** touch the §5 target text the change installs, and zero are +prose about either. Not present: a prose cluster, and **no require↔withdraw pair** — the near-miss +worth naming is that `1ba45be` declared conditions 1/5/6 "stay as they are" and pass 49 demands 6 +change, which is a pass challenging a stated **non-change**, not a demand for something an earlier +pass removed. + +**The "clearly stuck" exit is NOT available, and all three conjuncts fail.** No plateau — the Blocker +curve is **rising**, not flat. No affirmative coverage judgement is possible: the sixteenth revision's +sweep was explicitly bounded, and pass 49 finding five more sites of the same class is **direct +evidence that coverage is insufficient**. And the findings are **newly discovered at previously +unswept sites**, not regenerated from the repairs — pass 49 re-raised **nothing** against the three +blocks `1ba45be` actually repaired. + +**The severity procedure's unsettled question does not bite here, and this was checked rather than +assumed.** All six name an operational consumer (the executor acting on the plan) and a decision that +changes (which commit becomes base, reviewed head, or closing tip). The instrument carve-out is +symmetric, so findings 1/2/4 keep severity as **false greens** on a gate and finding 6 as a **false +red**. Nothing is demoted, so the per-pass counts and clusters above are unaffected by the open +question in `docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md`'s sibling +consolidation work. **Standing orders now in force** (also in the session memory): - **Subagents, parallel, by default.** Fan out read-only inspection; keep writing serial — one writer. @@ -54,6 +340,565 @@ commit), **steps 3–4** (version bump and battery, invariant 12), **step 6** (t condition** — `todos.md` "Review-loop usefulness" (Dark Factory vision §§4/7/11) is the formal version and activates no thresholds. +**WHERE THE SCOPE NOW STANDS — 2026-09-17, after Daniel's reviewer assessed pass 49.** The stop was +upheld; a blanket "repair all six, then pass 50" was **declined**. Two artifacts carry the result: + +- `.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-49-dispositions.md` — verdict + reason per + finding, and the three corrections to my pass-49 status report. +- `.context/plan-drafts/pass-49-repair-draft.md` — the coherent draft for findings **1, 2, 4, 5** + along the four axes (check subject · stored identity · consumer · recovery), finding **6** as a + separate one-line fix, and finding **3** dispositioned rather than repaired. **Applied: nothing.** + +**Three corrections to the pass-49 report, all validated before acceptance — do not re-adopt the +originals.** (1) "All six are paid for by the executor alone" is **wrong**: the standard is the +causal chain, and `workflow-init.md`, `codex-gate.sh` and `codex-gate.test.sh` ship inside the plugin +package, so findings 1, 2 and 4 can put unreviewed content into shipped files. (2) "No regeneration +from the repairs" is **too sweeping**: `1ba45be` has seven hunks across **four** regions, not three, +and the first is Resume's marker validation — the very thing finding 5 attacks, so finding 5 **is** +regeneration. The "clearly stuck" exit stays unavailable all the same: Blockers are rising and no +coverage judgement is possible. (3) The claim that the utility test was missing from the instruction +file is **false** — it is at line 572. The other half of that correction is right and is **my +defect**: the four priority areas appear **nowhere** in the instruction file, and my wrapper asserted +the prompt named them when it did not. **Pass 49 is therefore not a usable test of the sharpened +approach**, and any pass 50 must carry the four priorities in the prompt file itself. + +**A SEVENTEENTH BOUNDED REVISION was made on 2026-09-17 and all six findings are repaired in the +worktree. Nothing is committed; the plan is dirty.** Daniel released the scope with "weiter" and +instructed that questions like finding 3's go **through a gate, not to him as a menu**. + +- **Finding 1** — Preparation captures the starting revision **once**, runs branch, ancestry and the + three approved-blob predicates against that captured object, and records it to a new file + `.context/loop-rule-start`. Task 0 step 1 reads that file, **re-asserts `HEAD` still equals it + immediately before the first mutation**, records it as the base and retires the start file. The new + state file **ships with its recovery rule in the same change** — surviving start with no base is an + interrupted first entry and a precondition stop, never rebuilt, never resumed from — because a + state file introduced without one is precisely the defect pass 48 found in the revision before. +- **Finding 2** — the call's `headSha` is now **the exact content of `.context/loop-rule-reviewed-head`**, + quoted as a literal, never a fresh resolution. The repair is prose because **there is no command to + guard** — the value goes into an MCP argument — so the residual is stated: nothing mechanically + compares the argument sent against the file, and Close condition 1 authenticates the file. +- **Finding 3** — settled **through a Codex gate**. Verdict **NOT OWED**: condition 5 scopes its claim + to "when the closing invocation begins", nothing stated is violated, the wide window is already + closed, and an atomic guard means replacing `reset --soft` with `update-ref` plus separate index + handling — a different close mechanism, hence a new precondition, which this plan declines on the + same ground as pass 44's three escalations. Implemented as a **disclosed residual on condition 5**, + which names the unguarded span, says its width is unmeasured, and states that no later check would + catch it because the content comparison is the parked tree-equality condition. +- **Finding 4** — condition 6 captures **one object** and addresses subject, parent and body to it; + the tree check and the **cleanup are gated on a `HEAD = $CLOSED` re-assert**, so the cleanup gate is + a command instead of the sentence "all four pass, and only then". The cleanup block is merged into + condition 6's block; `1ba45be`'s claim that conditions 1, 5 and 6 alike "read `HEAD` on purpose" is + corrected in the same change. **Say what this is, not more: the four predicates now describe ONE + commit — it is not established that this commit is the one 8b created.** The plan discloses that + limit in its own text, and a summary reading "captures the closing commit" would hide it. +- **Finding 5** — the marker must now be **the newest records commit reachable from `HEAD`**, which is + what separates a stale-but-valid marker from the current pass; three outcomes are named separately + (ordinary · stale marker · marker absent with a records commit present), and `$NEWEST` empty with + status 0 is read as a real result. The marker is **carved out of the generic "rebuild from the + `$BASE` blobs" rule by name** — it is not a function of `$BASE`'s content and cannot be rebuilt at + all. **Simpler than the draft proposed**: no record format change, so step two's consumer is untouched. +- **Finding 6** — the fragment drops its leading `Codex is `, measured at 1/1 in both copies, and the + plan now states that a fragment is line-local by construction and is verified **before** it is + written down. + +**Finding 5 was corrected a second time, on Daniel's reviewer's report, and the defect was +reproduced before it was repaired.** The first repair asked +`git rev-list -n 1 --grep='^WIP: pass [0-9][0-9]* records$' HEAD` — **unbounded**: it searched the +whole reachable history for a generic subject, so a records commit from **another cycle** could be +taken for this cycle's newest and make a **perfectly valid marker read as stale**. Reproduced in a +disposable repository. The shipped check is now bounded **twice**, and **both bounds are needed** — +measured, not assumed: `$BASE..HEAD` alone still lets a foreign commit inside the range win; it is +the **nonce-scoped slot pathspec** that makes the answer this cycle's. + +**Verified by execution: 21 fixtures × `sh`/`dash`/`bash`, all green** — 10 for finding 1, 6 for +finding 4, 5 for finding 5 — blocks extracted verbatim, each finding carrying a control that +reproduces the old defect, including a foreign-cycle counter-case. `sh -n`/`dash -n` failures are +**6 and 7, identical to the pre-repair baseline**. **Three first-run fixture failures were all harness +bugs** (no `.gitignore` in the disposable repo; one miscalculated control assertion), recorded in the +dispositions file. + +**Two corrections to my own verification, both found by running it rather than reading it.** +(1) **"Fenced blocks 55 → 54" was wrong.** My extractor only matched fences at column 0, so it never +saw the new **indented** marker block — which also means that block went unchecked in the first +round. With the corrected extractor the count is **55 before and 55 after** (one top-level block +merged away, one indented block added), and the new block parses under all three shells. +(2) **The fixture results are self-reported.** Daniel's reviewer has not executed them and says so; +they are evidence about the cases they cover and **no claim of completeness**. + +**What is NOT established.** No sweep was run for a **seventh** site of the resolve-once class — +pass 49 found five after a sweep that believed itself complete, and nothing here rules out another. +The repairs are verified against fixtures, not reviewed. Resume's new marker branches and the +`loop-rule-start` recovery rule are **reader text, asserted as text and not executed**. + +### TWENTY-THIRD BOUNDED REVISION — 2026-09-17. The rerun path reconciled. + +The twenty-second revision fixed the first-pass order and **left the rerun contradicting itself**. +Two defects, both confirmed in the text before editing: + +- My rerun bullet said *"no commit is made between the candidate head being written and the call"*, + while **the very next bullet** still said *"Commit those records, resolve the new `HEAD`, and only + then issue the candidate final pass."* Directly contradictory. +- The rerun put **step 4b's reader checks after the capture**, although their result is a record that + must be committed. + +**The reconciled sequence, now written as six numbered steps and identical on both paths:** +1 repair and produce every record that must be committed · 2 **commit them** (step 7 step three's +commit — the last before the call) · 3 **capture the candidate** into +`.context/loop-rule-reviewed-head` · 4 final verification that writes no committed record — the +battery, then **4c**; non-zero or unresolved returns to 1 · 5 write the evidence entry into the +gitignored `closing-msg` · 6 issue the call with the captured head and that entry verbatim. + +**The recording duty kept its home rather than being deleted with the bullet.** The conflicting +bullet's two claims were split: its *recording* duty moved into step 1, and its *"only a clean +response against that exact `HEAD` closes the cycle"* rule is restated where the ordering now makes +it true. The complete-rerun list gained an explicit split — **record-producing work** (repair, 4b's +twelve items, every mechanical observation) versus **final verification** (battery, 4c) — which is +what orders the whole step. + +**A second correction, and it was an overclaim of mine.** I had written that the battery reads the +candidate identity. **It does not.** Step 4 reads `.context/loop-rule-baseref` and runs shellcheck, +the hook suites, the invariant checkers and `claude plugin validate` **over the working tree of the +current repository**; it takes no candidate id. Both sites now say so, and state the real binding: +the battery describes the candidate only because it runs immediately after the capture with no commit +in between, and **nothing enforces that the worktree is unmodified across that window** — step 6 +catches a moved `HEAD`, not a dirty tree. **A stated limit, not a guard**; no index or concurrency +policy was invented. + +**Walked, both paths.** *First pass:* 1–3 → 4b (items + repairs + records, its commit) → capture +(`:2991`) → battery → 4c → entry → call. *Rerun with a repair to a shipped file:* step one's records +commit → step two routes → step three stages the repair **and** the plan, commits (block line 31), +then captures `NEXTHEAD` (line 71) → battery → 4c → entry → call. *Rerun with no repair:* step three +commits nothing, `NEXTHEAD` = the records commit step one validated, capture → battery → 4c → entry → +call. **Exactly two writers of `reviewed-head` exist** (`:2991`, `:3576`), and in both the commit +precedes the capture. **No instruction after the capture requires another commit before the call.** + +**Validation limits.** The walk is a reading of the text; the shell blocks it names were re-run — +29 existing fixtures plus the 4c cases — all green under `sh`, `dash` and `bash`, with `sh -n`/`dash +-n` at **6 and 7, the baseline**. **No fixture exercises the six-step sequence end to end**; that +would need the whole of Task 15 executed, which is implementation. Fixture results remain +self-reported. + +**Minor and Nit remain collected.** No findings are claimed resolved and the cycle is not clean. + +### TWENTY-SECOND BOUNDED REVISION — 2026-09-17. One candidate flow for Task 15. + +**My own proposal was wrong and the reviewer took it apart correctly.** I offered a choice between +binding 4c late (in 7b) and binding it early without a record. **Option A was self-defeating**, and +two citations settle it: step 6 hands the reviewer *"the evidence entry quoted verbatim"* (`:3202`) +and 7b says that if that entry changes at revalidation **the candidate is over** (`:3640`). A binding +4c result first appearing at 7b therefore changes the entry the final reviewer judged and forces the +very next round it was meant to prevent. My claim that the early option means "a failure only shows +at CI" was also **false**: a binding pre-Gate-B 4c stops before Gate B. And no new carrier was ever +needed — step 5 already writes the entry into `.context/loop-rule-closing-msg`, which is +**gitignored**, so no pre-pass commit touches it. + +**What was implemented — the reviewer's sequence, unchanged:** + +1. **Step 4b moved to before step 4.** It keeps its letter (six places cite "step 4b"; a rename would + have to reach all of them) and the block carries a note that this file executes in **reading** + order. It is the last step that may alter content. +2. **One candidate identity, created once, at the end of 4b** — written to + `.context/loop-rule-reviewed-head`, the file that already means this. **No new state file, no new + recovery rule, no new cleanup entry.** +3. **Battery and 4c run against that identity**, read from the file. Not verified → stop, no entry, + no call. +4. **Step 5 writes the evidence entry including 4c's three ids, plugin diff and raw status**, into + the existing closing-msg, uncommitted. +5. **Step 6 resolves nothing.** It reads the candidate head, refuses if `HEAD` has moved since, and + passes that value as `headSha` with the entry verbatim. +6. **7b names how it preserves the entry** — the open design point. It reads the existing entry out + of closing-msg and **carries item 4 across** while rebuilding items 1–3; a missing entry is a + **stop**, not a rebuild, because the 4c result is read from a run and cannot be regenerated from + the record sections. The existing re-review rule for a changed entry is untouched. +7. **Merge direction reversed** — check out the base, merge the candidate in, the order + `refs/pull/N/merge` is built in. The intended checker is still pinned **from the candidate** by + object id before anything merges. +8. **Step 7's rerun bullet now states the same sequence as the first pass**, and that **no commit is + made between fixing the candidate head and issuing the call**. + +**Verified by execution — 8 cases for 4c after the direction and identity changes, plus the 29 +existing fixtures, all green under `sh`, `dash` and `bash`.** Added case: an unresolvable candidate +head read from the file → unresolved. `sh -n`/`dash -n` remain **6 and 7 — the baseline**. Step 6 +contains **zero** fresh resolutions of `HEAD` into the reviewed-head file. **The direction change +moved no verdict in these fixtures** — none uses a merge driver — so it is adopted because it matches +how CI builds the object, not on fixture evidence. + +**Not claimed:** that pass 51's findings are resolved, or that the cycle is clean. **The Minor +(unguarded plugin-diff pipeline) and the Nit (no `mktemp` cleanup) remain collected and unrepaired**, +as instructed — both still sit in the 4c block. + +### PASS 51 — RUN, VALID, UNCLEAN. Cycle stays OPEN. No repair authorized, none made. + +Reviewed the frozen candidate: plan blob **`974e223ed71d52b49c6b368695372ccb72e510f1`**, `HEAD` +`5fb3b942f1907f5d6dd25b32d15b497f855e05f8`, branch `loop-rule-consolidation`, diff +482/−64. +**Identity checked before and after the call — the blob did not move.** Tool: `mcp__codex__exec`; +the prompt opens with the required brainstorming line and names the four priorities plus the seven +4c-specific questions. + +**Result VALID** — terminator exact, 6 body lines, every line a finding line, 6 fields each, no +blanks, count matches. **6 findings: 2 Blockers, 2 Majors, 1 Minor, 1 Nit.** Per-finding verdicts and +the split between reviewer claim and my own verification are in +`.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-51-dispositions.md`; **all six confirmed**. + +**Floor 3**, derived from `docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md` read +fresh at this pass: **Risk `high`** (level 2) · **Security `none`** (0) · max 2 ≠ 0 → 3. One cited +story, profile resolves. Floor long satisfied; what is missing remains a **clean** pass. + +**The two Blockers, in one line each.** Step 4b **commits after** the battery and 4c have run, and +the first-pass path runs no complete rerun before step 6 — so Gate B can review and close a head that +never got either check, and a 4b repair to a shipped prompt or hook is absent from 4c's merge result +entirely. And 4c's `HEADID` is an operator placeholder while step 6 independently resolves live +`HEAD`: **nothing ties them**, so 4c can certify one candidate while Gate B reviews another. + +**Four of the six are in machinery I added in the last two revisions, and two of those are defect +classes this cycle had already swept.** Finding 5 is an unguarded pipeline whose status a following +`tr` consumes — the class of passes 43 and 44. Finding 6 is new state with no cleanup rule — the +class of pass 48. **Writing new code reintroduced both.** That is regeneration from the repairs, not +discovery in unswept ground, and the record says so rather than presenting it as fresh coverage. + +**Loop health: three of five tells, so stop-and-surface is mandatory** — independently of the +instruction to stop. Findings 42–51: **3, 5, 9, 2, 1, 2, 3, 6, 4, 6**; Blockers **2, 2, 2, 1, 1, 1, +3, 5, 2, 2**; Majors **1, 2, 7, 1, 0, 1, 0, 1, 1, 2**. (1) Finding count **rising**, 4 → 6. (2) +Blocker count **flat at 2 — failing to fall**. (3) **Instrument cluster**, and this time +concentrated: five of six land on step 4c and its interface. Not present: a prose cluster; and **no +require↔withdraw pair** — finding 4 asks to reverse a merge direction I chose freely, which is a +correction rather than a reversal of something an earlier pass removed. + +**The "clearly stuck" exit is still NOT available**, and this pass does not change that: the Blocker +curve is 3, 5, 2, 2 over the last four — no plateau at six or more passes — and **no affirmative +coverage judgement is possible**, since the last two revisions demonstrably added new unswept +machinery. Regeneration is now present, which it was not at pass 50; that is one conjunct of three, +and one is not the exit. + +**Nothing here closes the cycle**, and a clean pass would not close it by itself either: closure +needs the clean pass **plus** every other duty §5 names. + +### TWENTY-FIRST BOUNDED REVISION — 2026-09-17. Three corrections to step 4c itself. + +All three were reported as static findings against the new procedure, and **all three validated**. + +**1. The outcome separation did not hold.** `scripts/check-version-bump.sh` exits `1` from **both** +`fail()` (a policy violation, `rc=1`) and `die()` (an operational failure — unresolvable ref, failed +git call, unparseable manifest). Verified at lines 68 and 73. My 4c read every `1` as a version +violation. **Fixed without a parser and without touching the checker:** the output and the **raw** +status are preserved, non-zero is reported as **NOT VERIFIED with the cause undetermined**, and the +cause must be established from that output before anyone calls it a version violation. Both cases +stop progression, so nothing depends on guessing. + +**2. The byte check ran before the merge — and its reference was wrong too.** Comparing before the +merge establishes nothing about the bytes that then execute. **But fixing only the ordering was not +enough, and the fixture caught it:** the comparison read the checker from the source repository's +**worktree**, which can already sit on the base side and carry the very edit the check exists to +catch — so it agreed with itself and the modified checker ran. It now pins +`git show "$HEADID:scripts/check-version-bump.sh"` **before** merging and compares the merge result +against that, **after** the merge. A green check that could not go red is not evidence; this one +now goes red. + +**3. The integration contradicted the addition.** Step 7 still said checking the merge result was "a +scope question and not a change this plan makes" — now retracted **in place**, with the reason (a +different checking mechanism is not by itself a new requirement), and the only surviving occurrence +of that phrase is the quotation inside the retraction. The complete rerun now **names 4c +explicitly**, places it **last**, and says a non-zero or unresolved 4c stops progression. + +**The recording cycle is closed by ordering, not by machinery.** 4c runs *after* the rerun's records +are committed and the new `HEAD` is resolved, **against that head** — the one the candidate pass is +issued against — and **its own result goes into the closing commit body**, not into a further +pre-pass commit. Any scheme that committed 4c's result before the pass would move the head 4c had +just certified and demand another run, forever. + +**Validation — the block as extracted from the plan, eight cases, expected vs observed:** + +| case | want | got | +|---|---|---| +| valid version bump | 0 | **0** | +| same version + plugin diff | 1 | **1**, policy cause present in the preserved output | +| operational failure **inside** the checker | 1 | **1**, **not** labelled a version violation | +| conflict-free merge changing the checker's bytes | 2 | **2**, refused **before** the checker ran | +| merge conflict | 2 | **2** | +| unresolvable pinned head | 2 | **2** | +| no recorded base ref | 2 | **2** | +| source fixture repository afterwards | unchanged | **branch unchanged, 0 staged, 0 dirty** | + +The byte-change case **failed on the first run** — exit 0, checker executed — which is what exposed +the worktree-reference defect. It is recorded rather than quietly fixed. + +**Remaining limitations.** Fixture results are self-reported; the reviewer has not reproduced them. +**No claim of complete error classification is made** — 4c distinguishes *verified* from *not +verified* from *unresolved*, and deliberately does **not** classify the cause of a non-zero checker +status. The other standing limits are unchanged: Resume's three marker outcomes and the +`loop-rule-start` recovery rule text remain reader instructions, and no sweep establishes the +resolve-once class is discharged everywhere. + +### TWENTIETH BOUNDED REVISION — 2026-09-17. Step 4c: the merge-result version check. + +Scope from Daniel's reviewer, and it is narrow: **add one bounded, isolated check of the version +requirement against the merge result of a pinned head and a pinned base, validate it with the +unchanged checker, dispose the pass-50 version finding on evidence, report, stop.** No CI change, no +version policy, no pass 51, no implementation. + +**The reasoning error it corrects is mine.** I wrote that full closure "would be a scope decision, +not a repair". **That does not follow** — the same correction the reviewer made about the atomicity +case applies here: *a different checking mechanism is not automatically a new requirement*. An +accepted Blocker is not discharged by disclosing it, and `design.md` §7 — which requires the battery +green **at the Gate-B WIP commit** and promises nothing about a later merge commit — neither grants a +CI guarantee nor excuses the version duty. Verified at `design.md:215`. + +**What was added: Task 15 step 4c.** It clones the work repo `--shared --no-checkout` into a +disposable directory, checks out the **pinned implementation head**, merges the **pinned base** step 4 +already recorded, and runs the **unchanged** `scripts/check-version-bump.sh` against that merge +result. **The existing battery and the Gate-B review range are untouched.** + +**Three outcomes, deliberately distinct** — and this separation is itself a repair: `0` the rule +holds · `1` the rule is **violated** on the merge result · `2` **UNRESOLVED**, the check could not be +carried out. A single non-zero code would have made a merge conflict indistinguishable from a +rejection, so an unresolved check could have been filed as a failure — or, worse, a passing run +inferred from "no rejection". + +**Two traps the step closes by construction, both learned the hard way in this cycle.** The checker's +line 58 is `cd "$(dirname "$0")/.."`, so it is invoked **relatively from inside the clone** — an +absolute-path call silently runs it against its own source tree, which is exactly how my earlier +"could not be reproduced" was manufactured. And its bytes are compared against the source checker, so +a modified checker at the pinned head is UNRESOLVED rather than trusted. + +**Verified by execution against the plan's own extracted block — 7 cases × the real checker:** + +| case | expected | observed | +|---|---|---| +| same version both sides + plugin diff, **current** base | reject | **exit 1** | +| same pair against a **stale** base | false green | **exit 0** | +| genuinely bumped version | pass | **exit 0** | +| merge conflict | unresolved | **exit 2** | +| unresolvable pinned head | unresolved | **exit 2** | +| no recorded base ref | unresolved | **exit 2** | +| checker bytes differ at the pinned head | unresolved | **exit 2** | + +The work repo was checked after the run: **still on its branch, zero dirty lines** — the clone +borrows objects and writes nothing back. The rejection message names the fixture's own base commit, +which is the observable evidence that the checker read the fixture and not its source tree. + +**Finding disposition — pass-50 Blocker 2 (version check): RESOLVED BY ADDED VERIFICATION**, not by +disclosure. The failure chain is demonstrated, and the plan now performs a check that catches it for +the pinned pair. **It is not closed in general**, and the plan says so in its own text: a `main` that +advances after this check is a merge result that does not exist yet, and one script against one merge +result is not a check of CI as a whole. + +**One more self-inflicted syntax regression, caught by the check and fixed.** The new block opened +with a bare `HEADID=<…>`, which is a redirect, not an assignment — `sh -n` went 6 → 7. It is quoted +now and the baseline is restored. **This is the second time the same placeholder trap bit me in one +session**, which is worth recording as a pattern rather than an incident. + +**Remaining limits.** Fixture results are self-reported; Daniel's reviewer has not reproduced them. +Resume's three marker outcomes and the `loop-rule-start` recovery *rule text* remain reader +instructions asserted as text. No sweep has established that the resolve-once class is discharged +everywhere. + +### NINETEENTH BOUNDED REVISION — 2026-09-17. Two corrections, both against my own last report. + +Scope from Daniel's reviewer: **clean up the remaining no-repair contradiction, and put the version +case to the real checker with documented ids.** No CI change, no version policy, no pass 51. + +**Candidate state:** `HEAD` `5fb3b942f1907f5d6dd25b32d15b497f855e05f8` (unmoved), plan blob +**`17bb2a85ec90ded1aedb81c81396380de0c49abb`**, diff **+344 / −64**, 55 fenced blocks, +`sh -n`/`dash -n` **6 and 7 — baseline**. 29 fixtures × three shells green. Product files untouched. + +**1. The no-repair contradiction was still there. My claim that it was rewritten was false.** I had +rewritten the step-three *intro*; the paragraph that actually contradicted the new branch stood +untouched, asserting *"A non-closing pass that owes no repair still commits"* and *"There is no +'nothing to commit' branch here: the findings files are always something."* Both phrases now occur +**zero** times. The replacement states the true sequence: **step one commits the records**, so by +step three those files are tracked and clean, the pile-up the old paragraph feared cannot happen, and +the no-repair route is the ordinary one. + +**2. THE VERSION BLOCKER IS REAL. My "could not be reproduced" was an artefact of a broken harness, +and the reviewer's derivation is confirmed by execution.** + +`scripts/check-version-bump.sh` line 58 is `cd "$(dirname "$0")/.."`. Invoking it by **absolute +path** from a disposable repository therefore ran it **against this repository instead** — the +checker under test never saw the fixture, which is why it answered `ok` to everything and could not +resolve the fixture's refs. That is exactly the "wired so it could not fail" defect this plan warns +about, and it produced a confident negative. The corrected run copies the checker **into** the +fixture. + +Topology and results, with ids, as instructed: + +| | commit | version | +|---|---|---| +| common ancestor | `47c6193` | 0.11.0 | +| `main` advances alone | `c0df286` | 0.12.0 | +| branch head | `417e5e3` | 0.12.0 | +| **R, the merge commit CI checks out** | `44a2698` | 0.12.0 | + +Plugin diff `main`..`R`: `plugins/dev-workflow/extra.md`, `plugins/dev-workflow/file.md`. + +- against the **stale** base `47c6193` (0.11.0 ≠ 0.12.0) → **`ok`, exit 0** +- against the **current** base `c0df286` (0.12.0 = 0.12.0, plugin changed) → **rejected, exit 1** + +**And `.github/workflows/ci.yml` passes no `ref:` to `actions/checkout`**, so a `pull_request` job +checks the **merge commit**, not the branch head — verified by reading the workflow. + +**The proposed repair is insufficient, and the plan now says so instead of claiming a fix.** Naming +the fetch in the complete rerun is a precision improvement, but the battery runs against the **local +branch head** while CI evaluates the **merge commit**; no refresh of the base ref makes those the +same object. Closing it would mean checking the merge result — **a scope question, not a change this +plan makes**. Recorded as a disclosed gap. + +**What this says about my verification generally.** Two rounds in a row, the harness was the thing +that was wrong — first an extractor blind to indented fences, now a checker silently rebased onto its +own repository. **A green result is evidence only once the wiring has been shown capable of +producing a red one.** The version fixture now demonstrates both outcomes, which is why it can be +believed. + +### EIGHTEENTH BOUNDED REVISION — 2026-09-17, after pass 50. Cycle still OPEN. + +Scope was set by Daniel's reviewer and is deliberately narrow: **fix the two confirmed flow defects, +prove-or-dispose the version Blocker, collect the Minor, then report and stop. No pass 51.** + +**Candidate state:** `HEAD` `5fb3b942f1907f5d6dd25b32d15b497f855e05f8` (unmoved), plan blob +**`b1978d77396451781ad18db6ecb372e43c08fef3`**, diff **+315 / −59**, 55 fenced blocks, +`sh -n`/`dash -n` failures **6 and 7 — the pre-repair baseline**. Product files untouched. + +**Repaired, each with a counter-case and an unchanged normal case:** + +- **No-repair continue route (pass-50 Blocker 1).** Step three now branches on `REPAIR_OWED`, the + route step two established — **never on what the tree looks like**. Repair owed → stage, guarded + commit, pin check, head = the repair commit. **None owed → commit nothing** (step one already + committed the findings files, so no tracked change remains) and head = **the records commit step + one validated**, with a guard that `HEAD` still is it. The contradictory prose claiming those files + were still dirty is rewritten. +- **Start-file router (pass-50 Major — my own defect from the seventeenth revision).** Task 0 step 1 + now has **three** outcomes, not two: base present → Resume; **start present without base → + precondition stop, reported, file preserved**; neither → first entry. The recovery rule I added was + unreachable because the router tested only the base and Preparation overwrites the start file. + +**Version Blocker (pass-50 Blocker 2) — DISPOSED, not repaired, and the disposal is measured.** The +claim was that a stale base flips `check-version-bump.sh` from fail to pass. **Put to the real +checker in a disposable repository, it could not be reproduced.** The checker compares the +**merge-base** of the given ref with `HEAD`, and an advancing `origin/main` does not move it: +unrelated commits on `main`, `main` bumping the same plugin to the same version, and `main` merged +into the branch **all left the result at `ok`**. The one direction that did change the result was the +**opposite**: an older base ref made the check **fail**. So the *imprecise rerun instruction* — which +is real and confirmed — is made explicit (the complete rerun names step 4's guarded fetch and base-ref +recording), **without asserting the unproven consequence**. Topologies tested are named in the plan +and are **not exhaustive**. + +**Minor (marker absent from the Failure report) — collected, not repaired**, per instruction. + +**Two corrections to my pass-50 report, both validated before acceptance.** + +1. **That Minor is NOT mine.** I attributed it to this repair round. **Both the Failure-report + enumeration and the marker already exist in `1ba45be`** — the enumeration verbatim at line 492 of + that revision, the marker four times. It is **pre-existing**, and the attribution was wrong. +2. **"The curve turned" was interpretation, not a finding.** `CLAUDE.md:226` requires the Blocker + curve **across passes** and says in terms that *"one low count is a snapshot rather than a + plateau"* and that **neither curve measures coverage**. 6 → 4 findings and 5 → 2 Blockers are an + improvement **over pass 49**, not a demonstrated trend reversal, and "only one tell remains" is + read the same way. The instructed stop stands regardless. + +**Verified by execution: 29 fixtures × `sh`/`dash`/`bash`, all green** — 10 (finding 1), 6 +(finding 4), 5 (marker selection), 8 (router + no-repair route). Every repair carries a control that +reproduces the defect it fixes. **One syntax regression was introduced and caught by the check**: a +placeholder `if ` is not valid shell and took `sh -n` from 6 to 7; it is now a +quoted assignment and the baseline is restored. **Harness bugs again outnumbered plan defects** — +BSD `sed` reading `||` as a flag separator, and an `awk` skip count that left an orphan continuation +line; both fixed in the harness, neither in the plan. + +**Unverified, and stated rather than implied:** Resume's three marker outcomes and the +`loop-rule-start` recovery *rule text* are reader instructions, asserted as text. The router branch +that reaches that rule **is** now executed. Daniel's reviewer has not reproduced any fixture result. + +**No pass 51 is authorized, and none was run.** + +### PASS 50 — RUN, VALID, UNCLEAN. Cycle stays OPEN. + +Run 2026-09-17 against the pinned state below, on the reviewer's recommendation of **exactly one** +full pass with the four priorities **in the prompt file itself** — which pass 49's lacked. Result +**VALID** (terminator exact, 4 body lines, 6 fields each, no blanks, count matches). The plan blob +was **`af3e676d…` before and after the call**, so the artifact did not move under the reviewer this +time. + +**4 findings: 2 Blockers, 1 Major, 1 Minor** — file +`.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-50.md`. **No repair is authorized; none was made.** + +**Two of the four are mine, from the seventeenth revision, and that is the important half.** + +- **MAJOR — my `loop-rule-start` recovery rule is UNREACHABLE.** Resume defines "surviving start, no + base" as a precondition stop, but **Task 0 routes solely on whether the base exists**, so that state + selects first-entry Preparation, which never refuses an existing start file and **overwrites it**. + The recovery rule I added to satisfy the very class pass 48 punished cannot be reached — the record + of the validated revision is lost silently. **I added a state file and a rule, and did not wire the + router to it.** +- **MINOR — the Failure handoff report was not updated** when `loop-rule-records-commit` became + authoritative recovery state; it still enumerates base, reviewed head, reviewed tip and closing + message, so a failure report can omit that a recorded pass is still awaiting the ordering. + +**Two are pre-existing and were not created by this revision.** + +- **BLOCKER — step 7 step three deadlocks the no-repair continue route.** Continue explicitly permits + an unrevised artifact when no repair is owed, but step three always stages the unchanged plan and + attempts a commit; after step one committed the findings files there is **no tracked change**, so + the commit fails and the reviewed head is never advanced. A below-floor Minor-only pass or an + answered suspension **cannot issue its next Gate-B pass**. +- **BLOCKER (medium) — the complete pre-candidate rerun does not repeat step 4's guarded fetch and + base-ref recording**, so the battery reuses a stale base object while PR CI compares against a newer + one: the plan can close a commit that **fails invariant 12**. + +**Loop health: one tell of five, and the curve turned.** Findings across 42–50: **3, 5, 9, 2, 1, 2, +3, 6, 4**; Blockers **2, 2, 2, 1, 1, 1, 3, 5, 2**; Majors **1, 2, 7, 1, 0, 1, 0, 1, 1**. Finding +count **fell** 6 → 4; Blockers **fell** 5 → 2; no prose cluster; no require↔withdraw pair. The only +tell standing is the **instrument cluster**, for the thirteenth pass. **One tell is below the +mandatory-surface threshold of two** — this stop is instructed, not compelled by the tells. The +"clearly stuck" exit remains unavailable: there is no plateau, the curve is improving. + +**Nothing here closes the cycle.** Two Blockers and a Major are open and §5 requires each to resolve. + +### PINNED STATE for pass 50 — 2026-09-17 + +The plan stopped moving here. **Anything reviewing this revision must check these first**, because +two consecutive reviews were spent on a target that changed underneath them: + +- **HEAD** `5fb3b942f1907f5d6dd25b32d15b497f855e05f8` — unchanged throughout. +- **Last committed plan revision** `1ba45be`; the plan is **uncommitted-dirty** against it. +- **Worktree plan blob** `af3e676d17e6da280557e96e7db4a3657ab48d05`. +- **Diff** +205 / −26. Earlier reports of +174/−26 and +186/−26 were true of superseded states. +- **Fenced blocks** 55, unchanged from the baseline. +- Advisor files still dirty and untouched; `CLAUDE.md`, `AGENTS.md`, `plugins/`, `scripts/` untouched. + +### State reconciliation, 2026-09-17 — the handoff against the actual tree + +Daniel's reviewer read this work **while it was still moving** and was right to flag it. Recorded +plainly so no later reader trusts a stale line: + +- **HEAD is `5fb3b94` and has not moved.** The last *committed* plan revision is **`1ba45be`**. +- **The plan is uncommitted-dirty.** The reviewer saw blob `16922358…` mid-edit — a state that held + findings 2 and 6, then 4 and the cleanup list. That blob is **superseded**; the current worktree + blob is what `git hash-object` reports now, and it carries all six. +- **`.context/plan-drafts/pass-49-repair-draft.md` said "DRAFT ONLY / Nothing applied".** True when + written, false by the time it was read. **Corrected** — the file is now marked SUPERSEDED and keeps + its value as the reasoning record, not as a description of the tree. +- **No pass-50 findings file exists, and none should.** Pass 50 is not authorized. **A continuously + edited draft is not a finished repair**, and nothing here claims the cycle is clean: pass 49's + finding stands as recorded, the cycle is **still OPEN and UNCLEAN**, and neither the disposition of + finding 3 nor these repairs make pass 49 retroactively clean or close anything. + +**Two corrections adopted from that review, both validated first.** + +1. **"A different close mechanism ⇒ a new precondition" is not a valid general inference**, and the + installed wording no longer makes it. What is out of scope is the **atomicity guarantee against a + concurrently writing second actor**; needing another git command is a *consequence* of that demand, + not the reason. The residual also now states plainly that **the race is neither impossible nor + harmless** and that **its window is unmeasured**. +2. **The reviewer withdrew their own claim that the utility test was missing from pass 49's prompt** — + it was present at line 572. **The four priority areas were genuinely missing**, and that half + stands: it is why pass 49 is not a usable test of the sharpened approach. **They are now written + into `.context/gate-a-plan-prompt.md` itself**, as priorities that explicitly exclude nothing. + +**My own error, recorded because it is the cycle's own defect class.** I reported that a validation +subagent had invented the quotation *"The gap is real and this change does not close it"*. **It is +real**, at `target-text:1234`; my `grep -F` missed it because the sentence **wraps between `this` and +`change`**. That is finding 6's defect exactly, hit while repairing finding 6 — the third instance in +this cycle after carried `e9` (pass 10) and Task 7's fragment (pass 49). The accusation is withdrawn +in the dispositions file. + **After a clean close**, the real work starts: both prompt copies, the seven hook strings and their test expectations, version bump 0.11.0 → 0.12.0 + CHANGELOG, the quality battery, the evidence entry, then **Gate B with a fresh nonce** — this cycle's is Gate-A plan only. @@ -386,6 +1231,7 @@ worth checking before a pass rather than after. | — | — | — | — | — | — | **FIFTEENTH BOUNDED REVISION.** Step one persists the record commit's object id to `.context/loop-rule-records-commit` (added to the cleanup list); **step two reads both slots from that exact object**, and a moved `HEAD` is a reason to **stop and report**, never to re-resolve. 8b's removal is guarded — **that is the line that closes the demonstrated case**, and the prose credits it rather than the `test -f` after it, which **no demonstrated path reaches**. Commit `c2e4932` | | 48 | c2e4932 | 2→**3** | 1→**3** | 1→**0** | yes | **All three Blockers are ONE defect**: the plan resolved `HEAD` a second time as the identity of a commit whose properties it had already checked. Step 7 step one (tree checked on one commit, id persisted from another); **8a — the one with a shipped payer**: a moved commit could become the closing tip having satisfied **none** of condition 4's checks, and `reset --soft` folds it into the closing commit, publishing content no pass reviewed; and the `.context/loop-rule-records-commit` marker added one revision earlier with **no Resume rule**, which this plan requires of every state file | | — | — | — | — | — | — | **SIXTEENTH BOUNDED REVISION.** **Swept rather than patched per site** — repairing one of several sites is this cycle's most reliable defect. Each site now resolves the commit **once**, immediately after it lands, and uses that object id for every check and every record; 8a's whole condition-4 chain runs against it and writes it as the tip. **A third site the reviewer did not report** was found and fixed: step 7 step three checked the repair commit's tree and then re-resolved `HEAD` for the reviewed head. **Checked and deliberately unchanged:** step 4b resolves `HEAD` once and persists no identity, so a move makes its comparison fail rather than pass; conditions 1, 5 and 6 read `HEAD` on purpose, to detect a move. Resume now validates the marker and reports a recorded pass as **unrouted** unless the reconciliation shows step two's outcome — a precondition stop, no mutation, no call. Commit `1ba45be` | +| 49 | 1ba45be | 3→**6** | 3→**5** | 0→**1** | yes | **UNCLEAN — reported, no repair round opened; the reviewer's authorization covered exactly this one pass.** Nothing re-raised against the three blocks the sixteenth revision repaired. **Five of six are that same defect class at sites the sweep did not reach**, which is what makes this discovery rather than regeneration. **BLOCKER** Preparation resolves `HEAD` six times and Task 0 step 1 writes the base from a seventh — missed by the sweep, in neither of its lists. **BLOCKER** step 6 records the reviewed head, then specifies the call's `headSha` in prose as a fresh resolution never wired to that file; **not shell, so a shellcheck-verified sweep could not see it**. **BLOCKER** condition 6 composes subject, parent, tree and body from four live `HEAD` reads with no captured id, then cleanup erases the recovery state — **explicitly excluded by the sweep on a rationale that holds for conditions 1 and 5 and not for 6**. **BLOCKER, partial** 8b's `reset --soft` has no atomic expected-old-object guard, but condition 5 already closes the wide window; the residual is sub-second and needs a concurrent actor, and the gap is disclosed as parked. **BLOCKER, partial** the records-commit marker's new checks are satisfied trivially by any earlier already-routed records commit, and the rebuild rule cannot reconstruct it; the named case does have a real precondition stop. **MAJOR, demonstrated** Task 7's worked fragment spans the line break between `Codex is` and `advisory` in both copies, so its `grep -cF` returns 0 where the plan expects `parent=1 worktree=1` — a correct source fails its own check. Three tells stand (findings 3→6, Blockers 3→5, instrument cluster) → mandatory stop | | 26 | 05ec1b2 | **5** | **1** | 2 | yes | first pass on the revised plan. Resume audited progress by `WIP:` commits and called any non-`WIP` commit a stale base — the handoff leaves two other shapes, one of them `HEAD` **at** the base with everything in the index | | 27 | f69db7d | 5→**6** | 1→**3** | 2→**1** | yes | the close required the commit's **tree to equal the tip's**, which target §I **parks**; `reset --soft` leaves the index untouched, so the prose beside it was wrong about git; Resume still inferred staleness from a commit subject where §A3 names that state as reachable | | 28 | f2e5dda | 6→**6** | 3→**2** | 1→**2** | yes | pass 27's stale-base fix reached one site of three; its reset paragraph stated the index behaviour correctly and repeated the false claim four lines later; nothing checked the closing commit's **parent**; the message was validated in the file, where `commit-msg` hooks rewrite git's copy after `-F` reads it | diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-49-dispositions.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-49-dispositions.md new file mode 100644 index 0000000..0063bdf --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-49-dispositions.md @@ -0,0 +1,43 @@ +# Pass 49 dispositions — cycle `om0bdd7udh` + +> **APPLIED 2026-09-17 — the seventeenth bounded revision. All six are repaired in the worktree; +> nothing is committed.** Finding 3 was settled **through a Codex gate**, not by handing Daniel a +> menu: the gate returned **NOT OWED — record as a disclosed residual**, its citation of condition +> 5's wording was verified against the file, and it was implemented that way. Verification: **16 +> fixtures × `sh`/`dash`/`bash`, all green**, with the blocks extracted verbatim from the plan and a +> control per finding reproducing the old defect. `sh -n`/`dash -n` failures are **6 and 7, +> identical to the pre-repair baseline** — no new syntax failure. Fenced blocks 55 → 54, the one +> removal being the cleanup block merged into condition 6. +> +> **Three fixtures failed on the first run and all three were harness bugs, recorded rather than +> quietly fixed:** the disposable repo had no `.gitignore`, so `.context/` counted as a dirty tree +> (two failures); and a control asserted a pass-count that was simply miscalculated. The plan was +> not touched for any of them. +> +> **RETRACTED — the error was mine, and it is finding 6's own defect class.** An earlier version of +> this block said a subagent had supplied the quotation *"The gap is real and this change does not +> close it"* and that the string occurred in no file here. **The string is real**, at +> `docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md:1234`, inside the bullet +> beginning at 1231. My `grep -F` for it returned nothing **because the sentence wraps between `this` +> and `change`** — a literal search for a string that spans a line break, which is precisely the +> defect finding 6 repairs and which this cycle has now hit three times (carried `e9` at pass 10, +> Task 7's fragment at pass 49, and this). The subagent was right; the accusation was wrong and is +> withdrawn. **The lesson is the one finding 6 installs**: a literal count over multi-line prose +> proves nothing until the wrap is ruled out. + +Advisory companion per CLAUDE.md §5 "Optional companions". One line per finding: verdict + reason. +Not the findings file; participates in no pass validation. Verdicts recorded 2026-09-17 after +validation against the plan text (five read-only subagents) and, where marked, by execution. + +1. BLOCKER — Preparation / Task 0 step 1 base binding | ACCEPTED, into the fix set | Confirmed: six live `HEAD` resolutions establish the predicates, a seventh writes `.context/loop-rule-base`. Not in either of `1ba45be`'s lists — missed, not excluded. No guard, no disclosure. +2. BLOCKER — step 6 / step 7 call issuance `headSha` | ACCEPTED, into the fix set | Confirmed: the reviewed head is recorded to a file, then the call's `headSha` is specified in prose as a fresh resolution never tied to it. CLAUDE.md's own rule is satisfied; the plan's self-review item 28 is what demands more. Not shell, so the sweep could not have reached it. +3. BLOCKER — 8b `reset --soft`, no atomic expected-old-object guard | DISPUTED, NOT accepted into the fix set; Daniel's decision required | Condition 5 says "when the closing invocation begins" and the plan does exactly that, so no stated requirement is violated; the wide window is already closed; the residual needs a concurrent writer that nothing establishes exists; and the proposed guard replaces `reset --soft` with a different close mechanism, which is a new precondition rather than a repair. Full disposition in the repair draft. +4. BLOCKER — Close condition 6 and cleanup | ACCEPTED, into the fix set | Confirmed: subject, parent, tree and body read through four independent live `HEAD` resolutions with no captured id; the parent test admits any commit parented at `$BASE`; cleanup then erases the recovery state behind a prose-only gate. Excluded by `1ba45be` on a rationale that holds for conditions 1 and 5 and not for 6. +5. BLOCKER — records-commit marker recovery | ACCEPTED, into the fix set | Confirmed in both halves: the five added checks are satisfied trivially by any earlier already-routed records commit, and the generic "rebuild from the `$BASE` blobs" rule is a category error for an artifact that is not a function of `$BASE`'s content. Blunted only for the one case the plan names, which does carry a real precondition stop. +6. MAJOR — Task 7's worked carried fragment | ACCEPTED, separate small fix | Confirmed BY EXECUTION: the fragment spans the line break between `Codex is` and `advisory` in both copies, so `grep -cF` returns 0 where the plan expects `parent=1 worktree=1`. Dropping the leading `Codex is ` gives 1/1 in both copies, verified. The plan's second example, `Open a TodoWrite`, is 1/1 and is fine. + +## Corrections to the pass-49 status report, all three validated before acceptance + +- ACCEPTED — "all six are paid for by the executor alone" was wrong. The standard is the causal chain, not the first affected party. `plugins/dev-workflow/commands/workflow-init.md`, `codex-gate.sh` and `codex-gate.test.sh` all ship inside the plugin package, so findings 1, 2 and 4 — and 3 if it ever fires, and 6 by way of an executor "repairing" correct text — can put content into shipped files that no review covered. +- ACCEPTED — "no regeneration from the repairs" was too sweeping. `1ba45be` carries SEVEN hunks across FOUR regions, not three; the first (`@@ -628`) is Resume's marker validation, which is exactly what finding 5 attacks. Finding 5 IS regeneration. The "clearly stuck" exit remains unavailable: the plateau conjunct fails (Blockers rising 3 → 5) and no affirmative coverage judgement is possible. +- PARTLY REJECTED — the claim that the utility check was absent from the instruction file is false: it is present at line 572, "a finding is acted on when it names a bugfix, a security fix or a performance gain, and names who pays". The other half is right and is my defect: the four priority areas appear nowhere in that file, and the wrapper instruction asserted "the prompt names four priority areas" when it named none. The conclusion stands on that half alone — pass 49 is not a usable test of the sharpened approach. diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-49.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-49.md new file mode 100644 index 0000000..21921a0 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-49.md @@ -0,0 +1,7 @@ +BLOCKER | high | Preparation and Task 0 step 1 | Preparation tests the branch, ancestry, and three approved-input blobs through separate live HEAD resolutions, then Task 0 resolves HEAD again to create the base, so the recorded base is not bound to the commit whose properties were checked | A commit, reset, rebase, or checkout between those reads can make the plan implement from an unapproved revision or the wrong branch and later review and squash that different base | Resolve one full starting object id once, use it for every Preparation predicate and as the exact base value, and refuse to proceed unless both HEAD and the branch still match it before the first mutation +BLOCKER | high | Task 15 steps 6 and 7 call issuance | The plan records a reviewed head from one HEAD resolution but defines each review call's headSha as what HEAD resolves to at call time instead of explicitly passing the already recorded object id | A ref move between recording and issuing the call can make Gate B review one commit while Close condition 1 later authenticates another; if the ref returns to the recorded id, unreviewed content can close | Pass the exact object id written to loop-rule-reviewed-head as headSha for both branches of every call and never resolve HEAD again to choose that call's reviewed object +BLOCKER | high | Task 15 step 8b, Close condition 5 | Condition 5 compares HEAD with the closing tip and checks cleanliness, but the later reset --soft operates on whatever HEAD names then, with no atomic expected-old-object guard | A commit, amend, reset, rebase, or checkout in that window can put different content into the index after the cleanliness check; the closing commit can silently publish it because the parked tree-equality condition is the only later content comparison | Move the ref from the exact recorded tip to the base with an atomic expected-old-object check, preserving the soft-reset index semantics, and stop if the tip is no longer the object being moved +BLOCKER | high | Task 15 step 8, Close condition 6 and cleanup | The post-close procedure never captures the closing commit object and reads its subject, parent, and body through separate live HEAD resolutions before deleting all recovery state | A move or amend between those reads can combine properties from different commits or substitute another parented-by-base commit, let the split postconditions pass, and erase the evidence needed to recover | Capture and persist the closing commit object id immediately after the commit lands, use that immutable id for every condition-6 read, and verify HEAD still equals it before cleanup +BLOCKER | high | Resume scratch-artifact rule and Task 15 step 7 step one | loop-rule-records-commit is written only after the records commit lands; a failed write can leave it absent or pointing at an older records commit, while Resume accepts any ancestral commit with the two slot blobs and the generic instruction to rebuild an invalid artifact from base blobs cannot reconstruct this marker | Resume can mistake the earlier pass's routed marker for the new unrouted pass and issue a repair or review before the ordering speaks, or stop in a state the plan gives no executable recovery for | Give this marker its own recovery rule that detects a records commit newer than the marker, binds the marker to the exact current slot pair and pass, treats missing, stale, or invalid state as a precondition stop, and includes it in the failure handoff report +MAJOR | high | Task 7 step 4 | The claimed carried-condition fragment `Codex is advisory — validate before applying; dismissed finding → one-line why` spans the parent files' line break between `is` and `advisory`, so it counts zero in both parents even though the step says it is one of the pre-install fragments and expects parent=1 worktree=1 | Following the worked example makes the required preservation check fail on correct source text or leaves a21 and a22 without trustworthy recorded evidence | Replace the example with line-local fragments derived before installation for the affected carried conditions, or remove the example and rely on the governing derivation procedure +END OF FINDINGS (6 total) diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-50.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-50.md new file mode 100644 index 0000000..3851fbb --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-50.md @@ -0,0 +1,5 @@ +BLOCKER | high | Task 15 step 7 step three | Continue explicitly permits an unrevised artifact when no repair is owed, but the only step-three procedure always stages the unchanged plan and attempts a new repair commit; after step one already committed the two findings files there is no tracked change, so that commit fails and the reviewed-head file is never advanced to the records commit | A required no-repair continue route such as a below-floor Minor-only pass or an answered suspension cannot issue its next Gate-B pass, so the plan deadlocks on a state the installed ordering requires to progress | Split step three: when no repair is owed, remove any stale tip and record the captured records-commit object id as the next reviewed head without another commit; keep the guarded repair-commit path only when a repair or refreshed record changed tracked content +MAJOR | high | Preparation and Task 0 step 1 / Resume scratch-artifact validation | Resume defines a surviving loop-rule-start with no base as a precondition stop that reports the recorded revision and current HEAD without mutation, but Task 0 routes solely on base existence; that state selects first-entry Preparation, which never refuses the existing start file and overwrites it | The newly added recovery rule is unreachable, so an interrupted first entry silently loses the only record of the revision the earlier Preparation validated and skips the mandated handoff | Check loop-rule-start before entering Preparation; when it exists without the base, report and stop as Resume specifies, then allow a deliberate Preparation rerun to replace it only after the predicates are re-established +BLOCKER | medium | Task 15 step 7 complete pre-candidate rerun | The mandatory complete rerun names the whole step-4 battery but does not say to repeat step 4's guarded fetch and base-ref recording; the battery therefore reuses the object id persisted by the initial run even if origin/main advances during the Gate-B loop | The version-bump check can pass against a stale merge base while pull-request CI compares against a newer base carrying the same plugin version, so the plan can close a commit that fails invariant 12 | Make the guarded fetch, resolution, and recording of the exact base object part of every complete pre-candidate step-4 rerun, then pass that refreshed object id to check-version-bump.sh +MINOR | medium | Failure procedure report | The cycle-value report still enumerates only the base, reviewed head, reviewed tip, and closing message after loop-rule-records-commit became authoritative recovery state | A closing failure handoff can omit that a recorded pass is still awaiting the ordering, so the report no longer contains every cycle value Resume must reconcile | Add the records-commit marker's existence, object id, and newest-records comparison to the Failure handoff without mutating it +END OF FINDINGS (4 total) diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-51-dispositions.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-51-dispositions.md new file mode 100644 index 0000000..87b7ff3 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-51-dispositions.md @@ -0,0 +1,25 @@ +# Pass 51 dispositions — cycle `om0bdd7udh` + +Advisory companion per CLAUDE.md §5. One line per finding: verdict + reason. Not the findings file. +**No repair was made and none is authorized.** Verdicts recorded 2026-09-17 by checking each claim +against the worktree plan at blob `974e223ed71d52b49c6b368695372ccb72e510f1`. + +Reviewer claims are marked **[R]**; what I established myself is marked **[V]**. + +1. BLOCKER — step 4b commits after the battery and 4c | **CONFIRMED, accepted** | [V] Task 15 order is step 4 (`:2895`) → **4c (`:2945`)** → **4b (`:3074`)** → step 5 (`:3145`) → step 6 (`:3181`), and 4b's own heading says "and commit the result before Gate B". So a new WIP commit lands **after** 4c certified a head, and [R] the first-pass path runs no complete rerun before step 6 although a zero-finding first pass may close. A 4b repair to a shipped prompt or hook is then absent from 4c's merge result entirely. +2. BLOCKER — 4c's `HEADID` and step 6's reviewed head are not tied | **CONFIRMED, accepted** | [V] `HEADID` is an operator placeholder at `:2959`; step 6 independently runs `git rev-parse HEAD > .context/loop-rule-reviewed-head` at `:3188`. Nothing compares them. [R] A ref move between the two lets 4c certify one candidate while Gate B reviews another. +3. MAJOR — 4c has no durable record | **CONFIRMED, accepted** | [V] 4c only prints; the step 5 / 7b region (`:3145`–`:3260`) contains **zero** mentions of 4c, so the "result goes into the closing commit body" claim I wrote has **no mechanism behind it**. [R] The closing body can omit 4c or carry a stale result. Note the shape: this is the same defect class as pass 48's marker — new state with no record, recovery or cleanup rule. +4. MAJOR — the synthetic merge reverses CI's parent direction | **CONFIRMED as described, accepted** | [V] 4c checks out `HEADID` (`:2968`) and merges `BASEID` into it (`:2980`); CI's `refs/pull/N/merge` is the base with the head merged in. [R] With direction-sensitive merge behaviour the trees can differ. Reviewer's own confidence is `medium` and I did not construct a case where the trees actually diverge — **the direction mismatch is established, a resulting tree difference is not**. +5. MINOR — the plugin-diff pipeline is unguarded | **CONFIRMED, collected** | [V] `:2997` embeds `git diff --name-only … \| tr` in a command substitution, so the status is `tr`'s and a failed diff yields a silent empty evidence line. This is the exact defect class passes 43 and 44 swept for — **and I reintroduced it in new code**. +6. NIT — the `mktemp -d` directory is never removed | **CONFIRMED, collected** | [V] The only `trap` match in the plan is the English word in unrelated prose at `:1798`. No cleanup on any branch. + +## Provenance of these findings, which is the uncomfortable part + +**Four of six are in machinery I added in the last two revisions** (findings 3, 4, 5, 6 are all step +4c; finding 2 is 4c's interface to step 6). **Findings 5 and 6 are defect classes this cycle had +already swept for** — an unguarded pipeline whose status a following command consumes, and state +created without a cleanup rule. Writing new code reintroduced both. That is regeneration from the +repairs, not discovery in unswept territory, and it should be read that way. + +**Finding 1 is the most consequential and is not about 4c's internals at all** — it is about where 4c +sits in the sequence. Adding a verification step in the wrong place bought less than it appeared to. diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-51.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-51.md new file mode 100644 index 0000000..e3792dc --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-51.md @@ -0,0 +1,7 @@ +BLOCKER | high | Task 15 steps 4, 4c, 4b, and 6 | Step 4b necessarily creates a new WIP commit after both the battery and 4c have run, and the first-pass path contains no complete rerun before step 6 even though a zero-finding first Gate-B pass may close the cycle | Gate B can review and close a head that never received the required battery or merge-result version check; if step 4b repairs a shipped prompt or hook, that repaired plugin content is wholly absent from 4c's merge result | Move step 4b and its commit before the final verification phase, then run the complete battery and 4c after that commit and before the first Gate-B call +BLOCKER | high | Task 15 step 4c and step 6 | 4c takes `HEADID` from an operator-filled placeholder, while step 6 independently resolves live `HEAD` into `.context/loop-rule-reviewed-head`; no check establishes that these are the same object id | A commit, amend, reset, rebase, or checkout between the two resolutions lets 4c verify one candidate while Gate B reviews and the close authenticates another | Capture the candidate head once after the last pre-candidate commit, persist that exact id, use it as 4c's `HEADID` and Gate B's literal `headSha`, and only re-read live `HEAD` to reject movement +MAJOR | high | Task 15 step 4c, step 5, and step 7b | 4c only prints its ids, diff, output, raw status, and verdict to the terminal; it has no durable record, step 5's specified evidence sources omit it, and step 7b later rebuilds the closing message whole from current records | The closing commit body can omit 4c or carry a stale result for an earlier head/base pair, so the newly required merge-result verification is not auditable at close | Persist 4c's complete output and status in a cycle record keyed by the exact head/base/checker identities, give that state a recovery and cleanup rule, and make steps 5 and 7b consume only the latest matching record +MAJOR | medium | Task 15 step 4c | The synthetic merge checks out the candidate head and merges the base into it, reversing the parent direction of the pull-request merge result that CI checks out | With direction-sensitive merge behavior such as asymmetric merge drivers, 4c can run against a tree different from CI's merge tree and certify the wrong version-check result | Check out the pinned base and merge the pinned candidate head into it, while still pinning the intended checker bytes from the candidate head before the merge +MINOR | high | Task 15 step 4c | The plugin-diff record embeds `git diff --name-only ... \| tr ...` in a command substitution without checking the pipeline's `git diff` status | A failed diff is masked by `tr`, producing an empty or partial evidence line while 4c can continue to a verified exit | Capture and guard the `git diff` output before transforming or printing it, and treat an unestablished record as unresolved +NIT | high | Task 15 step 4c | The `mktemp -d` repository is never removed on success or on any failure branch | Every first pass and rerun leaves a shared clone and its object-alternates metadata in temporary storage | Install a cleanup trap immediately after `mktemp -d` and remove the directory on every exit +END OF FINDINGS (6 total) diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-52-dispositions.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-52-dispositions.md new file mode 100644 index 0000000..6c75f4d --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-52-dispositions.md @@ -0,0 +1,146 @@ +# Pass 52 dispositions — Gate-A plan cycle `om0bdd7udh` + +Advisory companion per CLAUDE.md §5. Not the findings file; not part of pass validation. + +**Snapshot.** `HEAD` `5fb3b942f1907f5d6dd25b32d15b497f855e05f8`, branch `loop-rule-consolidation`, +plan worktree blob `15a9b1c2ff4beb3c6dfa6ae21e00550210d4b969` — **identical before and after the +call**. Nothing staged. Tool: `mcp__codex__exec`, session `01a0b0ac-5d48-75d1-86b8-e196edda1f40`. + +**Pass VALID.** Terminator `END OF FINDINGS (4 total)` exact; 4 body lines; every line a finding +line; 6 fields each; no blanks; count matches. File +`.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-52.md`, untracked. + +**Floor 3**, derived from `docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md` +read fresh at this pass: **Risk `high`** (level 2) · **Security `none`** (0) · max 2 ≠ 0 → 3. One +cited story, from the plan's `**Story:**` header, which parses. Profile resolves; the profile log +holds one `adoption` entry, no override and no axis change. Floor long satisfied; what is missing is +a **clean** pass. + +**Method change, declared.** Pass 52's instruction was delivered by pointing Codex at +`.context/gate-a-plan-pass-52-instruction.md` (1006 lines) rather than inlining it in the tool +argument. The reason is fidelity: reproducing 100 KB verbatim into a parameter risks silent drift, +while the file on disk is exact. The findings-protocol contract was restated inline so it could not +be missed. **Whether Codex read the file in full is not established** — the reply and the findings +are consistent with having done so, and that is evidence rather than proof. + +## Findings + +### 1 — BLOCKER, Task 15 step 4b candidate capture. **CONFIRMED.** + +Verified in the text. The commit block is plan lines **2954–2967**: it pins `ITREE=$(git +write-tree)`, commits, then checks `test "$(git rev-parse "HEAD^{tree}")" = "$ITREE"` — one live +`HEAD` read. The capture is a **separate** block, plan lines **2987–2993**, whose whole body is +`git rev-parse HEAD > .context/loop-rule-reviewed-head` — a **second, independent** live `HEAD` +resolution. The plan states the binding rule against itself at line **3027**: "each fenced block +below runs in its own shell invocation". So the tree check authenticates one object and the capture +records whatever `HEAD` names at a later invocation, with nothing tying them. + +The capture block's own comment reads "Resolve it ONCE; every consumer below reads this file." That +is true of the consumers **below** it and false of the binding **above** it, which is the half the +finding attacks. + +**This is regeneration, not discovery in unswept ground.** It is the resolve-once class of passes +47–49, reintroduced by the twenty-second revision's own new capture block — the third time in this +cycle that new machinery has brought back a class the cycle had already swept. + +**Payer:** `workflow-init.md`, `codex-gate.sh` and `codex-gate.test.sh` ship inside the plugin +package, so a substituted candidate can put content into shipped files that never received the +pre-candidate checks. Not the executor alone. + +**In the assigned fix set.** It corrects the correction the last two revisions made, inside the same +candidate-identity machinery. Absorbable by the loop — **not repaired here**, because this +assignment stops before repairs. + +### 2 — BLOCKER, steps 4b–4 and step 6, battery bound to the worktree. **CONFIRMED AS FACT; OPENS A NEW STRUCTURAL QUESTION.** + +The fact is not in dispute and the plan **states it itself**, at lines 2977–2985: step 4 runs over +"the **working tree of the current repository**"; it "describes this candidate only because it runs +immediately after this capture with no commit in between"; and "**Nothing enforces that the worktree +stays unmodified in that window** — step 6 checks that `HEAD` has not moved, which catches a commit +and not a dirty tree. Stated as a limit rather than guarded." + +So Codex reports a real gap the plan discloses. **Disclosure does not discharge it** — the twentieth +revision settled that in this cycle ("an accepted Blocker is not discharged by disclosing it"), and +the consequence reaches shipped files for the same reason as finding 1. + +**But the fix it proposes is a new mechanism**, not a repair of an existing one: run the battery in +a detached disposable checkout of the recorded candidate, or establish and re-check that index and +tracked worktree equal the candidate on both sides of the battery. That is a second checking +apparatus with its own preconditions and its own failure modes. This cycle has declined such +additions before on the "new precondition" ground (pass 44's three escalations, finding 3 of pass +49) **and has also ruled the opposite way** — the twentieth and twenty-first revisions established +that "a different checking mechanism is not by itself a new requirement", which is what put step 4c +into the plan at all. **Both rulings are live and they point opposite ways here.** + +**That is a contract question, and §5 sends it to the user rather than letting the loop absorb it.** +Novelty of the question decides, not size. **This is the blocking decision.** + +### 3 — MINOR, Task 0 step 3 baseline site coverage. **CONFIRMED IN KIND; THE NAMED SITES UNVERIFIED. COLLECTED.** + +Confirmed in kind: step 3's body promises "Diff every **inventoried and changed** site" (line 1220), +but the artifact's schema says "one `site` record per **inventoried** site" (line 1247), the +completeness predicate says "**every inventoried site** has a `site` record" (line 686), and the +step's own heading is "the **inventoried** ranges" (line 1218). The words "and changed" have no +carrier anywhere — three statements of the rule say "inventoried" and one says more. That is the +same shape as the already-collected "locator steps stating a result for both copies while naming +one". + +**Not verified:** that §G and §F item 18's work-loop line are in fact outside the nine anchor pairs +in the `SITES` heredoc. Establishing that needs the design's §4 destination-block set cross-read +against those anchors, which this pass did not do. + +**Minor → collect, never iterate** (§5 Mechanics). Recorded, not repaired. + +### 4 — NIT, step 4b commit subject. **CONFIRMED BY DIRECT READ. COLLECTED.** + +Verified independently before reading the finding. The block's comment at plan line ~2957 reads +`# Where a repair was made, the message is: "WIP: prompt-standards repair + plan records"`, and the +only command in the block is `git commit -m "WIP: plan records"` — unconditional. The prose promises +a distinction the block cannot make. + +**Nit → collect, never iterate.** Recorded, not repaired. + +## Mechanical prechecks, run on this blob before the call + +New observations this session, not a restatement of earlier reports. They used a different extractor +from the plan's own reported baseline and **are not comparable to it**. + +- **60 fenced blocks, all balanced, none unclosed** — extractor matches indented fences, since one + blind to them already produced a wrong count in this cycle. 57 carry `bash`, 1 `python`, 2 none. +- **`sh -n` fails on 8 extracted blocks, `dash -n` on 9.** Every one is in an intended class: five + are `<…>` operator placeholders (`base<…>`, `NONCE=<…>`, `git add … `, the record-shape schema), four are `diff <(…)` process substitution inside `bash`-marked + fences, which `dash` cannot parse by design. **No new syntax regression.** +- **Every repo path the plan cites resolves.** The only non-resolving strings are bare basenames used + as shorthand in prose, two fixture filenames from the version-bump reproduction table, and + run-time state files. +- `.context/loop-rule-baseline-diff` is written `.tmp` and atomically `mv`-ed to `.txt` — one file, + not an inconsistency. + +## Loop health at pass 52 + +Findings 42–52: **3, 5, 9, 2, 1, 2, 3, 6, 4, 6, 4**. Blockers: **2, 2, 2, 1, 1, 1, 3, 5, 2, 2, 2**. +Majors: **1, 2, 7, 1, 0, 1, 0, 1, 1, 2, 0**. + +**Tells present — at least two, so stop-and-surface is mandatory, not discretionary.** + +1. Finding count rising — **NOT present**; 6 → 4, falling. +2. Blocker count failing to fall — **PRESENT**; flat at 2 for the third consecutive pass. +3. Instrument cluster — **PRESENT**, and total: all four findings are the plan's own execution + machinery. **Zero** touch the §5 target text this change installs. +4. Prose cluster — **partially present**, and reported rather than resolved: findings 3 and 4 are + both prose promising what the command beside it does not do, which is two of four. Whether that + counts as a cluster in its own right or as a subset of the instrument cluster is a judgement; it + changes nothing, since 2 and 3 alone already make the stop mandatory. +5. require↔withdraw pair — **NOT present**. Pass 51 asked for the candidate to be captured once and + persisted; the twenty-second revision did that; pass 52 asks for it to be **bound** to the commit + just made. A sharpening, not a demand for something an earlier pass removed. + +**The "clearly stuck" exit is NOT available.** Of three conjuncts: no plateau of six or more passes — +the Blocker curve over the last six is 1, 3, 5, 2, 2, 2, flat only for three; **no affirmative +coverage judgement is possible** — finding 1 is a defect the twenty-second revision itself created, +which is direct evidence that the new machinery is unswept; regeneration **is** present, which is one +conjunct of three, and one is not the exit. + +**Nothing here closes the cycle**, and a clean pass would not close it alone: closure needs the clean +pass **plus** every other duty §5 names. diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-52.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-52.md new file mode 100644 index 0000000..e9ca224 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-52.md @@ -0,0 +1,5 @@ +BLOCKER | high | Task 15 step 4b candidate capture | The plan checks the just-created commit through a live `HEAD^{tree}` read, then resolves `HEAD` again in a later fenced invocation to write `.context/loop-rule-reviewed-head`; a move between those resolutions can record a different commit from the one compared with `$ITREE`. | Step 4c, Gate B, and Close condition 1 then authenticate the substituted commit while the prompt-standards record and tree pin describe the earlier one, so shipped prompt or hook content can bypass the required pre-candidate checks. | Resolve the post-commit object once immediately after `git commit`, use that same object id for both the tree comparison and the reviewed-head write, and keep those operations in one guarded invocation. +BLOCKER | high | Task 15 steps 4b–4 and step 6 | Step 4 runs the quality battery over the mutable working tree rather than the captured candidate, and step 6 checks only that `HEAD` stayed fixed; no check binds the battery's before-and-after worktree bytes to `.context/loop-rule-reviewed-head`. | A dirty helper change can make the battery pass for bytes absent from the candidate and then be reverted before closing, allowing a candidate that fails CI or invariant checks to be reviewed and shipped on evidence from a different tree. | Run the battery in a detached disposable checkout of the recorded candidate, or establish and recheck that the index and tracked worktree equal that candidate before and after the battery. +MINOR | high | Task 0 step 3 | The step promises a baseline for every inventoried and changed site, but `.context/loop-rule-sites` covers only the ten inventoried a–j passages; changed prompt-copy sites outside that inventory, including §G and §F item 18's work-loop line, receive no `site` record, and Resume's completeness rule checks only inventoried sites. | The baseline can be accepted as complete without a parent comparison for changed blocks that Task 14 later classifies, so inherited and introduced differences at those omitted regions are no longer distinguishable by the artifact built for that purpose. | Derive baseline rows from the target text's full prompt-copy destination-block set as well as the a–j inventory, and validate completeness against that derived set. +NIT | high | Task 15 step 4b commit block | The prose says a prompt-standards repair uses the subject `WIP: prompt-standards repair + plan records`, but the only command always commits `WIP: plan records`. | A repair commit is indistinguishable by subject from a records-only commit, weakening the audit trail the preceding instruction promises even though routing still treats both as WIP snapshots. | Select the subject from whether repaired files were staged, or remove the contradictory subject requirement. +END OF FINDINGS (4 total) diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-53-dispositions.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-53-dispositions.md new file mode 100644 index 0000000..46a9180 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-53-dispositions.md @@ -0,0 +1,52 @@ +# Pass 53 dispositions — Gate-A plan cycle `om0bdd7udh` + +Advisory companion per CLAUDE.md §5. Not the findings file; not part of pass validation. + +**Snapshot.** `HEAD` `5fb3b942f1907f5d6dd25b32d15b497f855e05f8`, branch `loop-rule-consolidation`, +plan worktree blob `769fbc61553fcc8b1d6069c3343282d78028e35c` — identical before and after the call. +Tool: `mcp__codex__exec`, session `01a0d94e-ecfb-7390-8f8d-b299712092dc`. Instruction: +`.context/gate-a-plan-pass-53-instruction.md` (pointed at, contract restated inline; that Codex read +the file in full is not established). + +**Pass VALID.** Terminator `END OF FINDINGS (7 total)`, 7 body lines, 6 fields each, severity tokens +valid. **7 findings: 2 Blockers, 2 Majors, 2 Minors, 1 Nit.** + +**Floor 3**, read fresh: story `docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md`, +Risk `high` (2), Security `none` (0) → 3. One cited story, from the plan's `**Story:**` header. + +## Findings — validated against the plan text, not repaired (the release allowed pass 53 and a report, nothing more) + +1. **BLOCKER, 4b failure route (2918–2921). CONFIRMED — and created by revision 24.** The route still + demands "the whole step-4 battery" before committing the repair, but the battery now reads + `.context/loop-rule-reviewed-head`, which 4b writes only after its commit. First repair: file + absent → exit 2. Later repair: it names the previous candidate. This is regeneration: my Blocker-2 + fix changed the battery's input without re-ordering the one route that runs it before the commit. +2. **BLOCKER, rerun sequence (3724–3730) vs step three's no-repair route. CONFIRMED in the text.** The + six-step list says "Commit them" and "capture that commit" unconditionally; step three has an + explicit no-repair/no-commit route that takes the records commit. Pre-existing; not touched by + revision 24. +3. **MAJOR, Resume four-source block (785–792). CONFIRMED.** The fence uses `$BASE` and assigns it + nowhere; the plan states each fence is its own shell. Pre-existing. +4. **MAJOR, step two source-block release into Close (3496–3510). PLAUSIBLE, not fully traced.** The + claim is that step one's records commit already moved `HEAD` when a source-block release routes the + same pass to Close, so Close condition 1 rejects. Reading the ordering supports it; the full route + through Close conditions 1 and 2 was not walked end to end. +5. **MINOR, 4b rationale on `reset --soft` (2934–2939).** Collected. +6. **MINOR, 4c merge failure always reported as conflict (3133–3137).** Collected. +7. **NIT, 4c fixture table says eight cases, lists nine rows.** Collected. + +## Loop health at pass 53 (owed from pass 4) + +- **Trend.** Findings 42–53: 3, 5, 9, 2, 1, 2, 3, 6, 4, 6, 4, **7**. Blockers: 2, 2, 2, 1, 1, 1, 3, 5, + 2, 2, 2, **2**. Majors: 1, 2, 7, 1, 0, 1, 0, 1, 1, 2, 0, **2**. +- **Cluster.** All seven findings sit in the plan's execution machinery (Task 15 flow, Resume, 4c); + none touches the §5 target text the change installs. +- **require↔withdraw.** None found: finding 1 asks to move the battery after the capture, which + revision 24 did not forbid. + +**Tells present: 1 (count rising, 4 → 7), 2 (Blockers flat at 2 for four passes), 3 (instrument +cluster, total). Stop-and-surface is mandatory.** One Blocker (finding 1) regenerated from this +revision's own repair. The "clearly stuck" exit is not claimed: no affirmative coverage judgement is +possible while repairs keep creating the next Blocker. + +**Cycle stays OPEN. No repair made, no pass 54.** The release covered pass 53 and a report only. diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-53.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-53.md new file mode 100644 index 0000000..b7bdfd0 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-53.md @@ -0,0 +1,8 @@ +BLOCKER | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:2918-2921, Task 15 step 4b failure route | The failure route still requires the whole step-4 battery before committing the repaired text and refreshed record, but the revised battery reads the candidate file that this step creates only after its commit at lines 2963-2976. On the first such repair the candidate file does not exist; on a later repair it names the previous candidate. | The executor cannot complete the first prompt-standards repair in the prescribed order, and a later pre-commit battery run checks the old candidate rather than the repair. | Keep the affected record-producing checks before the commit, but move the failure route's required battery run after the guarded repair commit and candidate capture, matching the six-step sequence. +BLOCKER | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:3724-3730, Task 15 rerun sequence | The sequence asserted to govern every rerun unconditionally says to repair, commit the records, and capture that commit, despite step three's explicit no-repair route at lines 3517-3524 and 3586-3600 requiring no commit and selecting the existing records commit. A below-floor Minor-only pass can leave all rerun observations unchanged after step one has already committed its findings. | Literal execution of the later sequence recreates the empty-commit deadlock that pass 50 repaired; the executor must choose which of two mandatory procedures to disregard before issuing the next pass. | Express the sequence through step three's two obligation-based routes, explicitly preserving the no-repair/no-commit route and its records-commit candidate, and reconcile the record-refresh duty with that route without dropping the complete rerun checks. +MAJOR | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:785-792, Resume four-source reconciliation | This fenced block consumes BASE without loading it. Its assignments at lines 640 and 703 belong to earlier fences, while the plan explicitly treats fenced blocks as separate shell invocations. With BASE unset, the log range becomes HEAD..HEAD and git diff --cached receives an empty revision. | An ordinary valid re-entry reports an empty commit range and then stops on the staged-content read, so the executor cannot perform the required four-source reconciliation. | Read and guard .context/loop-rule-base inside this block before either BASE-dependent command, as the other independent consumers do. +MAJOR | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:3496-3510, Task 15 step two source-block release and clean-completion route | Step one has already advanced HEAD with the pass's records before step two can release a source block and route the same pass to Close. For example, a zero-finding pass held by an unreadable Story header can become closable when access to the unchanged header is restored, which target A1 explicitly permits without another pass; nevertheless Close condition 1 rejects because HEAD is now the records commit rather than the reviewed head, and condition 2 also expects findings that step one has already committed. | The executor cannot follow the specified same-pass source-block release into closure; the plan's own bookkeeping forces a stop or an otherwise unnecessary review even though the source repair moved no pass-cost value. | Reconcile recording with this route: determine whether the source-block release can close before making a non-closing records commit, or provide a continuation that validates and uses the existing records commit under the closing requirements. Do not rewrite the reviewed-head marker merely to make the mismatch disappear. +MINOR | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:2934-2939, Task 15 step 4b commit rationale | The paragraph says reset --soft stages only what discarded commits contained, and that staging after review breaks the reviewed-head condition. A soft reset leaves the index unchanged, including previously staged content, and staging alone does not change HEAD. Close condition 5 correctly states the former elsewhere. | The executor receives contradictory explanations of what reaches the squash and which check catches post-review staging; the dirty-set and clean-tree checks, rather than the asserted Git behavior, are what prevent that content from closing. | Replace this rationale with the actual index behavior and identify the dirty-set/clean-tree checks for staged post-review changes, while retaining the requirement to commit repairs before their review. +MINOR | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:3133-3137, Task 15 step 4c merge failure | Every nonzero merge result is reported as a conflict after both output streams have been discarded. The merge can also fail without a conflict, for example when inherited commit-signing configuration cannot sign the synthetic merge commit. | The executor gets a false conflict diagnosis and loses the diagnostic needed to distinguish a content conflict from a local setup failure, even though both correctly stop progression as unresolved. | Preserve the merge diagnostic and report an unresolved merge failure without asserting a conflict unless that cause was established. +NIT | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:3217-3231, Task 15 step 4c fixture report | The introduction claims eight cases, but the table has nine case rows, including both an unresolvable pinned head and an unresolvable candidate head read from the file. | The executor cannot tell whether the last row is an additional executed case or a duplicate description of an existing one, so the stated evidence count is ambiguous. | Deduplicate the two head-resolution rows if they describe one case, or update the count and distinguish their setups if they are separate cases. +END OF FINDINGS (7 total) diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-54-dispositions.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-54-dispositions.md new file mode 100644 index 0000000..1b8e540 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-54-dispositions.md @@ -0,0 +1,52 @@ +# Pass 54 dispositions — Gate-A plan cycle `om0bdd7udh` + +Advisory companion per CLAUDE.md §5. Not the findings file; not part of pass validation. + +**Snapshot.** `HEAD` `5fb3b942f1907f5d6dd25b32d15b497f855e05f8`, plan blob +`2d7eeff004758108222f9f921c78f72c4dacdd28` — identical before and after the call. Tool +`mcp__codex__exec`, session `01a0d9e6-2fca-73a3-b15f-a58682ac11a8`, instruction +`.context/gate-a-plan-pass-54-instruction.md` (pointed at; full read not established). + +**Pass VALID.** `END OF FINDINGS (7 total)`, 7 lines, 6 fields each. **0 Blockers, 4 Majors, 1 +Minor, 2 Nits.** Floor 3 (story: risk high, security none), read fresh. + +## Findings — checked against the plan and sources; nothing repaired (release: one pass, then report) + +1. **MAJOR — which rules govern this Gate-B cycle. CONFIRMED, and it is a contract question.** + Revision 25 says §5 at `$BASE` governs *and* keeps the reading against the installed ordering. They + conflict on a real case. At `$BASE`, `CLAUDE.md:266–268` makes any two tells a mandatory stop. Target + §A (`target-text:206–212`) lets tells on a closing pass go into the report without blocking the + close. I called the ordering read "additive"; it is not. **Deciding which governs is Daniel's call.** +2. **MAJOR — the no-repair branch says "no rerun", but the complete-set rule (and accounting row 27) + still requires a full rerun and a recommit before the candidate final pass. CONFIRMED; my own + sentence in revision 25 caused it.** In-set correction of the correction. +3. **MAJOR — the step-4b commit block's staging set names only the two prompt copies, the hook and its + test, plus the plan.** A spec repair (the "fix that changes specified behaviour updates the spec" + rule) or a version repair from 4c has no named path. **CONFIRMED** (plan 2836–2842). This comes from + revision 25 making 4b the only candidate writer. In-set. +4. **MAJOR — step 8 tells the executor to delete the working record before the closing block.** §5 + retires it at closure and makes it the recovery source while the cycle runs. An interruption + between the deletion and a successful commit loses the cycle's identity source. **CONFIRMED; + revision 25 caused it** (plan 420–422 and step 8's preamble). In-set. +5. **MINOR** — stale tip/8b/condition-6/"four predicates" wording in the reassessment table and the + postcondition notes. Collected. +6. **NIT** — accounting rows 7 and 40 still say "the only two deletions". Collected. +7. **NIT** — "six places cite step 4b" is stale. Collected. + +## Loop health at pass 54 + +- **Trend.** Findings 50–54: 4, 6, 4, 7, **7**. Blockers: 2, 2, 2, 2, **0**. Majors: 1, 2, 0, 2, **4**. +- **Cluster.** All seven findings are about the plan's own execution procedure (instrument), and three + of them are prose about it. None touches the §5 target text. +- **require↔withdraw.** Finding 4 asks to keep the working record that revision 25 told the executor to + delete. That undoes revision 25's own addition, not an earlier pass's. + +**Tells present:** instrument cluster (yes); prose cluster (partly, 3 of 7); count rising (no, flat +at 7); Blockers failing to fall (no, 2 → 0); require↔withdraw (no, per above). At least the instrument +tell is present, and a prose tell arguably too. **Finding 1 is in any case a new contract question, +which stops the loop under §5 on its own.** + +**Three of the four Majors were introduced by revision 25**, which is the regeneration pattern again, +at lower severity: Blockers went from 2 to 0. + +**Cycle stays OPEN. No repair made, no pass 55.** diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-54.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-54.md new file mode 100644 index 0000000..f1096e0 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-54.md @@ -0,0 +1,8 @@ +MAJOR | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:3236–3265; Close:400–407 | The plan makes §5 at $BASE govern this cycle, then requires the same pass to follow the incompatible ordering being installed. At $BASE, CLAUDE.md:266–268 makes any two tells a mandatory stop-and-surface; target §A:206–228 lets a successful clean completion outrank that stop. The retained bullets also put any open suspension ahead of clean completion while calling this the installed ordering. | A clean floor-reaching pass with two tells has conflicting instructions to close or await a user decision; calling the new reading additive does not resolve which duty governs. | Use the starting rules consistently for this cycle and identify which new practices are compatible additions; if a conflicting new rule is intended to govern, explicitly resolve that transition instead of claiming both rule sets apply. +MAJOR | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:3274–3279,3303–3321; accounting row 27 | The no-repair continuation says the record-producing set owes no rerun and nothing may be committed, but the retained final-pass obligation unconditionally requires the complete mechanical and reader set to be rerun, re-recorded and committed before the candidate final pass; row 27 still labels that obligation unchanged. | A below-floor pass carrying only Minors can continue without repair into the potentially final pass, but the executor must either skip a stated verification obligation or depart from the explicit no-rerun/no-commit route; refreshed dirty records would also fail Close condition 2. | Reconcile the complete-set obligation with the no-repair branch in its governing paragraph and accounting row, explicitly stating when existing records suffice and when a records-only candidate must be committed before verification. +MAJOR | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:2834–2851,3271–3274,3289–3293 | Every Gate-B repair must now commit through step 4b, whose prose and command comment restrict repaired files to the two prompt copies, hook and test, plus this plan. Step 7 nevertheless requires a behavior-changing repair to update the spec in the same commit and explicitly includes target-text repairs among possible fixes; the spec is outside that staging set. | A required spec repair cannot be committed through the sole authorized candidate writer as written, leaving it outside the reviewed range and dirty at closure or forcing the executor to disregard the staging restriction. A version repair identified by 4c has the same problem for the manifest and changelog. | Keep the guarded candidate commit block, but make its named staging inputs depend on the authorized repair route, including the spec and any required version files, while still staging only explicitly identified files. +MAJOR | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:420–422,3437–3439 | The plan requires deletion of this cycle's working record before the closing block runs. Governing CLAUDE.md:363–364 deletes it once the cycle closes, and :405–417 makes that record the positive identity source while the cycle runs, with history becoming a source only once the cycle's own commit exists. | An interruption after this deletion and before a successful closing commit, including a failed commit after reset --soft, deliberately removes a recoverable cycle's identity source while it is still open. Findings filenames and the ignored closing-message draft are not authorized substitutes under the governing recovery rule, so a resumed session can be forced to start a new cycle. This loss is caused by the normal procedure, beyond D1's disclosed exposure to an external .context clear. | Retain an existing working record until successful closure is established; narrowly account for that advisory path in the dirty-set/post-close cleanup ordering without staging it into the closing commit or introducing a new backup mechanism. +MINOR | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:317,3595–3608 | The current reassessment still says the closing tip survives as 8b's precondition and step 7 clears it, although D1 deleted both the tip and 8b. The postcondition explanation also points to a nonexistent condition-6 block and says four predicates where condition 4 now enumerates five. | The recovery and cleanup explanations direct an executor to deleted state and a deleted condition number, and misstate the verification actually performed by the retained block. | Update the reassessment to the surviving reviewed-head comparison and refer to condition 4's postcondition block and five checks; remove statements about writing or clearing the deleted tip. +NIT | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:248,281 | Accounting row 7 says rows 20 and 40 are still the only conditions deleted, and row 40 calls itself the second of two deletions, but rows 34 and 35 are now explicitly dropped under D1 and the paragraph at line 297 acknowledges them. | The condition-preservation audit contradicts its own current dispositions and understates what D1 removed. | Make these counts explicitly historical or remove them and reference the current dropped rows. +NIT | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:2788–2790 | The reason for retaining the name 4b says six places elsewhere cite it, but the current plan has more than six such references, including the battery, 4c, evidence assembly, call preparation, continuation, rerun duties and output section. | The stated rename scope is already stale and would understate the references an executor must update. | Remove the fixed count and retain the instruction to locate and update every reference if the step is renamed. +END OF FINDINGS (7 total) diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-55-dispositions.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-55-dispositions.md new file mode 100644 index 0000000..1e9fca2 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-55-dispositions.md @@ -0,0 +1,42 @@ +# Pass 55 dispositions — Gate-A plan cycle `om0bdd7udh` + +Advisory companion per CLAUDE.md §5. Not the findings file; not part of pass validation. + +**Snapshot.** `HEAD` `5fb3b942f1907f5d6dd25b32d15b497f855e05f8`, plan blob +`f34d03267b9a4387b8121b76b773c6eb7281d71a` — identical before and after the call. Tool +`mcp__codex__exec`, session `01a0da21-a62f-7d83-bb43-d0a00cecad82`, instruction +`.context/gate-a-plan-pass-55-instruction.md` (pointed at; that it was read in full is not established). +The tool reply listed no created file, but the slot exists (written 22:03) and validates. + +**Pass VALID.** `END OF FINDINGS (3 total)`, 3 lines, 6 fields each. **0 Blockers, 2 Majors, 1 +Minor.** Floor 3 (story: risk high, security none), read fresh. + +## Findings — checked against the text; nothing repaired (release: one pass, then report) + +1. **MAJOR — the no-repair branch stops on "anything else dirty"** (plan 3310). Under D1 this cycle's + findings files are always dirty during the loop, and so is the working record. The branch + therefore stops on the ordinary Minor-only continuation. **CONFIRMED; revision 26 caused it.** + In-set: correct the stop to name only paths outside the cycle's findings slots, the working + record and this plan. +2. **MAJOR — step 7's routing has no route for a clean or zero-finding pass that owes a re-review + because the evidence entry changed** (7b, plan ~3438) or because Failure says a closure condition + costs a pass. The repair branch needs Blocker/Major, the no-repair branch needs "below the floor", + and clean-at-floor leads back to 7b. **CONFIRMED.** Revision 26's routing list introduced the gap + (the old "continue" bullet covered it implicitly). In-set: add a route "a re-review is owed by the + evidence or a closure-condition rule → the candidate sequence, repairing only what that rule + requires". +3. **MINOR — Task 7 step 3's "four kinds" list omits `a13`'s first-sentence absence check** (plan + 1992 vs 1918–1921). Pre-existing, outside Task 15. Collected. + +## Loop health at pass 55 + +- **Trend.** Findings 51–55: 6, 4, 7, 7, **3**. Blockers: 2, 2, 2, 0, **0**. Majors: 2, 0, 2, 4, **2**. +- **Cluster.** Both Majors are in the plan's own loop procedure (instrument). The Minor is a + verification checklist of a product task. +- **require↔withdraw.** None. + +**Tells:** count falling (no), Blockers at 0 (no), instrument cluster (yes), prose cluster (no), +require↔withdraw (no). **One tell, so no mandatory stop from tells.** Both Majors came from the +last revision, at lower severity and smaller scope than before. + +**Cycle stays OPEN. No repair made, no pass 56** — the release covered one pass and a report. diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-55.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-55.md new file mode 100644 index 0000000..f715152 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-55.md @@ -0,0 +1,4 @@ +MAJOR | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:3310 (Task 15 step 7, no-repair candidate branch) | The instruction to stop on "Anything else dirty" exempts only this plan, although D1 requires the just-validated findings files to remain uncommitted throughout the loop; .gitignore explicitly exposes .context/codex-reviews/, and the optional working record and dispositions notes can also be present there. | An ordinary below-floor pass containing only Minor/Nit findings reaches the required no-repair continuation with dirty findings files and must stop instead of issuing the next pass; committing those files to clear the stop would violate D1. | Limit this stop to unauthorized dirty paths, explicitly allowing the exact current-cycle findings and permitted recovery/advisory artifacts to remain uncommitted; keep them out of the records-only commit. +MAJOR | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:3438 (Task 15 step 7b, changed-evidence route back to step 7) | A changed evidence entry invalidates the candidate and sends the same pass back through step 7, but that router has no route for an otherwise clean or zero-finding pass that owes another review because evidence changed. Its repair branch requires Blocker/Major findings, its no-repair continuation requires being below the floor, and its clean/zero branch leads straight back to 7b. | For a clean pass at the floor with changed evidence and no stop trigger, the required re-review has no executable route: following the bullets either loops through closure preparation or risks treating the rebuilt entry as validated without a new call, contrary to CLAUDE.md at BASE lines 725–730. Failure also returns to this router when a closure condition costs another pass. | After honoring the governing stop predicates, explicitly route a pass that owes re-review under an evidence or other closure-condition rule into the candidate sequence, authorizing only the repair that rule requires; permit Close only when no such pass is owed. +MINOR | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:1992 (Task 7 step 3, "four kinds" list) | The list claims to enumerate everything step 1b added but omits the standalone absence check for a13's first sentence, expressly required at lines 1918–1921. Its listed absence checks cover only the four moved conditions a17–a20. | The detailed execution checklist conflicts with the earlier obligation and can lead an executor following that list to leave the obsolete "This replaces the pass-count number and nothing else" sentence unobserved, although the general rerun duty still requires its check. | Include the a13 first-sentence absence observation in both copies with parent=1 and worktree=0, or replace the exhaustive four-kind claim with a direct instruction to run every observation derived in step 1b. +END OF FINDINGS (3 total) diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-56-dispositions.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-56-dispositions.md new file mode 100644 index 0000000..4534677 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-56-dispositions.md @@ -0,0 +1,28 @@ +# Pass 56 dispositions — Gate-A plan cycle `om0bdd7udh` + +Advisory companion per CLAUDE.md §5. Not the findings file; not part of pass validation. + +**Snapshot.** `HEAD` `5fb3b942f1907f5d6dd25b32d15b497f855e05f8`, plan blob +`79649d37e89ebd07b60e53d6be5f44a4ccb1cbb5` — identical before and after the call. Tool +`mcp__codex__exec`, session `01a0dc89-184b-7d10-86d4-d0d659130526`, instruction +`.context/gate-a-plan-pass-56-instruction.md` (pointed at; that it was read in full is not established). +The slot was confirmed absent before the call and was written at 09:14 during it. The tool reply +listed no created file. + +**Pass VALID and CLEAN.** The file is exactly `NO FINDINGS` followed by `END OF FINDINGS (0 total)`. +**0 findings.** Floor 3 (story: risk high, security none), read fresh; pass 56 is above the floor. + +## What this does and does not mean + +- **Closure-eligible under §5 as this cycle started**: a clean pass at or above the floor. Blockers + and Majors of every earlier pass are resolved or dispositioned in their dispositions files. + Collected Minors and Nits stay collected. +- **The cycle is NOT closed.** Closing a Gate-A plan cycle is a commit of the reviewed plan with its + provenance line and per-pass curve (§5 Mechanics). This release did not authorize a commit. The + curve covers passes 1–56 and has to be built from the tracked and untracked findings files. +- Pass 56 reviewed the plan **text**. It is not evidence that the plan executes as written. The + executed shell checks cover only the blocks named in the resume record. + +## Loop health at pass 56 + +Findings 52–56: 4, 7, 7, 3, **0**. Blockers: 2, 2, 0, 0, **0**. Majors: 0, 2, 4, 2, **0**. No tells. diff --git a/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-56.md b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-56.md new file mode 100644 index 0000000..8d781f4 --- /dev/null +++ b/.context/codex-reviews/gate-a-plan-om0bdd7udh-pass-56.md @@ -0,0 +1,2 @@ +NO FINDINGS +END OF FINDINGS (0 total) diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index bc68575..79649d3 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -28,7 +28,7 @@ It adds no task, prerequisite or closure condition to this plan. - **Invariant 11:** every prompt change passes all 12 items of `docs/prompt-standards.md`. Most at risk: item 6 (every constraint carries its reason in the same sentence) and item 8 (token-lean). - **Invariant 4 / the hook:** no control flow, counter, fingerprint computation, routing or event handling changes. Only `note` strings and their expectations. - **Gate-B cycle discipline:** snapshot commits are named `WIP: …`, and a non-`WIP` commit mid-cycle resets the hook's counters. **This change closes with `git reset --soft "$BASE"` followed by one commit, not with `--amend`** — Mechanics prescribes the reset shape wherever several WIP snapshots piled up, and this plan makes one per task. Task 15 step 8 is the operation. -- **Every `git add` in this plan is guarded, and a guard is where most of them stop.** A failed staging otherwise falls through to the `git commit` on the next line, which **can still succeed on whatever the index already held** — this plan states in Task 15 step 7 that the index may carry foreign staged content, and Resume deliberately admits a live staged index — so the block reports a snapshot it did not take. **A content comparison after the commit is added only where a later step reads that commit rather than the worktree**, and that is three places: step 7's step one, whose record step two then reads; step 7's step three and step 4b, whose commits bound the range the next Gate-B call reviews. The per-task snapshots have no such consumer — every check this plan runs afterwards reads the worktree, and Gate B reads the accumulated `$BASE..HEAD` range rather than any one commit — so they are guarded and not pinned. **8a is the exception that needs no addition**: Close condition 4 already pins its blobs before staging and compares them against the committed tree. +- **Every `git add` in this plan is guarded, and a guard is where most of them stop.** A failed staging otherwise falls through to the `git commit` on the next line, which **can still succeed on whatever the index already held** — this plan states in Task 15 step 7 that the index may carry foreign staged content, and Resume deliberately admits a live staged index — so the block reports a snapshot it did not take. **A content comparison after the commit is added only where a later step reads that commit rather than the worktree**, and that is one place: the step-4b commit block, which every Gate-B candidate uses, first pass and repairs alike, and whose commit bounds the range the next call reviews. The per-task snapshots have no such consumer — every check this plan runs afterwards reads the worktree or the step-4b candidate, and Gate B reads the accumulated `$BASE..HEAD` range rather than any one commit — so they are guarded and not pinned. **The closing commit is the other one**: Close condition 4 compares it against pins taken before staging. --- @@ -217,15 +217,13 @@ live in the home named. Where a task, a shell comment or an error string needs a | Branch, clean tree, no base file yet, `ba15e83` ancestral, approved inputs unchanged at `HEAD` | **Preparation** | Task 0 step 1, first-entry branch | | Branch; base file shape, self-resolving commit, ancestral; approved inputs unchanged at `$BASE`; scratch-artifact validity; how far the implementation got | **Resume** | Task 0 step 1, re-entry branch | | What a re-entry reads, and that no rule depends on the state fitting a named shape | **Resume** — its table is illustration, not a classification | accounting row 7; every state-reading rule | -| `HEAD` equals the reviewed head | **Close**, condition 1 | step 8a | -| The dirty set is exactly the candidate pass's findings files | **Close**, condition 2 | step 8a | -| The closing message is complete, and revalidated before it is consumed | **Close**, condition 3 | step 7b writes it; steps 8a and 8b check it | -| The record commit's paths, parent, blob identity, and the findings files' structure and eligibility | **Close**, condition 4 | step 8a | -| Clean tree and the closing tip when the closing invocation begins | **Close**, condition 5 | step 8b | -| After the closing commit: subject, parent, clean tree, landed body | **Close**, condition 6 | step 8b's post-close block | -| Two invocations, and why the hook requires them | **Close** | steps 8a and 8b | -| What happens on a failed closing operation or postcondition | **Failure** | every rejection in steps 8a and 8b | -| What a successful close removes | **Close**, its success recognition | step 8b's cleanup block | +| `HEAD` equals the reviewed head | **Close**, condition 1 | step 8, closing block | +| The dirty set is exactly this cycle's findings files | **Close**, condition 2 | step 8, closing block | +| The closing message is complete, and revalidated before it is consumed | **Close**, condition 3 | step 7b writes it; step 8's closing block checks and pins it | +| After the closing commit: subject, parent, clean tree, landed body, committed findings blobs | **Close**, condition 4 | step 8, postcondition block | +| One closing command with no `-m`, and why | **Close** | step 8, closing block | +| What happens on a failed closing operation or postcondition | **Failure** | every rejection after the reset in step 8 | +| What a successful close removes | **Close**, its success recognition | step 8's postcondition block | **Straight-line commands stay where the work happens.** Task 15 step 8 keeps the shell that runs these checks — it is verified and it is what an executor types — and each block names the condition @@ -262,24 +260,24 @@ independent reader did. | 19 | The baseline covers every inventoried site and is read whole | **kept**, and **strengthened**: per-site literal-uniqueness, escaped anchors and non-empty extraction, as Task 14 already requires (pass 21) | | 20 | Task 0 commits nothing | **proposed for deletion, and deleted.** It now commits one record — the fragment sweep (pass 21). A validation whose result is not committed cannot be told from one that never ran, and every other reader check in this plan already commits its record | | 21 | `baseSha` is `$BASE`, never `HEAD^` | **kept** — the rule, unchanged | -| 22 | `headSha` resolved to the full object name at that moment, kept with each branch's result | **kept**, and **extended**: it is persisted, because 8a needs it (pass 21) | +| 22 | `headSha` resolved to the full object name at that moment, kept with each branch's result | **kept**, and **extended**: it is persisted, because the close needs it (pass 21); its consumer is Close condition 1 | | 23 | A fresh nonce for the Gate-B cycle | **kept** — the rule, unchanged | -| 24 | Each fix is committed before the next review | **kept**, and **widened**: every non-closing pass commits, repair or not (pass 21) | +| 24 | Each fix is committed before the next review | **kept.** The pass-21 widening — every non-closing pass commits its findings, repair or not — is **DELIBERATELY DROPPED by Daniel's decision D1 of 2026-09-25**: findings files are committed once, at the close. Consequence stated at Close, *D1* | | 25 | Every affected check re-runs after each fix, mechanical and reader | **kept** — the rule, unchanged | | 26 | Re-run records are committed before the re-review | **kept** — the rule, unchanged | -| 27 | The complete set re-runs before the candidate final pass | **kept** — the rule, unchanged | +| 27 | The complete set re-runs before the candidate final pass | **kept, and made explicit** (pass 54): before every call, since any call can be final. Records that come out unchanged are not recommitted; records that change enter a records-only candidate before final verification | | 28 | Only a clean response against that exact `HEAD` closes | **kept**, and **made checkable**: the reviewed head is recorded, so "that exact `HEAD`" has a value to compare against (pass 21) | | 29 | Closure-eligible = clean at or above the floor **or** zero-finding | **kept** — the rule, unchanged | -| 30 | The final pass's own findings files are the sole permitted post-review addition | **kept** — the rule, unchanged | +| 30 | The final pass's own findings files are the sole permitted post-review addition | **changed by D1**: this cycle's findings files, every pass's, are the sole permitted addition at the close | | 31 | The closing message is rebuilt whole; exactly one provenance line and one curve | **kept** — the rule, unchanged | | 32 | Every owed record is present before the close | **kept** — the rule, unchanged | -| 33 | The dirty set is exactly the final pass's findings files, read from every porcelain record | **kept** as the close procedure's check, and **corrected**: it asks git which paths outside the expected pair are dirty instead of splitting status output into pathnames, which dropped a newline-bearing path invisibly and mangled a rename's second field (pass 34) | -| 34 | The record commit's changed-path set equals those files | **kept** as the close procedure's check | -| 35 | The tree is clean after the record commit | **kept** as the close procedure's check | -| 36 | `HEAD` equals the reviewed tip before the reset | **kept** as the close procedure's check, and **corrected**: the reviewed *head* and the closing *tip* are two values, not one (pass 21). With row 40 dropped the tip is **only** this precondition — nothing resets to it | +| 33 | The dirty set is exactly this cycle's findings files (changed by D1 from the final pair), read from every porcelain record | **kept** as the close procedure's check, and **corrected**: it asks git which paths outside the expected pair are dirty instead of splitting status output into pathnames, which dropped a newline-bearing path invisibly and mangled a rename's second field (pass 34) | +| 34 | The record commit's changed-path set equals those files | **dropped with its subject** (D1): there is no record commit. What it guarded — only findings files enter beside the reviewed tree — is Close condition 2 plus condition 4's blob check | +| 35 | The tree is clean after the record commit | **dropped with its subject** (D1): there is no record commit | +| 36 | `HEAD` equals the reviewed tip before the reset | **merged into Close condition 1** (D1): with no record commit the closing tip and the reviewed head are the same object, and condition 1 is read in the same block as the reset | | 37 | The closing commit's subject is not a snapshot | **kept** as the close procedure's check | | 38 | The tree is clean after the close | **kept** as the close procedure's check | -| 39 | 8a and 8b are separate invocations; no `-m` in the closing one | **kept**, and it is why the close procedure names two invocations | +| 39 | 8a and 8b are separate invocations; no `-m` in the closing one | **kept in part** (D1): no `-m` in the closing command. The split is dropped, because no `WIP:` commit shares a command string with the close any more | | 40 | Every rejection restores a tip by **mixed** reset, never `--hard` | **DELIBERATELY DROPPED**, by Daniel's decision, and the second of two deletions this table records. **It was never a requirement of the approved spec** — §A says *"re-establish every closure condition against the repository as it now stands"*, which reads the state rather than rewinding it, and a rejected commit at `HEAD` is a legitimate starting point for that reading. It was this plan's own implementation choice, and it had grown restore targets, phase selection, pre- and post-act captures and their own error handling, which four consecutive passes then found defects in. **What replaces it is a bounded handoff**: stop every further mutating action, report the failed step and the observed state, change nothing else. **The obligations §A does impose are unchanged** — surface the failure, re-establish every closure condition against the state as it stands, and take one of its three routes. **Not replaced by a generic backup mechanism**, which would be the same growth under another name | | 41 | The scratch files are removed only after a successful close | **kept** as the close procedure's last step | @@ -287,8 +285,8 @@ independent reader did. has exactly one home, named in the index above; what went is the **third copy**, the prose beside Task 15 step 8's shell that repeated in words what the block already ran and the condition already defined. Two rationales that existed only in that prose were moved into the conditions they belong -to: `reset --soft`'s index behaviour into condition 5, and why the dirty set can be exact into -condition 2. **The one correction is condition 6's landed-body check**, which compared bytes where +to: `reset --soft`'s index behaviour into the then condition 5, and why the dirty set can be exact into +condition 2. **The one correction is the landed-body check (then condition 6, now 4)**, which compared bytes where `git log --pretty=%B` adds a trailing newline the source file has none of — it rejected a *correct* close. **The predicate is unchanged and is now true as stated**: the body is extracted with `--pretty=format:%B`, which emits the stored message alone, and the closing commit is made with @@ -296,7 +294,7 @@ close. **The predicate is unchanged and is now true as stated**: the body is ext a disposable repository, in both directions. An intermediate revision normalized trailing blank lines on both sides instead; that made the check's own predicate false and is gone. -**Two conditions are dropped, 20 and 40, and each is named as a drop rather than lost.** Condition +**Rows 34 and 35 left with their subject, and row 24's pass-21 widening was dropped, by Daniel's decision D1 of 2026-09-25** (Close, *D1*). **Before that, two conditions were dropped, 20 and 40, and each is named as a drop rather than lost.** Condition 20 is replaced by a stronger obligation. **Condition 40 is dropped outright** — a guard this plan introduced itself, which the approved spec never asked for, removed by an explicit decision after it became the largest single source of findings in the cycle. **A guard the plan invented does not @@ -340,15 +338,35 @@ state permanently unresumable through this plan's own success path. approved inputs still hold their approved blobs at the revision the tasks are derived against; and that **no base file exists yet** — if one does, this is a re-entry and Resume owns it. +**Every one of those predicates takes the same captured revision as its subject**, and Preparation +records it to `.context/loop-rule-start`. That file is the whole point: without it the checks here +and the base Task 0 step 1 records are statements about **different resolutions of a movable ref**, +and nothing binds the revision that was approved to the revision the work proceeds from. + ```bash +# Resolve the starting revision ONCE and make it the subject of every predicate below. +# `HEAD` is a movable ref: checking the branch, the ancestry and the three approved +# blobs through separate live resolutions and then recording the base from yet another +# one binds the recorded base to NO commit whose properties were checked. Each fenced +# block is its own shell, so this value is written to a file for Task 0 step 1 rather +# than carried in a variable. +START=$(git rev-parse HEAD) \ + || { echo "cannot resolve the starting revision; stop"; exit 1; } +test "${#START}" -eq 40 || { echo "starting revision is not a full object name: $START"; exit 1; } +test "$(git rev-parse --verify "$START^{commit}")" = "$START" \ + || { echo "starting revision does not resolve to itself as a commit"; exit 1; } +# The symbolic check catches a detached HEAD, which prints `HEAD`; the ref check binds +# that branch to the captured object. Both, because neither implies the other. test "$(git rev-parse --abbrev-ref HEAD)" = loop-rule-consolidation || { echo "wrong branch"; exit 1; } +test "$(git rev-parse --verify refs/heads/loop-rule-consolidation)" = "$START" \ + || { echo "the branch is not at the starting revision"; exit 1; } # Two steps, because a FAILED `git status` prints nothing and `test -z ""` would read # that as a clean tree. The predicate is unchanged; only its establishment is. TREESTATE=$(git status --porcelain) \ || { echo "reading the tree state FAILED — cleanliness is unestablished; stop"; exit 1; } test -z "$TREESTATE" || { echo "tree not clean — if this is a re-entry, run Resume"; exit 1; } test ! -e .context/loop-rule-base || { echo "a base is already recorded — this is a re-entry, run Resume"; exit 1; } -git merge-base --is-ancestor ba15e83 HEAD || { echo "ba15e83 not in this history"; exit 1; } +git merge-base --is-ancestor ba15e83 "$START" || { echo "ba15e83 not in this history"; exit 1; } for f in target-text design condition-inventory; do p="docs/superpowers/specs/2026-09-10-loop-rule-consolidation-$f.md" # Resolve separately and check each status. `git rev-parse` without `--verify` @@ -356,13 +374,17 @@ for f in target-text design condition-inventory; do # exactly when the two revision expressions are equal — and either way the result # says nothing, because success must depend on both exit statuses. (Observed: # `git rev-parse HEAD:missing` prints `HEAD:missing`, status 128.) - HEAD_ID=$(git rev-parse "HEAD:$p") \ - || { echo "$p does not resolve at HEAD — the comparison is unestablished; stop"; exit 1; } + START_ID=$(git rev-parse "$START:$p") \ + || { echo "$p does not resolve at the starting revision — the comparison is unestablished; stop"; exit 1; } APPROVED_ID=$(git rev-parse "ba15e83:$p") \ || { echo "$p does not resolve at ba15e83 — the comparison is unestablished; stop"; exit 1; } - test "$HEAD_ID" = "$APPROVED_ID" \ - || { echo "$p differs at HEAD from its approved version"; exit 1; } + test "$START_ID" = "$APPROVED_ID" \ + || { echo "$p differs at the starting revision from its approved version"; exit 1; } done +# Hand the checked revision to Task 0 step 1. Guarded, and no `cat` afterwards: a failed +# redirect with a following `cat` prints a STALE value and the block still exits 0. +printf '%s\n' "$START" > .context/loop-rule-start \ + || { echo "recording the starting revision FAILED — Task 0 step 1 has nothing to record"; exit 1; } ``` **How success is recognised.** Every line above exits 0. **No claim is made here about a recorded @@ -375,99 +397,102 @@ an unapproved edit to a spec that the gates already closed. ### Close — after a closure-eligible pass +**Which rules govern this close.** This change's Gate-B cycle starts at the first `WIP:` commit and +ends at the closing commit, which is the commit that ships the new §5. `CLAUDE.md` "When these rules +bind" says a running cycle finishes under the rules it started with, so this cycle runs under **§5 as +it stands at `$BASE`** — read it with `git show "$BASE:CLAUDE.md"`, because the tasks rewrite the +working copy. That text owns the closing shape (`reset --soft` to the parent of the first `WIP:`, then +one commit), the provenance line, the curve and the evidence entry. What follows are this plan's own +conditions around that act. (The Gate-A plan cycle `om0bdd7udh` reviewing this plan is a different +cycle and keeps the rules it started under.) + **What is checked**, in this order, and these are the conditions rather than a script: -1. **`HEAD` equals the head the candidate pass was issued against** — the value recorded before that - call, in `.context/loop-rule-reviewed-head`. **That is the reviewed head.** It is not the same - value as the **closing tip** below, and conflating them is how an unreviewed commit reaches the - close. -2. **The only thing dirty is the candidate pass's own findings files** — every record counts, not - two status codes, because a staged tracked modification is invisible to a filter that reads only - `??` and `A `. **And the check does not split git's output into pathnames at all**: a pathname - may contain a newline and a rename record carries a second, prefix-less path, so a hand-rolled - split can mangle a name or drop a dirty path entirely — an extra file whose name begins with a - newline passed such a parser invisibly, observed in a disposable repository. **Ask git which - paths outside the expected pair are dirty**, naming only the two expected paths, which are this - plan's own slot names and carry no such bytes. **The set can be exact because step 7 commits - every record this plan collects before the candidate pass is issued**; anything else uncommitted - here arrived after the review and no pass has seen it. A full Gate-B pass writes **two** files, - so the pair is what is expected. +1. **`HEAD` equals `.context/loop-rule-reviewed-head`** — the head the candidate pass was issued + against. **That is the reviewed head.** Nothing is committed between that call and this check, + because findings files are not committed during the loop (*D1* below), so the reviewed head is + still `HEAD` unless something reached the repository no pass has seen. +2. **Nothing is dirty except this cycle's Gate-B findings files.** The set is the **exact slot + paths** `.context/codex-reviews/gate-b-spec--pass-

.md` and + `gate-b-quality--pass-

.md` for every pass number `p` the cycle issued, built from the + nonce and the pass numbers the cycle already holds — **never from a glob**, which would also take a + dispositions note for a findings file. The final pass's pair must exist; any other missing slot + must belong to a pass recorded as INCOMPLETE. **Advisory companions are not in the set.** A + dispositions note still in the tree stops the close here — §5 lets it be deleted. **The cycle's + working record, `.context/codex-reviews/gate-b--resume.md`, is the one exact path left out + of the check**: it stays untracked through the close, because §5 makes it the recovery source + while the cycle is open, and it is retired only after condition 4 has passed (pass 54, finding 4). + No other path is exempt. **Staged or not makes no + difference to a selected file**: every one is pinned, staged and committed. **Anything else dirty + stops the close — tracked, staged or untracked, under `.context/` included.** Ignored scratch + files do not appear in git's status at all, so no exemption is written for them. **Ask git which + paths outside the set are dirty**, naming only the set, which are this plan's own slot names and + carry no newline or rename field; a hand-rolled split of git's output can mangle a name or drop a + dirty path, observed in a disposable repository. **The set can be exact because findings files + are committed only here.** 3. **The closing message is complete** — rebuilt whole, exactly one provenance line, exactly one curve, every owed evidence entry, and either the applicable human-exception records or - `Human exceptions: none`. **This is checked here, before anything moves**, because an incomplete - message discovered after the record commit leaves `HEAD` moved for a reason no restoration was - owed for. **And it is re-read and revalidated at 8b, immediately before the commit consumes it** - — 8a runs a commit in between, hooks can rewrite anything, and the file is under `.context/`, - which is ignored, so no clean-tree check between the two invocations can see it change. A - difference or a malformed record there is a Failure handoff, not a repair. -4. **Those files are committed** in their own invocation, and that commit satisfies **three things, - not one**: - - it **changes exactly those paths**, and its parent is the reviewed head; - - **the committed blobs are the validated ones** — the object ids recorded before staging equal - the ids at `HEAD` afterwards. A path check cannot see this: a hook rewriting a staged file in - place leaves the pathname untouched, which the plan states elsewhere and must therefore guard - here; - - **the committed findings files still satisfy the findings-file structure and still make this - the eligible logical pass they were judged as** — every line before the terminator a finding - line, the terminator exact, the count matching, both branch files present, and the - clean-or-zero-finding reading unchanged. Subject: the committed content. Base: the protocol in - §5 and the eligibility this candidate was issued on. **Any difference at any of the three is a - Failure handoff, not a repair.** - - Record the resulting commit as the **closing tip**, in - `.context/loop-rule-reviewed-tip` — **a precondition value, not a restore target**: 8b refuses to - reset unless `HEAD` is still exactly it, which is how a commit landing between the two - invocations is caught. Nothing in this plan resets *to* it. -5. **The tree is clean**, and `HEAD` is still the closing tip, when the closing invocation begins — - **both read in that invocation**, not carried from the previous one. **`git reset --soft` moves - `HEAD` and leaves the index exactly as it was**: it stages nothing and unstages nothing, so the - index still holds every `WIP:` commit's content, which is what makes one closing commit carry the - whole change. **Anything staged and uncommitted at that moment is in the index too and lands in - the closing commit; only unstaged work stays out.** This condition is what closes that asymmetry — - without it a staged edit made between the two invocations, a rewritten findings file included, - ships unreviewed and leaves a clean tree behind it. -6. **After the closing commit**, four things: its **subject is not a snapshot**; its **parent is - `$BASE`**; the **tree is clean**; and its **body is identical to the bytes 8b pinned when it - revalidated them** — **not to the source file, which is the wrong oracle**: the same hooks that - rewrite git's copy *after* `-F` has read it can rewrite that ignored file too, and a comparison - against it would then pass on a body nobody validated. The check is not decoration either way — - a dropped provenance line, curve, exception marker or evidence entry would pass every other one. - **What this does not close**, stated rather than left to be found: a hook that also rewrites the - pinned copy defeats it, and nothing here detects that. What it closes is the ordinary case, where - the oracle was the very file the commit read. **Any of the four failing is a Failure handoff.** + `Human exceptions: none`. It is checked **in the closing block, immediately before the commit + consumes it**, and its bytes are pinned there as condition 4's oracle — the file is under + `.context/`, which is ignored, so no tree check can see it change. +4. **After the closing commit**, five things, all read from one resolved commit `$CLOSED`: its + **subject is not a snapshot**; its **parent is `$BASE`**; the **tree is clean** except the + working record's one exact path; its **body is identical to the pinned validated bytes** — **not to the source file, which is the wrong + oracle**: the same hooks that rewrite git's copy *after* `-F` has read it can rewrite that ignored + file too; and **each selected findings file's committed blob equals the blob pinned immediately + before staging**, with the final pass's pair re-read from `$CLOSED` against §5's *Accept a pass + only when* and the clean-or-zero-finding reading it was judged on. **What the blob check + compares, stated exactly:** bytes at pin time against bytes committed. It catches a change between + the pin and the commit — a staging hook, a concurrent writer — and **nothing earlier**: a findings + file edited between the pass's acceptance and the pin is pinned as edited and passes. **What the + body check does not close**: a hook that also rewrites the pinned copy defeats it. **Any of the + five failing is a Failure handoff.** **How success is recognised.** A single commit at `HEAD` whose subject is the real message, whose -parent is `$BASE`, **whose body is identical to the pinned bytes of the revalidated closing -message**, with a clean tree -behind it. **No tree comparison** — target §I parks a Gate-B tree-equality condition on -Daniel's decision of 2026-09-13, and an earlier draft of this line added one anyway, which would -have made the executor either invent an out-of-scope closure check or declare success without -establishing its own stated predicate. Then, and only then, the scratch -files are removed. - -**Two invocations, not one.** `codex-gate.sh`'s `is_wip_commit` +parent is `$BASE`, whose body is identical to the pinned bytes of the revalidated closing message, +which carries this cycle's findings files as pinned, with a clean tree behind it. **No tree +comparison** — target §I parks a Gate-B tree-equality condition on Daniel's decision of 2026-09-13. +Then, and only then, the scratch files and the working record are removed, and the tree is clean +without exception. + +**One closing command, with no `-m`.** `codex-gate.sh`'s `is_wip_commit` (`plugins/dev-workflow/hooks/codex-gate.sh:763`) tests the **whole command string** against -`-m[[:space:]]*['"]?[[:space:]]*wip`. The findings-file commit carries that pattern; the closing -commit must not share a command string with it, or the hook reads the close as cycle-internal and -carries this cycle's count and fingerprint into the next. - -**Which of the six are preconditions and which are postconditions**, because they do not all come -before a move and an earlier draft said they did: - -- **1, 2 and 3 are preconditions** — checkable while nothing has moved. **On deviation, stop; no - restoration is owed**, because nothing was changed. **The message check is among them - deliberately**: an earlier draft classified it as a precondition while listing it *after* the - record commit, so a literal executor could find an incomplete message with `HEAD` already moved - and then stop without restoring. -- **4, 5 and 6 straddle or follow a move** — 4 makes the record commit, 5 reads the state it left, - 6 reads the state after the closing commit. **On deviation, run `## The four procedures` · - Failure**, which stops, reports the state and hands over. +`-m[[:space:]]*['"]?[[:space:]]*wip`. The closing block makes no `WIP:` commit and its commit carries +`-F`, not `-m`, so the hook cannot read the close as cycle-internal. **That is why one invocation is +enough**; the earlier two-invocation split existed only because a `WIP:` record commit shared the +close. + +**Which conditions are preconditions.** **1, 2 and 3 are preconditions** — checkable while nothing +has moved; on deviation stop, and no restoration is owed. **A failed `git add` inside the block is +not yet a move of `HEAD`, but the index may then hold part of the set** — stop and report it; do not +reset. **4 follows the move**; on deviation run `## The four procedures` · Failure. + +**What this does NOT close, stated rather than left to be found.** The block reads `HEAD` and the tree +before `git reset --soft` a few lines later and claims nothing about the span between. A second writer +mutating the branch inside that span — another agent session, a background hook, an editor running +git — could put different content into the index after the checks, and **no later check would catch +it**: condition 4 compares subject, parent, cleanliness, body and the findings blobs, never the rest of +the content, and the content comparison that would have caught it is the Gate-B tree-equality +condition target §I **parks** (`docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md`, +§I). **The width of that span is not measured**, and no claim of negligibility is made. **It is not +guarded**: what a fix would demand is an atomicity guarantee against a concurrently writing second +actor, which nothing in the approved text agrees to — **a scope question for the user, not a defect +in this disclosure.** Anything committed or staged **before** the block runs fails condition 1 or 2. **A rejected commit left at `HEAD`, or a live soft-reset index, is a state this plan stops in and -reports** — it is not an error condition the procedure has to undo. §A re-establishes the closure -conditions **against the repository as it now stands**, so that state is a legitimate starting -point for the resume decision, and an earlier draft's claim that "a retry cannot start from it" was -wrong about the approved text. +reports** — it is not an error condition the procedure has to undo. Failure's resume rule +re-establishes the closure conditions **against the repository as it now stands**, so that state is a legitimate starting +point for the resume decision. + +**D1 — findings files are not committed during the loop. Daniel's decision, 2026-09-25.** Each +pass's findings files are saved on disk under this cycle's nonce and committed **together, at the +close**. Until then they are **not Git-durable**: a parked cycle has no bounded duration, and a +`.context/` clear can remove them together with the working record. **Counts that survive elsewhere, +and a curve's `?`, do not establish what a lost file said or that its findings were resolved.** If it +happens, §5's existing rules apply unchanged — a pass whose file is gone before it was validated is +INCOMPLETE, and a finding whose resolution can no longer be shown is not resolved: **stop and +surface**. No backup, copy or recovery mechanism is added. **What D1 buys:** no commit lands between a +review and the reading of its result, so the reviewed head stays `HEAD` until a repair is committed. ### Failure — the closing act did not complete @@ -489,8 +514,8 @@ first is what destroys the evidence that decision needs. --soft` had already run, say so, because the index then holds the whole change; - **what the tree holds**: `git status --porcelain --untracked-files=all`, `git diff --cached --stat`, `git diff --stat`; -- **which cycle values still exist** — the recovery base, the reviewed head, the reviewed tip, the - closing message — by listing them, not by assuming. +- **which cycle values still exist** — the recovery base, the reviewed head, the findings list and + its blob pin, the closing message, the working record — by listing them, not by assuming. **What stopping does not do, said plainly.** It prevents *further* change; it is **no guarantee that the failed operation itself destroyed nothing.** A hook that rejected after rewriting a staged file @@ -500,14 +525,15 @@ has already done that, and this procedure cannot undo it or prove it did not hap **How success is recognised.** The failure is surfaced with the four observations above, nothing further has been changed, and the cycle is waiting on a person. -**Resuming is a decision about the actual state**, taken with that report in hand, and then §A's -rules apply unchanged: +**Resuming is a decision about the actual state**, taken with that report in hand, and then these +rules apply. **They are this plan's own procedure, worded after target §A**; §5 at `$BASE` says nothing +about a failed closing act, so they add to it without contradicting it: - **Every closure condition is re-established against the repository as it now stands.** Where they all still hold, **perform the act again**. - **Where the attempt or its repair moved anything a condition is read from**, that condition has changed and **its own rule decides what it costs**, a further pass included, and the cycle is back - in the ordering with that pass owed. + in step 7 with that pass owed. - **Where the failure cannot be repaired at all** — a signing key nobody has, a permission nobody can grant — **surface it and leave the cycle parked**: open, not running, spending no passes, restarted by an explicit later continue. @@ -526,9 +552,10 @@ reconcile. **Resume reads the repository; it does not classify it first.** Every check below runs against the **actual** `HEAD`, index, worktree and surviving cycle records, and **none of them is gated on the -state matching a named shape.** The approved §A is what this follows — *"re-establish every closure -condition against the repository as it now stands"* — and it asks for the conditions to be read, -never for the history to be catalogued. **The table below is illustration and is explicitly not +state matching a named shape.** Its wording follows target §A — *"re-establish every closure +condition against the repository as it now stands"* — as a plan procedure, not as a rule governing +this cycle; §5 at `$BASE` says nothing about re-entry. It asks for the conditions to be read, never +for the history to be catalogued. **The table below is illustration and is explicitly not exhaustive**: no permission, no base-validity conclusion, no recovery operation and no continuation route may depend on the state fitting a row, and **a state that fits none is an ordinary input to Resume, not an error**. @@ -539,34 +566,26 @@ a precondition for resuming at all. **That was this plan's invention, never §A' twice in two passes in the same way: pass 36 found a reachable state no row described, and pass 37 found the row added for it describing only one shape of the several that reach it. **What survives unchanged is every governing duty** — the conditions, their object-id equality checks, the -precondition/handoff distinction, the scratch rules, the bounded handoff, and §A's three routes. +precondition/handoff distinction, the scratch rules, the bounded handoff, and Failure's three routes. | Illustration, not a classification | `HEAD` | `$BASE..HEAD` | Where the work is | |---|---|---|---| | **Normal, mid-implementation** | the last `WIP:` snapshot | this run's `WIP:` commits, and possibly a §A3 stray commit or amend | committed | -| **8a rejected, no commit landed** | the last `WIP:` snapshot, unchanged | this run's `WIP:` commits | committed, **plus whatever the failed attempt left in the index or worktree** | -| **8a rejected after its commit landed** | a `WIP:` findings commit above the cycle's `WIP:` chain | those commits | committed, **plus any delta the post-commit clean-tree check found** | -| **A closure condition rejected on `HEAD`** — no reset has run | **not the object id the condition expected** — the reviewed head at condition 1, 8a's findings commit at condition 5. **The direction is not part of the test**: a further commit, an amend, a rebase and a move backwards all fail the same equality | whatever the actual history holds | wherever the actual content is — and **anything the mismatch introduced is present but unreviewed** | -| **8b rejected** | either `$BASE` itself, if the closing commit never landed, **or one commit parented by `$BASE`**, if it landed and a postcondition refused it | **empty**, or that one commit — the `WIP:` chain is gone either way, squashed by `reset --soft` | the **index**, or that one commit's tree | - -**The 8b illustration is the one worth reading before any rule below.** `reset --soft` removes the -`WIP:` chain from the ancestry, so after 8b there is no chain to find; an empty `$BASE..HEAD` there -means the work is staged, not absent. **And the two 8a illustrations differ from each other**: -before its commit the history is untouched and the delta is loose, after it the findings commit sits -above the chain — pass 31 separated them because they need different content reconciliation. - -**Conditions 1 and 5 are equality tests and stay that way.** `test "$(git rev-parse HEAD)" = -"$HEADREV"` and the same against `$TIP` ask one question — is `HEAD` the object the cycle recorded — -and **any answer of no stops the closing sequence**, whatever git operation produced it. **Do not -rewrite the recorded value to make the test pass**: those files are what make "the head the -candidate pass was issued against" a fact rather than a claim, and editing one turns an unreviewed -tree into a closable one with nothing left to notice. - -**The two rejections differ in one way the rules below turn on.** Condition 1 is a **precondition**, -so nothing this plan does has moved: that is a plain stop, and no handoff is in progress. **Condition -5 rejects after 8a's record commit has landed**, so the Failure handoff *is* in progress — which is -what selects the `inside a handoff` branch of the scratch-artifact rule below, report and change -nothing, rather than delete and rebuild. +| **Close condition 1 rejected** — no reset has run | **not the reviewed head**. **The direction is not part of the test**: a further commit, an amend, a rebase and a move backwards all fail the same equality | whatever the actual history holds | wherever the actual content is — and **anything the mismatch introduced is present but unreviewed** | +| **The closing act rejected** | either `$BASE` itself, if the closing commit never landed, **or one commit parented by `$BASE`**, if it landed and a postcondition refused it | **empty**, or that one commit — the `WIP:` chain is gone either way, squashed by `reset --soft` | the **index**, or that one commit's tree | + +**The closing-act illustration is the one worth reading before any rule below.** `reset --soft` +removes the `WIP:` chain from the ancestry, so after it there is no chain to find; an empty +`$BASE..HEAD` there means the work is staged, not absent. **In every state, this cycle's findings +files are uncommitted until the close (D1)** — they are part of what the four sources show. + +**Condition 1 is an equality test and stays that way.** `test "$(git rev-parse HEAD)" = "$HEADREV"` +asks one question — is `HEAD` the object the cycle recorded — and **any answer of no stops the +closing sequence**, whatever git operation produced it. **Do not rewrite the recorded value to make +the test pass**: that file is what makes "the head the candidate pass was issued against" a fact +rather than a claim, and editing it turns an unreviewed tree into a closable one with nothing left to +notice. **A condition-1 rejection is a precondition stop**: nothing this plan did has moved, and no +handoff is in progress. A failure **after** `reset --soft` is a Failure handoff. - [ ] **Validate the base — Resume's own checks, not Preparation's** @@ -628,39 +647,41 @@ Each by its `base` line **and** its completeness predicate: condition in the five regions appears in exactly one `span` or one `cond`. - `.context/loop-rule-baseline-diff.txt` — every line parses as `base`, a `site` record, or diff output belonging to the site above it, and **every inventoried site has a `site` record**. -- `.context/loop-rule-records-commit` — **present means a pass was recorded and step two may not - have read it.** Step 7 step one writes it, step two consumes it, and nothing else touches it, so on - re-entry it is the one artifact that says *a pass is recorded and the ordering may not have - spoken*. Validate it: a full 40-character object name, resolving to a commit, with `ba15e83` as an - ancestor and the commit itself an ancestor of `HEAD`; and its two slot paths present as blobs - **in that commit**, not in the worktree. **This step establishes the marker; it decides nothing.** - -**A valid marker is a finding of the reconciliation, not an instruction.** It says the recorded pass -exists and reports it as **unrouted unless the reconciliation below shows step two's outcome in the -content** — a repair committed on an authorizing route, a park, a stop answer. **Where it cannot be -shown, this is a precondition stop: report the marker, the commit it names and what the four sources -hold, and take no mutation and issue no call.** Rerunning step two over a recorded pass is the -ordinary continuation and needs nothing new; **what must not happen is a repair or a next call while -the ordering has not spoken**, which is the loop §A forbids and the one this plan must not itself -execute. The marker is **retired by the close's cleanup**, with the other scratch values, and by -nothing else — an interrupted cycle keeps it precisely so this check can find it. +- `.context/loop-rule-start` — **written by Preparation, consumed and removed by Task 0 step 1, and + never present once a base exists.** It carries the revision Preparation's predicates were about. + Its recovery rule ships with it, because a cycle state file introduced without one is the defect + pass 48 found in the revision before this: **a surviving `loop-rule-start` with no + `.context/loop-rule-base` is an interrupted first entry** — Preparation completed, Task 0 step 1 + did not. That is a **precondition stop**: report the recorded revision and where `HEAD` now is, and + take no mutation. **Do not rebuild it and do not resume from it**, because the tree may have moved + since Preparation checked it and nothing here re-establishes those predicates; re-running + Preparation from a clean tree is the ordinary continuation, and it overwrites this file. **Both + present is not a conflict** — it is Task 0 step 1 interrupted between its redirect and its `rm`, and + the base is the authority; remove the start file and carry on. **Neither present** is an ordinary + first entry. **On failure the answer depends on why you are here.** Outside a handoff: delete and rebuild from the `$BASE` blobs — never reuse, never repair in place, because a same-base partial file is the one shape a `base` line alone cannot catch. **Inside a handoff, while a failed close waits on a person: -report the invalid artifact and change nothing.** Rebuilding it would remove evidence before anyone -chose a §A route. **This plan performs no cleanup after a failure — it does not claim the failed -operation left anything intact.** Enumerate which scratch values actually survive and validate each; +report the invalid artifact and change nothing.** + +**This plan performs no cleanup after a failure — it does not claim the failed operation left +anything intact.** Enumerate which scratch values actually survive and validate each; a value is trustworthy because it passed a check, never because cleanup was skipped. - [ ] **Establish how far the implementation got — from the content, not from the log** **Read all four sources, every time, and do not decide first which of them matters.** Which one -holds the work varies — after `reset --soft` it is the index and `$BASE..HEAD` is empty; after an 8a -handoff part of it is loose — and **routing on a guessed shape is how a source gets skipped**. So +holds the work varies — after `reset --soft` it is the index and `$BASE..HEAD` is empty; after a failed +close part of it may be loose, and this cycle's findings files are always uncommitted until the close — and **routing on a guessed shape is how a source gets skipped**. So read them all and reconcile against what they actually contain: ```bash +# $BASE is read HERE: each fenced block is its own shell, so an assignment in an +# earlier block is gone (pass 53, finding 3). +BASE=$(cat .context/loop-rule-base) \ + || { echo "cannot read the base — the reconciliation cannot start; stop"; exit 1; } +test -n "$BASE" || { echo "base empty — the reconciliation cannot start; stop"; exit 1; } # Guarded one by one, because this step's rule is that all four ARE read: unguarded, # a failed read is indistinguishable from a source that holds nothing, and the next # command's success carries the block to an exit 0 on an incomplete view. @@ -676,7 +697,7 @@ git status --porcelain --untracked-files=all \ A ticked checkbox is confirmed by the change being *present in that content*, wherever it lives. `git status` names paths and says nothing about what is in them, and `--stat` counts lines. **The -delta that caused an 8a handoff is precisely the one a path listing cannot describe** — a rewritten +delta that caused a failed close is precisely the one a path listing cannot describe** — a rewritten staged file keeps its name — so reconcile against `HEAD`, the index and the worktree **contents** together. @@ -691,20 +712,18 @@ reviewed** — a further commit, an amended one, a rebased one, or a tree the hi and a ticked checkbox does not make any of it reviewed. **Do not adopt it**: carrying it into a close is the defect Close condition 1 exists to stop, and that condition's own rule decides the cost — only a clean response issued against that exact `HEAD` closes the cycle. **Do not remove it**: -this plan restores nothing, and what you delete here is evidence a person has not yet chosen a §A -route on. **And do not rewrite `.context/loop-rule-reviewed-head` or `-reviewed-tip` to match**, +this plan restores nothing, and what you delete here is evidence a person has not yet chosen a Failure +route on. **And do not rewrite `.context/loop-rule-reviewed-head` to match**, which would make the equality pass by discarding the only record of what was reviewed. -**Route the mismatch by the check that rejected it, not by the fact that it was a mismatch.** The -no-adopt, no-remove and no-rewrite rules above hold for both; **where they go does not.** A -**condition 1** mismatch is a precondition rejection — nothing this plan did has moved, so it is a -**plain stop**: report the mismatch and the state, and **no handoff is in progress**, which is what -lets the scratch-artifact rule above take its ordinary *delete and rebuild* branch. A **condition 5** -mismatch rejects after 8a's record commit has landed, so it goes to `## The four procedures` · -Failure, is named in that report's "whether a commit landed" line, and **is** inside a handoff — so -the scratch rule's *report and change nothing* branch applies. Either way §A's three routes decide -what happens next. **An earlier revision sent every mismatch to Failure**, which would have put a -condition-1 stop inside a handoff it is not in and suppressed a rebuild that was owed. +**Route by where the close stopped.** A **condition-1** mismatch is a precondition rejection: +nothing this plan did has moved, so it is a **plain stop** — report the mismatch and the state, no +handoff is in progress, and the scratch-artifact rule above takes its ordinary *delete and rebuild* +branch. A failure **after `reset --soft`** goes to `## The four procedures` · Failure and **is** +inside a handoff, so the scratch rule's *report and change nothing* branch applies. Either way Failure's +three routes decide what happens next. **A pass that was validated but not yet routed** is found like +any other state: its findings files sit uncommitted in the four sources, and routing it through step +7 is the ordinary continuation. **How success is recognised.** The base passes Resume's own checks; the artifacts match it and are complete, or are reported as invalid and left alone; and the plan's task list has been reconciled @@ -890,13 +909,29 @@ preservation fragments that an earlier draft chose afterwards. - [ ] **Step 1: Decide which entry this is, then run that procedure** -**The base file decides it, and nothing else does:** +**Two files decide it, and there are three outcomes — not two.** The base file alone is not enough: +Preparation writes `.context/loop-rule-start` and Task 0 step 1 consumes it, so **a start file +surviving with no base is an interrupted first entry**, and Resume defines that state as a +precondition stop. Routing on the base alone sends it into Preparation, **which does not refuse an +existing start file and overwrites it** — destroying the only record of the revision the earlier +Preparation validated, and making that recovery rule unreachable. The route must test the file the +rule is about. ```bash if [ -e .context/loop-rule-base ]; then echo "a base file exists — this is a RE-ENTRY: run Resume, not Preparation" +elif [ -e .context/loop-rule-start ]; then + # Reachable exactly because it is tested BEFORE Preparation, which would overwrite it. + echo "a start file exists with no base — Preparation completed and Task 0 step 1 did not." + echo "This is a PRECONDITION STOP. Report, and take no mutation:" + echo " recorded starting revision: $(cat .context/loop-rule-start 2>/dev/null)" + echo " HEAD is now: $(git rev-parse HEAD 2>/dev/null)" + echo "Do NOT rebuild it and do NOT resume from it — the tree may have moved since those" + echo "predicates were established. Re-running Preparation from a clean tree is the" + echo "ordinary continuation, and it replaces this file." + exit 1 else - echo "no base file — this is a FIRST ENTRY: run Preparation, then record the base below" + echo "no base and no start file — this is a FIRST ENTRY: run Preparation, then record the base below" fi ``` @@ -909,16 +944,33 @@ or malformed base file is Resume's**, which fails it on its shape checks and sta overwrite one you did not just write:** ```bash -git rev-parse HEAD > .context/loop-rule-base -BASE=$(cat .context/loop-rule-base) +# The base is the revision PREPARATION CHECKED, never a fresh resolution of HEAD. +# Preparation ran the branch, ancestry and approved-blob predicates against this exact +# object and wrote it here; re-resolving `HEAD` would record a revision nothing checked. +test -e .context/loop-rule-start \ + || { echo "no starting revision recorded — Preparation did not complete; run it"; exit 1; } +BASE=$(cat .context/loop-rule-start) \ + || { echo "cannot read the starting revision; stop"; exit 1; } # Not "non-empty": a symbolic value such as HEAD passes every check below and # then RESOLVES DIFFERENTLY as WIP commits accrue, moving the reviewed range, # the parent counts and the final reset with the branch. case "$BASE" in [0-9a-f][0-9a-f][0-9a-f][0-9a-f][0-9a-f][0-9a-f][0-9a-f][0-9a-f]*) ;; *) echo "recorded base is not an object name: $BASE"; exit 1 ;; esac test "${#BASE}" -eq 40 || { echo "recorded base is not a full 40-character object name"; exit 1; } test "$(git rev-parse --verify "$BASE^{commit}")" = "$BASE" || { echo "recorded base does not resolve to itself as a commit"; exit 1; } +# The last moment before the first mutation. Preparation's checks are only about this +# repository if nothing has moved since; a move here is a stop, not something to record. +test "$(git rev-parse HEAD)" = "$BASE" \ + || { echo "HEAD moved between Preparation and recording the base — stop and report; do NOT record"; exit 1; } +printf '%s\n' "$BASE" > .context/loop-rule-base \ + || { echo "recording the base FAILED; stop"; exit 1; } +rm -f .context/loop-rule-start ``` +**The re-assert narrows the window; it does not remove it**, and saying otherwise would be the +overclaim `AGENTS.md` names. A move landing between that test and the redirect is still recorded. What +the change buys is that the base is now the object Preparation's predicates were **about**, which it +previously was not at all. + **Re-entry — run `## The four procedures` · Resume.** It owns the existing base, the scratch artifacts and how far the implementation got. @@ -1973,12 +2025,20 @@ strict-reading run, and every pair, presence check and parity comparison would s installed over it. Two of them, for orientation: ```bash -grep -cF 'Codex is advisory — validate before applying; dismissed finding → one-line why' CLAUDE.md +grep -cF 'advisory — validate before applying; dismissed finding → one-line why' CLAUDE.md grep -cF 'Open a TodoWrite' CLAUDE.md ``` Expected: `1` each in the worktree, and `parent=1 worktree=1` in each copy for every one of the nine. +**A fragment is line-local by construction, and that is established by running its `grep -cF` before +it is written down, never after.** The first example above began as `Codex is advisory — validate +before applying; …` and counted **zero in both copies**: the sentence wraps between `Codex is` and +`advisory`, so the literal never occurs. A *correct* source text failed its own preservation check. +That is the second time in this cycle a hand-written fragment crossed a line break — carried `e9`'s +clause was the first, at pass 10 — so the rule is stated here rather than left to the next author's +care. + - [ ] **Step 5: Parity** for all ten sites, each extracted by its own bounded region rather than a fixed line window. - [ ] **Step 6: Commit** @@ -2729,58 +2789,16 @@ git commit -m "WIP: bump dev-workflow to 0.12.0" `scripts/check-version-bump.sh` compares **commits**, so the bump must be committed before the battery runs. -- [ ] **Step 4: Run the full quality battery** - -**First fetch the base ref, because one step in the battery has a precondition the others do not.** -`check-version-bump.sh` compares *commits* against a base ref, so a stale base compares the bump -against a different commit than the pull request will, and the run is green about the wrong -comparison. **Fetch, then pass the fetched ref to the checker itself** — an earlier draft fetched -`origin/main` while the battery went on passing the local `main`, and recorded `origin/main` as -what supplied a comparison it had not supplied: - -```bash -# Both guarded, and no `cat`. A failed fetch leaves whatever an earlier one wrote in -# `origin/main`, and `git rev-parse` resolves that stale ref happily — the block would -# exit 0 having recorded a base the pull request does not have. A failed write is the -# same shape: the `cat` after it printed the older recorded value. -git fetch origin main \ - || { echo "fetch FAILED — origin/main is whatever an earlier fetch left; do not run the battery"; exit 1; } -git rev-parse origin/main > .context/loop-rule-baseref \ - || { echo "recording the base ref FAILED — do not run the battery"; exit 1; } -``` - -**Pass that object name to `check-version-bump.sh`, not `origin/main`.** A remote-tracking ref is -mutable: another fetch between the record and the run makes the evidence name one commit while the -checker resolves another, and the claim that the recorded revision is the argument the check -received would be false. The object name is the argument. - -**It goes to a file, not to a shell variable**, for the reason Task 0 gives about `$BASE`: each -fenced block below runs in its own shell invocation, so a `BASEREF=` assignment here is gone by the -time the battery runs and the checker would receive an empty argument. The battery reads it back. - -```bash -BASEREF=$(cat .context/loop-rule-baseref) -test -n "$BASEREF" || { echo "BASEREF empty — the fetch step did not run"; exit 1; } -shellcheck --shell=sh plugins/dev-workflow/hooks/codex-gate.sh && \ -shellcheck --shell=sh --exclude=SC2015 plugins/dev-workflow/hooks/codex-gate.test.sh && \ -shellcheck --shell=sh scripts/check-invariants.sh && \ -shellcheck --shell=sh --exclude=SC2015 scripts/check-invariants.test.sh && \ -shellcheck --shell=sh scripts/check-version-bump.sh && \ -shellcheck --shell=sh scripts/check-version-bump.test.sh && \ -HOOK_SH=sh sh plugins/dev-workflow/hooks/codex-gate.test.sh && \ -HOOK_SH=dash dash plugins/dev-workflow/hooks/codex-gate.test.sh && \ -sh scripts/check-invariants.test.sh && sh scripts/check-invariants.sh && \ -sh scripts/check-version-bump.test.sh && sh scripts/check-version-bump.sh "$BASEREF" && \ -claude plugin validate . --strict -``` - -Expected: exit 0. **`$BASEREF`, not `main` and not `origin/main`** — AGENTS.md's battery row writes -`main` because a human running it locally usually has one; here the base is the fetched commit, -resolved once, so the recorded revision and the argument the successful check received are the same -value by construction. - - [ ] **Step 4b: Apply all twelve `docs/prompt-standards.md` items to the installed text, and commit the result before Gate B** +**It keeps the letter `4b` and runs BEFORE step 4 — execute this file in reading order, not in +alphabetical order.** The name is unchanged because six places elsewhere cite "step 4b" and a rename +would have to reach every one of them. **The position is what changed, and it is the repair:** this +is the **last step that may alter content**, and until it has committed there is no candidate to +verify. It previously sat after the battery and after step 4c, so both of them certified a head that +4b's own commit then replaced — and a 4b repair to a shipped prompt or hook was absent from +everything that had already run. + **Nothing mechanical does this and no other task claims it.** The battery's three narrow checks are a floor — one `Target model:` spelling, one prose count claim, one severity vocabulary — and invariant 11 requires all twelve items of every skill, command, hook message and scaffolded template @@ -2794,11 +2812,12 @@ risk. **Expected result: all twelve pass. A failure stops this step.** Invariant 11 is not "record the score"; recording a failure and continuing ships a prompt change that violates it, which is the one -outcome this step exists to prevent. **On any failure: repair the text, then re-run every check the -repair invalidated** — the affected pairs, presence, absence and preservation counts, the parity -and untouched checks over the blocks touched, and the whole step-4 battery — and only then commit -the refreshed record. A repair made after the battery, with the battery not re-run, is the same -defect as a Gate-B fix with no re-run. +outcome this step exists to prevent. **On any failure: repair the text, then re-run every +record-producing check the repair invalidated** — the affected pairs, presence, absence and +preservation counts, the parity and untouched checks over the blocks touched — and commit the repair +with the refreshed record in the block below. **The battery is not among them here**: it runs on the +committed candidate, which does not exist until that block has run, and step 4 runs it right after +(pass 53, finding 1). **Record the subject set once, at the top of the output section**, and make each item's result refer to it: the §A–§H blocks installed in C, the same in W, and **ten hook prompt bodies, not @@ -2819,30 +2838,47 @@ dead-ends. Staging it *after* the review instead breaks the reviewed-head condit later moment that works. **Stage the repaired artifacts by name, and only those.** `git add -A` would carry unrelated work in -the tree into a range a reviewer is about to read as this change. The repairable set here is the -installed text this step judges: the two prompt copies, the hook and its test. **Only the ones the -repair actually touched go in**, together with this plan's refreshed record — and the re-run records +the tree into a range a reviewer is about to read as this change. On the first pass the repairable set is +the installed text this step judges: the two prompt copies, the hook and its test. **A step-7 repair +through this block names its own set** — which can include the spec or target text, since a fix that +changes specified behaviour updates the spec in the same commit, and `plugin.json` and +`CHANGELOG.md` where 4c's repair owes them (pass 54, finding 3). **Only the files the repair actually +touched go in**, together with this plan's refreshed record — and the re-run records must describe *that* commit, which is why the re-runs above come first. **Staging and commit are both guarded, and the commit is checked against what it was given**, because step 6 issues Gate B over the range this commit ends: a failed `git add` otherwise falls through to a commit that can still succeed on content the index already held, and the review then runs over a range -missing its repair. The pin is the same as step 7 step three's — **the whole index, unrelated paths -included**, demanding no presence, since a repair here may delete a file too. +missing its repair. The pin is **the whole index, unrelated paths included**, demanding no presence, +since a repair here may delete a file too. ```bash -# add only what the repair touched, from: CLAUDE.md, -# plugins/dev-workflow/commands/workflow-init.md, -# plugins/dev-workflow/hooks/codex-gate.sh, plugins/dev-workflow/hooks/codex-gate.test.sh -# Where a repair was made, the message is: "WIP: prompt-standards repair + plan records" +# Stage this plan plus exactly the files the authorized repair touched, each named — never +# `-A`. First pass: from CLAUDE.md, plugins/dev-workflow/commands/workflow-init.md, +# plugins/dev-workflow/hooks/codex-gate.sh, plugins/dev-workflow/hooks/codex-gate.test.sh. +# A step-7 repair also names the spec or target text where the fix changes specified +# behaviour, and plugin.json / CHANGELOG.md where 4c's repair owes them. +# Step 7 runs this block with its own LITERAL message — "WIP: fix " or +# "WIP: refreshed records" — written into the command itself, never through a variable: +# the gate hook reads the command string for `-m … wip`. git add docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md \ || { echo "staging FAILED — step 6 must not issue Gate B over this range"; exit 1; } ITREE=$(git write-tree) \ || { echo "cannot pin the staged index; do not proceed to step 6"; exit 1; } git commit -m "WIP: plan records" \ || { echo "commit FAILED — step 6 must not issue Gate B over this range"; exit 1; } -test "$(git rev-parse "HEAD^{tree}")" = "$ITREE" \ +# Resolve the new commit ONCE, and use that one id for the tree check AND the candidate +# record. Every candidate commit — this first one and every repair — uses this block. Checking `HEAD^{tree}` and then +# resolving `HEAD` again in a later block let the recorded candidate differ from the +# commit whose tree was checked (pass 52, Blocker 1). +CAND=$(git rev-parse --verify "HEAD^{commit}") \ + || { echo "cannot resolve the new commit — do not proceed to step 6"; exit 1; } +test "$(git rev-parse "$CAND^{tree}")" = "$ITREE" \ || { echo "the commit's tree is not the index that was pinned; do not proceed to step 6"; exit 1; } +# No `cat` after the redirect: a failed write with a following `cat` prints a stale value +# and the block still exits 0. +printf '%s\n' "$CAND" > .context/loop-rule-reviewed-head \ + || { echo "recording the candidate head FAILED — no verification and no call may run"; exit 1; } ``` **The same applies to every record this plan collects** — the sweep, the next-state table, the @@ -2850,6 +2886,269 @@ divergence list. **`check-invariants.sh` includes the prompt-conformance checks** — a `Target model:` line naming one recognized model, a prose checklist-count claim matching the checklist, and the finding-severity vocabulary as a closed set in both prompt copies. Those three are a floor, not coverage; invariant 11's other eleven items are judged by a reader. +**The commit block above also fixes the candidate, once — this is the only place a candidate +identity is created on the first pass.** It records the **same object id** whose tree it just +compared with the pinned index, in the same invocation; a separate block that resolved `HEAD` again +could record a different commit from the one checked (pass 52, Blocker 1). The consumers that take +an **object id** read it from this file and never resolve a ref again to obtain it: **step 4's +battery**, **step 4c**, **Gate B's `headSha`** and **Close condition 1** are four consumers of one +value rather than independent resolutions of a movable ref — the defect passes 47–49 kept finding +and pass 51 found again between 4c and step 6. + +**The file is `.context/loop-rule-reviewed-head`, which already exists and already means this.** It +needs no new recovery rule, no new cleanup entry and no new state: Close condition 1 already reads +it, this same block rewrites it for every repaired candidate in step 7, and the close removes it. +**Writing it at 4b rather than at step 6** is what keeps one candidate from being verified as one +object and reviewed as another — step 6 used to resolve `HEAD` itself. + +- [ ] **Step 4: Run the full quality battery** + +**First fetch the base ref, because one step in the battery has a precondition the others do not.** +`check-version-bump.sh` compares *commits* against a base ref, so a stale base compares the bump +against a different commit than the pull request will, and the run is green about the wrong +comparison. **Fetch, then pass the fetched ref to the checker itself** — an earlier draft fetched +`origin/main` while the battery went on passing the local `main`, and recorded `origin/main` as +what supplied a comparison it had not supplied: + +```bash +# Both guarded, and no `cat`. A failed fetch leaves whatever an earlier one wrote in +# `origin/main`, and `git rev-parse` resolves that stale ref happily — the block would +# exit 0 having recorded a base the pull request does not have. A failed write is the +# same shape: the `cat` after it printed the older recorded value. +git fetch origin main \ + || { echo "fetch FAILED — origin/main is whatever an earlier fetch left; do not run the battery"; exit 1; } +git rev-parse origin/main > .context/loop-rule-baseref \ + || { echo "recording the base ref FAILED — do not run the battery"; exit 1; } +``` + +**Pass that object name to `check-version-bump.sh`, not `origin/main`.** A remote-tracking ref is +mutable: another fetch between the record and the run makes the evidence name one commit while the +checker resolves another, and the claim that the recorded revision is the argument the check +received would be false. The object name is the argument. + +**It goes to a file, not to a shell variable**, for the reason Task 0 gives about `$BASE`: each +fenced block below runs in its own shell invocation, so a `BASEREF=` assignment here is gone by the +time the battery runs and the checker would receive an empty argument. The battery reads it back. + +```bash +HEADID=$(cat .context/loop-rule-reviewed-head) \ + || { echo "battery UNRESOLVED: no candidate head recorded — step 4b did not complete"; exit 2; } +test "${#HEADID}" -eq 40 || { echo "battery UNRESOLVED: candidate head is not a full object name"; exit 2; } +BASEREF=$(cat .context/loop-rule-baseref) \ + || { echo "battery UNRESOLVED: no recorded base ref — the fetch step did not run"; exit 2; } +test -n "$BASEREF" || { echo "battery UNRESOLVED: recorded base ref is empty"; exit 2; } + +BT=$(mktemp -d) || { echo "battery UNRESOLVED: cannot create the disposable repository"; exit 2; } +# The same disposable-clone shape as step 4c: --shared borrows objects and writes nothing +# back, so the work repo, its index, its worktree and its branches are untouched. +git clone -q --shared --no-checkout . "$BT/r" || { echo "battery UNRESOLVED: clone failed"; exit 2; } +cd "$BT/r" || { echo "battery UNRESOLVED: cannot enter the clone"; exit 2; } +git checkout -q --detach "$HEADID" \ + || { echo "battery UNRESOLVED: the candidate does not resolve in the clone"; exit 2; } +test "$(git rev-parse HEAD)" = "$HEADID" \ + || { echo "battery UNRESOLVED: the clone is not at the candidate"; exit 2; } +ST=$(git status --porcelain) || { echo "battery UNRESOLVED: cannot read the clone's status"; exit 2; } +test -z "$ST" || { echo "battery UNRESOLVED: the fresh checkout is not clean"; exit 2; } + +printf 'battery candidate=%s base=%s\n' "$HEADID" "$BASEREF" +shellcheck --shell=sh plugins/dev-workflow/hooks/codex-gate.sh && \ +shellcheck --shell=sh --exclude=SC2015 plugins/dev-workflow/hooks/codex-gate.test.sh && \ +shellcheck --shell=sh scripts/check-invariants.sh && \ +shellcheck --shell=sh --exclude=SC2015 scripts/check-invariants.test.sh && \ +shellcheck --shell=sh scripts/check-version-bump.sh && \ +shellcheck --shell=sh scripts/check-version-bump.test.sh && \ +HOOK_SH=sh sh plugins/dev-workflow/hooks/codex-gate.test.sh && \ +HOOK_SH=dash dash plugins/dev-workflow/hooks/codex-gate.test.sh && \ +sh scripts/check-invariants.test.sh && sh scripts/check-invariants.sh && \ +sh scripts/check-version-bump.test.sh && sh scripts/check-version-bump.sh "$BASEREF" && \ +claude plugin validate . --strict \ + || { echo "battery FAILED on candidate $HEADID"; exit 1; } +echo "battery PASSED on candidate $HEADID" +``` + +Expected: exit 0. **`$BASEREF`, not `main` and not `origin/main`** — AGENTS.md's battery row writes +`main` because a human running it locally usually has one; here the base is the fetched commit, +resolved once, so the recorded revision and the argument the successful check received are the same +value by construction. + +**The battery runs on the candidate, not on the working tree** (pass 52, Blocker 2; scope decided by +Daniel on 2026-09-25 as option B). It reads the candidate id from `.context/loop-rule-reviewed-head` +and runs in a fresh disposable clone checked out at that id — the same `--shared` clone shape step 4c +already uses — so uncommitted changes in the work repository cannot supply its inputs. **Three +outcomes:** `0` passed on the named candidate; `1` failed on it; `2` unresolved — the check could not +be set up, which is **neither a pass nor a failure** and, like a failure, stops here. The first +printed line names the candidate and the base it ran against; that line belongs to the evidence. + +**What this does not do.** A clone is not an immutable runtime: a test can still change files inside +its own clone while it runs. It checks the **candidate commit**, not the merge result CI builds — +that is step 4c's separate observation, with its own revision named — and a later CI run does not +replace it. The clone is left in a temporary directory and not cleaned up here. + +**Old conditions, accounted.** Kept: the fetch-and-record of the base ref, the base passed as an +object id, the full battery in AGENTS.md order, exit 0 as the pass, and the rule that a failure stops +before step 5. Changed deliberately: the subject — from "the working tree as it stands when it runs" +to the recorded candidate — and the disclosed limit "nothing enforces that the worktree stays +unmodified", which no longer applies to the battery's inputs. Added: exit `2` for a setup failure, +matching step 4c. + +- [ ] **Step 4c: Check the version requirement against the MERGE RESULT, not only the branch head** + +**Why this exists, stated as the comparison each side actually performs.** Step 4 runs +`check-version-bump.sh` against the **local branch head**. `.github/workflows/ci.yml` passes no +`ref:` to `actions/checkout`, so a `pull_request` job evaluates the **merge commit** instead. Those +are different objects, and **refreshing the base ref does not make them the same one** — which is why +step 4's fetch is a precision fix and not the repair for this. Demonstrated with the unchanged +checker: a branch that bumps to a version `main` has meanwhile reached **on its own**, while carrying +further plugin changes, **passes against the stale base and is rejected against the current one**. + +**This step does not change step 4's battery or the Gate-B review range.** It is an additional, +isolated observation; both existing checks keep their inputs and their meaning. + +```bash +# The candidate identity comes from the file step 4b wrote — NEVER from a placeholder and +# never from a fresh resolution. This is the same value the battery ran against, the same +# value step 6 passes as `headSha`, and the same value Close condition 1 authenticates. +HEADID=$(cat .context/loop-rule-reviewed-head) \ + || { echo "4c UNRESOLVED: no candidate head recorded — step 4b did not complete"; exit 2; } +test "${#HEADID}" -eq 40 || { echo "4c UNRESOLVED: candidate head is not a full object name"; exit 2; } +BASEID=$(cat .context/loop-rule-baseref) # the base object name step 4 already recorded +test -n "$BASEID" || { echo "4c: no recorded base ref — step 4 did not run"; exit 2; } + +WT=$(mktemp -d) || { echo "4c UNRESOLVED: cannot create the disposable repository"; exit 2; } +# --shared borrows objects and writes nothing back: the work repo, its index and its +# branches are untouched by everything below. +git clone -q --shared --no-checkout . "$WT/r" || { echo "4c UNRESOLVED: clone failed"; exit 2; } +cd "$WT/r" || { echo "4c UNRESOLVED: cannot enter the clone"; exit 2; } + +# Pin the INTENDED checker FROM THE CANDIDATE — the one step 4's battery ran with — before +# anything is merged. Read it by object id, not from a worktree: a worktree may sit on the +# base side and already carry the very edit this check exists to catch, in which case the +# comparison is against itself and always agrees. +git show "$HEADID:scripts/check-version-bump.sh" > "$WT/intended-checker" \ + || { echo "4c UNRESOLVED: cannot read the checker at the candidate head"; exit 2; } + +# DIRECTION: check out the BASE and merge the CANDIDATE into it. That is the order +# `refs/pull/N/merge` is built in, and it is what `actions/checkout` hands a +# `pull_request` job. The reverse order produces the same tree for an ordinary merge but +# not necessarily under direction-sensitive behaviour — merge drivers and `.gitattributes` +# among them — so the cheap way to be about CI's object is to build it the way CI does. +git checkout -q --detach "$BASEID" 2>/dev/null \ + || { echo "4c UNRESOLVED: recorded base does not resolve here"; exit 2; } + +# A conflict is an UNRESOLVED check, never a verdict. No merge strategy option, no manual +# resolution, no retry: resolving it here would invent a merge result nobody reviewed. +git -c user.email=v@v -c user.name=v merge --no-edit -m "4c merge result" "$HEADID" >/dev/null 2>&1 \ + || { git merge --abort >/dev/null 2>&1 + echo "4c UNRESOLVED: merging the candidate into the base conflicts — report and stop"; exit 2; } +R=$(git rev-parse HEAD) || { echo "4c UNRESOLVED: cannot resolve the merge result"; exit 2; } + +# THE BYTE CHECK BELONGS HERE, AFTER THE MERGE — these are the bytes that will actually +# run. Comparing before the merge establishes nothing about them: the base side can carry +# a different checker, and a merge that takes it succeeds without conflict. The checker +# must also run INSIDE this clone: its line 58 is `cd "$(dirname "$0")/.."`, so an +# ABSOLUTE-PATH invocation from elsewhere silently runs it against its own source tree +# and answers about the wrong repository — which is how an earlier attempt at this check +# produced a confident wrong negative. The relative call below is what binds it here. +cmp -s scripts/check-version-bump.sh "$WT/intended-checker" \ + || { echo "4c UNRESOLVED: the checker in the MERGE RESULT differs from the one at the pinned head"; exit 2; } + +# The record. Every id is a full object name; the diff and the status are observations. +printf '4c head=%s\n4c base=%s\n4c merge-result=%s\n' "$HEADID" "$BASEID" "$R" +printf '4c plugin diff base..R: %s\n' "$(git diff --name-only "$BASEID" "$R" -- plugins/ | tr '\n' ' ')" + +# Preserve the output and the RAW status. `check-version-bump.sh` exits 1 from BOTH +# `fail()` (a policy violation) and `die()` (an operational failure such as an +# unresolvable ref or a git call that errored), so a bare 1 does not establish which +# happened. Nonzero is therefore recorded as NOT VERIFIED with its cause undetermined; +# a person reads the preserved output to establish it. No parser is built here, and the +# checker's interface is not touched. +out=$(sh scripts/check-version-bump.sh "$BASEID" 2>&1); st=$? +printf '%s\n' "$out" +printf '4c raw checker exit=%s\n' "$st" +if [ "$st" -eq 0 ]; then + echo "4c VERIFIED: the version rule holds on the merge result"; exit 0 +fi +echo "4c NOT VERIFIED: the checker exited $st — cause UNDETERMINED from the status alone." +echo "4c exit 1 is BOTH a policy violation (fail) and an operational failure (die)." +echo "4c Read the preserved output above to establish which, before calling it either." +exit 1 +``` + +**Three outcomes, and what each one does and does not assert.** `0` — **verified**: the version rule +holds on this merge result. `1` — **not verified**: the checker exited nonzero, and **that is all the +status establishes**. `2` — **unresolved**: the check could not be carried out at all. + +**Non-zero is never reported as "the version is wrong" on the strength of the exit code.** +`scripts/check-version-bump.sh` exits `1` from `fail()` for a policy violation **and** from `die()` +for an operational failure — an unresolvable ref, a git call that errored, an unparseable manifest. +The two are indistinguishable by status, so 4c preserves the **output** and the **raw status** and +requires the cause to be established from that output before anyone describes it as a version +violation. **Both cases stop progression**, so nothing depends on guessing. **Unresolved is neither a +pass nor a failure**, and a single non-zero code would have made a merge conflict indistinguishable +from a checker verdict. + +**When it runs, and why it binds BEFORE Gate B.** 4c runs **after step 4b has fixed the candidate and +the battery has passed, and before step 5 writes the evidence entry** — on the first pass and on +every rerun, the same sequence. **A non-zero or unresolved 4c stops here**: no evidence entry is +written on it, and no Gate-B call is issued. + +**It cannot bind at 7b instead, and that was a real error in an earlier proposal.** Step 6 must hand +the reviewer **the evidence entry quoted verbatim**, and 7b states that if revalidation changes that +entry **the candidate is over**. A binding 4c result appearing first at 7b would therefore change the +entry the final reviewer had judged and force another candidate — the exact round this step exists to +avoid. + +**Where the result is carried, using a file that already exists.** Into +`.context/loop-rule-closing-msg`, which step 5 opens **before** Gate B and step 8 commits. **It is +gitignored, so writing it creates no commit** — which is what breaks the +record-commit/reverification cycle without any new state file, parser or recovery subsystem. 7b +rebuilds that message whole and **preserves this entry unchanged** unless revalidation genuinely +changes it, in which case the existing re-review rule applies untouched. + +**What the evidence covers, as an identity.** One pinned pair: the **candidate head** in +`.context/loop-rule-reviewed-head` and the **base object name** in `.context/loop-rule-baseref`. The +record names both, plus the merge result. **Re-run 4c whenever either side of that pair changes** — +a repair commit through the step-4b block writes a new candidate head, a fetch may move the recorded base — and whenever the +checker or a plugin manifest changes. **No commit is made between 4c and the call it precedes**, so +nothing can move the head it just certified. + +**The evidence boundary, stated rather than implied.** This verifies **the merge of one pinned pair +of commits**. It says nothing about a `main` that advances afterwards — that merge result does not +exist yet and cannot be checked here — and it is **not a check of CI as a whole**: it runs one script +against one merge result, while a CI run does more. `design.md` §7 requires the battery green at the +Gate-B WIP commit and **promises nothing about a later merge commit**; this step narrows that gap for +one pinned pair and does not close it. + +**Verified by execution of this block as extracted from this plan, with the unchanged +`scripts/check-version-bump.sh` inside each fixture repository** — eight cases, expected against +observed: + +| case | want | got | +|---|---|---| +| valid version bump | 0 verified | **0** | +| same version both sides + plugin diff | 1 not verified | **1**, and the policy cause is present in the preserved output | +| **operational failure inside the checker** (a plugin directory name outside its grammar → `die`) | 1 not verified | **1**, and **not** described as a version violation | +| **conflict-free merge that changes the checker's bytes** | 2 unresolved | **2**, refused **before** the checker ran | +| merge conflict | 2 unresolved | **2** | +| unresolvable pinned head | 2 unresolved | **2** | +| no recorded base ref | 2 unresolved | **2** | +| the source repository afterwards | unchanged | **branch unchanged, 0 staged, 0 dirty** | +| an unresolvable **candidate head** read from the file | 2 unresolved | **2** | + +**Re-run after the direction change and the identity change**, with the candidate read from +`.context/loop-rule-reviewed-head` and the merge built base-first: all of the above still hold, and +the stale-base control still reproduces the false green. The direction change moved no verdict in +these fixtures — which is expected, since none of them uses a merge driver — so it is adopted +**because it matches how CI builds the object**, not because a fixture showed a difference. + +**Two of those cases exist because earlier drafts of this step failed them.** The operational-failure +case is why non-zero is no longer read as a verdict. The byte-change case is why the comparison +happens **after** the merge and against the checker **at the pinned head** — an earlier draft compared +against the source repository's *worktree*, which can already sit on the base side and carry the very +edit the check exists to catch, so the comparison agreed with itself and the modified checker ran. +**A green result from a check that cannot go red is not evidence**, which is the rule this step is +built to respect. + - [ ] **Step 5: Write the evidence entry into `.context/loop-rule-closing-msg`** **Not into a WIP commit body.** Step 8 squashes with `git reset --soft`, which keeps the tree and @@ -2876,6 +3175,13 @@ It names: the battery run; **every pair this plan built, with its counts in each for those two reader records. **An entry assembled from one section would silently drop whatever the other five hold.** +**And the step-4c result, which is the one item read from this run rather than from a section.** +Record its three ids — candidate head, base, merge result — its plugin diff line and its raw checker +status. **It is written here because Gate B is handed this entry verbatim**: a merge-result check +whose result appears only after the review would change the entry the reviewer judged, and 7b says a +changed entry ends the candidate. **4c has already stopped the step if it was not verified**, so the +only status this entry can carry is a verified one. + **Every record shape belongs in the entry, not only pairs and presence.** A **dropped** condition's absence and a **moved** condition's absence are what prove an obsolete instruction was removed; a **carried** or span-less **kept** condition's preservation count is what proves @@ -2891,23 +3197,40 @@ an enumeration here** — an id list in this step was already stale once, naming ```bash BASE=$(cat .context/loop-rule-base) # the parent of the FIRST WIP, from Task 0 test -n "$BASE" || { echo "BASE empty or unreadable — Task 0 did not run"; exit 1; } -# The head THIS call is issued against. Guarded, and no `cat`: the redirect can fail -# while a following `cat` prints a STALE value and the block still exits 0. -git rev-parse HEAD > .context/loop-rule-reviewed-head \ - || { echo "recording the reviewed head FAILED — no call may be issued"; exit 1; } +# This step RESOLVES NOTHING. The candidate was fixed by step 4b and verified under that +# identity by the battery and step 4c; resolving `HEAD` again here is how the verified +# object and the reviewed object came apart. +HEADREV=$(cat .context/loop-rule-reviewed-head) \ + || { echo "no candidate head recorded — step 4b did not complete; no call may be issued"; exit 1; } +test "${#HEADREV}" -eq 40 || { echo "candidate head is not a full object name"; exit 1; } +# A move since the candidate was fixed is a STOP, never something to record over: the +# verification above describes $HEADREV and nothing else. +test "$(git rev-parse HEAD)" = "$HEADREV" \ + || { echo "HEAD has moved since the candidate was fixed — re-run 4b onward; no call may be issued"; exit 1; } ``` -**Every Gate-B call records the head it is issued against, this first one included**, and step 7 -rewrites the file before each later call. It is what makes "a clean response against that exact -`HEAD`" a comparison rather than a claim: without it the close procedure has nothing to hold `HEAD` -up against, and a commit landing after the response is indistinguishable from the reviewed state. +**The head is written by the step-4b commit block and read here.** A repair in step 7 runs that +same block — it is the only writer. It is what makes "a clean response against that +exact `HEAD`" a comparison rather than a claim: without it the close procedure has nothing to hold +`HEAD` up against, and a commit landing after the response is indistinguishable from the reviewed +state. **An earlier revision resolved `HEAD` here instead**, which gave the same candidate two +identities — the one 4c certified and the one Gate B reviewed — with nothing comparing them. **`baseSha` is `$BASE`, never `HEAD^`.** This plan makes a WIP commit per task, so `HEAD^` is the parent of the *last* one and the review range would hold the version bump alone — Gate B would close having reviewed none of the prompt or hook changes. `$BASE` is the parent of the first WIP and is the only value whose range contains the whole implementation. -`mcp__codex__review` with `reviewType: full`, `baseSha` = `$BASE`, `headSha` = the full 40-character object name `HEAD` resolves to **at that moment**, resolved once and kept with each branch's result. Carry the story path and the evidence entry quoted verbatim. Write findings to `.context/codex-reviews/gate-b---pass-

.md` — **draw a fresh nonce for this cycle**; it is a different cycle from `awsf1ec771`. +`mcp__codex__review` with `reviewType: full`, `baseSha` = `$BASE`, `headSha` = **the exact content of `.context/loop-rule-reviewed-head`**, read at the moment the call is built and passed as a literal value — **never a fresh resolution of `HEAD`, and never the symbolic `HEAD`**. The file is the identity; the call quotes it, and the same value is kept with each branch's result. **If you cannot show that the value you passed and the file's content are the same string, the call has not been issued** and the reviewed head is unestablished. Carry the story path and the evidence entry quoted verbatim. + +**Not guarded, and stated rather than implied.** Nothing mechanically compares the argument actually +sent to the reviewer against this file: the call is not shell, so there is no command to guard, and +`mcp__codex__review` reports no reviewed revision to read back. Close condition 1 therefore +authenticates **the file**, and the file is trustworthy only because the step above wrote it under a +guard and no other step touches it. A second resolution of `HEAD` here would break that chain +silently — Gate B would review one commit while the close authenticates another, and the content that +then closes includes this change's edits to `plugins/dev-workflow/commands/workflow-init.md` and +`codex-gate.sh`, **which ship inside the plugin package**. Write findings to `.context/codex-reviews/gate-b---pass-

.md` — **draw a fresh nonce for this cycle**; it is a different cycle from `awsf1ec771`. **Standing lens, every call:** "which existing statements does this diff falsify?" and **name what this diff changes the size, value or position of** — a list, a count, a version, an identifier, a cited line — then grep for where each is described elsewhere. @@ -2919,270 +3242,123 @@ Floor derives from the story profile: risk `high` → 2, security `none` → 0, floor**, or a **zero-finding logical pass** at any pass number — the early exit below the floor. Both are subject to every other closure condition. **An earlier draft named only the first**, so a first or second Gate-B pass whose two branch files both read `NO FINDINGS` would have been routed -into further passes — the implementation procedure overriding the very rule §A installs, and the -one case Task 13's table checks by name. A zero-finding pass is `NO FINDINGS` in **every** required +into further passes — overriding §5's own early exit (§5 at `$BASE`: *"The only early exit below +the floor is a pass with **zero** findings"*), which Task 13's table also checks by name in the +installed text. A zero-finding pass is `NO FINDINGS` in **every** required branch file; one branch clean and the other not is not it. -**Each fix is committed before the next review is issued**, or the re-review targets the unchanged -WIP tip while the repair sits in the worktree — and the final squash then publishes a fix no pass -reviewed: - -**Every pass that does not close ends the same way, whether or not it produced a repair — and it -ends in three steps, in this order.** Record the pass. Read it through the ordering this change -installs. **Then, and only on a route that authorizes it, repair.** The plan must not execute a loop -its own product forbids. - -**Step one — record the pass, and nothing else. This commits; it does not authorize anything.** - -**Do not apply the repair yet — not in the worktree either.** Deferring only the *commit* changes -nothing: this plan restores nothing, so an edit made before its route authorized it is just as -unauthorized, and a later decline has nothing that removes it. - -**A failed precondition, staging, commit or content check stops here; step two does not run.** No -command in this block is the last thing that happens — the commit is followed by a tree comparison, -and that by prose. A reader who saw a failure scroll past can still walk into step two, and the -ordering must not be read over a pass that was never recorded, nor over one whose findings files the -commit does not carry. - -**The plan is not staged here, and must be unchanged before this commit.** Step one produces no plan -output — the re-run records are written at step three — so a dirty plan here is a post-review edit, -and committing it puts a change in history before step two has read the ordering. Removing the path -from `git add` does not settle it: **`git commit` commits the index**, so an edit already staged -rides in regardless. The check therefore asks HEAD against the index first, then the index against -the worktree; a difference and a failed comparison both stop, and neither says what caused it, so -the answer to both is to **report the concrete state** — what is in the worktree, what is committed, -and which route had not yet spoken. This plan restores nothing and invents no rollback. The check -covers **this one path**; the residual below is unchanged. - -**A successful `git add` does not establish presence.** It stages a removal as readily as a content -change, so a tracked findings file that is absent from the worktree is staged as *gone* — and the -pinned index and the resulting commit then agree, both without it. **Tree equality preserves -presence only where presence was established in the pin**, which is why the two paths are checked -there, and checked as **blobs**: an existence test alone accepts a directory standing at the path, -one of the causes §5 already tells a reader to diagnose separately. Step one deletes nothing; an -authorized repair may, which is why step three and step 4b carry no such check. - -**What the pin is, and what it is not.** It is the index **as staged at that moment**, taken after -`git add` and before `git commit` — not the content the pass acceptance validated. Nothing between -that reading and this staging observes a change, so a findings file rewritten in between is pinned -as staged and passes here. Close condition 4 re-runs the structural check on committed content for -8a's record commit; **step one has no such re-run and this check does not stand in for one.** - -**Stage the two slot paths this pass was called with, spelled out.** Not `git add -A`, which sweeps -exactly the edit this step exists to keep out — and **not the review directory either**: `git add -.context/codex-reviews/` is a directory pathspec that stages every changed or untracked file beneath -it, so another cycle's findings, or an earlier pass's, ride into this commit and then into the -closing squash, where the exact-dirty-set guard can no longer see them because they were committed -before the next call. **`NONCE` and `P` are this cycle's nonce and this pass's number** — the same -two the call's slot paths were built from. - -```bash -NONCE=; P= -SPEC=".context/codex-reviews/gate-b-spec-$NONCE-pass-$P.md" -QUAL=".context/codex-reviews/gate-b-quality-$NONCE-pass-$P.md" -PLAN=docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md - -# The plan is not staged here and must be unchanged: HEAD against the index (what the -# commit takes), then the index against the worktree. A difference and a failed -# comparison both stop, and neither says what caused it — report the state. -git diff --quiet --cached HEAD -- "$PLAN" \ - || { echo "plan is not identical between HEAD and the index, or the comparison failed — stop and report the state"; exit 1; } -git diff --quiet -- "$PLAN" \ - || { echo "plan is not identical between the index and the worktree, or the comparison failed — stop and report the state"; exit 1; } - -git add "$SPEC" "$QUAL" \ - || { echo "staging FAILED — the pass is not recorded; step two does not run"; exit 1; } - -# Pinned after staging, before the commit: the WHOLE index as staged at that moment, -# unrelated paths included. -ITREE=$(git write-tree) \ - || { echo "cannot pin the staged index; step two does not run"; exit 1; } - -# A successful add does not mean these paths are present — it stages a removal too. -# Require a blob at each: an existence test would accept a directory at the path. -test "$(git cat-file -t "$ITREE:$SPEC" 2>/dev/null)" = blob \ - || { echo "$SPEC is not a blob in the staged index; the pass is not recorded, step two does not run"; exit 1; } -test "$(git cat-file -t "$ITREE:$QUAL" 2>/dev/null)" = blob \ - || { echo "$QUAL is not a blob in the staged index; the pass is not recorded, step two does not run"; exit 1; } - -git commit -m "WIP: pass $P records" \ - || { echo "records commit FAILED — stop here; step two does not run"; exit 1; } - -# Resolve the record commit ONCE, immediately after it lands, and use that object id for -# everything below. `HEAD` is a movable ref: resolving it a second time — for the tree -# check, and again to persist the identity — lets a commit, amend, reset, rebase or -# checkout in between hand those two commands DIFFERENT commits, which recreates the -# very defect this records the object id to close. The commit object is immutable; -# `HEAD` is not, and only the captured id carries that. -RECORD=$(git rev-parse HEAD) \ - || { echo "cannot resolve the records commit; step two does not run"; exit 1; } - -# Any divergence between the pinned index and THAT commit's tree stops here, whatever -# produced it. -test "$(git rev-parse "$RECORD^{tree}")" = "$ITREE" \ - || { echo "the commit's tree is not the index that was pinned; step two does not run"; exit 1; } - -# Persist the same id, because step two must read both slots from this exact commit — -# and a move between its two reads could otherwise take the branch files from different -# commits. -printf '%s\n' "$RECORD" > .context/loop-rule-records-commit \ - || { echo "recording the records commit's object id FAILED; step two does not run"; exit 1; } -``` - -**What this does not cover, disclosed rather than guarded:** naming the paths controls what this step -*adds* to the index; `git commit` still commits whatever the index already held. **A foreign path -staged before this step runs is carried in, and no check in this plan catches it** — the close's -dirty-set and changed-path checks read 8a's commit, not these. Nothing here fixes that. - -**Step two reads the findings out of step one's RECORD COMMIT, by object id — not the worktree -copies, and not through `HEAD`.** Target §A requires every finding-derived predicate to read the -**validated** findings file, and the worktree copy is the one thing that can still change after -validation: a `commit-msg` or `pre-commit` hook that rewrites only the worktree leaves step one's -checks entirely satisfied — the staged index, the pin and the commit tree all agree — while the file -a reader would open no longer holds what the pass was accepted on. **A commit object is immutable; -`HEAD` is a movable ref and is not**, so reading `HEAD:` would be redirected by any commit, amend, -reset, rebase or checkout in between, and a move between the two reads could take the two branch -files from different commits. The object id step one persisted is what makes the source one commit: - -```bash -NONCE=; P= -SPEC=".context/codex-reviews/gate-b-spec-$NONCE-pass-$P.md" -QUAL=".context/codex-reviews/gate-b-quality-$NONCE-pass-$P.md" -RECORDS=$(cat .context/loop-rule-records-commit) \ - || { echo "cannot read the records commit id — step one did not complete"; exit 1; } -test "${#RECORDS}" -eq 40 || { echo "records commit id is not a full object name"; exit 1; } -# Both reads come from this one object, so neither can be redirected and the two -# branch files cannot come from different commits. -git show "$RECORDS:$SPEC" \ - || { echo "cannot read the committed spec findings from the records commit"; exit 1; } -git show "$RECORDS:$QUAL" \ - || { echo "cannot read the committed quality findings from the records commit"; exit 1; } -# `HEAD` having moved is not a reason to re-resolve the files — it is a reason to stop -# and report, because something reached this repository between the record and the read. -test "$(git rev-parse HEAD)" = "$RECORDS" \ - || { echo "HEAD moved after the records commit — stop and report the concrete state"; exit 1; } -``` - -**What this does not establish, disclosed rather than guarded.** It does not prove those committed -bytes are the ones **pass acceptance** validated. Step one pins what is on disk when step one runs; -an edit between accepting the result and staging it is pinned as staged and passes every check here, -which this plan already says of its pin. Closing *that* gap would mean carrying the accepted blob -ids out of the acceptance step in a file, and **no such duty exists in the approved spec or in §5** — -it is named here as the residual it is, not repaired. **The reviewer's further demand — re-running -the structural and branch-eligibility checks on the committed content — is declined for the same -reason**: object identity is what step one can establish, and Close condition 4's third bullet is a -*closing*-commit duty the approved text does not place on a non-closing record commit. - -**Step two — read the pass against the ordering, before any next call exists:** - -- **A source block standing** → **the source rule decides what must be repaired or answered, and its - repair happens on that route** — §A sends the reader to the source and does not hold that repair - behind a continue. Then §A's release rule decides whether this pass is read again. **No further - pass while the block stands.** -- **Any suspension open** — a membership stop, a new-question stop, a two-tell stop, a clearly-stuck - surface — → **collect every answer and compose them.** **No repair to a finding whose membership - is unanswered**: §A answers membership against the fix set *as it stood for the pass that raised - the question*, so repairing first decides the question the stop exists to ask. Another pass only - where the composition yields **continue**. -- **A stop answer** → **park**: open, not running, **spending no passes**, restarted only by an - explicit later continue. **There is no "commit and carry on" from a stop**, and acceptance - criterion 4 requires that parked state to be distinct. **Nothing below runs on this route.** - **Parking is not a retroactive revocation**: §A defines stop as parking the open cycle and - prescribes no rollback, so a repair some earlier route had already authorized stays where it is. -- **The clean-completion branch** → **Close**, not another pass. -- **Continue** → step three. - -**Step three — repair where the route authorized one, then prepare the call.** Continue permits an -**unrevised** artifact only where no repair is owed; **where one is owed it comes before the -post-answer pass**, which is §A's rule and not this plan's. So on this route: apply the repair now, -re-run every check it invalidated, commit both, and only then record the head. - -**The commit is guarded because this block does not end with it.** Commands follow, and among the -first things they do is read `HEAD` — so an unguarded failure is masked by the `rev-parse` after it, -the *old* head is recorded, and the next call is issued against a tree the repair never reached. **A -failed commit stops the commands below it**: the previous candidate's tip is not removed, the -reviewed head is not rewritten, and no call is issued. It is **not** a claim that the failed attempt -left the tree as it was — a hook can rewrite anything before failing, which this plan already says of -the closing message file. **8a guards its record commit the same way**, and for the same reason. - -**The staging is guarded too, and the commit is checked against what it was given.** A failed `git -add` otherwise falls through to a commit that can still succeed on content the index already held. -The pin is `git write-tree` — **the whole index, unrelated paths included**, so any divergence between -it and the resulting commit tree stops this step, whatever produced it. That is broader than this -step's obligation and deliberately fail-closed. It is **not** a foreign-index check: content staged -before this step stands on both sides and passes. And it demands no presence, which is what step one -needs and this step must not have — **an authorized repair may delete a file**, and a staged deletion -is in the pin and in the commit alike. - -```bash -# The repaired files, spelled out, plus this plan — which carries the refreshed -# re-run records. The findings files went in at step one and are not re-staged. -# Neither `git add -A` nor the review directory: both carry in work no pass asked for. -git add \ - docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md \ - || { echo "staging FAILED — nothing committed, no call may be issued"; exit 1; } - -# Pinned after staging, before the commit: the whole index as staged at that moment. -ITREE=$(git write-tree) \ - || { echo "cannot pin the staged index; no call may be issued"; exit 1; } - -# Guard the commit. Everything below derives from HEAD, so an unguarded failure -# is masked by the `git rev-parse` that follows it: the old head gets recorded, -# the repair stays staged, and the next call reviews a tree it is not in. -git commit -m "WIP: fix " \ - || { echo "repair commit FAILED — tip and reviewed head left as they are; no call may be issued"; exit 1; } -# Resolve the repair commit ONCE and use that id for both the tree check and the head -# record. Resolving `HEAD` twice lets a move in between check one commit's tree and -# record another as the head the next call is issued against — the same defect step one -# records an object id to close, and not one the reviewer named here. -REPAIR=$(git rev-parse HEAD) \ - || { echo "cannot resolve the repair commit; no call may be issued"; exit 1; } -test "$(git rev-parse "$REPAIR^{tree}")" = "$ITREE" \ - || { echo "the commit's tree is not the index that was pinned; no call may be issued"; exit 1; } - -# The previous candidate's closing tip. Guarded: without the guard a failed removal -# leaves the stale tip standing beside the next candidate's reviewed head, so the next -# call is issued with two candidates' markers present. (It is NOT the value condition 5 -# reads in a correct sequence — 8a rewrites the tip before 8b runs.) -rm -f .context/loop-rule-reviewed-tip \ - || { echo "removing the previous candidate's tip FAILED — it survives; no call may be issued"; exit 1; } -# The head the NEXT call is issued against — the SAME id whose tree was checked above, -# never a fresh resolution. Guarded, and no `cat`: the redirect can fail while a -# following `cat` prints the STALE value and the block exits 0. -printf '%s\n' "$REPAIR" > .context/loop-rule-reviewed-head \ - || { echo "recording the reviewed head FAILED — no call may be issued"; exit 1; } -``` - -**The reviewed-head file is written on this route alone, and after the repair commit**, because -writing it is what makes a next call possible: recording it before the ordering has spoken is how a -pass gets issued over a standing source block, an unanswered suspension or a parked cycle — and -recording it before the repair lands would aim that call at a tree the repair is not in. - -**If a repair was already applied before its route authorized it, stop and report the concrete -state.** Do not design a way to take it back: this plan restores nothing, §A prescribes no rollback, -and inventing one here would be a new rule nobody approved. Report what is in the worktree, what is -committed, and which route had not yet spoken. - -**A non-closing pass that owes no repair still commits.** A Minor-only clean pass below the floor, -or an answered suspension that changes no artifact, leaves its findings files tracked and dirty — -and the next pass's files pile up beside them, after which the close procedure's dirty-set check can -never equal one pass's two files and a route the installed ordering requires cannot close through -this plan. **There is no "nothing to commit" branch here**: the findings files are always something. - -**The recorded head is what makes "a clean response against that exact `HEAD`" checkable.** Resolved -and written before the call, it is the value the close procedure compares against — without it, -a commit landing after the response is indistinguishable from the reviewed state. - -Re-review after every fix. **Revalidate the evidence entry before every re-review and before the -closing commit.** +**Each pass is issued, validated and read as §5 at `$BASE` says** — *Accept a pass only when*, +*Recovery: one attempt per pass*, *Severity*, re-review after every fix, and the evidence entry +revalidated before every re-review and before the closing commit. **Nothing here restates those.** +What this plan adds is where its records go and how a pass is routed. + +**Findings files are not committed during the loop (Close, *D1*).** They stay in +`.context/codex-reviews/` under this cycle's nonce and are committed once, all passes together, by +step 8. **No commit lands between a review and the reading of its result**, so the reviewed head +stays `HEAD` until a repair is committed. + +**After every valid pass, route it by §5 at `$BASE` — the rules this cycle started under** +(`CLAUDE.md` "When these rules bind"; Daniel's decision of 2026-09-25, option 1). **The ordering this +change installs does not steer this cycle.** It is the product under review, and Task 13 and design +§7 test it; applying it here as well would give this cycle two rule sets that disagree (pass 54, +finding 1). The routes, each with its §5 source: + +- **A new structural or contract question, or a correction outside the assigned fix set** → stop and + surface to Daniel (§5, *What a loop absorbs, and what stops it*). Nothing is repaired and no call is + issued until he answers; the loop resumes on the revised artifact. +- **From pass 4 on, two or more of the five tells** → stop and surface (§5: *"Any two present makes + stop-and-surface mandatory"*). **This holds for an otherwise closure-eligible pass too.** The + installed text's rule that tells never block a closing pass does not govern this cycle. +- **Clearly stuck**, all three of §5's conjuncts → stop and surface. +- **Blocker or Major findings inside the assigned fix set** → repair them (§5, *Severity*): the + candidate sequence below, repair branch. +- **Clean at or above the floor, or zero findings, no stop owed, but a re-review is owed** — the + evidence entry changed when 7b revalidated it (§5 at `$BASE`: the entry is revalidated before + every re-review and before the closing commit, and a changed entry is not covered by the pass that + judged the old one), or Failure's resume rule says a condition the failed attempt moved costs a + pass → **the candidate sequence below**, making only the change that rule requires: a changed + entry or refreshed records need no product change and take the no-repair branch; a moved closure + input takes the repair its own rule names. Then battery, 4c, evidence, call — a new review. This + pass does **not** close (pass 55, finding 2). +- **Clean at or above the floor, or zero findings, no stop owed and no re-review owed** → Close: + step 7b, then step 8 (§5: the final pass must be clean; the zero-finding early exit). +- **Free of Blocker and Major but below the floor** → collect the Minors and Nits and continue + without a product repair (§5: *"Minor · Nit → collect, never iterate"*, *"Below the floor nothing + closes"*): the candidate sequence below, no-repair branch. +- **Daniel's answer to a stop** decides what follows. Where he stops the cycle it waits, open and + spending no passes, until he says to continue; nothing closes it (§5: *"Surfacing does not close + the cycle"*). + +**What option 1 removed from this cycle's routing, accounted:** + +| Rule the plan applied to this cycle from the installed ordering | Now | +|---|---| +| Source-block route (release rule, re-read) | **deliberately removed for this cycle**: §5 at `$BASE` has no source-block concept. It stays product text, tested by Task 13 | +| Collect every suspension answer and compose them | **replaced by §5's form**: each stop is surfaced and answered | +| No repair to a finding whose membership is unanswered | **kept by §5**: an out-of-set or new-question finding stops before any repair | +| Stop answer → park, restarted only by an explicit continue | **kept in substance by §5**: surfacing does not close the cycle; it waits for Daniel | +| Clean completion takes precedence over tells | **deliberately removed for this cycle**: §5 at `$BASE` stops on two tells. It stays product text, tested by Task 13 | +| Continue → next pass | **kept**: the candidate sequence below | + +**Repair or continue → the candidate sequence.** It is the first pass's sequence — 4b's commit block, step 4, +step 4c, step 5, step 6 — with one branch, and the branch is on the **route's obligation**, never on +what the tree looks like: + +1. **A repair is owed on the route** → apply it, re-run every record-producing check it invalidated + (below), and commit the repair with the refreshed records **through the step-4b commit block**, + message `WIP: fix `. That block resolves the new commit once and writes it to + `.context/loop-rule-reviewed-head`. **No repair is owed** → no product file changes, but **the complete set + still runs before the call** (below). Where its records come out **identical** to the committed + ones — `git diff --quiet HEAD -- docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md` + succeeds — **commit nothing**; there is no empty commit. `HEAD` must then still equal + `.context/loop-rule-reviewed-head`; if it does not, something reached the repository between the + call and now — stop and report the concrete state, and never record over it (pass 53, finding 2). + Where the refreshed records **differ**, commit this plan alone through the step-4b commit block, + message `WIP: refreshed records` — a records-only candidate, and the only change this branch may + commit. **Anything else dirty is not an authorized record change**: stop and report it + (pass 54, finding 2). **"Else" is exact, not "everything but this plan"**: under D1 this cycle's + findings files are uncommitted throughout the loop, and its working record must survive for + recovery, so both are expected here and neither is committed (pass 55, finding 1). The check, + run before the records decision: + + ```bash + # NONCE and LASTP: this cycle's nonce and the pass just routed. Only this plan, this + # cycle's findings slots, its per-slot dispositions notes (§5's `-dispositions.md`) + # and its working record may be dirty. Exact paths, never a glob. + NONCE=; LASTP= + PLAN=docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md + set -- ":(exclude)$PLAN" ":(exclude).context/codex-reviews/gate-b-$NONCE-resume.md" + p=1 + while [ "$p" -le "$LASTP" ]; do + for b in spec quality; do + set -- "$@" ":(exclude).context/codex-reviews/gate-b-$b-$NONCE-pass-$p.md" \ + ":(exclude).context/codex-reviews/gate-b-$b-$NONCE-pass-$p-dispositions.md" + done + p=$((p + 1)) + done + # Two steps: a FAILED status prints nothing and would read as clean. + OUT=$(git status --porcelain -z --untracked-files=all -- . "$@") \ + || { echo "reading the dirty set FAILED — no call may be issued"; exit 1; } + test -z "$OUT" \ + || { echo "a path outside this plan and this cycle's own records is dirty — stop and report"; git status --porcelain --untracked-files=all; exit 1; } + ``` + + **Dispositions notes are allowed here and not at the close**: they are advisory, may sit beside + the findings during the loop, and step 8 asks for them to be removed first. +2. **Final verification on the candidate**: step 4's fetch and battery, then step 4c. Both read the + candidate from `.context/loop-rule-reviewed-head`. A non-zero or unresolved result stops here. +3. **Revalidate and write the evidence entry** (step 5). +4. **Issue the call** (step 6). + +**No commit is made between 1 and 4**, which keeps the verified object and the reviewed object the +same one. **If a repair was already applied before its route authorized it, stop and report the +concrete state** — this plan restores nothing, and §5 prescribes no rollback. **A fix that changes specified behaviour updates the spec in the same commit.** **The battery and the reader checks are not done when step 4 passed once.** A Gate-B fix can touch the hook, its test, either prompt copy or the target text, and step 4's run and step 4b's -twelve-item review both describe the tree as it was *before* that fix. Without re-running them the -plan reaches its closing commit on a tree that never passed its own quality battery — a repo unable -to pass CI, or an invariant-11 violation, published by a cycle that closed clean. +twelve-item review both describe the tree as it was *before* that fix. - **After each fix, re-run every check whose subject it changed — mechanical and reader alike.** The hook or its test means the suite under both shells; a shell file means `shellcheck`; either @@ -3192,27 +3368,69 @@ to pass CI, or an invariant-11 violation, published by a cycle that closed clean a hook message changed, Task 13's transitions and closure checks where §A changed, Task 14's parity and untouched checks wherever text moved at all. **Naming only the mechanical ones let a fix invalidate a reader record that then reached the closing commit unchanged.** -- **Persist every re-run record and commit it before the re-review**, so the reviewed range holds - it. A refreshed record left in the worktree is in neither the Gate-B range nor `reset --soft`. -- **Before the candidate final pass, re-run the complete set, and "complete" means mechanical as - well as reader.** The whole step-4 battery against the current `HEAD`; step 4b's twelve items; - every reader check above; **and every mechanical observation this plan built** — each - discriminating pair, each add-only presence check, each moved and dropped absence, each carried - and kept preservation count — re-run and re-recorded in `## Fragment evidence (per-task output)`. - **An earlier draft named only the battery and the reader checks**, so the clean pass could close - on a `HEAD` whose fragment evidence had never been re-established as a set, and the plan's - design §7 claim would rest on counts taken from an older tree. -- **Commit those records, resolve the new `HEAD`, and only then issue the candidate final pass.** - **Only a clean response issued against that exact `HEAD` closes the cycle.** A pass that was - clean against an earlier tree, plus records committed afterwards, closes on a tree no pass - reviewed — which is the same defect as reviewing the wrong range, arrived at from the other end. -- **One thing cannot exist before that pass: the pass's own findings files.** A `full` Gate-B pass - writes **two** — the spec and quality branch files — and the close's dirty-set check requires both, - so the **pair** is the sole permitted post-review addition, and step 8 asserts that it is the only - one — every other - path must already be in the reviewed `HEAD`. `.context/` moves none of the hook's fingerprint - inputs, so the file changes nothing the review looked at; what would be wrong is a *second* - delta riding along beside it. +- **Persist every re-run record and commit it before the re-review**, through the step-4b block, so + the reviewed range holds it. +- **Before every call, re-run the complete set, and "complete" means mechanical as well as + reader.** Every call can be the final one, because a zero-finding pass closes at any pass number; + on the first call Tasks 0–15 have just produced the set. The whole step-4 battery; step 4b's twelve items; every reader check above; + **and every mechanical observation this plan built** — each discriminating pair, each add-only + presence check, each moved and dropped absence, each carried and kept preservation count — re-run + and re-recorded in `## Fragment evidence (per-task output)`. **An earlier draft named only the + battery and the reader checks**, so the clean pass could close on a `HEAD` whose fragment evidence + had never been re-established as a set, and the plan's design §7 claim would rest on counts taken + from an older tree. + + **That set splits in two, and the split is what orders the candidate sequence above.** **Record-producing work** + writes into this plan and therefore **must be committed**: the repair itself, step 4b's twelve-item + result, every mechanical observation, and any other record a task owes. **Final verification** + writes **no** committed record: step 4's battery, whose outcome is named in the evidence entry + rather than in a plan section, and step 4c, whose result goes to `.context/loop-rule-closing-msg`. + **All record-producing work is finished and committed first, through the step-4b block, which + captures the candidate; only then does final verification run.** Reading it the other way is what left a rerun + requiring a commit after the candidate had already been fixed. + + **"The whole step-4 battery" includes step 4's guarded `git fetch` and its recording of the base + ref object id** — the battery is re-run end to end, not from the object id the first run happened + to persist. + + **A stale base really can turn a rejection into a pass, and it was demonstrated with the real + checker.** Topology, with the ids from the run: common ancestor `47c6193` carrying version + `0.11.0`; `main` advancing on its own to `c0df286` at `0.12.0`; a branch that also reaches `0.12.0` + **and carries further plugin changes**; and `R` = `44a2698`, the **merge commit**, which is what CI + checks out — `.github/workflows/ci.yml` passes no `ref:` to `actions/checkout`, so a + `pull_request` job checks the merge commit rather than the branch head. The plugin diff between + `main` and `R` is `plugins/dev-workflow/extra.md` and `plugins/dev-workflow/file.md`. Results: + + - `check-version-bump.sh old-A` → merge-base `47c6193`, base version `0.11.0` ≠ `0.12.0` → + **`ok`, exit 0**; + - `check-version-bump.sh main` → merge-base `c0df286`, base version `0.12.0` = `0.12.0` with the + plugin changed → **rejected, exit 1**. + + **An earlier attempt at this test concluded the opposite and was wrong.** The checker's line 58 is + `cd "$(dirname "$0")/.."`, so invoking it by absolute path from a disposable repository silently + ran it **against its own repository instead** — the checker under test never saw the fixture. That + is the "wired so it could not fail" defect this plan warns about, and it produced a confident + `could not be reproduced` that was an artefact of the harness. The run above copies the checker + **into** the fixture repository. + + **Re-fetching does not by itself close this**, and naming the fetch here is precision rather than + the repair: the battery runs against the **local branch head**, while CI evaluates the **merge + commit**, and no amount of refreshing the base ref makes those the same object. **The repair is + step 4c**, which checks the merge result of the pinned candidate head and the recorded base with + the unchanged checker. An earlier draft of this paragraph said checking the merge result was "a + scope question and not a change this plan makes" — **that is no longer true and was never a + conclusion the evidence supported**: a different checking mechanism is not by itself a new + requirement. + +- **What "the battery is bound to the candidate" does and does not mean.** The battery block reads + `.context/loop-rule-reviewed-head` and `.context/loop-rule-baseref`, and runs **shellcheck, the + hook suites, the invariant checkers and `claude plugin validate` in a disposable clone checked out + at that candidate id** (step 4). Its subject is the **committed candidate**, so a dirty work + repository cannot supply its inputs. It does **not** make the clone immutable while tests run, and + it does **not** check the merge result — step 4c does that, naming its own revision. Step 6 still + asserts that `HEAD` has not moved, which is what keeps the reviewed object the captured one. +- **Only a clean response issued against the recorded head closes the cycle** (Close condition 1). + This cycle's findings files are the sole permitted addition at the close (Close condition 2). - [ ] **Step 7b: After the clean pass, complete `.context/loop-rule-closing-msg` — this is the action step 5 defers to, and step 8 has no other source for these records** @@ -3226,10 +3444,10 @@ records.** The handoff changes nothing, but it **does not claim the failed opera as it was** — a hook can rewrite anything before failing — so "the failure preserves it" is not a statement this plan can make. -**And a failed act does not by itself owe a further pass.** §A retries the act where every closure +**And a failed act does not by itself owe a further pass.** Failure's resume rule retries the act where every closure condition still holds; a pass is owed only where the attempt or its repair moved something a condition is read from, **and that condition's own rule is what decides.** An earlier draft required -a fresh final pass unconditionally here, which forces a review §A does not ask for. Write the complete message from the current records at every candidate close, then +a fresh final pass unconditionally here, which forces a review neither §5 nor Failure asks for. Write the complete message from the current records at every candidate close, then assert **exactly one** provenance line and **exactly one** curve for this cycle before step 8. The message carries, in this order: @@ -3241,14 +3459,24 @@ The message carries, in this order: until the loop ends; 3. any **human-exception record**, and beside it the skip reason if a cycle was skipped; 4. the **revalidated evidence entry**, replacing step 5's draft if revalidation changed it — and - **if it changed, this candidate is over.** The final reviewer judged the entry it was handed + **if it changed, this candidate is over.** + + **"Rebuild from the current records" does not say how this item survives a rebuild, and that + gap is closed here.** Before writing anything, **read the existing entry out of + `.context/loop-rule-closing-msg` and keep that text**; it is the entry the final reviewer was + handed verbatim, and it carries the step-4c result, which is read from a run rather than from a + section and so cannot be regenerated by re-reading the record sections. Items 1, 2 and 3 are + rebuilt; **item 4 is carried across unless revalidation changes it.** Where the file is missing + or holds no entry, that is **not** a rebuild case: the candidate cannot be closed on an entry + nobody can produce — **stop and report**. Revalidation compares the carried text against the + current records; **equal means carry it, different means the candidate is over** and the existing + re-review rule below applies unchanged. The final reviewer judged the entry it was handed verbatim; a different entry in the closing commit is evidence no pass covered. **Take this pass - through step 7's three steps like any other non-closing pass** — record it, read it against the - ordering, and repair only on a route that authorizes one; a - standing source block, an open suspension or a stop answer binds here exactly as it does there, - and **only its continue result records a new reviewed head and issues another candidate.** An - earlier draft sent this branch straight to a new call, which is the plan's own Gate-B loop - stepping over the ordering it installs. **Revalidate before recording the candidate head and + through step 7 like any other non-closing pass** — route it, and repair only on a route that + authorizes one; a stop binds here exactly as it does there, and **only a repair, continue or + re-review route runs the candidate sequence and issues another call** — a changed entry takes the + re-review route. An earlier draft sent this branch straight + to a new call, which stepped over step 7's routing. **Revalidate before recording the candidate head and issuing the pass**, so that in the ordinary case this branch is never reached. **Items 1, 2 and 4 are owed unconditionally; item 3 is owed only where such a record exists.** @@ -3266,228 +3494,198 @@ write one, and a later reader could not tell it from an exception record that lo holds, `reset --soft` has already discarded every WIP body, and a closing commit without its provenance line or curve has not validly closed the cycle. -- [ ] **Step 8: Close the cycle — the close procedure, in two invocations** +- [ ] **Step 8: Close the cycle — one closing block, then the postcondition block** -**`## The four procedures` · Close states the six conditions, what success looks like, and why the -two invocations are separate.** Nothing here restates them. What follows is the shell that discharges -them, each block naming its condition; **the conditions are the obligation and this shell is one way +**`## The four procedures` · Close states the four conditions, what success looks like, and why one +closing command is enough.** Nothing here restates them. What follows is the shell that discharges +them, each part naming its condition; **the conditions are the obligation and this shell is one way to run them** — where the observed state is not one it expects, read the state and pick the operation, rather than extending the block. -**8a — record the candidate pass's findings files. Discharges conditions 1, 2, 3 (its pre-move -check) and 4.** +**Before it runs:** remove this cycle's dispositions notes from `.context/codex-reviews/` — §5 lets +them be deleted — or condition 2 stops the close on them. **Keep the working record**: the close +retires it last, after condition 4 has passed. + +**The closing block — conditions 1, 2 and 3, then the act.** ```bash -BASE=$(cat .context/loop-rule-base) -HEADREV=$(cat .context/loop-rule-reviewed-head) -test -n "$BASE" && test -n "$HEADREV" || { echo "BASE or reviewed head missing"; exit 1; } +BASE=$(cat .context/loop-rule-base) \ + || { echo "cannot read the base — NOT closing"; exit 1; } +HEADREV=$(cat .context/loop-rule-reviewed-head) \ + || { echo "cannot read the reviewed head — NOT closing"; exit 1; } +test -n "$BASE" && test -n "$HEADREV" || { echo "base or reviewed head empty — NOT closing"; exit 1; } +# NONCE and LASTP are this cycle's nonce and the candidate pass number — the values the +# calls' slot paths were built from. INCOMPLETE lists the pass numbers recorded as +# INCOMPLETE (space-separated, or empty); only their slots may be missing. +NONCE=; LASTP= +INCOMPLETE="" +# The working record: the one exact path left out of condition 2, never staged. +WR=".context/codex-reviews/gate-b-$NONCE-resume.md" # Condition 1. -test "$(git rev-parse HEAD)" = "$HEADREV" || { echo "HEAD is not the head the candidate pass was issued against"; exit 1; } - -# Condition 2. NONCE and P are this cycle's nonce and the candidate pass number, -# the same two the call's slot paths were built from. The exclusion pathspec is -# condition 2's "ask git which paths outside the expected pair are dirty". -NONCE=; P= -FINAL=".context/codex-reviews/gate-b-spec-$NONCE-pass-$P.md .context/codex-reviews/gate-b-quality-$NONCE-pass-$P.md" -set -- -for f in $FINAL; do set -- "$@" ":(exclude)$f"; done -for f in $FINAL; do - test -n "$(git status --porcelain --untracked-files=all -- "$f")" \ - || { echo "this pass's findings file is not dirty: $f"; exit 1; } -done -# Two steps: a FAILED `git status` prints nothing, and inline this reads as "nothing -# outside the findings files is dirty" — condition 2 satisfied by a comparison that -# never ran, and it is the only check between a foreign delta and the closing squash. -OUTSIDE=$(git status --porcelain -z --untracked-files=all -- "$@") \ - || { echo "reading the dirty set outside this pass's findings files FAILED"; exit 1; } -test -z "$OUTSIDE" \ - || { echo "something outside this pass's findings files is dirty:"; git status --porcelain --untracked-files=all; exit 1; } - -# Condition 3, its pre-move check — the last precondition, so nothing has moved if -# it fails. Reader check on the message's records; the expected result is -# condition 3's. 8b revalidates it immediately before the commit consumes it. -test -s .context/loop-rule-closing-msg || { echo "closing message missing or empty"; exit 1; } - -# Condition 4, first bullet: pin the validated blobs before staging. -# Guarded per iteration AND on the redirect. Only the last iteration's status survives -# a bare loop, and a failed redirect leaves an EARLIER attempt's pin standing — which, -# paired with the same failure below, makes the comparison compare two stale files and -# pass. The `>&2` matters: this compound's stdout IS the pin file. -for f in $FINAL; do - git hash-object "$f" \ - || { echo "pinning the validated blob for $f FAILED — nothing is staged" >&2; exit 1; } -done > .context/loop-rule-final-blobs \ - || { echo "writing the validated-blob pin FAILED — nothing is staged"; exit 1; } -# shellcheck disable=SC2086 -git add $FINAL \ - || { echo "staging FAILED — run Failure; do NOT commit"; exit 1; } -git commit -m "WIP: Gate-B findings files" || { echo "record commit FAILED — run Failure"; exit 1; } -# Resolve the record commit ONCE, here, and run every condition-4 check against that -# object id. `HEAD` is a movable ref: re-resolving it at each check — and again at the -# tail, where the tip is written — lets a move land a commit as the closing tip that -# satisfied none of the parent, path, blob-identity or eligibility checks. Condition 5 -# would then accept it, `reset --soft` would fold it into the closing commit, and the -# only later content check is the one this plan deliberately parks. -REC=$(git rev-parse HEAD) || { echo "cannot resolve the record commit — run Failure"; exit 1; } -test "$(git rev-parse "$REC^")" = "$HEADREV" || { echo "record commit's parent is not the reviewed head — run Failure"; exit 1; } -git diff --quiet "$HEADREV" "$REC" -- "$@" \ - || { echo "record commit changed paths beyond this pass's findings files — run Failure"; exit 1; } -# The polarity is inverted here: exit 0 means "no difference", i.e. NOT carried. With -# `&&` an execution error (status 2 or above) takes the same path as "carried" and the -# close proceeds on a comparison that could not run. Status 1 — a real difference — is -# the legitimate result this check wants. -for f in $FINAL; do - git diff --quiet "$HEADREV" "$REC" -- "$f" - case $? in - 0) echo "record commit did not carry $f — run Failure"; exit 1 ;; - 1) : ;; - *) echo "comparing $f between the reviewed head and the record commit FAILED — run Failure"; exit 1 ;; - esac +test "$(git rev-parse HEAD)" = "$HEADREV" \ + || { echo "HEAD is not the head the candidate pass was issued against — NOT closing"; exit 1; } + +# Condition 2, the set: the EXACT findings slots of passes 1..LASTP, never a glob, so a +# dispositions note or the working record is never taken for a findings file. +: > .context/loop-rule-final-paths || { echo "cannot write the findings list — NOT closing"; exit 1; } +p=1 +while [ "$p" -le "$LASTP" ]; do + for b in spec quality; do + f=".context/codex-reviews/gate-b-$b-$NONCE-pass-$p.md" + if [ -f "$f" ]; then + printf '%s\n' "$f" >> .context/loop-rule-final-paths \ + || { echo "cannot write the findings list — NOT closing"; exit 1; } + else + case " $INCOMPLETE " in + *" $p "*) test "$p" -ne "$LASTP" \ + || { echo "the candidate pass cannot be INCOMPLETE — NOT closing"; exit 1; } ;; + *) echo "missing findings file for a pass not recorded as INCOMPLETE: $f — NOT closing"; exit 1 ;; + esac + fi + done + p=$((p + 1)) done -# Condition 4, second bullet. -# Guarded the same way as the pin above, and for the same paired-stale reason. Note -# that on an unresolvable rev `git rev-parse` echoes its argument to stdout and exits -# 128, so without the guard this file can hold a junk line and the loop's status is gone. -for f in $FINAL; do - git rev-parse "$REC:$f" \ - || { echo "reading the committed blob for $f FAILED — run Failure" >&2; exit 1; } -done > .context/loop-rule-committed-blobs \ - || { echo "writing the committed-blob list FAILED — run Failure"; exit 1; } -diff .context/loop-rule-final-blobs .context/loop-rule-committed-blobs \ - || { echo "a findings file was rewritten between validation and commit — run Failure"; exit 1; } - -# Condition 4, third bullet: re-run the findings-file structural check on the -# COMMITTED content and re-establish this pass's eligibility. Reader check; the -# expected result is condition 4's. -TREESTATE=$(git status --porcelain) \ - || { echo "reading the tree state after the record commit FAILED — run Failure"; exit 1; } -test -z "$TREESTATE" || { echo "tree not clean after the record commit — run Failure"; exit 1; } - -# Condition 4's tail: the closing tip — the SAME object every check above ran against, -# never a fresh resolution, so a `HEAD` that moved after those checks cannot become the -# tip. Guarded even though it ends the block: a failed redirect leaves whatever an -# earlier attempt wrote, and 8b's condition 5 compares against that. (A failed write -# after a successful truncate leaves the file empty, which 8b's `test -n "$TIP"` does -# catch — the redirect failure is the half it cannot.) -printf '%s\n' "$REC" > .context/loop-rule-reviewed-tip \ - || { echo "recording the closing tip FAILED — run Failure; 8b must not run"; exit 1; } -``` - -**8b — reset and close. A separate invocation, carrying no `-m` option at all. Discharges condition -3's revalidation and conditions 5 and 6.** - -```bash -BASE=$(cat .context/loop-rule-base); TIP=$(cat .context/loop-rule-reviewed-tip) -test -n "$BASE" && test -n "$TIP" || { echo "BASE or TIP missing — 8a did not complete"; exit 1; } - -# Condition 5: the tip AND a clean tree, both read in THIS invocation. -test "$(git rev-parse HEAD)" = "$TIP" || { echo "HEAD has moved since 8a — NOT resetting"; exit 1; } -# Two steps, and this is the most expensive instance: the next command is the -# `reset --soft`. Both branches keep both markers — "NOT resetting", because nothing -# has moved in this invocation, and "run Failure", because 8a's record commit has. -TREESTATE=$(git status --porcelain) \ - || { echo "reading the tree state at 8b FAILED — NOT resetting; run Failure"; exit 1; } -test -z "$TREESTATE" || { echo "tree not clean at 8b — NOT resetting; run Failure"; exit 1; } - -git reset --soft "$BASE" || { echo "reset --soft FAILED — run Failure; do NOT commit"; exit 1; } - -# Condition 3's revalidation: re-read the message here, immediately before it is -# consumed. Reader check on its records; the expected result is condition 3's. -test -s .context/loop-rule-closing-msg || { echo "closing message missing or empty — run Failure"; exit 1; } -# Pin the bytes this invocation validated, BEFORE the commit runs. Condition 6's -# oracle must not be the same mutable path the commit reads: a `commit-msg` hook -# that rewrites git's copy AND this ignored source file would otherwise make the -# post-close comparison pass on a body nobody validated. -# Guarding the REMOVAL is what closes the demonstrated case. Unguarded, a DIRECTORY at -# this path let the block reach the commit: `rm -f` fails on a directory, `cp source -# dir` then succeeds by writing beneath it, both commands report success and condition -# 6 only finds the missing oracle after history has moved. Observed under sh, dash and -# bash — the old shape printed REACHED THE COMMIT with the path still a directory. -# The copy is guarded because it is the write, and a surviving earlier pin would be -# non-empty enough for condition 6's `test -s`. The `test -f` after them is a residual -# check on the destination's shape; no demonstrated path reaches it, since the removal -# guard fires first. None of this promises cleanup — a failed step leaves whatever is -# on disk and stops before the commit, which is the whole of what it guarantees. +# Condition 2, the check: anything dirty OUTSIDE that set stops the close — tracked, +# staged or untracked. Ignored scratch never appears in status, so nothing else is exempt. +# Two steps: a FAILED status prints nothing and would read as clean. +set -- +while IFS= read -r f; do set -- "$@" ":(exclude)$f"; done < .context/loop-rule-final-paths +OUT=$(git status --porcelain -z --untracked-files=all -- . "$@" ":(exclude)$WR") \ + || { echo "reading the dirty set FAILED — NOT closing"; exit 1; } +test -z "$OUT" \ + || { echo "something outside this cycle's findings files is dirty — NOT closing"; git status --porcelain --untracked-files=all; exit 1; } + +# Condition 3, immediately before the commit consumes it. Reader check on its records; +# the expected result is condition 3's. Then pin the validated bytes as condition 4's oracle. +test -s .context/loop-rule-closing-msg || { echo "closing message missing or empty — NOT closing"; exit 1; } rm -f .context/loop-rule-validated-msg \ - || { echo "removing the previous pin FAILED — run Failure; do NOT commit"; exit 1; } + || { echo "removing the previous message pin FAILED — NOT closing"; exit 1; } cp .context/loop-rule-closing-msg .context/loop-rule-validated-msg \ - || { echo "pinning the validated closing message FAILED — run Failure; do NOT commit"; exit 1; } -# A guarded `cp` is not enough on its own: where a DIRECTORY sits at the destination, -# `rm -f` fails and `cp source dir` succeeds by writing beneath it, so both commands -# report success, the commit runs, and condition 6 only discovers the missing oracle -# after history has moved. Require the destination to be the regular file just written. + || { echo "pinning the validated closing message FAILED — NOT closing"; exit 1; } test -f .context/loop-rule-validated-msg \ - || { echo "the pin path is not a regular file — run Failure; do NOT commit"; exit 1; } -# `--cleanup=verbatim` so the stored body is the validated bytes: git's default -# cleanup for -F strips trailing whitespace and collapses blank runs, and -# condition 6 compares bytes. It carries no `-m`, so `is_wip_commit` still misses it. -git commit --cleanup=verbatim -F .context/loop-rule-closing-msg + || { echo "the message pin path is not a regular file — NOT closing"; exit 1; } + +# Pin every selected findings file's bytes, staged or not: `git add` below stages exactly +# these worktree bytes, and condition 4 compares the commit with this pin. +while IFS= read -r f; do + b=$(git hash-object -- "$f") || { echo "pinning $f FAILED — NOT closing" >&2; exit 1; } + printf '%s\t%s\n' "$f" "$b" +done < .context/loop-rule-final-paths > .context/loop-rule-final-blobs \ + || { echo "writing the findings-blob pin FAILED — NOT closing"; exit 1; } + +# The act. A failed add has not moved HEAD, but the index may hold part of the set. +while IFS= read -r f; do + git add -- "$f" || { echo "staging $f FAILED — the index may hold part of the set; stop, do NOT reset" >&2; exit 1; } +done < .context/loop-rule-final-paths || exit 1 +git reset --soft "$BASE" || { echo "reset --soft FAILED — run Failure; do NOT commit"; exit 1; } +# `--cleanup=verbatim` so the stored body is the validated bytes; `-F`, never `-m`, so +# `is_wip_commit` cannot read the close as cycle-internal. +git commit --cleanup=verbatim -F .context/loop-rule-closing-msg \ + || { echo "closing commit FAILED — run Failure"; exit 1; } ``` -**Condition 6, in its own invocation — which is why it re-reads `$BASE`.** Every fenced block here is -a separate shell, so a variable assigned in 8b is empty in this one: +**The postcondition block — condition 4, in its own invocation, which is why it re-reads `$BASE`.** +Every fenced block here is a separate shell, so a variable assigned in the closing block is empty in +this one: ```bash BASE=$(cat .context/loop-rule-base) test -n "$BASE" || { echo "BASE empty or unreadable — cannot verify the close"; exit 1; } +NONCE= +WR=".context/codex-reviews/gate-b-$NONCE-resume.md" # retired last, below +# Resolve the closing commit ONCE and make it the subject of every predicate below. +# Condition 4 is not one test but five — subject, parent, tree, body, findings blobs — and reading +# `HEAD` separately for each lets a move or an amend in between hand them DIFFERENT +# commits while all five still pass. Condition 1 may read `HEAD` live because it +# compares ONE read against ONE recorded baseline; condition 4 has no baseline +# and composes, which is why the same licence does not extend to it. +CLOSED=$(git rev-parse HEAD) \ + || { echo "cannot resolve the closing commit — run Failure"; exit 1; } # Read under a guard: a failed `git log` yields an empty subject, which matches no # pattern, so the WIP arm never fires and the check is satisfied by a command that # failed. No `test -n` is added — a genuinely empty subject read successfully passes # today, and requiring one would be a new condition rather than a repair to this one. -SUBJECT=$(git log -1 --pretty=%s) \ +SUBJECT=$(git log -1 --pretty=%s "$CLOSED") \ || { echo "reading the closing commit's subject FAILED — run Failure"; exit 1; } case "$SUBJECT" in [Ww][Ii][Pp]:*) echo "closing commit still reads as a snapshot — run Failure"; exit 1 ;; esac -test "$(git rev-parse HEAD^)" = "$BASE" || { echo "closing commit's parent is not \$BASE — run Failure"; exit 1; } -TREESTATE=$(git status --porcelain) \ +test "$(git rev-parse "$CLOSED^")" = "$BASE" || { echo "closing commit's parent is not \$BASE — run Failure"; exit 1; } +TREESTATE=$(git status --porcelain -- . ":(exclude)$WR") \ || { echo "reading the tree state after the close FAILED — run Failure"; exit 1; } test -z "$TREESTATE" || { echo "tree dirty after the close — run Failure"; exit 1; } # `--pretty=format:%B` emits the stored message alone; the `%B` spelling appends a # trailing newline the source file has none of, which rejected a CORRECT close. -# 8b's `--cleanup=verbatim` is the other half — without it git stores its own +# The closing block's `--cleanup=verbatim` is the other half — without it git stores its own # tidied copy. Both observed in a disposable repository, so this diff is the -# byte equality condition 6 states, not a normalized stand-in for it. -test -s .context/loop-rule-validated-msg || { echo "no pinned message — 8b did not complete; run Failure"; exit 1; } -# Same shape as 8b's pin: remove the previous extraction first, then guard this one. +# byte equality condition 4 states, not a normalized stand-in for it. +test -s .context/loop-rule-validated-msg || { echo "no pinned message — the closing block did not complete; run Failure"; exit 1; } +# Same shape as the closing block's pin: remove the previous extraction first, then guard this one. # A failed redirect otherwise leaves an earlier body at the path and the `diff` below # masks the failure — it can even pass, when the earlier body was the same validated # message, so the close is accepted on bytes nobody read out of THIS commit. rm -f .context/loop-rule-landed-msg -git log -1 --pretty=format:%B > .context/loop-rule-landed-msg \ +git log -1 --pretty=format:%B "$CLOSED" > .context/loop-rule-landed-msg \ || { echo "extracting the committed body FAILED — run Failure; the comparison has no input"; exit 1; } diff .context/loop-rule-validated-msg .context/loop-rule-landed-msg \ || { echo "the committed body differs from the validated message — run Failure"; exit 1; } -``` - -**Run step 8's blocks under `sh` or `bash`, not `zsh`.** `for f in $FINAL` and the `git add $FINAL` -beside it rely on the unquoted variable **word-splitting into two paths** — which is why the -`shellcheck disable=SC2086` is there — and `zsh` does not split unquoted parameters by default, so -the whole string is taken as one filename and every command fails on a path that does not exist. -**Observed, not assumed**: the first verification run of this block was made under `zsh` and failed -exactly that way. **Nothing in step 8 needs `bash` specifically** — an earlier revision compared the -two message copies through process substitution, which `dash` has none of; the comparison now reads -two ordinary files. -**All four pass, and only then the cleanup Close's success recognition names:** - -```bash -rm -f .context/loop-rule-base .context/loop-rule-reviewed-tip .context/loop-rule-reviewed-head \ +# Condition 4, the findings blobs: every selected file, committed exactly as pinned just +# before staging. This compares pin-time bytes with committed bytes and nothing earlier. +while IFS="$(printf '\t')" read -r f b; do + test "$(git rev-parse "$CLOSED:$f" 2>/dev/null)" = "$b" \ + || { echo "committed findings file differs from its pin, or is missing: $f — run Failure"; exit 1; } +done < .context/loop-rule-final-blobs \ + || { echo "reading the findings-blob pin FAILED — run Failure"; exit 1; } +# The final pass's pair, re-read from $CLOSED against §5's "Accept a pass only when" and +# the clean-or-zero-finding reading it was judged on: reader check, expected result +# condition 4's. + +# The cleanup gate, as a COMMAND rather than a sentence. Everything above describes +# $CLOSED; the cleanup is about the repository as it stands now, so this is where the +# capture is re-asserted. A move between the close and this point must not be followed +# by deleting the recovery state, which is the only evidence of what was closed. +test "$(git rev-parse HEAD)" = "$CLOSED" \ + || { echo "HEAD moved during the close verification — run Failure; do NOT clean up"; exit 1; } + +rm -f .context/loop-rule-base .context/loop-rule-reviewed-head \ .context/loop-rule-baseref .context/loop-rule-closing-msg \ .context/loop-rule-untouched .context/loop-rule-baseline-diff.txt \ .context/loop-rule-sites .context/loop-rule-changed-sites \ .context/loop-rule-c.src .context/loop-rule-w.src \ .context/loop-rule-a.txt .context/loop-rule-b.txt .context/loop-rule-landed-msg \ .context/loop-rule-validated-msg \ - .context/loop-rule-final-blobs .context/loop-rule-committed-blobs \ - .context/loop-rule-records-commit + .context/loop-rule-final-blobs .context/loop-rule-final-paths \ + .context/loop-rule-start if ls .context/loop-rule-* >/dev/null 2>&1; then echo "cycle scratch survives the close:"; ls .context/loop-rule-*; exit 1 fi +# Retire the working record LAST — only now is the cycle's own commit its recovery source +# (§5) — and then require a clean tree with no exception. +rm -f "$WR" || { echo "retiring the working record FAILED — it survives; report it"; exit 1; } +TREESTATE=$(git status --porcelain) \ + || { echo "reading the tree state after retiring the working record FAILED"; exit 1; } +test -z "$TREESTATE" || { echo "tree not clean after the close's cleanup:"; git status --porcelain; exit 1; } ``` +**What the capture buys, and what it does not.** All four predicates now describe **one object**, +which is the defect it repairs. It does **not** establish that the object is the commit the closing block created: +a substituted commit would still have to be parented at `$BASE` and carry the validated body +byte-for-byte, which is much narrower but is not nothing. The re-assert before the cleanup closes the +window inside this block; it says nothing about the window between the closing commit and this block's first +line. + +**Run step 8's blocks under `sh`, `dash` or `bash`, not `zsh`.** The closing block builds its +exclusion list with `set --` inside a redirected `while` loop, which relies on POSIX shell +behaviour. **Nothing in step 8 needs `bash` specifically.** + +**The cleanup Close's success recognition names is the tail of the condition-6 block above**, not a +block of its own. It used to be separate, gated only by the sentence *"all four pass, and only +then"* — a prose gate in a different shell, which asserts an ordering rather than checking one. It +now runs after the `HEAD = $CLOSED` re-assert, so the gate is a command. + **The `ls` is what makes the removal checked rather than asserted**, and the list is every `loop-rule-*` name this plan writes. **The plan's own records are not among them** — they live in this plan and in `.context/codex-reviews/`, both tracked, both already inside the closing commit. From d26de4b40655a18c44c5cb3a3d8fd362943c8495 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sat, 26 Sep 2026 09:32:54 +0200 Subject: [PATCH 179/181] docs: record parked sparring-side work (dark-factory vision, todos, sfx field report) Committed on its own, before the loop-rule implementation starts, so that work is not left uncommitted in the tree the implementation's WIP commits and close checks read. Content unchanged; nothing here belongs to the loop-rule change. --- .../2026-09-17-sfx-review-loop-economics.md | 93 +++++++++++++++ .../specs/2026-08-30-dark-factory-vision.md | 111 +++++++++++++++++- todos.md | 74 ++++++++++++ 3 files changed, 276 insertions(+), 2 deletions(-) create mode 100644 docs/field-reports/2026-09-17-sfx-review-loop-economics.md diff --git a/docs/field-reports/2026-09-17-sfx-review-loop-economics.md b/docs/field-reports/2026-09-17-sfx-review-loop-economics.md new file mode 100644 index 0000000..f7d0d9f --- /dev/null +++ b/docs/field-reports/2026-09-17-sfx-review-loop-economics.md @@ -0,0 +1,93 @@ +# Field evidence — sfx-bricks-api-builder review-loop economics + +Observed 2026-09-17 at Daniel's request. Bounded, read-only inspection of the +consumer project; no tests, gates, instrumentation or updates were run there. +This report informs the [dashboard and usefulness requirements](../superpowers/specs/2026-08-30-dark-factory-vision.md) +and does not authorize changes to either project's review rules. + +## Snapshot and workflow provenance + +Consumer: `sfx-bricks-api-builder`, branch `calendar-prototype`, HEAD +`1f9f93234355b883da8c406d5a955f909437600e`. No tracked changes were reported; +16 untracked status entries were present, including review directories. +All consumer paths below are relative to that repository, at this snapshot +unless stated otherwise; its artifacts are not copied into the kit. + +- **Observed installation:** the user-scoped entry in + `~/.claude/plugins/installed_plugins.json` names + `dev-workflow@dev-workflow-kit` **0.9.1**, commit `baa75c1516dccff5e9fe04fd6b1b6bb3a5ad4dad`. + The cached plugin manifest independently says 0.9.1. User settings enable it; + inspected project/ancestor settings contain no plugin override. The kit's + source manifest currently says **0.11.0**. No plugin update was performed. +- **Observed local rules:** `CLAUDE.md` still requires a fixed three-pass floor + (line 80); it lacks the kit's newer five-tell stop procedure and detailed + consequence-based severity procedure. Its blob is + `d21bec36e998f151679bfaf2f384cb249873f9d2`, last changed by `b50f42c`. + Installed package version and project-local rule revision are separate facts. +- **Unverified:** which version the currently running agent process loaded, and + the exact package/rule combination used by each historical pass. The latest + inspected session log supplied no cache-path reference that settled this. + These cycles are not an effectiveness test of the current 0.11.0 rules. + +## Recounted observations + +Source directory: `docs/superpowers/reviews/2026-09-16-poi-sync-gate-a/`. +Counts below were recomputed from severity-prefixed finding lines in the +`gate-a-poi-sync-{spec,plan}-pass-N.md` files; each matched its final terminator. +This verifies counts, not finding validity, deduplication or pass acceptance. + +| Cycle | Findings by pass | Blockers by pass | Majors by pass | +|---|---|---|---| +| POI specification | 74 → 52 → 42 | 20 → 15 → 13 | 47 → 33 → 18 | +| Joint POI plan/specification | 46 → 33 → 19 → 30 | 11 → 5 → 4 → 6 | 29 → 23 → 13 → 18 | + +The plan pass-4 dispositions label 16 of 30 findings as induced by the previous +revision, with their item numbers. That attribution is the author's report; +it was not independently reconstructed across all revision diffs. The same +corpus records repairs to Minor/Nit findings too. This demonstrates recorded +repair scope, not that those repairs caused later findings or wasted time. +No attributable per-pass cost/duration series was established by this inspection. + +The primary findings mix very different consequences. Plan pass 4 item 2 concerns +an in-flight write reopening an abandoned run after tombstone deletion; item 5 +specifies a private method that another class must call; item 20 names unsupported +verification harnesses. Those are runtime design and verification consequences, +not merely textual polish because the findings live in a plan/specification. +Their historical reports were read; their failure scenarios were not executed. + +## Lessons adopted into the planning requirements + +1. **Separate origin, consequence and effort.** A repair-induced finding can + still expose a serious product defect. Record origin with evidence and + attribution confidence separately from severity, affected behaviour and + repair effort. Count optional Minor/Nit repairs separately from required + fixes. Declining raw counts or many applied repairs alone do not prove value. +2. **Classify instruments by their actual effects.** The source of + `docs/measurements/2026-09-12-request-attribution/restore.py` declares four + product targets and replaces their files; `instrument.py` transforms those + targets. Their location under `docs/` is not distance from product impact. + The package README and `.context/codex-reviews/gate-b-c1-group-6-quality-pass-5.md` + record earlier destructive restore and evidence defects. The present source + contains digest/preflight checks; no claim is made that the historical bugs + still exist or that the current tools were validated by execution here. + Earlier effort escalation for peripheral work must retain scrutiny of + destructive effects and invalid verification results. +3. **Version the measurement context.** Display installed workflow version, + observed loaded version when available, local rule revision and any explicit + overrides separately. Preserve unknowns. Compare cycles within documented + rule/profile/artifact contexts rather than treating all passes as equivalent. + An installation update must not be assumed to synchronize local rules. +4. **Separate implementation, verification and acceptance status.** At the + snapshot, `docs/HANDOVER.md`'s top section reports G1 implemented at WIP + `25c8838` with Gate B in progress; current HEAD is `1f9f932`. Older sections + still describe earlier unimplemented states. Those are historical reports, + not current verification of this HEAD. A dashboard must bind evidence to its + revision, preserve the historical account and show a mismatch as unknown or + needing reconciliation. It must not collapse implemented, reported tested, + reviewed, accepted and released into one green status. + +The consumer records deliberately unclean stops and owner decisions to proceed. +These are observations of that project's decisions, not a clean-gate result or +a proposed bypass for the kit. One project's historical curves under incompletely +attributed rules cannot set numeric utility thresholds or establish that stopping +earlier improves product outcomes. Step 2c still owes calibration and its limits. diff --git a/docs/superpowers/specs/2026-08-30-dark-factory-vision.md b/docs/superpowers/specs/2026-08-30-dark-factory-vision.md index 843ecbc..fb9e0b3 100644 --- a/docs/superpowers/specs/2026-08-30-dark-factory-vision.md +++ b/docs/superpowers/specs/2026-08-30-dark-factory-vision.md @@ -13,6 +13,10 @@ from an external video (four loop maturity stages; artifacts as the only handoff between fresh-context nodes; a "dark factory" as a repository that ships its own code, policed by an adversarial model with sampled human audit). +**Follow-up decision, 2026-09-17 (Daniel):** record generated HTML views and +review-loop usefulness assessment in the roadmap. This authorizes the planning +additions below, not their implementation or a change to current review rules. + ## 1. The vision Stories and ideas flow into a pool. The factory turns approved pool items into @@ -271,6 +275,94 @@ the inspiration: the adversarial verifier is a different model *family* LLM summaries only on demand. Rejected: a daemon/TUI and a hosted artifact page (account-bound). The wave plan and the dashboard share one renderer — the plan is the forecast, the dashboard is the now. +- **HTML is a generated presentation, not another maintained source + (decided 2026-09-17).** Existing Markdown todos, stories and planning + documents remain the authored sources; task lists, task details and progress + views are rendered from them and the available workflow records. Do not + replace them with hand-maintained HTML or maintain status in both formats. + The owning story must define stable item identities, explicit status, + priority, dependencies, parking triggers and evidence links. Missing values + remain unknown; an unchecked backlog box alone does not establish active + work, and raw checkbox counts are not a completion percentage. Every view + shows its source revision and generation time; live views also show the + snapshot/tick freshness above. A first static snapshot view may precede live + telemetry, provided it is labelled as a snapshot. This does not decide the + future pool's storage format (4a). + **Consumer evidence, 2026-09-17:** the [SFX field report](../../field-reports/2026-09-17-sfx-review-loop-economics.md) + adds two source requirements: display installed workflow version, observed + loaded version (or unknown) and project-local rule revision separately; + bind implementation, test evidence and review/acceptance status to their own + revisions. An older handover or a reported test result must not silently + become verification of the current HEAD, nor a single green completion state. +- **Review-loop usefulness is a separate dashboard requirement + (recorded 2026-09-17; not implemented).** Step 2c owns the measurement and + assessment design, using P8's evidence where available without widening P8. + Begin with visible dimensions and an explained traffic-light assessment; + a composite score follows only once its calibration is supported by data: + - yield: confirmed, distinct material findings, linked to their evidence; + - repair effects: recurrence, reopened findings and findings introduced by + a repair, with uncertain attribution labelled as such; + - effort: elapsed time, tokens and cost where attributable; + - coverage: reviewed scope and known gaps, including evidence limitations. + Few findings do not establish poor usefulness or sufficient coverage, and + many findings do not establish high usefulness. A clean verification pass + can be useful. An instrument finding is judged by its consequence, not + discounted solely because it concerns the instrument. + **Product impact governs review effort (Daniel, 2026-09-17).** The more + indirect the evidenced effect on the operating product, the less tolerance + there is for additional review and repair rounds without a concrete failure + consequence and a justified expected benefit. Distance means the causal path + to product behaviour, not the file extension: shipped prompts can act directly + on the product, and plan shell commands can invalidate its verification. + Product behaviour, security and data integrity receive the strongest scrutiny. + Execution and verification machinery is checked for reliable outcomes; + further hardening or optimization must justify its benefit. Explanatory prose, + presentation and hypothetical edge cases without an operative consequence + receive less effort and do not justify continued repair loops. + + **Lower thresholds mean earlier reassessment, not lower severity.** The + threshold design must distinguish direct product work, plan/execution + machinery and non-operative material, with earlier warnings and escalation + for repeated instrument work when its marginal benefit is unsubstantiated. + Unknown impact requires clarification, not an automatic low-risk label. + A real security or correctness defect is not discounted by its location; + false-green, false-red and valid-change-blocking failures retain the existing + consequence-based assessment. Purely hypothetical robustness gains and + performance tuning of rarely executed helpers need a demonstrated use case + and material cost to justify further rounds. + Show time and repair rounds spent on the instrument alongside evidenced + benefits and any observable delay to product work; a high instrument share + is a warning, not proof of waste. Escalation offers simplification, a different + review method or stopping the approach, without automatically continuing + repairs or advancing to implementation with unmet conditions. + **Calibration acceptance cases:** distinguish harmless explanatory polish + from a plan command that corrupts review evidence; distinguish speculative + helper optimization from an observed product bottleneck; keep a useful clean + verification pass distinct from a pass with unknown coverage. Record the + evidence and expected recommendation for each, rather than assigning priority + solely from the artifact's name. This is a future assessment requirement, + not the experimental one-repair-round cap parked in `todos.md`. + + **Field-informed calibration requirements:** the SFX report above supplies + recounted curves, not calibrated cutoffs. Record finding origin (including + repair-induced, with attribution evidence/confidence) separately from its + product consequence and effort; distinguish optional Minor/Nit repairs from + required fixes. Classify helper tools by the files/state they can change and + decisions they affect, not by a `docs/` path. Compare cycles under documented + workflow/rule revisions and profiles; preserve unknowns rather than claiming + an efficiency improvement from raw counts or mixing incompatible contexts. + + **Before implementation:** specify each metric's source, counting and + deduplication rules, comparison window and missing-data behaviour; define + candidate warning/escalation thresholds and any score weights; evaluate + them on recorded cycles and document their limitations. Link each proposed + threshold to an explained recommendation (continue, change review method, + or stop and surface). Missing evidence must not become a green assessment. + No numeric thresholds or weights are settled here. Existing pass floors, + mandatory tells, finding-resolution duties and closure conditions remain + unchanged: the assessment cannot waive them or close an unclean cycle. + Automatic actions or changes to those rules need a separately authorized + design; this entry introduces neither. - **Waves structure a new project.** Phase 0 assigns every initial pool item a wave mark (wave 1, 2, … or named milestones) — the deliberate "these subareas develop together first, those later" decision, usually aligned with @@ -346,7 +438,11 @@ proceeds, the breaking part waits on the meta-story). does not do: run analytics with cost and duration per step, a trace ID carried through every stage artifact, mechanical spec-delta capture, and live cost counters. That is new state and new instrumentation, so it - cannot ride inside 2b. + cannot ride inside 2b. It also owns the review-loop usefulness metrics, + threshold calibration and explainable assessment specified in §4; + dashboard presentation consumes those results. P8 alone cannot supply + all of these inputs. The generated HTML view has a separate owning-leaf + question below (§11), including a possible initial static snapshot. 3. Orchestrator story (codify the role as skill/agent per decision 1) — the role, the artifact handoff, and the *interface* of the tick execution plan. The working dry-run reads the pool, the projection, waves, lanes and the @@ -678,7 +774,18 @@ leaf yet owns. Those are marked as such rather than counted as decomposed. the live-status view in §4 (snapshot per tick, rendered by code into status.html / status line / optional menu bar, decision queue first, staleness visible); still open: the snapshot schema and the owning leaf, - fed by 2c's traces and analytics. [2c / dashboard] + fed by 2c's traces and analytics. The 2026-09-17 source/presentation decision + in §4 adds generated task and progress views with Markdown sources; choose + the initial static-view scope and source mapping in that leaf, without + requiring live telemetry for a labelled snapshot. [2c / dashboard] +- Review-loop usefulness — the requirement and dimensions are recorded in + §4; metric definitions, evidence availability, calibration sample, numeric + thresholds by evidenced product impact, any composite weights and + recommendation mapping remain to be designed. The direction is settled: + indirect impact requires earlier reassessment of further effort, not an + automatic severity demotion. Distinguish usefulness assessment from §10's proposed minimum + review cost/duration signal: spending longer or more is not proof of useful + review. Neither proposal changes current pass-validity rules. [2c] - Hooks and shortcuts — decided 2026-08-31 and recorded in §9 (shortcuts are shorter lanes, never side doors; hooks are subscribers to station-boundary events, passive or story-creating). Still open: the diff --git a/todos.md b/todos.md index 4aff5b1..659b42a 100644 --- a/todos.md +++ b/todos.md @@ -348,6 +348,39 @@ driven by recurrence rather than by enthusiasm. is still the shape recurring; excluding it would tune the count to who was watching. Four occurrences of the timing gap now, three of them benign; the dangerous `tree_hash` consumer above is still the one that decides this row's priority. +- [ ] **Hook repository context can differ from the operation's target worktree.** A **distinct + cause from the timing row above**, and that row's occurrence count is deliberately left + unchanged: timing is "the staged set was empty because `git add` had not run yet", + this is "the staged set was read in the wrong repository". **Reported observation + (2026-09-21):** a docs-only commit of two `docs/**.md` paths in the linked worktree + `dwk-claude-init` received the Gate-B STOP **although `git add` had run in a separate + Bash call**, which is the control the timing row's `29da026` data point relies on. The + index readings behind that account are reported, not re-reproduced here. + **Verified mechanism, read from the source:** the hook derives its root from **its own + Git process context** — `repo_root=$(git rev-parse --show-toplevel)` + (`plugins/dev-workflow/hooks/codex-gate.sh:42`) — and uses that root both for the state + paths (`state_dir="$repo_root/.context"`, `:43`, holding `codex-gate.gateB`, + `.passCount`, `.freshCount`, `.passCountA`) and for the staged-file inspection + (`files=$(git -C "$repo_root" diff --cached --name-only)`, `:914`). **Do not state this + as "it always uses the original session worktree"** — what is established is that the + root comes from the hook's own context, which need not be the operation's target. + **Two consequences derived from the code and NOT reproduced.** First, a **false + docs-only exemption**: where the foreign root's staged list is non-empty and passes + `is_docs_only` — which also requires clearing its `is_prompt_path` exclusions, so + arbitrary Markdown is not enough — the hook can report "Gate B N/A" while the actual + commit carries product files. That is the dangerous direction, so **invariant 2 is not a + blanket answer for this row**, even though the observed instance was a benign false + alarm. Second, a **foreign state reset**: the non-WIP commit branch runs + `rm -f "$state_file" "$count_file" "$fresh_file"` (`:886-888`) under the same + `$repo_root`, so it can clear another worktree's Gate-B fingerprint and counters. + **Operational precaution, not a repair:** starting a session directly in the intended + worktree keeps the two contexts aligned; a `cd` **inside** a Bash call does not, because + the hook runs before that command. **Implementation deliberately undecided** — the + payload's cwd, the hook process's cwd and an operation's explicit target must not be + assumed equivalent, and picking one is the design question, not a detail. **Trigger:** + Daniel explicitly selects a bounded reproduction-and-design task. That task should + separate the reported false alarm, the possible false exemption and the possible foreign + reset, and establish each before a repair is proposed. This entry starts none of them. - [ ] **No regression test for a `git add`/`write-tree` failure inside the throwaway index.** Derived from the code, not recalled: sections 24a-24e stub FIVE failure shapes — every checksum tool failing silently, a checksum printing a token then failing, the @@ -570,6 +603,47 @@ backlog. because it is followed. *Trigger: 3–5 real stories completed in a product project* — fewer than that and the state machine would be modelled on this repo's own atypical usage. +- [ ] **Generated status HTML — task, progress and KPI views.** Direction agreed + with Daniel on 2026-09-17: keep Markdown as the authored source and generate + HTML views, with source links, revision and freshness visible. Define item + identity and explicit status; parked, rejected and completed work must not + be reduced to a raw checkbox completion percentage. A labelled static + snapshot may precede live telemetry. Owner/scope is the dashboard leaf to + be defined in `docs/superpowers/specs/2026-08-30-dark-factory-vision.md` + §§4/11; this does not activate P1 or decide the future pool's storage. + *Trigger: Daniel explicitly selects the dashboard story for design.* + Recording this direction does not authorize implementation. + **SFX field input (2026-09-17):** expose installed/observed-loaded workflow + versions and local rule revision; bind status and evidence to their own + revisions instead of treating an older handover as current verification. +- [ ] **Review-loop usefulness — metrics, scoring and calibrated thresholds.** + Requirement recorded with Daniel on 2026-09-17; owned by vision step 2c, + presented in the dashboard. Define confirmed distinct finding yield, + recurrence/repair effects, effort and coverage evidence; start with + separate indicators and an explained traffic light. Before implementation, + specify data sources, deduplication, comparison windows, missing-data + handling, thresholds and any composite weights, and evaluate them against + recorded cycles. **Priority principle agreed 2026-09-17:** the more indirect + the evidenced product impact, the earlier further review effort must be + reassessed. Define lower warning/escalation thresholds for repeated plan + instrument work with unsubstantiated marginal benefit, not lower severity + by file type. Track instrument time/rounds, evidenced benefit and observable + delay to product work. Include §4's calibration cases; preserve the severity + of defects that invalidate product verification. This does not activate the + experimental one-round cap parked above. + **SFX field input (2026-09-17):** separate finding origin, consequence and + effort; record optional Minor/Nit repair scope and attribution confidence. + Instrument classification follows actual effects on product files/state. + Counts independently recounted; causal attribution and test outcomes remain + reported. Evidence and calibration limits: + `docs/field-reports/2026-09-17-sfx-review-loop-economics.md`. + Unknown evidence must remain visible. Recommendations + may support continuing, changing method or stopping to surface; they do + not waive floors, mandatory tells or closure conditions. P8 remains passive + and supplies only the evidence it has. Details and unresolved decisions: + `docs/superpowers/specs/2026-08-30-dark-factory-vision.md` §§4/7/11. + *Trigger: step 2c is explicitly picked up for design.* No thresholds are + activated by this entry; automatic actions require separate authorization. - [ ] **P7 — `workflow-doctor`, extracted from the `/workflow-init` preflight.** Not a second implementation of the same checks: the point is a **single shared check source** that both the initializer and the doctor call, or the two drift and the From 52a7aedde33edac667616d7d19d3038cb3b115cd Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sat, 26 Sep 2026 13:24:52 +0200 Subject: [PATCH 180/181] =?UTF-8?q?Install=20the=20=C2=A75=20closure=20ord?= =?UTF-8?q?ering=20and=20replace=20what=20it=20falsifies=20(dev-workflow?= =?UTF-8?q?=200.13.0)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit One closure ordering now decides how a review pass is read and when a cycle may close, installed in CLAUDE.md §5 and in /dev-workflow:workflow-init's template: source block, clean completion, suspension, continue; eligibility (clean at or above the floor, or zero findings) plus every closure condition plus the gate's closing act. The absorb paragraph owns the assigned fix set, Mechanics · Severity answers the demotion question it had handed over, the one-contract paragraph gains a membership test, and the twenty-three standing sentences the ordering falsified are replaced — sixteen in the two prompt copies, seven in the hook's gate reminders, which now report the hook's own state instead of a gate verdict. No hook logic changed; the hook suite's expectations move with the strings. Plan: docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md. cycle t57gp3hwu1; floor 3 per {docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md (level 2)}; hook reminder threshold absent cycle t57gp3hwu1; Gate B (passes 1-5, codex): Findings 10,10,13,13,15. Blockers 0,0,0,0,0. Majors 3,3,1,5,0. The curve counts logical passes and the reviewer's raw severities. Passes 1-3 were one `reviewType: full` call each; passes 4 and 5 were two sequential calls each (spec, then quality) against the same full base and head ids, so the hook counted seven calls for five logical passes. The pass-3 Major was demoted by the author and that demotion was later withdrawn as unsupported; the pass-4 Majors were one product regression (the dropped `--soft`, repaired) and four gaps in the plan's verification instrument, answered by Daniel's decision of 2026-09-26. Closed on pass 5: Blocker- and Major-free at the floor. Two loop-health tells were present from pass 4 on (a rising finding count, findings clustering on the instrument) and were surfaced to Daniel, who decided to close. Collected, not repaired: the observation procedure's unchecked Git exit statuses, its same-second start boundary and its reset-to-identical-content case; docs/getting-started.md and docs/coding-workflow.md still describe the old hook output and pass cadence; the Named residual's "standing sentences above"; the 9a/9 parity row; a trailing space in the plan. Human exceptions: none Evidence entry — docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md (mode read fresh from its header: battery+check+verification) Battery: the full AGENTS.md quality command, run in a disposable --shared clone checked out at the candidate 1bea70ed133bbe31a688e2bd700a2ad0d11d70eb with base c00f705768f2d1874bd3e4e09b3d10596f457931 (origin/main, fetched) — exit 0: both hook suites all passed (sh and dash), invariant checks 148 assertions, version-bump suite 36 assertions and check ok, claude plugin validate --strict passed. Step 4c (merge result): head 1bea70ed133bbe31a688e2bd700a2ad0d11d70eb, base c00f705768f2d1874bd3e4e09b3d10596f457931, merge result 1bea70ed133bbe31a688e2bd700a2ad0d11d70eb (the base is an ancestor, so the merge is the head); plugin diff base..R: plugin.json, CHANGELOG.md, commands/workflow-init.md, hooks/codex-gate.sh, hooks/codex-gate.test.sh; raw checker exit 0 — VERIFIED. Check (fails without the change) — the fragment observations, every one recorded with its counts in the plan's "## Fragment evidence (per-task output)", the OLD fragments in its fragment table (P1–P103, F1–F14, F7b), each counted in the worktree and in the base commit d26de4b40655a18c44c5cb3a3d8fd362943c8495. Every discriminating pair reads old/worktree=0 old/parent=1 new/worktree=1 new/parent=0: the old wording is present without the change and gone with it, the new wording absent without it and present with it. - Task 1 (§A, add-only, counterfactual ABSENT and claimed as absent): 3 presence checks x 2 copies. - Task 3 (§B): 9 pairs x 2 copies, 1 add-only presence x 2, 9 carried preservations x 2. - Task 4 (§C): 3 pairs x 2, 1 add-only presence x 2, 5 moved conditions (absence at the source + presence in §A) x 2, 7 preservations x 2 (3 kept prefix, 3 carried, the plateau rationale). - Task 5 (§D): 1 pair per copy (P5 C, P5w W), W pronoun presence, pointer presence x 2, 9 preservations (e11 C only). - Task 6 (§E): 2 pairs x 2, 3 dropped absences (g4 C only). - Task 7 (§G, §H): 14 pairs x 2, a13 first-sentence absence x 2, 4 moved (a17–a20) x 2, 5 add-only presences x 2, 20 preservations x 2 (9 carried, 11 kept in passage (i)). - Task 8 (§F items 1–9): 15 pairs x 2 (new/worktree total 15 per copy), 2 carried preservations x 2. - Task 9 (§F items 14, 18): 2 pairs x 2. - Task 10 (hook, §F items 10–13, 15–17): 8 pairs in codex-gate.sh; the hook suite's exact-match expectations are the second observation. - Task 11: codex-gate.test.sh swept, 91 sites recorded one line each; verdict vocabulary 0. - Task 0 / Task 2 / Task 14 step 4b: 7 untouched spans per copy — no difference; 4 kept conditions sharing a line with changed text — parent=1 worktree=1 in each copy. Verification (named) of the risk path — the closure ordering's transitions: the next-state table in the plan's "## Next-state table (Task 13 output)", 45 rows, each with one next state, every row passes the oracle; its claim width is transitions once the predicates are established, not how each predicate is derived, nor that the rows cover every reachable combination. Beside it, 13 per-condition closure checks and 2 separate checks (logical-pass-validated, conditions-held-at-act). They are defined and demonstrated in disposable repositories; none is applied to a real closing act here, since this cycle closes under §5 as at its base (Daniel's decision of 2026-09-26, recorded in design §7). The rows that read history (close-header-during-pass over the pass's own interval, close-profile-fixset over the longer window from every input of the set definition) report change observed / no change observed / source unreadable, never "held"; their procedure observe-header-changes was run on eight cases under sh and dash — linear, merge-only, amend-hidden and reset-and-back changes observed; a side-branch change and a change after the window not reported; an uncommitted edit-and-restore not observed, which is the stated blind spot. Parity (design §6): "## Divergence list (Task 14 output)" — 34 site regions C against W, 31 equal, 3 differing only in kept text beside a block and classified; one inherited divergence aligned (C's missing blank line before "Every pass report states"). b11/b13 equivalence: "## b11/b13 equivalence (Task 12 output)" — equivalent in both directions per copy; §A's continue-branch gloss now scopes each exemption to its own trigger (repaired at Gate-B pass 1 in both copies and in the target text). Reader records: "## Completeness sweep (Task 12b output)" (nothing found) and "## Prompt-standards result (Task 15 step 4b output)" (all twelve pass). Product repair at Gate-B pass 4: Finishing the cycle again names `git reset --soft ` (target §F item 5 and both copies); observed in a disposable repository, a mixed reset leaves nothing staged and the closing commit fails, a soft reset keeps the WIP tree. Deviation from the reviewed plan: version 0.12.0 -> 0.13.0 instead of 0.11.0 -> 0.12.0, by Daniel's decision recorded in 47e2d94 (main had reached 0.12.0). Process deviation: the pass-1 repair round and the pass-2 records refresh were made without Daniel's prior go; kept as the starting point by his decision of 2026-09-26, not retroactively authorized. Review provenance of passes 1–3 established from the original Codex transcripts (plan, "Gate-B provenance and deviation"). --- .../gate-b-quality-t57gp3hwu1-pass-1.md | 4 + .../gate-b-quality-t57gp3hwu1-pass-2.md | 5 + .../gate-b-quality-t57gp3hwu1-pass-3.md | 7 + .../gate-b-quality-t57gp3hwu1-pass-4.md | 8 + .../gate-b-quality-t57gp3hwu1-pass-5.md | 9 + .../gate-b-spec-t57gp3hwu1-pass-1.md | 8 + .../gate-b-spec-t57gp3hwu1-pass-2.md | 7 + .../gate-b-spec-t57gp3hwu1-pass-3.md | 8 + .../gate-b-spec-t57gp3hwu1-pass-4.md | 7 + .../gate-b-spec-t57gp3hwu1-pass-5.md | 8 + CLAUDE.md | 765 +++++++++++--- .../2026-09-14-loop-rule-consolidation.md | 953 +++++++++++++++++- ...26-09-10-loop-rule-consolidation-design.md | 9 + ...-10-loop-rule-consolidation-target-text.md | 6 +- .../dev-workflow/.claude-plugin/plugin.json | 2 +- plugins/dev-workflow/CHANGELOG.md | 30 + .../dev-workflow/commands/workflow-init.md | 760 +++++++++++--- plugins/dev-workflow/hooks/codex-gate.sh | 14 +- plugins/dev-workflow/hooks/codex-gate.test.sh | 182 ++-- 19 files changed, 2441 insertions(+), 351 deletions(-) create mode 100644 .context/codex-reviews/gate-b-quality-t57gp3hwu1-pass-1.md create mode 100644 .context/codex-reviews/gate-b-quality-t57gp3hwu1-pass-2.md create mode 100644 .context/codex-reviews/gate-b-quality-t57gp3hwu1-pass-3.md create mode 100644 .context/codex-reviews/gate-b-quality-t57gp3hwu1-pass-4.md create mode 100644 .context/codex-reviews/gate-b-quality-t57gp3hwu1-pass-5.md create mode 100644 .context/codex-reviews/gate-b-spec-t57gp3hwu1-pass-1.md create mode 100644 .context/codex-reviews/gate-b-spec-t57gp3hwu1-pass-2.md create mode 100644 .context/codex-reviews/gate-b-spec-t57gp3hwu1-pass-3.md create mode 100644 .context/codex-reviews/gate-b-spec-t57gp3hwu1-pass-4.md create mode 100644 .context/codex-reviews/gate-b-spec-t57gp3hwu1-pass-5.md diff --git a/.context/codex-reviews/gate-b-quality-t57gp3hwu1-pass-1.md b/.context/codex-reviews/gate-b-quality-t57gp3hwu1-pass-1.md new file mode 100644 index 0000000..80f7cca --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-t57gp3hwu1-pass-1.md @@ -0,0 +1,4 @@ +MINOR | high | docs/getting-started.md:60 | The walkthrough still quotes `✓ Codex Gate B satisfied (...)`, but this change replaces that output with `✓ Codex Gate B hook checks passed (...)` and deliberately removes the gate verdict | Readers following the walkthrough are shown an output the shipped hook no longer emits and the obsolete implication that its checks establish gate satisfaction | Update the quoted output and adjacent unsatisfied wording to the hook-state terminology, pointing to §5 for cycle closure +MINOR | high | docs/coding-workflow.md:96-99 | The methodology still says every pass reviews a revised artifact and only serious-tier findings force another iteration, while the installed ordering permits unrevised passes and requires a further pass after an assigned-fix-set change even when the accepted finding is Minor or Nit | The public explanation now contradicts the loop behavior this release installs and misstates when additional passes are owed | Replace these restatements with a reference to the closure ordering or describe unchanged retries and scope-driven pass costs accurately +MINOR | high | plugins/dev-workflow/CHANGELOG.md:33-34 | The release note says Gate A compares the artifact with text the final pass reviewed, whereas CLAUDE.md:546-555 compares it with text sent in the review request and explicitly disclaims evidence of reviewer consumption | The release note overstates what the new condition establishes, contrary to AGENTS.md's rule to name the actual gate comparison | Say the artifact equals the text sent in the final review request, without claiming what the reviewer consumed +END OF FINDINGS (3 total) diff --git a/.context/codex-reviews/gate-b-quality-t57gp3hwu1-pass-2.md b/.context/codex-reviews/gate-b-quality-t57gp3hwu1-pass-2.md new file mode 100644 index 0000000..ed1c04a --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-t57gp3hwu1-pass-2.md @@ -0,0 +1,5 @@ +MAJOR | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:3914 | close-profile-set-fixset applies the change-even-when-undone rule to the cited set throughout the window to the closing act, whereas installed CLAUDE.md:244-248 applies that rule to profiles and the assigned fix set and separately checks headers during the pass plus the current cited set at closure | A cited-set edit made and reverted after the final pass, with no profile or assigned-fix-set change, incorrectly fails this check and demands another pass; the named verification therefore tests a stricter policy than the installed ordering | Separate the predicates: observe every profile/fix-set change, reject governing-header changes during the pass, and compare the final pass's cited set with the current set at closure +MINOR | high | docs/getting-started.md:60 | The walkthrough still quotes `Codex Gate B satisfied` and describes content changes as returning the gate to unsatisfied, although this release changes the output to `Codex Gate B hook checks passed` and deliberately removes the gate verdict | Readers are shown output the shipped hook no longer emits and an obsolete interpretation of its advisory checks | Update the quoted output and adjacent wording to the hook-state terminology and point to CLAUDE.md section 5 for cycle closure +MINOR | high | docs/coding-workflow.md:96 | The methodology still says every pass reviews a revised artifact and only serious-tier findings force another iteration, while the installed ordering permits unrevised passes and charges another pass for an assigned-fix-set change even when the accepted finding is Minor or Nit | The public explanation contradicts this release's loop behavior and misstates when additional passes are owed | Reference the closure ordering or describe unrevised passes and scope-driven pass costs accurately +NIT | high | CLAUDE.md:156; plugins/dev-workflow/commands/workflow-init.md:363 | The installed Named residual says the standing sentences above require correcting contradictory reminders, but that reference came from target-text section F and has no corresponding list above it in either installed prompt copy | The downstream reader cannot resolve the cited authority; matching the approved target text preserves its dangling reference | Remove the trailing reference or replace it with a reference that exists in each installed prompt +END OF FINDINGS (4 total) diff --git a/.context/codex-reviews/gate-b-quality-t57gp3hwu1-pass-3.md b/.context/codex-reviews/gate-b-quality-t57gp3hwu1-pass-3.md new file mode 100644 index 0000000..4578fa4 --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-t57gp3hwu1-pass-3.md @@ -0,0 +1,7 @@ +MINOR | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:3916 | close-cited-set observes only the governing headers at the final pass's start and end, although its stated condition rejects a header changed during the pass | A header changed and restored during the pass has equal endpoints and passes this observation despite violating the condition; the named verification does not establish the condition it claims to check | Account for header changes throughout the pass, including reverted changes, while retaining the separate current-set comparison at the closing act +MINOR | high | docs/getting-started.md:60 | The walkthrough still quotes Codex Gate B satisfied and describes content changes as returning the gate to unsatisfied, although this release changes the output to Codex Gate B hook checks passed and removes the gate verdict | Readers see output the shipped hook no longer emits and an obsolete interpretation of its advisory checks | Update the quoted output and adjacent wording to the hook-state terminology and reference CLAUDE.md section 5 for cycle closure +MINOR | high | docs/coding-workflow.md:96 | The methodology still requires a revised artifact each pass and says only serious-tier findings force another iteration, whereas the installed ordering permits unrevised passes and charges a further pass for an assigned-fix-set change even when the accepted finding is Minor or Nit | The public explanation contradicts this release's loop behavior and misstates when additional passes are owed | Reference the closure ordering or describe unrevised passes and scope-driven pass costs accurately +NIT | high | CLAUDE.md:156; plugins/dev-workflow/commands/workflow-init.md:363 | The Named residual says the standing sentences above require correcting contradictory reminders, but that reference belongs to target-text section F and has no corresponding list above it in either installed prompt | Downstream readers cannot resolve the cited authority even though the installation matches the approved target | Remove the trailing reference or replace it with a reference available in both installed prompts +NIT | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:3887-3888 | Newly split rows 33 and 33b omit the Oracle column, while Task 13 step 3 requires a pass or fail mark per row and the summary says all 45 rows pass | The two new cases lack the per-row result recorded for every other case, leaving the verification record incomplete | Evaluate both rows and append their Oracle results +NIT | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:4099 | The new P1 fragment-evidence line ends with trailing whitespace | git diff --check for the reviewed range exits 2 | Remove the trailing space +END OF FINDINGS (6 total) diff --git a/.context/codex-reviews/gate-b-quality-t57gp3hwu1-pass-4.md b/.context/codex-reviews/gate-b-quality-t57gp3hwu1-pass-4.md new file mode 100644 index 0000000..d630984 --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-t57gp3hwu1-pass-4.md @@ -0,0 +1,8 @@ +MAJOR | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:3916 | The close-cited-set command observes reachable ancestry, not every committed header change during the pass; a committed A-to-B change followed by restoring the candidate HEAD produces an empty candidate..act-head range, and successive amendments can also hide B. Reproduced with a committed intermediate header, a clean worktree and HEAD equal to the candidate | The check can certify a pass as final despite a governing-header change during it; the stated limitation covers uncommitted edits but omits these committed changes | Observe header transitions during the pass independently of final ancestry, or explicitly bound this command to retained history and require a separate account for rewritten/discarded commits; exercise that case in the verification record +MAJOR | high | CLAUDE.md:1330; plugins/dev-workflow/commands/workflow-init.md:1519 | The replacement closing operation drops the previous explicit git reset --soft and now merely says reset to the parent and commit once | The ordinary git reset uses mixed mode: it unstages the reviewed change and makes newly added files untracked, so the following commit fails or an attempted commit -a omits those files; reproduced with two WIP commits | Restore the explicit --soft operation in both prompt copies and the matching target text so the closing sequence preserves the reviewed index +MAJOR | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:3915 | close-profile-fixset claims to observe every set change using the story profile log and membership answers, but the installed set definition also includes the scope assigned by every approved governing story or plan | A human-approved scope change, including one later undone, can change that union without changing a profile or answering a finding; this observation then misses the further pass owed by the absorb rule | Include changes to all governing scope inputs throughout the window in the named observation, alongside profile changes and membership answers, and record a scope-change counterexample +MINOR | high | docs/getting-started.md:60 | The walkthrough still quotes the removed Gate B satisfied message and describes the hook state as satisfied/unsatisfied although this diff deliberately replaces that verdict with hook checks passed | Readers following the current release encounter different output and an obsolete explanation of what its success message means | Use the new message and point to the closure ordering for whether the cycle may close +MINOR | high | docs/coding-workflow.md:97 | The methodology still requires a revised artifact every pass and says only serious findings force another iteration, while the new rules permit unrevised passes and charge a pass for accepting even a Minor into the fix set | The public workflow describes behavior that conflicts with the shipped loop rules | Replace these assertions with a reference to the closure ordering and distinguish repair severity from pass costs caused by scope changes +NIT | high | CLAUDE.md:156; plugins/dev-workflow/commands/workflow-init.md:363 | The transplanted Named residual points to the standing sentences above, but neither destination contains the target-text standing-sentence list above this paragraph | The justification has an unresolved reference in both the repository policy and every scaffolded copy | Remove the directional reference or cite the applicable prompt-consistency rule available in the destination +NIT | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:4099 | The new P1 fragment-sweep record has trailing whitespace | git diff --check d26de4b..8e620db exits 2 | Remove the trailing space +END OF FINDINGS (7 total) diff --git a/.context/codex-reviews/gate-b-quality-t57gp3hwu1-pass-5.md b/.context/codex-reviews/gate-b-quality-t57gp3hwu1-pass-5.md new file mode 100644 index 0000000..9550c9b --- /dev/null +++ b/.context/codex-reviews/gate-b-quality-t57gp3hwu1-pass-5.md @@ -0,0 +1,9 @@ +MINOR | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:3948; docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:3953 | observe-header-changes tests the status of grep rather than git log or git diff, so failed Git reads fall through to no change observed; reproduced under sh and dash with a failing configured textconv driver and an actual committed header change | The named observation reports a negative result where its source was unreadable, contradicting its three-result contract | Capture and check each Git command before inspecting its output, and report source unreadable on failure; include the failing-read case in the demonstrations +MINOR | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:3945; docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:3952 | The start boundary uses whole-second reflog timestamps: it chooses the newest entry at or before t0 and excludes every move stamped t0, including moves made after the request within that same second; a committed A-to-B-to-A sequence in that second returns no change observed under both sh and dash | The stated observation of every branch ref move in the pass window has an additional committed-change blind spot beyond the documented uncommitted and outside-branch cases | Capture a ref/reflog cursor when the window starts, or report timestamp-boundary ambiguity explicitly and document this limitation; demonstrate same-second changes +MINOR | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:3948 | The h0..h1 first-parent scan is a reachability range rather than a time window: resetting main during the pass to an existing side-branch tip with the identical A header reports change observed when that tip has an older A-to-B-to-A history predating the pass; reproduced under sh and dash | The observation attributes a historical change to the current pass even though the branch header never changed in that interval, and its result is documented as making the pass non-final | Distinguish imported historical ancestry from header transitions during the window, or bound the reported claim accordingly; add the reset-to-identical-content case +MINOR | high | docs/getting-started.md:60 | The walkthrough still quotes Codex Gate B satisfied and describes content changes as returning the gate to unsatisfied, although this release replaces that verdict with Codex Gate B hook checks passed | Readers encounter output the hook no longer emits and an obsolete explanation of its advisory state | Update the quoted output and adjacent terminology, and reference the closure ordering for whether the cycle may close +MINOR | high | docs/coding-workflow.md:97 | The methodology still requires a revised artifact every pass and says only serious-tier findings force another iteration, whereas the new rules permit unrevised passes and charge a further pass for accepting even a Minor into the assigned fix set | The public workflow contradicts the shipped rules about when to revise an artifact and when another pass is owed | Reference the closure ordering and distinguish repair severity from pass costs caused by scope changes +NIT | high | CLAUDE.md:156; plugins/dev-workflow/commands/workflow-init.md:363 | The Named residual cites the standing sentences above, but the target-text standing-sentence list is not present above this paragraph in either installed prompt | The justification contains an unresolved reference in the repository policy and scaffolded copies | Remove the directional reference or cite a consistency rule available in both destinations +NIT | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:4012 | The parity record combines target section F blocks 9a and 9 into one row despite Task 14 requiring one row per fenced destination block even when adjacent; 35 prompt blocks have only 34 recorded regions | The verification record does not provide the required block-by-block accounting, although both installed spans match | Record separate bounded comparisons and result rows for 9a and 9 and update the region count +NIT | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:4151 | The added P1 fragment-sweep record ends with trailing whitespace | git diff --check d26de4b..1bea70e exits 2 | Remove the trailing space +END OF FINDINGS (8 total) diff --git a/.context/codex-reviews/gate-b-spec-t57gp3hwu1-pass-1.md b/.context/codex-reviews/gate-b-spec-t57gp3hwu1-pass-1.md new file mode 100644 index 0000000..c1eee9e --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-t57gp3hwu1-pass-1.md @@ -0,0 +1,8 @@ +MAJOR | high | CLAUDE.md:421; plugins/dev-workflow/commands/workflow-init.md:628 | The continue-branch gloss says an already-declined finding raises no trigger, although target §B's b11 exempts only its membership trigger and expressly leaves a new structural/contract question reachable; Task 12 records equivalence by assuming a respective pairing the sentence never states | A declined finding that raises a new question can be treated as clean or continued without the question stop, violating story criteria 1 and 3 | Remove the gloss and defer to the absorb paragraph, or explicitly scope each exemption to its own trigger; align the approved target and both copies and rerun the bidirectional check. +MAJOR | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:3876 | Task 13 requires every row to specify all route predicates and yield exactly one next state, but row 21 leaves cleanliness, eligibility and closure conditions unspecified and lists four destinations; rows 7, 15-17 and 22 combine alternative answers, and row 27 asserts CONT without excluding a source block or suspension | The claimed 32 passing rows do not supply the required determinate risk-path verification and can endorse continuation where the ordering requires a stop | Split each alternative answer/state into its own row, state all predicates that select its route, rerun the oracle, and refresh the row count and evidence entry. +MAJOR | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:3903 | close-profile-set-fixset observes only header and fix-set values at pass start and at the closing act, while installed §A and §B require another pass for a relevant change even when it is later undone | Equal endpoints can pass this named closure check after an intervening change that still owes a pass, so the check does not establish its stated condition | Check the intervening changes over each source rule's required window as well as current values, and explicitly exercise the change-then-restore case in the verification record. +MINOR | high | docs/getting-started.md:55; docs/getting-started.md:60 | The walkthrough still describes fingerprint invalidation as making the gate unsatisfied and quotes the removed Gate B satisfied message as the cue to amend, whereas §F items 12, 15 and 16 deliberately replace gate verdicts with hook-state observations | New users are taught a message the release no longer emits and can mistake a successful hook check for permission to close | Update the quoted message and explain that the hook reports its checks; point to §5's closure ordering for permission and the closing operation. +MINOR | high | docs/coding-workflow.md:96 | The methodology still requires a revised artifact on every pass and says only serious findings force another iteration, although §F item 6 permits unrevised passes and §B makes accepting even a Minor into the assigned fix set cost another pass | The public workflow description contradicts the installed rules and encourages unnecessary edits or omission of a required scope-change pass | Update this paragraph to allow an unrevised artifact where no repair is owed and cite the closure ordering and assigned-fix-set rule for whether another pass is required. +MINOR | high | plugins/dev-workflow/CHANGELOG.md:33 | The new release entry describes Gate A's condition as equality with text the final pass reviewed; target §A2 defines equality with text supplied in the final review request and explicitly disclaims evidence of what the reviewer consumed | The release notes overstate what the new comparison establishes, contrary to AGENTS.md's rule against overstating gate mechanisms | Describe equality with the final request's artifact text, or link to the authoritative condition without restating it. +NIT | high | CLAUDE.md:156; plugins/dev-workflow/commands/workflow-init.md:363 | The installed Named residual cites the standing sentences above, but that reference belongs to target §F and no such standing-sentence list exists above this location in either installed prompt | The scaffold ships a dangling reference to specification-local material unavailable to its downstream reader | Remove the specification-local clause or replace it with a reference that exists in the installed policy, keeping the target and both copies aligned. +END OF FINDINGS (7 total) diff --git a/.context/codex-reviews/gate-b-spec-t57gp3hwu1-pass-2.md b/.context/codex-reviews/gate-b-spec-t57gp3hwu1-pass-2.md new file mode 100644 index 0000000..a2f0365 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-t57gp3hwu1-pass-2.md @@ -0,0 +1,7 @@ +MAJOR | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:3887 | Task 13 row 33 still derives a further pass merely from a failed closing act moving a condition input; it does not identify that input or establish a pass-cost predicate, although target §A2 permits Gate-A artifact edits restored byte for byte and §A1 then permits retrying the act when every condition holds | The row marks CONT as verified for inputs that also admit retrying the act without another pass, so the claimed determinate risk-path verification remains incomplete | State a concrete pass-cost change in the starting predicates, such as an assigned-fix-set change, and distinguish it from restoration of Gate A's current-equality condition; rerun the oracle and refresh the evidence. +MAJOR | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:3913 | Task 13's named closure checks omit evidence-entry revalidation and the obligation to re-review when it changes, which target §F items 9a/9 and installed CLAUDE.md:1190 require; checking profile/set changes and carrying owed records does not compare the revalidated entry with the entry supplied to the final pass | All ten listed checks can pass with unchanged profiles and a newly corrected evidence entry that no final pass covered, so conditions-held-at-act does not establish every required closure condition | Add a named check for each owed entry's final revalidation and its coverage by the final review request, including the changed-entry route that owes a new pass. +MINOR | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:3861 | Rows 7-9 and 28 mark K=yes while eligibility fails below the floor with nonzero findings, but K is defined as every closure condition holding and close-floor explicitly includes the unmet floor among those conditions | These rows use inconsistent predicate values while being reported as passing verification of reachable states | Mark the floor condition unmet in these rows, or explicitly define K as the remaining conditions excluding eligibility/floor and apply that definition consistently. +MINOR | high | docs/getting-started.md:55; docs/getting-started.md:60 | The walkthrough still describes fingerprint invalidation as making the gate unsatisfied and quotes the removed Gate B satisfied message as the cue to amend, whereas §F items 12, 15 and 16 replace gate verdicts with hook-state observations | Users are taught a message this release no longer emits and can mistake a successful hook check for permission to close | Update the quoted message and point to §5's closure ordering for permission and the closing operation. +MINOR | high | docs/coding-workflow.md:96 | The methodology still requires a revised artifact on every pass and says only serious findings force another iteration, although §F item 6 permits unrevised passes and §B makes accepting even a Minor into the assigned fix set cost another pass | The public workflow description contradicts the installed rules and encourages unnecessary edits or omission of a required scope-change pass | Allow an unrevised artifact where no repair is owed and cite the closure ordering and assigned-fix-set rule for whether another pass is required. +NIT | high | CLAUDE.md:156; plugins/dev-workflow/commands/workflow-init.md:363 | The installed Named residual cites the standing sentences above, but that reference belongs to target §F and no such standing-sentence list exists above this location in either installed prompt | The scaffold carries a reference to specification-local material unavailable to its downstream reader | Remove the specification-local clause or replace it with a reference that exists in the installed policy, keeping the target and both copies aligned. +END OF FINDINGS (6 total) diff --git a/.context/codex-reviews/gate-b-spec-t57gp3hwu1-pass-3.md b/.context/codex-reviews/gate-b-spec-t57gp3hwu1-pass-3.md new file mode 100644 index 0000000..dd4fc60 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-t57gp3hwu1-pass-3.md @@ -0,0 +1,8 @@ +MAJOR | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:3916 | close-cited-set observes only the governing headers at the final pass's start and end, although its stated condition and installed CLAUDE.md:138 disqualify a pass whenever a header changed during it; changing a header and restoring it before the end passes both comparisons | The named check can certify finality after the condition it is meant to verify failed, leaving Task 13's required per-condition verification incomplete | Check for every governing-header change during the pass as well as comparing the final pass's cited set with the current set at the act; explicitly reject a during-pass edit restored before the end. +MINOR | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:3880 | Rows 26, 28, 30 and 31 simultaneously give SB=yes and K=yes, but K now means every closure condition except eligibility holds, and close-no-source-block expressly includes absence of a source block among those conditions | These starting states remain internally inconsistent despite being reported as verified reachable transitions; it is unclear whether K describes the starting state or the state after repair | Separate pre-repair and post-repair values, or explicitly exclude the independently represented source-block condition from K and apply that definition consistently. +MINOR | high | docs/getting-started.md:55; docs/getting-started.md:60 | The walkthrough still describes fingerprint invalidation as making the gate unsatisfied and quotes the removed Gate B satisfied message as the cue to amend, whereas target §F items 12, 15 and 16 replace gate verdicts with hook-state observations | Users are taught a message this release no longer emits and can mistake a successful hook check for permission to close | Update the quoted message and point to §5's closure ordering for permission and the closing operation. +MINOR | high | docs/coding-workflow.md:96 | The methodology still requires a revised artifact on every pass and says only serious findings force another iteration, although target §F item 6 permits unrevised passes and §B makes accepting even a Minor into the assigned fix set cost another pass | The public workflow description contradicts the installed rules and encourages unnecessary edits or omission of a required scope-change pass | Allow an unrevised artifact where no repair is owed and cite the closure ordering and assigned-fix-set rule for whether another pass is required. +NIT | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:3887 | Newly split rows 33 and 33b omit their Oracle cells, although Task 13 step 3 requires a pass-or-fail mark for each row | The per-row record does not contain the two oracle results claimed by the all-45-rows-pass summary and evidence entry | Apply the oracle to both rows and record each result in its Oracle cell. +NIT | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:3960 | Task 14 combines target §F blocks 9a and 9 into one parity row, despite its explicit one-row-per-fenced-block requirement even for adjacent blocks; there are 35 prompt destination blocks but only 34 recorded regions | The parity record does not follow the requested block-by-block accounting, although both installed spans currently match | Give 9a and 9 separate bounded comparisons and result rows, then refresh the recorded region count. +NIT | high | CLAUDE.md:156; plugins/dev-workflow/commands/workflow-init.md:363 | The installed Named residual cites the standing sentences above, but that reference belongs to target §F and no such standing-sentence list exists above this location in either installed prompt | The scaffold carries a reference to specification-local material unavailable to its downstream reader | Remove the specification-local clause or replace it with a reference that exists in the installed policy, keeping the target and both copies aligned. +END OF FINDINGS (7 total) diff --git a/.context/codex-reviews/gate-b-spec-t57gp3hwu1-pass-4.md b/.context/codex-reviews/gate-b-spec-t57gp3hwu1-pass-4.md new file mode 100644 index 0000000..fb944c6 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-t57gp3hwu1-pass-4.md @@ -0,0 +1,7 @@ +MAJOR | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:3916 | close-cited-set requires git log -G to report every committed governing-header change, but the command does not enable merge diffs; a disposable-repository A -> B -> A sequence introduced and restored solely by merge commits returned no matches despite both header changes | The named check can certify finality after committed during-pass changes, so Task 13's required condition verification still has a false negative beyond its stated uncommitted-edit limitation | Inspect merge changes explicitly as well as ordinary commits, state the history coverage accurately, and verify the committed merge-and-restore case. +MAJOR | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:3916 | The during-pass check scans candidate..act-head rather than ending at pass completion, although target section A distinguishes headers changed during a pass from the current cited-set comparison at the act | A Story header edit made and restored after the pass, with profiles and the assigned fix set unchanged, fails this check even though this cited-set condition holds; the verification imposes a stricter rule than the installed ordering | Bound the change observation to the actual pass interval and retain the separate current-set comparison at the act; verify that an after-pass restoration does not itself fail this condition. +MINOR | high | docs/getting-started.md:55; docs/getting-started.md:60 | The walkthrough still describes fingerprint invalidation as making the gate unsatisfied and quotes the removed Gate B satisfied message as the cue to amend, whereas target section F items 12, 15 and 16 replace gate verdicts with hook-state observations | Users are taught a message this release no longer emits and can mistake a successful hook check for permission to close | Update the quoted message and point to section 5's closure ordering for permission and the closing operation. +MINOR | high | docs/coding-workflow.md:96 | The methodology still requires a revised artifact on every pass and says only serious findings force another iteration, although target section F item 6 permits unrevised passes and section B makes accepting even a Minor into the assigned fix set cost another pass | The public workflow description contradicts the installed rules and encourages unnecessary edits or omission of a required scope-change pass | Allow an unrevised artifact where no repair is owed and cite the closure ordering and assigned-fix-set rule for whether another pass is required. +NIT | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:3960 | Task 14 still combines target section F blocks 9a and 9 into one parity row despite its explicit one-row-per-fenced-block requirement even for adjacent blocks; 35 prompt destination blocks have only 34 recorded regions | The parity record does not follow the requested block-by-block accounting, although both installed spans currently match | Give 9a and 9 separate bounded comparisons and result rows, then refresh the recorded region count. +NIT | high | CLAUDE.md:156; plugins/dev-workflow/commands/workflow-init.md:363 | The installed Named residual cites the standing sentences above, but that reference belongs to target section F and no such standing-sentence list exists above this location in either installed prompt | The scaffold carries a reference to specification-local material unavailable to its downstream reader | Remove the specification-local clause or replace it with a reference that exists in the installed policy, keeping the target and both copies aligned. +END OF FINDINGS (6 total) diff --git a/.context/codex-reviews/gate-b-spec-t57gp3hwu1-pass-5.md b/.context/codex-reviews/gate-b-spec-t57gp3hwu1-pass-5.md new file mode 100644 index 0000000..bd8f313 --- /dev/null +++ b/.context/codex-reviews/gate-b-spec-t57gp3hwu1-pass-5.md @@ -0,0 +1,8 @@ +MINOR | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:3948 | observe-header-changes checks the grep status but discards failures from git log and git diff; unlike the specified three-way result, unreadable history can produce no change observed | Reproduced under sh and dash with a missing intermediate commit object: both history sources error and the procedure still reports no change observed, losing the source-unreadable disposition required by close-header-during-pass | Check each Git producer exit status separately before inspecting its output and report source unreadable on failure; demonstrate that case +MINOR | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:3945 | The start head is reconstructed from the newest reflog entry with timestamp <= t0 and the move scan excludes every entry timestamped t0; second-resolution timestamps cannot identify the actual request-start boundary | A committed A-to-B-to-A sequence after request construction but within its starting second is excluded from both reads and reports no change observed under sh and dash; this is an additional committed-change blind spot beyond the ones disclosed | Capture the starting ref and reflog position when constructing the request, or explicitly report boundary ambiguity and document and demonstrate this limitation +MINOR | high | docs/getting-started.md:60 | The public walkthrough still quotes the removed Gate B satisfied message and directs the real commit from that hook state; steps 3-5 also omit the newly required Gate-A closing acts | Readers following the walkthrough receive the old closure procedure instead of the installed ordering and can advance before a cycle has closed | Update the message spelling and point each phase transition and closing operation to the closure ordering, including the Gate-A closing acts +MINOR | high | docs/coding-workflow.md:96 | The methodology still requires the revised artifact on every pass and says only serious findings force another iteration, although target sections F6 and E now allow an unrevised pass and a Minor acceptance can cost a pass through the fix-set change | The public explanation contradicts the implemented cadence and can prompt unnecessary edits or omit a required pass after a scope answer | Describe conditional revision and defer next-pass and closure decisions to the ordering rather than to severity alone +MINOR | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:3986 | Task 14 requires one parity row per destination fenced block, but its output combines F9a and F9 into one adjacent region and records 34 regions instead of 35 blocks | The record does not implement the approved per-block reporting procedure, although independent inspection finds both blocks identical between copies | Record F9a and F9 separately and update the dependent region count +NIT | high | CLAUDE.md:156; plugins/dev-workflow/commands/workflow-init.md:363 | The installed Named residual cites the standing sentences above, but that enumeration exists in target-text section F and is absent from either installed prompt | Downstream readers cannot resolve the copied reference without this repository-specific design artifact | Remove the dangling directional reference or point to an applicable rule actually present in both installed copies, keeping the target text aligned +NIT | high | docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md:4151 | The newly added P1 fragment-sweep record ends with trailing whitespace | git diff --check for the requested range exits 2 | Remove the trailing space +END OF FINDINGS (7 total) diff --git a/CLAUDE.md b/CLAUDE.md index 0a43981..7696d1f 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -80,7 +80,7 @@ Strong success criteria let you loop independently. Weak criteria ("make it work Ground progress claims: before reporting a step as done, audit the claim against a tool result from this session ("tests green" needs a test run to point to). Report unverified work as unverified — this keeps status reports factual on long runs. -The work loop includes the review gates: **spec ready → Gate A (spec) → plan ready → Gate A (plan) → execute → tests green → Gate B → commit** (see §5). +The work loop includes the review gates: **spec ready → Gate A (spec) → Gate-A closing act → plan ready → Gate A (plan) → Gate-A closing act → execute → tests green → Gate B → Gate-B closing act** (see §5, which states when each act may be performed and what it is). ## 5. Cross-Model Review (Codex) — TWO MANDATORY GATES @@ -89,8 +89,9 @@ yours — a non-blocking hook (shipped by the `dev-workflow` plugin) reminds you each. Opt out per-workspace with `.context/codex-gate.off` (delete to re-enable); the gates still apply. -**Both gates are a LOOP with a HARD FLOOR: a minimum number of passes per run -(Blocker/Major only), derived from the cited story's profile.** +**Both gates are a LOOP with a HARD FLOOR: a minimum number of passes per run (a clean final pass +being what the floor is spent on, and cleanliness taking more than Blocker/Major), derived from the +cited story's profile.** The derivation is max(risk, security): a value of 0 gives a floor of 1; every resolvable profile above that, and an artifact citing no story, gives 3. Two levels, not three — `high` takes its rigor from lens sets and evidence mode, not from extra @@ -145,19 +146,15 @@ threshold that controls nothing.** The hook still counts passes, and it still ca findings or tell the spec run from the plan run (it resets at `writing-plans`), so Gate A — the spec run especially — is instruction-backed: a satisfied count is not a clean review, and a below-threshold reminder is noted in the pass report and disregarded -where the cycle's own closure rules are satisfied. This replaces the pass-count number -and nothing else. Every other rule stated here about how a cycle closes stands as -written, and none of them is restated — a summary is where their conditions would get -dropped. Nothing here writes the floor knob: it stays the user's, never written, never -removed, never read for this derivation. Open a TodoWrite "Codex pass N" per pass; fix Blocker/Major after each. Your -final pass must be clean — if the pass at the floor still finds Blocker/Major, keep going until -clean or clearly stuck → then STOP and surface to the user. The only early exit -below the floor is a pass with **zero** findings; don't manufacture findings to pad. Codex is -advisory — validate before applying; dismissed finding → one-line why. +where the cycle's own closure rules are satisfied. Every other rule stated **in this paragraph** about how a cycle closes stands as written, and +none of them is restated — a summary is where their conditions would get dropped. Nothing here writes the floor knob: it stays the user's, never written, never +removed, never read for this derivation. Open a TodoWrite "Codex pass N" per pass; resolve Blocker/Major after each as +Mechanics · Severity requires. What a clean final pass and the zero-finding early exit mean for closing is stated once in the +closure ordering. Codex is advisory — validate before applying; dismissed finding → one-line why. **Named residual:** the hook's messages state its own threshold as an obligation, so at a -floor of 1 they report a shortfall the cycle does not owe. Hook text is out of scope here -by decision; what makes that tolerable is the precedence rule above plus the hook exiting +floor of 1 they report a shortfall the cycle does not owe. **That particular overstatement is out of scope here by decision, and it is not a blanket exemption for hook text** — a reminder this change's own rules falsify is corrected in the same change, as the standing sentences above require. +What makes that tolerable is the precedence rule above plus the hook exiting 0 on every branch, not the reminder being harmless. **The gate-off surface — routes known today, not a complete list**, since an enumeration @@ -172,9 +169,16 @@ the cited profiles. **When these rules bind.** From the commit that ships them, and a cycle already running finishes under the rules it started with. Where a cycle's starting rules cannot be -established it takes the stricter reading of every part this change touches — at minimum -floor 3, severity classified without the demotion, the provenance-line duty owed, the curve -duty owed, and the nonce duties at their strictest — the cycle is treated as post-rule, so it +established it takes the stricter reading of every part this change touches — at minimum floor 3, severity classified without the demotion, the provenance-line duty owed, the curve duty owed, +the nonce duties at their strictest, **every suspension binding, since +starting rules that cannot be established cannot be read as having waived an open hold**, **the +repeated-dismissal cleanliness exclusion unavailable, a cycle that cannot establish its starting +rules being unable to establish that they contained it**, **the parked state binding after a closing act that cannot +be repaired, a cycle whose starting rules cannot be established being the last one that should be +left with no terminal transition**, and **every closure condition and +pass-cost rule this change ships owed rather than waived — Gate A's content condition and its +commit-carry duty, and the further pass an assigned-fix-set change costs — since a rule that cannot +be established as absent is cheaper to owe than to skip** — the cycle is treated as post-rule, so it owes a nonce, owes its provenance line and its curve or skip record, and uses that nonce in every cycle record it does write — which changes what a record is named, never whether one is owed, so the working record stays optional and a skipped cycle still writes no findings slots. Where it @@ -212,32 +216,450 @@ the predicate itself. In any such state nothing here resolves which rule governs have a human complete or revert the adoption, before running a gate under it.** What prompt text can do about downstream adoption is limited, and that limit is what this paragraph states. +**How a cycle ends — one ordering, stated here and referenced everywhere else.** A pass is read in +a fixed order, because every rule bearing on one decision — may this cycle close — otherwise +qualifies the others and the ranking survives only in a reader's head. + +**What this paragraph owns.** It owns **how a pass is evaluated and what follows from that**: the +evaluation order, what each predicate is read from, the eligibility test, the hold a surfaced +finding places and what discharges it, the composition of several suspensions, the pairs that +cannot co-occur, and — as a stated exception, because precedence is evaluation order — the +clearly-stuck precedence clause quoted into it below. **No number is put on that list.** Whether +the read model counts as part of the evaluation order or as a thing beside it is a question of +wording, not of authority. It also owns the **classification** of the standing duties below, and +**those duties bind both gates alike**: the derived floor and the Blocker/Major-resolve duty keep +their definitions at their own sources and are read here for every cycle whatever its gate, while +the hold and no-clean-credit are defined here, being properties of the evaluation itself. + +**What belongs to a gate is what the two gates do differently: the content condition its closure +requires, where it has one, and the closing act.** **Gate A has such a condition and Gate B has +none**; each paragraph below states its own and is read from there, and **neither gate's paragraph restates the conditions this one gives every cycle**, so +neither is a complete inventory on its own. **Everything else this paragraph names it cites**: the +scope triggers, the assigned fix set, every severity rule and every closure precondition with a +source of its own keep their one definition in the paragraph that owns them, and a reader who finds +one of *those* defined here has found a defect. + +**The conditions every cycle has, whatever its gate.** The duties classified below; and the +**profile**, the **cited set** and the **assigned fix set**, each gating as its own source says and +cited here without restatement. **The profile and the assigned fix set answer a change and not a +differing value**, which is why one of them changed and then undone still costs a pass. **The cited +set answers as its own source says**, and that source says two things: a governing header changed +**during** a pass makes that pass not final, and the final clean pass runs against the **current** +set. The contrast that matters is against a gate's own content condition, which may be a comparison +of current values and says so where it is stated. + +**What a pass is read from.** Every finding-derived predicate reads the validated findings file +**or files** of the logical pass as **the concatenation of their finding lines after each file has +been validated separately** — a `full` Gate-B pass has two, one branch alone is already an +incomplete pass, each file's terminator is not a finding line, and a branch whose body is +`NO FINDINGS` contributes an empty sequence rather than a line. **Which severity field each +predicate reads is settled in Mechanics · Severity**, which is where that split lives and is not +repeated here. Beyond the findings, closure reads the **derived floor**, the **resolve duty's +standing over this cycle**, any **hold still standing**, and **every closure condition this cycle +has** — those above and this cycle's gate's. The scope triggers read the **current assigned fix +set** as the absorb paragraph defines it, and the answers already given; the clearly-stuck reading +adds its own coverage judgement. **A line in one branch file and a line in the other are distinct +findings for holds and answers**, so a `full` pass asks twice rather than risk resuming over one it +never asked about. **Within one running cycle an answer binds to the finding or question as the +pass that raised it recorded them** — which is what an agent running the cycle can do with nothing +written down. Recognising the same finding or question **across a lost session** has no mandatory +identity, sameness or recovery rule in these rules; the optional `-dispositions.md` note is +advisory and authoritative for nothing. + +**The four branches are named, and a cross-reference anywhere in this section names the branch +rather than its position**, so reordering them breaks no reference. **They are read once, on the +pass, before any closing act is attempted** — so a pass that reached the act has already been +classified and is not classified again by what the act does. + +**The source-block branch, read first.** **Where any unmet closure condition's own source +prescribes stop-and-surface** — a profile present but unresolvable, governing headers that +disagree, a `Story:` header that cannot be read, an unobservable counterfactual, and **any other +source rule that prescribes it; the list is examples and not the set** — **the cycle stays open, that source decides what +must be repaired or answered, and no further pass runs while its block stands.** It is neither a +suspension nor a continue and needs no name and no procedure of its own: the source rule carries +both, and this ordering's part is to send the reader there rather than to run a pass over a cycle +another rule has stopped. **Once its source condition is repaired**, a block that stood **before any pass of this cycle was +read** — the ones its floor derivation and its cited set raise among them — leaves the cycle to run +its next pass, there being no pass to read again; a block that stood **on a pass already read** has +**that pass read again through the ordering**, no closing act +having been attempted on it — **and where that pass also carried a suspension, this branch +releases only its own block**: the composition rule still holds the cycle on every answer that +suspension asked for, a continue still leads to a pass run after the answer, and a stop still parks +the cycle until an explicit later continue — so the reread happens where no suspension of that +pass is outstanding, and otherwise the suspension's own route runs first — the read-once rule below is about a pass that reached the act, not +about one a block held before it, and without this a repair that moves no pass-cost value would +leave a clean eligible pass with no route to the act and none to a suspension. +**It is read first and it silences nothing.** Where the same pass also +carries a suspension, that suspension is surfaced with its reasons and its questions exactly as the +suspension branch requires and its answers are collected; what the block adds is that **no next +pass runs until its own source condition is repaired**, whatever those answers were. + +**Then the clean-completion branch, and eligibility is its own test.** §5 uses *clean* in two +senses and now says which is which: a **clean findings file** is the `NO FINDINGS` signal the +protocol defines, and a **clean pass** is the predicate here, read on the logical pass with every +required branch file combined, so one branch's clean file never establishes a clean pass. **A pass +is clean** when its findings carry no in-set Blocker or Major at effective severity and **no +scope-stop trigger** — the two the absorb paragraph defines, read there and not redefined here, +each already carrying the qualification **an answer given before that pass ran** puts on it. **A +pass's cleanliness is settled on what it found and on the answers standing when it ran**, and a +later answer never rewrites it: an answer discharges the holds it was asked for and leaves the pass +that raised them exactly as clean or unclean as it was, which is the same fact the duties paragraph +states of no-clean-credit. Reading a later answer back onto an earlier pass would let a cycle close +on a pass that was surfaced, answered and never re-run. A pass with **zero** findings is clean +whatever the floor, because a floor buys further looks at an artifact that keeps yielding findings, +and one yielding none has already given what those looks were for; don't manufacture findings to +pad. + +**A repetition of a finding this cycle has validly dismissed does not, on its own, make a pass +unclean.** Every part of this is required and the exclusion is narrow: the **dismissal was made in +an earlier pass of this cycle**, its stated reason is **still true of the artifact as it now +stands**, and the later finding **makes the same complaint and brings no new evidence** — no +observation the dismissal did not answer, and no change to the text the reason turned on. Where any +part fails — new evidence, relevant content changed, or genuine doubt that this finding is that +one — the finding is read afresh like any other, and **doubt never resolves in the exclusion's +favour**. Without it the standing duty *dismiss validly, then run another pass* cannot finish, +because a reviewer repeating its own refuted claim would decide whether that claim had been dealt +with. **It reaches only a dismissal this cycle made**, which is what an agent running the cycle +knows; nothing here ships a record, and recognising a dismissal across a lost session has no more +support than the paragraph above gives it. + +**It changes cleanliness and nothing else, which is what keeps it from being a waiver.** The +repetition is **still a finding**: it stands in its pass's findings file, and every loop-health +reading counts it exactly as it counts any other, so a loop spending passes on a point it keeps +refuting still shows up as one. It remains the clearly-stuck reading's re-raise condition. **No +earlier pass becomes clean in retrospect** — a pass's cleanliness is settled on what it found and +is never rewritten, which this exclusion leaves untouched: it decides the pass being read and no +other. And it is **not** a second dismissal; the resolve duty was discharged when the finding was +dismissed and there is nothing here to discharge again. + +**Eligibility is exactly this and nothing more: a clean pass at or above the derived floor, or a +zero-finding pass.** It is a property of the pass. **Closure is eligibility plus every closure +condition of this cycle holding plus this cycle's gate's closing act**, and those conditions are +properties of the *cycle*, not of the pass — so an eligible pass whose conditions are unmet closes +nothing, and is not thereby made unclean. Keeping the two apart is what lets the order decide the +case where they disagree, which the branches do. **The order inside closure is fixed: every +condition is established first, and only then is the closing act performed.** + +**An act that does not complete has not closed the cycle**, and it is not a branch: the branches +below decide what a *pass* is, and this cycle's pass already took the clean-completion branch and +reached closure. **A failed act returns to the closure step it failed in, not to the branches.** +**Nothing is assumed about what the attempt left behind** — a hook can modify and stage content +before failing, so the attempt itself can move what a condition is read from — so surface the +concrete command failure and **re-establish every closure condition against the repository as it +now stands.** Where they all still hold, perform the act again. Where the attempt or its repair +moved anything a condition is read from, **that condition has changed and its own rule decides what +it costs**, a further pass included, and the cycle is back in the ordering with that pass owed. +**Where the failure cannot be repaired at all** — a signing key nobody has, a permission nobody can +grant — **surface it and leave the cycle parked**: open, not running, spending no passes, restarted +by an explicit later continue, which is the state a stop answer already produces and is named here +rather than invented. A cycle that can neither close nor be parked is the outcome this sentence +exists to prevent. + +**Closure introduces no new kind of record, and it excuses none**: every other record this cycle +owes, a human-exception record among them, is owed and written exactly as before, and **a +human-exception record this cycle owes goes in the commit its closing act uses**, so the two never +land in different places. **No closure condition is read on the branch tip**, the conditions being +read where their sources say and not off the tip. **That is not a claim that the act publishes what +the pass read.** A commit is written from the effective index, so a staged edit the working tree +does not show lands in the closing commit; where the edit changes something a **source rule** +governs — a profile value, cited-set membership, the assigned fix set — **the change has happened +and that source's own rule applies**, so a further pass is owed and no condition is added here for +it. **Where it changes a review input no source rule governs** — a cited story's acceptance +criteria or settled decisions, say — **nothing here reaches it, and nothing in this section does**: +the artifact's own equality condition covers the artifact, and the inputs beside it have only the +rules their sources give them. + +A plateau or tells on **the pass that closes** go into **that pass's status report to the user** — +the carrier the three-line duty already names, and no second report form is introduced — and never +block it, because reporting "will not converge" on a converged loop is a false report; on an +eligible pass that does **not** close they go into the same place, a pass reporting what it read of +the loop whether or not it closes. **No other pass outcome makes a cycle eligible to close**, +because every other pass either leaves a required repair, a hold or a question outstanding, **or +has not reached the floor, or is itself unclean on its own findings** — and closing over any of +those is the failure this ordering exists to prevent. The floor is named separately because a +below-floor pass whose only findings are Minors leaves nothing outstanding and is still not +eligible; **pass-level uncleanliness is named separately** because an in-set Blocker or Major +repaired after the pass that raised it discharges the resolve duty without making that pass clean, +so a cycle can owe nothing and still hold no pass it may close on. The one termination that is not a pass outcome is the Gate-B triviality skip, which runs +no passes and is outside this ordering. + +**Then the suspension branch, which a pass reaches where the clean-completion branch did not take +it and a suspension applies to it.** **Clean completion outranks a suspension by taking the pass to +the closing act, not by eligibility alone**: a pass that took the branch above, met +every closure condition and had the closing act performed has ended the cycle, and a suspension has +nothing left to suspend. **A pass the clean-completion branch did not take reaches this branch +whatever its cleanliness, where a suspension applies to it** — a clean pass below the floor, and +equally an eligible pass whose unmet closure conditions kept that branch from taking it. Cleanliness is what this branch stops +asking about; whether a suspension applies is still what puts a pass here, and where none does the +continue branch has it. That is what makes "clean completion outranks the two-tell stop" +executable rather than asserted, and cleanliness alone never decides it. **What D2 and D3 forbid is +reporting "will not converge" on a loop that converged, and a loop still owing a repair, an answer +or a closure condition has not converged** — so a mandatory two-tell stop and the clearly-stuck +reading stay reachable exactly where the loop is still running, which is the only place their +question means anything. Three suspensions, by the names their +paragraphs use and read by those paragraphs: the **scope stop**, raised by either trigger above — a +**membership stop** by the first, a **question stop** by the second; the **clearly-stuck exit**; +and the **two-tell stop**. A suspension waives nothing. Any non-empty set of them can apply to one +pass: **one surface, every reason reported, every question asked**, because a reason left out is a +decision made by omission. A finding the clearly-stuck reading surfaces that also carries either +trigger takes the scope stop's answers at that same surface, so it is not asked twice; the two-tell +stop surfaces tells and not a finding. + +**Otherwise the continue branch, which no source block reaches: where none stands and neither of +the two branches above took the pass, it continues** — the loop +runs another pass on the **current** artifact, revised where the severity and scope rules require a +repair and unrevised where they do not. **An eligible pass with an unmet closure condition lands +here**, and like every other non-closing pass **only where no suspension applies to it**: clean +completion did not take it, so the loop continues on whatever the unmet condition requires — most +often a repair still owed from an earlier pass. A below-floor clean pass lands here on the same +terms; where a suspension does apply, the suspension branch has already taken it, because only +the clean-completion branch outranks a suspension. So does a pass whose only findings are Minors and +Nits **and which carries no scope-stop trigger** — those are collected and never iterated and may +leave nothing to revise, while a Minor or Nit **carrying a scope-stop trigger as the absorb +passage defines one** carries it like any other finding, is not clean, and has already been taken +by the suspension branch — **severity does not raise a trigger and does not suppress one**, and +which findings raise one is that passage's entire, an already-declined finding raising no +membership trigger and an already-answered question no question trigger. It is a branch and +not an inference, because "does not close" read alone says nothing about whether to run again. + +**The four standing duties, classified.** The **derived floor** is a **precondition on closure**: +it gates closing, discharged by the count of valid logical passes reaching it with the last of them +clean, or by the zero-finding exit. The **Blocker/Major-resolve duty** is a **precondition on +closure**: Mechanics · Severity, scoped to the assigned fix set, states what it demands and what +discharges it, at that source and not here. **It is discharged per finding and tracked across the +cycle, never inferred from a later pass.** A findings file establishes the **inventory** of what +that pass found and not the resolution of anything, so a later pass that does not mention an +earlier in-set Blocker or Major says nothing about whether it was resolved; reading its absence as +discharge would let an omission close a cycle. **The duty is not a second test on whether a pass is +clean, and the two are not run together.** A pass is clean on its own findings. Stated as the case +that separates them, because a reader who conflates them decides it wrongly: **pass 1 raises an +in-set Major; it is not resolved; pass 2 finds nothing.** Pass 2 **is** clean, and at or above the +floor it is **eligible** — and the cycle **still cannot close**, the duty being unmet; it takes the +continue branch until that Major is discharged. What the open Major does *not* do is make pass 2 +unclean. The **hold** a surfaced finding places on closure **participates in the ordering**: it +gates closing while it stands, and is discharged by the answers that surface requires. **It +attaches to every surfaced finding, whichever suspension surfaced it** — clean completion creates +none, because it wins before anything is surfaced. **No-clean-credit** — no pass carrying a +scope-stop trigger is credited as clean — also participates, and is a fact about that pass that +nothing discharges, a later pass being judged on its own findings. It is not a second test beside +the clean predicate but that predicate's second half, which is why it is stated in its words. + +**What a suspension asks, and what ends it.** **A hold ends when every answer its finding requires +has been given, in whichever direction each is given** — one **scope-stop** answer for a +single-trigger finding, both for one carrying both. That is **one rule with two parts**, how many +answers and which way each may go, and neither is a test the other has to pass. Where a health +suspension applies to the same pass, its shared continue-or-stop answer is **additional** to those +and not counted among them, the health readings asking about the loop rather than about this +finding. **The two health readings differ in what they surface, and therefore in what they hold.** +The **two-tell stop surfaces tells and no finding**, so it creates no hold; what it leaves +outstanding is its own continue-or-stop question, which the composition rule below holds the cycle +on until it is answered. The **clearly-stuck reading surfaces findings** — **the findings of the +pass being read that satisfy its regeneration condition, and only those**; earlier members of a +regeneration chain that were repaired or dismissed are history the reading consults and never +findings it re-surfaces, so no discharged finding takes a second hold. Each surfaced finding takes +a hold like any other. **A re-raised valid dismissal stays discharged for the resolve duty on +exactly the terms the clean predicate sets out above** — the same complaint, no new evidence, no +change to the text the dismissal turned on, and the dismissal's reason still true of the artifact. +The dismissal was the resolution and a reviewer repeating it does not undo it, so no second +dismissal is owed. **Where any of those fails the recurrence is an ordinary fresh finding** and is +handled as one — by the severity and scope rules at their own sources, which decide whether it is +in set and what it owes; reading the old dismissal as covering it would let a finding that has +since become true close a cycle. What a +qualifying recurrence creates is the **clearly-stuck hold**, ended by that reading's +continue-or-stop answer. Any trigger the recurrence independently carries raises its own +stop as usual. **Where one finding is surfaced by more than one route it carries a hold +component per route, and each is discharged by its own answer**: the **membership** component by +the membership answer **in either direction**, a decline releasing it exactly as an accept does; +the **question** component by the user's decision on that question; the **clearly-stuck** component +by that reading's continue-or-stop answer. **No answer discharges another route's component**, and +a finding surfaced by one route has one component, discharged by the one answer its surface asks +for. **Resumption is still the +composition rule's**, which waits for every outstanding answer. At a **membership stop** the answer is **accept** or **decline**; +**what each does to the assigned fix set is the absorb paragraph's, which defines that set from +these outcomes**, and what an in-set finding then owes is Mechanics · Severity's. A later answer +that contradicts a decline **does not reverse it**: the decline **remains binding** and the +contradiction is **surfaced to the user as information**, changing neither membership, nor the +cycle's state, nor any outstanding question — **whether the loop resumes is decided by the +composition rule below and by nothing here**, so a contradictory answer is never itself a +resumption and cannot step past a hold or a health question still awaiting its own answer. Nothing +here turns one answer into another, since that would let a finding be moved out of the set and back +into it to escape what it owes inside it. **There is no withdrawal inside the cycle that +declined**, a decline binding for the remainder of its cycle and admitting no exception; a +reconsideration is a later cycle's, where that decline has no effect at all and the finding takes +the ordinary route. Either answer is an **explicit, attributable decision on that specific +finding** — never silence, never a general remark about scope, never inferred, because a fix set +changed by inference is a fix set nobody chose. **Membership is answered against the set as the +absorb paragraph fixes it for the pass that raised the question**: a later broadening is a new fact +the **next** pass reads and never discharges a standing hold, a hold discharged by a scope change +being a hold nobody answered. At a **question stop** the answer is the user's decision on the +question and membership does not change; an out-of-set finding that opened one is a membership stop +as well. **Decline is available only at a membership stop**, that being the only stop whose +question is whether a finding belongs to the set. The **clearly-stuck and two-tell readings** ask +**continue or stop**. **Continue consumes the reading that raised the suspension**: a further +health suspension needs that reading recomputed over a pass run after the answer, which is new +data — so continue produces a distinct next state, and the same reading cannot return the same stop +unanswered. It permits an **unrevised** artifact **only where no repair is owed**; where effective +severity or scope requires one, that repair comes before the post-answer pass, since a pass run +over an unrepaired in-set Blocker or Major spends a look on text the rules already say must change. +**Stop parks the cycle**: open, not running, spending no passes, restarted only by an explicit +later continue — a distinct state from the suspended-awaiting-answer one it was in before the +answer. **That continue restarts the cycle and never skips an answer**: where any question the +suspension raised is still outstanding, it returns the cycle to suspended-awaiting-answer, and only +once every answer the composition rule requires has been given does the next pass run. So a cycle +parked with an unanswered membership or question stop cannot be continued into a pass, and cannot +sit parked with no transition either — the continue is always available and always moves it. +Nothing a parked cycle wrote is a closing commit, and a parked cycle nobody restarts is a human's +to resolve, exactly as the nonce rules already say of open cycles. + +**Composition, and what cannot happen.** Every **question** is answered on its own and the loop +resumes only when every answer resumes it — accept or decline at a membership stop, a decision at a +question stop, continue at the health readings; one stop answer parks the whole suspension, because +a loop resumed over an unanswered question decides it by running. **The clearly-stuck and two-tell +readings raise one question between them, not two**, both asking continue or stop, so one answer +carrying every reason ends both — an instance of the sentence before it, not an exception. **Two +pairings cannot occur**, and no rule ranks them: clean completion and a **scope stop**, since that +stop's triggers are the clean predicate's own second half, so a pass raising one is not clean; and +a zero-finding pass and any suspension, since it has nothing to surface, nothing regenerating, no +cluster and no require↔withdraw pair. **Clean completion and the clearly-stuck exit can**, and the +overlap is admitted rather than argued away: the two read different severity fields, as +Mechanics · Severity sets out, so an **in-set** Blocker or Major the ceiling demotes below Major +can regenerate across passes on a pass that is clean. **A declined finding is not a route into that +reading**: the third condition admits regeneration across repair attempts and a qualifying +re-raised validated dismissal, and a decline is neither — it is the user's decision that a *true* finding stays outside +the set, and it binds for the cycle. **The order decides it and no new rule is needed.** The +clearly-stuck paragraph's own precedence clause is stated here rather than there, because +precedence is evaluation order and this paragraph is where evaluation order is stated once; the +rationale that clause turns on stays beside the reading in that paragraph, which is where the +reading itself lives. **A clean completion takes precedence over this exit**: a Blocker/Major-free +pass **at or above the floor** has satisfied the clean-final-pass rule — collect the Minors and +Nits and close — and reporting "will not converge" on a converged loop is a false report. Below the +floor the pass **suspends**, the clean pass having failed eligibility. **That sentence ranks two +readings and licenses no closure**, its "close" being the closure this ordering defines and +carrying every condition that closure carries: a Minor or Nit bearing a scope-stop trigger makes +the pass unclean, so the sentence does not reach it, and an undischarged duty, a standing hold or +an unmet gate condition means the pass does not close — leaving it on the suspension branch where +one applies and the continue branch where none does, exactly as those branches say. + +**Gate A's content condition, and its closing act.** These are what Gate A adds to the conditions +the ordering states for every cycle; that list is there and is not repeated here. + +**Its content condition: the artifact as it now stands is identical to the text that went into the +final pass's review request.** Gate A hands the reviewer text rather than a git range, which is why +this condition is Gate A's and is written nowhere else. It is **current equality and deliberately +nothing more**: it does **not** say the artifact went untouched in between, and text edited and +then restored byte for byte satisfies it — stated here because the ordering's conditions answer a change +instead, and said of this condition rather than as a claim about every other. +That is a decision rather than an oversight: a content comparison cannot tell those two states +apart, and a condition nobody can check is a condition nobody applies. It likewise says nothing +about **what the reviewer consumed**: no part of this act is offered as evidence of the review +payload. + +**The act.** A Gate-A cycle has no WIP snapshot to replace, so it closes by **writing the closing +commit — or the closing message of one that already exists — carrying the records this cycle +already owes**: its provenance line and its per-pass curve, in the forms Mechanics fixes, neither +of them altered. **The content must survive into that commit, and carrying it there is part of the +act**: the equality above is read on the artifact, while the commit is written from the effective +index, so an act that does not carry that same content through has checked the condition without +performing it. **The commit that closes the cycle contains, at the artifact path, exactly the text +the condition was read against.** The safe command sequence and whatever demonstrates it belong to +the plan; **the duty belongs here**, and it needs no new fingerprint and no new record. + +**Two cases, told apart by the repository's current state rather than by which commit introduced +what; the safe git sequence for each belongs to the plan. That reading decides how a cycle closes, +never whether it may.** +- **`HEAD` already carries that text at the artifact path.** Close by **amending `HEAD`'s message** + to add the records, leaving that path as it stands. +- **`HEAD` does not, the reviewed text being still uncommitted.** **Commit it unchanged** and close + in that commit. This case had no answer before: an eligible pass over repairs nobody had + committed could neither close nor suspend. + +**No new revision of the artifact is made to close a Gate-A cycle**, a new revision being one no +pass has run against — and committing already-reviewed text that was never committed is not one. + +**Gate B's content condition, and its closing act.** These are what Gate B adds to the conditions +the ordering states for every cycle; that list is there and is not repeated here, so nothing below +is an inventory of what this gate requires. **This gate adds no content condition of its own**, +and the next paragraph says why. + +**Its content condition is not an artifact/request equality, and none is written for it.** Gate B +reviews a **diff** identified by `baseSha` and `headSha` rather than a text handed to the reviewer, +so there is no reviewed text to compare an artifact against, and **nothing is put in its place**: +content the final review request did not select can reach the closing commit, and **no rule in this +section reaches it**. **What this gate does have is every Gate-B duty this section already +states, at the paragraphs that state them, and none of them is summarised here** — a compressed +inventory is where a load-bearing part goes missing while the list still looks complete. + +**The act** is the one Mechanics · Finishing the cycle describes, in whichever shape that section +gives the repository's current state. **This paragraph says when it is performed and never which +shape it takes**: once the ordering reaches it, on an eligible pass with every closure condition +holding, **never a clean pass on its own**. + +**A commit the hook reads as cycle-closing is a Gate-B matter.** A non-`WIP` commit mid-cycle makes +the hook drop its Gate-B review state and that gate's counter — an observation about the counter, +since the cycle itself stays open until the conditions hold, so an accidental commit resets what +the hook reports and closes nothing; **what the reset erases is counter state, and it does not +invalidate a pass that already satisfied the validation rules this section states** — which is +where what makes a pass valid stays. **A stray commit that succeeds and does not amend also leaves the +`WIP:` snapshot as an ancestor**, an amend replacing the tip instead — +so in that one shape **the closing act still owes what Mechanics already requires of it: no `WIP:` +commit left in history.** Which git sequence reaches that from this state +belongs to the plan, as every other closing sequence does. **It does not reach a Gate-A cycle's count**, which the hook +clears at the skill boundaries that start a new Gate-A cycle rather than on any commit. + **What a loop absorbs, and what stops it — a question of scope, not of action.** A finding that corrects the correction you just made **and stays inside the assigned fix set** is **inside this loop's scope**: keep it here rather than handing it back, then act on it by its -severity exactly as Mechanics already says — Blocker/Major resolve, Minor/Nit collect and -never iterate. Ancestry decides where a finding belongs; it never decides what you do with -it, and it grants no Minor or Nit a repair round it would not otherwise get. **The assigned fix set is fixed before the pass you are answering: it is the -scope the approved story or plan assigns to this cycle, plus repair obligations you already -accepted in earlier passes.** A finding is in-set when repairing it stays inside that scope — -never merely because it arrived in the current pass, which would put every new finding in the -set by definition and leave the boundary deciding nothing. Where membership is genuinely -unclear treat the finding as **outside**, which costs a question and never a silent expansion. **A correction that leaves that set stops the -loop like any other out-of-scope finding**, even when it opens no new question at all — -absorbing it would grow the assigned work without anyone agreeing to that — and it resumes -the moment the user says whether the set now includes it. A finding -that opens a **new structural or contract question** stops the loop and goes to the user — -**size is not the test, novelty of the question is**, so a structural finding that is -genuinely small still stops it, while a long correction still aimed at the last correction -does not — provided that correction, too, stays inside the set, which its ancestry never -supplies on its own. **When a finding is both** — it corrects the last correction *and* opens a new -structural or contract question — **the new question wins and the loop stops**: novelty -overrides correction ancestry, because absorbing on ancestry is exactly how a contract -decision gets made without anyone choosing it. Stopping this way is **not an exit from the gate**: the floor, the -Blocker/Major filter and the clean-final-pass rule all stand, and the loop resumes on the -revised artifact once the question is answered — what the stop prevents is a loop -committing you to a design you never chose, which is a different failure from an -unfinished review. (Field-minted in `infinite-portfolio-canvas` and carried here because +severity exactly as Mechanics · Severity says. Ancestry decides where a finding belongs; it +never decides what you do with it, and it grants no Minor or Nit a repair round it would not +otherwise get. **The assigned fix set is fixed before the pass you are answering: it is the +union of the scope every approved story or plan governing this change assigns to this cycle, +plus every finding this cycle has accepted at a membership stop together with any repair +obligation accepted with it, minus every finding this cycle has declined.** **A decline excludes +the finding it answers and nothing else**: it does not cancel a repair obligation that the +approved scope, or another accepted finding, has independently put in the set. So where a `full` +Gate-B pass's two branch files carry the same complaint, **each line is a finding of its own, owes +its own explicit answer, and each answer binds only its own line** — an acceptance puts its own +finding in, a decline takes only its own finding out, and neither reads the other; the set is +whatever the definition above then computes. **A declined finding stays declined for the cycle**, +and performing a repair to discharge a different finding's obligation neither reverses that +decision nor returns it to the set. **This adds no deduplication, no reconciliation stop and no +rule that an acceptance overrides a decline** — the two answers are about different findings, and +until both are given the unanswered one's membership hold stands. **An accept puts the +finding in the set whatever its severity**: membership and the repair duty are different things, +so a Minor or Nit accepted into the set is in it though Mechanics · Severity asks no repair for +it, and a later pass that recomputed it as outside would raise the membership question a second +time and make the accept decide nothing. A finding is in-set when repairing it stays inside **the assigned fix set as +just defined** — never merely because it arrived in the current pass, which would put every new +finding in the set by definition and leave the boundary deciding nothing. Where membership is +genuinely unclear treat the finding as **outside**, which costs a question and never a silent +expansion. **A correction that leaves that set stops the loop like any other out-of-scope +finding** — except one this cycle has already declined, which is outside the set by that +decision and **raises no membership trigger on that account**, its membership being the one +question already answered — even when it opens no new question at all, and the membership +answer ends that finding's membership hold; **what the pass does next is the closure +ordering's**, which resumes only when every answer outstanding on that surface has been given. +A finding that opens a **new structural or contract question** — new meaning not already +answered in this cycle, so an answered question raised again stops nothing — stops the loop and +goes to the user — **size is not the test, novelty of the question is**, so a structural finding +that is genuinely small still stops it, while a long correction still aimed at the last +correction does not — provided that correction, too, stays inside the set, which its ancestry +never supplies on its own. **When a finding is both** — it corrects the last correction *and* +opens a new structural or contract question — **the new question wins and the loop stops**: +novelty overrides correction ancestry, because absorbing on ancestry is exactly how a contract +decision gets made without anyone choosing it. **Novelty overrides ancestry and nothing else: +where the finding is also out of set, both triggers hold and both answers are owed**, since a +question answered about a finding nobody placed in or out of the set leaves its membership +decided by default. Stopping this way is **not an exit from the gate**: it is a **suspension** +in the closure ordering's sense, the floor, the Blocker/Major filter and the clean-final-pass +rule all stand, and what the answer does is stated there — what the stop prevents is a loop +committing you to a design you never chose, which is a different failure from an unfinished +review. **A change to this set costs the cycle at least one further pass.** The set a pass was +begun under is the set its closure would rely on, so the window opens **when that set is fixed +for the pass**, as this paragraph defines it, and runs to the closing act; a change anywhere in +that window costs a further pass, **in either direction and whether or not the change is later +undone**, the set having governed the pass differently while it stood. That further pass must +itself be clean and every other closure duty must be satisfied; it is one more pass, not a +licence to close on the next one. (Field-minted in `infinite-portfolio-canvas` and carried here because the alternative was observed there: handing back a three-line repair-of-a-repair wastes a session, and absorbing a contract question spends a decision that was not the loop's to make.) @@ -246,24 +668,33 @@ make.) **Blocker curve across passes**, not any single pass's total — it is the better of the two signals, the total says less than it looks like, and one low count is a snapshot rather than a plateau. **Neither curve measures coverage:** a low Blocker count can sit beside an -entirely unreviewed subsystem. So this exit needs three things **together**, and a missing -one means keep going: a plateau visible across passes (six or more is where the field saw -one); an **affirmative judgement that coverage is sufficient**, stated — a known materially -unreviewed area forbids this exit outright, and disclosing it does not license it; and -**Blocker or Major findings that keep regenerating across genuine repair attempts**, each -round's fix producing the next. That third condition is what makes a plateau rather than a -finish, and it is why **a clean completion takes precedence over this exit**: a -Blocker/Major-free pass **at or above the floor** has satisfied the clean-final-pass rule — -collect the Minors and Nits and close — and reporting "will not converge" on a converged -loop is a false report. **Below the floor nothing closes**, and a zero-finding pass remains -the only exception, exactly as above; a Blocker/Major-free pass below the floor -carrying a Minor keeps -looping. -**Surfacing does not close the cycle, and that is what makes this reachable.** You surface -*with the finding still open* — the resolve rule is not waived, no pass is credited as -clean, and the loop resumes on whatever the user decides. Reading it as "stop instead of -fixing" would put the exit in competition with the rule that every Blocker and Major -resolves, and then nothing could satisfy both. +entirely unreviewed subsystem. So this exit needs three things **together**, and a missing one means only that *this* exit does not apply — what the pass does instead is the closure ordering's, read there in full: a plateau +visible across passes (six or more is where the field saw one); an **affirmative judgement that +coverage is sufficient**, stated — a known materially unreviewed area forbids this exit outright, +and disclosing it does not license it; and **Blocker or Major findings that keep regenerating +across genuine repair attempts**, each round's fix producing the next — **or a finding the author +has validly dismissed that the reviewer re-raises across passes on the terms the closure ordering +sets**, a recurrence failing them being an ordinary fresh finding and not a re-raise at all, the +re-raise standing in for +the regenerating fix, since a dismissal gets no repair and produces none, and a reviewer returning +to the same refuted point every pass says the same thing about the loop that a fix producing the +next finding says. That third condition is what makes a plateau rather than a finish. **Where this reading and a clean +completion both apply, the closure ordering decides it** — the precedence sentence lives there, +because precedence is evaluation order. +**Surfacing does not close the cycle, and that is what makes this reachable.** You surface *with +the cycle and the new hold still open* — the resolve rule stands over the finding exactly as +Mechanics · Severity states it, **which scopes it to the assigned fix set**, so a recurrence of +one already validly dismissed **stays resolved on the terms the closure ordering sets** and owes +neither a second dismissal nor a repair, while a recurrence failing any of them is an ordinary +fresh finding; +the hold stands until +its answers are given, and **what the answer does is the closure ordering's**. +**A pass is credited clean or not on its own findings**, as that ordering defines cleanliness; +surfacing a finding that carries a scope-stop trigger is what withholds the credit, and no +credit is withheld for surfacing alone. Reading this as "stop instead of fixing" would put the +exit in competition with the rule that every **in-set** Blocker and Major resolves, and then +nothing could satisfy both. + **Every pass report states three things about the floor**, from pass 1 onward: the derived floor, the risk and security values read, and the cited stories they were read from. A report giving the number alone leaves a reader unable to check the derivation @@ -283,15 +714,19 @@ pass demanding what an earlier pass had removed. Those three lines expose **five tells**: the finding count rising rather than falling; the Blocker count failing to fall; findings clustering on the **instrument** rather than on product behaviour; findings clustering on **prose about** either; and a require↔withdraw -pair. **Any two present makes stop-and-surface mandatory, not discretionary** — you report -the tells and hand the decision to the user, and the "clearly stuck" reading above is not a -precondition for it. A loop can be worth stopping long before it plateaus. +pair. **Any two present makes stop-and-surface mandatory, not discretionary** — read **after** the +clean-completion branch of the closure ordering, which outranks it **by taking the pass to the +closing act and only then** — and you report the tells and hand the decision to the user, and the "clearly stuck" +reading above is not a precondition for it. A loop can be worth stopping long before it plateaus. Recorded rationale, from the maintainer rather than from a measurement of this repo: in the Bricks consumer all five signals were measurable by **day two** of a week-long loop, and the cost was never detection — it was the absence of a duty to say so. That is why this is a reporting obligation with a mandatory threshold and not another heuristic to weigh. +**What the answer does** is the closure ordering's, which is where this stop's place among the +suspensions and what its answer produces are both stated. + **The two rules above do not compete**, and neither overrides the other: the absorb rule decides whether *a finding* is inside this loop's scope, this reading decides whether *the loop* can still converge. A small correction-of-a-correction that stays inside the assigned fix @@ -345,8 +780,8 @@ Append to the gate prompt: > Severity is one of exactly: BLOCKER | MAJOR | MINOR | NIT — no other token. > Every line before the terminator is exactly one finding line — no blank lines, > headings, prose or wrapped continuations. End the file with a final line reading -> exactly `END OF FINDINGS ( total)`, `` being the number of finding lines. A -> clean pass is the single body line `NO FINDINGS` with `END OF FINDINGS (0 total)`. +> exactly `END OF FINDINGS ( total)`, `` being the number of finding lines. +> A **clean findings file** is the single body line `NO FINDINGS` with `END OF FINDINGS (0 total)`. > > Then reply with ONLY one line per branch — ` | pass

| findings | ` > — or `INCOMPLETE | | ` if you could not write the file. An unwritten @@ -433,8 +868,9 @@ source while the cycle runs, and it is the one the candidate rules above apply t may be present and the run must decide which, if any, is its own. **History is the source once the cycle's own commit exists**, and there is no search there: the cycle is reading **its own commit body**, so kind and artifact are settled by which commit is being read, and the nonce is -taken from the provenance line and the curve, which must agree. A Gate-A cycle mid-run has no -such commit and therefore has only the working record. Recovering a single candidate from +taken from the provenance line and the curve, which must agree. A Gate-A cycle has such a commit only once its own closing commit exists — an already-committed +revision of the reviewed text is not one, carrying no provenance line and no curve for a nonce to +be taken from — so mid-run it has only the working record. Recovering a single candidate from **either** keeps identity **as far as the field can distinguish cycles** — two cycles sharing a nonce are one cycle to it. **No candidate, disagreeing sources, or more than one candidate → no identity: start a new cycle**, which costs @@ -570,27 +1006,35 @@ not all of `.context/`, which would strip the committed `codex-gate.on` adoption **spec** right after brainstorming (before `writing-plans`), then on the **plan** before `executing-plans`/`subagent-driven-development` — catching a spec flaw before it's baked into the plan. Tool: `mcp__codex__exec` (raw; - reviews the TEXT you pass, not the git tree). Use ONE broad prompt, re-run it - each pass over the revised artifact (don't narrow per-dimension; new findings - surface because the artifact changes between passes). The prompt MUST open with + reviews the TEXT you pass, not the git tree). Use ONE broad prompt: **its review question stays the same every pass, while the artifact text it + carries and the dimensions it asks for are always the current ones** — the lens sets the profiles + section derives are recomputed from the current profile and cited set each pass and appended, since + a profile or cited-set change changes what is owed. Re-running it over an **unrevised** artifact + is legitimate wherever no repair is owed, an edit made to justify a pass being no reason to run + one. Don't narrow per-dimension: new findings surface because the artifact changed, because an + answer given since the last pass changed what the rules require of it, or because a broad prompt + reaches what the last reading did not. The prompt MUST open with *"Use the superpowers:brainstorming skill to review this spec,"* (say "plan" on the plan run), then ask Codex to check it against our settled decisions and surface **contradictions/inconsistencies, missing requirements, unhandled state/edge/error/empty/concurrent paths, and risks to the Key Invariants (@AGENTS.md) — plus anything else** (coverage floor, not a cage). Append the intent + artifact text + which invariants it touches. Ask for **every** finding - with severity and confidence — you filter to Blocker/Major downstream, Codex - never does, because a model told to report only high severity drops real - findings silently (`docs/prompt-standards.md`, "coverage first, filter later"). - Ask for one line per finding and a literal `NO FINDINGS` when a pass is clean — - the explicit clean signal is what lets you exit the loop: + with severity and confidence — **you filter to Blocker/Major for what must be repaired and read + every line for everything else**, Codex never filters, because a model told to report only high + severity drops real findings silently (`docs/prompt-standards.md`, "coverage first, filter later"). + Ask for one line per finding and a literal `NO FINDINGS` when a pass found none — that explicit + signal is what lets a pass be read as clean without inspecting it, and a pass carrying only Minors + **and no scope-stop trigger** is clean too and could never produce that file: ``` MAJOR | high | §3 "Retry policy" | retry count unbounded | a poisoned job loops forever | cap at 5, then dead-letter NO FINDINGS ``` - Each pass: validate, revise, re-run. Before each read pass, settle mechanically what + Each pass: validate, revise **where a repair is required**, and re-run **where the closure + ordering selects its continue branch** — where it selects a suspension instead, the answer comes + first and that ordering says what the answer produces. Before each read pass, settle mechanically what the artifact asserts and a machine can decide without side effects — cited paths, quoted passages, stated counts, the syntax of standalone fenced blocks — because a read pass spends expensive judgement on what a parser settles in seconds and misses it @@ -614,9 +1058,11 @@ not all of `.context/`, which would strip the committed `codex-gate.on` adoption handed both artifacts could catch it — which is why this is a rule about what you commit, not a claim about what the gates detect. - Same coverage rule as Gate A: put "report every finding with severity and confidence; say - `NO FINDINGS` if clean" in `additionalContext`, with the same one-line format. - You filter to Blocker/Major, Codex never does. + Same coverage rule as Gate A: put "report every finding with severity and confidence; write + `NO FINDINGS` only when the branch found none" in `additionalContext`, with the same one-line + format. **You filter to Blocker/Major for what must be repaired, and read every line for + everything else** — cleanliness, the scope triggers, the assigned fix set and the loop-health + readings all take Minor and Nit lines. Codex never filters. **Standing lens, every Gate-B call: "which existing statements does this diff falsify?"** A change makes sentences wrong in files it never touches. Checks scoped to the edited @@ -669,10 +1115,8 @@ the reviewer, the other obliges the author. it; risk's *abuse* and security's *abuse paths* are **one lens carrying both labels**, not two questions. -Lenses are **different questions, not more passes** — they change what a pass asks, never -how many a cycle owes. The Blocker/Major filter, the file-first findings protocol and the -clean-final-pass rule are unchanged. The floor is not among them: it is no longer a fixed -number but derives from the profile and the cited set. +Lenses change what a pass asks, never how many a cycle owes: **the lens sets** leave every other +rule in this section alone. **Reading the profile — five cases, five answers:** 0. **The ordinary case**: every cited story is readable and its profile resolves → derive the @@ -743,10 +1187,12 @@ appear, or uses a fixture that never reaches the branch it covers reports succes of how it was wired, not because the thing it checks succeeded. **The evidence entry lives in the commit body** (see Mechanics), carries the **story path -and the named evidence but not the mode value**, and is **revalidated before every Gate-B -re-review and before the cycle-closing amend** — a fix changes the diff even when the -profile sits still. If revalidation changes the entry, the clean pass no longer covers what -is being committed: fix, re-review, close on the entry that pass validated. +and the named evidence but not the mode value**, and is **revalidated before every Gate-B re-review and before the commit its closing act +produces** — a fix changes the diff even when the profile sits still. If revalidation changes the entry, the pass was read against an entry that no longer stands: **fix the +entry and re-review on it**, and **which branch the pass takes meanwhile is the closure ordering's, +read there in full** — this paragraph states the repair and never the branch. The pass +that follows is read by that ordering like any other and closes only if it reaches closure, on the +entry revalidated for it. **Every Gate-B call and re-review carries the path of every cited story**, so the reviewer reads each profile itself, **plus the current evidence entry, quoted verbatim, for each @@ -768,9 +1214,11 @@ directions; an agent never moves it alone. On confirmation, correct the header a one profile-log line. Any axis change **voids every prior override**, raised or lowered, and `+abuse-path` follows the current security value. Passes already run under the lower profile **keep counting** toward the floor; only the **final clean pass** must run under -the current profile. Inside an active Gate-B cycle, fold the edit into the active `WIP:` -snapshot by amend — a non-`WIP` commit reads to the hook as the cycle closing and would -discard the accumulated passes. +the current profile. Inside an active Gate-B cycle, the edit must end up **in the content the next review reads** — +folded into the active `WIP:` snapshot by amend where that snapshot is the tip, and otherwise +reviewed as its own change, since no amend reaches a snapshot a stray commit has made an ancestor +and this section prescribes no operation that does. A non-`WIP` commit reads to the hook as the +cycle closing and would discard **the hook's count of** the accumulated passes. **While a gate is running, the floor derives from the current profile at each pass.** Passes already run keep counting; closing requires the floor as currently derived. These @@ -800,13 +1248,21 @@ header changed mid-call, or whether the lens sets were appended. This is instruc like the rest of §5; the detection is a reader comparing the pass against the story. ### Mechanics (reference) -- **Severity:** Blocker (wrong/unsafe/breaks invariant) · Major (design flaw → - rework) → both must resolve. Minor · Nit → collect, never iterate. +- **Severity:** Blocker (wrong/unsafe/breaks invariant) · Major (design flaw → rework) → both + must resolve, **for every finding in the assigned fix set as the absorb paragraph computes it**. + **A finding is resolved by a repair or by a validated dismissal** — the author's judgement, + carrying the one-line why this section already requires, that the finding is not true of the + artifact. A dismissal does not rewrite the pass that found it and the later clean pass is still + owed; **a dismissal is not a decline**, a dismissal saying the finding is false and a decline + being the user's decision that a **true** finding stays outside the set. Minor · Nit → collect; + **their severity buys no repair round and no further pass.** Where accepting one into the + assigned fix set costs a pass, that cost is the **set change's** and is stated at the absorb + paragraph, not this severity's. **Deciding severity — one procedure. The subject list is illustration, not a second rule.** Name what in the system consumes this text — whatever *acts* on it — and the - decision that act takes differently if the text is wrong. Both are required. If you - cannot name both, the finding is Minor or below: collect, never iterate. + decision that act takes differently if the text is wrong. Both are required. If you cannot name both, the finding is Minor or below: collect; **its severity buys no repair + round and no further pass**, and any pass a later scope decision costs is that decision's. The exclusions are contract, not commentary. The reader must consume the text in the system's *operation*, not in reviewing it — the review pass raising the finding is not @@ -827,12 +1283,31 @@ like the rest of §5; the detection is a reader comparing the pass against the s This is the finding-level analog of the path-level prose exemption: one principle at two granularities — text that *describes* the product versus text that *is* the product. - **How this demotion bears on the loop-health measures — the per-pass counts, the finding - clusters and the stop thresholds — is not settled here, and this change does not settle it. - Until it is, a pass whose outcome would turn on that question reports the question and - stops rather than deciding it** — the same answer any unresolved gate question gets. - That question is owned by the loop-rule consolidation work in - `docs/superpowers/stories/2026-08-29-loop-rule-consolidation-story.md`. + **The demotion changes what a cycle must resolve, never what it observes.** **Every loop-health + reading observes the findings as the reviewer produced them, before the ceiling is applied** — + so a demoted finding still counts in the finding total and in its cluster, and a Blocker demoted + to Minor is still a Blocker to the curve and still regeneration to the clearly-stuck reading's + third condition. Where such a reading uses severity at all it takes the **reader-normalized + pre-ceiling severity**, which is what the Reader paragraph above already produces from the + findings file — case-folded, and a non-empty unrecognized token read as `MAJOR` — never the raw + token, so a finding written `IMPORTANT` enters the curve as a Major. **The ceiling is applied + after that and only to what the cycle owes**: cleanliness and the resolve duty read the + effective severity, so the same finding can be a Major to the curve and a Minor **at effective + severity**, and that difference is the point rather than a discrepancy. It stays in the fix set + either way; the ceiling moves what the cycle owes for it and never whether it is in. The line is **what the cycle owes + versus what it observes about itself**, which is why no list of readings has to be kept complete + here. Two reasons for the split. The curve's **three numeric series** must stay derivable from + the validated findings files alone wherever those files remain available — the rest of the curve + is not and does not claim to be, its cycle field, pass ranges and model identifiers coming from + elsewhere, and an unrecoverable count being written `?` exactly as the standing grammar allows — + the finding total counts finding lines and the Blocker and Major series count the lines whose + normalized severity is each, which is the only thing that makes a self-reported curve checkable; + the subject clusters use no severity at all, being a judgement per finding that no count + reproduces. And the demotion is the author's judgement about **the finding's repair severity**, + never about which findings the fix set contains — a separate predicate the absorb paragraph + defines, and one this must not be read as touching; a loop spending passes on findings the + author keeps demoting is exactly what the prose-cluster tell exists to surface, and lowering the + counts by that same judgement would hide it. - **Tool routing:** docs (spec/plan, incl. code snippets) → `mcp__codex__exec`; implemented diff → `mcp__codex__review`. Never `review` a doc — it reads the git range, not the text. @@ -842,12 +1317,22 @@ like the rest of §5; the detection is a reader comparing the pass against the s pre-commit, `baseSha` = HEAD is an empty range (HEAD..HEAD) — make a WIP commit and set `baseSha` to its parent. **Name that commit `WIP: …`** — the hook treats a `wip`-prefixed commit message as cycle-internal, so it neither fires a Gate-B STOP - nor resets your pass counters. A pre-review snapshot named anything else reads as a - real commit and closes the cycle, discarding the passes you just accumulated. - **Finishing the cycle:** after the final clean pass, close it with - `git commit --amend -m ""` — that replaces the WIP commit, and the hook - reads the amend as the real cycle-closing commit. If several WIP snapshots piled up, - `git reset --soft ` first, then commit once. Amend rather than a + nor resets your pass counters. A pre-review snapshot named anything else reads to the hook as a real commit: the hook treats the + cycle as closed and **discards its count of the passes you just accumulated**, while the cycle + itself stays open until the closure ordering's conditions hold. **What the hook loses is its counter state**, and that + counter is not what makes a pass valid — so the reminder now understates what you hold, and no + close was intended or made. **What such a commit does to the repository, and what the closing act + then owes, is in the Gate-B closure paragraph**, not here. + **Finishing the cycle:** **when a Gate-B cycle's closing act is performed is the closure + ordering's, stated there entire**; this section gives only the operation. Close it with + `git commit --amend -m ""`, which replaces the WIP commit; **where a `WIP:` snapshot + would survive the amend** — several piled up, or a stray non-amending commit made one an + ancestor — **`git reset --soft `, then commit once instead**. This section is the only + place either shape is defined. **The hook treats any + non-`WIP` commit *attempt* as a Gate-B boundary and clears its state even where the command + fails**, so a failed closing act leaves that counter cleared — a fact about the + counter and not about the cycle. This section states the operation and never whether the cycle may + close. Amend rather than a follow-up commit for two reasons: a `WIP: …` commit left in history defeats the naming convention it exists for, and a follow-up commit has nothing to commit when the review produced no fixes. @@ -896,18 +1381,48 @@ like the rest of §5; the detection is a reader comparing the pass against the s different way, the last of them demonstrably so, and the rule for a claim needing a fourth correction is to delete it. Whoever needs to know why reads the file and the hook. - **These records are one contract, and a partial adoption breaks it.** The nonce, the slot - naming, the provenance line, the curve, this carry rule **and the unknown-start activation - semantics that say what a cycle owes when its starting rules cannot be established** depend on - one another, and the requirement is that the adopted definitions **agree**, not merely that all - of them are present: a curve - without a cycle field cannot be attributed, a slot rule without a nonce has nothing to key on, - and a carry rule naming records a project does not produce is inert. **A project whose text - carries some of them and not others, or carries all of them in versions that disagree, stops - and has a human complete, revert or reconcile the adoption before running a gate under it** — - disagreement is the harder case and gets the same stop, because a project holding two - definitions of a record has no single answer to what it owes — the same answer, and for the same reason, as a partial - adoption of the floor rule. + **These rules and records are one contract, and a partial adoption breaks it.** The nonce, the + slot naming, the provenance line, the curve, the carry rule, the unknown-start activation + semantics **and the closure ordering together with every rule it reads** depend on one another, + and the requirement is that the adopted definitions **agree**, not merely that all of them are + present. **Membership is decided by a test a reader can apply to the text in front of them, with + no list to consult, and the test reads what a rule states rather than what changing it would do: + a live rule belongs to this contract when what it says **defines the validity of an input the + closure ordering reads, or how that input is read**, which branch a pass takes, what a hold is or what discharges it, **what a suspension asks, or what state its answer or an + incomplete closing act produces**, **what a gate's closing act is**, **which version of these rules + governs a cycle**, + whether a cycle may close **or may terminate without running a pass at all, the Gate-B triviality + skip being the one such route and its eligibility test therefore a member**, or the production, + identity or transport of **any §5 cycle record, required or optional** — §5 entire and not the + Mechanics subsection this paragraph sits in, record duties being stated in both, and the optional + companions' slot rules carrying the nonce that keeps sibling cycles apart.** **Read it on the sentence, never on the section the sentence sits in.** + A sentence is a member when **it itself** fixes one of those things — what counts as a valid + finding line, which files or records are owed, what ends a hold. It is not a member when it only + shapes what a review produces, as the choice of reviewer, the lens set and **prompt wording that only frames the + review question** do: those change the findings without deciding what a finding *is* or what the + ordering may do with one. **Prompt wording that fixes a valid input is a member**, the gate-prompt + sentences defining a finding line and the `NO FINDINGS` signal being exactly that. **No paragraph is exempt as a paragraph** — a sentence inside a routing or prompt paragraph + that fixes a valid input or an owed file is a member, and a sentence anywhere that only influences + the findings is not. The examples follow the test; they do not stand in for it. + Asking instead what an imagined edit would do decides nothing, because + any rule can be edited into deciding a branch and none decides one when edited cosmetically, so + membership would follow the edit a reader pictured rather than the text in front of them. The + last clause is why the squash carry belongs: it moves no pass and + decides no branch, and a record that does not survive the merge is unreachable from the squash + commit and from `main`'s history. A curve without a cycle field cannot be reliably told from + another cycle's in every multi-cycle context — kind and surrounding context sometimes separate + them, which is why the standing rule calls missing attribution a limitation rather than a + disqualification — a slot rule without a nonce cannot keep sibling cycles + apart — the bare names staying reserved for the legacy single-cycle case they already serve — a + carry rule naming records a project does not produce is inert, and a clean predicate without the + fix-set boundary it reads decides membership by accident. + **A project whose text carries some of them and not others, or carries all of them in versions + that disagree, stops and has a human complete, revert or reconcile the adoption before running a + gate under it.** **The test classifies sentences that are present, and completeness is not among + what it establishes**: where an adoption drops a member together with every sentence that would + refer to it, what remains reads as coherent and nothing in it marks the absence. **That bounds + what this text lets a reader detect, never what the rule obliges** — a partial adoption is a stop + however it becomes known, and learning of it from outside this text is learning of it. **On squash-merge, copy every evidence entry, every human-exception record, the provenance lines, the curves and any skipped cycle's skip record TOGETHER WITH THE SKIP REASON IT POINTS AT in the squash range into the squash body — a skip record carried without its reason is a pointer into a body the squash has made unreachable — the squash commit is the only body the merge carries into `main`'s history, so anything left behind is unreachable from it.** @@ -955,8 +1470,9 @@ like the rest of §5; the detection is a reader comparing the pass against the s contributing to a split logical pass is listed**, joined by `+`, since recording one of two is the same loss as recording none. - **Majors are recorded as well as Findings and Blockers**, because the severity rule moves the - Blocker/Major line rather than the total, so totals and Blockers alone could not show even a + **Majors are recorded as well as Findings and Blockers**, because the three series are read + before the ceiling and the mix among them is what a later reader compares; the ceiling moves what + a cycle owes and leaves these counts alone, so totals and Blockers alone could not show even a change in the mix. **Subject categories are deliberately not recorded** — they are a judgement per finding rather than a count, and the findings files carry the material. @@ -1005,8 +1521,9 @@ like the rest of §5; the detection is a reader comparing the pass against the s Accepted because: ``` - **Which commit:** an ungated change records it in that commit; a Gate-A cycle in the spec or - plan commit; a Gate-B cycle in the WIP commit, restated by the closing amend. Several records + **Which commit:** an ungated change records it in that commit; a Gate-A cycle in the commit its + closing act uses; a Gate-B cycle in the WIP commit, restated by the commit its closing act + produces. Several records accumulate; order means nothing. **A decision made after its commit closed** — during PR review, say — goes in whichever of @@ -1030,9 +1547,11 @@ like the rest of §5; the detection is a reader comparing the pass against the s was already the human's to make about something genuinely optional. It is **never** the answer to a below-floor pass, an unclean final pass, a `STOP and surface`, a Gate-A or Gate-B obligation, or a profile-derived evidence requirement — and more generally **it authorizes - nothing that any mandatory rule in this file or in `AGENTS.md` requires.** Those have their - own terminal actions and this paragraph changes none of them: on a STOP you still stop, and - neither a human's assent nor this record lets an agent close or continue a cycle. + nothing that any mandatory rule in this file or in `AGENTS.md` requires.** Those have their own terminal actions and this paragraph changes none of them: on a STOP you + still stop, and **neither a human's general assent nor this record** lets an agent close or + continue a cycle. **The answers a suspension asks for are not assent of that kind**: they are the + answers the closure ordering prescribes, and both which answers those are and what they produce are + stated there. **"Mandatory" is not limited to this file.** A rule in `AGENTS.md`, a project doc, CI, a branch policy or the platform is equally out of reach — under **Wait for**, diff --git a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md index 79649d3..3c73db5 100644 --- a/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md +++ b/docs/superpowers/plans/2026-09-14-loop-rule-consolidation.md @@ -61,6 +61,91 @@ It adds no task, prerequisite or closure condition to this plan. | P16 | strict-reading list | `and the nonce duties at their strictest — the cycle` | 157 | 364 | | P17 | §F item 14, Named residual | `Hook text is out of scope here` | 139 | 346 | | P18 | §F item 18, work-loop line | `execute → tests green → Gate B → commit` | 63 | 262 | +| P19 | `a12`, kept, shares its line with `a13` (Task 0 preservation) | `where the cycle's own closure rules are satisfied.` | 148 | 355 | +| P20 | `a14`, kept, shares its line with `a13` (Task 0 preservation) | `Nothing here writes the floor knob: it stays the user's, never written, never` | 151 | 358 | +| P21 | `h6`, kept, shares its line with `h5` (Task 0 preservation) | `Several records` | 1009 | 1193 | +| P22 | `h18`, kept, shares its line with item 4's block (Task 0 preservation) | ``nothing that any mandatory rule in this file or in `AGENTS.md` requires.**`` | 1033 | 1217 | +| P23 | `b3`, replaced (Task 3 OLD) | `Blocker/Major resolve, Minor/Nit collect` | 608 | 815 | +| P24 | `b8`, replaced (Task 3 OLD) | `A finding is in-set when repairing it stays inside that scope` | 612 | 819 | +| P25 | `b11`, replaced (Task 3 OLD) | `like any other out-of-scope finding**, even when it opens no new question at all` | 616 | 823 | +| P26 | `b13`, replaced (Task 3 OLD) | `**new structural or contract question** stops the loop and goes to the user` | 619 | 826 | +| P27 | `b16`, replaced (Task 3 OLD) | `without anyone choosing it. Stopping this way is` | 626 | 833 | +| P28 | `b17`, replaced (Task 3 OLD) | `**not an exit from the gate**: the floor, the` | 626 | 833 | +| P29 | `b18`, replaced (Task 3 OLD) | `revised artifact once the question is answered` | 628 | 835 | +| P30 | `b1`, carried (Task 3 preservation) | `that corrects the correction you just made **and stays inside the assigned fix set** is` | 606 | 813 | +| P31 | `b2`, carried (Task 3 preservation) | `keep it here rather than handing it back` | 607 | 814 | +| P32 | `b4`, carried (Task 3 preservation) | `Ancestry decides where a finding belongs; it` | 609 | 816 | +| P33 | `b5`, carried (Task 3 preservation) | `it grants no Minor or Nit a repair round it would not` | 610 | 817 | +| P34 | `b6`, carried (Task 3 preservation) | `The assigned fix set is fixed before the pass you are answering` | 610 | 817 | +| P35 | `b9`, carried (Task 3 preservation) | `never merely because it arrived in the current pass` | 613 | 820 | +| P36 | `b10`, carried (Task 3 preservation) | `treat the finding as **outside**, which costs a question and never a silent` | 615 | 822 | +| P37 | `b14`, carried (Task 3 preservation) | `**size is not the test, novelty of the question is**` | 620 | 827 | +| P38 | `b15`, carried (Task 3 preservation) | `does not — provided that correction, too, stays inside the set` | 622 | 829 | +| P39 | `c1`, kept, passage (c) prefix (Task 4 preservation) | `**Blocker curve across passes**, not any single pass's total` | 664 | 868 | +| P40 | `c2`, kept, passage (c) prefix (Task 4 preservation) | `signals, the total says less than it looks like, and one low count is a snapshot rather` | 665 | 869 | +| P41 | `c3`, kept, passage (c) prefix (Task 4 preservation) | `**Neither curve measures coverage:** a low Blocker count can sit beside an` | 666 | 870 | +| P42 | `c4`, replaced (Task 4 OLD) | `one means keep going` | 668 | 872 | +| P43 | `c5`, carried (Task 4 preservation) | `visible across passes (six or more is where the field saw` | 668 | 872 | +| P44 | `c6`, carried (Task 4 preservation) | `coverage is sufficient**, stated` | 669 | 873 | +| P45 | `c7`, carried (Task 4 preservation) | `unreviewed area forbids this exit outright` | 670 | 874 | +| P46 | `c8`, replaced (Task 4 OLD) | `round's fix producing the next. That` | 672 | 876 | +| P47 | `c9`, moved to §A (Task 4 source absence) | `finish, and it is why **a clean completion takes precedence over this exit**` | 673 | 877 | +| P48 | `plateau rationale`, no id, stays (Task 4 preservation) | `That third condition is what makes a plateau rather than a` | 672 | 876 | +| P49 | `c10`, moved to §A (Task 4 source absence) | `Blocker/Major-free pass **at or above the floor** has satisfied the clean-final-pass rule —` | 674 | 878 | +| P50 | `c11`, moved to §A (Task 4 source absence) | `collect the Minors and Nits and close — and reporting` | 675 | 879 | +| P51 | `c12`, moved to §A (Task 4 source absence) | `**Below the floor nothing closes**` | 676 | 880 | +| P52 | `c13`, moved to §A (Task 4 source absence) | `the only exception, exactly as above;` | 677 | 881 | +| P53 | `e1`, kept, passage (e) (Task 5 preservation) | `Those three lines expose **five tells**` | 701 | 906 | +| P54 | `e2`, kept (Task 5 preservation) | `the finding count rising rather than falling` | 701 | 906 | +| P55 | `e3`, kept (Task 5 preservation) | `Blocker count failing to fall` | 702 | 907 | +| P56 | `e4`, kept (Task 5 preservation) | `findings clustering on the **instrument** rather than on` | 702 | 907 | +| P57 | `e5`, kept (Task 5 preservation) | `findings clustering on **prose about** either` | 703 | 908 | +| P58 | `e6`, kept (Task 5 preservation) | `either; and a require↔withdraw` | 703 | 908 | +| P59 | `e9`, carried (Task 5 preservation) | `hand the decision to the user, and the "clearly stuck"` | 705 | 910 | +| P60 | `e10`, kept, outside §D's block (Task 5 preservation) | `A loop can be worth stopping long before it plateaus.` | 706 | 911 | +| P61 | `e11`, kept, C only (Task 5 preservation) | `reporting obligation with a mandatory threshold and not another heuristic to weigh.` | 711 | — | +| P62 | `Severity resolve duty`, replaced (Task 6 OLD) | `rework) → both must resolve. Minor · Nit → collect, never iterate.` | 1226 | 1414 | +| P63 | `g2`, dropped (Task 6 absence) | `Until it is, a pass whose outcome would turn on that question reports the question and` | 1254 | 1442 | +| P64 | `g3`, dropped (Task 6 absence) | `the same answer any unresolved gate question gets` | 1255 | 1443 | +| P65 | `g4`, dropped, C only (Task 6 absence) | `That question is owned by the loop-rule consolidation work in` | 1256 | — | +| P66 | `c15`, carried (Task 7 preservation) | `**Surfacing does not close the cycle, and that is what makes this reachable.**` | 680 | 884 | +| P67 | `c16`, replaced (Task 7 OLD) | `*with the finding still open*` | 681 | 885 | +| P68 | `c17`, replaced (Task 7 OLD) | `the resolve rule is not waived` | 681 | 885 | +| P69 | `c19`, replaced (Task 7 OLD) | `the loop resumes on whatever the user decides` | 682 | 886 | +| P70 | `c20`, replaced (Task 7 OLD) | `with the rule that every Blocker and Major` | 683 | 887 | +| P71 | `a13, first sentence`, replaced; first sentence's own absence (Task 7) | `This replaces the pass-count number` | 148 | 355 | +| P72 | `a15`, carried (Task 7 preservation) | `Open a TodoWrite "Codex pass N" per pass;` | 152 | 359 | +| P73 | `a18`, moved to §A (Task 7 source absence) | `if the pass at the floor still finds Blocker/Major, keep going until` | 153 | 360 | +| P74 | `a19`, moved to §A (Task 7 source absence) | `The only early exit` | 154 | 361 | +| P75 | `a20`, moved to §A (Task 7 source absence) | `don't manufacture findings to pad` | 155 | 362 | +| P76 | `a21`, carried (Task 7 preservation) | `advisory — validate before applying` | 156 | 363 | +| P77 | `a22`, carried (Task 7 preservation) | `dismissed finding → one-line why` | 156 | 363 | +| P78 | `i4`, carried (Task 7 preservation) | `touches — at minimum` | 175 | 382 | +| P79 | `i5`, carried (Task 7 preservation) | `severity classified without the demotion` | 176 | 383 | +| P80 | `i6`, carried (Task 7 preservation) | `the provenance-line duty owed` | 176 | 383 | +| P81 | `i7`, carried (Task 7 preservation) | `owed, the curve` | 176 | 383 | +| P82 | `i8`, carried (Task 7 preservation) | `the nonce duties at their strictest` | 177 | 384 | +| P83 | `i1`, kept, passage (i) (Task 7 preservation) | `From the commit that ships them` | 173 | 380 | +| P84 | `i2`, kept (Task 7 preservation) | `finishes under the rules it started with.` | 174 | 381 | +| P85 | `i3`, kept, shares its sentence with the list (Task 7 preservation) | `established it takes the stricter reading of every part this change touches` | 175 | 382 | +| P86 | `i9`, kept (Task 7 preservation) | `the cycle is treated as post-rule, so it` | 177 | 384 | +| P87 | `i10`, kept (Task 7 preservation) | `the working record stays optional and a skipped cycle still writes no findings slots` | 180 | 387 | +| P88 | `i11`, kept (Task 7 preservation) | `` cannot recover a nonce it starts a new cycle rather than claiming `none (pre-rule)`, that reserved `` | 181 | 388 | +| P89 | `i12`, kept, discharged (Task 7 preservation) | `Each further rule this change ships adds its own strict` | 182 | 389 | +| P90 | `i13`, kept (Task 7 preservation) | `cycle a floor of 1 and skip passes on the strength of not knowing when it started` | 184 | 391 | +| P91 | `i14`, kept (Task 7 preservation) | `knob set above 3 is not lowered by this fallback` | 185 | 392 | +| P92 | `i15`, kept (Task 7 preservation) | `A revert is itself a shipping commit for` | 185 | 392 | +| P93 | `i16`, kept (Task 7 preservation) | `the old rules, and the activation rule wins wherever the start is determinable; the` | 186 | 393 | +| P94 | `a1`, carried inside item 8a (Task 8 preservation) | `**Both gates are a LOOP with a HARD FLOOR: a minimum number of passes per run` | 92 | 299 | +| P95 | `h3`, carried inside item 7 (Task 8 preservation) | `an ungated change records it in that commit` | 1499 | 1688 | +| P96 | `§F item 10`, hook only, the honesty claim (Task 10 OLD; `codex-gate.sh` line 967) | `this floor is the only thing keeping the spec review honest` | — | — | +| P97 | `§F item 10`, hook only, the tail (Task 10 OLD; `codex-gate.sh` line 967) | `Run more passes before executing` | — | — | +| P98 | `§F item 11`, hook only, the Gate-A clean definition (Task 10 OLD; line 973) | `Proceed only if your final pass was clean` | — | — | +| P99 | `§F item 12`, hook only, the Gate-B clean definition (Task 10 OLD; line 956) | `commit only if your final pass was clean — no new Blocker/Major.` | — | — | +| P100 | `§F item 13`, hook only, the WIP reminder (Task 10 OLD; line 908) | `then make the real commit when your final pass is clean` | — | — | +| P101 | `§F item 15`, hook only, the no-fingerprint reminder (Task 10 OLD; line 933) | `Run Gate B (mcp__codex__review) now; if this repeats` | — | — | +| P102 | `§F item 16`, hook only, the stale-fingerprint reminder (Task 10 OLD; line 945) | `one clean pass is the complete remedy for the staging and post-upgrade cases too` | — | — | +| P103 | `§F item 17`, hook only, the below-floor instruction (Task 10 OLD; line 947) | `or proceed only if $policy's skip rule applies to this change` | — | — | **The fourteen §F prompt-copy items (Task 8) — fifteen rows, because item 7 changes two clauses on two lines — derived from each item's cited lines and checked the same three ways.** Pass 2 found these deferred to the executor as `` placeholders, which put fourteen meaning-changing checks outside Gate A's reach; they are concrete now. The NEW halves stay deferred, for the reason the paragraph below gives. @@ -3710,39 +3795,411 @@ this plan and in `.context/codex-reviews/`, both tracked, both already inside th ## b11/b13 equivalence (Task 12 output) -*Empty until Task 12 runs. Task 12 replaces this entire section, carrying both predicates in full -as extracted and the result in both directions, per copy.* +**Subject.** The installed §A (`**How a cycle ends`) and the installed passage (b) +(`**What a loop absorbs`), read in C; W carries the same bytes over both (Task 1 and Task 3 parity: +§A no difference, §B's common span no difference), so the result is the same per copy. + +**What the ordering attributes to the absorb paragraph (step 1), in full:** +1. Clean predicate: a clean pass carries **no scope-stop trigger** — "the two the absorb paragraph + defines, read there and not redefined here, each already carrying the qualification **an answer + given before that pass ran** puts on it." +2. Suspension branch: the scope stop is "raised by either trigger above — a **membership stop** by + the first, a **question stop** by the second." +3. Continue branch: "which findings raise one is that passage's entire, **an already-declined finding + raising no membership trigger and an already-answered question no question trigger**." (Repaired + at Gate-B pass 1, cycle `t57gp3hwu1`, in both copies and in the target text: the earlier wording, + "an already-declined finding and an already-answered question raising none", also read as "a + declined finding raises no trigger at all", which `b11` rejects.) +4. Answers: "an out-of-set finding that opened one is a membership stop as well"; "**Decline is + available only at a membership stop**". + +**What the absorb paragraph states (step 2), in full:** +- `b11`: "A correction that leaves that set stops the loop like any other out-of-scope finding — + except one this cycle has already declined, which is outside the set by that decision and + **raises no membership trigger on that account**, its membership being the one question already + answered — even when it opens no new question at all, and the membership answer ends that + finding's membership hold; what the pass does next is the closure ordering's …" +- `b13`: "A finding that opens a **new structural or contract question** — new meaning not already + answered in this cycle, so an answered question raised again stops nothing — stops the loop and + goes to the user …", with "Novelty overrides ancestry and nothing else: where the finding is also + out of set, both triggers hold and both answers are owed." + +**Comparison (step 3).** +- **Block → source** (a condition the ordering attributes that the paragraph lacks): none. Items 1, + 2 and 4 match `b11`/`b13` term for term, and item 3 now scopes each exemption to its own trigger — + exactly `b11`'s "raises no membership trigger on that account" and `b13`'s already-answered + qualification. Before the repair, item 3 admitted the wider reading Gate-B pass 1 raised. +- **Source → block** (a condition the paragraph states that the ordering does not read): none. + `b11`'s "even when it opens no new question at all" and `b13`'s "size is not the test" are + trigger-internal and the ordering reads the triggers "there"; the both-triggers rule is item 4. + +**Result: equivalent in both directions, per copy (C and W).** The copies were not changed. --- ## Next-state table (Task 13 output) -*Empty until Task 13 runs. Task 13 replaces this entire section.* +**Read against the installed §A in `CLAUDE.md`** (W is byte-identical over §A). Predicates per row: +**SB** a source block stands · **CL** the pass is clean (no in-set Blocker/Major at effective +severity, no scope-stop trigger) · **EL** eligible (clean at or above the floor, or zero findings) · +**K** every closure condition of the cycle other than eligibility and the source block holds (the floor is read in EL, the source block in SB; a reread row gives K as it stands after the repair) · **SUS** which suspensions apply (M membership, Q +question, S clearly-stuck, T two-tell, — none). Next states: **CLOSED** (closing act performed and +completed) · **CONT** (continue branch: next pass on the current artifact, repaired only where a +repair is owed) · **SUSP** (suspended awaiting answers) · **PARKED** (open, not running, no passes, +restarted only by an explicit continue) · **BLOCKED** (source rule's stop; no pass runs). **Oracle:** +a row fails if its answer does not produce a distinct resumable or closed state — the same stop +returning with its reading unconsumed — or if it closes on anything but the stated route. + +| # | Starting state (SB · CL · EL · K · SUS) | Answer / event | Next state (route in §A) | Oracle | +|---|---|---|---|---| +| 1 | no · yes · yes · yes · — | — | CLOSED — clean-completion branch; conditions established first, then the act | pass | +| 2 | no · yes · yes · **no** (a repair owed from an earlier pass) · — | — | CONT, repair first — "an eligible pass with an unmet closure condition lands here" | pass | +| 3 | no · yes · yes · no (repair owed) · T | continue | CONT, repair first — the reading is consumed; a new one needs a post-answer pass | pass | +| 4 | no · yes · yes · no (repair owed) · T | stop | PARKED — "stop parks the cycle" | pass | +| 5 | **yes**, on this read pass (unresolvable profile) · yes · yes · yes once repaired · — | source repaired | the pass is read again → CLOSED (clean, eligible, every condition holds) | pass | +| 6 | **yes**, on this read pass · yes · yes · no (a non-source repair also owed) · — | source repaired | the pass is read again → CONT, repair first | pass | +| 7 | no · yes · **no** (below floor, only a Minor, no trigger) · yes · — | — | CONT, unrevised allowed — "a below-floor clean pass lands here" | pass | +| 8 | no · yes · no (below floor) · yes · T | continue | CONT — the suspension branch takes a below-floor clean pass; its answer continues | pass | +| 9 | no · yes · no (below floor) · yes · T | stop | PARKED | pass | +| 10 | no · yes · yes (**zero findings**, below floor) · yes · — | — | CLOSED — zero-finding eligibility; no suspension can co-occur | pass | +| 11 | no · yes · yes (zero findings) · **no** (an earlier in-set Major undischarged) · — | — | CONT, repair first — the pass-2-clean/Major-open case §A states | pass | +| 12 | no · **no** (an out-of-set finding) · no · — · M | accept | CONT — the finding enters the set; the set change costs a further pass; repair owed only if it is a Blocker/Major | pass | +| 13 | no · no (out-of-set) · no · — · M | decline | CONT — hold discharged; the finding stays outside for the cycle; this pass stays unclean | pass | +| 14 | no · no (a new-question finding) · no · — · Q | the user's decision | CONT on the artifact revised per the decision; membership unchanged | pass | +| 15 | no · no (out-of-set and new question) · no · — · M+Q | only one of the two answers given | SUSP — "resumes only when every answer resumes it" | pass | +| 16 | no · no (out-of-set and new question) · no · — · M+Q | both answers given | CONT | pass | +| 17 | no · no (an in-set Major, repair owed) · no · — · T | continue | CONT, repair first | pass | +| 18 | no · no (in-set Major) · no · — · T | stop | PARKED | pass | +| 19 | no · no (regenerating in-set Majors) · no · — · S | continue | CONT — the surfaced findings' clearly-stuck holds are discharged by the answer; repair first | pass | +| 20 | no · no (regenerating in-set Majors) · no · — · S | stop | PARKED | pass | +| 21 | no · no · no · — · S+T | one answer, continue, carrying every reason | CONT, repair first — one question between the two health readings | pass | +| 22 | no · no · no · — · S+T | one answer, stop | PARKED | pass | +| 23 | PARKED, a membership answer outstanding | explicit continue | SUSP — "that continue … never skips an answer" | pass | +| 24 | PARKED, nothing outstanding | explicit continue | CONT | pass | +| 25 | **yes**, raised before any pass was read | source repaired | CONT — the next pass runs; there is no pass to read again | pass | +| 26 | **yes**, on a read pass · yes · yes · yes · — | source repaired | read again → CLOSED | pass | +| 27 | **yes**, on a read pass · yes · yes · no (repair owed) · — | source repaired | read again → CONT, repair first | pass | +| 28 | **yes**, on a read pass · yes · no (below floor) · yes · — | source repaired | read again → CONT | pass | +| 29 | **yes**, on a read pass · no (in-set Major) · no · — · — | source repaired | read again → CONT, repair first | pass | +| 30 | **yes**, on a read pass that also carried T · yes · yes · yes · T | T answered continue, then source repaired | CONT — the suspension's route runs first; its continue needs a post-answer pass, so the pass is not closed on | pass | +| 31 | **yes**, on a read pass that also carried T · yes · yes · yes · T | T answered stop | PARKED — no reread while the suspension's stop stands | pass | +| 32 | row 1, closing act fails, nothing a condition reads moved | repaired | the act is performed again → CLOSED | pass | +| 33 | row 1, closing act fails; the repair changed the assigned fix set | repaired | the set change costs a further pass (the absorb paragraph) → CONT | pass | +| 33b | row 1 (Gate A), closing act fails; the repair edited the artifact and restored it byte for byte, every condition holding again | repaired | Gate A's condition is current equality, so it holds; the act is performed again → CLOSED | pass | +| 34 | row 1, closing act fails and cannot be repaired | — | PARKED — "surface it and leave the cycle parked" | pass | +| 35 | no · yes · yes · yes · — — Gate-B `full`: spec branch `NO FINDINGS`, quality branch a Minor, no trigger | — | CLOSED — the logical pass is the concatenation, and it is clean | pass | +| 36 | no · no · no · — · — — Gate-B `full`: one branch `NO FINDINGS`, the other an in-set Major | — | CONT, repair first — one branch's clean file never makes a clean pass | pass | +| 37 | no · no · no · — · M — the same out-of-set complaint in both branch files | accept / accept | CONT — both lines in the set | pass | +| 38 | no · no · no · — · M — same | accept / decline | CONT — the accepted line in, the declined line out; the accepted repair obligation stands | pass | +| 39 | no · no · no · — · M — same | decline / accept | CONT — mirror of the row above | pass | +| 40 | no · no · no · — · M — same | decline / decline | CONT — both lines out for the cycle; the pass stays unclean | pass | +| 41 | no · no · no · — · M — same | one line answered, the other not | SUSP — each line owes its own explicit answer | pass | +| 42 | no · yes · yes · yes · — — a finding this cycle declined recurs, no new question, nothing else found | — | CLOSED — b11: it raises no membership trigger; it is outside the set, so not an in-set Blocker/Major | pass | +| 43 | no · no · no · — · Q — a finding this cycle declined recurs and opens a new question | the user's decision | CONT — b11's exception is the membership trigger only; the question trigger reaches it | pass | +| 44 | no · yes · yes · **no** (the assigned fix set changed after it was fixed for this pass and was changed back) · — | — | CONT — the change costs a further pass even though undone; equal endpoints do not discharge it | pass | + +**45 rows, each with one next state; every row passes the oracle.** Revised at Gate-B pass 1 (cycle `t57gp3hwu1`): rows that combined alternative answers are split, the reread rows state every predicate, and three cases are added — a declined finding recurring without and with a new question, and a fix-set change undone before the act. Revised again at pass 2: K excludes eligibility (the floor is read in EL), and the failed-act row is split by which condition input the repair changed. Claim width, as the plan fixes it: the table covers +answer-state transitions **once the predicates producing them are established**; it does not show +how each predicate was derived, nor that these rows cover every reachable combination. + +### Per-condition closure checks (step 4) — one per condition §A states for closure (13) + +| Check | Condition, as §A states it | What is observed | +|---|---|---| +| `close-eligible` | the pass is eligible: clean at or above the derived floor, or zero findings | the pass's validated findings, its effective severities, the derived floor, the pass number | +| `close-floor` | the derived floor, a precondition, discharged by valid logical passes reaching it with the last clean, or by the zero-finding exit | count of valid logical passes; the last one's cleanliness | +| `close-resolve` | every in-set Blocker/Major discharged per finding (repair or validated dismissal), tracked across the cycle, never inferred from a later pass | the per-finding disposition record for every in-set Blocker/Major the cycle raised | +| `close-no-hold` | no hold standing — every surfaced finding's answers given | the answers recorded against every surfaced finding | +| `close-no-question` | every suspension question answered (composition), including a two-tell continue-or-stop | the suspension record of the cycle | +| `close-no-source-block` | no source block stands | each source rule's stop condition (profile resolvable, headers agree, `Story:` readable, counterfactual observable, …) | +| `close-header-during-pass` | a governing `**Story:**` header, or a cited story's profile header, changed **during** the final pass makes it not final — including a change restored before the pass ends | **Observation, not proof:** the named procedure `observe-header-changes` below, run over the pass's own interval — from building its review request to accepting its findings file(s), **not** to the act. It reports `change observed` (the pass is not final), `no change observed`, or `source unreadable` (treated like a change: not final). `no change observed` names what it read and what it cannot see: an edit made and undone without a commit, or by an actor outside the branch's ref moves, leaves no trace. | +| `close-cited-set-at-act` | the final clean pass runs against the **current** cited set | the set named by the governing `**Story:**` headers read at the act, against the set recorded in the final pass's request; unequal → not final. A header changed and restored **after** acceptance does not fail this row or the one above; a profile or fix-set change in that time is `close-profile-fixset`'s. | +| `close-profile-fixset` | a change to a cited story's profile or to the assigned fix set costs a further pass, in either direction and **even when undone**, from the moment the set is fixed for the final pass to the act | **Observation, not proof,** over that longer window, from every input the set definition reads: (a) profile headers — `observe-header-changes` on the cited stories; (b) the scope each governing story or plan assigns — every diff to those files in the window, read for a change to the scope they assign (other edits, such as appended verification records, are not scope changes); (c) membership answers — the cycle's recorded dispositions. Reports `change observed` (a further pass is owed), `no change observed`, or `source unreadable`. **Blind to** an approval given with no file change and no recorded answer, and to uncommitted edits. | +| `close-evidence-revalidated` | each owed evidence entry is revalidated before the commit the closing act produces; a changed entry owes a re-review on it, and the pass that follows closes only on the entry revalidated for it | the entry handed verbatim to the final pass against the entry revalidated at the act; equal, or else a further pass | +| `close-gate-content` | Gate A: the artifact equals the text in the final pass's review request; Gate B: no content condition | Gate A: byte comparison artifact ↔ request text; Gate B: none, stated | +| `close-order` | every condition established first, only then the act | the order of checks and act in the closing record | +| `close-act` | the gate's closing act performed and completed, carrying the records the cycle owes (provenance line, curve, human-exception records in the commit the act uses; Gate A: the reviewed text at the artifact path) | the closing commit and its body | + + +**What these checks are, stated once** (Daniel's decision, 2026-09-26): each row names **what is +observed and from which source**; where a row reports an observation it reports `change observed`, +`no change observed` or `source unreadable`, **never "held"** — silence in a source is not proof that +nothing changed. The checks are **defined and demonstrated in disposable repositories; none is +applied to a real closing act in this change**, because this change's own cycle closes under §5 as at +its base (option 1), not under the ordering it installs. + +**`observe-header-changes`** — the named procedure the header rows use (POSIX `sh`; shellcheck clean): + +```sh +# observe-header-changes ... +# Reports whether any governing header line changed on between two moments. +# Source 1: the branch's first-parent line between the heads it held at those moments, +# merges read against their first parent. +# Source 2: every move of the branch ref in that window (its reflog), each old->new pair compared. +# Prints: "change observed ()" or "no change observed" or "source unreadable". +obs() { + br=$1; t0=$2; t1=$3; shift 3 + pat='^[-+]\*\*(Story|Risk|Security|Validation):\*\*' + moves=$(git reflog show --date=unix --format='%H %gd' "$br" 2>/dev/null) || { echo "source unreadable (reflog)"; return; } + # heads at t0 and t1: newest reflog entry at or before each moment + h0=$(printf '%s\n' "$moves" | awk -v t="$t0" '{split($2,a,"[{}]"); if (a[2]<=t) {print $1; exit}}') + h1=$(printf '%s\n' "$moves" | awk -v t="$t1" '{split($2,a,"[{}]"); if (a[2]<=t) {print $1; exit}}') + [ -n "$h0" ] && [ -n "$h1" ] || { echo "source unreadable (no reflog entry for a moment)"; return; } + if git log --first-parent --diff-merges=first-parent -p "$h0..$h1" -- "$@" | grep -Eq "$pat"; then + echo "change observed (first-parent history)"; return; fi + # every ref move inside (t0, t1], oldest first, as consecutive pairs + prev=$h0 + for h in $(printf '%s\n' "$moves" | awk -v a="$t0" -v b="$t1" '{split($2,x,"[{}]"); if (x[2]>a && x[2]<=b) print NR, $1}' | sort -rn | cut -d' ' -f2); do + if git diff "$prev" "$h" -- "$@" | grep -Eq "$pat"; then echo "change observed (reflog move)"; return; fi + prev=$h + done + echo "no change observed" +} +``` + +**Demonstrated** (Gate-B pass-5 preparation; each case a fresh disposable repository, run under +`sh` and under `dash` with identical results): + +``` +a control, nothing changed no change observed +b A->B->A in two commits change observed (first-parent history) +c A->B->A carried only by merges into main change observed (first-parent history) +d side branch did A->B->A before; main's header never moved no change observed +e B committed, then amended away change observed (reflog move) +f main reset to an existing commit with B, then back change observed (reflog move) +g A->B->A after the window ends (acceptance) no change observed +h edited and restored, never committed no change observed <- the stated blind spot +``` + +### Separate named checks (step 5) + +| Check | What it establishes | +|---|---| +| `logical-pass-validated` | a logical pass was validated across **every** required branch file — both for a `full` Gate-B pass, each file separately against *Accept a pass only when* — before any finding-derived predicate read it | +| `conditions-held-at-act` | every closure condition above held at the moment the closing act was performed, re-established after any failed attempt against the repository as it then stood | --- ## Divergence list (Task 14 output) -*Empty until Task 14 runs. Task 14 replaces this entire section.* +**Sites compared, C against W** — one row per destination block, read off the target's markers +(§A three, §B one, §C one, §D two, §E two, §F sixteen prompt-copy items with 9a and 9 extracted as +one adjacent region, §G one, §H nine). Each region runs from the block's first installed line to its +last, both extracts non-empty, anchors unique in each copy. + +| Site | Parity | +|---|---| +| §A1 | no difference | +| §A2 | no difference | +| §A3 | no difference | +| §B | differs — see below | +| §C | no difference | +| §D e7 | no difference | +| §D pointer | no difference | +| §E Severity | no difference | +| §E answer | no difference | +| §F 1 | no difference | +| §F 2 | no difference | +| §F 3 | no difference | +| §F 4 | no difference | +| §F 5 | no difference | +| §F 6 | no difference | +| §F 7 | no difference | +| §F 8 | no difference | +| §F 7a | no difference | +| §F 8a | no difference | +| §F 8b | differs — see below | +| §F 9a+9 | no difference | +| §F 9b | no difference | +| §F 14 | no difference | +| §F 18 | no difference | +| §G | no difference | +| §H c18 | no difference | +| §H a13 | no difference | +| §H a16 | no difference | +| §H a17–a22 | no difference | +| §H clean signal | no difference | +| §H template | no difference | +| §H cadence | differs — see below | +| §H lens | no difference | +| §H strict list | no difference | + +**Every difference, classified (step 2):** + +| Difference | Bucket | Reason | +|---|---|---| +| §B: C's field-mint parenthetical after the common span | deliberate, stays | inventory passage (b) difference 3; §B stops short of it by design | +| §F 8b: C's kept `` (`docs/prompt-standards.md`, "coverage first, filter later") `` after the block | deliberate, stays | pre-existing C-only citation of a repo-local doc the scaffolded template cannot assume; outside the block | +| §H cadence: the kept remainder of the block's last line wraps differently (`what` / `what the`) | inherited, stays | pre-existing wrap difference in untouched Gate-A text; same words | +| passage (e): C's `e11` rationale paragraph | deliberate, stays | inventory passage (e) difference 2 | +| passage (f): `f5`–`f7` evidence framing | deliberate, stays | inventory passage (f) | +| passage (d): the three-lines sentence wraps differently | inherited, stays | same words; passage (d) is untouched by design, and a diff touching it is a defect | +| `**Findings go to a FILE` opening (`In the field, long finding` / `Long finding lists come back`) | inherited, stays | outside every inventoried and changed site | +| C had no blank line between the Surfacing paragraph and `**Every pass report states`; W had one | **not deliberate — aligned in this task** | C now carries the blank line; the region `**Recognizing "clearly stuck"` … `**Every pass report states` is byte-identical | + +**Not in the deliberate bucket:** `e8` (Task 5 gave W the pronoun; the passage-(e) extract now differs +only by `e11`) and `b3` (step 3: W reads `severity exactly as Mechanics · Severity says`, count 1; +Task 3's pair recorded it). `g4` is gone from C (Task 6), so passage (g) is byte-identical. + +**Step 4, re-run after the alignment:** the same three block-adjacent differences and nothing else. --- ## Completeness sweep (Task 12b output) -*Empty until Task 12b runs. Task 12b replaces this entire section.* +**What was looked for:** any live sentence outside §A that still answers a question the installed +ordering now decides — closure, eligibility, cleanliness, the hold and what discharges it, +composition, the two scope triggers, the suspensions and their answers, the two gates' closing +acts, the duty classification, the source-block branch and its reread routes, the +repeated-dismissal exclusion, and the parked state (the list read off the installed §A, which is +the complete statement). **How:** a grep for closure vocabulary (`final pass`, `clean pass`, `keep +going`, `never iterate`, `resumes`, `close it`, `closes the cycle`, `run more`, `exit the loop`, +`only early exit`, `until clean`, `make the real commit`) over both copies outside §A, then each hit +and its paragraph read; plus reading the sections below whole. + +Sections read, per copy (C; W by the same grep and by parity with C): + +``` +sweep CLAUDE.md §4 work-loop line — closing acts named, §5 cited (item 18 installed) → nothing further +sweep CLAUDE.md §5 HARD FLOOR paragraph — floor arithmetic, set comparison before a clean pass is final → a condition §A cites, not a competing rule +sweep CLAUDE.md §5 derived-floor paragraph — a13/a16/a17 pointers installed → nothing further +sweep CLAUDE.md §5 Named residual, gate-off surface, When these rules bind, Downstream → nothing that answers closure +sweep CLAUDE.md §5 absorb paragraph, clearly-stuck, surfacing, pass-report duties, five tells, two-rules → installed text; resumption points at §A +sweep CLAUDE.md §5 findings-file protocol, Accept a pass only when, Reader, Recovery, What this does not do → validity and recovery rules; "Spent and still incomplete → STOP and surface" is a source-rule stop §A's source-block branch reads +sweep CLAUDE.md Gate A / Gate B sections — broad prompt, clean signal, cadence, coverage instruction installed; "Re-review after every fix" and "A fix that changes specified behaviour updates the spec" are duties, not closure permissions → nothing further +sweep CLAUDE.md Profiles — "That further pass must itself be clean and every other closure duty must be satisfied … not a licence to close on the next one" and "the final clean pass runs against the current set" → conditions §A cites; consistent +sweep CLAUDE.md Mechanics · Severity, baseSha / Finishing the cycle, provenance line, curve, human exception, Timeout → installed text or record rules; no closure permission left +sweep codex-gate.sh eight gate reminders, both channels — seven replaced by §F; the docs-only notice (line 918) says Gate B does not apply to a docs-only commit and points at Gate A → states no closure permission; exclusion confirmed +``` + +**Found: nothing.** No live sentence in `CLAUDE.md`, `plugins/dev-workflow/commands/workflow-init.md` +or `plugins/dev-workflow/hooks/codex-gate.sh` was found still answering a question the ordering +decides, beyond the sites §F replaces. The sweep is a reader's judgement and nothing checks its +coverage. --- ## Prompt-standards result (Task 15 step 4b output) -*Empty until Task 15 runs. It replaces this entire section, one line per checklist item.* +**Subject set, read once and referred to by every line below:** (1) the §A–§H blocks as installed in +`CLAUDE.md`; (2) the same blocks as installed in `plugins/dev-workflow/commands/workflow-init.md` +(byte-identical to (1) except the classified divergences in the Task 14 list); (3) ten hook prompt +bodies in `plugins/dev-workflow/hooks/codex-gate.sh` — the seven `additionalContext` bodies items +10–13 and 15–17 replace (lines 908, 933, 945, 947, 956, 967, 973) and the three `systemMessage` +bodies items 12, 15 and 16 replace (`✓ Codex Gate B hook checks passed (…)`, `⚠ Codex Gate B: no +recorded fingerprint`, `⚠ Codex Gate B: cannot confirm reviewed content`). Checked against +`docs/prompt-standards.md` as it stands at this commit; a reader check, as the plan says. + +1. **Target model named — PASS.** (2) sits under W's `Target model: Claude via Claude Code` line; + (3) sits under the hook's `# Target model:` comments (lines 348, 862); (1) is this repository's + own `CLAUDE.md`, read by Claude via Claude Code, and adds no second, conflicting model claim. +2. **Success criteria explicit — PASS.** §A states closure as a checkable conjunction — eligibility + (clean at or above the floor, or zero findings) plus every closure condition plus a completed + closing act — and each hook body names the observed state (`hook checks passed`, `no fingerprint + is recorded`, `cannot confirm`). +3. **Stop conditions defined — PASS.** §A's source-block branch, the three suspensions, the parked + state and the unrepairable-act route; §B's two scope triggers; the hook bodies defer every next + step to the policy's closure ordering rather than issuing one. +4. **Output format with an example — PASS.** The changed text adds no new output format; the + findings-file format and its example block (`MAJOR | high | …`, `NO FINDINGS`) are unchanged + and the template's clean sentence still names the exact body line and terminator. +5. **Structured sections — PASS.** §A is three bolded-lead paragraphs placed as a unit before the + absorb paragraph; each §B–§H replacement stays inside the section and paragraph it replaced. +6. **Rules carry their why — PASS.** Each constraint in §A carries its reason clause (e.g. the + read-once rule, "so a pass that reached the act has already been classified"; zero-finding + eligibility, "because a floor buys further looks at an artifact that keeps yielding findings"); + §E states "Two reasons for the split"; each hook body states why it defers ("this reminder + decides none of it", "which restores no passes"). +7. **No contradictions with CLAUDE.md / AGENTS.md — PASS.** Superseded sentences are replaced in + the same change (§F, 23 sites), and Task 12b's completeness sweep found no remaining sentence + answering what the ordering decides. +8. **Token-lean — PASS, with the observation stated.** §A1 is long; it restates no rule owned + elsewhere — the scope triggers, fix set, severity rules and closure preconditions are cited + ("read there and not redefined here"), and every replaced entry point now points at the + ordering instead of carrying a copy. The hook bodies replace enumerations with a pointer. +9. **Positive instructions — PASS.** Instructions are phrased as what to do (answer, repair, + continue, park, perform the act); the prohibitions that remain ("no pass is credited…", + "nothing here turns one answer into another") draw a boundary that a positive restatement would + lose, the exception the item names. +10. **Diagnostic states name their causes — PASS.** The no-fingerprint body lists its three causes + (no review ran, the fingerprint could not be written or read back, a non-`WIP` commit attempt + cleared it) with the check and fix for the storage cause; the stale-fingerprint body lists + worktree/index change, staging only, upgraded format and failed compute/store, each with its + remedy and the machinery checks. +11. **Enforcement claims name their mechanism — PASS.** §G says it is "not a checker" and bounds + what a reader can detect; §A3 states "no rule in this section reaches it" for content the + final review request did not select; the Gate-B hook body says what the hook checked and no + more (`hook checks passed`), and the terse channel dropped `satisfied`. +12. **Calibrated emphasis — PASS.** The changed text adds no new MUST/CRITICAL; the one `MUST` + kept in the stale-fingerprint body ("MUST re-review after every fix") is the pre-existing §5 + gate language the item names as a deliberate exception. + +**Result: all twelve pass. No repair was made.** --- ## Fragment sweep (Task 0 output) -*Empty until Task 0 runs. Task 0 replaces this entire section: the revision the sweep ran at, one -line per fragment-table row with its three results, and the reading result for `F1`, `F2` and `F3`.* +**Ran at:** base `d26de4b40655a18c44c5cb3a3d8fd362943c8495` — both prompt copies equal to the base; +only this plan's table had gained rows P19–P22. **Mechanical checks:** the count in each copy the +row claims (a `grep -F` hit is one line, so a count of 1 is single-line and unique), and absence from +every fenced block of the target text with whitespace normalized. **Post-edit-region check** — an +OLD row must overlap wording its item removes; a kept row (P19–P22) must lie outside every +replacement: done by reading each row's live line against its item's block. Every OLD row runs into +wording its block changes; P19–P22 sit wholly in kept text beside a replacement. + +- `P1 no fragment (P1 by design) *(none — counterfactual ABSENT, presence only)*` — +- `P2 C=1 W=1 claims=CW single-line+unique=ok in-replacement-block=no` — `scope the approved story or plan assigns to this cycle, plus repair ob…` +- `P3 C=1 W=1 claims=CW single-line+unique=ok in-replacement-block=no` — `the moment the user says whether the set now includes it…` +- `P4 C=1 W=1 claims=CW single-line+unique=ok in-replacement-block=no` — `a Blocker/Major-free pass below the floor…` +- `P5 C=1 W=0 claims=C single-line+unique=ok in-replacement-block=no` — `not discretionary** — you report…` +- `P5w C=0 W=1 claims=W single-line+unique=ok in-replacement-block=no` — `not discretionary** — report the…` +- `P6 C=1 W=1 claims=CW single-line+unique=ok in-replacement-block=no` — `is not settled here, and this change does not settle it…` +- `P7 C=1 W=1 claims=CW single-line+unique=ok in-replacement-block=no` — `These records are one contract…` +- `P8 C=1 W=1 claims=CW single-line+unique=ok in-replacement-block=no` — `the resolve rule is not waived, no pass is credited as…` +- `P9 C=1 W=1 claims=CW single-line+unique=ok in-replacement-block=no` — `Every other rule stated here about how a cycle closes…` +- `P10 C=1 W=1 claims=CW single-line+unique=ok in-replacement-block=no` — `fix Blocker/Major after each…` +- `P11 C=1 W=1 claims=CW single-line+unique=ok in-replacement-block=no` — `final pass must be clean…` +- `P12 C=1 W=1 claims=CW single-line+unique=ok in-replacement-block=no` — `when a pass is clean…` +- `P13 C=1 W=1 claims=CW single-line+unique=ok in-replacement-block=no` — `clean pass is the single body line…` +- `P14 C=1 W=1 claims=CW single-line+unique=ok in-replacement-block=no` — `Each pass: validate, revise, re-run…` +- `P15 C=1 W=1 claims=CW single-line+unique=ok in-replacement-block=no` — `The Blocker/Major filter, the file-first findings protocol…` +- `P16 C=1 W=1 claims=CW single-line+unique=ok in-replacement-block=no` — `and the nonce duties at their strictest — the cycle…` +- `P17 C=1 W=1 claims=CW single-line+unique=ok in-replacement-block=no` — `Hook text is out of scope here…` +- `P18 C=1 W=1 claims=CW single-line+unique=ok in-replacement-block=no` — `execute → tests green → Gate B → commit…` +- `P19 C=1 W=1 claims=CW single-line+unique=ok in-replacement-block=no` — `where the cycle's own closure rules are satisfied.…` +- `P20 C=1 W=1 claims=CW single-line+unique=ok in-replacement-block=no` — `Nothing here writes the floor knob: it stays the user's, never written…` +- `P21 C=1 W=1 claims=CW single-line+unique=ok in-replacement-block=no` — `Several records…` +- `P22 C=1 W=1 claims=CW single-line+unique=ok in-replacement-block=no` — `` nothing that any mandatory rule in this file or in `AGENTS.md` require… `` +- `F1 C=1 W=1 claims=CW single-line+unique=ok in-replacement-block=no` — `nor resets your pass counters. A pre-review snapshot named anything el…` +- `F2 C=1 W=1 claims=CW single-line+unique=ok in-replacement-block=no` — `` `NO FINDINGS` if clean" in `additionalContext`, with the same one-line… `` +- `F3 C=1 W=1 claims=CW single-line+unique=ok in-replacement-block=no` — `**Majors are recorded as well as Findings and Blockers**, because the…` +- `F4 C=1 W=1 claims=CW single-line+unique=ok in-replacement-block=no` — `neither a human's assent nor this record…` +- `F5 C=1 W=1 claims=CW single-line+unique=ok in-replacement-block=no` — `**Finishing the cycle:** after the final clean pass, close it with…` +- `F6 C=1 W=1 claims=CW single-line+unique=ok in-replacement-block=no` — `reviews the TEXT you pass, not the git tree). Use ONE broad prompt, re…` +- `F7 C=1 W=1 claims=CW single-line+unique=ok in-replacement-block=no` — `**Which commit:** an ungated change records it in that commit; a Gate-…` +- `F7b C=1 W=1 claims=CW single-line+unique=ok in-replacement-block=no` — `restated by the closing amend…` +- `F8 C=1 W=1 claims=CW single-line+unique=ok in-replacement-block=no` — `` snapshot by amend — a non-`WIP` commit reads to the hook as the cycle… `` +- `F9 C=1 W=1 claims=CW single-line+unique=ok in-replacement-block=no` — `taken from the provenance line and the curve, which must agree. A Gate…` +- `F10 C=1 W=1 claims=CW single-line+unique=ok in-replacement-block=no` — `(Blocker/Major only), derived from the cited story's profile.**…` +- `F11 C=1 W=1 claims=CW single-line+unique=ok in-replacement-block=no` — `with severity and confidence — you filter to Blocker/Major downstream,…` +- `F12 C=1 W=1 claims=CW single-line+unique=ok in-replacement-block=no` — `before the cycle-closing amend…` +- `F13 C=1 W=1 claims=CW single-line+unique=ok in-replacement-block=no` — `cannot name both, the finding is Minor or below: collect, never iterat…` +- `F14 C=1 W=1 claims=CW single-line+unique=ok in-replacement-block=no` — `profile sits still. If revalidation changes the entry, the clean pass…` + +**Reading result for F1, F2, F3** (no §F line range, so checked by reading): F1 runs from the kept +`nor resets your pass counters.` into `reads as a`, which item 1 rewrites to `reads to the hook as` — +gone after install. F2 lies inside the sentence item 2 replaces whole. F3 runs into `because the +severity rule moves the`, which item 3 rewrites — gone after install. All three usable. + +**Rows added by Task 0:** P19 (`a12`), P20 (`a14`), P21 (`h6`), P22 (`h18`) — kept conditions that +share a line with changed text, recorded as `cond` rows in `.context/loop-rule-untouched`. + +**Inherited drift found by step 3, beyond the inventory's expected divergences** — recorded here and +carried to Task 14; no edit of this change is aimed at it: (1) passage (d)'s three-lines sentence +wraps differently in C and W, same words; (2) C has no blank line between the Surfacing paragraph and +`**Every pass report states`, W has one; (3) the opening of the `**Findings go to a FILE` paragraph +differs (`In the field, long finding` in C, `Long finding lists come back` in W). --- @@ -3782,3 +4239,483 @@ with the other five.* `presence` at its destination — and both are required for it to count as observed. **Every one of the six shapes goes into the closing evidence entry**; naming only pairs and presence leaves the absences, preservations and span results run but unrecorded.* + +### Task 1 + +§A is add-only: row P1, no OLD half. Three presence fragments, one per paragraph, each a single line +and unique in each copy: + +``` +presence §A1 `How a cycle ends — one ordering, stated here and referenced everywhere else` worktree=1 parent=0 C +presence §A1 `How a cycle ends — one ordering, stated here and referenced everywhere else` worktree=1 parent=0 W +presence §A2 `**Gate A's content condition, and its closing act.** These are what Gate A adds to the conditions` worktree=1 parent=0 C +presence §A2 `**Gate A's content condition, and its closing act.** These are what Gate A adds to the conditions` worktree=1 parent=0 W +presence §A3 `**Gate B's content condition, and its closing act.** These are what Gate B adds to the conditions` worktree=1 parent=0 C +presence §A3 `**Gate B's content condition, and its closing act.** These are what Gate B adds to the conditions` worktree=1 parent=0 W +``` + +Parity: the `**How a cycle ends` … `**What a loop absorbs` region extracted from both copies, both +non-empty, `diff` empty. The blocks were installed with the target's own line breaks, identical in +both copies. + +### Task 3 + +Passage (b) replaced whole by §B's common span, installed with the target's line breaks in both +copies; C keeps its field-mint parenthetical after the span. Step 1: P2 and P3 counted 1 in both +copies before the install. Step 1b appended P23–P38 (seven replaced OLD halves, nine carried +preservation fragments). + +``` +pair b3 `Blocker/Major resolve, Minor/Nit collect` (P23) `severity exactly as Mechanics · Severity says. Ancestry decides` 0 1 1 0 C +pair b3 (same) 0 1 1 0 W +pair b7 P2 `union of the scope every approved story or plan governing this change assigns to this cycle,` 0 1 1 0 C +pair b7 (same) 0 1 1 0 W +pair b8 P24 `A finding is in-set when repairing it stays inside **the assigned fix set as` 0 1 1 0 C +pair b8 (same) 0 1 1 0 W +pair b11 P25 `finding** — except one this cycle has already declined, which is outside the set by that` 0 1 1 0 C +pair b11 (same) 0 1 1 0 W +pair b12 P3 `which resumes only when every answer outstanding on that surface has been given.` 0 1 1 0 C +pair b12 (same) 0 1 1 0 W +pair b13 P26 `new meaning not already` 0 1 1 0 C +pair b13 (same) 0 1 1 0 W +pair b16 P27 `**Novelty overrides ancestry and nothing else:` 0 1 1 0 C +pair b16 (same) 0 1 1 0 W +pair b17 P28 `it is a **suspension**` 0 1 1 0 C +pair b17 (same) 0 1 1 0 W +pair b18 P29 `rule all stand, and what the answer does is stated there` 0 1 1 0 C +pair b18 (same) 0 1 1 0 W +presence closing-time set-change rule `**A change to this set costs the cycle at least one further pass.**` 1 0 C +presence closing-time set-change rule (same) 1 0 W +preservation b1 P30 1 1 C · 1 1 W +preservation b2 P31 1 1 C · 1 1 W +preservation b4 P32 1 1 C · 1 1 W +preservation b5 P33 1 1 C · 1 1 W +preservation b6 P34 1 1 C · 1 1 W +preservation b9 P35 1 1 C · 1 1 W +preservation b10 P36 1 1 C · 1 1 W +preservation b14 P37 1 1 C · 1 1 W +preservation b15 P38 1 1 C · 1 1 W +``` + +Pair columns are old/worktree old/parent new/worktree new/parent; preservation columns are parent +worktree. Parity (`**What a loop absorbs` … `**Recognizing "clearly stuck"`): the only difference is +C's field-mint parenthetical. Walk: the installed passage is §B's block verbatim; `b3` now reads +`Mechanics · Severity` in both copies, so W's old `the severity rule` is gone. + +### Task 4 + +§C's block installed over `So this exit needs three things` … `keeps looping.` in both copies, with +the target's line breaks; the kept prefix (`c1`–`c3`) and the Surfacing paragraph (Task 7's) are +untouched. Step 1 re-confirmed P4 at 1 in both copies and appended P39–P52. The precedence clause +counts exactly 1 per copy, inside §A (C 529, W 736). Pair columns old/worktree old/parent +new/worktree new/parent; the others worktree parent. Every line holds for C and for W. + +``` +pair c4 P42 `a missing one means only that *this* exit does not apply` 0 1 1 0 C,W +pair c8 P46 `a recurrence failing them being an ordinary fresh finding` 0 1 1 0 C,W +pair c14 (suspend) P4 `floor the pass **suspends**, the clean pass having failed eligibility.` 0 1 1 0 C,W +presence c14 (continue, add-only) `A below-floor clean pass lands here on the same` 1 0 C,W +absence c9 P47 0 1 C,W · presence §A `**A clean completion takes precedence over this exit**` 1 0 C,W +absence c10 P49 0 1 C,W · presence §A `pass **at or above the floor** has satisfied the clean-final-pass rule — collect the Minors and` 1 0 C,W +absence c11 P50 0 1 C,W · presence §A `Nits and close — and reporting "will not converge" on a converged loop is a false report. Below the` 1 0 C,W +absence c12 P51 0 1 C,W · presence §A `**Eligibility is exactly this and nothing more: a clean pass at or above the derived floor, or a` 1 0 C,W +absence c13 P52 0 1 C,W · presence §A `zero-finding pass.** It is a property of the pass.` 1 0 C,W +preservation c1 P39 · c2 P40 · c3 P41 (kept prefix) 1 1 C,W +preservation c5 P43 · c6 P44 · c7 P45 (carried) 1 1 C,W +preservation plateau rationale P48 (stays, no id) 1 1 C,W +``` + +Parity over `**Recognizing "clearly stuck"` … `**Surfacing does not`: no difference. + +### Task 5 + +§D's `e7` sentence installed over the live threshold sentence in both copies — W now carries C's +`you report the tells`, so the `e8` divergence is gone. §D's pointer paragraph added at the end of +passage (e): in C after the C-only `e11` paragraph, in W after `e10`, each followed by a blank line +before `**The two rules above`. Step 1 re-confirmed P5 (C) and P5w (W) at 1 and appended P53–P61. + +``` +pair e7 P5 `clean-completion branch of the closure ordering, which outranks it **by taking the pass to the` 0 1 1 0 C +pair e7 P5w (same NEW) 0 1 1 0 W +presence e8 alignment `you report the tells` 1 0 W +presence §D pointer `**What the answer does** is the closure ordering's, which is where this stop's place among the` 1 0 C,W +preservation e1 P53 · e2 P54 · e3 P55 · e4 P56 · e5 P57 · e6 P58 (kept) 1 1 C,W +preservation e9 P59 (carried) 1 1 C,W +preservation e10 P60 (kept, after the block) 1 1 C,W +preservation e11 P61 (kept, C only) 1 1 C · 0 0 W +``` + +Parity over `Those three lines expose` … `**The two rules above`: the only difference is C's `e11` +paragraph. + +### Task 6 + +§E's two blocks installed in both copies: the Severity bullet's first sentence, and the answer +paragraph in place of `**How this demotion bears…` — in C together with the `g4` ownership +sentence. Step 1 observed C `g1/g2: 1 g4: 1`, W `g1/g2: 1 g4: 0`, and appended P62–P65. + +``` +pair Severity resolve duty P62 `must resolve, **for every finding in the assigned fix set as the absorb paragraph computes it**.` 0 1 1 0 C,W +pair g1 P6 `**The demotion changes what a cycle must resolve, never what it observes.**` 0 1 1 0 C,W +absence g2 P63 1 0 C,W +absence g3 P64 1 0 C,W +absence g4 P65 1 0 C · 0 0 W (W never carried it) +``` + +Absence columns are parent worktree. Step 4: the story path counts 0 in `CLAUDE.md`. Step 5: parity +over `- **Severity:**` … `- **Tool routing:` — no difference. + +### Task 7 + +§G's block and §H's nine blocks installed in both copies, one site at a time; step 1 re-confirmed +P7–P16 at 1 in each copy (20 counts) and step 1b appended P66–P93. Blocks inside an indented list +take the host's two-space indent (§G, the Gate-A clean signal, the Gate-A cadence); the template +sentence takes the host's `> ` prefix. The lens paragraph is replaced whole, its first sentence +being restated by §H's block. The strict-reading block's first two lines are joined at +`owed, the curve duty owed,` so the carried `i7` fragment sits on one line (words unchanged). + +``` +pair §G P7 `**These rules and records are one contract, and a partial adoption breaks it.**` 0 1 1 0 C,W +pair c18 P8 `**A pass is credited clean or not on its own findings**` 0 1 1 0 C,W +pair a13 P9 `Every other rule stated **in this paragraph**` 0 1 1 0 C,W +pair a16 P10 `resolve Blocker/Major after each as` 0 1 1 0 C,W +pair a17–a22 pointer P11 `What a clean final pass and the zero-finding early exit mean for closing` 0 1 1 0 C,W +pair Gate-A clean signal P12 `and no scope-stop trigger** is clean too` 0 1 1 0 C,W +pair template clean sentence P13 `A **clean findings file** is the single body line` 0 1 1 0 C,W +pair Gate-A cadence P14 `revise **where a repair is required**` 0 1 1 0 C,W +pair lens unchanged-list P15 `**the lens sets** leave every other` 0 1 1 0 C,W +pair strict-reading list P16 `every suspension binding, since` 0 1 1 0 C,W +pair c16 P67 `the cycle and the new hold still open*` 0 1 1 0 C,W +pair c17 P68 `the resolve rule stands over the finding exactly as` 0 1 1 0 C,W +pair c19 P69 `its answers are given, and **what the answer does is the closure ordering's**.` 0 1 1 0 C,W +pair c20 P70 `exit in competition with the rule that every **in-set** Blocker and Major resolves, and then` 0 1 1 0 C,W +absence a13 first sentence P71 1 0 C,W +absence a17 P11 1 0 C,W · presence §A `discharged by the count of valid logical passes reaching it with the last of them` 1 0 C,W +absence a18 P73 1 0 C,W · presence §A `runs another pass on the **current** artifact, revised where the severity and scope rules require a` 1 0 C,W +absence a19 P74 1 0 C,W · presence §A `A pass with **zero** findings is clean` 1 0 C,W +absence a20 P75 1 0 C,W · presence §A `has already given what those looks were for; don't manufacture findings to` 1 0 C,W +presence §G membership test (add-only) `**Membership is decided by a test a reader can apply to the text in front of them, with` 1 0 C,W +presence strict tail 1 `starting rules that cannot be established cannot be read as having waived an open hold` 1 0 C,W +presence strict tail 2 `repeated-dismissal cleanliness exclusion unavailable` 1 0 C,W +presence strict tail 3 `the parked state binding after a closing act that cannot` 1 0 C,W +presence strict tail 4 `pass-cost rule this change ships owed rather than waived` 1 0 C,W +preservation carried c15 P66 · a15 P72 · a21 P76 · a22 P77 · i4 P78 · i5 P79 · i6 P80 · i7 P81 · i8 P82 1 1 C,W +preservation kept, passage (i) i1 P83 · i2 P84 · i3 P85 · i9 P86 · i10 P87 · i11 P88 · i12 P89 · i13 P90 · i14 P91 · i15 P92 · i16 P93 1 1 C,W +``` + +Pair columns old/worktree old/parent new/worktree new/parent; the rest parent worktree for absence +and preservation, worktree parent for presence. Parity per site (first installed line to last): §G, +c18/surfacing, a13+a16+a17, clean signal, template sentence, lens, strict list — no difference. +Cadence: the block is identical; the only difference is the kept remainder of its last line +(`settle mechanically what` in C, `what the` in W), an inherited wrap difference. + +### Task 8 + +§F items 1–9, 7a, 8a, 8b, 9a, 9b installed in both copies. Step 1: every quoted live sentence located +by its fragment row; F1–F14 and F7b counted 1 in each copy before the install; P94 (`a1`) and P95 +(`h3`) appended. Live sentences that wrap across lines, so a whole-sentence `grep -F` finds nothing: +all fourteen. Items 1, 4, 6, 7a, 8, 8b, 9a, 9 and 9b start mid-line and are joined to the kept text +before them; items 2, 3, 5, 7 and 8a start on their own line. W's item-2 sentence had no +`You filter to Blocker/Major, Codex never does.` line; the replacement covers both copies' forms. + +Pair columns old/worktree old/parent new/worktree new/parent; every line holds for C and for W. + +``` +pair F1 0 1 1 0 C +pair F1 0 1 1 0 W +pair F2 0 1 1 0 C +pair F2 0 1 1 0 W +pair F3 0 1 1 0 C +pair F3 0 1 1 0 W +pair F4 0 1 1 0 C +pair F4 0 1 1 0 W +pair F5 0 1 1 0 C +pair F5 0 1 1 0 W +pair F6 0 1 1 0 C +pair F6 0 1 1 0 W +pair F7 0 1 1 0 C +pair F7 0 1 1 0 W +pair F7b 0 1 1 0 C +pair F7b 0 1 1 0 W +pair F8 0 1 1 0 C +pair F8 0 1 1 0 W +pair F9 0 1 1 0 C +pair F9 0 1 1 0 W +pair F10 0 1 1 0 C +pair F10 0 1 1 0 W +pair F11 0 1 1 0 C +pair F11 0 1 1 0 W +pair F12 0 1 1 0 C +pair F12 0 1 1 0 W +pair F13 0 1 1 0 C +pair F13 0 1 1 0 W +pair F14 0 1 1 0 C +pair F14 0 1 1 0 W +``` + +NEW fragments: F1 `reads to the hook as a real commit: the hook treats the` · F2 `` `NO FINDINGS` only when +the branch found none `` · F3 `because the three series are read` · F4 `**neither a human's general +assent nor this record**` · F5 `**when a Gate-B cycle's closing act is performed is the closure` · F6 +`Use ONE broad prompt: **its review question stays the same every pass` · F7 `a Gate-A cycle in the +commit its` · F7b `restated by the commit its closing act` · F8 `the edit must end up **in the content +the next review reads**` · F9 `A Gate-A cycle has such a commit only once its own closing commit +exists` · F10 `passes per run (a clean final pass` · F11 `**you filter to Blocker/Major for what must be +repaired and read` · F12 `before the commit its closing act` · F13 `the finding is Minor or below: +collect; **its severity buys no` · F14 `the pass was read against an entry that no longer stands`. + +``` +preservation a1 P94 1 1 C,W +preservation h3 P95 1 1 C,W +``` + +New/worktree total: 15 in C, 15 in W — fifteen changed clauses from fourteen items. Parity per item, +first installed line to last: thirteen sites equal; item 8b's block is identical and its last line +differs only in C's kept parenthetical (`` (`docs/prompt-standards.md`, "coverage first, filter +later") ``), an inherited C-only divergence. + +### Task 9 + +§F items 14 and 18 installed in both copies. Item 14's live sentence wraps after `here`; the block +replaces `Hook text is out of scope here by decision;`, and the kept clause after the semicolon +becomes its own sentence (`What makes that tolerable is the precedence rule above plus the hook +exiting 0 on every branch, …`). Item 18 replaces the whole one-line work-loop sentence. + +``` +pair item 14 P17 `it is not a blanket exemption for hook text**` 0 1 1 0 C,W +pair item 18 P18 `Gate B → Gate-B closing act**` 0 1 1 0 C,W +``` + +Step 4: the residual still stands — the paragraph's first sentence still says the hook's messages +state its own threshold as an obligation at a floor of 1, and the new sentence names that +overstatement as out of scope by decision. Parity: both sites equal. + +### Task 10 + +Hazard probe: every §F block for items 10–13 and 15–17 (both channels for 12, 15, 16) printed +`clean`, read as data from the spec. `note "` hits: thirteen; the seven gate reminders are lines +908, 933, 945, 947, 956, 967, 973, and the docs-only notice (918) is untouched. Step 2b appended +P96–P103 from the live strings before the install. Step 4: seven zero-context hunks, one per +reminder line; every removed and added line read, and each is inside its `note` string — no +control flow, counter, fingerprint or routing change. `shellcheck --shell=sh` exit 0. + +``` +pair item 10 honesty claim P96 `Gate A has no content check in this hook; what the gate itself requires` 0 1 1 0 H +pair item 10 tail P97 `a further pass being one of its answers and not the only one` 0 1 1 0 H +pair item 11 P98 `Proceed only once this Gate-A cycle has closed under` 0 1 1 0 H +pair item 12 P99 `commit only if your final pass was clean and every other closure condition holds` 0 1 1 0 H +pair item 13 P100 `Use this commit as the review range; whether this cycle runs a review now` 0 1 1 0 H +pair item 15 P101 `Codex gate state: no fingerprint is recorded for this cycle` 0 1 1 0 H +pair item 16 P102 `A fresh Gate-B pass is the complete remedy for the staging and post-upgrade cases too` 0 1 1 0 H +pair item 17 P103 `skip rule decides only whether a cycle runs at all` 0 1 1 0 H +``` + +H = `plugins/dev-workflow/hooks/codex-gate.sh`. The suite is red between this commit and Task 11's, +as the plan states; the battery is not run here. + +### Task 11 + +Swept `plugins/dev-workflow/hooks/codex-gate.test.sh` with the three step-1 locators. Mapping used, +so each assertion tests the same hook state as before: `grep -q 'not satisfied'` matched exactly the +two old STOP messages and now reads `grep -q 'Codex gate state:'`, which matches exactly the two new +ones (items 15, 16); `grep -q 'Gate B satisfied'` now reads `grep -q 'Gate B hook checks passed'` +(item 12's terse channel); positive and negative `grep -q 'STOP'` checks now read +`'Codex gate state:'`. The three `expected_ctx` and three `expected_msg` assignments take items 16, +12 and 15 whole, with `$policy` rendered as `this project's review policy` and the counters the +fixture sets. Labels and comments now name the observed hook state. + +Suite: `HOOK_SH=sh sh …test.sh` exit 0 and `HOOK_SH=dash dash …test.sh` exit 0, both `all passed`. +`shellcheck --shell=sh --exclude=SC2015` exit 0. Step 5: `Gate B satisfied|Gate B not +satisfied|Gate A satisfied` counts 0. Remaining case-insensitive `satisfied` hits, disposed of: line +1352 and line 1536 are comments using the plain English verb about test rows, not a gate verdict; +line 1971 is the fixture for the hook's unknown-tool note (`… satisfied count as covering them`), +a message this change does not replace. + +Sites examined and what each became: + +``` +sweep codex-gate.test.sh 173 `# 2. Below floor (1/3) -> NOT satisfied yet; reaching floor (3/3) -> satisfied` → `# 2. Below floor (1/3) -> below-floor reminder; reaching floor (3/3) -> hook checks passed` +sweep codex-gate.test.sh 179 `printf '%s' "$out" | grep -q 'Gate B satisfied' && pass "3/3 passes, unchanged tree -> satisfied" || fail "3/3` → `printf '%s' "$out" | grep -q 'Gate B hook checks passed' && pass "3/3 passes, unchanged tree -> hook checks pa` +sweep codex-gate.test.sh 184 `# intent "a change means Gate B is not satisfied" is asserted at the BEHAVIOR level.)` → `# intent "a change means the hook cannot confirm the reviewed content" is asserted at the BEHAVIOR level.)` +sweep codex-gate.test.sh 189 `printf '%s' "$out" | grep -q 'not satisfied' && pass "Edit-tool change -> not satisfied" || fail "Edit-tool ch` → `printf '%s' "$out" | grep -q 'Codex gate state:' && pass "Edit-tool change -> gate-state reminder" || fail "Ed` +sweep codex-gate.test.sh 197 `printf '%s' "$out" | grep -q 'Gate B satisfied' && pass "setup: satisfied before bash edit" || fail "setup: sa` → `printf '%s' "$out" | grep -q 'Gate B hook checks passed' && pass "setup: hook checks passed before bash edit" ` +sweep codex-gate.test.sh 200 `printf '%s' "$out" | grep -q 'not satisfied' && pass "bash-modified file after review -> NOT satisfied (Findin` → `printf '%s' "$out" | grep -q 'Codex gate state:' && pass "bash-modified file after review -> gate-state remind` +sweep codex-gate.test.sh 206 `printf '%s' "$(commitpre)" | grep -q 'Gate B satisfied' && pass "setup: satisfied on clean tree" || fail "setu` → `printf '%s' "$(commitpre)" | grep -q 'Gate B hook checks passed' && pass "setup: hook checks passed on clean t` +sweep codex-gate.test.sh 209 `printf '%s' "$out" | grep -q 'not satisfied' && pass "untracked new file after review -> NOT satisfied" || fai` → `printf '%s' "$out" | grep -q 'Codex gate state:' && pass "untracked new file after review -> gate-state remind` +sweep codex-gate.test.sh 215 `printf '%s' "$(commitpre)" | grep -q 'Gate B satisfied' && pass "setup: satisfied with untracked file present"` → `printf '%s' "$(commitpre)" | grep -q 'Gate B hook checks passed' && pass "setup: hook checks passed with untra` +sweep codex-gate.test.sh 217 `printf '%s' "$(commitpre)" | grep -q 'not satisfied' && pass "edited untracked file -> NOT satisfied" || fail ` → `printf '%s' "$(commitpre)" | grep -q 'Codex gate state:' && pass "edited untracked file -> gate-state reminder` +sweep codex-gate.test.sh 223 `printf '%s' "$(commitpre)" | grep -q 'not satisfied' && pass "new untracked dir -> NOT satisfied" || fail "new` → `printf '%s' "$(commitpre)" | grep -q 'Codex gate state:' && pass "new untracked dir -> gate-state reminder" ||` +sweep codex-gate.test.sh 226 `printf '%s' "$(commitpre)" | grep -q 'not satisfied' && pass "edited file in untracked dir -> NOT satisfied" |` → `printf '%s' "$(commitpre)" | grep -q 'Codex gate state:' && pass "edited file in untracked dir -> gate-state r` +sweep codex-gate.test.sh 235 `printf '%s' "$(commitpre)" | grep -q 'not satisfied' && pass "new exotic-path untracked file -> NOT satisfied"` → `printf '%s' "$(commitpre)" | grep -q 'Codex gate state:' && pass "new exotic-path untracked file -> gate-state` +sweep codex-gate.test.sh 238 `printf '%s' "$(commitpre)" | grep -q 'not satisfied' && pass "edited exotic-path untracked file -> NOT satisfi` → `printf '%s' "$(commitpre)" | grep -q 'Codex gate state:' && pass "edited exotic-path untracked file -> gate-st` +sweep codex-gate.test.sh 247 `printf '%s' "$(commitpre)" | grep -q 'not satisfied' && pass "new untracked symlink -> NOT satisfied" || fail ` → `printf '%s' "$(commitpre)" | grep -q 'Codex gate state:' && pass "new untracked symlink -> gate-state reminder` +sweep codex-gate.test.sh 250 `printf '%s' "$(commitpre)" | grep -q 'not satisfied' && pass "retargeted untracked symlink -> NOT satisfied" |` → `printf '%s' "$(commitpre)" | grep -q 'Codex gate state:' && pass "retargeted untracked symlink -> gate-state r` +sweep codex-gate.test.sh 279 `# 3d. Reverting the tree back to the reviewed content -> satisfied again` → `# 3d. Reverting the tree back to the reviewed content -> hook checks pass again` +sweep codex-gate.test.sh 282 `printf '%s' "$out" | grep -q 'Gate B satisfied' && pass "revert to reviewed tree -> satisfied again" || fail "` → `printf '%s' "$out" | grep -q 'Gate B hook checks passed' && pass "revert to reviewed tree -> hook checks pass ` +sweep codex-gate.test.sh 287 `printf '%s' "$out" | grep -q 'Gate B satisfied' && pass ".context/ churn does not invalidate the hash" || fail` → `printf '%s' "$out" | grep -q 'Gate B hook checks passed' && pass ".context/ churn does not invalidate the hash` +sweep codex-gate.test.sh 291 `# review it just recorded -> a permanent stale STOP. The adoption marker is meant to be` → `# review it just recorded -> a permanent stale-fingerprint reminder. The adoption marker is meant to be` +sweep codex-gate.test.sh 296 `printf '%s' "$(commitpre)" | grep -q 'Gate B satisfied' && pass "tracked .context/ state does not invalidate t` → `printf '%s' "$(commitpre)" | grep -q 'Gate B hook checks passed' && pass "tracked .context/ state does not inv` +sweep codex-gate.test.sh 298 `printf '%s' "$(commitpre)" | grep -q 'Gate B satisfied' && pass "tracked .context/ churn stays satisfied" || f` → `printf '%s' "$(commitpre)" | grep -q 'Gate B hook checks passed' && pass "tracked .context/ churn still passes` +sweep codex-gate.test.sh 312 `printf '%s' "$(commitpre)" | grep -q 'Gate B satisfied' && pass "setup: satisfied with sidefile.ts staged" || ` → `printf '%s' "$(commitpre)" | grep -q 'Gate B hook checks passed' && pass "setup: hook checks passed with sidef` +sweep codex-gate.test.sh 316 `printf '%s' "$out" | grep -q 'Gate B satisfied' \` → `printf '%s' "$out" | grep -q 'Gate B hook checks passed' \` +sweep codex-gate.test.sh 324 `printf '%s' "$(commitpre)" | grep -q 'not satisfied' && pass "tracked .context/: real code change still invali` → `printf '%s' "$(commitpre)" | grep -q 'Codex gate state:' && pass "tracked .context/: real code change still in` +sweep codex-gate.test.sh 333 `# direction"), and the STOP message explains that staging alone can cause it.` → `# direction"), and the stale-fingerprint message explains that staging alone can cause it.` +sweep codex-gate.test.sh 339 `printf '%s' "$(commitpre)" | grep -q 'Gate B satisfied' && pass "setup: satisfied on unstaged change" || fail ` → `printf '%s' "$(commitpre)" | grep -q 'Gate B hook checks passed' && pass "setup: hook checks passed on unstage` +sweep codex-gate.test.sh 341 `printf '%s' "$(commitpre)" | grep -q 'not satisfied' \` → `printf '%s' "$(commitpre)" | grep -q 'Codex gate state:' \` +sweep codex-gate.test.sh 341 `&& pass "staging a reviewed tracked file -> NOT satisfied (spec §2 decision)" \` → `&& pass "staging a reviewed tracked file -> gate-state reminder (spec §2 decision)" \` +sweep codex-gate.test.sh 341 `|| fail "staging a reviewed tracked file -> NOT satisfied (spec §2 decision)"` → `|| fail "staging a reviewed tracked file -> gate-state reminder (spec §2 decision)"` +sweep codex-gate.test.sh 341 `# The old trailing assertion ("untracked file on a staged tree -> not satisfied") is` → `# The old trailing assertion ("untracked file on a staged tree -> gate-state reminder") is` +sweep codex-gate.test.sh 350 `# 4. Gate A exec must NOT satisfy Gate B (separate state)` → `# 4. Gate A exec must NOT count toward Gate B (separate state)` +sweep codex-gate.test.sh 438 `printf '%s' "$out" | grep -q 'Gate B satisfied' && pass "re-enable sees same counting semantics as gate-on, no` → `printf '%s' "$out" | grep -q 'Gate B hook checks passed' && pass "re-enable sees same counting semantics as ga` +sweep codex-gate.test.sh 447 `reset_all # state ABSENT -> would normally STOP on a code commit` → `reset_all # state ABSENT -> would normally emit the no-fingerprint reminder on a code commit` +sweep codex-gate.test.sh 453 `printf '%s' "$out" | grep -q 'STOP' && fail "docs-only must not STOP" || pass "docs-only does not STOP"` → `printf '%s' "$out" | grep -q 'Codex gate state:' && fail "docs-only must not emit the gate-state reminder" || ` +sweep codex-gate.test.sh 504 `printf '%s' "$out" | grep -q 'floor met' && pass "3/3 exec -> Gate A satisfied" || fail "3/3 exec -> Gate A sa` → `printf '%s' "$out" | grep -q 'floor met' && pass "3/3 exec -> Gate A floor met" || fail "3/3 exec -> Gate A fl` +sweep codex-gate.test.sh 504 `# FINDING 12: the Gate-A satisfied wording must NOT overstate — it counts calls only.` → `# FINDING 12: the Gate-A floor-met wording must NOT overstate — it counts calls only.` +sweep codex-gate.test.sh 504 `printf '%s' "$out" | grep -qE 'count only|COUNT ONLY' && pass "Gate A satisfied says 'count only' (Finding 12)` → `printf '%s' "$out" | grep -qE 'count only|COUNT ONLY' && pass "Gate A floor-met message says 'count only' (Fin` +sweep codex-gate.test.sh 523 `printf '%s' "$out" | grep -q '1/1' && pass "floor override 1 -> satisfied at 1 pass" || fail "floor override 1` → `printf '%s' "$out" | grep -q '1/1' && pass "floor override 1 -> hook checks pass at 1 pass" || fail "floor ove` +sweep codex-gate.test.sh 523 `printf '%s' "$out" | grep -q 'Gate B satisfied' && pass "floor override 1 -> reports satisfied" || fail "floor` → `printf '%s' "$out" | grep -q 'Gate B hook checks passed' && pass "floor override 1 -> reports hook checks pass` +sweep codex-gate.test.sh 540 `# 18. FINDING 9 — satisfied message distinguishes fresh passes from cycle passes` → `# 18. FINDING 9 — hook-checks-passed message distinguishes fresh passes from cycle passes` +sweep codex-gate.test.sh 556 `# 19. FINDING 11 — WIP commit is cycle-internal: gentle note, no STOP, no reset` → `# 19. FINDING 11 — WIP commit is cycle-internal: gentle note, no gate-state reminder, no reset` +sweep codex-gate.test.sh 561 `printf '%s' "$out" | grep -q 'STOP' && fail "WIP commit must not STOP" || pass "WIP commit does not STOP"` → `printf '%s' "$out" | grep -q 'Codex gate state:' && fail "WIP commit must not emit the gate-state reminder" ||` +sweep codex-gate.test.sh 633 `[ -z "$(commitpre)" ] && pass "non-adopted repo: unreviewed commit -> no STOP" || fail "non-adopted repo: unre` → `[ -z "$(commitpre)" ] && pass "non-adopted repo: unreviewed commit -> no reminder" || fail "non-adopted repo: ` +sweep codex-gate.test.sh 652 `# one — otherwise a stale-tree STOP would masquerade as non-adoption.)` → `# one — otherwise a stale-fingerprint reminder would masquerade as non-adoption.)` +sweep codex-gate.test.sh 660 `printf '%s' "$(commitpre)" | grep -q 'Gate B satisfied' && pass ".on marker alone -> adopted" || fail ".on mar` → `printf '%s' "$(commitpre)" | grep -q 'Gate B hook checks passed' && pass ".on marker alone -> adopted" || fail` +sweep codex-gate.test.sh 665 `printf '%s' "$(commitpre)" | grep -q 'Gate B satisfied' && pass "CLAUDE.md gate heading -> adopted" || fail "C` → `printf '%s' "$(commitpre)" | grep -q 'Gate B hook checks passed' && pass "CLAUDE.md gate heading -> adopted" |` +sweep codex-gate.test.sh 713 `printf '%s' "$out" | grep -q 'STOP' && pass "marker-only: still STOPs" || fail "marker-only: still STOPs"` → `printf '%s' "$out" | grep -q 'Codex gate state:' && pass "marker-only: still emits the gate-state reminder" ||` +sweep codex-gate.test.sh 732 `# 24. Failure contract: an uncomputable hash must never satisfy, and repeated failures` → `# 24. Failure contract: an uncomputable hash must never pass the hook checks, and repeated failures` +sweep codex-gate.test.sh 753 `printf '%s' "$out" | grep -q 'not satisfied' && pass "silent checksum -> not satisfied" || fail "silent checks` → `printf '%s' "$out" | grep -q 'Codex gate state:' && pass "silent checksum -> gate-state reminder" || fail "sil` +sweep codex-gate.test.sh 768 `printf '%s' "$out" | grep -q 'not satisfied' && pass "checksum prints then fails -> not satisfied" || fail "ch` → `printf '%s' "$out" | grep -q 'Codex gate state:' && pass "checksum prints then fails -> gate-state reminder" |` +sweep codex-gate.test.sh 786 `printf '%s' "$(PATH="$stub_dir:$PATH" commitpre)" | grep -q 'not satisfied' \` → `printf '%s' "$(PATH="$stub_dir:$PATH" commitpre)" | grep -q 'Codex gate state:' \` +sweep codex-gate.test.sh 786 `&& pass "seed-copy failure -> not satisfied" || fail "seed-copy failure -> not satisfied"` → `&& pass "seed-copy failure -> gate-state reminder" || fail "seed-copy failure -> gate-state reminder"` +sweep codex-gate.test.sh 800 `printf '%s' "$(PATH="$stub_dir:$PATH" commitpre)" | grep -q 'not satisfied' \` → `printf '%s' "$(PATH="$stub_dir:$PATH" commitpre)" | grep -q 'Codex gate state:' \` +sweep codex-gate.test.sh 800 `&& pass "git diff failure -> not satisfied" || fail "git diff failure -> not satisfied"` → `&& pass "git diff failure -> gate-state reminder" || fail "git diff failure -> gate-state reminder"` +sweep codex-gate.test.sh 812 `printf '%s' "$(PATH="$stub_dir:$PATH" commitpre)" | grep -q 'not satisfied' \` → `printf '%s' "$(PATH="$stub_dir:$PATH" commitpre)" | grep -q 'Codex gate state:' \` +sweep codex-gate.test.sh 812 `&& pass "unresolvable git-dir -> not satisfied" || fail "unresolvable git-dir -> not satisfied"` → `&& pass "unresolvable git-dir -> gate-state reminder" || fail "unresolvable git-dir -> gate-state reminder"` +sweep codex-gate.test.sh 836 `# first commit in a fresh repo STOPs forever. Spec §3 "But an absent index is not a` → `# first commit in a fresh repo gets the gate-state reminder forever. Spec §3 "But an absent index is not a` +sweep codex-gate.test.sh 851 `# ...and the FIRST commit must actually be able to reach satisfied. Hashing and` → `# ...and the FIRST commit must actually be able to reach hook checks passed. Hashing and` +sweep codex-gate.test.sh 851 `# self-matching is not enough: a consumer-side regression could still STOP every` → `# self-matching is not enough: a consumer-side regression could still remind on every` +sweep codex-gate.test.sh 855 `printf '%s' "$out" | grep -q 'Gate B satisfied' || exit 1` → `printf '%s' "$out" | grep -q 'Gate B hook checks passed' || exit 1` +sweep codex-gate.test.sh 855 `) && pass "unborn repo hashes, self-matches, and can reach satisfied" \` → `) && pass "unborn repo hashes, self-matches, and can reach hook checks passed" \` +sweep codex-gate.test.sh 855 `|| fail "unborn repo hashes, self-matches, and can reach satisfied"` → `|| fail "unborn repo hashes, self-matches, and can reach hook checks passed"` +sweep codex-gate.test.sh 866 `printf '%s' "$(commitpre)" | grep -q 'Gate B satisfied' && pass "setup: satisfied on clean tree" || fail "setu` → `printf '%s' "$(commitpre)" | grep -q 'Gate B hook checks passed' && pass "setup: hook checks passed on clean t` +sweep codex-gate.test.sh 869 `printf '%s' "$(commitpre)" | grep -q 'not satisfied' \` → `printf '%s' "$(commitpre)" | grep -q 'Codex gate state:' \` +sweep codex-gate.test.sh 869 `&& pass "staged-vs-worktree divergence -> NOT satisfied" \` → `&& pass "staged-vs-worktree divergence -> gate-state reminder" \` +sweep codex-gate.test.sh 869 `|| fail "staged-vs-worktree divergence -> NOT satisfied"` → `|| fail "staged-vs-worktree divergence -> gate-state reminder"` +sweep codex-gate.test.sh 874 `# 27. Ambient alternate index. Three shapes: a negative-only test would be satisfied by` → `# 27. Ambient alternate index. Three shapes: a negative-only test would be passed by` +sweep codex-gate.test.sh 874 `# an implementation that fires whenever GIT_INDEX_FILE is set — a permanent STOP.` → `# an implementation that fires whenever GIT_INDEX_FILE is set — a permanent gate-state reminder.` +sweep codex-gate.test.sh 884 `printf '%s' "$(GIT_INDEX_FILE="$alt_dir/alt" commitpre)" | grep -q 'not satisfied' \` → `printf '%s' "$(GIT_INDEX_FILE="$alt_dir/alt" commitpre)" | grep -q 'Codex gate state:' \` +sweep codex-gate.test.sh 884 `&& pass "ambient divergent alternate index -> NOT satisfied" \` → `&& pass "ambient divergent alternate index -> gate-state reminder" \` +sweep codex-gate.test.sh 884 `|| fail "ambient divergent alternate index -> NOT satisfied"` → `|| fail "ambient divergent alternate index -> gate-state reminder"` +sweep codex-gate.test.sh 884 `# 27b. stable: same unchanged alternate index across review AND commit -> satisfied,` → `# 27b. stable: same unchanged alternate index across review AND commit -> hook checks passed,` +sweep codex-gate.test.sh 896 `printf '%s' "$(GIT_INDEX_FILE="$alt_dir/alt" commitpre)" | grep -q 'Gate B satisfied' \` → `printf '%s' "$(GIT_INDEX_FILE="$alt_dir/alt" commitpre)" | grep -q 'Gate B hook checks passed' \` +sweep codex-gate.test.sh 896 `&& pass "ambient stable alternate index -> satisfied" \` → `&& pass "ambient stable alternate index -> hook checks passed" \` +sweep codex-gate.test.sh 896 `|| fail "ambient stable alternate index -> satisfied"` → `|| fail "ambient stable alternate index -> hook checks passed"` +sweep codex-gate.test.sh 928 `# the two constant empty-tree hashes MATCH — a false "satisfied" even though the` → `# the two constant empty-tree hashes MATCH — a false "hook checks passed" even though the` +sweep codex-gate.test.sh 943 `printf '%s' "$out" | grep -q 'not satisfied' \` → `printf '%s' "$out" | grep -q 'Codex gate state:' \` +sweep codex-gate.test.sh 943 `&& pass "relative ambient GIT_INDEX_FILE from a subdirectory -> NOT satisfied (Finding 1)" \` → `&& pass "relative ambient GIT_INDEX_FILE from a subdirectory -> gate-state reminder (Finding 1)" \` +sweep codex-gate.test.sh 943 `|| fail "relative ambient GIT_INDEX_FILE from a subdirectory -> NOT satisfied (Finding 1)"` → `|| fail "relative ambient GIT_INDEX_FILE from a subdirectory -> gate-state reminder (Finding 1)"` +sweep codex-gate.test.sh 1006 `expected_ctx="STOP — Codex Gate B not satisfied: the hook cannot confirm that the content you are about to com` → `expected_ctx="Codex gate state: the hook cannot confirm that the content you are about to commit is the conten` +sweep codex-gate.test.sh 1006 `expected_msg="⚠ Codex Gate B not satisfied (cannot confirm review)"` → `expected_msg="⚠ Codex Gate B: cannot confirm reviewed content"` +sweep codex-gate.test.sh 1012 `# 29b. SATISFIED branch: 3/3 passes this cycle, all 3 fresh (unchanged tree). The hook` → `# 29b. HOOK-CHECKS-PASSED branch: 3/3 passes this cycle, all 3 fresh (unchanged tree). The hook` +sweep codex-gate.test.sh 1019 `expected_ctx="Codex Gate B: 3/3 pass(es) this cycle, of which 3 cover the CURRENT content fingerprint (unchang` → `expected_ctx="Codex Gate B: 3/3 pass(es) this cycle, of which 3 cover the CURRENT content fingerprint (unchang` +sweep codex-gate.test.sh 1019 `expected_msg="✓ Codex Gate B satisfied (3/3 cycle, 3 on current fingerprint)"` → `expected_msg="✓ Codex Gate B hook checks passed (3/3 cycle, 3 on current fingerprint)"` +sweep codex-gate.test.sh 1019 `[ "$ctx" = "$expected_ctx" ] && pass "satisfied additionalContext matches exactly" || fail "satisfied addition` → `[ "$ctx" = "$expected_ctx" ] && pass "hook-checks-passed additionalContext matches exactly" || fail "hook-chec` +sweep codex-gate.test.sh 1019 `[ "$msg" = "$expected_msg" ] && pass "satisfied systemMessage matches exactly" || fail "satisfied systemMessag` → `[ "$msg" = "$expected_msg" ] && pass "hook-checks-passed systemMessage matches exactly" || fail "hook-checks-p` +sweep codex-gate.test.sh 1029 `expected_ctx="STOP — Codex Gate B not satisfied: no fingerprint is recorded for this cycle — either no mcp__co` → `expected_ctx="Codex gate state: no fingerprint is recorded for this cycle. The hook cannot tell why — no mcp__` +sweep codex-gate.test.sh 1029 `expected_msg="⚠ Codex Gate B: no recorded review"` → `expected_msg="⚠ Codex Gate B: no recorded fingerprint"` +sweep codex-gate.test.sh 1176 `# 7/9 Gate B satisfied` → `# 7/9 Gate B hook checks passed` +sweep codex-gate.test.sh 1178 `one_doc "Gate B satisfied" "$(commitpre)"` → `one_doc "Gate B hook checks passed" "$(commitpre)"` +``` + +### Task 14 + +Step 4b, after the last text edit — every `span` and `cond` row of `.context/loop-rule-untouched`, parent against worktree, anchors resolved separately in each tree: + +``` +span The derivation is max(risk, security) clean review, and a below-threshold remi CLAUDE.md no difference +span From pass 4 onward every pass report car Those three lines expose CLAUDE.md no difference +span The two rules above do not compete Findings go to a FILE CLAUDE.md no difference +span On squash-merge, copy every evidence ent On squash-merge, copy every evidence ent CLAUDE.md no difference +span Recording a human exception Accepted because: CLAUDE.md no difference +span accumulate; order means nothing. obligation, or a profile-derived evidenc CLAUDE.md no difference +span **"Mandatory" is not limited to this fil because writing it down makes it sound CLAUDE.md no difference +preservation a12 P19 parent=1 worktree=1 CLAUDE.md ok +preservation a14 P20 parent=1 worktree=1 CLAUDE.md ok +preservation h6 P21 parent=1 worktree=1 CLAUDE.md ok +preservation h18 P22 parent=1 worktree=1 CLAUDE.md ok +span The derivation is max(risk, security) clean review, and a below-threshold remi plugins/dev-workflow/commands/workflow-init.md no difference +span From pass 4 onward every pass report car Those three lines expose plugins/dev-workflow/commands/workflow-init.md no difference +span The two rules above do not compete Findings go to a FILE plugins/dev-workflow/commands/workflow-init.md no difference +span On squash-merge, copy every evidence ent On squash-merge, copy every evidence ent plugins/dev-workflow/commands/workflow-init.md no difference +span Recording a human exception Accepted because: plugins/dev-workflow/commands/workflow-init.md no difference +span accumulate; order means nothing. obligation, or a profile-derived evidenc plugins/dev-workflow/commands/workflow-init.md no difference +span **"Mandatory" is not limited to this fil because writing it down makes it sound plugins/dev-workflow/commands/workflow-init.md no difference +preservation a12 P19 parent=1 worktree=1 plugins/dev-workflow/commands/workflow-init.md ok +preservation a14 P20 parent=1 worktree=1 plugins/dev-workflow/commands/workflow-init.md ok +preservation h6 P21 parent=1 worktree=1 plugins/dev-workflow/commands/workflow-init.md ok +preservation h18 P22 parent=1 worktree=1 plugins/dev-workflow/commands/workflow-init.md ok +``` + +Failures: 0. + +### Re-run before Gate-B pass 2 + +Before Gate-B pass 2 (cycle `t57gp3hwu1`), after the pass-1 repairs (§A continue-branch gloss in C, +W and the target text; CHANGELOG Gate-A sentence; Task 12 and Task 13 records): every fragment-table +row P2–P103, F1–F14, F7b re-counted in the copies it claims, to its class's result (OLD, absence → +worktree 0 parent 1; preservation → 1 1), and every recorded NEW / presence fragment → worktree 1 +parent 0: 320 observations, 0 failures. The 18 moved-condition §A presences: 0 failures. The 14 +untouched spans and 8 `cond` rows: 0 failures. Parity over the 34 site regions: the same three +classified differences, nothing new. The prompt-standards items were re-read for the one changed §A +clause: all twelve still pass. + +### Re-run before Gate-B pass 3 + +Before Gate-B pass 3: the pass-2 amendments touched only this plan's Task 13 section (K defined without eligibility, the failed-act row split, three closure checks where one stood). Re-run of every fragment-table row and every recorded NEW/presence fragment: 320 observations, 0 failures; untouched spans and cond rows: 0 failures; parity: the same three classified differences. + +### Gate-B provenance and deviation (cycle t57gp3hwu1) + +**Process deviation, recorded:** the pass-1 repair round (commit `0382219`) and the pass-2 records +refresh (`f56fd50`) were made without Daniel's go, which the session handoff required for repair +rounds. Daniel kept them as the starting point on 2026-09-26 (the reviewer's recommendation he +forwarded); this is not a retroactive authorization. Pass 4 is released as one bounded pass. + +**Review provenance, from the original Codex session transcripts** (`~/.codex/sessions/2026/09/26/`, +the last write to each slot; times UTC): + +``` +pass 1 base d26de4b… head 8436f17… spec session 01a0dccf-4b26 quality session 01a0dccf-4b27 + 08:26:02 quality slot written by 01a0dccf-fbcd (subagent of the quality session) 3 lines — final + 08:26:39 spec slot written by 01a0dccf-cbdf (subagent of the QUALITY session) 3 lines — overwritten + 08:28:09 spec slot written by 01a0dccf-4b26 (the spec session itself) 7 lines — final +pass 2 base d26de4b… head 0382219… 08:42:06 quality 01a0dcdc-1eb4 4 lines · 08:42:09 spec 01a0dcdc-1ec7 6 lines — one write each +pass 3 base d26de4b… head f56fd50… 08:53:58 spec 01a0dce7-01a2 7 lines · 08:55:06 quality 01a0dce7-017e 6 lines — one write each +``` + +Pass 1's overwritten spec write carried three MAJOR findings — the §A decline gloss, the unsplit +next-state rows, the endpoint-only fix-set check — and the spec session's final seven-line file +carries all three as its first three lines, so no finding was lost; every final file is its own +branch's. Passes 1–3 each count toward the floor. Pass 4 runs the two branches as two sequential +calls with identical full base and head ids, each told its one slot. + +### Re-run before Gate-B pass 4 + +Before Gate-B pass 4: close-cited-set corrected (every governing-header commit in the window, demonstrated on A -> B -> A), K defined without eligibility and the source block, Oracle cells added to rows 33 and 33b, provenance and deviation recorded. Re-run: 320 fragment observations, 0 failures; untouched spans and cond rows, 0 failures; parity, the same three classified differences. + +### Re-run before Gate-B pass 5 + +Released by Daniel on 2026-09-26: the `--soft` repair, the observation-labelled checks, the design §7 +note, one pass 5. + +``` +pair §F item 5 `--soft` — OLD `reset to the parent of the first and commit once instead`: 1 at 8e620db, 0 now, C and W; + NEW `` `git reset --soft `, then commit once instead ``: 1 now, 0 at the base, C and W +parity of the Finishing-the-cycle block: no difference +``` + +The OLD half is counted against the previous candidate, not the base: the base carried a different +wording (`` `git reset --soft ` first, then commit once ``) and never this one. +Reset behaviour observed in a disposable repository: two WIP commits, the second adding a new file; +`git reset ` leaves 0 paths staged and the commit fails; `git reset --soft ` keeps both +staged and one commit has parent = base and tree = the last WIP tree. Re-run of the whole record +set: 320 fragment observations, 0 failures; untouched spans and `cond` rows, 0 failures; parity, +the same three classified differences. diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md index 32e93bc..8937f46 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-design.md @@ -276,6 +276,15 @@ set off the block and writes one check per condition**, and fails where the bloc condition the plan has no check for. **No fixture per predicate is built**; that question is parked in the story's §2 and is not reopened. +**How that duty is discharged in this change — Daniel's decision of 2026-09-26, at Gate-B pass 4.** +This change's own cycle closes under §5 as it stood before the change, so **no closing act under the +installed ordering happens here**, and "held at the closing act" cannot be shown for one. The duty is +discharged by **defining each check and demonstrating it in disposable repositories**, and the +evidence entry says so. A check that reads history **reports what it observed** — change observed, +no change observed, or source unreadable — **never that a condition held**: no source available to +it can show that nobody made and undid an edit outside its view, and a check that claimed so would +certify its own blind spot. + **That list is not exhaustive, and reading it as exhaustive is how the evidence entry would overclaim.** Two further things the table does not establish, named because they are the ones a reader would otherwise assume it covers: **how each predicate was derived** — the table takes a diff --git a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md index 67ceecd..1365376 100644 --- a/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md +++ b/docs/superpowers/specs/2026-09-10-loop-rule-consolidation-target-text.md @@ -252,8 +252,8 @@ Nits **and which carries no scope-stop trigger** — those are collected and nev leave nothing to revise, while a Minor or Nit **carrying a scope-stop trigger as the absorb passage defines one** carries it like any other finding, is not clean, and has already been taken by the suspension branch — **severity does not raise a trigger and does not suppress one**, and -which findings raise one is that passage's entire, an already-declined finding and an -already-answered question raising none. It is a branch and +which findings raise one is that passage's entire, an already-declined finding raising no +membership trigger and an already-answered question no question trigger. It is a branch and not an inference, because "does not close" read alone says nothing about whether to run again. **The four standing duties, classified.** The **derived floor** is a **precondition on closure**: @@ -760,7 +760,7 @@ W 1011–1012. ordering's, stated there entire**; this section gives only the operation. Close it with `git commit --amend -m ""`, which replaces the WIP commit; **where a `WIP:` snapshot would survive the amend** — several piled up, or a stray non-amending commit made one an -ancestor — **reset to the parent of the first and commit once instead**. This section is the only +ancestor — **`git reset --soft `, then commit once instead**. This section is the only place either shape is defined. **The hook treats any non-`WIP` commit *attempt* as a Gate-B boundary and clears its state even where the command fails**, so a failed closing act leaves that counter cleared — a fact about the diff --git a/plugins/dev-workflow/.claude-plugin/plugin.json b/plugins/dev-workflow/.claude-plugin/plugin.json index 7251d72..7b628ec 100644 --- a/plugins/dev-workflow/.claude-plugin/plugin.json +++ b/plugins/dev-workflow/.claude-plugin/plugin.json @@ -1,7 +1,7 @@ { "name": "dev-workflow", "displayName": "Cross-Model Review Workflow", - "version": "0.12.0", + "version": "0.13.0", "description": "Spec-driven workflow with two independent cross-model review gates, an append-only hardening ledger with an escalation ladder, and repo-enforced quality. Requires the superpowers plugin.", "author": { "name": "Daniel Sänger", diff --git a/plugins/dev-workflow/CHANGELOG.md b/plugins/dev-workflow/CHANGELOG.md index b6e0b9a..3289350 100644 --- a/plugins/dev-workflow/CHANGELOG.md +++ b/plugins/dev-workflow/CHANGELOG.md @@ -22,6 +22,36 @@ unambiguously, still fails. Deleting only a plugin's *manifest* while the direct keeps shipping fails too. AGENTS.md invariant 12 carries the complete list. +## 0.13.0 + +- **One closure ordering for §5**, installed in `CLAUDE.md`'s gate section and in + `/dev-workflow:workflow-init`'s template: a pass is read once, in a fixed order — source block, + clean completion, suspension, continue — and a cycle closes only on an eligible pass (clean at or + above the floor, or zero findings) with every closure condition holding and the gate's closing act + performed. It names the four standing duties, what a surfaced finding's hold is and what ends it, + how several suspensions compose, and what a failed closing act leads to (repair and retry, a + further pass, or a parked cycle). Gate A gains a content condition — the artifact equals the text + sent in the final pass's review request, which says nothing about what the reviewer read — and a + closing act; Gate B's closing act stays the one `Finishing the + cycle` defines. +- **The absorb paragraph owns the assigned fix set**: the union of every governing story's or plan's + scope plus findings accepted at a membership stop, minus findings declined; a change to the set + costs a further pass. A decline binds for the cycle. +- **Mechanics · Severity answers the question it had handed over**: the ceiling changes what a cycle + must resolve, never what the loop-health readings observe. The resolve duty is scoped to the fix + set, and a validated dismissal resolves a finding. +- **Twenty-three standing sentences the ordering falsified are replaced** — sixteen in the two + prompt copies, among them the `Finishing the cycle` lead-in, the Gate-A and Gate-B coverage + instructions, the HARD FLOOR parenthetical, the human-exception destination and scope sentence, + the Named residual's blanket exemption and the work-loop line — and **seven in the hook's gate + reminders**. The hook now reports what it checked (`hook checks passed`, `no recorded + fingerprint`, `cannot confirm reviewed content`, `Codex gate state: …`) instead of a gate + verdict, and points at the policy's closure ordering for what happens next. No hook logic, + counter, fingerprint or routing changed; `codex-gate.test.sh` moves its expectations with the + strings. +- **The one-contract paragraph** now includes the closure ordering and every rule it reads, with a + membership test a reader can apply sentence by sentence. + ## 0.12.0 - **New command `/dev-workflow:claude-init`** writes one `CLAUDE.md` of general working rules diff --git a/plugins/dev-workflow/commands/workflow-init.md b/plugins/dev-workflow/commands/workflow-init.md index 0d789c6..b7ef2ca 100644 --- a/plugins/dev-workflow/commands/workflow-init.md +++ b/plugins/dev-workflow/commands/workflow-init.md @@ -279,7 +279,7 @@ Strong success criteria let you loop independently. Weak criteria ("make it work Ground progress claims: before reporting a step as done, audit the claim against a tool result from this session ("tests green" needs a test run to point to). Report unverified work as unverified — this keeps status reports factual on long runs. -The work loop includes the review gates: **spec ready → Gate A (spec) → plan ready → Gate A (plan) → execute → tests green → Gate B → commit** (see §5). +The work loop includes the review gates: **spec ready → Gate A (spec) → Gate-A closing act → plan ready → Gate A (plan) → Gate-A closing act → execute → tests green → Gate B → Gate-B closing act** (see §5, which states when each act may be performed and what it is). ## 5. Cross-Model Review (Codex) — TWO MANDATORY GATES @@ -296,8 +296,9 @@ counters and Gate B reports "not run" forever. A mapped name must itself start w to the hook or, for `Bash`/`Skill`, hijacks a lifecycle event. Register the server as `codex` to place its tools there. -**Both gates are a LOOP with a HARD FLOOR: a minimum number of passes per run -(Blocker/Major only), derived from the cited story's profile.** +**Both gates are a LOOP with a HARD FLOOR: a minimum number of passes per run (a clean final pass +being what the floor is spent on, and cleanliness taking more than Blocker/Major), derived from the +cited story's profile.** The derivation is max(risk, security): a value of 0 gives a floor of 1; every resolvable profile above that, and an artifact citing no story, gives 3. Two levels, not three — `high` takes its rigor from lens sets and evidence mode, not from extra @@ -352,19 +353,15 @@ threshold that controls nothing.** The hook still counts passes, and it still ca findings or tell the spec run from the plan run (it resets at `writing-plans`), so Gate A — the spec run especially — is instruction-backed: a satisfied count is not a clean review, and a below-threshold reminder is noted in the pass report and disregarded -where the cycle's own closure rules are satisfied. This replaces the pass-count number -and nothing else. Every other rule stated here about how a cycle closes stands as -written, and none of them is restated — a summary is where their conditions would get -dropped. Nothing here writes the floor knob: it stays the user's, never written, never -removed, never read for this derivation. Open a TodoWrite "Codex pass N" per pass; fix Blocker/Major after each. Your -final pass must be clean — if the pass at the floor still finds Blocker/Major, keep going until -clean or clearly stuck → then STOP and surface to the user. The only early exit -below the floor is a pass with **zero** findings; don't manufacture findings to pad. Codex is -advisory — validate before applying; dismissed finding → one-line why. +where the cycle's own closure rules are satisfied. Every other rule stated **in this paragraph** about how a cycle closes stands as written, and +none of them is restated — a summary is where their conditions would get dropped. Nothing here writes the floor knob: it stays the user's, never written, never +removed, never read for this derivation. Open a TodoWrite "Codex pass N" per pass; resolve Blocker/Major after each as +Mechanics · Severity requires. What a clean final pass and the zero-finding early exit mean for closing is stated once in the +closure ordering. Codex is advisory — validate before applying; dismissed finding → one-line why. **Named residual:** the hook's messages state its own threshold as an obligation, so at a -floor of 1 they report a shortfall the cycle does not owe. Hook text is out of scope here -by decision; what makes that tolerable is the precedence rule above plus the hook exiting +floor of 1 they report a shortfall the cycle does not owe. **That particular overstatement is out of scope here by decision, and it is not a blanket exemption for hook text** — a reminder this change's own rules falsify is corrected in the same change, as the standing sentences above require. +What makes that tolerable is the precedence rule above plus the hook exiting 0 on every branch, not the reminder being harmless. **The gate-off surface — routes known today, not a complete list**, since an enumeration @@ -379,9 +376,16 @@ the cited profiles. **When these rules bind.** From the commit that ships them, and a cycle already running finishes under the rules it started with. Where a cycle's starting rules cannot be -established it takes the stricter reading of every part this change touches — at minimum -floor 3, severity classified without the demotion, the provenance-line duty owed, the curve -duty owed, and the nonce duties at their strictest — the cycle is treated as post-rule, so it +established it takes the stricter reading of every part this change touches — at minimum floor 3, severity classified without the demotion, the provenance-line duty owed, the curve duty owed, +the nonce duties at their strictest, **every suspension binding, since +starting rules that cannot be established cannot be read as having waived an open hold**, **the +repeated-dismissal cleanliness exclusion unavailable, a cycle that cannot establish its starting +rules being unable to establish that they contained it**, **the parked state binding after a closing act that cannot +be repaired, a cycle whose starting rules cannot be established being the last one that should be +left with no terminal transition**, and **every closure condition and +pass-cost rule this change ships owed rather than waived — Gate A's content condition and its +commit-carry duty, and the further pass an assigned-fix-set change costs — since a rule that cannot +be established as absent is cheaper to owe than to skip** — the cycle is treated as post-rule, so it owes a nonce, owes its provenance line and its curve or skip record, and uses that nonce in every cycle record it does write — which changes what a record is named, never whether one is owed, so the working record stays optional and a skipped cycle still writes no findings slots. Where it @@ -419,54 +423,481 @@ the predicate itself. In any such state nothing here resolves which rule governs have a human complete or revert the adoption, before running a gate under it.** What prompt text can do about downstream adoption is limited, and that limit is what this paragraph states. +**How a cycle ends — one ordering, stated here and referenced everywhere else.** A pass is read in +a fixed order, because every rule bearing on one decision — may this cycle close — otherwise +qualifies the others and the ranking survives only in a reader's head. + +**What this paragraph owns.** It owns **how a pass is evaluated and what follows from that**: the +evaluation order, what each predicate is read from, the eligibility test, the hold a surfaced +finding places and what discharges it, the composition of several suspensions, the pairs that +cannot co-occur, and — as a stated exception, because precedence is evaluation order — the +clearly-stuck precedence clause quoted into it below. **No number is put on that list.** Whether +the read model counts as part of the evaluation order or as a thing beside it is a question of +wording, not of authority. It also owns the **classification** of the standing duties below, and +**those duties bind both gates alike**: the derived floor and the Blocker/Major-resolve duty keep +their definitions at their own sources and are read here for every cycle whatever its gate, while +the hold and no-clean-credit are defined here, being properties of the evaluation itself. + +**What belongs to a gate is what the two gates do differently: the content condition its closure +requires, where it has one, and the closing act.** **Gate A has such a condition and Gate B has +none**; each paragraph below states its own and is read from there, and **neither gate's paragraph restates the conditions this one gives every cycle**, so +neither is a complete inventory on its own. **Everything else this paragraph names it cites**: the +scope triggers, the assigned fix set, every severity rule and every closure precondition with a +source of its own keep their one definition in the paragraph that owns them, and a reader who finds +one of *those* defined here has found a defect. + +**The conditions every cycle has, whatever its gate.** The duties classified below; and the +**profile**, the **cited set** and the **assigned fix set**, each gating as its own source says and +cited here without restatement. **The profile and the assigned fix set answer a change and not a +differing value**, which is why one of them changed and then undone still costs a pass. **The cited +set answers as its own source says**, and that source says two things: a governing header changed +**during** a pass makes that pass not final, and the final clean pass runs against the **current** +set. The contrast that matters is against a gate's own content condition, which may be a comparison +of current values and says so where it is stated. + +**What a pass is read from.** Every finding-derived predicate reads the validated findings file +**or files** of the logical pass as **the concatenation of their finding lines after each file has +been validated separately** — a `full` Gate-B pass has two, one branch alone is already an +incomplete pass, each file's terminator is not a finding line, and a branch whose body is +`NO FINDINGS` contributes an empty sequence rather than a line. **Which severity field each +predicate reads is settled in Mechanics · Severity**, which is where that split lives and is not +repeated here. Beyond the findings, closure reads the **derived floor**, the **resolve duty's +standing over this cycle**, any **hold still standing**, and **every closure condition this cycle +has** — those above and this cycle's gate's. The scope triggers read the **current assigned fix +set** as the absorb paragraph defines it, and the answers already given; the clearly-stuck reading +adds its own coverage judgement. **A line in one branch file and a line in the other are distinct +findings for holds and answers**, so a `full` pass asks twice rather than risk resuming over one it +never asked about. **Within one running cycle an answer binds to the finding or question as the +pass that raised it recorded them** — which is what an agent running the cycle can do with nothing +written down. Recognising the same finding or question **across a lost session** has no mandatory +identity, sameness or recovery rule in these rules; the optional `-dispositions.md` note is +advisory and authoritative for nothing. + +**The four branches are named, and a cross-reference anywhere in this section names the branch +rather than its position**, so reordering them breaks no reference. **They are read once, on the +pass, before any closing act is attempted** — so a pass that reached the act has already been +classified and is not classified again by what the act does. + +**The source-block branch, read first.** **Where any unmet closure condition's own source +prescribes stop-and-surface** — a profile present but unresolvable, governing headers that +disagree, a `Story:` header that cannot be read, an unobservable counterfactual, and **any other +source rule that prescribes it; the list is examples and not the set** — **the cycle stays open, that source decides what +must be repaired or answered, and no further pass runs while its block stands.** It is neither a +suspension nor a continue and needs no name and no procedure of its own: the source rule carries +both, and this ordering's part is to send the reader there rather than to run a pass over a cycle +another rule has stopped. **Once its source condition is repaired**, a block that stood **before any pass of this cycle was +read** — the ones its floor derivation and its cited set raise among them — leaves the cycle to run +its next pass, there being no pass to read again; a block that stood **on a pass already read** has +**that pass read again through the ordering**, no closing act +having been attempted on it — **and where that pass also carried a suspension, this branch +releases only its own block**: the composition rule still holds the cycle on every answer that +suspension asked for, a continue still leads to a pass run after the answer, and a stop still parks +the cycle until an explicit later continue — so the reread happens where no suspension of that +pass is outstanding, and otherwise the suspension's own route runs first — the read-once rule below is about a pass that reached the act, not +about one a block held before it, and without this a repair that moves no pass-cost value would +leave a clean eligible pass with no route to the act and none to a suspension. +**It is read first and it silences nothing.** Where the same pass also +carries a suspension, that suspension is surfaced with its reasons and its questions exactly as the +suspension branch requires and its answers are collected; what the block adds is that **no next +pass runs until its own source condition is repaired**, whatever those answers were. + +**Then the clean-completion branch, and eligibility is its own test.** §5 uses *clean* in two +senses and now says which is which: a **clean findings file** is the `NO FINDINGS` signal the +protocol defines, and a **clean pass** is the predicate here, read on the logical pass with every +required branch file combined, so one branch's clean file never establishes a clean pass. **A pass +is clean** when its findings carry no in-set Blocker or Major at effective severity and **no +scope-stop trigger** — the two the absorb paragraph defines, read there and not redefined here, +each already carrying the qualification **an answer given before that pass ran** puts on it. **A +pass's cleanliness is settled on what it found and on the answers standing when it ran**, and a +later answer never rewrites it: an answer discharges the holds it was asked for and leaves the pass +that raised them exactly as clean or unclean as it was, which is the same fact the duties paragraph +states of no-clean-credit. Reading a later answer back onto an earlier pass would let a cycle close +on a pass that was surfaced, answered and never re-run. A pass with **zero** findings is clean +whatever the floor, because a floor buys further looks at an artifact that keeps yielding findings, +and one yielding none has already given what those looks were for; don't manufacture findings to +pad. + +**A repetition of a finding this cycle has validly dismissed does not, on its own, make a pass +unclean.** Every part of this is required and the exclusion is narrow: the **dismissal was made in +an earlier pass of this cycle**, its stated reason is **still true of the artifact as it now +stands**, and the later finding **makes the same complaint and brings no new evidence** — no +observation the dismissal did not answer, and no change to the text the reason turned on. Where any +part fails — new evidence, relevant content changed, or genuine doubt that this finding is that +one — the finding is read afresh like any other, and **doubt never resolves in the exclusion's +favour**. Without it the standing duty *dismiss validly, then run another pass* cannot finish, +because a reviewer repeating its own refuted claim would decide whether that claim had been dealt +with. **It reaches only a dismissal this cycle made**, which is what an agent running the cycle +knows; nothing here ships a record, and recognising a dismissal across a lost session has no more +support than the paragraph above gives it. + +**It changes cleanliness and nothing else, which is what keeps it from being a waiver.** The +repetition is **still a finding**: it stands in its pass's findings file, and every loop-health +reading counts it exactly as it counts any other, so a loop spending passes on a point it keeps +refuting still shows up as one. It remains the clearly-stuck reading's re-raise condition. **No +earlier pass becomes clean in retrospect** — a pass's cleanliness is settled on what it found and +is never rewritten, which this exclusion leaves untouched: it decides the pass being read and no +other. And it is **not** a second dismissal; the resolve duty was discharged when the finding was +dismissed and there is nothing here to discharge again. + +**Eligibility is exactly this and nothing more: a clean pass at or above the derived floor, or a +zero-finding pass.** It is a property of the pass. **Closure is eligibility plus every closure +condition of this cycle holding plus this cycle's gate's closing act**, and those conditions are +properties of the *cycle*, not of the pass — so an eligible pass whose conditions are unmet closes +nothing, and is not thereby made unclean. Keeping the two apart is what lets the order decide the +case where they disagree, which the branches do. **The order inside closure is fixed: every +condition is established first, and only then is the closing act performed.** + +**An act that does not complete has not closed the cycle**, and it is not a branch: the branches +below decide what a *pass* is, and this cycle's pass already took the clean-completion branch and +reached closure. **A failed act returns to the closure step it failed in, not to the branches.** +**Nothing is assumed about what the attempt left behind** — a hook can modify and stage content +before failing, so the attempt itself can move what a condition is read from — so surface the +concrete command failure and **re-establish every closure condition against the repository as it +now stands.** Where they all still hold, perform the act again. Where the attempt or its repair +moved anything a condition is read from, **that condition has changed and its own rule decides what +it costs**, a further pass included, and the cycle is back in the ordering with that pass owed. +**Where the failure cannot be repaired at all** — a signing key nobody has, a permission nobody can +grant — **surface it and leave the cycle parked**: open, not running, spending no passes, restarted +by an explicit later continue, which is the state a stop answer already produces and is named here +rather than invented. A cycle that can neither close nor be parked is the outcome this sentence +exists to prevent. + +**Closure introduces no new kind of record, and it excuses none**: every other record this cycle +owes, a human-exception record among them, is owed and written exactly as before, and **a +human-exception record this cycle owes goes in the commit its closing act uses**, so the two never +land in different places. **No closure condition is read on the branch tip**, the conditions being +read where their sources say and not off the tip. **That is not a claim that the act publishes what +the pass read.** A commit is written from the effective index, so a staged edit the working tree +does not show lands in the closing commit; where the edit changes something a **source rule** +governs — a profile value, cited-set membership, the assigned fix set — **the change has happened +and that source's own rule applies**, so a further pass is owed and no condition is added here for +it. **Where it changes a review input no source rule governs** — a cited story's acceptance +criteria or settled decisions, say — **nothing here reaches it, and nothing in this section does**: +the artifact's own equality condition covers the artifact, and the inputs beside it have only the +rules their sources give them. + +A plateau or tells on **the pass that closes** go into **that pass's status report to the user** — +the carrier the three-line duty already names, and no second report form is introduced — and never +block it, because reporting "will not converge" on a converged loop is a false report; on an +eligible pass that does **not** close they go into the same place, a pass reporting what it read of +the loop whether or not it closes. **No other pass outcome makes a cycle eligible to close**, +because every other pass either leaves a required repair, a hold or a question outstanding, **or +has not reached the floor, or is itself unclean on its own findings** — and closing over any of +those is the failure this ordering exists to prevent. The floor is named separately because a +below-floor pass whose only findings are Minors leaves nothing outstanding and is still not +eligible; **pass-level uncleanliness is named separately** because an in-set Blocker or Major +repaired after the pass that raised it discharges the resolve duty without making that pass clean, +so a cycle can owe nothing and still hold no pass it may close on. The one termination that is not a pass outcome is the Gate-B triviality skip, which runs +no passes and is outside this ordering. + +**Then the suspension branch, which a pass reaches where the clean-completion branch did not take +it and a suspension applies to it.** **Clean completion outranks a suspension by taking the pass to +the closing act, not by eligibility alone**: a pass that took the branch above, met +every closure condition and had the closing act performed has ended the cycle, and a suspension has +nothing left to suspend. **A pass the clean-completion branch did not take reaches this branch +whatever its cleanliness, where a suspension applies to it** — a clean pass below the floor, and +equally an eligible pass whose unmet closure conditions kept that branch from taking it. Cleanliness is what this branch stops +asking about; whether a suspension applies is still what puts a pass here, and where none does the +continue branch has it. That is what makes "clean completion outranks the two-tell stop" +executable rather than asserted, and cleanliness alone never decides it. **What D2 and D3 forbid is +reporting "will not converge" on a loop that converged, and a loop still owing a repair, an answer +or a closure condition has not converged** — so a mandatory two-tell stop and the clearly-stuck +reading stay reachable exactly where the loop is still running, which is the only place their +question means anything. Three suspensions, by the names their +paragraphs use and read by those paragraphs: the **scope stop**, raised by either trigger above — a +**membership stop** by the first, a **question stop** by the second; the **clearly-stuck exit**; +and the **two-tell stop**. A suspension waives nothing. Any non-empty set of them can apply to one +pass: **one surface, every reason reported, every question asked**, because a reason left out is a +decision made by omission. A finding the clearly-stuck reading surfaces that also carries either +trigger takes the scope stop's answers at that same surface, so it is not asked twice; the two-tell +stop surfaces tells and not a finding. + +**Otherwise the continue branch, which no source block reaches: where none stands and neither of +the two branches above took the pass, it continues** — the loop +runs another pass on the **current** artifact, revised where the severity and scope rules require a +repair and unrevised where they do not. **An eligible pass with an unmet closure condition lands +here**, and like every other non-closing pass **only where no suspension applies to it**: clean +completion did not take it, so the loop continues on whatever the unmet condition requires — most +often a repair still owed from an earlier pass. A below-floor clean pass lands here on the same +terms; where a suspension does apply, the suspension branch has already taken it, because only +the clean-completion branch outranks a suspension. So does a pass whose only findings are Minors and +Nits **and which carries no scope-stop trigger** — those are collected and never iterated and may +leave nothing to revise, while a Minor or Nit **carrying a scope-stop trigger as the absorb +passage defines one** carries it like any other finding, is not clean, and has already been taken +by the suspension branch — **severity does not raise a trigger and does not suppress one**, and +which findings raise one is that passage's entire, an already-declined finding raising no +membership trigger and an already-answered question no question trigger. It is a branch and +not an inference, because "does not close" read alone says nothing about whether to run again. + +**The four standing duties, classified.** The **derived floor** is a **precondition on closure**: +it gates closing, discharged by the count of valid logical passes reaching it with the last of them +clean, or by the zero-finding exit. The **Blocker/Major-resolve duty** is a **precondition on +closure**: Mechanics · Severity, scoped to the assigned fix set, states what it demands and what +discharges it, at that source and not here. **It is discharged per finding and tracked across the +cycle, never inferred from a later pass.** A findings file establishes the **inventory** of what +that pass found and not the resolution of anything, so a later pass that does not mention an +earlier in-set Blocker or Major says nothing about whether it was resolved; reading its absence as +discharge would let an omission close a cycle. **The duty is not a second test on whether a pass is +clean, and the two are not run together.** A pass is clean on its own findings. Stated as the case +that separates them, because a reader who conflates them decides it wrongly: **pass 1 raises an +in-set Major; it is not resolved; pass 2 finds nothing.** Pass 2 **is** clean, and at or above the +floor it is **eligible** — and the cycle **still cannot close**, the duty being unmet; it takes the +continue branch until that Major is discharged. What the open Major does *not* do is make pass 2 +unclean. The **hold** a surfaced finding places on closure **participates in the ordering**: it +gates closing while it stands, and is discharged by the answers that surface requires. **It +attaches to every surfaced finding, whichever suspension surfaced it** — clean completion creates +none, because it wins before anything is surfaced. **No-clean-credit** — no pass carrying a +scope-stop trigger is credited as clean — also participates, and is a fact about that pass that +nothing discharges, a later pass being judged on its own findings. It is not a second test beside +the clean predicate but that predicate's second half, which is why it is stated in its words. + +**What a suspension asks, and what ends it.** **A hold ends when every answer its finding requires +has been given, in whichever direction each is given** — one **scope-stop** answer for a +single-trigger finding, both for one carrying both. That is **one rule with two parts**, how many +answers and which way each may go, and neither is a test the other has to pass. Where a health +suspension applies to the same pass, its shared continue-or-stop answer is **additional** to those +and not counted among them, the health readings asking about the loop rather than about this +finding. **The two health readings differ in what they surface, and therefore in what they hold.** +The **two-tell stop surfaces tells and no finding**, so it creates no hold; what it leaves +outstanding is its own continue-or-stop question, which the composition rule below holds the cycle +on until it is answered. The **clearly-stuck reading surfaces findings** — **the findings of the +pass being read that satisfy its regeneration condition, and only those**; earlier members of a +regeneration chain that were repaired or dismissed are history the reading consults and never +findings it re-surfaces, so no discharged finding takes a second hold. Each surfaced finding takes +a hold like any other. **A re-raised valid dismissal stays discharged for the resolve duty on +exactly the terms the clean predicate sets out above** — the same complaint, no new evidence, no +change to the text the dismissal turned on, and the dismissal's reason still true of the artifact. +The dismissal was the resolution and a reviewer repeating it does not undo it, so no second +dismissal is owed. **Where any of those fails the recurrence is an ordinary fresh finding** and is +handled as one — by the severity and scope rules at their own sources, which decide whether it is +in set and what it owes; reading the old dismissal as covering it would let a finding that has +since become true close a cycle. What a +qualifying recurrence creates is the **clearly-stuck hold**, ended by that reading's +continue-or-stop answer. Any trigger the recurrence independently carries raises its own +stop as usual. **Where one finding is surfaced by more than one route it carries a hold +component per route, and each is discharged by its own answer**: the **membership** component by +the membership answer **in either direction**, a decline releasing it exactly as an accept does; +the **question** component by the user's decision on that question; the **clearly-stuck** component +by that reading's continue-or-stop answer. **No answer discharges another route's component**, and +a finding surfaced by one route has one component, discharged by the one answer its surface asks +for. **Resumption is still the +composition rule's**, which waits for every outstanding answer. At a **membership stop** the answer is **accept** or **decline**; +**what each does to the assigned fix set is the absorb paragraph's, which defines that set from +these outcomes**, and what an in-set finding then owes is Mechanics · Severity's. A later answer +that contradicts a decline **does not reverse it**: the decline **remains binding** and the +contradiction is **surfaced to the user as information**, changing neither membership, nor the +cycle's state, nor any outstanding question — **whether the loop resumes is decided by the +composition rule below and by nothing here**, so a contradictory answer is never itself a +resumption and cannot step past a hold or a health question still awaiting its own answer. Nothing +here turns one answer into another, since that would let a finding be moved out of the set and back +into it to escape what it owes inside it. **There is no withdrawal inside the cycle that +declined**, a decline binding for the remainder of its cycle and admitting no exception; a +reconsideration is a later cycle's, where that decline has no effect at all and the finding takes +the ordinary route. Either answer is an **explicit, attributable decision on that specific +finding** — never silence, never a general remark about scope, never inferred, because a fix set +changed by inference is a fix set nobody chose. **Membership is answered against the set as the +absorb paragraph fixes it for the pass that raised the question**: a later broadening is a new fact +the **next** pass reads and never discharges a standing hold, a hold discharged by a scope change +being a hold nobody answered. At a **question stop** the answer is the user's decision on the +question and membership does not change; an out-of-set finding that opened one is a membership stop +as well. **Decline is available only at a membership stop**, that being the only stop whose +question is whether a finding belongs to the set. The **clearly-stuck and two-tell readings** ask +**continue or stop**. **Continue consumes the reading that raised the suspension**: a further +health suspension needs that reading recomputed over a pass run after the answer, which is new +data — so continue produces a distinct next state, and the same reading cannot return the same stop +unanswered. It permits an **unrevised** artifact **only where no repair is owed**; where effective +severity or scope requires one, that repair comes before the post-answer pass, since a pass run +over an unrepaired in-set Blocker or Major spends a look on text the rules already say must change. +**Stop parks the cycle**: open, not running, spending no passes, restarted only by an explicit +later continue — a distinct state from the suspended-awaiting-answer one it was in before the +answer. **That continue restarts the cycle and never skips an answer**: where any question the +suspension raised is still outstanding, it returns the cycle to suspended-awaiting-answer, and only +once every answer the composition rule requires has been given does the next pass run. So a cycle +parked with an unanswered membership or question stop cannot be continued into a pass, and cannot +sit parked with no transition either — the continue is always available and always moves it. +Nothing a parked cycle wrote is a closing commit, and a parked cycle nobody restarts is a human's +to resolve, exactly as the nonce rules already say of open cycles. + +**Composition, and what cannot happen.** Every **question** is answered on its own and the loop +resumes only when every answer resumes it — accept or decline at a membership stop, a decision at a +question stop, continue at the health readings; one stop answer parks the whole suspension, because +a loop resumed over an unanswered question decides it by running. **The clearly-stuck and two-tell +readings raise one question between them, not two**, both asking continue or stop, so one answer +carrying every reason ends both — an instance of the sentence before it, not an exception. **Two +pairings cannot occur**, and no rule ranks them: clean completion and a **scope stop**, since that +stop's triggers are the clean predicate's own second half, so a pass raising one is not clean; and +a zero-finding pass and any suspension, since it has nothing to surface, nothing regenerating, no +cluster and no require↔withdraw pair. **Clean completion and the clearly-stuck exit can**, and the +overlap is admitted rather than argued away: the two read different severity fields, as +Mechanics · Severity sets out, so an **in-set** Blocker or Major the ceiling demotes below Major +can regenerate across passes on a pass that is clean. **A declined finding is not a route into that +reading**: the third condition admits regeneration across repair attempts and a qualifying +re-raised validated dismissal, and a decline is neither — it is the user's decision that a *true* finding stays outside +the set, and it binds for the cycle. **The order decides it and no new rule is needed.** The +clearly-stuck paragraph's own precedence clause is stated here rather than there, because +precedence is evaluation order and this paragraph is where evaluation order is stated once; the +rationale that clause turns on stays beside the reading in that paragraph, which is where the +reading itself lives. **A clean completion takes precedence over this exit**: a Blocker/Major-free +pass **at or above the floor** has satisfied the clean-final-pass rule — collect the Minors and +Nits and close — and reporting "will not converge" on a converged loop is a false report. Below the +floor the pass **suspends**, the clean pass having failed eligibility. **That sentence ranks two +readings and licenses no closure**, its "close" being the closure this ordering defines and +carrying every condition that closure carries: a Minor or Nit bearing a scope-stop trigger makes +the pass unclean, so the sentence does not reach it, and an undischarged duty, a standing hold or +an unmet gate condition means the pass does not close — leaving it on the suspension branch where +one applies and the continue branch where none does, exactly as those branches say. + +**Gate A's content condition, and its closing act.** These are what Gate A adds to the conditions +the ordering states for every cycle; that list is there and is not repeated here. + +**Its content condition: the artifact as it now stands is identical to the text that went into the +final pass's review request.** Gate A hands the reviewer text rather than a git range, which is why +this condition is Gate A's and is written nowhere else. It is **current equality and deliberately +nothing more**: it does **not** say the artifact went untouched in between, and text edited and +then restored byte for byte satisfies it — stated here because the ordering's conditions answer a change +instead, and said of this condition rather than as a claim about every other. +That is a decision rather than an oversight: a content comparison cannot tell those two states +apart, and a condition nobody can check is a condition nobody applies. It likewise says nothing +about **what the reviewer consumed**: no part of this act is offered as evidence of the review +payload. + +**The act.** A Gate-A cycle has no WIP snapshot to replace, so it closes by **writing the closing +commit — or the closing message of one that already exists — carrying the records this cycle +already owes**: its provenance line and its per-pass curve, in the forms Mechanics fixes, neither +of them altered. **The content must survive into that commit, and carrying it there is part of the +act**: the equality above is read on the artifact, while the commit is written from the effective +index, so an act that does not carry that same content through has checked the condition without +performing it. **The commit that closes the cycle contains, at the artifact path, exactly the text +the condition was read against.** The safe command sequence and whatever demonstrates it belong to +the plan; **the duty belongs here**, and it needs no new fingerprint and no new record. + +**Two cases, told apart by the repository's current state rather than by which commit introduced +what; the safe git sequence for each belongs to the plan. That reading decides how a cycle closes, +never whether it may.** +- **`HEAD` already carries that text at the artifact path.** Close by **amending `HEAD`'s message** + to add the records, leaving that path as it stands. +- **`HEAD` does not, the reviewed text being still uncommitted.** **Commit it unchanged** and close + in that commit. This case had no answer before: an eligible pass over repairs nobody had + committed could neither close nor suspend. + +**No new revision of the artifact is made to close a Gate-A cycle**, a new revision being one no +pass has run against — and committing already-reviewed text that was never committed is not one. + +**Gate B's content condition, and its closing act.** These are what Gate B adds to the conditions +the ordering states for every cycle; that list is there and is not repeated here, so nothing below +is an inventory of what this gate requires. **This gate adds no content condition of its own**, +and the next paragraph says why. + +**Its content condition is not an artifact/request equality, and none is written for it.** Gate B +reviews a **diff** identified by `baseSha` and `headSha` rather than a text handed to the reviewer, +so there is no reviewed text to compare an artifact against, and **nothing is put in its place**: +content the final review request did not select can reach the closing commit, and **no rule in this +section reaches it**. **What this gate does have is every Gate-B duty this section already +states, at the paragraphs that state them, and none of them is summarised here** — a compressed +inventory is where a load-bearing part goes missing while the list still looks complete. + +**The act** is the one Mechanics · Finishing the cycle describes, in whichever shape that section +gives the repository's current state. **This paragraph says when it is performed and never which +shape it takes**: once the ordering reaches it, on an eligible pass with every closure condition +holding, **never a clean pass on its own**. + +**A commit the hook reads as cycle-closing is a Gate-B matter.** A non-`WIP` commit mid-cycle makes +the hook drop its Gate-B review state and that gate's counter — an observation about the counter, +since the cycle itself stays open until the conditions hold, so an accidental commit resets what +the hook reports and closes nothing; **what the reset erases is counter state, and it does not +invalidate a pass that already satisfied the validation rules this section states** — which is +where what makes a pass valid stays. **A stray commit that succeeds and does not amend also leaves the +`WIP:` snapshot as an ancestor**, an amend replacing the tip instead — +so in that one shape **the closing act still owes what Mechanics already requires of it: no `WIP:` +commit left in history.** Which git sequence reaches that from this state +belongs to the plan, as every other closing sequence does. **It does not reach a Gate-A cycle's count**, which the hook +clears at the skill boundaries that start a new Gate-A cycle rather than on any commit. + **What a loop absorbs, and what stops it — a question of scope, not of action.** A finding that corrects the correction you just made **and stays inside the assigned fix set** is **inside this loop's scope**: keep it here rather than handing it back, then act on it by its -severity exactly as the severity rule already says — Blocker/Major resolve, Minor/Nit collect -and never iterate. Ancestry decides where a finding belongs; it never decides what you do -with it, and it grants no Minor or Nit a repair round it would not otherwise get. **The assigned fix set is fixed before the pass you are answering: it is the -scope the approved story or plan assigns to this cycle, plus repair obligations you already -accepted in earlier passes.** A finding is in-set when repairing it stays inside that scope — -never merely because it arrived in the current pass, which would put every new finding in the -set by definition and leave the boundary deciding nothing. Where membership is genuinely -unclear treat the finding as **outside**, which costs a question and never a silent expansion. **A correction that leaves that set stops the -loop like any other out-of-scope finding**, even when it opens no new question at all — -absorbing it would grow the assigned work without anyone agreeing to that — and it resumes -the moment the user says whether the set now includes it. A finding -that opens a **new structural or contract question** stops the loop and goes to the user — -**size is not the test, novelty of the question is**, so a structural finding that is -genuinely small still stops it, while a long correction still aimed at the last correction -does not — provided that correction, too, stays inside the set, which its ancestry never -supplies on its own. **When a finding is both** — it corrects the last correction *and* opens a new -structural or contract question — **the new question wins and the loop stops**: novelty -overrides correction ancestry, because absorbing on ancestry is how a contract decision -gets made without anyone choosing it. Stopping this way is **not an exit from the gate**: the floor, the -Blocker/Major filter and the clean-final-pass rule all stand, and the loop resumes on the -revised artifact once the question is answered. What it prevents is a loop committing you -to a design nobody chose — a different failure from an unfinished review. +severity exactly as Mechanics · Severity says. Ancestry decides where a finding belongs; it +never decides what you do with it, and it grants no Minor or Nit a repair round it would not +otherwise get. **The assigned fix set is fixed before the pass you are answering: it is the +union of the scope every approved story or plan governing this change assigns to this cycle, +plus every finding this cycle has accepted at a membership stop together with any repair +obligation accepted with it, minus every finding this cycle has declined.** **A decline excludes +the finding it answers and nothing else**: it does not cancel a repair obligation that the +approved scope, or another accepted finding, has independently put in the set. So where a `full` +Gate-B pass's two branch files carry the same complaint, **each line is a finding of its own, owes +its own explicit answer, and each answer binds only its own line** — an acceptance puts its own +finding in, a decline takes only its own finding out, and neither reads the other; the set is +whatever the definition above then computes. **A declined finding stays declined for the cycle**, +and performing a repair to discharge a different finding's obligation neither reverses that +decision nor returns it to the set. **This adds no deduplication, no reconciliation stop and no +rule that an acceptance overrides a decline** — the two answers are about different findings, and +until both are given the unanswered one's membership hold stands. **An accept puts the +finding in the set whatever its severity**: membership and the repair duty are different things, +so a Minor or Nit accepted into the set is in it though Mechanics · Severity asks no repair for +it, and a later pass that recomputed it as outside would raise the membership question a second +time and make the accept decide nothing. A finding is in-set when repairing it stays inside **the assigned fix set as +just defined** — never merely because it arrived in the current pass, which would put every new +finding in the set by definition and leave the boundary deciding nothing. Where membership is +genuinely unclear treat the finding as **outside**, which costs a question and never a silent +expansion. **A correction that leaves that set stops the loop like any other out-of-scope +finding** — except one this cycle has already declined, which is outside the set by that +decision and **raises no membership trigger on that account**, its membership being the one +question already answered — even when it opens no new question at all, and the membership +answer ends that finding's membership hold; **what the pass does next is the closure +ordering's**, which resumes only when every answer outstanding on that surface has been given. +A finding that opens a **new structural or contract question** — new meaning not already +answered in this cycle, so an answered question raised again stops nothing — stops the loop and +goes to the user — **size is not the test, novelty of the question is**, so a structural finding +that is genuinely small still stops it, while a long correction still aimed at the last +correction does not — provided that correction, too, stays inside the set, which its ancestry +never supplies on its own. **When a finding is both** — it corrects the last correction *and* +opens a new structural or contract question — **the new question wins and the loop stops**: +novelty overrides correction ancestry, because absorbing on ancestry is exactly how a contract +decision gets made without anyone choosing it. **Novelty overrides ancestry and nothing else: +where the finding is also out of set, both triggers hold and both answers are owed**, since a +question answered about a finding nobody placed in or out of the set leaves its membership +decided by default. Stopping this way is **not an exit from the gate**: it is a **suspension** +in the closure ordering's sense, the floor, the Blocker/Major filter and the clean-final-pass +rule all stand, and what the answer does is stated there — what the stop prevents is a loop +committing you to a design you never chose, which is a different failure from an unfinished +review. **A change to this set costs the cycle at least one further pass.** The set a pass was +begun under is the set its closure would rely on, so the window opens **when that set is fixed +for the pass**, as this paragraph defines it, and runs to the closing act; a change anywhere in +that window costs a further pass, **in either direction and whether or not the change is later +undone**, the set having governed the pass differently while it stood. That further pass must +itself be clean and every other closure duty must be satisfied; it is one more pass, not a +licence to close on the next one. **Recognizing "clearly stuck", so that exit is a reading and not a mood.** Read the **Blocker curve across passes**, not any single pass's total — it is the better of the two signals, the total says less than it looks like, and one low count is a snapshot rather than a plateau. **Neither curve measures coverage:** a low Blocker count can sit beside an -entirely unreviewed subsystem. So this exit needs three things **together**, and a missing -one means keep going: a plateau visible across passes (six or more is where the field saw -one); an **affirmative judgement that coverage is sufficient**, stated — a known materially -unreviewed area forbids this exit outright, and disclosing it does not license it; and -**Blocker or Major findings that keep regenerating across genuine repair attempts**, each -round's fix producing the next. That third condition is what makes a plateau rather than a -finish, and it is why **a clean completion takes precedence over this exit**: a -Blocker/Major-free pass **at or above the floor** has satisfied the clean-final-pass rule — -collect the Minors and Nits and close — and reporting "will not converge" on a converged -loop is a false report. **Below the floor nothing closes**, and a zero-finding pass remains -the only exception, exactly as above; a Blocker/Major-free pass below the floor -carrying a Minor keeps -looping. -**Surfacing does not close the cycle, and that is what makes this reachable.** You surface -*with the finding still open* — the resolve rule is not waived, no pass is credited as -clean, and the loop resumes on whatever the user decides. Reading it as "stop instead of -fixing" would put the exit in competition with the rule that every Blocker and Major -resolves, and then nothing could satisfy both. +entirely unreviewed subsystem. So this exit needs three things **together**, and a missing one means only that *this* exit does not apply — what the pass does instead is the closure ordering's, read there in full: a plateau +visible across passes (six or more is where the field saw one); an **affirmative judgement that +coverage is sufficient**, stated — a known materially unreviewed area forbids this exit outright, +and disclosing it does not license it; and **Blocker or Major findings that keep regenerating +across genuine repair attempts**, each round's fix producing the next — **or a finding the author +has validly dismissed that the reviewer re-raises across passes on the terms the closure ordering +sets**, a recurrence failing them being an ordinary fresh finding and not a re-raise at all, the +re-raise standing in for +the regenerating fix, since a dismissal gets no repair and produces none, and a reviewer returning +to the same refuted point every pass says the same thing about the loop that a fix producing the +next finding says. That third condition is what makes a plateau rather than a finish. **Where this reading and a clean +completion both apply, the closure ordering decides it** — the precedence sentence lives there, +because precedence is evaluation order. +**Surfacing does not close the cycle, and that is what makes this reachable.** You surface *with +the cycle and the new hold still open* — the resolve rule stands over the finding exactly as +Mechanics · Severity states it, **which scopes it to the assigned fix set**, so a recurrence of +one already validly dismissed **stays resolved on the terms the closure ordering sets** and owes +neither a second dismissal nor a repair, while a recurrence failing any of them is an ordinary +fresh finding; +the hold stands until +its answers are given, and **what the answer does is the closure ordering's**. +**A pass is credited clean or not on its own findings**, as that ordering defines cleanliness; +surfacing a finding that carries a scope-stop trigger is what withholds the credit, and no +credit is withheld for surfacing alone. Reading this as "stop instead of fixing" would put the +exit in competition with the rule that every **in-set** Blocker and Major resolves, and then +nothing could satisfy both. **Every pass report states three things about the floor**, from pass 1 onward: the derived floor, the risk and security values read, and the cited stories they were read @@ -487,9 +918,14 @@ demanding what an earlier pass had removed. Those three lines expose **five tells**: the finding count rising rather than falling; the Blocker count failing to fall; findings clustering on the **instrument** rather than on product behaviour; findings clustering on **prose about** either; and a require↔withdraw -pair. **Any two present makes stop-and-surface mandatory, not discretionary** — report the -tells and hand the decision to the user, and the "clearly stuck" reading above is not a -precondition for it. A loop can be worth stopping long before it plateaus. +pair. **Any two present makes stop-and-surface mandatory, not discretionary** — read **after** the +clean-completion branch of the closure ordering, which outranks it **by taking the pass to the +closing act and only then** — and you report the tells and hand the decision to the user, and the "clearly stuck" +reading above is not a precondition for it. A loop can be worth stopping long before it plateaus. + +**What the answer does** is the closure ordering's, which is where this stop's place among the +suspensions and what its answer produces are both stated. + **The two rules above do not compete**, and neither overrides the other: the absorb rule decides whether *a finding* is inside this loop's scope, this reading decides whether *the loop* can still converge. A small correction-of-a-correction that stays inside the assigned fix @@ -539,8 +975,8 @@ because the response stops carrying the findings at all. Append to the gate prom > Severity is one of exactly: BLOCKER | MAJOR | MINOR | NIT — no other token. > Every line before the terminator is exactly one finding line — no blank lines, > headings, prose or wrapped continuations. End the file with a final line reading -> exactly `END OF FINDINGS ( total)`, `` being the number of finding lines. A -> clean pass is the single body line `NO FINDINGS` with `END OF FINDINGS (0 total)`. +> exactly `END OF FINDINGS ( total)`, `` being the number of finding lines. +> A **clean findings file** is the single body line `NO FINDINGS` with `END OF FINDINGS (0 total)`. > > Then reply with ONLY one line per branch — ` | pass

| findings | ` > — or `INCOMPLETE | | ` if you could not write the file. An unwritten @@ -627,8 +1063,9 @@ source while the cycle runs, and it is the one the candidate rules above apply t may be present and the run must decide which, if any, is its own. **History is the source once the cycle's own commit exists**, and there is no search there: the cycle is reading **its own commit body**, so kind and artifact are settled by which commit is being read, and the nonce is -taken from the provenance line and the curve, which must agree. A Gate-A cycle mid-run has no -such commit and therefore has only the working record. Recovering a single candidate from +taken from the provenance line and the curve, which must agree. A Gate-A cycle has such a commit only once its own closing commit exists — an already-committed +revision of the reviewed text is not one, carrying no provenance line and no curve for a nonce to +be taken from — so mid-run it has only the working record. Recovering a single candidate from **either** keeps identity **as far as the field can distinguish cycles** — two cycles sharing a nonce are one cycle to it. **No candidate, disagreeing sources, or more than one candidate → no identity: start a new cycle**, which costs @@ -762,26 +1199,35 @@ Add `/.context/codex-reviews/` to `.gitignore` — that entry specifically, not **spec** right after brainstorming (before `writing-plans`), then on the **plan** before `executing-plans`/`subagent-driven-development` — catching a spec flaw before it's baked into the plan. Tool: `mcp__codex__exec` (raw; - reviews the TEXT you pass, not the git tree). Use ONE broad prompt, re-run it - each pass over the revised artifact (don't narrow per-dimension; new findings - surface because the artifact changes between passes). The prompt MUST open with + reviews the TEXT you pass, not the git tree). Use ONE broad prompt: **its review question stays the same every pass, while the artifact text it + carries and the dimensions it asks for are always the current ones** — the lens sets the profiles + section derives are recomputed from the current profile and cited set each pass and appended, since + a profile or cited-set change changes what is owed. Re-running it over an **unrevised** artifact + is legitimate wherever no repair is owed, an edit made to justify a pass being no reason to run + one. Don't narrow per-dimension: new findings surface because the artifact changed, because an + answer given since the last pass changed what the rules require of it, or because a broad prompt + reaches what the last reading did not. The prompt MUST open with *"Use the superpowers:brainstorming skill to review this spec,"* (say "plan" on the plan run), then ask Codex to check it against our settled decisions and surface **contradictions/inconsistencies, missing requirements, unhandled state/edge/error/empty/concurrent paths, and risks to the Key Invariants (@AGENTS.md) — plus anything else** (coverage floor, not a cage). Append the intent + artifact text + which invariants it touches. Ask for **every** finding - with severity and confidence — you filter to Blocker/Major downstream, Codex - never does, because a model told to report only high severity drops real - findings silently. Ask for one line per finding and a literal `NO FINDINGS` - when a pass is clean — the explicit clean signal is what lets you exit the loop: + with severity and confidence — **you filter to Blocker/Major for what must be repaired and read + every line for everything else**, Codex never filters, because a model told to report only high + severity drops real findings silently. + Ask for one line per finding and a literal `NO FINDINGS` when a pass found none — that explicit + signal is what lets a pass be read as clean without inspecting it, and a pass carrying only Minors + **and no scope-stop trigger** is clean too and could never produce that file: ``` MAJOR | high | §3 "Retry policy" | retry count unbounded | a poisoned job loops forever | cap at 5, then dead-letter NO FINDINGS ``` - Each pass: validate, revise, re-run. Before each read pass, settle mechanically what the + Each pass: validate, revise **where a repair is required**, and re-run **where the closure + ordering selects its continue branch** — where it selects a suspension instead, the answer comes + first and that ordering says what the answer produces. Before each read pass, settle mechanically what the artifact asserts and a machine can decide without side effects — cited paths, quoted passages, stated counts, the syntax of standalone fenced blocks — because a read pass spends expensive judgement on what a parser settles in seconds and misses it anyway, @@ -805,8 +1251,11 @@ Add `/.context/codex-reviews/` to `.gitignore` — that entry specifically, not handed both artifacts could catch it — which is why this is a rule about what you commit, not a claim about what the gates detect. - Same coverage rule as Gate A: put "report every finding with severity and confidence; say - `NO FINDINGS` if clean" in `additionalContext`, with the same one-line format. + Same coverage rule as Gate A: put "report every finding with severity and confidence; write + `NO FINDINGS` only when the branch found none" in `additionalContext`, with the same one-line + format. **You filter to Blocker/Major for what must be repaired, and read every line for + everything else** — cleanliness, the scope triggers, the assigned fix set and the loop-health + readings all take Minor and Nit lines. Codex never filters. **Standing lens, every Gate-B call: "which existing statements does this diff falsify?"** A change makes sentences wrong in files it never touches. Checks scoped to the edited @@ -855,10 +1304,8 @@ the reviewer, the other obliges the author. it; risk's *abuse* and security's *abuse paths* are **one lens carrying both labels**, not two questions. -Lenses are **different questions, not more passes** — they change what a pass asks, never -how many a cycle owes. The Blocker/Major filter, the file-first findings protocol and the -clean-final-pass rule are unchanged. The floor is not among them: it is no longer a fixed -number but derives from the profile and the cited set. +Lenses change what a pass asks, never how many a cycle owes: **the lens sets** leave every other +rule in this section alone. **Reading the profile — five cases, five answers:** 0. **The ordinary case**: every cited story is readable and its profile resolves → derive the @@ -929,10 +1376,12 @@ uses a fixture that never reaches the branch it covers reports success because o wired, not because the thing it checks succeeded. **The evidence entry lives in the commit body** (see Mechanics), carries the **story path -and the named evidence but not the mode value**, and is **revalidated before every Gate-B -re-review and before the cycle-closing amend** — a fix changes the diff even when the -profile sits still. If revalidation changes the entry, the clean pass no longer covers what -is being committed: fix, re-review, close on the entry that pass validated. +and the named evidence but not the mode value**, and is **revalidated before every Gate-B re-review and before the commit its closing act +produces** — a fix changes the diff even when the profile sits still. If revalidation changes the entry, the pass was read against an entry that no longer stands: **fix the +entry and re-review on it**, and **which branch the pass takes meanwhile is the closure ordering's, +read there in full** — this paragraph states the repair and never the branch. The pass +that follows is read by that ordering like any other and closes only if it reaches closure, on the +entry revalidated for it. **Every Gate-B call and re-review carries the path of every cited story**, so the reviewer reads each profile itself, **plus the current evidence entry, quoted verbatim, for each @@ -954,9 +1403,11 @@ directions; an agent never moves it alone. On confirmation, correct the header a one profile-log line. Any axis change **voids every prior override**, raised or lowered, and `+abuse-path` follows the current security value. Passes already run under the lower profile **keep counting** toward the floor; only the **final clean pass** must run under -the current profile. Inside an active Gate-B cycle, fold the edit into the active `WIP:` -snapshot by amend — a non-`WIP` commit reads to the hook as the cycle closing and would -discard the accumulated passes. +the current profile. Inside an active Gate-B cycle, the edit must end up **in the content the next review reads** — +folded into the active `WIP:` snapshot by amend where that snapshot is the tip, and otherwise +reviewed as its own change, since no amend reaches a snapshot a stray commit has made an ancestor +and this section prescribes no operation that does. A non-`WIP` commit reads to the hook as the +cycle closing and would discard **the hook's count of** the accumulated passes. **While a gate is running, the floor derives from the current profile at each pass.** Passes already run keep counting; closing requires the floor as currently derived. These @@ -986,13 +1437,21 @@ header changed mid-call, or whether the lens sets were appended. This is instruc like the rest of §5; the detection is a reader comparing the pass against the story. ### Mechanics (reference) -- **Severity:** Blocker (wrong/unsafe/breaks invariant) · Major (design flaw → - rework) → both must resolve. Minor · Nit → collect, never iterate. +- **Severity:** Blocker (wrong/unsafe/breaks invariant) · Major (design flaw → rework) → both + must resolve, **for every finding in the assigned fix set as the absorb paragraph computes it**. + **A finding is resolved by a repair or by a validated dismissal** — the author's judgement, + carrying the one-line why this section already requires, that the finding is not true of the + artifact. A dismissal does not rewrite the pass that found it and the later clean pass is still + owed; **a dismissal is not a decline**, a dismissal saying the finding is false and a decline + being the user's decision that a **true** finding stays outside the set. Minor · Nit → collect; + **their severity buys no repair round and no further pass.** Where accepting one into the + assigned fix set costs a pass, that cost is the **set change's** and is stated at the absorb + paragraph, not this severity's. **Deciding severity — one procedure. The subject list is illustration, not a second rule.** Name what in the system consumes this text — whatever *acts* on it — and the - decision that act takes differently if the text is wrong. Both are required. If you - cannot name both, the finding is Minor or below: collect, never iterate. + decision that act takes differently if the text is wrong. Both are required. If you cannot name both, the finding is Minor or below: collect; **its severity buys no repair + round and no further pass**, and any pass a later scope decision costs is that decision's. The exclusions are contract, not commentary. The reader must consume the text in the system's *operation*, not in reviewing it — the review pass raising the finding is not @@ -1013,10 +1472,31 @@ like the rest of §5; the detection is a reader comparing the pass against the s This is the finding-level analog of the path-level prose exemption: one principle at two granularities — text that *describes* the product versus text that *is* the product. - **How this demotion bears on the loop-health measures — the per-pass counts, the finding - clusters and the stop thresholds — is not settled here, and this change does not settle it. - Until it is, a pass whose outcome would turn on that question reports the question and - stops rather than deciding it** — the same answer any unresolved gate question gets. + **The demotion changes what a cycle must resolve, never what it observes.** **Every loop-health + reading observes the findings as the reviewer produced them, before the ceiling is applied** — + so a demoted finding still counts in the finding total and in its cluster, and a Blocker demoted + to Minor is still a Blocker to the curve and still regeneration to the clearly-stuck reading's + third condition. Where such a reading uses severity at all it takes the **reader-normalized + pre-ceiling severity**, which is what the Reader paragraph above already produces from the + findings file — case-folded, and a non-empty unrecognized token read as `MAJOR` — never the raw + token, so a finding written `IMPORTANT` enters the curve as a Major. **The ceiling is applied + after that and only to what the cycle owes**: cleanliness and the resolve duty read the + effective severity, so the same finding can be a Major to the curve and a Minor **at effective + severity**, and that difference is the point rather than a discrepancy. It stays in the fix set + either way; the ceiling moves what the cycle owes for it and never whether it is in. The line is **what the cycle owes + versus what it observes about itself**, which is why no list of readings has to be kept complete + here. Two reasons for the split. The curve's **three numeric series** must stay derivable from + the validated findings files alone wherever those files remain available — the rest of the curve + is not and does not claim to be, its cycle field, pass ranges and model identifiers coming from + elsewhere, and an unrecoverable count being written `?` exactly as the standing grammar allows — + the finding total counts finding lines and the Blocker and Major series count the lines whose + normalized severity is each, which is the only thing that makes a self-reported curve checkable; + the subject clusters use no severity at all, being a judgement per finding that no count + reproduces. And the demotion is the author's judgement about **the finding's repair severity**, + never about which findings the fix set contains — a separate predicate the absorb paragraph + defines, and one this must not be read as touching; a loop spending passes on findings the + author keeps demoting is exactly what the prose-cluster tell exists to surface, and lowering the + counts by that same judgement would hide it. - **Tool routing:** docs (spec/plan, incl. code snippets) → `mcp__codex__exec`; implemented diff → `mcp__codex__review`. Never `review` a doc — it reads the git range, not the text. @@ -1026,12 +1506,22 @@ like the rest of §5; the detection is a reader comparing the pass against the s pre-commit, `baseSha` = HEAD is an empty range (HEAD..HEAD) — make a WIP commit and set `baseSha` to its parent. **Name that commit `WIP: …`** — the hook treats a `wip`-prefixed commit message as cycle-internal, so it neither fires a Gate-B STOP - nor resets your pass counters. A pre-review snapshot named anything else reads as a - real commit and closes the cycle, discarding the passes you just accumulated. - **Finishing the cycle:** after the final clean pass, close it with - `git commit --amend -m ""` — that replaces the WIP commit, and the hook - reads the amend as the real cycle-closing commit. If several WIP snapshots piled up, - `git reset --soft ` first, then commit once. Amend rather than a + nor resets your pass counters. A pre-review snapshot named anything else reads to the hook as a real commit: the hook treats the + cycle as closed and **discards its count of the passes you just accumulated**, while the cycle + itself stays open until the closure ordering's conditions hold. **What the hook loses is its counter state**, and that + counter is not what makes a pass valid — so the reminder now understates what you hold, and no + close was intended or made. **What such a commit does to the repository, and what the closing act + then owes, is in the Gate-B closure paragraph**, not here. + **Finishing the cycle:** **when a Gate-B cycle's closing act is performed is the closure + ordering's, stated there entire**; this section gives only the operation. Close it with + `git commit --amend -m ""`, which replaces the WIP commit; **where a `WIP:` snapshot + would survive the amend** — several piled up, or a stray non-amending commit made one an + ancestor — **`git reset --soft `, then commit once instead**. This section is the only + place either shape is defined. **The hook treats any + non-`WIP` commit *attempt* as a Gate-B boundary and clears its state even where the command + fails**, so a failed closing act leaves that counter cleared — a fact about the + counter and not about the cycle. This section states the operation and never whether the cycle may + close. Amend rather than a follow-up commit for two reasons: a `WIP: …` commit left in history defeats the naming convention it exists for, and a follow-up commit has nothing to commit when the review produced no fixes. @@ -1080,18 +1570,48 @@ like the rest of §5; the detection is a reader comparing the pass against the s different way, the last of them demonstrably so, and the rule for a claim needing a fourth correction is to delete it. Whoever needs to know why reads the file and the hook. - **These records are one contract, and a partial adoption breaks it.** The nonce, the slot - naming, the provenance line, the curve, this carry rule **and the unknown-start activation - semantics that say what a cycle owes when its starting rules cannot be established** depend on - one another, and the requirement is that the adopted definitions **agree**, not merely that all - of them are present: a curve - without a cycle field cannot be attributed, a slot rule without a nonce has nothing to key on, - and a carry rule naming records a project does not produce is inert. **A project whose text - carries some of them and not others, or carries all of them in versions that disagree, stops - and has a human complete, revert or reconcile the adoption before running a gate under it** — - disagreement is the harder case and gets the same stop, because a project holding two - definitions of a record has no single answer to what it owes — the same answer, and for the same reason, as a partial - adoption of the floor rule. + **These rules and records are one contract, and a partial adoption breaks it.** The nonce, the + slot naming, the provenance line, the curve, the carry rule, the unknown-start activation + semantics **and the closure ordering together with every rule it reads** depend on one another, + and the requirement is that the adopted definitions **agree**, not merely that all of them are + present. **Membership is decided by a test a reader can apply to the text in front of them, with + no list to consult, and the test reads what a rule states rather than what changing it would do: + a live rule belongs to this contract when what it says **defines the validity of an input the + closure ordering reads, or how that input is read**, which branch a pass takes, what a hold is or what discharges it, **what a suspension asks, or what state its answer or an + incomplete closing act produces**, **what a gate's closing act is**, **which version of these rules + governs a cycle**, + whether a cycle may close **or may terminate without running a pass at all, the Gate-B triviality + skip being the one such route and its eligibility test therefore a member**, or the production, + identity or transport of **any §5 cycle record, required or optional** — §5 entire and not the + Mechanics subsection this paragraph sits in, record duties being stated in both, and the optional + companions' slot rules carrying the nonce that keeps sibling cycles apart.** **Read it on the sentence, never on the section the sentence sits in.** + A sentence is a member when **it itself** fixes one of those things — what counts as a valid + finding line, which files or records are owed, what ends a hold. It is not a member when it only + shapes what a review produces, as the choice of reviewer, the lens set and **prompt wording that only frames the + review question** do: those change the findings without deciding what a finding *is* or what the + ordering may do with one. **Prompt wording that fixes a valid input is a member**, the gate-prompt + sentences defining a finding line and the `NO FINDINGS` signal being exactly that. **No paragraph is exempt as a paragraph** — a sentence inside a routing or prompt paragraph + that fixes a valid input or an owed file is a member, and a sentence anywhere that only influences + the findings is not. The examples follow the test; they do not stand in for it. + Asking instead what an imagined edit would do decides nothing, because + any rule can be edited into deciding a branch and none decides one when edited cosmetically, so + membership would follow the edit a reader pictured rather than the text in front of them. The + last clause is why the squash carry belongs: it moves no pass and + decides no branch, and a record that does not survive the merge is unreachable from the squash + commit and from `main`'s history. A curve without a cycle field cannot be reliably told from + another cycle's in every multi-cycle context — kind and surrounding context sometimes separate + them, which is why the standing rule calls missing attribution a limitation rather than a + disqualification — a slot rule without a nonce cannot keep sibling cycles + apart — the bare names staying reserved for the legacy single-cycle case they already serve — a + carry rule naming records a project does not produce is inert, and a clean predicate without the + fix-set boundary it reads decides membership by accident. + **A project whose text carries some of them and not others, or carries all of them in versions + that disagree, stops and has a human complete, revert or reconcile the adoption before running a + gate under it.** **The test classifies sentences that are present, and completeness is not among + what it establishes**: where an adoption drops a member together with every sentence that would + refer to it, what remains reads as coherent and nothing in it marks the absence. **That bounds + what this text lets a reader detect, never what the rule obliges** — a partial adoption is a stop + however it becomes known, and learning of it from outside this text is learning of it. **On squash-merge, copy every evidence entry, every human-exception record, the provenance lines, the curves and any skipped cycle's skip record TOGETHER WITH THE SKIP REASON IT POINTS AT in the squash range into the squash body — a skip record carried without its reason is a pointer into a body the squash has made unreachable — the squash commit is the only body the merge carries into `main`'s history, so anything left behind is unreachable from it.** @@ -1139,8 +1659,9 @@ like the rest of §5; the detection is a reader comparing the pass against the s contributing to a split logical pass is listed**, joined by `+`, since recording one of two is the same loss as recording none. - **Majors are recorded as well as Findings and Blockers**, because the severity rule moves the - Blocker/Major line rather than the total, so totals and Blockers alone could not show even a + **Majors are recorded as well as Findings and Blockers**, because the three series are read + before the ceiling and the mix among them is what a later reader compares; the ceiling moves what + a cycle owes and leaves these counts alone, so totals and Blockers alone could not show even a change in the mix. **Subject categories are deliberately not recorded** — they are a judgement per finding rather than a count, and the findings files carry the material. @@ -1189,8 +1710,9 @@ like the rest of §5; the detection is a reader comparing the pass against the s Accepted because: ``` - **Which commit:** an ungated change records it in that commit; a Gate-A cycle in the spec or - plan commit; a Gate-B cycle in the WIP commit, restated by the closing amend. Several records + **Which commit:** an ungated change records it in that commit; a Gate-A cycle in the commit its + closing act uses; a Gate-B cycle in the WIP commit, restated by the commit its closing act + produces. Several records accumulate; order means nothing. **A decision made after its commit closed** — during PR review, say — goes in whichever of @@ -1214,9 +1736,11 @@ like the rest of §5; the detection is a reader comparing the pass against the s was already the human's to make about something genuinely optional. It is **never** the answer to a below-floor pass, an unclean final pass, a `STOP and surface`, a Gate-A or Gate-B obligation, or a profile-derived evidence requirement — and more generally **it authorizes - nothing that any mandatory rule in this file or in `AGENTS.md` requires.** Those have their - own terminal actions and this paragraph changes none of them: on a STOP you still stop, and - neither a human's assent nor this record lets an agent close or continue a cycle. + nothing that any mandatory rule in this file or in `AGENTS.md` requires.** Those have their own terminal actions and this paragraph changes none of them: on a STOP you + still stop, and **neither a human's general assent nor this record** lets an agent close or + continue a cycle. **The answers a suspension asks for are not assent of that kind**: they are the + answers the closure ordering prescribes, and both which answers those are and what they produce are + stated there. **"Mandatory" is not limited to this file.** A rule in `AGENTS.md`, a project doc, CI, a branch policy or the platform is equally out of reach — under **Wait for**, diff --git a/plugins/dev-workflow/hooks/codex-gate.sh b/plugins/dev-workflow/hooks/codex-gate.sh index 396bd44..b7ae064 100755 --- a/plugins/dev-workflow/hooks/codex-gate.sh +++ b/plugins/dev-workflow/hooks/codex-gate.sh @@ -905,7 +905,7 @@ case "$event" in cmd=$(input_field command) if is_commit "$cmd"; then if is_wip_commit "$cmd"; then - note "WIP commit — cycle-internal, per $policy: this exists so mcp__codex__review has a non-empty range to read (baseSha = this commit's parent). Gate B is not evaluated here and your pass counters are preserved. Run the review against this commit, then make the real commit when your final pass is clean." "ℹ WIP commit (Codex cycle preserved)" + note "WIP commit — cycle-internal, per $policy: this exists so mcp__codex__review has a non-empty range to read (baseSha = this commit's parent). Gate B is not evaluated here and your pass counters are preserved. Use this commit as the review range; whether this cycle runs a review now, and when its closing act may be performed, are both $policy's closure ordering's, read there in full." "ℹ WIP commit (Codex cycle preserved)" else # Docs-only commits (spec/plan .md files) carry no code diff, # so Gate B (mcp__codex__review reviews a code diff) cannot apply — emit a @@ -930,7 +930,7 @@ case "$event" in # preserves the older, now-stale fingerprint, which reaches the STALE # branch below, not this one. The message names the absent FINGERPRINT, # not an absent review. - note "STOP — Codex Gate B not satisfied: no fingerprint is recorded for this cycle — either no mcp__codex__review has run, or the last one's fingerprint could not be written or read back. Per $policy you MUST reach a minimum of $floor passes per cycle. Run Gate B (mcp__codex__review) now; if this repeats, check that .context/ and the state file inside it are readable and writable, and if the file exists but is unreadable or empty, delete it and run a fresh pass." "⚠ Codex Gate B: no recorded review" + note "Codex gate state: no fingerprint is recorded for this cycle. The hook cannot tell why — no mcp__codex__review has run, the last one's fingerprint could not be written or read back, or a non-WIP commit attempt cleared it while the cycle itself stayed open. What this cycle does next, the floor it owes included, is $policy's closure ordering's, read there entire, and this reminder decides none of it. If this repeats, check that .context/ and the state file inside it are readable and writable; if the file exists but is unreadable or empty, delete it — which restores no passes, and lets the next pass the ordering permits record a fingerprint." "⚠ Codex Gate B: no recorded fingerprint" # `unavailable` on EITHER side is never a match: an uncomputable fingerprint # must read as unverified, and two of them must not cancel out. elif [ "$current" = unavailable ] || [ "$reviewed" = unavailable ] || @@ -942,9 +942,9 @@ case "$event" in # that the tree changed — under a repeated computation failure nothing # changed, and under a failed state write the content may be exactly what # was reviewed. - note "STOP — Codex Gate B not satisfied: the hook cannot confirm that the content you are about to commit is the content mcp__codex__review last saw ($passes recorded pass(es) this cycle). Usually that means the working tree or the index changed since the review. It can also mean you only staged already-reviewed content — the bytes are fine, but the hook cannot tell staging from editing; that this hook was upgraded and the recorded fingerprint uses the older format (see CHANGELOG); or that the fresh fingerprint could not be computed or could not be stored. Run Gate B (mcp__codex__review) now — one clean pass is the complete remedy for the staging and post-upgrade cases too. If a fresh pass leaves this unchanged with nothing edited in between, the fault is in the machinery rather than the code: check that .context/ is writable, that TMPDIR is writable, that a checksum tool (shasum, sha1sum or cksum) runs, that git status works, and that the disk is not full — then run one more pass to record a usable fingerprint. Per $policy you MUST re-review after every fix." "⚠ Codex Gate B not satisfied (cannot confirm review)" + note "Codex gate state: the hook cannot confirm that the content you are about to commit is the content mcp__codex__review last saw ($passes recorded pass(es) this cycle). Usually that means the working tree or the index changed since the review. It can also mean you only staged already-reviewed content — the bytes are fine, but the hook cannot tell staging from editing; that this hook was upgraded and the recorded fingerprint uses the older format (see CHANGELOG); or that the fresh fingerprint could not be computed or could not be stored. A fresh Gate-B pass is the complete remedy for the staging and post-upgrade cases too, and when this cycle may run one is $policy's closure ordering's, read there entire. If a fresh pass leaves this unchanged, check the worktree and the index first — the fingerprint moves when either does, staging included, and another hook can stage during the commit attempt. Where neither changed, the fault is in the machinery rather than the code: check that .context/ is writable, that TMPDIR is writable, that a checksum tool (shasum, sha1sum or cksum) runs, that git status works, and that the disk is not full — a store that fails again leaves the hook unable to confirm a fingerprint and may return either fingerprint-state diagnosis. Per $policy you MUST re-review after every fix." "⚠ Codex Gate B: cannot confirm reviewed content" elif [ "$passes" -lt "$floor" ]; then - note "Codex Gate B floor NOT met: only $passes/$floor mcp__codex__review pass(es) since the last commit. Per $policy the review is a LOOP with a hard minimum of $floor passes — run more (the ONLY early exit is a pass that returned zero findings), or proceed only if $policy's skip rule applies to this change — if you cannot locate and check that rule, run the remaining passes." "⚠ Codex Gate B below floor ($passes/$floor)" + note "Codex Gate B floor NOT met: only $passes/$floor mcp__codex__review pass(es) since the last commit. Per $policy the review is a LOOP with a hard minimum of $floor passes, and what this cycle does next is that policy's closure ordering's, read there entire; $policy's skip rule decides only whether a cycle runs at all, never whether one already running may stop short." "⚠ Codex Gate B below floor ($passes/$floor)" else # Distinguish the two counts (Finding 9): the cycle total includes passes # made BEFORE later edits, so they carry a different fingerprint. @@ -953,7 +953,7 @@ case "$event" in # establish that Codex read these bytes (spec §7, and the review-range row # in todos.md). The stronger phrasing was here and was removed; do not # restore it as a clarity improvement. - note "Codex Gate B: $passes/$floor pass(es) this cycle, of which $fresh cover the CURRENT content fingerprint (unchanged since that review). The floor counts the cycle; only the fresh pass(es) carry the same fingerprint as what you are committing. Per $policy, commit only if your final pass was clean — no new Blocker/Major." "✓ Codex Gate B satisfied ($passes/$floor cycle, $fresh on current fingerprint)" + note "Codex Gate B: $passes/$floor pass(es) this cycle, of which $fresh cover the CURRENT content fingerprint (unchanged since that review). The floor counts the cycle; only the fresh pass(es) carry the same fingerprint as what you are committing. Per $policy, commit only if your final pass was clean and every other closure condition holds, both as it defines them." "✓ Codex Gate B hook checks passed ($passes/$floor cycle, $fresh on current fingerprint)" fi fi fi @@ -964,13 +964,13 @@ case "$event" in superpowers:executing-plans | superpowers:subagent-driven-development) passesA=$(read_count "$countA_file") if [ "$passesA" -lt "$floor" ]; then - note "Codex Gate A floor NOT met: only $passesA/$floor mcp__codex__exec pass(es) on this spec/plan. Per $policy Gate A is a LOOP with a hard minimum of $floor passes (start each instruction with the superpowers:brainstorming directive; the ONLY early exit is a pass that returned zero findings). Gate A has no content check behind it — this floor is the only thing keeping the spec review honest. Run more passes before executing." "⚠ Codex Gate A below floor ($passesA/$floor)" + note "Codex Gate A floor NOT met: only $passesA/$floor mcp__codex__exec pass(es) on this spec/plan. Per $policy Gate A is a LOOP with a hard minimum of $floor passes (start each instruction with the superpowers:brainstorming directive; the ONLY early exit is a pass that returned zero findings). Gate A has no content check in this hook; what the gate itself requires of the reviewed artifact is stated in $policy and is instruction-backed. What this cycle does next is that policy's closure ordering's, a further pass being one of its answers and not the only one." "⚠ Codex Gate A below floor ($passesA/$floor)" else # Deliberately weaker wording than Gate B (Finding 12): countA counts # mcp__codex__exec CALLS, bound to no artifact. Hashing the artifact would # be wrong — a spec is SUPPOSED to change between passes — so the hook # cannot verify what was reviewed, and must not imply that it did. - note "Codex Gate A: $passesA/$floor mcp__codex__exec pass(es) on this spec/plan — floor met by COUNT ONLY. The hook counts calls; it cannot verify what was reviewed or that findings were addressed. Proceed only if your final pass was clean — no new Blocker/Major." "✓ Codex Gate A floor met ($passesA/$floor passes, count only)" + note "Codex Gate A: $passesA/$floor mcp__codex__exec pass(es) on this spec/plan — floor met by COUNT ONLY. The hook counts calls; it cannot verify what was reviewed or that findings were addressed. Proceed only once this Gate-A cycle has closed under $policy's closure ordering, read there entire." "✓ Codex Gate A floor met ($passesA/$floor passes, count only)" fi ;; esac diff --git a/plugins/dev-workflow/hooks/codex-gate.test.sh b/plugins/dev-workflow/hooks/codex-gate.test.sh index 9c22be6..6ed4d8c 100644 --- a/plugins/dev-workflow/hooks/codex-gate.test.sh +++ b/plugins/dev-workflow/hooks/codex-gate.test.sh @@ -170,23 +170,23 @@ rev [ -s "$state" ] && pass "state holds a tree hash (non-empty)" || fail "state holds a tree hash (non-empty)" [ "$(cat "$count" 2>/dev/null)" = 1 ] && pass "review bumps pass count to 1" || fail "review bumps pass count to 1" -# 2. Below floor (1/3) -> NOT satisfied yet; reaching floor (3/3) -> satisfied +# 2. Below floor (1/3) -> below-floor reminder; reaching floor (3/3) -> hook checks passed out=$(commitpre) printf '%s' "$out" | grep -qE 'below floor|floor NOT met' && pass "1/3 passes -> below floor" || fail "1/3 passes -> below floor" printf '%s' "$out" | grep -q 'hookSpecificOutput' && pass "emits JSON additionalContext" || fail "emits JSON additionalContext" rev; rev # reach the floor: 3 passes total, tree unchanged out=$(commitpre) -printf '%s' "$out" | grep -q 'Gate B satisfied' && pass "3/3 passes, unchanged tree -> satisfied" || fail "3/3 passes, unchanged tree -> satisfied" +printf '%s' "$out" | grep -q 'Gate B hook checks passed' && pass "3/3 passes, unchanged tree -> hook checks passed" || fail "3/3 passes, unchanged tree -> hook checks passed" # 3. FINDING 1 — content-based invalidation. # (Replaces the old event-based assertion `[ ! -f state ]` after an Edit. The state # file now legitimately SURVIVES a change — it holds the reviewed hash — so the -# intent "a change means Gate B is not satisfied" is asserted at the BEHAVIOR level.) +# intent "a change means the hook cannot confirm the reviewed content" is asserted at the BEHAVIOR level.) # 3a. Edit-tool change -> stale printf 'v2\n' >> app.ts run '{"hook_event_name":"PostToolUse","tool_name":"Edit","tool_input":{"file_path":"app.ts"}}' >/dev/null out=$(commitpre) -printf '%s' "$out" | grep -q 'not satisfied' && pass "Edit-tool change -> not satisfied" || fail "Edit-tool change -> not satisfied" +printf '%s' "$out" | grep -q 'Codex gate state:' && pass "Edit-tool change -> gate-state reminder" || fail "Edit-tool change -> gate-state reminder" printf '%s' "$out" | grep -q 'cannot confirm' && pass "Edit-tool change -> reported as unconfirmed" || fail "Edit-tool change -> reported as unconfirmed" # 3b. THE MAJOR: a file changed through BASH (no Edit/Write event at all) -> stale. @@ -194,36 +194,36 @@ printf '%s' "$out" | grep -q 'cannot confirm' && pass "Edit-tool change -> repor reset_all rev; rev; rev # 3 clean passes on the current tree out=$(commitpre) -printf '%s' "$out" | grep -q 'Gate B satisfied' && pass "setup: satisfied before bash edit" || fail "setup: satisfied before bash edit" +printf '%s' "$out" | grep -q 'Gate B hook checks passed' && pass "setup: hook checks passed before bash edit" || fail "setup: hook checks passed before bash edit" printf 'sed-style in-place edit\n' >> app.ts # NO hook event fires for this out=$(commitpre) -printf '%s' "$out" | grep -q 'not satisfied' && pass "bash-modified file after review -> NOT satisfied (Finding 1)" || fail "bash-modified file after review -> NOT satisfied (Finding 1)" +printf '%s' "$out" | grep -q 'Codex gate state:' && pass "bash-modified file after review -> gate-state reminder (Finding 1)" || fail "bash-modified file after review -> gate-state reminder (Finding 1)" # 3c. Untracked new file after review -> stale git checkout -- app.ts >/dev/null 2>&1 reset_all rev; rev; rev -printf '%s' "$(commitpre)" | grep -q 'Gate B satisfied' && pass "setup: satisfied on clean tree" || fail "setup: satisfied on clean tree" +printf '%s' "$(commitpre)" | grep -q 'Gate B hook checks passed' && pass "setup: hook checks passed on clean tree" || fail "setup: hook checks passed on clean tree" printf 'new\n' > brand-new.ts out=$(commitpre) -printf '%s' "$out" | grep -q 'not satisfied' && pass "untracked new file after review -> NOT satisfied" || fail "untracked new file after review -> NOT satisfied" +printf '%s' "$out" | grep -q 'Codex gate state:' && pass "untracked new file after review -> gate-state reminder" || fail "untracked new file after review -> gate-state reminder" # 3c-bis. Untracked CONTENT counts, not just the name. `git add f && git commit` is one # Bash call, so the hook sees `f` still untracked — a name-only hash would hand that # commit a stale ✓ on edited content (invariant 3). reset_all; rev; rev; rev # re-review with brand-new.ts present, so its name is known -printf '%s' "$(commitpre)" | grep -q 'Gate B satisfied' && pass "setup: satisfied with untracked file present" || fail "setup: satisfied with untracked file present" +printf '%s' "$(commitpre)" | grep -q 'Gate B hook checks passed' && pass "setup: hook checks passed with untracked file present" || fail "setup: hook checks passed with untracked file present" printf 'edited\n' > brand-new.ts # same name, different content -printf '%s' "$(commitpre)" | grep -q 'not satisfied' && pass "edited untracked file -> NOT satisfied" || fail "edited untracked file -> NOT satisfied" +printf '%s' "$(commitpre)" | grep -q 'Codex gate state:' && pass "edited untracked file -> gate-state reminder" || fail "edited untracked file -> gate-state reminder" # ...and a file inside a NEW untracked directory too: porcelain would collapse that to # a single `dir/` entry and never hash what is in it. reset_all; rev; rev; rev mkdir -p newdir && printf 'a\n' > newdir/f.ts -printf '%s' "$(commitpre)" | grep -q 'not satisfied' && pass "new untracked dir -> NOT satisfied" || fail "new untracked dir -> NOT satisfied" +printf '%s' "$(commitpre)" | grep -q 'Codex gate state:' && pass "new untracked dir -> gate-state reminder" || fail "new untracked dir -> gate-state reminder" reset_all; rev; rev; rev printf 'b\n' > newdir/f.ts -printf '%s' "$(commitpre)" | grep -q 'not satisfied' && pass "edited file in untracked dir -> NOT satisfied" || fail "edited file in untracked dir -> NOT satisfied" +printf '%s' "$(commitpre)" | grep -q 'Codex gate state:' && pass "edited file in untracked dir -> gate-state reminder" || fail "edited file in untracked dir -> gate-state reminder" rm -rf newdir # ...and paths git does not print literally. It C-quotes non-ASCII and control @@ -232,10 +232,10 @@ rm -rf newdir for name in "café ñ.ts" "$(printf 'tab\tnewline\nname.ts')"; do reset_all; rev; rev; rev printf 'a\n' > "$name" - printf '%s' "$(commitpre)" | grep -q 'not satisfied' && pass "new exotic-path untracked file -> NOT satisfied" || fail "new exotic-path untracked file -> NOT satisfied" + printf '%s' "$(commitpre)" | grep -q 'Codex gate state:' && pass "new exotic-path untracked file -> gate-state reminder" || fail "new exotic-path untracked file -> gate-state reminder" reset_all; rev; rev; rev printf 'b\n' > "$name" - printf '%s' "$(commitpre)" | grep -q 'not satisfied' && pass "edited exotic-path untracked file -> NOT satisfied" || fail "edited exotic-path untracked file -> NOT satisfied" + printf '%s' "$(commitpre)" | grep -q 'Codex gate state:' && pass "edited exotic-path untracked file -> gate-state reminder" || fail "edited exotic-path untracked file -> gate-state reminder" rm -f "$name" done @@ -244,10 +244,10 @@ done # referents, or block forever on a link to a FIFO. reset_all; rev; rev; rev ln -s absent-a link.ts -printf '%s' "$(commitpre)" | grep -q 'not satisfied' && pass "new untracked symlink -> NOT satisfied" || fail "new untracked symlink -> NOT satisfied" +printf '%s' "$(commitpre)" | grep -q 'Codex gate state:' && pass "new untracked symlink -> gate-state reminder" || fail "new untracked symlink -> gate-state reminder" reset_all; rev; rev; rev rm -f link.ts; ln -s absent-b link.ts -printf '%s' "$(commitpre)" | grep -q 'not satisfied' && pass "retargeted untracked symlink -> NOT satisfied" || fail "retargeted untracked symlink -> NOT satisfied" +printf '%s' "$(commitpre)" | grep -q 'Codex gate state:' && pass "retargeted untracked symlink -> gate-state reminder" || fail "retargeted untracked symlink -> gate-state reminder" rm -f link.ts # NOTE: sections 24-25 cover the "could not be computed" guard for checksum, seed-copy, @@ -276,26 +276,26 @@ fi rm -f brand-new.ts reset_all; rev; rev; rev -# 3d. Reverting the tree back to the reviewed content -> satisfied again +# 3d. Reverting the tree back to the reviewed content -> hook checks pass again # (content-based, so an edit-then-undo is correctly NOT stale) out=$(commitpre) -printf '%s' "$out" | grep -q 'Gate B satisfied' && pass "revert to reviewed tree -> satisfied again" || fail "revert to reviewed tree -> satisfied again" +printf '%s' "$out" | grep -q 'Gate B hook checks passed' && pass "revert to reviewed tree -> hook checks pass again" || fail "revert to reviewed tree -> hook checks pass again" # 3e. The hook's own .context/ churn must NOT change the hash (else it never matches itself) rev # writes state files out=$(commitpre) -printf '%s' "$out" | grep -q 'Gate B satisfied' && pass ".context/ churn does not invalidate the hash" || fail ".context/ churn does not invalidate the hash" +printf '%s' "$out" | grep -q 'Gate B hook checks passed' && pass ".context/ churn does not invalidate the hash" || fail ".context/ churn does not invalidate the hash" # 3e-bis. ...including when .context/ is COMMITTED. Filtering only the untracked list # leaves tracked state in `git diff HEAD`, where the hook's own writes invalidate the -# review it just recorded -> a permanent stale STOP. The adoption marker is meant to be +# review it just recorded -> a permanent stale-fingerprint reminder. The adoption marker is meant to be # shared, so a tracked .context/ is the normal case. git add -f .context >/dev/null 2>&1; git commit -qm "track .context" >/dev/null 2>&1 reset_all rev; rev; rev -printf '%s' "$(commitpre)" | grep -q 'Gate B satisfied' && pass "tracked .context/ state does not invalidate the hash" || fail "tracked .context/ state does not invalidate the hash" +printf '%s' "$(commitpre)" | grep -q 'Gate B hook checks passed' && pass "tracked .context/ state does not invalidate the hash" || fail "tracked .context/ state does not invalidate the hash" rev # more churn against the committed state -printf '%s' "$(commitpre)" | grep -q 'Gate B satisfied' && pass "tracked .context/ churn stays satisfied" || fail "tracked .context/ churn stays satisfied" +printf '%s' "$(commitpre)" | grep -q 'Gate B hook checks passed' && pass "tracked .context/ churn still passes the hook checks" || fail "tracked .context/ churn still passes the hook checks" # 3e-ter. GAP 1 — the INDEX-tree component must exclude .context/ too, not just the # diff-HEAD and worktree-tree components (each guarded by its own `:(exclude)` @@ -309,11 +309,11 @@ printf '%s' "$(commitpre)" | grep -q 'Gate B satisfied' && pass "tracked .contex # be the thing that (in)validates — only the .context staging can. printf 'side\n' > sidefile.ts; git add sidefile.ts >/dev/null 2>&1 rev; rev; rev -printf '%s' "$(commitpre)" | grep -q 'Gate B satisfied' && pass "setup: satisfied with sidefile.ts staged" || fail "setup: satisfied with sidefile.ts staged" +printf '%s' "$(commitpre)" | grep -q 'Gate B hook checks passed' && pass "setup: hook checks passed with sidefile.ts staged" || fail "setup: hook checks passed with sidefile.ts staged" rev # hook writes fresh state into .context/ git add -f .context >/dev/null 2>&1 # stage the hook's own churn (sidefile.ts untouched) out=$(commitpre) -printf '%s' "$out" | grep -q 'Gate B satisfied' \ +printf '%s' "$out" | grep -q 'Gate B hook checks passed' \ && pass "staged .context churn does not invalidate (index-tree .context exclusion)" \ || fail "staged .context churn does not invalidate (index-tree .context exclusion)" git rm -q --cached sidefile.ts >/dev/null 2>&1; rm -f sidefile.ts @@ -321,7 +321,7 @@ git reset -q -- .context >/dev/null 2>&1 # unstage; real index back to HEAD for # ...while a real code change is still caught printf 'code change\n' >> app.ts -printf '%s' "$(commitpre)" | grep -q 'not satisfied' && pass "tracked .context/: real code change still invalidates" || fail "tracked .context/: real code change still invalidates" +printf '%s' "$(commitpre)" | grep -q 'Codex gate state:' && pass "tracked .context/: real code change still invalidates" || fail "tracked .context/: real code change still invalidates" git checkout -- app.ts >/dev/null 2>&1 git rm -rq --cached .context >/dev/null 2>&1; git commit -qm "untrack .context" >/dev/null 2>&1 reset_all; rev; rev; rev @@ -330,24 +330,24 @@ reset_all; rev; rev; rev # already-reviewed content DOES invalidate. The hash covers the index tree, and # `git add` changes it. The bytes that would be committed are unchanged, so this is # a false invalidation — accepted under invariant 2 ("loose in the firing -# direction"), and the STOP message explains that staging alone can cause it. +# direction"), and the stale-fingerprint message explains that staging alone can cause it. # This test previously asserted the OPPOSITE as though it were a principle; the # behaviour was never decided, it fell out of an implementation choice. reset_all printf 'reviewed change\n' >> app.ts rev; rev; rev # 3 passes covering the modified (unstaged) tree -printf '%s' "$(commitpre)" | grep -q 'Gate B satisfied' && pass "setup: satisfied on unstaged change" || fail "setup: satisfied on unstaged change" +printf '%s' "$(commitpre)" | grep -q 'Gate B hook checks passed' && pass "setup: hook checks passed on unstaged change" || fail "setup: hook checks passed on unstaged change" git add app.ts >/dev/null 2>&1 # staging only — no content change -printf '%s' "$(commitpre)" | grep -q 'not satisfied' \ - && pass "staging a reviewed tracked file -> NOT satisfied (spec §2 decision)" \ - || fail "staging a reviewed tracked file -> NOT satisfied (spec §2 decision)" -# The old trailing assertion ("untracked file on a staged tree -> not satisfied") is +printf '%s' "$(commitpre)" | grep -q 'Codex gate state:' \ + && pass "staging a reviewed tracked file -> gate-state reminder (spec §2 decision)" \ + || fail "staging a reviewed tracked file -> gate-state reminder (spec §2 decision)" +# The old trailing assertion ("untracked file on a staged tree -> gate-state reminder") is # GONE on purpose: once staging alone invalidates, it passes regardless of the untracked # file and tests nothing. Test 3c already covers untracked content. git reset -q >/dev/null 2>&1; git checkout -- app.ts >/dev/null 2>&1 reset_all; rev; rev; rev -# 4. Gate A exec must NOT satisfy Gate B (separate state) +# 4. Gate A exec must NOT count toward Gate B (separate state) reset_all execp [ ! -f "$state" ] && pass "exec does not set Gate B" || fail "exec does not set Gate B" @@ -435,7 +435,7 @@ rm -f "$off" # re-enable reset_all rev; rev; rev out=$(commitpre) -printf '%s' "$out" | grep -q 'Gate B satisfied' && pass "re-enable sees same counting semantics as gate-on, not evidence of review" || fail "re-enable sees same counting semantics as gate-on, not evidence of review" +printf '%s' "$out" | grep -q 'Gate B hook checks passed' && pass "re-enable sees same counting semantics as gate-on, not evidence of review" || fail "re-enable sees same counting semantics as gate-on, not evidence of review" # 14. Below-floor reminder shows N/floor reset_all @@ -444,13 +444,13 @@ out=$(commitpre) printf '%s' "$out" | grep -q '2/3' && pass "below-floor reminder shows N/3" || fail "below-floor reminder shows N/3" # 14b. Docs-only commit -> gentle N/A note; mixed and undeterminable commits still fire. -reset_all # state ABSENT -> would normally STOP on a code commit +reset_all # state ABSENT -> would normally emit the no-fingerprint reminder on a code commit mkdir -p docs printf 'spec\n' > docs/plan.md; printf 'readme\n' > NOTES.md git add docs/plan.md NOTES.md >/dev/null 2>&1 out=$(commitpre) printf '%s' "$out" | grep -q 'docs-only commit' && pass "docs-only staged -> N/A note" || fail "docs-only staged -> N/A note" -printf '%s' "$out" | grep -q 'STOP' && fail "docs-only must not STOP" || pass "docs-only does not STOP" +printf '%s' "$out" | grep -q 'Codex gate state:' && fail "docs-only must not emit the gate-state reminder" || pass "docs-only does not emit the gate-state reminder" printf 'code\n' > extra.ts; git add extra.ts >/dev/null 2>&1 out=$(commitpre) printf '%s' "$out" | grep -q 'Gate B' && pass "mixed staged -> Gate B fires" || fail "mixed staged -> Gate B fires" @@ -501,9 +501,9 @@ printf '%s' "$out" | grep -qE 'below floor|floor NOT met' && pass "1/3 exec -> G execp execp out=$(run '{"hook_event_name":"PreToolUse","tool_name":"Skill","tool_input":{"skill":"superpowers:executing-plans"}}') -printf '%s' "$out" | grep -q 'floor met' && pass "3/3 exec -> Gate A satisfied" || fail "3/3 exec -> Gate A satisfied" -# FINDING 12: the Gate-A satisfied wording must NOT overstate — it counts calls only. -printf '%s' "$out" | grep -qE 'count only|COUNT ONLY' && pass "Gate A satisfied says 'count only' (Finding 12)" || fail "Gate A satisfied says 'count only' (Finding 12)" +printf '%s' "$out" | grep -q 'floor met' && pass "3/3 exec -> Gate A floor met" || fail "3/3 exec -> Gate A floor met" +# FINDING 12: the Gate-A floor-met wording must NOT overstate — it counts calls only. +printf '%s' "$out" | grep -qE 'count only|COUNT ONLY' && pass "Gate A floor-met message says 'count only' (Finding 12)" || fail "Gate A floor-met message says 'count only' (Finding 12)" run '{"hook_event_name":"PostToolUse","tool_name":"Skill","tool_input":{"skill":"superpowers:executing-plans"}}' >/dev/null [ ! -f "$countA" ] && pass "plan execution resets Gate A count" || fail "plan execution resets Gate A count" @@ -520,8 +520,8 @@ reset_all printf '1' > "$floorf" rev out=$(commitpre) -printf '%s' "$out" | grep -q '1/1' && pass "floor override 1 -> satisfied at 1 pass" || fail "floor override 1 -> satisfied at 1 pass" -printf '%s' "$out" | grep -q 'Gate B satisfied' && pass "floor override 1 -> reports satisfied" || fail "floor override 1 -> reports satisfied" +printf '%s' "$out" | grep -q '1/1' && pass "floor override 1 -> hook checks pass at 1 pass" || fail "floor override 1 -> hook checks pass at 1 pass" +printf '%s' "$out" | grep -q 'Gate B hook checks passed' && pass "floor override 1 -> reports hook checks passed" || fail "floor override 1 -> reports hook checks passed" reset_all printf '5' > "$floorf" rev; rev; rev @@ -537,7 +537,7 @@ for bad in 0 -2 three ""; do done rm -f "$floorf" -# 18. FINDING 9 — satisfied message distinguishes fresh passes from cycle passes +# 18. FINDING 9 — hook-checks-passed message distinguishes fresh passes from cycle passes reset_all rev; rev; rev # 3 passes on the current tree printf 'post-review rewrite\n' >> app.ts # big change AFTER the passes @@ -553,12 +553,12 @@ printf '%s' "$out" | grep -qF 'of which 1 cover the CURRENT content fingerprint' git checkout -- app.ts >/dev/null 2>&1 reset_all -# 19. FINDING 11 — WIP commit is cycle-internal: gentle note, no STOP, no reset +# 19. FINDING 11 — WIP commit is cycle-internal: gentle note, no gate-state reminder, no reset reset_all rev; rev # 2 passes accumulated wip() { run "{\"hook_event_name\":\"PreToolUse\",\"tool_name\":\"Bash\",\"tool_input\":{\"command\":\"$1\"}}"; } out=$(wip "git commit -m 'wip: pre-review snapshot'") -printf '%s' "$out" | grep -q 'STOP' && fail "WIP commit must not STOP" || pass "WIP commit does not STOP" +printf '%s' "$out" | grep -q 'Codex gate state:' && fail "WIP commit must not emit the gate-state reminder" || pass "WIP commit does not emit the gate-state reminder" printf '%s' "$out" | grep -q 'WIP commit' && pass "WIP commit -> gentle note" || fail "WIP commit -> gentle note" out=$(wip "git commit -m 'WIP: caps variant'") printf '%s' "$out" | grep -q 'WIP commit' && pass "WIP matcher is case-insensitive" || fail "WIP matcher is case-insensitive" @@ -630,7 +630,7 @@ rm -f "$on" rev; rev; rev [ -z "$(commitpre)" ] && pass "non-adopted repo: commit -> silent" || fail "non-adopted repo: commit -> silent" reset_all -[ -z "$(commitpre)" ] && pass "non-adopted repo: unreviewed commit -> no STOP" || fail "non-adopted repo: unreviewed commit -> no STOP" +[ -z "$(commitpre)" ] && pass "non-adopted repo: unreviewed commit -> no reminder" || fail "non-adopted repo: unreviewed commit -> no reminder" out=$(run '{"hook_event_name":"PreToolUse","tool_name":"Skill","tool_input":{"skill":"superpowers:executing-plans"}}') [ -z "$out" ] && pass "non-adopted repo: Gate A -> silent" || fail "non-adopted repo: Gate A -> silent" [ -z "$(codextool mcp__codex__codex)" ] && pass "non-adopted repo: unknown-tool note -> silent" || fail "non-adopted repo: unknown-tool note -> silent" @@ -649,7 +649,7 @@ rm -rf "$sub" # 22b. Either adoption marker is enough, and it takes effect without a restart. # (Each CLAUDE.md write is itself a tree change, so the passes are re-run after -# one — otherwise a stale-tree STOP would masquerade as non-adoption.) +# one — otherwise a stale-fingerprint reminder would masquerade as non-adoption.) reset_all printf '# p\n\n## 5. Something else entirely\n' > CLAUDE.md # a CLAUDE.md without the gates rm -f "$on" @@ -657,12 +657,12 @@ rev; rev; rev # inert: these [ -z "$(commitpre)" ] && pass "unrelated CLAUDE.md -> not adopted" || fail "unrelated CLAUDE.md -> not adopted" : > "$on" # the explicit marker rev; rev; rev # passes only count once adopted -printf '%s' "$(commitpre)" | grep -q 'Gate B satisfied' && pass ".on marker alone -> adopted" || fail ".on marker alone -> adopted" +printf '%s' "$(commitpre)" | grep -q 'Gate B hook checks passed' && pass ".on marker alone -> adopted" || fail ".on marker alone -> adopted" rm -f "$on" [ -z "$(commitpre)" ] && pass "removing the marker -> silent again" || fail "removing the marker -> silent again" printf '# p\n\n## 5. Cross-Model Review (Codex)\n' > CLAUDE.md # the committed, team-wide signal rev; rev; rev -printf '%s' "$(commitpre)" | grep -q 'Gate B satisfied' && pass "CLAUDE.md gate heading -> adopted" || fail "CLAUDE.md gate heading -> adopted" +printf '%s' "$(commitpre)" | grep -q 'Gate B hook checks passed' && pass "CLAUDE.md gate heading -> adopted" || fail "CLAUDE.md gate heading -> adopted" # 22bis. Adoption needs the gate SECTION, not the words. A substring grep adopts a # project on a passing mention — including one that says the opposite. @@ -710,7 +710,7 @@ rm -f CLAUDE.md; reset_all reset_all rm -f CLAUDE.md "$on"; : > "$on" # marker-only adoption, no CLAUDE.md at all out=$(commitpre) -printf '%s' "$out" | grep -q 'STOP' && pass "marker-only: still STOPs" || fail "marker-only: still STOPs" +printf '%s' "$out" | grep -q 'Codex gate state:' && pass "marker-only: still emits the gate-state reminder" || fail "marker-only: still emits the gate-state reminder" printf '%s' "$out" | grep -q 'CLAUDE.md' && fail "marker-only must not cite CLAUDE.md" || pass "marker-only: cites no CLAUDE.md" printf '%s' "$out" | grep -q "this project's review policy" && pass "marker-only: cites the project's policy generically" || fail "marker-only: cites the project's policy generically" out=$(run '{"hook_event_name":"PreToolUse","tool_name":"Skill","tool_input":{"skill":"superpowers:executing-plans"}}') @@ -729,7 +729,7 @@ rm -f CLAUDE.md git checkout -- . >/dev/null 2>&1 reset_all -# 24. Failure contract: an uncomputable hash must never satisfy, and repeated failures +# 24. Failure contract: an uncomputable hash must never pass the hook checks, and repeated failures # must never match each other. Spec §3 "The nonce goes away". # Each fault spans BOTH the stored and the recomputed fingerprint — with the fault # applied only at commit time, a mismatch proves nothing about the handling. @@ -750,7 +750,7 @@ for t in shasum sha1sum cksum; do done ( PATH="$stub_dir:$PATH" rev ) out=$(PATH="$stub_dir:$PATH" commitpre) -printf '%s' "$out" | grep -q 'not satisfied' && pass "silent checksum -> not satisfied" || fail "silent checksum -> not satisfied" +printf '%s' "$out" | grep -q 'Codex gate state:' && pass "silent checksum -> gate-state reminder" || fail "silent checksum -> gate-state reminder" # Absence is asserted by inverting the RESULT, not with `grep -v` — `grep -qv` means # "some line lacks the pattern", which is a different question and was observed to # return 1 regardless on the dev machine. @@ -765,7 +765,7 @@ done reset_all; printf '1' > "$floorf" ( PATH="$stub_dir:$PATH" rev ) out=$(PATH="$stub_dir:$PATH" commitpre) -printf '%s' "$out" | grep -q 'not satisfied' && pass "checksum prints then fails -> not satisfied" || fail "checksum prints then fails -> not satisfied" +printf '%s' "$out" | grep -q 'Codex gate state:' && pass "checksum prints then fails -> gate-state reminder" || fail "checksum prints then fails -> gate-state reminder" [ "$(cat "$fresh" 2>/dev/null || echo 0)" = 0 ] && pass "unhashable pass leaves freshCount 0" || fail "unhashable pass leaves freshCount 0" # 24b-bis. GAP 2 — a SECOND consecutive unhashable pass must not be treated as a @@ -783,8 +783,8 @@ printf '#!/bin/sh\nexit 1\n' > "$stub_dir/cp"; chmod +x "$stub_dir/cp" rm -f "$stub_dir/shasum" "$stub_dir/sha1sum" "$stub_dir/cksum" reset_all; printf '1' > "$floorf" ( PATH="$stub_dir:$PATH" rev ) -printf '%s' "$(PATH="$stub_dir:$PATH" commitpre)" | grep -q 'not satisfied' \ - && pass "seed-copy failure -> not satisfied" || fail "seed-copy failure -> not satisfied" +printf '%s' "$(PATH="$stub_dir:$PATH" commitpre)" | grep -q 'Codex gate state:' \ + && pass "seed-copy failure -> gate-state reminder" || fail "seed-copy failure -> gate-state reminder" rm -f "$stub_dir/cp" # 24d. a selective git wrapper that fails ONLY `diff` @@ -797,8 +797,8 @@ chmod +x "$stub_dir/git" REAL_GIT=$(command -v git); export REAL_GIT reset_all; printf '1' > "$floorf" ( PATH="$stub_dir:$PATH" rev ) -printf '%s' "$(PATH="$stub_dir:$PATH" commitpre)" | grep -q 'not satisfied' \ - && pass "git diff failure -> not satisfied" || fail "git diff failure -> not satisfied" +printf '%s' "$(PATH="$stub_dir:$PATH" commitpre)" | grep -q 'Codex gate state:' \ + && pass "git diff failure -> gate-state reminder" || fail "git diff failure -> gate-state reminder" # 24e. a selective git wrapper that fails ONLY `rev-parse --absolute-git-dir` cat > "$stub_dir/git" <<'STUB' @@ -809,8 +809,8 @@ STUB chmod +x "$stub_dir/git" reset_all; printf '1' > "$floorf" ( PATH="$stub_dir:$PATH" rev ) -printf '%s' "$(PATH="$stub_dir:$PATH" commitpre)" | grep -q 'not satisfied' \ - && pass "unresolvable git-dir -> not satisfied" || fail "unresolvable git-dir -> not satisfied" +printf '%s' "$(PATH="$stub_dir:$PATH" commitpre)" | grep -q 'Codex gate state:' \ + && pass "unresolvable git-dir -> gate-state reminder" || fail "unresolvable git-dir -> gate-state reminder" rm -f "$stub_dir/git" rm -rf "$stub_dir" # GAP 3 — stub_dir (mktemp -d) outlives its own guard otherwise unset REAL_GIT # GAP 3 — exported at 24d, never unset otherwise @@ -833,7 +833,7 @@ reset_all; rev; rev; rev [ ! -d "$stub_dir" ] && pass "stub_dir removed after section 24" || fail "stub_dir removed after section 24" # 25. Unborn repo: no commits and no .git/index must still hash and self-match, or the -# first commit in a fresh repo STOPs forever. Spec §3 "But an absent index is not a +# first commit in a fresh repo gets the gate-state reminder forever. Spec §3 "But an absent index is not a # failed copy". unborn=$(mktemp -d) ( @@ -848,13 +848,13 @@ unborn=$(mktemp -d) printf '%s' "$R" | "$HOOK_SH_BIN" "$HOOK" >/dev/null h2=$(cat .context/codex-gate.gateB 2>/dev/null) [ -n "$h1" ] && [ "$h1" != unavailable ] && [ "$h1" = "$h2" ] || exit 1 - # ...and the FIRST commit must actually be able to reach satisfied. Hashing and - # self-matching is not enough: a consumer-side regression could still STOP every + # ...and the FIRST commit must actually be able to reach hook checks passed. Hashing and + # self-matching is not enough: a consumer-side regression could still remind on every # first commit forever, which is the failure this fixture exists to catch. out=$(printf '%s' '{"hook_event_name":"PreToolUse","tool_name":"Bash","tool_input":{"command":"git commit --allow-empty -m init"}}' | "$HOOK_SH_BIN" "$HOOK") - printf '%s' "$out" | grep -q 'Gate B satisfied' || exit 1 -) && pass "unborn repo hashes, self-matches, and can reach satisfied" \ - || fail "unborn repo hashes, self-matches, and can reach satisfied" + printf '%s' "$out" | grep -q 'Gate B hook checks passed' || exit 1 +) && pass "unborn repo hashes, self-matches, and can reach hook checks passed" \ + || fail "unborn repo hashes, self-matches, and can reach hook checks passed" rm -rf "$unborn" # 26. THE DEFECT (spec §1): staged content diverging from the worktree. @@ -863,16 +863,16 @@ rm -rf "$unborn" # the divergence, making this test pass against the unfixed hook. reset_all rev; rev; rev -printf '%s' "$(commitpre)" | grep -q 'Gate B satisfied' && pass "setup: satisfied on clean tree" || fail "setup: satisfied on clean tree" +printf '%s' "$(commitpre)" | grep -q 'Gate B hook checks passed' && pass "setup: hook checks passed on clean tree" || fail "setup: hook checks passed on clean tree" printf 'v2\n' > app.ts; git add app.ts >/dev/null 2>&1 # index: v2 printf 'v1\n' > app.ts # worktree: back to HEAD bytes -printf '%s' "$(commitpre)" | grep -q 'not satisfied' \ - && pass "staged-vs-worktree divergence -> NOT satisfied" \ - || fail "staged-vs-worktree divergence -> NOT satisfied" +printf '%s' "$(commitpre)" | grep -q 'Codex gate state:' \ + && pass "staged-vs-worktree divergence -> gate-state reminder" \ + || fail "staged-vs-worktree divergence -> gate-state reminder" git reset -q >/dev/null 2>&1; git checkout -- app.ts >/dev/null 2>&1 -# 27. Ambient alternate index. Three shapes: a negative-only test would be satisfied by -# an implementation that fires whenever GIT_INDEX_FILE is set — a permanent STOP. +# 27. Ambient alternate index. Three shapes: a negative-only test would be passed by +# an implementation that fires whenever GIT_INDEX_FILE is set — a permanent gate-state reminder. alt_dir=$(mktemp -d) # 27a. divergent: fingerprint recorded WITHOUT the alternate index, alternate enabled # only for the commit check. @@ -881,10 +881,10 @@ rev cp .git/index "$alt_dir/alt" printf 'SNEAKY\n' > app.ts; GIT_INDEX_FILE="$alt_dir/alt" git add app.ts >/dev/null 2>&1 printf 'v1\n' > app.ts -printf '%s' "$(GIT_INDEX_FILE="$alt_dir/alt" commitpre)" | grep -q 'not satisfied' \ - && pass "ambient divergent alternate index -> NOT satisfied" \ - || fail "ambient divergent alternate index -> NOT satisfied" -# 27b. stable: same unchanged alternate index across review AND commit -> satisfied, +printf '%s' "$(GIT_INDEX_FILE="$alt_dir/alt" commitpre)" | grep -q 'Codex gate state:' \ + && pass "ambient divergent alternate index -> gate-state reminder" \ + || fail "ambient divergent alternate index -> gate-state reminder" +# 27b. stable: same unchanged alternate index across review AND commit -> hook checks passed, # with each index file byte-identical to its OWN pre-hook snapshot. cp .git/index "$alt_dir/default.before"; cp "$alt_dir/alt" "$alt_dir/alt.before" reset_all; printf '1' > "$floorf" @@ -893,9 +893,9 @@ reset_all; printf '1' > "$floorf" # prefix on an external command. Without containment, GIT_INDEX_FILE leaks out of # section 27 and corrupts every later section's fixtures (notably section 28). ( GIT_INDEX_FILE="$alt_dir/alt" rev ) -printf '%s' "$(GIT_INDEX_FILE="$alt_dir/alt" commitpre)" | grep -q 'Gate B satisfied' \ - && pass "ambient stable alternate index -> satisfied" \ - || fail "ambient stable alternate index -> satisfied" +printf '%s' "$(GIT_INDEX_FILE="$alt_dir/alt" commitpre)" | grep -q 'Gate B hook checks passed' \ + && pass "ambient stable alternate index -> hook checks passed" \ + || fail "ambient stable alternate index -> hook checks passed" cmp -s .git/index "$alt_dir/default.before" && pass "default index untouched" || fail "default index untouched" cmp -s "$alt_dir/alt" "$alt_dir/alt.before" && pass "alternate index untouched" || fail "alternate index untouched" # 27c. missing path: git treats a nonexistent GIT_INDEX_FILE as an EMPTY index, so the @@ -925,7 +925,7 @@ reset_all; rev; rev; rev # from. Run BOTH the review pass and the commit check from a SUBDIRECTORY with a # relative alt index that actually lives at the repo root: an unnormalized hook # can't find it either time, takes the same absent-index carve-out both times, and -# the two constant empty-tree hashes MATCH — a false "satisfied" even though the +# the two constant empty-tree hashes MATCH — a false "hook checks passed" even though the # alt index stages content the worktree does not have. # The alt index file lives under `.context/` — excluded from the diff-HEAD and # worktree-tree components by their own `:(exclude).context` pathspec — so it is @@ -940,9 +940,9 @@ printf 'SNEAKY\n' > app.ts GIT_INDEX_FILE="$work/.context/rel-idx" git add app.ts >/dev/null 2>&1 # stage into the ALT index only printf 'v1\n' > app.ts # worktree stays at the reviewed bytes out=$(cd sub && GIT_INDEX_FILE=.context/rel-idx commitpre) -printf '%s' "$out" | grep -q 'not satisfied' \ - && pass "relative ambient GIT_INDEX_FILE from a subdirectory -> NOT satisfied (Finding 1)" \ - || fail "relative ambient GIT_INDEX_FILE from a subdirectory -> NOT satisfied (Finding 1)" +printf '%s' "$out" | grep -q 'Codex gate state:' \ + && pass "relative ambient GIT_INDEX_FILE from a subdirectory -> gate-state reminder (Finding 1)" \ + || fail "relative ambient GIT_INDEX_FILE from a subdirectory -> gate-state reminder (Finding 1)" rm -f "$work/.context/rel-idx"; rm -rf sub git checkout -- app.ts >/dev/null 2>&1; git reset -q >/dev/null 2>&1 reset_all; rev; rev; rev @@ -1003,31 +1003,31 @@ run '{"hook_event_name":"PostToolUse","tool_name":"Edit","tool_input":{"file_pat out=$(commitpre) ctx=$(json_field "$out" additionalContext) msg=$(json_field "$out" systemMessage) -expected_ctx="STOP — Codex Gate B not satisfied: the hook cannot confirm that the content you are about to commit is the content mcp__codex__review last saw (1 recorded pass(es) this cycle). Usually that means the working tree or the index changed since the review. It can also mean you only staged already-reviewed content — the bytes are fine, but the hook cannot tell staging from editing; that this hook was upgraded and the recorded fingerprint uses the older format (see CHANGELOG); or that the fresh fingerprint could not be computed or could not be stored. Run Gate B (mcp__codex__review) now — one clean pass is the complete remedy for the staging and post-upgrade cases too. If a fresh pass leaves this unchanged with nothing edited in between, the fault is in the machinery rather than the code: check that .context/ is writable, that TMPDIR is writable, that a checksum tool (shasum, sha1sum or cksum) runs, that git status works, and that the disk is not full — then run one more pass to record a usable fingerprint. Per this project's review policy you MUST re-review after every fix." -expected_msg="⚠ Codex Gate B not satisfied (cannot confirm review)" +expected_ctx="Codex gate state: the hook cannot confirm that the content you are about to commit is the content mcp__codex__review last saw (1 recorded pass(es) this cycle). Usually that means the working tree or the index changed since the review. It can also mean you only staged already-reviewed content — the bytes are fine, but the hook cannot tell staging from editing; that this hook was upgraded and the recorded fingerprint uses the older format (see CHANGELOG); or that the fresh fingerprint could not be computed or could not be stored. A fresh Gate-B pass is the complete remedy for the staging and post-upgrade cases too, and when this cycle may run one is this project's review policy's closure ordering's, read there entire. If a fresh pass leaves this unchanged, check the worktree and the index first — the fingerprint moves when either does, staging included, and another hook can stage during the commit attempt. Where neither changed, the fault is in the machinery rather than the code: check that .context/ is writable, that TMPDIR is writable, that a checksum tool (shasum, sha1sum or cksum) runs, that git status works, and that the disk is not full — a store that fails again leaves the hook unable to confirm a fingerprint and may return either fingerprint-state diagnosis. Per this project's review policy you MUST re-review after every fix." +expected_msg="⚠ Codex Gate B: cannot confirm reviewed content" [ "$ctx" = "$expected_ctx" ] && pass "stale additionalContext matches exactly" || fail "stale additionalContext matches exactly" [ "$msg" = "$expected_msg" ] && pass "stale systemMessage matches exactly" || fail "stale systemMessage matches exactly" git checkout -- app.ts >/dev/null 2>&1 -# 29b. SATISFIED branch: 3/3 passes this cycle, all 3 fresh (unchanged tree). The hook +# 29b. HOOK-CHECKS-PASSED branch: 3/3 passes this cycle, all 3 fresh (unchanged tree). The hook # fingerprints disk; mcp__codex__review reads a git range (spec §7) — the exact # fixture below is what pins that the message never claims Codex read the bytes. reset_all; rev; rev; rev out=$(commitpre) ctx=$(json_field "$out" additionalContext) msg=$(json_field "$out" systemMessage) -expected_ctx="Codex Gate B: 3/3 pass(es) this cycle, of which 3 cover the CURRENT content fingerprint (unchanged since that review). The floor counts the cycle; only the fresh pass(es) carry the same fingerprint as what you are committing. Per this project's review policy, commit only if your final pass was clean — no new Blocker/Major." -expected_msg="✓ Codex Gate B satisfied (3/3 cycle, 3 on current fingerprint)" -[ "$ctx" = "$expected_ctx" ] && pass "satisfied additionalContext matches exactly" || fail "satisfied additionalContext matches exactly" -[ "$msg" = "$expected_msg" ] && pass "satisfied systemMessage matches exactly" || fail "satisfied systemMessage matches exactly" +expected_ctx="Codex Gate B: 3/3 pass(es) this cycle, of which 3 cover the CURRENT content fingerprint (unchanged since that review). The floor counts the cycle; only the fresh pass(es) carry the same fingerprint as what you are committing. Per this project's review policy, commit only if your final pass was clean and every other closure condition holds, both as it defines them." +expected_msg="✓ Codex Gate B hook checks passed (3/3 cycle, 3 on current fingerprint)" +[ "$ctx" = "$expected_ctx" ] && pass "hook-checks-passed additionalContext matches exactly" || fail "hook-checks-passed additionalContext matches exactly" +[ "$msg" = "$expected_msg" ] && pass "hook-checks-passed systemMessage matches exactly" || fail "hook-checks-passed systemMessage matches exactly" # 29c. EMPTY-STATE branch: no fingerprint recorded this cycle, default floor 3. reset_all out=$(commitpre) ctx=$(json_field "$out" additionalContext) msg=$(json_field "$out" systemMessage) -expected_ctx="STOP — Codex Gate B not satisfied: no fingerprint is recorded for this cycle — either no mcp__codex__review has run, or the last one's fingerprint could not be written or read back. Per this project's review policy you MUST reach a minimum of 3 passes per cycle. Run Gate B (mcp__codex__review) now; if this repeats, check that .context/ and the state file inside it are readable and writable, and if the file exists but is unreadable or empty, delete it and run a fresh pass." -expected_msg="⚠ Codex Gate B: no recorded review" +expected_ctx="Codex gate state: no fingerprint is recorded for this cycle. The hook cannot tell why — no mcp__codex__review has run, the last one's fingerprint could not be written or read back, or a non-WIP commit attempt cleared it while the cycle itself stayed open. What this cycle does next, the floor it owes included, is this project's review policy's closure ordering's, read there entire, and this reminder decides none of it. If this repeats, check that .context/ and the state file inside it are readable and writable; if the file exists but is unreadable or empty, delete it — which restores no passes, and lets the next pass the ordering permits record a fingerprint." +expected_msg="⚠ Codex Gate B: no recorded fingerprint" [ "$ctx" = "$expected_ctx" ] && pass "empty-state additionalContext matches exactly" || fail "empty-state additionalContext matches exactly" [ "$msg" = "$expected_msg" ] && pass "empty-state systemMessage matches exactly" || fail "empty-state systemMessage matches exactly" reset_all; rev; rev; rev @@ -1173,9 +1173,9 @@ git checkout -- app.ts >/dev/null 2>&1 # 6/9 below the Gate-B floor reset_all; rev one_doc "Gate B below floor" "$(commitpre)" -# 7/9 Gate B satisfied +# 7/9 Gate B hook checks passed reset_all; rev; rev; rev -one_doc "Gate B satisfied" "$(commitpre)" +one_doc "Gate B hook checks passed" "$(commitpre)" # 8/9 Gate A below floor reset_all one_doc "Gate A below floor" "$(run '{"hook_event_name":"PreToolUse","tool_name":"Skill","tool_input":{"skill":"superpowers:executing-plans"}}')" From bf4d2cf6b6e111f050d498e038fcb707b9b28253 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Sat, 26 Sep 2026 15:10:51 +0200 Subject: [PATCH 181/181] docs: describe the hook's new output and the closure ordering in the walkthrough and methodology MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit getting-started.md quoted the removed "Gate B satisfied" message and closed every cycle with git commit --amend; coding-workflow.md required a revised artifact on every pass and said only serious findings cost another pass. Both now point at CLAUDE.md §5's closure ordering, the Gate-A closing acts and Finishing the cycle. Explanatory docs only (docs/**.md): no Gate B. Answers Greptile on PR #28 (discussion_r4111298126). --- docs/coding-workflow.md | 18 ++++++++++++------ docs/getting-started.md | 22 ++++++++++++++-------- 2 files changed, 26 insertions(+), 14 deletions(-) diff --git a/docs/coding-workflow.md b/docs/coding-workflow.md index ee3d018..189b7f9 100644 --- a/docs/coding-workflow.md +++ b/docs/coding-workflow.md @@ -93,10 +93,13 @@ a wish list. author — reads the spec text and checks it for contradictions and internal inconsistencies, missing requirements, unhandled state/edge/error/empty/concurrent paths, and risks to the project's core invariants. It is run as a **loop with a -hard floor**: a minimum number of passes, re-running one broad review prompt over -the *revised* artifact each time (new findings surface precisely because the -artifact changed between passes). Only design-breaking findings — the serious -tiers — force another iteration; the final pass must come back clean. Catching a +hard floor**: a minimum number of passes, re-running one broad review prompt each +time — over the revised artifact where a finding required a repair, over the unchanged +one where none did. A finding's severity decides what must be repaired, not on its own +whether another pass is owed: the review policy's closure ordering decides that. A cycle +closes on an eligible pass — clean at or above the floor, or one with zero findings — only +when every other closure condition holds, and accepting even a minor finding into the +assigned work costs a further pass. Catching a flaw in the spec is far cheaper than catching it after it has been baked into the plan and the code. @@ -107,7 +110,9 @@ are restated at the top so they aren't lost mid-build. A plan at this resolution makes execution mechanical and the resulting diff traceable back to a requirement. **5. Gate A on the plan.** The same independent review, now applied to the plan — -so a plan-level flaw is caught before implementation, not during it. +so a plan-level flaw is caught before implementation, not during it. Each Gate-A cycle +ends with its own closing act — the reviewed artifact committed with the cycle's +records — before the next stage starts. **6. Execution.** The plan is implemented task by task, followed literally. Bounded subtasks can be delegated to cheaper models or subagents. Discipline @@ -124,7 +129,8 @@ checks is not enforcement. **8. Gate B on the code.** Before the change is committed, the independent reviewer reads the actual *diff* and checks it against the invariants file. It is re-run after every fix, because each fix changes the diff and invalidates the -prior review. Trivial changes may skip it, on terms that depend on the story: an +prior review, and it closes when the review policy's closure ordering says so — never +on a clean pass alone. Trivial changes may skip it, on terms that depend on the story: an unprofiled one keeps the judgement call, while a profiled one qualifies only at effective level 0 — trivial risk *and* no security relevance — so a trivial-looking change on security-relevant surface is not eligible. A skip removes the review, never diff --git a/docs/getting-started.md b/docs/getting-started.md index e5456e1..47d557a 100644 --- a/docs/getting-started.md +++ b/docs/getting-started.md @@ -5,7 +5,7 @@ assumed — see the README if not. Up front: you don't operate the workflow like a machine. You talk to Claude normally; the skills and gates structure *how Claude works*, and the hook reminds -both of you when a gate isn't satisfied. Your job is the decision points — +both of you what it has and has not seen of the gates. Your job is the decision points — answering questions, approving drafts, judging findings. **1. Capture the idea.** Say "users want to export their invoices as CSV" (or paste @@ -30,15 +30,19 @@ settled decisions with rationale, not a wish list. **3. Gate A on the spec.** Claude sends the spec text to Codex (`mcp__codex__exec`) — a different model family, so it doesn't share Claude's blind -spots. Blocker/Major findings get fixed, the review reruns on the revised spec: -the floor its profile derives, final pass clean — the one early exit is a pass that -comes back with zero findings. Hook messages like `⚠ Codex Gate A below floor (1/3)` are +spots. Blocker/Major findings get fixed and the review reruns — on the revised spec where a +repair was owed, on the same text where none was. When the cycle may close is `CLAUDE.md` +§5's closure ordering, not a rule of thumb: an eligible pass (clean at or above the floor +its profile derives, or a pass with zero findings) with every other closure condition +holding, then the **Gate-A closing act** — the reviewed spec committed with the cycle's +records. Planning starts after that act, not after the last clean pass. Hook messages like `⚠ Codex Gate A below floor (1/3)` are the counter, not an error. Your job: arbitrate disputed findings — Codex is advisory, and a dismissed finding needs a one-line reason. **4. Plan, and Gate A again.** `superpowers:writing-plans` turns the spec into a task-by-task plan (each task starts with a failing test); the same loop runs at the derived floor -on the plan. A flaw caught here never reaches code. +on the plan and ends with its own Gate-A closing act before execution starts. A flaw caught +here never reaches code. **5. Implement.** `superpowers:executing-plans` works through the plan, test-first, progress claims backed by test runs. If the hook's own threshold wasn't met, it says @@ -52,13 +56,15 @@ skipping locally only postpones the red. range to read; the hook knows WIP doesn't end the cycle), then loops `mcp__codex__review` the same way: the derived floor, final clean. Invalidation is by **content** — any change to included content present when the hook runs, even from a -formatter, flips it back to unsatisfied. What that proves is bounded, and the hook's own +formatter, makes the hook report that it cannot confirm the reviewed content. What that proves is bounded, and the hook's own source says so: the current fingerprint matches the one recorded on a counted call, which is not evidence that Codex read those bytes; `.context/` and untracked ignored paths are excluded, and staging counts, because the fingerprint covers the index and that is what a commit carries. On -`✓ Codex Gate B satisfied (/ cycle, on current fingerprint)` — three different numbers: the calls the hook counted this cycle, the hook's own reminder threshold, and the **consecutive** counted calls on the current fingerprint since it last changed. The first is not the calls you made: the hook withholds the count for a recognized failure envelope, the backgrounding notice, and a result it can get no text from. The third is a streak, not a tally — the hook keeps the last fingerprint and that streak, so a pass on a changed fingerprint restarts it and an earlier matching pass separated by a different fingerprint is not counted. None of the three is the floor §5 obliges — the real commit replaces -the WIP via `git commit --amend`. +`✓ Codex Gate B hook checks passed (/ cycle, on current fingerprint)` — three different numbers: the calls the hook counted this cycle, the hook's own reminder threshold, and the **consecutive** counted calls on the current fingerprint since it last changed. The first is not the calls you made: the hook withholds the count for a recognized failure envelope, the backgrounding notice, and a result it can get no text from. The third is a streak, not a tally — the hook keeps the last fingerprint and that streak, so a pass on a changed fingerprint restarts it and an earlier matching pass separated by a different fingerprint is not counted. None of the three is the floor §5 obliges, and the message is what the hook checked, not +permission to close: §5's closure ordering says when the Gate-B cycle may close, and its +*Finishing the cycle* operation says how — amend the WIP commit, or, where several `WIP:` +snapshots piled up, `git reset --soft` to the parent of the first and commit once. **8. PR and bots.** Open the PR as usual; once the bots have commented, run `/dev-workflow:process-pr-review`. Every comment is validated against code and