From c60b78bae9aff75e124d90cc9e299290f54a0b5d Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Thu, 1 Oct 2026 15:54:53 +0200 Subject: [PATCH 1/5] docs(specs): design passive metrics over the ledger and git; profile the story MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Dark-factory vision step 2b. A repo-local, read-only script reports fingerprint recurrence and rung holding from docs/hardening-log.md, and the review-cycle records (provenance line, curve, skip record) from commit bodies, side by side with the fic2 baseline. It computes no shares or verdicts. The story header now carries the profile Daniel confirmed on 2026-10-01 (standard / none / battery+check) and the placement decision (repo-local script, not shipped). Gate-A spec cycle closed. Pass 4 is clean at the derived floor: 0 Blockers and 0 Majors, and every earlier Major was resolved by a repair the next pass confirmed. Pass 4's six Minors and one Nit are collected, not iterated. The committed spec is byte-identical to the text pass 4 reviewed (sha256 60d108d30adcbe09c6dafcd8c0bb00389d94e98bd4d2d7f47aa0bd925d3eb039). Two earlier calls for pass 1 failed because the Codex app had set an unsupported model (gpt-6.1-sol). They were discarded and are not counted. Docs-only change (docs/**.md): Gate B is N/A per CLAUDE.md §5. cycle mbu2nfkahz; floor 3 per {docs/superpowers/stories/2026-08-04-passive-metrics-over-the-ledger-story.md (level 1)}; hook reminder threshold absent cycle mbu2nfkahz; Gate-A spec (passes 1-4, gpt-6-astra): Findings 24,19,13,7. Blockers 0,0,0,0. Majors 14,6,3,0. --- .../2026-10-01-passive-metrics-design.md | 265 ++++++++++++++++++ ...4-passive-metrics-over-the-ledger-story.md | 18 +- 2 files changed, 279 insertions(+), 4 deletions(-) create mode 100644 docs/superpowers/specs/2026-10-01-passive-metrics-design.md diff --git a/docs/superpowers/specs/2026-10-01-passive-metrics-design.md b/docs/superpowers/specs/2026-10-01-passive-metrics-design.md new file mode 100644 index 0000000..341ac64 --- /dev/null +++ b/docs/superpowers/specs/2026-10-01-passive-metrics-design.md @@ -0,0 +1,265 @@ +# Passive metrics over the ledger and git — design + +**Story:** `docs/superpowers/stories/2026-08-04-passive-metrics-over-the-ledger-story.md` — read the profile from its header at every gate call. + +Dark-factory vision step **2b** (`docs/superpowers/specs/2026-08-30-dark-factory-vision.md` §7). This spec +covers only what that story asks for. It stays read-only. It does not widen into step 2c's telemetry, +which the vision assigns to a separate story. + +**Design rule for this spec: report what the records say, compute as little as possible.** The script +counts and lists. It does not compute shares, ratios or verdicts, because every derived figure brings +its own edge cases (missing values, zero denominators), and the story only asks that the data become +readable and checkable. + +## §1 What is added or edited + +| Path | Change | +|---|---| +| `scripts/ledger-metrics.sh` | **New.** POSIX `sh` + `awk` + `git`. Reads, prints a report on stdout, writes nothing. | +| `scripts/ledger-metrics.test.sh` | **New.** Its regression suite, built on throwaway fixture repositories. | +| `AGENTS.md` | § Commands: the quality and lint rows gain shellcheck runs for **both** new files; the quality row (not the lint row) also gains a run of the suite. Layout tree: both files. **Boundaries** paragraph: the list of executable artifacts gains this script and its test. | +| `.github/workflows/ci.yml` | The shellcheck step gains both new files, and the checker-suite step gains the suite. Any step name or comment that enumerates the checkers is updated too. | +| `README.md` | The Contributing paragraph that counts the executables and checker suites is updated, and one line says how to run the report. | + +**Before editing, grep for every other place that enumerates the executables or checker suites** +(AGENTS.md Don't: "the layout tree above is part of the surface that drifts"). The plan carries the +exact grep. + +**Not shipped.** The script lives under `scripts/`, like `check-invariants.sh`. The plugin does not +contain it, so invariants 11 and 12 do not bind, and the plugin version does not change (Daniel, +2026-10-01). Shipping it to consumer projects is a later, separate decision. + +## §2 Interface + +``` +sh scripts/ledger-metrics.sh [] +``` + +- **The ref is resolved once.** The script runs `git rev-parse --verify ^{commit}` (default + `HEAD`) once at start and uses only the resulting 40-character SHA afterwards. The ledger is read as + `git show :docs/hardening-log.md`, and the commit bodies as `git log `. Every commit + reachable from that SHA is read, not just the first-parent chain, because an ordinary merge can carry + cycle records on its merged side. The working tree is never read. So uncommitted ledger edits are not + counted, and the report header says so. +- **Shallow clones.** If `git rev-parse --is-shallow-repository` prints `true`, the report header says + that history is truncated. Every "none found" statement in the cycle and checkpoint sections then + carries `(history truncated: absence not established)`. +- **Locale.** The script sets `LC_ALL=C` so sorting and output do not depend on the environment. +- **Output** goes to stdout only. Exit `0` when the report is complete. Exit `1` with a one-line + `ledger-metrics: ` on stderr when a source cannot be read. Each cause has its own message: not + a git repository, the ref does not resolve to a commit, the ledger is absent at that commit, the + ledger path is not a regular file there (checked with `git ls-tree`: mode `100644` or `100755`, type + `blob` — a directory or symlink is rejected), the ledger cannot be read, or `git log` fails. Malformed input lines are **not** an exit: they are counted and listed (§3, §4). +- **The script writes nothing** — no file, no git ref, no config. The suite checks the parts of this a + test can observe (§6). The rest is held by reading the script, and §6 says which part is which. +- **What the script does not control.** It runs ordinary read-only git commands. Git itself can still + write when the caller's environment or repository tells it to — for example trace variables such as + `GIT_TRACE` pointing at a file, or a partial clone fetching a missing object. The script does not + override the caller's git configuration, and the report header says so in one line. This is a stated + limit, not a guard. + +## §3 Ledger section — recurrence and rung holding + +**What a row is: exactly the `harden-finding` match.** A line is a row for fingerprint `F` if it matches +the skill's recurrence grep, `^\| *[0-9-]{10} *\| *F *\|`. The script uses that same pattern, so its +count for `F` equals the count the skill's grep gives. This answers the story's second open question: +there is one definition, not two. Cells are split on **unescaped** `|`. A matching row with a cell count +other than seven is **still counted**, as the grep counts it, and is also listed as `irregular` with +its line number. Irregular-width rows are **excluded from rung holding**, and their rung shows as `?` in +the recurrence line, because their rung cell cannot be located reliably. + +**Fingerprint spelling.** Taxonomy classes are kebab-case. A fingerprint cell that does not match +`^[a-z0-9]+(-[a-z0-9]+)*$` is listed as `irregular`. No verification command is printed for it, because +its text would be pasted into a shell command. + +`Superseded rows` entries are list lines, not rows, so they never match. The ledger header says this +itself: a superseded row "keeps matching the column-2 grep, and keeps counting". + +**Recurrence.** One line per fingerprint, sorted by count (descending), then by name: + +``` + lines ,,… rungs > > … +``` + +`lines` are ledger line numbers at the SHA, ascending. `rungs` are the rung cells in file order. After +the list comes the exact command a reader can run to check any one count: + +``` +git show :docs/hardening-log.md | grep -cE '^\| *[0-9-]{10} *\| * *\|' +``` + +**Rung holding.** It counts only rung cells that are one of the values the ledger uses: `1 prose`, +`2 lint`, `3 type`, `4 test`, `P std` and `pending` (the ledger on `main` uses five of these on +2026-10-01; `3 type` is the ladder's remaining rung). Any other rung value — empty, a typo, `0` — is +listed as `irregular rung` with its line number and is not counted. One line per counted rung value, +sorted by name. `pending` rows are listed separately as +`pending ` and are not counted as landed, because the `harden-finding` skill treats a `pending` +row as "no hardening landed yet". + +``` + landed followed-by-same-fingerprint last-of-fingerprint +``` + +**What this line means, printed beneath it.** "Followed" means a later row with the same fingerprint +exists in the file — nothing more. It does not mean this rung failed: a later row can record a different +sub-shape, an out-of-scope guard, or a prerequisite being resolved. And "last" does not mean it held: a +recurrence nobody hardened leaves no row. Judging whether a guard held means reading the rows. + +**Empty states.** If the ledger has no rows, the section says `no rows`. + +## §4 Git section — review cycles + +**Inputs are the pinned commit-body forms** from CLAUDE.md §5 Mechanics: the provenance line, the +per-pass curve, and the skip record (`; : skipped (see skip reason)`). + +**Candidate lines.** A body line is a candidate if it matches `^cycle [^ ;]*;` — the word `cycle`, one +token, then a semicolon — or starts with `cycle none (pre-rule);`. This catches malformed records whose +nonce or spacing after the semicolon is wrong. It does not catch prose that starts with the word +"cycle" (seen on `main`: "cycle closed on the zero-finding exit …"), because there the second word is +not followed by a semicolon. A candidate that fails the full grammar is listed as +**unparsed**, with its commit, and excluded. + +**Deduplication.** A record with a real nonce is keyed by its exact text. Text that appears in several +commits is counted once, and every commit carrying it is listed. `cycle none (pre-rule)` records are +**not** deduplicated, because identical text in two commits may be two different legacy cycles: each +occurrence is listed with its commit and flagged `may duplicate another pre-rule record`. + +**Ordering, everywhere in this section** — records, cycle lines, unparsed lines and conflict groups: +by the **earliest** committer date (`%ct`) among the commits carrying the record (for a group, among all +its records), then that commit's SHA, then the line text. + +**Grouping.** Records with the same real nonce are grouped. A group joins cleanly only if it has at most +one provenance line and at most one curve or skip record. A group with two different provenance lines, +two different curves, a curve and a skip record, or two different skip records is reported as a +**conflict** with all its records, and it is not listed as a cycle. Copying errors and nonce collisions both look like this. CLAUDE.md says the nonce is +collision-resistant, not collision-proof, so grouping by nonce is a strong default and not a guarantee, +and the report says so. `cycle none (pre-rule)` records are never grouped, because that field +identifies nothing. + +**Per cycle, one line,** sorted by first commit date, then nonce: + +``` + passes floor set findings blockers majors +``` + +- The `values` are **the recorded series, verbatim**, including `?`. No sums and no shares. +- After each series: `(? )`, the number of `?` values in it, so excluded values are counted per + series as story criterion 5 requires. +- `-` means the record has no provenance line. A provenance line with no curve and no skip record is + printed with `no curve`. +- A skip record is printed as `skipped`, followed by its reason: CLAUDE.md §5 says the reason is the text + immediately following the marker in the same commit body. The script takes the lines after the marker + up to the next blank line, joined with spaces. If the next non-empty line is itself a candidate, or + there is nothing after the marker, it prints `no reason found`. A skip record is keyed by marker + **and** reason, so two copies with different reasons are two records and form a conflict. + +**Empty states.** `no cycle records`, `no unparsed lines` and `no conflicts` are printed when they apply. + +## §5 The review-loop comparison (story criterion 5) + +The story asks whether profiled cycles under the new rules show a different severity mix than the +`fic2` baseline. **The script prints the two side by side and leaves the comparison to the reader.** +It computes no shares, so it cannot invent one from a missing or zero value. + +**Profiled cycles:** joined cycles with a curve whose provenance set contains at least one `(level N)` +entry. They are listed again here with their series and `(? )` counts, as in §4. If there are none, +the section says `no profiled cycles with a curve`. + +**The `fic2` baseline**, a constant in the script because its source commit `3cdd075` is not reachable +from `main` (checked 2026-10-01). Passes 1–5 come from the cycle-shape table in +`docs/field-reports/2026-08-26-fic2-cycle-evidence.md`; findings and blockers for passes 6–7 come from +the story's §1. The output names both sources: + +- findings `14,24,12,3,6,6,2` `(? 0)` +- blockers `3,4,0,0,0,0,0` `(? 0)` +- majors `5,13,6,2,5,?,?` `(? 2)` — passes 6 and 7 have no recorded Major count. + +**Printed with every report, always, not only when a comparison is possible:** + +1. The curves are author-written and unchecked. Nothing compares them against the validated pass + files, so they are self-reported and not measurement. +2. The cycles reviewed different artifacts, so a difference is evidence about the population as much + as about the rule. +3. No demotion figure is derivable. That would need one finding classified under both rules, and + nothing records that. + +**First checkpoint.** The story names "the first post-merge cycle whose cited set licenses floor 1". +CLAUDE.md §5 licenses floor 1 only when the set is non-empty and every member is profiled at level 0. +The script does not decide whether the checkpoint was reached. It lists the evidence, in the ordering +above: + +- every provenance line whose set licenses floor 1, with its recorded floor and commit +- every provenance line that records `floor 1`, with whether its set licenses it + +Either list may be empty, and then says `none`. On 2026-10-01 both are empty on `main`: all 19 +provenance lines say floor 3, and no set is all level 0. The reader decides which entry, if any, is the +story's "first post-merge" one. + +## §6 Testing — the `battery+check` evidence + +`scripts/ledger-metrics.test.sh` builds throwaway repositories under a `mktemp -d` directory. It commits +a fixture ledger and fixture commit bodies, then compares the script's output **line for line** against +expected text written in the test. Fixtures cover: + +- **Ledger:** two fingerprints with different counts; an escaped `\|` inside a finding; a supersession + list line (not counted); a row with too few cells (counted and listed as irregular); a fingerprint + with an uppercase letter (irregular, no command printed); a `pending` row; a row that is followed and + one that is last of its fingerprint; an empty ledger (`no rows`). +- **Cycles:** a provenance line and a curve with the same nonce (joined); a curve with no provenance + (`-`); a provenance line with no curve (`no curve`); a skip record; a `none (pre-rule)` pair (not + grouped); the same curve in two commits (once, both commits listed); two different curves under one + nonce, a curve and a skip record under one nonce, and two skip records with different reasons (each a + conflict, not listed as a cycle); a record only on the merged side of a merge commit (found); a prose line + starting with "cycle " (not a candidate); a candidate with a 7-character nonce (unparsed). +- **Grammar variants:** a quoted story path containing an escaped `\"`; a quoted model; discontiguous + pass ranges (`1-3,5`); per-pass models with `+`; a `?` in one series (printed verbatim, counted in + `(? n)`). Malformed variants listed as unparsed: a repeated story path; a count list longer than the + pass list; a per-pass model list missing a pass; a candidate with no space after the semicolon. +- **Skip records:** a skip record followed by a two-line reason (both lines printed); one with nothing + after it; one directly followed by another cycle record (both `no reason found`). +- **Rungs:** an empty rung cell and a rung `0` (each `irregular rung`, not counted). +- **Comparison and checkpoint:** the baseline lines exactly; the three caveats; a joined cycle with a + level-1 set and a curve (appears in the comparison); `no profiled cycles with a curve`; a level-0 set recorded with floor 3 (appears in the first list); a `floor 1` line with an + `(unprofiled)` entry (appears in the second list, not licensed); both lists empty. +- **Inputs:** a working-tree edit to the ledger that is not committed (no change in output); each + exit-1 cause, matched by its message, including the ledger path committed as a directory and as a + symlink; a shallow clone (header says history is truncated). + +**What the no-write check covers, and what it does not.** Before and after each run, the suite records +a listing of the whole fixture directory, `.git` included, with each file's mode and a checksum of its +contents. The two must be identical. That catches a write that leaves a file's content, mode or presence +different afterwards. It does not catch a rewrite with the same bytes, a change that is undone before the +run ends, a file created and deleted during the run, or a write elsewhere on the machine. For the +script's own commands those are held by reading it: it contains no file redirection other than to +`/dev/null`, and no git command that writes. That reading is a review check, not a test, and is labelled +as one. Writes git makes because of the caller's environment are outside both (§2). + +**The counterfactual — the observation against the prior state.** Before this change there is no +script, so nothing can produce the recurrence report. The suite's first case runs +`sh scripts/ledger-metrics.sh` in a fixture where the script path does not exist and checks that it +fails with "No such file". That is the prior state, observed. It is also why the check is weak on its +own: a missing-file error shows the tool is new, not that it counts right. + +**The negative control** shows the count assertion can fail. The suite runs its fingerprint-count case +against a temp copy of the script where `sed` changes the count by one. That run must **exit 0 from the +script** and fail **on the count assertion**, by its name. A broken copy that fails some other way, or +passes, makes the suite fail itself. Only then does it run the real script. + +## §7 What it cannot answer, printed in every report + +- **Unlogged recurrences.** The ledger only knows recurrences someone hardened. +- **Whether a guard held.** Row succession is not guard failure, and a missing later row is not + success (§3). +- **Cycles with no record at all.** Before 0.11.0 the curve was a habit, not a rule. A cycle whose + author wrote neither a provenance line nor a curve is invisible; a malformed record shows up as + unparsed. Provenance-only and skipped cycles are visible, as §4 describes. +- **Findings files.** The script does not read `.context/`. In this repository the findings files are + tracked under `.context/codex-reviews/`; in other projects they may be gitignored. Either way, the + report counts only what commit bodies say. +- **Whether a curve is true.** See §5, point 1. +- **Cost and duration.** That is step 2c's telemetry, out of scope here. + +## §8 Out of scope + +Writing anything back, any new state file, instrumentation, a plugin command, step 2c's telemetry, the +dashboard, computed shares or verdicts, and changing a ledger or commit-body format. diff --git a/docs/superpowers/stories/2026-08-04-passive-metrics-over-the-ledger-story.md b/docs/superpowers/stories/2026-08-04-passive-metrics-over-the-ledger-story.md index f3d6fe7..fc5ca0e 100644 --- a/docs/superpowers/stories/2026-08-04-passive-metrics-over-the-ledger-story.md +++ b/docs/superpowers/stories/2026-08-04-passive-metrics-over-the-ledger-story.md @@ -1,10 +1,20 @@ # Passive metrics, read-only over the ledger and git — Story **Date:** 2026-08-04 · **Size:** story +**Risk:** standard · **Security:** none · **Validation:** battery+check -**Unprofiled, deliberately** — a split from a designed round, so it bypassed -`dev-workflow:intake`, which excludes work already in solution design. A profile written now -would look confirmed without being confirmed; acceptance criterion 1 carries the debt instead. +**Profile log:** +- 2026-10-01 · adoption · proposed as `standard` / `none` / `battery+check` when design resumed, + per acceptance criterion 1; **confirmed by Daniel on 2026-10-01, exactly as proposed.** + Reason: the dark-factory vision (`docs/superpowers/specs/2026-08-30-dark-factory-vision.md` §8) + gates autonomy expansion on this story's evidence, so a wrong count would mislead a later + decision — bounded, not trivial. Security `none`: the analysis only reads. Derived floor 3. + +**Placement, decided by Daniel on 2026-10-01:** a repo-local script with its own test suite, like +the existing checkers — not shipped in the plugin. This answers §5's first open question. + +Originally unprofiled, deliberately — a split from a designed round, so it bypassed +`dev-workflow:intake`, which excludes work already in solution design. ## 1. Problem statement @@ -68,7 +78,7 @@ from the file rather than from recall — without the analysis writing anything ## 3. Acceptance criteria -- [ ] Before design resumes on this story, whoever picks it up proposes both axes and the mode +- [x] Before design resumes on this story, whoever picks it up proposes both axes and the mode derived from them, pauses for Daniel's confirmation, and writes the confirmed profile into this header. Design continues only after that. - [ ] The analysis reads the ledger and git and writes nothing — no new state file, no From 49b90f814c3f4c4268860913bc4ffd08cba20a42 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Thu, 1 Oct 2026 19:43:23 +0200 Subject: [PATCH 2/5] docs(specs): passive metrics in Python, after the plan review found awk unfit MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Gate-A plan pass 1 found 16 Majors, most of them awk failing to parse the quoted, escaped §5 record grammar. Daniel chose Python (3.8+, standard library); the suite stays POSIX sh. This revision changes the language and records the narrowings the implementation needed, marked (revision): one resolved SHA read with --no-replace-objects, UTF-8-forced log output, a literal column-2 fingerprint compare, a skip-reason excerpt bounded by a blank line or the next record, one ordering rule, control characters shown as \xNN, a conflict-only empty state, and suite isolation from the caller's git config. Gate-A spec cycle closed. Pass 3 is clean at the derived floor (0 Blockers, 0 Majors); pass 1's one Major (log output encoding) was repaired and pass 2 confirmed it. Passes 2 and 3 reviewed the same text. Their Minors are collected, not iterated, and the plan carries the ones that affect the implementation. The committed spec is byte-identical to the text pass 3 reviewed (sha256 4e6d08282a25005517bf691b9b528397a13b9bf7d2359874419e4592de0ba5a7). Docs-only change (docs/**.md): Gate B is N/A per CLAUDE.md §5. cycle p9yvzn4fvi; floor 3 per {docs/superpowers/stories/2026-08-04-passive-metrics-over-the-ledger-story.md (level 1)}; hook reminder threshold absent cycle p9yvzn4fvi; Gate-A spec (passes 1-3, gpt-6-astra): Findings 12,12,16. Blockers 0,0,0. Majors 1,0,0. --- .../2026-10-01-passive-metrics-design.md | 138 ++++++++++++------ 1 file changed, 95 insertions(+), 43 deletions(-) diff --git a/docs/superpowers/specs/2026-10-01-passive-metrics-design.md b/docs/superpowers/specs/2026-10-01-passive-metrics-design.md index 341ac64..deb9dc6 100644 --- a/docs/superpowers/specs/2026-10-01-passive-metrics-design.md +++ b/docs/superpowers/specs/2026-10-01-passive-metrics-design.md @@ -11,14 +11,20 @@ counts and lists. It does not compute shares, ratios or verdicts, because every its own edge cases (missing values, zero denominators), and the story only asks that the data become readable and checkable. +**Revision, 2026-10-01 (after Gate-A plan pass 1).** The first revision fixed the language as POSIX +`sh` + `awk`. Plan review found 16 Majors, and most of them were awk parsing the quoted, escaped §5 +record grammar: quote handling, control characters, and differences between awk implementations. +Daniel chose Python. This revision changes the language and records the narrowings the implementation +needed; every changed passage is marked *(revision)*. Nothing else changed. + ## §1 What is added or edited | Path | Change | |---|---| -| `scripts/ledger-metrics.sh` | **New.** POSIX `sh` + `awk` + `git`. Reads, prints a report on stdout, writes nothing. | -| `scripts/ledger-metrics.test.sh` | **New.** Its regression suite, built on throwaway fixture repositories. | -| `AGENTS.md` | § Commands: the quality and lint rows gain shellcheck runs for **both** new files; the quality row (not the lint row) also gains a run of the suite. Layout tree: both files. **Boundaries** paragraph: the list of executable artifacts gains this script and its test. | -| `.github/workflows/ci.yml` | The shellcheck step gains both new files, and the checker-suite step gains the suite. Any step name or comment that enumerates the checkers is updated too. | +| `scripts/ledger-metrics.py` | **New.** Python 3.8 or later, standard library only, plus `git`. Reads, prints a report on stdout, writes nothing. *(revision)* | +| `scripts/ledger-metrics.test.sh` | **New.** Its regression suite in POSIX `sh`, built on throwaway fixture repositories; it runs the script with `python3`. | +| `AGENTS.md` | § Commands: the quality and lint rows gain a shellcheck run for the suite; the quality row (not the lint row) also gains a run of the suite. The prerequisites paragraph names `python3` 3.8 or later, addressed by name like `dash` because it is the system interpreter, not a pinned tool. Layout tree: both files. **Boundaries** paragraph: the list of executable artifacts gains this script and its test. *(revision)* | +| `.github/workflows/ci.yml` | The shellcheck step gains the suite, and the checker-suite step runs it. *(revision)* Any step name or comment that enumerates the checkers is updated too. | | `README.md` | The Contributing paragraph that counts the executables and checker suites is updated, and one line says how to run the report. | **Before editing, grep for every other place that enumerates the executables or checker suites** @@ -32,24 +38,33 @@ contain it, so invariants 11 and 12 do not bind, and the plugin version does not ## §2 Interface ``` -sh scripts/ledger-metrics.sh [] +python3 scripts/ledger-metrics.py [] ``` - **The ref is resolved once.** The script runs `git rev-parse --verify ^{commit}` (default `HEAD`) once at start and uses only the resulting 40-character SHA afterwards. The ledger is read as - `git show :docs/hardening-log.md`, and the commit bodies as `git log `. Every commit + `git cat-file blob :docs/hardening-log.md`, and the commit bodies as + `git -c i18n.logOutputEncoding=UTF-8 log --encoding=UTF-8 -z --format='%H %ct%n%B' `. Forcing + UTF-8 output keeps the NUL separators intact whatever the repository's log encoding is; both sources + are decoded as UTF-8, and an undecodable byte becomes U+FFFD, which makes a record containing it + unparsed rather than silently different. Every git call runs with `--no-replace-objects`, so the + objects read are the ones the SHA names and not `git replace` substitutes. These two settings change + only how git reads, for this script's own calls; they are not a write and change no configuration. A + record header that is not a 40-hex SHA and a number is an exit-1 error. *(revision)* Every commit reachable from that SHA is read, not just the first-parent chain, because an ordinary merge can carry cycle records on its merged side. The working tree is never read. So uncommitted ledger edits are not counted, and the report header says so. - **Shallow clones.** If `git rev-parse --is-shallow-repository` prints `true`, the report header says that history is truncated. Every "none found" statement in the cycle and checkpoint sections then carries `(history truncated: absence not established)`. -- **Locale.** The script sets `LC_ALL=C` so sorting and output do not depend on the environment. +- **Ordering** compares Python strings and integers, so it does not depend on the locale. *(revision)* - **Output** goes to stdout only. Exit `0` when the report is complete. Exit `1` with a one-line `ledger-metrics: ` on stderr when a source cannot be read. Each cause has its own message: not a git repository, the ref does not resolve to a commit, the ledger is absent at that commit, the ledger path is not a regular file there (checked with `git ls-tree`: mode `100644` or `100755`, type - `blob` — a directory or symlink is rejected), the ledger cannot be read, or `git log` fails. Malformed input lines are **not** an exit: they are counted and listed (§3, §4). + `blob` — a directory or symlink is rejected; `--full-tree`, so the caller's subdirectory does not + matter), the ledger cannot be read, or `git log` fails. Git's own stderr is captured and not shown, + so the one line is the only error. *(revision)* Malformed input lines are **not** an exit: they are counted and listed (§3, §4). - **The script writes nothing** — no file, no git ref, no config. The suite checks the parts of this a test can observe (§6). The rest is held by reading the script, and §6 says which part is which. - **What the script does not control.** It runs ordinary read-only git commands. Git itself can still @@ -60,17 +75,19 @@ sh scripts/ledger-metrics.sh [] ## §3 Ledger section — recurrence and rung holding -**What a row is: exactly the `harden-finding` match.** A line is a row for fingerprint `F` if it matches -the skill's recurrence grep, `^\| *[0-9-]{10} *\| *F *\|`. The script uses that same pattern, so its -count for `F` equals the count the skill's grep gives. This answers the story's second open question: +**What a row is: the `harden-finding` match.** A line is a row if it matches +`^\| *[0-9-]{10} *\| * *\|`, the shape of the skill's recurrence grep. Its fingerprint is that +column-2 cell with spaces trimmed, compared **literally**. For every regular (kebab-case) fingerprint the +count equals the count the skill's grep gives. For an irregular one the grep would treat the text as a +pattern, so the two can differ, and no grep command is offered for it *(revision)*. This answers the story's second open question: there is one definition, not two. Cells are split on **unescaped** `|`. A matching row with a cell count other than seven is **still counted**, as the grep counts it, and is also listed as `irregular` with its line number. Irregular-width rows are **excluded from rung holding**, and their rung shows as `?` in the recurrence line, because their rung cell cannot be located reliably. **Fingerprint spelling.** Taxonomy classes are kebab-case. A fingerprint cell that does not match -`^[a-z0-9]+(-[a-z0-9]+)*$` is listed as `irregular`. No verification command is printed for it, because -its text would be pasted into a shell command. +`^[a-z0-9]+(-[a-z0-9]+)*$` is listed as `irregular`. The verification command below does not fit it, +and the report says so. `Superseded rows` entries are list lines, not rows, so they never match. The ledger header says this itself: a superseded row "keeps matching the column-2 grep, and keeps counting". @@ -82,10 +99,11 @@ itself: a superseded row "keeps matching the column-2 grep, and keeps counting". ``` `lines` are ledger line numbers at the SHA, ascending. `rungs` are the rung cells in file order. After -the list comes the exact command a reader can run to check any one count: +the list comes one command template a reader can run to check any regular fingerprint's count, with +`` being the SHA in the report header: ``` -git show :docs/hardening-log.md | grep -cE '^\| *[0-9-]{10} *\| * *\|' +git show :docs/hardening-log.md | grep -cE '^\| *[0-9-]{10} *\| * *\|' ``` **Rung holding.** It counts only rung cells that are one of the values the ledger uses: `1 prose`, @@ -120,23 +138,26 @@ not followed by a semicolon. A candidate that fails the full grammar is listed a **unparsed**, with its commit, and excluded. **Deduplication.** A record with a real nonce is keyed by its exact text. Text that appears in several -commits is counted once, and every commit carrying it is listed. `cycle none (pre-rule)` records are -**not** deduplicated, because identical text in two commits may be two different legacy cycles: each -occurrence is listed with its commit and flagged `may duplicate another pre-rule record`. +commits is counted once, and every commit carrying it is listed, oldest first. A real-nonce line +repeated inside one commit body counts once for that commit. `cycle none (pre-rule)` records are +**never** deduplicated, not even inside one body, because identical text may be two different legacy +cycles: every occurrence is listed and flagged `may duplicate another pre-rule record`. **Ordering, everywhere in this section** — records, cycle lines, unparsed lines and conflict groups: by the **earliest** committer date (`%ct`) among the commits carrying the record (for a group, among all -its records), then that commit's SHA, then the line text. +its records), then that commit's SHA, then the printed line. Committer dates compare as integers. +This one rule also orders the cycle lines; the format line below adds nothing to it. *(revision)* **Grouping.** Records with the same real nonce are grouped. A group joins cleanly only if it has at most one provenance line and at most one curve or skip record. A group with two different provenance lines, two different curves, a curve and a skip record, or two different skip records is reported as a **conflict** with all its records, and it is not listed as a cycle. Copying errors and nonce collisions both look like this. CLAUDE.md says the nonce is collision-resistant, not collision-proof, so grouping by nonce is a strong default and not a guarantee, -and the report says so. `cycle none (pre-rule)` records are never grouped, because that field -identifies nothing. +and the report's "cannot answer" list says so. `cycle none (pre-rule)` records are never grouped, +because that field identifies nothing. Provenance lines in a conflict group still appear in the +checkpoint evidence lists (§5), because a conflict does not make a floor claim disappear. -**Per cycle, one line,** sorted by first commit date, then nonce: +**Per cycle, one line:** ``` passes floor set findings blockers majors @@ -147,13 +168,28 @@ identifies nothing. series as story criterion 5 requires. - `-` means the record has no provenance line. A provenance line with no curve and no skip record is printed with `no curve`. -- A skip record is printed as `skipped`, followed by its reason: CLAUDE.md §5 says the reason is the text +- A skip record is printed as `skipped`, followed by `reason excerpt:` and its reason: CLAUDE.md §5 says the reason is the text immediately following the marker in the same commit body. The script takes the lines after the marker - up to the next blank line, joined with spaces. If the next non-empty line is itself a candidate, or - there is nothing after the marker, it prints `no reason found`. A skip record is keyed by marker + up to the next blank line **or the next candidate line**, joined with spaces, so a record directly + after the marker is never absorbed into the reason. It is labelled an **excerpt** because a reason can + run past a blank line and only the first paragraph is shown. If that leaves nothing, it prints + `no reason found`. *(revision)* A skip record is keyed by marker **and** reason, so two copies with different reasons are two records and form a conflict. -**Empty states.** `no cycle records`, `no unparsed lines` and `no conflicts` are printed when they apply. +**Empty states.** `no cycle records`, `no unparsed lines` and `no conflicts` are printed when they apply, +each with the shallow-history suffix from §2 when it applies. If valid records exist but every group is +a conflict, the cycle list says `no joined cycles (every record is in a conflict below)`. *(revision)* + +**Raw text in output.** Unparsed and conflicting lines are printed as they were written, except that +any control character is shown as `\xNN`, so a malformed record cannot break the report's lines or move +the terminal cursor. *(revision)* + +**Grammar details the parser enforces** (CLAUDE.md §5 Mechanics): a quoted path or model may use only +the escapes `\"` and `\\`, and any control character makes the record unparsed; a repeated story path +is compared after decoding quotes and escapes, so `a.md` and `"a.md"` are the same path; a pass spec +that expands to more than 10000 passes makes the record unparsed, so a malformed range cannot exhaust +memory (a single large pass number is fine); each count series must have exactly one value per expanded +pass; per-pass model keys must be exactly the expanded passes, in order. *(revision)* ## §5 The review-loop comparison (story criterion 5) @@ -199,22 +235,31 @@ story's "first post-merge" one. `scripts/ledger-metrics.test.sh` builds throwaway repositories under a `mktemp -d` directory. It commits a fixture ledger and fixture commit bodies, then compares the script's output **line for line** against -expected text written in the test. Fixtures cover: +expected text written in the test. **Isolation** *(revision)*: the suite sets `GIT_CONFIG_GLOBAL=/dev/null` +and `GIT_CONFIG_NOSYSTEM=1`, unsets inherited repository and configuration variables (`GIT_DIR`, +`GIT_INDEX_FILE`, `GIT_CONFIG_COUNT`, `GIT_CONFIG_PARAMETERS`, `GIT_REPLACE_REF_BASE` and similar), creates repositories with `--template=` and `--object-format=sha1`, and fixes identity and +dates per commit, so SHAs are reproducible and the developer's git setup is never read or touched. +Fixtures cover: - **Ledger:** two fingerprints with different counts; an escaped `\|` inside a finding; a supersession list line (not counted); a row with too few cells (counted and listed as irregular); a fingerprint - with an uppercase letter (irregular, no command printed); a `pending` row; a row that is followed and + with an uppercase letter (irregular; the report says the command template does not fit it); a `pending` row; a row that is followed and one that is last of its fingerprint; an empty ledger (`no rows`). - **Cycles:** a provenance line and a curve with the same nonce (joined); a curve with no provenance (`-`); a provenance line with no curve (`no curve`); a skip record; a `none (pre-rule)` pair (not grouped); the same curve in two commits (once, both commits listed); two different curves under one nonce, a curve and a skip record under one nonce, and two skip records with different reasons (each a - conflict, not listed as a cycle); a record only on the merged side of a merge commit (found); a prose line + conflict, not listed as a cycle); two different provenance lines under one nonce, one of them level 0 + (a conflict, and still listed in the checkpoint evidence); a history where every group is a conflict + (the cycle list says so); the same pre-rule curve twice in one body (listed twice); a record only on + the merged side of a merge commit (found); a prose line starting with "cycle " (not a candidate); a candidate with a 7-character nonce (unparsed). - **Grammar variants:** a quoted story path containing an escaped `\"`; a quoted model; discontiguous pass ranges (`1-3,5`); per-pass models with `+`; a `?` in one series (printed verbatim, counted in - `(? n)`). Malformed variants listed as unparsed: a repeated story path; a count list longer than the - pass list; a per-pass model list missing a pass; a candidate with no space after the semicolon. + `(? n)`); a quoted model containing `; `, `)` and `:` (valid). Malformed variants listed as unparsed: + a repeated story path, and the same path once bare and once quoted; a quoted model with an invalid + escape (`\q`); a quoted model containing a tab (shown as `\x09`); a count list longer than the pass + list; a per-pass model list missing a pass; a candidate with no space after the semicolon. - **Skip records:** a skip record followed by a two-line reason (both lines printed); one with nothing after it; one directly followed by another cycle record (both `no reason found`). - **Rungs:** an empty rung cell and a rung `0` (each `irregular rung`, not counted). @@ -222,37 +267,44 @@ expected text written in the test. Fixtures cover: level-1 set and a curve (appears in the comparison); `no profiled cycles with a curve`; a level-0 set recorded with floor 3 (appears in the first list); a `floor 1` line with an `(unprofiled)` entry (appears in the second list, not licensed); both lists empty. - **Inputs:** a working-tree edit to the ledger that is not committed (no change in output); each - exit-1 cause, matched by its message, including the ledger path committed as a directory and as a - symlink; a shallow clone (header says history is truncated). + exit-1 cause that a fixture can produce, matched by its message: not a repository, an unresolvable + ref, the ledger absent, committed as a directory, committed as a symlink, its blob missing (cannot be + read), and an ancestor commit missing (`git log` fails); a shallow clone (header says history is + truncated). The unexpected-record-header error is not reachable from a well-formed repository and is + covered by reading the script. -**What the no-write check covers, and what it does not.** Before and after each run, the suite records +**What the no-write check covers, and what it does not.** Before and after **every** run *(revision)*, the suite records a listing of the whole fixture directory, `.git` included, with each file's mode and a checksum of its contents. The two must be identical. That catches a write that leaves a file's content, mode or presence different afterwards. It does not catch a rewrite with the same bytes, a change that is undone before the run ends, a file created and deleted during the run, or a write elsewhere on the machine. For the -script's own commands those are held by reading it: it contains no file redirection other than to -`/dev/null`, and no git command that writes. That reading is a review check, not a test, and is labelled +script's own commands those are held by reading it: it opens no file for writing and runs no git +command that writes. That reading is a review check, not a test, and is labelled as one. Writes git makes because of the caller's environment are outside both (§2). **The counterfactual — the observation against the prior state.** Before this change there is no script, so nothing can produce the recurrence report. The suite's first case runs -`sh scripts/ledger-metrics.sh` in a fixture where the script path does not exist and checks that it -fails with "No such file". That is the prior state, observed. It is also why the check is weak on its +`python3 scripts/ledger-metrics.py` in a fixture where the script path does not exist and checks that +it fails. That is the prior state, observed. It is also why the check is weak on its own: a missing-file error shows the tool is new, not that it counts right. **The negative control** shows the count assertion can fail. The suite runs its fingerprint-count case against a temp copy of the script where `sed` changes the count by one. That run must **exit 0 from the -script** and fail **on the count assertion**, by its name. A broken copy that fails some other way, or -passes, makes the suite fail itself. Only then does it run the real script. +script**, and **the same comparison function every golden case uses** must reject its report, with the +count column showing the off-by-one value. *(revision)* A broken copy that +fails some other way, or whose report still matches, makes the suite fail itself. Only then does it run the real script. ## §7 What it cannot answer, printed in every report - **Unlogged recurrences.** The ledger only knows recurrences someone hardened. - **Whether a guard held.** Row succession is not guard failure, and a missing later row is not success (§3). -- **Cycles with no record at all.** Before 0.11.0 the curve was a habit, not a rule. A cycle whose - author wrote neither a provenance line nor a curve is invisible; a malformed record shows up as - unparsed. Provenance-only and skipped cycles are visible, as §4 describes. +- **Cycles with no record at all.** Before 0.11.0 the curve was a habit, not a rule. A cycle with no + provenance line, curve or skip record is invisible. A malformed record shows up as unparsed only when + its line still starts like a record (`cycle ;`). Provenance-only and skipped cycles are visible, + as §4 describes. +- **Nonce attribution.** Records are grouped by nonce, which is collision-resistant, not + collision-proof; two cycles that drew the same nonce read as one. - **Findings files.** The script does not read `.context/`. In this repository the findings files are tracked under `.context/codex-reviews/`; in other projects they may be gitignored. Either way, the report counts only what commit bodies say. From 9230e4fc5f063e4a523d6fd88347b7b5dbbd88d1 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Thu, 1 Oct 2026 20:56:22 +0200 Subject: [PATCH 3/5] docs(plans): add the passive-metrics implementation plan MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Gate-A plan cycle closed on a zero-finding pass 6. Pass 1 reviewed the earlier awk plan; its 16 Majors led to the Python revision of the spec (49b90f8). Passes 2-5 found 5, 1, 1 and 3 Majors in the Python plan, each repaired and confirmed by the next pass. The committed plan is byte-identical to the text pass 6 reviewed (sha256 c0965410c3f0519fc3a8625c2003ca43532fe3726cc9ba2b9a23c9a7ead06c56). Its embedded script and suite were run as a prototype before every pass; the suite stands at 30/30 under sh and dash. Docs-only change (docs/**.md): Gate B is N/A per CLAUDE.md §5. cycle mtf7ua7qze; floor 3 per {docs/superpowers/stories/2026-08-04-passive-metrics-over-the-ledger-story.md (level 1)}; hook reminder threshold absent cycle mtf7ua7qze; Gate-A plan (passes 1-6, gpt-6-astra): Findings 34,19,9,4,4,0. Blockers 0,0,0,0,0,0. Majors 16,5,1,1,3,0. --- .../plans/2026-10-01-passive-metrics.md | 1180 +++++++++++++++++ 1 file changed, 1180 insertions(+) create mode 100644 docs/superpowers/plans/2026-10-01-passive-metrics.md diff --git a/docs/superpowers/plans/2026-10-01-passive-metrics.md b/docs/superpowers/plans/2026-10-01-passive-metrics.md new file mode 100644 index 0000000..3da0b32 --- /dev/null +++ b/docs/superpowers/plans/2026-10-01-passive-metrics.md @@ -0,0 +1,1180 @@ +# Passive Metrics Implementation Plan + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. + +**Story:** `docs/superpowers/stories/2026-08-04-passive-metrics-over-the-ledger-story.md` — read the profile from its header at every gate call. + +**Goal:** Add `scripts/ledger-metrics.py`, a read-only report over the hardening ledger and the review-cycle records in commit bodies, with its POSIX-sh regression suite, and wire the suite into the battery, CI and the docs. + +**Architecture:** One Python 3.8+ standard-library script resolves one commit, reads the ledger blob and the `git log -z` bodies at it, parses the §5 record grammar with a small hand-written scanner (quoted strings, escapes, control characters), and prints the report. The suite builds throwaway, config-isolated repositories with fixed commit dates, so SHAs and output are deterministic, and compares the report with expected text. + +**Tech Stack:** Python 3.8+ (standard library), `git`, POSIX `sh` for the suite, `shellcheck` 0.11.0. + +**Spec:** `docs/superpowers/specs/2026-10-01-passive-metrics-design.md` — revision committed in `49b90f8` (Gate-A spec cycle `p9yvzn4fvi`, closed). It supersedes the awk revision (`c60b78b`). + +**This plan replaces the awk plan reviewed in Gate-A plan pass 1** (`.context/codex-reviews/gate-a-plan-mtf7ua7qze-pass-1.md`). That pass's findings were about awk parsing, mawk semantics and suite isolation; the table at the end says where each landed. + +## Global Constraints + +- Worktree `/Users/daniel/DEVELOPMENT/APPS/dwk-metrics`, branch `passive-metrics`. Paths are relative to it. +- The script is not shipped: no file under `plugins/` changes, so no version bump (spec §1). +- The script writes nothing itself: stdout only, and every git call is read-only (spec §2). +- Python 3.8 or later, standard library only; the suite passes `shellcheck --shell=sh --exclude=SC2015`, like the other suites. +- Every report states its limits verbatim: the three comparison caveats (spec §5) and the "cannot answer" list (spec §7). +- Unexpected repository state is a stop and a question to Daniel, never an automated stash, rebase or reset. + +## Rulings carried from the spec cycle's collected Minors + +The spec is closed and is not edited. Where the collected Minors of passes 2–3 (`.context/codex-reviews/gate-a-spec-p9yvzn4fvi-pass-2.md`, `-pass-3.md`) pointed at an implementation choice, the plan takes it as below; the Gate-B call names these. + +1. **No pass-count ceiling.** The expanded pass count must equal the length of the supplied count series, and that is checked **before** anything is allocated, so a huge range costs nothing and needs no cap. Python's integer-digit cap (3.11+) is lifted at start, so every Python accepts the same pass numbers. (Replaces the spec's "more than 10000" sentence, which pass 2 found stricter than §5's grammar.) +2. **Git reads ignore grafts, shallow-file overrides and signatures and cannot be ambiguous:** every call sets `GIT_GRAFT_FILE` to the null device and drops an inherited `GIT_SHALLOW_FILE`, besides `--no-replace-objects`; `git log` adds `--no-show-signature` and ends with ` --`. +3. **Shallow state is read before and after the walk**; either reading `true` marks the history truncated. A boundary that exists only between the two readings is not detected — a stated limit, printed in the report's "cannot answer" list. +4. **Output and the one-line error are UTF-8 whatever the locale** (`backslashreplace` for anything unencodable), and **every source-derived field** — fingerprints, rungs, story sets, skip excerpts, raw lines — goes through the same control-character display. +5. **The printed check command** is `git --no-replace-objects show …`, so it reads what the script read. +6. **A skip record's key is its marker plus the excerpt**, so two copies differing only after the first paragraph are one record. The excerpt label says the text is partial. +7. Left as stated limits, not changed: an undecodable byte becomes U+FFFD **before** grouping, so two fingerprints or story paths differing only in undecodable bytes count as one (this repository's ledger and commit bodies are UTF-8, so the case does not arise here); and `\xNN` display can coincide with a literal backslash sequence in what is printed. Both are printed in the report's "cannot answer" list. + +## Review Focus + +1. **A different Python on CI.** Ubuntu 24.04's `python3` (3.12) runs the suite in CI; locally 3.12 and 3.9 gave identical reports. → Task 3 Step 5 reads the CI log for `30 passed, 0 failed`. +2. **A real `main` with ~60 commits.** Expect a full report quickly and no unparsed lines. → Task 1 Step 6. +3. **A body line that starts with `cycle` in prose.** Expect it ignored. → fixture line "cycle closed on the zero-finding exit below the floor." +4. **Different git versions.** Fixture SHAs depend only on content, identity and dates; a mismatch on CI is a stop, not a re-record. +5. **Uncommitted ledger edits.** Expect them ignored. → suite case "working-tree ledger edit does not change the report". + +--- + +## File map + +| Path | Change | Task | +|---|---|---| +| `scripts/ledger-metrics.test.sh` | Create | 1 | +| `scripts/ledger-metrics.py` | Create | 1 | +| `AGENTS.md` | Layout tree, Boundaries, § Commands quality + lint rows, prerequisites | 2 | +| `.github/workflows/ci.yml` | Lint step, suite step, two comments | 2 | +| `README.md` | Contributing: wording + one paragraph on running the report | 2 | +| `CLAUDE.md` | §5 Mechanics: the sentence saying the metrics consumer "does not exist yet" | 2 | + +--- + +### Task 1: The suite, then the script + +**Files:** +- Create: `scripts/ledger-metrics.test.sh` +- Create: `scripts/ledger-metrics.py` + +**Interfaces:** +- Produces: `python3 scripts/ledger-metrics.py []` → report on stdout, exit 0; exit 1 with `ledger-metrics: ` on stderr. `sh scripts/ledger-metrics.test.sh` → `ok -`/`FAIL -` lines, then ` passed, failed`, exit 0 only when `m` is 0. + +- [ ] **Step 1: Write the suite** + +Create `scripts/ledger-metrics.test.sh` with exactly this content: + +````sh +#!/bin/sh +# Regression suite for ledger-metrics.py. +# +# Spec: docs/superpowers/specs/2026-10-01-passive-metrics-design.md §6. Every case builds +# a throwaway repository with fixed author/committer dates, so commit SHAs and therefore +# the report are deterministic, and compares the report against expected text. +# +# Order matters: the prior-state counterfactual and the negative control run FIRST, so +# a suite that cannot fail is caught before any green case is believed. +set -u + +SCRIPT="$(cd "$(dirname "$0")" && pwd)/ledger-metrics.py" +pass_n=0; fail_n=0 +pass() { pass_n=$((pass_n + 1)); printf 'ok - %s\n' "$1"; } +fail() { fail_n=$((fail_n + 1)); printf 'FAIL - %s\n' "$1"; } + +work=$(mktemp -d) || work='' +# Abort rather than continue with an empty $work: every path below is built under it. +if [ -z "$work" ] || [ ! -d "$work" ]; then + printf 'FAIL - could not create a temporary directory; refusing to run\n' >&2 + exit 1 +fi +trap 'rm -rf "$work"' EXIT +# An empty HOME and XDG_CONFIG_HOME: git finds no global attributes, ignore or config files +# there, so nothing from the developer's home can change what a fixture commit stores. +mkdir -p "$work/home/.config" && HOME="$work/home" && XDG_CONFIG_HOME="$work/home/.config" && + export HOME XDG_CONFIG_HOME + +# Isolation: no global or system git config, no inherited repository variables, no init +# templates. Identity and dates are fixed per invocation, so every SHA is reproducible and +# the suite never touches the developer's git setup. +GIT_CONFIG_GLOBAL=/dev/null; GIT_CONFIG_NOSYSTEM=1; GIT_ATTR_NOSYSTEM=1 +export GIT_CONFIG_GLOBAL GIT_CONFIG_NOSYSTEM GIT_ATTR_NOSYSTEM +unset GIT_DIR GIT_WORK_TREE GIT_INDEX_FILE GIT_OBJECT_DIRECTORY GIT_ALTERNATE_OBJECT_DIRECTORIES \ + GIT_CEILING_DIRECTORIES GIT_DEFAULT_HASH GIT_COMMON_DIR GIT_NAMESPACE \ + GIT_CONFIG_COUNT GIT_CONFIG_PARAMETERS GIT_REPLACE_REF_BASE GIT_SHALLOW_FILE \ + GIT_GRAFT_FILE GIT_NO_REPLACE_OBJECTS +tick=1700000000 +gitc() { + GIT_AUTHOR_NAME=t GIT_AUTHOR_EMAIL=t@example.com GIT_COMMITTER_NAME=t GIT_COMMITTER_EMAIL=t@example.com \ + GIT_AUTHOR_DATE="$tick +0000" GIT_COMMITTER_DATE="$tick +0000" \ + git -c user.name=t -c user.email=t@example.com -c commit.gpgsign=false \ + -c init.defaultBranch=main "$@" +} +# Commit everything with the given message (printf %b escapes), one second later each time. +# The message is an argument, never piped in: a pipeline runs `commit` in a subshell, +# which would lose the tick increment and give every commit the same date. +commit() { tick=$((tick + 1)); printf '%b' "$1" > "$work/msg"; gitc add -A && gitc commit -q --allow-empty -F "$work/msg"; } + +newrepo() { rm -rf "${work:?}/$1"; mkdir -p "$work/$1/docs" && (cd "$work/$1" && gitc init -q --template= --object-format=sha1 .); } +# Every run is wrapped in the no-write check: a run that changes the fixture leaves a marker +# outside it, and the suite fails at the end if any marker exists (spec §6). +# run DIR SCRIPT [REF]: stdout only, status preserved. +run() { + _b=$(snapshot "$1") + (cd "$1" && python3 "$2" ${3+"$3"}); _s=$? + [ "$_b" = "$(snapshot "$1")" ] || : > "$work/WROTE.$(basename "$1")" + return "$_s" +} +report() { run "$work/$1" "$2" ${3+"$3"} 2>&1; } + +# A whole-directory fingerprint: every path with its mode and size, plus a checksum of +# every file's contents. Equal before and after a run means the run left no file +# changed, added or removed in the fixture (spec §6 states what this does NOT catch). +snapshot() { (cd "$1" && ls -lAR . && find . -type f -exec cksum {} + | sort); } + +# The one comparison every golden case uses; the negative control calls it too, so a +# comparison that stopped comparing would be caught there. +same() { diff "$work/expected" "$work/actual" > "$work/diff"; } +expect() { # name, expected (stdin), actual + printf '%s\n' "$2" > "$work/actual" + cat > "$work/expected" + if same; then pass "$1" + else fail "$1"; sed 's/^/ /' "$work/diff"; fi +} + +# ---- fixture: the ledger -------------------------------------------------------- +LEDGER_HEAD='# Hardening log + +Columns: date, fingerprint, finding, source, severity, rung, ref. + +| date | fingerprint | finding | source | severity | rung | ref | +|------|-------------|---------|--------|----------|------|-----|' +mk_ledger_repo() { + newrepo L + { + printf '%s\n' "$LEDGER_HEAD" + printf '%s\n' '| 2026-01-01 | alpha-one | f1 | gate-a | major | 1 prose | r |' + printf '%s\n' '| 2026-01-02 | beta | a \| piped finding | bot | minor | P std | r |' + printf '%s\n' '| 2026-01-03 | alpha-one | f3 | bot | major | 2 lint | r |' + printf '%s\n' '| 2026-01-04 | alpha-one | short row | bot |' + printf '%s\n' '| 2026-01-05 | gamma | f5 | bot | nit | pending | r |' + printf '%s\n' '| 2026-01-06 | Bad-Case | f6 | bot | nit | 1 prose | r |' + printf '%s\n' '| 2026-01-07 | delta | f7 | bot | nit | | r |' + printf '%s\n' '| 2026-01-08 | delta | f8 | bot | nit | 0 | r |' + printf '%s\n' "- 2026-01-09 · supersedes 2026-01-01 \`alpha-one\` \"f1\" · wrong · see row" + } > "$work/L/docs/hardening-log.md" + (cd "$work/L" && commit 'ledger\n') +} + +ledger_section() { sed -n '/^== Ledger: recurrence/,/^== Git: review cycles/p' | sed '$d'; } +LEDGER_EXPECTED='== Ledger: recurrence (rows matched exactly as harden-finding greps column 2) +3 alpha-one lines 7,9,10 rungs 1 prose > 2 lint > ? +2 delta lines 13,14 rungs > 0 +1 Bad-Case lines 12 rungs 1 prose +1 beta lines 8 rungs P std +1 gamma lines 11 rungs pending +check one count: git --no-replace-objects show :docs/hardening-log.md | grep -cE '"'"'^\| *[0-9-]{10} *\| * *\|'"'"' + ( is the commit in the header; use a fingerprint from the list. No command fits an irregular fingerprint, because grep would read it as a pattern.) + +== Ledger: rung holding +1 prose landed 2 followed-by-same-fingerprint 1 last-of-fingerprint 1 +2 lint landed 1 followed-by-same-fingerprint 1 last-of-fingerprint 0 +P std landed 1 followed-by-same-fingerprint 0 last-of-fingerprint 1 +pending 1 +"followed" means a later row with the same fingerprint exists - nothing more. It does not mean +the rung failed (a later row can be a different sub-shape, an out-of-scope guard, or a resolved +prerequisite), and "last" does not mean it held (an unhardened recurrence leaves no row). + +== Ledger: irregular rows +irregular width: line 10 (4 cells) +irregular fingerprint: line 12 +irregular rung: line 13 (rung '"''"') +irregular rung: line 14 (rung '"'0'"')' + +# ---- 1. counterfactual: the prior state has no script -------------------------- +mk_ledger_repo +out=$(report L "$work/L/scripts/ledger-metrics.py"); st=$? +if [ "$st" -ne 0 ] && [ ! -e "$work/L/scripts/ledger-metrics.py" ]; then + pass "prior state: no script exists, so no report can be produced (exit $st)" +else fail "prior state: expected a failing run with no script, got exit $st"; fi + +# ---- 2. negative control: a count off by one must fail the count comparison ---------- +# The broken copy must still run to completion (exit 0), and its report must differ from +# the expected text in the count column - here alpha-one reads 4 instead of 3. A copy +# that crashes, or whose report still matches, proves nothing about the comparison. +sed 's/n = len(by_fp\[fp\])$/n = len(by_fp[fp]) + 1/' "$SCRIPT" > "$work/broken.py" +if cmp -s "$SCRIPT" "$work/broken.py"; then + fail "negative control: the mutation did not apply, so it proves nothing" +else + out=$(report L "$work/broken.py"); st=$? + printf '%s\n' "$LEDGER_EXPECTED" > "$work/expected" + printf '%s\n' "$out" | ledger_section > "$work/actual" + if [ "$st" -eq 0 ] && ! same && + grep -qx '4 alpha-one lines 7,9,10 rungs 1 prose > 2 lint > ?' "$work/actual"; then + pass "negative control: the broken copy exits 0 and the count comparison catches it" + else + fail "negative control: broken copy (exit $st) was not caught by the count comparison" + fi +fi + +# ---- 3. ledger section ------------------------------------------------------------ +out=$(report L "$SCRIPT"); st=$? +[ "$st" -eq 0 ] && pass "ledger fixture: exit 0" || fail "ledger fixture: exit $st" +expect "ledger: recurrence, rung holding and irregular rows" "$(printf '%s\n' "$out" | ledger_section)" </$c/; s//alpha-one/") +n=$(cd "$work/L" && sh -c "$cmd") +[ "$n" = 3 ] && pass "printed check command reproduces alpha-one = 3" || fail "printed check command gave '$n' ($cmd)" + +# The "cannot answer" footer is printed in full. +expect "footer: what the report cannot answer" "$(printf '%s\n' "$out" | sed -n '/^== What this report cannot answer/,$p')" <<'FOOTER' +== What this report cannot answer +- Unlogged recurrences: the ledger only knows recurrences someone hardened. +- Whether a guard held: row succession is not guard failure, and a missing later row is not success. +- Cycles with no record at all: before 0.11.0 the curve was a habit, not a rule. A cycle with no provenance line, curve or skip record is invisible; a malformed record shows up as unparsed only when its line still starts like a record ("cycle ;"). +- Nonce attribution: records are grouped by nonce, which is collision-resistant, not collision-proof; two cycles that drew the same nonce read as one. +- Findings files: this script does not read .context/. It counts only what commit bodies say. +- Undecodable bytes: both sources are read as UTF-8 and an invalid byte becomes U+FFFD before anything is grouped, so two fingerprints or paths differing only in invalid bytes count as one. Control characters are shown as \xNN, which can look like a literal backslash sequence. +- A shallow boundary that appears and disappears during the run: shallow state is checked before and after reading history, not during it. +- Whether a curve is true: see point 1 above. +- Cost and duration: that is vision step 2c's telemetry, out of scope here. +FOOTER + +# Uncommitted ledger edits are not read. +printf '%s\n' '| 2026-01-10 | alpha-one | uncommitted | bot | nit | 1 prose | r |' >> "$work/L/docs/hardening-log.md" +out2=$(report L "$SCRIPT"); st=$? +[ "$st" -eq 0 ] && [ "$out" = "$out2" ] && pass "working-tree ledger edit does not change the report" || fail "working-tree ledger edit changed the report" + +# Empty ledger. +newrepo E; printf '%s\n' "$LEDGER_HEAD" > "$work/E/docs/hardening-log.md"; (cd "$work/E" && commit 'e\n') +oute=$(report E "$SCRIPT"); st=$? +[ "$st" -eq 0 ] && pass "empty ledger: exit 0" || fail "empty ledger: exit $st" +printf '%s\n' "$oute" | grep -qx 'no rows' && pass "empty ledger prints 'no rows'" || fail "empty ledger" +printf '%s\n' "$oute" | grep -qx 'no cycle records' && pass "no records prints 'no cycle records'" || fail "no cycle records" +printf '%s\n' "$oute" | grep -qx ' no profiled cycles with a curve' && pass "comparison empty state" || fail "comparison empty state" + +# ---- 4. cycle records ------------------------------------------------------------------ +newrepo C; printf '%s\n' "$LEDGER_HEAD" > "$work/C/docs/hardening-log.md" +( + cd "$work/C" || exit 1 + commit 'ledger\n' + commit 'joined\n\ncycle aaaaaaaa; floor 3 per {docs/s.md (level 1)}; hook reminder threshold absent\ncycle aaaaaaaa; Gate B (passes 1-3,5, codex): Findings 4,3,?,1. Blockers 1,0,0,0. Majors 2,1,0,0.\n' + commit 'curve only\n\ncycle bbbbbbbb; Gate-A spec (passes 1, pass 1 codex+"my model"): Findings 2. Blockers 0. Majors 1.\n' + commit 'provenance only, quoted path\n\ncycle cccccccc; floor 1 per {"docs/q\\"x.md" (level 0)}; hook reminder threshold 3\ncycle closed on the zero-finding exit below the floor.\n' + commit 'skip with reason\n\ncycle dddddddd; Gate B: skipped (see skip reason)\nSkip reason: trivial,\nsecond line.\n\nafter the blank line\n' + commit 'skip without reason\n\ncycle eeeeeeee; Gate B: skipped (see skip reason)\n' + commit 'pre-rule\n\ncycle none (pre-rule); floor 3 per none; hook reminder threshold absent\ncycle none (pre-rule); Gate B (passes 1, codex): Findings 0. Blockers 0. Majors 0.\ncycle none (pre-rule); Gate B (passes 1, codex): Findings 0. Blockers 0. Majors 0.\n' + commit 'duplicate copy\n\ncycle bbbbbbbb; Gate-A spec (passes 1, pass 1 codex+"my model"): Findings 2. Blockers 0. Majors 1.\n' + commit 'conflicts\n\ncycle ffffffff; Gate B (passes 1, codex): Findings 1. Blockers 0. Majors 0.\ncycle ffffffff; Gate B (passes 1, codex): Findings 2. Blockers 0. Majors 0.\ncycle gggggggg; Gate B (passes 1, codex): Findings 1. Blockers 0. Majors 0.\ncycle gggggggg; Gate B: skipped (see skip reason)\nreason g\n\ncycle iiiiiiii; floor 3 per none; hook reminder threshold absent\ncycle iiiiiiii; floor 3 per {docs/s.md (level 0)}; hook reminder threshold absent\n' + commit 'skip reason A\n\ncycle hhhhhhhh; Gate B: skipped (see skip reason)\nreason A\n' + commit 'skip reason B\n\ncycle hhhhhhhh; Gate B: skipped (see skip reason)\nreason B\n' + commit 'malformed\n\ncycle abcdefg; floor 3 per none; hook reminder threshold absent\ncycle jjjjjjjj;floor 3 per none; hook reminder threshold absent\ncycle jjjjjjjj; floor 3 per {a.md (level 1),a.md (level 1)}; hook reminder threshold absent\ncycle jjjjjjjj; Gate B (passes 1-2, codex): Findings 1,2,3. Blockers 0,0. Majors 0,0.\ncycle jjjjjjjj; Gate B (passes 1-2, pass 1 codex): Findings 1,2. Blockers 0,0. Majors 0,0.\ncycle jjjjjjjj; Gate B (passes 1, "bad\\q"): Findings 1. Blockers 0. Majors 0.\ncycle jjjjjjjj; Gate B (passes 1, "tab\there"): Findings 1. Blockers 0. Majors 0.\ncycle jjjjjjjj; floor 3 per {a.md (level 1),"a.md" (level 1)}; hook reminder threshold absent\ncycle oooooooo; Gate B (passes 1, "a; b): c"): Findings 1. Blockers 0. Majors 0.\n' + commit 'skip then record\n\ncycle kkkkkkkk; Gate B: skipped (see skip reason)\ncycle llllllll; floor 1 per {docs/u.md (unprofiled)}; hook reminder threshold absent\n' + commit 'grammar 2\n\ncycle pppppppp; Gate-A plan (passes 1-2, pass 1 "model+variant"; pass 2 "a\\\\\\\\b"+codex): Findings 1,0. Blockers 0,0. Majors 1,0.\ncycle rrrrrrrr; floor 1 per {docs/a.md (level 0),docs/b.md (level 1)}; hook reminder threshold absent\ncycle ssssssss; Gate B (passes 1-100000000, codex): Findings 1. Blockers 0. Majors 0.\n' + gitc checkout -q -b side + commit 'side branch record\n\ncycle mmmmmmmm; Gate B (passes 1, codex): Findings 0. Blockers 0. Majors 0.\n' + gitc checkout -q main + tick=$((tick + 1)); gitc merge -q --no-ff -m 'merge side' side + commit 'licensed floor 1 recorded as floor 3\n\ncycle nnnnnnnn; floor 3 per {docs/z.md (level 0)}; hook reminder threshold absent\ncycle nnnnnnnn; Gate B (passes 1, codex): Findings 0. Blockers 0. Majors 0.\n' +) +git_section() { sed -n '/^== Git: review cycles/,/^== What this report cannot answer/p' | sed '1d;$d'; } +out=$(report C "$SCRIPT"); st=$? +[ "$st" -eq 0 ] && pass "cycle fixture: exit 0" || fail "cycle fixture: exit $st" +expect "cycles, conflicts, unparsed, comparison and checkpoint" "$(printf '%s\n' "$out" | git_section)" <<'EOF' +aaaaaaaa Gate B passes 1-3,5 floor 3 set {docs/s.md (level 1)} findings 4,3,?,1 (? 1) blockers 1,0,0,0 (? 0) majors 2,1,0,0 (? 0) +bbbbbbbb Gate-A spec passes 1 floor - set - findings 2 (? 0) blockers 0 (? 0) majors 1 (? 0) +cccccccc no curve floor 1 set {"docs/q\"x.md" (level 0)} +dddddddd Gate B skipped floor - set - reason excerpt: Skip reason: trivial, second line. commit ae01f24cfccb +eeeeeeee Gate B skipped floor - set - reason excerpt: no reason found commit 68042820934b +none (pre-rule) Gate B passes 1 floor - set - findings 0 (? 0) blockers 0 (? 0) majors 0 (? 0) commit 3333a18313c4 [may duplicate another pre-rule record] +none (pre-rule) Gate B passes 1 floor - set - findings 0 (? 0) blockers 0 (? 0) majors 0 (? 0) commit 3333a18313c4 [may duplicate another pre-rule record] +none (pre-rule) no curve floor 3 set none commit 3333a18313c4 [may duplicate another pre-rule record] +oooooooo Gate B passes 1 floor - set - findings 1 (? 0) blockers 0 (? 0) majors 0 (? 0) +kkkkkkkk Gate B skipped floor - set - reason excerpt: no reason found commit 13f09168f5c1 +llllllll no curve floor 1 set {docs/u.md (unprofiled)} +pppppppp Gate-A plan passes 1-2 floor - set - findings 1,0 (? 0) blockers 0,0 (? 0) majors 1,0 (? 0) +rrrrrrrr no curve floor 1 set {docs/a.md (level 0),docs/b.md (level 1)} +mmmmmmmm Gate B passes 1 floor - set - findings 0 (? 0) blockers 0 (? 0) majors 0 (? 0) +nnnnnnnn Gate B passes 1 floor 3 set {docs/z.md (level 0)} findings 0 (? 0) blockers 0 (? 0) majors 0 (? 0) +seen in several commits: cycle bbbbbbbb; Gate-A spec (passes 1, pass 1 codex+"my model"): Findings 2. Blockers 0. Majors 1. -> d61f00c143a8,34329462931c + +== Git: conflicts (records sharing a nonce that disagree; not listed as cycles) +ffffffff cycle ffffffff; Gate B (passes 1, codex): Findings 1. Blockers 0. Majors 0. +ffffffff cycle ffffffff; Gate B (passes 1, codex): Findings 2. Blockers 0. Majors 0. +gggggggg cycle gggggggg; Gate B (passes 1, codex): Findings 1. Blockers 0. Majors 0. +gggggggg cycle gggggggg; Gate B: skipped (see skip reason) | reason g +iiiiiiii cycle iiiiiiii; floor 3 per none; hook reminder threshold absent +iiiiiiii cycle iiiiiiii; floor 3 per {docs/s.md (level 0)}; hook reminder threshold absent +hhhhhhhh cycle hhhhhhhh; Gate B: skipped (see skip reason) | reason A +hhhhhhhh cycle hhhhhhhh; Gate B: skipped (see skip reason) | reason B + +== Git: unparsed candidate lines +7c5d9cd9db31 cycle abcdefg; floor 3 per none; hook reminder threshold absent +7c5d9cd9db31 cycle jjjjjjjj; Gate B (passes 1, "bad\q"): Findings 1. Blockers 0. Majors 0. +7c5d9cd9db31 cycle jjjjjjjj; Gate B (passes 1, "tab\x09here"): Findings 1. Blockers 0. Majors 0. +7c5d9cd9db31 cycle jjjjjjjj; Gate B (passes 1-2, codex): Findings 1,2,3. Blockers 0,0. Majors 0,0. +7c5d9cd9db31 cycle jjjjjjjj; Gate B (passes 1-2, pass 1 codex): Findings 1,2. Blockers 0,0. Majors 0,0. +7c5d9cd9db31 cycle jjjjjjjj; floor 3 per {a.md (level 1),"a.md" (level 1)}; hook reminder threshold absent +7c5d9cd9db31 cycle jjjjjjjj; floor 3 per {a.md (level 1),a.md (level 1)}; hook reminder threshold absent +7c5d9cd9db31 cycle jjjjjjjj;floor 3 per none; hook reminder threshold absent +4b1693862af7 cycle ssssssss; Gate B (passes 1-100000000, codex): Findings 1. Blockers 0. Majors 0. + +== Comparison with the fic2 baseline (story criterion 5) +profiled cycles (joined, with a curve, set has a (level N) entry): + aaaaaaaa Gate B passes 1-3,5 floor 3 set {docs/s.md (level 1)} findings 4,3,?,1 (? 1) blockers 1,0,0,0 (? 0) majors 2,1,0,0 (? 0) + nnnnnnnn Gate B passes 1 floor 3 set {docs/z.md (level 0)} findings 0 (? 0) blockers 0 (? 0) majors 0 (? 0) +fic2 baseline (docs/field-reports/2026-08-26-fic2-cycle-evidence.md, passes 1-5; story §1, passes 6-7): + findings 14,24,12,3,6,6,2 (? 0) blockers 3,4,0,0,0,0,0 (? 0) majors 5,13,6,2,5,?,? (? 2) +1. The curves are author-written and unchecked. Nothing compares them against the validated pass files, so they are self-reported and not measurement. +2. The cycles reviewed different artifacts, so a difference is evidence about the population as much as about the rule. +3. No demotion figure is derivable. That would need one finding classified under both rules, and nothing records that. + +== First checkpoint: evidence (the reader decides) +provenance lines whose set licenses floor 1: + cccccccc recorded floor 1 commit 4364bb042d4d + iiiiiiii recorded floor 3 commit 85b7cc11dd38 + nnnnnnnn recorded floor 3 commit 77b3167bb0fd +provenance lines recording floor 1: + cccccccc floor 1 set licenses it commit 4364bb042d4d + llllllll floor 1 set does not license it commit 13f09168f5c1 + rrrrrrrr floor 1 set does not license it commit 4b1693862af7 +EOF + +# A worktree file named like the commit SHA must not make `git log ` ambiguous, and a +# caller's log.showSignature must not leak signature text into the parsed stream. +sha=$(cd "$work/C" && git rev-parse HEAD); : > "$work/C/$sha" +out3=$(GIT_CONFIG_COUNT=1 GIT_CONFIG_KEY_0=log.showSignature GIT_CONFIG_VALUE_0=true report C "$SCRIPT"); st=$? +rm -f "$work/C/$sha" +[ "$st" -eq 0 ] && [ "$out3" = "$out" ] && + pass "SHA-named worktree file and log.showSignature leave the report unchanged" || + fail "SHA-named file / showSignature changed the report (exit $st)" + +# A 5001-digit pass number is valid under the §5 grammar on every Python (3.11+ caps int() +# at 4300 digits unless lifted), and a per-pass key that does not match its pass is unparsed. +newrepo T; printf '%s\n' "$LEDGER_HEAD" > "$work/T/docs/hardening-log.md" +long=$(printf '%05001d' 0 | tr 0 1) +(cd "$work/T" && commit "t\n\ncycle tttttttt; Gate B (passes $long, pass $long codex): Findings 1. Blockers 0. Majors 0.\ncycle uuuuuuuu; Gate B (passes 1, pass $long codex): Findings 1. Blockers 0. Majors 0.\n") +outt=$(report T "$SCRIPT"); st=$? +[ "$st" -eq 0 ] && printf '%s\n' "$outt" | grep -q '^tttttttt Gate B passes 1111' && + pass "5001-digit pass number: parsed as a cycle" || fail "5001-digit pass number (exit $st)" +printf '%s\n' "$outt" | grep -q '^[0-9a-f]\{12\} cycle uuuuuuuu; Gate B (passes 1, pass 1111' && + pass "per-pass key not matching its pass: unparsed, no crash" || fail "mismatched per-pass key" + +# Two conflict groups whose first records share a commit stay contiguous (group, then nonce). +newrepo O; printf '%s\n' "$LEDGER_HEAD" > "$work/O/docs/hardening-log.md" +( + cd "$work/O" || exit 1 + commit 'o1\n\ncycle aaaaaaaa; Gate B (passes 1, codex): Findings 1. Blockers 0. Majors 0.\ncycle bbbbbbbb; Gate B (passes 1, codex): Findings 1. Blockers 0. Majors 0.\n' + commit 'o2\n\ncycle bbbbbbbb; Gate B (passes 1, codex): Findings 2. Blockers 0. Majors 0.\n' + commit 'o3\n\ncycle aaaaaaaa; Gate B (passes 1, codex): Findings 2. Blockers 0. Majors 0.\n' +) +outo=$(report O "$SCRIPT"); st=$? +order=$(printf '%s\n' "$outo" | sed -n '/^== Git: conflicts/,/^$/p' | sed -n 's/^\([a-z]*\) cycle [a-z]*; Gate B (passes 1, codex): Findings \([0-9]\).*/\1\2/p' | tr '\n' ' ') +[ "$st" -eq 0 ] && [ "$order" = "aaaaaaaa1 aaaaaaaa2 bbbbbbbb1 bbbbbbbb2 " ] && + pass "tied conflict groups stay contiguous" || fail "conflict order: $order" + +# A pre-rule skip names its commit once. +newrepo Q; printf '%s\n' "$LEDGER_HEAD" > "$work/Q/docs/hardening-log.md" +(cd "$work/Q" && commit 'q\n\ncycle none (pre-rule); Gate B: skipped (see skip reason)\nwhy\n') +outq=$(report Q "$SCRIPT"); st=$? +n=$(printf '%s\n' "$outq" | grep '^none (pre-rule) Gate B skipped' | grep -o ' commit ' | wc -l | tr -d ' ') +[ "$st" -eq 0 ] && [ "$n" = 1 ] && pass "pre-rule skip names its commit once" || fail "pre-rule skip: exit $st, commit count $n" + +# An escaped final pipe does not close a row, and CRLF rows read like LF rows. +newrepo X +{ printf '%s\n' "$LEDGER_HEAD" + printf '%s\n' '| 2026-01-01 | esc | f | bot | nit | 1 prose | ref ends r\|' + printf '%s\r\n' '| 2026-01-02 | crlf | f | bot | nit | 2 lint | r |' +} > "$work/X/docs/hardening-log.md" +(cd "$work/X" && commit 'x\n') +outx=$(report X "$SCRIPT"); st=$? +[ "$st" -eq 0 ] && printf '%s\n' "$outx" | grep -qx '1 esc lines 7 rungs 1 prose' && + printf '%s\n' "$outx" | grep -qx '1 crlf lines 8 rungs 2 lint' && + printf '%s\n' "$outx" | sed -n '/^== Ledger: irregular rows/,+1p' | grep -qx 'none' && + pass "escaped final pipe and CRLF rows are regular" || fail "escaped pipe / CRLF rows" + +# The one-line error is UTF-8 whatever the locale says. +PYTHONIOENCODING=latin1 run "$work/L" "$SCRIPT" "$(printf 'caf\303\251')" > /dev/null 2> "$work/err"; st=$? +e=$(od -An -tx1 < "$work/err" | tr -d ' \n') +case $st:$e in 1:*c3a9*) pass "error message is UTF-8 under a latin1 locale (exit 1)" ;; *) fail "UTF-8 error: exit $st, bytes $e" ;; esac + +# Only conflicting records: the cycle list says why it is empty. +newrepo K; printf '%s\n' "$LEDGER_HEAD" > "$work/K/docs/hardening-log.md" +(cd "$work/K" && commit 'k\n\ncycle qqqqqqqq; Gate B (passes 1, codex): Findings 1. Blockers 0. Majors 0.\ncycle qqqqqqqq; Gate B (passes 1, codex): Findings 2. Blockers 0. Majors 0.\n') +outk=$(report K "$SCRIPT"); st=$? +[ "$st" -eq 0 ] && printf '%s\n' "$outk" | grep -qx 'no joined cycles (every record is in a conflict below)' && + pass "conflict-only history: the cycle list says why it is empty" || fail "conflict-only empty label" + +# ---- 5. exit-1 causes ----------------------------------------------------------------- +bad() { # name, expected stderr substring, dir, [ref] + e=$(run "$3" "$SCRIPT" ${4+"$4"} 2>&1 >/dev/null); s=$? + if [ "$s" -eq 1 ] && printf '%s' "$e" | grep -qF "$2"; then pass "$1"; else fail "$1 (exit $s: $e)"; fi +} +mkdir -p "$work/norepo" +bad "not a git repository" "ledger-metrics: not a git repository" "$work/norepo" +bad "unresolvable ref" "ledger-metrics: ref does not resolve to a commit: nope" "$work/L" nope +newrepo A; (cd "$work/A" && printf 'x\n' > x && commit 'x\n') +bad "ledger absent" "ledger-metrics: ledger absent at" "$work/A" +newrepo D; mkdir -p "$work/D/docs/hardening-log.md" && printf 'x\n' > "$work/D/docs/hardening-log.md/f"; (cd "$work/D" && commit 'd\n') +bad "ledger is a directory" "ledger-metrics: ledger is not a regular file" "$work/D" +newrepo S; printf '%s\n' "$LEDGER_HEAD" > "$work/S/real.md"; ln -s ../real.md "$work/S/docs/hardening-log.md"; (cd "$work/S" && commit 's\n') +bad "ledger is a symlink" "ledger-metrics: ledger is not a regular file" "$work/S" +newrepo G; printf '%s\n' "$LEDGER_HEAD" > "$work/G/docs/hardening-log.md" +(cd "$work/G" && commit 'one\n' && printf 'y\n' > y && commit 'two\n') +first=$(cd "$work/G" && git rev-parse HEAD~1) +rm -f "$work/G/.git/objects/$(printf '%s' "$first" | cut -c1-2)/$(printf '%s' "$first" | cut -c3-)" +bad "git log fails" "ledger-metrics: git log failed" "$work/G" +newrepo U; printf '%s\n' "$LEDGER_HEAD" > "$work/U/docs/hardening-log.md"; (cd "$work/U" && commit 'u\n') +blob=$(cd "$work/U" && git rev-parse HEAD:docs/hardening-log.md) +rm -f "$work/U/.git/objects/$(printf '%s' "$blob" | cut -c1-2)/$(printf '%s' "$blob" | cut -c3-)" +bad "ledger cannot be read" "ledger-metrics: ledger cannot be read" "$work/U" + +# ---- 6. shallow clone ------------------------------------------------------------------- +rm -rf "$work/shallow"; git clone -q --depth 1 "file://$work/C" "$work/shallow" 2>/dev/null +out=$(report shallow "$SCRIPT"); st=$? +[ "$st" -eq 0 ] && printf '%s\n' "$out" | grep -qx 'history: truncated (shallow clone): absence not established' && + pass "shallow clone: header says history is truncated" || fail "shallow clone header" + +set -- "$work"/WROTE.* +if [ -e "$1" ]; then fail "no-write: a run changed its fixture ($*)" +else pass "no-write: no run changed its fixture"; fi + +printf '\n%d passed, %d failed\n' "$pass_n" "$fail_n" +[ "$fail_n" -eq 0 ] +```` + +- [ ] **Step 2: Run it before the script exists** + +Run: `sh scripts/ledger-metrics.test.sh > /tmp/lm-red.log 2>&1; echo "exit=$?"; tail -1 /tmp/lm-red.log` +Expected: `exit=1`, and the last line reports failures. The prior-state case passes; the negative-control case fails (its mutation has nothing to apply to). + +- [ ] **Step 3: Write the script** + +Create `scripts/ledger-metrics.py` with exactly this content: + +````python +#!/usr/bin/env python3 +"""Read-only report over docs/hardening-log.md and the review-cycle records in git commit bodies. + +Spec: docs/superpowers/specs/2026-10-01-passive-metrics-design.md +Usage: python3 scripts/ledger-metrics.py [] (default HEAD) + +Reads the ledger and the commit bodies at ONE resolved commit, never the working tree, and +prints a report on stdout. It writes no file, ref or config itself; git can still write when +the caller's environment tells it to (spec §2), which the report header states. +Standard library only; Python 3.8 or later. +""" +import os +import re +import subprocess +import sys + +LEDGER = "docs/hardening-log.md" +ALLOWED_RUNGS = ("1 prose", "2 lint", "3 type", "4 test", "P std", "pending") +FIC2 = (" findings 14,24,12,3,6,6,2 (? 0) blockers 3,4,0,0,0,0,0 (? 0)" + " majors 5,13,6,2,5,?,? (? 2)") +CAVEATS = ( + "1. The curves are author-written and unchecked. Nothing compares them against the validated" + " pass files, so they are self-reported and not measurement.", + "2. The cycles reviewed different artifacts, so a difference is evidence about the population" + " as much as about the rule.", + "3. No demotion figure is derivable. That would need one finding classified under both rules," + " and nothing records that.", +) +CANNOT = """ +== What this report cannot answer +- Unlogged recurrences: the ledger only knows recurrences someone hardened. +- Whether a guard held: row succession is not guard failure, and a missing later row is not success. +- Cycles with no record at all: before 0.11.0 the curve was a habit, not a rule. A cycle with no provenance line, curve or skip record is invisible; a malformed record shows up as unparsed only when its line still starts like a record ("cycle ;"). +- Nonce attribution: records are grouped by nonce, which is collision-resistant, not collision-proof; two cycles that drew the same nonce read as one. +- Findings files: this script does not read .context/. It counts only what commit bodies say. +- Undecodable bytes: both sources are read as UTF-8 and an invalid byte becomes U+FFFD before anything is grouped, so two fingerprints or paths differing only in invalid bytes count as one. Control characters are shown as \\xNN, which can look like a literal backslash sequence. +- A shallow boundary that appears and disappears during the run: shallow state is checked before and after reading history, not during it. +- Whether a curve is true: see point 1 above. +- Cost and duration: that is vision step 2c's telemetry, out of scope here.""" + + +def die(msg): + sys.stderr.buffer.write(("ledger-metrics: %s\n" % msg).encode("utf-8", "backslashreplace")) + sys.stderr.flush() + sys.exit(1) + + +def git(*args): + """Run a read-only git command; return (returncode, stdout bytes). Stderr is not shown, + so the only error a caller sees is this script's own one-line message.""" + try: + # --no-replace-objects and an empty graft file: read the objects and history the SHA + # names, not `git replace` substitutes or legacy grafts. Both affect reading only. + env = dict(os.environ, GIT_GRAFT_FILE=os.devnull) + env.pop("GIT_SHALLOW_FILE", None) # the repository's own shallow file, not an override + p = subprocess.run(("git", "--no-replace-objects") + args, env=env, + stdout=subprocess.PIPE, stderr=subprocess.PIPE) + except OSError: + die("git could not be run") + return p.returncode, p.stdout + + +# ---- grammar (CLAUDE.md §5 Mechanics) ------------------------------------------------------ + +NONCE = r"[a-z0-9]{8,16}" +FIELD = r"(none \(pre-rule\)|" + NONCE + r")" +KIND = r"(Gate-A spec|Gate-A plan|Gate B)" +CANDIDATE = re.compile(r"^cycle (none \(pre-rule\)|[^ ;]*);") +BARE_PATH = re.compile(r"[A-Za-z0-9._/-]+") +BARE_MODEL = re.compile(r"[!#-'*\-./0-9<-~]+") # printable ASCII minus space " ( ) + , : ; +COUNT = re.compile(r"0|[1-9][0-9]*|\?") + + +def is_control(c): + """C0, DEL and the C1 range U+0080-U+009F.""" + o = ord(c) + return o < 0x20 or 0x7F <= o <= 0x9F + + +def read_quoted(s, i): + """s[i] == '"'. Return (decoded, next_index) or None. Escapes \\" and \\\\ only; any + control character makes the value unrepresentable, so the record is malformed.""" + out, i = [], i + 1 + while i < len(s): + c = s[i] + if is_control(c): + return None + if c == "\\": + if i + 1 < len(s) and s[i + 1] in "\"\\": + out.append(s[i + 1]) + i += 2 + continue + return None + if c == '"': + return ("".join(out), i + 1) if out else None + out.append(c) + i += 1 + return None + + +def parse_set(s, i): + """Parse at s[i:]. Return (entries, next_index) or None. entries is None for + `none`, else a list of (decoded_path, level) with level '0'..'2' or 'unprofiled'.""" + if s.startswith("none", i): + return None, i + 4 + if not s.startswith("{", i): + return None + i += 1 + entries, seen = [], set() + while True: + if i < len(s) and s[i] == '"': + q = read_quoted(s, i) + if q is None: + return None + path, i = q + else: + m = BARE_PATH.match(s, i) + if not m: + return None + path, i = m.group(0), m.end() + m = re.compile(r" \((level ([012])|unprofiled)\)").match(s, i) + if not m: + return None + level = m.group(2) if m.group(2) is not None else "unprofiled" + i = m.end() + if path in seen: # compared decoded, so "a.md" and a.md are the same path + return None + seen.add(path) + entries.append((path, level)) + if s.startswith(",", i): + i += 1 + continue + if s.startswith("}", i): + return entries, i + 1 + return None + + +def parse_model(s, i): + """One at s[i:]; return next index or None.""" + if i < len(s) and s[i] == '"': + q = read_quoted(s, i) + return None if q is None else q[1] + m = BARE_MODEL.match(s, i) + return m.end() if m else None + + +def expand(spec, limit): + """Expand "1-3,5" to [1, 2, 3, 5]. The total is checked against `limit` (the length of the + supplied count series) BEFORE anything is allocated, so a huge range costs nothing. Python's + integer-digit cap is lifted at start (see __main__); the ValueError guard is a last resort.""" + bounds, prev, total = [], 0, 0 + for part in spec.split(","): + m = re.fullmatch(r"([1-9][0-9]*)(?:-([1-9][0-9]*))?", part) + if not m: + return None + try: + a = int(m.group(1)) + b = int(m.group(2)) if m.group(2) else a + except ValueError: + return None + if a <= prev or b < a: + return None + total += b - a + 1 + if total > limit: + return None + bounds.append((a, b)) + prev = b + if total != limit: + return None + return [p for a, b in bounds for p in range(a, b + 1)] + + +def parse_curve(rest): + """rest is the text after '; '. Return dict or None.""" + m = re.compile(KIND + r" \(passes ([0-9,-]+), ").match(rest) + if not m: + return None + kind, spec, i = m.group(1), m.group(2), m.end() + tail = re.compile(r"\): Findings ([^ ]+)\. Blockers ([^ ]+)\. Majors ([^ ]+)\.$").search(rest) + if not tail: + return None + passes = expand(spec, len(tail.group(1).split(","))) + if passes is None: + return None + if rest.startswith("pass ", i): # per-pass models + for k, p in enumerate(passes): + if k: + if not rest.startswith("; ", i): + return None + i += 2 + m = re.compile(r"pass ([1-9][0-9]*) ").match(rest, i) + if not m or m.group(1) != str(p): # compared as text: no int() on untrusted digits + return None + i = m.end() + while True: + j = parse_model(rest, i) + if j is None: + return None + i = j + if rest.startswith("+", i): + i += 1 + continue + break + else: + j = parse_model(rest, i) + if j is None: + return None + i = j + m = re.compile(r"\): Findings ([^ ]+)\. Blockers ([^ ]+)\. Majors ([^ ]+)\.$").match(rest, i) + if not m: + return None + series = [] + for raw in m.groups(): + vals = raw.split(",") + if len(vals) != len(passes) or not all(COUNT.fullmatch(v) for v in vals): + return None + series.append(raw) + return {"kind": kind, "spec": spec, "f": series[0], "b": series[1], "m": series[2]} + + +def parse_record(line): + """Return (field, type, data) for a valid record, or None for a malformed candidate.""" + m = re.compile(r"^cycle " + FIELD + r"; ").match(line) + if not m: + return None + field, rest = m.group(1), line[m.end():] + m = re.compile(r"floor ([1-9][0-9]*) per ").match(rest) + if m: + parsed = parse_set(rest, m.end()) + if parsed is None: + return None + entries, i = parsed + if not re.compile(r"; hook reminder threshold (absent|unusable|[1-9][0-9]*)$").fullmatch(rest, i): + return None + return field, "prov", {"floor": m.group(1), "set": rest[m.end():i], "entries": entries} + m = re.fullmatch(KIND + r": skipped \(see skip reason\)", rest) + if m: + return field, "skip", {"kind": m.group(1)} + c = parse_curve(rest) + if c: + return field, "curve", c + return None + + +# ---- sections ------------------------------------------------------------------------------ + +def shown(text): + """Render a raw line for output: control characters become \\xNN, so a malformed record + cannot move the cursor or break the report's line structure.""" + return "".join("\\x%02x" % ord(c) if is_control(c) else c for c in text) + + +def ledger_section(text): + out = ["", "== Ledger: recurrence (rows matched exactly as harden-finding greps column 2)"] + rowre = re.compile(r"^\| *[0-9-]{10} *\| *([^|]*?) *\|") + rows, irregular = [], [] + for no, line in enumerate(text.split("\n"), 1): + line = line[:-1] if line.endswith("\r") else line # CRLF ledgers read like LF ones + m = rowre.match(line) + if not m: + continue + cells = re.split(r"(? 1 else m.group(1) + width_ok = len(cells) == 7 + rung = cells[5].strip(" ") if width_ok else None + rows.append((no, fp, rung)) + if not width_ok: + irregular.append("irregular width: line %d (%d cells)" % (no, len(cells))) + if not re.fullmatch(r"[a-z0-9]+(-[a-z0-9]+)*", fp): + irregular.append("irregular fingerprint: line %d" % no) + if not rows: + out.append("no rows") + return out + by_fp = {} + for no, fp, rung in rows: + by_fp.setdefault(fp, []).append((no, rung)) + order = sorted(by_fp, key=lambda f: (-len(by_fp[f]), f)) + for fp in order: + n = len(by_fp[fp]) + out.append("%d %s lines %s rungs %s" % ( + n, shown(fp), ",".join(str(no) for no, _ in by_fp[fp]), + " > ".join("?" if r is None else shown(r) for _, r in by_fp[fp]))) + out.append("check one count: git --no-replace-objects show :docs/hardening-log.md | grep -cE" + " '^\\| *[0-9-]{10} *\\| * *\\|'") + out.append(" ( is the commit in the header; use a fingerprint from the list. No command" + " fits an irregular fingerprint, because grep would read it as a pattern.)") + out += ["", "== Ledger: rung holding"] + last = {fp: max(no for no, _ in v) for fp, v in by_fp.items()} + landed, followed, pending = {}, {}, 0 + for no, fp, rung in rows: + if rung is None: + continue + if rung not in ALLOWED_RUNGS: + irregular.append("irregular rung: line %d (rung '%s')" % (no, shown(rung))) + continue + if rung == "pending": + pending += 1 + continue + landed[rung] = landed.get(rung, 0) + 1 + if no < last[fp]: + followed[rung] = followed.get(rung, 0) + 1 + for rung in sorted(landed): # only ALLOWED_RUNGS reach here, so no escaping is needed + f = followed.get(rung, 0) + out.append("%s landed %d followed-by-same-fingerprint %d last-of-fingerprint %d" + % (rung, landed[rung], f, landed[rung] - f)) + out.append("pending %d" % pending) + out.append('"followed" means a later row with the same fingerprint exists - nothing more. It does not mean') + out.append("the rung failed (a later row can be a different sub-shape, an out-of-scope guard, or a resolved") + out.append('prerequisite), and "last" does not mean it held (an unhardened recurrence leaves no row).') + out += ["", "== Ledger: irregular rows"] + irregular.sort(key=lambda s: (int(re.search(r"line (\d+)", s).group(1)), s)) + out += irregular or ["none"] + return out + + +def read_commits(raw): + """Parse `git log -z --format='%H %ct%n%B'` output into (sha, ct, body_lines).""" + commits = [] + for rec in raw.split("\0"): + if not rec.strip("\n"): + continue + head, _, body = rec.lstrip("\n").partition("\n") + m = re.fullmatch(r"([0-9a-f]{40}) ([0-9]+)", head) + if not m: + die("git log produced an unexpected record header") + commits.append((m.group(1), int(m.group(2)), body.split("\n"))) + return commits + + +def git_sections(commits, absence): + recs = {} # key -> record dict (key: line text; pre-rule keys also carry the sha) + unparsed = [] # (ct, sha, line) + for sha, ct, lines in commits: + seen_here = set() + i = 0 + while i < len(lines): + line = lines[i] + pos = i + i += 1 + if not CANDIDATE.match(line): + continue + parsed = parse_record(line) + if parsed is None: + unparsed.append((ct, sha, line)) + continue + field, typ, data = parsed + text = line + if typ == "skip": # the reason is the text right after the marker (CLAUDE.md §5) + reason = [] + while i < len(lines) and lines[i] != "" and not CANDIDATE.match(lines[i]): + reason.append(lines[i]) + i += 1 + data["reason"] = " ".join(reason) + text = line + " | " + data["reason"] # identity uses the excerpt as found + # pre-rule records identify nothing, so every occurrence stays separate (commit + position) + key = (text, sha, pos) if field == "none (pre-rule)" else (text, None, None) + if key in seen_here: + continue + seen_here.add(key) + r = recs.setdefault(key, {"field": field, "type": typ, "data": data, "text": text, + "commits": []}) + r["commits"].append((ct, sha)) + for r in recs.values(): + r["commits"].sort() + r["first"] = r["commits"][0] + + groups = {} + for key, r in recs.items(): + gkey = ("pre", key) if r["field"] == "none (pre-rule)" else ("n", r["field"]) + groups.setdefault(gkey, []).append(r) + + out = [] + cycles, conflicts, prof, lic, f1 = [], [], [], [], [] + for gkey, rs in groups.items(): + first = min(r["first"] for r in rs) + provs = [r for r in rs if r["type"] == "prov"] + rest = [r for r in rs if r["type"] != "prov"] + for p in provs: # checkpoint evidence comes from every valid provenance line, conflicts included + ents = p["data"]["entries"] + licensed = bool(ents) and all(lv == "0" for _, lv in ents) + sha = p["first"][1][:12] + if licensed: + lic.append((p["first"], "%s recorded floor %s commit %s" % (p["field"], p["data"]["floor"], sha))) + if p["data"]["floor"] == "1": + f1.append((p["first"], "%s floor 1 set %s commit %s" % ( + p["field"], "licenses it" if licensed else "does not license it", sha))) + if len(provs) > 1 or len(rest) > 1: + for r in rs: # groups sort by earliest record, then nonce; records inside by their own + conflicts.append((first, rs[0]["field"], r["first"], "%s %s" % (r["field"], shown(r["text"])))) + continue + p = provs[0] if provs else None + c = rest[0] if rest else None + fld = rs[0]["field"] + floor = p["data"]["floor"] if p else "-" + sset = p["data"]["set"] if p else "-" + if c is None: + line = "%s no curve floor %s set %s" % (fld, floor, shown(sset)) + elif c["type"] == "skip": + line = "%s %s skipped floor %s set %s reason excerpt: %s commit %s" % ( + fld, c["data"]["kind"], floor, shown(sset), + shown(c["data"]["reason"]) or "no reason found", c["first"][1][:12]) + else: + d = c["data"] + line = "%s %s passes %s floor %s set %s findings %s (? %d) blockers %s (? %d) majors %s (? %d)" % ( + fld, d["kind"], d["spec"], floor, shown(sset), d["f"], d["f"].split(",").count("?"), + d["b"], d["b"].split(",").count("?"), d["m"], d["m"].split(",").count("?")) + if gkey[0] == "pre": + if not (c and c["type"] == "skip"): # a skip line already names its commit + line += " commit %s" % rs[0]["first"][1][:12] + line += " [may duplicate another pre-rule record]" + cycles.append((first, line)) + if p and c and c["type"] == "curve" and p["data"]["entries"] and \ + any(lv != "unprofiled" for _, lv in p["data"]["entries"]): + prof.append((first, " " + line)) + + def emit(items, empty): + items.sort(key=lambda t: t[:-1] + (t[-1],)) + if not items: + out.append(empty) + out.extend(t[-1] for t in items) + + emit(cycles, ("no joined cycles (every record is in a conflict below)" if conflicts else + "no cycle records") + absence) + several = sorted((r["first"], shown(r["text"]), ",".join(s[:12] for _, s in r["commits"])) + for r in recs.values() if len(r["commits"]) > 1) + out.extend("seen in several commits: %s -> %s" % (t, c) for _, t, c in several) + out += ["", "== Git: conflicts (records sharing a nonce that disagree; not listed as cycles)"] + emit(conflicts, "no conflicts" + absence) + out += ["", "== Git: unparsed candidate lines"] + emit([((ct, sha), "%s %s" % (sha[:12], shown(line))) for ct, sha, line in unparsed], + "no unparsed lines" + absence) + out += ["", "== Comparison with the fic2 baseline (story criterion 5)", + "profiled cycles (joined, with a curve, set has a (level N) entry):"] + emit(prof, " no profiled cycles with a curve" + absence) + out.append("fic2 baseline (docs/field-reports/2026-08-26-fic2-cycle-evidence.md, passes 1-5; story §1, passes 6-7):") + out.append(FIC2) + out += list(CAVEATS) + out += ["", "== First checkpoint: evidence (the reader decides)", + "provenance lines whose set licenses floor 1:"] + emit([(k, " " + v) for k, v in lic], " none" + absence) + out.append("provenance lines recording floor 1:") + emit([(k, " " + v) for k, v in f1], " none" + absence) + return out + + +def is_shallow(): + rc, out = git("rev-parse", "--is-shallow-repository") + if rc != 0: + die("git rev-parse --is-shallow-repository failed") + return out.decode().strip() == "true" + + +def main(argv): + if len(argv) > 2: + die("usage: ledger-metrics.py []") + ref = argv[1] if len(argv) == 2 else "HEAD" + rc, _ = git("rev-parse", "--git-dir") + if rc != 0: + die("not a git repository") + rc, out = git("rev-parse", "--verify", "--quiet", ref + "^{commit}") + if rc != 0: + die("ref does not resolve to a commit: %s" % shown(ref)) + sha = out.decode().strip() + rc, out = git("ls-tree", "--full-tree", sha, "--", LEDGER) + if rc != 0: + die("git ls-tree failed at %s" % sha) + entry = out.decode("utf-8", "replace").strip() + if not entry: + die("ledger absent at %s: %s" % (sha, LEDGER)) + if not re.match(r"^100(644|755) blob ", entry): + die("ledger is not a regular file at %s: %s" % (sha, LEDGER)) + rc, out = git("cat-file", "blob", "%s:%s" % (sha, LEDGER)) + if rc != 0: + die("ledger cannot be read at %s: %s" % (sha, LEDGER)) + ledger = out.decode("utf-8", "replace") + shallow = is_shallow() + # Bodies re-encoded to UTF-8 whatever i18n.logOutputEncoding says, so NUL framing holds; + # no signature output, which would land in the same stream; `--` so a file named like the + # SHA cannot make the argument ambiguous. + rc, out = git("-c", "i18n.logOutputEncoding=UTF-8", "log", "--encoding=UTF-8", "--no-show-signature", + "-z", "--format=%H %ct%n%B", sha, "--") + if rc != 0: + die("git log failed at %s" % sha) + commits = read_commits(out.decode("utf-8", "replace")) + shallow = is_shallow() or shallow # checked before and after the walk: either means truncated + absence = " (history truncated: absence not established)" if shallow else "" + + lines = [ + "ledger-metrics — read-only report", + "commit %s (from %s)" % (sha, shown(ref)), + "ledger: %s at that commit; the working tree is not read" % LEDGER, + "git writes: this script only runs read-only git commands; git may still write when your" + " environment tells it to (e.g. GIT_TRACE to a file, partial-clone fetches)", + "history: " + ("truncated (shallow clone): absence not established" if shallow else "complete"), + ] + lines += ledger_section(ledger) + lines += ["", "== Git: review cycles (all commits reachable from %s)" % sha] + lines += git_sections(commits, absence) + # UTF-8 whatever the locale says, so a valid non-ASCII path can always be printed. + out = "\n".join(lines) + "\n" + CANNOT + "\n" + sys.stdout.buffer.write(out.encode("utf-8", "backslashreplace")) + sys.stdout.flush() + return 0 + + +if __name__ == "__main__": + # One integer limit on every Python: 3.11+ caps int() at 4300 digits by default, 3.8-3.10 + # do not, and the §5 grammar has no digit limit. Lift the cap where it exists. + if hasattr(sys, "set_int_max_str_digits"): + sys.set_int_max_str_digits(0) + sys.exit(main(sys.argv)) +```` + +- [ ] **Step 4: Run the suite under both shells, reading each exit status** + +Run: `sh scripts/ledger-metrics.test.sh > /tmp/lm-sh.log 2>&1; echo "sh=$?"; dash scripts/ledger-metrics.test.sh > /tmp/lm-dash.log 2>&1; echo "dash=$?"; tail -1 /tmp/lm-sh.log /tmp/lm-dash.log` +Expected: `sh=0`, `dash=0`, and `30 passed, 0 failed` in both logs (measured on the prototype 2026-10-01). The second run checks the suite's own portability; the script runs under `python3` either way. + +- [ ] **Step 5: Lint the suite** + +Run: `shellcheck --shell=sh --exclude=SC2015 scripts/ledger-metrics.test.sh; echo "lint=$?"` +Expected: `lint=0`. + +- [ ] **Step 6: Run it on the real repository** + +Run: `python3 scripts/ledger-metrics.py main > /tmp/lm-main.txt; echo "exit=$?"; sed -n '/^== Git: unparsed/,+1p' /tmp/lm-main.txt` +Expected: `exit=0` and `no unparsed lines`. + +No commit (Task 3 makes the single Gate-B snapshot). + +--- + +### Task 2: Battery, CI and docs + +**Files:** +- Modify: `AGENTS.md`, `CLAUDE.md`, `.github/workflows/ci.yml`, `README.md` + +- [ ] **Step 1: Find every place that enumerates the executables or suites, and every manifest claim** + +```sh +grep -rnE 'two repo-local|both checker|all three executables|three suites|two checkers|both checkers' --include='*.md' --include='*.yml' --include='*.sh' . | grep -vE 'source-files/|docs/superpowers/|\.context/|hardening-log|CHANGELOG' +grep -rniE 'declare[sd]?|convention[- ]load' --include='*.md' . | grep -v source-files/ +cat plugins/dev-workflow/.claude-plugin/plugin.json +grep -rn 'consumer does not exist' --include='*.md' . | grep -v source-files/ +``` +Expected (measured 2026-10-01): the first grep finds `README.md:160`, `README.md:163`, `.github/workflows/ci.yml:40`, `:46`, `:60`, `:92`. Lines 46 and 60 say "the two repo-local checkers" about **checkers** and stay true, because the report is not a checker. The second grep's hits are claims about the plugin manifest; read them against the `cat` output — this change adds nothing to the plugin, so each should stay true. The third grep finds `CLAUDE.md` (§5 Mechanics) and `plugins/dev-workflow/commands/workflow-init.md`: the repository's own `CLAUDE.md` sentence becomes false with this change and is edited below; the scaffolded template's copy stays, because a consumer project gets no metrics script and its consumer still does not exist. Any other first-grep hit is a stop: surface it before editing. + +- [ ] **Step 2: Apply the edits** + +Save this as a temporary file outside the repository (for example `$TMPDIR/docs-edit.py`) and run it from the repository root with `python3`. It writes nothing unless every replacement matches exactly once. + +```python +# Applies the Task 2 documentation and CI edits. Every replacement must match exactly +# once, or the script stops before writing anything. +import sys +edits = { + "AGENTS.md": [ + ("scripts/check-version-bump.test.sh # its regression suite — policy/operational/accept\n", + "scripts/check-version-bump.test.sh # its regression suite — policy/operational/accept\n" + "scripts/ledger-metrics.py # read-only report over the ledger + git (vision step 2b)\n" + "scripts/ledger-metrics.test.sh # its regression suite — fixture repos, expected text\n"), + ("repo-local CI checkers and their tests (`scripts/check-invariants.{sh,test.sh}` and\n" + "`scripts/check-version-bump.{sh,test.sh}`) — the hook ships in the plugin, the checkers\n" + "do not; everything else is text read by a model.", + "repo-local CI checkers and their tests (`scripts/check-invariants.{sh,test.sh}` and\n" + "`scripts/check-version-bump.{sh,test.sh}`), plus the read-only Python report\n" + "`scripts/ledger-metrics.py` and its suite `scripts/ledger-metrics.test.sh` — the hook ships\n" + "in the plugin, the checkers and the report do not; everything else is text read by a model."), + # quality row and lint row each contain one of these two substrings + ("shellcheck --shell=sh scripts/check-version-bump.test.sh && HOOK_SH=sh", + "shellcheck --shell=sh scripts/check-version-bump.test.sh && shellcheck --shell=sh --exclude=SC2015 scripts/ledger-metrics.test.sh && HOOK_SH=sh"), + ("shellcheck --shell=sh scripts/check-version-bump.test.sh` |", + "shellcheck --shell=sh scripts/check-version-bump.test.sh && shellcheck --shell=sh --exclude=SC2015 scripts/ledger-metrics.test.sh` |"), + ("sh scripts/check-version-bump.sh main && claude plugin validate . --strict` |", + "sh scripts/check-version-bump.sh main && sh scripts/ledger-metrics.test.sh && claude plugin validate . --strict` |"), + ("| typecheck | n/a — no typed sources (shell + markdown) |", + "| typecheck | n/a — no typed sources (shell, one untyped Python report, markdown) |"), + ("tool. Without it the second run cannot start, and dropping that run is what let a\n" + "`dash`-only defect ship once already.\n", + "tool. Without it the second run cannot start, and dropping that run is what let a\n" + "`dash`-only defect ship once already. The metrics report's suite also needs\n" + "**`python3`** 3.8 or later, addressed by name for the same reason as `dash`: it is the\n" + "system interpreter, not a pinned tool.\n"), + ], + "CLAUDE.md": [ + ("variant; anything quoting this form elsewhere quotes an instance of it, because the deferred\n" + " metrics work is intended to parse it — that consumer does not exist yet, and the form is pinned\n" + " now so that it can.\n", + "variant; anything quoting this form elsewhere quotes an instance of it, because tooling parses\n" + " it — in this repository `scripts/ledger-metrics.py` (dark-factory vision step 2b) — and the\n" + " form is pinned so that it can.\n"), + ], + ".github/workflows/ci.yml": [ + (" # The hook, the two checkers and their three suites are the only executables here,\n", + " # The hook, the two checkers, the Python metrics report and their four suites are\n" + " # the only executables here,\n"), + (" --shell=sh scripts/check-version-bump.test.sh\n", + " --shell=sh scripts/check-version-bump.test.sh\n" + " # The metrics report's suite (vision step 2b). The report itself is Python, so\n" + " # only its POSIX-sh suite is linted here; the suite runs the report under the\n" + " # runner's python3. Same `[ c ] && pass || fail` lines as the other suites,\n" + " # hence SC2015.\n" + " docker run --rm -v \"$PWD:/mnt\" -w /mnt koalaman/shellcheck:v0.11.0 \\\n" + " --shell=sh --exclude=SC2015 scripts/ledger-metrics.test.sh\n"), + (" - name: Invariant checks (pinning, manifest, prompt conformance) + both checker suites\n", + " - name: Invariant checks (pinning, manifest, prompt conformance) + checker and report suites\n"), + (" sh scripts/check-version-bump.test.sh\n\n", + " sh scripts/check-version-bump.test.sh\n sh scripts/ledger-metrics.test.sh\n\n"), + ], + "README.md": [ + ("PR and push to main: `shellcheck --shell=sh` over all three executables and their test\nfiles,", + "PR and push to main: `shellcheck --shell=sh` over the three shell executables and every\nshell test file,"), + ("both checkers' regression suites, and `claude plugin validate . --strict`.\n", + "both checkers' regression suites, the metrics report's suite, and\n`claude plugin validate . --strict`.\n\n" + "`python3 scripts/ledger-metrics.py` prints a read-only report over the hardening ledger\n" + "and the review-cycle records in commit bodies: which fingerprints recur, how rungs were\n" + "followed, and each cycle's recorded curve beside the `fic2` baseline. It writes nothing.\n"), + ], +} +new = {} +for path, reps in edits.items(): + s = open(path).read() + for old, rep in reps: + n = s.count(old) + if n != 1: + sys.exit(f"STOP: {path}: expected exactly 1 match, found {n}: {old[:70]!r}") + s = s.replace(old, rep) + new[path] = s +for path, s in new.items(): + open(path, "w").write(s) +print("edited:", ", ".join(new)) +``` + +Expected: `edited: AGENTS.md, CLAUDE.md, .github/workflows/ci.yml, README.md`; then `git diff --stat` lists those four files only. + +- [ ] **Step 3: Run the AGENTS.md quality row, verbatim** + +Copy the quality row from `AGENTS.md` § Commands as it now stands and run it into a log: `sh -c '' > /tmp/lm-quality.log 2>&1; echo "quality=$?"`. Expected: `quality=0`, and the log contains `30 passed, 0 failed` (measured on a copy of the repository 2026-10-01). + +No commit. + +--- + +### Task 3: Evidence, Gate B, close + +- [ ] **Step 1: Base check before the snapshot** + +Run each separately and read its status: `git fetch origin`; `git merge-base --is-ancestor origin/main HEAD; echo "anc=$?"`. +Expected: `anc=0`. `anc=1` is a stop: ask Daniel. No plugin path changes, so there is no version to reconcile. + +- [ ] **Step 2: Stage and snapshot** + +```sh +git add scripts/ledger-metrics.py scripts/ledger-metrics.test.sh AGENTS.md CLAUDE.md .github/workflows/ci.yml README.md && git diff --cached --name-only +``` +Expected: exactly those six paths. Then, **as its own one-line tool call from inside the repository** (the recommended form CLAUDE.md §5 Mechanics describes): + +```sh +git commit -m 'WIP: passive metrics candidate' +``` + +- [ ] **Step 3: Evidence run** + +Run each and read its result: `git rev-parse HEAD` (record it); `git status --porcelain --untracked-files=no` (must print nothing); `test "$(git rev-parse main)" = "$(git rev-parse origin/main)"; echo "main-fresh=$?"` (non-zero is a stop — ask Daniel to update local `main`); the AGENTS.md quality row verbatim into a log, exit 0; `dash scripts/ledger-metrics.test.sh > /tmp/lm-dash.log 2>&1; echo "dash=$?"` (must be `dash=0` with `30 passed, 0 failed`); `git ls-tree c60b78b -- scripts/ledger-metrics.py scripts/ledger-metrics.sh` (must print nothing: the prior state, observed at a named commit); then the first two again (same `HEAD`, still nothing). + +Evidence entry for the commit body: + +``` +Evidence — docs/superpowers/stories/2026-08-04-passive-metrics-over-the-ledger-story.md +Battery: AGENTS.md quality row, exit 0 at . +Check (counterfactual): at c60b78b no metrics script exists (git ls-tree prints +nothing), and the suite's first case observes that an absent script produces no +report. Negative control: a copy whose fingerprint count is off by one runs to +completion, and the same comparison every golden case uses rejects its report +(suite case 2). Suite 30/30 under sh (in the quality row) and under dash. +``` + +- [ ] **Step 4: Gate B** + +CLAUDE.md §5: new nonce; floor 3 (story `standard`/`none`). First check `mcp__codex__health` → `config.effective.model` is `gpt-6-astra`; anything else is a stop (the ChatGPT Codex app rewrites `~/.codex/config.toml`). `baseSha` = parent of the WIP commit. **Immediately before each of the two calls**, resolve `headSha` with `git rev-parse HEAD` and pass that full value; keep each branch's `baseSha`/`headSha` with its result and require them equal before summing. Two separate `mcp__codex__review` calls (`spec`, then `quality`), each with its own findings file. Each call carries: the story path; the evidence entry verbatim; the seven rulings above; the standing lens "which existing statements does this diff falsify?", naming what this diff changes — the executable and suite counts, the CI step name, the AGENTS.md command rows, typecheck row and prerequisites, and the `CLAUDE.md` consumer sentence. + +Fixes: `git add` the paths, then `git commit --amend -m 'WIP: passive metrics candidate'` as its own one-line tool call, then an evidence run, then re-review. + +- [ ] **Step 5: Close and PR** + +When the §5 closure ordering allows it: evidence run again; `git log --format='%h %s' origin/main..HEAD` must show the WIP commit on top of the plan commit, `49b90f8` and `c60b78b`, and nothing staged — any other shape is a stop. Then `git commit --amend -m ""`, whose body carries the evidence entry, the Gate-B provenance line, the Gate-B curve, and a prose line saying the spec and quality calls of each pass together counted as one logical pass. No trailers. Push, open a PR, then read CI's `quality` log for `30 passed, 0 failed`. + +--- + +## Where Gate-A plan pass 1's findings landed + +| Pass-1 finding (awk plan) | Now | +|---|---| +| mawk array-membership semantics; first-match model/curve parsing; control characters in quoted values; decoded duplicate paths; fields rebuilt from decorated keys; level/unprofiled substring search; unbounded range expansion; record-separator spoofing; unchecked awk/sort status | **Removed by the language change**: Python dictionaries, a quote-aware scanner, `read_quoted` rejecting control characters, decoded-path comparison, parsed fields kept as data, entry levels parsed per entry, a length check before expansion, `-z` NUL framing with a validated header, and `git()` checking every return code. | +| `rowre` without the column-2 cell | The row pattern includes the cell and its closing pipe. | +| Conflict groups skipped before the checkpoint lists | Every valid provenance line feeds the checkpoint lists, conflicts included (fixture `iiiiiiii`). | +| `tail` hiding the suite's status | Every command in this plan reads its own exit status. | +| `HOOK_SH=dash` not running the product under dash | The product is Python now; the suite runs under `sh` and `dash` for its own portability. | +| No-write snapshot around one run only | `report()` snapshots around every run; one marker per changed fixture, checked at the end. | +| Inherited git config, identity, templates | `GIT_CONFIG_GLOBAL=/dev/null`, `GIT_CONFIG_NOSYSTEM=1`, inherited variables unset, `--template=`, `--object-format=sha1`. | +| `headSha` reused across the two Gate-B calls | Resolved immediately before each call and kept with its result. | +| Plan pass 2 (Python plan): an `int()` on per-pass keys, error runs bypassing the no-write check, report statuses not asserted, inherited `GIT_AUTHOR_*`/`GIT_COMMITTER_*`, dash not in the evidence run | Keys compared as text (5001-digit fixture); `run()` wraps every invocation, errors included; each report's status asserted; identity variables fixed in `gitc`; dash run in Task 3 Step 3. Minors: fingerprint from the split cell, group-level conflict ordering, commit shown on pre-rule and skip lines, excerpt identity without the fallback text, control-safe ref, C1 controls, printed check command executed as printed, footer asserted, wider manifest grep, `CLAUDE.md` consumer sentence, typecheck row. | +| Plan pass 3: inherited `GIT_SHALLOW_FILE` broke four suite cases | Unset in the suite (with `GIT_GRAFT_FILE`, `GIT_NO_REPLACE_OBJECTS`) and dropped by the script; suite passes with it set. Minors: group-then-nonce conflict ordering, escaped final pipe, CRLF rows, UTF-8 stderr, one integer limit on every Python, pre-rule skip commit printed once. Left as stated limits: U+FFFD before grouping (ruling 7), a transient shallow boundary (ruling 3). | +| Plan pass 4: the pass-3 repairs had no regression cases | Six cases added, one per repair — tied conflict order, escaped final pipe and CRLF, UTF-8 stderr under a latin1 locale, a 5001-digit pass number parsed as a cycle, a pre-rule skip naming its commit once. Each was checked by reverting its repair in a copy: exactly that case failed. The two kept limits are now printed in the report. | +| Plan pass 5: the UTF-8 error case bypassed `run()`; two new cases ignored the script's exit status; global git attributes from `XDG_CONFIG_HOME` could change fixture commits | The UTF-8 case runs through `run()` and requires exit 1; the pre-rule-skip case asserts exit 0; the suite sets an empty `HOME` and `XDG_CONFIG_HOME` and `GIT_ATTR_NOSYSTEM=1` (suite green with a hostile `XDG_CONFIG_HOME/git/attributes`). Minor: several-commit notices sort by the printed text. | +| The Minors (escaped-pipe restore, pre-rule positions, ordering ties, `` placeholder, conflict/unparsed empty states and suffixes, nonce caveat, git stderr before the one-line error, `ls-tree` relative path, "cannot be read" case, Task 2 manifest-claim search, closing-message prose) | Each handled in the script, the suite or the steps above. | + +## Self-review (2026-10-01) + +- **Spec coverage:** §1 → Tasks 1–2. §2 → `main()`/`git()` and suite exit cases. §3 → `ledger_section` + ledger fixture. §4 → `parse_record`, `git_sections` + cycle fixture. §5 → comparison and checkpoint output + fixture. §6 → the suite (isolation, prior state, negative control, no-write around every run). §7 → `CANNOT`. §8 → nothing beyond. +- **Story ACs:** 1 done (profile, `c60b78b`); 2 → no-write check + script reading; 3 → printed command + suite case "check command reproduces"; 4 → printed "cannot answer"; 5 → comparison section, `(? n)` per series, caveats, checkpoint evidence. +- **Placeholders:** `` and `` are run-time values. From 0d3da0a484861796bdb3bc685497fb3bb52ccc1c Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Thu, 1 Oct 2026 21:03:07 +0200 Subject: [PATCH 4/5] Add a read-only metrics report over the ledger and git (vision step 2b) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit scripts/ledger-metrics.py (Python 3.8+, standard library) reads docs/hardening-log.md and the commit bodies at one resolved commit and prints which fingerprints recur, how rungs were followed, and every review-cycle record (provenance line, curve, skip record) beside the fic2 baseline. It computes no shares or verdicts, writes nothing, and prints what it cannot answer. scripts/ledger-metrics.test.sh is its POSIX-sh suite: 31 cases on config-isolated fixture repositories with fixed dates, a no-write check around every run, a prior-state case and a negative control. The suite joins the quality row, the lint row and CI. AGENTS.md, README.md and the CI comments now count it; CLAUDE.md no longer says the metrics consumer does not exist. The scaffolded template keeps that sentence, because consumer projects get no script. Not shipped, so no plugin bump. Executed natively from the plan. One deviation, ruled: Gate-B pass 1's quality Minor (the empty checkpoint lists were not asserted) was fixed by one added suite case (mutation-checked), so the suite has 31 cases where the plan has 30. Evidence — docs/superpowers/stories/2026-08-04-passive-metrics-over-the-ledger-story.md Battery: AGENTS.md quality row, exit 0 at d33b2424eee6e69a664a4343d24f0c7278c5337d. Check (counterfactual): at c60b78b no metrics script exists (git ls-tree prints nothing), and the suite's first case observes that an absent script produces no report. Negative control: a copy whose fingerprint count is off by one runs to completion, and the same comparison every golden case uses rejects its report (suite case 2). Suite 31/31 under sh (in the quality row) and under dash. cycle 08x02c2od1; floor 3 per {docs/superpowers/stories/2026-08-04-passive-metrics-over-the-ledger-story.md (level 1)}; hook reminder threshold absent cycle 08x02c2od1; Gate B (passes 1-2, gpt-6-astra): Findings 1,0. Blockers 0,0. Majors 0,0. Each Gate-B pass was one logical pass run as two calls (reviewType spec, then quality) against the same baseSha 9230e4fc5f063e4a523d6fd88347b7b5dbbd88d1 and the headSha resolved before each call (pass 1: 0a1304454a285a4e00ea4738e612aba0bcca1a35; pass 2: d33b2424eee6e69a664a4343d24f0c7278c5337d). Pass 2 found nothing in either branch, so the cycle closed on the zero-finding exit below the floor. Human exceptions: none --- .github/workflows/ci.yml | 12 +- AGENTS.md | 17 +- CLAUDE.md | 6 +- README.md | 11 +- scripts/ledger-metrics.py | 513 +++++++++++++++++++++++++++++++++ scripts/ledger-metrics.test.sh | 387 +++++++++++++++++++++++++ 6 files changed, 932 insertions(+), 14 deletions(-) create mode 100644 scripts/ledger-metrics.py create mode 100644 scripts/ledger-metrics.test.sh diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index b0499d4..fae1b98 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -37,7 +37,8 @@ jobs: # merge-base. ~50 commits, so the full history costs nothing here. fetch-depth: 0 - # The hook, the two checkers and their three suites are the only executables here, + # The hook, the two checkers, the Python metrics report and their four suites are + # the only executables here, # so lint plus those suites are the only mechanical gates we have. Everything else # here is a prompt, and prompts have no typechecker (see README, "Contributing"). @@ -69,6 +70,12 @@ jobs: --shell=sh scripts/check-version-bump.sh docker run --rm -v "$PWD:/mnt" -w /mnt koalaman/shellcheck:v0.11.0 \ --shell=sh scripts/check-version-bump.test.sh + # The metrics report's suite (vision step 2b). The report itself is Python, so + # only its POSIX-sh suite is linted here; the suite runs the report under the + # runner's python3. Same `[ c ] && pass || fail` lines as the other suites, + # hence SC2015. + docker run --rm -v "$PWD:/mnt" -w /mnt koalaman/shellcheck:v0.11.0 \ + --shell=sh --exclude=SC2015 scripts/ledger-metrics.test.sh # Two runs, because the hook has to be correct under both shells and the runner's # /bin/sh is dash while a contributor's may be bash. HOOK_SH selects the shell the @@ -89,11 +96,12 @@ jobs: # Each rule was prose first and each was violated anyway — a floating action ref # shipped in this very workflow, a duplicate hooks manifest key stopped the plugin # loading in 0.2.1, and the plugin shipped changes without a bump twice. - - name: Invariant checks (pinning, manifest, prompt conformance) + both checker suites + - name: Invariant checks (pinning, manifest, prompt conformance) + checker and report suites run: | sh scripts/check-invariants.test.sh sh scripts/check-invariants.sh sh scripts/check-version-bump.test.sh + sh scripts/ledger-metrics.test.sh # Invariant 12 itself needs a base to diff against, so it runs only on pull # requests — a push to main has no PR base, and diffing the push range would just diff --git a/AGENTS.md b/AGENTS.md index 032fb2f..fa54083 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -38,6 +38,8 @@ scripts/check-invariants.sh # invariants 5, 6 + prompt conformance (rung 2 scripts/check-invariants.test.sh # its regression suite — reject/accept pairs scripts/check-version-bump.sh # invariant 12, mechanically — PR-only (rung 2) scripts/check-version-bump.test.sh # its regression suite — policy/operational/accept +scripts/ledger-metrics.py # read-only report over the ledger + git (vision step 2b) +scripts/ledger-metrics.test.sh # its regression suite — fixture repos, expected text plugins/dev-workflow/ .claude-plugin/plugin.json # metadata only — no component keys (invariant 6) CHANGELOG.md # every manifest version, newest first @@ -62,8 +64,9 @@ source-files/ # the extraction seed this repo was built from **Boundaries.** `skills/`, `commands/`, `agents/` and `hooks/hooks.json` are loaded by convention from their paths. The executable artifacts are the hook and its test, plus the two repo-local CI checkers and their tests (`scripts/check-invariants.{sh,test.sh}` and -`scripts/check-version-bump.{sh,test.sh}`) — the hook ships in the plugin, the checkers -do not; everything else is text read by a model. `examples/` is reference material, +`scripts/check-version-bump.{sh,test.sh}`), plus the read-only Python report +`scripts/ledger-metrics.py` and its suite `scripts/ledger-metrics.test.sh` — the hook ships +in the plugin, the checkers and the report do not; everything else is text read by a model. `examples/` is reference material, outside the loaded surface: never scaffolded or copied into a user's project, though it does ship inside the plugin package. @@ -253,9 +256,9 @@ Every command below was run in this session and observed to exit 0. | Role | Command | |---|---| -| quality (the whole battery — what CI runs) | `shellcheck --shell=sh plugins/dev-workflow/hooks/codex-gate.sh && shellcheck --shell=sh --exclude=SC2015 plugins/dev-workflow/hooks/codex-gate.test.sh && shellcheck --shell=sh scripts/check-invariants.sh && shellcheck --shell=sh --exclude=SC2015 scripts/check-invariants.test.sh && shellcheck --shell=sh scripts/check-version-bump.sh && shellcheck --shell=sh scripts/check-version-bump.test.sh && HOOK_SH=sh sh plugins/dev-workflow/hooks/codex-gate.test.sh && HOOK_SH=dash dash plugins/dev-workflow/hooks/codex-gate.test.sh && sh scripts/check-invariants.test.sh && sh scripts/check-invariants.sh && sh scripts/check-version-bump.test.sh && sh scripts/check-version-bump.sh main && claude plugin validate . --strict` | -| typecheck | n/a — no typed sources (shell + markdown) | -| lint | `shellcheck --shell=sh plugins/dev-workflow/hooks/codex-gate.sh && shellcheck --shell=sh --exclude=SC2015 plugins/dev-workflow/hooks/codex-gate.test.sh && shellcheck --shell=sh scripts/check-invariants.sh && shellcheck --shell=sh --exclude=SC2015 scripts/check-invariants.test.sh && shellcheck --shell=sh scripts/check-version-bump.sh && shellcheck --shell=sh scripts/check-version-bump.test.sh` | +| quality (the whole battery — what CI runs) | `shellcheck --shell=sh plugins/dev-workflow/hooks/codex-gate.sh && shellcheck --shell=sh --exclude=SC2015 plugins/dev-workflow/hooks/codex-gate.test.sh && shellcheck --shell=sh scripts/check-invariants.sh && shellcheck --shell=sh --exclude=SC2015 scripts/check-invariants.test.sh && shellcheck --shell=sh scripts/check-version-bump.sh && shellcheck --shell=sh scripts/check-version-bump.test.sh && shellcheck --shell=sh --exclude=SC2015 scripts/ledger-metrics.test.sh && HOOK_SH=sh sh plugins/dev-workflow/hooks/codex-gate.test.sh && HOOK_SH=dash dash plugins/dev-workflow/hooks/codex-gate.test.sh && sh scripts/check-invariants.test.sh && sh scripts/check-invariants.sh && sh scripts/check-version-bump.test.sh && sh scripts/check-version-bump.sh main && sh scripts/ledger-metrics.test.sh && claude plugin validate . --strict` | +| typecheck | n/a — no typed sources (shell, one untyped Python report, markdown) | +| lint | `shellcheck --shell=sh plugins/dev-workflow/hooks/codex-gate.sh && shellcheck --shell=sh --exclude=SC2015 plugins/dev-workflow/hooks/codex-gate.test.sh && shellcheck --shell=sh scripts/check-invariants.sh && shellcheck --shell=sh --exclude=SC2015 scripts/check-invariants.test.sh && shellcheck --shell=sh scripts/check-version-bump.sh && shellcheck --shell=sh scripts/check-version-bump.test.sh && shellcheck --shell=sh --exclude=SC2015 scripts/ledger-metrics.test.sh` | | test | `HOOK_SH=sh sh plugins/dev-workflow/hooks/codex-gate.test.sh && HOOK_SH=dash dash plugins/dev-workflow/hooks/codex-gate.test.sh` — two runs; `HOOK_SH` selects the shell the HOOK runs under, and without it a dash invocation only exercises the harness | | invariant checks (5 pinning, 6 manifest, prompt conformance) | `sh scripts/check-invariants.test.sh && sh scripts/check-invariants.sh` | | invariant check (12 version bump) | `sh scripts/check-version-bump.test.sh && sh scripts/check-version-bump.sh main` | @@ -268,7 +271,9 @@ twice, once with the hook under `sh` and once under `dash`, because the hook has correct under both and Ubuntu's `/bin/sh` IS dash. Bump the first two deliberately, per invariant 5; `dash` is addressed by name because it is the system shell, not a pinned tool. Without it the second run cannot start, and dropping that run is what let a -`dash`-only defect ship once already. +`dash`-only defect ship once already. The metrics report's suite also needs +**`python3`** 3.8 or later, addressed by name for the same reason as `dash`: it is the +system interpreter, not a pinned tool. **The `--exclude=SC2015` on the test file** is a single-code exclusion, not a blanket disable: every other shellcheck rule still applies to that file. Its hits are all diff --git a/CLAUDE.md b/CLAUDE.md index 4cc4ad0..39fd417 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -1358,9 +1358,9 @@ like the rest of §5; the detection is a reader comparing the pass against the s **Every cycle records one provenance line in its closing commit body** — default floor or not, so an absent line is never ambiguous between "the default applied" and "someone forgot". **One line per cycle**, so a change running five cycles records five. There is no informal - variant; anything quoting this form elsewhere quotes an instance of it, because the deferred - metrics work is intended to parse it — that consumer does not exist yet, and the form is pinned - now so that it can. + variant; anything quoting this form elsewhere quotes an instance of it, because tooling parses + it — in this repository `scripts/ledger-metrics.py` (dark-factory vision step 2b) — and the + form is pinned so that it can. ; floor per ; hook reminder threshold diff --git a/README.md b/README.md index 256ec06..7598002 100644 --- a/README.md +++ b/README.md @@ -157,10 +157,15 @@ honest gap ([reasoning](docs/coding-workflow.md#adapting-it-to-another-project)) ## Contributing CI ([`.github/workflows/ci.yml`](.github/workflows/ci.yml)) runs four checks on every -PR and push to main: `shellcheck --shell=sh` over all three executables and their test -files, the hook's test suite, +PR and push to main: `shellcheck --shell=sh` over the three shell executables and every +shell test file, the hook's test suite, [`scripts/check-invariants.sh`](scripts/check-invariants.sh) (invariants 5 and 6, plus three prompt-conformance checks) plus -both checkers' regression suites, and `claude plugin validate . --strict`. +both checkers' regression suites, the metrics report's suite, and +`claude plugin validate . --strict`. + +`python3 scripts/ledger-metrics.py` prints a read-only report over the hardening ledger +and the review-cycle records in commit bodies: which fingerprints recur, how rungs were +followed, and each cycle's recorded curve beside the `fic2` baseline. It writes nothing. A fifth check runs **on pull requests only**: [`scripts/check-version-bump.sh`](scripts/check-version-bump.sh) (invariant 12), which diff --git a/scripts/ledger-metrics.py b/scripts/ledger-metrics.py new file mode 100644 index 0000000..3bd1257 --- /dev/null +++ b/scripts/ledger-metrics.py @@ -0,0 +1,513 @@ +#!/usr/bin/env python3 +"""Read-only report over docs/hardening-log.md and the review-cycle records in git commit bodies. + +Spec: docs/superpowers/specs/2026-10-01-passive-metrics-design.md +Usage: python3 scripts/ledger-metrics.py [] (default HEAD) + +Reads the ledger and the commit bodies at ONE resolved commit, never the working tree, and +prints a report on stdout. It writes no file, ref or config itself; git can still write when +the caller's environment tells it to (spec §2), which the report header states. +Standard library only; Python 3.8 or later. +""" +import os +import re +import subprocess +import sys + +LEDGER = "docs/hardening-log.md" +ALLOWED_RUNGS = ("1 prose", "2 lint", "3 type", "4 test", "P std", "pending") +FIC2 = (" findings 14,24,12,3,6,6,2 (? 0) blockers 3,4,0,0,0,0,0 (? 0)" + " majors 5,13,6,2,5,?,? (? 2)") +CAVEATS = ( + "1. The curves are author-written and unchecked. Nothing compares them against the validated" + " pass files, so they are self-reported and not measurement.", + "2. The cycles reviewed different artifacts, so a difference is evidence about the population" + " as much as about the rule.", + "3. No demotion figure is derivable. That would need one finding classified under both rules," + " and nothing records that.", +) +CANNOT = """ +== What this report cannot answer +- Unlogged recurrences: the ledger only knows recurrences someone hardened. +- Whether a guard held: row succession is not guard failure, and a missing later row is not success. +- Cycles with no record at all: before 0.11.0 the curve was a habit, not a rule. A cycle with no provenance line, curve or skip record is invisible; a malformed record shows up as unparsed only when its line still starts like a record ("cycle ;"). +- Nonce attribution: records are grouped by nonce, which is collision-resistant, not collision-proof; two cycles that drew the same nonce read as one. +- Findings files: this script does not read .context/. It counts only what commit bodies say. +- Undecodable bytes: both sources are read as UTF-8 and an invalid byte becomes U+FFFD before anything is grouped, so two fingerprints or paths differing only in invalid bytes count as one. Control characters are shown as \\xNN, which can look like a literal backslash sequence. +- A shallow boundary that appears and disappears during the run: shallow state is checked before and after reading history, not during it. +- Whether a curve is true: see point 1 above. +- Cost and duration: that is vision step 2c's telemetry, out of scope here.""" + + +def die(msg): + sys.stderr.buffer.write(("ledger-metrics: %s\n" % msg).encode("utf-8", "backslashreplace")) + sys.stderr.flush() + sys.exit(1) + + +def git(*args): + """Run a read-only git command; return (returncode, stdout bytes). Stderr is not shown, + so the only error a caller sees is this script's own one-line message.""" + try: + # --no-replace-objects and an empty graft file: read the objects and history the SHA + # names, not `git replace` substitutes or legacy grafts. Both affect reading only. + env = dict(os.environ, GIT_GRAFT_FILE=os.devnull) + env.pop("GIT_SHALLOW_FILE", None) # the repository's own shallow file, not an override + p = subprocess.run(("git", "--no-replace-objects") + args, env=env, + stdout=subprocess.PIPE, stderr=subprocess.PIPE) + except OSError: + die("git could not be run") + return p.returncode, p.stdout + + +# ---- grammar (CLAUDE.md §5 Mechanics) ------------------------------------------------------ + +NONCE = r"[a-z0-9]{8,16}" +FIELD = r"(none \(pre-rule\)|" + NONCE + r")" +KIND = r"(Gate-A spec|Gate-A plan|Gate B)" +CANDIDATE = re.compile(r"^cycle (none \(pre-rule\)|[^ ;]*);") +BARE_PATH = re.compile(r"[A-Za-z0-9._/-]+") +BARE_MODEL = re.compile(r"[!#-'*\-./0-9<-~]+") # printable ASCII minus space " ( ) + , : ; +COUNT = re.compile(r"0|[1-9][0-9]*|\?") + + +def is_control(c): + """C0, DEL and the C1 range U+0080-U+009F.""" + o = ord(c) + return o < 0x20 or 0x7F <= o <= 0x9F + + +def read_quoted(s, i): + """s[i] == '"'. Return (decoded, next_index) or None. Escapes \\" and \\\\ only; any + control character makes the value unrepresentable, so the record is malformed.""" + out, i = [], i + 1 + while i < len(s): + c = s[i] + if is_control(c): + return None + if c == "\\": + if i + 1 < len(s) and s[i + 1] in "\"\\": + out.append(s[i + 1]) + i += 2 + continue + return None + if c == '"': + return ("".join(out), i + 1) if out else None + out.append(c) + i += 1 + return None + + +def parse_set(s, i): + """Parse at s[i:]. Return (entries, next_index) or None. entries is None for + `none`, else a list of (decoded_path, level) with level '0'..'2' or 'unprofiled'.""" + if s.startswith("none", i): + return None, i + 4 + if not s.startswith("{", i): + return None + i += 1 + entries, seen = [], set() + while True: + if i < len(s) and s[i] == '"': + q = read_quoted(s, i) + if q is None: + return None + path, i = q + else: + m = BARE_PATH.match(s, i) + if not m: + return None + path, i = m.group(0), m.end() + m = re.compile(r" \((level ([012])|unprofiled)\)").match(s, i) + if not m: + return None + level = m.group(2) if m.group(2) is not None else "unprofiled" + i = m.end() + if path in seen: # compared decoded, so "a.md" and a.md are the same path + return None + seen.add(path) + entries.append((path, level)) + if s.startswith(",", i): + i += 1 + continue + if s.startswith("}", i): + return entries, i + 1 + return None + + +def parse_model(s, i): + """One at s[i:]; return next index or None.""" + if i < len(s) and s[i] == '"': + q = read_quoted(s, i) + return None if q is None else q[1] + m = BARE_MODEL.match(s, i) + return m.end() if m else None + + +def expand(spec, limit): + """Expand "1-3,5" to [1, 2, 3, 5]. The total is checked against `limit` (the length of the + supplied count series) BEFORE anything is allocated, so a huge range costs nothing. Python's + integer-digit cap is lifted at start (see __main__); the ValueError guard is a last resort.""" + bounds, prev, total = [], 0, 0 + for part in spec.split(","): + m = re.fullmatch(r"([1-9][0-9]*)(?:-([1-9][0-9]*))?", part) + if not m: + return None + try: + a = int(m.group(1)) + b = int(m.group(2)) if m.group(2) else a + except ValueError: + return None + if a <= prev or b < a: + return None + total += b - a + 1 + if total > limit: + return None + bounds.append((a, b)) + prev = b + if total != limit: + return None + return [p for a, b in bounds for p in range(a, b + 1)] + + +def parse_curve(rest): + """rest is the text after '; '. Return dict or None.""" + m = re.compile(KIND + r" \(passes ([0-9,-]+), ").match(rest) + if not m: + return None + kind, spec, i = m.group(1), m.group(2), m.end() + tail = re.compile(r"\): Findings ([^ ]+)\. Blockers ([^ ]+)\. Majors ([^ ]+)\.$").search(rest) + if not tail: + return None + passes = expand(spec, len(tail.group(1).split(","))) + if passes is None: + return None + if rest.startswith("pass ", i): # per-pass models + for k, p in enumerate(passes): + if k: + if not rest.startswith("; ", i): + return None + i += 2 + m = re.compile(r"pass ([1-9][0-9]*) ").match(rest, i) + if not m or m.group(1) != str(p): # compared as text: no int() on untrusted digits + return None + i = m.end() + while True: + j = parse_model(rest, i) + if j is None: + return None + i = j + if rest.startswith("+", i): + i += 1 + continue + break + else: + j = parse_model(rest, i) + if j is None: + return None + i = j + m = re.compile(r"\): Findings ([^ ]+)\. Blockers ([^ ]+)\. Majors ([^ ]+)\.$").match(rest, i) + if not m: + return None + series = [] + for raw in m.groups(): + vals = raw.split(",") + if len(vals) != len(passes) or not all(COUNT.fullmatch(v) for v in vals): + return None + series.append(raw) + return {"kind": kind, "spec": spec, "f": series[0], "b": series[1], "m": series[2]} + + +def parse_record(line): + """Return (field, type, data) for a valid record, or None for a malformed candidate.""" + m = re.compile(r"^cycle " + FIELD + r"; ").match(line) + if not m: + return None + field, rest = m.group(1), line[m.end():] + m = re.compile(r"floor ([1-9][0-9]*) per ").match(rest) + if m: + parsed = parse_set(rest, m.end()) + if parsed is None: + return None + entries, i = parsed + if not re.compile(r"; hook reminder threshold (absent|unusable|[1-9][0-9]*)$").fullmatch(rest, i): + return None + return field, "prov", {"floor": m.group(1), "set": rest[m.end():i], "entries": entries} + m = re.fullmatch(KIND + r": skipped \(see skip reason\)", rest) + if m: + return field, "skip", {"kind": m.group(1)} + c = parse_curve(rest) + if c: + return field, "curve", c + return None + + +# ---- sections ------------------------------------------------------------------------------ + +def shown(text): + """Render a raw line for output: control characters become \\xNN, so a malformed record + cannot move the cursor or break the report's line structure.""" + return "".join("\\x%02x" % ord(c) if is_control(c) else c for c in text) + + +def ledger_section(text): + out = ["", "== Ledger: recurrence (rows matched exactly as harden-finding greps column 2)"] + rowre = re.compile(r"^\| *[0-9-]{10} *\| *([^|]*?) *\|") + rows, irregular = [], [] + for no, line in enumerate(text.split("\n"), 1): + line = line[:-1] if line.endswith("\r") else line # CRLF ledgers read like LF ones + m = rowre.match(line) + if not m: + continue + cells = re.split(r"(? 1 else m.group(1) + width_ok = len(cells) == 7 + rung = cells[5].strip(" ") if width_ok else None + rows.append((no, fp, rung)) + if not width_ok: + irregular.append("irregular width: line %d (%d cells)" % (no, len(cells))) + if not re.fullmatch(r"[a-z0-9]+(-[a-z0-9]+)*", fp): + irregular.append("irregular fingerprint: line %d" % no) + if not rows: + out.append("no rows") + return out + by_fp = {} + for no, fp, rung in rows: + by_fp.setdefault(fp, []).append((no, rung)) + order = sorted(by_fp, key=lambda f: (-len(by_fp[f]), f)) + for fp in order: + n = len(by_fp[fp]) + out.append("%d %s lines %s rungs %s" % ( + n, shown(fp), ",".join(str(no) for no, _ in by_fp[fp]), + " > ".join("?" if r is None else shown(r) for _, r in by_fp[fp]))) + out.append("check one count: git --no-replace-objects show :docs/hardening-log.md | grep -cE" + " '^\\| *[0-9-]{10} *\\| * *\\|'") + out.append(" ( is the commit in the header; use a fingerprint from the list. No command" + " fits an irregular fingerprint, because grep would read it as a pattern.)") + out += ["", "== Ledger: rung holding"] + last = {fp: max(no for no, _ in v) for fp, v in by_fp.items()} + landed, followed, pending = {}, {}, 0 + for no, fp, rung in rows: + if rung is None: + continue + if rung not in ALLOWED_RUNGS: + irregular.append("irregular rung: line %d (rung '%s')" % (no, shown(rung))) + continue + if rung == "pending": + pending += 1 + continue + landed[rung] = landed.get(rung, 0) + 1 + if no < last[fp]: + followed[rung] = followed.get(rung, 0) + 1 + for rung in sorted(landed): # only ALLOWED_RUNGS reach here, so no escaping is needed + f = followed.get(rung, 0) + out.append("%s landed %d followed-by-same-fingerprint %d last-of-fingerprint %d" + % (rung, landed[rung], f, landed[rung] - f)) + out.append("pending %d" % pending) + out.append('"followed" means a later row with the same fingerprint exists - nothing more. It does not mean') + out.append("the rung failed (a later row can be a different sub-shape, an out-of-scope guard, or a resolved") + out.append('prerequisite), and "last" does not mean it held (an unhardened recurrence leaves no row).') + out += ["", "== Ledger: irregular rows"] + irregular.sort(key=lambda s: (int(re.search(r"line (\d+)", s).group(1)), s)) + out += irregular or ["none"] + return out + + +def read_commits(raw): + """Parse `git log -z --format='%H %ct%n%B'` output into (sha, ct, body_lines).""" + commits = [] + for rec in raw.split("\0"): + if not rec.strip("\n"): + continue + head, _, body = rec.lstrip("\n").partition("\n") + m = re.fullmatch(r"([0-9a-f]{40}) ([0-9]+)", head) + if not m: + die("git log produced an unexpected record header") + commits.append((m.group(1), int(m.group(2)), body.split("\n"))) + return commits + + +def git_sections(commits, absence): + recs = {} # key -> record dict (key: line text; pre-rule keys also carry the sha) + unparsed = [] # (ct, sha, line) + for sha, ct, lines in commits: + seen_here = set() + i = 0 + while i < len(lines): + line = lines[i] + pos = i + i += 1 + if not CANDIDATE.match(line): + continue + parsed = parse_record(line) + if parsed is None: + unparsed.append((ct, sha, line)) + continue + field, typ, data = parsed + text = line + if typ == "skip": # the reason is the text right after the marker (CLAUDE.md §5) + reason = [] + while i < len(lines) and lines[i] != "" and not CANDIDATE.match(lines[i]): + reason.append(lines[i]) + i += 1 + data["reason"] = " ".join(reason) + text = line + " | " + data["reason"] # identity uses the excerpt as found + # pre-rule records identify nothing, so every occurrence stays separate (commit + position) + key = (text, sha, pos) if field == "none (pre-rule)" else (text, None, None) + if key in seen_here: + continue + seen_here.add(key) + r = recs.setdefault(key, {"field": field, "type": typ, "data": data, "text": text, + "commits": []}) + r["commits"].append((ct, sha)) + for r in recs.values(): + r["commits"].sort() + r["first"] = r["commits"][0] + + groups = {} + for key, r in recs.items(): + gkey = ("pre", key) if r["field"] == "none (pre-rule)" else ("n", r["field"]) + groups.setdefault(gkey, []).append(r) + + out = [] + cycles, conflicts, prof, lic, f1 = [], [], [], [], [] + for gkey, rs in groups.items(): + first = min(r["first"] for r in rs) + provs = [r for r in rs if r["type"] == "prov"] + rest = [r for r in rs if r["type"] != "prov"] + for p in provs: # checkpoint evidence comes from every valid provenance line, conflicts included + ents = p["data"]["entries"] + licensed = bool(ents) and all(lv == "0" for _, lv in ents) + sha = p["first"][1][:12] + if licensed: + lic.append((p["first"], "%s recorded floor %s commit %s" % (p["field"], p["data"]["floor"], sha))) + if p["data"]["floor"] == "1": + f1.append((p["first"], "%s floor 1 set %s commit %s" % ( + p["field"], "licenses it" if licensed else "does not license it", sha))) + if len(provs) > 1 or len(rest) > 1: + for r in rs: # groups sort by earliest record, then nonce; records inside by their own + conflicts.append((first, rs[0]["field"], r["first"], "%s %s" % (r["field"], shown(r["text"])))) + continue + p = provs[0] if provs else None + c = rest[0] if rest else None + fld = rs[0]["field"] + floor = p["data"]["floor"] if p else "-" + sset = p["data"]["set"] if p else "-" + if c is None: + line = "%s no curve floor %s set %s" % (fld, floor, shown(sset)) + elif c["type"] == "skip": + line = "%s %s skipped floor %s set %s reason excerpt: %s commit %s" % ( + fld, c["data"]["kind"], floor, shown(sset), + shown(c["data"]["reason"]) or "no reason found", c["first"][1][:12]) + else: + d = c["data"] + line = "%s %s passes %s floor %s set %s findings %s (? %d) blockers %s (? %d) majors %s (? %d)" % ( + fld, d["kind"], d["spec"], floor, shown(sset), d["f"], d["f"].split(",").count("?"), + d["b"], d["b"].split(",").count("?"), d["m"], d["m"].split(",").count("?")) + if gkey[0] == "pre": + if not (c and c["type"] == "skip"): # a skip line already names its commit + line += " commit %s" % rs[0]["first"][1][:12] + line += " [may duplicate another pre-rule record]" + cycles.append((first, line)) + if p and c and c["type"] == "curve" and p["data"]["entries"] and \ + any(lv != "unprofiled" for _, lv in p["data"]["entries"]): + prof.append((first, " " + line)) + + def emit(items, empty): + items.sort(key=lambda t: t[:-1] + (t[-1],)) + if not items: + out.append(empty) + out.extend(t[-1] for t in items) + + emit(cycles, ("no joined cycles (every record is in a conflict below)" if conflicts else + "no cycle records") + absence) + several = sorted((r["first"], shown(r["text"]), ",".join(s[:12] for _, s in r["commits"])) + for r in recs.values() if len(r["commits"]) > 1) + out.extend("seen in several commits: %s -> %s" % (t, c) for _, t, c in several) + out += ["", "== Git: conflicts (records sharing a nonce that disagree; not listed as cycles)"] + emit(conflicts, "no conflicts" + absence) + out += ["", "== Git: unparsed candidate lines"] + emit([((ct, sha), "%s %s" % (sha[:12], shown(line))) for ct, sha, line in unparsed], + "no unparsed lines" + absence) + out += ["", "== Comparison with the fic2 baseline (story criterion 5)", + "profiled cycles (joined, with a curve, set has a (level N) entry):"] + emit(prof, " no profiled cycles with a curve" + absence) + out.append("fic2 baseline (docs/field-reports/2026-08-26-fic2-cycle-evidence.md, passes 1-5; story §1, passes 6-7):") + out.append(FIC2) + out += list(CAVEATS) + out += ["", "== First checkpoint: evidence (the reader decides)", + "provenance lines whose set licenses floor 1:"] + emit([(k, " " + v) for k, v in lic], " none" + absence) + out.append("provenance lines recording floor 1:") + emit([(k, " " + v) for k, v in f1], " none" + absence) + return out + + +def is_shallow(): + rc, out = git("rev-parse", "--is-shallow-repository") + if rc != 0: + die("git rev-parse --is-shallow-repository failed") + return out.decode().strip() == "true" + + +def main(argv): + if len(argv) > 2: + die("usage: ledger-metrics.py []") + ref = argv[1] if len(argv) == 2 else "HEAD" + rc, _ = git("rev-parse", "--git-dir") + if rc != 0: + die("not a git repository") + rc, out = git("rev-parse", "--verify", "--quiet", ref + "^{commit}") + if rc != 0: + die("ref does not resolve to a commit: %s" % shown(ref)) + sha = out.decode().strip() + rc, out = git("ls-tree", "--full-tree", sha, "--", LEDGER) + if rc != 0: + die("git ls-tree failed at %s" % sha) + entry = out.decode("utf-8", "replace").strip() + if not entry: + die("ledger absent at %s: %s" % (sha, LEDGER)) + if not re.match(r"^100(644|755) blob ", entry): + die("ledger is not a regular file at %s: %s" % (sha, LEDGER)) + rc, out = git("cat-file", "blob", "%s:%s" % (sha, LEDGER)) + if rc != 0: + die("ledger cannot be read at %s: %s" % (sha, LEDGER)) + ledger = out.decode("utf-8", "replace") + shallow = is_shallow() + # Bodies re-encoded to UTF-8 whatever i18n.logOutputEncoding says, so NUL framing holds; + # no signature output, which would land in the same stream; `--` so a file named like the + # SHA cannot make the argument ambiguous. + rc, out = git("-c", "i18n.logOutputEncoding=UTF-8", "log", "--encoding=UTF-8", "--no-show-signature", + "-z", "--format=%H %ct%n%B", sha, "--") + if rc != 0: + die("git log failed at %s" % sha) + commits = read_commits(out.decode("utf-8", "replace")) + shallow = is_shallow() or shallow # checked before and after the walk: either means truncated + absence = " (history truncated: absence not established)" if shallow else "" + + lines = [ + "ledger-metrics — read-only report", + "commit %s (from %s)" % (sha, shown(ref)), + "ledger: %s at that commit; the working tree is not read" % LEDGER, + "git writes: this script only runs read-only git commands; git may still write when your" + " environment tells it to (e.g. GIT_TRACE to a file, partial-clone fetches)", + "history: " + ("truncated (shallow clone): absence not established" if shallow else "complete"), + ] + lines += ledger_section(ledger) + lines += ["", "== Git: review cycles (all commits reachable from %s)" % sha] + lines += git_sections(commits, absence) + # UTF-8 whatever the locale says, so a valid non-ASCII path can always be printed. + out = "\n".join(lines) + "\n" + CANNOT + "\n" + sys.stdout.buffer.write(out.encode("utf-8", "backslashreplace")) + sys.stdout.flush() + return 0 + + +if __name__ == "__main__": + # One integer limit on every Python: 3.11+ caps int() at 4300 digits by default, 3.8-3.10 + # do not, and the §5 grammar has no digit limit. Lift the cap where it exists. + if hasattr(sys, "set_int_max_str_digits"): + sys.set_int_max_str_digits(0) + sys.exit(main(sys.argv)) diff --git a/scripts/ledger-metrics.test.sh b/scripts/ledger-metrics.test.sh new file mode 100644 index 0000000..27f9f80 --- /dev/null +++ b/scripts/ledger-metrics.test.sh @@ -0,0 +1,387 @@ +#!/bin/sh +# Regression suite for ledger-metrics.py. +# +# Spec: docs/superpowers/specs/2026-10-01-passive-metrics-design.md §6. Every case builds +# a throwaway repository with fixed author/committer dates, so commit SHAs and therefore +# the report are deterministic, and compares the report against expected text. +# +# Order matters: the prior-state counterfactual and the negative control run FIRST, so +# a suite that cannot fail is caught before any green case is believed. +set -u + +SCRIPT="$(cd "$(dirname "$0")" && pwd)/ledger-metrics.py" +pass_n=0; fail_n=0 +pass() { pass_n=$((pass_n + 1)); printf 'ok - %s\n' "$1"; } +fail() { fail_n=$((fail_n + 1)); printf 'FAIL - %s\n' "$1"; } + +work=$(mktemp -d) || work='' +# Abort rather than continue with an empty $work: every path below is built under it. +if [ -z "$work" ] || [ ! -d "$work" ]; then + printf 'FAIL - could not create a temporary directory; refusing to run\n' >&2 + exit 1 +fi +trap 'rm -rf "$work"' EXIT +# An empty HOME and XDG_CONFIG_HOME: git finds no global attributes, ignore or config files +# there, so nothing from the developer's home can change what a fixture commit stores. +mkdir -p "$work/home/.config" && HOME="$work/home" && XDG_CONFIG_HOME="$work/home/.config" && + export HOME XDG_CONFIG_HOME + +# Isolation: no global or system git config, no inherited repository variables, no init +# templates. Identity and dates are fixed per invocation, so every SHA is reproducible and +# the suite never touches the developer's git setup. +GIT_CONFIG_GLOBAL=/dev/null; GIT_CONFIG_NOSYSTEM=1; GIT_ATTR_NOSYSTEM=1 +export GIT_CONFIG_GLOBAL GIT_CONFIG_NOSYSTEM GIT_ATTR_NOSYSTEM +unset GIT_DIR GIT_WORK_TREE GIT_INDEX_FILE GIT_OBJECT_DIRECTORY GIT_ALTERNATE_OBJECT_DIRECTORIES \ + GIT_CEILING_DIRECTORIES GIT_DEFAULT_HASH GIT_COMMON_DIR GIT_NAMESPACE \ + GIT_CONFIG_COUNT GIT_CONFIG_PARAMETERS GIT_REPLACE_REF_BASE GIT_SHALLOW_FILE \ + GIT_GRAFT_FILE GIT_NO_REPLACE_OBJECTS +tick=1700000000 +gitc() { + GIT_AUTHOR_NAME=t GIT_AUTHOR_EMAIL=t@example.com GIT_COMMITTER_NAME=t GIT_COMMITTER_EMAIL=t@example.com \ + GIT_AUTHOR_DATE="$tick +0000" GIT_COMMITTER_DATE="$tick +0000" \ + git -c user.name=t -c user.email=t@example.com -c commit.gpgsign=false \ + -c init.defaultBranch=main "$@" +} +# Commit everything with the given message (printf %b escapes), one second later each time. +# The message is an argument, never piped in: a pipeline runs `commit` in a subshell, +# which would lose the tick increment and give every commit the same date. +commit() { tick=$((tick + 1)); printf '%b' "$1" > "$work/msg"; gitc add -A && gitc commit -q --allow-empty -F "$work/msg"; } + +newrepo() { rm -rf "${work:?}/$1"; mkdir -p "$work/$1/docs" && (cd "$work/$1" && gitc init -q --template= --object-format=sha1 .); } +# Every run is wrapped in the no-write check: a run that changes the fixture leaves a marker +# outside it, and the suite fails at the end if any marker exists (spec §6). +# run DIR SCRIPT [REF]: stdout only, status preserved. +run() { + _b=$(snapshot "$1") + (cd "$1" && python3 "$2" ${3+"$3"}); _s=$? + [ "$_b" = "$(snapshot "$1")" ] || : > "$work/WROTE.$(basename "$1")" + return "$_s" +} +report() { run "$work/$1" "$2" ${3+"$3"} 2>&1; } + +# A whole-directory fingerprint: every path with its mode and size, plus a checksum of +# every file's contents. Equal before and after a run means the run left no file +# changed, added or removed in the fixture (spec §6 states what this does NOT catch). +snapshot() { (cd "$1" && ls -lAR . && find . -type f -exec cksum {} + | sort); } + +# The one comparison every golden case uses; the negative control calls it too, so a +# comparison that stopped comparing would be caught there. +same() { diff "$work/expected" "$work/actual" > "$work/diff"; } +expect() { # name, expected (stdin), actual + printf '%s\n' "$2" > "$work/actual" + cat > "$work/expected" + if same; then pass "$1" + else fail "$1"; sed 's/^/ /' "$work/diff"; fi +} + +# ---- fixture: the ledger -------------------------------------------------------- +LEDGER_HEAD='# Hardening log + +Columns: date, fingerprint, finding, source, severity, rung, ref. + +| date | fingerprint | finding | source | severity | rung | ref | +|------|-------------|---------|--------|----------|------|-----|' +mk_ledger_repo() { + newrepo L + { + printf '%s\n' "$LEDGER_HEAD" + printf '%s\n' '| 2026-01-01 | alpha-one | f1 | gate-a | major | 1 prose | r |' + printf '%s\n' '| 2026-01-02 | beta | a \| piped finding | bot | minor | P std | r |' + printf '%s\n' '| 2026-01-03 | alpha-one | f3 | bot | major | 2 lint | r |' + printf '%s\n' '| 2026-01-04 | alpha-one | short row | bot |' + printf '%s\n' '| 2026-01-05 | gamma | f5 | bot | nit | pending | r |' + printf '%s\n' '| 2026-01-06 | Bad-Case | f6 | bot | nit | 1 prose | r |' + printf '%s\n' '| 2026-01-07 | delta | f7 | bot | nit | | r |' + printf '%s\n' '| 2026-01-08 | delta | f8 | bot | nit | 0 | r |' + printf '%s\n' "- 2026-01-09 · supersedes 2026-01-01 \`alpha-one\` \"f1\" · wrong · see row" + } > "$work/L/docs/hardening-log.md" + (cd "$work/L" && commit 'ledger\n') +} + +ledger_section() { sed -n '/^== Ledger: recurrence/,/^== Git: review cycles/p' | sed '$d'; } +LEDGER_EXPECTED='== Ledger: recurrence (rows matched exactly as harden-finding greps column 2) +3 alpha-one lines 7,9,10 rungs 1 prose > 2 lint > ? +2 delta lines 13,14 rungs > 0 +1 Bad-Case lines 12 rungs 1 prose +1 beta lines 8 rungs P std +1 gamma lines 11 rungs pending +check one count: git --no-replace-objects show :docs/hardening-log.md | grep -cE '"'"'^\| *[0-9-]{10} *\| * *\|'"'"' + ( is the commit in the header; use a fingerprint from the list. No command fits an irregular fingerprint, because grep would read it as a pattern.) + +== Ledger: rung holding +1 prose landed 2 followed-by-same-fingerprint 1 last-of-fingerprint 1 +2 lint landed 1 followed-by-same-fingerprint 1 last-of-fingerprint 0 +P std landed 1 followed-by-same-fingerprint 0 last-of-fingerprint 1 +pending 1 +"followed" means a later row with the same fingerprint exists - nothing more. It does not mean +the rung failed (a later row can be a different sub-shape, an out-of-scope guard, or a resolved +prerequisite), and "last" does not mean it held (an unhardened recurrence leaves no row). + +== Ledger: irregular rows +irregular width: line 10 (4 cells) +irregular fingerprint: line 12 +irregular rung: line 13 (rung '"''"') +irregular rung: line 14 (rung '"'0'"')' + +# ---- 1. counterfactual: the prior state has no script -------------------------- +mk_ledger_repo +out=$(report L "$work/L/scripts/ledger-metrics.py"); st=$? +if [ "$st" -ne 0 ] && [ ! -e "$work/L/scripts/ledger-metrics.py" ]; then + pass "prior state: no script exists, so no report can be produced (exit $st)" +else fail "prior state: expected a failing run with no script, got exit $st"; fi + +# ---- 2. negative control: a count off by one must fail the count comparison ---------- +# The broken copy must still run to completion (exit 0), and its report must differ from +# the expected text in the count column - here alpha-one reads 4 instead of 3. A copy +# that crashes, or whose report still matches, proves nothing about the comparison. +sed 's/n = len(by_fp\[fp\])$/n = len(by_fp[fp]) + 1/' "$SCRIPT" > "$work/broken.py" +if cmp -s "$SCRIPT" "$work/broken.py"; then + fail "negative control: the mutation did not apply, so it proves nothing" +else + out=$(report L "$work/broken.py"); st=$? + printf '%s\n' "$LEDGER_EXPECTED" > "$work/expected" + printf '%s\n' "$out" | ledger_section > "$work/actual" + if [ "$st" -eq 0 ] && ! same && + grep -qx '4 alpha-one lines 7,9,10 rungs 1 prose > 2 lint > ?' "$work/actual"; then + pass "negative control: the broken copy exits 0 and the count comparison catches it" + else + fail "negative control: broken copy (exit $st) was not caught by the count comparison" + fi +fi + +# ---- 3. ledger section ------------------------------------------------------------ +out=$(report L "$SCRIPT"); st=$? +[ "$st" -eq 0 ] && pass "ledger fixture: exit 0" || fail "ledger fixture: exit $st" +expect "ledger: recurrence, rung holding and irregular rows" "$(printf '%s\n' "$out" | ledger_section)" </$c/; s//alpha-one/") +n=$(cd "$work/L" && sh -c "$cmd") +[ "$n" = 3 ] && pass "printed check command reproduces alpha-one = 3" || fail "printed check command gave '$n' ($cmd)" + +# The "cannot answer" footer is printed in full. +expect "footer: what the report cannot answer" "$(printf '%s\n' "$out" | sed -n '/^== What this report cannot answer/,$p')" <<'FOOTER' +== What this report cannot answer +- Unlogged recurrences: the ledger only knows recurrences someone hardened. +- Whether a guard held: row succession is not guard failure, and a missing later row is not success. +- Cycles with no record at all: before 0.11.0 the curve was a habit, not a rule. A cycle with no provenance line, curve or skip record is invisible; a malformed record shows up as unparsed only when its line still starts like a record ("cycle ;"). +- Nonce attribution: records are grouped by nonce, which is collision-resistant, not collision-proof; two cycles that drew the same nonce read as one. +- Findings files: this script does not read .context/. It counts only what commit bodies say. +- Undecodable bytes: both sources are read as UTF-8 and an invalid byte becomes U+FFFD before anything is grouped, so two fingerprints or paths differing only in invalid bytes count as one. Control characters are shown as \xNN, which can look like a literal backslash sequence. +- A shallow boundary that appears and disappears during the run: shallow state is checked before and after reading history, not during it. +- Whether a curve is true: see point 1 above. +- Cost and duration: that is vision step 2c's telemetry, out of scope here. +FOOTER + +# Uncommitted ledger edits are not read. +printf '%s\n' '| 2026-01-10 | alpha-one | uncommitted | bot | nit | 1 prose | r |' >> "$work/L/docs/hardening-log.md" +out2=$(report L "$SCRIPT"); st=$? +[ "$st" -eq 0 ] && [ "$out" = "$out2" ] && pass "working-tree ledger edit does not change the report" || fail "working-tree ledger edit changed the report" + +# Empty ledger. +newrepo E; printf '%s\n' "$LEDGER_HEAD" > "$work/E/docs/hardening-log.md"; (cd "$work/E" && commit 'e\n') +oute=$(report E "$SCRIPT"); st=$? +[ "$st" -eq 0 ] && pass "empty ledger: exit 0" || fail "empty ledger: exit $st" +printf '%s\n' "$oute" | grep -qx 'no rows' && pass "empty ledger prints 'no rows'" || fail "empty ledger" +printf '%s\n' "$oute" | grep -qx 'no cycle records' && pass "no records prints 'no cycle records'" || fail "no cycle records" +printf '%s\n' "$oute" | grep -qx ' no profiled cycles with a curve' && pass "comparison empty state" || fail "comparison empty state" +expect "checkpoint: both evidence lists empty" "$(printf '%s\n' "$oute" | sed -n '/^== First checkpoint/,/^$/p' | sed '$d')" <<'CHECKPOINT' +== First checkpoint: evidence (the reader decides) +provenance lines whose set licenses floor 1: + none +provenance lines recording floor 1: + none +CHECKPOINT + +# ---- 4. cycle records ------------------------------------------------------------------ +newrepo C; printf '%s\n' "$LEDGER_HEAD" > "$work/C/docs/hardening-log.md" +( + cd "$work/C" || exit 1 + commit 'ledger\n' + commit 'joined\n\ncycle aaaaaaaa; floor 3 per {docs/s.md (level 1)}; hook reminder threshold absent\ncycle aaaaaaaa; Gate B (passes 1-3,5, codex): Findings 4,3,?,1. Blockers 1,0,0,0. Majors 2,1,0,0.\n' + commit 'curve only\n\ncycle bbbbbbbb; Gate-A spec (passes 1, pass 1 codex+"my model"): Findings 2. Blockers 0. Majors 1.\n' + commit 'provenance only, quoted path\n\ncycle cccccccc; floor 1 per {"docs/q\\"x.md" (level 0)}; hook reminder threshold 3\ncycle closed on the zero-finding exit below the floor.\n' + commit 'skip with reason\n\ncycle dddddddd; Gate B: skipped (see skip reason)\nSkip reason: trivial,\nsecond line.\n\nafter the blank line\n' + commit 'skip without reason\n\ncycle eeeeeeee; Gate B: skipped (see skip reason)\n' + commit 'pre-rule\n\ncycle none (pre-rule); floor 3 per none; hook reminder threshold absent\ncycle none (pre-rule); Gate B (passes 1, codex): Findings 0. Blockers 0. Majors 0.\ncycle none (pre-rule); Gate B (passes 1, codex): Findings 0. Blockers 0. Majors 0.\n' + commit 'duplicate copy\n\ncycle bbbbbbbb; Gate-A spec (passes 1, pass 1 codex+"my model"): Findings 2. Blockers 0. Majors 1.\n' + commit 'conflicts\n\ncycle ffffffff; Gate B (passes 1, codex): Findings 1. Blockers 0. Majors 0.\ncycle ffffffff; Gate B (passes 1, codex): Findings 2. Blockers 0. Majors 0.\ncycle gggggggg; Gate B (passes 1, codex): Findings 1. Blockers 0. Majors 0.\ncycle gggggggg; Gate B: skipped (see skip reason)\nreason g\n\ncycle iiiiiiii; floor 3 per none; hook reminder threshold absent\ncycle iiiiiiii; floor 3 per {docs/s.md (level 0)}; hook reminder threshold absent\n' + commit 'skip reason A\n\ncycle hhhhhhhh; Gate B: skipped (see skip reason)\nreason A\n' + commit 'skip reason B\n\ncycle hhhhhhhh; Gate B: skipped (see skip reason)\nreason B\n' + commit 'malformed\n\ncycle abcdefg; floor 3 per none; hook reminder threshold absent\ncycle jjjjjjjj;floor 3 per none; hook reminder threshold absent\ncycle jjjjjjjj; floor 3 per {a.md (level 1),a.md (level 1)}; hook reminder threshold absent\ncycle jjjjjjjj; Gate B (passes 1-2, codex): Findings 1,2,3. Blockers 0,0. Majors 0,0.\ncycle jjjjjjjj; Gate B (passes 1-2, pass 1 codex): Findings 1,2. Blockers 0,0. Majors 0,0.\ncycle jjjjjjjj; Gate B (passes 1, "bad\\q"): Findings 1. Blockers 0. Majors 0.\ncycle jjjjjjjj; Gate B (passes 1, "tab\there"): Findings 1. Blockers 0. Majors 0.\ncycle jjjjjjjj; floor 3 per {a.md (level 1),"a.md" (level 1)}; hook reminder threshold absent\ncycle oooooooo; Gate B (passes 1, "a; b): c"): Findings 1. Blockers 0. Majors 0.\n' + commit 'skip then record\n\ncycle kkkkkkkk; Gate B: skipped (see skip reason)\ncycle llllllll; floor 1 per {docs/u.md (unprofiled)}; hook reminder threshold absent\n' + commit 'grammar 2\n\ncycle pppppppp; Gate-A plan (passes 1-2, pass 1 "model+variant"; pass 2 "a\\\\\\\\b"+codex): Findings 1,0. Blockers 0,0. Majors 1,0.\ncycle rrrrrrrr; floor 1 per {docs/a.md (level 0),docs/b.md (level 1)}; hook reminder threshold absent\ncycle ssssssss; Gate B (passes 1-100000000, codex): Findings 1. Blockers 0. Majors 0.\n' + gitc checkout -q -b side + commit 'side branch record\n\ncycle mmmmmmmm; Gate B (passes 1, codex): Findings 0. Blockers 0. Majors 0.\n' + gitc checkout -q main + tick=$((tick + 1)); gitc merge -q --no-ff -m 'merge side' side + commit 'licensed floor 1 recorded as floor 3\n\ncycle nnnnnnnn; floor 3 per {docs/z.md (level 0)}; hook reminder threshold absent\ncycle nnnnnnnn; Gate B (passes 1, codex): Findings 0. Blockers 0. Majors 0.\n' +) +git_section() { sed -n '/^== Git: review cycles/,/^== What this report cannot answer/p' | sed '1d;$d'; } +out=$(report C "$SCRIPT"); st=$? +[ "$st" -eq 0 ] && pass "cycle fixture: exit 0" || fail "cycle fixture: exit $st" +expect "cycles, conflicts, unparsed, comparison and checkpoint" "$(printf '%s\n' "$out" | git_section)" <<'EOF' +aaaaaaaa Gate B passes 1-3,5 floor 3 set {docs/s.md (level 1)} findings 4,3,?,1 (? 1) blockers 1,0,0,0 (? 0) majors 2,1,0,0 (? 0) +bbbbbbbb Gate-A spec passes 1 floor - set - findings 2 (? 0) blockers 0 (? 0) majors 1 (? 0) +cccccccc no curve floor 1 set {"docs/q\"x.md" (level 0)} +dddddddd Gate B skipped floor - set - reason excerpt: Skip reason: trivial, second line. commit ae01f24cfccb +eeeeeeee Gate B skipped floor - set - reason excerpt: no reason found commit 68042820934b +none (pre-rule) Gate B passes 1 floor - set - findings 0 (? 0) blockers 0 (? 0) majors 0 (? 0) commit 3333a18313c4 [may duplicate another pre-rule record] +none (pre-rule) Gate B passes 1 floor - set - findings 0 (? 0) blockers 0 (? 0) majors 0 (? 0) commit 3333a18313c4 [may duplicate another pre-rule record] +none (pre-rule) no curve floor 3 set none commit 3333a18313c4 [may duplicate another pre-rule record] +oooooooo Gate B passes 1 floor - set - findings 1 (? 0) blockers 0 (? 0) majors 0 (? 0) +kkkkkkkk Gate B skipped floor - set - reason excerpt: no reason found commit 13f09168f5c1 +llllllll no curve floor 1 set {docs/u.md (unprofiled)} +pppppppp Gate-A plan passes 1-2 floor - set - findings 1,0 (? 0) blockers 0,0 (? 0) majors 1,0 (? 0) +rrrrrrrr no curve floor 1 set {docs/a.md (level 0),docs/b.md (level 1)} +mmmmmmmm Gate B passes 1 floor - set - findings 0 (? 0) blockers 0 (? 0) majors 0 (? 0) +nnnnnnnn Gate B passes 1 floor 3 set {docs/z.md (level 0)} findings 0 (? 0) blockers 0 (? 0) majors 0 (? 0) +seen in several commits: cycle bbbbbbbb; Gate-A spec (passes 1, pass 1 codex+"my model"): Findings 2. Blockers 0. Majors 1. -> d61f00c143a8,34329462931c + +== Git: conflicts (records sharing a nonce that disagree; not listed as cycles) +ffffffff cycle ffffffff; Gate B (passes 1, codex): Findings 1. Blockers 0. Majors 0. +ffffffff cycle ffffffff; Gate B (passes 1, codex): Findings 2. Blockers 0. Majors 0. +gggggggg cycle gggggggg; Gate B (passes 1, codex): Findings 1. Blockers 0. Majors 0. +gggggggg cycle gggggggg; Gate B: skipped (see skip reason) | reason g +iiiiiiii cycle iiiiiiii; floor 3 per none; hook reminder threshold absent +iiiiiiii cycle iiiiiiii; floor 3 per {docs/s.md (level 0)}; hook reminder threshold absent +hhhhhhhh cycle hhhhhhhh; Gate B: skipped (see skip reason) | reason A +hhhhhhhh cycle hhhhhhhh; Gate B: skipped (see skip reason) | reason B + +== Git: unparsed candidate lines +7c5d9cd9db31 cycle abcdefg; floor 3 per none; hook reminder threshold absent +7c5d9cd9db31 cycle jjjjjjjj; Gate B (passes 1, "bad\q"): Findings 1. Blockers 0. Majors 0. +7c5d9cd9db31 cycle jjjjjjjj; Gate B (passes 1, "tab\x09here"): Findings 1. Blockers 0. Majors 0. +7c5d9cd9db31 cycle jjjjjjjj; Gate B (passes 1-2, codex): Findings 1,2,3. Blockers 0,0. Majors 0,0. +7c5d9cd9db31 cycle jjjjjjjj; Gate B (passes 1-2, pass 1 codex): Findings 1,2. Blockers 0,0. Majors 0,0. +7c5d9cd9db31 cycle jjjjjjjj; floor 3 per {a.md (level 1),"a.md" (level 1)}; hook reminder threshold absent +7c5d9cd9db31 cycle jjjjjjjj; floor 3 per {a.md (level 1),a.md (level 1)}; hook reminder threshold absent +7c5d9cd9db31 cycle jjjjjjjj;floor 3 per none; hook reminder threshold absent +4b1693862af7 cycle ssssssss; Gate B (passes 1-100000000, codex): Findings 1. Blockers 0. Majors 0. + +== Comparison with the fic2 baseline (story criterion 5) +profiled cycles (joined, with a curve, set has a (level N) entry): + aaaaaaaa Gate B passes 1-3,5 floor 3 set {docs/s.md (level 1)} findings 4,3,?,1 (? 1) blockers 1,0,0,0 (? 0) majors 2,1,0,0 (? 0) + nnnnnnnn Gate B passes 1 floor 3 set {docs/z.md (level 0)} findings 0 (? 0) blockers 0 (? 0) majors 0 (? 0) +fic2 baseline (docs/field-reports/2026-08-26-fic2-cycle-evidence.md, passes 1-5; story §1, passes 6-7): + findings 14,24,12,3,6,6,2 (? 0) blockers 3,4,0,0,0,0,0 (? 0) majors 5,13,6,2,5,?,? (? 2) +1. The curves are author-written and unchecked. Nothing compares them against the validated pass files, so they are self-reported and not measurement. +2. The cycles reviewed different artifacts, so a difference is evidence about the population as much as about the rule. +3. No demotion figure is derivable. That would need one finding classified under both rules, and nothing records that. + +== First checkpoint: evidence (the reader decides) +provenance lines whose set licenses floor 1: + cccccccc recorded floor 1 commit 4364bb042d4d + iiiiiiii recorded floor 3 commit 85b7cc11dd38 + nnnnnnnn recorded floor 3 commit 77b3167bb0fd +provenance lines recording floor 1: + cccccccc floor 1 set licenses it commit 4364bb042d4d + llllllll floor 1 set does not license it commit 13f09168f5c1 + rrrrrrrr floor 1 set does not license it commit 4b1693862af7 +EOF + +# A worktree file named like the commit SHA must not make `git log ` ambiguous, and a +# caller's log.showSignature must not leak signature text into the parsed stream. +sha=$(cd "$work/C" && git rev-parse HEAD); : > "$work/C/$sha" +out3=$(GIT_CONFIG_COUNT=1 GIT_CONFIG_KEY_0=log.showSignature GIT_CONFIG_VALUE_0=true report C "$SCRIPT"); st=$? +rm -f "$work/C/$sha" +[ "$st" -eq 0 ] && [ "$out3" = "$out" ] && + pass "SHA-named worktree file and log.showSignature leave the report unchanged" || + fail "SHA-named file / showSignature changed the report (exit $st)" + +# A 5001-digit pass number is valid under the §5 grammar on every Python (3.11+ caps int() +# at 4300 digits unless lifted), and a per-pass key that does not match its pass is unparsed. +newrepo T; printf '%s\n' "$LEDGER_HEAD" > "$work/T/docs/hardening-log.md" +long=$(printf '%05001d' 0 | tr 0 1) +(cd "$work/T" && commit "t\n\ncycle tttttttt; Gate B (passes $long, pass $long codex): Findings 1. Blockers 0. Majors 0.\ncycle uuuuuuuu; Gate B (passes 1, pass $long codex): Findings 1. Blockers 0. Majors 0.\n") +outt=$(report T "$SCRIPT"); st=$? +[ "$st" -eq 0 ] && printf '%s\n' "$outt" | grep -q '^tttttttt Gate B passes 1111' && + pass "5001-digit pass number: parsed as a cycle" || fail "5001-digit pass number (exit $st)" +printf '%s\n' "$outt" | grep -q '^[0-9a-f]\{12\} cycle uuuuuuuu; Gate B (passes 1, pass 1111' && + pass "per-pass key not matching its pass: unparsed, no crash" || fail "mismatched per-pass key" + +# Two conflict groups whose first records share a commit stay contiguous (group, then nonce). +newrepo O; printf '%s\n' "$LEDGER_HEAD" > "$work/O/docs/hardening-log.md" +( + cd "$work/O" || exit 1 + commit 'o1\n\ncycle aaaaaaaa; Gate B (passes 1, codex): Findings 1. Blockers 0. Majors 0.\ncycle bbbbbbbb; Gate B (passes 1, codex): Findings 1. Blockers 0. Majors 0.\n' + commit 'o2\n\ncycle bbbbbbbb; Gate B (passes 1, codex): Findings 2. Blockers 0. Majors 0.\n' + commit 'o3\n\ncycle aaaaaaaa; Gate B (passes 1, codex): Findings 2. Blockers 0. Majors 0.\n' +) +outo=$(report O "$SCRIPT"); st=$? +order=$(printf '%s\n' "$outo" | sed -n '/^== Git: conflicts/,/^$/p' | sed -n 's/^\([a-z]*\) cycle [a-z]*; Gate B (passes 1, codex): Findings \([0-9]\).*/\1\2/p' | tr '\n' ' ') +[ "$st" -eq 0 ] && [ "$order" = "aaaaaaaa1 aaaaaaaa2 bbbbbbbb1 bbbbbbbb2 " ] && + pass "tied conflict groups stay contiguous" || fail "conflict order: $order" + +# A pre-rule skip names its commit once. +newrepo Q; printf '%s\n' "$LEDGER_HEAD" > "$work/Q/docs/hardening-log.md" +(cd "$work/Q" && commit 'q\n\ncycle none (pre-rule); Gate B: skipped (see skip reason)\nwhy\n') +outq=$(report Q "$SCRIPT"); st=$? +n=$(printf '%s\n' "$outq" | grep '^none (pre-rule) Gate B skipped' | grep -o ' commit ' | wc -l | tr -d ' ') +[ "$st" -eq 0 ] && [ "$n" = 1 ] && pass "pre-rule skip names its commit once" || fail "pre-rule skip: exit $st, commit count $n" + +# An escaped final pipe does not close a row, and CRLF rows read like LF rows. +newrepo X +{ printf '%s\n' "$LEDGER_HEAD" + printf '%s\n' '| 2026-01-01 | esc | f | bot | nit | 1 prose | ref ends r\|' + printf '%s\r\n' '| 2026-01-02 | crlf | f | bot | nit | 2 lint | r |' +} > "$work/X/docs/hardening-log.md" +(cd "$work/X" && commit 'x\n') +outx=$(report X "$SCRIPT"); st=$? +[ "$st" -eq 0 ] && printf '%s\n' "$outx" | grep -qx '1 esc lines 7 rungs 1 prose' && + printf '%s\n' "$outx" | grep -qx '1 crlf lines 8 rungs 2 lint' && + printf '%s\n' "$outx" | sed -n '/^== Ledger: irregular rows/,+1p' | grep -qx 'none' && + pass "escaped final pipe and CRLF rows are regular" || fail "escaped pipe / CRLF rows" + +# The one-line error is UTF-8 whatever the locale says. +PYTHONIOENCODING=latin1 run "$work/L" "$SCRIPT" "$(printf 'caf\303\251')" > /dev/null 2> "$work/err"; st=$? +e=$(od -An -tx1 < "$work/err" | tr -d ' \n') +case $st:$e in 1:*c3a9*) pass "error message is UTF-8 under a latin1 locale (exit 1)" ;; *) fail "UTF-8 error: exit $st, bytes $e" ;; esac + +# Only conflicting records: the cycle list says why it is empty. +newrepo K; printf '%s\n' "$LEDGER_HEAD" > "$work/K/docs/hardening-log.md" +(cd "$work/K" && commit 'k\n\ncycle qqqqqqqq; Gate B (passes 1, codex): Findings 1. Blockers 0. Majors 0.\ncycle qqqqqqqq; Gate B (passes 1, codex): Findings 2. Blockers 0. Majors 0.\n') +outk=$(report K "$SCRIPT"); st=$? +[ "$st" -eq 0 ] && printf '%s\n' "$outk" | grep -qx 'no joined cycles (every record is in a conflict below)' && + pass "conflict-only history: the cycle list says why it is empty" || fail "conflict-only empty label" + +# ---- 5. exit-1 causes ----------------------------------------------------------------- +bad() { # name, expected stderr substring, dir, [ref] + e=$(run "$3" "$SCRIPT" ${4+"$4"} 2>&1 >/dev/null); s=$? + if [ "$s" -eq 1 ] && printf '%s' "$e" | grep -qF "$2"; then pass "$1"; else fail "$1 (exit $s: $e)"; fi +} +mkdir -p "$work/norepo" +bad "not a git repository" "ledger-metrics: not a git repository" "$work/norepo" +bad "unresolvable ref" "ledger-metrics: ref does not resolve to a commit: nope" "$work/L" nope +newrepo A; (cd "$work/A" && printf 'x\n' > x && commit 'x\n') +bad "ledger absent" "ledger-metrics: ledger absent at" "$work/A" +newrepo D; mkdir -p "$work/D/docs/hardening-log.md" && printf 'x\n' > "$work/D/docs/hardening-log.md/f"; (cd "$work/D" && commit 'd\n') +bad "ledger is a directory" "ledger-metrics: ledger is not a regular file" "$work/D" +newrepo S; printf '%s\n' "$LEDGER_HEAD" > "$work/S/real.md"; ln -s ../real.md "$work/S/docs/hardening-log.md"; (cd "$work/S" && commit 's\n') +bad "ledger is a symlink" "ledger-metrics: ledger is not a regular file" "$work/S" +newrepo G; printf '%s\n' "$LEDGER_HEAD" > "$work/G/docs/hardening-log.md" +(cd "$work/G" && commit 'one\n' && printf 'y\n' > y && commit 'two\n') +first=$(cd "$work/G" && git rev-parse HEAD~1) +rm -f "$work/G/.git/objects/$(printf '%s' "$first" | cut -c1-2)/$(printf '%s' "$first" | cut -c3-)" +bad "git log fails" "ledger-metrics: git log failed" "$work/G" +newrepo U; printf '%s\n' "$LEDGER_HEAD" > "$work/U/docs/hardening-log.md"; (cd "$work/U" && commit 'u\n') +blob=$(cd "$work/U" && git rev-parse HEAD:docs/hardening-log.md) +rm -f "$work/U/.git/objects/$(printf '%s' "$blob" | cut -c1-2)/$(printf '%s' "$blob" | cut -c3-)" +bad "ledger cannot be read" "ledger-metrics: ledger cannot be read" "$work/U" + +# ---- 6. shallow clone ------------------------------------------------------------------- +rm -rf "$work/shallow"; git clone -q --depth 1 "file://$work/C" "$work/shallow" 2>/dev/null +out=$(report shallow "$SCRIPT"); st=$? +[ "$st" -eq 0 ] && printf '%s\n' "$out" | grep -qx 'history: truncated (shallow clone): absence not established' && + pass "shallow clone: header says history is truncated" || fail "shallow clone header" + +set -- "$work"/WROTE.* +if [ -e "$1" ]; then fail "no-write: a run changed its fixture ($*)" +else pass "no-write: no run changed its fixture"; fi + +printf '\n%d passed, %d failed\n' "$pass_n" "$fail_n" +[ "$fail_n" -eq 0 ] From 33f23d911de279c2112edd55c30fe8b8a0719d2c Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20S=C3=A4nger?= <20968534+dsnger@users.noreply.github.com> Date: Thu, 1 Oct 2026 21:55:49 +0200 Subject: [PATCH 5/5] Fix PR #33 review findings: spec pass rule, plan counts, header test MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Greptile found three true claims on PR #33: - The spec still described a 10000-pass limit that plan ruling 1 had removed. The grammar paragraph now states the implemented rule, marked as updated after implementation. - The executed plan still expected 30 suite cases. It now carries a dated note: its counts describe the suite as approved, the committed suite has 32, and "the spec is not edited" was the planning-time decision. - The suite never compared the report header. A golden five-line header case was added; it was mutation-checked by changing the history line. docs/hardening-log.md records the ninth docs-drift occurrence. Evidence — docs/superpowers/stories/2026-08-04-passive-metrics-over-the-ledger-story.md Battery: AGENTS.md quality row, exit 0 at d5e9dcb4bcd4af13964c6c08b508f4a35eab3062. Check (counterfactual): the new header case fails against a copy whose history line is changed, and passes against the real script. Suite 32/32 under sh (in the quality row) and under dash. cycle p4hht73863; floor 3 per {docs/superpowers/stories/2026-08-04-passive-metrics-over-the-ledger-story.md (level 1)}; hook reminder threshold absent cycle p4hht73863; Gate B (passes 1-2, gpt-6-astra): Findings 5,0. Blockers 0,0. Majors 0,0. Each Gate-B pass was one logical pass run as two calls (spec, then quality) against baseSha 0d3da0a484861796bdb3bc685497fb3bb52ccc1c and the headSha resolved before each call (pass 1: 8f2a4b741b06dc0649203ce538756380065b76ee; pass 2: d5e9dcb4bcd4af13964c6c08b508f4a35eab3062). Pass 1's spec branch found 3 findings and its quality branch 2, two of them the same complaints; per §5 the branches are summed, so the pass records 5. Pass 2 found nothing in either branch, so the cycle closed on the zero-finding exit. Human exceptions: none --- docs/hardening-log.md | 1 + docs/superpowers/plans/2026-10-01-passive-metrics.md | 2 ++ .../specs/2026-10-01-passive-metrics-design.md | 8 +++++--- scripts/ledger-metrics.test.sh | 10 ++++++++++ 4 files changed, 18 insertions(+), 3 deletions(-) diff --git a/docs/hardening-log.md b/docs/hardening-log.md index 2bfa02d..68702c9 100644 --- a/docs/hardening-log.md +++ b/docs/hardening-log.md @@ -127,3 +127,4 @@ escape `\|`, one line), `source` (gate-a|gate-b|bot|manual), | 2026-09-02 | verification-masks-failure | fifth occurrence: PR #26 (CodeRabbit) — Plan C task 25's completion assert tested only that the provisional wording was ABSENT (`grep -c … -eq 0`). A replacement that deleted the provisional passage and wrote no closed record at all would have passed it, which is precisely the outcome the task exists to prevent. The task did run correctly this cycle, so the mask never fired; it was found by reading, not by failing | bot | major | 4 test | The assert gained a positive arm beside the negative one: the closed curves must be PRESENT, with a cardinality floor (`grep -cE '^Findings( +[0-9]+,?)+' … -ge 3`). PRIOR ROWS under this fingerprint all share one shape — a check whose only assertion is that something is gone. WHAT GENERALIZES: an absence assert is half a check whenever the edit it guards is a replacement rather than a deletion, and the missing half is always the same one. This row's remedy is specific to task 25; the general form belongs to the loop-rule consolidation story, which owns the plan-assert conventions | | 2026-09-25 | missing-input-validation | first row of this base class here: PR #27 (Greptile P1, thread 4091660211) — `/dev-workflow:claude-init` filled an unconstrained project name into the create command's shell here-document; a name with a line break could put the fixed delimiter on its own line, end the document early and run the rest of the name and the template as shell. Confirmed as transport under sh and dash with a harmless marker, not as an end-to-end command run | bot | blocker | 1 prose | claude-init.md *Writing* step 2 now states the rule (one line, no Unicode `Cc` character, every name source, ask on failure, nothing sanitized) and gives a Python check for a directory-derived name; spec §2 carries it. AGENTS.md Don'ts gains the class-level rule: data spliced into shell source a prompt tells the agent to run must be bounded, with the rule, the check and the failure path stated. Does NOT guard: the rule is agent-followed — the create script does not validate the name, and a user-given name is checked by reading. NO DETERMINISTIC RUNG EXISTS: nothing mechanical can tell which prompt text becomes shell source; it is a reading check | | 2026-09-30 | docs-drift | eighth occurrence: PR #32 (Greptile P2) — the sparring-skill spec said "The eight walkthrough scenarios … the last two were added by pass 1" above a list numbering twelve; the second Gate-A repair round added 9–12 and left the count sentence. Gate-A spec pass 3 had already found it and collected it as a NIT, so detection held and only the collect-only triage let it ship. Fixed in the same PR (the sentence now says twelve and names which round added which). | bot | nit | 1 prose | No new text: the existing guard already covers it — CLAUDE.md §5 Gate-B lens, "Name what this diff changes the size, value or position of … and grep for where each is described elsewhere", which a Gate-A repair that grows an enumerated list should apply to its own count sentence. NO DETERMINISTIC RUNG EXISTS for this shape: the rung-2 count lint (2026-07-26 row) matches only the "all N checklist items" spelling, and docs/superpowers/ is excluded from it on purpose as dated records. Recorded so the recurrence count stays true; escalate if a count-vs-list drift ships from a shipped prompt rather than a historical artifact. PRIOR ROW: 2026-09-02 docs-drift (1 prose). | +| 2026-10-01 | docs-drift | ninth occurrence: PR #33 (Greptile) — the passive-metrics spec still described a 10000-pass limit that the implementation had removed under plan ruling 1, and the executed plan still expected `30 passed` after Gate B added suite cases. Both were statements about this same change, left behind by later steps of it. Fixed in the same PR: the spec now states the implemented rule, marked as updated after implementation, and the plan carries a dated note that its counts describe the suite as approved. | bot | minor | 1 prose | No new text: CLAUDE.md §5 Gate B already says a fix that changes specified behaviour updates the spec in the same commit, and that rule covers this case. What slipped was applying it to a ruling made in the plan rather than a Gate-B fix. NO DETERMINISTIC RUNG: nothing can tell that a ruling contradicts spec prose. PRIOR ROW: 2026-09-30 docs-drift (1 prose). | diff --git a/docs/superpowers/plans/2026-10-01-passive-metrics.md b/docs/superpowers/plans/2026-10-01-passive-metrics.md index 3da0b32..a12ca8a 100644 --- a/docs/superpowers/plans/2026-10-01-passive-metrics.md +++ b/docs/superpowers/plans/2026-10-01-passive-metrics.md @@ -14,6 +14,8 @@ **This plan replaces the awk plan reviewed in Gate-A plan pass 1** (`.context/codex-reviews/gate-a-plan-mtf7ua7qze-pass-1.md`). That pass's findings were about awk parsing, mawk semantics and suite isolation; the table at the end says where each landed. +**Executed 2026-10-01 (PR #33).** Gate B and PR review added suite cases after this plan was approved, so the committed `scripts/ledger-metrics.test.sh` has more cases than the copy embedded below. The counts below (`30 passed`) describe the suite as the plan approved it; after PR #33's review fixes the committed suite has 32 cases. Ruling 1 below was also carried into the spec afterwards (its grammar paragraph is marked "updated after implementation"), so "the spec is not edited" describes the decision at planning time. + ## Global Constraints - Worktree `/Users/daniel/DEVELOPMENT/APPS/dwk-metrics`, branch `passive-metrics`. Paths are relative to it. diff --git a/docs/superpowers/specs/2026-10-01-passive-metrics-design.md b/docs/superpowers/specs/2026-10-01-passive-metrics-design.md index deb9dc6..466476a 100644 --- a/docs/superpowers/specs/2026-10-01-passive-metrics-design.md +++ b/docs/superpowers/specs/2026-10-01-passive-metrics-design.md @@ -187,9 +187,11 @@ the terminal cursor. *(revision)* **Grammar details the parser enforces** (CLAUDE.md §5 Mechanics): a quoted path or model may use only the escapes `\"` and `\\`, and any control character makes the record unparsed; a repeated story path is compared after decoding quotes and escapes, so `a.md` and `"a.md"` are the same path; a pass spec -that expands to more than 10000 passes makes the record unparsed, so a malformed range cannot exhaust -memory (a single large pass number is fine); each count series must have exactly one value per expanded -pass; per-pass model keys must be exactly the expanded passes, in order. *(revision)* +must expand to exactly as many passes as the Findings series has values, and that is checked before any +pass list is built, so a malformed range cannot exhaust memory and no pass-count ceiling is needed; the +Blockers and Majors series are then checked against the same pass count +*(updated after implementation, plan ruling 1; it replaces an earlier 10000-pass limit that §5's +grammar does not have)*; each count series must have exactly one value per expanded pass; per-pass model keys must be exactly the expanded passes, in order. *(revision)* ## §5 The review-loop comparison (story criterion 5) diff --git a/scripts/ledger-metrics.test.sh b/scripts/ledger-metrics.test.sh index 27f9f80..10de28a 100644 --- a/scripts/ledger-metrics.test.sh +++ b/scripts/ledger-metrics.test.sh @@ -152,6 +152,16 @@ fi # ---- 3. ledger section ------------------------------------------------------------ out=$(report L "$SCRIPT"); st=$? [ "$st" -eq 0 ] && pass "ledger fixture: exit 0" || fail "ledger fixture: exit $st" +# The report header, compared line for line: commit and ref, the working-tree disclosure, +# the git-writes limit and the history state. +csha=$(cd "$work/L" && git rev-parse HEAD) +expect "header: commit, ref, sources and history state" "$(printf '%s\n' "$out" | sed -n '1,5p')" <