Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 5 additions & 1 deletion .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -37,7 +37,7 @@ jobs:
# merge-base. ~50 commits, so the full history costs nothing here.
fetch-depth: 0

# The hook, the two checkers, the Python metrics report and their four suites are
# The hook, the two checkers, the two Python reports and their five suites are
# the only executables here,
# so lint plus those suites are the only mechanical gates we have. Everything else
# here is a prompt, and prompts have no typechecker (see README, "Contributing").
Expand Down Expand Up @@ -76,6 +76,9 @@ jobs:
# hence SC2015.
docker run --rm -v "$PWD:/mnt" -w /mnt koalaman/shellcheck:v0.11.0 \
--shell=sh --exclude=SC2015 scripts/ledger-metrics.test.sh
# The run-analytics suite (vision step 2c, part 1), linted the same way.
docker run --rm -v "$PWD:/mnt" -w /mnt koalaman/shellcheck:v0.11.0 \
--shell=sh --exclude=SC2015 scripts/run-analytics.test.sh

# Two runs, because the hook has to be correct under both shells and the runner's
# /bin/sh is dash while a contributor's may be bash. HOOK_SH selects the shell the
Expand All @@ -102,6 +105,7 @@ jobs:
sh scripts/check-invariants.sh
sh scripts/check-version-bump.test.sh
sh scripts/ledger-metrics.test.sh
sh scripts/run-analytics.test.sh

# Invariant 12 itself needs a base to diff against, so it runs only on pull
# requests — a push to main has no PR base, and diffing the push range would just
Expand Down
18 changes: 11 additions & 7 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -40,6 +40,8 @@ scripts/check-version-bump.sh # invariant 12, mechanically — PR-only (rung
scripts/check-version-bump.test.sh # its regression suite — policy/operational/accept
scripts/ledger-metrics.py # read-only report over the ledger + git (vision step 2b)
scripts/ledger-metrics.test.sh # its regression suite — fixture repos, expected text
scripts/run-analytics.py # gate-call effort from local logs (vision step 2c, part 1)
scripts/run-analytics.test.sh # its regression suite — fixture home + repos, expected text
plugins/dev-workflow/
.claude-plugin/plugin.json # metadata only — no component keys (invariant 6)
CHANGELOG.md # every manifest version, newest first
Expand All @@ -64,9 +66,10 @@ source-files/ # the extraction seed this repo was built from
**Boundaries.** `skills/`, `commands/`, `agents/` and `hooks/hooks.json` are loaded by convention
from their paths. The executable artifacts are the hook and its test, plus the two
repo-local CI checkers and their tests (`scripts/check-invariants.{sh,test.sh}` and
`scripts/check-version-bump.{sh,test.sh}`), plus the read-only Python report
`scripts/ledger-metrics.py` and its suite `scripts/ledger-metrics.test.sh` — the hook ships
in the plugin, the checkers and the report do not; everything else is text read by a model. `examples/` is reference material,
`scripts/check-version-bump.{sh,test.sh}`), plus two Python reports with their suites:
`scripts/ledger-metrics.{py,test.sh}` (read-only) and `scripts/run-analytics.{py,test.sh}`
(writes only its own store under `.context/telemetry/`) — the hook ships in the plugin, the
checkers and the reports do not; everything else is text read by a model. `examples/` is reference material,
outside the loaded surface: never scaffolded or copied into a user's project, though it
does ship inside the plugin package.

Expand Down Expand Up @@ -256,9 +259,9 @@ Every command below was run in this session and observed to exit 0.

| Role | Command |
|---|---|
| quality (the whole battery — what CI runs) | `shellcheck --shell=sh plugins/dev-workflow/hooks/codex-gate.sh && shellcheck --shell=sh --exclude=SC2015 plugins/dev-workflow/hooks/codex-gate.test.sh && shellcheck --shell=sh scripts/check-invariants.sh && shellcheck --shell=sh --exclude=SC2015 scripts/check-invariants.test.sh && shellcheck --shell=sh scripts/check-version-bump.sh && shellcheck --shell=sh scripts/check-version-bump.test.sh && shellcheck --shell=sh --exclude=SC2015 scripts/ledger-metrics.test.sh && HOOK_SH=sh sh plugins/dev-workflow/hooks/codex-gate.test.sh && HOOK_SH=dash dash plugins/dev-workflow/hooks/codex-gate.test.sh && sh scripts/check-invariants.test.sh && sh scripts/check-invariants.sh && sh scripts/check-version-bump.test.sh && sh scripts/check-version-bump.sh main && sh scripts/ledger-metrics.test.sh && claude plugin validate . --strict` |
| typecheck | n/a — no typed sources (shell, one untyped Python report, markdown) |
| lint | `shellcheck --shell=sh plugins/dev-workflow/hooks/codex-gate.sh && shellcheck --shell=sh --exclude=SC2015 plugins/dev-workflow/hooks/codex-gate.test.sh && shellcheck --shell=sh scripts/check-invariants.sh && shellcheck --shell=sh --exclude=SC2015 scripts/check-invariants.test.sh && shellcheck --shell=sh scripts/check-version-bump.sh && shellcheck --shell=sh scripts/check-version-bump.test.sh && shellcheck --shell=sh --exclude=SC2015 scripts/ledger-metrics.test.sh` |
| quality (the whole battery — what CI runs) | `shellcheck --shell=sh plugins/dev-workflow/hooks/codex-gate.sh && shellcheck --shell=sh --exclude=SC2015 plugins/dev-workflow/hooks/codex-gate.test.sh && shellcheck --shell=sh scripts/check-invariants.sh && shellcheck --shell=sh --exclude=SC2015 scripts/check-invariants.test.sh && shellcheck --shell=sh scripts/check-version-bump.sh && shellcheck --shell=sh scripts/check-version-bump.test.sh && shellcheck --shell=sh --exclude=SC2015 scripts/ledger-metrics.test.sh && shellcheck --shell=sh --exclude=SC2015 scripts/run-analytics.test.sh && HOOK_SH=sh sh plugins/dev-workflow/hooks/codex-gate.test.sh && HOOK_SH=dash dash plugins/dev-workflow/hooks/codex-gate.test.sh && sh scripts/check-invariants.test.sh && sh scripts/check-invariants.sh && sh scripts/check-version-bump.test.sh && sh scripts/check-version-bump.sh main && sh scripts/ledger-metrics.test.sh && sh scripts/run-analytics.test.sh && claude plugin validate . --strict` |
| typecheck | n/a — no typed sources (shell, two untyped Python reports, markdown) |
| lint | `shellcheck --shell=sh plugins/dev-workflow/hooks/codex-gate.sh && shellcheck --shell=sh --exclude=SC2015 plugins/dev-workflow/hooks/codex-gate.test.sh && shellcheck --shell=sh scripts/check-invariants.sh && shellcheck --shell=sh --exclude=SC2015 scripts/check-invariants.test.sh && shellcheck --shell=sh scripts/check-version-bump.sh && shellcheck --shell=sh scripts/check-version-bump.test.sh && shellcheck --shell=sh --exclude=SC2015 scripts/ledger-metrics.test.sh && shellcheck --shell=sh --exclude=SC2015 scripts/run-analytics.test.sh` |
| test | `HOOK_SH=sh sh plugins/dev-workflow/hooks/codex-gate.test.sh && HOOK_SH=dash dash plugins/dev-workflow/hooks/codex-gate.test.sh` — two runs; `HOOK_SH` selects the shell the HOOK runs under, and without it a dash invocation only exercises the harness |
| invariant checks (5 pinning, 6 manifest, prompt conformance) | `sh scripts/check-invariants.test.sh && sh scripts/check-invariants.sh` |
| invariant check (12 version bump) | `sh scripts/check-version-bump.test.sh && sh scripts/check-version-bump.sh main` |
Expand All @@ -273,7 +276,8 @@ invariant 5; `dash` is addressed by name because it is the system shell, not a p
tool. Without it the second run cannot start, and dropping that run is what let a
`dash`-only defect ship once already. The metrics report's suite also needs
**`python3`** 3.8 or later, addressed by name for the same reason as `dash`: it is the
system interpreter, not a pinned tool.
system interpreter, not a pinned tool. `scripts/run-analytics.py` also needs git 2.36 or later
and checks for it.

**The `--exclude=SC2015` on the test file** is a single-code exclusion, not a blanket
disable: every other shellcheck rule still applies to that file. Its hits are all
Expand Down
7 changes: 6 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -160,13 +160,18 @@ CI ([`.github/workflows/ci.yml`](.github/workflows/ci.yml)) runs four checks on
PR and push to main: `shellcheck --shell=sh` over the three shell executables and every
shell test file, the hook's test suite,
[`scripts/check-invariants.sh`](scripts/check-invariants.sh) (invariants 5 and 6, plus three prompt-conformance checks) plus
both checkers' regression suites, the metrics report's suite, and
both checkers' regression suites, the two reports' suites, and
`claude plugin validate . --strict`.

`python3 scripts/ledger-metrics.py` prints a read-only report over the hardening ledger
and the review-cycle records in commit bodies: which fingerprints recur, how rungs were
followed, and each cycle's recorded curve beside the `fic2` baseline. It writes nothing.

`python3 scripts/run-analytics.py` reads the Claude Code transcripts and Codex session logs
on this machine and reports how many gate calls each review cycle and story took, how long
they ran and how many tokens they used, including effort no closed cycle accounts for. It
keeps numbers and identifiers only, in `.context/telemetry/` of this clone, for 365 days.

A fifth check runs **on pull requests only**:
[`scripts/check-version-bump.sh`](scripts/check-version-bump.sh) (invariant 12), which
needs a base branch to diff against. Its *suite* runs unconditionally with the others;
Expand Down
1 change: 1 addition & 0 deletions docs/hardening-log.md
Original file line number Diff line number Diff line change
Expand Up @@ -128,3 +128,4 @@ escape `\|`, one line), `source` (gate-a|gate-b|bot|manual),
| 2026-09-25 | missing-input-validation | first row of this base class here: PR #27 (Greptile P1, thread 4091660211) — `/dev-workflow:claude-init` filled an unconstrained project name into the create command's shell here-document; a name with a line break could put the fixed delimiter on its own line, end the document early and run the rest of the name and the template as shell. Confirmed as transport under sh and dash with a harmless marker, not as an end-to-end command run | bot | blocker | 1 prose | claude-init.md *Writing* step 2 now states the rule (one line, no Unicode `Cc` character, every name source, ask on failure, nothing sanitized) and gives a Python check for a directory-derived name; spec §2 carries it. AGENTS.md Don'ts gains the class-level rule: data spliced into shell source a prompt tells the agent to run must be bounded, with the rule, the check and the failure path stated. Does NOT guard: the rule is agent-followed — the create script does not validate the name, and a user-given name is checked by reading. NO DETERMINISTIC RUNG EXISTS: nothing mechanical can tell which prompt text becomes shell source; it is a reading check |
| 2026-09-30 | docs-drift | eighth occurrence: PR #32 (Greptile P2) — the sparring-skill spec said "The eight walkthrough scenarios … the last two were added by pass 1" above a list numbering twelve; the second Gate-A repair round added 9–12 and left the count sentence. Gate-A spec pass 3 had already found it and collected it as a NIT, so detection held and only the collect-only triage let it ship. Fixed in the same PR (the sentence now says twelve and names which round added which). | bot | nit | 1 prose | No new text: the existing guard already covers it — CLAUDE.md §5 Gate-B lens, "Name what this diff changes the size, value or position of … and grep for where each is described elsewhere", which a Gate-A repair that grows an enumerated list should apply to its own count sentence. NO DETERMINISTIC RUNG EXISTS for this shape: the rung-2 count lint (2026-07-26 row) matches only the "all N checklist items" spelling, and docs/superpowers/ is excluded from it on purpose as dated records. Recorded so the recurrence count stays true; escalate if a count-vs-list drift ships from a shipped prompt rather than a historical artifact. PRIOR ROW: 2026-09-02 docs-drift (1 prose). |
| 2026-10-01 | docs-drift | ninth occurrence: PR #33 (Greptile) — the passive-metrics spec still described a 10000-pass limit that the implementation had removed under plan ruling 1, and the executed plan still expected `30 passed` after Gate B added suite cases. Both were statements about this same change, left behind by later steps of it. Fixed in the same PR: the spec now states the implemented rule, marked as updated after implementation, and the plan carries a dated note that its counts describe the suite as approved. | bot | minor | 1 prose | No new text: CLAUDE.md §5 Gate B already says a fix that changes specified behaviour updates the spec in the same commit, and that rule covers this case. What slipped was applying it to a ruling made in the plan rather than a Gate-B fix. NO DETERMINISTIC RUNG: nothing can tell that a ruling contradicts spec prose. PRIOR ROW: 2026-09-30 docs-drift (1 prose). |
| 2026-10-02 | verification-masks-failure | sixth occurrence, two cases in one suite: PR #34 (Greptile P2) — `scripts/run-analytics.test.sh` accepted its no-text negative control when the mutated collector wrote **no** store (`grep … \|\| ! [ -s "$STORE" ]`), so a crash read as "leak caught"; and a "prior state" counterfactual ran a deliberately nonexistent `run-analytics.missing`, which fails whatever the code under test does. Both survived three Gate-A plan passes and three Gate-B passes that named the first one only as a collected Minor. The prior row's rung (a positive arm per assert) holds per site and does not travel to a new suite; no deterministic rung decides "could this check have failed" across suites, so the recurrence count is the signal | bot | minor | 4 test | fecb397: the control now passes only when the marker is found in the mutant's store; the vacuous case is removed. A new skip-reason test was checked against the pre-fix code and fails there ("skip reasons: confirmed") |
Loading