Skip to content

Apply corrections, typeset the report, check the ledgers in CI - #3

Merged
lvelho merged 11 commits into
mainfrom
feat/typeset-report
Sep 9, 2026
Merged

lvelho merged 11 commits into
mainfrom
feat/typeset-report

Conversation

@lvelho

@lvelho lvelho commented Sep 8, 2026

Copy link
Copy Markdown
Contributor

Four parts specified; three done, one blocked. Four commits.

⚠ Part B could not be done

The specification reads "Insert the text the maintainer supplies" followed by the literal line [PASTE THE SECTION TEXT HERE]. The text was never pasted. Nothing was inserted and nothing was guessed at. The insertion point is untouched and ready.

Its checkable claim was checked anyway, and it fails as worded. Part B was to assert that no specification text survives anywhere in bio-3d-vision at 5f8e39f. Three fragments do:

location what survives
exp007/preregistration.md:38 quotes it directly — the specification says "six arms, one stimulus"
exp009/preregistration.md:19 quotes a clause, then argues it is wrong
exp007/findings.md:46 "which neither exp005 nor this specification stated"

The weaker claim is true and is the one to make: no specification was ever committed.

The pattern sharpens the argument rather than weakening it. Every surviving fragment is text an artifact was arguing with. Specification text reaches the record only where a preregistration or findings file had to disagree with it to justify what it did instead — that is, only where the specification was executed. A rejected specification produces no artifact, so nothing is ever in a position to quote it. The survival mechanism is parasitic on the work happening, which is exactly why it cannot preserve a rejection. gap-004 carried the same overstatement and is corrected in place.

Part A — corrections

All 13 proposed replacements applied across 20 passages (five claims recur). Count matches the tally: 5 CORRECTED + 8 UNDER-QUALIFIED.

Running the fixed check after applying found one the record had not named: claim 5's figure recurs a third time in §6 — "here are the five open questions", sixty lines below the cited sentence. The check has now earned its place twice, both times on the same failure: a figure corrected where it was noticed and left where it was not.

Judgement calls: "five working days" applied to the four elapsed-time sentences, not the two register descriptors the entry says are right; claims 31 and 41 offered alternatives and the first was taken in both.

The six UNSUPPORTED claims are unchanged, and the draft now says so in the document rather than only in the ledger.

gap-003 closed — but the closure is the re-read, not the pin: claim 2 is a verbatim quotation cited to active-stereo CLAUDE.md:141, and the line was read again at 3f7a263 and still says what the claim says.

spec-defects.md gains two entries: the specified one (entry 4), and entry 5 — this specification's own unfilled placeholder, under the file's rule that a defect not recorded is paid for twice. The file now separates them as a second family: defects of assembly, where each clause reads correctly alone and the error is a relation between two, so every check in the family must be mechanical.

Part C — typeset · 22 pages, zero overfull or underfull boxes

Nine files, one per section plus appendix. Converted by a re-runnable script, not by hand.

The transit is lossless and checked rather than asserted: 11,628 words in, 11,628 out, zero differing blocks, similarity 1.0000.

  • No bibliography — the draft cites no published work (checked: no author name, no "et al.", no year-in-parentheses). Citations I judge would be wanted are named in refs.bib as comments — four @software entries at the SHAs now in the ledger — and deliberately not written in.
  • No abstract — writing one would mean composing prose that has never been through a claims pass, in the passage the template names as a report's most dangerous.
  • What read worse in LaTeX: one monospace path. \texttt{report/claims-verified.md} cannot hyphenate, 17.6pt over; Markdown does not justify, so it reads fine there. \allowbreak was tried and did not take — recorded, because the obvious fix failing is the useful half. \sloppy scoped to that block does.
  • CI needs no new packages.

Part D — od-004 closed, with no new dependency

A ledgers job parses both ledgers with ruby -ryaml. Ruby, not Python, is the whole reason this was open: YAML is not in Python's stdlib and PyYAML is not guaranteed on the runner, so the obvious check needed a pip install. Ruby ships Psych in its stdlib and is preinstalled. The smallest dependency that works is none.

It was watched failing before it was committed — and the command was extracted from the workflow file and run, not hand-typed, so the YAML quoting was part of what got tested.

Falsifiers

  1. PASSES. Clean git clone → make -C report → 22 pages, exit 0, no manual install. CI compiles it.
  2. PASSES. No figure stated two ways. Worth noting: "thirteen" now carries three referents that are all genuinely 13 (experiments, foreclosures, and the 13 applied corrections) — a collision, not an inconsistency. One false positive: the check reads LaTeX comments, where "twenty-two pages" appears.
  3. FAILS — reported above, in detail, without the section text.
  4. PASSES. 13 of 13 applied; count matches.

🤖 Generated with Claude Code

lvelho and others added 6 commits September 8, 2026 19:59
…efects

PART A of four.

A1/A2. All thirteen proposed replacements applied — every CORRECTED and every
UNDER-QUALIFIED claim — across TWENTY passages, because five of the claims
recur. The count matches claims-verified.md's tally: 5 corrected + 8
under-qualified = 13.

A1's :397 was the third occurrence of the eighteen-claims figure, left standing
when :874 and :1228 were amended. It is now in line. That inconsistency lived in
the paragraph introducing the repeated-claim check, which is the check that
found it.

RUNNING THE FIXED CHECK AFTER APPLYING FOUND ONE THE RECORD HAD NOT NAMED:
claim 5's figure recurs a third time in §6 — "here are the five open questions",
sixty lines below the sentence the entry cited. Brought into line. The check
earned its place twice now, and both times on the same failure: a figure
corrected where it was noticed and left where it was not.

Judgement calls, both recorded in claims-verified.md:
- Claim 1: "five working days" applied to the four elapsed-time sentences, NOT
  to the two register descriptors (§1, §7), which the entry says are right.
- Claims 31 and 41 offered alternatives; the first was taken in both. For 41 the
  second option requires editing the template repository.

THE SIX UNSUPPORTED CLAIMS ARE UNCHANGED. No replacement can be written from the
record, and the draft's status block now says so in the document rather than only
in the ledger.

A3. gap-003 closed. active-stereo pinned at 3f7a263 and bioeye at e908170. The
closure is the RE-READ, not the pin: claim 2 is a verbatim quotation cited to
active-stereo CLAUDE.md:141, and the line was read again at the pinned SHA and
still says what the claim says.

A4. docs/spec-defects.md gains two entries, not one. Entry 4 is the specified
defect: a setup step enabling "require a pull request before merging" against a
later step in the same instruction saying "commit direct to main". It cost a
round trip and left a permanently false sentence in a commit message, since
correcting it would need a force-push.

Entry 5 is observed in the specification for THIS task, under the file's own
rule that a defect not recorded is paid for twice: Part B reads "[PASTE THE
SECTION TEXT HERE]" and the text was never pasted. Part B is not done and is
reported rather than guessed at.

The two are one family with entry 3, and the file now says so: defects of
ASSEMBLY rather than of finish conditions, where each clause reads correctly in
isolation and the error is a relation between two of them. Re-reading cannot
catch them; every check in the family is mechanical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Part B could not be done: the specification reads "Insert the text the
maintainer supplies" followed by the literal line [PASTE THE SECTION TEXT HERE],
and the text was never pasted. Nothing was inserted and nothing was guessed at.
Recorded as docs/spec-defects.md entry 5 in the previous commit.

The one thing Part B could be checked on without its text was checked, since the
specification named it as checkable and said to stop and report if it failed.

IT FAILS AS WORDED. Part B was to assert that no specification text survives
anywhere in bio-3d-vision at 5f8e39f. Three fragments do:

  experiments/exp007_rendered_policy_sweep/preregistration.md:38
    quotes it directly -- the specification says "six arms, one stimulus"
  experiments/exp009_epipolar_cost/preregistration.md:19
    quotes a clause -- "the specification says to compose
    rectification_rotation into eye_camera_poses" -- and then argues it is wrong
  experiments/exp007_rendered_policy_sweep/findings.md:46
    records a reading "which neither exp005 nor this specification stated"

The weaker claim is true and is the one to make: no specification was ever
COMMITTED. No file in the instance is a specification, and no filename matches
spec, task, brief, instruction or prompt.

THE PATTERN SHARPENS THE ARGUMENT RATHER THAN WEAKENING IT, which is why this is
worth the commit. Every surviving fragment is text an artifact was ARGUING WITH.
Specification text reaches the record only where a preregistration or findings
file had to disagree with it to justify what it did instead -- that is, only
where the specification was EXECUTED and produced an artifact. A specification
rejected before work begins produces no artifact, so nothing is ever in a
position to quote it. The survival mechanism is parasitic on the work happening,
which is exactly why it cannot preserve a rejection.

gap-004 carried the overstatement too and is corrected in place, with the
original left standing beside it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
PART C of four. 22 pages, zero overfull or underfull boxes.

Nine files, one per numbered section plus the appendix, wired into main.tex.
Converted by a script rather than by hand or by paste, committed at
scratchpad/md2tex.py's shape and re-runnable: the draft stays the prose of
record and the .tex is its typeset form, so the diff between them stays the
place a mangled figure shows.

THE TRANSIT IS LOSSLESS, and checked rather than asserted: 11,628 words in,
11,628 out, zero differing blocks under difflib, similarity 1.0000. Em dashes,
en dashes, section marks and curly quotes all arrive as ---, --, \S and ``''.
Code spans become \texttt, including where bold wraps them.

NO BIBLIOGRAPHY, and it is a finding rather than an omission. The draft cites no
published work -- checked, not assumed: no author name, no "et al.", no
year-in-parentheses anywhere in it. biblatex is commented out and refs.bib holds
its rule and no entries.

CITATIONS I JUDGE WOULD BE WANTED, named as the specification asked and NOT
written in: four @software entries, one per repository in Appendix A, each at
the SHA now recorded in the ledger -- bio-3d-vision 5f8e39f, math-ai-method
81a638e, active-stereo 3f7a263, bioeye e908170. They are listed in refs.bib as
comments. Adding a bibliography the document does not reference would be the
same defect as a check that cannot fail.

NO ABSTRACT. The draft has none. Writing one would mean composing prose that has
never been through a claims pass, in the passage the template names as the most
dangerous in a report. The draft's own status note is typeset in its place.

THE ONE THING THAT READ WORSE IN LATEX THAN IN MARKDOWN was a monospace path.
\texttt{report/claims-verified.md} cannot hyphenate and has no break point TeX
will take; it came out 17.6pt over. Markdown does not justify, so it reads fine
there. \allowbreak after the slash was tried and did NOT take -- recorded,
because the obvious fix failing is the useful half. \sloppy scoped to that one
block does. Scoped, not global: one box in twenty-two pages does not justify
loosening how the whole document is set.

CI needs no new packages: the document loads article, geometry, fontenc,
lmodern, graphicx, booktabs and hyperref, all already in report.yml's install
list.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
PART D of four.

A `ledgers` job in report.yml loads docs/state.yaml and
docs/inherited-measurements.yaml and fails if either does not parse or does not
come back a mapping. Two files, one command, no install step.

RUBY, NOT PYTHON, AND THAT IS THE WHOLE REASON THIS WAS OPEN. YAML is not in
Python's standard library and PyYAML is not guaranteed on the runner, so the
obvious check would have needed a `pip install` -- the dependency question
od-004 was waiting on an answer to. Ruby ships Psych in its standard library and
is preinstalled on ubuntu-latest. The smallest dependency that works is none.

`YAML.load` rather than `safe_load_file`, which is Ruby 3.x only. The ledgers
contain no anchors or aliases -- checked -- so the portable form is both safe
and version-independent.

A SEPARATE JOB, not a step in `build`: a one-character YAML error should not be
reported from behind a 44-second TeX install, and the record should still be
checked when the document fails to compile.

IT WAS WATCHED FAILING BEFORE IT WAS COMMITTED. An unclosed flow sequence
appended to each ledger in turn produced exit 1 naming the file, line and
column. The command was EXTRACTED FROM THE WORKFLOW FILE and run, not
hand-typed, so what was tested is what CI will run -- the quoting inside the
YAML block is part of what could break. A check nobody has seen fail is a check
nobody knows can fail.

Scope recorded on the entry: it tests that the files parse. It does not test
that ids are unique, that a superseded entry names its successor, or that a
cited id resolves.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The check committed in the previous commit FAILED ON ITS FIRST CI RUN, with
"Tried to load unspecified class: Date". It had passed locally.

CAUSE. `YAML.load` is unsafe-by-default in Psych 3 (Ruby 2.6, this machine) and
SAFE-by-default in Psych 4 (Ruby 3.x, ubuntu-latest), where it refuses to
instantiate Date. Both ledgers carry unquoted dates like 2026-09-08, which YAML
resolves to a Date.

FIX. Not permitting the class -- `YAML.parse`, which returns the node tree and
constructs nothing. It behaves identically on every Psych version, and it is a
more exact test of what od-004 asks: not "does this load into Ruby objects" but
"is this parseable". The mapping-root assertion moves to Psych::Nodes::Mapping.

Re-verified on three axes, not one: good ledgers pass; a broken one exits 1
naming file and line; a non-mapping root is rejected; an unquoted date parses
without loading a class.

WHY THIS IS RECORDED ON od-004 RATHER THAN QUIETLY FIXED. The previous commit
claimed "it was watched failing before it was committed", and that claim was
true and insufficient -- it was watched failing on the one axis its author
thought of, and never run anywhere but the author's machine. That is the exact
shape .gitignore already warns about in this repository: something that passes
locally and fails everywhere else because the author's environment is not the
runner's. "I watched it fail" is not the same claim as "I watched it fail where
it runs."

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Inserted verbatim between "The devil's advocate was agreed and never written
down" and "And now the one that matters". Not edited.

ITS FIVE CHECKABLE CLAIMS WERE VERIFIED BEFORE INSERTION, all confirmed, and
recorded as A1-A5 in report/claims-verified.md. New prose making factual claims
is what that document exists to catch, and it does not get an exemption for
being new.

THE FLAGGED STOP CONDITION DID NOT FIRE. The specification warned that the
section asserts no specification text survives anywhere in bio-3d-vision at
5f8e39f, and said to stop if that failed. IT WOULD HAVE FAILED -- three
fragments survive, quoted inside artifacts that were being committed anyway.
The supplied text does not say that. It says the specification "is the only
artifact in the loop that is never committed", which is the true and weaker
claim, and which re-checks clean: zero tracked files are specifications and no
filename matches spec, task, brief, instruction, relay or prompt. The
distinction is load-bearing and the section gets it right. Recorded as A3 so
nobody re-derives the stronger version later.

A5 -- "five of the six unsupported claims" -- checks out under the selector the
sentence states, and DISAGREES WITH gap-004's "four", which used a stricter one.
Both are right. Recorded as two counts with their selectors rather than resolved
by picking one, which is exactly what mn-001 and mn-002 required and for the
same reason. The odd one out under either selector is claim 38: not an event but
a proportion of a document that still exists, so it is unsupported and
MEASURABLE -- the only one of the six a successor could close by going to look.

Typeset with the rest: sections regenerated, transit re-verified lossless
(12,184 words in, 12,184 out, zero differing blocks), 22 pages, zero overfull or
underfull boxes. The section lands as §6.3 and does not disturb §6's opening
claim that "the last one is the serious one" -- "And now the one that matters"
is still last.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@lvelho

lvelho commented Sep 8, 2026

Copy link
Copy Markdown
Contributor Author

Part B is now in — inserted verbatim as §6.3, between the devil's-advocate subsection and "And now the one that matters". PR now six commits; CI green on both jobs.

The flagged stop condition did not fire — but it would have

The specification warned that the section asserts no specification text survives anywhere in bio-3d-vision at 5f8e39f, and said to stop if that failed. It would have failed — three fragments survive, quoted inside artifacts that were being committed anyway.

The supplied text does not say that. It says the specification "is the only artifact in the loop that is never committed" — the true and weaker claim, which re-checks clean: zero tracked files are specifications, and no filename matches spec, task, brief, instruction, relay or prompt. The distinction is load-bearing and the section gets it right. Recorded as A3 so nobody re-derives the stronger version later.

Its five checkable claims were verified before insertion

All CONFIRMED, recorded as A1–A5 in claims-verified.md. New prose making factual claims doesn't get an exemption for being new. mn-004's tally is unchanged — these are numbered separately.

A5 disagrees with my own ledger, and both are right

The section says "five of the six unsupported claims"; gap-004 said four. Different selectors:

  • Four — unrecordable by construction, the workflow guarantees no artifact (17 and 20 live in specifications; 12 is a rate over conversational turns; 25 is a cleared session's search).
  • Five — "describes an event the workflow guarantees will leave no trace", the section's own selector, which also admits 45: the afternoons a foreclosure cost were events, and nothing logs duration.

Recorded as two counts with their selectors rather than resolved by picking one — which is exactly what mn-001/mn-002 required, for the same reason.

Claim 38 is the odd one out under either selector: not an event, but a proportion of a document that still exists. So it is unsupported and measurable — the only one of the six a successor could close by going to look.

Unchanged

22 pages, zero overfull/underfull boxes. Transit re-verified lossless: 12,184 words in, 12,184 out, zero differing blocks. The section lands as §6.3 and does not disturb §6's opening claim that "the last one is the serious one" — And now the one that matters is still last.

Section replaced in report/draft.md and regenerated into
report/sections/06-where-it-failed.tex, so both stay committed and the
translation stays diffable. 23 pages, zero overfull or underfull boxes, transit
re-verified lossless: 12,292 words in, 12,292 out, zero differing blocks.

VERIFIED AT SOURCE: seven checkable claims, A1-A7 in claims-verified.md. Five
CONFIRMED, one CORRECTED, one UNDER-QUALIFIED. The two that did not pass:

A6 -- "THREE passages in the instance quote specification text" is FOUR, and one
of the three named is not one of them. exp007/findings.md:46 records what the
specification did NOT state, which is an absence, not a quotation. Meanwhile
exp007/findings.md:123 -- "The specification said 18 steps for A, A', D and E",
under a heading reading "A deviation from the specification, and why" -- is the
clearest instance of the section's own thesis and is not among the three.

A7 -- "Every one is text that an artifact was arguing with" is UNDER-QUALIFIED.
Three of the four depart from the clause they quote; exp007/findings.md:32
AGREES with it: "The specification asked for this to be visible rather than
inferred, and it is."

THE CONCLUSION IS UNAFFECTED, WHICH IS WHY THE CORRECTION IS CHEAP. The
load-bearing condition is the next sentence -- the survival mechanism is
parasitic on the work happening -- and it holds for all four. Agreement or
disagreement, the fragment exists because the specification was EXECUTED and
produced an artifact. A rejected specification produces no artifact either way.
"Arguing with" is too narrow a mechanism for a conclusion that only needs
"executed". Both proposed replacements are in the addendum; the prose is
unedited, per the standing practice.

A6 IS MY ERROR PROPAGATING, and it is named as such. The search behind the
previous round returned FOUR hits and my report said THREE; findings.md:123 was
in the output and never opened, and the section was written from that summary.

docs/spec-defects.md gains entry 6: a verification search written from the
expected finding rather than from the claim. It records BOTH instances -- Chat's
grep for REPOSITORY STATE ASSUMED, which could only have found a committed
specification and never a quoted one, and mine, which was better aimed and still
read from the expectation rather than from the output. Two stations, same error,
three iterations apart. The check is two clauses: search for what would falsify
the claim, and open every hit before reporting a count.

AND THE od-003 CHECK ITSELF HAD THE SAME DEFECT. It was case-sensitive: it saw
"three passages" and missed "Three passages". Measured over the typeset sources,
252 hits against 301 -- missing 49, or 16%, every sentence-initial figure. Fixed
with -i and a tolower in the grouping, in the documented command and on od-003.
It had been run three times and reported as passing. Found by hand, not by the
check.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@lvelho

lvelho commented Sep 9, 2026

Copy link
Copy Markdown
Contributor Author

Revised Part B is in — replaced in draft.md and regenerated into 06-where-it-failed.tex, so both stay committed and the translation stays diffable. 23 pages, zero overfull/underfull boxes, CI green on both jobs.

Falsifiers

  1. PASSES. Clean git clone → make -C report → 23 pages, exit 0, no manual install. CI compiles it.
  2. PASSES. No figure stated two ways. "Thirteen" carries three referents (experiments, foreclosures, applied corrections), "eighteen" two (claims, plan steps), "three" two (failures that arrived late, rejections) — all genuinely those values, all contextually distinguished.
  3. The rewritten claim holds. No specification was ever committed re-checks clean: zero tracked files are specifications, no filename matches spec/task/brief/instruction/relay/prompt. A6 and A7 do not hold — below.
  4. PASSES. 13 of 13 applied earlier; mn-004's tally unchanged, and the new section's verdicts are numbered A1–A7 separately.

Two claims in the new text did not check out

A6 — "Three passages quote specification text" → CORRECTED to four, and one of the three named isn't one of them.

passage reports spec content?
exp007/preregistration.md:38 yes — quotes "six arms, one stimulus"
exp009/preregistration.md:19 yes — quotes, then argues it's wrong
exp007/findings.md:123 yes — and it's the best example, omitted. "The specification said 18 steps for A, A′, D and E", under a heading reading "A deviation from the specification, and why"
exp007/findings.md:32 yes
exp007/findings.md:46 no — records what the specification did not state. An absence, not a quotation.

A7 — "Every one is text that an artifact was arguing with" → UNDER-QUALIFIED. exp007/findings.md:32 agrees: "The specification asked for this to be visible rather than inferred, and it is."

The conclusion is unaffected, which is why this is cheap. The load-bearing sentence is the next one — the survival mechanism is parasitic on the work happening — and it holds for all four. Agreement or disagreement, the fragment exists because the specification was executed. "Arguing with" is too narrow a mechanism for a conclusion that only needs "executed". Both replacements proposed; prose unedited.

A6 is my error propagating

My earlier search returned four hits and my report said three. findings.md:123 was in the output and never opened, and your text was written from that summary.

spec-defects.md §6 records both instances — Chat's grep for REPOSITORY STATE ASSUMED, which could only ever have found a committed specification and never a quoted one, and mine, better-aimed and still read from the expectation rather than the output. Two stations, same error, three iterations apart.

And the od-003 check had the same defect

It was case-sensitive — saw "three passages", missed "Three passages". Measured: 252 hits against 301, missing 49 (16%), and the missed class was every sentence-initial figure. Fixed with -i plus tolower in the grouping (without the second, "Three" and "three" count as two figures). It had been run three times and reported as passing. Found by hand, not by the check — which is §6's defect inside the tool built to catch §6's defect.

lvelho and others added 3 commits September 8, 2026 23:30
Two straight double quotes reached the typeset output, both inside verbatim
quotations: the active-stereo constitution passage in §2 and the findings-file
phrase in §6. LaTeX renders " as two RIGHT-facing marks, so the opening quote of
each pointed the wrong way. Markdown renders both forms identically, which is
why the defect is invisible at the source and visible only after the transit --
which is what the draft/typeset separation exists to surface, so the practice is
now recorded as having fired rather than argued for.

FIXED IN THE CONVERTER, NOT IN THE .TEX. The section files are generated from
report/draft.md; a hand-edit would have been reverted by the next regeneration,
silently, and the document would have been correct exactly until someone re-ran
the conversion. Straight pairs now become `` and '' before output.

AND THE CONVERTER NOW FAILS RATHER THAN WARNS. It refuses to write output
containing a straight double quote, an unconverted curly quote, an em or en dash
or a section mark. Tested by removing one closing quote from the draft: the
conversion exits non-zero and names the file and the count. The first attempt at
that test was worthless -- it produced a BALANCED pair, which the converter
correctly handles -- and was redone with a genuinely unpaired quote.

FULL AUDIT OF report/sections/, reported in claims-verified.md. Straight quotes:
zero. Apostrophes: 49, all correct -- LaTeX renders ' as a right single quote,
which is what an apostrophe is. Five more on section references (\S4.1's).
Opening single quotes, ellipses, stray backticks, unconverted unicode: none. One
straight quote remains in main.tex, inside a LaTeX comment, never rendered --
reported rather than changed, because changing it would suggest it mattered.

CHECKED BECAUSE IT WOULD HAVE BEEN SILENT: the quote conversion also runs over
code-span contents, so a " inside a `...` span would have been turned into
typographic marks inside a literal \texttt. There are none in this draft. That
is a property of this draft, not of the converter.

23 pages, zero boxes, transit re-verified lossless at 12,292 words.

od-007 opened, because this fix exposed a gap it did not create: the converter
is NOT IN THE REPOSITORY. The regeneration a reader is told to perform cannot be
performed by them, and the guard that now protects this class of defect protects
nobody else. The obvious fix reopens fc-002 and fc-003 by putting Python back,
so the three options and their costs are recorded rather than one being taken.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…n 2)

tools/md2tex.py generates report/sections/*.tex from report/draft.md. It was
outside the repository, which made report/main.tex's claim that the .tex is
regenerable true and unverifiable at the same time. od-007 closed.

fc-002 AND fc-003 ARE ANNOTATED, NOT REVERSED, per CLAUDE.md's rule. What came
back is the tooling fc-002's scope said would come back with pyproject.toml:
ruff, ruff format, mypy and pytest, as a `code` job in CI. Taking the code and
leaving the gates would have been taking the convenient half of a decision.

WHAT DID NOT COME BACK. No [build-system], no hatchling, no package -- there is
still no library, and a build backend for a package that does not exist is
scaffolding that gets believed. No dependencies: the converter is standard
library only. src/ and experiments/ stay closed.

THE ORDER fc-003 PRESCRIBED WAS FOLLOWED. Its scope says "restore the
[pinned-clone] guard first, before the others", and that guard is the first test
in tests/test_scaffold.py for that reason and no other. IT FIRED IMMEDIATELY, on
the docstrings of the two files that restored it, which named the directory in
prose. The prose was rewritten and the guard left crude. Loosening a check to
admit its first true positive is how a check stops being one.

THE SECOND GUARD IS WHAT EARNS THE REOPENING.
test_typeset_sections_match_the_draft regenerates the sections into a temporary
directory and compares bytes, so main.tex's claim that the .tex is generated
from the draft can now fail. Both guards were watched failing before being
committed -- one by appending a comment to a .tex, one by adding a real read of
the pinned clones -- and the sync test writes only to a temp directory, so a
failing test can never repair the thing it is checking.

fc-002's SCOPE NAMED THE WRONG TRIGGER, and that is the finding worth keeping.
It said the entry stops holding "the moment one generated figure is wanted". No
figure is wanted; a converter arrived. fc-003 said "the moment any .py file
lands" and covered this exactly. The difference is that fc-002's scope described
the USE it was picturing and fc-003's described the CONDITION. A scope clause
that names an instance will be silent on the case that actually occurs.

COSTS PAID, recorded because they are most of the cost. The script was
scratch-quality -- semicolons, percent formatting -- and had to be rewritten to
pass the gates it was arriving under; its output is verified byte-identical
before and after, and again after ruff format. Five statements elsewhere became
false and are corrected in this commit: CLAUDE.md's Commands section and its
definition-of-done clause 2, .gitignore's note that the guard was gone,
README.md's "no Python", and report/figures/README.md's restore instructions.

Unchanged: 23 pages, zero boxes, and `make -C report` still needs no Python --
the .tex files are committed and `make -C report sections` is a separate step
that never runs during a build.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The `code` job added in the previous commit FAILED ON ITS FIRST CI RUN. It ran
`pip install ".[dev]"`, which asks pip to build this directory as a package.
There is no build backend by design, so setuptools was guessed at and rejected
pyproject.toml: "`project` must contain ['version'] properties".

IT PASSED LOCALLY, BECAUSE THE INSTALL STEP NEVER RAN THERE -- the tools were
already in the environment. That is od-004's failure again, three commits later:
local success meant nothing because the author's machine had already done the
step CI has to do from scratch.

THE FIX IS NOT A VERSION AND A BUILD BACKEND. That would manufacture the package
pyproject.toml exists to avoid. [dependency-groups] (PEP 735) is the table for
development tooling belonging to something that is not a package, and
`pip install --group dev` never builds anything. Versions stay in pyproject.toml
rather than the workflow, so local and runner cannot drift.

THEN RE-RAN THE GATES AGAINST THE VERSIONS CI ACTUALLY INSTALLS -- ruff 0.16.6,
mypy 2.3.1, pytest 9.1.1, all newer than the local ones -- rather than against
the local ones again. All four pass. Checking locally alone would have made the
same mistake a third time.

Recorded on od-007.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@lvelho

lvelho commented Sep 9, 2026

Copy link
Copy Markdown
Contributor Author

Option 2 done — converter committed, tooling back with it. od-007 closed. All three CI jobs green on de575e4; 23 pages, zero boxes; clean checkout still builds with TeX alone.

What landed

tools/md2tex.py generates report/sections/*.tex from report/draft.md. It lived outside the repo, which made main.tex's claim that the .tex is regenerable true and unverifiable at the same time.

fc-002 and fc-003 are annotated, not reversed, per CLAUDE.md. What came back is what fc-002's scope said would: ruff, ruff format, mypy, pytest, as a code job. What did not: no [build-system], no package, no dependencies, and src//experiments/ stay closed.

fc-003's order was followed. Its scope says restore the pinned-clone guard first, and it is the first test for that reason. It fired immediately — on the docstrings of the two files that restored it, which named the directory in prose. I rewrote the prose and left the guard crude. Loosening a check to admit its first true positive is how a check stops being one.

The second guard is what earns the reopening. test_typeset_sections_match_the_draft regenerates into a temp dir and compares bytes, so main.tex's standing claim can now fail. Both guards were watched failing before commit, and the sync test writes only to a temp dir — a failing test can never repair what it checks.

Two findings worth your attention

fc-002's scope named the wrong trigger. It said the entry stops holding "the moment one generated figure is wanted". No figure was wanted; a converter arrived. fc-003 said "the moment any .py file lands" and covered it exactly. fc-002 described the use it was picturing; fc-003 described the condition. A scope clause that names an instance is silent on the case that actually occurs — and looks wrong when it isn't.

The gates failed on their first CI run, and it was od-004's failure again. pip install ".[dev]" asks pip to build the directory as a package; there's no build backend by design, so setuptools rejected pyproject.toml. It passed locally because the install step never ran there — the tools were already present. Same shape as the Psych 3/Psych 4 split three commits earlier: local success meant nothing.

Fixed with PEP 735 [dependency-groups] and pip install --group dev — not by adding a version and a backend, which would manufacture the package the file exists to avoid. Then I re-ran all four gates against the versions CI actually installs (ruff 0.16.6, mypy 2.3.1, pytest 9.1.1 — all newer than local). Checking locally again would have made the same mistake a third time.

Costs paid

The script was scratch-quality (semicolons, % formatting) and had to be rewritten to pass the gates it arrived under — output verified byte-identical before and after, and again after ruff format.

Five statements elsewhere became false and are corrected in the same commit: CLAUDE.md's Commands section and definition-of-done clause 2, .gitignore's note that the guard was gone, README's "no Python", and report/figures/README.md's restore instructions. A decision that invalidates five statements is one whose cost is mostly in the prose.

A6 -- "THREE passages in the instance quote specification text" -> FOUR, with
the list rebuilt. exp007/findings.md:46 is dropped: it records what the
specification did NOT state, which is an absence and not a quotation.
exp007/findings.md:123 and :32 are added, the first being the clearest instance
of the section's own thesis. All four re-read at 5f8e39f before the sentence was
written, and each characterisation checked against its text.

A7 -- "Every one is text that an artifact was arguing with" -> "had to reckon
with -- mostly to depart from it, once to confirm it." exp007/findings.md:32
agrees with its specification rather than arguing: "The specification asked for
this to be visible rather than inferred, and it is."

AND A THIRD SENTENCE, WHICH HAD NO VERDICT OF ITS OWN. The paragraph after A7
restated the same overreach in different words -- "reaches the record only where
a pre-registration or a findings file had to DISAGREE with it in order to
justify doing something else". A7's verdict named one sentence; the defect was
in two, sixty words apart, in the same subsection. Corrected to "had to ACCOUNT
FOR it -- usually to justify departing from it, once to note that it was met",
which keeps the mechanism and drops the false universal.

THE MECHANICAL CHECK COULD NOT HAVE FOUND THAT ONE. It repeats an IDEA, not a
number, and every word in the restatement differs. Recorded in
claims-verified.md as a practice: when a characterisation is corrected, read
forward for its restatements, exactly as a corrected figure is grepped for its
other occurrences.

The section's conclusion is unchanged. The load-bearing condition was always
that the specification had to be EXECUTED and produce an artifact, and that
holds for all four passages; "arguing with" was a narrower mechanism than the
conclusion needed.

ALSO FIXED, AND IT WAS MY SLIP: the draft status block and main.tex still said
the section's "five checkable claims" were "all confirmed, recorded as A1-A5".
That was written for the first version of the section and never updated when the
revision took it to seven claims with two non-confirmed. Both now say A1-A7,
five confirmed, one corrected, one under-qualified, both since closed.

gap-004 already carried the corrected count of four, so the ledger and the prose
now agree. Before this change they did not, and the ledger was right.

23 pages, zero boxes, transit lossless, all gates green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@lvelho

lvelho commented Sep 9, 2026

Copy link
Copy Markdown
Contributor Author

A6 and A7 closed. All three CI jobs green on d8c7052; 23 pages, zero boxes, transit lossless at 12,335 words.

claim was now
A6 "Three passages … quote specification text", listing one that records an absence Four passages, list rebuilt. exp007/findings.md:46 dropped; :123 and :32 added. All four re-read at 5f8e39f and each characterisation checked against its text before the sentence was written.
A7 "Every one is text that an artifact was arguing with" "…had to reckon with — mostly to depart from it, once to confirm it."
(no verdict) "reaches the record only where … had to disagree with it" "had to account for it — usually to justify departing from it, once to note that it was met."

The third row is the finding

A7's verdict named one sentence; the defect was in two — sixty words apart, in the same subsection. The paragraph after A7 restated the same overreach in different words.

The mechanical repeated-figure check could not have caught it. It repeats an idea, not a number, and every word in the restatement differs. The only thing that finds this class is applying a correction and then re-reading the paragraph around it for the same claim in other words. Recorded in claims-verified.md as a practice: when a characterisation is corrected, read forward for its restatements — the same move as grepping a corrected figure for its other occurrences, but nothing automates it.

Your conclusion is untouched. The load-bearing condition was always that the specification had to be executed and produce an artifact, and that holds for all four passages; "arguing with" was a narrower mechanism than the conclusion needed.

One slip of mine, also fixed

The draft status block and main.tex still read "its five checkable claims … all confirmed, recorded as A1–A5". That was written for the first version of the section and never updated when the revision took it to seven claims with two non-confirmed. Both now say A1–A7 — five confirmed, one corrected, one under-qualified, both since closed.

gap-004 already carried the corrected count of four, so ledger and prose now agree. Before this change they didn't, and the ledger was right — which is the third time in this PR that the record caught the prose rather than the other way round.

Where the report stands

Every CORRECTED and UNDER-QUALIFIED claim in the document is now closed — 13 from the main pass, 2 from the addendum. The six UNSUPPORTED claims remain, unchanged and flagged in the prose itself, because no replacement can be written from a record that does not contain them. That is gap-004 and od-006, and it is the subject of §6.3.

@lvelho
lvelho merged commit 295207f into main Sep 9, 2026
3 checks passed
@lvelho
lvelho deleted the feat/typeset-report branch September 9, 2026 03:54
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant