Skip to content

feat(tri pr ready): --why compares what a pre-existing failure printed - #5853

Merged
gHashTag merged 3 commits into
masterfrom
feat/tri-pr-ready-why
Oct 4, 2026
Merged

gHashTag merged 3 commits into
masterfrom
feat/tri-pr-ready-why

Conversation

@gHashTag

@gHashTag gHashTag commented Oct 4, 2026

Copy link
Copy Markdown
Owner

What

tri pr ready --why: a failure called pre-existing because a check of the same name is red on master (or a merged PR, or its workflow's newest master run) is now also compared by what its failing step printed. A different step, or a line here whose shape is not printed there, is NEW REASON -- its own verdict, exit 7.

Closes #5852
Refs #5837 #5839 (stacked on #5839: base is feat/tri-pr-ready-workflow-baseline)

Why

The Corpus ratchet is red on master for four conflicted type names. A pull request that adds a fifth printed the same four ##[error] lines and read as "pre-existing". The difference is in the step's output above the annotations.

How

  • Jobs API -> the first failing step of this PR's job and of the job the baseline used (the master walk's check-run, the workflow fallback's run, or the first merged PR failing that name). Logs API -> that step's output, ##[endgroup] to the first ##[error].
  • Masked: timestamps, colour codes, durations, shas (7-40 hex with a digit and a letter), ids of 9+ digits.
  • Each side's last 60 lines are looked up in the other side's whole output (a dropped line earlier does not shift the window). A line not found as itself is looked up by shape (digits read as #) and printed as ~ here / there, not judged: a count moves when the cause does not.
  • Precedence: WAIT 2 > CANNOT TELL 3 > DO NOT 1 > NEW REASON 7 > safe 0. --merge refuses on 7. A comparison that cannot run (not Actions, log expired) is listed as not established and leaves the verdict alone.

Measured (real runs, 2026-10-04, against the committed binary or the same code)

tri pr ready 5781 --repo gHashTag/t27 --why:

  Corpus ratchet (expected-failure ledger)
      also failing in 5 other place(s) — pre-existing
      why, here:  step `A type name may not gain a second definition`, job 111256330898, 5 output line(s)
      why, there: step `A type name may not gain a second definition`, master at 6e3322918, job 111324093808, 8 output line(s)
      NEW REASON -- the same name, failing differently:
        +     + ModuleInterface  NEW conflict
        4 line(s) there are not printed here
        ~ here:  ledger 77 name(s), observed 78
          there: ledger 77 name(s), observed 81
...
VERDICT: NEW REASON -- 1 failure(s) are red elsewhere too, but not for the same reason:
  - Corpus ratchet (expected-failure ledger)

exit 7.

#5663 (one of master's four names, observed 78): SAME REASON, fewer -- its last 5 line(s) are printed there too, the 78/81 line printed under (1 of them with other numbers -- read, not judged:), exit 0. The first draft compared counts as words and would have called this NEW -- found by scanning the 24 open PRs with a failing ratchet before trusting the rule.

#5812: NEW REASON, the failing step differs: Run the corpus ratchet here, exit 7. #5849, #5839, #5828, #5824: every comparison SAME (master walk, merged-PR and workflow-fallback paths each exercised on #5839). About 85 s per PR.

Tests

  • 6 new in prcheck::why_tests, on a fixture shaped like job 111329604306's log; the verdict test now calls the real verdict_code. cargo test -p tri prcheck: 59 passed.
  • 16 mutations, each turns the suite red: no timestamp strip, no colour strip, no duration / sha / id mask, sha mask without the digit rule, script echo kept, annotations kept, tail compared to tail, step ignored, fewer read as new, 7 above 1, 7 never set, no dedup, numbers judged as words, renumbered counted as new.
  • tri census pin --gate: fetches 72 -> 74 lines, 31 -> 33 fetch sites, --paginate 9 -> 11 (failing_jobs_on, failing_job_in_run), blessed in this commit. tri gates tests --gate and tri now check pass.

Not established

  • That the same text is the same cause; a NEW REASON is two outputs to read, not proof this change caused it.
  • A step whose whole reason is a count: two runs differing only in a number read SAME (the numbers are printed for a person).
  • Output after the first ##[error]; a reason only in an artifact or summary; logs past retention (read as cannot compare).
  • The full cargo test -p tri was not run locally (its fpga Lean tests clone mathlib4); CI runs it.

🤖 Generated with Claude Code

gHashTag and others added 2 commits October 4, 2026 06:05
… missed it

`tri pr ready 5826` said CANNOT TELL (exit 3) on fpga-conformance: "did not
run on any recent master commit". It had: master's newest run of FPGA E2E
Build at e7ed379 failed that job, 21 commits back; the walk reads 15.

Only for a failure neither the walk nor the merged-PR baseline observed:
details_url -> run -> workflow -> its newest completed default-branch runs
(page 10, page-fill guarded), read one at a time until a job of the same
name reached success/failure/timed_out. Cancelled and skipped are passed
over. Red there is pre-existing, green there is new here, nothing there
stays NO BASELINE and says how far it looked. An API error leaves the
check without a baseline: CANNOT TELL, never safe. The walk is unchanged.

Real run: tri pr ready 5826 -> exit 0, citing e7ed379 and its age.
8 new tests (53 in prcheck), 7 mutations each red.

Census: fetches moved 69 -> 72 lines, fetch sites 28 -> 31 (--paginate
7 -> 9, page-fill guarded 10 -> 11) -- the three new reads, every one
complete or guarded; unguarded buckets unchanged. Blessed here.

Closes #5837
Refs #5786

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A check red on master under the same NAME was called pre-existing,
whatever its step printed. A pull request that adds a fifth conflicted
type name to the Corpus ratchet read exactly like one that adds nothing.

--why (off by default) fetches the failing step's own output from both
jobs -- this pull request's and the one the baseline used -- masks
timestamps, colour codes, durations, shas and long ids, and compares
each side's last 60 lines against the other side's whole output. A line
not found as itself is looked for by its shape (digits read as #): it is
printed for a person, not judged, because a count moves when the cause
does not (#5663: observed 78 vs 81, a subset of master's names).

NEW REASON is exit 7; precedence 2 > 3 > 1 > 7 > 0; --merge refuses.
Live: #5781 NEW REASON (+ ModuleInterface), exit 7; #5663 SAME REASON,
fewer, exit 0; #5812 NEW REASON (another step), exit 7.

Tests: 6 new in prcheck::why_tests + the verdict test now calls the
real verdict_code (59 in prcheck). 16 mutations, each red.

Census: fetches 72 -> 74 lines, 31 -> 33 fetch sites, --paginate
9 -> 11 (failing_jobs_on, failing_job_in_run). Blessed here.

Closes #5852
Refs #5837 #5839

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Oct 4, 2026

Copy link
Copy Markdown
Contributor

📓 NotebookLM Notebook linked to this PR

This notebook contains session context, decisions, and artifacts for this work.

This was referenced Oct 4, 2026
#5839 (this branch's base) was squash-merged. This merge takes master as it
is and re-applies only this branch's own change over 37836b4, so the PR
diff against master is exactly the --why work. tools/census re-blessed with
`tri census pin --bless`: fetches 72 -> 74 sites, 31 -> 33 fetch sites,
--paginate 9 -> 11, the same numbers this PR pinned over its base.
`tri census pin --gate` PASS; `cargo test -p tri prcheck` 59 passed.

Refs #5852

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@gHashTag
gHashTag changed the base branch from feat/tri-pr-ready-workflow-baseline to master October 4, 2026 08:22
@github-actions

github-actions Bot commented Oct 4, 2026

Copy link
Copy Markdown
Contributor

📓 NotebookLM Notebook linked to this PR

This notebook contains session context, decisions, and artifacts for this work.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

tri pr ready: a red check's name is not its reason -- compare why it fails (--why)

1 participant