feat(t27b): corpus retries a timed-out file once, sequentially (Closes #6310) - #6311
Merged
Merged
Conversation
Closes #6310 Refs #6063 On the Railway lab (--jobs 24, qemu, 60 s) fast files timed out under contention: bilstm.t27 (12 ms natively) in run e2fb1d8 and ops.railway-cli.t27 (3.6 ms natively, reference passes) in run 0851055. t27b corpus now gives each timed-out file one more run, alone, after the parallel pass, with the same timeout; the retry's verdict is final, so genuine infinite loops stay timeouts. Records carry "retried_after_timeout": true, totals carry timeout_retried, and the lab summary counts them. Local corpus totals unchanged except the new field (mismatch 0). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Contributor
6 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #6310
Refs #6063
Why
The Railway lab runs
qemu-aarch64 t27b corpus specs --json --runner qemu-aarch64 --timeout-ms 60000 --jobs 24. Two consecutive runs each lost a fast file to a timeout:specs/ml/recurrent/bilstm.t27timeout (rejected natively in 12 ms; blocked again the next run)specs/trinity/capabilities/ops.railway-cli.t27timeout while the reference passes (t27b passes natively in 3.6 ms)What
blockers::retry_timeouts_once: each item that timed out is re-run once, in order, one at a time; the retry's verdict replaces it; returns which were retried.t27b corpuscalls it after the parallel pass with the same timeout and runner. Per-file JSON gets"retried_after_timeout": true(only on retried records), totals gettimeout_retried, text output printsRETRIEDlines and atimed out, retried oncerow. All other totals count the retry's verdict; pass/fail/mismatch semantics otherwise unchanged.contrib/railway/t27b-lab/lab.py: summary gainstimeout_retried(read from the per-file records). The lab copiestotalsand records through unchanged, andtri t27breads named keys only, so the new fields break nothing.lower.rsuntouched.Evidence
cargo test -p t27b(target /tmp/t27b-retry-target): all green, incl. newa_timeout_is_retried_exactly_once_and_the_retry_decides(times-out-once -> retry verdict; always-times-out -> stays timeout; never retried twice; non-timeouts never re-run, negative control panics if re-run).t27b corpus specs --jobs 6(load ~7.6), master binary vs this branch:timeout_retried 3(axi4_tb, clock_domain_tb, gf16_accel_tb: genuine loops, still timeout). No per-file record changed. mismatch 0.retried_after_timeout: true; master binary on the same runner -> 2 timeouts.python3 scripts/ci/test_the_t27b_lab_heals_its_clone.py: PASS.Cost on the lab: the 3 genuine loops now take one extra 60 s each, sequentially (~3 min per run).
🤖 Generated with Claude Code