Skip to content

fix(t27b-lab): LAB-IMAGE-STALE, 600 s poll, reference cache keyed by use closure + t27c + zig (Closes #6443) - #6535

Merged
gHashTag merged 1 commit into
masterfrom
claude/t27b-lab-image-sha
Oct 5, 2026
Merged

gHashTag merged 1 commit into
masterfrom
claude/t27b-lab-image-sha

Conversation

@gHashTag

@gHashTag gHashTag commented Oct 5, 2026

Copy link
Copy Markdown
Owner

Closes #6443
Refs #6063

What

  1. The lab says which image it is.
    • lab.py publishes image in status.json, in the run JSON (image) and in the run log. image holds the git blob sha of the running lab.py and Dockerfile, the image build time (written by the Dockerfile) and the Railway deployment id.
    • tri t27b doctor gets a new code, LAB-IMAGE-STALE. It fires when the running lab does not report an image (an image from before this change), or when the reported sha differs from master's contrib/railway/t27b-lab.
    • The decision is image_code in specs/tri/t27b/steward.t27, with 4 tests. Its C is generated with t27c gen-c.
  2. The poll defaults to 600 s, down from 1800 s.
    • An idle poll costs one git ls-remote, and a run starts only when the master sha has moved. The unchanged-sha path was already there, so a shorter interval adds no idle work.
    • T27_POLL_SECONDS still overrides the default.
  3. Reference cache, keyed correctly.
    • The key covers the spec's bytes plus the bytes of its transitive use closure (t27c splices imported declarations in), the t27c binary's content and the zig version.
    • A timeout is never cached. Neither is a t27c killed by a signal or one that could not start.
    • lab.py: the cache is persistent at /srv/refcache.json on the volume. Until now every run re-ran all ~720 reference specs, which took 320 s on the last run. The cache also carries t27b: per-test reference differential -- mismatch 0 today compares the JIT with an interpreter over the same IR #6441's per-test verdicts. A pass or fail cached without them is a miss when they are wanted.
    • cli/t27b/src/blockers.rs RefRunner: the same key (toolchain_stamp, use_closure). Timeouts are no longer cached, and old timeout rows are dropped on load.
    • Run steps now report steps.reference.cache = {hits, misses, not_cached}.

Owner exception

On 2026-10-05 at about 16:00Z the owner approved hand-written code in languages other than t27 for this fix ("fix it and do not stop, do what is best"). The lab and the doctor are still Python, and the reference runner is Rust. This adds to the debt that #6198 (port the t27b tooling to t27) has to repay. Only the decision moved into t27 (image_code).

file kind
specs/tri/t27b/steward.t27 t27
gen/c/tri/t27b/steward.c generated (t27c gen-c, byte-reproducible)
contrib/railway/t27b-lab/lab.py hand-written Python (owner exception)
contrib/railway/t27b-lab/Dockerfile hand-written Dockerfile (owner exception)
scripts/tri_loop/t27b.py, scripts/tri_loop/t27b_rules.py hand-written Python (owner exception)
cli/t27b/src/blockers.rs, cli/t27b/tests/refcache.rs hand-written Rust (owner exception)
scripts/ci/test_a_t27b_tick_reads_before_it_acts.py, scripts/ci/test_the_t27b_lab_heals_its_clone.py hand-written Python tests (owner exception)

Verified locally

  • Generated C: cc -DT27_TEST_MAIN gen/c/tri/t27b/steward.c reports All 124 tests passed. The steward test is count-agnostic since feat(t27b): per-test reference verdicts, reference_disagree apart from jit/interp mismatch #6530.

  • test_a_t27b_tick_reads_before_it_acts.py: PASS. Its image fixtures are current, unreported, differs, redeployed and no-master. A mutation control on (same == false) is caught.

  • test_the_t27b_lab_heals_its_clone.py: PASS. The refcache checks are:

    • a cold run;
    • a pass reused while a timeout is re-run;
    • a change to an imported spec misses;
    • a rebuilt t27c misses;
    • per-test verdicts are round-tripped.

    Mutation controls (dropping the closure, caching timeouts) each fail it.

  • cargo test --release -p t27b --test refcache --test blockers: 10 + 7 passed.

  • Doctor against the live lab: before this change it reported LAB-IMAGE-STALE the lab does not report its image (deployed before #6443).

After the merge, the lab is redeployed from master with railway up. It has no git source, so a merge does not redeploy it (S37).

🤖 Generated with Claude Code

…he keyed by use closure + t27c + zig, timeouts never cached

Closes #6443
Refs #6063

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

PR Dashboard

Generated at: 2026-10-05 16:18:12 UTC

Summary

Status Count
Total Open PRs 50
PRs with Failing Checks 40
PRs with All Checks Green 10
READY 5
FAILING 40
PENDING 0
NO CHECKS YET 1

These columns do not partition: 5 + 40 + 0 + 1 = 46, and there are 50 open PRs. A PR is being counted twice or not at all.

Seal Status

  • ⚠️ STALE -- sha256(compiler.rs)=9f2c8a4829f6 != manifest seal=87e5cbd3ad94.
    The committed NMSE numbers were certified against an older compiler.rs.
    Run scripts/reseal-check.sh locally for the two-step reseal command (advisory; not a merge gate).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

owner-approved-foreign Owner-approved exception to the only-t27 rule: hand-written foreign code allowed in this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

t27b lab: stale deployed image (no ratchet step), 1800 s poll, reference cache misses imports and caches timeouts

1 participant