Skip to content

fix(tri): per-pid temp scratch cannot outlive its run -- Drop guard plus dead-pid sweep - #5990

Merged
gHashTag merged 4 commits into
masterfrom
claude/fpga-m2l-scratch-cleanup
Oct 4, 2026
Merged

gHashTag merged 4 commits into
masterfrom
claude/fpga-m2l-scratch-cleanup

Conversation

@gHashTag

@gHashTag gHashTag commented Oct 4, 2026 •

Copy link
Copy Markdown
Owner

Pull Request Checklist

  • PR title follows semantic convention
  • PR body includes Closes #5982
  • docs/now/2026-10-04-tri-per-pid-temp-scratch-cannot-outlive-its-run.md added (tri now add ... --closes 5982)
  • Tests added: 12 new tests, all run locally (the full ./scripts/tri test was NOT run, because of the disk limit below)
  • Specs changed: none

Round 2: what the review of 7bc4843 found, and what changed

The review refused 7bc4843 on deletion safety. pid_is_running read any ps that exits 1 with empty stdout as "no such process", including a ps that exits 1 with an error on stderr. A fake ps first on PATH (echo err >&2; exit 1) made pid 1 read as dead, and sweep_dead removed probe_pkg_1. A ps that rejects -p or -o answers the same way. On such a machine every pid reads as dead, so runs delete each other's live 7.6 GB packages. The old module doc promised the opposite of what the code did.

New commits on the same branch, no force-push:

  • 82a1e8ae5 fix(tri): a pid reads dead only on a silent exit 1 from ps
    • Dead now needs exit status 1 and empty stdout and empty stderr. Everything else reads as running:
      • the probe cannot start;
      • the probe was killed by a signal;
      • any other exit status;
      • exit 1 with even one byte on either stream;
      • no probe at all.
    • The probe is a parameter, as build already is for the lake check:
      • pid_is_running_by(probe, pid);
      • sweep_dead_by(parent, prefix, running);
      • pid_is_running and sweep_dead are those with the real ps.
    • Pid 0 and this process's own pid never ask the probe.
    • The module doc now says exactly this.
    • Item 5, the tests' own parents: tri_piddir_<tag>_<pid> and tri_m2l_scratch_<tag>_<pid> are swept by the same rule before each test makes its own. Two tests pin it.
  • 9e9d4f094 merge origin/master: picks up fix(ci): the orphan-ceiling ledger names cli/t27b #5989 (e84927b60), the fix for the inherited "Every source file is reachable" failure. A clean merge, no conflicts. After the merge, tri census pin --gate printed PASS: no pinned census moved.

What ps answers on this Mac (macOS, /bin/ps, 2026-10-04)

Command Exit stdout stderr Reads as
ps -p 99999 -o pid= (no such pid) 1 0 bytes 0 bytes dead
ps -p 0 -o pid= 1 0 bytes 0 bytes running, because pid 0 never asks the probe
ps -p 4194303 -o pid= 1 0 bytes 34 bytes: ps: process id too large: 4194303 running; the old rule read this as dead
ps -p 1 -o bogus= (a keyword ps rejects) 1 497 bytes (the keyword list) 68 bytes, starting ps: bogus: keyword not found running under both rules: this ps happens to print on stdout too. One that prints its error only on stderr is the review's case
ps -p 1 -o pid= 0 6 bytes ( 1) 0 bytes running

Linux procps was NOT verified. No Linux container was available on this machine: the Docker daemon was not running and colima was not started, and no image was pulled. If procps answers a missing pid in any way other than a silent exit 1, the sweep keeps everything there. That costs disk, never a live run.

Description

test_measured_to_lean_standalone_builds_in_temp_lake_package (in cli/tri/src/fpga.rs) builds a lake package in temp_dir()/tri_m2l_standalone_pkg_<pid>. Its require trinity pulls mathlib into the package's own .lake/packages, about 7.6 GB. The dir was removed only on the line after assert!(status.success()):

  • a failed lake build panicked past that cleanup;
  • a killed run never reached it.

On 2026-10-04 the leftovers filled the owner's disk twice. Free space fell to about 300 MB and other runs died with ENOSPC. Side files from 8 dead pids (24 files) are still in $TMPDIR on that machine.

The fix has two parts, because each covers what the other cannot:

  1. Drop guard (PidPath in the new cli/tri/src/piddir.rs).
    • It removes its file or dir when dropped, which happens on return and during a panic unwind.
    • Every path the test writes is now guarded, and each guard is created before anything can fail.
  2. Dead-pid sweep (sweep_dead(parent, prefix)). kill -9 runs no destructor, so the next run removes <prefix><pid>[.ext] entries whose pid is not running.
    • Liveness comes from ps -p <pid> -o pid=. That needs no unsafe and no new dependency; this crate has neither libc nor tempfile.
    • Exactly one answer reads as dead: exit status 1 with nothing on stdout and nothing on stderr.
    • Everything else reads as running: a probe that cannot start, a signal, any other status, or exit 1 with any byte on either stream.
    • Pid 0 and this process's own pid are running without asking.
    • A name must start with the prefix, and after it carry only digits, optionally followed by .ext.
    • So a live run's dir is never removed on an unclear answer. A dead one that gets kept costs only disk until a sweep that can tell.

The test body is now m2l_standalone_lake_check(parent, build). Taking the build command as a parameter lets the failure path be driven with false instead of a 7.6 GB download.

Item 4: other per-pid temp_dir() paths in the crate, cleaned only on success

Fixed in this PR (the big or never-cleaned ones):

Where Path Size Was Now
fpga.rs m2l lake test tri_m2l_standalone_pkg_<pid> + 3 side files ~7.6 GB success-only guarded + swept
fpga.rs test_measured_to_lean_standalone_outputs_consumable_lean tri_standalone_lake_pkg_<pid> (+ _in_, _generated_) empty dir never removed guarded + swept
census.rs Scratch (tri census explain) tri-census-explain-<pid> whole working-tree copy, a few hundred MB had a Drop; a killed run leaked it swept at the next Scratch::new
this PR's own test parents tri_piddir_<tag>_<pid>, tri_m2l_scratch_<tag>_<pid> KB guarded; a killed run leaked them also swept at the next run of the same test

Left as is. All are KB-sized, and a leak would need a panic or a kill:

  • Runtime, where an early ? or bail skips the trailing cleanup:
    • vsim.rs:141 tri-vsim-<pid>
    • oneaway.rs:187 tri-oneaway-<pid>
    • misread.rs:272 tri-misread-<pid>
    • types_dup.rs:889 tri-redef-probe-<pid>
    • modreach.rs:184 tri-mods-selfcheck-<pid>
  • Tests, cleaned only on success:
    • gates.rs: tri-sha- (4408), tri_t119_ (4836), tri_t112_ (4882), tri_t111_ (4975, a git init), tri_gates_empty_ (6105), tri_copy_from_/tri_copy_to_ (6348)
    • hooks.rs:449 now_gate_
    • issues.rs:1173
    • nownote.rs:524
    • seals.rs:810 w719-
    • modreach.rs:532 tri-modreach-bin-
    • fpga.rs: tri_sweep_report_json_, tri_cold_por_mock_, tri_cold_por_synthetic_, tri_sweep_report_synthetic_, tri_smoke_gate_validate_standalone(_snapshot)_ (JSON logs)
  • Not per-pid, so a new run overwrites them rather than piling up:
    • unparsed.rs: tri-unparsed-probes, tri-unparsed-counters, tri-locate-probe.t27
    • prose.rs: tri-prose-<spec>.t27

Any of these can adopt PidPath with one line when someone touches it.

Item 3: shared, locked lake package cache. Deferred.

Free space was 9.0 GiB at the start of round 1, and 5.7 GiB at its lowest during that build. In round 2 it read between 8.2 and 11 GiB. The rule for a real lake build is at least 20 GiB, so no before/after measurement was possible. Without one, the cache cannot be shown to be safe:

  • A permanent 7.6 GB dir is an owner policy question. It stops the disk filling repeatedly, but it costs 7.6 GB forever on every machine that ever runs the test.
  • Partial clones. A run killed mid-clone leaves a half-populated cache that later runs would trust. Fixing that needs a completeness marker written last, and a wipe when the marker is missing. Untested.
  • Concurrency. Two cargo test processes would need a cross-process lock around lake build. std::fs::File::lock is available in this toolchain. Whether lake tolerates one packagesDir shared across different root packages is unverified.
  • Toolchain. A shared cache built under the wrong toolchain would be shared breakage (see the finding below).

Measurement plan, for when there is room (>= 20 GiB free):

  1. Write lean-toolchain into the temp package first, copied from proofs/lean4/lean-toolchain.
  2. Cold run: /usr/bin/time -l cargo test -p tri -- --exact fpga::tests::test_measured_to_lean_standalone_builds_in_temp_lake_package, taking du -sh of the package dir before the guard drops it.
  3. Shared cache: tri_m2l_standalone_cache/ behind a File::locked .lock and a .complete marker, with lake pointed at it through an absolute packagesDir.
    • First verify that lake accepts an absolute packagesDir.
    • Then measure a cold run, a warm run, and two concurrent runs.
  4. kill -9 in the middle of a cold clone. Confirm that the next run wipes the cache and rebuilds, rather than trusting it.

Finding (not changed here): why the 8 runs probably failed

The temp package has no lean-toolchain, so elan uses its default toolchain there. On this machine that is v4.34.1 (lake 5.0.0). Meanwhile:

  • proofs/lean4/lean-toolchain pins leanprover/lean4:v4.31.0;
  • the mathlib rev in proofs/lean4/lake-manifest.json (800238935) is built for v4.31.0.

That mismatch is the likely reason lake build failed and the dirs leaked. Copying the toolchain file into the temp package is a one-line fix. It is unmeasured for the same disk reason, so it belongs with item 3.

Testing

CARGO_TARGET_DIR=/private/tmp/bee-t27-m2l-target cargo test -p tri --offline --no-run
# then, in cli/tri, with CARGO_MANIFEST_DIR set:
<test binary> piddir m2l_scratch test_measured_to_lean_standalone_outputs_consumable_lean
# test result: ok. 13 passed; 0 failed; 0 ignored; 0 measured; 849 filtered out

The same 13 passed again after the merge of master.

Test Shows
m2l_scratch_is_gone_after_a_failed_build build = false. It asserts that the panic message is the build step's own, so the test cannot pass by failing earlier, before the package existed. It then asserts that nothing survived.
m2l_scratch_is_gone_after_a_passing_build build = true: the success path leaves nothing either
m2l_scratch_sweeps_a_dead_runs_package_and_keeps_a_live_one a reaped child's package and side file are swept; ..._pkg_1 (pid 1, always running) is kept
m2l_scratch_parent_sweeps_a_killed_runs_parent a dead pid's tri_m2l_scratch_killed_<pid> is gone once the next run makes its own parent
piddir_pid_of_reads_only_prefix_digits_and_an_extension <prefix><pid>[.ext] only. Rejected: overflow, empty, signed and foreign names, x_pkg_123_x (suffix) and zx_pkg_123 (prefix not at the start)
piddir_probe_dead_only_on_a_silent_exit_1 fake probes: a silent exit 1 is dead. Running: exit 1 + stderr, exit 1 + stdout, exit 2, killed by a signal, a missing program, an empty probe
piddir_pid_zero_and_self_are_running_whatever_the_probe_says pid_is_running(0) with the real ps; pid 0 and self with a probe that says dead
piddir_liveness_running_and_reaped real ps: self and pid 1 are running, a reaped child is not
piddir_sweep_removes_dead_pids_and_keeps_live_ones reports exactly what it removed. Keeps self, pid 1, unparseable names, other prefixes, probe_pkg_<dead>_x and zprobe_pkg_<dead>
piddir_sweep_removes_nothing_when_the_probe_is_unclear the review's case. With each of the five unclear probes, probe_pkg_1, probe_pkg_<dead> and probe_pkg_<dead>.json all survive and nothing is reported
piddir_test_parent_sweeps_a_killed_runs_parent a dead pid's tri_piddir_killed_<pid> is gone once the next run makes its own parent
piddir_guard_removes_its_path_on_unwind a guarded dir and a guarded file do not outlive a panic

Mutants. The fix was committed first (82a1e8ae5). Then each mutant was applied to the file by a script that checked the diff was non-empty, rebuilt, and ran the 13 tests. Each was reverted with git checkout --, and git status --porcelain was checked to be empty after every one (it was, 16 of 16).

All runs used a private TMPDIR, so the mutants that break liveness could not reach anyone else's temp files. 16 of 16 went red:

Mutant Change Red First message
R1 pid_of splits at _ as well as . (accepts a suffix) pid_of_reads_only..., sweep_removes_dead_pids... the sweep must report exactly what it removed
R3 probe cannot start -> dead probe_dead_only..., sweep_removes_nothing... a probe that cannot start read as dead
R4 stdout check dropped probe_dead_only..., sweep_removes_nothing... exit 1 with stdout read as dead
R5 code == Some(1) -> !status.success() (any non-zero) probe_dead_only..., sweep_removes_nothing... exit 2 read as dead
R8 prefix matched anywhere (name.find(prefix)) pid_of_reads_only..., sweep_removes_dead_pids... the sweep must report exactly what it removed
R12 pid-0 guard removed pid_zero_and_self... pid 0, real ps. Killed, not argued: real ps -p 0 exits 1 silently on macOS; the fake-probe assert covers other systems
RERR stderr check dropped (the review's finding) probe_dead_only..., sweep_removes_nothing... probe ["sh", "-c", "echo err >&2; exit 1", "fake-ps"] swept [.../probe_pkg_1, ...]
RSELF self-pid guard removed pid_zero_and_self... this process, a probe saying dead
RSIG code().unwrap_or(1) == 1 (a signal reads as exit 1) probe_dead_only..., sweep_removes_nothing... a probe killed by a signal read as dead
RNONE an empty probe -> dead probe_dead_only... no probe at all read as dead
B sweep ignores liveness sweep_removes_nothing..., sweep_removes_dead_pids..., m2l_scratch_sweeps... the sweep must report exactly what it removed
A2 PidPath::drop skips while std::thread::panicking() guard_removes_its_path_on_unwind, m2l_scratch_is_gone_after_a_failed_build the guarded dir outlived a panic
A package dir removed only on success (ManuallyDrop + remove_dir_all after the assert!, the old shape) m2l_scratch_is_gone_after_a_failed_build a failed build left ["tri_m2l_standalone_pkg_88482"] behind
C the lake check sweeps nothing m2l_scratch_sweeps... a dead run's package survived the sweep
H1 parent() helper sweep removed piddir_test_parent_sweeps... a killed run's test parent survived
H2 m2l_scratch_parent() helper sweep removed m2l_scratch_parent_sweeps... a killed run's probe parent survived

Not run: the real test_measured_to_lean_standalone_builds_in_temp_lake_package. It needs 7.6 GB, and free space never reached the 20 GiB rule (at most 11 GiB). It still compiles and is listed.

In the wild (round 1): running test_measured_to_lean_standalone_outputs_consumable_lean swept two real dead-pid leftovers, tri_standalone_lake_pkg_34565 and tri_standalone_lake_pkg_53435.

Census: fetches moved in round 1, files read 46 -> 47. That count is the number of cli/tri/src/*.rs files the census reads, and the new piddir.rs is the extra one. It was re-blessed in the same commit, and no fetch site moved. Round 2 added no file; the hook printed PASS: no pinned census moved.

Checks that were red, and why

  • cli-tri: red at 7bc4843, inherited from master (the modreach cli/t27b ceiling, 6765d6379). Passes at 9e9d4f0 after the merge.
  • Every source file is reachable: red at 7bc4843, inherited from master and fixed there by fix(ci): the orphan-ceiling ledger names cli/t27b #5989. Passes at 9e9d4f0 after the merge.
  • fpga-conformance: the elaboration ratchet failed at 7bc4843 with elaboration errors: 178 (baseline 176), and the log notes the baseline came from iverilog 13.0 while the run used 12.0. The same job was red on all 8 of the most recent finished FPGA E2E Build runs listed at the time: 7 other branches plus this one. This PR touches only tri test scaffolding and temp-dir cleanup.

Review Notes

  • The census call site (Scratch::new) is a three-line sweep using the tested sweep_dead. The call site itself is not exercised by a test, because tri census explain only builds a Scratch when a pinned census has moved.
  • A pid recycled by an unrelated live process keeps a dead run's dir alive until that process exits. This is deliberately conservative.
  • ps is looked up on PATH. A hostile ps that exits 1 silently for every pid would still make every pid read as dead; no output-only rule can tell that apart from the real answer. A merely broken or unfamiliar ps (one that errors, prints, or exits with another code) now reads as running.

Closes #5982

🤖 Generated with Claude Code

gHashTag and others added 2 commits October 4, 2026 16:45
…lus dead-pid sweep

The m2l standalone lake test built a ~7.6 GB package in
temp_dir()/tri_m2l_standalone_pkg_<pid> and removed it on the line after
assert!(status.success()). A failed `lake build` panicked past the cleanup and
a killed run never reached it; on 2026-10-04 the leftovers filled the owner's
disk twice (about 300 MB free, other runs died with ENOSPC). Side files of 8
dead pids are still in $TMPDIR on that machine.

- cli/tri/src/piddir.rs: PidPath, a guard that removes its file or dir on
  drop (return and unwind alike), and sweep_dead(parent, prefix), which
  removes `<prefix><pid>[.ext]` entries whose pid is not running. Liveness is
  `ps -p`; anything short of a clear "no such process" counts as running, so
  a live run's scratch is never removed. No new dependency, no unsafe.
- fpga.rs: the lake test body is m2l_standalone_lake_check(parent, build).
  Every path is guarded before anything can fail; dead runs' leftovers are
  swept first. `build` is a parameter so the failure path is driven with
  `false` instead of a 7.6 GB download.
- fpga.rs: test_measured_to_lean_standalone_outputs_consumable_lean created
  tri_standalone_lake_pkg_<pid> and never removed it; now guarded and swept.
- census.rs: Scratch (a whole working-tree copy, a few hundred MB) already had
  a Drop; a killed run's copy is now swept by the next Scratch::new.

Census: `fetches` moved, files read 46 -> 47. That count is the number of
cli/tri/src/*.rs files it reads, and piddir.rs is the new one; no fetch site
moved. Re-blessed in this commit; quiet and shell did not move.

Tests: m2l_scratch_is_gone_after_a_failed_build (asserts the panic came from
the build step, then that nothing survived), ..._after_a_passing_build,
m2l_scratch_sweeps_a_dead_runs_package_and_keeps_a_live_one, and four piddir
unit tests (name parsing, liveness of self / pid 1 / a reaped child, sweep
keeps live and unparseable names, guard on unwind).

Closes #5982

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Closes #5982

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Oct 4, 2026

Copy link
Copy Markdown
Contributor

PR Dashboard

Generated at: 2026-10-04 10:04:24 UTC

Summary

Status Count
Total Open PRs 50
PRs with Failing Checks 44
PRs with All Checks Green 6
READY 1
FAILING 44
PENDING 0
NO CHECKS YET 0

These columns do not partition: 1 + 44 + 0 + 0 = 45, and there are 50 open PRs. A PR is being counted twice or not at all.

Seal Status

  • ⚠️ STALE -- sha256(compiler.rs)=2bb99f4036f1 != manifest seal=87e5cbd3ad94.
    The committed NMSE numbers were certified against an older compiler.rs.
    Run scripts/reseal-check.sh locally for the two-step reseal command (advisory; not a merge gate).

@github-actions

github-actions Bot commented Oct 4, 2026

Copy link
Copy Markdown
Contributor

📓 NotebookLM Notebook linked to this PR

This notebook contains session context, decisions, and artifacts for this work.

The review of #5990 refused it: pid_is_running treated any `ps` that exits
1 with empty stdout as "no such process", including a `ps` that exits 1
with an error on stderr. A fake `ps` first on PATH (`echo err >&2; exit 1`)
made pid 1 dead, and sweep_dead removed probe_pkg_1. A `ps` that rejects
`-p` or `-o` answers the same way, so on such a machine every pid reads
dead and runs delete each other's live 7.6 GB lake packages. The module
doc promised the opposite.

- Dead now needs exit status 1 AND empty stdout AND empty stderr. A probe
  that cannot start, a signal, any other status, and exit 1 with any byte
  on either stream all read as running. Checked on this Mac:
  `ps -p 99999 -o pid=` exits 1 with 0 bytes on both streams, while
  `ps -p 4194303 -o pid=` exits 1 with "process id too large" on stderr,
  which the old rule read as dead.
- The probe is a parameter (pid_is_running_by, sweep_dead_by), as `build`
  is for the lake check. Tests drive a fake per case: silent exit 1 (dead),
  exit 1 + stderr, exit 1 + stdout, exit 2, a signal, a missing program,
  no program (all running), and the reviewer's sweep: with every unclear
  probe, nothing planted is removed.
- Pid 0 and this process never ask the probe; a test pins both against a
  probe that says dead, plus pid_is_running(0) on the real ps.
- pid_of tests pin `x_pkg_123_x` (suffix) and `zx_pkg_123` (prefix not at
  the start) as rejected, and the sweep test keeps both shapes on disk.
- The tests' own parents (tri_piddir_<tag>_<pid>, tri_m2l_scratch_<tag>_<pid>)
  are swept by the same rule before each test makes its own, so a killed
  test run is cleaned by the next one. Two tests pin it, one per helper.

Refs #5982

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Picks up #5989 (e84927b), which fixes the "Every source file is
reachable" check this PR inherited red from master. Clean merge, no
conflicts.

Refs #5982

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Oct 4, 2026

Copy link
Copy Markdown
Contributor

PR Dashboard

Generated at: 2026-10-04 11:20:22 UTC

Summary

Status Count
Total Open PRs 50
PRs with Failing Checks 43
PRs with All Checks Green 7
READY 5
FAILING 43
PENDING 0
NO CHECKS YET 0

These columns do not partition: 5 + 43 + 0 + 0 = 48, and there are 50 open PRs. A PR is being counted twice or not at all.

Seal Status

  • ⚠️ STALE -- sha256(compiler.rs)=ecf0f8f43791 != manifest seal=87e5cbd3ad94.
    The committed NMSE numbers were certified against an older compiler.rs.
    Run scripts/reseal-check.sh locally for the two-step reseal command (advisory; not a merge gate).

This was referenced Oct 4, 2026
@gHashTag gHashTag added the bee-reviewed A reviewer bee reviewed and verified this PR at its current head; the only merge signal (#5525) label Oct 4, 2026
@gHashTag
gHashTag merged commit 75bc8a9 into master Oct 4, 2026
38 of 39 checks passed
@gHashTag
gHashTag deleted the claude/fpga-m2l-scratch-cleanup branch October 4, 2026 17:35
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bee-reviewed A reviewer bee reviewed and verified this PR at its current head; the only merge signal (#5525)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fpga.rs m2l lake test leaks a 7.6 GB temp package when lake build fails or the run is killed

1 participant