Skip to content

fix(t27b): signal a timed-out group with kill(2), not procps kill - #6362

Merged
gHashTag merged 1 commit into
masterfrom
claude/t27b-killpg
Oct 5, 2026
Merged

gHashTag merged 1 commit into
masterfrom
claude/t27b-killpg

Conversation

@gHashTag

@gHashTag gHashTag commented Oct 5, 2026

Copy link
Copy Markdown
Owner

Closes #6361 (Refs #6063)

What

run_capture (cli/t27b/src/blockers.rs) signals a timed-out child's process group with kill(2) directly. Before, it spawned kill -KILL -<pgid>. It also refuses pgid 0 and 1.

Why

The lab image ships kill from procps-ng 4.0.2, which reads a negative pid that follows a signal as -1. I checked this on the lab container:

  • /bin/kill -CONT -999999 gives rc 0, although no such group exists.
  • /bin/kill -CONT -- -999999 gives rc 1, No such process.

So the first 60 s corpus timeout on the lab sent SIGKILL to every process except PID 1, including the corpus driver. That is why runs d8375ad and 84e478b show t27b corpus exited -9 and wrote no JSON at about 62-68 s.

Evidence (lab at 84e478b)

  • With a logging kill wrapper that execs /bin/kill, the driver died right after -KILL -1165936.
  • With a logging wrapper that does nothing, the driver survived three timeout kills and went on to retrying 2 timed-out file(s).
  • Not OOM: cgroup oom_kill 0, memory.peak 5.4 GB of 24 GB.
  • Not a restart: PID 1 has been up since 2026-10-04 17:55Z.

Tests

  • New test a_timeout_kills_the_group_and_spares_the_caller. The group includes a grandchild holding the pipe, and the test asserts it dies.
  • cargo test -p t27b --test blockers: 7 passed, run locally on macOS.
  • I am running a lab cross-build of this branch (qemu, procps 4.0.2) and will add the result in a comment.

🤖 Generated with Claude Code

procps-ng 4.0.2 (the lab's Debian bookworm image) reads
`kill -KILL -<pgid>` as `kill -KILL -1`, so every corpus timeout on the
lab killed every process but PID 1, the corpus driver included (lab runs
d8375ad and 84e478b: exit -9, no JSON). Call kill(2) with a negative
pid directly and refuse pgid 0 and 1.

Closes #6361

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

PR Dashboard

Generated at: 2026-10-05 04:56:59 UTC

Summary

Status Count
Total Open PRs 46
PRs with Failing Checks 34
PRs with All Checks Green 12
READY 11
FAILING 34
PENDING 0
NO CHECKS YET 0

These columns do not partition: 11 + 34 + 0 + 0 = 45, and there are 46 open PRs. A PR is being counted twice or not at all.

Seal Status

  • ⚠️ STALE -- sha256(compiler.rs)=8597b6ded596 != manifest seal=87e5cbd3ad94.
    The committed NMSE numbers were certified against an older compiler.rs.
    Run scripts/reseal-check.sh locally for the two-step reseal command (advisory; not a merge gate).

@gHashTag

gHashTag commented Oct 5, 2026

Copy link
Copy Markdown
Owner Author

Lab check of this branch (8258b73), cross-built on the t27b lab container in a separate target dir. The run used qemu-aarch64 with --jobs 24 --timeout-ms 60000 against specs at 84e478b, with the real procps-ng 4.0.2 /bin/kill on PATH.

  • The driver survived its timeout kills. Its stderr shows retrying 4 timed-out file(s) one at a time.
  • It wrote JSON after 200.3 s: pass 370, pass_vacuous 278, fail 8, blocked 559, frontend 40, timeout 2, crash 0, mismatch 0.
  • The master binary at 84e478b, run the same way, died with -9 right after its first timeout kill.

@gHashTag
gHashTag merged commit df7b627 into master Oct 5, 2026
30 of 32 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

t27b: timeout kill via procps kill -KILL -<pgid> kills the corpus driver on the lab

1 participant