Skip to content

Let a runner pause on its channel without taking a job - #1122

Merged
epompeii merged 1 commit into
develfrom
u/ep/runner-protocol/pause-wire
Oct 8, 2026
Merged

epompeii merged 1 commit into
develfrom
u/ep/runner-protocol/pause-wire

Conversation

@epompeii

@epompeii epompeii commented Oct 7, 2026 •

Copy link
Copy Markdown
Member

What

Add a paused runner message, so an idle runner that should take no Job says so, says why, and stays connected.

  • paused carries reasons, each tagged by kind, plus since, when the pause began by the runner's clock, and every ready field.
    • raid: the array, its sync_action, and its progress from sync_completed and sync_speed.
    • maintenance: whether the maintenance marker is present and whether another process holds the job lock.
  • The server answers paused as it answers ready, without the claim. It runs the same self-update check, holds the poll until the runner's poll_timeout, and answers no_job. The hold is bounded by the poll timeout, not the heartbeat timeout, and the next ready on the same channel claims as usual.
  • During a Job, paused is warned about and acked like ready, and it does not reset the heartbeat.
  • A reason kind or sync action this server does not know parses as other, so a newer runner's message still parses. A message that fails to parse closes the channel.
  • The server reads at most 16 reasons and caps each string in them to 64 bytes on a character boundary, before any log line.
  • It logs Runner pause began, with the reasons and since, and Runner pause ended at the transitions only. The runner.state gauge gains paused, and a new runner.pause counter counts each pause start once per reason kind, under a pause.reason attribute.
  • RunnerMessage::Ready becomes a newtype over JsonReady, with an unchanged wire format, and paused flattens the same struct. Callers build it with JsonReady::new, so a later field touches no call site.
  • The runner protocol reference lists paused in all 9 languages.

Across versions:

  • No runner sends paused yet. This is the server half, and it must be live before any runner that sends it.
  • A server without this change cannot parse paused, so it closes the channel and the runner reconnects.
  • A runner without this change never sends paused, and its ready reads exactly as before.

Why

An idle runner that should take no Job, because an md array is resyncing or scrubbing or host maintenance is about to run, has two choices today. It can send ready and risk a Job measured while its disks are busy, or it can go quiet, which the server ends at its idle timeout by closing the channel, and self-update goes with it. paused keeps the runner connected and updating at its usual poll cadence, and it takes Jobs again within one poll of its reasons clearing.

Holding the poll, rather than answering no_job at once, reuses the timers ready already has. The runner needs no sleep of its own, and nothing has to stay under a server heartbeat timeout that the runner cannot see.

The reasons in each message carry the detail, and the counter shows a runner that pauses often.

Verification

  • Each new test fails under a mutation of the behavior it guards:
    • channel_paused_claims_no_job: a paused runner with a matching pending Job gets no_job no sooner than its poll, the Job stays pending, and ready on the same channel then claims it. Kills claiming on a pause, and answering no_job at once.
    • channel_paused_stale_runner_gets_update and paused_stable_version_mismatch_sends_update: a stale stable runner that pauses gets update, with the checksum a test listener serves. Kills skipping the update check on a pause.
    • channel_pause_outlasts_the_heartbeat_timeout: a 6 s pause against a 5 s heartbeat timeout, then a 1 s pause, then ready, each answered no_job on one channel. Kills a hold bounded by the heartbeat timeout.
    • paused_reasons_are_capped and paused_reason_strings_are_capped: 17 reasons become 16, and a 1,000 byte array name and sync action become 64 bytes. Kills reading the reasons uncapped.
    • channel_paused_during_job_does_not_reset_heartbeat: a paused during a Job is warned about and acked, and the Job still times out on its heartbeat. Kills a paused that resets the heartbeat.
    • ready_wire_format_is_unchanged pins today's full ready JSON both ways, paused_wire_format_is_pinned pins paused, and unknown_pause_reason_is_other and unknown_sync_action_is_other parse unknown values. Kills renaming a ready field, dropping the flatten, and dropping either catch-all.
  • A scratch test parsed and reserialized 13 runner messages on the previous server and on this one. Every ready is byte-identical: bare, with poll_timeout, with runner metadata, and with a canary channel and checksum in both key orders. So are heartbeat, running, and canceled. The previous server rejects paused with unknown variant 'paused' and closes the channel.
  • A scratch test sent a paused with a 5,000 byte array name and sync action and found no log line carrying 65 or more bytes of either.
  • The 11 new tests and the rest of bencher_json, bencher_schema, bencher_otel, api_runners, and bencher_runner pass. cargo gen-spec and typeshare reproduce the committed files, and clippy passes with -D warnings for Linux x86_64, Linux aarch64, and macOS.

@github-actions

github-actions Bot commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

馃惏 Bencher Report

ProjectBencher
Branchu/ep/runner-protocol/pause-wire
Testbedintel-v1
Click to view all benchmark results
BenchmarkLatencyBenchmark Result
microseconds (碌s)
(Result 螖%)
Upper Boundary
microseconds (碌s)
(Limit %)
Adapter::Json馃搱 view plot
馃毞 view threshold
5.17 碌s
(+0.71%)Baseline: 5.14 碌s
6.09 碌s
(84.98%)
Adapter::Magic (JSON)馃搱 view plot
馃毞 view threshold
5.02 碌s
(+1.30%)Baseline: 4.96 碌s
5.74 碌s
(87.46%)
Adapter::Magic (Rust)馃搱 view plot
馃毞 view threshold
24.57 碌s
(-8.84%)Baseline: 26.96 碌s
30.53 碌s
(80.50%)
Adapter::Rust馃搱 view plot
馃毞 view threshold
4.48 碌s
(-0.19%)Baseline: 4.49 碌s
6.39 碌s
(70.07%)
Adapter::RustBench馃搱 view plot
馃毞 view threshold
4.47 碌s
(-0.24%)Baseline: 4.48 碌s
6.39 碌s
(69.97%)
馃惏 View full continuous benchmarking report in Bencher

@epompeii
epompeii marked this pull request as ready for review October 7, 2026 14:55
@epompeii
epompeii added this pull request to stack #1126 October 7, 2026 14:56
A runner that should take no job sends `paused` instead of `ready`,
with its reasons (`raid`, `maintenance`, or a kind this server does not
know) and when the pause began. The server runs the same self-update
check as for `ready`, holds the poll without claiming a job, and
answers `no_job`. `ready` becomes a newtype over `JsonReady`, which
`paused` carries too, with an unchanged wire format.

The API caps the runner's strings in the reasons, logs when a pause
begins and ends, counts pause starts by reason, and shows paused
runners in the `runner.state` gauge. The runner protocol reference
lists the new message.
@epompeii
epompeii marked this pull request as draft October 8, 2026 17:21
@epompeii
epompeii force-pushed the u/ep/runner-protocol/pause-wire branch from e71c91c to a7336fa Compare October 8, 2026 17:21
@epompeii
epompeii marked this pull request as ready for review October 8, 2026 19:01
@epompeii
epompeii merged commit 17b521c into devel Oct 8, 2026
65 checks passed
@epompeii
epompeii deleted the u/ep/runner-protocol/pause-wire branch October 8, 2026 19:02
epompeii added a commit that referenced this pull request Oct 10, 2026
## What

Make `runner up` probe its host's md arrays before each `ready`, and
send `paused` instead while any array syncs, so no Job starts on disks
that a resync, recovery, or scrub keeps busy. This is the runner half of
`paused` (#1122).

- **The probe** runs only at the idle decision point: after any pending
result is resent, before each `ready`, never during a Job. It reads
`sync_action`, `sync_completed`, and `sync_speed` for each
`/sys/block/md*` entry with an `md/` directory.
- An array is busy while `sync_action` is `resync`, `recover`, `check`,
`repair`, or `reshape` and `sync_completed` is not `none`, which covers
a running sync (`N / M`) and one queued behind another array
(`delayed`).
- `idle`, `frozen`, an action this runner does not know, an array with
no sync thread such as raid0, and a host with no arrays never pause.
- **`paused`** carries one `raid` reason per busy array, with its
action, progress, and speed, plus `since` and every `ready` field.
`since` is when the pause began by the runner's clock, and it holds for
the whole pause, whatever its reasons become.
- **After `paused`** the runner waits for `no_job` or `update` exactly
as after `ready`, then probes again. The server holds the poll, so a
paused runner stays connected, still updates itself, and takes Jobs
again within one poll of the sync ending.
- **A Job sent to a paused runner** is failed unrun, with `The runner
was paused for a RAID sync on md0, so it ran no Job`, through the usual
terminal message, ACK, and retry, and no `running`.
- **Reconnects while paused** first hold off one poll timeout, ahead of
the usual 5 to 10 s delay, on every path that reconnects. The hold
sleeps in slices of at most 1 s, so a stop ends it within a second.
- **Logs:** one `Paused for a RAID sync` record per busy array when a
pause begins (`array`, `action`, `percent`, `speed_kib`), `Pause ended`
when it ends, and `Holding off the reconnect while paused` with
`hold_secs`. A pause whose reasons change mid-pause is not logged again.
- **`--no-raid-pause`** on `runner up` keeps taking Jobs through a sync.
The pause is on by default, and `raid_pause` is in the `Runner starting`
record.
- The `runner up` reference documents the flag in all 9 languages, and
the changelog gains a line.

With #1125 and #1130 on the server, the pause shows on the admin runner
endpoints, and a pause past 12 hours mails the server admins.

Across versions:
- A server with #1122 holds the paused poll and answers `no_job` on the
same channel.
- A server without it cannot parse `paused`, so it closes the channel.
The runner then reconnects no faster than once per poll, so such a
server sees no more connects than polls, and no self-update reaches the
runner until its pause ends. The first `ready` after the pause takes
Jobs as before.

## Why

A sync reads or writes the whole array, and a Job measured during one
measures the disks' contention as much as the benchmark. Probing only
between Jobs leaves a running Job alone: a sync that starts mid-Job does
not stop it, and the next Job waits.

`sync_action` alone is not enough. It reads `recover` while a recovery
is needed but cannot run, such as on a degraded array with no spare,
which would pause the runner until someone adds a disk. A running or
queued sync thread is what keeps the disks busy, so `sync_completed` has
to say one exists.

Holding the poll on the server, rather than sleeping on the runner,
reuses the timers `ready` already has, so the runner needs no sleep of
its own for the pause.

The reconnect hold is for a server without `paused`. Without it, a
runner whose channel closes on every `paused` reconnects every 5 to 10
s, about 480 times an hour, and trips a runner connect limit of 256 an
hour after about half an hour; connects then fail even after the pause
ends, until the hour drains. With the hold, a runner reconnects at most
about `3600 / (poll + 7.5)` times an hour: about 58 at the default 55 s
poll. A poll under about 7 s would still pass 256 an hour against such a
server.

The server never claims a Job for a paused runner, so a Job that arrives
anyway is failed rather than measured on busy disks.

## Verification

- Each new test fails under a mutation of the behavior it guards:
- `a_running_sync_of_each_kind_pauses`: a busy set that misses an
action, or counts `frozen`, `idle`, or an unknown action as busy.
- `a_sync_action_with_no_sync_thread_does_not_pause`: busy on
`sync_action` alone (`recover` with `sync_completed` `none`).
  - `a_queued_sync_pauses_with_no_progress`: `delayed` not busy.
- `a_pause_reason_carries_the_progress_and_speed`: swapped or dropped
progress and speed, and an array named after anything but its block
device.
- `a_host_with_no_arrays_never_pauses` and
`an_array_with_no_sync_thread_is_listed_but_never_busy`: a disk with no
`md/` read as an array, a host with no `block` directory failing, and
raid0 dropped or busy.
- `a_busy_probe_pauses_and_a_clear_one_sends_ready`: the probe's reasons
ignored, a `paused` without `ready`'s fields (the server's update check
needs the runner metadata), and a `no_job` that does not probe again.
- `a_pause_keeps_its_start_and_is_logged_once_each_way`: a `since` that
restarts each probe or outlives the pause, and a start or end logged
every poll.
- `a_paused_runner_still_updates` and
`a_job_after_a_pause_fails_without_running`: a pause that ignores
`update`, and a paused runner that runs a Job.
- `reconnects_hold_off_a_poll_only_while_paused`, at a 55 s poll, drives
every path that can reconnect while paused: the channel closed, a
refused connect, a receive timeout, a failed self-update, a Job failed
unrun whose result is lost with the connection, and each of the 3 ways
its resend can fail at the next connect. Removing the hold at any one of
those 8 sites fails it, and so do a fixed 30 s hold and a hold that
outlives the pause. The ninth reconnect site, a connection lost during a
Job, cannot be reached while paused.
- `a_paused_reconnect_holds_off_for_the_whole_hold` and
`a_stop_ends_a_hold_within_one_slice` run the hold through the driver's
effect: a hold that does not sleep, and one a stop cannot end.
- `no_raid_pause_takes_jobs_through_a_sync` and
`the_raid_pause_is_on_unless_turned_off`: the flag ignored by the probe,
inverted, or off by default.
- None of the new tests compiles on the parent, which has no probe. The
14 runner tests and the CLI test pass 20 of 20 as an ordinary user and
as root.
- **In a disposable VM**, a root `runner up` at a 5 s poll ran against a
local API, with a RAID1 over two 2 GiB loop devices:
- `echo check > sync_action` paused the runner within 4 s, with `Paused
for a RAID sync` (`md0`, `check`), and the server stored `paused` with
the `raid` reason, its progress, and its speed.
  - A Job submitted then read `pending` at each of 12 checks over 60 s.
- `echo idle` ended the pause at the next probe, 1 s later, and the Job
was claimed that second and processed 2 s after.
- The whole run used one channel. A second run at a 7 s poll behaved the
same.
- **Against a server built from the commit before #1122**, with its
runner connect limiter on at 16 a minute and 256 an hour, a root `runner
up` at a 20 s poll was held paused by a maintenance marker from the next
layer in this stack; the hold does not depend on the reason.
- Each `paused` closed the channel, and the runner logged `Holding off
the reconnect while paused` with `hold_secs` 20.
- Over 627 s it made 24 connects, 25.0 to 30.0 s apart: at most 3 in any
minute, and 132 an hour. No connect drew a 429, while a burst of 20
connects with the same key at the end drew 5.
- Once the marker was gone, the next connect probed clear and a waiting
Job was claimed 9 s after the marker came off.
- A hold that does not sleep reconnected every 5.0 to 10.0 s against the
same server, 483 an hour.
  - SIGTERM 2.5 s into a 20 s hold ended `runner up` in 0.50 s.
- The runner unit tests pass as an ordinary user and as root, 427 of
427, the CLI tests pass, 14 of 14, the `runner_ops` tests pass, 86 of
86, `cargo test-runner scenarios --build-only` passes, `Cargo.lock` is
unchanged, and clippy passes with `-D warnings` for Linux x86_64, Linux
aarch64, and macOS on the runner, its CLI, `test_runner`, and
`runner_ops`. The scenario suite passes as root, 99 of 99, and clippy
passes for the whole workspace as CI runs it, on the top of this stack.
epompeii added a commit that referenced this pull request Oct 10, 2026
## What

Make the runner and its host's maintenance take turns through a lock and
markers in `/run/bencher`, so maintenance never runs during a Job, and
the runner takes no Job while maintenance waits or runs.

- **The job lock** is `/run/bencher/job.lock`. Before each `ready`, the
probe takes it with a nonblocking `flock` on a close-on-exec descriptor,
and the runner holds it through the poll and any Job the poll brings. It
is released on `no_job`, `update`, a receive timeout, a lost connection,
a `running` that fails to send, and when the Job ends, before its result
goes out.
- **A marker** is a file in `/run/bencher` named `maintenance` or
`maintenance.<name>`. Maintenance creates its own before it waits for
the lock and removes it when done.
- **The probe pauses for maintenance:**
- while a marker exists, the runner sends `paused` with a `maintenance`
reason and `marker` true, and leaves the lock free;
- while another process holds the lock, it sends a `maintenance` reason
with `lock` true;
- only a probe that finds no reason to pause keeps the lock, so a runner
paused for a RAID sync leaves it free too.
- **The plain `maintenance` marker drains a runner.** It finishes its
current poll and any Job that poll brings, then takes no Job until the
marker is gone. A unit uses a marker of its own, so it never removes the
drain.
- **`runner run`** waits for the job lock, trying it every 250 ms and
logging `Waiting for host maintenance to finish` once, and a stop ends
the wait. It warns but runs through a RAID sync (`Running during a RAID
sync`) or a marker (`Running although host maintenance is due`).
- **`/run/bencher` is root's alone:** the runner now creates it 0700
rather than 0755, for the runner lock and the job lock alike, and each
lock file 0600. A directory an older runner created keeps its mode until
a reboot or a drop-in's `install -d -m 0700`.
- **A runner that is not root** can open neither the lock nor the 0700
directory, so it takes Jobs without them, and `runner up` warns once at
startup that maintenance may run during a Job. A root runner that cannot
open or take the lock pauses with `lock` true rather than run unguarded.
- **`bencher_runner::maintenance`**, public and outside the `plus`
feature, holds the directory, lock, and marker names and
`marker_present`, which the probe and `runner run` both use, so host
tooling can share them.
- **Docs:** a "Host Maintenance" section on the self-hosted runners page
in all 9 languages: the lock, the markers, an example `fstrim` drop-in,
and how to drain a Runner. The changelog gains a line.

Across versions:
- Like the RAID pause in the layer below, this needs a server with
#1122. A server without it closes the channel on `paused`, and the
runner reconnects no faster than once per poll.
- A runner without this change never takes the lock and ignores the
markers, so a unit that waits on the lock runs at once, as it does
today.

## Why

Package upgrades, `fstrim`, and the monthly RAID check run on timers
that know nothing about Jobs, so any of them can land during a Job and
move its measurement.

The runner takes the lock before `ready` rather than when a Job arrives.
Taken when a Job arrives, the lock could already be maintenance's, which
would hold a Job the server has handed out while that Job's clock ran.
Taken before `ready`, maintenance waits at most one poll plus one Job.

A waiting `flock` holds nothing the runner can see, and on each release
the runner's own nonblocking try can win the lock back. The marker tells
the runner that maintenance waits, so it steps aside and the waiter gets
the lock at the next release. Markers of their own let two units
overlap, such as a scrub and the morning upgrade: with one shared
marker, the first to finish would remove it while the second still
waits, and any unit would remove an operator's drain.

`flock(2)` locks whatever mode a file was opened in, so any user who can
open the lock file can hold it and keep the runner paused. A root-only
directory and 0600 files keep that to root. The descriptor is
close-on-exec, so the jailer, the VMM, and every other child never keep
the lock past the Job.

`runner run` is started by hand at a moment someone chose, so it warns
about a sync or a marker and runs, but it never runs beside maintenance
that holds the lock. Trying the lock in a loop, rather than blocking in
`flock(2)`, lets a stop end the wait at once.

## Verification

- Lock tests use two independent opens of a file in a temporary
directory, which `flock` treats as two holders even in one process. Each
new test fails under a mutation of the behavior it guards:
- `a_free_lock_is_the_runners_turn` and
`a_lock_held_elsewhere_pauses_for_maintenance`: a turn that never takes
the lock, and a held lock taken as free.
- `a_marker_pauses_without_taking_the_lock`, with `maintenance` and
`maintenance.fstrim.service`: a marker ignored, the lock kept past a
marker (which starves waiting maintenance), and only the plain marker
counted.
- `only_the_maintenance_names_are_markers`: any file, or any name
starting `maintenance`, read as a marker.
- `dropping_the_turn_releases_the_lock` and
`a_child_never_holds_the_lock`: a leaked descriptor, and one that is not
close-on-exec.
- `the_lock_and_its_directory_are_root_only`: `job.lock` at 0644 or
0640, and `/run/bencher` at 0755 or 0750.
- `a_runner_that_is_not_root_takes_its_turn_without_the_lock`: a runner
that is not root failing closed, and a root runner running on without
the lock.
- `runner_run_waits_for_maintenance_then_takes_the_lock` and
`a_stop_ends_the_wait_for_maintenance`: a `runner run` that does not
wait, gives up once maintenance holds the lock, or keeps waiting past a
stop.
- `the_job_lock_is_held_from_ready_until_the_job_finishes` and
`the_job_lock_is_released_whenever_no_job_follows`: a release when the
Job arrives, none when it ends, and none on `no_job`, `update`, a
receive timeout, a lost connection, or a failed `running` send.
- `the_job_lock_is_kept_only_by_a_clear_probe`: the lock kept through a
pause, or dropped by a clear probe.
- `the_driver_releases_the_job_lock_when_told`: a lock the probe took
must be free to a second open after the driver runs the release. Kills a
driver that ignores it, which would hold maintenance off for a whole
server outage, since no probe runs until a connect succeeds.
- None of the new tests compiles on the parent, which has no job lock.
The 14 tests pass 20 of 20 as an ordinary user and as root.
- **In a disposable VM**, a root `runner up` at a 5 s poll ran against a
local API, with `/run/bencher` absent as after a boot:
  - The runner created `/run/bencher` 0700 and `job.lock` 0600.
- A unit shaped like the next layer's drop-in, started 1 s into a 20 s
Job, created its marker at once and waited. Its command started 1 ms
after the runner released the lock at the Job's end, while the runner
paused with `marker` and `lock` true. A second Job, submitted during the
first, was claimed after the maintenance.
- `systemctl kill -s KILL` of such a unit, while it waited and while it
ran, left the unit `failed` and its marker gone, and the runner's next
probe ended its pause.
- `touch /run/bencher/maintenance` paused the runner once its current
poll ended, with `marker` true and `lock` false. A Job submitted then
read `pending` for 30 s and was claimed within 2 s of the `rm`.
- **`runner run` behind a held lock**, in the same VM, sandboxed with
Firecracker, while another process held `job.lock` for 20 s:
- It logged `Waiting for host maintenance to finish` once, started no VM
until the release, then ran its benchmark.
- While Firecracker ran, `job.lock` was open only in the runner, and
`/proc/locks` showed the runner's pid holding it.
- SIGTERM during the wait ended the run in 0.24 s, with no VM started
and the holder's lock untouched.
- **The scenario `runner_run_takes_its_turn_at_the_job_lock`** puts both
halves of `runner run`'s turn in CI, in about 7 s:
- The harness holds `job.lock` and starts a sandboxed `runner run`. The
scenario requires one `Waiting for host maintenance to finish`, then no
jail for the 2 s until the release, then the Job. Once the VM is up,
neither Firecracker nor the jailer it came from may hold the lock.
- The harness then creates `/run/bencher/maintenance` and runs the image
on the host. The scenario requires one `Running although host
maintenance is due`, no wait, and a successful Job whose benchmark, the
runner's own child, lists no descriptor of `job.lock`. The host run
carries the close-on-exec check, since the jailer closes every
descriptor it inherits.
  - The marker and the harness's lock are gone on every path.
- It fails when `runner run` never waits, when its descriptor is not
close-on-exec, and when a marker makes it refuse, wait, or run without
the warning. It fails on the parent, which has no job lock, and passes
20 of 20 as root.
- The runner unit tests pass as an ordinary user and as root, 441 of
441, the CLI tests pass, 14 of 14, the `runner_ops` tests pass, 86 of
86, `cargo test-runner scenarios --build-only` passes, `Cargo.lock` is
unchanged, and clippy passes with `-D warnings` for Linux x86_64, Linux
aarch64, and macOS on the runner, its CLI, `test_runner`, and
`runner_ops`. The scenario suite passes as root, 100 of 100, and clippy
passes for the whole workspace as CI runs it, on the top of this stack.
epompeii added a commit that referenced this pull request Oct 10, 2026
## What

Make the runner send its host's disk health with every `ready` and
`paused`, in the `health` field that #1125 stores, so admins see it on
the runner endpoints and #1130 mails them when it is news.

- **md, at every probe:** each array's `array_state`, `degraded`, and
`sync_action`, read with the RAID pause's own read. An array with no
sync thread, such as raid0, reads `idle`. A degraded array is a
`degraded` finding that is `failing`.
- **NVMe, at most hourly:** each controller under `/sys/class/nvme`,
with its model, its serial, and the SMART / health log fields that set a
finding: critical warning, temperature, spare and its threshold,
percentage used, and media errors. The page is read with Get Log Page
(log `0x02`, all namespaces, 512 bytes) through the kernel's admin
passthrough, `NVME_IOCTL_ADMIN_CMD`, on a read-only open of
`/dev/<controller>`. The read is cached for an hour, and only the probe
reads it, so never during a Job.
- **States:**
- `warning`: spare at or below twice its threshold, percentage used over
80, the temperature critical warning bit, or any media and data
integrity error;
- `failing`: any other critical warning bit, spare below its threshold,
or a degraded array;
  - the host's state is its worst finding, and `ok` with none.
- **A runner that is not root** cannot use the passthrough, which needs
`CAP_SYS_ADMIN`. It lists each controller with no log, warns once with
`NVMe health not read, so it is reported unreadable`, and still reports
md.
- Health is always sent: a host with no arrays and no controllers
reports `ok` with empty lists.
- **Nothing pauses on health.** A degraded array keeps taking Jobs; only
a sync pauses, as before.
- `AdminCommand` is the kernel's `struct nvme_admin_cmd`, pinned at 72
bytes by a const assertion. `nix` gains its `ioctl` feature, which adds
no package to `Cargo.lock`.
- The changelog gains a line.

Across versions:
- A server without #1125 ignores `health`, since the runner messages
accept unknown fields.
- `paused` needs a server with #1122, as in the layers below.

## Why

A degraded array or a failing NVMe drive on a runner host is silent
until the next disk fails too. The server half stores health and mails
admins on news; this is what sends it.

md costs a few sysfs reads, so it is read at every probe. Each NVMe read
is an admin command to the drive, so it runs at most hourly, and only
between Jobs.

The temperature bit is a warning, not failing, because it clears by
itself. Media errors warn while their count is nonzero, rather than when
it rises, which would read `ok` again at the next hourly read.

Health is reported, never a pause. An idle degraded array barely moves a
benchmark, and the mail reaches whoever can replace the disk. Health
still rides on `paused`, because the rebuild after a disk is replaced is
a pause, and that is when health matters most.

## Verification

- Each new test fails under a mutation of the rule it guards:
- The NVMe parser runs on fixture pages.
`a_page_parses_little_endian_at_its_offsets` and
`a_media_error_count_past_u64_saturates`: swapped temperature bytes,
media errors read big-endian or from the wrong counter, spare read from
the threshold byte, and a 128-bit count that wraps.
- `a_healthy_page_has_no_findings` and
`each_critical_warning_bit_but_temperature_is_failing`, over bits 0 and
2 to 7: a rule that fires on a healthy page, any one critical bit
missed, and the temperature bit failing or ignored.
- `spare_at_twice_its_threshold_warns_and_below_it_fails`,
`percentage_used_over_80_warns`, and `any_media_error_warns`: either
spare boundary or the endurance boundary moved by one, and media errors
ignored or failing.
- `an_unreadable_log_still_lists_the_controller`: an unreadable
controller dropped, and a log read from another device.
- `a_degraded_array_is_failing` and `the_worst_finding_sets_the_state`:
a degraded array that reads healthy, every array failing, and the state
taken from the first finding rather than the worst.
- `nvme_is_read_at_most_hourly`, on an injected clock and reader: read
once at the start, served from the cache 1 s before the hour, and read
again at the hour. Kills a read every probe, and one never repeated.
- `an_unreadable_nvme_still_reports_md` and
`every_probe_reports_md_health`: a runner that is not root reporting
nothing, and a probe that sends no health, or none with the RAID pause
off.
- `the_probes_health_rides_on_ready_and_paused`: health dropped from
either message.
- None of the new tests compiles on the parent, which reads no health.
The 14 tests pass 20 of 20 as an ordinary user and as root.
- A C probe against Ubuntu 24.04's `linux/nvme_ioctl.h` gives
`sizeof(struct nvme_admin_cmd)` 72, `NVME_IOCTL_ADMIN_CMD` `0xc0484e41`,
and `addr`, `data_len`, and `cdw10` at offsets 24, 36, and 40, which the
Rust struct and `ioctl_readwrite!(b'N', 0x41)` match.
- **In a disposable VM**, an NVMe controller on the kernel's `nvme-loop`
target, backed by a file, stood in for a drive:
- As root, the runner's own read returned the page with status 0, and
the page's data units read and host read commands equal `nvme
smart-log`'s values for that controller.
- As an ordinary user, the read fails with `EACCES`, the controller is
listed with no log, and md is still reported.
- A dword count off by one fails there with status `0x600f`, which the
fixture pages cannot show.
- **In the same kind of VM**, a root `runner up` ran against a local
API, with a RAID1 over two loop devices:
- Healthy, the server stored health `ok` with `md0` `clean`, 0 degraded,
`idle`.
- `mdadm --fail` on one member: the next report read `failing`, with
`md0` `degraded` `failing`, and the runner still `ready`. A Job
submitted then ran.
- `mdadm --remove` and `--add`: the runner paused for the `recover` sync
with health still `failing`. Once the recovery finished, the pause
ended, and the next report read `ready` and `ok`.
- The runner unit tests pass as an ordinary user and as root, 455 of
455, the CLI tests pass, 14 of 14, the `runner_ops` tests pass, 86 of
86, `cargo test-runner scenarios --build-only` passes, `Cargo.lock` is
unchanged, and clippy passes with `-D warnings` for Linux x86_64, Linux
aarch64, and macOS on the runner, its CLI, `test_runner`, and
`runner_ops`. The scenario suite passes as root, 100 of 100, and clippy
passes for the whole workspace as CI runs it, on the top of this stack.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant