Repository navigation
Let a runner pause on its channel without taking a job - #1122
Merged
Merged
Conversation
Contributor
|
| Project | Bencher |
| Branch | u/ep/runner-protocol/pause-wire |
| Testbed | intel-v1 |
Click to view all benchmark results
| Benchmark | Latency | Benchmark Result microseconds (碌s) (Result 螖%) | Upper Boundary microseconds (碌s) (Limit %) |
|---|---|---|---|
| Adapter::Json | 馃搱 view plot 馃毞 view threshold | 5.17 碌s(+0.71%)Baseline: 5.14 碌s | 6.09 碌s (84.98%) |
| Adapter::Magic (JSON) | 馃搱 view plot 馃毞 view threshold | 5.02 碌s(+1.30%)Baseline: 4.96 碌s | 5.74 碌s (87.46%) |
| Adapter::Magic (Rust) | 馃搱 view plot 馃毞 view threshold | 24.57 碌s(-8.84%)Baseline: 26.96 碌s | 30.53 碌s (80.50%) |
| Adapter::Rust | 馃搱 view plot 馃毞 view threshold | 4.48 碌s(-0.19%)Baseline: 4.49 碌s | 6.39 碌s (70.07%) |
| Adapter::RustBench | 馃搱 view plot 馃毞 view threshold | 4.47 碌s(-0.24%)Baseline: 4.48 碌s | 6.39 碌s (69.97%) |
epompeii
marked this pull request as ready for review
October 7, 2026 14:55
epompeii
added this pull request to stack #1126
October 7, 2026 14:56
A runner that should take no job sends `paused` instead of `ready`, with its reasons (`raid`, `maintenance`, or a kind this server does not know) and when the pause began. The server runs the same self-update check as for `ready`, holds the poll without claiming a job, and answers `no_job`. `ready` becomes a newtype over `JsonReady`, which `paused` carries too, with an unchanged wire format. The API caps the runner's strings in the reasons, logs when a pause begins and ends, counts pause starts by reason, and shows paused runners in the `runner.state` gauge. The runner protocol reference lists the new message.
epompeii
marked this pull request as draft
October 8, 2026 17:21
epompeii
force-pushed
the
u/ep/runner-protocol/pause-wire
branch
from
October 8, 2026 17:21
e71c91c to
a7336fa
Compare
epompeii
marked this pull request as ready for review
October 8, 2026 19:01
This was referenced Oct 9, 2026
epompeii
added a commit
that referenced
this pull request
Oct 10, 2026
## What Make `runner up` probe its host's md arrays before each `ready`, and send `paused` instead while any array syncs, so no Job starts on disks that a resync, recovery, or scrub keeps busy. This is the runner half of `paused` (#1122). - **The probe** runs only at the idle decision point: after any pending result is resent, before each `ready`, never during a Job. It reads `sync_action`, `sync_completed`, and `sync_speed` for each `/sys/block/md*` entry with an `md/` directory. - An array is busy while `sync_action` is `resync`, `recover`, `check`, `repair`, or `reshape` and `sync_completed` is not `none`, which covers a running sync (`N / M`) and one queued behind another array (`delayed`). - `idle`, `frozen`, an action this runner does not know, an array with no sync thread such as raid0, and a host with no arrays never pause. - **`paused`** carries one `raid` reason per busy array, with its action, progress, and speed, plus `since` and every `ready` field. `since` is when the pause began by the runner's clock, and it holds for the whole pause, whatever its reasons become. - **After `paused`** the runner waits for `no_job` or `update` exactly as after `ready`, then probes again. The server holds the poll, so a paused runner stays connected, still updates itself, and takes Jobs again within one poll of the sync ending. - **A Job sent to a paused runner** is failed unrun, with `The runner was paused for a RAID sync on md0, so it ran no Job`, through the usual terminal message, ACK, and retry, and no `running`. - **Reconnects while paused** first hold off one poll timeout, ahead of the usual 5 to 10 s delay, on every path that reconnects. The hold sleeps in slices of at most 1 s, so a stop ends it within a second. - **Logs:** one `Paused for a RAID sync` record per busy array when a pause begins (`array`, `action`, `percent`, `speed_kib`), `Pause ended` when it ends, and `Holding off the reconnect while paused` with `hold_secs`. A pause whose reasons change mid-pause is not logged again. - **`--no-raid-pause`** on `runner up` keeps taking Jobs through a sync. The pause is on by default, and `raid_pause` is in the `Runner starting` record. - The `runner up` reference documents the flag in all 9 languages, and the changelog gains a line. With #1125 and #1130 on the server, the pause shows on the admin runner endpoints, and a pause past 12 hours mails the server admins. Across versions: - A server with #1122 holds the paused poll and answers `no_job` on the same channel. - A server without it cannot parse `paused`, so it closes the channel. The runner then reconnects no faster than once per poll, so such a server sees no more connects than polls, and no self-update reaches the runner until its pause ends. The first `ready` after the pause takes Jobs as before. ## Why A sync reads or writes the whole array, and a Job measured during one measures the disks' contention as much as the benchmark. Probing only between Jobs leaves a running Job alone: a sync that starts mid-Job does not stop it, and the next Job waits. `sync_action` alone is not enough. It reads `recover` while a recovery is needed but cannot run, such as on a degraded array with no spare, which would pause the runner until someone adds a disk. A running or queued sync thread is what keeps the disks busy, so `sync_completed` has to say one exists. Holding the poll on the server, rather than sleeping on the runner, reuses the timers `ready` already has, so the runner needs no sleep of its own for the pause. The reconnect hold is for a server without `paused`. Without it, a runner whose channel closes on every `paused` reconnects every 5 to 10 s, about 480 times an hour, and trips a runner connect limit of 256 an hour after about half an hour; connects then fail even after the pause ends, until the hour drains. With the hold, a runner reconnects at most about `3600 / (poll + 7.5)` times an hour: about 58 at the default 55 s poll. A poll under about 7 s would still pass 256 an hour against such a server. The server never claims a Job for a paused runner, so a Job that arrives anyway is failed rather than measured on busy disks. ## Verification - Each new test fails under a mutation of the behavior it guards: - `a_running_sync_of_each_kind_pauses`: a busy set that misses an action, or counts `frozen`, `idle`, or an unknown action as busy. - `a_sync_action_with_no_sync_thread_does_not_pause`: busy on `sync_action` alone (`recover` with `sync_completed` `none`). - `a_queued_sync_pauses_with_no_progress`: `delayed` not busy. - `a_pause_reason_carries_the_progress_and_speed`: swapped or dropped progress and speed, and an array named after anything but its block device. - `a_host_with_no_arrays_never_pauses` and `an_array_with_no_sync_thread_is_listed_but_never_busy`: a disk with no `md/` read as an array, a host with no `block` directory failing, and raid0 dropped or busy. - `a_busy_probe_pauses_and_a_clear_one_sends_ready`: the probe's reasons ignored, a `paused` without `ready`'s fields (the server's update check needs the runner metadata), and a `no_job` that does not probe again. - `a_pause_keeps_its_start_and_is_logged_once_each_way`: a `since` that restarts each probe or outlives the pause, and a start or end logged every poll. - `a_paused_runner_still_updates` and `a_job_after_a_pause_fails_without_running`: a pause that ignores `update`, and a paused runner that runs a Job. - `reconnects_hold_off_a_poll_only_while_paused`, at a 55 s poll, drives every path that can reconnect while paused: the channel closed, a refused connect, a receive timeout, a failed self-update, a Job failed unrun whose result is lost with the connection, and each of the 3 ways its resend can fail at the next connect. Removing the hold at any one of those 8 sites fails it, and so do a fixed 30 s hold and a hold that outlives the pause. The ninth reconnect site, a connection lost during a Job, cannot be reached while paused. - `a_paused_reconnect_holds_off_for_the_whole_hold` and `a_stop_ends_a_hold_within_one_slice` run the hold through the driver's effect: a hold that does not sleep, and one a stop cannot end. - `no_raid_pause_takes_jobs_through_a_sync` and `the_raid_pause_is_on_unless_turned_off`: the flag ignored by the probe, inverted, or off by default. - None of the new tests compiles on the parent, which has no probe. The 14 runner tests and the CLI test pass 20 of 20 as an ordinary user and as root. - **In a disposable VM**, a root `runner up` at a 5 s poll ran against a local API, with a RAID1 over two 2 GiB loop devices: - `echo check > sync_action` paused the runner within 4 s, with `Paused for a RAID sync` (`md0`, `check`), and the server stored `paused` with the `raid` reason, its progress, and its speed. - A Job submitted then read `pending` at each of 12 checks over 60 s. - `echo idle` ended the pause at the next probe, 1 s later, and the Job was claimed that second and processed 2 s after. - The whole run used one channel. A second run at a 7 s poll behaved the same. - **Against a server built from the commit before #1122**, with its runner connect limiter on at 16 a minute and 256 an hour, a root `runner up` at a 20 s poll was held paused by a maintenance marker from the next layer in this stack; the hold does not depend on the reason. - Each `paused` closed the channel, and the runner logged `Holding off the reconnect while paused` with `hold_secs` 20. - Over 627 s it made 24 connects, 25.0 to 30.0 s apart: at most 3 in any minute, and 132 an hour. No connect drew a 429, while a burst of 20 connects with the same key at the end drew 5. - Once the marker was gone, the next connect probed clear and a waiting Job was claimed 9 s after the marker came off. - A hold that does not sleep reconnected every 5.0 to 10.0 s against the same server, 483 an hour. - SIGTERM 2.5 s into a 20 s hold ended `runner up` in 0.50 s. - The runner unit tests pass as an ordinary user and as root, 427 of 427, the CLI tests pass, 14 of 14, the `runner_ops` tests pass, 86 of 86, `cargo test-runner scenarios --build-only` passes, `Cargo.lock` is unchanged, and clippy passes with `-D warnings` for Linux x86_64, Linux aarch64, and macOS on the runner, its CLI, `test_runner`, and `runner_ops`. The scenario suite passes as root, 99 of 99, and clippy passes for the whole workspace as CI runs it, on the top of this stack.
epompeii
added a commit
that referenced
this pull request
Oct 10, 2026
## What Make the runner and its host's maintenance take turns through a lock and markers in `/run/bencher`, so maintenance never runs during a Job, and the runner takes no Job while maintenance waits or runs. - **The job lock** is `/run/bencher/job.lock`. Before each `ready`, the probe takes it with a nonblocking `flock` on a close-on-exec descriptor, and the runner holds it through the poll and any Job the poll brings. It is released on `no_job`, `update`, a receive timeout, a lost connection, a `running` that fails to send, and when the Job ends, before its result goes out. - **A marker** is a file in `/run/bencher` named `maintenance` or `maintenance.<name>`. Maintenance creates its own before it waits for the lock and removes it when done. - **The probe pauses for maintenance:** - while a marker exists, the runner sends `paused` with a `maintenance` reason and `marker` true, and leaves the lock free; - while another process holds the lock, it sends a `maintenance` reason with `lock` true; - only a probe that finds no reason to pause keeps the lock, so a runner paused for a RAID sync leaves it free too. - **The plain `maintenance` marker drains a runner.** It finishes its current poll and any Job that poll brings, then takes no Job until the marker is gone. A unit uses a marker of its own, so it never removes the drain. - **`runner run`** waits for the job lock, trying it every 250 ms and logging `Waiting for host maintenance to finish` once, and a stop ends the wait. It warns but runs through a RAID sync (`Running during a RAID sync`) or a marker (`Running although host maintenance is due`). - **`/run/bencher` is root's alone:** the runner now creates it 0700 rather than 0755, for the runner lock and the job lock alike, and each lock file 0600. A directory an older runner created keeps its mode until a reboot or a drop-in's `install -d -m 0700`. - **A runner that is not root** can open neither the lock nor the 0700 directory, so it takes Jobs without them, and `runner up` warns once at startup that maintenance may run during a Job. A root runner that cannot open or take the lock pauses with `lock` true rather than run unguarded. - **`bencher_runner::maintenance`**, public and outside the `plus` feature, holds the directory, lock, and marker names and `marker_present`, which the probe and `runner run` both use, so host tooling can share them. - **Docs:** a "Host Maintenance" section on the self-hosted runners page in all 9 languages: the lock, the markers, an example `fstrim` drop-in, and how to drain a Runner. The changelog gains a line. Across versions: - Like the RAID pause in the layer below, this needs a server with #1122. A server without it closes the channel on `paused`, and the runner reconnects no faster than once per poll. - A runner without this change never takes the lock and ignores the markers, so a unit that waits on the lock runs at once, as it does today. ## Why Package upgrades, `fstrim`, and the monthly RAID check run on timers that know nothing about Jobs, so any of them can land during a Job and move its measurement. The runner takes the lock before `ready` rather than when a Job arrives. Taken when a Job arrives, the lock could already be maintenance's, which would hold a Job the server has handed out while that Job's clock ran. Taken before `ready`, maintenance waits at most one poll plus one Job. A waiting `flock` holds nothing the runner can see, and on each release the runner's own nonblocking try can win the lock back. The marker tells the runner that maintenance waits, so it steps aside and the waiter gets the lock at the next release. Markers of their own let two units overlap, such as a scrub and the morning upgrade: with one shared marker, the first to finish would remove it while the second still waits, and any unit would remove an operator's drain. `flock(2)` locks whatever mode a file was opened in, so any user who can open the lock file can hold it and keep the runner paused. A root-only directory and 0600 files keep that to root. The descriptor is close-on-exec, so the jailer, the VMM, and every other child never keep the lock past the Job. `runner run` is started by hand at a moment someone chose, so it warns about a sync or a marker and runs, but it never runs beside maintenance that holds the lock. Trying the lock in a loop, rather than blocking in `flock(2)`, lets a stop end the wait at once. ## Verification - Lock tests use two independent opens of a file in a temporary directory, which `flock` treats as two holders even in one process. Each new test fails under a mutation of the behavior it guards: - `a_free_lock_is_the_runners_turn` and `a_lock_held_elsewhere_pauses_for_maintenance`: a turn that never takes the lock, and a held lock taken as free. - `a_marker_pauses_without_taking_the_lock`, with `maintenance` and `maintenance.fstrim.service`: a marker ignored, the lock kept past a marker (which starves waiting maintenance), and only the plain marker counted. - `only_the_maintenance_names_are_markers`: any file, or any name starting `maintenance`, read as a marker. - `dropping_the_turn_releases_the_lock` and `a_child_never_holds_the_lock`: a leaked descriptor, and one that is not close-on-exec. - `the_lock_and_its_directory_are_root_only`: `job.lock` at 0644 or 0640, and `/run/bencher` at 0755 or 0750. - `a_runner_that_is_not_root_takes_its_turn_without_the_lock`: a runner that is not root failing closed, and a root runner running on without the lock. - `runner_run_waits_for_maintenance_then_takes_the_lock` and `a_stop_ends_the_wait_for_maintenance`: a `runner run` that does not wait, gives up once maintenance holds the lock, or keeps waiting past a stop. - `the_job_lock_is_held_from_ready_until_the_job_finishes` and `the_job_lock_is_released_whenever_no_job_follows`: a release when the Job arrives, none when it ends, and none on `no_job`, `update`, a receive timeout, a lost connection, or a failed `running` send. - `the_job_lock_is_kept_only_by_a_clear_probe`: the lock kept through a pause, or dropped by a clear probe. - `the_driver_releases_the_job_lock_when_told`: a lock the probe took must be free to a second open after the driver runs the release. Kills a driver that ignores it, which would hold maintenance off for a whole server outage, since no probe runs until a connect succeeds. - None of the new tests compiles on the parent, which has no job lock. The 14 tests pass 20 of 20 as an ordinary user and as root. - **In a disposable VM**, a root `runner up` at a 5 s poll ran against a local API, with `/run/bencher` absent as after a boot: - The runner created `/run/bencher` 0700 and `job.lock` 0600. - A unit shaped like the next layer's drop-in, started 1 s into a 20 s Job, created its marker at once and waited. Its command started 1 ms after the runner released the lock at the Job's end, while the runner paused with `marker` and `lock` true. A second Job, submitted during the first, was claimed after the maintenance. - `systemctl kill -s KILL` of such a unit, while it waited and while it ran, left the unit `failed` and its marker gone, and the runner's next probe ended its pause. - `touch /run/bencher/maintenance` paused the runner once its current poll ended, with `marker` true and `lock` false. A Job submitted then read `pending` for 30 s and was claimed within 2 s of the `rm`. - **`runner run` behind a held lock**, in the same VM, sandboxed with Firecracker, while another process held `job.lock` for 20 s: - It logged `Waiting for host maintenance to finish` once, started no VM until the release, then ran its benchmark. - While Firecracker ran, `job.lock` was open only in the runner, and `/proc/locks` showed the runner's pid holding it. - SIGTERM during the wait ended the run in 0.24 s, with no VM started and the holder's lock untouched. - **The scenario `runner_run_takes_its_turn_at_the_job_lock`** puts both halves of `runner run`'s turn in CI, in about 7 s: - The harness holds `job.lock` and starts a sandboxed `runner run`. The scenario requires one `Waiting for host maintenance to finish`, then no jail for the 2 s until the release, then the Job. Once the VM is up, neither Firecracker nor the jailer it came from may hold the lock. - The harness then creates `/run/bencher/maintenance` and runs the image on the host. The scenario requires one `Running although host maintenance is due`, no wait, and a successful Job whose benchmark, the runner's own child, lists no descriptor of `job.lock`. The host run carries the close-on-exec check, since the jailer closes every descriptor it inherits. - The marker and the harness's lock are gone on every path. - It fails when `runner run` never waits, when its descriptor is not close-on-exec, and when a marker makes it refuse, wait, or run without the warning. It fails on the parent, which has no job lock, and passes 20 of 20 as root. - The runner unit tests pass as an ordinary user and as root, 441 of 441, the CLI tests pass, 14 of 14, the `runner_ops` tests pass, 86 of 86, `cargo test-runner scenarios --build-only` passes, `Cargo.lock` is unchanged, and clippy passes with `-D warnings` for Linux x86_64, Linux aarch64, and macOS on the runner, its CLI, `test_runner`, and `runner_ops`. The scenario suite passes as root, 100 of 100, and clippy passes for the whole workspace as CI runs it, on the top of this stack.
epompeii
added a commit
that referenced
this pull request
Oct 10, 2026
## What Make the runner send its host's disk health with every `ready` and `paused`, in the `health` field that #1125 stores, so admins see it on the runner endpoints and #1130 mails them when it is news. - **md, at every probe:** each array's `array_state`, `degraded`, and `sync_action`, read with the RAID pause's own read. An array with no sync thread, such as raid0, reads `idle`. A degraded array is a `degraded` finding that is `failing`. - **NVMe, at most hourly:** each controller under `/sys/class/nvme`, with its model, its serial, and the SMART / health log fields that set a finding: critical warning, temperature, spare and its threshold, percentage used, and media errors. The page is read with Get Log Page (log `0x02`, all namespaces, 512 bytes) through the kernel's admin passthrough, `NVME_IOCTL_ADMIN_CMD`, on a read-only open of `/dev/<controller>`. The read is cached for an hour, and only the probe reads it, so never during a Job. - **States:** - `warning`: spare at or below twice its threshold, percentage used over 80, the temperature critical warning bit, or any media and data integrity error; - `failing`: any other critical warning bit, spare below its threshold, or a degraded array; - the host's state is its worst finding, and `ok` with none. - **A runner that is not root** cannot use the passthrough, which needs `CAP_SYS_ADMIN`. It lists each controller with no log, warns once with `NVMe health not read, so it is reported unreadable`, and still reports md. - Health is always sent: a host with no arrays and no controllers reports `ok` with empty lists. - **Nothing pauses on health.** A degraded array keeps taking Jobs; only a sync pauses, as before. - `AdminCommand` is the kernel's `struct nvme_admin_cmd`, pinned at 72 bytes by a const assertion. `nix` gains its `ioctl` feature, which adds no package to `Cargo.lock`. - The changelog gains a line. Across versions: - A server without #1125 ignores `health`, since the runner messages accept unknown fields. - `paused` needs a server with #1122, as in the layers below. ## Why A degraded array or a failing NVMe drive on a runner host is silent until the next disk fails too. The server half stores health and mails admins on news; this is what sends it. md costs a few sysfs reads, so it is read at every probe. Each NVMe read is an admin command to the drive, so it runs at most hourly, and only between Jobs. The temperature bit is a warning, not failing, because it clears by itself. Media errors warn while their count is nonzero, rather than when it rises, which would read `ok` again at the next hourly read. Health is reported, never a pause. An idle degraded array barely moves a benchmark, and the mail reaches whoever can replace the disk. Health still rides on `paused`, because the rebuild after a disk is replaced is a pause, and that is when health matters most. ## Verification - Each new test fails under a mutation of the rule it guards: - The NVMe parser runs on fixture pages. `a_page_parses_little_endian_at_its_offsets` and `a_media_error_count_past_u64_saturates`: swapped temperature bytes, media errors read big-endian or from the wrong counter, spare read from the threshold byte, and a 128-bit count that wraps. - `a_healthy_page_has_no_findings` and `each_critical_warning_bit_but_temperature_is_failing`, over bits 0 and 2 to 7: a rule that fires on a healthy page, any one critical bit missed, and the temperature bit failing or ignored. - `spare_at_twice_its_threshold_warns_and_below_it_fails`, `percentage_used_over_80_warns`, and `any_media_error_warns`: either spare boundary or the endurance boundary moved by one, and media errors ignored or failing. - `an_unreadable_log_still_lists_the_controller`: an unreadable controller dropped, and a log read from another device. - `a_degraded_array_is_failing` and `the_worst_finding_sets_the_state`: a degraded array that reads healthy, every array failing, and the state taken from the first finding rather than the worst. - `nvme_is_read_at_most_hourly`, on an injected clock and reader: read once at the start, served from the cache 1 s before the hour, and read again at the hour. Kills a read every probe, and one never repeated. - `an_unreadable_nvme_still_reports_md` and `every_probe_reports_md_health`: a runner that is not root reporting nothing, and a probe that sends no health, or none with the RAID pause off. - `the_probes_health_rides_on_ready_and_paused`: health dropped from either message. - None of the new tests compiles on the parent, which reads no health. The 14 tests pass 20 of 20 as an ordinary user and as root. - A C probe against Ubuntu 24.04's `linux/nvme_ioctl.h` gives `sizeof(struct nvme_admin_cmd)` 72, `NVME_IOCTL_ADMIN_CMD` `0xc0484e41`, and `addr`, `data_len`, and `cdw10` at offsets 24, 36, and 40, which the Rust struct and `ioctl_readwrite!(b'N', 0x41)` match. - **In a disposable VM**, an NVMe controller on the kernel's `nvme-loop` target, backed by a file, stood in for a drive: - As root, the runner's own read returned the page with status 0, and the page's data units read and host read commands equal `nvme smart-log`'s values for that controller. - As an ordinary user, the read fails with `EACCES`, the controller is listed with no log, and md is still reported. - A dword count off by one fails there with status `0x600f`, which the fixture pages cannot show. - **In the same kind of VM**, a root `runner up` ran against a local API, with a RAID1 over two loop devices: - Healthy, the server stored health `ok` with `md0` `clean`, 0 degraded, `idle`. - `mdadm --fail` on one member: the next report read `failing`, with `md0` `degraded` `failing`, and the runner still `ready`. A Job submitted then ran. - `mdadm --remove` and `--add`: the runner paused for the `recover` sync with health still `failing`. Once the recovery finished, the pause ended, and the next report read `ready` and `ok`. - The runner unit tests pass as an ordinary user and as root, 455 of 455, the CLI tests pass, 14 of 14, the `runner_ops` tests pass, 86 of 86, `cargo test-runner scenarios --build-only` passes, `Cargo.lock` is unchanged, and clippy passes with `-D warnings` for Linux x86_64, Linux aarch64, and macOS on the runner, its CLI, `test_runner`, and `runner_ops`. The scenario suite passes as root, 100 of 100, and clippy passes for the whole workspace as CI runs it, on the top of this stack.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Add a
pausedrunner message, so an idle runner that should take no Job says so, says why, and stays connected.pausedcarriesreasons, each tagged bykind, plussince, when the pause began by the runner's clock, and everyreadyfield.raid: the array, itssync_action, and its progress fromsync_completedandsync_speed.maintenance: whether the maintenance marker is present and whether another process holds the job lock.pausedas it answersready, without the claim. It runs the same self-update check, holds the poll until the runner'spoll_timeout, and answersno_job. The hold is bounded by the poll timeout, not the heartbeat timeout, and the nextreadyon the same channel claims as usual.pausedis warned about and acked likeready, and it does not reset the heartbeat.other, so a newer runner's message still parses. A message that fails to parse closes the channel.Runner pause began, with the reasons andsince, andRunner pause endedat the transitions only. Therunner.stategauge gainspaused, and a newrunner.pausecounter counts each pause start once per reason kind, under apause.reasonattribute.RunnerMessage::Readybecomes a newtype overJsonReady, with an unchanged wire format, andpausedflattens the same struct. Callers build it withJsonReady::new, so a later field touches no call site.pausedin all 9 languages.Across versions:
pausedyet. This is the server half, and it must be live before any runner that sends it.paused, so it closes the channel and the runner reconnects.paused, and itsreadyreads exactly as before.Why
An idle runner that should take no Job, because an md array is resyncing or scrubbing or host maintenance is about to run, has two choices today. It can send
readyand risk a Job measured while its disks are busy, or it can go quiet, which the server ends at its idle timeout by closing the channel, and self-update goes with it.pausedkeeps the runner connected and updating at its usual poll cadence, and it takes Jobs again within one poll of its reasons clearing.Holding the poll, rather than answering
no_jobat once, reuses the timersreadyalready has. The runner needs no sleep of its own, and nothing has to stay under a server heartbeat timeout that the runner cannot see.The reasons in each message carry the detail, and the counter shows a runner that pauses often.
Verification
channel_paused_claims_no_job: a paused runner with a matching pending Job getsno_jobno sooner than its poll, the Job stays pending, andreadyon the same channel then claims it. Kills claiming on a pause, and answeringno_jobat once.channel_paused_stale_runner_gets_updateandpaused_stable_version_mismatch_sends_update: a stale stable runner that pauses getsupdate, with the checksum a test listener serves. Kills skipping the update check on a pause.channel_pause_outlasts_the_heartbeat_timeout: a 6 s pause against a 5 s heartbeat timeout, then a 1 s pause, thenready, each answeredno_jobon one channel. Kills a hold bounded by the heartbeat timeout.paused_reasons_are_cappedandpaused_reason_strings_are_capped: 17 reasons become 16, and a 1,000 byte array name and sync action become 64 bytes. Kills reading the reasons uncapped.channel_paused_during_job_does_not_reset_heartbeat: apausedduring a Job is warned about and acked, and the Job still times out on its heartbeat. Kills apausedthat resets the heartbeat.ready_wire_format_is_unchangedpins today's fullreadyJSON both ways,paused_wire_format_is_pinnedpinspaused, andunknown_pause_reason_is_otherandunknown_sync_action_is_otherparse unknown values. Kills renaming areadyfield, dropping the flatten, and dropping either catch-all.readyis byte-identical: bare, withpoll_timeout, with runner metadata, and with a canary channel and checksum in both key orders. So areheartbeat,running, andcanceled. The previous server rejectspausedwithunknown variant 'paused'and closes the channel.pausedwith a 5,000 byte array name and sync action and found no log line carrying 65 or more bytes of either.bencher_json,bencher_schema,bencher_otel,api_runners, andbencher_runnerpass.cargo gen-specand typeshare reproduce the committed files, and clippy passes with-D warningsfor Linux x86_64, Linux aarch64, and macOS.