Skip to content

Add a runner ops audit that compares runners - #1124

Merged
epompeii merged 1 commit into
u/ep/runner/ops-runner-keyfrom
u/ep/runner/ops-audit
Oct 9, 2026
Merged

epompeii merged 1 commit into
u/ep/runner/ops-runner-keyfrom
u/ep/runner/ops-audit

Conversation

@epompeii

@epompeii epompeii commented Oct 7, 2026

Copy link
Copy Markdown
Member

What

Add cargo ops audit <runner> [--against <reference-runner>], which takes a read-only snapshot of each runner over one SSH connection and checks it.

  • Snapshot: 20 sections, each filled by one read-only command, with its exit status. 17 are compared across runners: OS, kernel, kernel packages, kernel cmdline, hardware, SMT, RAID, the runner binary's sha256, the runner unit, sshd, apt, GRUB, ufw, enabled units, timezone, installed packages, and manually installed packages. The runner service, a pending reboot, and auto-removable packages feed only the health checks. The audit prints how many sections and lines it read.
  • Health checks: every section has output and its command succeeded, the runner service is active, no reboot is pending, RAID arrays are active and neither syncing nor degraded, no packages are auto-removable, and the isolation boot args are set.
  • Diff: with --against, it prints the differing lines of each section, lines that are reordered or repeated, and non-zero exit statuses, after normalizing what must differ: the runner name, the root UUID, and the RAID member order.
  • The key stays on the runner: the runner unit is read through an allowlist of its lines, never systemctl cat, and any line holding a runner key is dropped on the server and again on parse.
  • It exits non-zero on a failed health check or any difference.
  • AGENTS.md documents the audit and the steps to add a runner, which takes no Jobs until its RAID1 initial sync completes.

Why

Runners drift apart in ways that change benchmark results or break Jobs: a newer kernel installed by unattended upgrades and waiting on a reboot, a package that autoremove took from one runner, a missing boot arg, or a RAID array that is degraded or still syncing. Without the audit, comparing runners means reading each one by hand. The audit only reads, so it is safe to run against a live runner.

Verification

  • runner_unit_command_reads_only_allowlisted_lines runs the real section command with sh against a fixture unit and drop-in, and asserts the exact output: the allowlisted lines stay, and the key lines are dropped, plain, quoted, and as ExecStart ... --key. It fails if the key filter matches nothing or the allowlist matches every line.
  • normalizer_strips_runner_key fails if the parse stops dropping lines that hold the key prefix.
  • The health checks are tested for each failing state: an inactive service, a pending reboot, a resync, a recovery, a failed member, an inactive array, an unreadable /proc/mdstat, an empty or failed section, auto-removable packages, and missing isolation args. The diff tests ignore normalized values and package order, and report differing, reordered, and repeated lines and exit statuses.
  • The runner_ops unit tests pass in 20 of 20 runs, cargo ops --help and every subcommand's --help run, and clippy passes with -D warnings for Linux x86_64, Linux aarch64, and macOS.

@epompeii
epompeii added this pull request to stack #1115 October 7, 2026 14:56
@github-actions

github-actions Bot commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

🐰 Bencher Report

ProjectBencher
Branchu/ep/runner/ops-audit
Testbedintel-v1
Click to view all benchmark results
BenchmarkLatencyBenchmark Result
microseconds (µs)
(Result Δ%)
Upper Boundary
microseconds (µs)
(Limit %)
Adapter::Json📈 view plot
🚷 view threshold
5.19 µs
(-0.09%)Baseline: 5.20 µs
5.83 µs
(89.07%)
Adapter::Magic (JSON)📈 view plot
🚷 view threshold
5.02 µs
(+0.20%)Baseline: 5.01 µs
5.54 µs
(90.66%)
Adapter::Magic (Rust)📈 view plot
🚷 view threshold
24.39 µs
(-8.86%)Baseline: 26.77 µs
31.66 µs
(77.05%)
Adapter::Rust📈 view plot
🚷 view threshold
4.47 µs
(-2.71%)Baseline: 4.60 µs
5.86 µs
(76.40%)
Adapter::RustBench📈 view plot
🚷 view threshold
4.47 µs
(-2.73%)Baseline: 4.59 µs
5.85 µs
(76.33%)
🐰 View full continuous benchmarking report in Bencher

@epompeii
epompeii marked this pull request as ready for review October 8, 2026 04:20
- `cargo ops audit <runner> [--against <reference-runner>]` collects a
  read-only snapshot of each runner over one SSH connection, with each
  section's exit status, and prints how many sections and lines it
  audited
- Health checks: every section has output and its command succeeded,
  the runner service is active, no reboot is pending, RAID arrays are
  active and neither syncing nor degraded, no packages are
  auto-removable, and the isolation boot args are set
- The diff of the normalized snapshots reports differing lines,
  reordered or repeated lines, and exit statuses per section
- The runner unit is read through an allowlist of its lines, never
  `systemctl cat`, and any line holding a runner key is dropped on the
  server and again on parse
- Exits non-zero on a failed health check or any difference
- Document adding a runner, which takes no jobs until its RAID1 initial
  sync completes
@epompeii
epompeii marked this pull request as draft October 9, 2026 09:17
@epompeii
epompeii force-pushed the u/ep/runner/ops-audit branch from d043375 to 63046d5 Compare October 9, 2026 09:18
@epompeii
epompeii marked this pull request as ready for review October 9, 2026 14:22
@epompeii
epompeii merged commit 202ed17 into devel Oct 9, 2026
52 of 53 checks passed
@epompeii
epompeii deleted the u/ep/runner/ops-audit branch October 9, 2026 14:53
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant