Skip to content

Run task containers at the lowest priority so they can't destabilize the host - #430

Merged
tildesrc merged 1 commit into
mainfrom
panopticon/container-priority
Sep 23, 2026
Merged

tildesrc merged 1 commit into
mainfrom
panopticon/container-priority

Conversation

@tildesrc

Copy link
Copy Markdown
Contributor

A task container is where unbounded work happens — an agent running a repo's whole test suite, a build inside a dind task, six tasks at once — on the same machine as the operator's editor, their shell, and panopticon's own control plane. Every container is now spawned deprioritized: --cpu-shares 2 and --blkio-weight 10 (docker's floors, so a task loses CPU and disk races against normally-weighted processes) plus --oom-score-adj 500 (under real memory pressure the kernel kills a task before anything of the operator's). This is priority, not a cap — nothing limits what a task may use when nobody else wants it, so an idle host still runs tasks at full speed. A task that does get OOM-killed already explains itself: exit_reason reads OOMKilled off docker inspect before the exit code.

Plan: the task's plan.md artifact.

Two wrinkles the new sessionservice/priority module exists for

docker exec does not inherit the container's OOM score. oom_score_adj is a per-process attribute inherited across fork, and the agent pane forks from the Docker daemon rather than the container's PID 1 — so docker run --oom-score-adj covers the entrypoint and everything it starts but misses the one process that actually eats memory. docker exec has no flag for it, so the pane's command raises its own score and then execs the launcher (raising needs no privilege; only lowering does), keeping the process tree and signal behaviour identical. The cgroup levers need no such trick — an exec'd process joins the container's cgroup.

Some daemons reject the cgroup flags outright. A nested Docker daemon whose cgroup is in threaded mode answers --cpu-shares/--blkio-weight with unable to apply cgroup configuration, which would otherwise break every spawn on such a host. A failed flagged run is retried once without those flags (keeping --oom-score-adj, which needs no controller), and the answer is latched so the discovery costs one extra run per process, with one warning. The latch is set only if the retry succeeds, so an ordinary failure — a bad image, say — still surfaces and doesn't quietly disable priority control for later spawns.

Also in here

  • The workspace-cleanup sweep (an unbounded delete over a whole checkout) runs deprioritized and degrades the same way; it can run before any spawn has discovered the daemon's answer.
  • A shell workflow's session runs on the host with no cgroup around it, so it gets nice -n 19.
  • Every value is a per-host PANOPTICON_CONTAINER_* / PANOPTICON_HOST_NICE env knob over the shipped default, and any of them set to off emits exactly the argv panopticon emitted before this existed. Bad values warn and fall back rather than failing a spawn; values are clamped to what docker and the kernel accept, and a negative OOM adjustment — shielding a task at the host's expense — is refused.
  • An opt-in --cgroup-parent knob, unset by default: with docker's systemd cgroup driver, weights are compared between siblings, so weight 1 on a container's scope deprioritizes it within system.slice but doesn't by itself make it lose to the operator's user.slice. docs/container.md shows the low-weight slice for operators who want that, rather than panopticon installing a systemd unit itself.

…the host

A task container is where unbounded work happens — an agent running a repo's
whole test suite, a build inside a dind task, six tasks at once — on the same
machine as the operator's editor and panopticon's own control plane. Spawn every
one of them deprioritized: `--cpu-shares 2` and `--blkio-weight 10` (docker's
floors, so a task loses CPU and disk races against normally-weighted processes)
plus `--oom-score-adj 500` (the kernel kills a task before anything of the
operator's). Priority, not a cap — an idle host still runs tasks at full speed.

Two wrinkles the new `sessionservice/priority` module exists for:

- `docker exec` does not inherit the container's OOM score. `oom_score_adj` is
  per-process and inherited across *fork*, and the agent pane forks from the
  Docker daemon rather than the container's PID 1 — leaving the one process that
  actually eats memory at the host's default score. `docker exec` has no flag for
  it, so the pane's command raises its own score before `exec`ing the launcher
  (raising needs no privilege; only lowering does). The cgroup levers need no
  such trick — an exec'd process joins the container's cgroup.
- Some daemons reject the cgroup flags outright: a nested daemon whose cgroup is
  in threaded mode answers any of them with `unable to apply cgroup
  configuration`. A spawn must never be lost to a host that merely can't
  deprioritize, so a failed flagged run is retried once without those flags
  (keeping `--oom-score-adj`, which needs no controller) and the answer is
  latched, with one warning. The latch is set only if the retry *succeeds*, so an
  ordinary failure still surfaces and doesn't disable the flags.

The workspace-cleanup sweep runs deprioritized and degrades the same way, and a
shell workflow's session — which runs on the host with no cgroup around it — gets
`nice -n 19`. Every value is a per-host `PANOPTICON_CONTAINER_*` env knob over the
shipped default, switchable to `off` to emit exactly the pre-priority argv; bad
values warn and fall back rather than failing a spawn, and a negative OOM
adjustment (shielding a task at the host's expense) is refused. An opt-in
`--cgroup-parent` knob is there for operators who want enforcement at the
hierarchy level, since with docker's systemd cgroup driver weights compare
between siblings and `user.slice` is not one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@tildesrc
tildesrc marked this pull request as ready for review September 23, 2026 21:53
@tildesrc
tildesrc merged commit 4c27472 into main Sep 23, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant