Run task containers at the lowest priority so they can't destabilize the host - #430
Merged
Merged
Conversation
…the host A task container is where unbounded work happens — an agent running a repo's whole test suite, a build inside a dind task, six tasks at once — on the same machine as the operator's editor and panopticon's own control plane. Spawn every one of them deprioritized: `--cpu-shares 2` and `--blkio-weight 10` (docker's floors, so a task loses CPU and disk races against normally-weighted processes) plus `--oom-score-adj 500` (the kernel kills a task before anything of the operator's). Priority, not a cap — an idle host still runs tasks at full speed. Two wrinkles the new `sessionservice/priority` module exists for: - `docker exec` does not inherit the container's OOM score. `oom_score_adj` is per-process and inherited across *fork*, and the agent pane forks from the Docker daemon rather than the container's PID 1 — leaving the one process that actually eats memory at the host's default score. `docker exec` has no flag for it, so the pane's command raises its own score before `exec`ing the launcher (raising needs no privilege; only lowering does). The cgroup levers need no such trick — an exec'd process joins the container's cgroup. - Some daemons reject the cgroup flags outright: a nested daemon whose cgroup is in threaded mode answers any of them with `unable to apply cgroup configuration`. A spawn must never be lost to a host that merely can't deprioritize, so a failed flagged run is retried once without those flags (keeping `--oom-score-adj`, which needs no controller) and the answer is latched, with one warning. The latch is set only if the retry *succeeds*, so an ordinary failure still surfaces and doesn't disable the flags. The workspace-cleanup sweep runs deprioritized and degrades the same way, and a shell workflow's session — which runs on the host with no cgroup around it — gets `nice -n 19`. Every value is a per-host `PANOPTICON_CONTAINER_*` env knob over the shipped default, switchable to `off` to emit exactly the pre-priority argv; bad values warn and fall back rather than failing a spawn, and a negative OOM adjustment (shielding a task at the host's expense) is refused. An opt-in `--cgroup-parent` knob is there for operators who want enforcement at the hierarchy level, since with docker's systemd cgroup driver weights compare between siblings and `user.slice` is not one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
tildesrc
marked this pull request as ready for review
September 23, 2026 21:53
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A task container is where unbounded work happens — an agent running a repo's whole test suite, a build inside a dind task, six tasks at once — on the same machine as the operator's editor, their shell, and panopticon's own control plane. Every container is now spawned deprioritized:
--cpu-shares 2and--blkio-weight 10(docker's floors, so a task loses CPU and disk races against normally-weighted processes) plus--oom-score-adj 500(under real memory pressure the kernel kills a task before anything of the operator's). This is priority, not a cap — nothing limits what a task may use when nobody else wants it, so an idle host still runs tasks at full speed. A task that does get OOM-killed already explains itself:exit_reasonreadsOOMKilledoffdocker inspectbefore the exit code.Plan: the task's
plan.mdartifact.Two wrinkles the new
sessionservice/prioritymodule exists fordocker execdoes not inherit the container's OOM score.oom_score_adjis a per-process attribute inherited across fork, and the agent pane forks from the Docker daemon rather than the container's PID 1 — sodocker run --oom-score-adjcovers the entrypoint and everything it starts but misses the one process that actually eats memory.docker exechas no flag for it, so the pane's command raises its own score and thenexecs the launcher (raising needs no privilege; only lowering does), keeping the process tree and signal behaviour identical. The cgroup levers need no such trick — an exec'd process joins the container's cgroup.Some daemons reject the cgroup flags outright. A nested Docker daemon whose cgroup is in threaded mode answers
--cpu-shares/--blkio-weightwithunable to apply cgroup configuration, which would otherwise break every spawn on such a host. A failed flagged run is retried once without those flags (keeping--oom-score-adj, which needs no controller), and the answer is latched so the discovery costs one extra run per process, with one warning. The latch is set only if the retry succeeds, so an ordinary failure — a bad image, say — still surfaces and doesn't quietly disable priority control for later spawns.Also in here
nice -n 19.PANOPTICON_CONTAINER_*/PANOPTICON_HOST_NICEenv knob over the shipped default, and any of them set tooffemits exactly the argv panopticon emitted before this existed. Bad values warn and fall back rather than failing a spawn; values are clamped to what docker and the kernel accept, and a negative OOM adjustment — shielding a task at the host's expense — is refused.--cgroup-parentknob, unset by default: with docker's systemd cgroup driver, weights are compared between siblings, so weight 1 on a container's scope deprioritizes it withinsystem.slicebut doesn't by itself make it lose to the operator'suser.slice.docs/container.mdshows the low-weight slice for operators who want that, rather than panopticon installing a systemd unit itself.