Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
13 changes: 12 additions & 1 deletion AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -42,7 +42,11 @@ src/panopticon/
# workflows; the spawner routes on it, skipping the image + the clone unless the
# workflow opts in via clone_repo); images.py = ADR-0005 composed images
# (base→workflow→repo); provisioner.py = host-side provisioning
# (ADR 0011: branch the per-task clone on slug, record it back); clones.py =
# (ADR 0011: branch the per-task clone on slug, record it back); priority.py =
# host-side resource priority (env → argv: the deprioritizing `docker run` flags
# every task container gets — cpu/blkio weight floors + a raised OOM score — the
# agent pane's own oom_score_adj, `nice` for shell tasks, and the cgroup-flag
# strip the runner degrades through on a daemon that refuses them); clones.py =
# per-repo clone cache; spawn.py = spawn-prep (clone --local the per-task
# checkout, mounted rw at /workspace; point origin at the forge, then init any
# submodules — in that order, since relative .gitmodules URLs resolve against
Expand Down Expand Up @@ -251,6 +255,13 @@ on every PR (the same commands the Makefile wraps).
unresumable-transcript recognizer: the SDK marker shapes captured off the two tasks this broke,
the newest-first resume order, and the regression that an SDK-only project now yields a
first-run argv instead of the `--continue` that killed the pane.
- `tests/sessionservice/test_priority.py` — resource priority (env → argv): the shipped defaults
every task container is spawned with, each per-host override, the `off` switch that returns the
argv to its pre-priority form, the clamps (a negative OOM adjustment refused), a bad value falling
back rather than failing a spawn, and the cgroup-flag strip. `test_local_runner.py` covers the
wiring — the flags on `docker run`, the pane wrapper that raises the exec'd agent's OOM score
(`docker exec` doesn't inherit the container's), and the retry-without-cgroup-flags fallback plus
its latch; `test_shell_runner.py` covers the `nice` prefix on a shell task's host session.
- `tests/test_prefill.py` — the input-box prefill poller: unit tests drive `prefill_pane` with a
fake tmux runner + injected `sleep`/raw-log — pin the `pipe-pane`/`load-buffer`/`paste-buffer -p`
commands when the box becomes ready, and every best-effort give-up (empty prompt, timeout,
Expand Down
67 changes: 67 additions & 0 deletions docs/container.md
Original file line number Diff line number Diff line change
Expand Up @@ -116,6 +116,73 @@ For what a slug, branch, clone, and `provisioned` mean as task concepts, see [ta
- **Auth.** The agent authenticates from a `CLAUDE_CODE_OAUTH_TOKEN` injected from the repo's
env-file — see [auth](auth.md).

## Resource priority: tasks yield to you

A task container is where unbounded work happens — an agent running a repo's whole test suite, a
`docker build` inside a dind task, six tasks at once — on the same machine as your editor, your
shell, and panopticon's own control plane. So every container is spawned **deprioritized**: it
loses CPU and disk races against normally-weighted processes, and under real memory pressure the
kernel kills it before anything of yours.

This is *priority*, **not a cap**. Nothing limits how much CPU or RAM a task may use when nobody
else wants it, so on an idle host tasks run at full speed.

Concretely, each container gets `--cpu-shares 2` (docker's floor — cgroup v2 `cpu.weight` 1),
`--blkio-weight 10` (the floor; disk contention is what actually makes a desktop stutter) and
`--oom-score-adj 500`. The agent pane raises its own OOM score too: `oom_score_adj` is
per-process and inherited across *fork*, and the pane is started with `docker exec` (forked from
the Docker daemon, not from the container's PID 1), so the flag on `docker run` would otherwise
miss the one process that actually eats memory. The cgroup levers need no such trick — an exec'd
process joins the container's cgroup. The workspace-cleanup sweep runs deprioritized as well.

If a task does get OOM-killed you'll see it: the runner reads `OOMKilled` off `docker inspect`
before the exit code, so the task shows `failed` with an out-of-memory detail rather than a
mystery.

### Tuning it per host

The knobs are read from the **runner host's** environment, so a big build box and a laptop can
differ. Set any of them to `off` (or empty) to drop that flag entirely — the argv is then exactly
what panopticon emitted before any of this existed.

| Variable | Default | What it sets |
|---|---|---|
| `PANOPTICON_CONTAINER_CPU_SHARES` | `2` | CPU weight (`--cpu-shares`) |
| `PANOPTICON_CONTAINER_BLKIO_WEIGHT` | `10` | block-IO weight (`--blkio-weight`) |
| `PANOPTICON_CONTAINER_OOM_SCORE_ADJ` | `500` | OOM-killer preference, for the container *and* the agent pane |
| `PANOPTICON_CONTAINER_CGROUP_PARENT` | unset | parent cgroup (`--cgroup-parent`) — see below |
| `PANOPTICON_HOST_NICE` | `19` | `nice` for a shell task's host session (no container to weight) |

A bad value warns and falls back to the default instead of failing the spawn, and values are
clamped to what docker and the kernel accept (a *negative* OOM adjustment — shielding a task at
the host's expense — is refused).

**Hosts that refuse the flags.** Some daemons reject `--cpu-shares`/`--blkio-weight` outright: a
nested Docker daemon whose cgroup is in threaded mode answers any of them with `unable to apply
cgroup configuration`. A spawn is never lost to that — the runner retries once without the
cgroup-backed flags (keeping `--oom-score-adj`, which needs no controller), logs one warning, and
remembers the answer for the rest of its life.

**Going further with a slice.** With docker's systemd cgroup driver, containers land in
`system.slice/docker-<id>.scope` while your own processes sit in `user.slice`, and cgroup weights
are compared **between siblings** — so weight 1 on a container's scope deprioritizes it *within*
`system.slice` but doesn't by itself make it lose to `user.slice`. If you want that enforced at the
hierarchy level, create a low-weight slice and point panopticon at it:

```ini
# /etc/systemd/system/panopticon.slice
[Slice]
CPUWeight=1
IOWeight=1
```

```sh
sudo systemctl daemon-reload
export PANOPTICON_CONTAINER_CGROUP_PARENT=panopticon.slice # on the runner host
```

Installing a unit is your call, so per-container weights stay the zero-setup default.

## When it goes wrong

A container can disappear out from under a live task — an OOM kill, a host reboot, a `docker rm`.
Expand Down
91 changes: 88 additions & 3 deletions src/panopticon/sessionservice/local_runner.py
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,7 @@

from __future__ import annotations

import logging
import os
import shlex
import subprocess
Expand All @@ -21,8 +22,17 @@
from panopticon.core.dirs import credential_dir_path, secrets_file_path
from panopticon.core.features import CODEX_FLAG, codex_enabled
from panopticon.core.models import DEFAULT_AGENT_CLI, LifecyclePhase
from panopticon.sessionservice.priority import (
BLKIO_WEIGHT_VAR,
CPU_SHARES_VAR,
container_oom_score_adj,
container_priority_flags,
strip_cgroup_flags,
)
from panopticon.sessionservice.runner import Runner

_log = logging.getLogger(__name__)

#: The container home the per-CLI config dir lives under (the base image's ``panopticon`` user).
CONTAINER_HOME = "/home/panopticon"

Expand Down Expand Up @@ -130,6 +140,29 @@ def _invoking_user() -> str:
return f"{os.getuid()}:{os.getgid()}"


def _deprioritized(command: Sequence[str]) -> list[str]:
"""``command``, wrapped so it raises its own OOM score before exec'ing (or unchanged when off).

The agent pane is started with ``docker exec``, and ``oom_score_adj`` is a **per-process**
attribute inherited across fork — the pane forks from the Docker daemon, not from the
container's PID 1, so ``docker run --oom-score-adj`` does *not* reach it. That leaves the one
process that actually eats memory (the agent CLI and its children) at the host's default score,
which defeats the point. ``docker exec`` has no flag for it, so the command writes
``/proc/self/oom_score_adj`` itself — raising your own score needs no privilege (only
*lowering* it does) — and then ``exec``s, keeping the process tree and signal behaviour
identical to running the command directly. Best-effort: a host that refuses the write is
tolerated (``|| true``) rather than losing the agent. The cgroup-based flags need no such
treatment — an exec'd process joins the container's cgroup like any other."""
adj = container_oom_score_adj()
if adj is None:
return list(command)
return [
"sh",
"-c",
f"echo {adj} > /proc/self/oom_score_adj 2>/dev/null || true; exec {shlex.join(command)}",
]


class LocalRunner(Runner):
"""Runs task containers + host tmux on the local Docker daemon (one host)."""

Expand Down Expand Up @@ -161,12 +194,54 @@ def __init__(
self._agent_command = list(agent_command)
self._tmux_socket = tmux_socket # isolate panopticon's tmux server when set (-L)
self._extra_env = dict(extra_env or {})
# Set once a `docker run` proves this daemon can't apply the cgroup priority flags, so the
# flagless retry in `_run_container` is paid at most once per process, not per spawn.
self._cgroup_flags_unsupported = False
self._run = run

def _tmux(self, *args: str) -> list[str]:
prefix = ["tmux", *(["-L", self._tmux_socket] if self._tmux_socket else [])]
return [*prefix, *args]

def _priority_flags(self) -> list[str]:
"""The deprioritizing ``docker run`` flags for a task container (see
:mod:`panopticon.sessionservice.priority`), minus the cgroup ones once this daemon has
proved it can't apply them."""
flags = container_priority_flags()
return strip_cgroup_flags(flags) if self._cgroup_flags_unsupported else flags

def _run_container(self, docker_run: Sequence[str], container: str | None = None) -> None:
"""``docker run``, degrading rather than failing on a daemon that refuses the cgroup
priority flags.

Some daemons reject ``--cpu-shares``/``--blkio-weight`` outright — a nested daemon whose
cgroup is in threaded mode answers with ``unable to apply cgroup configuration`` — and no
work may be lost to a host that merely can't deprioritize. So a failed flagged run is
retried once without those flags; when the run named a container (a task spawn), the
half-created one is cleared first, since the name is taken. The "this daemon can't do it"
latch is set **only if the retry succeeds**: an ordinary failure (a bad image, say) then
propagates as before, with later runs still asking for full priority control."""
stripped = strip_cgroup_flags(docker_run)
if list(stripped) == list(docker_run): # nothing to fall back to
self._run(docker_run)
return
try:
self._run(docker_run)
return
except Exception as exc: # any runner failure is worth one flagless retry
first_error = exc
if container is not None:
self._run(["docker", "rm", "--force", container], check=False)
self._run(stripped) # still failing => a real error, raised as it would have been anyway
self._cgroup_flags_unsupported = True
_log.warning(
"docker refused the cgroup priority flags (%s) — running task containers without "
"them; they keep --oom-score-adj. Set %s/%s to `off` to silence this.",
first_error,
CPU_SHARES_VAR,
BLKIO_WEIGHT_VAR,
)

def spawn(
self,
task_id: str,
Expand Down Expand Up @@ -252,6 +327,11 @@ def _report(phase: LifecyclePhase) -> None:
f"panopticon.task={task_id}",
"--add-host",
HOST_GATEWAY,
# Run the task at the lowest priority we can: it yields CPU and disk to the operator's
# own processes and is the kernel's first OOM pick, so a runaway agent or a heavy
# in-task build can't destabilize the host (see `priority`). Not a cap — an idle host
# still runs the task at full speed.
*self._priority_flags(),
]
if (
docker_in_docker
Expand Down Expand Up @@ -288,7 +368,7 @@ def _report(phase: LifecyclePhase) -> None:
self._run(self._tmux("kill-session", "-t", container), check=False)
self._run(["docker", "rm", "--force", container], check=False)
_report(LifecyclePhase.STARTING) # docker run + the tmux session coming up
self._run(docker_run)
self._run_container(docker_run, container)
# The pane is a host shell command: wait (bounded) for the entrypoint's READY_MARKER — the
# remap-complete signal — then exec in as the unprivileged `panopticon` user, so `tmux
# attach` and the agent's `whoami` see that named user, not root (and never the pre-remap
Expand All @@ -303,7 +383,7 @@ def _report(phase: LifecyclePhase) -> None:
"--user",
CONTAINER_USER,
container,
*self._agent_command,
*_deprioritized(self._agent_command),
]
)
pane = (
Expand Down Expand Up @@ -383,7 +463,7 @@ def delete_workspace_contents(self, path: str) -> None:
it, so the daemon can then ``rmtree`` the now-empty directory. Overrides the
panopticon entrypoint (which would remap uid) so the container runs as root and can
reach files it created. Raises on nonzero docker exit."""
self._run(
self._run_container(
[
"docker",
"run",
Expand All @@ -392,6 +472,11 @@ def delete_workspace_contents(self, path: str) -> None:
"/bin/sh",
"--volume",
f"{path}:/cleanup",
# An unbounded delete over a whole checkout is exactly the IO storm the priority
# flags exist for, so the cleanup sweep yields to the host like a task does — and it
# degrades the same way on a daemon that won't take them (a spawn may not have
# discovered that yet). The container is `--rm` and unnamed, so no name to clear.
*self._priority_flags(),
self._image,
"-c",
"find /cleanup -mindepth 1 -delete",
Expand Down
Loading
Loading