From 834d4a090a0206700e72b2526d94e84235899eca Mon Sep 17 00:00:00 2001 From: Charlie Scherer Date: Wed, 23 Sep 2026 23:01:10 +0000 Subject: [PATCH 1/2] Stop a snoozed task's container (and respawn it on wake) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Snoozing a task muted it on the dashboard while its container kept running, holding CPU, RAM and a model session for as long as the operator ignored it. `Spawner.pause` now stops the container of a task this runner claims that is actively snoozed, and releases the claim — which also clears the reported lifecycle phase, so the task composes `queued`. Waking needs no path of its own: a lapsed (or cleared) deadline leaves an unclaimed, un-snoozed task, exactly what `spawn_one` claims and spawns, with the per-task clone and the CLI session history in its config volume intact. `snoozed_until` stays a recorded fact the task service never compares to a clock; the arithmetic moves to `core/snooze.py` (`now` passed in) and is shared by the two callers that own a clock — the dashboard, which mutes the row, and the session service, which stops the container. Matching gates keep `spawn_one`, `reconcile` and `heal`/`mark_healing` off a paused task, so nothing in the same pass undoes the stop or reports the deliberate kill as a failure. Shell tasks are never paused: their script runs once, so stopping it would be a cancel. Co-Authored-By: Claude Opus 5 --- AGENTS.md | 26 ++- docs/dashboard.md | 4 +- src/panopticon/core/snooze.py | 49 +++++ src/panopticon/sessionservice/host.py | 14 +- src/panopticon/sessionservice/spawner.py | 86 +++++++- src/panopticon/terminal/dashboard.py | 48 ++-- tests/core/test_snooze.py | 49 +++++ tests/sessionservice/test_host.py | 62 ++++++ tests/sessionservice/test_spawner.py | 266 ++++++++++++++++++++++- 9 files changed, 558 insertions(+), 46 deletions(-) create mode 100644 src/panopticon/core/snooze.py create mode 100644 tests/core/test_snooze.py diff --git a/AGENTS.md b/AGENTS.md index c5a0443f..9d1922f3 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -51,7 +51,9 @@ src/panopticon/ # checkout, mounted rw at /workspace; point origin at the forge, then init any # submodules — in that order, since relative .gitmodules URLs resolve against # origin); spawner.py = the spawn loop (claim an unclaimed task → spawn its - # container; prefills claude's input box with the task memo on a first spawn); + # container; prefills claude's input box with the task memo on a first spawn; + # `pause` stops a snoozed task's container + releases its claim, and the same + # snooze gate keeps spawn/reconcile/heal off it until the deadline lapses); # prefill.py = the detached input-box prefill # poller (mirrors cloude-cade: pipe-pane watch for ESC[?2004h → paste-buffer the # description, unsent); daemon.py = the provision-only pull loop; @@ -195,6 +197,9 @@ on every PR (the same commands the Makefile wraps). - `tests/test_discovery.py` — workflow discovery (Slice 8): the built-in package + an optional path are scanned for `Workflow` subclasses; a dropped-in module registers with no core change; underscored/non-workflow files are ignored; duplicate names are rejected. +- `tests/core/test_snooze.py` — the snooze predicate both clock-owners share: an active vs lapsed + deadline, the sticky sentinel, naive datetimes, and an unreadable fact reading *inactive* rather + than raising (a bad recorded value must never stall a spawn). - `tests/test_git.py` — local git ops: unit tests pin the emitted `git` commands and slug-gating for `GitWorktrees` and the per-task-clone ops `GitClones` (clone/branch/set-origin, ADR 0011); a `skipif` integration test creates a real worktree. @@ -220,10 +225,15 @@ on every PR (the same commands the Makefile wraps). `exit_reason` — OOM kill/exit code — including for an already-`down` task), `heal` **self-heal** (a claimed-by-us non-terminal task whose tmux session is gone → respawn via the idempotent spawn path; skips healthy/unclaimed/terminal tasks; the crash-loop cap — surfaced once as `failed` - with a press-R detail — + survivor-window budget reset), and the `spawnable_tasks` filter; an - integration test claims + spawns against the real task service over REST (fake git/runner). + with a press-R detail — + survivor-window budget reset), `pause` **snooze → stop** (a snoozed + task's container is stopped and its claim released — even mid-turn; no-op when unclaimed/terminal/ + shell/awake/already stopped — and the matching gates that keep `spawn_one`/`reconcile`/`heal` off a + paused task within the same pass), and the `spawnable_tasks` filter; integration tests claim + + spawn against the real task service over REST (fake git/runner) and drive the whole snooze cycle + (stopped + `queued` → the deadline lapses on an injected clock → claimed + respawned). - `tests/test_host.py` — the unified per-host daemon (ADR 0008/0011): a unit test isolates a - failing task and another pins that each pass also `heal`s every task; an integration test drives + failing task, another pins that each pass also `heal`s every task, and another that it `pause`s + every task **before** spawning it (so nothing re-spawns what a snooze just stopped); an integration test drives spawn → set slug → provision against the real task service over REST (claimed + spawned, then branched, no re-spawn). - `tests/test_daemon.py` — the observe-and-provision loop + its launch: unit tests drive @@ -326,6 +336,14 @@ on every PR (the same commands the Makefile wraps). in — but, being a transition, a free move runs through an agent skill (the user directs the agent), not the dashboard. `force_transition` is the engine primitive (e.g. going back to coding is just `set_state(ITERATING)` — not a named operation). +- **Snooze** — `Task.snoozed_until`, an operator-owned "not now" deadline (dashboard `e` = 12h, + `E` = the sticky sentinel, a second `e` clears it). The task service stores it **verbatim** and + never compares it to a clock (the determinism invariant); the arithmetic is `core/snooze.py` + (`is_snoozed(until, now)` — pure, `now` passed in), shared by the two callers that own a clock: + the dashboard mutes the row, and the **session service stops the container** (`Spawner.pause` → + stop + release the claim, so the task reads `queued`). Waking needs no separate path — a lapsed + or cleared deadline leaves an unclaimed, un-snoozed task, which `spawn_one` claims and respawns + with its per-task clone and CLI session history intact. - **Turn-flip / blocked** — the live `Task.turn` flips *within* a state via `PUT /tasks/{id}/turn` (the agnostic agent↔user ball tracking). The **contract** for the in-container hooks: the agent's stop hook sets `turn=user` (**unless a background task — a diff --git a/docs/dashboard.md b/docs/dashboard.md index be8d980b..17179c5e 100644 --- a/docs/dashboard.md +++ b/docs/dashboard.md @@ -18,8 +18,8 @@ kept in sync with the `HOTKEYS` keymap and the modal `BINDINGS` in | `R` | Respawn a down task (releases its claim so the runner re-spawns it) | | `p` | Open the task's URL in the browser | | `w` | Open the task's workdir (its per-task clone) in the host's file manager | -| `e` | Snooze the highlighted task for 12 hours | -| `E` | Snooze the highlighted task indefinitely | +| `e` | Snooze the highlighted task for 12 hours (stops its container) | +| `E` | Snooze the highlighted task indefinitely (stops its container) | | `a` | List the task's artifacts | | `A` | List the task's **repo's** artifacts (shared by every task in it) | | `g` | Open the repo config screen | diff --git a/src/panopticon/core/snooze.py b/src/panopticon/core/snooze.py new file mode 100644 index 00000000..379a0862 --- /dev/null +++ b/src/panopticon/core/snooze.py @@ -0,0 +1,49 @@ +"""The operator snooze predicate — pure arithmetic over a recorded deadline (no clock read). + +``Task.snoozed_until`` is a recorded fact: the task service stores the operator's deadline verbatim +and never compares it to a clock (the determinism invariant). Deciding whether a deadline is *active* +is therefore the caller's job, and two callers need the same answer: + +- the **dashboard**, to mute a snoozed row (its display clock), and +- the **session service**, to stop a snoozed task's container (its host clock, ADR 0008). + +So the arithmetic lives here, in ``core``, with ``now`` passed **in** — LLM-free, I/O-free, and +clock-free like the rest of the package. +""" + +from __future__ import annotations + +from datetime import UTC, datetime + +#: The reserved "sticky" deadline: a snooze that never expires, cleared only explicitly. Recorded by +#: the dashboard's `E` hotkey; defined here so the value has exactly one definition. +INDEFINITE_UNTIL = "9999-12-31T23:59:59+00:00" + + +def snooze_remaining(snoozed_until: str | None, now: datetime) -> float | None: + """Active seconds remaining on ``snoozed_until`` at ``now``, else ``None``. + + ``float("inf")`` for :data:`INDEFINITE_UNTIL`; ``None`` when there's no deadline, when it has + already passed, or when it isn't a parseable ISO-8601 string (an unreadable fact is inactive + rather than an error — it must never take down a display or stall a spawn). Naive datetimes on + either side are read as UTC. + """ + if not isinstance(snoozed_until, str): + return None + if snoozed_until == INDEFINITE_UNTIL: + return float("inf") + try: + deadline = datetime.fromisoformat(snoozed_until) + except ValueError: + return None + if deadline.tzinfo is None: + deadline = deadline.replace(tzinfo=UTC) + if now.tzinfo is None: + now = now.replace(tzinfo=UTC) + seconds = (deadline - now).total_seconds() + return seconds if seconds > 0 else None + + +def is_snoozed(snoozed_until: str | None, now: datetime) -> bool: + """Whether ``snoozed_until`` is an **active** snooze at ``now`` (see :func:`snooze_remaining`).""" + return snooze_remaining(snoozed_until, now) is not None diff --git a/src/panopticon/sessionservice/host.py b/src/panopticon/sessionservice/host.py index cc0c698b..4fabd265 100644 --- a/src/panopticon/sessionservice/host.py +++ b/src/panopticon/sessionservice/host.py @@ -69,10 +69,15 @@ def __init__( self._interval = interval def tick(self, tasks: list[JsonObj]) -> None: - """One pass over a task snapshot: spawn each spawnable task, provision each slugged one, - publish each pending push, reconcile each claimed one's container-lifecycle status - (down-detection), and heal each orphan (a claimed task whose tmux session is gone → - respawn). All self-gate, so re-running over an unchanged snapshot is a no-op. + """One pass over a task snapshot: pause each snoozed task (stop its container, release the + claim), spawn each spawnable task, provision each slugged one, publish each pending push, + reconcile each claimed one's container-lifecycle status (down-detection), and heal each orphan + (a claimed task whose tmux session is gone → respawn). All self-gate, so re-running over an + unchanged snapshot is a no-op. + + ``pause`` runs **first**: it's what frees a snoozed task's resources, and every step after it + skips a snoozed task, so nothing in the same pass undoes it. Waking needs no step of its own — + a task whose deadline has lapsed is unclaimed and un-snoozed, so ``spawn_one`` brings it back. ``publish`` runs **before** ``cleanup``: a task can request its push and reach a terminal state in the same breath, and cleanup deletes the per-task clone the push reads from. @@ -89,6 +94,7 @@ def tick(self, tasks: list[JsonObj]) -> None: _log.warning("flagging heal failed for task %s", task.get("id"), exc_info=True) for task in tasks: try: + self._spawner.pause(task) # snoozed → stop the container, release the claim self._spawner.spawn_one(task) self._provisioner.provision(task) self._publisher.publish(task) diff --git a/src/panopticon/sessionservice/spawner.py b/src/panopticon/sessionservice/spawner.py index 1401d86e..de36fc55 100644 --- a/src/panopticon/sessionservice/spawner.py +++ b/src/panopticon/sessionservice/spawner.py @@ -4,7 +4,9 @@ that are **unclaimed** and **non-terminal**, **claims** one for this host (the claim is the spawn gate — exactly one runner owns it; a lost race is a 409 we skip), prepares its writable per-task clone (`prepare_workspace`), and spawns the container via the runner with the repo's secrets + the -``/workspace`` mount. Provisioning (slug → branch) is the sibling loop (`ProvisionDaemon`); the +``/workspace`` mount. It also **pauses** a task the operator snoozed (stop the container, release the +claim) — the deadline is a recorded fact, and reading it against a clock is the host's job, not the +control plane's. Provisioning (slug → branch) is the sibling loop (`ProvisionDaemon`); the unified host daemon runs both. LLM-free. """ @@ -17,6 +19,7 @@ import subprocess import time from collections.abc import Callable +from datetime import UTC, datetime from pathlib import Path import httpx @@ -25,11 +28,12 @@ from panopticon.core.dirs import hook_file_path from panopticon.core.features import require_available_agent_cli from panopticon.core.models import ContainerStatus, LifecyclePhase, resolve_agent_cli +from panopticon.core.snooze import is_snoozed from panopticon.core.state import TERMINAL_LABELS from panopticon.sessionservice.clones import CloneCache from panopticon.sessionservice.executions import WorkflowExecutions from panopticon.sessionservice.images import ImageBuilder -from panopticon.sessionservice.local_runner import LocalRunner, base_image +from panopticon.sessionservice.local_runner import LocalRunner, base_image, session_name from panopticon.sessionservice.shell_runner import ShellRunner from panopticon.sessionservice.spawn import cleanup_workspace, prepare_workspace @@ -104,6 +108,7 @@ def __init__( rmtree: Callable[[str], None] = shutil.rmtree, docker_cleanup: Callable[[str], None] | None = None, now: Callable[[], float] = time.monotonic, + utcnow: Callable[[], datetime] = lambda: datetime.now(UTC), max_respawns: int = MAX_RESPAWNS, respawn_reset: float = RESPAWN_RESET_SECONDS, ) -> None: @@ -131,6 +136,10 @@ def __init__( docker_cleanup if docker_cleanup is not None else runner.delete_workspace_contents ) self._now = now + #: Wall clock for the snooze predicate (:meth:`pause`). Distinct from ``now`` (monotonic, for + #: the crash-loop window): a snooze deadline is an absolute time, so it needs a real clock. + #: Read **here**, on the host — the control plane records the deadline and never compares it. + self._utcnow = utcnow self._max_respawns = max_respawns self._respawn_reset = respawn_reset #: task_id → (respawns in the current burst, monotonic time of the last respawn), the @@ -147,9 +156,16 @@ def spawn_one(self, task: JsonObj) -> str | None: Reports each spawn phase to the task service as it goes (``CLAIMING`` → ``PREPARING`` → ``BUILDING`` → ``STARTING`` → ``AWAITING``) so the dashboard can surface the steps to becoming live; a step raising is reported as ``FAILED`` (with the error) before re-raising, so the - host daemon's per-task isolation still applies but the failure is visible, not silent.""" + host daemon's per-task isolation still applies but the failure is visible, not silent. + + An **actively snoozed** task is skipped: the operator muted it, so it must not come up (and a + task :meth:`pause` just released must not be re-spawned in the same pass). Once the deadline + lapses — or the operator clears it — this is also what wakes the task, spawning it through the + full visible lifecycle with its clone and CLI session history intact.""" if task["state"] in TERMINAL_LABELS or task.get("claimed_by"): return None + if self._is_snoozed(task): + return None # muted by the operator — don't bring it up (see `pause`) try: self._client.claim(task["id"], self._runner_id) except httpx.HTTPStatusError as exc: @@ -306,6 +322,49 @@ def _report(self, task_id: str, phase: LifecyclePhase, detail: str | None = None with contextlib.suppress(httpx.HTTPError): self._client.report_lifecycle(task_id, self._runner_id, phase.value, detail) + def _is_snoozed(self, task: JsonObj) -> bool: + """Whether the operator's snooze on ``task`` is **active** right now (this host's clock). + + ``snoozed_until`` is a recorded fact the task service never interprets (the determinism + invariant): reading it against a clock is the caller's job, and here the caller is the runner. + The arithmetic is shared with the dashboard (:mod:`panopticon.core.snooze`), so the row you see + muted and the container we stop are decided by the same rule.""" + return is_snoozed(task.get("snoozed_until"), self._utcnow()) + + def pause(self, task: JsonObj) -> bool: + """Stop the container of a task the operator snoozed, and release its claim. Did we stop one? + + Snoozing is "not now": the task is muted on the dashboard, so its container shouldn't sit + there holding CPU, RAM and a model session. We stop it and **release the claim**, which also + clears any reported lifecycle phase, so the task composes ``queued`` — and waking it needs no + separate path: once the deadline lapses (or the operator clears it) :meth:`spawn_one` claims + and spawns it again like any queued task, with its per-task clone and the CLI session history + in its config volume intact, so the agent resumes where it left off. + + Stopping is unconditional — even mid-turn. ``/workspace`` is a host bind mount, so the + checkout survives; what's lost is the in-container process (an in-flight tool call), not work + on disk. + + Self-gates so calling it on every task each pass is safe: only a task **this** runner claims, + non-terminal, actively snoozed, and with a container or session still up (so a task already + paused isn't released again every pass). **Shell** tasks are skipped — their script runs once, + so killing it would be a cancel, not a pause (the same reason :meth:`heal` skips them).""" + task_id = task["id"] + if task.get("claimed_by") != self._runner_id or task["state"] in TERMINAL_LABELS: + return False + if not self._is_snoozed(task): + return False + if self._executions.is_shell(task.get("workflow")): + return False # a shell script is run once — stopping it is a cancel, not a pause + runner = self._runner_for(task) + if not (runner.is_running(task_id) or runner.has_session(task_id)): + return False # nothing left to stop — already paused + _log.info("task %s: snoozed — stopping its container and releasing the claim", task_id) + runner.stop(session_name(task_id)) + self._respawns.pop(task_id, None) # a deliberate stop is not a crash — forget the burst + self._client.release(task_id) # → unclaimed, phase cleared → composes `queued` + return True + def reconcile(self, task: JsonObj) -> None: """Reconcile a task this runner claims into the right lifecycle status (down-detection). @@ -322,6 +381,8 @@ def reconcile(self, task: JsonObj) -> None: still running is left to keep coming up.""" if task.get("claimed_by") != self._runner_id: return # not ours (or unclaimed) — spawn_one handles the unclaimed case + if self._is_snoozed(task): + return # we stopped it on purpose (`pause`) — not a death to explain status = task.get("container_status") if status not in _IN_PROGRESS and status != ContainerStatus.DOWN.value: return # live / failed / queued / disconnected — nothing to reconcile @@ -347,6 +408,8 @@ def _is_orphan(self, task: JsonObj) -> bool: operator cancelling), not a crash to respawn — so re-running it would be wrong.""" if task.get("claimed_by") != self._runner_id or task["state"] in TERMINAL_LABELS: return False + if self._is_snoozed(task): + return False # deliberately stopped by `pause` — a paused task is not an orphan if self._executions.is_shell(task.get("workflow")): return False return not self._runner.has_session(task["id"]) @@ -502,12 +565,23 @@ def _compose_image(self, workflow: str, repo: JsonObj, agent_cli: str) -> str: return self._images.build(workflow, repo["id"], layers, agent_cli=agent_cli, verbose=True) -def spawnable_tasks(client: TaskServiceClient) -> Callable[[], list[JsonObj]]: - """This host's spawn candidates: unclaimed, non-terminal tasks (the runner claims-then-spawns). +def spawnable_tasks( + client: TaskServiceClient, *, utcnow: Callable[[], datetime] = lambda: datetime.now(UTC) +) -> Callable[[], list[JsonObj]]: + """This host's spawn candidates: unclaimed, non-terminal, un-snoozed tasks (the runner + claims-then-spawns). + + An actively snoozed task is no candidate — the operator muted it, and :meth:`Spawner.pause` stops + the container of one that's already up (the same gate :meth:`Spawner.spawn_one` applies, so the + list and the spawn can't disagree). For M1 (single host) that's every such task the service knows; scoping to this runner's own assignments is an M5 refinement. """ return lambda: [ - t for t in client.list_tasks() if not t["claimed_by"] and t["state"] not in TERMINAL_LABELS + t + for t in client.list_tasks() + if not t["claimed_by"] + and t["state"] not in TERMINAL_LABELS + and not is_snoozed(t.get("snoozed_until"), utcnow()) ] diff --git a/src/panopticon/terminal/dashboard.py b/src/panopticon/terminal/dashboard.py index ae5f97bc..8339c2ed 100644 --- a/src/panopticon/terminal/dashboard.py +++ b/src/panopticon/terminal/dashboard.py @@ -115,6 +115,7 @@ from panopticon.core.dirs import ARTIFACTS_DIR from panopticon.core.features import codex_enabled from panopticon.core.models import resolve_agent_cli +from panopticon.core.snooze import INDEFINITE_UNTIL, snooze_remaining from panopticon.core.state import TERMINAL_LABELS from panopticon.sessionservice.local_runner import session_name from panopticon.taskservice.artifacts_fs import FilesystemArtifactStore @@ -363,29 +364,18 @@ def _matches(task: JsonObj, query: str) -> bool: # Fixed operator snooze controls: `e` means "not today"; `E` records the reserved sticky value. # A snooze always mutes until it expires — there is no attention/piercing here (that field was # deliberately dropped from this fork), so the turn-column precedence is just: snoozed > normal. +# Snoozing also **stops the task's container** (the session service reads the same deadline off its +# own clock, `core.snooze`); clearing it — or the deadline lapsing — respawns the task. _SNOOZE_DURATION = timedelta(hours=12) -_INDEFINITE_SNOOZE_UNTIL = "9999-12-31T23:59:59+00:00" +_INDEFINITE_SNOOZE_UNTIL = INDEFINITE_UNTIL def _snooze_remaining(task: JsonObj, now: datetime) -> float | None: - """Active seconds remaining; +inf for the reserved sticky deadline; None if inactive.""" - raw = task.get("snoozed_until") - if not isinstance(raw, str): - return None - if raw == _INDEFINITE_SNOOZE_UNTIL: - return float("inf") - try: - deadline = datetime.fromisoformat(raw) - except ValueError: - return None - if deadline.tzinfo is None: - deadline = deadline.replace(tzinfo=UTC) - if now.tzinfo is None: - now = now.replace(tzinfo=UTC) - seconds = (deadline - now).total_seconds() - if seconds <= 0: - return None - return seconds + """Active seconds remaining; +inf for the reserved sticky deadline; None if inactive. + + The arithmetic itself lives in :mod:`panopticon.core.snooze` — shared with the session service, + which gates stopping the container on the same predicate.""" + return snooze_remaining(task.get("snoozed_until"), now) def _snooze_label(task: JsonObj, now: datetime) -> str | None: @@ -1927,12 +1917,18 @@ def binding(self) -> Binding: "Open the task's workdir (its per-task clone) in the host's file manager", show=False, ), - Hotkey("e", "snooze", "Snooze", "Snooze the highlighted task for 12 hours", show=False), + Hotkey( + "e", + "snooze", + "Snooze", + "Snooze the highlighted task for 12 hours (stops its container)", + show=False, + ), Hotkey( "E", "snooze_indefinitely", "Snooze sticky", - "Snooze the highlighted task indefinitely", + "Snooze the highlighted task indefinitely (stops its container)", show=False, ), Hotkey("g", "repos", "Repos", "Repo config (list / create / edit repos)", show=False), @@ -2465,8 +2461,11 @@ def action_drop(self) -> None: def action_snooze(self) -> None: """`e`: toggle a fixed twelve-hour operator snooze on the highlighted task. - Snoozing a task mutes it (dims the row, shows `snoozed · Nh left`) until the deadline; a - second `e` while it's active clears it. The 12h window is a hard constant.""" + Snoozing a task mutes it (dims the row, shows `snoozed · Nh left`) until the deadline **and + stops its container** — the session service reads the same deadline off its own clock, stops + the container and releases the claim, so the task reads `queued` until it wakes. A second `e` + while it's active clears it, and the next daemon pass respawns the task (the agent resumes + its session). The 12h window is a hard constant.""" task_id = self._current if task_id is None: return @@ -2482,7 +2481,8 @@ def action_snooze(self) -> None: self.action_refresh() def action_snooze_indefinitely(self) -> None: - """`E`: record the reserved sticky snooze deadline (mute until explicitly un-snoozed).""" + """`E`: record the reserved sticky snooze deadline (mute — and stop the container — until + explicitly un-snoozed).""" task_id = self._current if task_id is None: return diff --git a/tests/core/test_snooze.py b/tests/core/test_snooze.py new file mode 100644 index 00000000..28792f71 --- /dev/null +++ b/tests/core/test_snooze.py @@ -0,0 +1,49 @@ +"""The operator snooze predicate (``core.snooze``) — pure arithmetic, clock passed in. + +Two callers gate on this — the dashboard (mute the row) and the session service (stop the +container) — so it's pinned on its own: an unreadable or lapsed deadline must read *inactive* +rather than raise, or a bad recorded fact would stall a spawn. +""" + +from __future__ import annotations + +from datetime import UTC, datetime, timedelta + +from panopticon.core.snooze import INDEFINITE_UNTIL, is_snoozed, snooze_remaining + +NOW = datetime(2026, 8, 6, 12, 0, tzinfo=UTC) + + +def test_finite_deadline_in_the_future_is_active() -> None: + until = (NOW + timedelta(hours=3)).isoformat() + assert snooze_remaining(until, NOW) == 3 * 3600 + assert is_snoozed(until, NOW) + + +def test_lapsed_deadline_is_inactive() -> None: + """The wake signal: expiry is clock-only (no event), so the predicate alone reports it.""" + until = (NOW - timedelta(seconds=1)).isoformat() + assert snooze_remaining(until, NOW) is None + assert not is_snoozed(until, NOW) + assert not is_snoozed(NOW.isoformat(), NOW) # the deadline itself is no longer in the future + + +def test_sticky_sentinel_never_expires() -> None: + assert snooze_remaining(INDEFINITE_UNTIL, NOW) == float("inf") + assert is_snoozed(INDEFINITE_UNTIL, datetime(9999, 1, 1, tzinfo=UTC)) + + +def test_no_deadline_is_inactive() -> None: + assert snooze_remaining(None, NOW) is None + assert not is_snoozed(None, NOW) + + +def test_unparseable_deadline_is_inactive_not_an_error() -> None: + assert snooze_remaining("not a timestamp", NOW) is None + assert snooze_remaining(12345, NOW) is None # type: ignore[arg-type] # a non-string fact + + +def test_naive_datetimes_are_read_as_utc() -> None: + """Either side may be naive (a hand-written fact, a naive display clock) — both mean UTC.""" + assert is_snoozed("2026-08-06T15:00:00", NOW) # naive deadline vs aware now + assert is_snoozed((NOW + timedelta(hours=1)).isoformat(), NOW.replace(tzinfo=None)) diff --git a/tests/sessionservice/test_host.py b/tests/sessionservice/test_host.py index dc017499..780a4bdb 100644 --- a/tests/sessionservice/test_host.py +++ b/tests/sessionservice/test_host.py @@ -114,6 +114,9 @@ class _Spawner: def mark_healing(self, task: JsonObj) -> None: return None + def pause(self, task: JsonObj) -> None: + return None + def spawn_one(self, task: JsonObj) -> None: seen.append(task["id"]) if task["id"] == "t1": @@ -145,6 +148,9 @@ class _Spawner: def mark_healing(self, task: JsonObj) -> None: return None + def pause(self, task: JsonObj) -> None: + return None + def spawn_one(self, task: JsonObj) -> None: return None @@ -177,6 +183,9 @@ class _Spawner: def mark_healing(self, task: JsonObj) -> None: events.append(f"mark:{task['id']}") + def pause(self, task: JsonObj) -> None: + return None + def spawn_one(self, task: JsonObj) -> None: return None @@ -199,6 +208,41 @@ def provision(self, task: JsonObj) -> None: assert events == ["mark:t1", "mark:t2", "heal:t1", "heal:t2"] # all marks precede any respawn +def test_tick_pauses_each_task_before_spawning_it() -> None: + # Snoozing stops the container: `pause` runs over every task, and *before* `spawn_one` — it + # releases the claim of a snoozed task, and every later step self-gates on the same snooze, so + # nothing in the same pass brings back what it just stopped. + events: list[str] = [] + + class _Spawner: + def mark_healing(self, task: JsonObj) -> None: + return None + + def pause(self, task: JsonObj) -> None: + events.append(f"pause:{task['id']}") + + def spawn_one(self, task: JsonObj) -> None: + events.append(f"spawn:{task['id']}") + + def reconcile(self, task: JsonObj) -> None: + return None + + def heal(self, task: JsonObj) -> None: + return None + + def cleanup(self, task: JsonObj) -> None: + return None + + class _Provisioner: + def provision(self, task: JsonObj) -> None: + return None + + HostDaemon(_FakeClient([]), _Spawner(), _Provisioner(), _NoopPublisher()).tick( + [{"id": "t1"}, {"id": "t2"}] + ) # type: ignore[arg-type] + assert events == ["pause:t1", "spawn:t1", "pause:t2", "spawn:t2"] + + def test_tick_publishes_each_task_before_cleaning_it_up() -> None: # Both halves matter. A task can request its push and reach a terminal state in the same # breath, and `cleanup` deletes the per-task clone the push reads from — so publishing has to @@ -209,6 +253,9 @@ class _Spawner: def mark_healing(self, task: JsonObj) -> None: return None + def pause(self, task: JsonObj) -> None: + return None + def spawn_one(self, task: JsonObj) -> None: return None @@ -242,6 +289,9 @@ class _Spawner: def mark_healing(self, task: JsonObj) -> None: return None + def pause(self, task: JsonObj) -> None: + return None + def spawn_one(self, task: JsonObj) -> None: return None @@ -283,6 +333,9 @@ def startup_reclaim(self, tasks: list[JsonObj]) -> None: def mark_healing(self, task: JsonObj) -> None: return None + def pause(self, task: JsonObj) -> None: + return None + def spawn_one(self, task: JsonObj) -> None: return None @@ -329,6 +382,9 @@ def startup_reclaim(self, tasks: list[JsonObj]) -> None: def mark_healing(self, task: JsonObj) -> None: return None + def pause(self, task: JsonObj) -> None: + return None + def spawn_one(self, task: JsonObj) -> None: self.seen.append(task["id"]) @@ -371,6 +427,9 @@ def startup_reclaim(self, tasks: list[JsonObj]) -> None: def mark_healing(self, task: JsonObj) -> None: return None + def pause(self, task: JsonObj) -> None: + return None + def spawn_one(self, task: JsonObj) -> None: return None @@ -534,6 +593,9 @@ class _Spawner: def mark_healing(self, task: JsonObj) -> None: return None + def pause(self, task: JsonObj) -> None: + return None + def spawn_one(self, task: JsonObj) -> None: return None diff --git a/tests/sessionservice/test_spawner.py b/tests/sessionservice/test_spawner.py index c16ee8d7..3447d33a 100644 --- a/tests/sessionservice/test_spawner.py +++ b/tests/sessionservice/test_spawner.py @@ -1,11 +1,13 @@ """Host-side spawn loop (ADR 0008): claim an unclaimed task, then spawn its container. Unit tests -drive `spawn_one`/`spawnable_tasks` with fakes; an integration test runs against the real task -service over REST. No Docker, no LLM — `git`/runner are fakes.""" +drive `spawn_one`/`spawnable_tasks`/`pause` with fakes (an injected wall clock for the snooze +predicate); integration tests run against the real task service over REST. No Docker, no LLM — +`git`/runner are fakes.""" from __future__ import annotations import asyncio from collections.abc import Callable +from datetime import UTC, datetime, timedelta from pathlib import Path import httpx @@ -15,6 +17,7 @@ from panopticon.client import JsonObj, TaskServiceClient from panopticon.core.git import GitClones from panopticon.core.models import LifecyclePhase, Repo +from panopticon.core.snooze import INDEFINITE_UNTIL from panopticon.sessionservice.clones import CloneCache from panopticon.sessionservice.spawner import Spawner, spawnable_tasks from panopticon.taskservice.api import create_app @@ -37,6 +40,7 @@ def __init__( self, *, running: bool = True, session: bool = True, exit_reason: str | None = None ) -> None: self.spawned: list[dict[str, object]] = [] + self.stopped: list[str] = [] self._running = running self._session = session self._exit_reason = exit_reason @@ -85,7 +89,7 @@ def exit_reason(self, task_id: str) -> str | None: return self._exit_reason def stop(self, container_id: str) -> None: - pass + self.stopped.append(container_id) def delete_workspace_contents(self, path: str) -> None: pass @@ -162,7 +166,17 @@ def release(self, task_id: str) -> JsonObj: return {"id": task_id} -def _spawner(client: object, runner: object, images: object = None) -> Spawner: +#: The fixed "now" the injected clock reports, so snooze deadlines in these tests are exact. +NOW = datetime(2026, 8, 6, 12, 0, tzinfo=UTC) + + +def _spawner( + client: object, + runner: object, + images: object = None, + *, + now: datetime = NOW, +) -> Spawner: cache = CloneCache("/cache", run=_no_op_run, exists=lambda _p: True, makedirs=lambda _p: None) # type: ignore[arg-type] return Spawner( client, @@ -173,6 +187,7 @@ def _spawner(client: object, runner: object, images: object = None) -> Spawner: git=GitClones(run=_no_op_run), images=images or _FakeImageBuilder(), # type: ignore[arg-type] makedirs=lambda _p: None, + utcnow=lambda: now, ) @@ -308,6 +323,7 @@ class _FakeShellRunner: def __init__(self, *, session: bool = True) -> None: self.spawned: list[dict[str, object]] = [] + self.stopped: list[str] = [] self._session = session def spawn( @@ -346,11 +362,16 @@ def exit_reason(self, task_id: str) -> str | None: return None # a shell task has no container to inspect (mirrors ShellRunner) def stop(self, session_id: str) -> None: - pass + self.stopped.append(session_id) def _shell_spawner( - client: object, runner: object, shell_runner: object, made: list[str] | None = None + client: object, + runner: object, + shell_runner: object, + made: list[str] | None = None, + *, + now: datetime = NOW, ) -> Spawner: cache = CloneCache("/cache", run=_no_op_run, exists=lambda _p: True, makedirs=lambda _p: None) # type: ignore[arg-type] return Spawner( @@ -363,6 +384,7 @@ def _shell_spawner( git=GitClones(run=_no_op_run), images=_FakeImageBuilder(), # type: ignore[arg-type] makedirs=(made.append if made is not None else (lambda _p: None)), + utcnow=lambda: now, ) @@ -1088,6 +1110,187 @@ def list_tasks(self) -> list[JsonObj]: assert [t["id"] for t in spawnable_tasks(_Lister())()] == ["a"] # type: ignore[arg-type] +# --- snooze → stop the container (`pause`) ----------------------------------------------------- +# `snoozed_until` is a recorded fact the control plane never compares to a clock, so the runner +# reads it against its own (injected here). A snoozed task's container is stopped and its claim +# released; every other step must then leave it alone, and waking it is just `spawn_one` again. + +_SNOOZED = (NOW + timedelta(hours=12)).isoformat() # what the dashboard's `e` records + + +def _task(**over: object) -> JsonObj: + task: JsonObj = { + "id": "t1", + "repo_id": "r1", + "workflow": "spike", + "state": "ITERATING", + "claimed_by": "host-1", + "snoozed_until": _SNOOZED, + } + task.update(over) + return task + + +def test_pause_stops_the_container_and_releases_the_claim_of_a_snoozed_task() -> None: + client, runner = _FakeClient(repo=_REPO), _FakeRunner(running=True, session=True) + assert _spawner(client, runner).pause(_task()) is True + assert runner.stopped == ["panopticon-t1"] # tmux session + container, by session name + assert client.releases == ["t1"] # → unclaimed, phase cleared → composes `queued` + + +def test_pause_stops_a_task_the_agent_is_still_working_on() -> None: + # Unconditional by design: /workspace is a host bind mount, so the checkout survives — only the + # in-container process dies, and the respawn resumes the CLI session from its config volume. + client, runner = _FakeClient(repo=_REPO), _FakeRunner() + assert _spawner(client, runner).pause(_task(turn="agent", container_status="live")) is True + assert runner.stopped == ["panopticon-t1"] + + +def test_pause_stops_a_task_snoozed_indefinitely() -> None: + client, runner = _FakeClient(repo=_REPO), _FakeRunner() + assert _spawner(client, runner).pause(_task(snoozed_until=INDEFINITE_UNTIL)) is True + assert runner.stopped == ["panopticon-t1"] + + +def test_pause_forgets_the_crash_loop_budget() -> None: + # A deliberate stop is not a crash: it clears the task's respawn budget, so a snooze/wake cycle + # can never push a healthy task over the self-heal cap. With the cap at 1, a second heal only + # respawns because `pause` reset the burst. + client, runner = _FakeClient(repo=_REPO), _FakeRunner(session=False) + spawner = Spawner( + client, + runner, + runner_id="host-1", # type: ignore[arg-type] + cache=CloneCache( + "/cache", run=_no_op_run, exists=lambda _p: True, makedirs=lambda _p: None + ), + tasks_root="/tasks", + git=GitClones(run=_no_op_run), + images=_FakeImageBuilder(), # type: ignore[arg-type] + makedirs=lambda _p: None, + now=lambda: 0.0, # frozen well inside the survivor window, so only `pause` can reset it + utcnow=lambda: NOW, + max_respawns=1, + ) + awake = _task(snoozed_until=None) + spawner.heal(awake) # the one respawn the cap allows + spawner.heal(awake) # capped — no second respawn + assert len(runner.spawned) == 1 + spawner.pause(_task()) # snoozed → stopped, budget forgotten + spawner.heal(awake) # woken → respawnable again + assert len(runner.spawned) == 2 + + +def test_pause_skips_an_unsnoozed_task() -> None: + client, runner = _FakeClient(repo=_REPO), _FakeRunner() + spawner = _spawner(client, runner) + assert spawner.pause(_task(snoozed_until=None)) is False + lapsed = (NOW - timedelta(seconds=1)).isoformat() + assert spawner.pause(_task(snoozed_until=lapsed)) is False # the deadline has passed → awake + assert runner.stopped == [] and client.releases == [] + + +def test_pause_skips_tasks_not_claimed_by_this_runner() -> None: + client, runner = _FakeClient(repo=_REPO), _FakeRunner() + spawner = _spawner(client, runner) + assert spawner.pause(_task(claimed_by=None)) is False + assert spawner.pause(_task(claimed_by="host-9")) is False + assert runner.stopped == [] and client.releases == [] + + +def test_pause_skips_terminal_tasks() -> None: + client, runner = _FakeClient(repo=_REPO), _FakeRunner() + assert _spawner(client, runner).pause(_task(state="COMPLETE")) is False + assert runner.stopped == [] and client.releases == [] + + +def test_pause_is_a_no_op_once_the_task_is_already_stopped() -> None: + # Called on every task each pass, so an already-paused task must not be released again (which + # would churn the change feed forever). + client, runner = _FakeClient(repo=_REPO), _FakeRunner(running=False, session=False) + assert _spawner(client, runner).pause(_task(claimed_by="host-1")) is False + assert runner.stopped == [] and client.releases == [] + + +def test_pause_stops_a_container_whose_session_is_gone() -> None: + # Either half still being up is enough to stop: a container with no session (the orphan case) + # is exactly what we must not leave running. + client, runner = _FakeClient(repo=_REPO), _FakeRunner(running=True, session=False) + assert _spawner(client, runner).pause(_task()) is True + assert runner.stopped == ["panopticon-t1"] + + +def test_pause_never_stops_a_shell_task() -> None: + # A shell script runs once — killing it is a cancel, not a pause (the same reason heal skips it). + client = _FakeClient(repo=_REPO, runner_type="shell", shell_script="echo hi") + runner, shell = _FakeRunner(), _FakeShellRunner(session=True) + task = _task(workflow="setup-repo", state="RUNNING") + assert _shell_spawner(client, runner, shell).pause(task) is False + assert shell.stopped == [] and runner.stopped == [] and client.releases == [] + + +def test_spawn_one_skips_a_snoozed_task() -> None: + # Both the never-spawned case and the just-paused one: a snoozed task must not come up, and the + # release `pause` just did must not be undone later in the same pass. + client, runner = _FakeClient(repo=_REPO), _FakeRunner() + assert _spawner(client, runner).spawn_one(_task(claimed_by=None)) is None + assert client.claims == [] and runner.spawned == [] + + +def test_spawn_one_wakes_a_task_once_its_snooze_lapses() -> None: + # Waking needs no path of its own: the deadline passing leaves an unclaimed, un-snoozed task, + # which is exactly what spawn_one claims and spawns — the full visible lifecycle, same clone. + client, runner = _FakeClient(repo=_REPO), _FakeRunner() + spawner = _spawner(client, runner, now=NOW + timedelta(hours=13)) + assert spawner.spawn_one(_task(claimed_by=None)) == "panopticon-t1" + assert client.claims == [("t1", "host-1")] + assert [phase for _t, phase, _d in client.phases] == [ + "claiming", + "preparing", + "building", + "starting", + "awaiting", + ] + + +def test_reconcile_leaves_a_paused_task_alone() -> None: + # Within the same pass, reconcile still sees the pre-pause snapshot (claimed by us, `live`) with + # the container now gone. Without the snooze gate it would report `failed: container exited …` + # for a container we killed on purpose. + client = _FakeClient(repo=_REPO) + runner = _FakeRunner(running=False, exit_reason="container exited (exit 137)") + _spawner(client, runner).reconcile(_task(container_status="live")) + assert client.phases == [] and client.cleared == [] + + +def test_heal_does_not_respawn_a_paused_task() -> None: + # Same stale-snapshot window: claimed by us with no session is the orphan shape, but a task we + # stopped on purpose is not an orphan — respawning it would defeat the snooze entirely. + client, runner = _FakeClient(repo=_REPO), _FakeRunner(session=False) + spawner = _spawner(client, runner) + assert spawner.heal(_task()) is None + spawner.mark_healing(_task()) + assert runner.spawned == [] and client.phases == [] + + +def test_spawnable_tasks_filters_snoozed_tasks() -> None: + class _Lister: + def list_tasks(self) -> list[JsonObj]: + return [ + {"id": "a", "state": "ITERATING", "claimed_by": None, "snoozed_until": None}, + {"id": "b", "state": "ITERATING", "claimed_by": None, "snoozed_until": _SNOOZED}, + { + "id": "c", # a deadline that has already passed — awake again + "state": "ITERATING", + "claimed_by": None, + "snoozed_until": (NOW - timedelta(hours=1)).isoformat(), + }, + ] + + candidates = spawnable_tasks(_Lister(), utcnow=lambda: NOW)() # type: ignore[arg-type] + assert [t["id"] for t in candidates] == ["a", "c"] + + def test_spawn_runs_repo_hook_with_correct_args() -> None: calls: list[tuple[str, str, str, str]] = [] @@ -1385,3 +1588,54 @@ def test_spawner_against_the_real_service(tmp_path: Path) -> None: assert spawner.spawn_one(task) == f"panopticon-{task_id}" assert client.get_task(task_id)["claimed_by"] == "host-1" # claim recorded on the service assert spawnable_tasks(client)() == [] # now claimed → no longer spawnable + + +def test_snooze_stops_the_container_and_waking_respawns_it(tmp_path: Path) -> None: + """The whole cycle over REST: snooze → stopped + `queued`, wake → claimed + respawned.""" + service = TaskService(SqlAlchemyStore(), {"spike": Spike()}, FilesystemArtifactStore(tmp_path)) + asyncio.run(service.init()) + asyncio.run( + service.create_repo(Repo(id="r1", name="acme/widgets", git_url="https://forge/r1.git")) + ) + with TestClient(create_app(service)) as http: + client = TaskServiceClient(http) + task_id = client.create_task("r1", "spike")["id"] + runner = _FakeRunner() + clock = {"now": NOW} + + def _spawner_at() -> Spawner: + return Spawner( + client, + runner, + runner_id="host-1", # type: ignore[arg-type] + cache=CloneCache( + "/cache", run=_no_op_run, exists=lambda _p: True, makedirs=lambda _p: None + ), + tasks_root="/tasks", + git=GitClones(run=_no_op_run), + images=_FakeImageBuilder(), # type: ignore[arg-type] + makedirs=lambda _p: None, + utcnow=lambda: clock["now"], + ) + + spawner = _spawner_at() + spawner.spawn_one(client.get_task(task_id)) + assert client.get_task(task_id)["claimed_by"] == "host-1" # up and owned by this host + + client.set_snooze(task_id, _SNOOZED) # the dashboard's `e` + assert spawner.pause(client.get_task(task_id)) is True + assert runner.stopped == [f"panopticon-{task_id}"] + paused = client.get_task(task_id) + assert paused["claimed_by"] is None # claim released + # Unclaimed composes `queued` — releasing cleared the reported phase along with the claim, + # so nothing stale is left behind to read as a spawn still in flight. + assert paused["container_status"] == "queued" + assert spawner.spawn_one(paused) is None # still snoozed — stays down + assert len(runner.spawned) == 1 + + clock["now"] = NOW + timedelta(hours=13) # the deadline lapses — no event, just the clock + woken = _spawner_at() + assert woken.pause(client.get_task(task_id)) is False + assert woken.spawn_one(client.get_task(task_id)) == f"panopticon-{task_id}" + assert client.get_task(task_id)["claimed_by"] == "host-1" + assert len(runner.spawned) == 2 # respawned — same clone, same config volume From f303960c7ccf2da26bb153ed77ab40c7676bda27 Mon Sep 17 00:00:00 2001 From: Charlie Scherer Date: Wed, 23 Sep 2026 23:09:13 +0000 Subject: [PATCH 2/2] Leave the snooze hotkey descriptions as they were Review: the dashboard's help text doesn't need to spell out the container side-effect. Reverts the `e`/`E` legend strings (and the matching rows in the keybinding reference) to their original wording. Co-Authored-By: Claude Opus 5 --- docs/dashboard.md | 4 ++-- src/panopticon/terminal/dashboard.py | 10 ++-------- 2 files changed, 4 insertions(+), 10 deletions(-) diff --git a/docs/dashboard.md b/docs/dashboard.md index 17179c5e..be8d980b 100644 --- a/docs/dashboard.md +++ b/docs/dashboard.md @@ -18,8 +18,8 @@ kept in sync with the `HOTKEYS` keymap and the modal `BINDINGS` in | `R` | Respawn a down task (releases its claim so the runner re-spawns it) | | `p` | Open the task's URL in the browser | | `w` | Open the task's workdir (its per-task clone) in the host's file manager | -| `e` | Snooze the highlighted task for 12 hours (stops its container) | -| `E` | Snooze the highlighted task indefinitely (stops its container) | +| `e` | Snooze the highlighted task for 12 hours | +| `E` | Snooze the highlighted task indefinitely | | `a` | List the task's artifacts | | `A` | List the task's **repo's** artifacts (shared by every task in it) | | `g` | Open the repo config screen | diff --git a/src/panopticon/terminal/dashboard.py b/src/panopticon/terminal/dashboard.py index 8339c2ed..ec18907f 100644 --- a/src/panopticon/terminal/dashboard.py +++ b/src/panopticon/terminal/dashboard.py @@ -1917,18 +1917,12 @@ def binding(self) -> Binding: "Open the task's workdir (its per-task clone) in the host's file manager", show=False, ), - Hotkey( - "e", - "snooze", - "Snooze", - "Snooze the highlighted task for 12 hours (stops its container)", - show=False, - ), + Hotkey("e", "snooze", "Snooze", "Snooze the highlighted task for 12 hours", show=False), Hotkey( "E", "snooze_indefinitely", "Snooze sticky", - "Snooze the highlighted task indefinitely (stops its container)", + "Snooze the highlighted task indefinitely", show=False, ), Hotkey("g", "repos", "Repos", "Repo config (list / create / edit repos)", show=False),