Skip to content

A watchdog for the one question no instrument asked - #4495

Merged
gHashTag merged 1 commit into
masterfrom
feat/watchdog
Sep 21, 2026
Merged

gHashTag merged 1 commit into
masterfrom
feat/watchdog

Conversation

@gHashTag

Copy link
Copy Markdown
Owner

Closes #4494

$ curl -s .../queen/status
{"status":"error","code":502,"message":"Application failed to respond"}

Nine hours of that, with the whole swarm stopped. The agent server
crash-looped, spent Railway's restartPolicyMaxRetries: 10 — a budget, not a
promise — and the platform stopped restarting it.

Every instrument built this week reads the swarm's own numbers. The pusher
exits 2 when the endpoint is silent: one red scheduled run, nothing else.
Nothing asked whether it answers at all.

What it refuses

refusal why
one bad probe a deploy answers 502 too — two probes a minute apart must both fail
no token it alarms and says it restarted nothing, rather than pretend
twice in a row a redeploy that did not help is not fixed by another

The alarm issue closes itself when the swarm answers again.

$ python3 tools/queen/watchdog.py --self-test
ok: 4 answers, including three a live swarm never sends

One of those three is the platform's own 502 body: valid JSON, and not a swarm.
A 200 carrying workers is the only yes.

Both ledgers a new workflow moves are updated in this same commit.

To have it restart the service by itself, three secrets are needed —
RAILWAY_TOKEN, RAILWAY_SERVICE_ID, RAILWAY_ENVIRONMENT_ID.

🤖 Generated with Claude Code

The agent server crash-looped, spent Railway's restartPolicyMaxRetries budget,
and the platform stopped restarting it. `/queen/status` then answered
`{"status":"error","code":502,"message":"Application failed to respond"}` for
NINE HOURS with the whole swarm stopped.

Every instrument built this week reads the swarm's OWN numbers. The pusher
exits 2 when the endpoint is silent - one red scheduled run, nothing else.
Nothing asked whether it answers at all.

watchdog.py asks every five minutes and is the only thing here allowed to
restart the service. It refuses on one bad probe, because a deploy answers 502
too; refuses without a token, rather than pretending to have healed anything;
and refuses to ask twice in a row, because a redeploy that did not help is not
fixed by another. The alarm issue closes itself when the swarm answers.

The probe is tested against the platform's own 502 body, which is valid JSON
and is not a swarm: a 200 carrying `workers` is the only yes.

CENSUS: quiet 60 -> 61 workflow files; shell 60 -> 61 files, 81 -> 82 jobs,
273 -> 275 run-steps. Classified in check_pr_branch_filters.py in this commit.

Closes #4494

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@gHashTag
gHashTag enabled auto-merge (squash) September 21, 2026 04:56
@github-actions

Copy link
Copy Markdown
Contributor

PR Dashboard

Generated at: 2026-09-21 05:00:20 UTC

Summary

Status Count
Total Open PRs 21
PRs with Failing Checks 18
PRs with All Checks Green 3
READY 2
FAILING 18
PENDING 0
NO CHECKS YET 0

These columns do not partition: 2 + 18 + 0 + 0 = 20, and there are 21 open PRs. A PR is being counted twice or not at all.

Seal Status

  • ⚠️ STALE -- sha256(compiler.rs)=403499176a5d != manifest seal=87e5cbd3ad94.
    The committed NMSE numbers were certified against an older compiler.rs.
    Run scripts/reseal-check.sh locally for the two-step reseal command (advisory; not a merge gate).

@github-actions

Copy link
Copy Markdown
Contributor

📓 NotebookLM Notebook linked to this PR

This notebook contains session context, decisions, and artifacts for this work.

@gHashTag
gHashTag merged commit a152c20 into master Sep 21, 2026
27 of 30 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

The swarm answered 502 for nine hours and no instrument read it

2 participants