Skip to content

The sweep is bounded by a share of the tick, not four minutes - #493

Merged
gHashTag merged 5 commits into
feat/queen-supervisorfrom
fix/sweep-timebox
Sep 21, 2026
Merged

gHashTag merged 5 commits into
feat/queen-supervisorfrom
fix/sweep-timebox

Conversation

@gHashTag

Copy link
Copy Markdown
Owner

runRound awaits the review sweep before it reaps and before it
dispatches. Its deadline is therefore how long the swarm is willing to hand out
no work at all — and it was an absolute four minutes against a sixty-second
tick
: four rounds of silence, with the lease heartbeat reporting health.

Measured today

worker model changed   →  8.5 → 80 dispatches/hour
review sweep           →  3 a round, unchanged
unreviewed             →  16–20 all afternoon

Raising the count to 8 reviews and 6 measurements killed the container five
minutes later. A count bounds worktrees, not time: each measurement cuts a
temporary worktree and runs up to twenty commands.

So the bound is time

sweepDeadlineMs(60)   = 45_000     45 s of a 60 s round
sweepDeadlineMs(5)    = 20_000     floor — a short tick must not make review impossible
sweepDeadlineMs(3600) = 240_000    the old ceiling, for a deployment ticking in minutes

The count can be generous now, because the clock protects the dispatcher.

The deadline applies to reviews as well as measurements. A review is a
provider call and can sit on its own timeout; bounding only the measurements
left the dispatcher waiting on the half that was never counted.

bun test queen-tick-sweep queen-review-unjudged queen-adversarial-review \
         queen-criteria-run queend-choose
92 pass, 0 fail

🤖 Generated with Claude Code

TWO FAILURES WITH ONE SYMPTOM: an empty branch and a turn that reads as a model
failure. 80 of 101 stuck issues on 2026-09-20 had exactly that shape.

THE CHECKOUT. Measured in the running deployment's log:

  error: Your local changes to the following files would be overwritten by checkout:
  error: unable to create file specs/port/tools/gft_deep_demo.t27: Permission denied

$WORKSPACE_DIR was owned by the bee, so the one-time ownership walk was skipped,
while files underneath it were not - left by a root-run git from an older image.
`find ! -user -print -quit` stops at the FIRST wrong file, so the healthy case
costs one stat and the 45 GB walk that once outlasted the 300 s healthcheck
cannot come back. The repair walks the checkout only, never the worktrees
beside it, and changes only what is wrong.

THE FETCH. Measured at concurrency four on one key against
integrate.api.nvidia.com: two answers 200, one 503, and one socket that never
answered at all. The retry wrapper handled 429 and 5xx and the 200-carrying-an-
error case, and could not see the fourth - there is no response to branch on,
so it reached the agent loop as a terminal error and ended the turn. A throw is
now retried on the same backoff. An abort the CALLER asked for is re-thrown at
once: retrying a cancelled request outlives the thing that cancelled it.

  bun test apps/server/src/lib/overload-retry-fetch.test.ts
  10 pass, 0 fail

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Every tick for six minutes chose an issue and then refused to start it:

  Queen tick chose an issue but the container cannot carry another bee
    issue=4438 resource="disk"

Twenty lanes open, 684 candidates waiting, and zero bees - because the volume
had filled with bee worktrees nobody could use. An earlier reading of the same
volume found 41 of them holding 45 GB and three million inodes.

At entrypoint time no bee is running - this process is what starts the server
that starts them - so every directory under .worktrees/ belongs to a container
that is already gone. `git worktree prune` alone does not do it: it drops the
admin records for directories already removed, and these are still there.

Free space is printed before and after, so the next reader sees what it bought.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A refusal costs a provider call and returns nothing, and the provider is the
ceiling. Measured on the running deployment 2026-09-20 in one 400-line window:
18 of 19 filesystem tool failures were this refusal, on files of 159 and 500
lines - ordinary specs - while the same endpoint answered `Service temporarily
overloaded` 76 times in that same window. Every one of those refusals spent a
round trip to be told to ask again.

filesystem_read now returns the lines that fit under the character limit and
names the exact offset to continue from, which is what the caller would have
asked for on its second call. Room is kept for that note, so the answer can
always carry one.

A single line longer than the whole budget still throws - there is nothing to
hand back - and now says so in those words, pointing at filesystem_grep.

  bun test apps/server/tests/tools/filesystem/read.test.ts
  17 pass, 0 fail

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…utes

`runRound` awaits the review sweep before it reaps and before it dispatches, so
the sweep's deadline is how long the swarm is willing to hand out no work at
all. It was an absolute four minutes against a sixty-second tick - four rounds
of silence, with the lease heartbeat reporting health throughout.

Measured 2026-09-20: the swarm went from 8.5 to 80 dispatches an hour when the
worker model changed, and the review sweep - three a round, unchanged since it
was written - fell behind, sixteen to twenty unreviewed all afternoon. Raising
the COUNT to eight reviews and six measurements killed the container five
minutes later, because a count bounds worktrees and not time: each measurement
cuts a temporary worktree and runs up to twenty commands.

So the bound is time, and a fraction of the tick: 45 s of a 60 s round, floored
at 20 s so a very short tick cannot make review impossible, and ceilinged at
the old four minutes for a deployment whose tick is minutes long. The count can
now be generous because the clock is what protects the dispatcher.

The deadline applies to REVIEWS as well as measurements. A review is a provider
call and can sit on its own timeout; bounding only the measurements left the
dispatcher waiting on the half that was never counted.

  bun test queen-tick-sweep queen-review-unjudged queen-adversarial-review
  queen-criteria-run queend-choose        92 pass, 0 fail

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Sep 20, 2026

Copy link
Copy Markdown

❌ Tests failed — 11/2436 failed

Suite Passed Failed Skipped
agent 87/87 0 0
build 9/9 0 0
cdp-protocol 5/5 0 0
eval 93/93 0 0
server-agent 272/272 0 0
server-api 1252/1315 6 57
server-browser 6/6 0 0
server-integration 10/11 0 1
server-lib 279/279 0 0
server-pglive 1/3 2 0
server-root 68/68 0 0
server-skills 31/31 0 0
server-tools 240/243 3 0
shared 14/14 0 0
Failed tests
  • server-apithe note says who committed the branch, where it can be read > puts the salvage fact where the 1500-character cap cannot cut it
  • server-apia turn the provider killed does not spend the issue > is judged and sent back, but charges neither counter
  • server-apideriving work the repository already measured > carries the command that produced it
  • server-apideriving work the repository already measured > gives every candidate a path that can be a boundary
  • server-apideriving work the repository already measured > proposes nothing from code this project does not own
  • server-apisalvaging a turn that ended with its work uncommitted > refuses a worktree holding an unmerged path
  • server-pglivethe migration block, applied to a real PostgreSQL > creates every object it promises, and survives a second boot
  • server-pglivethe live gate > ran, or its absence is on the record
  • server-toolsnavigation tools > new_hidden_page opens a hidden tab
  • server-toolsnavigation tools > show_page restores a hidden page to visible
  • server-toolswindow tools > create_hidden_window creates and closes a hidden window

View workflow run

The emergency sweep keeps any worktree whose branch carries commits the base
does not, because the container holds no push credential and such a commit
would live in exactly one place. That is right, and it stops being true the
moment the branch is on origin at the same commit.

Measured 2026-09-20: the entrypoint found TWENTY-ONE worktrees at boot, every
one of them kept by that branch of the check. The volume filled, the process
was killed, the entrypoint cleared them, the swarm ran for a few minutes and it
happened again - six restarts in one afternoon, at four lanes as readily as at
twenty, because the leak is per FINISHED bee and not per running one.

So the sweep now reads the remote. A branch on origin at the same SHA is
published, and its worktree is disk. Reading the remote needs no push
credential; if it cannot be read the tree stays, because unreachable is not
published for the same reason unreadable is not clean.

  bun test apps/server/tests/api/queen-dispatch.test.ts   96 pass, 0 fail

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@gHashTag
gHashTag merged commit ca79519 into feat/queen-supervisor Sep 21, 2026
12 of 16 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants