Repository navigation
fix(darwin): route the ruvllm mutator through a Loom shim; fail uniformity only on zero real mutations - #26
Merged
Merged
Conversation
…s happen
Every darwin@0.10.2 night since 2026-09-07 scored all mutants 0.765 because
no mutation was ever applied. darwin's RuvllmMutator sends no loom_options,
so the Ontology Loom answers each call in ~15 ms with a verbatim ontology
page ("no model generation was performed"); darwin's validator rejects it,
every variant stays byte-identical to the baseline, and the deterministic
mock sandbox scores them all alike. With the scaffold off the model takes
~60 s, past darwin's hard-coded 30 s fetch abort (RUVLLM_TIMEOUT_MS is never
read), which is a second silent no-op.
scripts/darwin-entrypoint.sh now starts a loopback shim for --mutator ruvllm
that forces loom_options.scaffold=false, floors max_tokens at 8192, sends
headers at once (darwin's abort only covers the header wait) and enforces
RUVLLM_TIMEOUT_MS upstream. Scaffold pages, truncated completions, errors
and timeouts become {"choices":[]} so darwin records a no-op, never a bogus
file; a per-verdict summary lands in the receipt.
The script also refuses --mutator ruvllm without --ruvllm-url, and the
--flag=value spellings darwin silently ignores (--sandbox=mock ran the real
sandbox). ADR-0007 records the corrected diagnosis.
Co-Authored-By: jjohare <github@thedreamlab.uk>
Co-Authored-By: jjohare <github@thedreamlab.uk>
…s made
The Loom shim now rewrites its per-verdict counts to a stats file after
every mutator call, and the entrypoint passes it to verify-entrypoint as
--shim-stats. checkDarwinBounds takes the shim's ok count as
realMutations:
- ruvllm, ok > 0, uniform: exit 0 with the receipt note
"no improvement found: N real mutations, all scored S (= baseline)"
- ruvllm, ok = 0, uniform: exit 3 ("0 real mutations: the mutator made
no edit")
- ruvllm, summary missing or unreadable: counted as 0, so a broken shim
cannot pass silently
- deterministic mutator (no shim, nothing to count): strict ADR-0006
rule unchanged
The check exists to catch an inert pipeline; an honest "no score-moving
edit" run is not a failure (owner decision, ADR-0007 §3b).
Co-Authored-By: jjohare <github@thedreamlab.uk>
ADR-0007 §3b records the owner's rule (fail only on zero real mutations) and why: the check catches an inert pipeline, which hid the scaffold defect for a month. §3c lists the three @metaharness/darwin asks. ADR-0006 status and the index point at the amendment. Co-Authored-By: jjohare <github@thedreamlab.uk>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why every darwin night scored 0.765
No model scores darwin's mutants. Under
--sandbox mock, darwin@0.10.2 scores each variant deterministically from its own surface files; the model is only the mutator. Every mutant scored 0.765 because every mutant was byte-identical to the baseline:RuvllmMutatorsends noloom_options. The Loom answered all 20 calls in about 15 ms with an ontology page (_Served verbatim from the Ontology Loom … no model generation was performed._), and darwin's validator rejected each one:mutation rejected by validator; surface unchanged (secret handling). That explains the 1.5 s nightly duration.timeoutMs ?? 30_000) and never readsRUVLLM_TIMEOUT_MS, so the180000indream.config.jsonwas dead config and each call became a silent no-op.ADR-0006 assumed the scorer was returning a default; ADR-0007 corrects that.
Part 1: the Loom shim (ADR-0007 §2)
When the command uses
--mutator ruvllm,scripts/darwin-entrypoint.shstarts a loopback shim (packages/cli/src/loomShim.ts) and points darwin at it. The shim:loom_options.scaffold=falseand amax_tokensfloor of 8192;RUVLLM_TIMEOUT_MSupstream, capped at 280 s;{"choices":[]}, which darwin records as a no-op, and logs one line per call plus a per-verdict summary into the receipt.The script also refuses
--mutator ruvllmwithout--ruvllm-url, and the--flag=valuespellings. darwin reads only--flag value, so--sandbox=mockused to pass the guard while darwin ran the real sandbox.Part 2: amended uniformity rule (owner decision, ADR-0007 §3b)
ADR-0006's check exists to catch an inert pipeline, which is what hid the scaffold defect for a month. A run where the model really edited the surfaces but nothing moved the mock score is an honest "no score-moving edit" result, not a failure.
The shim rewrites its counts to a stats file after every call. The entrypoint passes that file as
verify-entrypoint --shim-stats, andcheckDarwinBounds(rows, 0, { realMutations })applies:okno improvement found: N real mutations, all scored S (= baseline)0 real mutations: the mutator made no editEvidence
loom-shim: summary ok=20 scaffold=0 truncated=0 empty=0 upstream-error=0. The leaderboard is still uniform at 0.765: the only score-relevant edits (slice(0,30)→40,maxAttempts 3→4) fall between mock-ladder rungs, which need width 50 for 0.875 and 70 for 0.985 (checked with darwin's own scorer).--shim-statsok=20: exit 0,darwin: no improvement found: 20 real mutations, all scored 0.765 (= baseline). With the summary missing: exit 3.Tests
Tests were written failing first.
vitest run: 223/223 pass;eslintis clean.loomShim.test.ts: request rewrite, response classification, early headers, timeout.darwinEntrypoint.test.ts: the real script with a stub darwin and a fake Loom, covering scaffold off, a slow answer outliving the header abort, refusals, and wiring cases (a)ok=0+ uniform → 3, (b)ok>0+ uniform → 0 with the note, (c) no summary + uniform → 3, (d) varied → 0.index.test.tsanddarwinBounds.test.ts: the same table at the CLI and pure-function levels, plus an unreadable summary and the generation bounds still enforced alongside a pass.Not changed / follow-ups
dream.config.jsonis unchanged; nightly runs keep--mutator ruvllm.loom_options, readRUVLLM_TIMEOUT_MS, and pass stdout traces to the mutator.🤖 Generated by Claude Code