Repository navigation
feat(darwin): flag a leaderboard where every mutant scores the same - #21
Merged
Merged
Conversation
Wire the score-uniformity check into the darwin bound check (ADR-0006). parseLeaderboardRows now reads each row's score, replacing the second parser (SCORED_ROW_RE). checkDarwinBounds reports two or more scored mutants with one identical score as a violation, so verify-entrypoint darwin and scripts/darwin-entrypoint.sh exit 3 under "darwin bounds VIOLATED", the channel ADR-0005 bound breaches already use. Fewer than two scored mutants is not a breach. The real 2026-09-07 fixture (all 0.765) is now the uniformity breach case; a discriminating fixture is the compliant ceiling. Co-Authored-By: jjohare <github@thedreamlab.uk>
jjohare
force-pushed
the
dream/security-adversarial-2026-10-03
branch
from
October 6, 2026 13:49
34f973d to
49d9efd
Compare
jjohare
marked this pull request as ready for review
October 6, 2026 13:49
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
A darwin leaderboard where every scored mutant has the identical score means the scorer is not discriminating (a default value, a fallback answer, or a parse error). The 2026-09-07 run and every pinned
@metaharness/darwin@0.10.2receipt since look like this: all mutants at 0.765,Delta over baseline: +0.000. Those runs still exited 0. This PR makes them fail the darwin evaluator. Decision recorded as ADR-0006.How
parseLeaderboardRowsnow reads each row's score. The second row parser (SCORED_ROW_RE) is gone.checkDarwinBoundscallscheckDarwinScoreUniformity(rows). If two or more scored mutants (non-baseline rows) all share one score, it adds the violationscore uniformity: all N mutants score S, equal to baseline — scorer not discriminating. When the mutants agree with each other but not with the baseline, the message readsbaseline Binstead. Fewer than two scored mutants is not a breach.verify-entrypoint darwinprints the violation underdarwin bounds VIOLATEDand exits 3.scripts/darwin-entrypoint.shis the configured darwin evaluator and passes that exit code through, so the annexe records the REQUIRED evaluator as FAILED. The nightly script needed no change, because it never parses leaderboards itself.darwin-leaderboard-varied.txt) is the compliant 4 + 5 ceiling.six-in-g2is rebuilt on top of it, so it isolates the candidate bound.Consequence
Until the scorer or the mutator changes, darwin nights will exit 3 and carry the uniformity line in their receipt. Before this change they passed on a leaderboard that meant nothing.
Tests
darwinBounds.test.ts: the parser reads scores. Uniform-equal-to-baseline and uniform-not-baseline cases are flagged. One differing mutant is clean, a lone mutant is clean, and hand-built rows without scores are ignored.index.test.ts:verify-entrypoint darwinexits 3 on the uniform leaderboard and 0 on the varied one.darwinEntrypoint.test.ts: the real script against the stub darwin covers the same two cases.npm ci,npm run build,npm run lint,npx vitest run --coverageall pass (10 files, 191 tests).The PR does not touch
docs/dream-cycle/LEDGER.md.🤖 Generated by Claude Code