Skip to content

feat(mcp_readability): decide what a run can reuse from its baseline - #623

Merged
akangsha7 merged 1 commit into
mainfrom
evalbench/mcp-readability-carry-forward
Oct 1, 2026
Merged

akangsha7 merged 1 commit into
mainfrom
evalbench/mcp-readability-carry-forward

Conversation

@akangsha7

@akangsha7 akangsha7 commented Sep 21, 2026 •

Copy link
Copy Markdown
Collaborator

Stacked on #637, which moved these modules into scorers/mcp_readability/. Follows #622 and #615, both now merged.

Summary
#622 records the judge's inputs and #615 reads the previous run's back. This is the logic between them: given a baseline and this run's fingerprints, decide which tools still hold, which must be re-judged, and what explains the difference. Pure functions — nothing calls them yet, so a run still judges every endpoint in full.

The invariant: same endpoint, same per-tool fingerprints, same judge fingerprint means byte-identical feedback and zero model calls. Every departure gets a named reason from an exact set-difference, not a guess.

Changes

  1. scorers/mcp_readability/carry_forward.py (new) — decide returns full, partial or carried for one endpoint; decide_with_components names which judge input changed. A changed judge fingerprint invalidates every finding, not just the ones for changed tools.
  2. Merging — merge_feedback takes the judge's entries only for changed or added tools and carries the rest. The judge is still shown the whole man page, since its severity calibration is relative to the full surface, so the filter rather than the prompt is what provides the guarantee.
  3. Finding identity — mint_finding_ids gives each finding a stable id that survives carrying, which is what makes new versus resolved countable. Tool plus rule_id is not unique, so a locator and the normalised title disambiguate.
  4. Provenance — build_provenance records where each tool's findings came from, which are new versus resolved, and the baseline job and timestamp being compared against. Persisted to mcp_readability_feedback_provenance_json only.
  5. test/mcp_carry_forward_test.py (new) — covers both failure directions: carrying a finding that should have been re-judged, and re-judging a tool nobody touched.

@akangsha7
akangsha7 force-pushed the evalbench/mcp-readability-baseline-store branch from 638cb78 to 1909c0a Compare September 28, 2026 19:24
@akangsha7
akangsha7 force-pushed the evalbench/mcp-readability-carry-forward branch 2 times, most recently from 9a4e63f to 96c0e72 Compare September 28, 2026 23:33
@akangsha7
akangsha7 force-pushed the evalbench/mcp-readability-baseline-store branch from 1909c0a to 17df0c4 Compare September 29, 2026 03:34
@akangsha7
akangsha7 force-pushed the evalbench/mcp-readability-carry-forward branch from 96c0e72 to 29831c6 Compare September 29, 2026 03:34
@akangsha7
akangsha7 force-pushed the evalbench/mcp-readability-baseline-store branch from 17df0c4 to 04a7f60 Compare September 29, 2026 03:44
@akangsha7
akangsha7 force-pushed the evalbench/mcp-readability-carry-forward branch from 29831c6 to efbd5a7 Compare September 29, 2026 03:44
@akangsha7
akangsha7 force-pushed the evalbench/mcp-readability-baseline-store branch from 04a7f60 to d16f9c7 Compare September 29, 2026 04:32
@akangsha7
akangsha7 force-pushed the evalbench/mcp-readability-carry-forward branch from efbd5a7 to 79cc805 Compare September 29, 2026 04:32
@akangsha7
akangsha7 force-pushed the evalbench/mcp-readability-baseline-store branch from d16f9c7 to 0b9373e Compare September 29, 2026 04:46
@akangsha7
akangsha7 force-pushed the evalbench/mcp-readability-carry-forward branch from 79cc805 to b78f588 Compare September 29, 2026 04:46
@akangsha7
akangsha7 force-pushed the evalbench/mcp-readability-baseline-store branch from 0b9373e to d9aca3e Compare September 29, 2026 05:11
@akangsha7
akangsha7 force-pushed the evalbench/mcp-readability-carry-forward branch from b78f588 to f9e455e Compare September 29, 2026 05:11
@akangsha7 akangsha7 self-assigned this Sep 29, 2026
@akangsha7
akangsha7 force-pushed the evalbench/mcp-readability-carry-forward branch from f9e455e to a555915 Compare September 29, 2026 19:18
Base automatically changed from evalbench/mcp-readability-baseline-store to main September 29, 2026 19:42
@akangsha7
akangsha7 force-pushed the evalbench/mcp-readability-carry-forward branch 10 times, most recently from baabe21 to 39b447a Compare September 30, 2026 01:06
@akangsha7
akangsha7 force-pushed the evalbench/mcp-readability-carry-forward branch from 39b447a to 0ba1389 Compare September 30, 2026 03:02

@helloeve helloeve left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It seems to me that MCP readability scorers are having multiple modules now, can we move the dependent modules into a dedicated directory under scorers? (similar to the dataset_quality one) and under scorers there is a top-level mcp-readability scorer which reference those modules?

@akangsha7
akangsha7 force-pushed the evalbench/mcp-readability-carry-forward branch from 0ba1389 to 34f3752 Compare September 30, 2026 19:59
@akangsha7
akangsha7 changed the base branch from main to evalbench/mcp-readability-scorers-package September 30, 2026 19:59
Base automatically changed from evalbench/mcp-readability-scorers-package to main October 1, 2026 19:24
@akangsha7
akangsha7 requested a review from helloeve October 1, 2026 19:25
Given a baseline and this run's fingerprints, work out which tools still hold
and which have to be re-judged, merge the carried findings with the fresh ones,
and record where each finding came from.

Pure functions only: nothing calls them yet, so a run still judges every
endpoint in full.
@akangsha7
akangsha7 force-pushed the evalbench/mcp-readability-carry-forward branch from 44d8651 to 8409098 Compare October 1, 2026 19:30
@akangsha7
akangsha7 merged commit f401b3d into main Oct 1, 2026
12 of 13 checks passed
@akangsha7
akangsha7 deleted the evalbench/mcp-readability-carry-forward branch October 1, 2026 20:56
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants