Skip to content

feat(mcp_readability): reuse unchanged findings so counts track tool changes - #604

Closed
akangsha7 wants to merge 1 commit into
GoogleCloudPlatform:mainfrom
akangsha7:evalbench/mcp-readability-deterministic-feedback
Closed

akangsha7 wants to merge 1 commit into
GoogleCloudPlatform:mainfrom
akangsha7:evalbench/mcp-readability-deterministic-feedback

Conversation

@akangsha7

@akangsha7 akangsha7 commented Sep 14, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

The style judge is a thinking model, so an identical tool surface can return different findings from one run to the next. Issue counts moved with no explanation and consumers stopped trusting them. Temperature is already 0, so the residual variance cannot be prompted away.

Rather than trying to make the judge repeatable, this stops re-judging what did not change. Given the same endpoint, the same per-tool fingerprints and the same judge fingerprint, the emitted feedback is byte-identical to the previous run and zero model calls are made. Consistency stops being a behaviour we hope a thinking model exhibits and becomes a property of the system.

Off by default. Without a baseline: block the run behaves exactly as it does today, while still writing the fingerprint columns a later opt-in needs, so the google3-mirrored run config keeps working untouched.

Changes

  1. scorers/mcp_fingerprint.py (new) — hashes every input that can change the feedback. Each tool is hashed by its rendered man-page section rather than its schema fields, so "the fingerprints match" and "the judge would read identical bytes" are the same statement by construction; hashing fields would need a hand-maintained list of which fields matter, kept in sync with the renderer. The style guide, prompt, model, product and waivers form a separate judge fingerprint, returned with its component dict so a mismatch can name the component that moved.

  2. scorers/mcp_carry_forward.py (new) — decides what to re-judge and merges carried findings with fresh ones. The judge already groups findings per tool, so the skip is per tool: only tools whose rendered surface changed go back to the model, the rest carry forward, and a team's numbers move for the tools they actually touched. Findings the model volunteers for unchanged tools are dropped by the filter rather than discouraged by the prompt. This sits under scorers/ and not evaluator/mcp_readability/ on purpose: that package's __init__ imports the orchestrator, which imports the scorers, so a scorer importing from it is a circular import.

  3. evaluator/mcp_readability/baseline.py (new) — reads the previous run's judgement back. NullBaselineStore is the default and never finds one, so every run is a full judge, exactly as today. LocalResultsBaselineStore scans results/*/evals.csv, bounded by max_runs because that directory grows without limit. BigQueryBaselineStore queries the shared results table; it is the repo's first BigQuery read path, and its lookback filters on the readability timestamp rather than the shared run_time column, which readability rows never populate. A read failure degrades to a full re-judge and never aborts — a deliberate exception to the orchestrator's fail-fast design, since a baseline is an optimisation rather than a measurement, and the writer identity may lack read permission.

  4. scorers/mcp_style_readability.py, evaluator/mcp_readability/orchestrator.py, scorers/mcp_readability_scoring.py — wire the baseline into the scorer and report every deviation in mcp_readability_change_reason (unchanged, tools_changed, style_guide_changed, model_changed, ...), derived from an exact set-difference over fingerprint components rather than inferred. The HTML review and the result columns state why a number moved, including the affirmative "identical to the previous run" case, which matters as much as the exceptions. EndpointContext.baseline is typed loosely and defaulted last so existing positional construction keeps working.

  5. generators/models/mcp_tool_formatter.py — format_tool_section() exposes the exact bytes the judge sees for one tool, so the fingerprint can hash the rendered section instead of re-deriving it.

  6. datasets/mcp_readability/run_config.yaml — a documented, commented-out baseline: block. Escape hatches ship with it, since carry-forward would otherwise entrench a hallucinated finding: staggered expiry (max_age_days, spread across endpoints so they do not all flip on the same day), global and per-product force_refresh (EVALBENCH_MCP_FORCE_REFRESH=1 for one-offs), and a recorded origin job so a long-carried finding stays visible rather than becoming silently authoritative. Worth knowing: waiving a rule in exceptions.yaml changes the judge fingerprint and so re-judges that whole endpoint.

Known limit

The reason is exact about which input changed and which findings appeared or disappeared. It does not explain the model's reasoning — you get "the model changed, 3 new, 2 gone", not "the new model is less strict about parameter naming".

Follow-ups

Delta columns (p0_delta, and a fixed/new/carried breakdown) with the consumer-facing "numbers went down" view, and deriving readability_score from finding counts to remove the last nondeterministic scalar.

@akangsha7
akangsha7 force-pushed the evalbench/mcp-readability-deterministic-feedback branch 2 times, most recently from 4013a23 to 0da3b44 Compare September 15, 2026 17:31
@akangsha7
akangsha7 force-pushed the evalbench/mcp-readability-deterministic-feedback branch from 0da3b44 to 184bf31 Compare September 15, 2026 19:52
@akangsha7 akangsha7 self-assigned this Sep 30, 2026
@akangsha7

Copy link
Copy Markdown
Collaborator Author

Superseded. This landed as a stack instead: #615 (baseline read-back), #622 (judge inputs), #623 (carry-forward logic), and #624 (wiring, open). The man-page format_tool_section split is already on main via 43d6bcf. Leaving the branch in place.

@akangsha7 akangsha7 closed this Oct 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant