Conversation
akangsha7
force-pushed
the
evalbench/mcp-readability-deterministic-feedback
branch
2 times, most recently
from
September 15, 2026 17:31
4013a23 to
0da3b44
Compare
akangsha7
force-pushed
the
evalbench/mcp-readability-deterministic-feedback
branch
from
September 15, 2026 19:52
0da3b44 to
184bf31
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The style judge is a thinking model, so an identical tool surface can return different findings from one run to the next. Issue counts moved with no explanation and consumers stopped trusting them. Temperature is already 0, so the residual variance cannot be prompted away.
Rather than trying to make the judge repeatable, this stops re-judging what did not change. Given the same endpoint, the same per-tool fingerprints and the same judge fingerprint, the emitted feedback is byte-identical to the previous run and zero model calls are made. Consistency stops being a behaviour we hope a thinking model exhibits and becomes a property of the system.
Off by default. Without a
baseline:block the run behaves exactly as it does today, while still writing the fingerprint columns a later opt-in needs, so the google3-mirrored run config keeps working untouched.Changes
scorers/mcp_fingerprint.py(new) — hashes every input that can change the feedback. Each tool is hashed by its rendered man-page section rather than its schema fields, so "the fingerprints match" and "the judge would read identical bytes" are the same statement by construction; hashing fields would need a hand-maintained list of which fields matter, kept in sync with the renderer. The style guide, prompt, model, product and waivers form a separate judge fingerprint, returned with its component dict so a mismatch can name the component that moved.scorers/mcp_carry_forward.py(new) — decides what to re-judge and merges carried findings with fresh ones. The judge already groups findings per tool, so the skip is per tool: only tools whose rendered surface changed go back to the model, the rest carry forward, and a team's numbers move for the tools they actually touched. Findings the model volunteers for unchanged tools are dropped by the filter rather than discouraged by the prompt. This sits underscorers/and notevaluator/mcp_readability/on purpose: that package's__init__imports the orchestrator, which imports the scorers, so a scorer importing from it is a circular import.evaluator/mcp_readability/baseline.py(new) — reads the previous run's judgement back.NullBaselineStoreis the default and never finds one, so every run is a full judge, exactly as today.LocalResultsBaselineStorescansresults/*/evals.csv, bounded bymax_runsbecause that directory grows without limit.BigQueryBaselineStorequeries the shared results table; it is the repo's first BigQuery read path, and its lookback filters on the readability timestamp rather than the sharedrun_timecolumn, which readability rows never populate. A read failure degrades to a full re-judge and never aborts — a deliberate exception to the orchestrator's fail-fast design, since a baseline is an optimisation rather than a measurement, and the writer identity may lack read permission.scorers/mcp_style_readability.py,evaluator/mcp_readability/orchestrator.py,scorers/mcp_readability_scoring.py— wire the baseline into the scorer and report every deviation inmcp_readability_change_reason(unchanged,tools_changed,style_guide_changed,model_changed, ...), derived from an exact set-difference over fingerprint components rather than inferred. The HTML review and the result columns state why a number moved, including the affirmative "identical to the previous run" case, which matters as much as the exceptions.EndpointContext.baselineis typed loosely and defaulted last so existing positional construction keeps working.generators/models/mcp_tool_formatter.py—format_tool_section()exposes the exact bytes the judge sees for one tool, so the fingerprint can hash the rendered section instead of re-deriving it.datasets/mcp_readability/run_config.yaml— a documented, commented-outbaseline:block. Escape hatches ship with it, since carry-forward would otherwise entrench a hallucinated finding: staggered expiry (max_age_days, spread across endpoints so they do not all flip on the same day), global and per-productforce_refresh(EVALBENCH_MCP_FORCE_REFRESH=1for one-offs), and a recorded origin job so a long-carried finding stays visible rather than becoming silently authoritative. Worth knowing: waiving a rule inexceptions.yamlchanges the judge fingerprint and so re-judges that whole endpoint.Known limit
The reason is exact about which input changed and which findings appeared or disappeared. It does not explain the model's reasoning — you get "the model changed, 3 new, 2 gone", not "the new model is less strict about parameter naming".
Follow-ups
Delta columns (
p0_delta, and a fixed/new/carried breakdown) with the consumer-facing "numbers went down" view, and derivingreadability_scorefrom finding counts to remove the last nondeterministic scalar.