EvalPort interop: PairedScore/EvalRecord → TestCase/ResultSet #3
Replies: 2 comments
|
Hi Sahi, thanks for taking the time to read the actual models rather than pitching from the README. The field names you cite on I'm going to pass on adding an EvalPort emitter or importer to EvalShift itself for now. Two reasons. 1. The per-side scores are not portable on their own. You correctly keep
Only the structural evaluators score each side independently. So "each half of a pair is a normal
2. The suite side is lossier than it looks. Where an adapter could live. Every run writes If EvalPort picks up independent consumers I'm happy to revisit a native emitter. Thanks again for the careful write-up. |
|
Thanks for the genuinely detailed read, Lukas — this is exactly the kind of review that's more useful than a yes. You're right on both counts, and I don't think there's a version of this that holds up. The pairwise-relative scoring (semantic similarity pinning source at 1.0, LLM-as-judge's (0,1)/(0.5,0.5) shape, tool-call provenance pinning) means a "source ResultSet" would misrepresent what actually happened for most of your evaluators, not just lose some metadata — that's a correctness problem, not a portability inconvenience, and I'd rather EvalPort not be the reason a false "source scored 0" ends up in someone's dashboard. Same read on Your suggestion is the right shape for an external adapter if I build one: read Appreciate you engaging with the actual code instead of waving it off — and the offer to revisit if independent consumers show up. No further ask from me here. |
Uh oh!
There was an error while loading. Please reload this page.
EvalPort interop:
PairedScore/EvalRecord→ EvalPortTestCase/ResultSetHi — I maintain EvalPort, an Apache-2.0 open
spec (with Python/TS SDKs) for portable LLM eval test cases, graders, suites, and
results. I read through
src/evalshift/evaluators/base.pyand the README, and wantedto flag a concrete interop path rather than a vague "you should support X".
What I'm not proposing: EvalPort has no native notion of a paired source/target
comparison today — its
ResultSetis the output of one suite run against one model.So I'm not claiming
PairedScorealready "maps onto" something that exists. What Ithink does map cleanly, without inventing anything, is the golden-suite side:
EvalShift's golden examples (
prompt_id,example_id, input vars, expected output,and — for the agent-migration path — the recorded toolset) line up almost 1:1 with
EvalPort's
TestCaseschema:id,input,expected_output,tools_called,expected_tools. That last pairexists specifically for agent/tool-call evaluation, which is the same thing your
README calls out as the killer scenario (an agent silently dropping
notify_security_teamafter a migration, text-eval green, tool-call eval CRITICAL).On the output side, each half of a pair — the source-model run and the
target-model run — is a normal EvalPort
ResultSet:results[].grader_results[]already carriesgrader_id,type,score,passed,reason.PairedScore.source_score/PairedScore.target_scoreare just two ofthose, one per model run, keyed by the same
test_case_id.PairedScore.deltaandEvalRecord.blockingare genuinely EvalShift-specific — they belong to youranalysis layer (Cohen's d, paired tests, BH correction across prompt×evaluator×slice),
not to a generic result format, and I wouldn't want to force them into the spec.
Rough sketch, using your real field names, of what an adapter would look like (this
is illustrative, not a PR):
Why this might be worth 20 minutes of thought rather than nothing: if EvalShift's
capture sync/ golden suites spokeTestCase, andevalshift run/reportcouldoptionally emit a
ResultSetper model alongside your ownscores.jsonl, the samecaptured golden suite would become runnable through anything else that already
consumes EvalPort — for a concrete, verified sense of what "anything else" means
today, two of the 37 merged adapters are
ragas-openeval-adapter(RAG-focused paired-style metrics, closest in spirit to your semantic-similarity
evaluator) and
mlflow-openeval-adapter(run-artifact based, closest in spirit to your
.evalshift/runs/<run-id>/layout).None of that requires EvalShift to change its own analysis or statistics — it's
purely an additional output format at the edges (
capture syncin,report/bundleout), so it wouldn't touch the AGPL-licensed statistical core.
If this isn't a direction you're interested in, no worries at all — I mostly wanted
to make sure the specific field-level claim was accurate rather than hand-wavy before
raising it. Happy to sketch a real converter if there's interest; I'd rather do that
as a follow-up than guess at scope here.
— Sahi, independent contributor (not affiliated with this project)
All reactions