Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -36,7 +36,7 @@ Instead of spoonfeeding agents in the prompt, move details into seed data to let

Prefer deterministic checks where possible because they are cheaper, faster, repeatable, and easier to debug. Avoid being overly prescriptive with the process an agent takes to reach a solution (unless critical to the scenario), prefer checking the end state by inspecting the project or filesystem.

Reserve LLM-as-a-judge checks via `judge()` for semantic or free-form outcomes where multiple valid forms make exact checks brittle.
Reserve LLM-as-a-judge checks via `ctx.judge()` for semantic or free-form outcomes where multiple valid forms make exact checks brittle.

Prefer building checks declaratively and returning the list in one place instead of accumulating checks within branching logic, so the list remains stable if one path fails.

Expand Down Expand Up @@ -66,7 +66,7 @@ Use this checklist after the author has followed [Submitting evals for review](#

- Read [Eval criteria](#eval-criteria) and [Writing prompts](#writing-prompts). The `motivation:` should cite a real user problem, with the respective external link or Linear issue ID, and the prompt should have one clear, observable goal without leaking commands, schema details, or rubric language.
- Review [Writing scorers](#writing-scorers) assertion by assertion. Each check should map to a user requirement, avoid known false passes and false fails, and measure behavior instead of one implementation path. Don't turn incidental diagnostics into pass/fail checks.
- Ask for a deterministic or fixture-based check whenever it can replace `judge()`. Reserve judges for semantic or free-form outcomes where exact checks would be brittle.
- Ask for a deterministic or fixture-based check whenever it can replace `ctx.judge()`. Reserve judges for semantic or free-form outcomes where exact checks would be brittle.
- Require refreshed CI results, inspect at least one success run and each distinct failure shape, and confirm the experiment matches the intended suite, skills, runtime, hosted or local state, and any CLI or Docker assumptions.
- Classify each failure before requesting changes: agent gap, product gap, scorer bug, or harness/runtime failure. Ideally, show that the scorer rejects at least one known-bad solution or counterexample with a scorer unit test or a failing run linked in the PR.

Expand Down
9 changes: 9 additions & 0 deletions apps/framework/harness/run-eval.ts
Original file line number Diff line number Diff line change
Expand Up @@ -34,6 +34,7 @@ import { buildSystemPrompt } from './system-prompt.js';
import {
buildDocsResult,
buildSkillResult,
createJudgeRecorder,
evalSuiteSchema,
rehydrateTruncatedDocsResults,
getExperimentDisplayMetadata,
Expand All @@ -44,6 +45,7 @@ import type {
ExperimentConfig,
EvalInterface,
EvalManifest,
JudgeCall,
EvalMode,
EvalSuite,
ToolScorer,
Expand Down Expand Up @@ -357,6 +359,7 @@ async function runOne(
runIndex: number
): Promise<
ScoreResult & {
judgeCalls: JudgeCall[];
run: number;
skills: SkillResult;
docs: DocsResult;
Expand Down Expand Up @@ -474,11 +477,13 @@ async function runOne(
copiedWithheldTests = true;
};

const judges = createJudgeRecorder();
const last = await (scorer as LocalStackScorer)({
...session.scoringContext,
toolCalls: run.toolCalls,
transcript: run.transcript,
agentReport: run.agentReport,
judge: judges.judge,
hostWorkspace,
runViteBuild: () => viteBuild(hostWorkspace),
runVitest: () => {
Expand All @@ -497,6 +502,7 @@ async function runOne(

return {
...last,
judgeCalls: [...judges.finish()],
run: runIndex,
skills: buildSkillResult(availableSkills, run.toolCalls),
docs: buildDocsResult(run.toolCalls),
Expand Down Expand Up @@ -554,11 +560,13 @@ async function runOne(
sessionArchivePath: sessionArchivePath(expName, ev.id, runIndex),
});
const agentRunEndedAt = Date.now();
const judges = createJudgeRecorder();
const last = await (scorer as ToolScorer)({
...session.scoringContext,
toolCalls: run.toolCalls,
transcript: run.transcript,
agentReport: run.agentReport,
judge: judges.judge,
});
const scoringEndedAt = Date.now();

Expand All @@ -568,6 +576,7 @@ async function runOne(

return {
...last,
judgeCalls: [...judges.finish()],
run: runIndex,
skills: buildSkillResult(availableSkills, run.toolCalls),
docs: buildDocsResult(run.toolCalls),
Expand Down
3 changes: 2 additions & 1 deletion apps/framework/harness/types.ts
Original file line number Diff line number Diff line change
Expand Up @@ -28,14 +28,15 @@ export type {
LocalStackScorer,
ExperimentConfig,
} from '@supabase-evals/core';
export { judge, serializeTranscript } from '@supabase-evals/core';
export { serializeTranscript } from '@supabase-evals/core';
export type {
EvalInterface,
EvalMetadata,
EvalProduct,
EvalStage,
EvalSuite,
ExperimentSuite,
JudgeCall,
} from '@supabase-evals/core/eval-metadata';

export type EvalMode = 'tools' | 'local-stack';
Expand Down
3 changes: 3 additions & 0 deletions apps/framework/scripts/smoke-framework.ts
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,7 @@ import assert from 'node:assert/strict';
import { cpSync, mkdirSync, rmSync, writeFileSync } from 'node:fs';
import { dirname, join } from 'node:path';
import { fileURLToPath, pathToFileURL } from 'node:url';
import { createJudgeRecorder } from '@supabase-evals/core';
import { bootPlatformBackend } from '../harness/platform-backend.js';
import { viteBuild, vitestRun } from '../harness/project-runner.js';
import type {
Expand Down Expand Up @@ -51,6 +52,7 @@ function scorerCtx(
toolCalls: [],
transcript: extra?.transcript ?? [],
agentReport: extra?.agentReport,
judge: createJudgeRecorder().judge,
};
}

Expand Down Expand Up @@ -191,6 +193,7 @@ function serviceRoleBypassCtx(
},
toolCalls: [],
transcript: [],
judge: createJudgeRecorder().judge,
};
}

Expand Down
77 changes: 77 additions & 0 deletions apps/framework/scripts/upload-braintrust.test.ts
Original file line number Diff line number Diff line change
@@ -1,6 +1,7 @@
import { describe, expect, it } from 'vitest';
import {
experimentName,
judgeMetrics,
logTranscript,
runViewUrl,
unwrapShell,
Expand Down Expand Up @@ -135,6 +136,7 @@ describe('logTranscript', () => {
prompt: 'go',
agentReport: 'done',
checks: [],
judgeCalls: [],
passed: true,
modelId: 'm',
startTime: 100,
Expand Down Expand Up @@ -247,6 +249,7 @@ describe('logTranscript', () => {
prompt: 'go',
agentReport: '',
checks: [],
judgeCalls: [],
passed: true,
modelId: 'm',
startTime: 100,
Expand Down Expand Up @@ -280,6 +283,7 @@ describe('logTranscript', () => {
prompt: '',
agentReport: '',
checks: [],
judgeCalls: [],
passed: false,
startTime: 100,
endTime: 120,
Expand Down Expand Up @@ -311,6 +315,7 @@ describe('logTranscript', () => {
prompt: '',
agentReport: '',
checks: [],
judgeCalls: [],
passed: true,
startTime: 100,
endTime: 170,
Expand All @@ -328,6 +333,78 @@ describe('logTranscript', () => {
});
});

describe('judge spans', () => {
it('nests judge calls under the score span', () => {
const spans: Record<string, unknown>[] = [];
const sink = (parentName?: string): SpanSink => ({
startSpan(args) {
const span: Record<string, unknown> = { ...args, parentName };
spans.push(span);
return {
...sink(args?.name),
log: (event) => Object.assign(span, event),
end: (end) => Object.assign(span, end),
};
},
log() {},
end() {},
});
logTranscript(sink(), {
prompt: '',
agentReport: '',
checks: [],
passed: true,
startTime: 100,
endTime: 150,
agentEndTime: 110,
scoringEndTime: 150,
toolLabels: [],
transcript: [],
judgeCalls: [
{
provider: 'openai',
system: 'sys',
prompt: 'Rubric:\nr',
output: { passed: true, notes: 'ok' },
model: 'gpt-6-sol',
usage: { inputTokens: 100, outputTokens: 7, reasoningTokens: 5 },
startedAt: 120_000,
durationMs: 4000,
},
],
});
expect(spans.find((span) => span.name === 'passed')).toMatchObject({
type: 'score',
spanAttributes: { purpose: 'scorer' },
});
expect(spans.find((span) => span.name === 'gpt-6-sol')).toMatchObject({
parentName: 'passed',
type: 'llm',
spanAttributes: { purpose: 'scorer' },
startTime: 120,
endTime: 124,
input: [
{ role: 'system', content: 'sys' },
{ role: 'user', content: 'Rubric:\nr' },
],
output: { passed: true, notes: 'ok' },
metrics: {
prompt_tokens: 100,
completion_tokens: 7,
completion_reasoning_tokens: 5,
tokens: 107,
},
metadata: { model: 'gpt-6-sol', provider: 'openai' },
});
});

it('omits unreported judge token counts', () => {
expect(judgeMetrics({ outputTokens: 12 })).toEqual({
completion_tokens: 12,
});
});
});

describe('unwrapShell', () => {
it("drops Codex's bash -lc wrapper and leaves other commands alone", () => {
expect(unwrapShell(`/bin/bash -lc "ls -la && cat 'a b'"`)).toBe(
Expand Down
54 changes: 53 additions & 1 deletion apps/framework/scripts/upload-braintrust.ts
Original file line number Diff line number Diff line change
Expand Up @@ -11,8 +11,10 @@ import { basename, dirname, join, resolve } from 'node:path';
import { fileURLToPath } from 'node:url';
import { z } from 'zod';
import {
judgeCallSchema,
modelUsageSchema,
type AgentUsage,
type JudgeCall,
} from '@supabase-evals/core/eval-metadata';
import {
normalizeExperimentName,
Expand Down Expand Up @@ -64,6 +66,7 @@ const transcriptPartSchema = z.discriminatedUnion('type', [
}),
]);
const transcriptSchema = z.array(transcriptPartSchema).catch([]);
const judgeCallsSchema = z.array(judgeCallSchema).catch([]);
type TranscriptPart = z.infer<typeof transcriptPartSchema>;

interface PendingRow {
Expand All @@ -72,6 +75,7 @@ interface PendingRow {
agentReport: string;
passed: boolean;
checks: unknown;
judgeCalls: JudgeCall[];
modelId?: string;
modelProvider?: string;
transcript: TranscriptPart[];
Expand Down Expand Up @@ -200,6 +204,28 @@ export function tokenMetrics(
return metrics;
}

/**
* Maps only the counts the provider reported, so missing usage stays missing.
* https://github.com/braintrustdata/braintrust-spec/blob/b068e39112e081e45b6070e035877f1e2e83f9b7/skills/instrumentation-spec/references/features/token-and-cost-metrics.md#canonical-metrics
*/
export function judgeMetrics(
usage: JudgeCall['usage']
): Record<string, number> {
const metrics: Record<string, number> = {};
const set = (key: string, value: number | undefined) => {
if (value !== undefined) metrics[key] = value;
};
set('prompt_tokens', usage.inputTokens);
set('prompt_cached_tokens', usage.cacheReadInputTokens);
set('prompt_cache_creation_tokens', usage.cacheWriteInputTokens);
set('completion_tokens', usage.outputTokens);
set('completion_reasoning_tokens', usage.reasoningTokens);
if (usage.inputTokens !== undefined && usage.outputTokens !== undefined) {
metrics.tokens = usage.inputTokens + usage.outputTokens;
}
return metrics;
}

function toolLabels(toolCalls: unknown): (string | undefined)[] {
if (!Array.isArray(toolCalls)) {
return [];
Expand Down Expand Up @@ -291,6 +317,7 @@ async function collectRows(
typeof result.agentReport === 'string' ? result.agentReport : '',
passed: result.passed === true,
checks: result.checks,
judgeCalls: judgeCallsSchema.parse(result.judgeCalls),
modelId: display?.modelId,
modelProvider: display?.modelProvider,
transcript,
Expand Down Expand Up @@ -318,8 +345,10 @@ async function collectRows(
...(promptData?.product ?? result.product ?? []),
...(promptData?.topic ?? result.topic ?? []),
].map(String),
// Braintrust sums tokens across a trace's spans, so only LLM spans carry
// them. `aiSdkAgent` records no per-request usage, so its rows show none.
// https://braintrust.dev/docs/reference/sql/query-structure#summary
metrics: {
...tokenMetrics(result.usage),
// Preserve the harness's own step and tool-call counts.
...(typeof result.stepCount === 'number'
? { step_count: result.stepCount }
Expand Down Expand Up @@ -410,6 +439,7 @@ export interface SpanSink {
* │ └─ llm (text) 31s → 38s
* ├─ teardown 38s → 40s CLI exit until `agent.run()` returns
* └─ passed (score) 40s → 50s workspace export, checks, judges
* └─ gpt-6-sol (llm) 44s → 48s one per judge call
*
* Without a prompt time or a leading non-assistant message there is no setup
* span, and the first LLM span starts at the run start. Parts without a
Expand All @@ -423,6 +453,7 @@ export function logTranscript(
| 'prompt'
| 'agentReport'
| 'checks'
| 'judgeCalls'
| 'passed'
| 'modelId'
| 'modelProvider'
Expand Down Expand Up @@ -606,15 +637,36 @@ export function logTranscript(
}
const scoreStart =
row.agentEndTime === undefined ? (row.endTime ?? latestTime) : agentEnd;
// `purpose: 'scorer'` keeps judge cost out of Braintrust's preset cost charts.
// https://braintrust.dev/docs/kb/total-llm-cost-preset-requirements#what-is-happening
const scorer = parent.startSpan({
name: 'passed',
type: 'score',
spanAttributes: { purpose: 'scorer' },
startTime: scoreStart,
});
scorer.log({
output: row.checks,
scores: { passed: row.passed ? 1 : 0 },
});
for (const call of row.judgeCalls) {
const span = scorer.startSpan({
name: call.model,
type: 'llm',
spanAttributes: { purpose: 'scorer' },
startTime: call.startedAt / 1000,
});
span.log({
input: [
{ role: 'system', content: call.system },
{ role: 'user', content: call.prompt },
],
output: call.output,
metrics: judgeMetrics(call.usage),
metadata: { model: call.model, provider: call.provider },
});
span.end({ endTime: (call.startedAt + call.durationMs) / 1000 });
}
scorer.end({ endTime: latest(scoreStart, row.scoringEndTime) });
}

Expand Down
Original file line number Diff line number Diff line change
@@ -1,5 +1,4 @@
import {
judge,
serializeTranscript,
type CheckResult,
type SupabaseClient,
Expand Down Expand Up @@ -261,7 +260,7 @@ async function checkUserBCannotUploadIntoUserAFolder(
async function checkPrivateAccessConfiguration(
ctx: ToolEvalContext
): Promise<CheckResult> {
const verdict = await judge({
const verdict = await ctx.judge({
input: serializeTranscript(ctx.transcript, {
includeToolCallInputs: true,
}),
Expand Down
Loading