Examples live under examples/.
Current example directories:
basic-jsontaxonomy-complete-fixtureagent-quality-presetllm-agent-evalsfinancial-domain-ollama-evalsvitest-evalsjest-custom-reporternode-plain-evalpython-pytest-evalslangchain-evalsgithub-actionsazure-devopsstatic-html-dashboardcustom-reporter-pluginbenchmark-packsscreenshot-fixturereport-power-artifacts
Start with:
eval-dashboards report --input=examples/basic-json --reporter=html --report-dir=eval-reportUse this when you already have eval-report/v1 JSON and only want to run report/check/publish.
eval-dashboards report --input=examples/basic-json --reporter=html --reporter=text --report-dir=eval-report
eval-dashboards publish --target=dir --input=examples/basic-json --report-dir=eval-report --out-dir=published-eval-reportUse a passing fixture for a "green" gate example:
eval-dashboards check --input=examples/agent-quality-preset/artifacts --allow-blocked-baselineexamples/basic-json/run-trace-links.json includes rows[].trace with traceId, spanId, traceUrl, and spanUrl.
eval-dashboards report --input=examples/basic-json --run-id=run-trace-links --reporter=html --report-dir=eval-reportOpen eval-report/index.html, find row agent/tool-timeout-001, then follow the row-level trace/span links from the details panel. This demonstrates dashboard row -> trace deep link triage.
Use this for concrete, local-openable artifacts demonstrating history, progress, gate outcomes, and row-level detail analysis.
./scripts/generate-report-power-artifacts.shArtifacts produced under examples/report-power-artifacts/:
report/history.json(history trends)report/summary.json(progress + detailed comparison)gates/check-pass.json(passing gate result)gates/check-fail.json(failing gate result)report/index.htmlandreport/summary.md(human-readable detail views)
Use this when starting an agent-eval program and you want scaffolded suite/rubric/dataset conventions first.
eval-dashboards init --preset=agent-quality
eval-dashboards teach
eval-dashboards init --preset=agent-quality --write --dry-run
eval-dashboards lint --input=examples/agent-quality-preset/artifacts
eval-dashboards check --input=examples/agent-quality-preset/artifacts --allow-blocked-baseline
eval-dashboards report --input=examples/agent-quality-preset/artifacts --reporter=html --reporter=json-summary --report-dir=eval-reportNote: examples/agent-quality-preset/artifacts includes two reports by design:
run-agent-quality-template.json(current run)run-agent-quality-calibration.json(independent calibration evidence,run.kind="calibration"so it is excluded from automatic baseline selection)
This keeps blocking calibration preflight examples runnable now that same-run calibration rows no longer satisfy blocking checks.
Use this as a local reference for agent/chat eval rows with tool-call and judge evidence.
pnpm example:llm-agent-reportGenerated artifacts are written under examples/llm-agent-evals/.evals_output/.
- Replace the local runner call with your runtime call.
- Keep scenario ids stable across runs.
- Record
promptVersionandagentVersionper row. - Put provider request ids and trace refs in
metadata/tracefields.
Use this for a real local-model eval run (Ollama) in a regulated-domain style workflow.
pnpm example:financial-domain-ollama-reportIf Ollama is unavailable, treat this suite as opt-in infrastructure-dependent coverage.
Use this when you want versioned suite bundles for safety, tool-routing, and groundedness planning.
Compatibility guidance: docs/benchmark-packs.md.
cat examples/benchmark-packs/safety-pack.v1.json
cat examples/benchmark-packs/tool-routing-pack.v1.json
cat examples/benchmark-packs/groundedness-pack.v1.jsonThese are template inputs for suite/dataset/rubric planning. Emit normal eval-report/v1 run artifacts after applying them.
Runnable local examples:
vitest-evalsjest-custom-reporternode-plain-evalpython-pytest-evalslangchain-evalsllm-agent-evalsfinancial-domain-ollama-evalsagent-quality-presetbasic-jsontaxonomy-complete-fixture
CI templates and workflow examples:
github-actionsazure-devops
Reference/template examples:
static-html-dashboardcustom-reporter-pluginbenchmark-packsscreenshot-fixturereport-power-artifacts
Use docs/teach-labs/README.md for delivery-stage labs and the FDE role workflow.
- Local dev loop:
docs/teach-labs/01-local-dev-loop.md - Pre-PR gating:
docs/teach-labs/02-pre-pr-gating.md - PR review triage:
docs/teach-labs/03-pr-review-triage.md - Release readiness:
docs/teach-labs/04-release-readiness.md - Post-release monitoring:
docs/teach-labs/05-post-release-monitoring.md - FDE role analysis + workflow:
docs/teach-labs/fde-role-workflow.md