Update this file as features are implemented. Keep it honest: mark an item done only after the relevant code and tests exist.
See also: ROADMAP.md for the prioritized improvement plan.
-
Create independent project directory.
-
Add npm package metadata for
@icodenet/eval-dashboards. -
Add TypeScript, tsup, and Vitest project skeleton.
-
Add CLI binary entry point (
eval-dashboards). -
Add PRP documentation.
-
Add versioned
eval-report/v1model and validation. -
Add first-class optional agent and LLM judge report fields.
-
Add portable suite manifest, gate policy, and rubric contract fields.
-
Add baseline compatibility assessment for dataset and rubric version drift.
-
Add basic report discovery and history building.
-
Add latest-vs-previous comparison.
-
Add initial gate checking.
-
Add starter
text,json-summary,markdown-summary, andhtmlreporters. -
Add local directory publish target.
-
Implement publishing adapters for GitHub Pages and Azure Storage, plus Azure Static Web Apps dry-run validation mode.
-
Add required example directories and starter artifacts.
-
Add runnable provider-free agent/chat eval example that generates artifacts and reports.
-
Add focused starter tests.
-
Add focused governance and baseline compatibility tests.
-
Implement
eval-dashboards initconfig generation. -
Implement full config loading from
eval-dashboards.config.ts,eval-dashboards.config.js, andpackage.json. -
Align documentation: sweep and replace old
@icodenet/eval-reports/eval-reportsnaming with@icodenet/eval-dashboards/eval-dashboards. -
Add GitHub Actions CI workflow (typecheck, test, build, example smoke tests).
-
Export JSON Schema for
eval-report/v1toschemas/eval-report-v1.schema.json(published and maintained in-repo). -
Create comprehensive taxonomy teaching documentation at
docs/taxonomy.md(definitions, examples, checklist, FAQ). -
Create taxonomy-complete init fixture at
examples/taxonomy-complete-fixture/run-complete.jsonwith README demonstrating best practices. -
Create runner cookbook: Vitest example with README and patterns.
-
Create runner cookbook: Jest custom reporter example with README and patterns.
-
Create runner cookbook: Node plain eval example with README and use cases.
-
Refresh README for adoption messaging, schema-first positioning, and cookbook links.
-
Implement HTML grouping by dataset and scenario with collapsible sections.
-
Add taxonomy completeness score (0–100%) to row display with visual indicators.
-
Add kind badges (deterministic, agent, llm-judge, human-review) to row display.
-
Add "All rows (by dataset & scenario)" section with full grouping.
-
Implement Azure Static Web Apps dry-run validation path (non-dry-run execution path pending).
-
Implement Azure Storage static website publishing adapter (including dry-run validation path).
-
Add persistent failure detection with
analyzeRowStability()function. -
Add flaky row classification based on pass/fail history across runs.
-
Create CONTRIBUTING.md with development workflow, project structure, and commit guidelines.
-
Create CODE_OF_CONDUCT.md (Contributor Covenant-based).
-
Create GitHub issue templates (bug report, feature request).
-
Create CHANGELOG.md with semantic versioning guidance.
-
Implement explicit baseline selection by run id (enhancement to
compareRuns). -
Enhance HTML dashboard with sparklines and pass-rate trends in history view.
-
Create README screenshot gallery (light and dark themes) — visual proof of UI.
-
Add concrete report-power artifact fixture with tracked history/progress/gate/detail outputs and deterministic regeneration script.
-
Add teach delivery-stage labs plus FDE role workflow guidance grounded in report artifacts and required evidence outputs.
-
Add npm publishing workflow and semantic version tagging (GitHub Actions).
- Audit 2026-09-15:
publish.yml/release.ymlareworkflow_dispatch-only — there is nopush: tagstrigger, so tag-triggered publishing is not implemented yet; manual dispatch works.
- Audit 2026-09-15:
- The branch history was rewritten to reflect the current codebase.
- Only the latest post-rewrite release should be treated as the valid reference for the current implementation.
- Earlier release artifacts are superseded and should not be used to evaluate the present code state.
- Create community feedback loop infrastructure and early-runner outreach tracker (Phase 4 readiness).
- Added weekly metrics loop via
pnpm metrics:adoptionand snapshot output indocs/adoption-metrics/latest.json. - Added manual signal tracker at
docs/adoption-metrics/manual-signals.json. - Added partnership log and outreach stages in
docs/community-partnership-log.md. - External adoption outcomes continue as live KPIs, not static checklist items.
- Added weekly metrics loop via
-
reference integration: inspect existing eval runner, dataset shape, scoring, and CI workflow.
-
reference integration: add
@icodenet/eval-dashboards@0.3.0as an explicit dev dependency.- Audit 2026-09-15: unverifiable from this repo — no external commit/PR link recorded; ROADMAP.md tracks this item as unchecked/blocked for the same reason. Corrected here to match.
-
reference integration: map current eval output into
eval-report/v1without replacing the existing runner. -
reference integration: emit
.evals_output/*.jsonartifacts with suite summaries and row-level evidence. -
reference integration: add suite manifests, dataset versions, rubric versions, and dashboard gates.
-
reference integration: wire
eval-dashboards lint,check, andreportinto local/CI eval commands. -
reference integration: add rubric contracts plus row provenance and lifecycle metadata.
-
reference integration: surface the generated
/eval-dashboard/report in the learning UI instead of the old bespoke summary dashboard.- Audit 2026-09-15: contradicted by
docs/case-studies/assistant-ui/README.md— adapter and CLI wiring exist in a local worktree only, no PR opened; ROADMAP.md tracks this item as unchecked/blocked for the same reason. Corrected here to match.
- Audit 2026-09-15: contradicted by
-
reference integration: create first published dashboard baseline and document quality gaps.
-
eval-dashboards: publish TypeScript declaration files and package metadata so downstream imports resolve public types.
- Audit 2026-09-15:
dist/*.d.tsship correctly, but the packed-package consumer smoke test named in this slice is still missing — the downstream compile was a one-off manual check, not a checked-in test (matches ROADMAP.md's caveat on the same item).
- Audit 2026-09-15:
-
eval-dashboards: define agent-quality suite presets (
retrieval-recall,answer-groundedness,answer-quality,refusal-safety,prompt-injection-resilience,mcp-routing,content-coverage,regression-incidents,judge-calibration). -
eval-dashboards: decide which setup concepts belong in schema fields/enums, preset files, examples, or docs.
-
eval-dashboards: design setup scaffolding for common agent eval programs, such as
init --preset agent-quality. -
eval-dashboards: add starter dataset/rubric templates with versioning, provenance, lifecycle, and judge calibration examples.
-
eval-dashboards: document how presets map to
riskArea,target,graders, gate policies, and rubric contracts. -
eval-dashboards: add a repo-context glossary explaining
eval-report/v1, suite, dataset, rubric, runner, and row terminology. -
eval-dashboards: add runner-adapter primitives so teams with an existing eval runner can map local results into
eval-report/v1without hand-writing aggregate, manifest, rubric, and output-cleanup boilerplate. -
Cross-feed: use reference integration learnings to amend eval-dashboards roadmap, templates, and docs before stabilizing setup-layer APIs.
- Captured so far: prefer directory inputs over config globs, require rubric versions for blocking suites, clean generated artifact directories before writing, make suite summaries row-complete, and expose/embed the generated static dashboard instead of duplicating it with host-app summary cards.
- Type packaging captured: emit declarations and expose them with
main,types, andexports; verified withpnpm build,npm pack, and a temporary downstream TypeScript compile against the packed tarball. - Adapter boundary captured: keep project-specific dataset rows local, but move repeated artifact assembly mechanics into public eval-dashboards helpers.
- Approval-gate pattern captured: document
eval-resultsbranch layout,pr-meta.jsonwiring, commit-status contract (eval/quality-gate), environment approval flow, and cleanup workflow templates for closed PRs. - Planning slices captured in ROADMAP: dataset governance, versioned rubrics, judge calibration, CI quality tiers, suite templates, setup scaffolding, and schema/taxonomy decision rules.
-
Research and publish industry coverage audit for suites/datasets/rubrics.
- Added docs/industry-coverage-audit.md with external-source mapping and local coverage matrix.
- Identified P0 additions:
goal-success,intent-resolution,task-adherence,sensitive-disclosure, andagency-boundarypresets.
-
Implement P0 industry coverage suites in presets, dataset templates, rubrics, artifact template, and init scaffold.
- Scope tracked in docs/industry-coverage-audit.md.
- Must include parity updates across docs + examples +
src/cli/init-scaffold.ts. - Added starter multi-turn trajectory coverage (
multiturn-trajectory) with preset guidance, dataset case, rubric axes, template artifact row, and scaffold output.
-
Phase 4H docs navigability: added
docs/README.mdaudience index (4H.1), regrouped README's Documentation Index into the same audience groups with a Start-here block (4H.2), addeddocs/teach-exercises/README.mdmatching theteach-labs/README.mdpattern (4H.3), and added adocs/ROADMAP.mdtable of contents plus a phase-ordering note (4H.4).
- 4B.1 Expand existing
initwith composable setup/runner/ci flags while preserving current behavior. - 4B.2 Generate checked-in local-agent setup playbook output with verify-before-merge command block.
- 4B.3 Add
importcommand adapters (Promptfoo, DeepEval, AgentEvals first) reusing adapter helper normalization. - 4B.6 Add optional portable trace-reference fields and reporter links (additive schema extension).
- 4B.5 Add guardrail-focused report profile aligned to industry-audit safety taxonomy.
- 4B.4 Add optional statistical gating mode after stable identity/sample-size prerequisites.
- 4B.7 Add human adjudication export/import package flow.
- 4B.8 Add cost-quality frontier and benchmark-pack templates.
- 4C.1 Stand up static product docs site on GitHub Pages.
- 4C.2 Publish CLI-first onboarding flow centered on
init+ local-agent setup prompts. - 4C.3 Publish interoperability guides (Promptfoo, DeepEval, OpenEvals/AgentEvals, trace stacks).
- 4C.4 Add assistant-ui reference integration case study with reproducible commands.
- 4C.5 Add docs-adoption measurement loop and friction backlog.
- 4C.6 Run docs truth-sync sweep across README/ROADMAP/STATUS/help/publishing/examples.
- 4C.7 Expand interoperability docs for supplementary eval toolchains and operations stack guidance.
- 4C.8 Add integration risk register (runtime/version drift, sidecar dependencies, schema drift, cloud coupling, synthetic overfitting).
- 4C.9 Add trace-first evidence hardening guidance and end-to-end example.
- 4C.10 Add adopt-now docs path and candidate existing-runner adoption map (
docs-site/v1/adopt-now.html,docs/adoption-map.md). - 4C.docs-stack Record and apply docs-site stack decision for current milestone (
docs/docs-site-stack-decision.md).
- 4D.1 Add schema generation + schema-drift CI guard and remove stale count/completion claims.
- 4D.2 Add prioritized execution backlog for validation hardening, CI-native outputs, metrics path, Python adoption, interop expansion, calibration, and OTel guidance.
- 4D.3 Author the 14-day window (exactly 6 items) with owner/dependency/acceptance criteria tracking.
- 4D.4 Author the 45-day window (exactly 8 items) with owner/dependency/acceptance criteria tracking.
- 4D.5 Add critical/high risk register entries with trigger signals and mitigations.
- 4D.6 Keep README/ROADMAP/STATUS/help/publishing/examples discoverability and claim consistency synchronized.
Immediate next implementation slices
- 4E.5 dataset governance hardening is complete:
lintnow warns on duplicate dataset case ids across suites, orphanedscenarioIdreferences (withoutdatasetId), and low per-category coverage for dataset-governed suites, with strictness controls via--strictand--fail-on-warning-code. - CI-native machine output slice is complete:
check --json-out,--junit-out,--sarif-out,--github-annotations-out, and--heartbeat-outemit machine-readable gate outputs, row anchors, and run-status heartbeats for CI monitoring. - 4E.1 alerting adapters are complete:
check --notifysupports Slack webhook, Teams webhook (adaptive-card payload), and email/SMTP adapters with CLI/env/config precedence plus machine-readable notification diagnostics incheck-result.json. - 4E.3 calibration preflight is complete: when
judge-calibrationis configured (or force-enabled with--calibration-preflight/gates.calibration.enabled: true),checkenforces independent calibration evidence for blocking judge-scored suites, treats missingjudgeModelvalues as blocking failures, warns in report-only mode, matches against the configured calibration-suite rubric contract, validates configurable recency (--calibration-max-age-hours), and documents explicit override paths (--allow-stale-calibration,--calibration-preflight,--no-calibration-preflight). - Breaking behavior changes:
- blocking suites no longer self-certify against current-run calibration rows; teams that previously relied on same-run evidence must supply an in-window independent calibration report in the input set.
- calibration preflight now fails fast with exit code
2when the configured calibration suite metadata omitsrubricVersionin both suite manifests and rubric contracts. - dataset-governance lint checks now add warning volume (
duplicate-dataset-case-id,orphan-scenario-reference,low-category-coverage), which can turn existinglint --strictand warning-budgetedcheckpolicies red until fixtures are adjusted or warning budgets are updated.
- Phase 4F (evidence, confidentiality, and org rollout) is in progress. 4F.1 (two-tier artifact split), 4F.2 (redaction profile and publish preflight), 4F.3 (
eval-check-result/v2provenance) and 4F.4 (artifact digest + detached signature viasign/verify) shipped previously. 4F.5 (waiver and exception register) shipped:eval-dashboards check --waiver-file=<path>(orwaiverFilein the config file) reads aneval-waiver-register/v1JSON file, treats matched/non-expired waivers' rows as passed for gating and reports them prominently asACTIVE WAIVER ...diagnostics pluswaivers.active[]in--json-out/--json-v2-out, and always fails the gate for an expired waiver even if the row still fails (src/gates/waivers.ts,src/cli/index.ts,test/waivers.test.ts,test/cli-check-waivers.test.ts,docs/gates.md). 4F.6 (threshold-change detection) shipped:eval-dashboards check --baseline-gate-config=<path>(orbaselineGateConfigFilein the config file) compares the resolved gate config for this run against a recorded baselineGateConfigJSON file and fails the gate on any unapproved loosening ofminPassRate,maxNewFailures,zeroCritical,requiredPassingSuites,failOnWarningCodes,maxWarningsByCode, statistical, and calibration thresholds, with an explicit--allow-gate-loosening(orgates.allowLoosening) escape hatch; every changed field is surfaced incheck --json-out/--json-v2-outasthresholdChanges[]regardless of whether the loosening was approved (src/gates/threshold-change.ts,src/cli/index.ts,test/threshold-change.test.ts,test/cli-check-threshold-change.test.ts,docs/gates.md). 4F.7 (heartbeat verifier) shipped:eval-dashboards heartbeat-verify --heartbeat=<path> --max-age-hours=<n>is a scheduled, pipeline-independent check that fails closed (exit 1) when a release'seval-check-heartbeat/v1file is missing, malformed, stale beyond the window, or reportsgateRunStatusofskipped/errored— turning a deleted/skipped gate step into an alert instead of silent success (src/gates/heartbeat.ts,src/cli/index.ts,test/heartbeat-verify.test.ts,docs/gates.md). 4F.8 (static org rollup index) shipped:eval-dashboards org-rollup --input=<dir> --out=<path>recursively finds already-published per-repohistory.jsonfiles (the same filehistory/report --reporter=htmlwrite) under a local directory and renders one static, offline HTML page ranking repos by regression status (pass-rate drop, new critical failures, new failures), with pass-rate delta, critical/new/persistent failure counts, run count, and last-run date per repo — no ingestion API, no auth, no server, and no cross-repo network calls; it only reads local files a human or existing CI publish step copies into place (src/history/org-rollup.ts,src/reporters/org-rollup.ts,src/cli/index.ts,test/org-rollup.test.ts,test/cli-org-rollup.test.ts). 4F.9 (bypass accounting) shipped:check --bypass-log=<path>andpublish --bypass-log=<path>(orbypassLogFilein the config file) append oneeval-bypass-log-entry/v1JSON-lines record per invocation naming which of--allow-blocked-baseline/--allow-stale-calibration/--allow-gate-loosening/--allow-sensitive-publishwere used;check --json-out/--json-v2-outalways includes abypassUsage: { flags, used, count }field so a clean run is a verifiablecount: 0rather than an absent field;history --bypass-log=<path>joins the log by run id into each history entry'sbypassUsage;org-rollupsurfaces a per-repobypassCountcolumn and an org-widetotalBypassCountsummary card, turning gate erosion into a visible trend instead of an audit discovery (src/gates/bypass-accounting.ts,src/history/history.ts,src/history/org-rollup.ts,src/reporters/org-rollup.ts,src/cli/index.ts,test/bypass-accounting.test.ts,test/cli-bypass-accounting.test.ts). 4F.10 (PR-subset vs full-suite tiering with cost budget) shipped:suiteManifests[].tier(pr|full|both, defaultboth) tags suites for fast PR gating vs a scheduled full run;check --tier=pr|fullfilters the report's suites/rows to that tier before all other gates run, and--max-pr-cost-usd/--max-pr-duration-msfail the gate when the tier's summedmetadata.costUsd/durationMsexceed an explicit budget;check --json-out/--json-v2-outalways includes aprTiercost/runtime summary when--tieris passed (src/gates/pr-tiering.ts,src/model/eval-report-v1.ts,src/cli/index.ts,test/pr-tiering.test.ts,test/cli-pr-tiering.test.ts). Remaining: 4F.11 (evidence export bundle). Seedocs/ROADMAP.mdPhase 4F. - Roadmap claim audit, 2026-09-14: all 140 completed roadmap items were re-verified against code. Five were downgraded to not-done (external reference-integration dev dependency, host-app report surfacing, example-alignment process claim, docs consistency checklist, interop adapter expansion for Ragas/Langfuse/Phoenix/Braintrust/OpenAI) and twelve were reworded to match what actually shipped. Every corrected item carries an inline
Audit 2026-09-14note indocs/ROADMAP.mdexplaining the gap. Two corrections matter for adoption claims: coded import adapters exist only for promptfoo, deepeval and agentevals/openevals, and npm publishing is manualworkflow_dispatchrather than tag-triggered. - The email notification channel depends on
nodemailer, which is anoptionalDependency.--notify=emailfails at runtime in installs that skip optional dependencies; document this alongside the notify flags. - Docs-navigability review, 2026-09-15:
docs/has grown to 94 files (29 top-level) with no index, no audience-based grouping, and two competing "canonical" surfaces (docs-site/v1/*.html, 11 curated pages, vs. rawdocs/*.md). Added Phase 4H (docs/ROADMAP.md) to add adocs/README.mdindex, regroup README's flat Documentation Index, add adocs/teach-exercises/README.mdmatching the existingteach-labs/README.mdpattern, add a ROADMAP table of contents, and generalize the ad hoc stale-embedded-help-text check into an automated sweep across all docs. - 4F.11 (evidence export bundle) shipped:
eval-dashboards evidence-export --report=<path> --check-result=<path> [--waiver-file=<path>] [--bypass-log=<path>] [--approval-trail=<path>] [--signature=<path>] [--run-id=<id>] [--baseline-run-id=<id>] --out=<path>reads each referenced file, hashes it (sha256), and embeds its exact bytes verbatim into a singleeval-evidence-bundle/v1JSON file alongside a top-level digest over all per-entry digests;eval-dashboards evidence-verify --bundle=<path>independently recomputes both and fails closed (exit 1) on any mismatch, needing nothing but the bundle file itself (src/evidence/bundle.ts,src/cli/index.ts,test/cli-evidence-export.test.ts,docs/cli-help/evidence-export.txt,docs/cli-help/evidence-verify.txt). This closes out the original 4F.1–4F.11 slice. Since then, four more items were added to Phase 4F from continuous competitive scanning and shipped: 4F.12 (CI dependency-audit gate —pnpm audit --audit-level=highas a blocking CI step,.github/workflows/ci.yml,docs/gates.md), 4F.13 (client-side compare for the static HTML report — offline file-picker + inline diff script, no network calls,src/reporters/render.ts,test/report-compare-client.test.ts), 4F.14 (compliance-framework tagging — optional free-formrows[].complianceRefs?: string[]andsuiteManifests[].complianceFrameworks?: string[], additive/no canonical enum,src/model/eval-report-v1.ts,docs/taxonomy.md), and 4F.15 (free-form run-leveltags?: Record<string,string>on the top-level run,src/model/eval-report-v1.ts,test/render.test.ts). Phase 4F (4F.1–4F.15) was done at that point; later scout items continued the phase through 4F.24 (structured per-row assertion evidence,checks[]). Seedocs/ROADMAP.mdfor the per-item evidence.
Phase 4I (competitive parity on CI/PR ergonomics) also shipped: 4I.1 (GitHub PR-comment publish target — sticky create-then-update-in-place comment via a hidden HTML marker, dry-run by default outside CI, src/publish/publish.ts, test/publish.test.ts) and 4I.2 (a versioned composite GitHub Action wrapping report+check+optional PR-comment publish, action.yml). Phase 4I (4I.1–4I.2) is done.
Phase 4J (OpenTelemetry GenAI evaluation-span mapping): 4J.1 shipped. eval-dashboards import --from=otel-genai reads an OTLP/JSON export (span events or log-record events; single object or collector JSONL) and maps each gen_ai.evaluation.result event to an llm-judge row with category/judgeCategory, judgeVerdict, score, reason/judgeReasoning and trace.traceId/trace.spanId; error.type becomes a failed row (src/cli/import-adapters.ts, test/import-adapters-otel-genai.test.ts, docs/integrations/otel-genai.md). --case-id-attribute=<key> gives stable <caseId>:<metric> row ids across runs so baselines report persistent failures (default <spanId>:<metric> ids change per run). Known gaps: gen_ai.response.id correlation is not mapped; the numeric-score fallback assumes higher-is-better on a 0–1 scale.
External Phase 4: Shipping & Adoption
- Announce on Reddit, HN, AI communities, eval-focused newsletters
- Expand real-world runner partnerships and external integration examples
- Open scoped "good first issues" for contributors
- Continue the feedback loop with early adopters
Phase 5+: Long-term (post-v1.0)
- Plugin system for custom reporters
- Optional local web-server mode for interactive exploration
- Richer risk-area and tool-routing views
- Cost/token/latency aggregation
- Diff views between any two runs
- AI-powered suggestions for suite manifests and rubric versions