Skip to content

Latest commit

 

History

History
186 lines (156 loc) · 24.5 KB

File metadata and controls

186 lines (156 loc) · 24.5 KB

Implementation Status

Update this file as features are implemented. Keep it honest: mark an item done only after the relevant code and tests exist.

See also: ROADMAP.md for the prioritized improvement plan.

Done

  • Create independent project directory.

  • Add npm package metadata for @icodenet/eval-dashboards.

  • Add TypeScript, tsup, and Vitest project skeleton.

  • Add CLI binary entry point (eval-dashboards).

  • Add PRP documentation.

  • Add versioned eval-report/v1 model and validation.

  • Add first-class optional agent and LLM judge report fields.

  • Add portable suite manifest, gate policy, and rubric contract fields.

  • Add baseline compatibility assessment for dataset and rubric version drift.

  • Add basic report discovery and history building.

  • Add latest-vs-previous comparison.

  • Add initial gate checking.

  • Add starter text, json-summary, markdown-summary, and html reporters.

  • Add local directory publish target.

  • Implement publishing adapters for GitHub Pages and Azure Storage, plus Azure Static Web Apps dry-run validation mode.

  • Add required example directories and starter artifacts.

  • Add runnable provider-free agent/chat eval example that generates artifacts and reports.

  • Add focused starter tests.

  • Add focused governance and baseline compatibility tests.

  • Implement eval-dashboards init config generation.

  • Implement full config loading from eval-dashboards.config.ts, eval-dashboards.config.js, and package.json.

  • Align documentation: sweep and replace old @icodenet/eval-reports / eval-reports naming with @icodenet/eval-dashboards / eval-dashboards.

  • Add GitHub Actions CI workflow (typecheck, test, build, example smoke tests).

  • Export JSON Schema for eval-report/v1 to schemas/eval-report-v1.schema.json (published and maintained in-repo).

  • Create comprehensive taxonomy teaching documentation at docs/taxonomy.md (definitions, examples, checklist, FAQ).

  • Create taxonomy-complete init fixture at examples/taxonomy-complete-fixture/run-complete.json with README demonstrating best practices.

  • Create runner cookbook: Vitest example with README and patterns.

  • Create runner cookbook: Jest custom reporter example with README and patterns.

  • Create runner cookbook: Node plain eval example with README and use cases.

  • Refresh README for adoption messaging, schema-first positioning, and cookbook links.

  • Implement HTML grouping by dataset and scenario with collapsible sections.

  • Add taxonomy completeness score (0–100%) to row display with visual indicators.

  • Add kind badges (deterministic, agent, llm-judge, human-review) to row display.

  • Add "All rows (by dataset & scenario)" section with full grouping.

  • Implement Azure Static Web Apps dry-run validation path (non-dry-run execution path pending).

  • Implement Azure Storage static website publishing adapter (including dry-run validation path).

  • Add persistent failure detection with analyzeRowStability() function.

  • Add flaky row classification based on pass/fail history across runs.

  • Create CONTRIBUTING.md with development workflow, project structure, and commit guidelines.

  • Create CODE_OF_CONDUCT.md (Contributor Covenant-based).

  • Create GitHub issue templates (bug report, feature request).

  • Create CHANGELOG.md with semantic versioning guidance.

  • Implement explicit baseline selection by run id (enhancement to compareRuns).

  • Enhance HTML dashboard with sparklines and pass-rate trends in history view.

  • Create README screenshot gallery (light and dark themes) — visual proof of UI.

  • Add concrete report-power artifact fixture with tracked history/progress/gate/detail outputs and deterministic regeneration script.

  • Add teach delivery-stage labs plus FDE role workflow guidance grounded in report artifacts and required evidence outputs.

  • Add npm publishing workflow and semantic version tagging (GitHub Actions).

    • Audit 2026-09-15: publish.yml/release.yml are workflow_dispatch-only — there is no push: tags trigger, so tag-triggered publishing is not implemented yet; manual dispatch works.

Release Notes

  • The branch history was rewritten to reflect the current codebase.
  • Only the latest post-rewrite release should be treated as the valid reference for the current implementation.
  • Earlier release artifacts are superseded and should not be used to evaluate the present code state.

Ongoing (live KPIs)

  • Create community feedback loop infrastructure and early-runner outreach tracker (Phase 4 readiness).
    • Added weekly metrics loop via pnpm metrics:adoption and snapshot output in docs/adoption-metrics/latest.json.
    • Added manual signal tracker at docs/adoption-metrics/manual-signals.json.
    • Added partnership log and outreach stages in docs/community-partnership-log.md.
    • External adoption outcomes continue as live KPIs, not static checklist items.

Planned / In Assessment

Parallel reference integration + eval-dashboards workstreams

  • reference integration: inspect existing eval runner, dataset shape, scoring, and CI workflow.

  • reference integration: add @icodenet/eval-dashboards@0.3.0 as an explicit dev dependency.

    • Audit 2026-09-15: unverifiable from this repo — no external commit/PR link recorded; ROADMAP.md tracks this item as unchecked/blocked for the same reason. Corrected here to match.
  • reference integration: map current eval output into eval-report/v1 without replacing the existing runner.

  • reference integration: emit .evals_output/*.json artifacts with suite summaries and row-level evidence.

  • reference integration: add suite manifests, dataset versions, rubric versions, and dashboard gates.

  • reference integration: wire eval-dashboards lint, check, and report into local/CI eval commands.

  • reference integration: add rubric contracts plus row provenance and lifecycle metadata.

  • reference integration: surface the generated /eval-dashboard/ report in the learning UI instead of the old bespoke summary dashboard.

    • Audit 2026-09-15: contradicted by docs/case-studies/assistant-ui/README.md — adapter and CLI wiring exist in a local worktree only, no PR opened; ROADMAP.md tracks this item as unchecked/blocked for the same reason. Corrected here to match.
  • reference integration: create first published dashboard baseline and document quality gaps.

  • eval-dashboards: publish TypeScript declaration files and package metadata so downstream imports resolve public types.

    • Audit 2026-09-15: dist/*.d.ts ship correctly, but the packed-package consumer smoke test named in this slice is still missing — the downstream compile was a one-off manual check, not a checked-in test (matches ROADMAP.md's caveat on the same item).
  • eval-dashboards: define agent-quality suite presets (retrieval-recall, answer-groundedness, answer-quality, refusal-safety, prompt-injection-resilience, mcp-routing, content-coverage, regression-incidents, judge-calibration).

  • eval-dashboards: decide which setup concepts belong in schema fields/enums, preset files, examples, or docs.

  • eval-dashboards: design setup scaffolding for common agent eval programs, such as init --preset agent-quality.

  • eval-dashboards: add starter dataset/rubric templates with versioning, provenance, lifecycle, and judge calibration examples.

  • eval-dashboards: document how presets map to riskArea, target, graders, gate policies, and rubric contracts.

  • eval-dashboards: add a repo-context glossary explaining eval-report/v1, suite, dataset, rubric, runner, and row terminology.

  • eval-dashboards: add runner-adapter primitives so teams with an existing eval runner can map local results into eval-report/v1 without hand-writing aggregate, manifest, rubric, and output-cleanup boilerplate.

  • Cross-feed: use reference integration learnings to amend eval-dashboards roadmap, templates, and docs before stabilizing setup-layer APIs.

    • Captured so far: prefer directory inputs over config globs, require rubric versions for blocking suites, clean generated artifact directories before writing, make suite summaries row-complete, and expose/embed the generated static dashboard instead of duplicating it with host-app summary cards.
    • Type packaging captured: emit declarations and expose them with main, types, and exports; verified with pnpm build, npm pack, and a temporary downstream TypeScript compile against the packed tarball.
    • Adapter boundary captured: keep project-specific dataset rows local, but move repeated artifact assembly mechanics into public eval-dashboards helpers.
    • Approval-gate pattern captured: document eval-results branch layout, pr-meta.json wiring, commit-status contract (eval/quality-gate), environment approval flow, and cleanup workflow templates for closed PRs.
    • Planning slices captured in ROADMAP: dataset governance, versioned rubrics, judge calibration, CI quality tiers, suite templates, setup scaffolding, and schema/taxonomy decision rules.
  • Research and publish industry coverage audit for suites/datasets/rubrics.

    • Added docs/industry-coverage-audit.md with external-source mapping and local coverage matrix.
    • Identified P0 additions: goal-success, intent-resolution, task-adherence, sensitive-disclosure, and agency-boundary presets.
  • Implement P0 industry coverage suites in presets, dataset templates, rubrics, artifact template, and init scaffold.

    • Scope tracked in docs/industry-coverage-audit.md.
    • Must include parity updates across docs + examples + src/cli/init-scaffold.ts.
    • Added starter multi-turn trajectory coverage (multiturn-trajectory) with preset guidance, dataset case, rubric axes, template artifact row, and scaffold output.
  • Phase 4H docs navigability: added docs/README.md audience index (4H.1), regrouped README's Documentation Index into the same audience groups with a Start-here block (4H.2), added docs/teach-exercises/README.md matching the teach-labs/README.md pattern (4H.3), and added a docs/ROADMAP.md table of contents plus a phase-ordering note (4H.4).

Next Phases

Phase 4B setup automation checklist (complete)

  • 4B.1 Expand existing init with composable setup/runner/ci flags while preserving current behavior.
  • 4B.2 Generate checked-in local-agent setup playbook output with verify-before-merge command block.
  • 4B.3 Add import command adapters (Promptfoo, DeepEval, AgentEvals first) reusing adapter helper normalization.
  • 4B.6 Add optional portable trace-reference fields and reporter links (additive schema extension).
  • 4B.5 Add guardrail-focused report profile aligned to industry-audit safety taxonomy.
  • 4B.4 Add optional statistical gating mode after stable identity/sample-size prerequisites.
  • 4B.7 Add human adjudication export/import package flow.
  • 4B.8 Add cost-quality frontier and benchmark-pack templates.

Phase 4C docs-site adoption checklist (complete)

  • 4C.1 Stand up static product docs site on GitHub Pages.
  • 4C.2 Publish CLI-first onboarding flow centered on init + local-agent setup prompts.
  • 4C.3 Publish interoperability guides (Promptfoo, DeepEval, OpenEvals/AgentEvals, trace stacks).
  • 4C.4 Add assistant-ui reference integration case study with reproducible commands.
  • 4C.5 Add docs-adoption measurement loop and friction backlog.
  • 4C.6 Run docs truth-sync sweep across README/ROADMAP/STATUS/help/publishing/examples.
  • 4C.7 Expand interoperability docs for supplementary eval toolchains and operations stack guidance.
  • 4C.8 Add integration risk register (runtime/version drift, sidecar dependencies, schema drift, cloud coupling, synthetic overfitting).
  • 4C.9 Add trace-first evidence hardening guidance and end-to-end example.
  • 4C.10 Add adopt-now docs path and candidate existing-runner adoption map (docs-site/v1/adopt-now.html, docs/adoption-map.md).
  • 4C.docs-stack Record and apply docs-site stack decision for current milestone (docs/docs-site-stack-decision.md).

Phase 4D trusted confidence + adoption execution checklist (planned)

  • 4D.1 Add schema generation + schema-drift CI guard and remove stale count/completion claims.
  • 4D.2 Add prioritized execution backlog for validation hardening, CI-native outputs, metrics path, Python adoption, interop expansion, calibration, and OTel guidance.
  • 4D.3 Author the 14-day window (exactly 6 items) with owner/dependency/acceptance criteria tracking.
  • 4D.4 Author the 45-day window (exactly 8 items) with owner/dependency/acceptance criteria tracking.
  • 4D.5 Add critical/high risk register entries with trigger signals and mitigations.
  • 4D.6 Keep README/ROADMAP/STATUS/help/publishing/examples discoverability and claim consistency synchronized.

Immediate next implementation slices

  • 4E.5 dataset governance hardening is complete: lint now warns on duplicate dataset case ids across suites, orphaned scenarioId references (without datasetId), and low per-category coverage for dataset-governed suites, with strictness controls via --strict and --fail-on-warning-code.
  • CI-native machine output slice is complete: check --json-out, --junit-out, --sarif-out, --github-annotations-out, and --heartbeat-out emit machine-readable gate outputs, row anchors, and run-status heartbeats for CI monitoring.
  • 4E.1 alerting adapters are complete: check --notify supports Slack webhook, Teams webhook (adaptive-card payload), and email/SMTP adapters with CLI/env/config precedence plus machine-readable notification diagnostics in check-result.json.
  • 4E.3 calibration preflight is complete: when judge-calibration is configured (or force-enabled with --calibration-preflight / gates.calibration.enabled: true), check enforces independent calibration evidence for blocking judge-scored suites, treats missing judgeModel values as blocking failures, warns in report-only mode, matches against the configured calibration-suite rubric contract, validates configurable recency (--calibration-max-age-hours), and documents explicit override paths (--allow-stale-calibration, --calibration-preflight, --no-calibration-preflight).
  • Breaking behavior changes:
    • blocking suites no longer self-certify against current-run calibration rows; teams that previously relied on same-run evidence must supply an in-window independent calibration report in the input set.
    • calibration preflight now fails fast with exit code 2 when the configured calibration suite metadata omits rubricVersion in both suite manifests and rubric contracts.
    • dataset-governance lint checks now add warning volume (duplicate-dataset-case-id, orphan-scenario-reference, low-category-coverage), which can turn existing lint --strict and warning-budgeted check policies red until fixtures are adjusted or warning budgets are updated.
  • Phase 4F (evidence, confidentiality, and org rollout) is in progress. 4F.1 (two-tier artifact split), 4F.2 (redaction profile and publish preflight), 4F.3 (eval-check-result/v2 provenance) and 4F.4 (artifact digest + detached signature via sign/verify) shipped previously. 4F.5 (waiver and exception register) shipped: eval-dashboards check --waiver-file=<path> (or waiverFile in the config file) reads an eval-waiver-register/v1 JSON file, treats matched/non-expired waivers' rows as passed for gating and reports them prominently as ACTIVE WAIVER ... diagnostics plus waivers.active[] in --json-out/--json-v2-out, and always fails the gate for an expired waiver even if the row still fails (src/gates/waivers.ts, src/cli/index.ts, test/waivers.test.ts, test/cli-check-waivers.test.ts, docs/gates.md). 4F.6 (threshold-change detection) shipped: eval-dashboards check --baseline-gate-config=<path> (or baselineGateConfigFile in the config file) compares the resolved gate config for this run against a recorded baseline GateConfig JSON file and fails the gate on any unapproved loosening of minPassRate, maxNewFailures, zeroCritical, requiredPassingSuites, failOnWarningCodes, maxWarningsByCode, statistical, and calibration thresholds, with an explicit --allow-gate-loosening (or gates.allowLoosening) escape hatch; every changed field is surfaced in check --json-out/--json-v2-out as thresholdChanges[] regardless of whether the loosening was approved (src/gates/threshold-change.ts, src/cli/index.ts, test/threshold-change.test.ts, test/cli-check-threshold-change.test.ts, docs/gates.md). 4F.7 (heartbeat verifier) shipped: eval-dashboards heartbeat-verify --heartbeat=<path> --max-age-hours=<n> is a scheduled, pipeline-independent check that fails closed (exit 1) when a release's eval-check-heartbeat/v1 file is missing, malformed, stale beyond the window, or reports gateRunStatus of skipped/errored — turning a deleted/skipped gate step into an alert instead of silent success (src/gates/heartbeat.ts, src/cli/index.ts, test/heartbeat-verify.test.ts, docs/gates.md). 4F.8 (static org rollup index) shipped: eval-dashboards org-rollup --input=<dir> --out=<path> recursively finds already-published per-repo history.json files (the same file history/report --reporter=html write) under a local directory and renders one static, offline HTML page ranking repos by regression status (pass-rate drop, new critical failures, new failures), with pass-rate delta, critical/new/persistent failure counts, run count, and last-run date per repo — no ingestion API, no auth, no server, and no cross-repo network calls; it only reads local files a human or existing CI publish step copies into place (src/history/org-rollup.ts, src/reporters/org-rollup.ts, src/cli/index.ts, test/org-rollup.test.ts, test/cli-org-rollup.test.ts). 4F.9 (bypass accounting) shipped: check --bypass-log=<path> and publish --bypass-log=<path> (or bypassLogFile in the config file) append one eval-bypass-log-entry/v1 JSON-lines record per invocation naming which of --allow-blocked-baseline/--allow-stale-calibration/--allow-gate-loosening/--allow-sensitive-publish were used; check --json-out/--json-v2-out always includes a bypassUsage: { flags, used, count } field so a clean run is a verifiable count: 0 rather than an absent field; history --bypass-log=<path> joins the log by run id into each history entry's bypassUsage; org-rollup surfaces a per-repo bypassCount column and an org-wide totalBypassCount summary card, turning gate erosion into a visible trend instead of an audit discovery (src/gates/bypass-accounting.ts, src/history/history.ts, src/history/org-rollup.ts, src/reporters/org-rollup.ts, src/cli/index.ts, test/bypass-accounting.test.ts, test/cli-bypass-accounting.test.ts). 4F.10 (PR-subset vs full-suite tiering with cost budget) shipped: suiteManifests[].tier (pr|full|both, default both) tags suites for fast PR gating vs a scheduled full run; check --tier=pr|full filters the report's suites/rows to that tier before all other gates run, and --max-pr-cost-usd/--max-pr-duration-ms fail the gate when the tier's summed metadata.costUsd/durationMs exceed an explicit budget; check --json-out/--json-v2-out always includes a prTier cost/runtime summary when --tier is passed (src/gates/pr-tiering.ts, src/model/eval-report-v1.ts, src/cli/index.ts, test/pr-tiering.test.ts, test/cli-pr-tiering.test.ts). Remaining: 4F.11 (evidence export bundle). See docs/ROADMAP.md Phase 4F.
  • Roadmap claim audit, 2026-09-14: all 140 completed roadmap items were re-verified against code. Five were downgraded to not-done (external reference-integration dev dependency, host-app report surfacing, example-alignment process claim, docs consistency checklist, interop adapter expansion for Ragas/Langfuse/Phoenix/Braintrust/OpenAI) and twelve were reworded to match what actually shipped. Every corrected item carries an inline Audit 2026-09-14 note in docs/ROADMAP.md explaining the gap. Two corrections matter for adoption claims: coded import adapters exist only for promptfoo, deepeval and agentevals/openevals, and npm publishing is manual workflow_dispatch rather than tag-triggered.
  • The email notification channel depends on nodemailer, which is an optionalDependency. --notify=email fails at runtime in installs that skip optional dependencies; document this alongside the notify flags.
  • Docs-navigability review, 2026-09-15: docs/ has grown to 94 files (29 top-level) with no index, no audience-based grouping, and two competing "canonical" surfaces (docs-site/v1/*.html, 11 curated pages, vs. raw docs/*.md). Added Phase 4H (docs/ROADMAP.md) to add a docs/README.md index, regroup README's flat Documentation Index, add a docs/teach-exercises/README.md matching the existing teach-labs/README.md pattern, add a ROADMAP table of contents, and generalize the ad hoc stale-embedded-help-text check into an automated sweep across all docs.
  • 4F.11 (evidence export bundle) shipped: eval-dashboards evidence-export --report=<path> --check-result=<path> [--waiver-file=<path>] [--bypass-log=<path>] [--approval-trail=<path>] [--signature=<path>] [--run-id=<id>] [--baseline-run-id=<id>] --out=<path> reads each referenced file, hashes it (sha256), and embeds its exact bytes verbatim into a single eval-evidence-bundle/v1 JSON file alongside a top-level digest over all per-entry digests; eval-dashboards evidence-verify --bundle=<path> independently recomputes both and fails closed (exit 1) on any mismatch, needing nothing but the bundle file itself (src/evidence/bundle.ts, src/cli/index.ts, test/cli-evidence-export.test.ts, docs/cli-help/evidence-export.txt, docs/cli-help/evidence-verify.txt). This closes out the original 4F.1–4F.11 slice. Since then, four more items were added to Phase 4F from continuous competitive scanning and shipped: 4F.12 (CI dependency-audit gate — pnpm audit --audit-level=high as a blocking CI step, .github/workflows/ci.yml, docs/gates.md), 4F.13 (client-side compare for the static HTML report — offline file-picker + inline diff script, no network calls, src/reporters/render.ts, test/report-compare-client.test.ts), 4F.14 (compliance-framework tagging — optional free-form rows[].complianceRefs?: string[] and suiteManifests[].complianceFrameworks?: string[], additive/no canonical enum, src/model/eval-report-v1.ts, docs/taxonomy.md), and 4F.15 (free-form run-level tags?: Record<string,string> on the top-level run, src/model/eval-report-v1.ts, test/render.test.ts). Phase 4F (4F.1–4F.15) was done at that point; later scout items continued the phase through 4F.24 (structured per-row assertion evidence, checks[]). See docs/ROADMAP.md for the per-item evidence.

Phase 4I (competitive parity on CI/PR ergonomics) also shipped: 4I.1 (GitHub PR-comment publish target — sticky create-then-update-in-place comment via a hidden HTML marker, dry-run by default outside CI, src/publish/publish.ts, test/publish.test.ts) and 4I.2 (a versioned composite GitHub Action wrapping report+check+optional PR-comment publish, action.yml). Phase 4I (4I.1–4I.2) is done.

Phase 4J (OpenTelemetry GenAI evaluation-span mapping): 4J.1 shipped. eval-dashboards import --from=otel-genai reads an OTLP/JSON export (span events or log-record events; single object or collector JSONL) and maps each gen_ai.evaluation.result event to an llm-judge row with category/judgeCategory, judgeVerdict, score, reason/judgeReasoning and trace.traceId/trace.spanId; error.type becomes a failed row (src/cli/import-adapters.ts, test/import-adapters-otel-genai.test.ts, docs/integrations/otel-genai.md). --case-id-attribute=<key> gives stable <caseId>:<metric> row ids across runs so baselines report persistent failures (default <spanId>:<metric> ids change per run). Known gaps: gen_ai.response.id correlation is not mapped; the numeric-score fallback assumes higher-is-better on a 0–1 scale.

External Phase 4: Shipping & Adoption

  • Announce on Reddit, HN, AI communities, eval-focused newsletters
  • Expand real-world runner partnerships and external integration examples
  • Open scoped "good first issues" for contributors
  • Continue the feedback loop with early adopters

Phase 5+: Long-term (post-v1.0)

  • Plugin system for custom reporters
  • Optional local web-server mode for interactive exploration
  • Richer risk-area and tool-routing views
  • Cost/token/latency aggregation
  • Diff views between any two runs
  • AI-powered suggestions for suite manifests and rubric versions