- Python-native evaluator ecosystem for LLM quality checks.
- Broad evaluator set for correctness, relevance, safety, and custom metrics.
- Natural fit for teams already in pytest/Python pipelines.
- Output schemas vary across versions and custom evaluators.
- Python dependency footprint can diverge from Node-only CI environments.
Use the built-in import adapter.
- Required: DeepEval JSON export (
test_resultsorresults). - Output: normalized
eval-report/v1rows with suite/pass metadata.
eval-dashboards import --from=deepeval --input=./deepeval-results.json --out=.evals_output/import-deepeval.json
eval-dashboards lint --input=.evals_output
eval-dashboards report --input=.evals_output --reporter=html --report-dir=eval-reportRelated risks: integration risk register