Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 2 additions & 1 deletion AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -60,7 +60,8 @@ Detailed operating rules — commands, repo map, hard rules — live in
uv venv --python 3.11 && source .venv/bin/activate
uv pip install -e ".[dev]"
pytest # full suite
pre-commit run --all-files # exactly what CI runs
make ci # exactly what CI runs (also the pre-push hook)
pre-commit run --all-files # commit-stage hooks only (file hygiene, ruff --fix, format, mypy)
```

- Config models use `extra="forbid"`; any new `evalshift.yaml` field needs docs
Expand Down
28 changes: 28 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -83,6 +83,34 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
whole source tree. The floor sits below the measured 94% on purpose, so a
real regression fails CI while an honest refactor does not.

- An audit of the docs against the code found them wrong in a few dozen
places, and all of them now say what the CLI does. Nothing about CLI
behaviour changed. The one that matters most concerns slices: the docs
said a top-level `slices:` block picks examples by tag, names the slice,
and scopes it to prompts. It does none of that. The block is validated
(the reserved `overall` name is still rejected) and recorded in the run
bundle, but analysis never reads it. Slices come from example `tags`, one
per distinct tag plus `all`, and per-slice budgets are keyed by tag under
`migration_policy.slices`. The docs now say the block is not applied, and
so does the `SliceConfig` docstring, which described `filter` as a Python
expression. Imported agent traces (`traces import`) stay local; the bundle
carries only the replay's own tool-call trace, without tool results or
`model_call` events. The response cache serves only tool-less examples,
so every `run` of an agent suite is live and full price. `--resume`
hashes the suite's path, not its contents. `push <run-id>` uploads an
existing bundle as-is instead of rebuilding it, and `bundle` needs a git
SHA. `--policy-gate` also fails when no `migration_policy` is configured.
The `init` profile table had the wrong `model-upgrade` numbers and no
tool-divergence column. The multi-turn suite example failed to load
because it had no `tools`. The failure-label list was missing
`TOOL_GROUND_TRUTH_MISS`. Upstream model-call failures and evaluator
failures are handled the same way, as errored rows excluded from the
statistics. The GitHub Action docs gained `require-policy` and the other
missing inputs. `record_model_call` examples now pass the required
`tools=`. DOCS.md's header said version 1.0.1; a new check in
`tests/unit/test_docs_currency.py` keeps the version in DOCS.md and
llms-full.txt equal to the package's.

## [1.1.0] - 2026-09-19

Shipped as a minor deliberately. The `thresholds` removal below is breaking by
Expand Down
3 changes: 2 additions & 1 deletion CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -35,7 +35,8 @@ pytest -m "not integration" # unit tests only
ruff check . # lint
ruff format . # auto-format
mypy --strict src/evalshift_cli # type-check
pre-commit run --all-files # everything pre-commit runs
pre-commit run --all-files # commit-stage hooks only
make ci # exactly what CI runs (also the pre-push hook)
```

## Style
Expand Down
97 changes: 51 additions & 46 deletions DOCS.md

Large diffs are not rendered by default.

7 changes: 4 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -226,9 +226,10 @@ The short version:
statistics, the analysis, the migration decision, economics, and the
machine-written insights narrative.
* **Never uploads**: provider API keys, prompt bodies and system prompts,
suite conversation histories, tool definitions/schemas, `raw.jsonl`, the
response cache, captures, and `report.html`. Prompt and dataset content is
replaced by SHA-256 hashes so diffs still align across runs.
suite conversation histories, tool definitions/schemas, `raw.jsonl`,
imported agent traces, the response cache, captures, and `report.html`.
Prompt and dataset content is replaced by SHA-256 hashes so diffs still
align across runs.
* **Can still be sensitive**: inputs, expected outputs, model outputs, and
traces carry whatever content your suite or your models put in them. Redact
at capture time (see the SDK's redaction boundary) and inspect before
Expand Down
12 changes: 7 additions & 5 deletions docs/agents.md
Original file line number Diff line number Diff line change
Expand Up @@ -46,7 +46,7 @@ row carrying its own `toolset_ref`:
cd examples/agent
export GOOGLE_API_KEY=<google-api-key>
evalshift run --yes --from gemini-2.5-flash --to gemini-3.1-flash-lite-preview
RUN_ID=$(ls .evalshift/runs/ | head -1)
RUN_ID=$(ls -t .evalshift/runs/ | head -1)
evalshift evaluate "$RUN_ID"
evalshift analyze "$RUN_ID"
evalshift report "$RUN_ID" --open
Expand Down Expand Up @@ -136,7 +136,7 @@ Nothing in `evalshift.yaml` wires a toolset to a prompt — dispatch reads it
off each golden-suite *example* instead (`toolset_ref` or inline `tools`, see
[Suite ground truth](#suite-ground-truth) below), so the same prompt can
legitimately dispatch some examples with tools and others without, in one
run. `tools.yaml` above is just this project's human-readable record of what
run. [`examples/agent/tools.yaml`](https://github.com/babaliauskas/evalshift-cli/blob/main/examples/agent/tools.yaml) is just this project's human-readable record of what
those tools are; it accepts either Anthropic-shape (`name` / `description` /
`input_schema`) or OpenAI-shape (`{ "type": "function", "function": {...}
}`) entries — `evalshift run` serialises whatever a toolset resolves to in
Expand Down Expand Up @@ -507,9 +507,11 @@ expectation are ignored here. Which keys are compared follows each expectation's
`match_strategy`: `exact` also flags arguments the ground truth did not record,
while `subset` (what `capture promote` writes) scores the recorded keys only.

Check `evalshift doctor` first. If it warns about **tool argument shape**, your
recorded arguments use keys the declared schema does not have, and every
ground-truth comparison will score 0 for a reason that is not the model's fault.
If every ground-truth comparison scores 0, check whether your recorded
arguments use keys the declared schema does not have — a capture of a decorated
function's parameters rather than the model's own arguments scores 0 for a
reason that is not the model's fault. See
[Wrapper arguments are unwrapped](#wrapper-arguments-are-unwrapped-legacy-captures).

A ground-truth field that **neither** model produced is dropped from that call's
denominator on both sides and disclosed as `unmeasured_fields` in the record's
Expand Down
Loading
Loading