Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 3 additions & 3 deletions .github/workflows/release.yml
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,7 @@ name: release
# repo, this workflow file, and the `pypi` environment as a trusted
# publisher. No token secret is stored anywhere.
# - `EVALSHIFT_ACTION_DISPATCH_TOKEN` (optional): a fine-grained PAT with
# contents: write on babaliauskas/evalshift-action. Without it the
# contents: write on evalshift/evalshift-action. Without it the
# dispatch step is skipped and the action repo's daily poll picks the
# release up within a day.
#
Expand Down Expand Up @@ -84,7 +84,7 @@ jobs:
needs: [build, publish]
runs-on: ubuntu-latest
steps:
- name: Dispatch evalshift-cli-release to babaliauskas/evalshift-action
- name: Dispatch evalshift-cli-release to evalshift/evalshift-action
env:
GH_TOKEN: ${{ secrets.EVALSHIFT_ACTION_DISPATCH_TOKEN }}
VERSION: ${{ needs.build.outputs.version }}
Expand All @@ -94,7 +94,7 @@ jobs:
echo "::notice::EVALSHIFT_ACTION_DISPATCH_TOKEN is not set; skipping the dispatch. The action repo's daily PyPI poll will pick up ${VERSION} within a day."
exit 0
fi
gh api "repos/babaliauskas/evalshift-action/dispatches" \
gh api "repos/evalshift/evalshift-action/dispatches" \
-f event_type=evalshift-cli-release \
-f "client_payload[version]=${VERSION}"
echo "Dispatched evalshift-cli-release for ${VERSION}; bump-cli-pin will open the pin PR."
2 changes: 1 addition & 1 deletion AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -38,7 +38,7 @@ repos (`evalshift-sdk`, `evalshift-action`); never edit them from here.
two packages never call each other. The CLI imports as `evalshift_cli` and
depends on the SDK, so both install into one environment. See
[docs/sdk.md](docs/sdk.md).
- **GitHub Action** — `babaliauskas/evalshift-action@v0`. Runs the pipeline on
- **GitHub Action** — `evalshift/evalshift-action@v0`. Runs the pipeline on
pull requests, pushes the run, keeps one PR comment updated, sets the
`evalshift/regression` commit status. See [docs/github-action.md](docs/github-action.md).
- **Hosted server** — API at `https://api.evalshift.dev`, web app at
Expand Down
13 changes: 13 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,19 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

## [Unreleased]

### Changed

- The EvalShift repositories moved from the `babaliauskas` GitHub account to
the `evalshift` organization: <https://github.com/evalshift/evalshift-cli>,
`evalshift-sdk` and `evalshift-action`. Links, package metadata and the
workflow `evalshift init --ci` scaffolds now use
`evalshift/evalshift-action@v0`. GitHub redirects the old names, so existing
workflows keep working.
- The CI pin check (`init`, `validate`, `doctor`, `capture sync`) recognizes
`evalshift/evalshift-action` steps as well as the old
`babaliauskas/evalshift-action` name, matching owner and repository names
case-insensitively as GitHub does.

## [1.2.1] - 2026-10-01

### Changed
Expand Down
8 changes: 4 additions & 4 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,7 +10,7 @@ We use [`uv`](https://docs.astral.sh/uv/) for environment and dependency managem

```bash
# Clone and enter the repo
git clone https://github.com/babaliauskas/EvalShift.git
git clone https://github.com/evalshift/evalshift-cli.git
cd evalshift

# Create a virtualenv with Python 3.11 and install dev deps
Expand Down Expand Up @@ -72,15 +72,15 @@ A release is a commit, a tag, and nothing else — CI does the publishing.
asserts the tag matches `pyproject.toml`, builds with `uv build`, publishes
to PyPI via [trusted publishing](https://docs.pypi.org/trusted-publishers/)
(no token secret), and sends a `repository_dispatch` to
`babaliauskas/evalshift-action` so its `bump-cli-pin` workflow opens the
`evalshift/evalshift-action` so its `bump-cli-pin` workflow opens the
pin-bump PR immediately.

One-time setup, held outside the repo:

- **PyPI trusted publisher** on the `evalshift` project: repository
`babaliauskas/evalshift-cli`, workflow `release.yml`, environment `pypi`.
`evalshift/evalshift-cli`, workflow `release.yml`, environment `pypi`.
- **`EVALSHIFT_ACTION_DISPATCH_TOKEN`** (optional repo secret): fine-grained
PAT with contents: write on `babaliauskas/evalshift-action`. Without it the
PAT with contents: write on `evalshift/evalshift-action`. Without it the
dispatch is skipped and the action repo's daily PyPI poll picks the release
up within a day.

Expand Down
16 changes: 8 additions & 8 deletions DOCS.md
Original file line number Diff line number Diff line change
Expand Up @@ -58,7 +58,7 @@ uv pip install evalshift
From source:

```bash
git clone https://github.com/babaliauskas/evalshift-cli
git clone https://github.com/evalshift/evalshift-cli
cd evalshift-cli
uv venv --python 3.11
source .venv/bin/activate
Expand Down Expand Up @@ -89,8 +89,8 @@ EvalShift is four pieces. Each is released and documented independently; each ow
| Piece | Distribution | What it does | Reference for humans | Reference for AI tools |
| --- | --- | --- | --- | --- |
| **CLI** | PyPI `evalshift` (import `evalshift_cli`) | Runs the suite on two models, scores, analyses, reports, bundles, pushes. | this document | <https://www.evalshift.dev/cli-llms-full.txt> |
| **SDK** | PyPI `evalshift-sdk` (import `evalshift`) | In-process capture: records your agent's model/tool calls to `.evalshift/captures/`. | [docs/sdk.md](docs/sdk.md), [SDK repo](https://github.com/babaliauskas/evalshift-sdk) | <https://www.evalshift.dev/sdk-llms-full.txt> |
| **GitHub Action** | `babaliauskas/evalshift-action@v0` | Runs the pipeline on PRs, pushes the run, maintains one PR comment, sets the `evalshift/regression` status. | [docs/github-action.md](docs/github-action.md), [action repo](https://github.com/babaliauskas/evalshift-action) | <https://www.evalshift.dev/ci-llms-full.txt> |
| **SDK** | PyPI `evalshift-sdk` (import `evalshift`) | In-process capture: records your agent's model/tool calls to `.evalshift/captures/`. | [docs/sdk.md](docs/sdk.md), [SDK repo](https://github.com/evalshift/evalshift-sdk) | <https://www.evalshift.dev/sdk-llms-full.txt> |
| **GitHub Action** | `evalshift/evalshift-action@v0` | Runs the pipeline on PRs, pushes the run, maintains one PR comment, sets the `evalshift/regression` status. | [docs/github-action.md](docs/github-action.md), [action repo](https://github.com/evalshift/evalshift-action) | <https://www.evalshift.dev/ci-llms-full.txt> |
| **Hosted server** | service — API `https://api.evalshift.dev`, web app `https://evalshift.dev` | Stores pushed run bundles, diffs runs across branches, serves the web app, drives PR comments and gating. | [docs/hosted.md](docs/hosted.md) | covered by the CLI reference (`push`/`bundle` contract) |

Data flow is one-directional: **SDK captures → CLI runs and bundles → server stores and diffs → web app displays.** The SDK and CLI never call each other — the interface is files under `.evalshift/captures/`. The CLI (import `evalshift_cli`) depends on the SDK (import `evalshift`), so one environment holds both.
Expand All @@ -113,7 +113,7 @@ Doing it by hand, the mapping is:

## Quickstart

Point EvalShift at a real project. `evalshift init` writes a capture-first config, the [evalshift-sdk](https://github.com/babaliauskas/evalshift-sdk) records what your agent actually does, and `capture sync` turns those recordings into a golden suite:
Point EvalShift at a real project. `evalshift init` writes a capture-first config, the [evalshift-sdk](https://github.com/evalshift/evalshift-sdk) records what your agent actually does, and `capture sync` turns those recordings into a golden suite:

```bash
evalshift init
Expand Down Expand Up @@ -176,7 +176,7 @@ evalshift capture clean # delete promoted capture files + sweep o
7. Skips captures whose turn recorded an `error` event — a turn that died before the agent acted is not ground truth, and promoting it would assert `expected_no_tools: true` on a question that needed a tool. `--allow-errored` promotes it anyway (still never asserting `expected_no_tools`). `capture promote` exits non-zero on the same condition. Separately and unconditionally — `--allow-errored` does not help — a capture whose first `model_call` has no `toolset_ref` is refused: the SDK did not record what tools were offered, so there is nothing to carry, and re-capturing with a current `evalshift-sdk` is the only fix.
8. Warns when two captures claim the same `(conversation_id, turn_index)` (a retried turn), and when a promoted turn contains a failed tool result (`error`, or `{"success": false}`). Both stay warnings — see [Agent evals → What does not belong in a golden suite](docs/agents.md).
9. Writes `.evalshift/suites/<suite>/golden.jsonl` and rewrites the managed `suites:` block in `evalshift.yaml` (between the `>>> evalshift suites` markers).
10. After the write (or after printing the block for you to paste), checks the CI pin: if a workflow under `.github/workflows/` uses `babaliauskas/evalshift-action` with an `evalshift-version` older than this CLI, with no pin at all, or with pins that are all newer than this CLI, it prints a warning naming the workflow and job plus the fix — the exact `evalshift-version: "<this version>"` line to set, or, for a newer pin, `pip install -U evalshift` locally. Advisory only — sync never edits a workflow and the exit code is unchanged. See [Pin drift](#pin-drift).
10. After the write (or after printing the block for you to paste), checks the CI pin: if a workflow under `.github/workflows/` uses `evalshift/evalshift-action` with an `evalshift-version` older than this CLI, with no pin at all, or with pins that are all newer than this CLI, it prints a warning naming the workflow and job plus the fix — the exact `evalshift-version: "<this version>"` line to set, or, for a newer pin, `pip install -U evalshift` locally. Advisory only — sync never edits a workflow and the exit code is unchanged. See [Pin drift](#pin-drift).

Strictness knobs for the derived tool expectations: `--strict-args` (exact argument matches), `--names-only` (ignore arguments), `--tool-count` (also pin the call count, scoped the same way as `expected_tools`), `--rounds {first,all}` (which agent rounds become ground truth, default `first`). `--tag` attaches extra slice tags; `--print` previews the `suites:` block without writing.

Expand Down Expand Up @@ -794,15 +794,15 @@ A suite wired under `suites:` is therefore selected by name — which is what `i

`evalshift.yaml` is `extra="forbid"` everywhere, so the CLI that *reads* the config in CI must be at least as new as the CLI that *wrote* it locally — a newer `capture sync` or `init` can add keys an older release rejects outright. The action installs an exact version (`evalshift-version`, or its own default when the input is absent), which is where drift creeps in: you upgrade locally, re-sync, and CI still installs last month's release.

The CLI checks for this wherever it writes or validates config — `capture sync`, `init` (without `--ci`, next to a workflow it didn't write — `init --ci` pins the scaffolding CLI itself and does not warn about the file it just wrote), `doctor` (a `ci pin` row), and `validate` — by parsing every `.github/workflows/*.yml` for `babaliauskas/evalshift-action` steps and comparing their `evalshift-version` with its own:
The CLI checks for this wherever it writes or validates config — `capture sync`, `init` (without `--ci`, next to a workflow it didn't write — `init --ci` pins the scaffolding CLI itself and does not warn about the file it just wrote), `doctor` (a `ci pin` row), and `validate` — by parsing every `.github/workflows/*.yml` for `evalshift/evalshift-action` steps and comparing their `evalshift-version` with its own:

- **stale** — a literal pin is older than the local CLI. Fix: set `evalshift-version: "<local version>"` on the step.
- **unpinned** — a step has no `evalshift-version`, so the action default applies and may lag. Fix: add the pin.
- **ahead** — every pin is newer than the local CLI. Fix: `pip install -U evalshift`.

Equal pins, `${{ }}` expressions, unparseable versions, and an editable install without metadata (`0.0.0+unknown`) are silent. The check is advisory: it never edits a workflow and never changes an exit code, and in CI it is a no-op by construction (the running CLI *is* the pin). Config `version: 1` is not bumped for additive fields, nor for a removal that fails the load with a message naming the key — see [Configuration](docs/configuration.md#config-version-policy).

Secrets needed: a provider API key matching your config's models, and `EVALSHIFT_TOKEN` — a service account key from Settings → API tokens → Service accounts, scoped to `run:create` + `run:read` + `policy:read`, stored as an encrypted repository or environment secret. Not a personal token, never a literal in the workflow YAML, and never reachable from `pull_request_target`. Rotate by minting the successor first (24h grace), updating the secret, confirming a green run, then letting the old key expire. One thing a scoped key can't do, by design: auto-create the project (`project:create` is owner-only — pre-create it and set `create-project: false`). Full guidance: the action's [README](https://github.com/babaliauskas/evalshift-action#readme).
Secrets needed: a provider API key matching your config's models, and `EVALSHIFT_TOKEN` — a service account key from Settings → API tokens → Service accounts, scoped to `run:create` + `run:read` + `policy:read`, stored as an encrypted repository or environment secret. Not a personal token, never a literal in the workflow YAML, and never reachable from `pull_request_target`. Rotate by minting the successor first (24h grace), updating the secret, confirming a green run, then letting the old key expire. One thing a scoped key can't do, by design: auto-create the project (`project:create` is owner-only — pre-create it and set `create-project: false`). Full guidance: the action's [README](https://github.com/evalshift/evalshift-action#readme).

---

Expand Down Expand Up @@ -993,6 +993,6 @@ Nowhere, by default. Model inputs/outputs go to the providers you configured (th
- [docs/](docs/) — the mkdocs site: [getting-started](docs/getting-started.md), [configuration](docs/configuration.md), [evaluators](docs/evaluators.md), [methodology](docs/methodology.md), [agents](docs/agents.md), [conversations](docs/conversations.md), [traces](docs/traces.md), [sdk](docs/sdk.md), [hosted](docs/hosted.md), [github-action](docs/github-action.md), [faq](docs/faq.md)
- [AGENTS.md](AGENTS.md) — repo orientation for AI coding agents; [CLAUDE.md](CLAUDE.md) — contributor workflow rules
- [CHANGELOG.md](CHANGELOG.md) — release history
- [evalshift-sdk](https://github.com/babaliauskas/evalshift-sdk) — the in-process capture SDK
- [evalshift-sdk](https://github.com/evalshift/evalshift-sdk) — the in-process capture SDK
- [llms-full.txt](llms-full.txt) — dense single-file reference for AI coding tools, hosted at <https://www.evalshift.dev/cli-llms-full.txt>
- License: Apache-2.0
14 changes: 7 additions & 7 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@

Open-source LLM migration and regression testing for AI agents.

[![CI](https://github.com/babaliauskas/evalshift-cli/actions/workflows/ci.yml/badge.svg)](https://github.com/babaliauskas/evalshift-cli/actions/workflows/ci.yml)
[![CI](https://github.com/evalshift/evalshift-cli/actions/workflows/ci.yml/badge.svg)](https://github.com/evalshift/evalshift-cli/actions/workflows/ci.yml)
[![License: Apache 2.0](https://img.shields.io/badge/License-Apache_2.0-blue.svg)](LICENSE)
[![Python 3.11+](https://img.shields.io/badge/python-3.11%2B-blue.svg)](https://www.python.org/downloads/)
[![PyPI](https://img.shields.io/pypi/v/evalshift.svg)](https://pypi.org/project/evalshift/)
Expand All @@ -27,7 +27,7 @@ statistics**: paired tests, Cohen's d, 95% CIs, and Benjamini-Hochberg
correction across every (prompt x evaluator x slice) comparison.

An eval is only worth the examples in it. That is why the
[capture SDK](https://github.com/babaliauskas/evalshift-sdk) is part of the
[capture SDK](https://github.com/evalshift/evalshift-sdk) is part of the
product rather than an add-on: it records real production runs — model calls,
tool calls, final outputs — to disk, and `evalshift capture sync` promotes them
into golden suites. Hand-written suites are fully supported too, but captured
Expand All @@ -44,7 +44,7 @@ Four pieces, released and documented independently:
| --- | --- | --- |
| **SDK** — PyPI `evalshift-sdk` | Records what your agent actually did in production — model calls, tool calls, final output — as capture files on disk. Those captures become your golden suite. | [docs/sdk.md](docs/sdk.md) |
| **CLI** — this repo, PyPI `evalshift` | Replays the suite on two models, scores, analyses, reports, bundles, pushes. | [DOCS.md](DOCS.md) |
| **GitHub Action** — `babaliauskas/evalshift-action@v0` | Runs the pipeline on pull requests, pushes the run, posts one PR comment, sets the `evalshift/regression` status. | [docs/github-action.md](docs/github-action.md) |
| **GitHub Action** — `evalshift/evalshift-action@v0` | Runs the pipeline on pull requests, pushes the run, posts one PR comment, sets the `evalshift/regression` status. | [docs/github-action.md](docs/github-action.md) |
| **Hosted server** — `api.evalshift.dev`, web app at `evalshift.dev` | Optional. Stores pushed run bundles, diffs them across branches, drives PR comments and gating. | [docs/hosted.md](docs/hosted.md) |

The SDK and the CLI never call each other — the interface is files under
Expand Down Expand Up @@ -95,7 +95,7 @@ uv pip install evalshift-sdk # or: pip install evalshift-sdk
From source (for contributors):

```bash
git clone https://github.com/babaliauskas/evalshift-cli.git
git clone https://github.com/evalshift/evalshift-cli.git
cd evalshift-cli
uv venv --python 3.11
source .venv/bin/activate
Expand All @@ -113,7 +113,7 @@ candidate model to it.
evalshift init # minimal capture-first evalshift.yaml
```

Instrument the agent with [evalshift-sdk](https://github.com/babaliauskas/evalshift-sdk)
Instrument the agent with [evalshift-sdk](https://github.com/evalshift/evalshift-sdk)
— stdlib-only, Python 3.10+, installed with the CLI or on its own:

```python
Expand Down Expand Up @@ -242,7 +242,7 @@ The short version:

`evalshift init --ci` scaffolds a production-shaped workflow: it discovers
every committed suite under `.evalshift/suites/`, evaluates each on every
pull request via [`babaliauskas/evalshift-action@v0`](https://github.com/babaliauskas/evalshift-action)
pull request via [`evalshift/evalshift-action@v0`](https://github.com/evalshift/evalshift-action)
(one matrix job per suite), pushes the runs to hosted EvalShift, compares
against the latest compatible base-branch run, posts one PR comment, and
gates merges on your `migration_policy` through a single required
Expand Down Expand Up @@ -368,7 +368,7 @@ Runnable projects under [`examples/`](examples/):

[Apache-2.0](LICENSE). Free for any use, commercial included — no share-back
requirement, and an explicit patent grant. The capture SDK
([`evalshift-sdk`](https://github.com/babaliauskas/evalshift-sdk)), the piece
([`evalshift-sdk`](https://github.com/evalshift/evalshift-sdk)), the piece
you import into your own application, is MIT.

Earlier releases stay under the license they shipped with: `0.3.0` and earlier
Expand Down
Loading
Loading