Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
27 changes: 27 additions & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,27 @@
# Working on Judgement

Judgement is a small ESM JavaScript library and CLI. Read README.md and
[docs/agents.md](docs/agents.md) for the public workflow and command semantics.

## Ownership

Keep generic checks, fixture loading, isolated Git setup, testing, calibration,
threshold selection, reporting, and cancellation in this library. Consuming
projects own their rules, labeled examples, and any custom inference adapter.
Generated calibration reports belong in local output or CI artifacts.

## Changes and validation

- Keep `src/index.d.ts` aligned with public JavaScript exports.
- Keep CLI help, README, and the packaged agent guide aligned with actual behavior.
- Run `npm run check` for code changes. Use `npm pack --dry-run --json` when changing
package files or documentation intended for installed-package users.
- Test Git behavior with disposable repositories. Preserve staged/unstaged
separation, base-policy enforcement, cancellation, and bounded concurrency.
- Unit tests use deterministic evaluators. Live model evaluations must be explicit,
use synthetic data, and report incomplete and failed judgments accurately.
- Do not automatically run calibration during normal hooks, overwrite user rules,
or silently drop examples that fail. Runtime and test commands have different
exit semantics; document and test both.
- Preserve the project's pinned dependencies. Do not publish as a side effect of
testing or preparing a PR.
176 changes: 149 additions & 27 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,23 +12,30 @@ Check code changes against your team's business rules, written in plain language
"context": ["src/audit/**", "src/auth/**"]
},
{
"id": "payment-owner",
"rule": "Only workspace owners may change payment settings."
}
]
}
```

Save this as `JUDGE.json` at the repository root. Run Judgement before a commit
Save this as `.judgement/rules.json` in your repository. Run Judgement before a commit
or in CI. It uses [Jev](https://typesafe.ai) to make focused judgments and reads
related source files from the exact Git snapshot being checked.

## Agent instructions

Coding agents: read [the agent guide](docs/agents.md) before adding or tuning rules.
It is included in the npm package and covers fixtures, commands, uncertainty, and
validation. Projects should keep their rules and examples, not copy the runner.

## Install

Node 22.20+ or 24+ and Git are required. Glob matching uses the small `picomatch` dependency.
Install the published package from npm.

```sh
npm install --save-dev --save-exact @roo-code/judgement@0.1.5
npm install --save-dev --save-exact @roo-code/judgement@0.2.0
export TYPESAFE_API_KEY=...
./node_modules/.bin/judgement check --staged
```
Expand Down Expand Up @@ -71,12 +78,12 @@ PR, resolve the merge base explicitly and pass its SHA. Checkout full history.
Use policy from a trusted base and provide the API key through CI secrets.
Run untrusted contributions without write-capable repository tokens.

| Exit | Strict check | `--hook` |
| --- | --- | --- |
| 0 | Completed without a blocking finding | Completed or explicitly incomplete |
| 1 | Confirmed violation | Confirmed violation |
| 2 | Invalid configuration or invocation | Invalid configuration or invocation |
| 3 | Incomplete, uncertain, service/Git failure | Reported, commit allowed |
| Exit | Strict check | `--hook` |
| ---- | ------------------------------------------ | ----------------------------------- |
| 0 | Completed without a blocking finding | Completed or explicitly incomplete |
| 1 | Confirmed violation | Confirmed violation |
| 2 | Invalid configuration or invocation | Invalid configuration or invocation |
| 3 | Incomplete, uncertain, service/Git failure | Reported, commit allowed |

`--advisory` reports all outcomes without blocking. `--timeout-ms` changes the
budget (strict mode defaults to 120 seconds). Use JSON `status` to distinguish
Expand Down Expand Up @@ -107,6 +114,129 @@ Existing policy comes from the base tree. If no policy exists there, a newly
staged policy can bootstrap checking. Editing or deleting a policy does not
weaken the check for the same commit. Judgement never rewrites your policy.

## Project layout

```text
.judgement/
rules.json
examples/
wording.json
billing-audit.json
```

Give each rule a stable `id`. Calibration uses that ID to find its examples.
Policy files and `.judgement/examples/` are excluded from normal rule checks, so
labeled counterexamples do not trigger your commit hook.
Save rules in `.judgement/rules.json`. File and context globs are relative to the
repository root.

## Test all rules

```sh
judgement test --dry-run
judgement test --rule wording --repeats 3
judgement test --format json > /tmp/judgement-tests.json
judgement calibrate --all --format json > /tmp/judgement-calibration.json
```

`test` uses configured thresholds; a wrong, incomplete, or failed judgment exits

1. Dry runs and suites whose results all match their labels exit 0. Setup errors
exit 2. `calibrate --all` compares candidates for every rule and exits 3 if any rule
has no recommendation. Both commands require examples for every selected rule;
missing or invalid examples fail before inference. Use `--examples-dir <path>` for
a separate suite. Rules run sequentially, with bounded fixture concurrency within
each rule. JSON suite reports include hashes of the normalized policy and exact
fixture text used. Keep generated output local or in CI artifacts.

The library exposes `testRules(options)`, `calibrateRules(options)`,
`formatExampleSuite(report)`, and `exampleSuiteExitCode(report)`. Options include
an optional `ruleId` (omitting it selects all rules), `examplesDirectory`, and the
same custom `evaluate` callback used by `calibrate`. `testRules` always uses the
configured thresholds. This keeps project-specific inference adapters small.

## Calibration command

Save labeled examples in `.judgement/examples/wording.json` for a rule whose
`id` is `wording`:

```json
{
"ruleId": "wording",
"examples": [
{
"name": "Mid-sentence capital",
"path": "docs/guide.md",
"before": "Review the remaining checks.\n",
"after": "Review the remaining Session checks.\n",
"expected": "violation"
},
{
"name": "Sentence beginning",
"path": "docs/guide.md",
"before": "History is available.\n",
"after": "Session history is available.\n",
"expected": "pass"
}
]
}
```

Each example has a unique name and different `before`/`after` text. Use `null`
for the absent side of an addition or deletion. `path` defaults to `example.md`
and must match the rule's `files`. Optional `context` maps relative paths to
unchanged supporting file contents; the rule's `context` globs select which
ones the checker sees. Paths cannot escape the fixture or replace its policy
or Git metadata. Label valid exceptions and inapplicable changes as `pass`.

```sh
judgement calibrate --rule wording --dry-run
judgement calibrate --rule wording --repeats 3 --thresholds 0.8,0.85,0.9,0.96
judgement calibrate --rule wording --format json > calibration.json
```

The command reads your **working-tree policy**, then commits each candidate
policy in a disposable Git repository and stages its example there. Your real
policy, working tree, and index remain untouched. Checks use the full evaluator,
including context expansion, with caching disabled. Calibration is opt-in and
is never run by a commit hook. Real runs make paid model requests.

Defaults: three repetitions, two concurrent fixture checks, a three-second
per-check deadline, and thresholds `0.8`, `0.85`, `0.9`, `0.95`, plus the rule's
current threshold. Use `--examples <path>` to select a different examples file,
`--concurrency <1–8>` to limit parallel checks, `--timeout-ms <ms>` to match your
check budget, and `--verbose` for progress on stderr. JSON remains on stdout.

The report includes per-example answers and scores, caught violations, false
blocks, incorrect passes, incomplete results, failures, and timing. It recommends
only candidates that caught every labeled violation with zero false blocks,
provided there were no operational failures anywhere in the run. Among those
candidates it prefers more valid passes, then the highest threshold. A dataset
must contain both labels to receive a recommendation. An incomplete valid case
can remain even at the recommended threshold; inspect those rows before adopting
it. Nothing automatically rewrites your policy.

Calibration exits `0` when a candidate is recommended (or for a dry run), `3`
when none can be recommended, and `2` for invalid configuration or fixture setup.
Unlike `check`, an expected violation is a successful calibration observation.
Use a held-out examples file to verify the selected threshold before saving it.

Applications with their own inference configuration can use the same library
harness and supply their existing evaluator:

```js
import { calibrate, formatCalibrationReport } from '@roo-code/judgement';

const report = await calibrate({
cwd: process.cwd(),
ruleId: 'wording',
evaluate: yourExistingEvaluator,
});
console.log(formatCalibrationReport(report));
```

The [wording example](examples/wording/.judgement/) includes a policy and fixtures.

## Calibrating a rule's confidence threshold

Choose a threshold from labeled examples of your rule. Confidence measures how
Expand All @@ -122,7 +252,7 @@ for the distinction. The default `0.85` is a starting point, not a universal cut
surrounding context, and small edits to large files. Reserve some examples
to validate your choice after tuning.
2. **Use an isolated test repository for each candidate policy.** Commit the
candidate `JUDGE.json` as the baseline, then stage a representative change.
candidate `.judgement/rules.json` as the baseline, then stage a representative change.
Merely editing or staging a new threshold in an existing repository will
still evaluate against its base policy. Do not overwrite your real staged
work to run calibration fixtures. Isolating one rule also makes its results
Expand All @@ -143,6 +273,7 @@ for the distinction. The default `0.85` is a starting point, not a universal cut
backend/model version, outcomes, confidence scores, final statuses, and time.
Avoid concurrent checks against the same index; use separate fixtures or
run their repetitions sequentially.

4. **Compare candidate thresholds.** Count violations that would block, valid
changes that would incorrectly block, and incomplete checks in each group.
Also track outright incorrect passes. A valid change reported as incomplete
Expand All @@ -154,21 +285,6 @@ for the distinction. The default `0.85` is a starting point, not a universal cut
examples, and hook deadlines before adopting it. Recalibrate after changing
the rule wording, evidence selection, model, or backend.

For example, a wording-rule calibration used 14 labeled changes with three runs
per change. Initial diff-only scores gave this comparison:

| Threshold | Violation runs that would block | Valid runs that would incorrectly block |
| --- | --- | --- |
| 0.96 | 14/18 | 0/24 |
| 0.90 | 17/18 | 0/24 |
| 0.85 | 18/18 | 0/24 |

The full checker at `0.85` then blocked all 18 violation runs, with no false
blocks across 24 valid runs. However, 12 valid runs remained incomplete. That
supported lowering the threshold for this wording rule, while also revealing
uncertainty around permitted exceptions. These are observations from a small
sample, not an accuracy guarantee or a recommended threshold for every rule.

The same threshold applies to `pass`, `not_applicable`, and `violation` answers.
An answer below the threshold remains incomplete; `unclear` remains incomplete
regardless of confidence. Lowering the threshold cannot fix missing credentials,
Expand Down Expand Up @@ -251,9 +367,15 @@ Judgement's library report:
import { fail, warn } from 'danger';
import { check, formatReport } from '@roo-code/judgement';

if (!process.env.REVIEW_BASE) throw new Error('Set REVIEW_BASE to the trusted base SHA');
const report = await check({ base: process.env.REVIEW_BASE, head: 'HEAD', cache: false });
if (report.status === 'violation' || report.status === 'invalid') fail(formatReport(report));
if (!process.env.REVIEW_BASE)
throw new Error('Set REVIEW_BASE to the trusted base SHA');
const report = await check({
base: process.env.REVIEW_BASE,
head: 'HEAD',
cache: false,
});
if (report.status === 'violation' || report.status === 'invalid')
fail(formatReport(report));
else if (report.status === 'incomplete') warn(formatReport(report));
```

Expand Down
Loading
Loading