Skip to content

Test Judgement presets and calibrate binary violation cutoffs - #3308

Merged
mrubens merged 7 commits into
developfrom
codex/judgement-tester-presets
Sep 29, 2026
Merged

mrubens merged 7 commits into
developfrom
codex/judgement-tester-presets

Conversation

@mrubens

@mrubens mrubens commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Load 47 labeled repository-rule examples into the decision tester as exact Judgement request presets. Admins can inspect initial or expanded evidence, edit inputs, repeat requests, and compare violation probabilities against each rule's cutoff. The last ten runs retain their exact inputs and metadata. Expected labels never enter model input.

Ask whether changed lines need correction to satisfy the rule, using Judgement’s binary Noul question. Flag when the violation probability meets the cutoff. Lower probabilities produce no finding; missing evidence and inference failures remain incomplete. Remove the async efficiency and privacy rules and their fixtures. Preserve stable IDs for the six contextual rules and add synthetic validation and prose confirmation fixtures. Focus the worker rule on worker-only credentials crossing into task-controlled code; allow safe environment copying, task-scoped credentials, and trusted worker-only use. Clarify the prose rule's exemption for security requirements.

Judgement owns evidence preparation. Roomote bundles synthetic request inputs with source hashes and question-parity tests so the deployed tester needs no Git checkout. Observed scores and calibration reports stay out of the repository.

Cutoffs

Rule Violation cutoff
Worker credential isolation 0.75
Sanitized errors 0.80
Agent setup prose 0.85
UI copy 0.85
Ordinary noun capitalization 0.75
Direct current prose 0.75

The correction question was selected from eight variants. The five rules other than worker credential isolation have a recorded full-check regression covering 61 examples, three repetitions, at the configured cutoffs above: 69/69 violation runs caught and 114/114 valid runs passed, with no incomplete checks or operational failures.

The narrowed worker rule was checked through the full checker on 14 synthetic examples repeated five times at 0.75: 25/25 credential-leak runs caught and 45/45 valid runs passed, with no incomplete checks or operational failures. Cases include credential-bearing environment copies, renamed credentials, safe environment copies, task-scoped credentials, and trusted worker-only use. No real credential values are included.

Fresh capitalization/prose confirmation used 20 examples repeated five times: 40/40 violation runs caught and 60/60 valid runs passed at 0.75. These are bounded synthetic tests, not guarantees. Harder validation cases remain checked in even when they fall below the cutoff. Expanded validation and confirmation fixtures preserve the cases used in prompt selection and confirmation; observed reports remain local.

Dependency

Use the published @roo-code/judgement@0.3.1 in both worker and cloud-agents, with an exact first-party release-age exception. Includes the correction question from RooCodeInc/judgement#3 and its cache protocol update.

Validation

  • 11 focused tests passed against the updated library for preset provenance/question parity and local/server adapters. The tester's four interaction tests passed during its implementation.
  • Regenerated 47 example presets without inference; source-hash and exact-question parity test passed.
  • pnpm check-types:fast: 27 tasks passed.
  • pnpm lint:fast and pnpm knip passed.

@roomote-community

roomote-community Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

No new code issues found. See task

  • packages/cloud-agents/package.json:103 — The pinned binary Judgement revision can mark partial screen evidence as a strict-check pass instead of incomplete. — dismissed: oversized evidence is recorded as unresolved before partial screens are evaluated.

Reviewed 8e52c09

@mrubens mrubens changed the title Load Judgement examples into the decision tester Test Judgement presets and calibrate binary violation cutoffs Sep 29, 2026
Comment thread packages/cloud-agents/package.json Outdated
"@modelcontextprotocol/sdk": "^1.29.0",
"@opencode-ai/sdk": "1.18.30",
"@roo-code/judgement": "0.2.0",
"@roo-code/judgement": "github:RooCodeInc/judgement#354078180dccc3912f9324106bfcb1edb854cde1",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The pinned Judgement revision no longer records a lower-than-cutoff binary answer as unresolved. For an oversized file, check() sends complete: false screen requests; if none reaches the cutoff, the binary consume() path leaves no unresolved entry, so the rule becomes pass and a strict check exits 0. That contradicts this PR's guarantee that incomplete evidence remains incomplete, and lets partial evidence approve a file. Preserve an unresolved result for sub-cutoff screen requests in Judgement before pinning this revision (and add an integration regression here).

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Rechecked against the published 0.3.1 source. Before evaluating partial screens, check() records Evidence exceeds the file budget, which keeps the rule incomplete when no screen reaches the violation cutoff. The existing large-file regression covers this path, so this finding is invalid.

@mrubens
mrubens marked this pull request as ready for review September 29, 2026 21:27
@mrubens
mrubens marked this pull request as draft September 29, 2026 22:45
@mrubens
mrubens marked this pull request as ready for review September 29, 2026 22:50
@mrubens
mrubens merged commit 5e41a40 into develop Sep 29, 2026
18 checks passed
@mrubens
mrubens deleted the codex/judgement-tester-presets branch September 29, 2026 23:16
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant