Repository navigation
Test Judgement presets and calibrate binary violation cutoffs - #3308
Conversation
|
No new code issues found. See task
Reviewed 8e52c09 |
| "@modelcontextprotocol/sdk": "^1.29.0", | ||
| "@opencode-ai/sdk": "1.18.30", | ||
| "@roo-code/judgement": "0.2.0", | ||
| "@roo-code/judgement": "github:RooCodeInc/judgement#354078180dccc3912f9324106bfcb1edb854cde1", |
There was a problem hiding this comment.
The pinned Judgement revision no longer records a lower-than-cutoff binary answer as unresolved. For an oversized file, check() sends complete: false screen requests; if none reaches the cutoff, the binary consume() path leaves no unresolved entry, so the rule becomes pass and a strict check exits 0. That contradicts this PR's guarantee that incomplete evidence remains incomplete, and lets partial evidence approve a file. Preserve an unresolved result for sub-cutoff screen requests in Judgement before pinning this revision (and add an integration regression here).
There was a problem hiding this comment.
Rechecked against the published 0.3.1 source. Before evaluating partial screens, check() records Evidence exceeds the file budget, which keeps the rule incomplete when no screen reaches the violation cutoff. The existing large-file regression covers this path, so this finding is invalid.
Summary
Load 47 labeled repository-rule examples into the decision tester as exact Judgement request presets. Admins can inspect initial or expanded evidence, edit inputs, repeat requests, and compare violation probabilities against each rule's cutoff. The last ten runs retain their exact inputs and metadata. Expected labels never enter model input.
Ask whether changed lines need correction to satisfy the rule, using Judgement’s binary Noul question. Flag when the violation probability meets the cutoff. Lower probabilities produce no finding; missing evidence and inference failures remain incomplete. Remove the async efficiency and privacy rules and their fixtures. Preserve stable IDs for the six contextual rules and add synthetic validation and prose confirmation fixtures. Focus the worker rule on worker-only credentials crossing into task-controlled code; allow safe environment copying, task-scoped credentials, and trusted worker-only use. Clarify the prose rule's exemption for security requirements.
Judgement owns evidence preparation. Roomote bundles synthetic request inputs with source hashes and question-parity tests so the deployed tester needs no Git checkout. Observed scores and calibration reports stay out of the repository.
Cutoffs
The correction question was selected from eight variants. The five rules other than worker credential isolation have a recorded full-check regression covering 61 examples, three repetitions, at the configured cutoffs above: 69/69 violation runs caught and 114/114 valid runs passed, with no incomplete checks or operational failures.
The narrowed worker rule was checked through the full checker on 14 synthetic examples repeated five times at 0.75: 25/25 credential-leak runs caught and 45/45 valid runs passed, with no incomplete checks or operational failures. Cases include credential-bearing environment copies, renamed credentials, safe environment copies, task-scoped credentials, and trusted worker-only use. No real credential values are included.
Fresh capitalization/prose confirmation used 20 examples repeated five times: 40/40 violation runs caught and 60/60 valid runs passed at 0.75. These are bounded synthetic tests, not guarantees. Harder validation cases remain checked in even when they fall below the cutoff. Expanded validation and confirmation fixtures preserve the cases used in prompt selection and confirmation; observed reports remain local.
Dependency
Use the published
@roo-code/judgement@0.3.1in both worker and cloud-agents, with an exact first-party release-age exception. Includes the correction question from RooCodeInc/judgement#3 and its cache protocol update.Validation
pnpm check-types:fast: 27 tasks passed.pnpm lint:fastandpnpm knippassed.