[Feat] Add shadow-only Auto approval mode for integration tool policies - #3052
Closed
roomote-roomote[bot] wants to merge 6 commits into
Closed
roomote-roomote[bot] wants to merge 6 commits into
roomote-roomote[bot] wants to merge 6 commits into
Conversation
…new transcript hooks in tests
…efresh instances instead of rebuilding sessions; refuse colliding flattened tool keys
Add a fourth per-tool approval mode, auto, on top of the integration tool approvals experiment. Auto compiles to the same native ask rule as ask and human approval remains mandatory: the configured judgment model only records an advisory would_approve/would_ask recommendation (with reason, available confidence, provider, and the actual model id requested) on the approval row, bound to the exact call fingerprint and the effective instruction, correlated with the eventual human decision on the same row. Evaluation fails closed to the human decision on timeout, invalid output, or an unconfigured backend, and tool descriptions/arguments are treated as untrusted, redacted data. The settings UI labels Auto clearly as a shadow preview with an optional instruction, and the approval card shows the preview recommendation.
Contributor
|
1 issue outstanding. See task
Reviewed 63acfab |
| integrationId: string; | ||
| integrationName: string; | ||
| policies: Map<string, IntegrationToolPolicyMode>; | ||
| policies: Map<string, IntegrationToolPolicyMetadata>; |
Contributor
There was a problem hiding this comment.
The branch is based on 1a6a312, but the target is now ee99d41; GitHub correctly reports this PR as unmergeable. A three-way merge conflicts here because the target moved per-tool controls into McpToolManagementDialog, and it also conflicts on the Drizzle 0109 snapshot. Please rebase and port the Auto controls to the current dialog, then regenerate the migration with the next available number; taking this stale component wholesale would undo the target branch's UI refactor.
Base automatically changed from
roomote/integration-tool-approvals-experiment
to
develop
September 21, 2026 20:12
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Related issue
Stacked follow-up on #3001 (targets its branch, not
develop).Why this PR exists
What changed
Adds a fourth per-tool approval mode, Auto, to the
integrationToolApprovalsexperiment — strictly as a shadow preview: every Auto-gated call still pauses on the exact same native Ask asAsk every time, and the Session owner's decision remains the only authorization. What Auto adds is an advisory evaluation: the configured judgment model (TypeSafe/Jev via the existing judgment client) records what it would have recommended for that exact call, so deployments can measure agreement before ever trusting automation.autocompiles to the identical nativeaskrule asask— no new OpenCode surface, no change to Allow/Ask/Reject, and the evaluator has no path to issue an allow, change grants, or override areject/askpolicy. Provider, actor, and admin authorization ceilings stay authoritative.would_approve/would_ask) over the exact tool identity, the same secret-stripped redacted arguments shown on the approval card, a bounded excerpt of the requester's current intent, and the effective policy instruction (per-policy optional, default: "Only actions clearly requested by the user that are not destructive or irreversible"). Tool descriptions and arguments are passed explicitly as untrusted data, never instructions.would_askwith the reason, and the ordinary human Ask flow proceeds unchanged. Cancellation, experiment-disable, expiry, and consume-once semantics from [Feat] Add per-tool integration approvals experiment for code-mode sessions #3001 are untouched; switching a policy betweenaskandautoproduces the same compiled rules hash, so it never triggers a session instance refresh.jev-latestalias, recorded as requested rather than claimed immutable), the effective instruction, and the call's SHA-256 fingerprint — on the same row as the eventual human decision (status,decidedByUserId,decidedAt), so agreement between recommendation and decision is directly measurable. The evaluation never authorizes anything; confidence alone never approves a call.How it was tested
would_askon unconfigured backend and on evaluator errors (invalid output/timeout); per-policy instruction override; untrusted-data phrasing; 3s timeout.autocompiles to the same nativeaskrule; the native ask stays open with no reply while the human decision is pending even onwould_approve; a human approval after awould_approveevaluation relaysonceexactly once (consume-before-relay); a fail-closed evaluation is recorded and still asks; non-auto policies never invoke the evaluator (Allow/Ask/Reject unchanged); a pre-existing mock leak in the sharedbeforeEach(consume mock left atfalseby an earlier test) was fixed.pnpm lint:fast,pnpm check-types:fast, andpnpm knipgreen.Known scope limits: no live end-to-end run of local mock model → native Ask → shadow recommendation → human decision → execution (the sandbox has no configured judgment backend or connected integration, and the base PR's own remaining live gaps apply here too); helper-subagent coverage rides the same shared bridge code path and is covered by the unit tests rather than a separate live run.
Checklist
[Fix],[Feat],[Improve],[Refactor],[Docs], or[Chore]followed by a user-facing descriptionpnpm lintandpnpm check-typespass locallypnpm changesetScreenshots