Skip to content

[Feat] Add shadow-only Auto approval mode for integration tool policies - #3052

Closed
roomote-roomote[bot] wants to merge 6 commits into
developfrom
roomote/jev-auto-approval-shadow
Closed

roomote-roomote[bot] wants to merge 6 commits into
developfrom
roomote/jev-auto-approval-shadow

Conversation

@roomote-roomote

Copy link
Copy Markdown
Contributor

​Opened on behalf of @daniel-lxs. Follow up by mentioning @roomote-roomote, in the web UI, or in Slack.

Related issue

Stacked follow-up on #3001 (targets its branch, not develop).

Why this PR exists

  • A maintainer explicitly invited this PR in the linked issue or discussion
  • I am a maintainer / this is internal Roomote work

What changed

Adds a fourth per-tool approval mode, Auto, to the integrationToolApprovals experiment — strictly as a shadow preview: every Auto-gated call still pauses on the exact same native Ask as Ask every time, and the Session owner's decision remains the only authorization. What Auto adds is an advisory evaluation: the configured judgment model (TypeSafe/Jev via the existing judgment client) records what it would have recommended for that exact call, so deployments can measure agreement before ever trusting automation.

  • auto compiles to the identical native ask rule as ask — no new OpenCode surface, no change to Allow/Ask/Reject, and the evaluator has no path to issue an allow, change grants, or override a reject/ask policy. Provider, actor, and admin authorization ceilings stay authoritative.
  • The bridge evaluates one bounded choice question (would_approve / would_ask) over the exact tool identity, the same secret-stripped redacted arguments shown on the approval card, a bounded excerpt of the requester's current intent, and the effective policy instruction (per-policy optional, default: "Only actions clearly requested by the user that are not destructive or irreversible"). Tool descriptions and arguments are passed explicitly as untrusted data, never instructions.
  • Fail closed everywhere: timeout (3s, the client's default), transport/HTTP error, invalid output, or no configured judgment backend all record would_ask with the reason, and the ordinary human Ask flow proceeds unchanged. Cancellation, experiment-disable, expiry, and consume-once semantics from [Feat] Add per-tool integration approvals experiment for code-mode sessions #3001 are untouched; switching a policy between ask and auto produces the same compiled rules hash, so it never triggers a session instance refresh.
  • Audit: the approval row stores the shadow evaluation — recommendation, reason, available confidence, provider, the actual model id requested (the TypeSafe direct path is the floating jev-latest alias, recorded as requested rather than claimed immutable), the effective instruction, and the call's SHA-256 fingerprint — on the same row as the eventual human decision (status, decidedByUserId, decidedAt), so agreement between recommendation and decision is directly measurable. The evaluation never authorizes anything; confidence alone never approves a call.
  • Settings UI: the per-tool dropdown gains "Auto (shadow preview — still asks you)" with a disclosure that redacted arguments and a short request excerpt leave the deployment to the configured judgment provider, plus an optional per-policy instruction field. The approval card shows the preview recommendation ("would have approved" / "would have asked", with confidence when available) while still requiring the human decision.

How it was tested

  • New shadow evaluator unit tests: would_approve/would_ask mapping with confidence, provider, and actual model id; fail-closed would_ask on unconfigured backend and on evaluator errors (invalid output/timeout); per-policy instruction override; untrusted-data phrasing; 3s timeout.
  • Bridge tests: auto compiles to the same native ask rule; the native ask stays open with no reply while the human decision is pending even on would_approve; a human approval after a would_approve evaluation relays once exactly once (consume-before-relay); a fail-closed evaluation is recorded and still asks; non-auto policies never invoke the evaluator (Allow/Ask/Reject unchanged); a pre-existing mock leak in the shared beforeEach (consume mock left at false by an earlier test) was fixed.
  • DB lifecycle tests: per-policy instruction stored/cleared with mode changes; shadow evaluation persisted on the pending row, bound to the request fingerprint, and still visible after the human decision for correlation.
  • Suites: db 17 tests, cloud-agents fast-agent 1015 tests, web settings 679 tests, feature-flags 17 tests — all green; pnpm lint:fast, pnpm check-types:fast, and pnpm knip green.
  • Settings UI verified in the browser: section copy and the Auto option/disclosure/instruction field with real policy + instruction persistence. The per-tool screenshot used a disclosed simulated tool list (no connected integrations exist in the local dev database) on the real settings page; it proves rendering and interaction of the new controls, not a live integration.

Known scope limits: no live end-to-end run of local mock model → native Ask → shadow recommendation → human decision → execution (the sandbox has no configured judgment backend or connected integration, and the base PR's own remaining live gaps apply here too); helper-subagent coverage rides the same shared bridge code path and is covered by the unit tests rather than a separate live run.

Checklist

  • The PR title follows the repo convention: [Fix], [Feat], [Improve], [Refactor], [Docs], or [Chore] followed by a user-facing description
  • This PR is small and scoped to one change
  • pnpm lint and pnpm check-types pass locally
  • I added tests or included a clear manual validation note above
  • I removed secrets, tokens, private keys, and customer data from code, logs, and screenshots
  • If this change should appear in the changelog, I ran pnpm changeset

Screenshots

Experimental settings: integration tool approvals section with Auto shadow preview copy

Per-tool Auto (shadow preview) mode selected with disclosure and instruction field — simulated tool list on the real settings page

…efresh instances instead of rebuilding sessions; refuse colliding flattened tool keys
Add a fourth per-tool approval mode, auto, on top of the integration tool
approvals experiment. Auto compiles to the same native ask rule as ask and
human approval remains mandatory: the configured judgment model only records
an advisory would_approve/would_ask recommendation (with reason, available
confidence, provider, and the actual model id requested) on the approval row,
bound to the exact call fingerprint and the effective instruction, correlated
with the eventual human decision on the same row. Evaluation fails closed to
the human decision on timeout, invalid output, or an unconfigured backend,
and tool descriptions/arguments are treated as untrusted, redacted data. The
settings UI labels Auto clearly as a shadow preview with an optional
instruction, and the approval card shows the preview recommendation.
@roomote-community

roomote-community Bot commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

1 issue outstanding. See task

  • Rebase onto roomote/integration-tool-approvals-experiment and port the Auto controls to its current Manage tools dialog; the branch currently conflicts in apps/web/src/components/settings/IntegrationToolApprovalsExperimentalSetting.tsx and the Drizzle 0109 snapshot.

Reviewed 63acfab

integrationId: string;
integrationName: string;
policies: Map<string, IntegrationToolPolicyMode>;
policies: Map<string, IntegrationToolPolicyMetadata>;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The branch is based on 1a6a312, but the target is now ee99d41; GitHub correctly reports this PR as unmergeable. A three-way merge conflicts here because the target moved per-tool controls into McpToolManagementDialog, and it also conflicts on the Drizzle 0109 snapshot. Please rebase and port the Auto controls to the current dialog, then regenerate the migration with the next available number; taking this stale component wholesale would undo the target branch's UI refactor.

Base automatically changed from roomote/integration-tool-approvals-experiment to develop September 21, 2026 20:12
@daniel-lxs daniel-lxs closed this Sep 22, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant