[Feat] Automatic approvals: a decision model risk-assesses tool calls nobody has made a choice about - #3084
Merged
Conversation
…he decision model answer Ask first calls
Contributor
|
No new code issues found. See task
Reviewed 6ed52fd |
…t made of a call on the card
…tions page with Roomote's copy
Bruno's per-tool controls (#3103) stay as they are: a toggle-button row with tooltips, Always allow / Always ask / Disable, legacy availability folded into Disable, bulk saves through setMany. The only change to them is the one this branch is for: Auto is no longer a button of its own on a tool row. It is the default, shown as nothing pressed; pressing the selected choice again returns to it. The group row keeps an Auto button so a whole group can be returned to it at once. The native in-process MCP guard from develop now reads the branch's approvals shape, so default tools follow Auto mode there exactly as at the proxy: shadow-assessed, or held for a claim while Auto is on. The proxy's shadow call no longer reads a body on GET requests.
This was referenced Sep 22, 2026
7 of 8 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Part of the experimental per-tool approvals (
integrationToolApprovals). Auto mode is a deployment-wide setting that lets a decision model risk-assess tool calls nobody has made a choice about, run the routine ones, and ask a person about the risky ones. A manual choice always wins.Per-tool modes
Each tool row offers three stored choices: Always allow (Auto never looks), Always ask (always a person), Disable (hidden and refused). Auto is the default and shows as nothing pressed; pressing the selected choice again returns the tool to Auto. The group row carries the same buttons plus Auto, so a whole group can go back to it in one click. Personal policies still only tighten. The tool rows themselves are the toggle-button rows from #3103, unchanged apart from Auto no longer being a button of its own.
Auto mode (its own card on Settings → Integrations, admins only, shown while the experiment is on)
Customer-facing copy, no mechanics: "Let Roomote handle routine work and ask before anything risky. Your other tool choices stay the same." Off: "Keep each tool’s current choice." On: "Handle routine work automatically and ask before anything risky." The guidance field sits behind an "Additional instructions" disclosure, named like the other guidance fields in the app. When On can't be enabled, the card just says "Auto mode isn’t available yet."
integration_tool_auto_evaluationstable (migration0114, additive), so the model's judgment can be reviewed against real traffic before turning On.The assessment
Following the checked-in
typesafe-aiskill: one Score question over five described situations (reads only → easy to undo → sends to people → spends or grants access → destroys data), plus Noul questions for "matches what the user asked" (only when a request is known), "steered by untrusted content", and "the guidance flags calls like this" (only when guidance is set). The guidance rides in the state, not the questions. The decision is in code: run only at the lowest risk level with confidence, matching the request, not steered, not flagged. Otherwise ask. The model never rejects; no model, an error, or a timeout all ask.Enforcement
askso the bridge can assess each call; a routine call runs through the reservation-and-claim path asauto_approvedwith no decider and the model's answers on the row.<server>_*: ask; at start the worker lists each Auto-gated server's tools and maps every native key to the real tool name (a key two tools share, or a server that cannot be listed, gets its asks refused), and sends the user's latest request with the ask; the server decides who answers from the governing policies (never from what the worker says); a routine call is leftapprovedfor the integration proxy to claim for the exact arguments, once. The proxy gates every default tool call of a task while On.Testing
Unit and real-database tests across types, db, cloud-agents, sdk, api, worker, and web. Smoke-tested on a local stack with Jev via OpenRouter as the judgment model:
read_wiki_structurecall ran untouched and a shadow row was recorded (risk 0, confidence 1).Screenshots
The Auto mode card on Settings → Integrations, collapsed and with the additional instructions opened:
The per-tool choices in Manage tools: Auto is the default and shows as nothing pressed, the group row can return every tool to it, and pressing a choice again returns that tool to Auto:
An approval card after Auto judged a call risky and asked a person: