Skip to content

[Feat] Automatic approvals: a decision model risk-assesses tool calls nobody has made a choice about - #3084

Merged
daniel-lxs merged 22 commits into
developfrom
feat/tool-approvals-auto-mode
Sep 22, 2026
Merged

daniel-lxs merged 22 commits into
developfrom
feat/tool-approvals-auto-mode

Conversation

@daniel-lxs

@daniel-lxs daniel-lxs commented Sep 22, 2026 •

Copy link
Copy Markdown
Member

Part of the experimental per-tool approvals (integrationToolApprovals). Auto mode is a deployment-wide setting that lets a decision model risk-assess tool calls nobody has made a choice about, run the routine ones, and ask a person about the risky ones. A manual choice always wins.

Per-tool modes

Each tool row offers three stored choices: Always allow (Auto never looks), Always ask (always a person), Disable (hidden and refused). Auto is the default and shows as nothing pressed; pressing the selected choice again returns the tool to Auto. The group row carries the same buttons plus Auto, so a whole group can go back to it in one click. Personal policies still only tighten. The tool rows themselves are the toggle-button rows from #3103, unchanged apart from Auto no longer being a button of its own.

Auto mode (its own card on Settings → Integrations, admins only, shown while the experiment is on)

Customer-facing copy, no mechanics: "Let Roomote handle routine work and ask before anything risky. Your other tool choices stay the same." Off: "Keep each tool’s current choice." On: "Handle routine work automatically and ask before anything risky." The guidance field sits behind an "Additional instructions" disclosure, named like the other guidance fields in the app. When On can't be enabled, the card just says "Auto mode isn’t available yet."

  • Off: tools run as they always have; Always ask still asks a person.
  • On: every call to a tool left on Auto is assessed first. Routine → runs; risky → the approval card, with "Auto flagged this call as risky and asked you." Disable, Always allow, Always ask, and session overrides are never touched.
  • Additional instructions (collapsed by default): what this deployment treats as routine or risky, in the admin's words. Given to the model as context, not as rules.
  • On requires a hosted judgment model. The helper-model fallback would be an LLM call per tool call, so the On option is disabled until one is configured, and the control says why.
  • Shadow is not an option: while Off, if a hosted judgment model is configured, every default tool call is still assessed in the background at the integration proxy and recorded in a new integration_tool_auto_evaluations table (migration 0114, additive), so the model's judgment can be reviewed against real traffic before turning On.

The assessment

Following the checked-in typesafe-ai skill: one Score question over five described situations (reads only → easy to undo → sends to people → spends or grants access → destroys data), plus Noul questions for "matches what the user asked" (only when a request is known), "steered by untrusted content", and "the guidance flags calls like this" (only when guidance is set). The guidance rides in the state, not the questions. The decision is in code: run only at the lowest risk level with confidence, matching the request, not steered, not flagged. Otherwise ask. The model never rejects; no model, an error, or a timeout all ask.

Enforcement

  • Sessions: with Auto On, default tools compile to a native ask so the bridge can assess each call; a routine call runs through the reservation-and-claim path as auto_approved with no decider and the model's answers on the row.
  • Tasks: default tools compile to <server>_*: ask; at start the worker lists each Auto-gated server's tools and maps every native key to the real tool name (a key two tools share, or a server that cannot be listed, gets its asks refused), and sends the user's latest request with the ask; the server decides who answers from the governing policies (never from what the worker says); a routine call is left approved for the integration proxy to claim for the exact arguments, once. The proxy gates every default tool call of a task while On.
  • A tool asked under a rule compiled while Auto was on, after it went off, simply runs.

Testing

Unit and real-database tests across types, db, cloud-agents, sdk, api, worker, and web. Smoke-tested on a local stack with Jev via OpenRouter as the judgment model:

  • Off + hosted model (shadow): a Session's read_wiki_structure call ran untouched and a shadow row was recorded (risk 0, confidence 1).
  • On: the same call to a tool with no policy row was auto-approved by the model (risk 0, matches request 0.98, steered 0.08); no card.
  • On + manual Always ask: the same tool set to Always ask showed the card with no model view on it.
  • Earlier: guidance flagging a call made it ask (guidance 0.99).

Screenshots

The Auto mode card on Settings → Integrations, collapsed and with the additional instructions opened:

Auto mode card, collapsed Auto mode card with additional instructions opened

The per-tool choices in Manage tools: Auto is the default and shows as nothing pressed, the group row can return every tool to it, and pressing a choice again returns that tool to Auto:

Manage tools dialog with every tool on Auto: nothing pressed on the tool rows, Auto pressed on the group row Manage tools dialog with one tool set to Always ask, the group row now showing a mixed group

An approval card after Auto judged a call risky and asked a person:

Approval card with the Auto line

@roomote-community

roomote-community Bot commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

No new code issues found. See task

  • Preserve an explicit session-level ask override instead of allowing Auto to bypass it.
  • Expose and render the stored Auto evaluation on approval cards for calls that still need a person.
  • Keep Caddy's local API and web proxy targets aligned with the development stack's published ports.
  • Preserve the real MCP tool identity for task Auto wildcard asks so evaluation and approval claims cannot cross sanitized-name collisions, including collisions between different server/tool pairs.
  • Hide the Auto mode section from non-admin members so it does not make unauthorized settings requests or render empty.

Reviewed 6ed52fd

Comment thread packages/cloud-agents/src/server/fast-agent/fast-agent-tool-approvals.ts Outdated
Comment thread packages/db/src/lib/integration-tool-approvals.ts
Comment thread .docker/caddy/Caddyfile Outdated
@daniel-lxs daniel-lxs changed the title [Feat] Auto mode for tool approvals: a deployment setting that lets the decision model answer Ask first calls [Feat] Automatic approvals: a decision model risk-assesses tool calls nobody has made a choice about Sep 22, 2026
Comment thread apps/worker/src/sandbox-server/lib/harnesses/opencode-server/tool-approvals.ts Outdated
Comment thread apps/web/src/components/settings/IntegrationToolAutoModeSection.tsx
Bruno's per-tool controls (#3103) stay as they are: a toggle-button row
with tooltips, Always allow / Always ask / Disable, legacy availability
folded into Disable, bulk saves through setMany. The only change to them
is the one this branch is for: Auto is no longer a button of its own on a
tool row. It is the default, shown as nothing pressed; pressing the
selected choice again returns to it. The group row keeps an Auto button
so a whole group can be returned to it at once.

The native in-process MCP guard from develop now reads the branch's
approvals shape, so default tools follow Auto mode there exactly as at
the proxy: shadow-assessed, or held for a claim while Auto is on. The
proxy's shadow call no longer reads a body on GET requests.
@daniel-lxs
daniel-lxs merged commit fefad24 into develop Sep 22, 2026
17 checks passed
@daniel-lxs
daniel-lxs deleted the feat/tool-approvals-auto-mode branch September 22, 2026 14:39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant