From b2e7f274aacadeda5826d63399f46597d300a862 Mon Sep 17 00:00:00 2001 From: "@mrubens" <2600+mrubens@users.noreply.github.com> Date: Wed, 23 Sep 2026 03:25:17 +0000 Subject: [PATCH] docs: correct recent user-facing guidance --- apps/docs/cost-analytics.mdx | 7 ++++++ apps/docs/integrations/index.mdx | 41 +++++++++++++++++++++++++++----- apps/docs/models.mdx | 36 +++++++++++++++++++++++++--- apps/docs/tasks.mdx | 7 ++++++ 4 files changed, 82 insertions(+), 9 deletions(-) diff --git a/apps/docs/cost-analytics.mdx b/apps/docs/cost-analytics.mdx index fedae926e3..4f016bf30f 100644 --- a/apps/docs/cost-analytics.mdx +++ b/apps/docs/cost-analytics.mdx @@ -19,6 +19,7 @@ recorded inference usage. Use it to answer questions such as: - which task and automation types account for the most spend - how much session orchestration and memory contribute to spend +- how much judgment-model routing and review pre-screening contribute to usage - which product features or runtime sources account for inference spend - whether a particular environment or model is driving costs - how usage differs between people and automations @@ -64,6 +65,12 @@ pricing nor token usage cannot contribute to those metrics. Use provider and model filters to narrow an investigation when totals do not match an external provider invoice. +Judgment-model calls appear as non-task inference under the `judgment_model` +source. Primary decisions and optional shadow evaluations are separate requests, +so both contribute when shadow mode is enabled. Roomote records provider-reported +tokens and cost when available; missing metadata leaves those values at zero +without changing the decision or its fallback behavior. + Deleting a task does not remove its recorded inference usage from Cost Analytics. The historical usage still contributes to totals, task counts, and averages, but its detail row becomes **Deleted task** and no longer exposes the diff --git a/apps/docs/integrations/index.mdx b/apps/docs/integrations/index.mdx index 3bef4029a2..e386878023 100644 --- a/apps/docs/integrations/index.mdx +++ b/apps/docs/integrations/index.mdx @@ -62,12 +62,41 @@ silently bypassed with another connection method. ## Manage available tools Some connected integrations expose a **Manage tools** action in -**Settings > Integrations**. Admins can use it to restrict the types of -operations Roomote can make with the service. - -Use this as a coarse permissions system, for example allowing only read -operations or restricting access to a certain type of entity. The list varies -depending on the integration. +**Settings > Integrations**. The dialog groups tools into read-only and +write/delete sections when the integration supplies that metadata. + +Admins can enable **Integration tool approvals** under **Settings > +Experimental**. While it is enabled, each tool has one of four modes: + +- **Auto** uses the deployment's Auto mode. With Auto mode on, the configured + judgment model runs routine calls and asks the session owner before risky + ones. With Auto mode off, the tool runs as it did before approvals were + enabled. +- **Always allow** runs without an Auto assessment or approval prompt. +- **Always ask** pauses every call for the session owner. +- **Disable** hides the tool from agents and rejects direct calls at the + Roomote proxy. + +Configure deployment-wide Auto mode under **Settings > Integrations**. Turning +it on requires a configured [judgment model](/models#judgment-model). Admins can +also add plain-language guidance describing what their deployment considers +routine or risky. Auto can only run a call or ask; it never rejects one, and it +does not widen the integration, actor, or provider permissions that decide +whether the tool is available in the first place. + +An approval request shows the integration, tool, and sanitized arguments. The +session owner can allow the call once, allow that tool for the rest of the +session, or deny it. Unanswered requests expire and fail closed. A delegated +task asks its parent session's owner; an unattended task, such as an +automation-started task, cannot run an **Always ask** tool when nobody can +answer. + +Deployment policies are admin-managed and apply from the next session turn. +Members can manage personal policies for their own personal integrations and +MCP servers. When both deployment and personal policies govern a built-in +integration, the stricter choice wins, so a personal policy cannot loosen an +admin restriction. When the experiment is off, the existing enabled/disabled +tool controls remain in effect. Some integrations can only list tools after the first user links their account from [Personal Settings](/personal-settings). diff --git a/apps/docs/models.mdx b/apps/docs/models.mdx index 0c88dd1ee9..2754f1c31d 100644 --- a/apps/docs/models.mdx +++ b/apps/docs/models.mdx @@ -142,6 +142,21 @@ in **Available Models**, and assign it to a role or select it for a session or task. Adding these routes does not change existing model defaults; access still depends on the connected provider account or plan. +**Claude Opus 5.5** replaces Claude Opus 5 in the recommended catalog. It is +available through managed Roomote inference, Anthropic, Azure OpenAI, Azure AI +Foundry, Amazon Bedrock, OpenRouter, Vercel AI Gateway, Requesty, OpenCode Zen, +and GitHub Copilot. New recommended mappings use it for planning where those +providers previously used Claude Opus 5. + +**GPT-6 Sol** and **GPT-6 Luna** replace GPT 5.6 Sol and Luna in the +recommended catalog for managed Roomote inference, OpenAI, Azure OpenAI, Azure +AI Foundry, Amazon Bedrock, OpenRouter, Vercel AI Gateway, Requesty, OpenCode +Zen, GitHub Copilot, and ChatGPT Subscription. New OpenAI and ChatGPT +connections use Sol as the coding default, while new GitHub Copilot connections +use Luna. Applying the current provider presets uses the GPT-6 mappings, but +existing saved role mappings do not change automatically. Provider and account +availability still apply. + **DeepSeek V4.1 Flash** is recommended through [DeepSeek](/providers/inference/deepseek), [OpenRouter](/providers/inference/openrouter), @@ -194,9 +209,9 @@ You can also split model roles with env vars: R_ORCHESTRATION_MODEL=openrouter/google/gemini-3.8-flash R_ORCHESTRATION_MODEL_REASONING_EFFORT=low R_SMALL_MODEL=openrouter/openai/gpt-4.1-mini -R_VISION_MODEL=openrouter/openai/gpt-5.6-sol -R_CODE_REVIEW_MODEL=openrouter/openai/gpt-5.6-sol -R_EXPLORE_MODEL=openrouter/openai/gpt-5.6-luna +R_VISION_MODEL=openrouter/openai/gpt-6-sol +R_CODE_REVIEW_MODEL=openrouter/openai/gpt-6-sol +R_EXPLORE_MODEL=openrouter/openai/gpt-6-luna ``` Roomote automatically makes common provider configuration available to the @@ -358,6 +373,16 @@ When a judgment model is on, Roomote asks it first for: - ranking a task's own changed hunks by defect risk during its self-review, before it pushes or opens a pull request (the same pre-screen pull request reviews use), as advisory hints the agent re-reads +- ranking specific changed hunks before a pull-request review, with up to three + advisory hints for the reviewer; the review still examines the complete diff, + and a hunk that is not highlighted is not treated as cleared + +Admins can separately enable **Jev Session communication** under **Settings > +Experimental**. It uses the judgment model for high-confidence task reports, +blocked-input detection, redundant-update suppression, and evidenced completion +reporting. Low-confidence, unavailable, or policy-sensitive cases fall back to +the regular model path. Steering, permissions, cancellation, lifecycle, and +destructive actions remain outside the experiment. Roomote acts on the judgment model only when it is confident. When it is unsure, unavailable, or returns an error, Roomote keeps the behavior it has @@ -376,6 +401,11 @@ helper model never opts into these batches, so deployments without a judgment model keep the existing conservative behavior and do not incur new high-volume language-model calls. +Primary and shadow judgment-model requests are recorded as non-task inference +under the `judgment_model` source in [Cost Analytics](/cost-analytics). A +provider that omits token or price metadata can still make a successful +decision, but the corresponding analytics row may have zero tokens or cost. + ## Reasoning settings Some models expose reasoning controls. Roomote lets admins set reasoning levels diff --git a/apps/docs/tasks.mdx b/apps/docs/tasks.mdx index 67258b8eda..12ae8b8689 100644 --- a/apps/docs/tasks.mdx +++ b/apps/docs/tasks.mdx @@ -202,6 +202,13 @@ that prior workspace instead of starting over. Snapshot retention depends on the sandbox provider: Modal-backed snapshots do not have Roomote's seven-day application expiry, while other providers may use a bounded window. +If you send a follow-up while the sandbox runtime is still starting, Roomote +accepts it into the task's queue and delivers it when the runtime can receive +instructions. Follow-ups stay in order, including later messages sent after the +runtime becomes reachable, and a retried send with the same client message ID +is not delivered twice. If the run finishes before Roomote can admit the +message, the request is rejected instead of reviving the settled run. + When a task delegated from a session asks for input, reply naturally in the originating session. If one input request is pending, Roomote applies a matching reply to that request instead of treating it as a separate