Promote v1.15.1 to production - #3376
Merged
Merged
Conversation
Co-authored-by: @mrubens <2600+mrubens@users.noreply.github.com>
…guments (#3347) * [Fix] Approvals no longer assess parallel calls with another call's arguments * Add changeset for parallel call approvals * Address review: share assessments for identical parallel calls, keyed by tool * Ask about parallel calls on one batch card instead of refusing them * Address review: share decisions only among waiting calls; namespace the batch marker * Address review: batch cards for identical parallel calls; record a batch as a list
* [Feat] Auto pauses a Session when calls can't be checked * Address review: lowercase session in prose; one pause notice per turn; honor Auto off after a pause
) * [Improve] Auto runs calls the owner asked for or already approved * Money checks judge the call, not the user's description of it * Address review: lowercase session in prose; earlier approvals alone reach the authorization question * Auto runs the rest of approved work and plans the owner agreed to * Auto decides routine reads with a read-only question instead of risk confidence * Keep a same-tool rejection in force for the whole session The recent outcomes the model sees are capped, so a rejection could drop out after a few newer decisions and let an authorization path approve the rejected tool. Look the rejection up session-wide and treat a lookup failure as a rejection. * Recheck a same-tool rejection right before Auto runs a call A call assessed while the owner rejected another call to the same tool could still auto-run on the stale lookup. Check again before reserving the auto-approval and ask instead. * Check a same-tool rejection inside the Auto approval claim The recheck before reserving still left a window before the claim. The claim now fails, and its reservation is cancelled, when the requester has rejected a call to this tool in the session, so a rejection that commits before the call runs is never missed. * Serialize same-tool rejections with Auto's approval claim A rejection and a guarded Auto claim of the same session tool now take the same transaction lock, so the claim's check always sees a rejection that committed first. * Take the rejection lock before writing the rejection
* Auto accepts a slightly less certain authorization when the call matches the request A write the owner asked for often scored just under the authorization cutoff (0.6-0.79) and asked. When the call also clearly matches the request (>= 0.8), an authorization of 0.75 or more now counts. Calls that must ask stay under it on one signal or the other across repeated runs. * No authorization counts for a call that plainly does not match the request A call continuing an approval for a different job (approved merging one PR, the next call merges another) reached the continuation cutoff in a rare run and ran. It scores about 0.2 on matching the request, while authorized calls score 0.71 or more, so require at least 0.5 when there is a request to match. * Apply the request-match floor to continuation only A plan the owner agreed to with a short reply need not match the request by itself, so the floor no longer applies to it or to a direct request. It guards the one signal it was measured for: continuing an approval.
…idance to its scope (#3350) * Auto questions cover batch items, plan items, checks, and guidance scope - Continuing an approval and agreeing to a plan: an identifier the model cannot read meaning into is taken as one of the items when the work covers a set, unless the call or the session shows otherwise. - Matching the request: a check on a step just taken, and looking up an identifier, label, or setting the request has to use, count. - Guidance limited to a place or kind of thing does not flag the same action elsewhere. - Granting more access than asked for is named as a stronger action. * Keep plan items to what the plan named An identifier the model cannot read is given the benefit of the doubt only when continuing a batch the owner already approved a call of. A plan's items must be ones the plan named or clearly included. * Do not assume an unreadable identifier is the next batch item Continuing an approved call again requires an item the session shows the request covers, as before this branch. The remaining wording changes are check-reads, guidance scope, and naming extra access as a stronger action.
* fix: keep human command-style prompts visible * fix: mark task prompt visibility explicitly * fix: mark Slack task prompts visible * fix: scope Slack prompt visibility to launch origin * fix: keep generated chat suggestion prompts hidden * fix: hide Discord automation task prompts --------- Co-authored-by: @mrubens <2600+mrubens@users.noreply.github.com>
Co-authored-by: Roomote <roomote@roomote.dev>
* feat: add configurable model fallbacks * fix: preserve model fallback recovery paths * fix: clear unsafe fallback state * fix: sanitize direct fallback notices * fix: isolate direct fallback notices * fix: distinguish fallback channel notices * test: stabilize worker test execution --------- Co-authored-by: Roomote <roomote@roomote.dev>
…3355) * Auto runs a call that only reads without asking whether it was requested A read changes nothing. Auto asks about actions that could be harmful and that the owner did not approve; what the agent then does with what it read is judged on the call that does it. A read still asks when it follows planted instructions, sends private data out, or the deployment's guidance names it, and a read of a task another session launched still has to match the request. * Trigger a new review run * Trigger a new review run * Trigger a new review run * Trigger a new review run
…oser (#3361) * Auto is turned on per session, under the composer, and starts off * Keep relaying permission asks after an instance refresh; align the Auto switch with the composer * Lowercase session in added prose * Choose the session's tool approvals mode from a composer chip * Describe what the nightly Auto toggle now adds * State current behavior directly in comments * Check model availability when a session starts in Auto; retry the pending-ask lookup after a refresh * Confirm the reopened event stream is live before catching up; stop the prompt when it cannot reconnect
Co-authored-by: @mrubens <2600+mrubens@users.noreply.github.com>
Contributor
* [Improve] Auto judges a task's calls with its session's context * Keep what a task read across harness prompts * A failed tool listing is not retried on every ask
| taskId: session.taskId, | ||
| ownerUserId: session.ownerUserId, | ||
| integrationId: input.integrationId, | ||
| toolName: input.toolName, |
Contributor
There was a problem hiding this comment.
This rejection lookup is only a snapshot before the model decision. If the owner rejects a same-tool call after this check but before the task's proxy consumes the resulting approved row, claimTaskIntegrationToolCall has no rejection guard and will execute this call anyway. The Fast path added an advisory-lock-backed claim guard for this exact race; apply the equivalent guard to model-approved task calls (or re-check atomically at proxy claim time).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Promote v1.15.1
Candidate refreshed to
bd80c46805e01edc239582e594693421d6f39c0afromdevelop.v1.15.1and triggers the existing GHCRv*image publish (latestchannel).release/v1.15.1branch can be deleted after this PR merges.Changelog
1.15.1 (2026-10-02)
Roomote 1.15.1 adds resilient model fallbacks, reduces Auto approval prompts, and fixes approval, sessions board, and task-detail edge cases.
Highlights
Patch changes