Repository navigation
[Improve] Auto runs calls the owner asked for or already approved - #3335
Merged
Merged
Conversation
Contributor
|
No new code issues found. See task
Reviewed 5e94589 |
…each the authorization question
This was referenced Sep 30, 2026
7 of 8 tasks
The recent outcomes the model sees are capped, so a rejection could drop out after a few newer decisions and let an authorization path approve the rejected tool. Look the rejection up session-wide and treat a lookup failure as a rejection.
A call assessed while the owner rejected another call to the same tool could still auto-run on the stale lookup. Check again before reserving the auto-approval and ask instead.
The recheck before reserving still left a window before the claim. The claim now fails, and its reservation is cancelled, when the requester has rejected a call to this tool in the session, so a rejection that commits before the call runs is never missed.
A rejection and a guarded Auto claim of the same session tool now take the same transaction lock, so the claim's check always sees a rejection that committed first.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Related issue
Internal follow-up to #3195; no separate issue is linked.
Why this PR exists
What changed
Auto used to run only read-only calls on its own. Anything that changed something asked the owner, even when they had just asked for exactly that call ("delete old-draft.docx") or had approved the first item of a batch. Earlier approvals were deliberately treated as history, never as consent.
Auto now also runs a call when the owner authorized it:
One new question (
userAuthorized) covers both. Earlier decisions now carry the redacted arguments the owner saw on the card, so the model can tell "the next draft" from "main". A different target, a wider scope, a stronger action (sending instead of drafting), a withdrawn request, a rejected call of the same kind, and anything that comes from content the agent read still ask.Authorization never overrides the existing safety checks. The call still asks if it carries out a planted instruction, sends private data out, or is flagged by the deployment's guidance. Calls that move money always ask, because amounts can't be checked reliably. A second question (
movesMoney) decides that from the call itself, whatever the conversation says about the payment.How it was tested
Variants were compared with real Jev against labeled cases, with labels fixed before any scores were seen:
Known follow-up, unchanged by this PR: the deployment-guidance question over-flags some calls. For example, "deletes need a person" also flags closing an issue.
Checklist
[Fix],[Feat],[Improve],[Refactor],[Docs], or[Chore]followed by a user-facing descriptionpnpm lintandpnpm check-typespass locallypnpm changeset(not needed: Auto is a nightly-only experiment)Combined local smoke (2026-09-30)
All six Auto/approval PRs (#3332, #3334, #3335, #3337, #3338, #3339) were merged together locally and run on a local stack. It used a real web session with Auto on, Jev through OpenRouter, and a local MCP server whose tools only log (
list_files,delete_file,pay_invoice). The PRs merge cleanly into develop in any order, except #3332 and #3337: both add an export next to each other inpackages/types, and whichever lands second needs a trivial rebase. The combined build passes uncachedpnpm check-types,pnpm lint, and the approval suites (cloud-agents 131, sdk 23, worker 13, db 32).delete_fileauto-approved (authorization 0.97, highest risk level). Alist_fileswith no arguments now shows a card instead of failing the insert.list_fileswas assessed by Auto (auto-approved row).Not covered live: chat surfaces (Slack, Telegram, Discord) and the task (sandbox) path, because there's no local worker. Both are covered by unit tests.
Review follow-up: approvals with no retained messages
An earlier approval now reaches the authorization question even when no human messages remain in context. With real Jev this is safe (0/50 unauthorized calls ran), but it rarely helps: without the request text, Jev scores "continues an approved call" at 0.29 to 0.59, so those calls still ask. In normal sessions the recent messages are kept, and the batch case works (S7).
Update: continuation, "go ahead", and routine reads (2026-09-30)
Follow-up work on the same goal: a risky call runs when the owner's intent is clear, and asks when it isn't. Approving one item, or saying "yes, go ahead" to a plan, shouldn't mean clicking through every item.
What changed since the first version
continuesApprovedCallasks whether this call repeats an approved call for the next item of the same work, with every argument except the item unchanged. It is asked only when code finds an owner approval of this tool in the session and no rejection of it.agentMessageRepliedTo, the agent's visible replies since the owner's previous message). A separateagreedToPlanquestion asks whether the owner agreed to a plan and this call is one of its actions. It is asked only when there is such a message.How it was chosen: a new labeled set of 32 development and 16 held-out cases (written before scoring), covering continuation, direct requests, "go ahead", near-misses and unsafe calls. Six approaches were compared on the development set, and the pick was confirmed on the held-out set with real Jev.
The old suites are unchanged: broad and held-out suites (only writes the user asked for now run, plus one known pre-existing read miss), injection lab 0/80 unsafe, session cases 0/140 wrong.
Rejected along the way
Live on a local stack (develop + this PR):
Known gaps
Final live check (with the model-down pause merged)
Run on the local stack against develop + this PR, with Auto on, a mock file-store MCP server, and the decision model via OpenRouter:
agreedToPlan0.90–0.92)userAuthorizedwas 0.95, because the rejection stays in forceRemaining known gap (not new here): a verification re-list right after a delete sometimes scores
matchesRequestbelow 0.8 and asks.