Repository navigation
Pub/Sub trigger: a redelivered message re-runs the agent in a new session, repeating tool side effects that already happened #7322
Description
Activity
@FurkhanShaikh. Reproduced on main with your script (500 then 200, two payments) and it's the same on every release since the trigger endpoints landed in v1.29.0.
Until #7325 is in a workaround that worked: a small ASGI middleware copies subscription:messageId into message.attributes, a before_agent_callback puts it in state and the tool passes it to the provider as an idempotency key. Two deliveries, one charge and a new messageId still gets charged. Could you try that in your harness especially the ack-deadline case?
@devansh173 #7325 looks good. One thing for the docstring: state is new on every delivery so the dedupe has to happen at the provider or in an external store not in state. Since messageId is only unique per topic maybe also suggest including the subscription in the key.
Please re-run the trigger tests and your E2E once more after that before asking for a review. The docs part (ack deadline, dead-letter policy) would be great as a follow-up in adk-docs.
Thanks for reproducing it and for the review. I pushed the changes to #7325:
- The
TRIGGER_DELIVERY_STATE_KEYdocstring now says session state is fresh on every delivery, so deduplication has to happen at the provider or in an external store. It recommends keying onsubscription+idfor Pub/Sub andevent_source+idfor Eventarc. - The redelivery unit test now keys on subscription + messageId in a ledger outside the session.
Re-ran everything: all 78 trigger tests pass. In the E2E, the idempotent tool pays twice for one messageId on
mainand once on this branch, and a new messageId is still charged. The logs are in the PR description.I'll follow up with an adk-docs PR covering the ack deadline, dead-letter policy and idempotent tools.
- The
0db8ecb reads well @devansh173. The docstring now says clearly where the dedupe has to happen and how to build the key. Re-checked: the 78 trigger tests pass and an idempotent tool keyed on trigger_delivery (subscription + id, ledger outside the session) pays once across the 500/200 redelivery while a new messageId is still charged. The same tool on main pays twice.
From my side the PR is ready for maintainer review. Will ask for the workflow approvals. Please link the adk-docs PR here once it's up.
@FurkhanShaikh whenever you get a chance it would be good to hear whether this (or the middleware workaround on current releases) holds up in your harness especially the ack-deadline case where the first run is still in flight.
Option 2 (deterministic session id) can go to a separate follow-up issue if maintainers want it so it doesn't hold this one up.
The adk-docs follow-up is up: google/adk-docs#2274. It covers the ack deadline, dead-letter topic and idempotent tools using
trigger_delivery, and asks to hold merging until #7325 is released.adk-docs#2274 reads well @devansh173 and matches #7325 (fields, key format and the "dedupe outside session state" point). Would take up your offer to split out the ack-deadline + dead-letter part. It applies to every release since v1.29.0 today while trigger_delivery waits on #7325 (still at 0db8ecb, not merged, main unchanged).
Two small additions while you're there. The ack deadline only helps if the Cloud Run request timeout is at least as long: the default is 300s so a 10-minute run gets cut at 5 min and redelivered anyway. Worth adding gcloud run services update --timeout=600 next to the ack-deadline step. For the dead-letter step the Pub/Sub service agent needs publisher on the dead-letter topic and subscriber on the subscription or nothing gets forwarded.
Might also be worth a line that for Eventarc the same settings go on the trigger's underlying Pub/Sub subscription. Please re-run mkdocs build --strict after the changes.
@FurkhanShaikh still keen to hear how the workaround or #7325 does in your harness especially the ack-deadline case.
- added a commit that references this issue
on Sep 28, 2026 Thanks! I split the ack-deadline and dead-letter part out as google/adk-docs#2275, which applies to current releases. While there I added the Cloud Run
--timeout=600step next to the ack deadline, the service agent's Publisher and Subscriber roles for the dead-letter topic, and a note for Eventarc Standard with a command to find the trigger's subscription. google/adk-docs#2274 now builds on it and only adds thetrigger_deliverypart, which waits on #7325.mkdocs build --strictpasses on both with no warnings.The split is exactly what I had in mind @devansh173. adk-docs#2275 covers the Cloud Run timeout, the service agent roles and the Eventarc subscription lookup and #2274 now only adds the trigger_delivery part on top of it.
One thing I noticed in the raw file (ecf96ea, also in b517386): the multi-line gcloud commands lost their trailing \ continuations so the dead-letter update, the two add-iam-policy-binding commands and the eventarc describe each render as one long line with runs of spaces. They still work when pasted which is probably why mkdocs --strict stays quiet but restoring the \ line breaks would match the rest of the page.
Once that's fixed #2275 is ready for docs review from my side and can go out independently. #2274 waits for #7325 (still 0db8ecb, not merged; main unchanged so the redelivery issue still applies from v1.29.0 through v2.10.0).
@FurkhanShaikh the offer still stands to try the workaround or #7325 in your harness whenever you get a chance.
- added a commit that references this issue
on Sep 28, 2026 Thanks for spotting this, @surajksharma07. The trailing
\continuations got lost when I split the section. Fixed in google/adk-docs#2275 (ba0c416): the dead-letter update, bothadd-iam-policy-bindingcommands and the Eventarcdescribeare split across lines again, matching the rest of the page. google/adk-docs#2274 is rebased on top of it. I re-ranmkdocs build --stricton both with no warnings, and checked the raw file: all 8 continuation lines end in\again.Checked ba0c416 @devansh173. All 8 continuations are back and no gcloud command is left on one long line in either #2275 or #2274. #2274 is cleanly on top of it and its trigger_delivery table and example still match #7325.
With that adk-docs#2275 is good to go from my side. Will ask for a docs reviewer and the workflow approvals so it can ship independently for current releases.
On the code side #7325 (0db8ecb) still merges cleanly onto today's main (e4c0d94) and the 78 trigger tests pass on the merged tree. Just waiting on a maintainer review now, nothing more needed from you there.
#2274 stays on hold until #7325 is released. Until then the redelivery behaviour applies to v1.29.0 through v2.10.0.
Reacted by Devansh BhargavaThanks both, and sorry for the delay.
I ran #7325 (
0db8ecb) and the middleware workaround (on 2.10.0)
through the harness, including the ack-deadline case.Short version: the key works on both paths, and it prevents the duplicate
when the dedupe is atomic at the effect.A provider idempotency key held in every case, including a real overlap.
One common reading of "check it against an external store" still pays twice.Setup.
- Stock
LlmAgenton the stockGeminiclass, pointed at a local fake Gemini
endpoint. - A local Pub/Sub push emulator. In the two ack-deadline rows the deadline was
scaled to 4 s (60 s elsewhere), and there was a 2 s wait before resending
after a nack. - A local payment provider with Stripe-style idempotency: the first request
with a key charges; a concurrent request with the same key gets 409; a later
one gets the stored reply. - Nothing in ADK was patched.
The tool builds
f"{subscription}:{id}"fromtrigger_delivery(#7325), or
from the attribute the middleware adds (workaround). It then dedupes one of
three ways:provider_key: passes the key as the provider's idempotency key;check_then_act: checks an external store, pays, then records the key;reserve: atomically inserts the key into the store before paying, and
keeps it on failure.
Charges, identical on #7325 and on the workaround:
Scenario no dedupe provider_keycheck_then_actreservemodel 503 after the payment (500 → 200) 2 1 1 1 provider commits, reply exceeds the tool timeout 2 1 2 1 ack deadline passes while payment 1 is in flight, redelivery on a 2nd instance 2 1 2 1 same, one instance with an asynctool2 1 2 1 ack deadline passes before the first run reaches the tool 2 1 1 1 provider fails before committing (503) 1 1 1 0 two different messages 2 2 2 2 In the overlap cases with
provider_key, the provider sawcharged, then
409for each redelivery that arrived while payment 1 was in flight. The tool
raised on the 409, so the run failed and the message was nacked; a later
delivery got the original reply replayed, and was acked with one charge.Things that might be worth a line in the #7325 docstring or adk-docs#2274:
- Prefer the provider's idempotency key. It was the only strategy that
held in every row, because the provider is the one place that knows
whether the effect happened. - "Check an external store" needs to be atomic, and to happen before the
effect. Check → act → record pays twice after an ambiguous failure (a
timeout after the provider committed), and again when a redelivery
overlaps a run still in flight. - A reservation taken before acting avoids the duplicate, but turns a
failure before commit into a skipped effect. In the last-but-one row the
message was acked and nothing was paid. So a store-based dedupe needs a way
to reconcile an ambiguous attempt, not just a flag. - Let a 409 (key in use) fail the run rather than returning it to the
model as a tool result. That is what our tool did; we didn't test the
alternative. The reasoning: the nack keeps the message alive until a
delivery sees the original result, whereas acking on a 409 would lose the
payment if the in-flight attempt then failed before committing.
One side observation on the ack-deadline case. With a synchronous tool,
the trigger path (runner.run_async) calls the tool inline on the event loop
(FunctionTool._invoke_callable, 2.10.0 lines 457–460).
RunConfig.tool_thread_pool_configdoesn't change this: it is only read in
the live path (_process_function_live_helper). While the first payment's
reply was held, that instance could not start the redelivered run at all:- the redeliveries queued;
- further deadlines expired;
- they then ran back to back (3 charges with no dedupe);
- so on one instance with a sync tool, the runs never overlapped in our tests.
It does happen across two instances, or with an async tool, which is why
those two rows are there. Probably a separate topic, but it seemed relevant
to the ack-deadline advice in adk-docs#2275: one slow sync tool call can stall
every trigger request on that instance.Limits:
- The provider is a local model of the documented idempotency behaviour, not
a real payment API. A provider whose keys don't cover concurrent in-flight
requests would behave likecheck_then_actin the overlap rows. - Pub/Sub is emulated, with scaled timings.
- Each message asked for one payment, and our fake model sent the same
arguments on every delivery, so the key wassubscription:idalone.
adk-docs#2274's example also appendsinvoice_id, which a message with
several payments needs. We didn't test a model that changes that argument
between deliveries; if it did, the key would change with it.
- Stock
Really useful matrix @FurkhanShaikh. Good to see the key behaves the same on #7325 and the middleware and that provider_key was the only one that held in every row, overlap cases included.
@devansh173 could you fold that into the TRIGGER_DELIVERY_STATE_KEY docstring and the adk-docs#2274 example? Prefer the provider's idempotency key. If you use a store reserve the key atomically before the effect and plan for reconciling an ambiguous attempt since check → act → record pays twice. And raise on a "key in use" 409 so the message is nacked rather than handing it back to the model. Worth noting the PR's E2E ledger is check-and-add in memory so it shouldn't read as the recommended pattern.
On the sync tool point: in 2.10.0 non-live runs do honor tool_thread_pool_config (_caller.py:924-950) but the trigger route calls runner.run_async() without a run_config (trigger_routes.py:382) so a sync tool runs on the event loop and blocks that instance. Until that's configurable an async tool (or asyncio.to_thread around the blocking call) avoids it. Opening a separate issue for that so it doesn't hold this one up.
#7325 (0db8ecb) is still unmerged and 2.10.0 is unchanged so it's still with the maintainers for review.
Thanks @surajksharma07 and @FurkhanShaikh. I've folded this in:
- fix(cli): expose trigger delivery id so tools can detect redeliveries #7325 (
c937675): theTRIGGER_DELIVERY_STATE_KEYdocstring now prefers the provider's idempotency key. It says a store has to reserve the key atomically before the effect, with a plan for reconciling an ambiguous attempt, since check → act → record pays twice. It also says to raise on a key-in-use 409 so the message is nacked. The PR description now notes that the E2EDONEledger is an in-memory check-and-add stand-in, not the recommended pattern. The change is docstring only, and the 78 trigger tests pass. - docs(runtime): show deduplicating trigger redeliveries by delivery ID adk-docs#2274 (
2114346): the example says the same, and lets a 409 propagate instead of returning it to the model.mkdocs build --strictpasses.
google/adk-docs#2275 is unchanged.
- fix(cli): expose trigger delivery id so tools can detect redeliveries #7325 (
Checked c937675 and adk-docs#2274 (2114346) @devansh173. The docstring and the example now say the same thing as FurkhanShaikh's matrix: provider key first, atomic reservation if it's a store and a 409 is raised rather than handed back to the model. The change is docstring only so behaviour is the same as 0db8ecb.
One thing to sort before review: main has moved under #7325. 14f4154 (max_llm_calls, in 2.11.0) added test_eventarc_passes_max_llm_calls at the exact spot where TestTriggerDeliveryIdentity goes so test_trigger_routes.py no longer applies cleanly. The source side still applies and after resolving that by hand I get 86 passed on the merged tree. Could you rebase onto current main and re-run the trigger tests?
On versions: 2.11.0 still creates a uuid4() session per delivery and doesn't pass messageId through so the redelivery behaviour still applies from v1.29.0 through v2.11.0. The trigger route now passes a RunConfig but it only sets max_llm_calls so the sync-tool-on-the-event-loop point from before still holds.
After the rebase it's back with the maintainers. adk-docs#2275 can still go out on its own and #2274 stays on hold until #7325 ships.
- added a commit that references this issue
on Oct 5, 2026 Thanks @surajksharma07. I rebased #7325 onto the current main (
ab9fffc), and it's now at54e1f19. All three commits applied without conflicts:test_eventarc_passes_max_llm_callsstays inTestTriggerEventarc, withTestTriggerDeliveryIdentityafter it, and_run_agentpasses both the newRunConfigand the session state. The trigger tests give 86 passed on the rebased branch, and the E2E repro still charges once across the 500 → 200 redelivery on 2.11.0.Checked 54e1f19 @devansh173. The rebase is clean on ab9fffc and main has moved 23 commits since but none of them touch trigger_routes.py or its tests so it still merges without conflicts. On the merged tree I get 86 passed including the 7 TestTriggerDeliveryIdentity cases.
Also ran the repro from this issue against it. A tool that ignores trigger_delivery still charges twice across 500 → 200 which is expected since the PR exposes the id and the dedup stays in the tool as the docstring now says. The middleware workaround from earlier still passes on the same tree so anyone on v1.29.0 through 2.11.0 can keep using it until this ships.
Nothing else pending from my side. #7325 is ready for maintainer review (the 2 workflows still need approval) and adk-docs#2274 stays on hold until it's released.
- added a commit that references this issue
on Oct 8, 2026 - addedtools[Component] This issue is related to tools[Component] This issue is related to toolscli[Component] This issue is related to cli[Component] This issue is related to cli
on Oct 9, 2026
🔴 Required Information
Describe the Bug:
Suppose a Pub/Sub trigger run fails after a tool has already caused an external
side effect, such as a payment, an email or a ticket.
the first run, so the side effect happens again.
Nothing inside the agent can prevent this. Pub/Sub keeps the
messageIdthesame across redeliveries ("A redelivered message retains the same message ID
between redelivery attempts"). ADK logs it but never passes it to the agent or
its tools, so a tool has nothing to deduplicate on.
In
src/google/adk/cli/trigger_routes.py(v2.10.0; byte-identical in 2.9.2):
session_id = str(uuid.uuid4()), once per HTTP delivery.the agent receives
{"data": ..., "attributes": ...}. ThemessageIdisnot included.
the
messageIdis only logged.HTTP 500 for a
TransientErrorafter retries, and for any other exception.the endpoint's own description: "errors trigger Pub/Sub retry".
The ambient agents docs describe the
design: "Each redelivery creates a new session. Trigger workloads are stateless
by design." They don't say what this means for tools with side effects.
The in-process 429 retry is fine: it reuses the session, so the model can see
the tool call that already completed. The problem is the hand-off to Pub/Sub.
Steps to Reproduce:
pip install google-adk==2.10.0repro.py.python repro.py.It needs no network access and no API key.
LlmAgenton ADK'sGeminiclass, withbase_urlpointed at a local fake Gemini endpoint.pay_invoicecall.It answers the request after the tool with HTTP 503 UNAVAILABLE ("The model
is overloaded") once, then behaves normally.
the same push envelope again, with the same
messageId.Expected Behavior:
One payment for one message. Failing that, a way for the agent or its tools to
recognise a redelivery, so they can skip a side effect that already happened.
Observed Behavior:
For delivery 1, ADK logs
Error processing Pub/Sub message: 503 UNAVAILABLE. {'error': {'code': 503, 'status': 'UNAVAILABLE', 'message': 'The model is overloaded. Please try again later.'}}.Environment Details:
trigger_routes.pyis byte-identical.Model Information:
gemini-2.5-flashthrough ADK'sGeminiclass,answered by a local fake endpoint. No real model was called.
🟡 Optional Information
Regression:
No. The same code is in the commit that added the trigger endpoints
(e2d970f).
Additional Context:
The same duplicate has other routes. They were measured with the same kind of
local harness: a fake Gemini endpoint, and Pub/Sub emulated from its documented
redelivery rule.
Two rows deserve a note:
subscription's acknowledgement deadline defaults to 10 seconds, and "if you
send a negative acknowledgment or the acknowledgment deadline expires,
Pub/Sub resends the message". The docs give "10 minutes (ack deadline)" as
the maximum processing time, but don't mention that a subscription created
with defaults gets 10 seconds.
repeats after the side effect repeats the side effect on every redelivery,
until the message expires. The docs recommend a dead-letter queue but don't
say why it matters here.
Two related observations:
Eventarc (Python): direct CloudEvents do put
ce-idin the agent'sinput (L617-627); Pub/Sub-wrapped events don't.
google/adk-goappears to share the design. I read its source but didnot run it:
RetriableRunner.RunAgentcreates one session per delivery (its comment:"One session per delivery").
messageContentFromPubSubleaves out the message ID.I'm happy to open a companion issue there.
Possible changes, roughly by cost:
messageId(and
ce-id) in session state or in the attributes the agent receives, soa tool can derive an idempotency key for its provider. This is small, and
lets applications fix the problem themselves.
session_idfrom the delivery (e.g.uuid5of subscription andmessageId) instead ofuuid4().calls, as the in-process retry already does.
final response, which covers the acknowledgement-deadline case.
a guard.
happened. Recommend an acknowledgement deadline above the expected run time
(up to 600 s), a dead-letter policy, and idempotent tools. Show how a
publisher can carry an idempotency key in the message attributes, which do
reach the agent today.
I'm happy to be told this is intended. In that case, (3) alone would still help.
Found while auditing retry behaviour across agent frameworks.
Minimal Reproduction Code:
How often has this issue occurred?:
Every time, under the conditions above. The reproduction is deterministic.