fix(knowledge): give document processing per-tenant queue lanes - #7937
Merged
Merged
Conversation
Document processing ran through one Trigger.dev queue with a global concurrency limit and no concurrency key, so a single connector backfill could hold every slot and leave every other tenant's uploads queued behind it for hours. Dispatch now names a lane — interactive for work a person is waiting on, backfill for connector-driven ingestion — and keys each run by the entity that owns the knowledge base, so the limit applies per tenant copy of the queue rather than fleet-wide. - Key on the owning entity, not the workspace: an organization-scoped knowledge base carries a null workspace id and would otherwise share one bucket with every other such tenant. Owned scopes reuse resourceScopeKey so tenant identity keeps one spelling. - Stamp the lane on the payload so quota and capacity continuations resume in the lane they were admitted against instead of promoting themselves out of the backfill ceiling. - Parse an absent or unrecognized lane as backfill rather than throwing, so payloads written before the lanes existed and payloads stamped by a newer version mid-rollout cannot burn a run's retry budget. - Keep interactive work on the pre-existing queue name; a queue the running worker has not registered parks its runs in PENDING_VERSION, and the app deploys separately from the worker. - Stop absorbing a rejected backfill chunk in-process: those documents are reported failed and reclaimed by the stuck-document sweep instead of outlasting the connector lease they run under.
|
The latest updates on your projects. Learn more about Vercel for GitHub. |
Contributor
|
This branch was previously deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
interactivefor work a person is waiting on,backfillfor connector-driven ingestion — and sets aconcurrencyKeyfor the entity that owns the knowledge base, so the limit applies per tenant copy of the queue instead of fleet-wide.resourceScopeKeyso tenant identity keeps one spelling.backfillinstead of throwing, so payloads written before the lanes existed — and payloads stamped by a newer version mid-rollout — cannot burn a run's retry budget.Throughput
Both lanes default to the limit the single shared queue carried, so nothing drains slower than before:
KB_CONFIG_BACKFILL_CONCURRENCY_LIMITis a separate variable fromKB_CONFIG_CONCURRENCY_LIMITso backfill can be dialled down without slowing person-facing work. Worth noting for review: the fleet-wide ceiling is now the Trigger.dev environment concurrency limit rather than this queue's own limit, so aggregate ingestion can rise above what it could reach before. Lowering the backfill variable is the lever if that budget gets tight.Deploy note
Interactive work deliberately stays on the existing queue name. A run naming a queue the running Trigger.dev worker has not registered parks in
PENDING_VERSION, and the app deploys separately from the worker — so only the new backfill queue can be caught by that window, and stranded backfill work is exactly what the stuck-document sweep already recovers.Type of Change
Testing
apps/simsuite: 53,005 passed, 1 pre-existing failure that needsDATABASE_URLbun run type-checkclean,bun run lintclean, all 46 audits pass, docs manifest in syncChecklist