Skip to content

Commit bad0ce4

Browse files
authored
chore(db): tune document autovacuum for connector sync churn (#7962)
* chore(db): tune document autovacuum for connector sync churn Connector syncs update this table in bursts that outrun autovacuum at the default 20% scale factor, which on a table of this size only triggers once dead rows are already substantial. Sustained bloat stops index-only scans trusting the visibility map, so the knowledge base listing's token aggregate degrades into per-row heap fetches and can exceed its statement timeout. The factors are less aggressive than 0357's on purpose. That one is a small queue table, while this is one of the largest here and shares a small autovacuum worker pool with many other large tables, so triggering too eagerly would risk holding a worker continuously and starving its peers. Cost limit and cost delay are left unset so the table stays inside cross-worker I/O balancing. * chore(docs): keep production scale out of published text Comments describing measured behaviour carried absolute production sizing — index byte sizes, chunk and row totals — which the repo's publishing rules exclude and which a public reader does not need. The measurements that justify each decision stay; only the figures that size production are replaced. The ship scrub missed these because its grep matched identities and IDs but never numbers, and because code and migration comments do not feel like publishing even though they are. It now states the distinction explicitly and greps the diff for byte sizes, k/M-scale entity counts, and seven-figure totals, while deliberately still allowing ordinary engineering numbers.
1 parent fe1c91f commit bad0ce4

6 files changed

Lines changed: 27802 additions & 12 deletions

File tree

‎.agents/skills/ship/SKILL.md‎

Lines changed: 8 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -112,12 +112,19 @@ The repo is public. **Everything you publish — title, description, commit mess
112112

113113
Describe the bug by its mechanism, not by how you found it. "Expired OAuth credentials fail to refresh in the worker" — not "the Sheets canary failed at 16:31Z for workspace abc-123". Aggregate counts are fine once detached from the tenant ("1,379 PDFs failed"); the same number attributed to a named customer is not. Replace real examples with placeholders (`<real sheet name>`) rather than cutting them — the illustration is usually the useful part.
114114

115-
**Scrub before publishing, not after** — a leak is public the instant it posts, and editing later does not unsend the notification email. This applies to every PR you open, including ones created directly with `gh pr create` rather than through this skill. Grep the title, body, and `git log origin/staging..HEAD` before publishing:
115+
**Measurements are not the problem; absolute production scale is.** Keep the numbers that justify a change — durations, ratios, before/after timings, test and audit counts. They are the evidence a reviewer needs, and stripping them makes the rationale unfalsifiable. What does not belong is anything that sizes production or a tenant: table and index byte sizes, row/chunk/document totals, dead-tuple counts, buffer and heap-fetch counts, worker or instance counts. "Visiting four times as many tuples took 5.1s and 9.7s on consecutive runs" is fine; "on a 132k-chunk index" or "reclaims ~19 GB" is not. The same rule applies to code comments and migration comments, which are published exactly like a PR body — this is the most commonly missed case, because they do not feel like publishing.
116+
117+
**Scrub before publishing, not after** — a leak is public the instant it posts, and editing later does not unsend the notification email. This applies to every PR you open, including ones created directly with `gh pr create` rather than through this skill. Grep the title, body, `git log origin/staging..HEAD`, AND the diff itself before publishing:
116118

117119
```bash
120+
# identities, IDs, infrastructure
118121
grep -niE 'customer-or-company-name|@[a-z0-9.-]+\.(com|io|ai)|[0-9a-f]{8}-[0-9a-f]{4}-|\.sharepoint\.com|arn:aws|https?://[a-z0-9.-]*\.internal'
122+
# absolute production scale — byte sizes, k/M-scale entity counts, 7-figure totals
123+
grep -niE '[0-9][0-9.,]* ?(TB|GB)\b|[0-9]+(\.[0-9]+)?[kKmM][- ](row|chunk|document|vector|tuple|doc)|[0-9]{1,3}(,[0-9]{3}){2,}'
119124
```
120125

126+
The second pattern deliberately allows ordinary engineering numbers (`5.1s`, `46 audits`, `2,921 tests`) and flags only production sizing.
127+
121128
## PR Description Format
122129

123130
Use this exact template in the user's voice (concise, bullet points):

‎apps/sim/lib/knowledge/search/queries.ts‎

Lines changed: 9 additions & 8 deletions
Original file line numberDiff line numberDiff line change
@@ -53,9 +53,10 @@ const MAX_AUTHORIZED_SEARCH_CANDIDATES = 20_000
5353
* Bounds a permission-starved graph walk, which returns fewer candidates rather than widening.
5454
* This approximate iterative-visit threshold excludes pgvector's initial scan; it is not a row limit.
5555
*
56-
* Raising it trades recall for latency far more steeply than its size suggests: on a 132k-chunk
57-
* index at 10% visibility, visiting 6.5k tuples instead of 1.5k took 5.1s and 9.7s on consecutive
58-
* identical runs, against ~115ms for the bounded walk. Re-measure before changing it.
56+
* Raising it trades recall for latency far more steeply than its size suggests. Measured on a
57+
* search index where visibility admitted a tenth of the corpus, visiting four times as many tuples
58+
* took 5.1s and 9.7s on consecutive identical runs, against ~115ms for the bounded walk — the walk
59+
* degrades superlinearly with depth, and unpredictably. Re-measure before changing it.
5960
*/
6061
const CANDIDATE_HNSW_MAX_SCAN_TUPLES = '1000'
6162
const CANDIDATE_HNSW_EF_SEARCH = '1000'
@@ -911,11 +912,11 @@ async function selectVectorResults(params: SearchParams): Promise<SearchResult[]
911912
*
912913
* An underfilled traversal yields fewer candidates rather than widening the search. Widening
913914
* it has no affordable form here: rescoring the projection exhaustively is O(corpus) and a
914-
* deeper `hnsw.max_scan_tuples` is worse still — measured on a 132k-chunk index at 10%
915-
* visibility, the exhaustive rescan took 1.9s while scanning 6.5k tuples instead of 1.5k took
916-
* 5.1s and 9.7s on consecutive identical runs. Both exceed the retrieval budget on a corpus
917-
* an order of magnitude larger, and a leg that exceeds its budget returns nothing at all, so
918-
* fewer candidates strictly beats every widening strategy available.
915+
* deeper `hnsw.max_scan_tuples` is worse still. Measured where visibility admitted a tenth of
916+
* the corpus, the exhaustive rescan took 1.9s while visiting four times as many tuples took
917+
* 5.1s and 9.7s on consecutive identical runs. Both exceed the retrieval budget once the
918+
* corpus grows, and a leg that exceeds its budget returns nothing at all, so fewer candidates
919+
* strictly beats every widening strategy available.
919920
*/
920921
const identities = await withVectorScanSettings(
921922
(executor) =>

‎apps/sim/lib/table/planner.ts‎

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -72,9 +72,9 @@ export async function withReadGuards<T>(
7272
* (`->>` extraction, `@>` containment, lateral `jsonb_each_text`) are opaque to
7373
* the planner — it estimates a handful of matching rows and picks a parallel
7474
* seq scan over the entire shared `user_table_rows` relation (every tenant's
75-
* rows) instead of the tenant's own index. Measured on a 1M-row table inside a
76-
* 12M-row relation: filtered count 12.7s → 1.0s, sorted page 9.7s → 0.76s,
77-
* filtered bulk select 14.4s → tenant-bounded. The flag only penalizes the plan
75+
* rows) instead of the tenant's own index. Measured on a large tenant inside a
76+
* far larger shared relation: filtered count 12.7s → 1.0s, sorted page
77+
* 9.7s → 0.76s, filtered bulk select 14.4s → tenant-bounded. The flag only penalizes the plan
7878
* shape: if no index plan exists, the seq scan still runs (and the timeout caps it).
7979
*/
8080
export async function withSeqscanOff<T>(fn: (trx: DbTransaction) => Promise<T>): Promise<T> {
Lines changed: 12 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,12 @@
1+
-- migration-safe: table-local maintenance settings only; no row rewrite or change to old readers and writers.
2+
-- Connector syncs update this table in bursts large enough to outrun autovacuum at the default 20%
3+
-- scale factor, which on a table of this size only triggers once dead rows are already substantial.
4+
-- Sustained bloat stops index-only scans from trusting the visibility map, so the knowledge base
5+
-- listing's token aggregate degrades into per-row heap fetches and can exceed its statement timeout.
6+
--
7+
-- These factors are deliberately less aggressive than 0357's. That one is a small queue table, while
8+
-- this is one of the largest tables here and shares a small autovacuum worker pool with many other
9+
-- large tables, so triggering too eagerly would risk holding a worker continuously and starving its
10+
-- peers. Cost limit and cost delay are deliberately left unset: PostgreSQL excludes tables carrying
11+
-- either from cross-worker I/O balancing, which is not a trade worth making on a table this busy.
12+
ALTER TABLE "document" SET (autovacuum_vacuum_scale_factor = 0.05, autovacuum_analyze_scale_factor = 0.02);

0 commit comments

Comments
 (0)