Skip to content

perf(db): pick the snapshot list page before reading wide columns - #3670

Closed
jakubno wants to merge 1 commit into
e2b-dev:mainfrom
jakubno:perf/snapshot-list-page-before-laterals
Closed

jakubno wants to merge 1 commit into
e2b-dev:mainfrom
jakubno:perf/snapshot-list-page-before-laterals

Conversation

@jakubno

@jakubno jakubno commented Oct 1, 2026

Copy link
Copy Markdown
Member

Description

In production, GetSnapshotsWithCursor (the paused half of GET /sandboxes) accounts for essentially all of the database's temp-file writes. It launches about 1.8 parallel workers per call, and its slowest calls take over a minute. The temp-file log shows the query's leader and worker each spilling 80–115 MB on a parallel bitmap scan.

In that plan Postgres sorts every matching snapshot of the team, with all columns including metadata and config, and only then runs the build and alias lookups and applies LIMIT. A LIMIT can't bound a sort through the lateral joins above it, so the sort holds the whole team.

This PR splits each of the four cursor queries in two:

  • A page subquery picks the rows. It has the same filters, joins and order as before, but returns only id, the sort keys and the build columns. The ready-build lookup stays in this subquery because it drops rows and has to run before LIMIT. Moving it after LIMIT would let snapshots without a ready build take page slots.
  • The outer query reads the full snapshot row by primary key, and the aliases, for the page's rows only. It orders by the page's columns so the planner can keep the subquery's order instead of sorting full rows again.

The generated row structs, parameters and scan order are unchanged, so callers don't change.

Plans

These come from a Postgres 18 test container with one team holding 100k snapshots (~1 KB of config each) and work_mem at 4 MB.

Default settings: both queries walk idx_snapshots_team_time_id in order, and nothing is sorted. This change doesn't affect that plan (1–1.6 ms at LIMIT 101).

Production's plan shape, forced with enable_indexscan = off (parallel scan → Sort → Gather Merge):

before after
LIMIT 101 sorts spill 46.7 + 38.0 + 36.5 MB, 110 ms 3.6 MB spilled (workers sort in memory), 35 ms
LIMIT 5101 sorts spill 40.1 + 40.5 + 40.5 MB, 194 ms 3.2 + 3.1 MB spilled, 176 ms *

* There's also a 6.7 MB re-sort of the page at the top. That's an artifact of disabling index scans: the primary-key lookup became a hash join. With index scans available it's a nested loop that keeps the page order, as in the LIMIT 101 plan.

I couldn't reproduce production's planner statistics locally, so the "after" numbers are for the forced plan shape. After deploy, temp_blks_written in pg_stat_statements for the new query, and the temporary file log lines, will show whether it holds in production.

Tests

  • TestSnapshotCursorQueriesShareOneProjection now compares the shared text around each query's WHERE and ORDER BY. It also checks that each query's outer ORDER BY matches its page ORDER BY.
  • New TestGetSnapshotsWithCursor_OrdersNewestFirstAndPaginates covers newest-first order, keyset pages with no gaps or overlaps, and a snapshot without a ready build not taking a page slot.

Not in this PR: the handler still adds one to the LIMIT for each running sandbox it excludes (sandboxes_list.go), which keeps the page large for teams with many running sandboxes.

The snapshot cursor queries sorted every matching snapshot of a team with all
of its columns, metadata and config included, before LIMIT could apply: a
LIMIT cannot bound a sort through the lateral joins above it. For large teams
the planner's parallel bitmap plan spilled that sort to disk on every call.

Each query now picks the page in a subquery that returns only the id, the sort
keys and the build columns, then reads the full row and the aliases for the
page alone. The ready-build lookup stays in the subquery because it drops rows
and must run before LIMIT. Generated row types and parameters are unchanged.
@cursor

cursor Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

PR Summary

Medium Risk
Touches the production-hot path behind paused sandbox listing; wrong pagination or ready-build filtering would be user-visible, though semantics are pinned by tests and the public query API is unchanged.

Overview
Restructures the four snapshot keyset-list queries so Postgres applies LIMIT on a narrow page (ids, sort keys, and build fields) before loading full snapshot rows (metadata, config) and alias aggregates. The ready-build lateral join stays inside that inner page so snapshots still in snapshotting are excluded before LIMIT, which fixes cases where they could otherwise consume page slots. Call-facing sqlc types and scan shape are unchanged; behavior is covered by an extended projection guard (outer ORDER BY must match the page subquery) and a new descending pagination test.

Reviewed by Cursor Bugbot for commit 571b78f. Bugbot is set up for automated code reviews on this repo. Configure here.

@jakubno

jakubno commented Oct 1, 2026

Copy link
Copy Markdown
Member Author

Opened in the wrong repository; moving to belt.

@jakubno jakubno closed this Oct 1, 2026
@jakubno
jakubno deleted the perf/snapshot-list-page-before-laterals branch October 1, 2026 14:14
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant