Skip to content

feat(sweep): ✨ out-of-core SRC sweep on one GPU - #42

Merged
Panadestein merged 19 commits into
mainfrom
feat/src-out-of-core
Oct 7, 2026
Merged

Panadestein merged 19 commits into
mainfrom
feat/src-out-of-core

Conversation

@robertodr

Copy link
Copy Markdown
Member

🤖 AI text below 🤖

Phase 1 of running SRC on stacks too large for one GPU, such as N . V . M . U with a large bond in M. The sweep now reads the input cores one site at a time, batches every contraction to a memory budget, and keeps the sketched environments on the GPU, in host memory or on node-local disk. The mathematics and the random draws are unchanged, so a seed gives the same operator as before, up to rounding.

Stacked on #41.

What changes

  • New Resources argument on src, apply and compress: GPU and host memory budgets and a scratch directory. Budgets left unset are detected: free device memory minus max(10%, 1 GiB), and MemAvailable minus 10%.
  • Lazy inputs: a train can be any sequence of array-likes with shape, dtype and np.asarray, e.g. np.memmap or zarr/HDF5 datasets. Each core is read when the sweep reaches its site, and a background thread prefetches the next one.
  • New private modules:
    • _plan.py: a pure planner that picks per-site batch sizes and environment tiers, newest environments on the fastest tier.
    • _kernels.py: the batched contractions, plus a peak-memory estimate from the opt_einsum path.
    • _store.py: environment storage, with pinned staging and writer/reader threads.
    • _sites.py: the site source.
  • _sweep.py becomes a driver over these modules. It logs the plan and the time spent waiting on data.
  • CuPy's memory pool is capped at the GPU budget during the sweep, so an underestimate fails at once. The previous limit is restored afterwards.

Design and plan: docs/superpowers/specs/2026-09-25-src-out-of-core-design.md and docs/superpowers/plans/2026-09-25-src-out-of-core.md. The spec lists five refinements made while prototyping.

Testing

  • uv run prek run --all-files and uv run pytest pass on the CPU: 153 tests, including the benchmarks. That covers batching, disk spilling, lazy and memmap inputs, bra stacks of lazy sites, and error and cleanup paths. Small problems still plan to one batch on the device, and the benchmarks run as fast as before.
  • Not yet run on a GPU. The GPU-only paths (streams and events, pinned staging, ndarray.get(blocking=False), the pool cap) and the new tests in tests/test_gpu_backend.py have not been run anywhere yet, since the development machine has no CUDA device.
  • Cluster acceptance is pending. benches/large/ targets 50 sites, D_M = 4000, chi_out = 2000, complex128 on one A100-40GB, with at least 300 GB of host memory and local NVMe. Its README lists the four pass criteria; the results will be added there.

@robertodr
robertodr added this pull request to stack #43 September 25, 2026 13:59
@Panadestein
Panadestein force-pushed the feat/src-out-of-core branch from f4de686 to a8a50be Compare October 5, 2026 14:48
Base automatically changed from feat/src-stack to main October 5, 2026 15:12
Phase 1 of the multi-GPU plan: stream the input cores, batch the
contractions to a memory budget and keep the environments on the GPU,
in host memory or on local disk. Targets N.V.M.U with D_M = 4000 and
chi_out = 2000 in complex128 on one A100-40GB.

Assisted-by: Pi:claude-opus-5-5
Eleven tasks from the backend helpers to the cluster acceptance run,
each test-first. The code was prototyped and every intermediate state
passes the CPU suite; five refinements of the spec are listed up front.

Assisted-by: Pi:claude-opus-5-5
Host-to-device copies on the compute stream, batch sizes by binary
search over the walked path, the promoted working dtype, lazy sites in
bra stacks, and tests on dense operators rather than cores.

Assisted-by: Pi:claude-opus-5-5
@Panadestein
Panadestein force-pushed the feat/src-out-of-core branch from a8a50be to 2bfa815 Compare October 5, 2026 15:13
@robertodr
robertodr marked this pull request as ready for review October 6, 2026 09:08
Resolve the conflicts with stdlib logging (#52), input-precision
sketches (#36), validation (#38), ty enforcement (#37, #55) and the
Fumadocs site (#56):

- the sweep draws its sketches with gaussian_sketch and logs its plan,
  pass times and stalls at DEBUG with %-style arguments;
- the working dtype is promoted over every site, not just the first;
- docs/large-problems.md moves to features/large-problems.mdx and the
  out-of-core testing notes to contributing/testing.mdx;
- bench_large.py logs through the stdlib and gains --debug.

Type lazily read sites: the entry points accept any array-like with
shape, dtype, ndim and np.asarray support but were typed
Sequence[NDArray]. Add the SiteLike protocol and the Site alias, use
them from src/apply/compress down to the site source, and export
SiteLike. Also fix the other ty findings: the Tier list in the planner,
known_kind() for validated trains, padded_shape without a type: ignore,
cupyx as an allowed unresolved import, and typed fakes in the tests.

Assisted-by: pi:claude-opus-5.5
The pool limit counts every block CuPy's pool holds, including split
blocks that are only partly in use, while the plan counts the bytes in
use. With the cap at the budget, fragmentation alone tripped it: the
first GPU run failed at a pool of 29.9 MB under a 36.6 MB cap with
about 13-21 MB in use.

Cap the pool at the budget plus the max(10%, 1 GiB) margin that
detection already keeps free for workspaces and fragmentation, for
explicit budgets too; with a detected budget the cap is the free device
memory. Budgets gains device_cap (None on the CPU backend).

Assisted-by: pi:claude-opus-5.5
@github-actions

github-actions Bot commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Test Results

  5 files  ±  0    5 suites  ±0   3m 6s ⏱️ +35s
206 tests + 71  205 ✅ + 71  1 💤 ±0  0 ❌ ±0 
986 runs  +346  983 ✅ +346  3 💤 ±0  0 ❌ ±0 

Results for commit 9d0563f. ± Comparison against base commit 6bcedfe.

♻️ This comment has been updated with latest results.

The plan was a step-by-step guide for building the PR; the design spec
stays as the record of the design and its risks.

Assisted-by: pi:claude-opus-5.5
The user docs, docstrings and the benchmark README cover the behaviour;
the spec was a working document for building the PR.

Assisted-by: pi:claude-opus-5.5
gpu-nvidia now installs cupy-cuda13x[ctk]>=14, whose ctk extra pulls
the CUDA libraries from NVIDIA's cuda-toolkit wheels, in place of
cupy-cuda12x and the hand-listed nvidia-*-cu12 packages. CUDA 13 needs
an NVIDIA driver >= 580; the install docs say so.

Assisted-by: pi:claude-opus-5.5

@Panadestein Panadestein left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you @robertodr !

@Panadestein
Panadestein merged commit 286d9e9 into main Oct 7, 2026
15 checks passed
@Panadestein
Panadestein deleted the feat/src-out-of-core branch October 7, 2026 11:50
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants