Repository navigation
feat(sweep): ✨ out-of-core SRC sweep on one GPU - #42
Merged
Merged
Conversation
robertodr
added this pull request to stack #43
September 25, 2026 13:59
This was referenced Sep 27, 2026
robertodr
force-pushed
the
feat/src-out-of-core
branch
from
September 30, 2026 19:55
00a8c92 to
f4de686
Compare
Panadestein
force-pushed
the
feat/src-out-of-core
branch
from
October 5, 2026 14:48
f4de686 to
a8a50be
Compare
Phase 1 of the multi-GPU plan: stream the input cores, batch the contractions to a memory budget and keep the environments on the GPU, in host memory or on local disk. Targets N.V.M.U with D_M = 4000 and chi_out = 2000 in complex128 on one A100-40GB. Assisted-by: Pi:claude-opus-5-5
Eleven tasks from the backend helpers to the cluster acceptance run, each test-first. The code was prototyped and every intermediate state passes the CPU suite; five refinements of the spec are listed up front. Assisted-by: Pi:claude-opus-5-5
Host-to-device copies on the compute stream, batch sizes by binary search over the walked path, the promoted working dtype, lazy sites in bra stacks, and tests on dense operators rather than cores. Assisted-by: Pi:claude-opus-5-5
Assisted-by: Pi:claude-opus-5-5
Assisted-by: Pi:claude-opus-5-5
Assisted-by: Pi:claude-opus-5-5
Assisted-by: Pi:claude-opus-5-5
Assisted-by: Pi:claude-opus-5-5
…disk Assisted-by: Pi:claude-opus-5-5
Assisted-by: Pi:claude-opus-5-5
Assisted-by: Pi:claude-opus-5-5
Assisted-by: Pi:claude-opus-5-5
Assisted-by: Pi:claude-opus-5-5
Assisted-by: Pi:claude-opus-5-5
Panadestein
force-pushed
the
feat/src-out-of-core
branch
from
October 5, 2026 15:13
a8a50be to
2bfa815
Compare
robertodr
marked this pull request as ready for review
October 6, 2026 09:08
Resolve the conflicts with stdlib logging (#52), input-precision sketches (#36), validation (#38), ty enforcement (#37, #55) and the Fumadocs site (#56): - the sweep draws its sketches with gaussian_sketch and logs its plan, pass times and stalls at DEBUG with %-style arguments; - the working dtype is promoted over every site, not just the first; - docs/large-problems.md moves to features/large-problems.mdx and the out-of-core testing notes to contributing/testing.mdx; - bench_large.py logs through the stdlib and gains --debug. Type lazily read sites: the entry points accept any array-like with shape, dtype, ndim and np.asarray support but were typed Sequence[NDArray]. Add the SiteLike protocol and the Site alias, use them from src/apply/compress down to the site source, and export SiteLike. Also fix the other ty findings: the Tier list in the planner, known_kind() for validated trains, padded_shape without a type: ignore, cupyx as an allowed unresolved import, and typed fakes in the tests. Assisted-by: pi:claude-opus-5.5
The pool limit counts every block CuPy's pool holds, including split blocks that are only partly in use, while the plan counts the bytes in use. With the cap at the budget, fragmentation alone tripped it: the first GPU run failed at a pool of 29.9 MB under a 36.6 MB cap with about 13-21 MB in use. Cap the pool at the budget plus the max(10%, 1 GiB) margin that detection already keeps free for workspaces and fragmentation, for explicit budgets too; with a detected budget the cap is the free device memory. Budgets gains device_cap (None on the CPU backend). Assisted-by: pi:claude-opus-5.5
Contributor
The plan was a step-by-step guide for building the PR; the design spec stays as the record of the design and its risks. Assisted-by: pi:claude-opus-5.5
The user docs, docstrings and the benchmark README cover the behaviour; the spec was a working document for building the PR. Assisted-by: pi:claude-opus-5.5
gpu-nvidia now installs cupy-cuda13x[ctk]>=14, whose ctk extra pulls the CUDA libraries from NVIDIA's cuda-toolkit wheels, in place of cupy-cuda12x and the hand-listed nvidia-*-cu12 packages. CUDA 13 needs an NVIDIA driver >= 580; the install docs say so. Assisted-by: pi:claude-opus-5.5
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
🤖 AI text below 🤖
Phase 1 of running SRC on stacks too large for one GPU, such as
N . V . M . Uwith a large bond inM. The sweep now reads the input cores one site at a time, batches every contraction to a memory budget, and keeps the sketched environments on the GPU, in host memory or on node-local disk. The mathematics and the random draws are unchanged, so a seed gives the same operator as before, up to rounding.Stacked on #41.
What changes
Resourcesargument onsrc,applyandcompress: GPU and host memory budgets and a scratch directory. Budgets left unset are detected: free device memory minusmax(10%, 1 GiB), andMemAvailableminus 10%.shape,dtypeandnp.asarray, e.g.np.memmapor zarr/HDF5 datasets. Each core is read when the sweep reaches its site, and a background thread prefetches the next one._plan.py: a pure planner that picks per-site batch sizes and environment tiers, newest environments on the fastest tier._kernels.py: the batched contractions, plus a peak-memory estimate from theopt_einsumpath._store.py: environment storage, with pinned staging and writer/reader threads._sites.py: the site source._sweep.pybecomes a driver over these modules. It logs the plan and the time spent waiting on data.Design and plan:
docs/superpowers/specs/2026-09-25-src-out-of-core-design.mdanddocs/superpowers/plans/2026-09-25-src-out-of-core.md. The spec lists five refinements made while prototyping.Testing
uv run prek run --all-filesanduv run pytestpass on the CPU: 153 tests, including the benchmarks. That covers batching, disk spilling, lazy and memmap inputs, bra stacks of lazy sites, and error and cleanup paths. Small problems still plan to one batch on the device, and the benchmarks run as fast as before.ndarray.get(blocking=False), the pool cap) and the new tests intests/test_gpu_backend.pyhave not been run anywhere yet, since the development machine has no CUDA device.benches/large/targets 50 sites,D_M = 4000,chi_out = 2000, complex128 on one A100-40GB, with at least 300 GB of host memory and local NVMe. Its README lists the four pass criteria; the results will be added there.