Skip to content

Commit 00a8c92

Browse files
committed
docs: 📝 large problems, Resources and the out-of-core test knobs
Assisted-by: Pi:claude-opus-5-5
1 parent 3518dd6 commit 00a8c92

4 files changed

Lines changed: 117 additions & 0 deletions

File tree

‎README.md‎

Lines changed: 20 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -87,6 +87,26 @@ The result then pairs with a ket by plain contraction, with no further conjugati
8787

8888
At runtime, pass `device="gpu"` to use GPU acceleration. The library handles backend dispatch automatically.
8989

90+
### Large Problems
91+
92+
Stacks that do not fit on the GPU, or in host memory, still run: the sweep reads the input cores one site at a time (from NumPy arrays, `np.memmap`, zarr or HDF5 datasets), batches every contraction to a memory budget, and keeps the sketched environments on the GPU, in host memory or on local disk. The budgets are detected, or set explicitly with `Resources`:
93+
94+
```python
95+
from src_method import Resources, src
96+
97+
out = src(
98+
N,
99+
V,
100+
M,
101+
U,
102+
chi_out=2000,
103+
device="gpu",
104+
resources=Resources(gpu_memory="36GB", scratch_dir="/local/scratch"),
105+
)
106+
```
107+
108+
See [Large problems](docs/large-problems.md) for the budgets, the scratch directory and how to read the logged plan.
109+
90110
## Installation
91111

92112
```bash

‎docs/developer-guide/testing.md‎

Lines changed: 13 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -253,6 +253,19 @@ def test_helper_namespace():
253253
assert pytest.helpers.foo(True) is True
254254
```
255255

256+
### Exercising the out-of-core paths
257+
258+
The batched contractions and the host and disk tiers of the sweep only engage when
259+
the budgets are tight, so tests force them with tiny budgets: on the CPU,
260+
`Resources(host_memory="1MB", scratch_dir=tmp_path)` gives small batches and spills
261+
the environments of a depth-4 stack with bonds of 4 to `tmp_path` (see
262+
`tests/test_stack.py`). Compare the dense operator of the result with that of a
263+
default run, not the cores: batching changes the rounding, and with it the cores of
264+
an ill-conditioned sketch, but not the operator they represent. The planner is a
265+
pure function, so `tests/test_plan.py` checks batch sizes and tiers from shapes
266+
alone, and `tests/test_gpu_backend.py` derives GPU budgets from a plan made with
267+
`make_plan` to reach every tier.
268+
256269
### How to use logging with tests
257270

258271
Within `aurora`, we adopt a `pytest` configuration that allows to see the output

‎docs/large-problems.md‎

Lines changed: 83 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,83 @@
1+
# Large problems
2+
3+
A single SRC sweep keeps, besides the input and output trains, one sketched
4+
environment per site. For stacks with a large bond, such as `N . V . M . U` with a
5+
bond of several thousand in `M`, these environments and the intermediates of the
6+
contractions no longer fit on a GPU, and often not in host memory either.
7+
`src_method` then plans the sweep to the memory at hand:
8+
9+
- the input cores are read one site at a time, so a train can live on disk;
10+
- every contraction runs in batches of sketch columns (or rows), sized to the GPU
11+
budget;
12+
- each environment is kept on the GPU, in host memory or in a scratch directory,
13+
the most recent ones on the fastest tier.
14+
15+
None of this changes the result beyond floating-point rounding: the random draws
16+
and the mathematics are those of a sweep that fits in memory.
17+
18+
## Budgets
19+
20+
The budgets are set with `Resources`, passed to `src`, `apply` or `compress`:
21+
22+
```python
23+
from src_method import Resources, src
24+
25+
out = src(
26+
N, V, M, U,
27+
chi_out=2000,
28+
dtype=np.complex128,
29+
device="gpu",
30+
resources=Resources(gpu_memory="36GB", scratch_dir="/local/scratch"),
31+
)
32+
```
33+
34+
Every field left unset is detected when the call starts:
35+
36+
| Field | Default |
37+
|---|---|
38+
| `gpu_memory` | free device memory, plus the free bytes of CuPy's pool, minus `max(10%, 1 GiB)` |
39+
| `host_memory` | `MemAvailable` from `/proc/meminfo`, minus 10% |
40+
| `scratch_dir` | `tempfile.gettempdir()`, which honours `$TMPDIR` |
41+
42+
Sizes are byte counts or strings: `"36GB"` is `36 * 10**9` bytes, `"36GiB"` is
43+
`36 * 2**30`. On the CPU there is a single budget, `host_memory`. Host memory is
44+
measured when the call starts, so inputs already held in memory are not counted
45+
twice.
46+
47+
During the sweep, CuPy's default memory pool is capped at the GPU budget, so that
48+
an estimate that falls short fails at once rather than when some other allocation
49+
does; the previous limit is restored afterwards. If a single site does not fit the
50+
budget even with batches of one column, `src` raises `MemoryError` naming the site.
51+
52+
## Inputs on disk
53+
54+
A train is any sequence of per-site array-likes with `shape`, `dtype` and
55+
`np.asarray` support: NumPy arrays, `np.memmap`, zarr or HDF5 datasets. Each core
56+
is read when the sweep reaches its site, once per pass, and a background thread
57+
reads the next site while the current one is computed. Raw `.npy` files opened with
58+
`np.load(path, mmap_mode="r")` are the fastest option; compressed formats may be
59+
limited by decompression.
60+
61+
## The scratch directory
62+
63+
Environments that fit in neither budget are written to a per-process directory,
64+
`<scratch_dir>/src-<pid>-<id>/`, one file per site. Use node-local disk: at the
65+
reference size of 50 sites, `D_M = 4000` and `chi_out = 2000` in complex128, up to
66+
410 GB are written and read back once. The directory is removed when the call
67+
returns, also after an error; a process killed with `SIGKILL` leaves it behind, and
68+
its name identifies the process.
69+
70+
## Reading the plan
71+
72+
Every sweep logs its plan at `info`:
73+
74+
- `tiers`: where the environment of each site lives (`device`, `host` or `disk`);
75+
- `batches`: for each site, the batch of the environment, sketch and projection
76+
steps;
77+
- `device_peak_bytes`, `host_peak_bytes`, `disk_bytes`: the planned peaks;
78+
- `prefetch`: whether the next site is read ahead.
79+
80+
At the end, `SRC stalls` reports the seconds the sweep waited for input cores
81+
(`site_seconds`) and for environments (`environment_seconds`). If they are a large
82+
fraction of the run, the disk is too slow for the compute. Set
83+
`LOG_LEVEL_SRC=DEBUG` to also get the time of each pass.

‎mkdocs.yml‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -5,6 +5,7 @@ watch:
55
- src/src_method
66
nav:
77
- Home: index.md
8+
- 'Large Problems': large-problems.md
89
- 'Developer Guide':
910
- 'Versioning Scheme': developer-guide/versioning.md
1011
- 'Managing Dependencies': developer-guide/dependencies.md

0 commit comments

Comments
 (0)