|
| 1 | +# Large problems |
| 2 | + |
| 3 | +A single SRC sweep keeps, besides the input and output trains, one sketched |
| 4 | +environment per site. For stacks with a large bond, such as `N . V . M . U` with a |
| 5 | +bond of several thousand in `M`, these environments and the intermediates of the |
| 6 | +contractions no longer fit on a GPU, and often not in host memory either. |
| 7 | +`src_method` then plans the sweep to the memory at hand: |
| 8 | + |
| 9 | +- the input cores are read one site at a time, so a train can live on disk; |
| 10 | +- every contraction runs in batches of sketch columns (or rows), sized to the GPU |
| 11 | + budget; |
| 12 | +- each environment is kept on the GPU, in host memory or in a scratch directory, |
| 13 | + the most recent ones on the fastest tier. |
| 14 | + |
| 15 | +None of this changes the result beyond floating-point rounding: the random draws |
| 16 | +and the mathematics are those of a sweep that fits in memory. |
| 17 | + |
| 18 | +## Budgets |
| 19 | + |
| 20 | +The budgets are set with `Resources`, passed to `src`, `apply` or `compress`: |
| 21 | + |
| 22 | +```python |
| 23 | +from src_method import Resources, src |
| 24 | + |
| 25 | +out = src( |
| 26 | + N, V, M, U, |
| 27 | + chi_out=2000, |
| 28 | + dtype=np.complex128, |
| 29 | + device="gpu", |
| 30 | + resources=Resources(gpu_memory="36GB", scratch_dir="/local/scratch"), |
| 31 | +) |
| 32 | +``` |
| 33 | + |
| 34 | +Every field left unset is detected when the call starts: |
| 35 | + |
| 36 | +| Field | Default | |
| 37 | +|---|---| |
| 38 | +| `gpu_memory` | free device memory, plus the free bytes of CuPy's pool, minus `max(10%, 1 GiB)` | |
| 39 | +| `host_memory` | `MemAvailable` from `/proc/meminfo`, minus 10% | |
| 40 | +| `scratch_dir` | `tempfile.gettempdir()`, which honours `$TMPDIR` | |
| 41 | + |
| 42 | +Sizes are byte counts or strings: `"36GB"` is `36 * 10**9` bytes, `"36GiB"` is |
| 43 | +`36 * 2**30`. On the CPU there is a single budget, `host_memory`. Host memory is |
| 44 | +measured when the call starts, so inputs already held in memory are not counted |
| 45 | +twice. |
| 46 | + |
| 47 | +During the sweep, CuPy's default memory pool is capped at the GPU budget, so that |
| 48 | +an estimate that falls short fails at once rather than when some other allocation |
| 49 | +does; the previous limit is restored afterwards. If a single site does not fit the |
| 50 | +budget even with batches of one column, `src` raises `MemoryError` naming the site. |
| 51 | + |
| 52 | +## Inputs on disk |
| 53 | + |
| 54 | +A train is any sequence of per-site array-likes with `shape`, `dtype` and |
| 55 | +`np.asarray` support: NumPy arrays, `np.memmap`, zarr or HDF5 datasets. Each core |
| 56 | +is read when the sweep reaches its site, once per pass, and a background thread |
| 57 | +reads the next site while the current one is computed. Raw `.npy` files opened with |
| 58 | +`np.load(path, mmap_mode="r")` are the fastest option; compressed formats may be |
| 59 | +limited by decompression. |
| 60 | + |
| 61 | +## The scratch directory |
| 62 | + |
| 63 | +Environments that fit in neither budget are written to a per-process directory, |
| 64 | +`<scratch_dir>/src-<pid>-<id>/`, one file per site. Use node-local disk: at the |
| 65 | +reference size of 50 sites, `D_M = 4000` and `chi_out = 2000` in complex128, up to |
| 66 | +410 GB are written and read back once. The directory is removed when the call |
| 67 | +returns, also after an error; a process killed with `SIGKILL` leaves it behind, and |
| 68 | +its name identifies the process. |
| 69 | + |
| 70 | +## Reading the plan |
| 71 | + |
| 72 | +Every sweep logs its plan at `info`: |
| 73 | + |
| 74 | +- `tiers`: where the environment of each site lives (`device`, `host` or `disk`); |
| 75 | +- `batches`: for each site, the batch of the environment, sketch and projection |
| 76 | + steps; |
| 77 | +- `device_peak_bytes`, `host_peak_bytes`, `disk_bytes`: the planned peaks; |
| 78 | +- `prefetch`: whether the next site is read ahead. |
| 79 | + |
| 80 | +At the end, `SRC stalls` reports the seconds the sweep waited for input cores |
| 81 | +(`site_seconds`) and for environments (`environment_seconds`). If they are a large |
| 82 | +fraction of the run, the disk is too slow for the compute. Set |
| 83 | +`LOG_LEVEL_SRC=DEBUG` to also get the time of each pass. |
0 commit comments