Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -31,7 +31,7 @@ build*/
# Executables
*.exe
*.out
!benches/leonardo/logs/*.out
!benches/*/leonardo/logs/*.out
*.app

### Python ###
Expand Down
30 changes: 29 additions & 1 deletion benches/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,8 @@ different SRC variants. They are grouped by the kind of workload exercised:
See [`primitives/README.md`](primitives/README.md).
- [`stack/`](stack/) — One-shot SRC over stacks of trains against sequential
pairwise application, by stack depth: accuracy against dense references and
wall time. See [`stack/README.md`](stack/README.md).
wall time on small chains, and quench dynamics of 50 and 100 qubits on the GPU.
See [`stack/README.md`](stack/README.md).
- [`large/`](large/) — Out-of-core SRC of `N . V . M . U` with a large `M` read
from disk, on one GPU: plan, wall time, memory peaks and stall time. See
[`large/README.md`](large/README.md).
Expand All @@ -19,3 +20,30 @@ The scripts need the `bench` dependency group (included in `dev`):
```bash
uv sync --group bench
```

## Leonardo

Each folder has a `leonardo/` subfolder with a Slurm script for the Booster
partition of [Leonardo](https://docs.hpc.cineca.it/general/getting_started.html)
at CINECA (one node: 32 Xeon 8358 cores, 4 A100-64GB, 512 GB of memory), and the
logs of the recorded runs. Set up the environment once, from the repository
root on a login node:

```bash
# Keep the uv cache and Python installs out of $HOME.
export UV_CACHE_DIR=$PWD/.uv-cache UV_PYTHON_INSTALL_DIR=$PWD/.uv-python
uv sync --all-groups --extra gpu-nvidia

# The Booster driver (535) predates CUDA 13: unpack NVIDIA's forward-compat libcuda.
mkdir -p .cuda-compat && cd .cuda-compat
curl -sSfO https://developer.download.nvidia.com/compute/cuda/repos/rhel8/x86_64/cuda-compat-13-4-615.71.09-1.el8.x86_64.rpm
rpm2cpio cuda-compat-13-4-*.rpm | cpio -idm && cd ..
```

The scripts prepend `.cuda-compat/usr/local/cuda-13.4/compat` to
`LD_LIBRARY_PATH` (override with `CUDA_COMPAT`) and keep CuPy's kernel cache in
`.cupy-cache/`. Submit from the `leonardo/` folder, passing the benchmark's own
arguments, e.g. `sbatch run.sh --n-sites 25 --chi-out 1000 --device gpu`. The
default QoS allows long runs but can queue for hours; for jobs under 30 minutes
add `--qos=boost_qos_dbg --time=00:30:00`, which starts quickly but takes at
most two jobs per user.
29 changes: 29 additions & 0 deletions benches/large/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -31,3 +31,32 @@ memory and node-local NVMe:
budget.
3. The stall time is below 10% of the wall time.
4. `compare` at `D_M = 1000`, `chi_out = 500` reports a distance below `1e-10`.

## Leonardo

The [`leonardo/`](leonardo/) folder holds the Slurm script and the logs; see the
[Leonardo section](../README.md#leonardo) for the environment. Booster nodes have
no local disk and a 10 GB `/tmp`, so the script points `TMPDIR` at
`$CINECA_SCRATCH`, and the stacks live there too (Lustre, not NVMe):

```bash
D=$CINECA_SCRATCH/src-large
sbatch run.sh generate $D/m4000 --bond-m 4000
sbatch run.sh --debug run $D/m4000 --chi-out 2000 --scratch-dir $D/scratch
sbatch run.sh generate $D/m1000 --bond-m 1000
sbatch run.sh --debug compare $D/m1000 --chi-out 500 --small 4GB --large 60GB \
--scratch-dir $D/scratch
```

`generate` writes the `D_M = 1000` stack (12 GB) in 51 s.

`compare` at `D_M = 1000`, `chi_out = 500`, complex128, on one A100-64GB:

| GPU budget | Wall time (s) | Pool (GB) | Planned device peak (GB) | Host peak (GB) | Tiers (device/host/disk) |
|---|---|---|---|---|---|
| 4GB | 129.8 | 5.46 | 3.86 | 40.6 | 1 / 49 / 0 |
| 60GB | 109.9 | 60.81 | 58.84 | 40.6 | 50 / 0 / 0 |

The two plans keep opposite tiers yet give bit-identical outputs (relative distance
`0.000e+00`). Spilling 49 of 50 environments to the host costs 18 % of the wall
time. Both pools stay within their caps (budget plus 10 % of the card, 6.9 GB).
15 changes: 15 additions & 0 deletions benches/large/leonardo/logs/large_59625448.out
Original file line number Diff line number Diff line change
@@ -0,0 +1,15 @@
NODE: lrdn0113, CPUS: 32, ARGS: generate /leonardo_scratch/large/userexternal/rpanades/src-large/m1000 --bond-m 1000
NVIDIA A100-SXM-64GB, 535.274.02
NVIDIA A100-SXM-64GB, 535.274.02
NVIDIA A100-SXM-64GB, 535.274.02
NVIDIA A100-SXM-64GB, 535.274.02
Layer N written: bond 4
Layer V written: bond 4
Layer M written: bond 1000
Layer U written: bond 4
JobID JobName Elapsed MaxRSS MaxDiskRead MaxDiskWrite ExitCode
------------ ---------- ---------- ---------- ------------ ------------ --------
59625448 large 00:01:01 0:0
59625448.ba+ batch 00:01:01 0:0
59625448.ex+ extern 00:01:01 0:0
59625448.0 python 00:00:51 5.74G 0.01G 0.02M 0:0
28 changes: 28 additions & 0 deletions benches/large/leonardo/logs/large_compare_59625755.out
Original file line number Diff line number Diff line change
@@ -0,0 +1,28 @@
NODE: lrdn2346, CPUS: 32, ARGS: --debug compare /leonardo_scratch/large/userexternal/rpanades/src-large/m1000 --chi-out 500 --small 4GB --large 60GB --scratch-dir /leonardo_scratch/large/userexternal/rpanades/src-large/scratch
NVIDIA A100-SXM-64GB, 535.274.02
NVIDIA A100-SXM-64GB, 535.274.02
NVIDIA A100-SXM-64GB, 535.274.02
NVIDIA A100-SXM-64GB, 535.274.02
Starting SRC: n_sites=50, depth=4, output=mpo, device=cupy
SRC plan: prefetch=1, device peak=3855012864 B, host peak=30958244864 B, disk=0 B, scratch=None, tiers=['device', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host'], batches (env, sketch, project)=[(480, 0, 480), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (0, 480, 480)]
Left-to-right sweep: 38.243 s
Right-to-left sweep: 69.762 s
SRC stalls: sites 10.108 s, environments 31.839 s
Device pool: 5460450304 B
SRC complete
Run complete: 129.8 s, pool 5460450304 B, host peak 40588197888 B, planned device peak 3855012864 B, planned host peak 30958244864 B, disk 0 B, tiers {'device': 1, 'host': 49, 'disk': 0}, sketch batches [32, 480]
Starting SRC: n_sites=50, depth=4, output=mpo, device=cupy
SRC plan: prefetch=1, device peak=58841624576 B, host peak=3904164864 B, disk=0 B, scratch=None, tiers=['device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device'], batches (env, sketch, project)=[(480, 0, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (0, 480, 480)]
Left-to-right sweep: 13.859 s
Right-to-left sweep: 95.954 s
SRC stalls: sites 13.835 s, environments 0.000 s
Device pool: 60811395072 B
SRC complete
Run complete: 109.9 s, pool 60811395072 B, host peak 40588197888 B, planned device peak 58841624576 B, planned host peak 3904164864 B, disk 0 B, tiers {'device': 50, 'host': 0, 'disk': 0}, sketch batches [480]
Relative distance between the runs: 0.000e+00
JobID JobName Elapsed MaxRSS MaxDiskRead MaxDiskWrite ExitCode
------------ ---------- ---------- ---------- ------------ ------------ --------
59625755 large_com+ 00:04:47 0:0
59625755.ba+ batch 00:04:47 0:0
59625755.ex+ extern 00:04:47 0:0
59625755.0 python 00:04:44 0:0
15 changes: 15 additions & 0 deletions benches/large/leonardo/logs/large_generate_59625937.out
Original file line number Diff line number Diff line change
@@ -0,0 +1,15 @@
NODE: lrdn2918, CPUS: 32, ARGS: generate /leonardo_scratch/large/userexternal/rpanades/src-large/m4000 --bond-m 4000
NVIDIA A100-SXM-64GB, 535.274.02
NVIDIA A100-SXM-64GB, 535.274.02
NVIDIA A100-SXM-64GB, 535.274.02
NVIDIA A100-SXM-64GB, 535.274.02
Layer N written: bond 4
Layer V written: bond 4
Layer M written: bond 4000
Layer U written: bond 4
JobID JobName Elapsed MaxRSS MaxDiskRead MaxDiskWrite ExitCode
------------ ---------- ---------- ---------- ------------ ------------ --------
59625937 large_gen+ 00:11:35 0:0
59625937.ba+ batch 00:11:35 0:0
59625937.ex+ extern 00:11:35 0:0
59625937.0 python 00:11:31 2.42G 0.01G 0.02M 0:0
41 changes: 41 additions & 0 deletions benches/large/leonardo/run.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,41 @@
#!/usr/bin/env -S bash -l
# Usage, from this folder: sbatch run.sh [--debug] generate|run|compare [options]
# e.g. sbatch run.sh --debug run "$CINECA_SCRATCH/src-large/m4000" --chi-out 2000
# Add --qos=boost_qos_dbg (30 min, 2 nodes) to sbatch for short test runs.
#SBATCH --nodes=1
#SBATCH --ntasks-per-node=1
#SBATCH --job-name=large
#SBATCH --account=EUHPC_D30_139
#SBATCH --partition=boost_usr_prod
#SBATCH --time=04:00:00
#SBATCH --cpus-per-task=32
#SBATCH --gres=gpu:1
#SBATCH --mem=0
#SBATCH --exclusive
#SBATCH --output=logs/%x_%j.out

set -euo pipefail

export OMP_NUM_THREADS=$SLURM_CPUS_PER_TASK
export OMP_PLACES=cores
export OMP_PROC_BIND=spread
# Booster nodes have no local disk and /tmp is a small tmpfs: spill to Lustre scratch.
export TMPDIR=$CINECA_SCRATCH/src-large/tmp
mkdir -p "$TMPDIR"

REPO=$(git -C "$SLURM_SUBMIT_DIR" rev-parse --show-toplevel)
source "$REPO/.venv/bin/activate"
export CUPY_CACHE_DIR=$REPO/.cupy-cache
# The Booster driver (535) predates CUDA 13: load NVIDIA's forward-compat libcuda.
CUDA_COMPAT=${CUDA_COMPAT:-$REPO/.cuda-compat/usr/local/cuda-13.4/compat}
if [[ -d $CUDA_COMPAT ]]; then
export LD_LIBRARY_PATH=$CUDA_COMPAT${LD_LIBRARY_PATH:+:$LD_LIBRARY_PATH}
fi

echo "NODE: $SLURMD_NODENAME, CPUS: $SLURM_CPUS_PER_TASK, ARGS: $*"
nvidia-smi --query-gpu=name,driver_version --format=csv,noheader

srun python "$REPO/benches/large/bench_large.py" "$@"

sacct --format=JobID,JobName,Elapsed,MaxRSS,MaxDiskRead,MaxDiskWrite,ExitCode \
--units=G --jobs="$SLURM_JOB_ID"
43 changes: 33 additions & 10 deletions benches/primitives/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,18 +8,41 @@ controlled bond dimension, and compare to a `quimb` reference.
## Leonardo

The Leonardo [supercomputer](https://docs.hpc.cineca.it/general/getting_started.html) at CINECA is
a pre-exascale Tier-0 EuroHPC system, the 10th fastest in the world as of June 2025. In the
corresponding folder you can find the scripts to run the benchmarks on it, and some example
results. It is assumed that the user has built the `uv` virtual environment in the root
of the repository. This can be done by running:

```bash
uv sync --all-groups --all-extras
```
a pre-exascale Tier-0 EuroHPC system. The [`leonardo/`](leonardo/) folder holds the Slurm
script and the logs; see the [Leonardo section](../README.md#leonardo) for the environment.

### MPO-MPO benchmark

The benchmark contracts two length-`n_sites` matrix product operators (MPOs) whose initial bond dimension is `chi_out` (unless `chi_id` is specified for the second MPO), then compresses the resulting MPO back to bond dimension `chi_out`. One MPO is fully random; the other is an identity MPO perturbed by summing it with another random MPO with small elements, making the contraction highly compressible. The table below compares the performance of `quimb` versus `src` in terms of CPU time, total memory usage (maximum resident set size, which is the peak memory usage of the job), relative speedup, and accuracy.
The benchmark contracts two length-`n_sites` matrix product operators (MPOs) whose initial bond dimension is `chi_out` (unless `chi_id` is specified for the second MPO), then compresses the resulting MPO back to bond dimension `chi_out`. One MPO is fully random; the other is an identity MPO perturbed by summing it with another random MPO with small elements, making the contraction highly compressible. The table below compares `quimb` (CPU, `rsvd` compression) with `src` on the CPU and on one GPU: wall time of the contraction-compression, peak host memory of the job (MaxRSS), speedup over `quimb`, and relative distance to the uncompressed product.

One Booster node (32 cores, A100-64GB), complex128 unless noted, `chi_id=4` at `chi=1000`:

| Experiment | Library | Time (s) | MaxRSS (GB) | Speedup | Distance to reference |
|-------------------|---------------------|----------|-------------|---------|-----------------------|
| 50 sites chi=50 | quimb | 37.76 | 0.67 | | 2.6e-08 |
| | src CPU | 0.23 | 0.57 | 164x | 1.5e-08 |
| | src GPU | 0.57 | 1.01 | 66x | 1.5e-08 |
| 20 sites chi=100 | quimb | 17.86 | 0.73 | | 0.0 |
| | src CPU | 0.25 | 0.74 | 71x | 1.5e-08 |
| | src GPU | 0.46 | 1.22 | 39x | 2.1e-08 |
| 25 sites chi=1000 | quimb | 422.57 | 24.64 | | |
| | src CPU | 60.47 | 3.85 | 7x | |
| | src GPU | 2.32 | 3.42 | 182x | |
| | src GPU (complex64) | 1.90 | 2.20 | 222x | |
| 50 sites chi=1000 | quimb | 1040.72 | 50.02 | | |
| | src CPU | 137.37 | 6.83 | 7.6x | |
| | src GPU | 4.47 | 6.40 | 233x | |
| | src GPU (complex64) | 3.03 | 3.17 | 344x | |

The distances sit at the floor of quimb's `distance` (about `sqrt(eps)` of the
dtype), so they show agreement rather than resolve the error. GPU times exclude a
one-off warm-up (CUDA context and kernel compilation, about 50 s on the first run
with an empty kernel cache). The CuPy pool peaks at 2.8 GB (25 sites) and 5.1 GB
(50 sites) at chi=1000 in complex128, half that in complex64.

### Earlier results (v0.3.2)

Before the refactoring, on Booster (32 CPUs) and DCGP (112 CPUs) nodes:

| Experiment | CPUs | Library | Time (s) | MaxRSS (GB) | Speedup | Distance to reference |
|-------------------|------|---------------|----------|-------------|---------|-----------------------|
Expand All @@ -35,7 +58,7 @@ The benchmark contracts two length-`n_sites` matrix product operators (MPOs) who
| 50 sites chi=1000 | 112 | quimb (rsvd) | 1187.82 | 49.47 | | |
| chi_id=4 | | src | 150.11 | 7.01 | 8x | |

### MPO-MPO strong scaling
### MPO-MPO strong scaling (v0.3.2, DCGP)

For this benchmark, we fix the problem size to `n_sites=25`, `chi=1000` and `chi_id=4`, and vary the number of CPUs in a single node. We tweak the OpenMP environment variables in the `run.sh` script to optimize performance, for instance `OMP_PLACES` and `OMP_PROC_BIND`, so the NUMA domain is taken into account. We do this both manually and automatically.

Expand Down
Loading
Loading