diff --git a/.gitignore b/.gitignore index ca75eea..862de93 100644 --- a/.gitignore +++ b/.gitignore @@ -31,7 +31,7 @@ build*/ # Executables *.exe *.out -!benches/leonardo/logs/*.out +!benches/*/leonardo/logs/*.out *.app ### Python ### diff --git a/benches/README.md b/benches/README.md index 8a49f83..051a290 100644 --- a/benches/README.md +++ b/benches/README.md @@ -9,7 +9,8 @@ different SRC variants. They are grouped by the kind of workload exercised: See [`primitives/README.md`](primitives/README.md). - [`stack/`](stack/) — One-shot SRC over stacks of trains against sequential pairwise application, by stack depth: accuracy against dense references and - wall time. See [`stack/README.md`](stack/README.md). + wall time on small chains, and quench dynamics of 50 and 100 qubits on the GPU. + See [`stack/README.md`](stack/README.md). - [`large/`](large/) — Out-of-core SRC of `N . V . M . U` with a large `M` read from disk, on one GPU: plan, wall time, memory peaks and stall time. See [`large/README.md`](large/README.md). @@ -19,3 +20,30 @@ The scripts need the `bench` dependency group (included in `dev`): ```bash uv sync --group bench ``` + +## Leonardo + +Each folder has a `leonardo/` subfolder with a Slurm script for the Booster +partition of [Leonardo](https://docs.hpc.cineca.it/general/getting_started.html) +at CINECA (one node: 32 Xeon 8358 cores, 4 A100-64GB, 512 GB of memory), and the +logs of the recorded runs. Set up the environment once, from the repository +root on a login node: + +```bash +# Keep the uv cache and Python installs out of $HOME. +export UV_CACHE_DIR=$PWD/.uv-cache UV_PYTHON_INSTALL_DIR=$PWD/.uv-python +uv sync --all-groups --extra gpu-nvidia + +# The Booster driver (535) predates CUDA 13: unpack NVIDIA's forward-compat libcuda. +mkdir -p .cuda-compat && cd .cuda-compat +curl -sSfO https://developer.download.nvidia.com/compute/cuda/repos/rhel8/x86_64/cuda-compat-13-4-615.71.09-1.el8.x86_64.rpm +rpm2cpio cuda-compat-13-4-*.rpm | cpio -idm && cd .. +``` + +The scripts prepend `.cuda-compat/usr/local/cuda-13.4/compat` to +`LD_LIBRARY_PATH` (override with `CUDA_COMPAT`) and keep CuPy's kernel cache in +`.cupy-cache/`. Submit from the `leonardo/` folder, passing the benchmark's own +arguments, e.g. `sbatch run.sh --n-sites 25 --chi-out 1000 --device gpu`. The +default QoS allows long runs but can queue for hours; for jobs under 30 minutes +add `--qos=boost_qos_dbg --time=00:30:00`, which starts quickly but takes at +most two jobs per user. diff --git a/benches/large/README.md b/benches/large/README.md index 24e15f4..7cc9eee 100644 --- a/benches/large/README.md +++ b/benches/large/README.md @@ -31,3 +31,32 @@ memory and node-local NVMe: budget. 3. The stall time is below 10% of the wall time. 4. `compare` at `D_M = 1000`, `chi_out = 500` reports a distance below `1e-10`. + +## Leonardo + +The [`leonardo/`](leonardo/) folder holds the Slurm script and the logs; see the +[Leonardo section](../README.md#leonardo) for the environment. Booster nodes have +no local disk and a 10 GB `/tmp`, so the script points `TMPDIR` at +`$CINECA_SCRATCH`, and the stacks live there too (Lustre, not NVMe): + +```bash +D=$CINECA_SCRATCH/src-large +sbatch run.sh generate $D/m4000 --bond-m 4000 +sbatch run.sh --debug run $D/m4000 --chi-out 2000 --scratch-dir $D/scratch +sbatch run.sh generate $D/m1000 --bond-m 1000 +sbatch run.sh --debug compare $D/m1000 --chi-out 500 --small 4GB --large 60GB \ + --scratch-dir $D/scratch +``` + +`generate` writes the `D_M = 1000` stack (12 GB) in 51 s. + +`compare` at `D_M = 1000`, `chi_out = 500`, complex128, on one A100-64GB: + +| GPU budget | Wall time (s) | Pool (GB) | Planned device peak (GB) | Host peak (GB) | Tiers (device/host/disk) | +|---|---|---|---|---|---| +| 4GB | 129.8 | 5.46 | 3.86 | 40.6 | 1 / 49 / 0 | +| 60GB | 109.9 | 60.81 | 58.84 | 40.6 | 50 / 0 / 0 | + +The two plans keep opposite tiers yet give bit-identical outputs (relative distance +`0.000e+00`). Spilling 49 of 50 environments to the host costs 18 % of the wall +time. Both pools stay within their caps (budget plus 10 % of the card, 6.9 GB). diff --git a/benches/large/leonardo/logs/large_59625448.out b/benches/large/leonardo/logs/large_59625448.out new file mode 100644 index 0000000..a976da7 --- /dev/null +++ b/benches/large/leonardo/logs/large_59625448.out @@ -0,0 +1,15 @@ +NODE: lrdn0113, CPUS: 32, ARGS: generate /leonardo_scratch/large/userexternal/rpanades/src-large/m1000 --bond-m 1000 +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +Layer N written: bond 4 +Layer V written: bond 4 +Layer M written: bond 1000 +Layer U written: bond 4 +JobID JobName Elapsed MaxRSS MaxDiskRead MaxDiskWrite ExitCode +------------ ---------- ---------- ---------- ------------ ------------ -------- +59625448 large 00:01:01 0:0 +59625448.ba+ batch 00:01:01 0:0 +59625448.ex+ extern 00:01:01 0:0 +59625448.0 python 00:00:51 5.74G 0.01G 0.02M 0:0 diff --git a/benches/large/leonardo/logs/large_compare_59625755.out b/benches/large/leonardo/logs/large_compare_59625755.out new file mode 100644 index 0000000..483c5ef --- /dev/null +++ b/benches/large/leonardo/logs/large_compare_59625755.out @@ -0,0 +1,28 @@ +NODE: lrdn2346, CPUS: 32, ARGS: --debug compare /leonardo_scratch/large/userexternal/rpanades/src-large/m1000 --chi-out 500 --small 4GB --large 60GB --scratch-dir /leonardo_scratch/large/userexternal/rpanades/src-large/scratch +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +Starting SRC: n_sites=50, depth=4, output=mpo, device=cupy +SRC plan: prefetch=1, device peak=3855012864 B, host peak=30958244864 B, disk=0 B, scratch=None, tiers=['device', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host', 'host'], batches (env, sketch, project)=[(480, 0, 480), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (64, 32, 16), (0, 480, 480)] +Left-to-right sweep: 38.243 s +Right-to-left sweep: 69.762 s +SRC stalls: sites 10.108 s, environments 31.839 s +Device pool: 5460450304 B +SRC complete +Run complete: 129.8 s, pool 5460450304 B, host peak 40588197888 B, planned device peak 3855012864 B, planned host peak 30958244864 B, disk 0 B, tiers {'device': 1, 'host': 49, 'disk': 0}, sketch batches [32, 480] +Starting SRC: n_sites=50, depth=4, output=mpo, device=cupy +SRC plan: prefetch=1, device peak=58841624576 B, host peak=3904164864 B, disk=0 B, scratch=None, tiers=['device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device'], batches (env, sketch, project)=[(480, 0, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (480, 480, 480), (0, 480, 480)] +Left-to-right sweep: 13.859 s +Right-to-left sweep: 95.954 s +SRC stalls: sites 13.835 s, environments 0.000 s +Device pool: 60811395072 B +SRC complete +Run complete: 109.9 s, pool 60811395072 B, host peak 40588197888 B, planned device peak 58841624576 B, planned host peak 3904164864 B, disk 0 B, tiers {'device': 50, 'host': 0, 'disk': 0}, sketch batches [480] +Relative distance between the runs: 0.000e+00 +JobID JobName Elapsed MaxRSS MaxDiskRead MaxDiskWrite ExitCode +------------ ---------- ---------- ---------- ------------ ------------ -------- +59625755 large_com+ 00:04:47 0:0 +59625755.ba+ batch 00:04:47 0:0 +59625755.ex+ extern 00:04:47 0:0 +59625755.0 python 00:04:44 0:0 diff --git a/benches/large/leonardo/logs/large_generate_59625937.out b/benches/large/leonardo/logs/large_generate_59625937.out new file mode 100644 index 0000000..7c1ba1f --- /dev/null +++ b/benches/large/leonardo/logs/large_generate_59625937.out @@ -0,0 +1,15 @@ +NODE: lrdn2918, CPUS: 32, ARGS: generate /leonardo_scratch/large/userexternal/rpanades/src-large/m4000 --bond-m 4000 +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +Layer N written: bond 4 +Layer V written: bond 4 +Layer M written: bond 4000 +Layer U written: bond 4 +JobID JobName Elapsed MaxRSS MaxDiskRead MaxDiskWrite ExitCode +------------ ---------- ---------- ---------- ------------ ------------ -------- +59625937 large_gen+ 00:11:35 0:0 +59625937.ba+ batch 00:11:35 0:0 +59625937.ex+ extern 00:11:35 0:0 +59625937.0 python 00:11:31 2.42G 0.01G 0.02M 0:0 diff --git a/benches/large/leonardo/run.sh b/benches/large/leonardo/run.sh new file mode 100755 index 0000000..55d9037 --- /dev/null +++ b/benches/large/leonardo/run.sh @@ -0,0 +1,41 @@ +#!/usr/bin/env -S bash -l +# Usage, from this folder: sbatch run.sh [--debug] generate|run|compare [options] +# e.g. sbatch run.sh --debug run "$CINECA_SCRATCH/src-large/m4000" --chi-out 2000 +# Add --qos=boost_qos_dbg (30 min, 2 nodes) to sbatch for short test runs. +#SBATCH --nodes=1 +#SBATCH --ntasks-per-node=1 +#SBATCH --job-name=large +#SBATCH --account=EUHPC_D30_139 +#SBATCH --partition=boost_usr_prod +#SBATCH --time=04:00:00 +#SBATCH --cpus-per-task=32 +#SBATCH --gres=gpu:1 +#SBATCH --mem=0 +#SBATCH --exclusive +#SBATCH --output=logs/%x_%j.out + +set -euo pipefail + +export OMP_NUM_THREADS=$SLURM_CPUS_PER_TASK +export OMP_PLACES=cores +export OMP_PROC_BIND=spread +# Booster nodes have no local disk and /tmp is a small tmpfs: spill to Lustre scratch. +export TMPDIR=$CINECA_SCRATCH/src-large/tmp +mkdir -p "$TMPDIR" + +REPO=$(git -C "$SLURM_SUBMIT_DIR" rev-parse --show-toplevel) +source "$REPO/.venv/bin/activate" +export CUPY_CACHE_DIR=$REPO/.cupy-cache +# The Booster driver (535) predates CUDA 13: load NVIDIA's forward-compat libcuda. +CUDA_COMPAT=${CUDA_COMPAT:-$REPO/.cuda-compat/usr/local/cuda-13.4/compat} +if [[ -d $CUDA_COMPAT ]]; then + export LD_LIBRARY_PATH=$CUDA_COMPAT${LD_LIBRARY_PATH:+:$LD_LIBRARY_PATH} +fi + +echo "NODE: $SLURMD_NODENAME, CPUS: $SLURM_CPUS_PER_TASK, ARGS: $*" +nvidia-smi --query-gpu=name,driver_version --format=csv,noheader + +srun python "$REPO/benches/large/bench_large.py" "$@" + +sacct --format=JobID,JobName,Elapsed,MaxRSS,MaxDiskRead,MaxDiskWrite,ExitCode \ + --units=G --jobs="$SLURM_JOB_ID" diff --git a/benches/primitives/README.md b/benches/primitives/README.md index b13682a..33bbc76 100644 --- a/benches/primitives/README.md +++ b/benches/primitives/README.md @@ -8,18 +8,41 @@ controlled bond dimension, and compare to a `quimb` reference. ## Leonardo The Leonardo [supercomputer](https://docs.hpc.cineca.it/general/getting_started.html) at CINECA is -a pre-exascale Tier-0 EuroHPC system, the 10th fastest in the world as of June 2025. In the -corresponding folder you can find the scripts to run the benchmarks on it, and some example -results. It is assumed that the user has built the `uv` virtual environment in the root -of the repository. This can be done by running: - -```bash -uv sync --all-groups --all-extras -``` +a pre-exascale Tier-0 EuroHPC system. The [`leonardo/`](leonardo/) folder holds the Slurm +script and the logs; see the [Leonardo section](../README.md#leonardo) for the environment. ### MPO-MPO benchmark -The benchmark contracts two length-`n_sites` matrix product operators (MPOs) whose initial bond dimension is `chi_out` (unless `chi_id` is specified for the second MPO), then compresses the resulting MPO back to bond dimension `chi_out`. One MPO is fully random; the other is an identity MPO perturbed by summing it with another random MPO with small elements, making the contraction highly compressible. The table below compares the performance of `quimb` versus `src` in terms of CPU time, total memory usage (maximum resident set size, which is the peak memory usage of the job), relative speedup, and accuracy. +The benchmark contracts two length-`n_sites` matrix product operators (MPOs) whose initial bond dimension is `chi_out` (unless `chi_id` is specified for the second MPO), then compresses the resulting MPO back to bond dimension `chi_out`. One MPO is fully random; the other is an identity MPO perturbed by summing it with another random MPO with small elements, making the contraction highly compressible. The table below compares `quimb` (CPU, `rsvd` compression) with `src` on the CPU and on one GPU: wall time of the contraction-compression, peak host memory of the job (MaxRSS), speedup over `quimb`, and relative distance to the uncompressed product. + +One Booster node (32 cores, A100-64GB), complex128 unless noted, `chi_id=4` at `chi=1000`: + +| Experiment | Library | Time (s) | MaxRSS (GB) | Speedup | Distance to reference | +|-------------------|---------------------|----------|-------------|---------|-----------------------| +| 50 sites chi=50 | quimb | 37.76 | 0.67 | | 2.6e-08 | +| | src CPU | 0.23 | 0.57 | 164x | 1.5e-08 | +| | src GPU | 0.57 | 1.01 | 66x | 1.5e-08 | +| 20 sites chi=100 | quimb | 17.86 | 0.73 | | 0.0 | +| | src CPU | 0.25 | 0.74 | 71x | 1.5e-08 | +| | src GPU | 0.46 | 1.22 | 39x | 2.1e-08 | +| 25 sites chi=1000 | quimb | 422.57 | 24.64 | | | +| | src CPU | 60.47 | 3.85 | 7x | | +| | src GPU | 2.32 | 3.42 | 182x | | +| | src GPU (complex64) | 1.90 | 2.20 | 222x | | +| 50 sites chi=1000 | quimb | 1040.72 | 50.02 | | | +| | src CPU | 137.37 | 6.83 | 7.6x | | +| | src GPU | 4.47 | 6.40 | 233x | | +| | src GPU (complex64) | 3.03 | 3.17 | 344x | | + +The distances sit at the floor of quimb's `distance` (about `sqrt(eps)` of the +dtype), so they show agreement rather than resolve the error. GPU times exclude a +one-off warm-up (CUDA context and kernel compilation, about 50 s on the first run +with an empty kernel cache). The CuPy pool peaks at 2.8 GB (25 sites) and 5.1 GB +(50 sites) at chi=1000 in complex128, half that in complex64. + +### Earlier results (v0.3.2) + +Before the refactoring, on Booster (32 CPUs) and DCGP (112 CPUs) nodes: | Experiment | CPUs | Library | Time (s) | MaxRSS (GB) | Speedup | Distance to reference | |-------------------|------|---------------|----------|-------------|---------|-----------------------| @@ -35,7 +58,7 @@ The benchmark contracts two length-`n_sites` matrix product operators (MPOs) who | 50 sites chi=1000 | 112 | quimb (rsvd) | 1187.82 | 49.47 | | | | chi_id=4 | | src | 150.11 | 7.01 | 8x | | -### MPO-MPO strong scaling +### MPO-MPO strong scaling (v0.3.2, DCGP) For this benchmark, we fix the problem size to `n_sites=25`, `chi=1000` and `chi_id=4`, and vary the number of CPUs in a single node. We tweak the OpenMP environment variables in the `run.sh` script to optimize performance, for instance `OMP_PLACES` and `OMP_PROC_BIND`, so the NUMA domain is taken into account. We do this both manually and automatically. diff --git a/benches/primitives/leonardo/bench_mpo_mpo.py b/benches/primitives/leonardo/bench_mpo_mpo.py index 9651171..3d3d090 100644 --- a/benches/primitives/leonardo/bench_mpo_mpo.py +++ b/benches/primitives/leonardo/bench_mpo_mpo.py @@ -18,42 +18,73 @@ app = cyclopts.App(help="Run SRC benchmark.") +def _log_distance( + H_ref: qtn.MatrixProductOperator, H: qtn.MatrixProductOperator +) -> None: + # quimb expands the norm of the difference, so cancellation floors this at + # about sqrt(eps) of the dtype. + distance = H_ref.distance(H) + logger.info( + " - Distance to reference: %s (relative %s)", + distance, + distance / abs(H_ref.norm()), + ) + + @app.default def main( n_sites: int = 50, chi_out: int = 20, + chi_id: int = 4, run: str = "src", compare: str = "yes", + device: str = "cpu", + dtype: str = "complex128", + seed: int = 0, ) -> None: """Main benchmarking function. Args: n_sites: Number of sites in the MPO chain. - chi_out: Output bond dimension for compression. + chi_out: Bond dimension of the random MPO and of the compressed output. + chi_id: Bond dimension of the perturbed identity MPO. run: Which library to run ('quimb', 'src'). compare: Whether to compare results to a reference ('yes', 'no'). + device: Where SRC runs ('cpu', 'gpu'); quimb always runs on the CPU. + dtype: Data type of the MPOs ('complex128', 'complex64'). + seed: Seed of the inputs and of the SRC sketch. """ phys_dim = 2 - array_type = np.complex128 + array_type = np.dtype(dtype) logger.info( - "benchmark_start: n_sites=%d, chi_out=%d, phys_dim=%d, dtype=complex128, " - "run=%s, compare=%s", + "benchmark_start: n_sites=%d, chi_out=%d, chi_id=%d, phys_dim=%d, dtype=%s, " + "run=%s, compare=%s, device=%s", n_sites, chi_out, + chi_id, phys_dim, + dtype, run, compare, + device, ) # Generate a random MPO and a perturbed identity MPO logger.info("Generating MPOs...") - H1 = qtn.MPO_rand(n_sites, bond_dim=chi_out, phys_dim=phys_dim, dtype=array_type) + H1 = qtn.MPO_rand( + n_sites, bond_dim=chi_out, phys_dim=phys_dim, dtype=array_type, seed=seed + ) + # The identity adds one to the bond: a near-identity, low-rank operator. H2 = qtn.MPO_identity( n_sites, phys_dim=phys_dim, dtype=array_type ) + 1e-8 * qtn.MPO_rand( - n_sites, bond_dim=3, phys_dim=phys_dim, dtype=array_type - ) # Total bond dimension is 4: a near-identity, low-rank operator + n_sites, + bond_dim=chi_id - 1, + phys_dim=phys_dim, + dtype=array_type, + seed=seed + 1, + ) # Quimb's contraction reference if compare == "yes": @@ -71,21 +102,33 @@ def main( tms = perf_counter_ns() - tms logger.info(" Quimb's contraction-compression took %s s", tms * 1e-9) if compare == "yes": - logger.info(" - Distance to reference: %s", H_ref.distance(H_quimb)) + _log_distance(H_ref, H_quimb) # SRC's contraction compression if run == "src": + if device == "gpu": + # CUDA context, cuBLAS handles and kernel compilation stay out of the timing. + logger.info("Warming up the GPU...") + warm = qtn.MPO_rand(4, bond_dim=2, phys_dim=phys_dim, dtype=array_type) + apply(warm.arrays, warm.arrays, chi_out=2, device=device) logger.info("Computing SRC's MPO-MPO contraction (with compression)...") tms = perf_counter_ns() # src_method takes and returns plain lists of site arrays; quimb is only # used here to build the inputs and to measure the distance. H_src = qtn.MatrixProductOperator( - apply(H1.arrays, H2.arrays, chi_out=chi_out, dtype=array_type) + apply(H1.arrays, H2.arrays, chi_out=chi_out, seed=seed, device=device) ) tms = perf_counter_ns() - tms logger.info(" SRC's contraction-compression took %s s", tms * 1e-9) + if device == "gpu": + import cupy # noqa: PLC0415 (optional GPU dependency) + + logger.info( + " - CuPy pool high-water mark: %.2f GB", + cupy.get_default_memory_pool().total_bytes() / 1e9, + ) if compare == "yes": - logger.info(" - Distance to reference: %s", H_ref.distance(H_src)) + _log_distance(H_ref, H_src) logger.info("benchmark_end") diff --git a/benches/primitives/leonardo/logs/mpo_20_100_quimb_cpu_59625646.out b/benches/primitives/leonardo/logs/mpo_20_100_quimb_cpu_59625646.out new file mode 100644 index 0000000..bea1898 --- /dev/null +++ b/benches/primitives/leonardo/logs/mpo_20_100_quimb_cpu_59625646.out @@ -0,0 +1,25 @@ +NODE: lrdn2278, CPUS: 32, ARGS: --n-sites 20 --chi-out 100 --compare yes --run quimb --device cpu +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +0: 2026-10-08 04:34:38,570 INFO __main__ benchmark_start: n_sites=20, chi_out=100, chi_id=4, phys_dim=2, dtype=complex128, run=quimb, compare=yes, device=cpu +0: 2026-10-08 04:34:38,570 INFO __main__ Generating MPOs... +0: 2026-10-08 04:34:38,780 INFO __main__ Computing reference contraction (no compression)... +0: 2026-10-08 04:34:38,911 INFO __main__ Reference contraction took 0.13113203 s +0: 2026-10-08 04:34:38,911 INFO __main__ Computing Quimb's MPO-MPO contraction (with compression)... +0: 2026-10-08 04:34:56,770 INFO __main__ Quimb's contraction-compression took 17.858333241 s +0: 2026-10-08 04:34:57,396 INFO __main__ - Distance to reference: 0.0 (relative 0.0) +0: 2026-10-08 04:34:57,396 INFO __main__ benchmark_end +JobID JobName NCPUS AveCPU Elapsed ExitCode +------------ ---------- ---------- ---------- ---------- -------- +59625646 mpo_20_10+ 32 00:00:53 0:0 +59625646.ba+ batch 32 00:00:53 0:0 +59625646.ex+ extern 32 00:00:53 0:0 +59625646.0 python 32 00:09:58 00:00:50 0:0 +JobID JobName MaxRSS MaxRSSNode +------------ ---------- ---------- ---------- +59625646 mpo_20_10+ +59625646.ba+ batch +59625646.ex+ extern +59625646.0 python 0.73G lrdn2278 diff --git a/benches/primitives/leonardo/logs/mpo_20_100_src_cpu_59626091.out b/benches/primitives/leonardo/logs/mpo_20_100_src_cpu_59626091.out new file mode 100644 index 0000000..4db44b7 --- /dev/null +++ b/benches/primitives/leonardo/logs/mpo_20_100_src_cpu_59626091.out @@ -0,0 +1,32 @@ +NODE: lrdn3007, CPUS: 32, ARGS: --n-sites 20 --chi-out 100 +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +0: 2026-10-07 15:51:38,906 INFO __main__ benchmark_start: n_sites=20, chi_out=100, chi_id=4, phys_dim=2, dtype=complex128, run=src, compare=yes, device=cpu +0: 2026-10-07 15:51:38,906 INFO __main__ Generating MPOs... +0: 2026-10-07 15:51:39,734 INFO __main__ Computing reference contraction (no compression)... +0: 2026-10-07 15:51:39,867 INFO __main__ Reference contraction took 0.13266686700000002 s +0: 2026-10-07 15:51:39,867 INFO __main__ Computing SRC's MPO-MPO contraction (with compression)... +0: 2026-10-07 15:51:39,867 DEBUG src_method.stack Starting SRC: n_sites=20, depth=2, output=mpo, device=numpy +0: 2026-10-07 15:51:39,879 DEBUG src_method._sweep SRC plan: prefetch=1, device peak=37439744 B, host peak=14089472 B, disk=0 B, scratch=None, tiers=['device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device'], batches (env, sketch, project)=[(96, 0, 96), (96, 96, 96), (96, 96, 96), (96, 96, 96), (96, 96, 96), (96, 96, 96), (96, 96, 96), (96, 96, 96), (96, 96, 96), (96, 96, 96), (96, 96, 96), (96, 96, 96), (96, 96, 96), (96, 96, 96), (96, 96, 96), (96, 96, 96), (96, 96, 96), (96, 96, 96), (96, 96, 96), (0, 96, 96)] +0: 2026-10-07 15:51:39,999 DEBUG src_method._sweep Left-to-right sweep: 0.120 s +0: 2026-10-07 15:51:40,113 DEBUG src_method._sweep Right-to-left sweep: 0.114 s +0: 2026-10-07 15:51:40,113 DEBUG src_method._sweep SRC stalls: sites 0.000 s, environments 0.000 s +0: 2026-10-07 15:51:40,114 DEBUG src_method._sweep Device pool: 0 B +0: 2026-10-07 15:51:40,114 DEBUG src_method.stack SRC complete +0: 2026-10-07 15:51:40,114 INFO __main__ SRC's contraction-compression took 0.247298909 s +0: 2026-10-07 15:51:40,612 INFO __main__ - Distance to reference: 1.4901161193847656e-08 (relative 1.4901161193847666e-08) +0: 2026-10-07 15:51:40,612 INFO __main__ benchmark_end +JobID JobName NCPUS AveCPU Elapsed ExitCode +------------ ---------- ---------- ---------- ---------- -------- +59626091 mpo_20_10+ 32 00:00:24 0:0 +59626091.ba+ batch 32 00:00:24 0:0 +59626091.ex+ extern 32 00:00:24 0:0 +59626091.0 python 32 00:00:20 0:0 +JobID JobName MaxRSS MaxRSSNode +------------ ---------- ---------- ---------- +59626091 mpo_20_10+ +59626091.ba+ batch +59626091.ex+ extern +59626091.0 python diff --git a/benches/primitives/leonardo/logs/mpo_20_100_src_gpu_59626976.out b/benches/primitives/leonardo/logs/mpo_20_100_src_gpu_59626976.out new file mode 100644 index 0000000..fdfcc88 --- /dev/null +++ b/benches/primitives/leonardo/logs/mpo_20_100_src_gpu_59626976.out @@ -0,0 +1,41 @@ +NODE: lrdn0473, CPUS: 32, ARGS: --n-sites 20 --chi-out 100 --device gpu +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +0: 2026-10-07 16:03:39,120 INFO __main__ benchmark_start: n_sites=20, chi_out=100, chi_id=4, phys_dim=2, dtype=complex128, run=src, compare=yes, device=gpu +0: 2026-10-07 16:03:39,120 INFO __main__ Generating MPOs... +0: 2026-10-07 16:03:39,298 INFO __main__ Computing reference contraction (no compression)... +0: 2026-10-07 16:03:39,441 INFO __main__ Reference contraction took 0.143532198 s +0: 2026-10-07 16:03:39,441 INFO __main__ Warming up the GPU... +0: 2026-10-07 16:03:56,097 DEBUG src_method.stack Starting SRC: n_sites=4, depth=2, output=mpo, device=cupy +0: 2026-10-07 16:03:57,085 DEBUG src_method._sweep SRC plan: prefetch=1, device peak=4224 B, host peak=2432 B, disk=0 B, scratch=None, tiers=['device', 'device', 'device', 'device'], batches (env, sketch, project)=[(2, 0, 2), (2, 2, 2), (2, 2, 2), (0, 2, 2)] +0: 2026-10-07 16:04:11,566 DEBUG src_method._sweep Left-to-right sweep: 14.230 s +0: 2026-10-07 16:04:23,004 DEBUG src_method._sweep Right-to-left sweep: 11.439 s +0: 2026-10-07 16:04:23,004 DEBUG src_method._sweep SRC stalls: sites 0.000 s, environments 0.000 s +0: 2026-10-07 16:04:23,004 DEBUG src_method._sweep Device pool: 48640 B +0: 2026-10-07 16:04:23,005 DEBUG src_method.stack SRC complete +0: 2026-10-07 16:04:23,005 INFO __main__ Computing SRC's MPO-MPO contraction (with compression)... +0: 2026-10-07 16:04:23,005 DEBUG src_method.stack Starting SRC: n_sites=20, depth=2, output=mpo, device=cupy +0: 2026-10-07 16:04:23,017 DEBUG src_method._sweep SRC plan: prefetch=1, device peak=23350272 B, host peak=14089472 B, disk=0 B, scratch=None, tiers=['device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device'], batches (env, sketch, project)=[(96, 0, 96), (96, 96, 96), (96, 96, 96), (96, 96, 96), (96, 96, 96), (96, 96, 96), (96, 96, 96), (96, 96, 96), (96, 96, 96), (96, 96, 96), (96, 96, 96), (96, 96, 96), (96, 96, 96), (96, 96, 96), (96, 96, 96), (96, 96, 96), (96, 96, 96), (96, 96, 96), (96, 96, 96), (0, 96, 96)] +0: 2026-10-07 16:04:23,148 DEBUG src_method._sweep Left-to-right sweep: 0.131 s +0: 2026-10-07 16:04:23,459 DEBUG src_method._sweep Right-to-left sweep: 0.311 s +0: 2026-10-07 16:04:23,460 DEBUG src_method._sweep SRC stalls: sites 0.000 s, environments 0.000 s +0: 2026-10-07 16:04:23,460 DEBUG src_method._sweep Device pool: 24724992 B +0: 2026-10-07 16:04:23,460 DEBUG src_method.stack SRC complete +0: 2026-10-07 16:04:23,460 INFO __main__ SRC's contraction-compression took 0.45528803300000004 s +0: 2026-10-07 16:04:23,460 INFO __main__ - CuPy pool high-water mark: 0.02 GB +0: 2026-10-07 16:04:23,978 INFO __main__ - Distance to reference: 2.1073424255447017e-08 (relative 2.107342425544703e-08) +0: 2026-10-07 16:04:23,979 INFO __main__ benchmark_end +JobID JobName NCPUS AveCPU Elapsed ExitCode +------------ ---------- ---------- ---------- ---------- -------- +59626976 mpo_20_10+ 32 00:01:14 0:0 +59626976.ba+ batch 32 00:01:14 0:0 +59626976.ex+ extern 32 00:01:14 0:0 +59626976.0 python 32 00:00:33 00:01:10 0:0 +JobID JobName MaxRSS MaxRSSNode +------------ ---------- ---------- ---------- +59626976 mpo_20_10+ +59626976.ba+ batch +59626976.ex+ extern +59626976.0 python 1.22G lrdn0473 diff --git a/benches/primitives/leonardo/logs/mpo_25_1000_quimb_cpu_59626117.out b/benches/primitives/leonardo/logs/mpo_25_1000_quimb_cpu_59626117.out new file mode 100644 index 0000000..be789e3 --- /dev/null +++ b/benches/primitives/leonardo/logs/mpo_25_1000_quimb_cpu_59626117.out @@ -0,0 +1,22 @@ +NODE: lrdn2680, CPUS: 32, ARGS: --n-sites 25 --chi-out 1000 --compare no --run quimb +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +0: 2026-10-07 15:52:23,165 INFO __main__ benchmark_start: n_sites=25, chi_out=1000, chi_id=4, phys_dim=2, dtype=complex128, run=quimb, compare=no, device=cpu +0: 2026-10-07 15:52:23,165 INFO __main__ Generating MPOs... +0: 2026-10-07 15:52:28,198 INFO __main__ Computing Quimb's MPO-MPO contraction (with compression)... +0: 2026-10-07 15:59:30,772 INFO __main__ Quimb's contraction-compression took 422.574737481 s +0: 2026-10-07 15:59:30,773 INFO __main__ benchmark_end +JobID JobName NCPUS AveCPU Elapsed ExitCode +------------ ---------- ---------- ---------- ---------- -------- +59626117 mpo_25_10+ 32 00:07:31 0:0 +59626117.ba+ batch 32 00:07:31 0:0 +59626117.ex+ extern 32 00:07:31 0:0 +59626117.0 python 32 03:17:05 00:07:27 0:0 +JobID JobName MaxRSS MaxRSSNode +------------ ---------- ---------- ---------- +59626117 mpo_25_10+ +59626117.ba+ batch +59626117.ex+ extern +59626117.0 python 24.64G lrdn2680 diff --git a/benches/primitives/leonardo/logs/mpo_25_1000_src_cpu_59626629.out b/benches/primitives/leonardo/logs/mpo_25_1000_src_cpu_59626629.out new file mode 100644 index 0000000..2ab181b --- /dev/null +++ b/benches/primitives/leonardo/logs/mpo_25_1000_src_cpu_59626629.out @@ -0,0 +1,29 @@ +NODE: lrdn2680, CPUS: 32, ARGS: --n-sites 25 --chi-out 1000 --compare no +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +0: 2026-10-07 16:00:06,993 INFO __main__ benchmark_start: n_sites=25, chi_out=1000, chi_id=4, phys_dim=2, dtype=complex128, run=src, compare=no, device=cpu +0: 2026-10-07 16:00:06,993 INFO __main__ Generating MPOs... +0: 2026-10-07 16:00:12,023 INFO __main__ Computing SRC's MPO-MPO contraction (with compression)... +0: 2026-10-07 16:00:12,023 DEBUG src_method.stack Starting SRC: n_sites=25, depth=2, output=mpo, device=numpy +0: 2026-10-07 16:00:12,045 DEBUG src_method._sweep SRC plan: prefetch=1, device peak=4409414144 B, host peak=1728067072 B, disk=0 B, scratch=None, tiers=['device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device'], batches (env, sketch, project)=[(992, 0, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (0, 992, 992)] +0: 2026-10-07 16:00:31,730 DEBUG src_method._sweep Left-to-right sweep: 19.685 s +0: 2026-10-07 16:01:12,496 DEBUG src_method._sweep Right-to-left sweep: 40.766 s +0: 2026-10-07 16:01:12,496 DEBUG src_method._sweep SRC stalls: sites 0.000 s, environments 0.000 s +0: 2026-10-07 16:01:12,496 DEBUG src_method._sweep Device pool: 0 B +0: 2026-10-07 16:01:12,497 DEBUG src_method.stack SRC complete +0: 2026-10-07 16:01:12,497 INFO __main__ SRC's contraction-compression took 60.47406152000001 s +0: 2026-10-07 16:01:12,497 INFO __main__ benchmark_end +JobID JobName NCPUS AveCPU Elapsed ExitCode +------------ ---------- ---------- ---------- ---------- -------- +59626629 mpo_25_10+ 32 00:01:30 0:0 +59626629.ba+ batch 32 00:01:30 0:0 +59626629.ex+ extern 32 00:01:30 0:0 +59626629.0 python 32 00:01:26 0:0 +JobID JobName MaxRSS MaxRSSNode +------------ ---------- ---------- ---------- +59626629 mpo_25_10+ +59626629.ba+ batch +59626629.ex+ extern +59626629.0 python 3.85G lrdn2680 diff --git a/benches/primitives/leonardo/logs/mpo_25_1000_src_gpu_59627165.out b/benches/primitives/leonardo/logs/mpo_25_1000_src_gpu_59627165.out new file mode 100644 index 0000000..8fa0f82 --- /dev/null +++ b/benches/primitives/leonardo/logs/mpo_25_1000_src_gpu_59627165.out @@ -0,0 +1,38 @@ +NODE: lrdn2796, CPUS: 32, ARGS: --n-sites 25 --chi-out 1000 --compare no --device gpu +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +0: 2026-10-07 16:05:28,788 INFO __main__ benchmark_start: n_sites=25, chi_out=1000, chi_id=4, phys_dim=2, dtype=complex128, run=src, compare=no, device=gpu +0: 2026-10-07 16:05:28,788 INFO __main__ Generating MPOs... +0: 2026-10-07 16:05:34,012 INFO __main__ Warming up the GPU... +0: 2026-10-07 16:05:54,418 DEBUG src_method.stack Starting SRC: n_sites=4, depth=2, output=mpo, device=cupy +0: 2026-10-07 16:05:55,873 DEBUG src_method._sweep SRC plan: prefetch=1, device peak=4224 B, host peak=2432 B, disk=0 B, scratch=None, tiers=['device', 'device', 'device', 'device'], batches (env, sketch, project)=[(2, 0, 2), (2, 2, 2), (2, 2, 2), (0, 2, 2)] +0: 2026-10-07 16:06:12,531 DEBUG src_method._sweep Left-to-right sweep: 16.401 s +0: 2026-10-07 16:06:21,998 DEBUG src_method._sweep Right-to-left sweep: 9.467 s +0: 2026-10-07 16:06:21,998 DEBUG src_method._sweep SRC stalls: sites 0.000 s, environments 0.000 s +0: 2026-10-07 16:06:21,998 DEBUG src_method._sweep Device pool: 48640 B +0: 2026-10-07 16:06:21,998 DEBUG src_method.stack SRC complete +0: 2026-10-07 16:06:21,998 INFO __main__ Computing SRC's MPO-MPO contraction (with compression)... +0: 2026-10-07 16:06:21,998 DEBUG src_method.stack Starting SRC: n_sites=25, depth=2, output=mpo, device=cupy +0: 2026-10-07 16:06:22,015 DEBUG src_method._sweep SRC plan: prefetch=1, device peak=2681347072 B, host peak=1728067072 B, disk=0 B, scratch=None, tiers=['device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device'], batches (env, sketch, project)=[(992, 0, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (0, 992, 992)] +0: 2026-10-07 16:06:22,462 DEBUG src_method._sweep Left-to-right sweep: 0.406 s +0: 2026-10-07 16:06:24,322 DEBUG src_method._sweep Right-to-left sweep: 1.859 s +0: 2026-10-07 16:06:24,322 DEBUG src_method._sweep SRC stalls: sites 0.188 s, environments 0.000 s +0: 2026-10-07 16:06:24,322 DEBUG src_method._sweep Device pool: 2795244544 B +0: 2026-10-07 16:06:24,322 DEBUG src_method.stack SRC complete +0: 2026-10-07 16:06:24,323 INFO __main__ SRC's contraction-compression took 2.3248113050000003 s +0: 2026-10-07 16:06:24,323 INFO __main__ - CuPy pool high-water mark: 2.80 GB +0: 2026-10-07 16:06:24,323 INFO __main__ benchmark_end +JobID JobName NCPUS AveCPU Elapsed ExitCode +------------ ---------- ---------- ---------- ---------- -------- +59627165 mpo_25_10+ 32 00:01:24 0:0 +59627165.ba+ batch 32 00:01:24 0:0 +59627165.ex+ extern 32 00:01:24 0:0 +59627165.0 python 32 00:01:21 0:0 +JobID JobName MaxRSS MaxRSSNode +------------ ---------- ---------- ---------- +59627165 mpo_25_10+ +59627165.ba+ batch +59627165.ex+ extern +59627165.0 python diff --git a/benches/primitives/leonardo/logs/mpo_25_1000_src_gpu_c64_59627591.out b/benches/primitives/leonardo/logs/mpo_25_1000_src_gpu_c64_59627591.out new file mode 100644 index 0000000..9e3d140 --- /dev/null +++ b/benches/primitives/leonardo/logs/mpo_25_1000_src_gpu_c64_59627591.out @@ -0,0 +1,38 @@ +NODE: lrdn0096, CPUS: 32, ARGS: --n-sites 25 --chi-out 1000 --compare no --device gpu --dtype complex64 +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +0: 2026-10-07 16:09:10,881 INFO __main__ benchmark_start: n_sites=25, chi_out=1000, chi_id=4, phys_dim=2, dtype=complex64, run=src, compare=no, device=gpu +0: 2026-10-07 16:09:10,881 INFO __main__ Generating MPOs... +0: 2026-10-07 16:09:13,471 INFO __main__ Warming up the GPU... +0: 2026-10-07 16:09:34,542 DEBUG src_method.stack Starting SRC: n_sites=4, depth=2, output=mpo, device=cupy +0: 2026-10-07 16:09:35,964 DEBUG src_method._sweep SRC plan: prefetch=1, device peak=2112 B, host peak=1216 B, disk=0 B, scratch=None, tiers=['device', 'device', 'device', 'device'], batches (env, sketch, project)=[(2, 0, 2), (2, 2, 2), (2, 2, 2), (0, 2, 2)] +0: 2026-10-07 16:09:57,185 DEBUG src_method._sweep Left-to-right sweep: 20.986 s +0: 2026-10-07 16:10:09,659 DEBUG src_method._sweep Right-to-left sweep: 12.474 s +0: 2026-10-07 16:10:09,659 DEBUG src_method._sweep SRC stalls: sites 0.000 s, environments 0.000 s +0: 2026-10-07 16:10:09,659 DEBUG src_method._sweep Device pool: 27648 B +0: 2026-10-07 16:10:09,659 DEBUG src_method.stack SRC complete +0: 2026-10-07 16:10:09,659 INFO __main__ Computing SRC's MPO-MPO contraction (with compression)... +0: 2026-10-07 16:10:09,659 DEBUG src_method.stack Starting SRC: n_sites=25, depth=2, output=mpo, device=cupy +0: 2026-10-07 16:10:09,676 DEBUG src_method._sweep SRC plan: prefetch=1, device peak=1340673536 B, host peak=864033536 B, disk=0 B, scratch=None, tiers=['device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device'], batches (env, sketch, project)=[(992, 0, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (0, 992, 992)] +0: 2026-10-07 16:10:10,116 DEBUG src_method._sweep Left-to-right sweep: 0.418 s +0: 2026-10-07 16:10:11,554 DEBUG src_method._sweep Right-to-left sweep: 1.438 s +0: 2026-10-07 16:10:11,554 DEBUG src_method._sweep SRC stalls: sites 0.146 s, environments 0.000 s +0: 2026-10-07 16:10:11,554 DEBUG src_method._sweep Device pool: 1397658112 B +0: 2026-10-07 16:10:11,554 DEBUG src_method.stack SRC complete +0: 2026-10-07 16:10:11,555 INFO __main__ SRC's contraction-compression took 1.895542328 s +0: 2026-10-07 16:10:11,555 INFO __main__ - CuPy pool high-water mark: 1.40 GB +0: 2026-10-07 16:10:11,555 INFO __main__ benchmark_end +JobID JobName NCPUS AveCPU Elapsed ExitCode +------------ ---------- ---------- ---------- ---------- -------- +59627591 mpo_25_10+ 32 00:01:32 0:0 +59627591.ba+ batch 32 00:01:32 0:0 +59627591.ex+ extern 32 00:01:32 0:0 +59627591.0 python 32 00:00:50 00:01:26 0:0 +JobID JobName MaxRSS MaxRSSNode +------------ ---------- ---------- ---------- +59627591 mpo_25_10+ +59627591.ba+ batch +59627591.ex+ extern +59627591.0 python 2.20G lrdn0096 diff --git a/benches/primitives/leonardo/logs/mpo_50_1000_quimb_cpu_59625936.out b/benches/primitives/leonardo/logs/mpo_50_1000_quimb_cpu_59625936.out new file mode 100644 index 0000000..5d6d519 --- /dev/null +++ b/benches/primitives/leonardo/logs/mpo_50_1000_quimb_cpu_59625936.out @@ -0,0 +1,22 @@ +NODE: lrdn2427, CPUS: 32, ARGS: --n-sites 50 --chi-out 1000 --compare no --run quimb +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +0: 2026-10-08 04:35:27,338 INFO __main__ benchmark_start: n_sites=50, chi_out=1000, chi_id=4, phys_dim=2, dtype=complex128, run=quimb, compare=no, device=cpu +0: 2026-10-08 04:35:27,338 INFO __main__ Generating MPOs... +0: 2026-10-08 04:35:38,108 INFO __main__ Computing Quimb's MPO-MPO contraction (with compression)... +0: 2026-10-08 04:52:58,831 INFO __main__ Quimb's contraction-compression took 1040.7233631410002 s +0: 2026-10-08 04:52:58,831 INFO __main__ benchmark_end +JobID JobName NCPUS AveCPU Elapsed ExitCode +------------ ---------- ---------- ---------- ---------- -------- +59625936 mpo_50_10+ 32 00:18:05 0:0 +59625936.ba+ batch 32 00:18:05 0:0 +59625936.ex+ extern 32 00:18:05 0:0 +59625936.0 python 32 00:18:02 0:0 +JobID JobName MaxRSS MaxRSSNode +------------ ---------- ---------- ---------- +59625936 mpo_50_10+ +59625936.ba+ batch +59625936.ex+ extern +59625936.0 python diff --git a/benches/primitives/leonardo/logs/mpo_50_1000_src_cpu_59626808.out b/benches/primitives/leonardo/logs/mpo_50_1000_src_cpu_59626808.out new file mode 100644 index 0000000..593270b --- /dev/null +++ b/benches/primitives/leonardo/logs/mpo_50_1000_src_cpu_59626808.out @@ -0,0 +1,30 @@ +NODE: lrdn2517, CPUS: 32, ARGS: --n-sites 50 --chi-out 1000 --compare no +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +0: 2026-10-07 16:02:04,119 INFO __main__ benchmark_start: n_sites=50, chi_out=1000, chi_id=4, phys_dim=2, dtype=complex128, run=src, compare=no, device=cpu +0: 2026-10-07 16:02:04,119 INFO __main__ Generating MPOs... +0: 2026-10-07 16:02:15,292 INFO __main__ Computing SRC's MPO-MPO contraction (with compression)... +0: 2026-10-07 16:02:15,293 DEBUG src_method.stack Starting SRC: n_sites=50, depth=2, output=mpo, device=numpy +0: 2026-10-07 16:02:15,328 DEBUG src_method._sweep SRC plan: prefetch=1, device peak=7609414144 B, host peak=3328067072 B, disk=0 B, scratch=None, tiers=['device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device'], batches (env, sketch, project)=[(992, 0, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (9 +0: 92, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (0, 992, 992)] +0: 2026-10-07 16:02:56,795 DEBUG src_method._sweep Left-to-right sweep: 41.467 s +0: 2026-10-07 16:04:32,664 DEBUG src_method._sweep Right-to-left sweep: 95.868 s +0: 2026-10-07 16:04:32,664 DEBUG src_method._sweep SRC stalls: sites 0.001 s, environments 0.000 s +0: 2026-10-07 16:04:32,664 DEBUG src_method._sweep Device pool: 0 B +0: 2026-10-07 16:04:32,665 DEBUG src_method.stack SRC complete +0: 2026-10-07 16:04:32,665 INFO __main__ SRC's contraction-compression took 137.37281321700002 s +0: 2026-10-07 16:04:32,665 INFO __main__ benchmark_end +JobID JobName NCPUS AveCPU Elapsed ExitCode +------------ ---------- ---------- ---------- ---------- -------- +59626808 mpo_50_10+ 32 00:02:52 0:0 +59626808.ba+ batch 32 00:02:52 0:0 +59626808.ex+ extern 32 00:02:52 0:0 +59626808.0 python 32 00:50:14 00:02:48 0:0 +JobID JobName MaxRSS MaxRSSNode +------------ ---------- ---------- ---------- +59626808 mpo_50_10+ +59626808.ba+ batch +59626808.ex+ extern +59626808.0 python 6.83G lrdn2517 diff --git a/benches/primitives/leonardo/logs/mpo_50_1000_src_gpu_59627373.out b/benches/primitives/leonardo/logs/mpo_50_1000_src_gpu_59627373.out new file mode 100644 index 0000000..312b3f2 --- /dev/null +++ b/benches/primitives/leonardo/logs/mpo_50_1000_src_gpu_59627373.out @@ -0,0 +1,39 @@ +NODE: lrdn3032, CPUS: 32, ARGS: --n-sites 50 --chi-out 1000 --compare no --device gpu +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +0: 2026-10-07 16:07:18,364 INFO __main__ benchmark_start: n_sites=50, chi_out=1000, chi_id=4, phys_dim=2, dtype=complex128, run=src, compare=no, device=gpu +0: 2026-10-07 16:07:18,364 INFO __main__ Generating MPOs... +0: 2026-10-07 16:07:29,145 INFO __main__ Warming up the GPU... +0: 2026-10-07 16:07:42,169 DEBUG src_method.stack Starting SRC: n_sites=4, depth=2, output=mpo, device=cupy +0: 2026-10-07 16:07:43,471 DEBUG src_method._sweep SRC plan: prefetch=1, device peak=4224 B, host peak=2432 B, disk=0 B, scratch=None, tiers=['device', 'device', 'device', 'device'], batches (env, sketch, project)=[(2, 0, 2), (2, 2, 2), (2, 2, 2), (0, 2, 2)] +0: 2026-10-07 16:07:59,920 DEBUG src_method._sweep Left-to-right sweep: 16.131 s +0: 2026-10-07 16:08:11,732 DEBUG src_method._sweep Right-to-left sweep: 11.812 s +0: 2026-10-07 16:08:11,732 DEBUG src_method._sweep SRC stalls: sites 0.000 s, environments 0.000 s +0: 2026-10-07 16:08:11,732 DEBUG src_method._sweep Device pool: 48640 B +0: 2026-10-07 16:08:11,733 DEBUG src_method.stack SRC complete +0: 2026-10-07 16:08:11,733 INFO __main__ Computing SRC's MPO-MPO contraction (with compression)... +0: 2026-10-07 16:08:11,733 DEBUG src_method.stack Starting SRC: n_sites=50, depth=2, output=mpo, device=cupy +0: 2026-10-07 16:08:11,752 DEBUG src_method._sweep SRC plan: prefetch=1, device peak=4281347072 B, host peak=3328067072 B, disk=0 B, scratch=None, tiers=['device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device'], batches (env, sketch, project)=[(992, 0, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (9 +0: 92, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (0, 992, 992)] +0: 2026-10-07 16:08:12,392 DEBUG src_method._sweep Left-to-right sweep: 0.604 s +0: 2026-10-07 16:08:16,200 DEBUG src_method._sweep Right-to-left sweep: 3.809 s +0: 2026-10-07 16:08:16,200 DEBUG src_method._sweep SRC stalls: sites 0.398 s, environments 0.000 s +0: 2026-10-07 16:08:16,200 DEBUG src_method._sweep Device pool: 5080812544 B +0: 2026-10-07 16:08:16,201 DEBUG src_method.stack SRC complete +0: 2026-10-07 16:08:16,201 INFO __main__ SRC's contraction-compression took 4.468574902 s +0: 2026-10-07 16:08:16,201 INFO __main__ - CuPy pool high-water mark: 5.08 GB +0: 2026-10-07 16:08:16,201 INFO __main__ benchmark_end +JobID JobName NCPUS AveCPU Elapsed ExitCode +------------ ---------- ---------- ---------- ---------- -------- +59627373 mpo_50_10+ 32 00:01:30 0:0 +59627373.ba+ batch 32 00:01:30 0:0 +59627373.ex+ extern 32 00:01:30 0:0 +59627373.0 python 32 00:05:16 00:01:26 0:0 +JobID JobName MaxRSS MaxRSSNode +------------ ---------- ---------- ---------- +59627373 mpo_50_10+ +59627373.ba+ batch +59627373.ex+ extern +59627373.0 python 6.40G lrdn3032 diff --git a/benches/primitives/leonardo/logs/mpo_50_1000_src_gpu_c64_59627813.out b/benches/primitives/leonardo/logs/mpo_50_1000_src_gpu_c64_59627813.out new file mode 100644 index 0000000..58d0fc1 --- /dev/null +++ b/benches/primitives/leonardo/logs/mpo_50_1000_src_gpu_c64_59627813.out @@ -0,0 +1,45 @@ +NODE: lrdn0096, CPUS: 32, ARGS: --n-sites 50 --chi-out 1000 --compare no --device gpu --dtype complex64 +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +0: 2026-10-07 16:10:38,169 INFO __main__ benchmark_start: n_sites=50, chi_out=1000, chi_id=4, phys_dim=2, dtype=complex64, run=src, compare=no, device=gpu +0: 2026-10-07 16:10:38,169 INFO __main__ Generating MPOs... +0: /leonardo_work/ALQ_prod2526/rpanades/src-method/.venv/lib/python3.14/site-packages/autoray/autoray.py:96: RuntimeWarning: overflow encountered in matmul +0: return func(*args, **kwargs) +0: /leonardo_work/ALQ_prod2526/rpanades/src-method/.venv/lib/python3.14/site-packages/autoray/autoray.py:96: RuntimeWarning: invalid value encountered in matmul +0: return func(*args, **kwargs) +0: /leonardo_work/ALQ_prod2526/rpanades/src-method/.venv/lib/python3.14/site-packages/quimb/tensor/tensor_core.py:5119: RuntimeWarning: invalid value encountered in scalar power +0: return self.multiply_(other**-1) +0: 2026-10-07 16:10:42,593 INFO __main__ Warming up the GPU... +0: 2026-10-07 16:10:42,861 DEBUG src_method.stack Starting SRC: n_sites=4, depth=2, output=mpo, device=cupy +0: 2026-10-07 16:10:42,960 DEBUG src_method._sweep SRC plan: prefetch=1, device peak=2112 B, host peak=1216 B, disk=0 B, scratch=None, tiers=['device', 'device', 'device', 'device'], batches (env, sketch, project)=[(2, 0, 2), (2, 2, 2), (2, 2, 2), (0, 2, 2)] +0: 2026-10-07 16:10:43,041 DEBUG src_method._sweep Left-to-right sweep: 0.049 s +0: 2026-10-07 16:10:43,069 DEBUG src_method._sweep Right-to-left sweep: 0.027 s +0: 2026-10-07 16:10:43,069 DEBUG src_method._sweep SRC stalls: sites 0.000 s, environments 0.000 s +0: 2026-10-07 16:10:43,069 DEBUG src_method._sweep Device pool: 27648 B +0: 2026-10-07 16:10:43,069 DEBUG src_method.stack SRC complete +0: 2026-10-07 16:10:43,069 INFO __main__ Computing SRC's MPO-MPO contraction (with compression)... +0: 2026-10-07 16:10:43,069 DEBUG src_method.stack Starting SRC: n_sites=50, depth=2, output=mpo, device=cupy +0: 2026-10-07 16:10:43,088 DEBUG src_method._sweep SRC plan: prefetch=1, device peak=2140673536 B, host peak=1664033536 B, disk=0 B, scratch=None, tiers=['device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device'], batches (env, sketch, project)=[(992, 0, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (9 +0: 92, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (992, 992, 992), (0, 992, 992)] +0: 2026-10-07 16:10:43,492 DEBUG src_method._sweep Left-to-right sweep: 0.385 s +0: 2026-10-07 16:10:46,100 DEBUG src_method._sweep Right-to-left sweep: 2.608 s +0: 2026-10-07 16:10:46,100 DEBUG src_method._sweep SRC stalls: sites 0.326 s, environments 0.000 s +0: 2026-10-07 16:10:46,100 DEBUG src_method._sweep Device pool: 2540442112 B +0: 2026-10-07 16:10:46,101 DEBUG src_method.stack SRC complete +0: 2026-10-07 16:10:46,101 INFO __main__ SRC's contraction-compression took 3.032048248 s +0: 2026-10-07 16:10:46,101 INFO __main__ - CuPy pool high-water mark: 2.54 GB +0: 2026-10-07 16:10:46,101 INFO __main__ benchmark_end +JobID JobName NCPUS AveCPU Elapsed ExitCode +------------ ---------- ---------- ---------- ---------- -------- +59627813 mpo_50_10+ 32 00:00:13 0:0 +59627813.ba+ batch 32 00:00:13 0:0 +59627813.ex+ extern 32 00:00:13 0:0 +59627813.0 python 32 00:00:10 0:0 +JobID JobName MaxRSS MaxRSSNode +------------ ---------- ---------- ---------- +59627813 mpo_50_10+ +59627813.ba+ batch +59627813.ex+ extern +59627813.0 python diff --git a/benches/primitives/leonardo/logs/mpo_50_50_quimb_cpu_59625938.out b/benches/primitives/leonardo/logs/mpo_50_50_quimb_cpu_59625938.out new file mode 100644 index 0000000..135acb3 --- /dev/null +++ b/benches/primitives/leonardo/logs/mpo_50_50_quimb_cpu_59625938.out @@ -0,0 +1,25 @@ +NODE: lrdn2926, CPUS: 32, ARGS: --n-sites 50 --chi-out 50 --run quimb +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +0: 2026-10-07 15:49:42,313 INFO __main__ benchmark_start: n_sites=50, chi_out=50, chi_id=4, phys_dim=2, dtype=complex128, run=quimb, compare=yes, device=cpu +0: 2026-10-07 15:49:42,314 INFO __main__ Generating MPOs... +0: 2026-10-07 15:49:42,502 INFO __main__ Computing reference contraction (no compression)... +0: 2026-10-07 15:49:42,557 INFO __main__ Reference contraction took 0.05452135 s +0: 2026-10-07 15:49:42,557 INFO __main__ Computing Quimb's MPO-MPO contraction (with compression)... +0: 2026-10-07 15:50:20,313 INFO __main__ Quimb's contraction-compression took 37.755962594 s +0: 2026-10-07 15:50:20,696 INFO __main__ - Distance to reference: 2.5809568279517847e-08 (relative 2.5809568279517943e-08) +0: 2026-10-07 15:50:20,696 INFO __main__ benchmark_end +JobID JobName NCPUS AveCPU Elapsed ExitCode +------------ ---------- ---------- ---------- ---------- -------- +59625938 mpo_50_50+ 32 00:01:04 0:0 +59625938.ba+ batch 32 00:01:04 0:0 +59625938.ex+ extern 32 00:01:04 0:0 +59625938.0 python 32 00:16:51 00:01:00 0:0 +JobID JobName MaxRSS MaxRSSNode +------------ ---------- ---------- ---------- +59625938 mpo_50_50+ +59625938.ba+ batch +59625938.ex+ extern +59625938.0 python 0.67G lrdn2926 diff --git a/benches/primitives/leonardo/logs/mpo_50_50_src_cpu_59626058.out b/benches/primitives/leonardo/logs/mpo_50_50_src_cpu_59626058.out new file mode 100644 index 0000000..a2b35e9 --- /dev/null +++ b/benches/primitives/leonardo/logs/mpo_50_50_src_cpu_59626058.out @@ -0,0 +1,33 @@ +NODE: lrdn1912, CPUS: 32, ARGS: --n-sites 50 --chi-out 50 +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +0: 2026-10-07 15:50:58,923 INFO __main__ benchmark_start: n_sites=50, chi_out=50, chi_id=4, phys_dim=2, dtype=complex128, run=src, compare=yes, device=cpu +0: 2026-10-07 15:50:58,923 INFO __main__ Generating MPOs... +0: 2026-10-07 15:50:59,079 INFO __main__ Computing reference contraction (no compression)... +0: 2026-10-07 15:50:59,134 INFO __main__ Reference contraction took 0.054777736 s +0: 2026-10-07 15:50:59,134 INFO __main__ Computing SRC's MPO-MPO contraction (with compression)... +0: 2026-10-07 15:50:59,134 DEBUG src_method.stack Starting SRC: n_sites=50, depth=2, output=mpo, device=numpy +0: 2026-10-07 15:50:59,149 DEBUG src_method._sweep SRC plan: prefetch=1, device peak=18300544 B, host peak=8326272 B, disk=0 B, scratch=None, tiers=['device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device'], batches (env, sketch, project)=[(32, 0, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 3 +0: 2), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (0, 32, 32)] +0: 2026-10-07 15:50:59,239 DEBUG src_method._sweep Left-to-right sweep: 0.090 s +0: 2026-10-07 15:50:59,364 DEBUG src_method._sweep Right-to-left sweep: 0.125 s +0: 2026-10-07 15:50:59,364 DEBUG src_method._sweep SRC stalls: sites 0.000 s, environments 0.000 s +0: 2026-10-07 15:50:59,364 DEBUG src_method._sweep Device pool: 0 B +0: 2026-10-07 15:50:59,365 DEBUG src_method.stack SRC complete +0: 2026-10-07 15:50:59,366 INFO __main__ SRC's contraction-compression took 0.23186999500000002 s +0: 2026-10-07 15:50:59,711 INFO __main__ - Distance to reference: 1.4901161193847656e-08 (relative 1.4901161193847712e-08) +0: 2026-10-07 15:50:59,711 INFO __main__ benchmark_end +JobID JobName NCPUS AveCPU Elapsed ExitCode +------------ ---------- ---------- ---------- ---------- -------- +59626058 mpo_50_50+ 32 00:00:25 0:0 +59626058.ba+ batch 32 00:00:25 0:0 +59626058.ex+ extern 32 00:00:25 0:0 +59626058.0 python 32 00:00:21 0:0 +JobID JobName MaxRSS MaxRSSNode +------------ ---------- ---------- ---------- +59626058 mpo_50_50+ +59626058.ba+ batch +59626058.ex+ extern +59626058.0 python 0.57G lrdn1912 diff --git a/benches/primitives/leonardo/logs/mpo_50_50_src_gpu_59626787.out b/benches/primitives/leonardo/logs/mpo_50_50_src_gpu_59626787.out new file mode 100644 index 0000000..01d3916 --- /dev/null +++ b/benches/primitives/leonardo/logs/mpo_50_50_src_gpu_59626787.out @@ -0,0 +1,42 @@ +NODE: lrdn2682, CPUS: 32, ARGS: --n-sites 50 --chi-out 50 --device gpu +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +NVIDIA A100-SXM-64GB, 535.274.02 +0: 2026-10-07 16:02:04,119 INFO __main__ benchmark_start: n_sites=50, chi_out=50, chi_id=4, phys_dim=2, dtype=complex128, run=src, compare=yes, device=gpu +0: 2026-10-07 16:02:04,119 INFO __main__ Generating MPOs... +0: 2026-10-07 16:02:04,239 INFO __main__ Computing reference contraction (no compression)... +0: 2026-10-07 16:02:04,294 INFO __main__ Reference contraction took 0.054615671000000005 s +0: 2026-10-07 16:02:04,294 INFO __main__ Warming up the GPU... +0: 2026-10-07 16:02:26,535 DEBUG src_method.stack Starting SRC: n_sites=4, depth=2, output=mpo, device=cupy +0: 2026-10-07 16:02:27,853 DEBUG src_method._sweep SRC plan: prefetch=1, device peak=4224 B, host peak=2432 B, disk=0 B, scratch=None, tiers=['device', 'device', 'device', 'device'], batches (env, sketch, project)=[(2, 0, 2), (2, 2, 2), (2, 2, 2), (0, 2, 2)] +0: 2026-10-07 16:02:50,464 DEBUG src_method._sweep Left-to-right sweep: 22.354 s +0: 2026-10-07 16:03:00,804 DEBUG src_method._sweep Right-to-left sweep: 10.340 s +0: 2026-10-07 16:03:00,804 DEBUG src_method._sweep SRC stalls: sites 0.000 s, environments 0.000 s +0: 2026-10-07 16:03:00,804 DEBUG src_method._sweep Device pool: 48640 B +0: 2026-10-07 16:03:00,804 DEBUG src_method.stack SRC complete +0: 2026-10-07 16:03:00,804 INFO __main__ Computing SRC's MPO-MPO contraction (with compression)... +0: 2026-10-07 16:03:00,804 DEBUG src_method.stack Starting SRC: n_sites=50, depth=2, output=mpo, device=cupy +0: 2026-10-07 16:03:00,816 DEBUG src_method._sweep SRC plan: prefetch=1, device peak=9974272 B, host peak=8326272 B, disk=0 B, scratch=None, tiers=['device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device', 'device'], batches (env, sketch, project)=[(32, 0, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32 +0: ), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (32, 32, 32), (0, 32, 32)] +0: 2026-10-07 16:03:00,896 DEBUG src_method._sweep Left-to-right sweep: 0.080 s +0: 2026-10-07 16:03:01,376 DEBUG src_method._sweep Right-to-left sweep: 0.480 s +0: 2026-10-07 16:03:01,376 DEBUG src_method._sweep SRC stalls: sites 0.000 s, environments 0.000 s +0: 2026-10-07 16:03:01,376 DEBUG src_method._sweep Device pool: 11751424 B +0: 2026-10-07 16:03:01,376 DEBUG src_method.stack SRC complete +0: 2026-10-07 16:03:01,376 INFO __main__ SRC's contraction-compression took 0.572353006 s +0: 2026-10-07 16:03:01,376 INFO __main__ - CuPy pool high-water mark: 0.01 GB +0: 2026-10-07 16:03:01,721 INFO __main__ - Distance to reference: 1.4901161193847656e-08 (relative 1.4901161193847712e-08) +0: 2026-10-07 16:03:01,721 INFO __main__ benchmark_end +JobID JobName NCPUS AveCPU Elapsed ExitCode +------------ ---------- ---------- ---------- ---------- -------- +59626787 mpo_50_50+ 32 00:01:26 0:0 +59626787.ba+ batch 32 00:01:26 0:0 +59626787.ex+ extern 32 00:01:26 0:0 +59626787.0 python 32 00:01:22 0:0 +JobID JobName MaxRSS MaxRSSNode +------------ ---------- ---------- ---------- +59626787 mpo_50_50+ +59626787.ba+ batch +59626787.ex+ extern +59626787.0 python diff --git a/benches/primitives/leonardo/run.sh b/benches/primitives/leonardo/run.sh index 28962b1..e5d46ea 100755 --- a/benches/primitives/leonardo/run.sh +++ b/benches/primitives/leonardo/run.sh @@ -1,63 +1,38 @@ #!/usr/bin/env -S bash -l +# Usage, from this folder: sbatch run.sh [bench_mpo_mpo.py options] +# e.g. sbatch run.sh --n-sites 25 --chi-out 1000 --compare no --device gpu +# Add --qos=boost_qos_dbg (30 min, 2 nodes) to sbatch for short test runs. #SBATCH --nodes=1 #SBATCH --ntasks-per-node=1 -#SBATCH --job-name=src_bench -#SBATCH --account=ALQ_prod2526_0 -#SBATCH --partition=dcgp_usr_prod -#SBATCH --time=01:00:00 -#SBATCH --cpus-per-task=28 -#SBATCH --mem=440G +#SBATCH --job-name=mpo_mpo +#SBATCH --account=EUHPC_D30_139 +#SBATCH --partition=boost_usr_prod +#SBATCH --time=02:00:00 +#SBATCH --cpus-per-task=32 +#SBATCH --gres=gpu:1 +#SBATCH --mem=0 #SBATCH --exclusive -#SBATCH --output=logs/1000_%j.out +#SBATCH --output=logs/%x_%j.out + +set -euo pipefail -# Automatic OpenMP binding export OMP_NUM_THREADS=$SLURM_CPUS_PER_TASK export OMP_PLACES=cores export OMP_PROC_BIND=spread -# Python uv environment -source ../../.venv/bin/activate - -# Problem size -NSITES=25 -CHI=1000 -COMPARE="no" # "yes", "no" -RUN_TYPE="src" # "src", "quimb", "both" - -# Report -echo "OPENMP THREADS: $OMP_NUM_THREADS" -echo "NODE: $SLURMD_NODENAME" -echo "CPUS: $SLURM_CPUS_PER_TASK" - -# Run types -if [ "$RUN_TYPE" = "both" ]; then - srun --verbose --label python bench_mpo_mpo.py --run quimb --compare $COMPARE --n_sites $NSITES --chi_out $CHI - srun --verbose --label python bench_mpo_mpo.py --run src --compare $COMPARE --n_sites $NSITES --chi_out $CHI -elif [ "$RUN_TYPE" = "src" ]; then - srun --verbose --label python bench_mpo_mpo.py --run src --compare $COMPARE --n_sites $NSITES --chi_out $CHI -elif [ "$RUN_TYPE" = "quimb" ]; then - srun --verbose --label python bench_mpo_mpo.py --run quimb --compare $COMPARE --n_sites $NSITES --chi_out $CHI -else - echo "Unknown run type: $RUN_TYPE" - exit 1 +REPO=$(git -C "$SLURM_SUBMIT_DIR" rev-parse --show-toplevel) +source "$REPO/.venv/bin/activate" +export CUPY_CACHE_DIR=$REPO/.cupy-cache +# The Booster driver (535) predates CUDA 13: load NVIDIA's forward-compat libcuda. +CUDA_COMPAT=${CUDA_COMPAT:-$REPO/.cuda-compat/usr/local/cuda-13.4/compat} +if [[ -d $CUDA_COMPAT ]]; then + export LD_LIBRARY_PATH=$CUDA_COMPAT${LD_LIBRARY_PATH:+:$LD_LIBRARY_PATH} fi -# Report stats -echo -e "\n\n#------------------------#" -echo "Task and CPU usage stats:" -sacct --format=JobID,JobName,NCPUS,NNodes,NTasks,AveCPU,MinCPU,MinCPUNode,MinCPUTask,Elapsed,ExitCode --jobs="${SLURM_JOBID}" - -echo "Memory usage stats:" -sacct --format=JobID,JobName,AveRSS%-50,MaxRSS%-50,MaxRSSNode,MaxRSSTask,AvePages,MaxPages,MaxPagesNode,MaxPagesTask --units=G --jobs="${SLURM_JOBID}" - -echo "Disk usage stats:" -sacct --format=JobID,JobName,AveDiskRead,MaxDiskRead,MaxDiskReadNode,MaxDiskReadTask,AveDiskWrite,MaxDiskWrite,MaxDiskWriteNode,MaxDiskWriteTask --units=G --jobs="${SLURM_JOBID}" - -echo "Trackable resources usage stats:" -sacct --format=JobID,JobName,AllocNodes,AllocCPUS,AllocTRES%-100 --units=G --jobs="${SLURM_JOBID}" +echo "NODE: $SLURMD_NODENAME, CPUS: $SLURM_CPUS_PER_TASK, ARGS: $*" +nvidia-smi --query-gpu=name,driver_version --format=csv,noheader -echo "Trackable resources (ingress) usage stats:" -sacct --format=JobID,JobName,TRESUsageInTot%-70,TRESUsageInAve%-70,TRESUsageInMax%-70,TRESUsageInMaxNode%-70,TRESUsageInMaxTask%-70 --units=G --jobs="${SLURM_JOBID}" +srun --label python bench_mpo_mpo.py "$@" -echo "Trackable resources (egress) usage stats:" -sacct --format=JobID,JobName,TRESUsageOutTot%-70,TRESUsageOutAve%-70,TRESUsageOutMax%-70,TRESUsageOutMaxNode%-70,TRESUsageOutMaxTask%-70 --units=G --jobs="${SLURM_JOBID}" +sacct --format=JobID,JobName,NCPUS,AveCPU,Elapsed,ExitCode --jobs="$SLURM_JOB_ID" +sacct --format=JobID,JobName,MaxRSS,MaxRSSNode --units=G --jobs="$SLURM_JOB_ID" diff --git a/benches/stack/README.md b/benches/stack/README.md index 62589c8..5f6b4e1 100644 --- a/benches/stack/README.md +++ b/benches/stack/README.md @@ -14,6 +14,8 @@ check. left). - `timing`: best-of-3 wall time on 30-site chains, MPS bond 64, `chi = 64`. `ratio` is one-shot time over sequential time. +- `evolve`: quench dynamics of 50+ qubits on the GPU (see + [Leonardo](#leonardo-quench-of-50-and-100-qubits)). Families: `random` are complex Gaussian MPOs of bond 3 (flat spectra); `trotter-