Skip to content

chore(bench): 📈 refresh Leonardo benchmarks and add a GPU quench experiment - #59

Draft
Panadestein wants to merge 1 commit into
mainfrom
chore/leonardo-benchmarks
Draft

Panadestein wants to merge 1 commit into
mainfrom
chore/leonardo-benchmarks

Conversation

@Panadestein

Copy link
Copy Markdown
Member

🤖 AI text below 🤖

Closes #25.

Refreshes the benchmarks on Leonardo after the refactoring, before the PyPI release. All runs use one Booster node (32 cores, one A100-64GB) and the EUHPC_D30_139 budget. The CPU-only DCGP budgets are exhausted, so the old 112-core rows are kept but labelled as v0.3.2. Results are added progressively; the open items are listed at the end.

Setup

  • The Booster driver is 535, which only supports up to CUDA 12.2. CUDA 13 (cupy-cuda13x, unchanged in pyproject.toml) runs through NVIDIA's forward-compatibility libcuda. That library is unpacked into .cuda-compat/ and the job scripts put it first on LD_LIBRARY_PATH. The setup steps are in benches/README.md, and docs/features/gpu.mdx gets a short note.
  • The uv cache, Python installs and CuPy kernel cache stay inside the repository, not $HOME.
  • New leonardo/ folders with a job script and logs for stack/ and large/. The primitives script is rewritten to pass arguments through to the benchmark and to cover both CPU and GPU.
  • .gitignore now keeps benches/*/leonardo/logs/*.out; the old pattern pointed at a path that no longer exists.
  • The default QoS currently queues for hours, so runs under 30 minutes used boost_qos_dbg.

MPO-MPO primitive

bench_mpo_mpo.py gains --device, --dtype, --chi-id and --seed. The GPU warm-up (about 50 s of kernel compilation) is left out of the timings.

Experiment quimb CPU (s) src CPU (s) src GPU (s) src GPU c64 (s) Speedup (GPU)
50 sites, χ=50 37.76 0.23 0.57 66x
20 sites, χ=100 17.86 0.25 0.46 39x
25 sites, χ=1000, χ_id=4 422.57 60.47 2.32 1.90 182x
50 sites, χ=1000, χ_id=4 1040.72 137.37 4.47 3.03 233x

At χ=1000, SRC's peak host memory is about 1/7 of quimb's (6.8 GB vs 50.0 GB at 50 sites).

Quench dynamics, 50 and 100 qubits (new bench_depth.py evolve)

Quench of the mixed-field Ising chain (Bañuls–Cirac–Hastings point) from |0…0⟩ to t=8, in 80 Trotter steps. SRC fuses k steps into one sweep on the GPU, and errors are measured against a χ=1024 reference. Selected rows (the full table is in benches/stack/README.md):

Qubits χ Method Time (s) Infidelity
50 64 quimb, 32 cores 731 0.037
50 64 src GPU, k=1 6.87 0.062
50 128 src GPU, k=4 3.01 0.022
50 512 src GPU, k=2 19.2 6.9e-4
100 512 src GPU, k=2 40.6 1.8e-3
  • SRC on the GPU is 100–250x faster than quimb at the same χ. quimb's exact SVD is more accurate at equal χ, but SRC at 2χ is both more accurate and still more than 90x faster.
  • Fusing 2 or 4 steps per sweep is 1.8–2.8x faster than one at a time, for at most 7 % more infidelity.

The small accuracy/timing tables are laptop-sized and stay as they were, now labelled as such.

Out-of-core

compare at D_M=1000, χ=500, with a 4 GB and a 60 GB GPU budget. The two plans keep 49 of 50 and 0 of 50 environments on the host, and give bit-identical outputs. Keeping environments on the host costs 18 % of the wall time, and both pools stay within their caps.

Still to add

  • large run at D_M=4000 (the 184 GB stack is generated), χ=2000: queued on the default QoS.
  • quimb at χ ≥ 128 for the quench: each run would take over an hour.

…riment

- Rerun the MPO-MPO primitives on one Booster node (CPU and A100, complex64).
- Add `bench_depth.py evolve`: 50/100-qubit mixed-field Ising quench on the GPU,
  fusing Trotter steps per SRC sweep, against quimb and a chi=1024 reference.
- Add leonardo/ job scripts and logs for stack and large; record `compare`.
- Use CUDA 13 on driver 535 through NVIDIA's forward-compat libcuda.

Refs #25

Assisted-by: GitHub Copilot:Claude Opus 5.5

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Update primitive benches after open refactoring issues are resolved

1 participant