Repository navigation
chore(bench): 📈 refresh Leonardo benchmarks and add a GPU quench experiment - #59
Draft
Panadestein wants to merge 1 commit into
Draft
Panadestein wants to merge 1 commit into
Panadestein wants to merge 1 commit into
Conversation
…riment - Rerun the MPO-MPO primitives on one Booster node (CPU and A100, complex64). - Add `bench_depth.py evolve`: 50/100-qubit mixed-field Ising quench on the GPU, fusing Trotter steps per SRC sweep, against quimb and a chi=1024 reference. - Add leonardo/ job scripts and logs for stack and large; record `compare`. - Use CUDA 13 on driver 535 through NVIDIA's forward-compat libcuda. Refs #25 Assisted-by: GitHub Copilot:Claude Opus 5.5
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
🤖 AI text below 🤖
Closes #25.
Refreshes the benchmarks on Leonardo after the refactoring, before the PyPI release. All runs use one Booster node (32 cores, one A100-64GB) and the
EUHPC_D30_139budget. The CPU-only DCGP budgets are exhausted, so the old 112-core rows are kept but labelled as v0.3.2. Results are added progressively; the open items are listed at the end.Setup
cupy-cuda13x, unchanged inpyproject.toml) runs through NVIDIA's forward-compatibilitylibcuda. That library is unpacked into.cuda-compat/and the job scripts put it first onLD_LIBRARY_PATH. The setup steps are inbenches/README.md, anddocs/features/gpu.mdxgets a short note.$HOME.leonardo/folders with a job script and logs forstack/andlarge/. Theprimitivesscript is rewritten to pass arguments through to the benchmark and to cover both CPU and GPU..gitignorenow keepsbenches/*/leonardo/logs/*.out; the old pattern pointed at a path that no longer exists.boost_qos_dbg.MPO-MPO primitive
bench_mpo_mpo.pygains--device,--dtype,--chi-idand--seed. The GPU warm-up (about 50 s of kernel compilation) is left out of the timings.At χ=1000, SRC's peak host memory is about 1/7 of quimb's (6.8 GB vs 50.0 GB at 50 sites).
Quench dynamics, 50 and 100 qubits (new
bench_depth.py evolve)Quench of the mixed-field Ising chain (Bañuls–Cirac–Hastings point) from |0…0⟩ to t=8, in 80 Trotter steps. SRC fuses k steps into one sweep on the GPU, and errors are measured against a χ=1024 reference. Selected rows (the full table is in
benches/stack/README.md):The small
accuracy/timingtables are laptop-sized and stay as they were, now labelled as such.Out-of-core
compareat D_M=1000, χ=500, with a 4 GB and a 60 GB GPU budget. The two plans keep 49 of 50 and 0 of 50 environments on the host, and give bit-identical outputs. Keeping environments on the host costs 18 % of the wall time, and both pools stay within their caps.Still to add
largerunat D_M=4000 (the 184 GB stack is generated), χ=2000: queued on the default QoS.