Conversation
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Hi Saqib, PantheomSim looks absolutely brilliant! I can see the simulated T4 tests passed 🎉 Given you emulate the CUDA runtime, am I right to think this will also work for testing our cuQuatum integrations? Very neat and a wonderful contribution, thanks very much! I'll commit to this PR to add you to our external contributor lists. |
|
Thank you so much! I really appreciate the kind words and for adding me to the external contributor list. Yes, that’s the idea. Since PantheonSim emulates the CUDA runtime/API behavior, cuQuantum integrations that rely on the CUDA runtime should also be testable through PantheonSim. That said, I’m still expanding and validating coverage across different CUDA/cuQuantum functionality, so I’d be very interested to hear if you run into any gaps or specific integration cases that don’t work as expected. Thanks again for the support and for testing it! Best, |
|
Hi @TysonRayJones, a follow-up on cuQuantum, now that I've tested it properly. NVIDIA's libcustatevec can't run on the simulator as it is: it carries its own statically linked CUDA runtime, which talks to the driver through internal interfaces the simulator doesn't provide, so With it, QuEST's devel branch built with One thing to be clear about: this tests QuEST's cuQuantum code paths (the calls, arguments, layouts and results) against our implementation of cuStateVec, not NVIDIA's kernels. I checked our implementation against NVIDIA's library on a real GPU, but a bug inside NVIDIA's library itself wouldn't show up here. Once it's in a setup-pantheonsim release, I'm happy to add a cuQuantum variant of the job to this PR, if you'd like one. |
test_paid.ymlruns the CUDA unit tests on the paid T4 runner, but only on manual dispatch, so a GPU regression can sit indeveluntil someone remembers to run it. This addstest_simulated_gpu.yml, which runs the same v4 unit tests with the CUDA backend on a freeubuntu-24.04runner, in every PR and push tomain/devel.How: PantheonSim (via
pantheongpu/setup-pantheonsim@v0) simulates a Tesla T4, the GPU of your paid runner, on the runner's CPU. The compiled kernels are executed as they are, and results are real. There is no timing model, so it checks correctness, never speed, andtest_paid.ymlstays the final word on real hardware (and on cuQuantum and multi-GPU, which this doesn't cover).Changes: one new workflow file, modeled on the CUDA job of
test_paid.yml: the same configure flags (precision 2,cuda_arch75, 10 qubit permutations, no deployment sweep) and the same test environment variables. Differences:-DCMAKE_CUDA_RUNTIME_LIBRARY=Shared, because the simulator stands in for the CUDA runtime at run time.ctest -E "Trotter|density evolution": Trotter is excluded as intest_paid.yml, and the integration test as intest_free.yml. Simulated on a CPU, its 14-qubit density matrix runs for more than 20 minutes.GPU-accelerated: 1(ctest hides a passing test's output).Nothing else in the repo is touched.
What it gives, from a run of this workflow on my fork:
Disclosure: I maintain PantheonSim. If the job is ever flaky or wrong, please open an issue at pantheongpu/pantheonsim and I'll fix it on our side.
🤖 Generated with Claude Code