Skip to content

CI: run the CUDA unit tests on a simulated T4 on free runners - #826

Open
saqibkh wants to merge 1 commit into
QuEST-Kit:develfrom
saqibkh:ci-simulated-gpu
Open

saqibkh wants to merge 1 commit into
QuEST-Kit:develfrom
saqibkh:ci-simulated-gpu

Conversation

@saqibkh

@saqibkh saqibkh commented Sep 28, 2026

Copy link
Copy Markdown

test_paid.yml runs the CUDA unit tests on the paid T4 runner, but only on manual dispatch, so a GPU regression can sit in devel until someone remembers to run it. This adds test_simulated_gpu.yml, which runs the same v4 unit tests with the CUDA backend on a free ubuntu-24.04 runner, in every PR and push to main/devel.

How: PantheonSim (via pantheongpu/setup-pantheonsim@v0) simulates a Tesla T4, the GPU of your paid runner, on the runner's CPU. The compiled kernels are executed as they are, and results are real. There is no timing model, so it checks correctness, never speed, and test_paid.yml stays the final word on real hardware (and on cuQuantum and multi-GPU, which this doesn't cover).

Changes: one new workflow file, modeled on the CUDA job of test_paid.yml: the same configure flags (precision 2, cuda_arch 75, 10 qubit permutations, no deployment sweep) and the same test environment variables. Differences:

  • -DCMAKE_CUDA_RUNTIME_LIBRARY=Shared, because the simulator stands in for the CUDA runtime at run time.
  • ctest -E "Trotter|density evolution": Trotter is excluded as in test_paid.yml, and the integration test as in test_free.yml. Simulated on a CPU, its 14-qubit density matrix runs for more than 20 minutes.
  • A one-test step that prints QuEST's environment report, so the log shows GPU-accelerated: 1 (ctest hides a passing test's output).

Nothing else in the repo is touched.

What it gives, from a run of this workflow on my fork:

  • 275 of 275 tests pass.
  • About 17 minutes in total: 6 to build the simulator, 3 to compile QuEST, 8 for the tests.

Disclosure: I maintain PantheonSim. If the job is ever flaky or wrong, please open an issue at pantheongpu/pantheonsim and I'll fix it on our side.

🤖 Generated with Claude Code

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@TysonRayJones

Copy link
Copy Markdown
Member

Hi Saqib,

PantheomSim looks absolutely brilliant! I can see the simulated T4 tests passed 🎉 Given you emulate the CUDA runtime, am I right to think this will also work for testing our cuQuatum integrations?

Very neat and a wonderful contribution, thanks very much! I'll commit to this PR to add you to our external contributor lists.

@saqibkh

saqibkh commented Sep 30, 2026

Copy link
Copy Markdown
Author

Hi @TysonRayJones

Thank you so much! I really appreciate the kind words and for adding me to the external contributor list.

Yes, that’s the idea. Since PantheonSim emulates the CUDA runtime/API behavior, cuQuantum integrations that rely on the CUDA runtime should also be testable through PantheonSim. That said, I’m still expanding and validating coverage across different CUDA/cuQuantum functionality, so I’d be very interested to hear if you run into any gaps or specific integration cases that don’t work as expected.

Thanks again for the support and for testing it!

Best,
Saqib

@saqibkh

saqibkh commented Oct 1, 2026

Copy link
Copy Markdown
Author

Hi @TysonRayJones, a follow-up on cuQuantum, now that I've tested it properly.

NVIDIA's libcustatevec can't run on the simulator as it is: it carries its own statically linked CUDA runtime, which talks to the driver through internal interfaces the simulator doesn't provide, so custatevecCreate fails. So I've added a cuStateVec of PantheonSim's own, written from NVIDIA's documentation, covering the functions QuEST calls (pantheongpu/pantheonsim#252, now merged). Programs linked against NVIDIA's library pick it up unchanged.

With it, QuEST's devel branch built with QUEST_ENABLE_CUQUANTUM=ON passes all 275 unit tests on a simulated T4 (the same selection as this PR's job), in about three minutes.

One thing to be clear about: this tests QuEST's cuQuantum code paths (the calls, arguments, layouts and results) against our implementation of cuStateVec, not NVIDIA's kernels. I checked our implementation against NVIDIA's library on a real GPU, but a bug inside NVIDIA's library itself wouldn't show up here.

Once it's in a setup-pantheonsim release, I'm happy to add a cuQuantum variant of the job to this PR, if you'd like one.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants