Skip to content

Add guarded GPU support for geometry engines - #125

Merged
NCCU-Schultz-Lab merged 2 commits into
mainfrom
codex/gpu-offload-pyfock-gpu
Sep 19, 2026
Merged

NCCU-Schultz-Lab merged 2 commits into
mainfrom
codex/gpu-offload-pyfock-gpu

Conversation

@NCCU-Schultz-Lab

@NCCU-Schultz-Lab NCCU-Schultz-Lab commented Sep 13, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • Repair PySCF GPU offload during ASE geometry optimization by applying the migration path on every SCF/force evaluation.
  • Record GPU device provenance in optimization and engine results.
  • Add independently guarded PyFock GPU support through CuPy.
  • Add engine-specific GPU diagnostics, optional CUDA extras, UI capability handling, CPU fallback, documentation, and regression tests.
  • Add opt-in real-hardware validation gates for PySCF geometry optimization and PyFock GPU execution.

Validation

  • Focused GPU, PyFock, CLI, worker-payload, and diagnostics tests pass.
  • Black and Ruff checks pass.
  • Real NVIDIA checks are opt-in and require a CUDA-capable host.
  • PyFock GPU geometry optimization uses numerical forces because the current PyFock analytical gradient path is CPU-only.

NCCU-Schultz-Lab and others added 2 commits September 13, 2026 00:23
… probe tests

tests/test_est_cross_device_probe.py::test_force_cpu_true_sets_disable_gpu_env
calls quantui.benchmarks._calibration_worker in-process to check that
force_cpu=True sets QUANTUI_DISABLE_GPU=1 before any PySCF import. The
worker sets that var directly on the real os.environ (by design, so a
freshly-imported gpu_offload module sees it) rather than through
monkeypatch, so nothing reverted it once the test's spy short-circuited
the rest of the worker body. Every later test in the same pytest-xdist
worker process inherited a permanently GPU-disabled environment.

That was latent and harmless until this PR's new
tests/test_pyfock_gpu.py added two tests that assert a fake CuPy device
is detected — they never touch QUANTUI_DISABLE_GPU themselves, so a
leaked "1" from the earlier test silently changed their expected
outcome. Only Python 3.10's CI job happened to schedule the leaking
test and the new probe tests onto the same xdist worker, which is why
3.9/3.11 stayed green while 3.10 failed with
"QUANTUI_DISABLE_GPU is set in the environment" instead of the fake
device tuple.

Fix: wrap the worker call in try/finally and pop the var afterward.
Plain os.environ.pop, not a second monkeypatch.delenv — delenv would
have recorded the leaked "1" as the value to restore at teardown,
reintroducing the exact same leak (verified both ways in isolation
before landing this). Also hardened the two new PyFock probe tests to
delenv the var defensively on entry, so this class of cross-test
pollution can't silently change their result again.

Verified: ran the leaking test immediately followed by the new PyFock
probe tests in a single serial process — reproduced the exact CI
failure before this change, confirmed clean after. Full suite (pytest
-m "not network", Python 3.10, matching CI) passes with no other
regressions.

Contributions:
- Claude (Sonnet 5): code edits, review, and conceptual discussion
- Jonathan Schultz: overall vision, planning, review, and orchestration

Co-authored-by: Jonathan Schultz <nccu-schultz-lab@users.noreply.github.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
@NCCU-Schultz-Lab
NCCU-Schultz-Lab merged commit ef5c7d0 into main Sep 19, 2026
6 checks passed
@NCCU-Schultz-Lab
NCCU-Schultz-Lab deleted the codex/gpu-offload-pyfock-gpu branch September 19, 2026 13:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant