Reviewer quickstart for the ParticleGS paper (SC26, pap525):
3D Gaussian Splatting for Scientific Particle Data Compression and Rendering.
Full details are in the separately-submitted Artifact Description (AD). This README is the short path to verifying the badge.
The paper in one paragraph: scientific simulations such as the HACC cosmology code output snapshots of hundreds of millions of particles; exploring them means moving multi-GB files and re-rendering them in a visualization tool (ParaView) at seconds per frame. ParticleGS instead trains a 3D Gaussian Splatting (3DGS) model of the snapshot's rendered appearance and uses it as a compressed, directly renderable stand-in for the data: the 280 M-particle HACC snapshot (3.4 GB) compresses ~290× and renders interactively (measured 2525× faster than ParaView on the raw particles), leading the error-bounded compressor SZ3 by +7.7 dB PSNR at a matched compression ratio. A small MLP ("VizMapper") lets the one trained model follow user-chosen visualization parameters (particle radius/opacity) at inference; KD-tree block partitioning with a merge/finetune stage scales training; and approximate particle positions can be recovered back from the Gaussians (density correlation 0.92). The experiments below reproduce each of these claims from the raw particles.
Fast path for reviewers — verifies the 18 scored metrics. Validated end-to-end
on a Chameleon Cloud gpu_rtx_6000 node in ~7 h (fits the ~8 h AE budget;
details and caveats in §4):
# 1. Chameleon Cloud (CHI@UC): reserve + launch a bare-metal RTX 6000 node.
# Skip this step if you already have a Linux box with a CUDA GPU (see §2).
openstack reservation lease create --end-date "<YYYY-MM-DD HH:MM>" \
--reservation min=1,max=1,resource_type=physical:host,resource_properties='["=","$node_type","gpu_rtx_6000"]' \
rtx6000-lease
openstack server create --image CC-Ubuntu24.04-CUDA --flavor baremetal \
--key-name <your-keypair> --network sharednet1 \
--hint reservation=<reservation id from `lease show rtx6000-lease`> \
rtx6000-node
# ... attach a floating IP, SSH in as user `cc`.
# 2. Clone.
git clone https://github.com/BoJiang03/ParticleGS && cd ParticleGS
# 3. Run. One command on a bare node: installs the env, fetches data, runs the
# experiments (~7 h on 1x RTX 6000; ~2.2 h on 2x RTX PRO 6000 — set
# --num_gpus to your GPU count), then verifies.
bash scripts/reproduce_ae.sh --num_gpus 1
python verify_results.py --ae # PASS/FAIL vs the reference valuesFull reproduction — retrains everything, all 26 metrics (~11–15 h):
bash scripts/reproduce.sh --num_gpus 1 && python verify_results.py
The fast path ships the pre-trained E25 single-block model and the 4 sub-block models (the two slowest trainings), trains only the 4-block finetune live (~17 min for the whole unit: merge + 60k-iter finetune + eval; the timed finetune itself is 14.4 min), and re-renders all ground truth on your node. The full path ships nothing and retrains from the raw particles.
Target badges: Results Reproduced (primary), Artifacts Evaluated — Functional, Artifacts Available (Zenodo DOI).
verify_results.py scores each metric in one of three classes:
- Hardware-independent (PSNR, Gaussian count, model size, compression ratio) — must match our reference within small tolerances on any GPU.
- Trend (e.g. 3DGS > 100× ParaView FPS) — enforced, but only as an order-of-magnitude rule, never as an exact figure.
- Hardware-dependent (absolute FPS, wall-clock, peak VRAM) — reported only.
| GPU | 1× CUDA GPU, ≥ 16 GB VRAM, compute ≥ 7.5 (Turing or newer). A graphics-class card (RTX PRO 6000 / RTX 6000 Ada / L40) is strongly preferred: ground-truth generation renders 280 M point-gaussians in ParaView, a fill-rate workload, so compute cards (A100/H100) are much slower here. --num_gpus N spreads rendering/training across N GPUs. |
| CPU / RAM / disk | 32 GB RAM, ~40 GB free disk. |
| OS / driver | Linux; NVIDIA driver supporting CUDA ≥ 12.4. |
| Validated (AE) | Chameleon Cloud, CHI@UC site, gpu_rtx_6000 node type (1× Quadro RTX 6000, Turing, 24 GB; 2× Intel Xeon Gold 6126 @ 2.60 GHz, 24 cores / 48 threads; 192 GiB RAM; image CC-Ubuntu24.04-CUDA): fast path 18/18 in ~7.0 h from a bare node. This is the exact reviewer recipe — see §4. The CPU matters: EXP-7's 280 M-particle comparison is CPU-bound and accounts for a large share of that wall-clock, so a node with fewer cores will be slower even with the same GPU. |
| Authors' reference | 1× RTX PRO 6000 Blackwell (96 GB) — the machine all absolute FPS/time/VRAM numbers were measured on. Fast path from a fresh clone: 18/18 in ~2.2 h on 2× RTX PRO 6000. |
Software. Everything installs into a conda env named particlegs. You don't
need conda beforehand — reproduce_ae.sh installs Miniforge and builds the env
(matching your driver's CUDA version) if it's missing. To build manually:
bash install.sh (~15 min: PyTorch cu130 + 3 CUDA extensions + SZ3/LCP), then
conda activate particlegs.
Raw particle data is not shipped and is fetched automatically on first run from the companion Zenodo record 10.5281/zenodo.22693099:
| Dataset | Used by | Size |
|---|---|---|
| HACC 280 M subset (HACC/EXASKY via SDRbench) | all main results | 3.1 GiB |
| FIRE-2 L172 snapshot 010 (CC BY 4.0) | full run only (FIRE-2 generalization) | 3.0 GiB |
HACC 1 B (--big) |
HACC-region generalization only | 12.0 GiB |
Each dataset is served as three headerless little-endian float32 arrays, one
per coordinate axis, which is the form the pipelines read — so there is nothing
to unpack, and the download is smaller than the upstream archives it came from.
Every file is checked against a known md5 before use.
The upstream hosts remain as a fallback, but note that the SDRbench endpoint has
been returning HTTP 404 for anonymous requests since ~2026-09, which is why the
Zenodo record exists. To fetch from somewhere else entirely (a shared filesystem,
a local mirror), set PARTICLEGS_DATA_MIRROR to a base URL or a file:// path.
Please cite the original datasets, not this mirror, when using the data itself.
bash scripts/reproduce_ae.sh --num_gpus 1 # set N to your GPU count; --no-setup if env builtRuns EXP-1/4/6/7/8/11/14 → 18 scored metrics, then verifies them. It drops
the two render-heaviest units of the full run (FIRE-2 retrain, EXP-4's 2-block
config) and the LCP baseline, and — using the shipped models — trains only the
4-block finetune live. Measured end-to-end from a fresh clone (18/18 both):
~7.0 h on 1× RTX 6000 (Turing, single GPU — fits the ~8 h AE budget), or
~2.2 h on 2× RTX PRO 6000. It is render-bound, so a compute-class GPU
(A100/H100) is slower. Runs on a single GPU; more/faster GPUs cut wall-clock.
Flags: --no-setup (skip env build), --sequential (disable parallel
scheduling), --gpu B (base GPU).
We validated the fast path end-to-end on Chameleon; reproducing our exact
setup takes three steps on CHI@UC (chi.uc.chameleoncloud.org):
-
Reserve a bare-metal lease for node type
gpu_rtx_6000(1× Quadro RTX 6000, Turing, 24 GB — check the availability calendar a few days ahead; GPU nodes are popular). Via the CLI (Blazar, times in UTC):openstack reservation lease create \ --end-date "<YYYY-MM-DD HH:MM>" \ --reservation min=1,max=1,resource_type=physical:host,resource_properties='["=","$node_type","gpu_rtx_6000"]' \ rtx6000-lease
-
Launch it with the
CC-Ubuntu24.04-CUDAimage. The image matters: the plainCC-Ubuntu*images ship no NVIDIA driver andreproduce_ae.shwill fail on them; the-CUDAimage preinstalls a driver new enough for the cu130 build (no conda needed — the script installs Miniforge).openstack server create \ --image CC-Ubuntu24.04-CUDA \ --flavor baremetal \ --key-name <your-keypair> \ --network sharednet1 \ --hint reservation=<reservation id from `lease show rtx6000-lease`> \ rtx6000-node
Then attach a floating IP and SSH in as user
cc. -
git clone https://github.com/BoJiang03/ParticleGS && cd ParticleGS, then run the TL;DR:bash scripts/reproduce_ae.sh --num_gpus 1.
On this node the fast path completed in ~7 h 01 min with 18/18 metrics passing. Avoid Chameleon's P100/V100 node types (Pascal/Volta — CUDA 13.0 dropped them; Turing cc 7.5 is the floor).
bash scripts/reproduce.sh --num_gpus 2Retrains from the raw particles (E25, the 2/4-block configs, FIRE-2) and runs the full SZ3/LCP rate-distortion sweeps → all 26 metrics. ~11 h on 2 GPUs; ~15 h on a single GPU (EXP-4 block training runs serially). (The paper's 8/16-block rows and Fig. 7 scaling are outside AE scope.)
The fast path ships E25 pre-trained, so reviewers never see the single-block training cost. To observe it, run this optional, supplementary script — it trains E25 live and reports the wall-clock (~1.5 h on 1× RTX 6000; run it separately from the fast path, not back-to-back within the 8 h budget). The time is graphics-hardware-specific; for the exact paper number, contact the authors to schedule time on the authors' workstation.
verify_results.py [--ae] compares runs/<exp>/results.json against
reference_results.json (captured on the authors' RTX PRO 6000). The full list
of expected values + tolerances for manual cross-checking is in
AE_EXPECTED.md. Headline claims:
| Metric | Expected | Tol / rule | Paper |
|---|---|---|---|
| ParticleGS E25 — masked PSNR @ CR 290× | 26.28 dB | ± 0.3 dB | Tab. VI / Fig. 8 ¹ |
| SZ3 at matched CR (~292×) — masked PSNR | 18.57 dB | ± 0.1 dB | R-D fig ¹ |
| → ParticleGS lead at iso-CR | +7.7 dB | — | headline |
| 4-block finetuned — PSNR / #G / size | 27.5 dB / 606k / 39.3 MB | ± 0.3 dB, ± 3 % | Tab. III ¹ |
| Particle recovery, 4-block — density corr | 0.923 | > 0.9 | recovery |
| 3DGS vs ParaView render speedup | 2525× | > 100× | Tab. VII ² |
| Generalization, out-of-range radius | 20.67 dB | > 18 dB | gen. |
¹ The artifact enforces masked PSNR (foreground pixels, the stricter
metric); the paper's tables print full-image PSNR — Tab. VI single-block
28.80 dB, Tab. III 4-block 29.94 dB, Tab. IV FIRE-2 29.27 dB. Both metrics
come from the same renders; the masked references here are what
verify_results.py scores.
² Re-measured on the artifact's 4-block model; paper Tab. VII prints 2386× for
the 8-block model. Only the > 100× trend is scored.
Training carries ±3 % Gaussian-count / ±0.3 dB PSNR stochastic noise; the SZ3
baseline is deterministic (hence the tight tolerance). The full run adds the
LCP baseline, the 2-block config, and the FIRE-2 row (26 metrics total).
Hardware-dependent references (reported, not scored): 3DGS 803 FPS, ParaView
0.32 FPS, training peak 10.5 GB (total nvidia-smi device usage; paper Tab. VII
prints 8.5 GB, the training process's allocator peak), finetune 14.4 min,
raw→VTP 6.85 min.
Full numerical tables land in runs/summary/.
- Multi-GPU is optional — everything runs on one GPU;
--num_gpus Nonly cuts wall-clock by parallelizing rendering/training. pvbatchon the wrong GPU?common.pyauto-probes the EGL→CUDA mapping; see theEGL device N → CUDA device Mlog line.- CUDA OOM? Lower
resolution_scalein the stage config (e.g.2halves each side). - Rasterizer built wrong (training dies at iter 0 with a nonsense multi-TiB
alloc): the env must build against its own pinned CUDA 13.0 toolchain, not a
host
/usr/local/cuda.install.shself-checks this viascripts/check_rasterizer.py; re-runbash install.shfrom a clean env.
If you use this artifact or build on ParticleGS, please cite the SC26 paper. The paper is accepted and in press (published November 2026), so the entry below carries no DOI or page numbers yet; the preprint is on arXiv and is readable now:
@inproceedings{jiang2026particlegs,
title = {{3D} {Gaussian} {Splatting} for Scientific Particle Data
Compression and Rendering},
author = {Jiang, Bo and Liu, Youyuan and Yang, Taolue and
Di, Sheng and Jin, Sian},
booktitle = {SC26: International Conference for High Performance
Computing, Networking, Storage and Analysis},
year = {2026},
organization = {IEEE},
eprint = {2607.22956},
archivePrefix = {arXiv},
primaryClass = {cs.GR},
}Preprint: arXiv:2607.22956
(DOI 10.48550/arXiv.2607.22956).
Please cite the SC26 proceedings version above rather than the preprint alone.
Every field in that entry is final except the DOI and page numbers, which are
assigned at publication and will be added here then; if you copied the entry
earlier, adding them is the only change. GitHub's "Cite this repository" button
always reflects the current CITATION.cff, so it picks the update up on its
own.
To cite this archived artifact snapshot itself rather than the paper, use the
Zenodo DOI below. CITATION.cff carries the same metadata in machine-readable
form, which is what GitHub's "Cite this repository" button reads.
Repository: https://github.com/BoJiang03/ParticleGS ·
Archival DOI: 10.5281/zenodo.22052139
(the sc26-final snapshot; 10.5281/zenodo.22052138
always resolves to the latest version) ·
Contact: Bo Jiang <bo.jiang@temple.edu>
License: see the top-level LICENSE file for the full layering.
In short: authors' code under the Gaussian-Splatting Research License (INRIA,
non-commercial research), inherited from diff-gaussian-rasterization / simple-knn;
third-party components keep their own licenses (fused-ssim MIT, glm MIT,
SZ3/LCP BSD).