Regime-aware sparse matrix multiplication for GNN workloads on GPUs (Future Generation Computer Systems, Elsevier). Six structural regimes, an interpretable seven-rule router, and five specialized kernel families over native CSR and Tensor Cores.
reproducible-research gpu cuda high-performance-computing sparse-matrix cusparse graph-neural-networks spmm tensor-cores kernel-selection
-
Updated
Jul 31, 2026 - Python