Skip to content

cuda: gate (SwiGLU) fused PTQ1_0 mat-vec for 2-4 columns - #310

Open
sb32445 wants to merge 4 commits into
PrismML-Eng:prismfrom
sb32445:pr/ptq1-gate-up-fuse-mc
Open

sb32445 wants to merge 4 commits into
PrismML-Eng:prismfrom
sb32445:pr/ptq1-gate-up-fuse-mc

Conversation

@sb32445

@sb32445 sb32445 commented Oct 4, 2026

Copy link
Copy Markdown

Overview

Two commits. The dedicated PTQ1_0 kernel (mul_mat_vec_ptq1_0_pt) is templated for gate + up + SwiGLU fusion with any column count, but the graph check, an assert in ggml_cuda_mul_mat_vec_q and the launcher only allowed ncols_dst == 1. So gate + up + SwiGLU ran as three kernels in every FFN block of a 2-4 token verify step (speculative decoding, small batches).

  1. cuda: gate (SwiGLU) fused PTQ1_0 mat-vec for 2-4 columns: allow the fused path for the dedicated kernel (plain 2D, K a multiple of 128, shared memory within the limit; the graph check calls the same helper, ggml_cuda_mmvq_ptq1_0_can_fuse_mc). Other types keep the one-column rule.
  2. cuda: keep the gate/up SwiGLU fusion when the GLU output overlaps the q8 rows: in 24 of 64 FFN blocks of the Bonsai 2 verify graph the fusion was refused because ggml_cuda_check_fusion_memory_ranges saw the GLU output overlap src1 (ggml-alloc hands it the block src1 released). The overlap is real when the q8_1 rows written by ggml_cuda_try_fwht_q8 live in that buffer, so the check was right. Now ggml_cuda_try_fwht_q8 puts the rows into a pool block when a following GLU of a consumer pair overlaps the buffer (the mechanism it already has for out_aliases_in), and the check ignores an overlap with an input whose q8 rows are registered in a pool block. All 64 FFN blocks fuse.

RTX 4070, Bonsai 2 27B PTQ1_0 + MTP head (n-max 2), q4_0 K/V, 4 greedy prompts, interleaved A/B pairs of the same binary (env switch), outputs identical in all runs: commit 1 105.56 -> 106.65 tok/s (+1.01 %, 95 % CI [+0.95, +1.06] %, 8 pairs), commit 2 106.74 -> 107.14 tok/s (+0.33 %, CI [+0.24, +0.43] %), together +1.28 %.

Additional information

  • The branch has four commits: the first two are the two changes with the environment switches GGML_CUDA_PTQ1_FUSE_MC and GGML_CUDA_FWHT_GLU_POOL (=0 restores the previous behavior) that were used for the measurements below, the third resets g_fwht_q8_ctx after the graph evaluation through a small scope guard (it is read only while a graph is evaluated), the last one removes the two switches. To reproduce a measurement, build the second commit.
  • test-backend-ops MUL_MAT_VEC_FUSION 57/57 with and without the path (incl. 19 local cases with 2-4 columns and K 512/5120/17408; the cases are not part of this PR), MUL_MAT, MUL_MAT_ID, MUL_MAT_VEC_FUSION, GLU, SWIGLU, RMS_NORM, CONCAT 3416/3416 for commit 2.
  • What the tests cover: test-backend-ops exercises the fused kernel (2 to 4 columns) but not the buffer aliasing handled by the second commit, because that depends on how ggml-alloc places the tensors of a real graph. For the aliasing I rely on identical outputs of the full model (below). The original comment in ggml_cuda_try_fwht_q8 describes the failure mode (rows stored early over input that other blocks still read: sporadic garbage in the recurrent state), so please review the second commit with that in mind.
  • Only measured on one Ada GPU and one model family.
  • The first two commits belong together (2 refines 1), review as one PR.

Test results

  • Hardware / software: RTX 4070 12 GB (AD104, cc 8.9, 100 KB shared memory per SM), Linux 6.18, NVIDIA driver 615.71, CUDA 13.4, GCC 16.2; Release build, -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=89.
  • Base: speed numbers were measured on prism at 88c4bc60b; the four commits since (SYCL, WebGPU and cuda: fused FWHT quantizer for 64-wide warps (#303)) do not touch these code paths. The branch is rebased on 2459f68b5, builds, and test-backend-ops was repeated on it.
  • Model: Ternary-Bonsai-2-27B (PTQ1_0) with a community MTP draft head (Q8_0), --spec-type draft-mtp --spec-draft-n-max 2 (the verify step runs 3 columns), q4_0 K/V cache, one slot; 64 FFN blocks.
  • Method: the same binary, one switch flipped through an environment variable (GGML_CUDA_PTQ1_FUSE_MC, GGML_CUDA_FWHT_GLU_POOL; in the first two commits of the branch), 4 greedy prompts x 256 tokens, alternating runs, 8 pairs, paired differences with a 95 % bootstrap interval; run-to-run noise about 0.04 to 0.12 %.
  • Commit 1 (fusion for 2 to 4 columns): 105.56 -> 106.65 tok/s, +1.01 % (95 % CI [+0.95, +1.06] %), outputs identical in all runs.
  • Commit 2 (keep the fusion when the GLU output overlaps the q8 rows): 106.74 -> 107.14 tok/s, +0.33 % (CI [+0.24, +0.43] %), outputs identical in all runs. Before it, 24 of 64 FFN blocks of the verify graph refused the fusion; now all 64 fuse. Together +1.28 % (CI [+1.21, +1.36] %).
  • test-backend-ops test -b CUDA0 -o MUL_MAT,MUL_MAT_VEC_FUSION,GLU,SWIGLU,RMS_NORM: 2167/2167 passed on the rebased branch (CUDA0 against CPU). On the earlier base, MUL_MAT_VEC_FUSION 57/57 with and without the path (including 19 local cases with 2 to 4 columns and K 512 / 5120 / 17408, not part of this PR) and MUL_MAT, MUL_MAT_ID, MUL_MAT_VEC_FUSION, GLU, SWIGLU, RMS_NORM, CONCAT 3416/3416 for commit 2.
  • Full-model outputs: identical text in every A/B run of the two commits (greedy benchmark above). Later runs of a stack that contains both commits (greedy functional test suite, 80 cases: 80 of 80 identical answers with the same pass/fail verdict, compared with the stack before three later, unrelated patches) do not isolate these two commits; the A/B runs above do. I did not prove by inspection that the fused SwiGLU epilogue is bitwise equal to the separate GLU kernel; I only observed identical outputs.
  • Not tested: other GPUs, more than 4 columns (they keep the old path), HIP (the helper returns false there), models other than PTQ1_0.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: The patches were developed with Claude Code (Anthropic's coding agent): it wrote the code, the measurement scripts and the first drafts of the commit messages and PR texts. I decided what to work on (which kernels and host paths to optimise, based on profiles of my own decode setup). The measurements and checks listed in the PR texts were run in the Claude Code sessions; I did not re-run them independently. I will maintain the changes. Commits where Claude Code was used carry a Co-Authored-By trailer.

sb32445 and others added 4 commits October 4, 2026 14:21
The dedicated PTQ1_0 kernel (mul_mat_vec_ptq1_0_pt) is templated for fusion with any column count, but the
graph check, an assert in ggml_cuda_mul_mat_vec_q and the launcher only allowed ncols_dst == 1. So gate + up +
SwiGLU ran as three kernels in every FFN block of a 2-4 token verify step (speculative decoding, small batches).
Allow the fused path for the dedicated kernel (plain 2D, K multiple of 128, shared memory within the limit; the
graph check calls the same helper, ggml_cuda_mmvq_ptq1_0_can_fuse_mc). Other types keep the one-column rule.

RTX 4070, Bonsai 2 27B PTQ1_0 + MTP head (n-max 2), q4_0 K/V, 4 greedy prompts, 8 interleaved A/B pairs of the same binary:
105.56 -> 106.65 tok/s (+1.01 %, 95 % CI [+0.95, +1.06] %, A/A noise: sd 0.13 % per pair). Outputs identical in all runs.
test-backend-ops MUL_MAT_VEC_FUSION 57/57 with and without the path (incl. 19 local cases, 2-4 columns, K 512/5120/17408).

GGML_CUDA_PTQ1_FUSE_MC=0 restores the old choice.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
… q8 rows

The fused gate/up mat-vec of the previous commit was refused in 24 of 64 FFN blocks of the Bonsai 2 verify graph:
ggml_cuda_check_fusion_memory_ranges saw the GLU output (ffn_swiglu) overlap src1, the Hadamard transform output,
because ggml-alloc hands it the block that src1 released. The overlap is real when the q8_1 rows written by
ggml_cuda_try_fwht_q8 live in that buffer (blocks would overwrite rows others still read), so the check was right.

Now ggml_cuda_try_fwht_q8 puts the rows into a pool block when a following GLU of a consumer pair overlaps the buffer
(the mechanism it already has for out_aliases_in), and the fusion memory check ignores an overlap with an input whose
q8 rows are registered in a pool block (g_fwht_q8_ctx, set while the graph is evaluated). All 64 FFN blocks fuse.

RTX 4070, Bonsai 2 27B PTQ1_0 + MTP head (n-max 2), q4_0 K/V, 8 interleaved A/B pairs of the same binary:
106.74 -> 107.14 tok/s (+0.33 %, 95 % CI [+0.24, +0.43] %). Outputs identical in all runs.
test-backend-ops MUL_MAT, MUL_MAT_ID, MUL_MAT_VEC_FUSION, GLU, SWIGLU, RMS_NORM, CONCAT: 3416/3416 with and without it.

GGML_CUDA_FWHT_GLU_POOL=0 restores the previous behavior.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
g_fwht_q8_ctx is read by the fusion memory check only while a graph is evaluated.
Set it through a small scope guard so that it does not point to a context that
was freed after the evaluation.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…witches

The environment switches of the two commits above were only there to measure
them; both paths are now taken whenever their conditions hold.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@sb32445 sb32445 changed the title cuda: use the transposing concat kernel on all GPUs, not only on GB10 cuda: gate (SwiGLU) fused PTQ1_0 mat-vec for 2-4 columns Oct 4, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant