Repository navigation
Conversation
The dedicated PTQ1_0 kernel (mul_mat_vec_ptq1_0_pt) is templated for fusion with any column count, but the graph check, an assert in ggml_cuda_mul_mat_vec_q and the launcher only allowed ncols_dst == 1. So gate + up + SwiGLU ran as three kernels in every FFN block of a 2-4 token verify step (speculative decoding, small batches). Allow the fused path for the dedicated kernel (plain 2D, K multiple of 128, shared memory within the limit; the graph check calls the same helper, ggml_cuda_mmvq_ptq1_0_can_fuse_mc). Other types keep the one-column rule. RTX 4070, Bonsai 2 27B PTQ1_0 + MTP head (n-max 2), q4_0 K/V, 4 greedy prompts, 8 interleaved A/B pairs of the same binary: 105.56 -> 106.65 tok/s (+1.01 %, 95 % CI [+0.95, +1.06] %, A/A noise: sd 0.13 % per pair). Outputs identical in all runs. test-backend-ops MUL_MAT_VEC_FUSION 57/57 with and without the path (incl. 19 local cases, 2-4 columns, K 512/5120/17408). GGML_CUDA_PTQ1_FUSE_MC=0 restores the old choice. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
… q8 rows The fused gate/up mat-vec of the previous commit was refused in 24 of 64 FFN blocks of the Bonsai 2 verify graph: ggml_cuda_check_fusion_memory_ranges saw the GLU output (ffn_swiglu) overlap src1, the Hadamard transform output, because ggml-alloc hands it the block that src1 released. The overlap is real when the q8_1 rows written by ggml_cuda_try_fwht_q8 live in that buffer (blocks would overwrite rows others still read), so the check was right. Now ggml_cuda_try_fwht_q8 puts the rows into a pool block when a following GLU of a consumer pair overlaps the buffer (the mechanism it already has for out_aliases_in), and the fusion memory check ignores an overlap with an input whose q8 rows are registered in a pool block (g_fwht_q8_ctx, set while the graph is evaluated). All 64 FFN blocks fuse. RTX 4070, Bonsai 2 27B PTQ1_0 + MTP head (n-max 2), q4_0 K/V, 8 interleaved A/B pairs of the same binary: 106.74 -> 107.14 tok/s (+0.33 %, 95 % CI [+0.24, +0.43] %). Outputs identical in all runs. test-backend-ops MUL_MAT, MUL_MAT_ID, MUL_MAT_VEC_FUSION, GLU, SWIGLU, RMS_NORM, CONCAT: 3416/3416 with and without it. GGML_CUDA_FWHT_GLU_POOL=0 restores the previous behavior. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
g_fwht_q8_ctx is read by the fusion memory check only while a graph is evaluated. Set it through a small scope guard so that it does not point to a context that was freed after the evaluation. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…witches The environment switches of the two commits above were only there to measure them; both paths are now taken whenever their conditions hold. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Open
2 tasks done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
Two commits. The dedicated PTQ1_0 kernel (
mul_mat_vec_ptq1_0_pt) is templated for gate + up + SwiGLU fusion with any column count, but the graph check, an assert inggml_cuda_mul_mat_vec_qand the launcher only allowedncols_dst == 1. So gate + up + SwiGLU ran as three kernels in every FFN block of a 2-4 token verify step (speculative decoding, small batches).cuda: gate (SwiGLU) fused PTQ1_0 mat-vec for 2-4 columns: allow the fused path for the dedicated kernel (plain 2D, K a multiple of 128, shared memory within the limit; the graph check calls the same helper,ggml_cuda_mmvq_ptq1_0_can_fuse_mc). Other types keep the one-column rule.cuda: keep the gate/up SwiGLU fusion when the GLU output overlaps the q8 rows: in 24 of 64 FFN blocks of the Bonsai 2 verify graph the fusion was refused becauseggml_cuda_check_fusion_memory_rangessaw the GLU output overlapsrc1(ggml-alloc hands it the blocksrc1released). The overlap is real when the q8_1 rows written byggml_cuda_try_fwht_q8live in that buffer, so the check was right. Nowggml_cuda_try_fwht_q8puts the rows into a pool block when a following GLU of a consumer pair overlaps the buffer (the mechanism it already has forout_aliases_in), and the check ignores an overlap with an input whose q8 rows are registered in a pool block. All 64 FFN blocks fuse.RTX 4070, Bonsai 2 27B PTQ1_0 + MTP head (n-max 2), q4_0 K/V, 4 greedy prompts, interleaved A/B pairs of the same binary (env switch), outputs identical in all runs: commit 1 105.56 -> 106.65 tok/s (+1.01 %, 95 % CI [+0.95, +1.06] %, 8 pairs), commit 2 106.74 -> 107.14 tok/s (+0.33 %, CI [+0.24, +0.43] %), together +1.28 %.
Additional information
GGML_CUDA_PTQ1_FUSE_MCandGGML_CUDA_FWHT_GLU_POOL(=0restores the previous behavior) that were used for the measurements below, the third resetsg_fwht_q8_ctxafter the graph evaluation through a small scope guard (it is read only while a graph is evaluated), the last one removes the two switches. To reproduce a measurement, build the second commit.test-backend-opsexercises the fused kernel (2 to 4 columns) but not the buffer aliasing handled by the second commit, because that depends on how ggml-alloc places the tensors of a real graph. For the aliasing I rely on identical outputs of the full model (below). The original comment inggml_cuda_try_fwht_q8describes the failure mode (rows stored early over input that other blocks still read: sporadic garbage in the recurrent state), so please review the second commit with that in mind.Test results
-DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=89.prismat88c4bc60b; the four commits since (SYCL, WebGPU andcuda: fused FWHT quantizer for 64-wide warps (#303)) do not touch these code paths. The branch is rebased on2459f68b5, builds, andtest-backend-opswas repeated on it.--spec-type draft-mtp --spec-draft-n-max 2(the verify step runs 3 columns), q4_0 K/V cache, one slot; 64 FFN blocks.GGML_CUDA_PTQ1_FUSE_MC,GGML_CUDA_FWHT_GLU_POOL; in the first two commits of the branch), 4 greedy prompts x 256 tokens, alternating runs, 8 pairs, paired differences with a 95 % bootstrap interval; run-to-run noise about 0.04 to 0.12 %.test-backend-ops test -b CUDA0 -o MUL_MAT,MUL_MAT_VEC_FUSION,GLU,SWIGLU,RMS_NORM: 2167/2167 passed on the rebased branch (CUDA0 against CPU). On the earlier base,MUL_MAT_VEC_FUSION57/57 with and without the path (including 19 local cases with 2 to 4 columns and K 512 / 5120 / 17408, not part of this PR) andMUL_MAT, MUL_MAT_ID, MUL_MAT_VEC_FUSION, GLU, SWIGLU, RMS_NORM, CONCAT3416/3416 for commit 2.Requirements