Conversation
The generic mmvq kernel loses DRAM throughput from 3 columns on (RTX 4070, Bonsai 2 27B shapes: 46% of peak at 4 columns, 36% at 8). Add a dedicated kernel for plain 2D PQ2_0 calls with 3-8 columns on Ada. It uses a new activation layout, GGML_CUDA_Q8_1_PQ2: the qs bytes of a column are permuted inside each 16-element group and the half2 (d, int sum) per 32-block follow. With that layout (code_word >> 2k) & 0x03030303 pairs the raw 2-bit codes with the activation bytes, so dp4a needs no per-weight decode, and the digit bias is one integer subtraction. A warp handles 4 rows and reuses the activation slice; registers are capped at 168 (3 blocks per SM). GGML_CUDA_PQ2_MULTICOL=0 turns it off. Other architectures, 1-2 columns, ids and batched calls keep the generic kernel. Tested on RTX 4070 (sm_89): - test-backend-ops MUL_MAT pq2_0: 219/219 pass - kernel, cold L2 (ncu), m=5120 k=17408: n=4 105.6 -> 54.0 us, n=8 136.8 -> 78.1 us - Ternary-Bonsai-2-27B-PQ2_0, llama-batched-bench decode: 4 seq +17%, 6 seq +25%, 8 seq +29% - perplexity (256 ctx, 6 chunks, -ub 8/5/4): 6.3784/6.3784/6.3776 -> 6.3774/6.3774/6.3820 Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The environment switch of the previous commit was only there to measure the change; the multi-column kernel is now selected by the conditions alone. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
Adds a mat-vec kernel for PQ2_0 weights with 3-8 activation columns on Ada (cc 8.9). The generic mmvq kernel drops to 36-66 % of DRAM bandwidth from 4 columns on. The new kernel feeds the raw 2-bit codes straight into
dp4awith an integer correction term (new activation layoutGGML_CUDA_Q8_1_PQ2), 4 rows per warp,__launch_bounds__(128, 3).RTX 4070, m=5120 k=17408 (cold, ncu): 4 columns 105.6 -> 54.0 us, 8 columns 136.8 -> 78.1 us.
Ternary-Bonsai-2-27B-PQ2_0, parallel decode sequences: 4: +17 %, 6: +25 %, 8: +29 %.
Additional information
GGML_CUDA_PQ2_MULTICOL(=0switches the kernel off) that was used for the measurements below, the last one removes it. To reproduce a measurement, build the first commit.cc == GGML_CUDA_CC_ADA_LOVELACEexactly (8.9), the registers (168, 3 blocks per SM) are tuned for that GPU.ggml_cuda_q8_1_layout_hostgot aplain_2dargument; the layout choice for the quantizer and the kernel choice come from the samey_layout, and the PQ2 branch asserts that no ids, batch dims or fusion are involved (fusion with more than one column does not exist for PQ2_0).Test results
-DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=89.prismat88c4bc60b; the four commits since (SYCL, WebGPU andcuda: fused FWHT quantizer for 64-wide warps (#303)) touchquantize.cuonly in the 64-wide-warp FWHT path, not the code of this PR. The branch is rebased on2459f68b5, builds, andtest-backend-opswas repeated on it.test-backend-ops test -b CUDA0 -o MUL_MAT: 1586/1586 passed on the rebased branch (CUDA0 against CPU; includes the PQ2_0 cases with odd row counts). On the earlier base the PQ2_0 subset alone was 219/219.llama-batched-bench): 4 sequences +17 %, 6 sequences +25 %, 8 sequences +29 %.-ub 8 / 5 / 4) before -> after: 6.3784 / 6.3784 / 6.3776 -> 6.3774 / 6.3774 / 6.3820, a difference of at most 0.07 %, far inside the reported error of +-0.53.-np 4, 4 different prompts sent at once, 192 tokens, q4_0 K/V, two rounds per setting): with the new kernel both rounds are identical for all 4 prompts. With the generic kernel (GGML_CUDA_PQ2_MULTICOL=0) two identical runs already differ for 3 of 4 prompts (the arithmetic of the generic kernel depends on the number of columns, and the batch composition varies slightly between runs). Old against new: 2 of 4 prompts identical, the others first differ at character 161 of 848 and 407 of 838.ids) models, column counts outside 3 to 8 (they take the old path), a task-level quality metric.Requirements