Skip to content

cuda: prefetch the next PTQ1_0 mat-vec's weights into L2 from the last CTAs - #314

Open
sb32445 wants to merge 2 commits into
PrismML-Eng:prismfrom
sb32445:pr/ptq1-l2-prefetch
Open

sb32445 wants to merge 2 commits into
PrismML-Eng:prismfrom
sb32445:pr/ptq1-l2-prefetch

Conversation

@sb32445

@sb32445 sb32445 commented Oct 4, 2026

Copy link
Copy Markdown

Overview

Every PTQ1_0 mat-vec in the decode graph pays about 2.5 us of ramp-up and ramp-down in which DRAM is idle (399 mat-vecs per verify step). The CUDA node loop now looks ahead for the next PTQ1_0 MUL_MAT (a gate/up pair that feeds one GLU counts as one op), stores its weight pointer and size in a thread-local hint (l2-hint.cuh), and the launcher of mul_mat_vec_ptq1_0_pt passes it to the kernel. The last 46 CTAs issue prefetch.global.L2 for the first 50 % (at most 16 MiB) of those weights right after their own loads. Values are not changed.

The numbers matter: more CTAs or more bytes compete with the running kernel for DRAM and are slower (184 CTAs: -2.3 %; 60 % from 92 CTAs: +0.6 %; 50 % from 46 CTAs: +2.0 %). RTX 4070, Bonsai 2 27B PTQ1_0 + MTP head (n-max 2), 8 interleaved A/B pairs of the same binary, outputs identical in all runs: greedy benchmark 108.31 -> 110.50 tok/s (+2.00 %, CI [+1.92, +2.08] %), agent-like setup (context 114688, thinking sampling, reasoning budget) 101.11 -> 103.22 tok/s (+2.10 %, CI [+2.04, +2.13] %); depth 0 / 16k / 65k +2.2 / +2.1 / +1.7 % (one run each, identical output hashes).

Additional information

  • The branch has two commits: the first contains the environment switches GGML_CUDA_L2_PREFETCH_PCT (=0 turns it off), _CTAS and _MAX_KB (defaults 50 / 46 / 16384) that were used for the tuning and measurements below, the last one fixes the measured values as constants and removes the switches. To reproduce a measurement, build the first commit.
  • Tuned on one GPU (RTX 4070, 46 SMs, 48 MB L2). Other GPUs may want other values; I am happy to gate it behind cc 8.9 or default it to off if you prefer.
  • A micro benchmark with a chain of streaming kernels showed the same pattern (1-1.6 us saved per kernel, DRAM saturates at ~94 % of peak).
  • On Hopper and newer, programmatic dependent launch would be the cleaner tool; it does not exist on cc 8.9 (ptxas: griddepcontrol requires sm_90). ggml_cuda_kernel_launch already opts into PDL there, so the idle window this PR fills may already be covered on those GPUs; I did not test whether the prefetch helps or hurts there, the default applies to all GPUs, and I can restrict it to cc 8.9.
  • The hint is a thread_local set by the node loop before each node and read by the launcher; the kernel receives the pointer, the number of lines per CTA and the number of prefetching CTAs as three extra arguments. The pointer and byte count come from the next PTQ1_0 weight tensor (at most 50 % of it, at most 16 MiB, so never beyond the tensor). The "last CTAs" are assumed to be the ones that run last, which the hardware does in practice but does not guarantee; if not, only the speed-up shrinks.
  • Weights are not modified, no value changes; outputs are identical.
  • test-backend-ops MUL_MAT, MUL_MAT_ID, MUL_MAT_VEC_FUSION, GLU, SWIGLU, RMS_NORM, FLASH_ATTN_EXT, CONCAT: 6373/6373 on CUDA0 for the full stack.

Test results

  • Hardware / software: RTX 4070 12 GB (AD104, cc 8.9, 46 SMs, 504 GB/s, 48 MB L2), Linux 6.18, NVIDIA driver 615.71, CUDA 13.4, GCC 16.2; Release build, -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=89.
  • Base: the speed numbers were measured on prism at 88c4bc60b plus my other patches (this PR is independent of them in code); the four commits since (SYCL, WebGPU and cuda: fused FWHT quantizer for 64-wide warps (#303)) do not touch these code paths. The branch is rebased on 2459f68b5, builds, and test-backend-ops was repeated on it.
  • Model: Ternary-Bonsai-2-27B (PTQ1_0) with a community MTP draft head (Q8_0), --spec-type draft-mtp --spec-draft-n-max 2, q4_0 K/V cache, one slot.
  • Method: the same binary, one switch (GGML_CUDA_L2_PREFETCH_PCT=0 against the default, first commit of the branch) flipped through an environment variable, 4 greedy prompts x 256 tokens, alternating runs A B B A, 8 pairs, paired differences with a 95 % bootstrap interval; run-to-run noise about 0.04 to 0.12 %. The baseline drifts by up to ~1 % between sessions, only paired numbers count.
  • Tuning, greedy benchmark, 4 pairs per setting (prefetched fraction of the next weights / number of prefetching CTAs): 60 % / 92: +0.6 % (6 pairs); 100 % / 92: +0.14 % (n.s.); 60 % / 184: -2.27 %; 30 % / 92: +1.14 %; 20 % / 92: +1.19 %; 15 % / 92: +1.54 %; 30 % / 138: +0.07 % (n.s.); 30 % / 46: +1.80 %; 15 % / 46: +1.61 %; 30 % / 23: +1.76 %; 15 % / 23: +1.48 %; 30 % / 12: +1.39 %; 50 % / 46: +2.02 %; 70 % / 46: +2.00 %; 100 % / 46: +1.78 %; 50 % / 69: +1.71 %; 70 % / 34: +1.74 %. Outputs identical in all of them.
  • Confirmation with the defaults (50 % / 46 CTAs / 16 MiB), 8 pairs each: greedy benchmark 108.31 -> 110.50 tok/s, +2.00 % (95 % CI [+1.92, +2.08] %); agent-like setup (context 114688, --reasoning-format deepseek --reasoning-budget 16384, thinking sampling 1.0 / 0.95 / 20 / 0.05, fixed seed) 101.11 -> 103.22 tok/s, +2.10 % (CI [+2.04, +2.13] %). Outputs identical in all 16 runs, acceptance unchanged.
  • By depth (one run each, so only indicative; the output hashes are identical): depth 0 / 16384 / 65536: 119.6 -> 122.2 (+2.2 %), 109.0 -> 111.3 (+2.1 %), 92.6 -> 94.2 (+1.7 %) tok/s.
  • Micro benchmark (chain of dependent streaming kernels, not part of the PR): 1.0 to 1.6 us saved per kernel for 6.9 to 39 MB of weights, DRAM saturating at ~94 % of peak with prefetch, the same pattern of "fewer prefetching CTAs is better".
  • test-backend-ops test -b CUDA0 -o MUL_MAT,MUL_MAT_VEC_FUSION,GLU,SWIGLU: 2116/2116 passed on the rebased branch (CUDA0 against CPU). For the stack with all my patches MUL_MAT, MUL_MAT_ID, MUL_MAT_VEC_FUSION, GLU, SWIGLU, RMS_NORM, FLASH_ATTN_EXT, CONCAT: 6373/6373.
  • Not tested: other GPUs, multiple GPUs / tensor split, weights in host memory or unified memory (the prefetch addresses come from the weight tensor's data), several slots, models other than PTQ1_0 (only this kernel reads the hint).

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: The patches were developed with Claude Code (Anthropic's coding agent): it wrote the code, the measurement scripts and the first drafts of the commit messages and PR texts. I decided what to work on (which kernels and host paths to optimise, based on profiles of my own decode setup). The measurements and checks listed in the PR texts were run in the Claude Code sessions; I did not re-run them independently. I will maintain the changes. Commits where Claude Code was used carry a Co-Authored-By trailer.

sb32445 and others added 2 commits October 4, 2026 14:21
…t CTAs

Every PTQ1_0 mat-vec in the decode graph pays about 2.5 us of ramp-up
and ramp-down in which the DRAM is not busy. The node loop now looks
ahead for the next PTQ1_0 MUL_MAT (a gate/up pair feeding one GLU counts
as one op) and hands its weight pointer to the launcher. The last 46
CTAs of the running kernel issue prefetch.global.L2 for the first 50 %
(at most 16 MiB) of those weights right after their own loads. More
CTAs or more bytes compete with the running kernel and get slower
(184 CTAs: -2.3 %). Values are not changed.

GGML_CUDA_L2_PREFETCH_PCT=0 turns it off; _CTAS and _MAX_KB tune it.

Decode with MTP n-max 2, same binary, outputs identical: +2.00 %
(greedy benchmark), +2.10 % (Hermes setup, ctx 114688), +2.2 / +2.1 /
+1.7 % at depth 0 / 16k / 65k.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
GGML_CUDA_L2_PREFETCH_PCT, _MAX_KB and _CTAS of the previous commit were only
there to tune and measure the change. The measured values (50 % of the next
tensor, at most 16 MiB, 46 CTAs) are now constants.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant