Conversation
…t CTAs Every PTQ1_0 mat-vec in the decode graph pays about 2.5 us of ramp-up and ramp-down in which the DRAM is not busy. The node loop now looks ahead for the next PTQ1_0 MUL_MAT (a gate/up pair feeding one GLU counts as one op) and hands its weight pointer to the launcher. The last 46 CTAs of the running kernel issue prefetch.global.L2 for the first 50 % (at most 16 MiB) of those weights right after their own loads. More CTAs or more bytes compete with the running kernel and get slower (184 CTAs: -2.3 %). Values are not changed. GGML_CUDA_L2_PREFETCH_PCT=0 turns it off; _CTAS and _MAX_KB tune it. Decode with MTP n-max 2, same binary, outputs identical: +2.00 % (greedy benchmark), +2.10 % (Hermes setup, ctx 114688), +2.2 / +2.1 / +1.7 % at depth 0 / 16k / 65k. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
GGML_CUDA_L2_PREFETCH_PCT, _MAX_KB and _CTAS of the previous commit were only there to tune and measure the change. The measured values (50 % of the next tensor, at most 16 MiB, 46 CTAs) are now constants. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
Every PTQ1_0 mat-vec in the decode graph pays about 2.5 us of ramp-up and ramp-down in which DRAM is idle (399 mat-vecs per verify step). The CUDA node loop now looks ahead for the next PTQ1_0
MUL_MAT(a gate/up pair that feeds one GLU counts as one op), stores its weight pointer and size in a thread-local hint (l2-hint.cuh), and the launcher ofmul_mat_vec_ptq1_0_ptpasses it to the kernel. The last 46 CTAs issueprefetch.global.L2for the first 50 % (at most 16 MiB) of those weights right after their own loads. Values are not changed.The numbers matter: more CTAs or more bytes compete with the running kernel for DRAM and are slower (184 CTAs: -2.3 %; 60 % from 92 CTAs: +0.6 %; 50 % from 46 CTAs: +2.0 %). RTX 4070, Bonsai 2 27B PTQ1_0 + MTP head (n-max 2), 8 interleaved A/B pairs of the same binary, outputs identical in all runs: greedy benchmark 108.31 -> 110.50 tok/s (+2.00 %, CI [+1.92, +2.08] %), agent-like setup (context 114688, thinking sampling, reasoning budget) 101.11 -> 103.22 tok/s (+2.10 %, CI [+2.04, +2.13] %); depth 0 / 16k / 65k +2.2 / +2.1 / +1.7 % (one run each, identical output hashes).
Additional information
GGML_CUDA_L2_PREFETCH_PCT(=0turns it off),_CTASand_MAX_KB(defaults 50 / 46 / 16384) that were used for the tuning and measurements below, the last one fixes the measured values as constants and removes the switches. To reproduce a measurement, build the first commit.ptxas:griddepcontrolrequires sm_90).ggml_cuda_kernel_launchalready opts into PDL there, so the idle window this PR fills may already be covered on those GPUs; I did not test whether the prefetch helps or hurts there, the default applies to all GPUs, and I can restrict it to cc 8.9.thread_localset by the node loop before each node and read by the launcher; the kernel receives the pointer, the number of lines per CTA and the number of prefetching CTAs as three extra arguments. The pointer and byte count come from the next PTQ1_0 weight tensor (at most 50 % of it, at most 16 MiB, so never beyond the tensor). The "last CTAs" are assumed to be the ones that run last, which the hardware does in practice but does not guarantee; if not, only the speed-up shrinks.Test results
-DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=89.prismat88c4bc60bplus my other patches (this PR is independent of them in code); the four commits since (SYCL, WebGPU andcuda: fused FWHT quantizer for 64-wide warps (#303)) do not touch these code paths. The branch is rebased on2459f68b5, builds, andtest-backend-opswas repeated on it.--spec-type draft-mtp --spec-draft-n-max 2, q4_0 K/V cache, one slot.GGML_CUDA_L2_PREFETCH_PCT=0against the default, first commit of the branch) flipped through an environment variable, 4 greedy prompts x 256 tokens, alternating runs A B B A, 8 pairs, paired differences with a 95 % bootstrap interval; run-to-run noise about 0.04 to 0.12 %. The baseline drifts by up to ~1 % between sessions, only paired numbers count.--reasoning-format deepseek --reasoning-budget 16384, thinking sampling 1.0 / 0.95 / 20 / 0.05, fixed seed) 101.11 -> 103.22 tok/s, +2.10 % (CI [+2.04, +2.13] %). Outputs identical in all 16 runs, acceptance unchanged.test-backend-ops test -b CUDA0 -o MUL_MAT,MUL_MAT_VEC_FUSION,GLU,SWIGLU: 2116/2116 passed on the rebased branch (CUDA0 against CPU). For the stack with all my patchesMUL_MAT, MUL_MAT_ID, MUL_MAT_VEC_FUSION, GLU, SWIGLU, RMS_NORM, FLASH_ATTN_EXT, CONCAT: 6373/6373.data), several slots, models other than PTQ1_0 (only this kernel reads the hint).Requirements