Skip to content

cuda: use the transposing concat kernel on all GPUs, not only on GB10 - #309

Open
sb32445 wants to merge 2 commits into
PrismML-Eng:prismfrom
sb32445:pr/concat-transpose-all-gpus
Open

sb32445 wants to merge 2 commits into
PrismML-Eng:prismfrom
sb32445:pr/concat-transpose-all-gpus

Conversation

@sb32445

@sb32445 sb32445 commented Oct 4, 2026

Copy link
Copy Markdown

Overview

build_conv_state() concatenates the conv state [3, C] with a transposed [C, n_tokens] view in every Gated-Delta-Net layer. The shared-memory transpose kernel concat_dim0_transpose_u32 already exists but was enabled only for cc 12.1 (DGX Spark). Other GPUs use the generic non-contiguous kernel (one block per output row, about 6 useful elements per block). This enables the transpose kernel on all GPUs for the same shape conditions.

RTX 4070 (cc 8.9), C=10240, f32, per call: 2 and 3 tokens 8.8 -> 2.5 us (test-backend-ops, local test case). Bonsai 2 27B with MTP (n-max 2, 48 GDN layers): 104.4 -> 105.6 tok/s (+1.2 %), 3 alternating runs, identical outputs. No change for 1 token (contiguous case).

Additional information

  • The branch has two commits: the first contains the environment switch GGML_CUDA_CONCAT_TRANSPOSE (=0 restores the old choice, on cc 12.1 the kernel stays on as before) that was used for the A/B measurements below, the last one removes it. To reproduce a measurement, build the first commit.
  • test-backend-ops CONCAT: 195/195 on CUDA0 (incl. 18 local cases with odd sizes, C=7 and C=100, 1-33 tokens; the cases are not part of this PR).
  • Only measured on one Ada GPU. The kernel itself is unchanged.

Test results

  • Hardware / software: RTX 4070 12 GB (AD104, cc 8.9, 504 GB/s, 48 MB L2), Linux 6.18, NVIDIA driver 615.71, CUDA 13.4, GCC 16.2; Release build, -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=89.
  • Base: speed numbers were measured on prism at 88c4bc60b. The four commits since (SYCL, WebGPU and cuda: fused FWHT quantizer for 64-wide warps (#303)) do not touch the code paths of this PR; the branch is rebased on 2459f68b5, builds, and the test-backend-ops runs below were repeated on it.
  • Model: Ternary-Bonsai-2-27B (PTQ1_0) with a community MTP draft head (Q8_0), speculative decoding with --spec-type draft-mtp --spec-draft-n-max 2, q4_0 K/V cache, one slot.
  • Method: the same binary, one switch (first commit of the branch) flipped through an environment variable, alternating runs A B B A, paired differences with a 95 % bootstrap interval. Run-to-run noise is about 0.04 to 0.12 %, so effects above ~0.3 % are reliable.
  • cmake --build build -j --target test-backend-ops && build/bin/test-backend-ops test -b CUDA0 -o CONCAT: CONCAT 177/177 passed on the rebased branch (CUDA0 against the CPU backend). On the earlier base I also ran 18 additional local cases with odd sizes (C=7, C=100, 1-33 tokens), 195/195; these cases are not part of this PR.
  • Per call, C=10240, f32, 2 and 3 tokens (local test cases, test-backend-ops perf): 8.8 -> 2.5 us. 1 token (contiguous case) is unchanged because it takes the other kernel.
  • Decode with MTP (n-max 2, 48 GDN layers), 4 greedy prompts x 256 tokens, 3 alternating A/B runs with GGML_CUDA_CONCAT_TRANSPOSE=0 / =1: 104.4 -> 105.6 tok/s (+1.2 %), outputs identical in all runs.
  • Not tested: other GPUs (on cc 12.1 the kernel was already enabled before this change, so nothing changes there), HIP and MUSA builds (concat.cu is shared with them, so this change enables the kernel there too; the kernel only uses __syncthreads and a 32x33 shared tile, but I have not run it on those backends and can restrict it to NVIDIA if you prefer), batch sizes above 33 tokens.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: The patches were developed with Claude Code (Anthropic's coding agent): it wrote the code, the measurement scripts and the first drafts of the commit messages and PR texts. I decided what to work on (which kernels and host paths to optimise, based on profiles of my own decode setup). The measurements and checks listed in the PR texts were run in the Claude Code sessions; I did not re-run them independently. I will maintain the changes. Commits where Claude Code was used carry a Co-Authored-By trailer.

sb32445 and others added 2 commits October 4, 2026 14:21
build_conv_state() concatenates the conv state [3, C] with a transposed
[C, n_tokens] view for every Gated-Delta-Net layer. The shared-memory transpose
kernel concat_dim0_transpose_u32 already exists but was enabled only for cc 12.1
(DGX Spark); other GPUs used the generic non-contiguous kernel with one block
per output row and about 6 useful elements per block.

RTX 4070 (cc 8.9), C=10240, f32: 2 tokens 8.8 -> 2.5 us, 3 tokens 8.8 -> 2.5 us
per call (test-backend-ops, local test case). Bonsai 2 27B with MTP (n-max 2),
48 GDN layers: 104.4 -> 105.6 tok/s (+1.2 %, 3 alternating runs, outputs
identical). No change for 1 token (contiguous case, other kernel).

GGML_CUDA_CONCAT_TRANSPOSE=0 restores the old choice. test-backend-ops CONCAT
passed on CUDA0 (195 cases incl. odd sizes C=7 and C=100, up to 33 tokens).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The environment switch of the previous commit was only there to measure the
change; the transposing kernel is now used on all GPUs unconditionally.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant