Conversation
build_conv_state() concatenates the conv state [3, C] with a transposed [C, n_tokens] view for every Gated-Delta-Net layer. The shared-memory transpose kernel concat_dim0_transpose_u32 already exists but was enabled only for cc 12.1 (DGX Spark); other GPUs used the generic non-contiguous kernel with one block per output row and about 6 useful elements per block. RTX 4070 (cc 8.9), C=10240, f32: 2 tokens 8.8 -> 2.5 us, 3 tokens 8.8 -> 2.5 us per call (test-backend-ops, local test case). Bonsai 2 27B with MTP (n-max 2), 48 GDN layers: 104.4 -> 105.6 tok/s (+1.2 %, 3 alternating runs, outputs identical). No change for 1 token (contiguous case, other kernel). GGML_CUDA_CONCAT_TRANSPOSE=0 restores the old choice. test-backend-ops CONCAT passed on CUDA0 (195 cases incl. odd sizes C=7 and C=100, up to 33 tokens). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The environment switch of the previous commit was only there to measure the change; the transposing kernel is now used on all GPUs unconditionally. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
build_conv_state()concatenates the conv state [3, C] with a transposed [C, n_tokens] view in every Gated-Delta-Net layer. The shared-memory transpose kernelconcat_dim0_transpose_u32already exists but was enabled only for cc 12.1 (DGX Spark). Other GPUs use the generic non-contiguous kernel (one block per output row, about 6 useful elements per block). This enables the transpose kernel on all GPUs for the same shape conditions.RTX 4070 (cc 8.9), C=10240, f32, per call: 2 and 3 tokens 8.8 -> 2.5 us (test-backend-ops, local test case). Bonsai 2 27B with MTP (n-max 2, 48 GDN layers): 104.4 -> 105.6 tok/s (+1.2 %), 3 alternating runs, identical outputs. No change for 1 token (contiguous case).
Additional information
GGML_CUDA_CONCAT_TRANSPOSE(=0restores the old choice, on cc 12.1 the kernel stays on as before) that was used for the A/B measurements below, the last one removes it. To reproduce a measurement, build the first commit.Test results
-DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=89.prismat88c4bc60b. The four commits since (SYCL, WebGPU andcuda: fused FWHT quantizer for 64-wide warps (#303)) do not touch the code paths of this PR; the branch is rebased on2459f68b5, builds, and thetest-backend-opsruns below were repeated on it.--spec-type draft-mtp --spec-draft-n-max 2, q4_0 K/V cache, one slot.cmake --build build -j --target test-backend-ops && build/bin/test-backend-ops test -b CUDA0 -o CONCAT: CONCAT 177/177 passed on the rebased branch (CUDA0 against the CPU backend). On the earlier base I also ran 18 additional local cases with odd sizes (C=7, C=100, 1-33 tokens), 195/195; these cases are not part of this PR.test-backend-ops perf): 8.8 -> 2.5 us. 1 token (contiguous case) is unchanged because it takes the other kernel.GGML_CUDA_CONCAT_TRANSPOSE=0/=1: 104.4 -> 105.6 tok/s (+1.2 %), outputs identical in all runs.concat.cuis shared with them, so this change enables the kernel there too; the kernel only uses__syncthreadsand a 32x33 shared tile, but I have not run it on those backends and can restrict it to NVIDIA if you prefer), batch sizes above 33 tokens.Requirements