Repository navigation
Conversation
The fused add_rms_norm_f32 path (d8d96cf) was unreachable outside GB10 for two independent reasons: 1. The gate `cc == GGML_CUDA_CC_DGX_SPARK` excluded every other NVIDIA device, although the kernel uses no architecture-specific instructions. 2. Even with that gate relaxed, ggml_cuda_check_fusion_memory_ranges rejects the fusion for the residual pattern ggml's allocator actually produces: the residual ADD is in-place (add->data == add->src[1]->data) and the norm output reuses the other input's buffer (mul->data == add->src[0]->data). Both are exact aliases - same base pointer, type, shape and contiguous layout - which the range check treats like a partial overlap. The kernel is elementwise within a row: every thread reads a[col] and b[col] before writing sum[col], and block_reduce synchronizes the block before dst[col] is written. Exact aliases are therefore safe for it, so the range check gains an opt-in `allow_exact_alias` that this fusion passes. Partial overlaps are still rejected for all callers, and the default is unchanged for existing callers. The norm weight is additionally required not to alias either output, as it is broadcast across rows. Verified on RTX 4070 Ti SUPER (sm_89) with Ternary-Bonsai-2-27B PQ2_0: - test-backend-ops -o ADD_RMS_NORM: 25/25 - fusion now fires (96 add_rms_norm_f32 launches per 32 decode tokens, rms_norm_f32<1024> launches drop by the same 96) - throughput unchanged within noise (tg128 63.88 vs 63.89 t/s, pp512 1704 vs 1701, interleaved 3x4 runs): at batch 1 the fusion saves one n_embd-row read per firing, which is negligible against weight traffic. Claude-Session: https://claude.ai/code/session_01UpAMXrmoyeC1gjM77dReJo
bri-prism
left a comment
There was a problem hiding this comment.
Agent review: posted by the maintainer's coding agent at their request.
No findings in this source pass. I traced the exact-alias predicate through the fused kernel's row accesses and reduction, and checked the exclusion for the broadcast norm weight.
CUDA execution was not performed. The remaining validation is a fused-versus-unfused comparison for in-place residual input, norm output reusing an input buffer, and rejected partial overlap, including a non-GB10 NVIDIA device newly enabled by this patch.
Reviewed commit: b989116e1ceef747a659d146b1977709e97342fc.
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
The RMS output can still be misused as an unwritten weight, and the fused alias path lacks regression coverage.
Get a fresh assessment by requesting another Copilot review.
Review effort: Balanced
Findings: 1
Open (2)
|
Agent benchmark follow-up, posted at the maintainer's request. Pinned head
Selected CPU-reference backend checks passed on the compared arms. The selected ADD_RMS_NORM checks were included alongside the matrix-multiplication checks. These are short-context throughput measurements. The percentages describe these paired runs; small changes should not be interpreted as established improvements. No long-context, multi-slot serving, or end-to-end logit-parity claim is made. PQ2 control measurements (paired pp512 / tg128 changes):
|
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
|
@cklxx @sb32445, #209 and #310 both change |


Summary
The fused
add_rms_norm_f32path added in d8d96cf (#135) never fires outside GB10 — and, I believe, doesn't fire on GB10 either for the standard residual pattern. Two independent gates block it:Arch gate.
cc == GGML_CUDA_CC_DGX_SPARKexcludes every other NVIDIA device, butadd_rms_norm_f32uses no architecture-specific instructions (plain loads,block_reduce,rsqrtf). Relaxed toGGML_CUDA_CC_IS_NVIDIA(cc).Memory-range check. With the arch gate relaxed, every candidate still failed
ggml_cuda_check_fusion_memory_ranges. Instrumenting it shows why — for this graph ggml's allocator produces exact aliases:Same base pointer, type, shape and contiguous layout, which the check treats like a partial overlap and rejects.
Since that aliasing comes from the graph allocator rather than the device, the same rejection should happen on GB10 — worth a quick check on your side (a temporary print in the range check is enough).
Change
ggml_cuda_check_fusion_memory_rangesgains an opt-inallow_exact_alias(defaultfalse, existing callers unchanged). An exact alias — identicaldata, type, shape, contiguous — is accepted; partial overlaps are still rejected for everyone.a[col],b[col]before writingsum[col], andblock_reducesyncs the block beforedst[col]is written.Verification (RTX 4070 Ti SUPER, sm_89, Ternary-Bonsai-2-27B PQ2_0)
test-backend-ops -o ADD_RMS_NORM: 25/25;RMS_NORM_MUL_ADD,ADD,RMS_NORM,MULall pass.nsys: fusion now fires — 96
add_rms_norm_f32launches per 32 decode tokens, andrms_norm_f32<1024>launches drop by exactly 96 (258 → 162).Greedy generation output coherent.
Throughput: no measurable change (interleaved base/patched, 3 rounds × 4 reps):
Expected: at batch 1 the fusion saves one
n_embd-row read (~20 KB) per firing, 3 firings per token, against ~6.7 GB of weight traffic per token. The value here is that the kernel actually runs, not a speedup on this workload.Not changed
While instrumenting the range check I also saw the upstream
MUL_MAT, MUL_MAT, GLUfusion rejected 48/128 times on this model. That one is a real partial overlap (GLU output, 69632 B, starts at the same address as the 20480 B activation input) and unsafe to relax — the check is doing its job. Noting it in case it's useful; only fixable at the allocator level.https://claude.ai/code/session_01UpAMXrmoyeC1gjM77dReJo