Repository navigation
Add opt-in fused Qwen3.5 inference kernels - #29
Open
BruceSheng1202 wants to merge 3 commits into
Open
BruceSheng1202 wants to merge 3 commits into
BruceSheng1202 wants to merge 3 commits into
Conversation
LoadOptions(fused_kernels=True), JEVANY_FUSED_KERNELS=1 or --fused-kernels rewrites a merged Qwen3.5 backbone in place for batch-1 inference: one FLA kernel per RMSNorm; the MLP gate/up, Gated DeltaNet and attention input projections as single GEMMs; the DeltaNet gate, beta sigmoid and grouped value heads inside FLA's chunked kernel; FLA's fused gated RMSNorm; and transposed weight storage so cuBLAS reads both GEMM operands in their natural layout. The original modules keep views of the fused weights, so memory does not grow. Calls with a cache, an attention mask or packed-sequence metadata use the original forwards. The option needs the LoRA merged on one CUDA device and flash-linear-attention>=0.5. It rejects compile_mode, device_map and choice readout, and describe() reports what was fused. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EpJAgfodeE1m4Fk2tDiLfB
Add JEVANY_FUSED_KERNELS / --fused-kernels to the acceleration table, describe what the option rewrites and requires, and record the A100 JevAny-Qwen3.5-4B latency and Transfer-v9 parity. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EpJAgfodeE1m4Fk2tDiLfB
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EpJAgfodeE1m4Fk2tDiLfB
Human Review
Human Review: see the findings and review evidence. AI semantic review has not completed.
Reviewed head |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Opt-in fast path for Qwen3.5-architecture backbones (
model_typeqwen3_5, including the Qwen3.8 bases):LoadOptions(fused_kernels=True),JEVANY_FUSED_KERNELS=1, or--fused-kernelsonjevany serve,jevany decideandscripts/benchmark_latency.py. Off by default.fuse_qwen3_5rewrites the merged, eval-mode backbone in place:A_log,dt_bias), beta sigmoid, q/k L2 norm and grouped value heads itself; FLA's fused gated RMSNorm follows;The original
nn.Linearmodules keep views of the fused matrices, so memory does not grow (peak 9.93 GiB with and without). The DeltaNet path runs only without a cache, mask or packed-sequence metadata, attention only without a cache; other calls and training use the original forwards. Loading requires CUDA, a merged LoRA (JEVANY_MERGE_BF16=1for BF16 checkpoints), one device and flash-linear-attention >= 0.5, and rejectscompile_mode,device_mapand choice readout.describe()reports what was fused.Measurements
A100-SXM4-40GB, one node: JevAny-Qwen3.5-4B-LoRA, BF16 merged LoRA, SDPA, FLA 0.5.2, causal-conv1d 1.7.0, torch 2.8.0, transformers 5.17.0.
python -m scripts.benchmark_latency --records 0 --repeats 3 --serving-kernels --cuda-graphs [--fused-kernels], public JevBench (231 requests x 3), forward ms:perf/cuda-graph-flash-attention, 4096 limitQuality with graphs and serving kernels (
jevany.benchmark): Transfer-v9 dev 823 -> 824 / 1,046 correct, NLL 0.5869 -> 0.5875, ECE 0.035 -> 0.029; JevBench 180 -> 180 / 231; 5 of 1,264 Transfer and 0 of 231 JevBench argmax decisions changed; largest probability change 0.044. On a random 4-layer Qwen3.5 the fused backbone matches transformers to 1.2e-5 relative error in fp32 and 1.2e-2 in BF16. Not measured: H200 and 27B.Tests
nn.Linearoutputs and share storage; backbone, device, CLI and choice-readout checks (transformers 5.17 and 5.18).pytest tests -m "not server"on the A100 node: 31 failures vs 30 onmain; the extratest_moe_public_trainer_and_checkpoint[bf16-glm]also fails onmainwhen rerun alone there.Not in this PR: other ideas from "Running VLAs at Real-time Speed" (arXiv 2510.26742)
With both branches, a 128-token replay is ~10.8 ms of GPU time: 199 GEMMs 7.3 ms (1.6x the 4.6 ms weight-read floor), ~420 elementwise/copy kernels 1.5 ms, FLA 1.3 ms, norms 0.4 ms.
Measured but left out: readout head and softmax in the graph with one pinned copy each way (-1.5%); residual add fused into the post-attention norm (-0.5%).
🤖 Generated with Claude Code
https://claude.ai/code/session_01EpJAgfodeE1m4Fk2tDiLfB