Skip to content

fix(turbomind): SM75 (pre-SM80) fixes for MTP speculative decoding [stacked on #5006] - #5019

Open
bltcn wants to merge 39 commits into
InternLM:mainfrom
bltcn:feat/turbomind-mtp-pr5006
Open

bltcn wants to merge 39 commits into
InternLM:mainfrom
bltcn:feat/turbomind-mtp-pr5006

Conversation

@bltcn

@bltcn bltcn commented Oct 3, 2026

Copy link
Copy Markdown
Contributor

Summary

Stacked on #5006 (TurboMind speculative decoding). This PR adds the two fixes required to make MTP speculative decoding actually work on pre-SM80 GPUs (Turing, e.g. RTX 2080Ti / sm_75). Both bugs are invisible on the CI's A100/SM80+ because they only manifest with a float16 compute dtype.

Once #5006 is merged, the diff of this PR shrinks to exactly the two fix commits.

Two fixes

1. qwen3_5_mtp method alias never reaches the engine (ffb4c96e)

The CLI exposes --speculative-algorithm qwen3_5_mtp (and hy3_mtp, deepseek_mtp), but TurboMind's C++ side registers the MTP model under the plain name mtp (TM_REGISTER_SPECULATIVE_MODEL("mtp", ...)). With no normalization in between:

  • build_draft_model() only special-cased method == 'mtp', so qwen3_5_mtp fell into the EAGLE DRAFT_WEIGHT_SPECS lookup and raised
  • EngineConfig.spec_method was passed verbatim to C++, which rejects the unknown name

Add normalize_spec_method() and apply it in both places so every MTP alias resolves to the registered mtp.

2. GDN recurrent-state commit kernel missing half_t (00feb403)

Qwen3.5/Ornith are hybrid-attention models (gated delta-net linear layers + full-attention layers); the MTP draft model carries the linear layers too. After the speculative round accepts tokens, CommitAcceptedRecurrentStateKernel writes the recurrent state back. Its state-dtype dispatch only listed (bfloat16_t, float):

[TM][FATAL][gdn_state_transaction.cu:385] unsupported type: f16
  [1] ModelExecutor::Impl::ForwardSpeculativeRound()

On pre-SM80 there is no bf16 tensor core, so the engine runs float16 and the GDN state is f16 -> hard abort on the first inference request (verified: Aborted (core dumped), 4/4 ranks). The kernel body already computes in float with generic ToFloat/FromFloat conversion, and the sibling input-side dispatch in the same function already accepts half_t, so adding half_t to the state dispatch is safe and minimal (1 line).

SM75 validation (RTX 2080Ti x4, tp=4)

Model: Qwen3.8-27B-FP8 (qwen3_5 arch, 22 mtp.* tensors), --speculative-algorithm qwen3_5_mtp --speculative-num-draft-tokens 3.

  • Before fix 2: first request -> unsupported type: f16 FATAL, process core-dumps (restart loop)
  • After: 50/50 requests succeed, service stays up
  • MTP acceptance stats (vLLM-style): draft acceptance ~55.6%, mean acceptance length ~2.67 (of 3)
  • Short-context decode benchmark (c4, 50 req, 512 tok, thinking off): MTP ON 251.0 tok/s vs OFF 152.9 tok/s = 1.64x
  • Long-context agent benchmark (128k/256k, c8, no max_tokens, needle-in-context tool call): recall 8/8 both ON and OFF (no regression); 256k x c8 peaks at 98.6% KV with no OOM

Files

  • lmdeploy/turbomind/spec_decode.py — normalize_spec_method() + MTP branch
  • lmdeploy/turbomind/turbomind.py — spec_method normalized before C++
  • src/turbomind/kernels/linear_attn/gdn_state_transaction.cu — state dispatch + half_t

bltcn and others added 30 commits December 7, 2025 00:35
actions-user and others added 9 commits February 4, 2026 06:45
- Fixed-chain speculator base with EAGLE3 and Qwen3.5 MTP draft models,
  target-native speculation, and method-keyed draft weight registry
- Block speculative verification: paged CuTe verification attention,
  SM90 GDR verify and commit kernels with persistent pipelining
- Unified parameterized batch-op pipeline; executor forward split into
  target pass and speculative round; host batch ops moved engine-side
- Scheduler: unified required admission, simplified rollback, submitted
  rows carry producer-set effects instead of a kind tag
- Op-level pytest suites: verification attention, draft carry,
  speculative sampling, speculative sequence, target hidden projection,
  copy
…d 'mtp'

The CLI exposes qwen3_5_mtp/hy3_mtp/deepseek_mtp, but TurboMind registers the
MTP head under the C++ name "mtp" (TM_REGISTER_SPECULATIVE_MODEL("mtp", ...)).
Without normalization, --speculative-algorithm qwen3_5_mtp hit
build_draft_model's ValueError and the C++ 'unknown speculative method' check.
Add normalize_spec_method() and apply it at both the draft-model builder and the
EngineConfig.spec_method assignment so the registered name reaches C++.
…TP on 2080Ti)

ForwardSpeculativeRound -> CommitAcceptedRecurrentStateKernel dispatches the
recurrent-state dtype through TM_DISPATCH_DTYPES with only (bfloat16_t, float).
On pre-SM80 GPUs (no bf16 tensor core) the engine dtype is float16, so the
GDN recurrent state is f16 and the speculative round aborts with
'unsupported type: f16'. The kernel body computes in float and only touches
StateT via ToFloat/FromFloat, and the input side already dispatches half_t, so
adding half_t is a correctness fix for SM75 MTP. Not caught by CI (A100/bf16
path never hits this branch).
…idth

CI lint's docformatter hook (v1.7.7, --wrap-descriptions 120) rejects the
82-char one-line summary. Shortened to 58 chars; ruff + docformatter clean.
…ixes

Inherited CI blockers from InternLM#5006 (this PR is stacked on it):
- unit_test (exit 2, collection error): 3 top-level 'import _turbomind' fail
  under CI 'pip install -e .' because the bare module name is not on sys.path.
  Switch to 'from lmdeploy.turbomind import _tm' (canonical, matches
  test_linear.py) and function-level imports to the package name.
- lint (ruff): F841 dead var, F401 unused imports, I001 import sorting.

Verified in CI-matched env (0.18.0 + cu12.8 + built _tm.so + torch):
308 tests collected, 0 collection errors (was 213 collected, 2 errors).

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants