Conversation
- Fixed-chain speculator base with EAGLE3 and Qwen3.5 MTP draft models, target-native speculation, and method-keyed draft weight registry - Block speculative verification: paged CuTe verification attention, SM90 GDR verify and commit kernels with persistent pipelining - Unified parameterized batch-op pipeline; executor forward split into target pass and speculative round; host batch ops moved engine-side - Scheduler: unified required admission, simplified rollback, submitted rows carry producer-set effects instead of a kind tag - Op-level pytest suites: verification attention, draft carry, speculative sampling, speculative sequence, target hidden projection, copy
…d 'mtp'
The CLI exposes qwen3_5_mtp/hy3_mtp/deepseek_mtp, but TurboMind registers the
MTP head under the C++ name "mtp" (TM_REGISTER_SPECULATIVE_MODEL("mtp", ...)).
Without normalization, --speculative-algorithm qwen3_5_mtp hit
build_draft_model's ValueError and the C++ 'unknown speculative method' check.
Add normalize_spec_method() and apply it at both the draft-model builder and the
EngineConfig.spec_method assignment so the registered name reaches C++.
…TP on 2080Ti) ForwardSpeculativeRound -> CommitAcceptedRecurrentStateKernel dispatches the recurrent-state dtype through TM_DISPATCH_DTYPES with only (bfloat16_t, float). On pre-SM80 GPUs (no bf16 tensor core) the engine dtype is float16, so the GDN recurrent state is f16 and the speculative round aborts with 'unsupported type: f16'. The kernel body computes in float and only touches StateT via ToFloat/FromFloat, and the input side already dispatches half_t, so adding half_t is a correctness fix for SM75 MTP. Not caught by CI (A100/bf16 path never hits this branch).
…idth CI lint's docformatter hook (v1.7.7, --wrap-descriptions 120) rejects the 82-char one-line summary. Shortened to 58 chars; ruff + docformatter clean.
…ixes Inherited CI blockers from InternLM#5006 (this PR is stacked on it): - unit_test (exit 2, collection error): 3 top-level 'import _turbomind' fail under CI 'pip install -e .' because the bare module name is not on sys.path. Switch to 'from lmdeploy.turbomind import _tm' (canonical, matches test_linear.py) and function-level imports to the package name. - lint (ruff): F841 dead var, F401 unused imports, I001 import sorting. Verified in CI-matched env (0.18.0 + cu12.8 + built _tm.so + torch): 308 tests collected, 0 collection errors (was 213 collected, 2 errors).
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Stacked on #5006 (TurboMind speculative decoding). This PR adds the two fixes required to make MTP speculative decoding actually work on pre-SM80 GPUs (Turing, e.g. RTX 2080Ti / sm_75). Both bugs are invisible on the CI's A100/SM80+ because they only manifest with a float16 compute dtype.
Once #5006 is merged, the diff of this PR shrinks to exactly the two fix commits.
Two fixes
1.
qwen3_5_mtpmethod alias never reaches the engine (ffb4c96e)The CLI exposes
--speculative-algorithm qwen3_5_mtp(andhy3_mtp,deepseek_mtp), but TurboMind's C++ side registers the MTP model under the plain namemtp(TM_REGISTER_SPECULATIVE_MODEL("mtp", ...)). With no normalization in between:build_draft_model()only special-casedmethod == 'mtp', soqwen3_5_mtpfell into the EAGLEDRAFT_WEIGHT_SPECSlookup and raisedEngineConfig.spec_methodwas passed verbatim to C++, which rejects the unknown nameAdd
normalize_spec_method()and apply it in both places so every MTP alias resolves to the registeredmtp.2. GDN recurrent-state commit kernel missing
half_t(00feb403)Qwen3.5/Ornith are hybrid-attention models (gated delta-net linear layers + full-attention layers); the MTP draft model carries the linear layers too. After the speculative round accepts tokens,CommitAcceptedRecurrentStateKernelwrites the recurrent state back. Its state-dtype dispatch only listed(bfloat16_t, float):On pre-SM80 there is no bf16 tensor core, so the engine runs float16 and the GDN state is f16 -> hard abort on the first inference request (verified:
Aborted (core dumped), 4/4 ranks). The kernel body already computes in float with genericToFloat/FromFloatconversion, and the sibling input-side dispatch in the same function already acceptshalf_t, so addinghalf_tto the state dispatch is safe and minimal (1 line).SM75 validation (RTX 2080Ti x4, tp=4)
Model: Qwen3.8-27B-FP8 (
qwen3_5arch, 22mtp.*tensors),--speculative-algorithm qwen3_5_mtp --speculative-num-draft-tokens 3.unsupported type: f16FATAL, process core-dumps (restart loop)Files
lmdeploy/turbomind/spec_decode.py—normalize_spec_method()+ MTP branchlmdeploy/turbomind/turbomind.py—spec_methodnormalized before C++src/turbomind/kernels/linear_attn/gdn_state_transaction.cu— state dispatch+ half_t