A vLLM fork that focuses on running frontier models on older cards like A6000, 3090 and A100.
Status:
| Supported Models | Quantization | Status |
|---|---|---|
DeepSeek-v4-Flash-0731 |
Native FP4 | Fully Supported (0.6.0+) |
Qwen3.8-27B |
BF16, AWQ W4A16 | Fully Supported (v0.8.0+) |
Qwen3.8-Flash-Next |
FP8, AWQ W4A16 | Fully Supported (v0.9.0+) |
GLM-5.3-Flash |
AWQ W4A16 | Fully Supported (v0.11.2+) |
Note we have a paired LMCache fork for production kvcache serving, which is also built into our docke images.
Prebuilt images are published to Docker Hub on every push:
| Image | Target GPUs |
|---|---|
lazymio/vllm-backport:latest-sm86 (also :latest) |
Ampere sm86 (A6000, RTX 30xx) |
lazymio/vllm-backport:latest-sm80 |
Ampere sm80 (A100) |
lazymio/vllm-backport:latest-sm89 |
Ada sm89 (RTX 4090, L40S) |
lazymio/vllm-backport:v0.11.2-sm86 / -sm80 / -sm89 |
pinned release builds |
Images are single-arch builds (no FA3/Hopper kernels), so pick the tag matching your GPU. The entrypoint is vllm serve and lmcache is also available within the same image!
:latest* tags track the main branch; each release also ships versioned tags like :v0.11.2-sm86 if you want to pin. Check Dockerhub or Github for available latest tags.
Note the sample below includes LMCache. If you do not have enough RAM or disk, you could remove the --kv-transfer-config and the full LMCache section.
services:
vllm:
image: lazymio/vllm-backport:latest
depends_on: [lmcache] # Remove this line if you remove the lmcache section below
command:
- deepseek-ai/DeepSeek-V4-Flash-0731
- '--kv-transfer-config={"kv_connector":"LMCacheMPConnector","kv_connector_module_path":"lmcache.integration.vllm.lmcache_mp_connector","kv_role":"kv_both","kv_connector_extra_config":{"lmcache.mp.host":"127.0.0.1","lmcache.mp.port":5556}}'
- ... (model specific commands see below)
ports:
- "8000:8000"
volumes:
- ~/.cache/huggingface:/root/.cache/huggingface
environment:
- HUGGING_FACE_HUB_TOKEN=${HUGGING_FACE_HUB_TOKEN:-}
ipc: host
restart: unless-stopped
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]
lmcache:
image: lazymio/vllm-backport:latest
entrypoint: ["lmcache", "server"]
command:
- --host=127.0.0.1
- --port=5556
- --http-port=18556
- --chunk-size=800
- --separate-object-groups
- --l1-size-gb=1024
- --eviction-policy=LRU
- --max-workers 16
- --l2-adapter '{"type":"fs_native","base_path":"path/to/l2","num_workers":32,"max_capacity_gb":1024,"eviction":{"eviction_policy":"LRU","trigger_watermark":0.8,"eviction_ratio":0.2}}'
environment:
- PYTHONUNBUFFERED=1
- CUDA_DEVICE_ORDER=PCI_BUS_ID
restart: unless-stopped
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]Then:
docker compose up -d
curl http://localhost:8000/v1/modelsFULL_AND_PIECEWISEcaptures the whole decode step — attention, MoE dispatch, NCCL all-reduce and the DSpark draft loop — into one CUDA graph. Measured on 4x/8x A6000: single-stream decode 45.6 -> 67-70 tok/s (+47%); prefill is unchanged (compute-bound). Two prerequisites:- Pin NCCL with
NCCL_ALGO=Ring NCCL_PROTO=Simple(and prefer--disable-custom-all-reduce). Graph replay must re-issue the exact captured collective; NCCL's size-adaptive algorithm switching is what made FULL capture "crash on Ampere" — Ampere itself is fine. - Bound
cudagraph_capture_sizesas shown. FULL graphs keep private memory pools; capturing every batch size up to--max-num-seqscan cost >800 MB per GPU and OOM warmup at high--gpu-memory-utilization.
- Pin NCCL with
- Adjust your TP (--tensor-parallel-size), PP (--pipeline-parallel-size) and EP (--enable-expert-parallel) accordingly.
NCCL_ALGO=Ring NCCL_PROTO=Simple
vllm serve wtdcode/GLM-5.3-Flash-AWQ-W4A16 \
--host 0.0.0.0 \
--port 8000 \
--served-model-name glm-5.3-flash \
--tensor-parallel-size 4 \
--max-model-len 524288 \
--gpu-memory-utilization 0.92 \
--max-num-seqs 16 \
--max-num-batched-tokens 8192 \
--enable-prefix-caching \
--enable-prompt-tokens-details \
--disable-custom-all-reduce \
--compilation-config '{"cudagraph_mode":"FULL_AND_PIECEWISE","cudagraph_capture_sizes":[1,2,4,8,16],"max_cudagraph_capture_size":16}' \
--speculative-config '{"method":"mtp","num_speculative_tokens":3}' \
--enable-auto-tool-choice \
--tool-call-parser glm47 \
--reasoning-parser glm45- Verified with MTP-3 on 4x A100-80GB (sm80). On 8x RTX A6000 (sm86) use
--tensor-parallel-size 8and drop--gpu-memory-utilizationto 0.85 — with the MTP draft loaded, 0.9+ OOMs during cudagraph warmup on 48 GB cards. - Serving directly from the Hugging Face repository ID works; no manual snapshot-path resolution is needed.
VLLM_PLE_CPU_OFFLOAD=1 VLLM_ALLOW_LONG_MAX_MODEL_LEN=1
vllm serve /path/to/your/qwen3.8 \
--host=0.0.0.0 \
--port=19088 \
--served-model-name=qwen3.8-flash-next \
--tensor-parallel-size=4 \
--enable-expert-parallel \
--max-model-len=1000000 \
--max-num-seqs=8 \
--max-num-batched-tokens=2048 \
--gpu-memory-utilization=0.9 \
--disable-custom-all-reduce \
--speculative-config '{"method":"mtp","num_speculative_tokens":3}' \
--compilation-config '{"cudagraph_mode":"FULL_AND_PIECEWISE"}' \
--hf-overrides '{"text_config": {"rope_parameters": {"mrope_interleaved": true, "mrope_section": [11, 11, 10], "rope_type": "yarn", "rope_theta": 10000000, "partial_rotary_factor": 0.25, "factor": 4.0, "original_max_position_embeddings": 262144}}}' \
--enable-auto-tool-choice \
--tool-call-parser=qwen3_xml \
--reasoning-parser=qwen3 \
--enable-prompt-tokens-details \
--enable-prefix-caching \
--mamba-cache-mode=align \VLLM_PLE_CPU_OFFLOAD=1keeps the 51B n-gram embedding (fp8, ~51 GiB) in pinned host RAM via a separatePleOffloadWorkerprocess. Without it the TP-sharded embedding adds ~12.8 GiB per GPU and KV memory goes negative on 48 GB cards.--enable-expert-parallelis required, not optional: with plain TP the 640-wide expert intermediate becomes 160 per rank, which is not a multiple of the 128x128 fp8 block, and vLLM then forces the Triton fp8 MoE kernel (no fp8 tensor cores on sm86). With EP the experts stay whole and the Marlin W8A16 backend is used.- AWQ W4A16 (
wtdcode/Qwen3.8-Flash-Next-AWQ-W4A16, compressed-tensorspack-quantized, routed experts INT4 g128, everything else BF16): MTP speculative decoding works (the BF16 MTP draft is kept unquantized automatically). Verified on 4x A100-80GB:VLLM_PLE_CPU_OFFLOAD=1 vllm serve wtdcode/Qwen3.8-Flash-Next-AWQ-W4A16 --tensor-parallel-size 4 --enable-expert-parallel --compilation-config '{"mode":0,"cudagraph_mode":"FULL_DECODE_ONLY"}' --speculative-config '{"method":"mtp","num_speculative_tokens":3}'. This achieves up to 936 tps.
vllm serve /path/to/your/deepseek \
--tensor-parallel-size 8 \
--max-model-len 1048576 \
--gpu-memory-utilization 0.90 \
--kv-cache-dtype fp8_ds_mla \
--trust-remote-code \
--disable-custom-all-reduce \
--compilation-config '{"cudagraph_mode":"FULL_AND_PIECEWISE","cudagraph_capture_sizes":[1,2,4,8,16,32,64],"max_cudagraph_capture_size":64}' \
--speculative-config '{"method":"dspark","num_speculative_tokens":5}' \
--enable-auto-tool-choice --tool-call-parser deepseek_v4 \
--host 0.0.0.0 --port 8000 \
--served-model-name deepseek-v4-flash- Keep
num_speculative_tokensat 5 on Ampere. Values below 5 (the checkpoint'sdspark_block_size) are rejected, and 7 needs ~200 KB of shared memory vs the 163 KB Ampere limit (triton OutOfResourceserror). 6 does start, but draft positions past the native block are almost never accepted (3–13% in our measurements), so it only wastes draft compute — output quality and speed are the same as 5. - Requests that set neither
thinkingnorreasoning_effortnow get thinking mode with high effort, matching the official 0731 API mapping (reasoning_effort: "none"restores plain chat mode). Agentic/tool-calling clients should pass areasoning_effortexplicitly from the first turn of a session — sessions that run without the effort prefix gradually stop thinking and can enter self-reinforcing reasoning loops. --hf-overrides '{"head_dtype": "float32"}'once helped reduce garbage outputs by improving precisions but might be not compulsory.