Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions autotest/configs/Qwen/Qwen2.5-7B-Instruct.yml
Original file line number Diff line number Diff line change
Expand Up @@ -63,10 +63,12 @@ h:
pytorch:
suites:
- toolcall
- hard_schema
extra:
tool-call-parser: qwen2d5
turbomind:
suites:
- toolcall
- hard_schema
extra:
tool-call-parser: qwen2d5
2 changes: 2 additions & 0 deletions autotest/configs/Qwen/Qwen3-8B-FP8.yml
Original file line number Diff line number Diff line change
Expand Up @@ -17,13 +17,15 @@ h:
pytorch:
suites:
- toolcall
- hard_schema
- reasoning
extra:
tool-call-parser: qwen3
reasoning-parser: default
turbomind:
suites:
- toolcall
- hard_schema
- reasoning
extra:
tool-call-parser: qwen3
Expand Down
14 changes: 14 additions & 0 deletions autotest/configs/Qwen/Qwen3.5-27B.yml
Original file line number Diff line number Diff line change
Expand Up @@ -75,3 +75,17 @@ h:
quantization:
turbomind:
- awq
- model_type:
- chat
- vl
engine_config:
tp: 2
extra:
prefix-cache-state-budget: 256
max-prefill-token-num: 64
backends:
- name: pytorch
communicators:
- nccl
test_coverage:
- prefix_cache
59 changes: 13 additions & 46 deletions autotest/configs/Qwen/Qwen3.5-35B-A3B-FP8.yml
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# HF model id: Qwen/Qwen3.5-35B-A3B-FP8

a100:
h:
- model_type:
- chat
- vl
Expand All @@ -10,6 +10,9 @@ a100:
- name: turbomind
communicators:
- nccl
- name: pytorch
communicators:
- nccl
test_coverage:
- benchmark
- evaluate
Expand All @@ -18,42 +21,23 @@ a100:
- longtext_evaluate
- mllm_evaluate
interface:
turbomind:
pytorch:
- suites:
- base
- logprob
- experts
- toolcall
- reasoning
extra:
logprobs-mode: raw_logprobs
enable-return-routed-experts: true
tool-call-parser: qwen3coder
reasoning-parser: default
- suites:
- anthropic
extra:
logprobs-mode: raw_logprobs
- suites:
- abort
extra:
enable-abort-handling: true
h:
- model_type:
- chat
- vl
engine_config:
tp: 1
backends:
- name: turbomind
communicators:
- nccl
test_coverage:
- benchmark
- evaluate
- func
- longtext_benchmark
- longtext_evaluate
- mllm_evaluate
interface:
enable-return-routed-experts: true
turbomind:
- suites:
- base
Expand All @@ -68,29 +52,20 @@ h:
- anthropic
extra:
logprobs-mode: raw_logprobs
- suites:
- abort
extra:
enable-abort-handling: true
- model_type:
- chat
- vl
engine_config:
tp: 2
tp: 1
extra:
prefix-cache-state-budget: 256
max-prefill-token-num: 64
backends:
- name: pytorch
communicators:
- nccl
test_coverage:
- func
interface:
pytorch:
- suites:
- abort
extra:
enable-abort-handling: true
- suites:
- sleep
- prefix_cache
- model_type:
- chat
- vl
Expand All @@ -111,11 +86,3 @@ h:
- longtext_benchmark
- longtext_evaluate
- mllm_evaluate
interface:
pytorch:
- suites:
- abort
extra:
enable-abort-handling: true
- suites:
- sleep
23 changes: 21 additions & 2 deletions autotest/configs/Qwen/Qwen3.5-35B-A3B.yml
Original file line number Diff line number Diff line change
Expand Up @@ -23,6 +23,7 @@ a100:
- logprob
- experts
- toolcall
- hard_schema
- reasoning
extra:
logprobs-mode: raw_logprobs
Expand All @@ -45,6 +46,7 @@ a100:
- base
- logprob
- toolcall
- hard_schema
- reasoning
extra:
logprobs-mode: raw_logprobs
Expand All @@ -58,6 +60,9 @@ a100:
- abort
extra:
enable-abort-handling: true
gen_config:
chat-template-kwargs:
enable_thinking: true
quantization:
turbomind:
- kvint8
Expand Down Expand Up @@ -116,6 +121,7 @@ h:
- logprob
- experts
- toolcall
- hard_schema
- reasoning
extra:
logprobs-mode: raw_logprobs
Expand All @@ -138,6 +144,7 @@ h:
- base
- logprob
- toolcall
- hard_schema
- reasoning
extra:
logprobs-mode: raw_logprobs
Expand All @@ -156,6 +163,20 @@ h:
- awq
pytorch:
- fp8
- model_type:
- chat
- vl
engine_config:
tp: 2
extra:
prefix-cache-state-budget: 256
max-prefill-token-num: 64
backends:
- name: pytorch
communicators:
- nccl
test_coverage:
- prefix_cache
- model_type:
- chat
- vl
Expand Down Expand Up @@ -205,8 +226,6 @@ ascend:
- func
- longtext_benchmark
- mllm_evaluate
- prefix_cache
- mllm_evaluate
- model_type:
- chat
- vl
Expand Down
8 changes: 4 additions & 4 deletions autotest/configs/Qwen/Qwen3.5-397B-A17B-FP8.yml
Original file line number Diff line number Diff line change
Expand Up @@ -61,11 +61,11 @@ h:
ep: 8
extra:
max-batch-size: 256
cache-max-entry-count: 0.7
prefix-cache-decode-state-interval: 1024
cache-max-entry-count: 0.9
prefix-cache-state-budget: 256
max-prefill-token-num: 64
backends:
- name: turbomind
- name: pytorch
communicators:
- nccl
test_coverage:
Expand All @@ -87,7 +87,7 @@ h:
communicators:
- nccl
test_coverage:
- prefix_cache
- prefix_cache_evaluate
- model_type:
- chat
- vl
Expand Down
25 changes: 18 additions & 7 deletions autotest/configs/Qwen/Qwen3.5-397B-A17B.yml
Original file line number Diff line number Diff line change
Expand Up @@ -71,20 +71,32 @@ h:
extra:
max-batch-size: 256
cache-max-entry-count: 0.7
prefix-cache-decode-state-interval: 1024
prefix-cache-state-budget: 256
max-prefill-token-num: 64
backends:
- name: turbomind
- name: pytorch
communicators:
- nccl
test_coverage:
- prefix_cache
- model_type:
- chat
- vl
engine_config:
tp: 2
dp: 4
ep: 8
extra:
max-batch-size: 256
cache-max-entry-count: 0.7
prefix-cache-decode-state-interval: 1024
prefix-cache-state-budget: 256
backends:
- name: pytorch
communicators:
- nccl
test_coverage:
- prefix_cache
quantization:
pytorch:
- fp8
- prefix_cache_evaluate
- model_type:
- chat
- vl
Expand All @@ -105,7 +117,6 @@ h:
- longtext_benchmark
- longtext_evaluate
- mllm_evaluate
- prefix_cache
ascend:
- model_type:
- chat
Expand Down
14 changes: 14 additions & 0 deletions autotest/configs/Qwen/Qwen3.5-4B.yml
Original file line number Diff line number Diff line change
Expand Up @@ -15,3 +15,17 @@ h:
- nccl
test_coverage:
- func
- model_type:
- chat
- vl
engine_config:
tp: 1
extra:
prefix-cache-state-budget: 256
max-prefill-token-num: 64
backends:
- name: pytorch
communicators:
- nccl
test_coverage:
- prefix_cache
25 changes: 24 additions & 1 deletion autotest/configs/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -96,12 +96,16 @@ Available keys:
- `longtext_evaluate`
- `mllm_evaluate`
- `prefix_cache`
- `prefix_cache_evaluate`
- `quantization`

Rules:

- Keep MTP (speculative decoding) on its own row with `speculative-algorithm` in `engine_config.extra`; use `func` and/or `evaluate` in `test_coverage`.
- Use `prefix_cache` in `test_coverage`; do not add `enable-prefix-caching` manually to `engine_config.extra`.
- Use `prefix_cache` / `prefix_cache_evaluate` in `test_coverage`; do not add `enable-prefix-caching` manually to `engine_config.extra`.
- SSM models (Qwen3.5 / Intern-S2): **separate** dedicated rows, **pytorch only** (TurboMind has no serve CLI for checkpoint interval; do not list turbomind on this row). Do not put these knobs on shared `func`/`evaluate` rows. Do not inject extras in Python.
- **Func** (`tools/restful` / pipeline): `test_coverage: [prefix_cache]`; `max-prefill-token-num: 64` + `prefix-cache-state-budget`.
- **Accuracy** (`evaluate` `*_prefix_cache_*` / `submit_prefix_accuracy_job`): own row with `test_coverage: [prefix_cache_evaluate]`; `prefix-cache-state-budget` + `prefix-cache-decode-state-interval`(**不要** `max-prefill-token-num: 64`).
- Use `quantization` in `test_coverage` only for runtime weight-quant rows (`awq`, `gptq`, `w8a8`).

## `interface` (REST interface coverage)
Expand Down Expand Up @@ -189,10 +193,29 @@ is enabled.
Suites:

- `base` — chat/completions basic cases; generate without logprob/experts

- `logprob` — generate logprob cases

- `experts` — generate routed-experts cases
- `anthropic` — Anthropic Messages HTTP + SDK smoke (chat protocol models from yaml). Share a profile with `base`/`logprob` when `extra` matches; otherwise its own profile **without** `tool-call-parser` / `reasoning-parser`.
- `toolcall` — `interface/restful/tool_parser/` (requires `tool-call-parser` in yaml `extra`; add `enable-return-routed-experts: true` when toolcall includes `@experts` cases)

- `hard_schema` — walle MFJS tool-call schema validation (`tool_parser/test_tool_call_json_schema.py`; requires `tool-call-parser`; for thinking models set `gen_config.chat-template-kwargs.enable_thinking: true` (not `interface.extra`); runs full walle case set). Kimi-Vendor-Verifier checkout: `eval_resource/Kimi-Vendor-Verifier` (see `kimi_vendor_verifier_path` in `env_paths.yml`), overridable via `KIMI_VENDOR_VERIFIER_ROOT`; cases override via `WALLE_CASE_DIR`. One representative model per parser (see table below); all `internlm/*` configs with `toolcall` also enable `hard_schema`.

| `tool-call-parser` | Representative model |
| ------------------ | ------------------------------------------------------- |
| `qwen3coder` | `Qwen/Qwen3.5-35B-A3B` |
| `qwen3` | `Qwen/Qwen3-8B-FP8` |
| `qwen2d5` | `Qwen/Qwen2.5-7B-Instruct` |
| `llama3` | `meta-llama/Llama-3.1-70B-Instruct` |
| `glm47` | `zai-org/GLM-4.7-Flash` |
| `kimi-k2` | `moonshotai/Kimi-K2-Instruct-0905` |
| `deepseek-v32` | `deepseek-ai/DeepSeek-V3` |
| `deepseek-v4` | `deepseek-ai/DeepSeek-V4-Flash-0731` |
| `intern-s1` | `internlm/Intern-S1` (+ all Intern-S1 / Intern-S1-mini) |
| `interns2-preview` | `internlm/Intern-S2-Preview` (+ all Intern-S2 variants) |
| `internlm` | `internlm/internlm3-8b-instruct` |

- `reasoning` — `interface/restful/reasoning_parser/` (requires `reasoning-parser` in yaml `extra`)

Notes:
Expand Down
12 changes: 12 additions & 0 deletions autotest/configs/deepseek-ai/DeepSeek-V3.yml
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,18 @@ h:
- nccl
test_coverage:
- func
interface:
pytorch:
- suites:
- base
- logprob
- toolcall
- hard_schema
extra:
logprobs-mode: raw_logprobs
enable-return-routed-experts: true
tool-call-parser: deepseek-v32
reasoning-parser: deepseek-v3
- model_type: chat
engine_config:
tp: 16
Expand Down
Loading
Loading