Skip to content

[Klaud Cold] kimik2.6-fp4-b200-dynamo-vllm: remove no-enable-flashinfer-autotune, target b200-new / 移除 no-enable-flashinfer-autotune - #2445

Closed
xinli-sw wants to merge 3 commits into
mainfrom
feat/kimik2.6-b200-dynamo-vllm-b200-new
Closed

[Klaud Cold] kimik2.6-fp4-b200-dynamo-vllm: remove no-enable-flashinfer-autotune, target b200-new / 移除 no-enable-flashinfer-autotune#2445
xinli-sw wants to merge 3 commits into
mainfrom
feat/kimik2.6-b200-dynamo-vllm-b200-new

Conversation

@xinli-sw

@xinli-sw xinli-sw commented Aug 1, 2026

Copy link
Copy Markdown
Collaborator

Summary

Remove no-enable-flashinfer-autotune: true from all 7 disagg-b200-* recipe YAMLs under benchmarks/multi_node/srt-slurm-recipes/vllm/kimi-k2.6/b200-fp4/8k1k/.

中文说明

benchmarks/multi_node/srt-slurm-recipes/vllm/kimi-k2.6/b200-fp4/8k1k/ 下所有 7 个 disagg-b200-* 配方 YAML 中移除 no-enable-flashinfer-autotune: true

@github-actions

github-actions Bot commented Aug 1, 2026

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

1 similar comment
@github-actions

github-actions Bot commented Aug 1, 2026

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

@xinli-sw xinli-sw changed the title [Klaud Cold] kimik2.6-fp4-b200-dynamo-vllm: remove no-enable-flashinfer-autotune, target b200-new / 移除 no-enable-flashinfer-autotune,切换至 b200-new runner [Klaud Cold] kimik2.6-fp4-b200-dynamo-vllm: remove no-enable-flashinfer-autotune, target b200-new / 移除 no-enable-flashinfer-autotune, Aug 1, 2026
@xinli-sw xinli-sw changed the title [Klaud Cold] kimik2.6-fp4-b200-dynamo-vllm: remove no-enable-flashinfer-autotune, target b200-new / 移除 no-enable-flashinfer-autotune, [Klaud Cold] kimik2.6-fp4-b200-dynamo-vllm: remove no-enable-flashinfer-autotune, target b200-new / 移除 no-enable-flashinfer-autotun Aug 1, 2026
@xinli-sw xinli-sw changed the title [Klaud Cold] kimik2.6-fp4-b200-dynamo-vllm: remove no-enable-flashinfer-autotune, target b200-new / 移除 no-enable-flashinfer-autotun [Klaud Cold] kimik2.6-fp4-b200-dynamo-vllm: remove no-enable-flashinfer-autotune, target b200-new / 移除 no-enable-flashinfer-autotune Aug 1, 2026
@github-actions

github-actions Bot commented Aug 1, 2026

Copy link
Copy Markdown
Contributor

Comment thread configs/nvidia-master.yaml Outdated
model: nvidia/Kimi-K2.6-NVFP4
model-prefix: kimik2.6
runner: b200-multinode
runner: b200-new

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 This PR retargets kimik2.6-fp4-b200-dynamo-vllm from runner: b200-multinode to runner: b200-new, but no such runner label exists anywhere in the repo — configs/runners.yaml only defines b200, b200-dsv4, and b200-multinode for Blackwell. Since this label is passed verbatim into the GitHub Actions runs-on, the multi-node job will queue indefinitely with no matching self-hosted runner instead of failing fast, which defeats the full-sweep-fail-fast validation this PR is labeled with.

Extended reasoning...

The core functional change in this PR is configs/nvidia-master.yaml:4984, which changes runner: b200-multinode to runner: b200-new for the kimik2.6-fp4-b200-dynamo-vllm entry. I grepped the entire repository and b200-new appears in exactly two places: this new master-config line and the corresponding perf-changelog.yaml:5362 entry describing the change. It is not defined anywhere else.

configs/runners.yaml is the source of truth for valid runner labels. Its labels: map defines only b200, b200-dsv4, and b200-multinode for B200 hardware (lines 53–82) — there is no b200-new key, and there is no corresponding hardware: entry or runners/launch_b200-new*.sh launcher script either. Per AGENTS.md's "Adding a runner" section, introducing a new runner requires three things: (1) an entry in configs/runners.yaml, (2) a launcher script under runners/, and (3) the master-config reference. This PR only does step (3).

I traced how this value is consumed. generate_sweep_configs.py only validates runner labels against runner_labels() when the CLI's --runner-type filter is explicitly passed (lines ~421-427) — it never validates the config's own runner: field during normal generation. For the multi-node fixed-seq-len path, the runner value is read unvalidated and emitted directly as the entry's runner list (runners_for_entry = [runner] when no node filter is set). That value then flows into .github/workflows/benchmark-multinode-tmpl.yml, which builds runs-on: ["self-hosted", inputs.runner, ...] verbatim. Since no self-hosted runner pool is registered with the b200-new label, the resulting GitHub Actions job has no runner to match against.

The practical impact: instead of failing fast (which the full-sweep-fail-fast label on this PR is meant to guarantee), the multi-node job for kimik2.6-fp4-b200-dynamo-vllm will simply queue forever with no runner available to pick it up. This is worse than an explicit failure — it silently hangs, consuming a queued job slot and giving no clear signal to the PR author about what went wrong, unlike a validation error that would point directly at the missing runner registration.

Concrete walkthrough:

  1. configs/nvidia-master.yaml:4984 now reads runner: b200-new for kimik2.6-fp4-b200-dynamo-vllm.
  2. generate_sweep_configs.py generates the multi-node sweep entry for this config, copying runner: b200-new through unchanged (no validation trips).
  3. The GitHub Actions workflow for this entry sets runs-on: ["self-hosted", "b200-new", ...].
  4. GitHub Actions looks for a registered self-hosted runner advertising the b200-new label. None exists (configs/runners.yaml only advertises b200, b200-dsv4, b200-multinode), so the job sits in "Queued" indefinitely.
  5. Even if a runner pool were manually tagged b200-new outside of configs/runners.yaml to work around this, the multi-node SLURM launch would still need runners/launch_b200-new*.sh, which doesn't exist.

To fix this, the author should either revert to runner: b200-multinode (if the intent was just to remove no-enable-flashinfer-autotune, which is the PR's stated purpose) or, if b200-new is an intentional new runner pool, complete the registration by adding it to configs/runners.yaml with its node list and adding the launcher script per AGENTS.md's "Adding a runner" instructions.

@functionstackx functionstackx left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

thanks for the contribution! any chance that there can not be an 100 tok/s/user gap between 55 and 155

Image

@github-actions

github-actions Bot commented Aug 2, 2026

Copy link
Copy Markdown
Contributor

@xinli-sw

xinli-sw commented Aug 2, 2026

Copy link
Copy Markdown
Collaborator Author

looking

@xinli-sw

xinli-sw commented Aug 2, 2026

Copy link
Copy Markdown
Collaborator Author

closed in favor of #2438

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Development

Successfully merging this pull request may close these issues.

2 participants