Skip to content

[AgentX][sglang] DeepSeek-V4 B300 MTP recipe, conc 1 / 新增 DeepSeek-V4 B300 SGLang MTP AgentX 配方(conc 1) - #2452

Open
yhyang201 wants to merge 4 commits into
mainfrom
dsv4-b300-sglang-agentic-mtp
Open

[AgentX][sglang] DeepSeek-V4 B300 MTP recipe, conc 1 / 新增 DeepSeek-V4 B300 SGLang MTP AgentX 配方(conc 1)#2452
yhyang201 wants to merge 4 commits into
mainfrom
dsv4-b300-sglang-agentic-mtp

Conversation

@yhyang201

@yhyang201 yhyang201 commented Aug 2, 2026

Copy link
Copy Markdown
Collaborator

Smoke test for a SGLang EAGLE/MTP AgentX recipe on DeepSeek-V4-Pro FP4 B300. One point only: TP8, no KV offload, conc 1.

Acceptance is pinned to the golden AL 2.49 (golden_al_distribution/dsv4_mtp.yaml, thinking_on, K=3).

中文:DeepSeek-V4-Pro FP4 B300 上 SGLang EAGLE/MTP AgentX 配方的冒烟测试,只跑一个点:TP8、无 KV offload、conc 1。接受长度按黄金 AL 2.49 固定。

中文:新增 DSv4 B300 SGLang AgentX MTP 配方。
@github-actions

github-actions Bot commented Aug 2, 2026

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

1 similar comment
@github-actions

github-actions Bot commented Aug 2, 2026

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

中文:回填 DSv4 B300 SGLang MTP changelog 条目的 PR 链接。
@yhyang201 yhyang201 changed the title [AgentX][sglang] DeepSeek-V4 B300 MTP smoke test / DeepSeek-V4 B300 SGLang MTP 冒烟测试 [AgentX][sglang] DeepSeek-V4 B300 MTP recipe, conc 1 / 新增 DeepSeek-V4 B300 SGLang MTP AgentX 配方(conc 1) Aug 2, 2026
@yhyang201 yhyang201 added the agentx-fast Run AgentX throughput with 1 warmup request per lane and a 20-minute profile; not reusable label Aug 2, 2026
Comment thread perf-changelog.yaml Outdated
- "Pin simulated acceptance to the committed golden AL: SGLANG_SIMULATE_ACC_LEN=2.49 (golden_al_distribution/dsv4_mtp.yaml, thinking_on, K=3), SGLANG_SIMULATE_ACC_METHOD=match-expected, SGLANG_SIMULATE_ACC_TOKEN_MODE=real-draft-token; EVAL_ONLY runs keep real verification"
- "Bump image from lmsysorg/sglang:nightly-dev-cu13-20260707-b4155233 to lmsysorg/sglang:nightly-dev-cu13-20260802-131bd51b, which reads SGLANG_SIMULATE_ACC_TOKEN_MODE (added in v0.5.16)"
- "Search space is one point: TP8, kv-offloading none, conc 1 — same topology as the STP entry TP8 row and the vLLM dsv4-fp4-b300-vllm-agentic-mtp TP8 arm"
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/PENDING

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 perf-changelog.yaml:5366 sets pr-link to .../pull/PENDING, but this PR is #2452. This isn't just a stale link — utils/validate_perf_changelog.py's validate_added_pr_link() only accepts the canonical pull/<pr_number> link or the XXX placeholder, so PENDING will fail changelog validation on merge/ingest. Please change it to pull/2452 (or XXX if the number wasn't known yet).

Extended reasoning...

The newly appended changelog entry for dsv4-fp4-b300-sglang-agentic-mtp sets:

pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/PENDING

Every other entry in perf-changelog.yaml uses a real numeric PR link (e.g. pull/2213, pull/2312, pull/2360), and AGENTS.md's changelog template only shows pull/XXX as a fill-in-the-blank placeholder — never as a literal committed value. PENDING is neither of those things.

Why this actually breaks validation (not just cosmetic):

utils/validate_perf_changelog.py defines:

CANONICAL_PR_LINK = re.compile(r"https://github\.com/SemiAnalysisAI/InferenceX/pull/\d+")
PR_LINK_PLACEHOLDERS = {"XXX", "https://github.com/SemiAnalysisAI/InferenceX/pull/XXX"}

and validate_added_pr_link(link, pr_number):

def validate_added_pr_link(link: str, pr_number: int | None) -> None:
    if pr_number is None:
        if not CANONICAL_PR_LINK.fullmatch(link):
            raise ChangelogValidationError(...)
        return
    expected = f"https://github.com/SemiAnalysisAI/InferenceX/pull/{pr_number}"
    if link not in PR_LINK_PLACEHOLDERS and link != expected:
        raise ChangelogValidationError(
            f"new PR entry must use {expected!r} or an XXX placeholder; found {link!r}"
        )

This is invoked from compare_entries() for every newly-added entry, which in turn is called by prepare_perf_changelog_merge.py's canonicalize_appended_links() and resolve_conflict_bytes() — i.e. it runs on both PR-time validation and at merge/ingest time.

Step-by-step proof:

  1. This PR is [AgentX][sglang] DeepSeek-V4 B300 MTP recipe, conc 1 / 新增 DeepSeek-V4 B300 SGLang MTP AgentX 配方(conc 1) #2452, so on a PR run pr_number = 2452 and expected = "https://github.com/SemiAnalysisAI/InferenceX/pull/2452".
  2. The committed value is "https://github.com/SemiAnalysisAI/InferenceX/pull/PENDING".
  3. Check link not in PR_LINK_PLACEHOLDERS: PENDING is not "XXX" nor ".../pull/XXX" → condition is True.
  4. Check link != expected: ".../pull/PENDING" != ".../pull/2452" → condition is True.
  5. Both conditions are True, so ChangelogValidationError is raised: "new PR entry must use '.../pull/2452' or an XXX placeholder; found '.../pull/PENDING'".
  6. Separately, on a main-branch run (pr_number is None), CANONICAL_PR_LINK.fullmatch() requires \d+ after pull/, which PENDING also fails to match.

Either path — PR-time validation or main-branch ingest — raises and blocks the changelog step.

Fix: change line 5366 to pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/2452 (the actual PR number), or use the literal XXX placeholder form if the number is genuinely unknown at commit time.

@github-actions

github-actions Bot commented Aug 2, 2026

Copy link
Copy Markdown
Contributor

Enable the FP4 indexer, cache reporting and the golden-AL thinking mode,
size the decode CUDA graph to max-running-requests, and add a TP4 arm at
conc 1 alongside TP8.

中文:为 DSv4 B300 SGLang MTP 配方启用 FP4 indexer、缓存上报和黄金 AL 对应的
thinking 模式,decode CUDA graph 按 max-running-requests 取值,并在 conc 1 上
新增 TP4 臂。
@yhyang201 yhyang201 added full-sweep-fail-fast-no-canary Full sweep, no canary gate; first failure in a matrix cancels that matrix agentx AgentX benchmarks, recipes, and infrastructure and removed agentx-fast Run AgentX throughput with 1 warmup request per lane and a 20-minute profile; not reusable labels Aug 2, 2026
@github-actions

github-actions Bot commented Aug 2, 2026

Copy link
Copy Markdown
Contributor

@yhyang201

Copy link
Copy Markdown
Collaborator Author

At CONC=1 the AgentX subagent fan-out puts more than one request in flight — peak 5 in the TP8 run on this branch. How should recipes handle that?

中文:CONC=1 时 AgentX 的 subagent 分裂会让实际在途请求数超过 1,本分支 TP8 那轮峰值为 5。配方应该如何处理?

@github-actions

github-actions Bot commented Aug 2, 2026

Copy link
Copy Markdown
Contributor

… SGLang MTP

中文:为 DSv4 B300 SGLang MTP 配方新增 conc 8 的有/无 HiCache 两个点。
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

agentx AgentX benchmarks, recipes, and infrastructure full-sweep-fail-fast-no-canary Full sweep, no canary gate; first failure in a matrix cancels that matrix

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

1 participant