Skip to content

Count reasoning_content in the restful chat benchmark - #5020

Open
David-Wu1119 wants to merge 1 commit into
InternLM:mainfrom
David-Wu1119:fix/restful-bench-reasoning-ttft
Open

David-Wu1119 wants to merge 1 commit into
InternLM:mainfrom
David-Wu1119:fix/restful-bench-reasoning-ttft

Conversation

@David-Wu1119

Copy link
Copy Markdown
Contributor

Motivation

With --backend lmdeploy-chat, benchmark/profile_restful_api.py only treats delta.content as output. A server started with --reasoning-parser streams the model's thinking as delta.reasoning_content, so:

  • TTFT is measured to the first answer token after all of the reasoning, or stays 0.0 when the whole max_tokens budget goes to reasoning (the request still counts as successful, so the 0 is averaged in);
  • ITL drops every reasoning token, and the retokenized output count only covers the answer.

Against a fake OpenAI-compatible server streaming 64 tokens per request (first token after 100 ms, then one every 10 ms), 20 prompts:

reasoning tokens of 64 Mean TTFT (ms) Mean TPOT (ms) Mean ITL (ms) retokenized output
64 before 0.00 13.57 0.00 0 of 1280
64 after 111.60 11.79 11.79 1280 of 1280
40 before 586.75 4.39 12.04 480 of 1280
40 after 106.55 12.07 12.07 1280 of 1280

lmdeploy's own benchmark/benchmark_chat_completion.py already counts content or reasoning_content as the first token, and sglang's bench_serving.py (which this script follows) counts reasoning_content as content too.

Modification

In async_request_openai_chat_completions, count reasoning_content + content of each delta as output, reading both with or '' since either can be null. The same lines also stop assuming choices is non-empty, so a usage-only chunk with choices: [] no longer raises IndexError.

No open issue or PR covers this (searched for reasoning_content, profile_restful_api reasoning, TTFT reasoning, reasoning-parser benchmark).

BC-breaking (Optional)

No. Non-reasoning output is measured as before.

Checklist

  1. pre-commit (ruff, docformatter, the copyright hook and the other configured hooks) passes on the changed file.
  2. There are no unit tests for the benchmark scripts; the table above comes from the fake-server runs.

Written with AI assistance (Claude Code).

🤖 Generated with Claude Code

With --backend lmdeploy-chat, profile_restful_api.py only treated
delta.content as output. A server started with --reasoning-parser streams
the model's thinking as delta.reasoning_content, so TTFT was measured to
the first answer token after all the reasoning, or stayed 0.0 when the
whole budget went to reasoning, and ITL and the retokenized output count
dropped those tokens. Count reasoning_content as output too, as
benchmark_chat_completion.py and sglang's bench_serving do, and tolerate
chunks with an empty choices list.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Copilot AI balanced review requested due to automatic review settings October 3, 2026 15:19

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants