Skip to content

docs(api): deepseek-v4-flash on /responses, qwen3.8-flash 1M context, two new 400s - #120

Merged
sre-helmcode merged 6 commits into
mainfrom
docs/responses-deepseek-and-qwen38-context
Sep 29, 2026
Merged

sre-helmcode merged 6 commits into
mainfrom
docs/responses-deepseek-and-qwen38-context

Conversation

@sre-helmcode

@sre-helmcode sre-helmcode commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

What

  1. /responses supports deepseek-v4-flash (description + model enum; new stream property). Streaming now documented per model: deepseek-v4-flash streams incrementally with the full event sequence; qwen3.6 and gemma4 hold the answer and send it in one burst at the end. The stale "single terminal event" sentence is gone from the overview, the /responses description and the Codex guide (EN/ES).
  2. qwen3.8-flash context is 1,048,576, not 262,144. Model card, choose-a-model, cline, opencode (limit.context 1048576, output stays 32768), pi (contextWindow), vscode (maxInputTokens 1015808 = window minus output), API catalog, home table (modelos.json), modelCatalog.ts. EN and ES. Other models untouched.
  3. Two new 400s:
    • Tool parameters, when present, must have root "type": "object" (any model). Tool schema pins it (required: [type], type enum ["object"]) and says so in prose.
    • Context overflow: 400, not retried automatically (errors table, shared BadRequest, max_tokens). The exact Context length exceeded for model '...' text is promised only on qwen3.6/gemma4 (measured); on deepseek-v4-flash and other external models it is "usually" that text, since a floor-induced overflow was measured returning the generic Invalid request. Check your request parameters.. A max_tokens above the model's limit is documented as a generic 400. The errors table's 400 row now shows code: "400" with type: invalid_request_error. Tool.parameters gets additionalProperties: true. On deepseek-v4-flash a smaller max_tokens is raised to 16384, so a prompt within 16384 tokens of the window overflows (spec + deepseek card EN/ES).

Deploy note: the tool-parameters 400 is rolling out in a separate PR; merge this after that one is live.

  1. deepseek-v4-flash window is 1048576 (deployment max_input_tokens set 2026-09-29): opencode limit.context, pi contextWindow and vscode maxInputTokens (1015808) updated from the one-token-stale 1048575, EN/ES; docsClientConfigs expects the new value.

  2. Structured output guarantees + deepseek-v4-flash recipe. ResponseFormat and the deepseek-v4-flash card (EN/ES) say what each mode guarantees (json_object: valid JSON only, not the shape; json_schema strict: fields, types, required keys, on qwen3.6/gemma4). New examples.md section (EN/ES): one strict function tool forced with tool_choice, curl + Python; both snippets run as written on 2026-09-29 and returned arguments matching the schema. Tool.function declares strict.

Evidence (measured 2026-09-29, community synthetic key, read-only)

  • POST /v1/responses deepseek-v4-flash, non-streaming: 200, status: completed, output = reasoning item + message item.
  • Same with stream: true: 200, response.created at ~3.0s, reasoning summary deltas, response.output_text.delta spread from 3.3s to 5.5s, response.completed at the end.
  • qwen3.6 / gemma4 with stream: true: 200, but all deltas and response.completed arrive within ~10ms at the end (32.5s / 56.9s).
  • qwen3.8-flash: community proxy LiteLLM_ProxyModelTable max_input_tokens = 1048576; rate-limit hook window 1_048_576.

Tests

  • openapiSpec.test.ts: 3 new describe blocks (responses models/streaming + Codex guide; qwen3.8-flash 1M across every surface in both locales; tool-parameters and context-overflow 400s). Mutation-checked: reverting each touched file to origin/main turns them red.
  • docsClientConfigs.test.ts: expected qwen3.8-flash window 1_048_576 (nan#53 resolved).
  • npm test: 66 files, 1320 tests passing. npm run build: complete. npx astro check: 0 errors.

🤖 Generated with Claude Code

barckcode and others added 6 commits September 29, 2026 20:43
… two new 400s

Measured 2026-09-29 against api.nan.builders with the community synthetic key.

- /responses: deepseek-v4-flash answers 200 (reasoning item + message) and,
  with stream: true, streams incrementally with the full event sequence
  (response.created ... response.output_text.delta ... response.completed).
  qwen3.6 and gemma4 still hold the answer and send it in one burst at the
  end. Added it to the description and the model enum, added a `stream`
  property, and replaced the stale "single terminal event" wording in the
  overview and on the Codex guide (EN/ES).
- qwen3.8-flash is served at 1,048,576 tokens (deployment max_input_tokens
  1048576, rate-limit hook window 1_048_576), not 262,144: fixed the model
  card, choose-a-model, cline, opencode, pi and vscode configs (EN/ES), the
  API catalog, the home table and modelCatalog. VS Code maxInputTokens is
  window minus output (1015808). Other models untouched.
- 400s: tool `parameters` must have a root "type": "object" (schema pins it
  with const); context overflow answers 400 "Context length exceeded for
  model '...'" and is not retried elsewhere; on deepseek-v4-flash a smaller
  max_tokens is raised to 16384, so a prompt within 16384 tokens of the
  window overflows.
- Tests: openapiSpec.test.ts pins all of the above; docsClientConfigs
  expects the 1M window for qwen3.8-flash (nan#53 resolved).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ck passes

SchemaLike in openapiToText has no `const`, so the spec stopped type-checking.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e and pi

The deployment declares max_input_tokens 1048576 since 2026-09-29, so the
1048575 figure is one token stale. VS Code maxInputTokens is window minus
output (1015808). docsClientConfigs expects the new window.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
On deepseek-v4-flash an overflow caused by the 16384 max_tokens floor
answers the generic 400 "Invalid request. Check your request parameters."
(measured 2026-09-29); gemma4 returns "Context length exceeded". The
errors table, BadRequest, max_tokens and the deepseek card (EN/ES) now
promise the exact text only on qwen3.6/gemma4 and say "usually" elsewhere.

Also: a max_tokens above the model's limit answers a generic 400; the
errors table's 400 row shows code "400" with type invalid_request_error,
which is what the API returns; "not retried automatically" replaces the
deployment wording; Tool.parameters gets additionalProperties: true.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…recipe for deepseek-v4-flash

- ResponseFormat and the deepseek-v4-flash card (EN/ES): json_object is
  valid JSON only, not its shape (describe the schema in the prompt; the
  word JSON is required on deepseek-v4-flash); json_schema with strict
  constrains fields, types and required keys, on qwen3.6 and gemma4.
- examples.md (EN/ES): structured output on deepseek-v4-flash via one
  strict function tool forced with tool_choice, curl + Python. Both
  snippets run as written 2026-09-29 against the community proxy and
  returned arguments matching the schema.
- Tool.function declares `strict`.
- Tests pin the copy in both locales and parse the curl body.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…pports it

Decode-time enforcement on deepseek-v4-flash could not be proven (it follows
the schema with strict false too), so Tool.function.strict and the example
prose (EN/ES) no longer read as a guarantee on every model.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@sre-helmcode
sre-helmcode merged commit 770a406 into main Sep 29, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants