docs(api): deepseek-v4-flash on /responses, qwen3.8-flash 1M context, two new 400s - #120
Merged
Merged
Conversation
… two new 400s Measured 2026-09-29 against api.nan.builders with the community synthetic key. - /responses: deepseek-v4-flash answers 200 (reasoning item + message) and, with stream: true, streams incrementally with the full event sequence (response.created ... response.output_text.delta ... response.completed). qwen3.6 and gemma4 still hold the answer and send it in one burst at the end. Added it to the description and the model enum, added a `stream` property, and replaced the stale "single terminal event" wording in the overview and on the Codex guide (EN/ES). - qwen3.8-flash is served at 1,048,576 tokens (deployment max_input_tokens 1048576, rate-limit hook window 1_048_576), not 262,144: fixed the model card, choose-a-model, cline, opencode, pi and vscode configs (EN/ES), the API catalog, the home table and modelCatalog. VS Code maxInputTokens is window minus output (1015808). Other models untouched. - 400s: tool `parameters` must have a root "type": "object" (schema pins it with const); context overflow answers 400 "Context length exceeded for model '...'" and is not retried elsewhere; on deepseek-v4-flash a smaller max_tokens is raised to 16384, so a prompt within 16384 tokens of the window overflows. - Tests: openapiSpec.test.ts pins all of the above; docsClientConfigs expects the 1M window for qwen3.8-flash (nan#53 resolved). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ck passes SchemaLike in openapiToText has no `const`, so the spec stopped type-checking. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e and pi The deployment declares max_input_tokens 1048576 since 2026-09-29, so the 1048575 figure is one token stale. VS Code maxInputTokens is window minus output (1015808). docsClientConfigs expects the new window. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
On deepseek-v4-flash an overflow caused by the 16384 max_tokens floor answers the generic 400 "Invalid request. Check your request parameters." (measured 2026-09-29); gemma4 returns "Context length exceeded". The errors table, BadRequest, max_tokens and the deepseek card (EN/ES) now promise the exact text only on qwen3.6/gemma4 and say "usually" elsewhere. Also: a max_tokens above the model's limit answers a generic 400; the errors table's 400 row shows code "400" with type invalid_request_error, which is what the API returns; "not retried automatically" replaces the deployment wording; Tool.parameters gets additionalProperties: true. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…recipe for deepseek-v4-flash - ResponseFormat and the deepseek-v4-flash card (EN/ES): json_object is valid JSON only, not its shape (describe the schema in the prompt; the word JSON is required on deepseek-v4-flash); json_schema with strict constrains fields, types and required keys, on qwen3.6 and gemma4. - examples.md (EN/ES): structured output on deepseek-v4-flash via one strict function tool forced with tool_choice, curl + Python. Both snippets run as written 2026-09-29 against the community proxy and returned arguments matching the schema. - Tool.function declares `strict`. - Tests pin the copy in both locales and parse the curl body. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…pports it Decode-time enforcement on deepseek-v4-flash could not be proven (it follows the schema with strict false too), so Tool.function.strict and the example prose (EN/ES) no longer read as a guarantee on every model. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
deepseek-v4-flash(description +modelenum; newstreamproperty). Streaming now documented per model:deepseek-v4-flashstreams incrementally with the full event sequence;qwen3.6andgemma4hold the answer and send it in one burst at the end. The stale "single terminal event" sentence is gone from the overview, the/responsesdescription and the Codex guide (EN/ES).qwen3.8-flashcontext is 1,048,576, not 262,144. Model card, choose-a-model, cline, opencode (limit.context1048576, output stays 32768), pi (contextWindow), vscode (maxInputTokens1015808 = window minus output), API catalog, home table (modelos.json),modelCatalog.ts. EN and ES. Other models untouched.parameters, when present, must have root"type": "object"(any model). Tool schema pins it (required: [type],typeenum["object"]) and says so in prose.BadRequest,max_tokens). The exactContext length exceeded for model '...'text is promised only onqwen3.6/gemma4(measured); ondeepseek-v4-flashand other external models it is "usually" that text, since a floor-induced overflow was measured returning the genericInvalid request. Check your request parameters.. Amax_tokensabove the model's limit is documented as a generic 400. The errors table's 400 row now showscode: "400"withtype: invalid_request_error.Tool.parametersgetsadditionalProperties: true. Ondeepseek-v4-flasha smallermax_tokensis raised to 16384, so a prompt within 16384 tokens of the window overflows (spec + deepseek card EN/ES).Deploy note: the tool-parameters 400 is rolling out in a separate PR; merge this after that one is live.
deepseek-v4-flashwindow is 1048576 (deploymentmax_input_tokensset 2026-09-29): opencodelimit.context, picontextWindowand vscodemaxInputTokens(1015808) updated from the one-token-stale 1048575, EN/ES;docsClientConfigsexpects the new value.Structured output guarantees + deepseek-v4-flash recipe.
ResponseFormatand the deepseek-v4-flash card (EN/ES) say what each mode guarantees (json_object: valid JSON only, not the shape;json_schemastrict: fields, types, required keys, on qwen3.6/gemma4). Newexamples.mdsection (EN/ES): one strict function tool forced withtool_choice, curl + Python; both snippets run as written on 2026-09-29 and returned arguments matching the schema.Tool.functiondeclaresstrict.Evidence (measured 2026-09-29, community synthetic key, read-only)
POST /v1/responsesdeepseek-v4-flash, non-streaming: 200,status: completed, output =reasoningitem +messageitem.stream: true: 200,response.createdat ~3.0s, reasoning summary deltas,response.output_text.deltaspread from 3.3s to 5.5s,response.completedat the end.qwen3.6/gemma4withstream: true: 200, but all deltas andresponse.completedarrive within ~10ms at the end (32.5s / 56.9s).qwen3.8-flash: community proxyLiteLLM_ProxyModelTablemax_input_tokens= 1048576; rate-limit hook window1_048_576.Tests
openapiSpec.test.ts: 3 new describe blocks (responses models/streaming + Codex guide; qwen3.8-flash 1M across every surface in both locales; tool-parameters and context-overflow 400s). Mutation-checked: reverting each touched file toorigin/mainturns them red.docsClientConfigs.test.ts: expected qwen3.8-flash window 1_048_576 (nan#53 resolved).npm test: 66 files, 1320 tests passing.npm run build: complete.npx astro check: 0 errors.🤖 Generated with Claude Code